Python Basic Interview Questions
Python Basic Interview Questions 🐍
Top 30 basic Python interview questions for Data Analysts — Variables, Data Types, Strings, Lists, Dictionaries, Loops, Functions, File Handling, aur NumPy/Pandas introduction. English answers with Hinglish explanations aur real code examples. Data Insights par.
📑 Is Blog Mein Kya Sikhenge:
- 🟢 Q1–Q6: Python Fundamentals — Variables, Data Types, Operators
- 🟢 Q7–Q12: Strings, Lists, Tuples, Dictionaries, Sets
- 🟢 Q13–Q18: Loops, Conditions, Functions
- 🟢 Q19–Q24: File Handling, Error Handling, Modules
- 🟢 Q25–Q30: NumPy & Pandas Introduction
- 💡 Pro Tips: Interview mein exactly kya bolna chahiye
🟢 Category 1: Python Fundamentals (Q1–Q6)
Q1: What is Python and why is it used in Data Analysis?
Answer: Python is a high-level, interpreted, general-purpose programming language known for its simple and readable syntax. It is widely used in Data Analysis because of its rich ecosystem of libraries — Pandas for data manipulation, NumPy for numerical computing, Matplotlib/Seaborn for visualization, Scikit-learn for machine learning, and Jupyter Notebooks for interactive analysis. Python is free, open-source, has a massive community, and is the most in-demand language for data roles across industries.
🎯 Explain: Python ek programming language hai jo simple English jaisi dikhti hai — beginners ke liye easy. Data Analysis mein isliye popular hai kyunki Pandas, NumPy, Matplotlib jaise powerful libraries free mein available hain. Excel mein jo 100 formulas mein hota hai woh Python mein 5 lines mein ho jaata hai. Companies jaise Google, Netflix, Amazon sab Python use karti hain data analysis ke liye. Interview mein "Python's simplicity and rich library ecosystem make it the top choice for data analysis" — yeh key line hai.
Q2: What are Variables in Python? How do you declare them?
Answer: Variables are named containers that store data values. In Python, you do not need to declare the data type explicitly — Python is dynamically typed, meaning the type is determined automatically at runtime based on the assigned value. Variable names must start with a letter or underscore, cannot start with a number, are case-sensitive, and cannot use reserved keywords. No special keyword (like var, int, let) is needed — just assign a value with the = operator.
🎯 Explain: Variable = ek naam dete ho value ko store karne ke liye. Python mein type declare nahi karna padta — x = 10 likho, Python khud samajh jayega ki yeh integer hai. x = "Hello" likho — ab string ho gaya. Dynamically typed hai — type runtime pe decide hota hai. Rules: naam letter ya underscore se start ho, numbers se nahi. Case-sensitive hai — Name aur name alag hain. Interview mein "Python is dynamically typed — no explicit type declaration needed" — yeh important point hai.
# Variable declaration — no type needed
name = "Aarav" # String
age = 25 # Integer
salary = 55000.50 # Float
is_active = True # Boolean
# Multiple assignment
a, b, c = 10, 20, 30
# Check type
print(type(name)) # <class 'str'>
print(type(salary)) # <class 'float'>
Q3: What are the basic data types in Python?
Answer: Python has several built-in data types: (1) int — whole numbers (10, -5, 0). (2) float — decimal numbers (3.14, -0.5). (3) str — text strings ("Hello", 'Python'). (4) bool — Boolean values (True, False). (5) NoneType — represents absence of value (None). (6) complex — complex numbers (3+4j). Additionally, Python has collection types: list, tuple, dict, set. The type() function checks any variable's data type. Understanding data types is fundamental for data manipulation and analysis.
🎯 Explain: Python mein basic types: int (poore numbers), float (decimal), str (text), bool (True/False), None (kuch nahi). Collections: list (ordered, changeable), tuple (ordered, unchangeable), dict (key-value pairs), set (unique values). type() function se kisi bhi variable ka type check karo. Data Analysis mein mostly int, float, str aur Pandas ke special types (datetime, category) use hote hain. Interview mein "Python has dynamic typing with int, float, str, bool as primitives and list, dict, tuple, set as collections" — complete answer.
# Basic Data Types
x = 10 # int
y = 3.14 # float
name = "Python" # str
flag = True # bool
nothing = None # NoneType
# Type conversion
a = int("25") # String to Integer → 25
b = str(100) # Integer to String → "100"
c = float("3.14") # String to Float → 3.14
d = bool(0) # Integer to Bool → False
e = bool(1) # Integer to Bool → True
Q4: What are the different types of Operators in Python?
Answer: Python operators include: (1) Arithmetic — +, -, *, /, // (floor division), % (modulus), ** (power). (2) Comparison — ==, !=, >, <, >=, <=. (3) Logical — and, or, not. (4) Assignment — =, +=, -=, *=, /=. (5) Identity — is, is not (checks if same object in memory). (6) Membership — in, not in (checks if value exists in sequence). (7) Bitwise — &, |, ^, ~, <<, >>. For data analysis, arithmetic, comparison, logical, and membership operators are most frequently used.
🎯 Explain: Operators = symbols jo operations perform karte hain. Arithmetic: +, -, *, / (basic math). // = floor division (decimal hata do). ** = power (2**3 = 8). % = remainder (10%3 = 1). Comparison: ==, !=, >, < (True/False return karte hain). Logical: and, or, not (conditions combine karo). Membership: "in" — "a" in "Aarav" → True. Identity: "is" — same object check. Data Analysis mein mostly arithmetic aur comparison use hote hain Pandas ke saath.
# Arithmetic
print(10 / 3) # 3.333... (true division)
print(10 // 3) # 3 (floor division)
print(10 % 3) # 1 (modulus/remainder)
print(2 ** 3) # 8 (power)
# Comparison
print(10 == 10) # True
print(10 != 5) # True
# Membership
print("a" in "Aarav") # True
print(5 in [1,2,3]) # False
# Logical
print(10 > 5 and 10 < 20) # True
print(not True) # False
Q5: What is the difference between = and == in Python?
Answer: = is the assignment operator — it assigns a value to a variable (x = 10 stores 10 in x). == is the comparison/equality operator — it checks if two values are equal and returns True or False (x == 10 returns True if x is 10). This is one of the most common beginner mistakes — using = instead of == in if conditions. Python will raise a syntax error if = is used where == is expected in modern versions.
🎯 Explain: = matlab "yeh value daal do" — x = 10 (x mein 10 store karo). == matlab "kya yeh barabar hai?" — x == 10 (kya x 10 hai? True/False). Beginners aksar if condition mein = likh dete hain == ki jagah — galat results aate hain. Python mein if x = 10 error dega — but kuch languages mein silently accept ho jaata hai. Interview mein yeh difference clearly batao — basic but important concept hai.
Q6: What is the difference between input() and print() functions?
Answer: print() is an output function that displays values to the console/screen. It can accept multiple arguments separated by commas, supports formatting through f-strings, and has parameters like sep (separator) and end (line ending). input() is an input function that takes user input from the keyboard — it always returns a string, so you need to convert it using int() or float() for numeric input. print() is used extensively in data analysis for displaying results, while input() is more common in standalone scripts.
🎯 Explain: print() = screen pe dikhao. input() = user se value lo. print("Hello") → Hello dikhega. name = input("Enter name: ") → user type karega aur name mein store hoga. Important: input() hamesha string return karta hai — age = input("Age: ") → "25" string aayega, 25 number nahi. int(input("Age: ")) — yeh sahi tarika hai number lene ka. Data Analysis mein print() zyada use hota hai results display karne ke liye.
# print() with f-string formatting
name = "Aarav"
salary = 55000
print(f"{name} earns ₹{salary}")
# Output: Aarav earns ₹55000
# print() with sep and end
print("A", "B", "C", sep="-") # A-B-C
print("Hello", end=" ")
print("World") # Hello World (same line)
# input() — always returns string
age = input("Enter age: ") # Returns "25" (string)
age = int(input("Enter age: ")) # Returns 25 (integer)
🟢 Category 2: Strings, Lists, Tuples, Dictionaries (Q7–Q12)
Q7: What are Strings in Python and how do you manipulate them?
Answer: Strings are immutable sequences of characters enclosed in single (''), double (""), or triple (''' ''' or """ """) quotes. Key operations: indexing (s[0]), slicing (s[1:4]), concatenation (+), repetition (*), and built-in methods — upper(), lower(), strip(), split(), join(), replace(), find(), count(), startswith(), endswith(). Strings are immutable — you cannot change individual characters; you must create a new string. F-strings (f"...{variable}...") provide the most readable way to format strings.
🎯 Explain: String = text data — quotes mein likha jaata hai. "Hello" ya 'Hello' dono same hain. Immutable hai — ek baar bana toh character change nahi kar sakte, naya banana padega. Slicing bahut important hai — s[0:3] pehle 3 characters. Methods: upper() se CAPS, lower() se small, strip() se spaces hatao, split() se words alag karo, join() se words jodo. Data cleaning mein strings bahut kaam aati hain — names clean karo, emails parse karo. Interview mein 5-6 string methods yaad rakho.
# String basics
s = "Hello Python"
print(s[0]) # H (indexing)
print(s[0:5]) # Hello (slicing)
print(s.upper()) # HELLO PYTHON
print(s.lower()) # hello python
print(s.split()) # ['Hello', 'Python']
print(s.replace("Python", "World")) # Hello World
print(len(s)) # 12 (length)
# String cleaning (data analysis use case)
dirty = " Aarav Sharma "
print(dirty.strip()) # "Aarav Sharma"
print(dirty.strip().title()) # "Aarav Sharma"
Q8: What is a List in Python?
Answer: A List is an ordered, mutable (changeable) collection that can hold items of different data types. Lists are defined using square brackets []. Key operations: indexing, slicing, append(), insert(), remove(), pop(), sort(), reverse(), extend(), count(), index(). Lists support nested lists (list of lists). They are the most commonly used data structure in Python — equivalent to arrays in other languages but more flexible. List comprehensions provide concise syntax for creating lists.
🎯 Explain: List = ordered collection — items ka order fixed hai, change kar sakte ho, duplicates allowed. [1, 2, 3, "Hello", True] — mixed types rakh sakte ho. Mutable hai — append() se add karo, remove() se hatao, sort() se order karo. List Comprehension — [x*2 for x in range(5)] — ek line mein list banao. Data Analysis mein bahut use hota hai — column names ki list, values ki list, filtered results. Interview mein list comprehension ka example do — impressive lagta hai.
# List operations
fruits = ["apple", "banana", "cherry"]
fruits.append("mango") # Add at end
fruits.insert(1, "grape") # Add at index 1
fruits.remove("banana") # Remove by value
fruits.sort() # Sort alphabetically
# List Comprehension — powerful one-liner
squares = [x**2 for x in range(1, 6)]
print(squares) # [1, 4, 9, 16, 25]
# Filtered list comprehension
salaries = [55000, 72000, 48000, 80000]
high = [s for s in salaries if s > 60000]
print(high) # [72000, 80000]
Q9: What is the difference between a List and a Tuple?
Answer: Lists are mutable (can be changed after creation) — defined with []. Tuples are immutable (cannot be changed after creation) — defined with (). Lists use more memory and are slightly slower. Tuples use less memory and are faster. Lists are used when data needs to be modified — adding, removing items. Tuples are used for fixed data — like coordinates (x, y), database records, dictionary keys. Tuples are hashable (can be dictionary keys), lists are not.
🎯 Explain: List = [] changeable. Tuple = () unchangeable. List: [1,2,3] → append karo, remove karo, modify karo — sab possible. Tuple: (1,2,3) → ek baar bana diya toh change nahi hoga — error aayega. Tuple kab use karo? Fixed data ke liye — (latitude, longitude), (name, age), function se multiple values return karo. Tuple faster hai aur kam memory leta hai. Dictionary key ke roop mein tuple use kar sakte ho, list nahi. Interview mein "Lists for mutable data, Tuples for immutable fixed data" — clear difference.
# List vs Tuple
my_list = [1, 2, 3]
my_tuple = (1, 2, 3)
my_list[0] = 99 # ✅ Works — list is mutable
# my_tuple[0] = 99 # ❌ TypeError — tuple is immutable
# Tuple unpacking
point = (10, 20)
x, y = point
print(x, y) # 10 20
Q10: What is a Dictionary in Python?
Answer: A Dictionary is an unordered (Python 3.7+ maintains insertion order), mutable collection of key-value pairs. Defined with curly braces {} or dict(). Keys must be unique and immutable (strings, numbers, tuples). Values can be any type. Access values by key — d["name"]. Key methods: keys(), values(), items(), get(), update(), pop(). Dictionaries are used extensively in data analysis — JSON data is essentially nested dictionaries, and Pandas DataFrames can be created from dictionaries.
🎯 Explain: Dictionary = key-value pairs ka collection. {"name": "Aarav", "age": 25, "dept": "IT"}. Key se value access karo — d["name"] → "Aarav". get() method safe hai — d.get("salary", 0) → agar key nahi hai toh 0 return karega, error nahi. JSON data dictionaries mein aata hai Python mein. Pandas DataFrame bana sakte ho dictionary se — pd.DataFrame({"Name": [...], "Salary": [...]}). Interview mein "Dictionaries are the backbone of JSON processing and DataFrame creation" — data context do.
# Dictionary operations
emp = {"name": "Aarav", "dept": "IT", "salary": 55000}
print(emp["name"]) # Aarav
print(emp.get("age", 0)) # 0 (key not found, default returned)
emp["city"] = "Delhi" # Add new key-value
print(emp.keys()) # dict_keys(['name','dept','salary','city'])
print(emp.values()) # dict_values(['Aarav','IT',55000,'Delhi'])
# Dictionary to DataFrame
import pandas as pd
data = {"Name": ["Aarav","Ishita"], "Salary": [55000,72000]}
df = pd.DataFrame(data)
print(df)
Q11: What is a Set in Python?
Answer: A Set is an unordered, mutable collection of unique elements — duplicates are automatically removed. Defined with {} or set(). Key operations: add(), remove(), discard(), union (|), intersection (&), difference (-), symmetric_difference (^). Sets are used for removing duplicates from lists, finding common elements between two datasets, membership testing (very fast O(1) lookup). Sets cannot contain mutable items like lists or dictionaries.
🎯 Explain: Set = unique values ka collection — duplicates automatically hat jaate hain. {1,2,3,3,2} → {1,2,3}. List mein duplicates hatane ka sabse fast tarika — set(my_list). Operations: union (dono ke sab elements), intersection (common elements), difference (ek mein hai dusre mein nahi). Data Analysis mein: unique customers nikalo, common products dhundho, missing IDs find karo. Interview mein "I use sets for fast duplicate removal and finding common elements between datasets."
# Set operations
a = {1, 2, 3, 4}
b = {3, 4, 5, 6}
print(a | b) # {1,2,3,4,5,6} — Union
print(a & b) # {3,4} — Intersection
print(a - b) # {1,2} — Difference
# Remove duplicates from list
names = ["Aarav","Ishita","Aarav","Kabir","Ishita"]
unique = list(set(names))
print(unique) # ['Aarav', 'Ishita', 'Kabir']
Q12: What is the difference between List, Tuple, Set, and Dictionary?
Answer: List — ordered, mutable, allows duplicates, defined with []. Tuple — ordered, immutable, allows duplicates, defined with (). Set — unordered, mutable, no duplicates, defined with {}. Dictionary — ordered (3.7+), mutable, no duplicate keys, key-value pairs, defined with {key:value}. Use Lists for general ordered data, Tuples for fixed data, Sets for unique values and fast membership testing, Dictionaries for key-value mappings and JSON-like data.
🎯 Explain: Yeh comparison table interview mein bahut pucha jaata hai. List = flexible, ordered, duplicates OK. Tuple = fixed, fast, duplicates OK. Set = unique only, unordered, fast lookup. Dictionary = key-value mapping, ordered (3.7+). Data Analysis context: List — column values. Tuple — function return. Set — unique IDs. Dictionary — JSON data, DataFrame creation. Interview mein comparison table mentally ready rakho — 4 properties pe compare karo: ordered, mutable, duplicates, syntax.
📊 Comparison Table:
| Feature | List | Tuple | Set | Dictionary |
|---|---|---|---|---|
| Syntax | [ ] | ( ) | { } | {k:v} |
| Ordered | ✅ Yes | ✅ Yes | ❌ No | ✅ Yes (3.7+) |
| Mutable | ✅ Yes | ❌ No | ✅ Yes | ✅ Yes |
| Duplicates | ✅ Allowed | ✅ Allowed | ❌ No | ❌ Keys No |
| Best For | General data | Fixed data | Unique values | Key-value mapping |
🟢 Category 3: Loops, Conditions, Functions (Q13–Q18)
Q13: How do if-elif-else conditions work in Python?
Answer: Python uses if-elif-else for conditional branching. if checks the first condition — if True, its block executes. elif (else if) checks additional conditions — only if previous conditions were False. else executes when all conditions are False. Python uses indentation (4 spaces) instead of curly braces to define code blocks. Conditions can use comparison operators, logical operators (and, or, not), and membership operators (in, not in).
🎯 Explain: if-elif-else = agar yeh toh woh karo, warna yeh karo. Python mein curly braces {} nahi — indentation (4 spaces) se block define hota hai. if salary > 75000: print("Executive") elif salary > 60000: print("Senior") else: print("Junior"). Conditions top to bottom check hoti hain — pehli TRUE wali execute hoti hai. Data Analysis mein conditions bahut use hote hain — data categorization, filtering, flag creation. Interview mein indentation importance mention karo — "Python uses indentation instead of braces for code blocks."
# if-elif-else
salary = 72000
if salary >= 75000:
print("Executive")
elif salary >= 60000:
print("Senior")
elif salary >= 45000:
print("Mid-Level")
else:
print("Junior")
# Output: Senior
# Ternary (one-line if-else)
status = "High" if salary > 60000 else "Low"
print(status) # High
Q14: What is the difference between for loop and while loop?
Answer: A for loop iterates over a sequence (list, tuple, string, range, dictionary) — the number of iterations is known or determined by the sequence length. A while loop repeats as long as a condition is True — the number of iterations is unknown and depends on when the condition becomes False. for loops are more common in Python due to the iterable-based approach. while loops are used for indefinite iteration — like reading until end of file or waiting for user input. break exits the loop, continue skips to next iteration.
🎯 Explain: for loop = sequence pe iterate karo — "har item ke liye yeh karo". for name in names: print(name). while loop = condition True hai tab tak chalo — while count < 10: count += 1. Python mein for loop zyada use hota hai — lists, DataFrames, files sab pe iterate karte hain. break = loop se bahar niklo. continue = current iteration skip karo, next pe jao. Data Analysis mein for loop common hai — rows iterate karo, files process karo. Interview mein "for for known iterations, while for unknown — I mostly use for loops with Pandas iterrows."
# for loop
departments = ["IT", "HR", "Finance"]
for dept in departments:
print(f"Department: {dept}")
# for with range
for i in range(1, 6):
print(i, end=" ") # 1 2 3 4 5
# while loop
count = 0
while count < 3:
print(count)
count += 1
# 0, 1, 2
# enumerate — index + value
for i, dept in enumerate(departments):
print(f"{i}: {dept}")
# 0: IT, 1: HR, 2: Finance
Q15: What are Functions in Python?
Answer: Functions are reusable blocks of code that perform a specific task. Defined using the def keyword followed by the function name and parameters in parentheses. Functions can accept arguments (positional, keyword, default, *args, **kwargs), perform operations, and return values using return statement. If no return is specified, the function returns None. Functions promote code reusability, modularity, and readability. Lambda functions are anonymous single-expression functions.
🎯 Explain: Function = code ka reusable block — ek baar likho, baar baar use karo. def function_name(parameters): likho, andar logic likho, return se value return karo. Default parameters de sakte ho — def greet(name="User"). *args multiple positional arguments le sakta hai. **kwargs multiple keyword arguments le sakta hai. Data Analysis mein custom functions bahut use hoti hain — data cleaning functions, calculation functions. Lambda = chhoti anonymous function — lambda x: x*2. Interview mein def function + lambda dono ka example do.
# Function with default parameter
def calc_bonus(salary, rate=0.10):
return salary * rate
print(calc_bonus(55000)) # 5500.0 (default 10%)
print(calc_bonus(55000, 0.15)) # 8250.0 (custom 15%)
# Function returning multiple values
def get_stats(numbers):
return min(numbers), max(numbers), sum(numbers)/len(numbers)
low, high, avg = get_stats([55000, 72000, 48000, 80000])
print(f"Min: {low}, Max: {high}, Avg: {avg}")
# Lambda function
double = lambda x: x * 2
print(double(25)) # 50
# Lambda with Pandas apply (preview)
# df['Bonus'] = df['Salary'].apply(lambda x: x * 0.10)
Q16: What is the difference between *args and **kwargs?
Answer: *args allows a function to accept any number of positional arguments — they are collected into a tuple. **kwargs allows a function to accept any number of keyword arguments — they are collected into a dictionary. *args is used when you do not know how many arguments will be passed. **kwargs is used when you want named arguments with flexibility. Both can be used together — def func(*args, **kwargs). They provide maximum flexibility for function design.
🎯 Explain: *args = multiple positional arguments tuple mein. def add(*args): return sum(args) → add(1,2,3) = 6, add(1,2,3,4,5) = 15 — kitne bhi numbers do. **kwargs = multiple keyword arguments dictionary mein. def info(**kwargs): → info(name="Aarav", age=25) → {"name": "Aarav", "age": 25}. Real-world: library functions mein bahut use hota hai — matplotlib, pandas ke functions *args aur **kwargs accept karte hain flexibility ke liye.
# *args — variable positional arguments
def total(*args):
return sum(args)
print(total(10, 20, 30)) # 60
print(total(5, 10, 15, 20, 25)) # 75
# **kwargs — variable keyword arguments
def employee_info(**kwargs):
for key, value in kwargs.items():
print(f"{key}: {value}")
employee_info(name="Aarav", dept="IT", salary=55000)
# name: Aarav, dept: IT, salary: 55000
Q17: What is a Lambda function?
Answer: A Lambda function is a small anonymous (unnamed) function defined using the lambda keyword. It can take any number of arguments but can only have one expression. Syntax: lambda arguments: expression. Lambda functions are commonly used with map(), filter(), sorted(), and Pandas apply() for quick inline operations. They are not meant to replace regular functions — use them for simple, short operations where defining a full function would be overkill.
🎯 Explain: Lambda = ek line ki chhoti function — naam nahi hota. lambda x: x*2 — input x lo, x*2 return karo. Pandas mein bahut use hota hai: df['Bonus'] = df['Salary'].apply(lambda x: x*0.10). map() ke saath: list(map(lambda x: x.upper(), names)) — sab names uppercase. filter() ke saath: list(filter(lambda x: x>60000, salaries)) — sirf 60000+ salaries. Interview mein Pandas apply() ke saath lambda ka example do — sabse practical use case hai.
# Lambda with map
salaries = [55000, 72000, 48000, 80000]
bonuses = list(map(lambda x: x * 0.10, salaries))
print(bonuses) # [5500.0, 7200.0, 4800.0, 8000.0]
# Lambda with filter
high_sal = list(filter(lambda x: x > 60000, salaries))
print(high_sal) # [72000, 80000]
# Lambda with sorted
employees = [("Aarav",55000), ("Ishita",72000), ("Meera",48000)]
sorted_emp = sorted(employees, key=lambda x: x[1], reverse=True)
print(sorted_emp)
# [('Ishita',72000), ('Aarav',55000), ('Meera',48000)]
Q18: What is the difference between range() and enumerate()?
Answer: range() generates a sequence of numbers — range(start, stop, step). It is used with for loops when you need a counter or need to iterate a specific number of times. enumerate() adds an automatic counter to an iterable — it returns both the index and the value as pairs. It is used when you need both the position and the element while iterating. enumerate() is more Pythonic than using range(len(list)) for indexed iteration.
🎯 Explain: range(5) → 0,1,2,3,4 — sirf numbers generate karta hai. enumerate(list) → (0, "IT"), (1, "HR"), (2, "Finance") — index + value dono deta hai. Purana tarika: for i in range(len(names)): print(i, names[i]) — ugly. Pythonic tarika: for i, name in enumerate(names): print(i, name) — clean. Interview mein "I prefer enumerate over range(len()) for cleaner, more Pythonic code" — coding style awareness dikhata hai.
🟢 Category 4: File Handling, Error Handling, Modules (Q19–Q24)
Q19: How do you read and write files in Python?
Answer: Python uses the open() function with modes: 'r' (read), 'w' (write — overwrites), 'a' (append), 'r+' (read+write). The with statement is the recommended approach — it automatically closes the file. Read methods: read() (entire file), readline() (one line), readlines() (list of lines). Write methods: write() (string), writelines() (list of strings). For data analysis, Pandas provides pd.read_csv(), pd.read_excel() for structured data files — much more practical than raw file handling.
🎯 Explain: File handling = files read/write karna. with open("file.txt", "r") as f: — yeh best tarika hai — file automatically close hoti hai. "r" = read, "w" = write (purana data delete), "a" = append (add karo). Data Analysis mein mostly Pandas use hota hai — pd.read_csv("data.csv") ek line mein CSV load. Raw file handling scripts, log files, text processing mein use hota hai. Interview mein "I use with statement for safe file handling and Pandas for structured data" — practical answer.
# Writing to file
with open("output.txt", "w") as f:
f.write("Hello Python\n")
f.write("Data Analysis\n")
# Reading from file
with open("output.txt", "r") as f:
content = f.read()
print(content)
# Pandas way (for data analysis)
import pandas as pd
df = pd.read_csv("employees.csv")
df.to_csv("output.csv", index=False)
Q20: What is Exception Handling in Python (try-except)?
Answer: Exception Handling prevents programs from crashing when errors occur. The try block contains code that might raise an error. except catches specific exceptions and handles them gracefully. else executes when no exception occurs. finally always executes regardless of exceptions — used for cleanup operations like closing files. Common exceptions: ValueError, TypeError, KeyError, IndexError, FileNotFoundError, ZeroDivisionError. Best practice: catch specific exceptions, not generic Exception.
🎯 Explain: try-except = error handle karo bina program crash kiye. try mein risky code likho, except mein error handle karo. Example: user ne number ki jagah text daal diya — ValueError aayega — except mein "Please enter a number" dikhao. Data Analysis mein: file nahi mili (FileNotFoundError), wrong column name (KeyError), type mismatch (TypeError). finally mein cleanup karo — file close, connection close. Interview mein "I use try-except for robust data pipelines — handling missing files and invalid data gracefully."
# try-except-else-finally
try:
result = 10 / 0
except ZeroDivisionError:
print("Cannot divide by zero!")
except TypeError:
print("Wrong data type!")
else:
print("Success!")
finally:
print("Cleanup done")
# Data Analysis use case
try:
df = pd.read_csv("data.csv")
except FileNotFoundError:
print("File not found! Check the path.")
except pd.errors.EmptyDataError:
print("File is empty!")
Q21: What is the difference between a Module, Package, and Library?
Answer: A Module is a single Python file (.py) containing functions, classes, and variables — imported using import statement. A Package is a collection of related modules organized in a directory with an __init__.py file — like a folder of modules. A Library is a collection of packages that provides functionality for a specific domain — like Pandas (data manipulation), NumPy (numerical computing), Matplotlib (visualization). Install libraries using pip install library_name.
🎯 Explain: Module = ek .py file — math module mein math.sqrt() hai. Package = modules ka folder — pandas package mein multiple modules hain. Library = packages ka collection — NumPy library, Pandas library. import pandas as pd — Pandas library import kari "pd" shortname se. pip install pandas — library install karo. Data Analysis ke key libraries: Pandas, NumPy, Matplotlib, Seaborn, Scikit-learn. Interview mein "Module is a file, Package is a directory of modules, Library is a collection of packages" — hierarchy clear batao.
Q22: What is pip in Python?
Answer: pip (Pip Installs Packages) is Python's default package manager used to install, update, and remove third-party libraries from PyPI (Python Package Index). Key commands: pip install package_name (install), pip install --upgrade package_name (update), pip uninstall package_name (remove), pip list (show installed packages), pip freeze (show versions for requirements.txt). For data analysis, pip install pandas numpy matplotlib seaborn installs the core data science stack.
🎯 Explain: pip = Python ka app store — libraries install karo terminal se. pip install pandas — Pandas install ho jayega. pip install --upgrade pandas — latest version update karo. pip list — sab installed packages dekho. pip freeze > requirements.txt — sab packages aur versions ek file mein save karo — team ke saath share karo taaki sab same versions use karein. Interview mein "I manage dependencies using pip and requirements.txt for reproducible environments."
Q23: What is the difference between import and from...import?
Answer: import module imports the entire module — you access functions using module.function() syntax. from module import function imports specific functions directly — you use function() without the module prefix. from module import * imports everything (not recommended — can cause naming conflicts). Best practices: use import pandas as pd for standard aliases, from datetime import datetime for specific classes, avoid import * in production code.
🎯 Explain: import math → math.sqrt(16) — full module import, prefix lagana padta hai. from math import sqrt → sqrt(16) — direct use, prefix nahi chahiye. from math import * — sab import (risky — naam clash ho sakta hai). Best practice: import pandas as pd (standard alias), from datetime import datetime (specific class). Interview mein "I use standard aliases like pd for Pandas, np for NumPy, plt for Matplotlib" — industry standard follow karo.
# Standard imports for Data Analysis
import pandas as pd # Standard alias
import numpy as np # Standard alias
import matplotlib.pyplot as plt # Standard alias
from datetime import datetime # Specific import
# Usage
df = pd.DataFrame({"A": [1,2,3]})
arr = np.array([1,2,3])
now = datetime.now()
Q24: What are List Comprehensions and Dictionary Comprehensions?
Answer: Comprehensions provide concise, one-line syntax for creating collections. List Comprehension: [expression for item in iterable if condition]. Dictionary Comprehension: {key: value for item in iterable if condition}. Set Comprehension: {expression for item in iterable}. Comprehensions are more Pythonic, readable, and often faster than equivalent for loops. They are extensively used in data preprocessing — creating filtered lists, transforming data, and building lookup dictionaries.
🎯 Explain: Comprehension = ek line mein collection banao. List: [x*2 for x in range(5)] → [0,2,4,6,8]. With filter: [x for x in salaries if x>60000]. Dictionary: {name: len(name) for name in names}. Yeh for loop se zyada clean aur fast hai. Data Analysis mein: column names clean karo — [col.strip().lower() for col in df.columns]. Interview mein comprehension ka practical data cleaning example do — shows Pythonic coding style.
# List Comprehension
squares = [x**2 for x in range(1,6)] # [1,4,9,16,25]
# Filtered List Comprehension
salaries = [55000,72000,48000,80000]
high = [s for s in salaries if s > 60000] # [72000,80000]
# Dictionary Comprehension
names = ["Aarav","Ishita","Kabir"]
name_len = {n: len(n) for n in names}
print(name_len) # {'Aarav':5, 'Ishita':6, 'Kabir':5}
# Practical: Clean column names
cols = [" Name ", " SALARY", "Dept "]
clean = [c.strip().lower() for c in cols]
print(clean) # ['name', 'salary', 'dept']
🟢 Category 5: NumPy & Pandas Introduction (Q25–Q30)
Q25: What is NumPy and why is it important?
Answer: NumPy (Numerical Python) is a fundamental library for numerical computing in Python. It provides the ndarray (N-dimensional array) object which is faster and more memory-efficient than Python lists for numerical operations. NumPy supports vectorized operations (no explicit loops needed), broadcasting, linear algebra, random number generation, and mathematical functions. Pandas is built on top of NumPy — understanding NumPy arrays is essential for efficient data manipulation.
🎯 Explain: NumPy = Python mein fast numerical computing ka foundation. Python lists slow hain large data ke liye — NumPy arrays 50x faster hain kyunki C language mein internally implement hain. Vectorized operations — loop likhne ki zaroorat nahi — arr * 2 se saare elements double ho jaayenge ek shot mein. Pandas internally NumPy use karta hai. Interview mein "NumPy provides the performance backbone for Pandas — I use it for vectorized operations that avoid slow Python loops."
import numpy as np
# Create array
arr = np.array([55000, 72000, 48000, 80000])
# Vectorized operations — no loops!
print(arr * 0.10) # [5500. 7200. 4800. 8000.] — 10% bonus
print(arr > 60000) # [False True False True] — boolean mask
print(arr[arr > 60000]) # [72000 80000] — filtered array
# Statistics
print(np.mean(arr)) # 63750.0
print(np.std(arr)) # Standard deviation
print(np.median(arr)) # 63500.0
Q26: What is Pandas and what are its main data structures?
Answer: Pandas is the primary Python library for data manipulation and analysis. It provides two main data structures: (1) Series — a one-dimensional labeled array (like a single column). (2) DataFrame — a two-dimensional labeled table (like a spreadsheet or SQL table). DataFrames are the workhorse of data analysis in Python — they support reading CSV/Excel files, filtering, grouping, merging, pivoting, handling missing values, and exporting results. Pandas is built on NumPy for performance.
🎯 Explain: Pandas = Python ka Excel/SQL. Series = ek column. DataFrame = poori table (rows + columns). pd.read_csv("file.csv") se data load karo, df.head() se top 5 rows dekho, df.describe() se statistics dekho, df.groupby() se summarize karo. Excel mein jo Pivot Table karta hai woh Pandas mein groupby() karta hai. SQL mein jo JOIN karta hai woh Pandas mein merge() karta hai. Interview mein "Pandas is my primary tool for data analysis — it replaces Excel formulas and SQL queries in Python."
import pandas as pd
# Create DataFrame from dictionary
data = {
"Name": ["Aarav", "Ishita", "Kabir", "Diya"],
"Dept": ["IT", "HR", "Finance", "IT"],
"Salary": [55000, 72000, 65000, 58000]
}
df = pd.DataFrame(data)
print(df.head()) # First 5 rows
print(df.shape) # (4, 3) — 4 rows, 3 columns
print(df.dtypes) # Column data types
print(df.describe()) # Statistical summary
print(df.info()) # Column info + null counts
Q27: How do you select and filter data in Pandas?
Answer: Data selection methods: (1) Column selection — df['column'] or df.column (single), df[['col1','col2']] (multiple). (2) Row selection — df.loc[label] (label-based), df.iloc[index] (integer position-based). (3) Filtering — df[df['Salary'] > 60000] creates a boolean mask and returns matching rows. (4) Multiple conditions — use & (and), | (or) with parentheses. loc is label-based inclusive, iloc is position-based exclusive. These are the most fundamental Pandas operations for data analysis.
🎯 Explain: Column select: df['Name'] — ek column. df[['Name','Salary']] — multiple columns. Row select: df.loc[0] — label 0 ki row. df.iloc[0:3] — pehli 3 rows (position based). Filter: df[df['Salary'] > 60000] — salary 60000+ wale rows. Multiple conditions: df[(df['Dept']=="IT") & (df['Salary']>50000)] — IT + 50000+ salary. Note: & use karo aur parentheses lagao — and nahi chalega Pandas mein. Interview mein loc vs iloc difference clearly batao.
# Column selection
print(df['Name']) # Single column (Series)
print(df[['Name','Salary']]) # Multiple columns (DataFrame)
# Row selection
print(df.loc[0]) # Row by label
print(df.iloc[0:2]) # First 2 rows by position
# Filtering
high_sal = df[df['Salary'] > 60000]
print(high_sal)
# Multiple conditions (AND)
filtered = df[(df['Dept'] == "IT") & (df['Salary'] > 50000)]
print(filtered)
Q28: What is the difference between loc and iloc in Pandas?
Answer: loc is label-based indexing — it uses row labels (index names) and column names for selection. Slicing with loc is inclusive of both start and end. iloc is integer position-based indexing — it uses numerical positions (0, 1, 2...). Slicing with iloc is exclusive of the end position (like Python standard slicing). loc raises KeyError for missing labels, iloc raises IndexError for out-of-range positions. Use loc when you know the label/name, iloc when you know the position.
🎯 Explain: loc = naam se access karo. df.loc[0:2] — label 0 se 2 tak (inclusive — 3 rows). iloc = position se access karo. df.iloc[0:2] — position 0 se 2 tak (exclusive — 2 rows). loc mein column naam use hota hai — df.loc[0, 'Name']. iloc mein column number — df.iloc[0, 1]. Key difference: loc inclusive hai, iloc exclusive hai slicing mein. Interview mein "loc is label-based inclusive, iloc is position-based exclusive" — yeh one-liner yaad rakho.
Q29: How do you handle missing values in Pandas?
Answer: Missing values in Pandas are represented as NaN (Not a Number) or None. Key functions: isnull() / isna() — detect missing values. notnull() / notna() — detect non-missing values. dropna() — remove rows/columns with missing values. fillna() — replace missing values with specified value (mean, median, mode, forward fill, backward fill). sum() with isnull() — count missing values per column. Handling missing data is the most critical step in data cleaning and analysis.
🎯 Explain: Missing values = NaN — data mein kuch cells blank hain. df.isnull().sum() — har column mein kitne NaN hain. df.dropna() — NaN wali rows hatao. df.fillna(0) — NaN ko 0 se replace karo. df['Salary'].fillna(df['Salary'].mean()) — NaN ko average salary se bharo. Strategy depends on data — numerical mein mean/median, categorical mein mode ya "Unknown". Interview mein "I check missing values first with isnull().sum(), then decide strategy — dropna for few nulls, fillna with median for numerical columns."
# Check missing values
print(df.isnull().sum()) # Count NaN per column
# Drop rows with any NaN
df_clean = df.dropna()
# Fill NaN with mean
df['Salary'].fillna(df['Salary'].mean(), inplace=True)
# Fill NaN with forward fill
df.fillna(method='ffill', inplace=True)
# Check percentage of missing
print(df.isnull().sum() / len(df) * 100) # % missing
Q30: How do you read a CSV file and display basic information using Pandas?
Answer: pd.read_csv('filename.csv') loads a CSV file into a DataFrame. Essential exploration functions: head() — first 5 rows. tail() — last 5 rows. shape — (rows, columns) count. columns — column names list. dtypes — data types per column. info() — complete summary (columns, types, non-null counts, memory usage). describe() — statistical summary (count, mean, std, min, max, quartiles). value_counts() — frequency count for categorical columns. These functions are the first steps in any data analysis workflow.
🎯 Explain: CSV load karo: df = pd.read_csv("data.csv"). Pehla step hamesha data explore karna hota hai: df.head() — top 5 rows dekho. df.shape — kitne rows aur columns hain. df.info() — data types, null counts, memory — complete picture. df.describe() — numerical columns ki statistics (mean, std, min, max). df['Dept'].value_counts() — department-wise count. Interview mein "My first steps with any dataset are head(), shape, info(), describe(), and isnull().sum() — gives me a complete understanding of the data."
import pandas as pd
# Load CSV
df = pd.read_csv("employees.csv")
# Essential exploration — FIRST STEPS ALWAYS
print(df.head()) # First 5 rows
print(df.shape) # (rows, columns)
print(df.columns.tolist()) # Column names
print(df.dtypes) # Data types
print(df.info()) # Complete summary
print(df.describe()) # Statistics
print(df.isnull().sum()) # Missing values
print(df['Dept'].value_counts()) # Category counts
📋 Quick Revision Table — 30 Questions at a Glance
| Q# | Question | One-Line Answer |
|---|---|---|
| Q1 | What is Python? | High-level language — Pandas, NumPy for data analysis |
| Q2 | Variables? | Named containers — dynamically typed, no declaration needed |
| Q3 | Data types? | int, float, str, bool, None + list, tuple, dict, set |
| Q4 | Operators? | Arithmetic, Comparison, Logical, Membership (in), Identity (is) |
| Q5 | = vs ==? | = assigns value, == compares equality |
| Q6 | print() vs input()? | print=output, input=user input (always returns string) |
| Q7 | Strings? | Immutable text — upper, lower, strip, split, replace |
| Q8 | Lists? | Ordered, mutable, duplicates allowed — [] syntax |
| Q9 | List vs Tuple? | List=mutable [], Tuple=immutable () |
| Q10 | Dictionary? | Key-value pairs — {key:value} — JSON & DataFrame creation |
| Q11 | Set? | Unique values only — fast duplicate removal & membership |
| Q12 | List vs Tuple vs Set vs Dict? | Compare: ordered, mutable, duplicates, syntax |
| Q13 | if-elif-else? | Conditional branching — indentation-based blocks |
| Q14 | for vs while? | for=known iterations, while=condition-based unknown |
| Q15 | Functions? | Reusable code blocks — def, parameters, return |
| Q16 | *args vs **kwargs? | *args=positional tuple, **kwargs=keyword dictionary |
| Q17 | Lambda? | Anonymous one-line function — used with map, filter, apply |
| Q18 | range vs enumerate? | range=numbers only, enumerate=index+value pairs |
| Q19 | File handling? | with open() — auto-close, Pandas for CSV/Excel |
| Q20 | try-except? | Error handling — catch specific exceptions gracefully |
| Q21 | Module vs Package vs Library? | Module=file, Package=folder, Library=collection |
| Q22 | pip? | Package manager — install, update, remove libraries |
| Q23 | import vs from...import? | import=full module, from=specific function/class |
| Q24 | Comprehensions? | One-line collection creation — list, dict, set |
| Q25 | NumPy? | Fast numerical computing — vectorized operations, no loops |
| Q26 | Pandas? | Data manipulation library — Series & DataFrame |
| Q27 | Pandas filtering? | Boolean indexing — df[df['col'] > value] with & | operators |
| Q28 | loc vs iloc? | loc=label-based inclusive, iloc=position-based exclusive |
| Q29 | Missing values? | isnull(), dropna(), fillna() — check, remove, replace |
| Q30 | Read CSV & explore? | read_csv → head, shape, info, describe, isnull |
Thanks for Reading! 🙏
Thanks for reading! Data Insights par aur bhi Power BI, Excel, SQL, Python topics available hain — explore karo aur apni analytics journey strong banao! Happy Learning & Keep Exploring! 🚀
— JatinAnalytics
💬 Comments (0)
Loading comments...