<DataInsights />
  • 🏠 Home
  • πŸ“Š SQL
  • 🐍 Python
  • πŸ“ˆ Power BI
  • πŸ“— Excel
  • πŸ’Ό Career
  • 🎯 Interview Q&A
  • πŸ“ Case Study
  • πŸ“₯ Downloads
  • πŸš€ My Portfolio
<DataInsights />

Practical Data Analytics tutorials covering SQL, Python, Power BI, Excel and career guidance for aspiring analysts β€” 100% free.

Topics

  • SQL Tutorials
  • Python Guide
  • Power BI
  • Excel Tips
  • Career Guide

Quick Links

  • πŸ› οΈ All Tools
  • πŸ—“οΈ Archive
  • πŸ“¬ Contact
  • πŸ” Search
  • Portfolio
  • Kaggle
  • GitHub

Legal & Info

  • About
  • Contact
  • Privacy Policy
  • Disclaimer
  • Terms & Conditions
  • DMCA
  • Sitemap
Copyright © 2026 Data Insights by Jatin Kumar. All Rights Reserved.Built with ❀️ for Data Analysts
Home/Interview Q&A/Python Intermediate Interview Questions...

Python Intermediate Interview Questions

A
August 25, 2026 Jatin Kumar 40 min read Interview Q&A
Data Insights Python β€” Interview Preparation (Intermediate)

Python Intermediate Interview Questions 🟑

Top 30 intermediate Python interview questions for Data Analysts β€” Pandas Deep Dive, GroupBy, Merge, Pivot, Apply, Data Cleaning, NumPy Advanced, Matplotlib & Seaborn visualization. Real code examples with sample data. Data Insights par.

πŸ“‘ Is Blog Mein Kya Sikhenge:

  • 🟑 Q1–Q6: Pandas Deep Dive β€” DataFrame, Series, Indexing, Sorting
  • 🟑 Q7–Q12: Data Manipulation β€” GroupBy, Merge, Pivot, Apply, Map
  • 🟑 Q13–Q18: Data Cleaning β€” Missing Values, Duplicates, Type Conversion
  • 🟑 Q19–Q24: NumPy Advanced β€” Arrays, Broadcasting, Operations
  • 🟑 Q25–Q30: Data Visualization β€” Matplotlib & Seaborn
  • πŸ’‘ Pro Tips: Interview mein exactly kya bolna chahiye

πŸ“Š Sample Data β€” Is Blog Ke Examples Isi Par Based Hain

import pandas as pd
import numpy as np

df = pd.DataFrame({
    "EmpID": [101,102,103,104,105,106,107,108],
    "Name": ["Aarav","Ishita","Kabir","Diya","Rohan","Meera","Arjun","Kavya"],
    "Dept": ["IT","HR","Finance","IT","Marketing","HR","Finance","Marketing"],
    "Salary": [55000,72000,65000,58000,80000,48000,70000,62000],
    "Sales": [85000,92000,45000,78000,65000,52000,88000,71000],
    "Region": ["North","South","East","North","West","South","East","West"],
    "JoinDate": pd.to_datetime(["2021-01-01","2020-03-15","2019-07-22","2022-11-10",
                                  "2018-06-05","2023-09-18","2021-02-28","2022-08-14"])
})
EmpID Name Dept Salary Sales Region JoinDate
101AaravIT5500085000North2021-01-01
102IshitaHR7200092000South2020-03-15
103KabirFinance6500045000East2019-07-22
104DiyaIT5800078000North2022-11-10
105RohanMarketing8000065000West2018-06-05
106MeeraHR4800052000South2023-09-18
107ArjunFinance7000088000East2021-02-28
108KavyaMarketing6200071000West2022-08-14

🟑 Category 1: Pandas Deep Dive (Q1–Q6)

Q1: What is the difference between a Pandas Series and a DataFrame?
Answer: A Series is a one-dimensional labeled array β€” like a single column with an index. A DataFrame is a two-dimensional labeled table β€” a collection of Series sharing the same index, similar to a spreadsheet or SQL table. Series has one axis (index), DataFrame has two axes (index + columns). When you select a single column from a DataFrame, you get a Series. When you select multiple columns, you get a DataFrame. Both support label-based indexing, vectorized operations, and missing value handling.
🎯 Explain: Series = ek column. df['Salary'] β†’ Series return hota hai β€” ek list with index. DataFrame = poori table β€” multiple columns. df[['Name','Salary']] β†’ DataFrame return hota hai. Series pe .mean(), .sum() directly call kar sakte ho. DataFrame pe column specify karke β€” df['Salary'].mean(). Interview mein "Series is 1D like a column, DataFrame is 2D like a table β€” selecting one column returns Series, multiple returns DataFrame."

# Series β€” single column
sal_series = df['Salary']
print(type(sal_series))  # pandas.core.series.Series

# DataFrame β€” multiple columns
sub_df = df[['Name', 'Salary']]
print(type(sub_df))      # pandas.core.frame.DataFrame

# Series operations
print(sal_series.mean())   # 63750.0
print(sal_series.max())    # 80000
print(sal_series.value_counts())  # Frequency count

Q2: How do you add, rename, and delete columns in a DataFrame?
Answer: Add column: df['new_col'] = values or df.assign(new_col=values). Rename: df.rename(columns={'old':'new'}) or df.columns = ['list','of','names']. Delete: df.drop(columns=['col']) or del df['col'] or df.pop('col'). drop() with inplace=True modifies the original DataFrame. When adding calculated columns, use vectorized operations β€” df['Bonus'] = df['Salary'] * 0.10 β€” much faster than loops.
🎯 Explain: Add: df['Bonus'] = df['Salary'] * 0.10 β€” naya column ban jayega. Rename: df.rename(columns={'Dept':'Department'}) β€” column naam change. Delete: df.drop(columns=['Bonus']) β€” column hatao. Tip: drop() default mein original change nahi karta β€” inplace=True lagao ya result save karo. Interview mein "I add calculated columns using vectorized operations and use rename() for standardizing column names."

# Add column
df['Bonus'] = df['Salary'] * 0.10

# Rename columns
df = df.rename(columns={'Dept': 'Department'})

# Delete column
df = df.drop(columns=['Bonus'])

# Conditional column
df['SalaryLevel'] = np.where(df['Salary'] > 60000, 'High', 'Low')

# Multiple conditions
conditions = [df['Salary']>=75000, df['Salary']>=60000, df['Salary']<60000]
choices = ['Executive', 'Senior', 'Mid-Level']
df['Level'] = np.select(conditions, choices, default='Junior')

Q3: How do you sort data in Pandas?
Answer: sort_values() sorts by column values β€” df.sort_values('Salary', ascending=False) for descending. Multiple columns: df.sort_values(['Dept','Salary'], ascending=[True,False]). sort_index() sorts by row index. nlargest(n, 'col') returns top N rows by column β€” faster than full sort for large data. nsmallest(n, 'col') for bottom N. All return new DataFrames by default β€” use inplace=True to modify original.
🎯 Explain: sort_values = column ke basis pe sort karo. df.sort_values('Salary', ascending=False) β€” highest salary first. Multiple columns: pehle Dept alphabetically, phir Salary highest first within each dept. nlargest(3, 'Salary') β€” top 3 salaries directly β€” full sort se fast hai large data pe. Interview mein "I use nlargest/nsmallest for top-N analysis instead of full sort β€” better performance on large datasets."

# Sort by salary descending
print(df.sort_values('Salary', ascending=False))

# Multi-column sort
print(df.sort_values(['Dept','Salary'], ascending=[True,False]))

# Top 3 salaries β€” faster than full sort
print(df.nlargest(3, 'Salary'))
# Rohan 80000, Ishita 72000, Arjun 70000

Q4: What is the difference between copy() and assignment in Pandas?
Answer: Direct assignment (df2 = df) creates a reference β€” both variables point to the same data, so changes to df2 also affect df. copy() creates an independent copy β€” df2 = df.copy() β€” changes to df2 do not affect df. This is crucial for data analysis β€” always use copy() when you want to modify a subset without affecting the original. Pandas also shows SettingWithCopyWarning when modifying a slice β€” using copy() avoids this warning.
🎯 Explain: df2 = df β€” yeh copy nahi hai, reference hai β€” df2 change karo toh df bhi change hoga. df2 = df.copy() β€” independent copy β€” safe hai. SettingWithCopyWarning bahut common warning hai β€” jab filtered DataFrame modify karte ho. Solution: filtered = df[df['Dept']=="IT"].copy() β€” pehle copy karo, phir modify karo. Interview mein "I always use .copy() when creating subsets to avoid unintended modifications to the original DataFrame."

# Reference (dangerous)
df2 = df
df2['Salary'] = 0  # ⚠️ df also changes!

# Safe copy
df2 = df.copy()
df2['Salary'] = 0  # βœ… df remains unchanged

# Filtered subset β€” always copy
it_team = df[df['Dept'] == 'IT'].copy()

Q5: How do you set and reset index in Pandas?
Answer: set_index('column') makes a column the row index β€” df.set_index('EmpID'). This enables faster label-based lookups with loc. reset_index() converts the index back to a regular column and creates a default integer index. reset_index(drop=True) discards the old index completely. Setting meaningful indexes is important for efficient data access, merging operations, and time-series analysis (DatetimeIndex).
🎯 Explain: set_index('EmpID') β€” EmpID ko row label banao β€” ab df.loc[103] se directly Kabir ki row access hogi. reset_index() β€” index wapas column ban jayega. Time series mein date ko index set karna standard practice hai β€” df.set_index('JoinDate'). GroupBy ke baad reset_index() zaroor karo β€” grouped result ka index clean ho jaaye. Interview mein "I set meaningful indexes for faster lookups and always reset_index after GroupBy operations."

# Set index
df_indexed = df.set_index('EmpID')
print(df_indexed.loc[103])  # Kabir's row directly

# Reset index
df_reset = df_indexed.reset_index()  # EmpID back as column

Q6: What is the difference between apply(), map(), and applymap()?
Answer: map() works on a Series only β€” applies a function or mapping dictionary element-wise. apply() works on both Series and DataFrame β€” for Series it applies function element-wise, for DataFrame it applies function along an axis (row or column). applymap() (renamed to map() in Pandas 2.1+) applies function element-wise to every cell in a DataFrame. Use map() for Series transformations, apply() for row/column-level operations, applymap() for cell-level operations across the entire DataFrame.
🎯 Explain: map() = Series pe β€” df['Name'].map(str.upper). apply() = Series ya DataFrame pe β€” df['Salary'].apply(lambda x: x*1.1) ya df.apply(sum, axis=0). applymap() = poore DataFrame ke har cell pe β€” df.applymap(str). Most used: apply() with lambda β€” df['Tax'] = df['Salary'].apply(lambda x: x*0.30 if x>75000 else x*0.10). Interview mein "I use apply with lambda for conditional column calculations and map for dictionary-based value replacement."

# map β€” Series only (dictionary mapping)
region_map = {'North':'N', 'South':'S', 'East':'E', 'West':'W'}
df['RegionCode'] = df['Region'].map(region_map)

# apply β€” Series with lambda
df['Tax'] = df['Salary'].apply(lambda x: x*0.30 if x>70000 else x*0.10)

# apply β€” DataFrame row-wise
df['Total'] = df.apply(lambda row: row['Salary'] + row['Sales'], axis=1)
πŸ’‘ Pro Tip: Pandas Deep Dive ka question aaye toh np.where aur np.select mention karo: "For simple binary conditions I use np.where β€” df['Level'] = np.where(df['Salary']>60000, 'High', 'Low'). For multiple conditions I use np.select with conditions list β€” cleaner than nested np.where or multiple apply calls." Yeh vectorized approach hai β€” loops se 100x faster β€” performance awareness dikhata hai.

🟑 Category 2: Data Manipulation (Q7–Q12)

Q7: How does GroupBy work in Pandas?
Answer: GroupBy splits data into groups based on column values, applies an aggregation function to each group, and combines results. Syntax: df.groupby('column')['value_column'].agg_function(). Supports multiple aggregations: agg({'Salary':['mean','sum'], 'Sales':'max'}). Multiple grouping columns: df.groupby(['Dept','Region']). After GroupBy, use reset_index() to convert the grouped result back to a regular DataFrame. GroupBy is equivalent to SQL GROUP BY and Excel Pivot Tables.
🎯 Explain: GroupBy = SQL ka GROUP BY, Excel ka Pivot Table. df.groupby('Dept')['Salary'].mean() β€” department-wise average salary. Multiple agg: .agg({'Salary':'mean', 'Sales':'sum'}) β€” ek mein mean, dusre mein sum. reset_index() karo baad mein taaki clean DataFrame mile. GroupBy bahut powerful hai β€” pivot table jaisa kaam ek line mein. Interview mein "GroupBy is my primary tool for aggregation β€” equivalent to SQL GROUP BY and Excel Pivot Tables."

# Simple GroupBy
print(df.groupby('Dept')['Salary'].mean())

# Multiple aggregations
summary = df.groupby('Dept').agg({
    'Salary': ['mean', 'sum', 'count'],
    'Sales': ['sum', 'max']
}).reset_index()

# Multi-level GroupBy
print(df.groupby(['Dept','Region'])['Salary'].mean().reset_index())

# Named aggregation (cleaner)
result = df.groupby('Dept').agg(
    AvgSalary=('Salary', 'mean'),
    TotalSales=('Sales', 'sum'),
    HeadCount=('EmpID', 'count')
).reset_index()

Q8: How do you merge/join DataFrames in Pandas?
Answer: pd.merge() combines DataFrames based on common columns β€” similar to SQL JOIN. Syntax: pd.merge(left, right, on='key', how='inner'). Join types: inner (default β€” matching only), left (all left + matching right), right (all right + matching left), outer (all from both). When key column names differ: left_on='col1', right_on='col2'. For index-based joins: left_index=True, right_index=True. concat() stacks DataFrames vertically (like SQL UNION) or horizontally.
🎯 Explain: merge() = SQL JOIN. pd.merge(sales, products, on='ProductID', how='left') β€” Sales ke saath Products join karo. inner = sirf matching. left = sab left wale + matching right. outer = dono ke sab. Different column names: left_on='EmpID', right_on='EmployeeID'. pd.concat([df1, df2]) β€” vertically stack karo (UNION). Interview mein "merge for horizontal joins with how parameter, concat for vertical stacking β€” equivalent to SQL JOIN and UNION."

# Create second DataFrame for merge
dept_info = pd.DataFrame({
    'Dept': ['IT','HR','Finance','Marketing'],
    'Manager': ['Raj','Priya','Vikram','Neha']
})

# Left join
merged = pd.merge(df, dept_info, on='Dept', how='left')

# Concat β€” stack vertically
df_combined = pd.concat([df1, df2], ignore_index=True)

Q9: What is pivot_table() in Pandas?
Answer: pivot_table() creates a spreadsheet-style pivot table from a DataFrame. Syntax: df.pivot_table(values='Salary', index='Dept', columns='Region', aggfunc='mean'). It groups data by index and columns, then applies the aggregation function to values. Supports multiple aggfuncs: aggfunc=['mean','sum','count']. fill_value=0 replaces NaN. margins=True adds row/column totals. It is the Pandas equivalent of Excel Pivot Tables β€” powerful for cross-tabulation analysis.
🎯 Explain: pivot_table = Excel Pivot Table ka Python version. df.pivot_table(values='Salary', index='Dept', columns='Region', aggfunc='mean') β€” Department rows mein, Region columns mein, Salary ka average values mein. margins=True se Grand Total aayega. fill_value=0 se NaN ki jagah 0. Multiple values aur aggfuncs support karta hai. Interview mein "pivot_table is my go-to for cross-tabulation β€” it replicates Excel Pivot Tables in Python."

# Pivot table β€” Dept vs Region, Avg Salary
pivot = df.pivot_table(
    values='Salary',
    index='Dept',
    columns='Region',
    aggfunc='mean',
    fill_value=0,
    margins=True
)
print(pivot)

Q10: What is the difference between merge(), join(), and concat()?
Answer: merge() combines DataFrames based on column values β€” like SQL JOIN. join() combines DataFrames based on their index β€” df1.join(df2). concat() stacks DataFrames vertically (axis=0, like UNION) or horizontally (axis=1, like column binding). Key differences: merge uses column keys, join uses index, concat uses positional alignment. merge is most flexible with how parameter for join types. In practice, merge() is used most frequently for relational data combinations.
🎯 Explain: merge = column-based join (SQL JOIN). join = index-based join. concat = stack karo (vertically ya horizontally). 90% kaam merge se hota hai β€” on='key', how='left/inner/outer'. concat tab use hota hai jab multiple DataFrames ko stack karna ho β€” monthly files combine karo. join tab use hota hai jab index pe join karna ho. Interview mein "merge for column-based joins, concat for stacking, join for index-based β€” I primarily use merge."

Q11: How does the transform() function work in Pandas?
Answer: transform() applies a function to each group and returns a result with the same shape as the input β€” unlike agg() which reduces data. It is used with GroupBy to add group-level calculations back to individual rows. Example: df['DeptAvgSalary'] = df.groupby('Dept')['Salary'].transform('mean') adds department average salary as a new column for every row while maintaining row-level detail. This enables calculations like "how much above/below department average" without separate merge.
🎯 Explain: transform = GroupBy ka result har row mein wapas daal do. GroupBy('Dept')['Salary'].mean() β†’ 4 values (ek per dept). transform('mean') β†’ 8 values (har row ke liye uski dept ka average). Bahut powerful hai β€” employee ki salary uski dept average se compare karo bina merge ke. df['AboveAvg'] = df['Salary'] - df.groupby('Dept')['Salary'].transform('mean'). Interview mein "I use transform for adding group-level metrics back to row-level data β€” like department averages alongside individual salaries."

# transform β€” add dept average to each row
df['DeptAvg'] = df.groupby('Dept')['Salary'].transform('mean')

# How much above/below dept average
df['VsAvg'] = df['Salary'] - df['DeptAvg']

# Percentage of department total
df['PctOfDept'] = df['Salary'] / df.groupby('Dept')['Salary'].transform('sum') * 100

Q12: What is the melt() function and when do you use it?
Answer: melt() transforms wide-format data into long-format (tall) data β€” the reverse of pivot. It unpivots columns into rows. Syntax: pd.melt(df, id_vars=['ID'], value_vars=['Q1','Q2','Q3'], var_name='Quarter', value_name='Sales'). id_vars are columns to keep as identifiers, value_vars are columns to unpivot. melt() is essential for preparing data for visualization libraries (Seaborn expects long format), time-series analysis, and normalizing cross-tab reports.
🎯 Explain: melt = wide data ko tall banao (Excel ka Unpivot). Products rows mein hain, Q1, Q2, Q3 columns mein β€” melt karo toh Quarter aur Sales columns ban jayenge. Seaborn ke liye long format chahiye β€” melt se convert karo. Cross-tab report mil rahi hai database se β€” analysis ke liye melt karo. Interview mein "I use melt for converting wide cross-tab data to long format β€” essential for Seaborn visualization and consistent data structure."

# Wide format data
wide = pd.DataFrame({
    'Name': ['Aarav','Ishita'],
    'Q1': [85000,92000],
    'Q2': [78000,88000],
    'Q3': [91000,95000]
})

# Melt to long format
long = pd.melt(wide, id_vars=['Name'], 
              value_vars=['Q1','Q2','Q3'],
              var_name='Quarter', 
              value_name='Sales')
print(long)
# Name  Quarter  Sales
# Aarav   Q1     85000
# Ishita  Q1     92000
# Aarav   Q2     78000 ...
πŸ’‘ Pro Tip: GroupBy aur merge ka question aaye toh practical workflow batao: "For summarization I use groupby with named aggregation β€” .agg(AvgSalary=('Salary','mean')). For enriching data I use merge with how='left'. For adding group metrics to row-level data I use transform β€” it's like a window function in SQL. For reshaping, melt converts wide to long and pivot_table converts long to wide." SQL parallels draw karo β€” interviewer ko lagta hai ki tum cross-tool thinking karte ho.

🟑 Category 3: Data Cleaning (Q13–Q18)

Q13: How do you detect and handle duplicate rows?
Answer: duplicated() returns a Boolean Series marking duplicate rows (True for duplicates). drop_duplicates() removes duplicate rows. Parameters: subset=['col1','col2'] checks duplicates based on specific columns only. keep='first' (default β€” keeps first occurrence), 'last' (keeps last), False (removes all duplicates). Always check duplicates before analysis β€” they can inflate counts, sums, and averages. For investigating duplicates: df[df.duplicated(keep=False)] shows all duplicate rows including originals.
🎯 Explain: df.duplicated().sum() β€” kitne duplicate rows hain. df.drop_duplicates() β€” duplicates hatao. subset=['EmpID'] β€” sirf EmpID ke basis pe check karo. keep='first' pehla rakho baaki hatao. Duplicates pehle investigate karo β€” df[df.duplicated(subset=['EmpID'], keep=False)] β€” sab duplicates dekho with originals β€” samjho kyu duplicate hain. Interview mein "I always investigate duplicates before removing β€” sometimes duplicates indicate data quality issues that need root cause analysis."

# Check duplicates
print(df.duplicated().sum())        # Count of duplicate rows

# Show all duplicates (including originals)
print(df[df.duplicated(keep=False)])

# Remove duplicates
df_clean = df.drop_duplicates()

# Duplicates based on specific column
df_clean = df.drop_duplicates(subset=['EmpID'], keep='last')

Q14: How do you handle missing values strategically?
Answer: Strategy depends on data type and percentage missing: (1) Less than 5% β€” dropna() is acceptable. (2) Numerical columns β€” fillna with mean (normal distribution) or median (skewed distribution). (3) Categorical columns β€” fillna with mode or "Unknown". (4) Time series β€” fillna with method='ffill' (forward fill) or 'bfill' (backward fill). (5) High percentage missing (>50%) β€” consider dropping the column. (6) interpolate() for numerical trends. Always analyze the pattern β€” MCAR (random), MAR, or MNAR (systematic) β€” before deciding strategy.
🎯 Explain: Missing values ki strategy data pe depend karti hai. Pehle check karo: df.isnull().sum() / len(df) * 100 β€” percentage dekho. <5% hai toh dropna(). Numerical mein median better hai mean se (outliers affect nahi karte). Categorical mein mode ya "Unknown". Time series mein forward fill (previous value carry forward). >50% missing toh column hi hata do. Interview mein "I analyze the missing pattern first β€” if random and under 5%, I drop. For numerical, I prefer median over mean to handle outliers. For categorical, I fill with mode or a placeholder."

# Missing value analysis
print(df.isnull().sum())                    # Count per column
print(df.isnull().sum() / len(df) * 100)  # Percentage

# Strategic filling
df['Salary'].fillna(df['Salary'].median(), inplace=True)  # Numerical
df['Dept'].fillna('Unknown', inplace=True)             # Categorical
df['Sales'].fillna(method='ffill', inplace=True)       # Forward fill

# Group-wise fill (dept average for missing salary)
df['Salary'] = df.groupby('Dept')['Salary'].transform(
    lambda x: x.fillna(x.median())
)

Q15: How do you change data types in Pandas?
Answer: astype() converts column data types β€” df['col'].astype(int), astype(str), astype('category'). pd.to_numeric() converts to numeric with error handling β€” errors='coerce' turns invalid values to NaN. pd.to_datetime() converts to datetime. Category type saves memory for repeated string values. Common conversions: string to numeric (imported data), object to datetime, string to category. Always check dtypes after loading data β€” wrong types cause calculation errors.
🎯 Explain: CSV se data load hota hai toh aksar types galat aate hain β€” numbers string mein, dates object mein. astype(int) β€” integer banao. pd.to_numeric(df['col'], errors='coerce') β€” non-numeric values NaN ban jayengi instead of error. pd.to_datetime(df['Date']) β€” datetime banao. astype('category') β€” repeated strings ke liye memory save hota hai. Interview mein "After loading any CSV, I always check dtypes and convert β€” to_numeric for numbers, to_datetime for dates, category for categorical columns."

# Check current types
print(df.dtypes)

# Convert types
df['EmpID'] = df['EmpID'].astype(str)         # to string
df['Salary'] = pd.to_numeric(df['Salary'], errors='coerce')
df['JoinDate'] = pd.to_datetime(df['JoinDate'])
df['Dept'] = df['Dept'].astype('category')   # saves memory

Q16: How do you clean string/text data in Pandas?
Answer: Pandas provides .str accessor for vectorized string operations on Series. Key methods: str.strip() removes whitespace, str.lower()/upper() for case, str.replace() for substitution, str.contains() for pattern matching, str.split() for splitting, str.extract() for regex extraction, str.len() for length. Common cleaning tasks: standardize names (strip+title), remove special characters, extract parts from combined fields, and normalize categories. Always clean text data before grouping or merging β€” "IT" and " IT " are different to Pandas.
🎯 Explain: .str accessor = Series pe string operations. df['Name'].str.strip() β€” spaces hatao. df['Name'].str.lower() β€” lowercase. df['Dept'].str.contains('IT') β€” IT wale rows filter. df['Email'].str.split('@').str[1] β€” domain extract. Data mein "IT", " IT", "it " β€” yeh sab different hain Pandas ke liye β€” pehle clean karo: df['Dept'] = df['Dept'].str.strip().str.upper(). Interview mein "I always standardize text columns with strip() and consistent casing before any groupby or merge operation."

# String cleaning pipeline
df['Name'] = df['Name'].str.strip().str.title()
df['Dept'] = df['Dept'].str.strip().str.upper()

# Pattern matching
it_emp = df[df['Dept'].str.contains('IT', case=False, na=False)]

# Replace values
df['Region'] = df['Region'].str.replace('North', 'N')

# Clean column names
df.columns = df.columns.str.strip().str.lower().str.replace(' ', '_')

Q17: How do you detect and handle outliers?
Answer: Outlier detection methods: (1) IQR Method β€” Q1 = 25th percentile, Q3 = 75th percentile, IQR = Q3-Q1. Outliers are below Q1-1.5*IQR or above Q3+1.5*IQR. (2) Z-Score β€” values with z-score above 3 or below -3 are outliers. (3) Visual β€” box plots and histograms. Handling: (1) Remove outliers if they are data errors. (2) Cap/clip at boundaries (winsorization). (3) Transform using log/sqrt. (4) Keep if they represent genuine extreme values. Decision depends on domain knowledge.
🎯 Explain: IQR Method sabse common hai. Q1 = df['Salary'].quantile(0.25), Q3 = quantile(0.75), IQR = Q3-Q1. Lower bound = Q1 - 1.5*IQR, Upper bound = Q3 + 1.5*IQR. Bounds ke bahar = outlier. Z-score: (value - mean) / std β€” 3 se zyada toh outlier. Handling: remove karo agar error hai, clip karo agar extreme but valid hai. Interview mein "I use IQR method for outlier detection and decide handling based on domain context β€” sometimes outliers are genuine high performers."

# IQR Method
Q1 = df['Salary'].quantile(0.25)
Q3 = df['Salary'].quantile(0.75)
IQR = Q3 - Q1
lower = Q1 - 1.5 * IQR
upper = Q3 + 1.5 * IQR

outliers = df[(df['Salary'] < lower) | (df['Salary'] > upper)]

# Clip outliers (winsorize)
df['Salary_clipped'] = df['Salary'].clip(lower=lower, upper=upper)

Q18: What is the difference between replace() and map() for value substitution?
Answer: replace() substitutes specific values β€” works on both Series and DataFrame. df['Dept'].replace('IT', 'Information Technology') or df.replace({'IT':'InfoTech', 'HR':'Human Resources'}). map() applies a function or dictionary to a Series β€” replaces ALL values based on the mapping, unmapped values become NaN. Key difference: replace() keeps unmapped values unchanged, map() turns unmapped values to NaN. Use replace() for selective substitution, map() for complete remapping.
🎯 Explain: replace() = specific values change karo, baaki unchanged. map() = dictionary se sab map karo, jo dictionary mein nahi woh NaN ban jayega. df['Dept'].replace({'IT':'Tech'}) β€” sirf IT change hoga, HR/Finance same rahenge. df['Dept'].map({'IT':'Tech', 'HR':'People'}) β€” ITβ†’Tech, HRβ†’People, Financeβ†’NaN (not in mapping). Interview mein "replace for selective changes keeping other values, map for complete column remapping β€” I prefer replace for partial updates."

πŸ’‘ Pro Tip: Data Cleaning ka question aaye toh structured pipeline batao: "My cleaning pipeline: 1) Check shape, dtypes, info. 2) Handle missing values β€” isnull().sum(), strategy based on percentage and type. 3) Remove duplicates with investigation. 4) Fix data types β€” to_numeric, to_datetime. 5) Clean text β€” strip, lower, standardize. 6) Detect outliers with IQR. 7) Validate with describe() and value_counts()." Yeh 7-step pipeline professional data cleaning approach dikhata hai.

🟑 Category 4: NumPy Advanced (Q19–Q24)

Q19: What is the difference between a Python List and a NumPy Array?
Answer: Python Lists are general-purpose containers that can hold mixed data types, are stored as separate objects in memory, and require explicit loops for element-wise operations. NumPy Arrays are homogeneous (single data type), stored in contiguous memory blocks, and support vectorized operations without loops. NumPy arrays are 10-100x faster than lists for numerical operations due to C-level optimizations, fixed data types, and memory efficiency. NumPy uses less memory because it stores raw data without Python object overhead.
🎯 Explain: List = slow, flexible, mixed types. NumPy Array = fast, fixed type, contiguous memory. List mein [1,2,3] * 2 = [1,2,3,1,2,3] (repeat). Array mein np.array([1,2,3]) * 2 = [2,4,6] (element-wise multiply). List mein loop lagana padta hai har element pe operation ke liye β€” NumPy mein directly array pe operation karo. 1 million numbers pe sum karna ho β€” List loop 100ms, NumPy 1ms. Interview mein "NumPy arrays are 50-100x faster than lists due to contiguous memory and vectorized C operations."

import numpy as np

# List vs Array behavior
py_list = [1, 2, 3]
np_arr = np.array([1, 2, 3])

print(py_list * 2)   # [1,2,3,1,2,3] β€” repetition
print(np_arr * 2)    # [2,4,6] β€” element-wise multiply

# Speed comparison
big_list = list(range(1000000))
big_arr = np.arange(1000000)
# np.sum(big_arr) is ~100x faster than sum(big_list)

# Memory comparison
import sys
print(sys.getsizeof(big_list))  # ~8MB
print(big_arr.nbytes)          # ~4MB β€” half the memory

Q20: What is Broadcasting in NumPy?
Answer: Broadcasting is NumPy's ability to perform operations on arrays of different shapes by automatically expanding the smaller array to match the larger one β€” without actually copying data. Rules: (1) If arrays have different dimensions, the smaller one is padded with 1s on the left. (2) Arrays with size 1 along a dimension are stretched to match the other array. (3) If sizes do not match and neither is 1, an error occurs. Broadcasting enables concise, loop-free code for operations between arrays of different sizes.
🎯 Explain: Broadcasting = alag size ke arrays pe operation karo bina loop ke. np.array([1,2,3]) + 10 β†’ [11,12,13] β€” scalar 10 automatically broadcast ho gaya har element pe. 2D array (3Γ—3) + 1D array (3,) β†’ 1D array har row pe broadcast hoga. Yeh loop likhne se bachata hai β€” fast aur clean. Pandas internally broadcasting use karta hai β€” df['Salary'] * 0.10 mein 0.10 broadcast hota hai. Interview mein "Broadcasting eliminates explicit loops β€” a scalar or smaller array is automatically expanded to match dimensions."

# Scalar broadcasting
arr = np.array([55000, 72000, 48000])
print(arr * 1.10)  # [60500. 79200. 52800.] β€” 10% raise

# 2D + 1D broadcasting
matrix = np.array([[1,2,3],
                   [4,5,6]])
row = np.array([10, 20, 30])
print(matrix + row)
# [[11, 22, 33],
#  [14, 25, 36]]  β€” row added to each row of matrix

Q21: How do you create different types of NumPy arrays?
Answer: Array creation functions: np.array([list]) β€” from Python list. np.zeros((rows,cols)) β€” all zeros. np.ones((rows,cols)) β€” all ones. np.full((rows,cols), value) β€” filled with specific value. np.arange(start,stop,step) β€” evenly spaced values (like range). np.linspace(start,stop,num) β€” exactly num evenly spaced values. np.random.rand(rows,cols) β€” random values 0-1. np.random.randint(low,high,size) β€” random integers. np.eye(n) β€” identity matrix. np.empty((rows,cols)) β€” uninitialized array (fast allocation).
🎯 Explain: np.zeros((3,4)) β€” 3Γ—4 zero matrix. np.arange(0,10,2) β†’ [0,2,4,6,8]. np.linspace(0,1,5) β†’ [0, 0.25, 0.5, 0.75, 1.0] β€” exactly 5 values. np.random.randint(1,100,10) β€” 10 random integers 1-99. Testing aur prototyping mein bahut use hota hai β€” dummy data generate karo. Interview mein "I use np.random for generating test data and np.linspace for creating evenly spaced ranges for plotting."

# Common array creation
print(np.zeros((2,3)))          # 2x3 zeros
print(np.ones((3,3)))           # 3x3 ones
print(np.arange(0, 10, 2))      # [0, 2, 4, 6, 8]
print(np.linspace(0, 1, 5))    # [0. 0.25 0.5 0.75 1.]
print(np.random.randint(1,100,5)) # 5 random ints 1-99
print(np.eye(3))                # 3x3 identity matrix

Q22: What are np.where() and np.select() and how are they used in data analysis?
Answer: np.where(condition, value_if_true, value_if_false) is a vectorized if-else β€” applies condition to entire array/Series at once. Equivalent to Excel IF function. np.select(conditions_list, choices_list, default) handles multiple conditions β€” equivalent to nested IF or IFS. Both are much faster than apply() with lambda for conditional column creation in Pandas. np.where for 2 outcomes, np.select for 3+ outcomes.
🎯 Explain: np.where = vectorized IF β€” df['Level'] = np.where(df['Salary']>60000, 'High', 'Low'). Ek line mein poora column create β€” apply+lambda se 10x faster. np.select = multiple conditions β€” conditions = [sal>=75000, sal>=60000, sal<60000], choices = ['Executive','Senior','Mid']. Default value bhi set karo. Interview mein "I prefer np.where over apply+lambda for binary conditions and np.select for multiple conditions β€” both are vectorized and significantly faster."

# np.where β€” binary condition
df['SalaryLevel'] = np.where(df['Salary'] > 60000, 'High', 'Low')

# np.select β€” multiple conditions
conditions = [
    df['Salary'] >= 75000,
    df['Salary'] >= 60000,
    df['Salary'] < 60000
]
choices = ['Executive', 'Senior', 'Mid-Level']
df['Level'] = np.select(conditions, choices, default='Junior')

Q23: What is array reshaping and when do you use it?
Answer: Reshaping changes the dimensions of an array without changing its data. reshape(rows, cols) β€” total elements must match. flatten() β€” converts any dimensional array to 1D (returns copy). ravel() β€” similar to flatten but returns a view (no copy). transpose() or .T β€” swaps rows and columns. Use -1 in reshape to auto-calculate one dimension β€” arr.reshape(3, -1). Reshaping is essential for preparing data for machine learning models, matrix operations, and data structure transformations.
🎯 Explain: reshape = array ka shape change karo β€” data same rahega. np.arange(12).reshape(3,4) β€” 1D array of 12 elements β†’ 3Γ—4 matrix. reshape(-1,1) β€” column vector banao (ML mein input ke liye). flatten() β€” kisi bhi shape ko 1D banao. .T β€” transpose β€” rows ↔ columns swap. Interview mein "I use reshape for preparing features for ML models and flatten for converting multi-dimensional results to 1D."

# Reshape
arr = np.arange(12)
print(arr.reshape(3, 4))    # 3 rows, 4 cols
print(arr.reshape(4, -1))   # 4 rows, auto cols (3)
print(arr.reshape(-1, 1))   # Column vector (12,1)

# Flatten & Transpose
matrix = arr.reshape(3,4)
print(matrix.flatten())     # Back to 1D
print(matrix.T)              # Transpose (4,3)

Q24: What are common NumPy statistical functions used in data analysis?
Answer: Key statistical functions: np.mean() β€” average. np.median() β€” middle value. np.std() β€” standard deviation. np.var() β€” variance. np.percentile(arr, q) β€” percentile value. np.corrcoef(x, y) β€” correlation matrix. np.cumsum() β€” cumulative sum. np.cumprod() β€” cumulative product. np.unique(arr, return_counts=True) β€” unique values with frequencies. These functions work on NumPy arrays and Pandas Series β€” forming the basis of descriptive statistics in data analysis.
🎯 Explain: np.mean() = average. np.median() = middle value (outlier resistant). np.std() = standard deviation (spread measure). np.percentile(arr, 75) = 75th percentile. np.corrcoef(salary, sales) = correlation between two variables. np.cumsum() = running total β€” useful for cumulative analysis. Pandas internally yeh sab NumPy se karta hai β€” df['Salary'].mean() internally np.mean() call karta hai. Interview mein "I use NumPy statistical functions for quick analysis and Pandas describe() for comprehensive summary."

salaries = np.array([55000,72000,65000,58000,80000,48000,70000,62000])

print(f"Mean: {np.mean(salaries)}")        # 63750.0
print(f"Median: {np.median(salaries)}")    # 63500.0
print(f"Std: {np.std(salaries):.2f}")      # Standard deviation
print(f"25th: {np.percentile(salaries,25)}")  # Q1
print(f"75th: {np.percentile(salaries,75)}")  # Q3

# Cumulative sum
print(np.cumsum(salaries))  # Running total

# Correlation
sales = np.array([85000,92000,45000,78000,65000,52000,88000,71000])
print(np.corrcoef(salaries, sales))  # Correlation matrix
πŸ’‘ Pro Tip: NumPy ka question aaye toh performance angle se answer do: "I use np.where and np.select instead of apply+lambda for conditional columns β€” they're vectorized and 10x faster on large datasets. For statistical analysis, I use NumPy functions directly or through Pandas describe(). Broadcasting eliminates explicit loops β€” a key performance optimization." Vectorization awareness = senior-level thinking.

🟑 Category 5: Data Visualization β€” Matplotlib & Seaborn (Q25–Q30)

Q25: What is Matplotlib and how do you create basic charts?
Answer: Matplotlib is Python's foundational plotting library β€” most other visualization libraries are built on top of it. The pyplot module (imported as plt) provides a MATLAB-like interface for creating charts. Basic workflow: create figure β†’ plot data β†’ customize (title, labels, legend) β†’ show/save. Key chart types: plt.plot() for line, plt.bar() for bar, plt.scatter() for scatter, plt.hist() for histogram, plt.pie() for pie. plt.figure(figsize=(width,height)) sets chart size. plt.savefig('chart.png') saves to file.
🎯 Explain: Matplotlib = Python ka OG charting library. import matplotlib.pyplot as plt. plt.bar(x, y) β€” bar chart. plt.title('Title') β€” title lagao. plt.xlabel('X') β€” x-axis label. plt.show() β€” chart dikhao. Jupyter Notebook mein %matplotlib inline lagao β€” inline charts dikhenge. Sabse basic hai β€” Seaborn iske upar built hai. Interview mein "Matplotlib for full customization control, Seaborn for statistical charts with better defaults."

import matplotlib.pyplot as plt

# Bar Chart β€” Department-wise average salary
dept_avg = df.groupby('Dept')['Salary'].mean()

plt.figure(figsize=(8,5))
plt.bar(dept_avg.index, dept_avg.values, color='steelblue')
plt.title('Average Salary by Department', fontsize=14)
plt.xlabel('Department')
plt.ylabel('Average Salary')
plt.tight_layout()
plt.show()

# Line Chart
plt.plot(df['Name'], df['Sales'], marker='o', color='green')
plt.title('Sales by Employee')
plt.xticks(rotation=45)
plt.show()

Q26: What is Seaborn and how is it different from Matplotlib?
Answer: Seaborn is a statistical visualization library built on top of Matplotlib that provides a high-level interface for attractive, informative statistical graphics. Key differences: (1) Seaborn has better default styling β€” charts look professional without customization. (2) Seaborn works directly with Pandas DataFrames β€” pass column names as strings. (3) Built-in statistical charts β€” heatmaps, violin plots, pair plots, box plots. (4) Automatic handling of categorical variables. (5) Built-in themes: set_style('whitegrid'). Seaborn for statistical analysis, Matplotlib for full custom control.
🎯 Explain: Seaborn = Matplotlib ka beautiful version. import seaborn as sns. sns.barplot(x='Dept', y='Salary', data=df) β€” ek line mein bar chart with error bars. Matplotlib mein 5-6 lines lagti hain same chart ke liye. Seaborn directly DataFrame accept karta hai β€” column names as strings. Heatmap, boxplot, violinplot, pairplot β€” sab built-in. Themes set karo β€” sns.set_style('whitegrid') β€” professional look. Interview mein "Seaborn for quick statistical visualizations with DataFrame integration, Matplotlib when I need pixel-level control."

import seaborn as sns

sns.set_style('whitegrid')

# Bar plot with Seaborn
sns.barplot(x='Dept', y='Salary', data=df, palette='viridis')
plt.title('Salary by Department')
plt.show()

# Box plot β€” salary distribution by department
sns.boxplot(x='Dept', y='Salary', data=df)
plt.title('Salary Distribution by Department')
plt.show()

# Scatter plot with hue
sns.scatterplot(x='Salary', y='Sales', hue='Dept', data=df, s=100)
plt.title('Salary vs Sales')
plt.show()

Q27: How do you create a Correlation Heatmap?
Answer: A correlation heatmap visualizes the pairwise correlation between numerical columns. Steps: (1) Calculate correlation matrix β€” df.corr(). (2) Plot using sns.heatmap(corr_matrix, annot=True, cmap='coolwarm'). annot=True shows correlation values on cells. cmap sets the color palette. Values range from -1 (negative correlation) to +1 (positive correlation). 0 means no correlation. Heatmaps are essential for exploratory data analysis β€” identifying which variables are related before building models.
🎯 Explain: Correlation heatmap = variables ke beech relationship dikhao visually. df.corr() β€” correlation matrix banao. sns.heatmap() β€” matrix ko color-coded chart mein dikhao. Red = positive correlation, Blue = negative. annot=True se numbers bhi dikhenge cells mein. Interview mein "I create correlation heatmaps as the first step in EDA to identify relationships between features β€” it guides which variables to investigate further."

# Correlation Heatmap
corr = df[['Salary','Sales']].corr()

plt.figure(figsize=(6,4))
sns.heatmap(corr, annot=True, cmap='coolwarm', 
           fmt='.2f', linewidths=1)
plt.title('Correlation Heatmap')
plt.tight_layout()
plt.show()

Q28: How do you create subplots (multiple charts in one figure)?
Answer: plt.subplots(nrows, ncols) creates a grid of charts in one figure. It returns a figure and axes objects. Access individual axes using indexing β€” axes[0] for first, axes[1] for second, or axes[row, col] for 2D grids. fig, axes = plt.subplots(2, 2, figsize=(12,8)) creates a 2Γ—2 grid. Each axis is an independent chart β€” you can plot different chart types in each. plt.tight_layout() prevents overlap. Subplots are essential for dashboard-style analysis showing multiple views simultaneously.
🎯 Explain: Subplots = ek figure mein multiple charts. fig, axes = plt.subplots(1,3) β€” ek row mein 3 charts. axes[0].bar(...) β€” pehla chart bar. axes[1].plot(...) β€” dusra line. axes[2].hist(...) β€” teesra histogram. 2D grid: plt.subplots(2,2) β€” 2Γ—2 = 4 charts. Dashboard presentations mein bahut useful. Interview mein "I use subplots for comparative analysis β€” showing salary distribution, department comparison, and trend analysis side by side."

# 1 row, 3 columns subplots
fig, axes = plt.subplots(1, 3, figsize=(15,5))

# Chart 1: Bar
axes[0].bar(df['Name'], df['Salary'], color='steelblue')
axes[0].set_title('Salary')
axes[0].tick_params(axis='x', rotation=45)

# Chart 2: Histogram
axes[1].hist(df['Salary'], bins=5, color='salmon', edgecolor='black')
axes[1].set_title('Salary Distribution')

# Chart 3: Scatter
axes[2].scatter(df['Salary'], df['Sales'], color='green')
axes[2].set_title('Salary vs Sales')

plt.tight_layout()
plt.show()

Q29: What chart types are best for different data scenarios?
Answer: Chart selection guide: (1) Comparison between categories β†’ Bar chart (vertical/horizontal). (2) Trend over time β†’ Line chart. (3) Distribution of single variable β†’ Histogram, KDE plot. (4) Relationship between 2 variables β†’ Scatter plot. (5) Part-to-whole composition β†’ Pie chart (use sparingly), stacked bar. (6) Distribution comparison across groups β†’ Box plot, Violin plot. (7) Correlation matrix β†’ Heatmap. (8) Multiple variable relationships β†’ Pair plot. (9) Geographic data β†’ Choropleth map. Choosing the right chart type is a critical data analyst skill.
🎯 Explain: Sahi chart select karna bahut important hai. Comparison chahiye β†’ Bar chart. Trend dikhana hai β†’ Line chart. Distribution β†’ Histogram ya Box plot. Relationship β†’ Scatter plot. Composition β†’ Pie (lekin sirf 3-5 categories ke liye). Outliers dhundhne hain β†’ Box plot. Correlation β†’ Heatmap. Interview mein "I choose chart types based on the analytical question β€” comparison uses bar, trend uses line, distribution uses histogram, and relationship uses scatter."

πŸ“Š Chart Selection Guide:

Purpose Best Chart Seaborn Function
Category ComparisonBar Chartsns.barplot()
Trend Over TimeLine Chartsns.lineplot()
DistributionHistogram / KDEsns.histplot() / sns.kdeplot()
RelationshipScatter Plotsns.scatterplot()
Distribution by GroupBox / Violin Plotsns.boxplot() / sns.violinplot()
CorrelationHeatmapsns.heatmap()
Pairwise ExplorationPair Plotsns.pairplot()
Count of CategoriesCount Plotsns.countplot()

Q30: How do you customize and save charts professionally?
Answer: Professional chart customization includes: (1) Figure size β€” plt.figure(figsize=(10,6)). (2) Titles and labels β€” plt.title(), xlabel(), ylabel() with fontsize. (3) Color palettes β€” Seaborn's palette parameter or custom colors. (4) Legend β€” plt.legend(). (5) Grid β€” plt.grid(True, alpha=0.3). (6) Annotations β€” plt.annotate() for highlighting specific points. (7) Axis formatting β€” plt.xlim(), ylim(), xticks(rotation=45). (8) Style β€” sns.set_style('whitegrid'). (9) Save β€” plt.savefig('chart.png', dpi=300, bbox_inches='tight'). dpi=300 for print quality, bbox_inches='tight' removes extra whitespace.
🎯 Explain: Professional charts = clear title, labeled axes, proper colors, clean layout. figsize se size set karo. fontsize se text readable banao. sns.set_style('whitegrid') β€” clean professional look. plt.savefig('chart.png', dpi=300, bbox_inches='tight') β€” high quality save. Reports aur presentations ke liye 300 dpi zaroor use karo. tight_layout() se charts overlap nahi honge. Interview mein "I follow a standard chart template β€” consistent sizing, labeled axes, professional color palettes, and always save at 300 DPI for presentations."

# Professional chart template
sns.set_style('whitegrid')
fig, ax = plt.subplots(figsize=(10,6))

sns.barplot(x='Dept', y='Salary', data=df, 
           palette='viridis', ax=ax)

ax.set_title('Average Salary by Department', fontsize=16, fontweight='bold')
ax.set_xlabel('Department', fontsize=12)
ax.set_ylabel('Salary (β‚Ή)', fontsize=12)

# Add value labels on bars
for p in ax.patches:
    ax.annotate(f'β‚Ή{p.get_height():.0f}',
               (p.get_x() + p.get_width()/2, p.get_height()),
               ha='center', va='bottom', fontsize=10)

plt.tight_layout()
plt.savefig('salary_chart.png', dpi=300, bbox_inches='tight')
plt.show()
πŸ’‘ Pro Tip: Visualization ka question aaye toh chart selection framework batao: "I follow the question-first approach β€” what am I trying to show? Comparison β†’ bar, Trend β†’ line, Distribution β†’ histogram/box, Relationship β†’ scatter, Correlation β†’ heatmap." Phir add karo: "I use Seaborn for quick EDA with sns.pairplot for exploring all variable relationships at once, and Matplotlib subplots for presentation-ready dashboards." Chart pe value labels add karna mention karo β€” professional touch hai.

πŸ“‹ Quick Revision Table β€” 30 Questions at a Glance

Q# Question One-Line Answer
Q1Series vs DataFrame?Series=1D column, DataFrame=2D table
Q2Add/Rename/Delete columns?df['new']=val, rename(), drop(columns=[])
Q3Sorting?sort_values(), nlargest(), nsmallest()
Q4copy() vs assignment?Assignment=reference, copy()=independent clone
Q5set_index / reset_index?Column→index for fast lookup, reset after groupby
Q6apply vs map vs applymap?map=Series, apply=Series/DF axis, applymap=all cells
Q7GroupBy?Split-apply-combine β€” SQL GROUP BY equivalent
Q8merge/join?merge=column JOIN, concat=stack, join=index JOIN
Q9pivot_table?Cross-tabulation β€” Excel Pivot Table in Python
Q10merge vs join vs concat?merge=column key, join=index, concat=stack
Q11transform()?Group metric back to row level β€” like SQL window function
Q12melt()?Wide→Long format — unpivot columns to rows
Q13Duplicate handling?duplicated(), drop_duplicates() β€” investigate first
Q14Missing value strategy?dropna/fillna β€” median for numerical, mode for categorical
Q15Data type conversion?astype(), to_numeric(errors='coerce'), to_datetime()
Q16String cleaning?.str accessor β€” strip, lower, contains, replace
Q17Outlier detection?IQR method or Z-score β€” clip/remove based on context
Q18replace() vs map()?replace=keeps unmapped, map=unmapped→NaN
Q19List vs NumPy Array?Array=50-100x faster, homogeneous, vectorized
Q20Broadcasting?Auto-expand smaller array to match larger β€” no loops
Q21Array creation?zeros, ones, arange, linspace, random, eye
Q22np.where vs np.select?where=binary IF, select=multiple conditions β€” vectorized
Q23Reshaping?reshape(), flatten(), .T β€” change dimensions
Q24NumPy statistics?mean, median, std, percentile, corrcoef, cumsum
Q25Matplotlib?Foundation plotting β€” bar, line, scatter, hist, pie
Q26Seaborn?Statistical charts β€” better defaults, DataFrame integration
Q27Correlation Heatmap?df.corr() + sns.heatmap(annot=True, cmap='coolwarm')
Q28Subplots?plt.subplots(rows,cols) β€” multiple charts in one figure
Q29Chart selection?Comparison→bar, Trend→line, Distribution→hist, Relation→scatter
Q30Chart customization?figsize, titles, labels, savefig(dpi=300), tight_layout

Thanks for Reading! πŸ™

Thanks for reading! Data Insights par aur bhi Power BI, Excel, SQL, Python topics available hain β€” explore karo aur apni analytics journey strong banao! Happy Learning & Keep Exploring! πŸš€

β€” JatinAnalytics

πŸ‘€
Jatin Kumar
Data Analyst & Educator

Python, SQL, Power BI aur Excel mein practical tutorials likhta hoon β€” taaki data analytics seekhna aasan ho. Portfolio: jatinanalytics.co.in

Portfolio LinkedIn GitHub Kaggle All Articles
Share:

πŸ’¬ Comments (0)

Spam/links allowed nahi hain β€” respectful comments welcome!

Loading comments...

Was this article helpful?