Basic EDA And Structure Functions in Pandas
Basic EDA & Structure Functions in Pandas: Complete Guide
Data Analysis ka sabse pehla step hota hai apne dataset ko samajhna — kitni rows hain, columns kya hain, data types kya hain, missing values kahan hain. Sikhiye 8 fundamental EDA functions jo har data scientist ka toolkit ka base hain.
📑 Is Masterclass Guide Mein Aap Kya Sikhenge:
Dataset ko samajhne aur explore karne ke 8 fundamental EDA tools:
- Quick Preview: head() / tail() — Dataset ke starting aur ending records dekhna
- Dimensions Check: shape — Total rows aur columns count
- Data Types Inspection: dtypes — Har column ka data type check karna
- Complete Summary: info() — One-shot dataset health report
- Statistical Overview: describe() — Numerical columns ki complete statistics
- Column Listing: columns — Saare column names access karna
- Unique Values: nunique() / unique() — Distinct values count aur list
- Value Distribution: value_counts() — Category-wise frequency count
1. head() / tail() — Dataset Ka Quick Preview
🔍 Kya Hai: head(n) DataFrame ki pehli n rows return karta hai (default 5) aur tail(n) last n rows return karta hai. Yeh dataset load karne ke baad sabse pehla function hai jo data ka quick preview deta hai — columns kya hain, data kaisa dikhta hai, format kya hai.
🎯 Kyu Use Hota Hai: Lakho rows ka poora dataset screen par dekhna impractical hai. head() se turant pata lagta hai ki data kaisa structured hai, column names kya hain, values ka format kya hai, aur koi obvious issue toh nahi hai. tail() se last records check hote hain — file truncation ya sorting verify karne ke liye.
💡 Kab Use Hota Hai: Dataset load karne ke turant baad (Step 1 of every EDA), data transformation ke baad result verify karne ke liye, aur debugging mein intermediate outputs check karne ke liye.
💻 Real-World Code Examples:
Example 1: Employee dataset load karke pehle 5 aur last 3 records preview karna.
import pandas as pd
employees = pd.read_csv("employees.csv")
# First 5 rows (default)
print(employees.head())
# Last 3 rows
print(employees.tail(3))
Example 2: E-commerce dataset mein top 10 records dekhna aur specific columns select karke preview karna.
ecommerce = pd.read_csv("orders.csv")
# Top 10 rows with selected columns
print(ecommerce[["OrderID", "CustomerName", "Revenue"]].head(10))
# Last 5 rows to verify data end
print(ecommerce.tail())
📊 Expected Output:
# head() Output:
# EmployeeID Name Age Department Salary
# 0 E001 Rahul 28 IT 85000
# 1 E002 Priya 34 HR 65000
# 2 E003 Amit 42 Finance 92000
# 3 E004 Sneha 26 IT 55000
# 4 E005 Vikram 38 Marketing 72000
# tail(3) Output:
# 9997 E9998 Kavita 31 HR 58000
# 9998 E9999 Deepak 45 Finance 95000
# 9999 E10000 Neha 29 IT 62000
✅ Best Practices:
- head() ko sirf "dekhne" ke liye use karein, analysis ke liye nahi. head() ka output misleading ho sakta hai agar data sorted hai ya first few rows representative nahi hain.
- Random sample chahiye toh
df.sample(10)use karein jo random 10 rows dikhata hai — yeh head() se zyada representative hota hai large datasets mein. - Specific columns select karke head() chalayein jab dataset mein 50+ columns hain — screen par clearly dikhega aur analysis focus hoga.
💬 Crack the Interview:
Q1: head() aur sample() mein kya difference hai aur kab kaunsa use karein?
Ans: head() hamesha first n rows return karta hai — sorted data par biased ho sakta hai. sample(n) random rows return karta hai jo data diversity better represent karta hai. EDA start mein head() quick structure dekhne ke liye, lekin data distribution samajhne ke liye sample() better hai.
Q2: Negative value head(-5) dene se kya hoga?
Ans: head(-5) last 5 rows EXCLUDE karke baaki sabhi rows return karta hai. Agar 100 rows hain toh 95 rows milenge. Similarly tail(-5) first 5 rows exclude karke baaki sabhi return karega. Yeh lesser-known but useful feature hai.
Q3: Large dataset mein head() slow ho sakta hai kya?
Ans: head() itself fast hai kyunki sirf n rows return karta hai. Lekin agar pehle heavy computation ho (jaise groupby ke baad head()) toh computation slow hogi, head() nahi. Jupyter notebooks mein df.head() display karna always instant hai.
2. shape — Dataset Dimensions Check Karna
🔍 Kya Hai: shape ek DataFrame attribute hai (function nahi, isliye parentheses nahi lagte) jo tuple format mein (rows, columns) return karta hai. Yeh turant batata hai ki dataset mein kitne records hain aur kitne features/columns hain.
🎯 Kyu Use Hota Hai: Data loading verification ke liye — expected rows load hui ya nahi. Data cleaning ke baad — kitni rows drop hui. Feature engineering ke baad — kitne naye columns add hue. Merge operations ke baad — row count verify karna ki unexpected duplicates toh nahi aaye.
💡 Kab Use Hota Hai: Har major data operation ke baad shape check karna standard practice hai — load ke baad, cleaning ke baad, merge ke baad, feature engineering ke baad, train-test split ke baad.
💻 Real-World Code Examples:
Example 1: Data load karke dimensions verify karna aur cleaning ke baad compare karna.
employees = pd.read_csv("employees.csv")
print(f"Original Shape: {employees.shape}")
# Cleaning — drop duplicates and NaN rows
employees_clean = employees.drop_duplicates().dropna()
print(f"After Cleaning: {employees_clean.shape}")
print(f"Rows Lost: {employees.shape[0] - employees_clean.shape[0]}")
Example 2: Merge operation ke baad unexpected row increase detect karna.
print(f"Orders: {ecommerce.shape}")
print(f"Customers: {customers.shape}")
merged = ecommerce.merge(customers, on="CustomerID", how="left")
print(f"Merged: {merged.shape}")
# Check if rows increased (indicates duplicate keys)
if merged.shape[0] > ecommerce.shape[0]:
print("⚠️ WARNING: Rows increased after merge — check for duplicate keys!")
📊 Expected Output:
# Original Shape: (10000, 8) # 10K rows, 8 columns
# After Cleaning: (9542, 8) # 458 rows removed
# Rows Lost: 458
# Orders: (15000, 6)
# Customers: (5000, 4)
# Merged: (15000, 9) # Rows same — merge is clean!
✅ Best Practices:
- shape attribute hai, method nahi —
df.shapelikhein,df.shape()nahi. Parentheses lagane se TypeError aayega. - Sirf rows chahiye toh
df.shape[0]ya faster alternativelen(df)use karein. Sirf columns count ke liyedf.shape[1]. - Har major transformation step ke baad shape print karein — yeh debugging ka sabse basic aur effective technique hai.
💬 Crack the Interview:
Q1: shape, size, aur len() mein kya difference hai?
Ans: shape = (rows, columns) tuple. size = rows × columns (total elements count). len(df) = sirf rows count. Example: 100 rows, 5 columns → shape=(100,5), size=500, len=100.
Q2: shape kyun method nahi hai? Parentheses kyun nahi lagte?
Ans: shape ek computed property/attribute hai jo directly NumPy array ki underlying shape access karta hai — koi computation nahi hoti isliye function call ki zaroorat nahi. Yeh instantaneous hai regardless of dataset size.
Q3: Merge ke baad rows increase hone ka practical impact kya hota hai?
Ans: Row increase matlab duplicate keys hain right DataFrame mein — yeh many-to-many join create karta hai. Isse downstream analysis mein double counting hogi. shape check se yeh immediately detect hota hai aur data quality issues fix hote hain.
3. dtypes — Har Column Ka Data Type Check Karna
🔍 Kya Hai: dtypes attribute har column ka data type (int64, float64, object, datetime64, bool, category) show karta hai. Yeh batata hai ki Pandas ne har column ko kis format mein store kiya hai — string, number, ya date.
🎯 Kyu Use Hota Hai: CSV load karne ke baad Pandas automatically dtypes assign karta hai — lekin aksar galat assign karta hai. Numeric columns string mein aa jaate hain, dates object mein rehti hain, boolean columns integer mein store hote hain. dtypes check se yeh issues turant dikhte hain aur fix kiye ja sakte hain.
💡 Kab Use Hota Hai: Data loading ke baad (verify correct parsing), mathematical operations se pehle (numeric dtype confirm karna), memory optimization (int64 → int32 downcast), aur ML model input preparation (encoding decisions).
💻 Real-World Code Examples:
Example 1: Dataset load karke dtypes check karna aur wrong types identify karke fix karna.
employees = pd.read_csv("employees.csv")
print(employees.dtypes)
# Fix wrong dtypes
employees["Salary"] = employees["Salary"].astype("float64")
employees["JoiningDate"] = pd.to_datetime(employees["JoiningDate"])
employees["Department"] = employees["Department"].astype("category")
print("\nAfter Fix:")
print(employees.dtypes)
Example 2: Memory usage reduce karne ke liye dtypes downcast karna.
print(f"Before: {employees.memory_usage(deep=True).sum() / 1024:.2f} KB")
# Downcast integers and floats
employees["Age"] = pd.to_numeric(employees["Age"], downcast="integer")
employees["Salary"] = pd.to_numeric(employees["Salary"], downcast="float")
print(f"After: {employees.memory_usage(deep=True).sum() / 1024:.2f} KB")
📊 Expected Output:
# Before Fix:
# EmployeeID object # OK
# Name object # OK
# Age int64 # OK
# Salary object # ❌ Should be float!
# JoiningDate object # ❌ Should be datetime!
# Department object # Can optimize to category
# After Fix:
# Salary float64
# JoiningDate datetime64[ns]
# Department category # Memory saved!
# Memory: Before: 845.20 KB → After: 412.50 KB # ~51% reduction!
✅ Best Practices:
- 'object' dtype ka matlab usually string hai — lekin mixed types bhi ho sakte hain (kuch cells number, kuch string). Aise columns ko carefully handle karein.
- Low cardinality string columns (jaise Department, Gender, City) ko 'category' dtype mein convert karein — memory 90%+ save hoti hai.
- dtypes check EDA ka mandatory Step 2 hai (Step 1 = head()). Bina dtypes fix kiye analysis start mat karein.
💬 Crack the Interview:
Q1: Pandas mein 'object' dtype ka exactly kya matlab hai?
Ans: 'object' dtype matlab column mein Python objects stored hain — usually strings. Lekin mixed types (kuch cells int, kuch string) bhi object dikhaata hai. Pandas 2.0+ mein 'string' dtype introduce hua hai jo explicitly strings ke liye hai aur better performance deta hai.
Q2: astype() aur pd.to_numeric() mein kya difference hai type conversion ke liye?
Ans: astype() direct type cast karta hai — invalid values par error deta hai. pd.to_numeric() mein errors='coerce' option hai jo invalid values ko NaN bana deta hai. Dirty data ke liye pd.to_numeric() safer hai.
Q3: Category dtype memory kaise save karta hai internally?
Ans: Category dtype unique values ko ek baar store karta hai aur baaki cells mein sirf integer codes rakhta hai. Jaise "Male" 10000 baar store karne ki jagah ek baar "Male" store ho aur 10000 cells mein sirf 0 ya 1 code ho — massive memory saving.
4. info() — One-Shot Complete Dataset Health Report
🔍 Kya Hai: info() Pandas ka single most powerful EDA function hai jo ek call mein complete dataset summary deta hai — total rows, columns list, har column ke non-null counts (missing values indicator), data types, aur total memory usage. Yeh ek medical checkup report jaisa hai dataset ka.
🎯 Kyu Use Hota Hai: Ek hi function se sab kuch pata chal jaata hai — columns kitne hain, kaunse column mein missing values hain (non-null count se), dtypes kya hain, memory kitni use ho rahi hai. head(), shape, dtypes, isnull().sum() — yeh sab info() mein summarized hai.
💡 Kab Use Hota Hai: Dataset load karne ke baad sabse pehle. Client ko dataset health summary deni ho. Data quality audit karna ho. New team member ko dataset samjhana ho. Har EDA notebook mein mandatory pehla cell.
💻 Real-World Code Examples:
Example 1: Employee dataset ki complete health report generate karna.
employees = pd.read_csv("employees.csv")
employees.info()
Example 2: Deep memory analysis enable karke exact memory per column check karna.
# Deep memory analysis (includes object column actual size)
employees.info(memory_usage="deep")
# Per-column memory breakdown
print("\nPer-Column Memory (KB):")
print((employees.memory_usage(deep=True) / 1024).round(2))
📊 Expected Output:
# <class 'pandas.core.frame.DataFrame'>
# RangeIndex: 10000 entries, 0 to 9999
# Data columns (total 8 columns):
# # Column Non-Null Count Dtype
# --- ------ -------------- -----
# 0 EmployeeID 10000 non-null object
# 1 Name 10000 non-null object
# 2 Age 9850 non-null float64 # 150 missing!
# 3 Department 10000 non-null object
# 4 Salary 9720 non-null float64 # 280 missing!
# 5 JoiningDate 10000 non-null object
# 6 Rating 9900 non-null float64 # 100 missing!
# 7 City 10000 non-null object
# dtypes: float64(3), object(5)
# memory usage: 625.0+ KB
✅ Best Practices:
- Non-Null Count ko total rows se compare karein — agar Non-Null < Total Rows toh woh column mein missing values hain. Example: 9850 non-null out of 10000 = 150 missing values.
memory_usage="deep"lagayein — default mein object columns ka actual memory show nahi hota. Deep mode exact bytes dikhata hai jo large datasets mein critical hai.- info() ka output directly capture nahi hota variable mein (None return karta hai). File mein save karna ho toh
bufparameter use karein:df.info(buf=open('info.txt','w')).
💬 Crack the Interview:
Q1: info() None kyun return karta hai — output variable mein capture kyun nahi hota?
Ans: info() directly stdout (console) par print karta hai, DataFrame ya string return nahi karta. Variable mein capture karne ke liye: import io; buf = io.StringIO(); df.info(buf=buf); info_str = buf.getvalue() use karein.
Q2: Very large datasets (100+ columns) mein info() truncated output deta hai — kaise fix karein?
Ans: df.info(max_cols=200) pass karein jo default 100 column limit override karta hai. Ya df.info(verbose=True) se force complete output milta hai sabhi columns ke saath.
Q3: info() se missing values percentage kaise quickly calculate karein?
Ans: info() sirf counts dikhata hai. Percentage ke liye: (df.isnull().sum() / len(df) * 100).round(2) use karein. Yeh har column ki missing percentage ek line mein de deta hai. info() ke baad yeh second step hona chahiye.
5. describe() — Statistical Summary of Numerical Columns
🔍 Kya Hai: describe() numerical columns ki complete statistical summary ek call mein return karta hai — count, mean, std (standard deviation), min, 25th percentile (Q1), 50th percentile (median), 75th percentile (Q3), aur max. Yeh ek X-ray hai data distribution ka.
🎯 Kyu Use Hota Hai: Sirf head() dekhne se data ki actual distribution nahi samajh aati. describe() turant batata hai — salary ka average kitna hai, minimum aur maximum kya hai, data symmetric hai ya skewed, outliers hain ya nahi (min/max aur mean ka comparison karke).
💡 Kab Use Hota Hai: Initial EDA mein data distribution samajhne ke liye, outlier detection ke liye (min/max vs mean compare karke), data quality check ke liye (impossible values jaise negative age), aur feature scaling decisions ke liye.
💻 Real-World Code Examples:
Example 1: Employee dataset ke numerical columns ki complete statistical overview.
# Default — only numerical columns
print(employees.describe())
# Custom percentiles
print(employees.describe(percentiles=[0.01, 0.05, 0.95, 0.99]))
Example 2: Categorical columns ki summary dekhna (include='object') aur complete summary (include='all').
# Categorical columns summary
print(employees.describe(include="object"))
# All columns — numeric + categorical
print(employees.describe(include="all"))
📊 Expected Output:
# Numerical Summary:
# Age Salary Rating
# count 9850.00 9720.00 9900.00
# mean 35.42 68450.25 3.65
# std 10.15 28500.80 0.92
# min 18.00 15000.00 1.00
# 25% 27.00 48000.00 3.00
# 50% 34.00 62000.00 4.00
# 75% 43.00 85000.00 4.50
# max 65.00 850000.00 5.00 # ⚠️ Max salary looks like outlier!
# Categorical Summary:
# Name Department City
# count 10000 10000 10000
# unique 9985 5 12
# top Rahul IT Mumbai
# freq 3 3200 2850
✅ Best Practices:
- Mean aur 50% (median) ka big gap = skewed data. Max aur mean ka big gap = outliers. Yeh patterns describe() se immediately identify hote hain.
- Outlier detection ke liye custom percentiles add karein:
describe(percentiles=[.01, .05, .95, .99])— extreme ends par values clearly dikhti hain. include='all'pass karein complete picture ke liye — default sirf numerical columns dikhata hai jo misleading ho sakta hai agar important categorical columns hain.
💬 Crack the Interview:
Q1: describe() se outliers kaise detect karein without any visualization?
Ans: Max vs 75th percentile compare karein — agar max bahut bada hai Q3 se toh upper outlier hai. Min vs 25th percentile compare karein. Mean vs median (50%) ka gap bada hai toh data skewed hai aur outliers influence kar rahe hain mean ko.
Q2: describe(include='object') mein 'top' aur 'freq' ka kya matlab hai?
Ans: 'top' sabse frequently occurring value hai (mode). 'freq' us value ka count hai. 'unique' total distinct values hai. Agar freq bahut high hai ek value ke liye toh woh column low variance hai aur ML mein useless ho sakta hai.
Q3: describe() ka output DataFrame hota hai — isko further analyze kaise karein?
Ans: Haan, describe() DataFrame return karta hai! stats = df.describe(); stats.loc['max','Salary'] se specific values access kar sakte hain. stats.loc['max'] - stats.loc['min'] se har column ka range ek line mein mil jaata hai.
6. columns — Saare Column Names Access Karna
🔍 Kya Hai: columns attribute DataFrame ke saare column names ka Index object return karta hai. Isse aap column names list kar sakte hain, rename kar sakte hain, filter kar sakte hain, ya programmatically columns select kar sakte hain.
🎯 Kyu Use Hota Hai: Large datasets mein 50-100+ columns hote hain — sabke names yaad rakhna impossible hai. columns se exact names pata chalte hain jo select, rename, drop operations mein zaroori hain. Column naming inconsistencies (spaces, uppercase) bhi identify hoti hain.
💡 Kab Use Hota Hai: Column names listing, programmatic column selection (regex-based), column renaming pipeline, feature names extraction for ML, aur dataset documentation mein.
💻 Real-World Code Examples:
Example 1: Column names list karna, clean karna (lowercase, spaces remove), aur rename karna.
print("Original Columns:", employees.columns.tolist())
# Clean column names — lowercase, replace spaces with underscore
employees.columns = employees.columns.str.lower().str.replace(" ", "_")
print("Cleaned Columns:", employees.columns.tolist())
Example 2: Programmatically specific pattern wale columns select karna.
# Select columns containing "date" keyword
date_cols = [col for col in ecommerce.columns if "date" in col.lower()]
print(f"Date Columns: {date_cols}")
# Select only numeric columns programmatically
numeric_cols = employees.select_dtypes(include=["int64", "float64"]).columns.tolist()
print(f"Numeric Columns: {numeric_cols}")
📊 Expected Output:
# Original Columns: ['Employee ID', 'Full Name', 'Age', 'Department', 'Salary']
# Cleaned Columns: ['employee_id', 'full_name', 'age', 'department', 'salary']
# Date Columns: ['order_date', 'delivery_date', 'return_date']
# Numeric Columns: ['age', 'salary', 'rating', 'experience']
✅ Best Practices:
- Dataset load karte hi columns ko standardize karein — lowercase, spaces replace with underscore, special characters remove. Isse baad mein typos aur KeyError avoid hote hain.
columns.tolist()se Python list milti hai jo list comprehension, loops, aur function arguments mein directly use hoti hai.select_dtypes()se programmatically numeric ya categorical columns select karna manual column listing se better hai — naye columns add hone par bhi kaam karega.
💬 Crack the Interview:
Q1: columns attribute vs select_dtypes() — kab kaunsa use karein?
Ans: columns sabhi column names deta hai regardless of type. select_dtypes() specific dtype ke columns filter karta hai. ML pipeline mein select_dtypes() better hai kyunki automatically naye numeric/categorical columns bhi include ho jaate hain bina hardcoding ke.
Q2: Column names mein trailing whitespace hone ka kya impact padta hai?
Ans: df['Salary'] kaam karega lekin df['Salary '] (trailing space) KeyError dega kyunki exact match hota hai. df.columns = df.columns.str.strip() se saari leading/trailing whitespace remove hoti hai — CSV files mein yeh bahut common issue hai.
Q3: columns ko directly assign karke rename karna vs rename() method — kab kaunsa?
Ans: Direct assignment df.columns = ['a','b','c'] se sabhi columns ek saath rename hote hain — exactly same count zaroori hai. df.rename(columns={'old':'new'}) se specific columns rename hote hain baaki unchanged rehte hain. Selective rename ke liye rename() better hai.
7. nunique() / unique() — Unique Values Count Aur List
🔍 Kya Hai: nunique() har column mein kitne unique (distinct) values hain woh count return karta hai. unique() actual unique values ka array return karta hai. Dono milke column ki cardinality (variety/diversity) batate hain.
🎯 Kyu Use Hota Hai: Feature analysis mein cardinality critical hai — agar ek column mein sirf 2 unique values hain (Male/Female) toh woh binary feature hai. Agar 10000 mein se 9985 unique hain (Names) toh woh high-cardinality hai aur encoding differently handle hogi. ID columns identify karna, constant columns detect karna — sab nunique() se hota hai.
💡 Kab Use Hota Hai: Feature selection mein (constant columns remove karna), encoding strategy decide karne mein (low vs high cardinality), ID columns identify karne mein, aur data quality check mein (unexpected unique values detect karna).
💻 Real-World Code Examples:
Example 1: Har column ki cardinality check karna aur low/high cardinality classify karna.
# All columns unique count
print(employees.nunique())
# Identify low cardinality columns (good for encoding)
low_cardinality = [col for col in employees.columns
if employees[col].nunique() 15]
print(f"Low Cardinality Columns: {low_cardinality}")
Example 2: Specific column ke actual unique values dekhna aur unexpected values identify karna.
# Actual unique values in Department column
print("Departments:", employees["Department"].unique())
print(f"Total: {employees['Department'].nunique()}")
# Check for unexpected Gender values
print("Gender values:", employees["Gender"].unique())
# Might reveal: ['Male', 'Female', 'M', 'F', nan] — inconsistency!
📊 Expected Output:
# nunique() per column:
# EmployeeID 10000 # High — likely ID column
# Name 9985 # High — near unique
# Age 48 # Medium — 18 to 65
# Department 5 # Low — perfect for encoding
# Salary 4520 # High — continuous
# City 12 # Low — good for encoding
# Low Cardinality: ['Department', 'City', 'Rating']
# Departments: ['IT', 'HR', 'Finance', 'Marketing', 'Operations']
# Gender
values: ['Male', 'Female', 'M', 'F', nan] # ⚠️ Inconsistency!
✅ Best Practices:
- nunique() == 1 wale columns ko drop karein — single constant value wale columns ML models ko koi information nahi dete.
constant_cols = [col for col in df.columns if df[col].nunique() == 1] - nunique() == len(df) wale columns potential ID columns hain — inhe features mein mat rakhein ML models mein.
- unique() ke output par set operations bhi kar sakte hain — do datasets ke columns ke unique values compare karna validation mein useful hai.
💬 Crack the Interview:
Q1: nunique() NaN values ko count karta hai ya nahi?
Ans: Default mein nunique() NaN ko count NAHI karta. NaN include karna ho toh nunique(dropna=False) use karein. unique() by default NaN ko include karta hai array mein.
Q2: High cardinality categorical columns ko ML mein kaise handle karein?
Ans: One-Hot Encoding se 1000+ columns ban jayenge jo memory aur model performance kharab karega. Solutions: Target Encoding, Frequency Encoding, ya top N categories rakhein aur baaki ko "Other" mein merge karein. nunique() se pehle hi decide hota hai kaunsi strategy use karni hai.
Q3: nunique() aur value_counts() mein kya relationship hai?
Ans: nunique() sirf count deta hai (single integer). value_counts() har unique value ka frequency deta hai (detailed breakdown). nunique() = len(value_counts()). Quick overview ke liye nunique(), detailed distribution ke liye value_counts().
8. value_counts() — Category-Wise Frequency Distribution
🔍 Kya Hai: value_counts() kisi column ki har unique value ka frequency count (kitni baar appear hua) descending order mein return karta hai. Yeh categorical data analysis ka most used function hai — data distribution instantly dikhata hai.
🎯 Kyu Use Hota Hai: Category distribution samajhna zaroori hai — kaunsa department sabse bada hai, kaunsi city mein zyada customers hain, ratings ka distribution kaisa hai. Imbalanced classes detect karna (90% Positive, 10% Negative) ML mein critical hai aur value_counts() se turant dikhta hai.
💡 Kab Use Hota Hai: Target variable balance check (classification problems), category distribution analysis, data quality check (unexpected values), customer segmentation distribution, aur feature importance intuition build karne ke liye.
💻 Real-World Code Examples:
Example 1: Department-wise employee count dekhna — absolute aur percentage dono mein.
# Absolute frequency
print(employees["Department"].value_counts())
# Percentage (normalized)
print(employees["Department"].value_counts(normalize=True).round(3) * 100)
Example 2: ML classification target variable ka class imbalance check karna aur NaN bhi count karna.
# Target variable class distribution
print("Churn Distribution:")
print(bank["Churned"].value_counts())
print(f"\nClass Ratio: {bank['Churned'].value_counts()[0] / bank['Churned'].value_counts()[1]:.1f}:1")
# Include NaN in count
print("\nWith NaN:")
print(bank["Churned"].value_counts(dropna=False))
📊 Expected Output:
# Department Distribution:
# IT 3200 32.0%
# HR 2400 24.0%
# Finance 2000 20.0%
# Marketing 1500 15.0%
# Operations 900 9.0%
# Churn Distribution:
# No 8500
# Yes 1500
# Class Ratio: 5.7:1 # ⚠️ Imbalanced — needs handling!
# With NaN:
# No 8500
# Yes 1500
# NaN 200 # 200 unlabeled records!
✅ Best Practices:
normalize=Truese percentage distribution milta hai jo absolute counts se zyada insightful hota hai — "3200 employees IT mein" vs "32% employees IT mein".- Classification target variable par hamesha value_counts() check karein — 5:1 se zyada imbalance ho toh SMOTE, undersampling ya class weights use karne padte hain.
dropna=Falselagake NaN values bhi count karein — missing values ki exact frequency pata chalti hai.
💬 Crack the Interview:
Q1: value_counts() numerical columns par kaise use karein bins ke saath?
Ans: df['Age'].value_counts(bins=5) numerical column ko 5 bins mein divide karke frequency count deta hai — pd.cut() + value_counts() ek step mein. Histogram data generate karne ka quickest way hai.
Q2: value_counts() ascending order mein kaise karein?
Ans: df['col'].value_counts(ascending=True) — smallest frequency pehle dikhega. Rare categories identify karne ke liye useful hai jo merge ya remove karne padein.
Q3: Multiple columns ka combined value_counts() kaise karein?
Ans: df[['Department','City']].value_counts() — Pandas 1.1+ mein DataFrame par directly value_counts() support karta hai jo cross-tabulation jaisa output deta hai. Har combination ki frequency milti hai jaise "IT + Mumbai: 450".
Conclusion: EDA Functions Quick Reference Matrix
EDA ke standard workflow mein in functions ka recommended order:
| Step | Function | Purpose |
|---|---|---|
| Step 1 | head() / tail() |
Quick data preview — structure samajhna |
| Step 2 | shape |
Dimensions verify — rows & columns count |
| Step 3 | info() |
Complete health report — dtypes, nulls, memory |
| Step 4 | dtypes |
Data type verification & correction |
| Step 5 | describe() |
Statistical distribution — outliers, skewness |
| Step 6 | columns |
Column names clean & standardize |
| Step 7 | nunique() / unique() |
Cardinality check — encoding decisions |
| Step 8 | value_counts() |
Category distribution — class imbalance detection |
Next Post Preview: Masterclass Part 9
Next masterclass mein hum cover karenge: Statistics & Aggregation Functions — mean(), median(), sum(), groupby(), agg(), pivot_table() aur 10 powerful statistical tools jo data summarization ko professional level par le jayenge.
Happy Coding & Stay Analytically Pure! 🚀
💬 Comments (0)
Loading comments...