<DataInsights />
  • 🏠 Home
  • 📊 SQL
  • 🐍 Python
  • 📈 Power BI
  • 📗 Excel
  • 💼 Career
  • 🎯 Interview Q&A
  • 📁 Case Study
  • 📥 Downloads
  • 🚀 My Portfolio
<DataInsights />

Practical Data Analytics tutorials covering SQL, Python, Power BI, Excel and career guidance for aspiring analysts — 100% free.

Topics

  • SQL Tutorials
  • Python Guide
  • Power BI
  • Excel Tips
  • Career Guide

Quick Links

  • 🛠️ All Tools
  • 🗓️ Archive
  • 📬 Contact
  • 🔍 Search
  • Portfolio
  • Kaggle
  • GitHub

Legal & Info

  • About
  • Contact
  • Privacy Policy
  • Disclaimer
  • Terms & Conditions
  • DMCA
  • Sitemap
Copyright © 2026 Data Insights by Jatin Kumar. All Rights Reserved.Built with ❤️ for Data Analysts
Home/Python/Basic EDA And Structure Functions in Pandas...

Basic EDA And Structure Functions in Pandas

A
August 2, 2026 Jatin Kumar 22 min read Python
Data Insights Masterclass — Part 8

Basic EDA & Structure Functions in Pandas: Complete Guide

Data Analysis ka sabse pehla step hota hai apne dataset ko samajhna — kitni rows hain, columns kya hain, data types kya hain, missing values kahan hain. Sikhiye 8 fundamental EDA functions jo har data scientist ka toolkit ka base hain.

📑 Is Masterclass Guide Mein Aap Kya Sikhenge:

Dataset ko samajhne aur explore karne ke 8 fundamental EDA tools:

  • Quick Preview: head() / tail() — Dataset ke starting aur ending records dekhna
  • Dimensions Check: shape — Total rows aur columns count
  • Data Types Inspection: dtypes — Har column ka data type check karna
  • Complete Summary: info() — One-shot dataset health report
  • Statistical Overview: describe() — Numerical columns ki complete statistics
  • Column Listing: columns — Saare column names access karna
  • Unique Values: nunique() / unique() — Distinct values count aur list
  • Value Distribution: value_counts() — Category-wise frequency count

1. head() / tail() — Dataset Ka Quick Preview

🔍 Kya Hai: head(n) DataFrame ki pehli n rows return karta hai (default 5) aur tail(n) last n rows return karta hai. Yeh dataset load karne ke baad sabse pehla function hai jo data ka quick preview deta hai — columns kya hain, data kaisa dikhta hai, format kya hai.

🎯 Kyu Use Hota Hai: Lakho rows ka poora dataset screen par dekhna impractical hai. head() se turant pata lagta hai ki data kaisa structured hai, column names kya hain, values ka format kya hai, aur koi obvious issue toh nahi hai. tail() se last records check hote hain — file truncation ya sorting verify karne ke liye.

💡 Kab Use Hota Hai: Dataset load karne ke turant baad (Step 1 of every EDA), data transformation ke baad result verify karne ke liye, aur debugging mein intermediate outputs check karne ke liye.

💻 Real-World Code Examples:

Example 1: Employee dataset load karke pehle 5 aur last 3 records preview karna.

import pandas as pd
employees = pd.read_csv("employees.csv")
# First 5 rows (default)
print(employees.head())
# Last 3 rows
print(employees.tail(3))

Example 2: E-commerce dataset mein top 10 records dekhna aur specific columns select karke preview karna.

ecommerce = pd.read_csv("orders.csv")
# Top 10 rows with selected columns
print(ecommerce[["OrderID", "CustomerName", "Revenue"]].head(10))
# Last 5 rows to verify data end
print(ecommerce.tail())

📊 Expected Output:

# head() Output:
#    EmployeeID  Name     Age  Department  Salary
# 0  E001        Rahul    28   IT          85000
# 1  E002        Priya    34   HR          65000
# 2  E003        Amit     42   Finance     92000
# 3  E004        Sneha    26   IT          55000
# 4  E005        Vikram   38   Marketing   72000

# tail(3) Output:
# 9997  E9998   Kavita   31   HR          58000
# 9998  E9999   Deepak   45   Finance     95000
# 9999  E10000  Neha     29   IT          62000

✅ Best Practices:

  • head() ko sirf "dekhne" ke liye use karein, analysis ke liye nahi. head() ka output misleading ho sakta hai agar data sorted hai ya first few rows representative nahi hain.
  • Random sample chahiye toh df.sample(10) use karein jo random 10 rows dikhata hai — yeh head() se zyada representative hota hai large datasets mein.
  • Specific columns select karke head() chalayein jab dataset mein 50+ columns hain — screen par clearly dikhega aur analysis focus hoga.

💬 Crack the Interview:

Q1: head() aur sample() mein kya difference hai aur kab kaunsa use karein?
Ans: head() hamesha first n rows return karta hai — sorted data par biased ho sakta hai. sample(n) random rows return karta hai jo data diversity better represent karta hai. EDA start mein head() quick structure dekhne ke liye, lekin data distribution samajhne ke liye sample() better hai.

Q2: Negative value head(-5) dene se kya hoga?
Ans: head(-5) last 5 rows EXCLUDE karke baaki sabhi rows return karta hai. Agar 100 rows hain toh 95 rows milenge. Similarly tail(-5) first 5 rows exclude karke baaki sabhi return karega. Yeh lesser-known but useful feature hai.

Q3: Large dataset mein head() slow ho sakta hai kya?
Ans: head() itself fast hai kyunki sirf n rows return karta hai. Lekin agar pehle heavy computation ho (jaise groupby ke baad head()) toh computation slow hogi, head() nahi. Jupyter notebooks mein df.head() display karna always instant hai.

2. shape — Dataset Dimensions Check Karna

🔍 Kya Hai: shape ek DataFrame attribute hai (function nahi, isliye parentheses nahi lagte) jo tuple format mein (rows, columns) return karta hai. Yeh turant batata hai ki dataset mein kitne records hain aur kitne features/columns hain.

🎯 Kyu Use Hota Hai: Data loading verification ke liye — expected rows load hui ya nahi. Data cleaning ke baad — kitni rows drop hui. Feature engineering ke baad — kitne naye columns add hue. Merge operations ke baad — row count verify karna ki unexpected duplicates toh nahi aaye.

💡 Kab Use Hota Hai: Har major data operation ke baad shape check karna standard practice hai — load ke baad, cleaning ke baad, merge ke baad, feature engineering ke baad, train-test split ke baad.

💻 Real-World Code Examples:

Example 1: Data load karke dimensions verify karna aur cleaning ke baad compare karna.

employees = pd.read_csv("employees.csv")
print(f"Original Shape: {employees.shape}")
# Cleaning — drop duplicates and NaN rows
employees_clean = employees.drop_duplicates().dropna()
print(f"After Cleaning: {employees_clean.shape}")
print(f"Rows Lost: {employees.shape[0] - employees_clean.shape[0]}")

Example 2: Merge operation ke baad unexpected row increase detect karna.

print(f"Orders: {ecommerce.shape}")
print(f"Customers: {customers.shape}")
merged = ecommerce.merge(customers, on="CustomerID", how="left")
print(f"Merged: {merged.shape}")
# Check if rows increased (indicates duplicate keys)
if merged.shape[0] > ecommerce.shape[0]:
    print("⚠️ WARNING: Rows increased after merge — check for duplicate keys!")

📊 Expected Output:

# Original Shape: (10000, 8)    # 10K rows, 8 columns
# After Cleaning: (9542, 8)     # 458 rows removed
# Rows Lost: 458

# Orders: (15000, 6)
# Customers: (5000, 4)
# Merged: (15000, 9)           # Rows same — merge is clean!

✅ Best Practices:

  • shape attribute hai, method nahi — df.shape likhein, df.shape() nahi. Parentheses lagane se TypeError aayega.
  • Sirf rows chahiye toh df.shape[0] ya faster alternative len(df) use karein. Sirf columns count ke liye df.shape[1].
  • Har major transformation step ke baad shape print karein — yeh debugging ka sabse basic aur effective technique hai.

💬 Crack the Interview:

Q1: shape, size, aur len() mein kya difference hai?
Ans: shape = (rows, columns) tuple. size = rows × columns (total elements count). len(df) = sirf rows count. Example: 100 rows, 5 columns → shape=(100,5), size=500, len=100.

Q2: shape kyun method nahi hai? Parentheses kyun nahi lagte?
Ans: shape ek computed property/attribute hai jo directly NumPy array ki underlying shape access karta hai — koi computation nahi hoti isliye function call ki zaroorat nahi. Yeh instantaneous hai regardless of dataset size.

Q3: Merge ke baad rows increase hone ka practical impact kya hota hai?
Ans: Row increase matlab duplicate keys hain right DataFrame mein — yeh many-to-many join create karta hai. Isse downstream analysis mein double counting hogi. shape check se yeh immediately detect hota hai aur data quality issues fix hote hain.

3. dtypes — Har Column Ka Data Type Check Karna

🔍 Kya Hai: dtypes attribute har column ka data type (int64, float64, object, datetime64, bool, category) show karta hai. Yeh batata hai ki Pandas ne har column ko kis format mein store kiya hai — string, number, ya date.

🎯 Kyu Use Hota Hai: CSV load karne ke baad Pandas automatically dtypes assign karta hai — lekin aksar galat assign karta hai. Numeric columns string mein aa jaate hain, dates object mein rehti hain, boolean columns integer mein store hote hain. dtypes check se yeh issues turant dikhte hain aur fix kiye ja sakte hain.

💡 Kab Use Hota Hai: Data loading ke baad (verify correct parsing), mathematical operations se pehle (numeric dtype confirm karna), memory optimization (int64 → int32 downcast), aur ML model input preparation (encoding decisions).

💻 Real-World Code Examples:

Example 1: Dataset load karke dtypes check karna aur wrong types identify karke fix karna.

employees = pd.read_csv("employees.csv")
print(employees.dtypes)
# Fix wrong dtypes
employees["Salary"]      = employees["Salary"].astype("float64")
employees["JoiningDate"] = pd.to_datetime(employees["JoiningDate"])
employees["Department"]  = employees["Department"].astype("category")
print("\nAfter Fix:")
print(employees.dtypes)

Example 2: Memory usage reduce karne ke liye dtypes downcast karna.

print(f"Before: {employees.memory_usage(deep=True).sum() / 1024:.2f} KB")
# Downcast integers and floats
employees["Age"] = pd.to_numeric(employees["Age"], downcast="integer")
employees["Salary"] = pd.to_numeric(employees["Salary"], downcast="float")
print(f"After: {employees.memory_usage(deep=True).sum() / 1024:.2f} KB")

📊 Expected Output:

# Before Fix:
# EmployeeID     object    # OK
# Name           object    # OK  
# Age            int64     # OK
# Salary         object    # ❌ Should be float!
# JoiningDate    object    # ❌ Should be datetime!
# Department     object    # Can optimize to category

# After Fix:
# Salary         float64
# JoiningDate    datetime64[ns]
# Department     category    # Memory saved!

# Memory: Before: 845.20 KB → After: 412.50 KB  # ~51% reduction!

✅ Best Practices:

  • 'object' dtype ka matlab usually string hai — lekin mixed types bhi ho sakte hain (kuch cells number, kuch string). Aise columns ko carefully handle karein.
  • Low cardinality string columns (jaise Department, Gender, City) ko 'category' dtype mein convert karein — memory 90%+ save hoti hai.
  • dtypes check EDA ka mandatory Step 2 hai (Step 1 = head()). Bina dtypes fix kiye analysis start mat karein.

💬 Crack the Interview:

Q1: Pandas mein 'object' dtype ka exactly kya matlab hai?
Ans: 'object' dtype matlab column mein Python objects stored hain — usually strings. Lekin mixed types (kuch cells int, kuch string) bhi object dikhaata hai. Pandas 2.0+ mein 'string' dtype introduce hua hai jo explicitly strings ke liye hai aur better performance deta hai.

Q2: astype() aur pd.to_numeric() mein kya difference hai type conversion ke liye?
Ans: astype() direct type cast karta hai — invalid values par error deta hai. pd.to_numeric() mein errors='coerce' option hai jo invalid values ko NaN bana deta hai. Dirty data ke liye pd.to_numeric() safer hai.

Q3: Category dtype memory kaise save karta hai internally?
Ans: Category dtype unique values ko ek baar store karta hai aur baaki cells mein sirf integer codes rakhta hai. Jaise "Male" 10000 baar store karne ki jagah ek baar "Male" store ho aur 10000 cells mein sirf 0 ya 1 code ho — massive memory saving.

4. info() — One-Shot Complete Dataset Health Report

🔍 Kya Hai: info() Pandas ka single most powerful EDA function hai jo ek call mein complete dataset summary deta hai — total rows, columns list, har column ke non-null counts (missing values indicator), data types, aur total memory usage. Yeh ek medical checkup report jaisa hai dataset ka.

🎯 Kyu Use Hota Hai: Ek hi function se sab kuch pata chal jaata hai — columns kitne hain, kaunse column mein missing values hain (non-null count se), dtypes kya hain, memory kitni use ho rahi hai. head(), shape, dtypes, isnull().sum() — yeh sab info() mein summarized hai.

💡 Kab Use Hota Hai: Dataset load karne ke baad sabse pehle. Client ko dataset health summary deni ho. Data quality audit karna ho. New team member ko dataset samjhana ho. Har EDA notebook mein mandatory pehla cell.

💻 Real-World Code Examples:

Example 1: Employee dataset ki complete health report generate karna.

employees = pd.read_csv("employees.csv")
employees.info()

Example 2: Deep memory analysis enable karke exact memory per column check karna.

# Deep memory analysis (includes object column actual size)
employees.info(memory_usage="deep")
# Per-column memory breakdown
print("\nPer-Column Memory (KB):")
print((employees.memory_usage(deep=True) / 1024).round(2))

📊 Expected Output:

# <class 'pandas.core.frame.DataFrame'>
# RangeIndex: 10000 entries, 0 to 9999
# Data columns (total 8 columns):
#  #   Column       Non-Null Count  Dtype
# ---  ------       --------------  -----
#  0   EmployeeID   10000 non-null  object
#  1   Name         10000 non-null  object
#  2   Age          9850 non-null   float64   # 150 missing!
#  3   Department   10000 non-null  object
#  4   Salary       9720 non-null   float64   # 280 missing!
#  5   JoiningDate  10000 non-null  object
#  6   Rating       9900 non-null   float64   # 100 missing!
#  7   City         10000 non-null  object
# dtypes: float64(3), object(5)
# memory usage: 625.0+ KB

✅ Best Practices:

  • Non-Null Count ko total rows se compare karein — agar Non-Null < Total Rows toh woh column mein missing values hain. Example: 9850 non-null out of 10000 = 150 missing values.
  • memory_usage="deep" lagayein — default mein object columns ka actual memory show nahi hota. Deep mode exact bytes dikhata hai jo large datasets mein critical hai.
  • info() ka output directly capture nahi hota variable mein (None return karta hai). File mein save karna ho toh buf parameter use karein: df.info(buf=open('info.txt','w')).

💬 Crack the Interview:

Q1: info() None kyun return karta hai — output variable mein capture kyun nahi hota?
Ans: info() directly stdout (console) par print karta hai, DataFrame ya string return nahi karta. Variable mein capture karne ke liye: import io; buf = io.StringIO(); df.info(buf=buf); info_str = buf.getvalue() use karein.

Q2: Very large datasets (100+ columns) mein info() truncated output deta hai — kaise fix karein?
Ans: df.info(max_cols=200) pass karein jo default 100 column limit override karta hai. Ya df.info(verbose=True) se force complete output milta hai sabhi columns ke saath.

Q3: info() se missing values percentage kaise quickly calculate karein?
Ans: info() sirf counts dikhata hai. Percentage ke liye: (df.isnull().sum() / len(df) * 100).round(2) use karein. Yeh har column ki missing percentage ek line mein de deta hai. info() ke baad yeh second step hona chahiye.

5. describe() — Statistical Summary of Numerical Columns

🔍 Kya Hai: describe() numerical columns ki complete statistical summary ek call mein return karta hai — count, mean, std (standard deviation), min, 25th percentile (Q1), 50th percentile (median), 75th percentile (Q3), aur max. Yeh ek X-ray hai data distribution ka.

🎯 Kyu Use Hota Hai: Sirf head() dekhne se data ki actual distribution nahi samajh aati. describe() turant batata hai — salary ka average kitna hai, minimum aur maximum kya hai, data symmetric hai ya skewed, outliers hain ya nahi (min/max aur mean ka comparison karke).

💡 Kab Use Hota Hai: Initial EDA mein data distribution samajhne ke liye, outlier detection ke liye (min/max vs mean compare karke), data quality check ke liye (impossible values jaise negative age), aur feature scaling decisions ke liye.

💻 Real-World Code Examples:

Example 1: Employee dataset ke numerical columns ki complete statistical overview.

# Default — only numerical columns
print(employees.describe())
# Custom percentiles
print(employees.describe(percentiles=[0.01, 0.05, 0.95, 0.99]))

Example 2: Categorical columns ki summary dekhna (include='object') aur complete summary (include='all').

# Categorical columns summary
print(employees.describe(include="object"))
# All columns — numeric + categorical
print(employees.describe(include="all"))

📊 Expected Output:

# Numerical Summary:
#          Age        Salary       Rating
# count    9850.00    9720.00      9900.00
# mean     35.42      68450.25     3.65
# std      10.15      28500.80     0.92
# min      18.00      15000.00     1.00
# 25%      27.00      48000.00     3.00
# 50%      34.00      62000.00     4.00
# 75%      43.00      85000.00     4.50
# max      65.00      850000.00    5.00  # ⚠️ Max salary looks like outlier!

# Categorical Summary:
#          Name       Department   City
# count    10000      10000        10000
# unique   9985       5            12
# top      Rahul      IT           Mumbai
# freq     3          3200         2850

✅ Best Practices:

  • Mean aur 50% (median) ka big gap = skewed data. Max aur mean ka big gap = outliers. Yeh patterns describe() se immediately identify hote hain.
  • Outlier detection ke liye custom percentiles add karein: describe(percentiles=[.01, .05, .95, .99]) — extreme ends par values clearly dikhti hain.
  • include='all' pass karein complete picture ke liye — default sirf numerical columns dikhata hai jo misleading ho sakta hai agar important categorical columns hain.

💬 Crack the Interview:

Q1: describe() se outliers kaise detect karein without any visualization?
Ans: Max vs 75th percentile compare karein — agar max bahut bada hai Q3 se toh upper outlier hai. Min vs 25th percentile compare karein. Mean vs median (50%) ka gap bada hai toh data skewed hai aur outliers influence kar rahe hain mean ko.

Q2: describe(include='object') mein 'top' aur 'freq' ka kya matlab hai?
Ans: 'top' sabse frequently occurring value hai (mode). 'freq' us value ka count hai. 'unique' total distinct values hai. Agar freq bahut high hai ek value ke liye toh woh column low variance hai aur ML mein useless ho sakta hai.

Q3: describe() ka output DataFrame hota hai — isko further analyze kaise karein?
Ans: Haan, describe() DataFrame return karta hai! stats = df.describe(); stats.loc['max','Salary'] se specific values access kar sakte hain. stats.loc['max'] - stats.loc['min'] se har column ka range ek line mein mil jaata hai.

6. columns — Saare Column Names Access Karna

🔍 Kya Hai: columns attribute DataFrame ke saare column names ka Index object return karta hai. Isse aap column names list kar sakte hain, rename kar sakte hain, filter kar sakte hain, ya programmatically columns select kar sakte hain.

🎯 Kyu Use Hota Hai: Large datasets mein 50-100+ columns hote hain — sabke names yaad rakhna impossible hai. columns se exact names pata chalte hain jo select, rename, drop operations mein zaroori hain. Column naming inconsistencies (spaces, uppercase) bhi identify hoti hain.

💡 Kab Use Hota Hai: Column names listing, programmatic column selection (regex-based), column renaming pipeline, feature names extraction for ML, aur dataset documentation mein.

💻 Real-World Code Examples:

Example 1: Column names list karna, clean karna (lowercase, spaces remove), aur rename karna.

print("Original Columns:", employees.columns.tolist())
# Clean column names — lowercase, replace spaces with underscore
employees.columns = employees.columns.str.lower().str.replace(" ", "_")
print("Cleaned Columns:", employees.columns.tolist())

Example 2: Programmatically specific pattern wale columns select karna.

# Select columns containing "date" keyword
date_cols = [col for col in ecommerce.columns if "date" in col.lower()]
print(f"Date Columns: {date_cols}")
# Select only numeric columns programmatically
numeric_cols = employees.select_dtypes(include=["int64", "float64"]).columns.tolist()
print(f"Numeric Columns: {numeric_cols}")

📊 Expected Output:

# Original Columns: ['Employee ID', 'Full Name', 'Age', 'Department', 'Salary']
# Cleaned Columns:  ['employee_id', 'full_name', 'age', 'department', 'salary']

# Date Columns: ['order_date', 'delivery_date', 'return_date']
# Numeric Columns: ['age', 'salary', 'rating', 'experience']

✅ Best Practices:

  • Dataset load karte hi columns ko standardize karein — lowercase, spaces replace with underscore, special characters remove. Isse baad mein typos aur KeyError avoid hote hain.
  • columns.tolist() se Python list milti hai jo list comprehension, loops, aur function arguments mein directly use hoti hai.
  • select_dtypes() se programmatically numeric ya categorical columns select karna manual column listing se better hai — naye columns add hone par bhi kaam karega.

💬 Crack the Interview:

Q1: columns attribute vs select_dtypes() — kab kaunsa use karein?
Ans: columns sabhi column names deta hai regardless of type. select_dtypes() specific dtype ke columns filter karta hai. ML pipeline mein select_dtypes() better hai kyunki automatically naye numeric/categorical columns bhi include ho jaate hain bina hardcoding ke.

Q2: Column names mein trailing whitespace hone ka kya impact padta hai?
Ans: df['Salary'] kaam karega lekin df['Salary '] (trailing space) KeyError dega kyunki exact match hota hai. df.columns = df.columns.str.strip() se saari leading/trailing whitespace remove hoti hai — CSV files mein yeh bahut common issue hai.

Q3: columns ko directly assign karke rename karna vs rename() method — kab kaunsa?
Ans: Direct assignment df.columns = ['a','b','c'] se sabhi columns ek saath rename hote hain — exactly same count zaroori hai. df.rename(columns={'old':'new'}) se specific columns rename hote hain baaki unchanged rehte hain. Selective rename ke liye rename() better hai.

7. nunique() / unique() — Unique Values Count Aur List

🔍 Kya Hai: nunique() har column mein kitne unique (distinct) values hain woh count return karta hai. unique() actual unique values ka array return karta hai. Dono milke column ki cardinality (variety/diversity) batate hain.

🎯 Kyu Use Hota Hai: Feature analysis mein cardinality critical hai — agar ek column mein sirf 2 unique values hain (Male/Female) toh woh binary feature hai. Agar 10000 mein se 9985 unique hain (Names) toh woh high-cardinality hai aur encoding differently handle hogi. ID columns identify karna, constant columns detect karna — sab nunique() se hota hai.

💡 Kab Use Hota Hai: Feature selection mein (constant columns remove karna), encoding strategy decide karne mein (low vs high cardinality), ID columns identify karne mein, aur data quality check mein (unexpected unique values detect karna).

💻 Real-World Code Examples:

Example 1: Har column ki cardinality check karna aur low/high cardinality classify karna.

# All columns unique count
print(employees.nunique())
# Identify low cardinality columns (good for encoding)
low_cardinality = [col for col in employees.columns 
                   if employees[col].nunique() 15]
print(f"Low Cardinality Columns: {low_cardinality}")

Example 2: Specific column ke actual unique values dekhna aur unexpected values identify karna.

# Actual unique values in Department column
print("Departments:", employees["Department"].unique())
print(f"Total: {employees['Department'].nunique()}")
# Check for unexpected Gender values
print("Gender values:", employees["Gender"].unique())
# Might reveal: ['Male', 'Female', 'M', 'F', nan] — inconsistency!

📊 Expected Output:

# nunique() per column:
# EmployeeID    10000  # High — likely ID column
# Name          9985   # High — near unique
# Age           48     # Medium — 18 to 65
# Department    5      # Low — perfect for encoding
# Salary        4520   # High — continuous
# City          12     # Low — good for encoding

# Low Cardinality: ['Department', 'City', 'Rating']
# Departments: ['IT', 'HR', 'Finance', 'Marketing', 'Operations']
# Gender
values: ['Male', 'Female', 'M', 'F', nan]  # ⚠️ Inconsistency!

✅ Best Practices:

  • nunique() == 1 wale columns ko drop karein — single constant value wale columns ML models ko koi information nahi dete. constant_cols = [col for col in df.columns if df[col].nunique() == 1]
  • nunique() == len(df) wale columns potential ID columns hain — inhe features mein mat rakhein ML models mein.
  • unique() ke output par set operations bhi kar sakte hain — do datasets ke columns ke unique values compare karna validation mein useful hai.

💬 Crack the Interview:

Q1: nunique() NaN values ko count karta hai ya nahi?
Ans: Default mein nunique() NaN ko count NAHI karta. NaN include karna ho toh nunique(dropna=False) use karein. unique() by default NaN ko include karta hai array mein.

Q2: High cardinality categorical columns ko ML mein kaise handle karein?
Ans: One-Hot Encoding se 1000+ columns ban jayenge jo memory aur model performance kharab karega. Solutions: Target Encoding, Frequency Encoding, ya top N categories rakhein aur baaki ko "Other" mein merge karein. nunique() se pehle hi decide hota hai kaunsi strategy use karni hai.

Q3: nunique() aur value_counts() mein kya relationship hai?
Ans: nunique() sirf count deta hai (single integer). value_counts() har unique value ka frequency deta hai (detailed breakdown). nunique() = len(value_counts()). Quick overview ke liye nunique(), detailed distribution ke liye value_counts().

8. value_counts() — Category-Wise Frequency Distribution

🔍 Kya Hai: value_counts() kisi column ki har unique value ka frequency count (kitni baar appear hua) descending order mein return karta hai. Yeh categorical data analysis ka most used function hai — data distribution instantly dikhata hai.

🎯 Kyu Use Hota Hai: Category distribution samajhna zaroori hai — kaunsa department sabse bada hai, kaunsi city mein zyada customers hain, ratings ka distribution kaisa hai. Imbalanced classes detect karna (90% Positive, 10% Negative) ML mein critical hai aur value_counts() se turant dikhta hai.

💡 Kab Use Hota Hai: Target variable balance check (classification problems), category distribution analysis, data quality check (unexpected values), customer segmentation distribution, aur feature importance intuition build karne ke liye.

💻 Real-World Code Examples:

Example 1: Department-wise employee count dekhna — absolute aur percentage dono mein.

# Absolute frequency
print(employees["Department"].value_counts())
# Percentage (normalized)
print(employees["Department"].value_counts(normalize=True).round(3) * 100)

Example 2: ML classification target variable ka class imbalance check karna aur NaN bhi count karna.

# Target variable class distribution
print("Churn Distribution:")
print(bank["Churned"].value_counts())
print(f"\nClass Ratio: {bank['Churned'].value_counts()[0] / bank['Churned'].value_counts()[1]:.1f}:1")
# Include NaN in count
print("\nWith NaN:")
print(bank["Churned"].value_counts(dropna=False))

📊 Expected Output:

# Department Distribution:
# IT           3200    32.0%
# HR           2400    24.0%
# Finance      2000    20.0%
# Marketing    1500    15.0%
# Operations    900     9.0%

# Churn Distribution:
# No     8500
# Yes    1500
# Class Ratio: 5.7:1  # ⚠️ Imbalanced — needs handling!

# With NaN:
# No     8500
# Yes    1500
# NaN     200    # 200 unlabeled records!

✅ Best Practices:

  • normalize=True se percentage distribution milta hai jo absolute counts se zyada insightful hota hai — "3200 employees IT mein" vs "32% employees IT mein".
  • Classification target variable par hamesha value_counts() check karein — 5:1 se zyada imbalance ho toh SMOTE, undersampling ya class weights use karne padte hain.
  • dropna=False lagake NaN values bhi count karein — missing values ki exact frequency pata chalti hai.

💬 Crack the Interview:

Q1: value_counts() numerical columns par kaise use karein bins ke saath?
Ans: df['Age'].value_counts(bins=5) numerical column ko 5 bins mein divide karke frequency count deta hai — pd.cut() + value_counts() ek step mein. Histogram data generate karne ka quickest way hai.

Q2: value_counts() ascending order mein kaise karein?
Ans: df['col'].value_counts(ascending=True) — smallest frequency pehle dikhega. Rare categories identify karne ke liye useful hai jo merge ya remove karne padein.

Q3: Multiple columns ka combined value_counts() kaise karein?
Ans: df[['Department','City']].value_counts() — Pandas 1.1+ mein DataFrame par directly value_counts() support karta hai jo cross-tabulation jaisa output deta hai. Har combination ki frequency milti hai jaise "IT + Mumbai: 450".

Conclusion: EDA Functions Quick Reference Matrix

EDA ke standard workflow mein in functions ka recommended order:

Step Function Purpose
Step 1 head() / tail() Quick data preview — structure samajhna
Step 2 shape Dimensions verify — rows & columns count
Step 3 info() Complete health report — dtypes, nulls, memory
Step 4 dtypes Data type verification & correction
Step 5 describe() Statistical distribution — outliers, skewness
Step 6 columns Column names clean & standardize
Step 7 nunique() / unique() Cardinality check — encoding decisions
Step 8 value_counts() Category distribution — class imbalance detection

Next Post Preview: Masterclass Part 9

Next masterclass mein hum cover karenge: Statistics & Aggregation Functions — mean(), median(), sum(), groupby(), agg(), pivot_table() aur 10 powerful statistical tools jo data summarization ko professional level par le jayenge.

Happy Coding & Stay Analytically Pure! 🚀

👤
Jatin Kumar
Data Analyst & Educator

Python, SQL, Power BI aur Excel mein practical tutorials likhta hoon — taaki data analytics seekhna aasan ho. Portfolio: jatinanalytics.co.in

Portfolio LinkedIn GitHub Kaggle All Articles
Share:

💬 Comments (0)

Spam/links allowed nahi hain — respectful comments welcome!

Loading comments...

Was this article helpful?