Basic Interview Question of Statistics
Statistics Basic Interview Questions π
Top 30 basic Statistics interview questions for Data Analysts β Mean, Median, Mode, Standard Deviation, Variance, Probability basics, Data Types, Distributions, Percentiles aur Charts. Real examples with formulas aur Python code. Data Insights par.
π Is Blog Mein Kya Sikhenge:
- π’ Q1βQ6: Statistics Fundamentals β Types, Population vs Sample
- π’ Q7βQ12: Central Tendency β Mean, Median, Mode, Weighted Mean
- π’ Q13βQ18: Dispersion β Range, Variance, Standard Deviation, IQR
- π’ Q19βQ24: Data Types, Scales & Distribution Basics
- π’ Q25βQ30: Probability Basics, Percentiles & Charts
- π‘ Pro Tips: Interview mein exactly kya bolna chahiye
π Sample Data β Is Blog Ke Examples Isi Par Based Hain
| Employee | Department | Salary (βΉ) | Experience (Yrs) | Sales (βΉ) | Rating |
|---|---|---|---|---|---|
| Aarav | IT | 55,000 | 3 | 85,000 | 4 |
| Ishita | HR | 72,000 | 5 | 92,000 | 5 |
| Kabir | Finance | 65,000 | 4 | 45,000 | 3 |
| Diya | IT | 58,000 | 2 | 78,000 | 4 |
| Rohan | Marketing | 80,000 | 7 | 65,000 | 4 |
| Meera | HR | 48,000 | 1 | 52,000 | 3 |
| Arjun | Finance | 70,000 | 6 | 88,000 | 5 |
| Kavya | Marketing | 62,000 | 3 | 71,000 | 4 |
π’ Category 1: Statistics Fundamentals (Q1βQ6)
Q1: What is Statistics and why is it important in Data Analysis?
Answer: Statistics is the science of collecting, organizing, analyzing, interpreting, and presenting data. It provides methods to summarize data (descriptive statistics), make predictions (inferential statistics), and draw conclusions from data with measurable confidence. In data analysis, statistics is essential for understanding data distributions, identifying patterns, testing hypotheses, making data-driven decisions, building predictive models, and communicating insights with quantified uncertainty.
π― Explain: Statistics = data ko samajhne ka science. Descriptive statistics data summarize karti hai β mean, median, mode. Inferential statistics predictions aur decisions banati hai β hypothesis testing, confidence intervals. Data Analyst ko statistics isliye zaroori hai kyunki: average salary kya hai? Sales badh rahi hain ya nahi? Yeh campaign effective hai ya coincidence? β sab statistics se answer hote hain. Interview mein "Statistics provides the mathematical foundation for turning raw data into actionable business insights."
Q2: What is the difference between Descriptive and Inferential Statistics?
Answer: Descriptive Statistics summarizes and describes the characteristics of a dataset β mean, median, mode, standard deviation, charts, tables. It answers "What happened?" β no predictions or generalizations beyond the data. Inferential Statistics uses sample data to make predictions or generalizations about a larger population β hypothesis testing, confidence intervals, regression. It answers "What can we conclude?" with a measured level of confidence. Descriptive is about the data you have, Inferential is about the data you don't have.
π― Explain: Descriptive = data describe karo β "average salary 63,750 hai, highest 80,000." β jo data hai usi ke baare mein baat. Inferential = data se conclusion nikalo β "sample ke basis pe puri company ki average salary 60,000-68,000 ke beech hogi 95% confidence ke saath." β sample se population ke baare mein predict karo. Interview mein "Descriptive answers what happened, Inferential answers what can we conclude about the larger population."
π Comparison Table:
| Feature | Descriptive | Inferential |
|---|---|---|
| Purpose | Summarize data | Make predictions |
| Scope | Current dataset only | Generalize to population |
| Examples | Mean, Median, Charts | Hypothesis Testing, CI |
| Uncertainty | No uncertainty β exact | Includes confidence level |
| Question | What happened? | What can we conclude? |
Q3: What is the difference between Population and Sample?
Answer: Population is the entire group you want to study β ALL employees in a company, ALL customers, ALL transactions. Sample is a subset of the population selected for analysis. Population parameters are denoted by Greek letters (ΞΌ for mean, Ο for standard deviation). Sample statistics are denoted by Latin letters (xΜ for mean, s for standard deviation). We use samples because studying the entire population is often impractical, expensive, or impossible. A good sample should be representative β reflecting the population's characteristics.
π― Explain: Population = poora group β company ke sab 10,000 employees. Sample = chhota group β 500 employees randomly selected. Population study karna mushkil hai (expensive, time-consuming) β isliye sample lete hain. Sample se population ke baare mein estimate karte hain. Population mean = ΞΌ (mu), Sample mean = xΜ (x-bar). Key rule: sample representative hona chahiye β biased sample = galat conclusions. Interview mein "Population is the entire group of interest, Sample is a representative subset used for practical analysis."
Q4: What is the difference between a Parameter and a Statistic?
Answer: A Parameter is a numerical measure that describes a characteristic of the entire population β it is fixed but usually unknown (e.g., population mean ΞΌ). A Statistic is a numerical measure calculated from a sample β it is known but varies from sample to sample (e.g., sample mean xΜ). Statistics are used to estimate parameters. Example: average salary of ALL employees in India (parameter β unknown) is estimated using a survey of 10,000 employees (statistic β known, calculated).
π― Explain: Parameter = population ki value β fixed hai but pata nahi (poore India ki average salary). Statistic = sample ki value β pata hai but har sample mein alag hogi. Statistics se parameters estimate karte hain. Population mean ΞΌ = parameter. Sample mean xΜ = statistic. Interview mein "Parameter describes the population and is usually unknown, Statistic describes the sample and is calculated to estimate the parameter."
Q5: What are the different types of data in Statistics?
Answer: Data is classified into two main types: (1) Qualitative (Categorical) β describes qualities/categories. Subtypes: Nominal (no order β gender, color, department) and Ordinal (natural order β ratings, education level). (2) Quantitative (Numerical) β describes quantities/numbers. Subtypes: Discrete (countable whole numbers β number of employees, orders) and Continuous (measurable, any value β salary, height, temperature). The data type determines which statistical methods and visualizations are appropriate.
π― Explain: Data 2 types ka hota hai β Categorical aur Numerical. Categorical: Nominal = koi order nahi β Department (IT, HR, Finance) β alphabet order meaningless. Ordinal = natural order hai β Rating (1,2,3,4,5) β order matters. Numerical: Discrete = countable β employees count (1,2,3 β 2.5 employees nahi hote). Continuous = measurable β salary 55,000.50 ho sakti hai. Data type decide karta hai kaunsa chart use karo, kaunsa test lagao. Interview mein "Nominal has no order, Ordinal has order, Discrete is countable, Continuous is measurable."
π Data Types Hierarchy:
| Type | Subtype | Example from Data | Best Chart |
|---|---|---|---|
| Categorical | Nominal | Department (IT, HR) | Bar, Pie |
| Categorical | Ordinal | Rating (1-5) | Bar (ordered) |
| Numerical | Discrete | Experience (1,2,3 yrs) | Bar, Count |
| Numerical | Continuous | Salary (55000.50) | Histogram, Box |
Q6: What are the four Scales of Measurement?
Answer: Four scales in increasing order of information: (1) Nominal β categories with no order (Gender: M/F, Department: IT/HR). Only mode applicable. (2) Ordinal β categories with natural order but unequal intervals (Rating: 1-5, Education: High School < Bachelor < Master). Mode and median applicable. (3) Interval β ordered with equal intervals but no true zero (Temperature in Β°C, IQ scores). Mean, median, mode applicable. (4) Ratio β ordered, equal intervals, true zero point (Salary, Height, Weight). All statistical operations applicable.
π― Explain: NOIR β Nominal, Ordinal, Interval, Ratio β yaad rakho. Nominal = naam β IT, HR (sirf count/mode). Ordinal = order β Rating 1<2<3 (but difference equal nahi zaroor). Interval = equal gaps, no true zero β 0Β°C matlab "no temperature" nahi (mean possible). Ratio = true zero β Salary βΉ0 matlab "no salary" (sab operations possible). Salary ratio scale hai β βΉ80000 is βΉ40000 ka double β meaningful ratio. Interview mein "NOIR hierarchy β each level adds more mathematical operations."
π’ Category 2: Central Tendency (Q7βQ12)
Q7: What is Mean (Average)?
Answer: Mean (Arithmetic Mean) is the sum of all values divided by the number of values. Formula: Mean = Ξ£x / n. It is the most commonly used measure of central tendency. Advantages: uses all data points, good for symmetric distributions. Disadvantages: heavily affected by outliers β a single extreme value can significantly shift the mean. For salary data [55K, 72K, 65K, 58K, 80K, 48K, 70K, 62K]: Mean = 510,000 / 8 = 63,750.
π― Explain: Mean = sabka total Γ· count. Salaries: (55+72+65+58+80+48+70+62)/8 = 510/8 = 63.75K. Sabse common measure hai β "average salary kitni hai?" β yeh mean hai. Problem: outlier aaye toh mean bias ho jaata hai. Agar ek CEO ki salary 5,00,000 add ho jaaye toh mean bahut upar chala jayega β misleading hoga. Isliye skewed data mein median better hai. Interview mein "Mean uses all data points but is sensitive to outliers β for skewed data, I prefer median."
# Mean calculation
salaries = [55000, 72000, 65000, 58000, 80000, 48000, 70000, 62000]
# Manual
mean = sum(salaries) / len(salaries)
print(f"Mean: βΉ{mean:,.0f}") # Mean: βΉ63,750
# NumPy
import numpy as np
print(np.mean(salaries)) # 63750.0
# Outlier effect
with_outlier = salaries + [500000] # CEO salary added
print(np.mean(with_outlier)) # 112,222 β hugely inflated!
Q8: What is Median?
Answer: Median is the middle value when data is sorted in ascending order. For odd number of values, it is the exact middle value. For even number of values, it is the average of the two middle values. Formula (even n): Median = (n/2 th value + (n/2+1) th value) / 2. Advantages: not affected by outliers β robust measure. Used when data is skewed or contains extreme values. For our data sorted: [48K, 55K, 58K, 62K, 65K, 70K, 72K, 80K] β Median = (62K + 65K) / 2 = 63,500.
π― Explain: Median = beech wali value. Pehle sort karo β [48, 55, 58, 62, 65, 70, 72, 80]. 8 values hain (even) β 4th aur 5th ka average β (62+65)/2 = 63.5K. Outlier aaye toh bhi median zyada nahi badlega β robust hai. Salary data mein median better representation deta hai β "half employees isse zyada kamate hain, half isse kam". Real estate prices, income data β skewed data mein hamesha median use karo. Interview mein "Median is robust to outliers β I use it for skewed distributions like income and property prices."
# Median calculation
sorted_sal = sorted(salaries)
print(f"Sorted: {sorted_sal}")
# [48000, 55000, 58000, 62000, 65000, 70000, 72000, 80000]
print(f"Median: βΉ{np.median(salaries):,.0f}") # Median: βΉ63,500
# Median vs Mean with outlier
print(f"Mean with outlier: βΉ{np.mean(with_outlier):,.0f}") # βΉ112,222
print(f"Median with outlier: βΉ{np.median(with_outlier):,.0f}") # βΉ65,000 β barely changed!
Q9: What is Mode?
Answer: Mode is the most frequently occurring value in a dataset. A dataset can be unimodal (one mode), bimodal (two modes), multimodal (multiple modes), or have no mode (all values unique). Mode is the only measure of central tendency applicable to nominal (categorical) data. For ratings [4, 5, 3, 4, 4, 3, 5, 4]: Mode = 4 (appears 4 times). For departments [IT, HR, Finance, IT, Marketing, HR, Finance, Marketing]: Modes = IT, HR, Finance, Marketing (bimodal if two max).
π― Explain: Mode = sabse zyada baar aane wali value. Ratings: [4,5,3,4,4,3,5,4] β Mode = 4 (4 baar aaya). Departments: IT=2, HR=2, Finance=2, Marketing=2 β sab equal β no mode ya multimodal. Mode categorical data ke liye best hai β "sabse common department kaunsa hai?" β mode se pata chalega. Numerical data mein mean/median zyada useful β mode tab use karo jab most frequent value jaanni ho. Interview mein "Mode is the only central tendency measure for nominal data β I use it for finding most common categories."
from scipy import stats
ratings = [4, 5, 3, 4, 4, 3, 5, 4]
print(f"Mode: {stats.mode(ratings, keepdims=False).mode}") # Mode: 4
# Pandas value_counts β better for mode analysis
import pandas as pd
s = pd.Series(ratings)
print(s.value_counts()) # 4: 4 times, 3: 2, 5: 2
print(s.mode()) # 4
Q10: When should you use Mean, Median, or Mode?
Answer: Use Mean when: data is symmetric (normal distribution), no significant outliers, interval/ratio scale. Use Median when: data is skewed, outliers present, ordinal or continuous data β income, property prices, response times. Use Mode when: categorical/nominal data β most popular product, most common department. In practice: report both mean and median β if they are close, data is symmetric; if mean >> median, data is right-skewed; if mean << median, left-skewed.
π― Explain: Selection rule: Symmetric data β Mean (most informative). Skewed data / outliers β Median (robust). Categorical data β Mode (only option). Quick check: Mean β Median β symmetric. Mean > Median β right skewed (tail right mein β salary data). Mean < Median β left skewed. Real-world: salary reports mein companies median report karti hain β "median salary βΉ65,000" β mean misleading hota hai high earners ki wajah se. Interview mein "I compare mean and median to detect skewness β if they differ significantly, I report median for central tendency."
π When to Use What:
| Measure | Best For | Avoid When | Example |
|---|---|---|---|
| Mean | Symmetric, no outliers | Skewed, extreme values | Test scores |
| Median | Skewed, outliers present | Small categorical data | Salary, house prices |
| Mode | Categorical data | Continuous unique values | Most sold product |
Q11: What is Weighted Mean?
Answer: Weighted Mean assigns different importance (weights) to different values β values with higher weights contribute more to the average. Formula: Weighted Mean = Ξ£(value Γ weight) / Ξ£(weights). Used when: calculating GPA (different credit hours), portfolio returns (different investment amounts), survey averages (different sample sizes). Regular mean treats all values equally, weighted mean accounts for relative importance.
π― Explain: Weighted Mean = kuch values ko zyada importance do. Example: 3 subjects β Maths 60 marks (weight 4), English 80 marks (weight 2), Science 70 marks (weight 3). Simple mean = (60+80+70)/3 = 70. Weighted mean = (60Γ4 + 80Γ2 + 70Γ3) / (4+2+3) = (240+160+210)/9 = 610/9 = 67.78. Maths ka zyada weightage hai toh result neeche aaya. Real-world: CGPA, stock portfolio returns, satisfaction surveys. Interview mein "Weighted mean accounts for varying importance β I use it for GPA calculations and portfolio-weighted returns."
# Weighted Mean
marks = [60, 80, 70]
weights = [4, 2, 3]
# Manual
w_mean = sum(m*w for m,w in zip(marks,weights)) / sum(weights)
print(f"Weighted Mean: {w_mean:.2f}") # 67.78
# NumPy
print(np.average(marks, weights=weights)) # 67.78
# Simple mean comparison
print(f"Simple Mean: {np.mean(marks):.2f}") # 70.00
Q12: What is the relationship between Mean, Median, and Mode in different distributions?
Answer: In a perfectly symmetric (normal) distribution: Mean = Median = Mode β all three are at the center. In a right-skewed distribution (positive skew): Mode < Median < Mean β the tail pulls the mean to the right. In a left-skewed distribution (negative skew): Mean < Median < Mode β the tail pulls the mean to the left. This relationship is a quick diagnostic for distribution shape. The empirical rule: Mean - Mode β 3 Γ (Mean - Median) for moderately skewed data.
π― Explain: Symmetric distribution: Mean = Median = Mode β perfectly balanced. Right skewed (income data): tail right mein hai β Mean sabse zyada, phir Median, phir Mode (Mode < Median < Mean). Left skewed (exam scores β most score high): tail left mein β Mean sabse kam (Mean < Median < Mode). Quick check: Mean vs Median compare karo β Mean > Median toh right skewed. Salary data usually right skewed hota hai β few high earners mean ko upar kheenchte hain. Interview mein "In right-skewed data like salary, Mean > Median β I report median for a more representative central value."
π’ Category 3: Dispersion / Spread (Q13βQ18)
Q13: What is Range?
Answer: Range is the simplest measure of dispersion β the difference between the maximum and minimum values. Formula: Range = Max - Min. For salaries: Range = 80,000 - 48,000 = 32,000. Advantages: easy to calculate and understand. Disadvantages: uses only two extreme values, ignores all other data points, heavily affected by outliers. Range gives a quick sense of data spread but is not reliable for detailed analysis.
π― Explain: Range = Max - Min. Salaries mein: 80,000 - 48,000 = 32,000. Matlab sabse zyada aur sabse kam ke beech βΉ32,000 ka gap hai. Problem: sirf 2 values use hoti hain β baaki sab ignore. Ek extreme outlier aaye toh range bahut badh jayega β misleading. Quick overview ke liye theek hai, detailed analysis ke liye Variance, Standard Deviation, IQR better hain. Interview mein "Range is a quick spread indicator but I rely on IQR and standard deviation for robust dispersion measurement."
salaries = [55000, 72000, 65000, 58000, 80000, 48000, 70000, 62000]
range_val = max(salaries) - min(salaries)
print(f"Range: βΉ{range_val:,}") # Range: βΉ32,000
# NumPy
print(np.ptp(salaries)) # 32000 (peak-to-peak)
Q14: What is Variance?
Answer: Variance measures the average squared deviation of each data point from the mean. Population Variance (ΟΒ²) = Ξ£(x - ΞΌ)Β² / N. Sample Variance (sΒ²) = Ξ£(x - xΜ)Β² / (n-1). The (n-1) denominator in sample variance is called Bessel's correction β it corrects for the bias of estimating population variance from a sample. Variance tells how spread out the data is from the mean. Higher variance = more spread. Variance is in squared units β hard to interpret directly, so we take its square root to get Standard Deviation.
π― Explain: Variance = har value mean se kitni door hai uska squared average. Steps: (1) Mean nikalo (63,750). (2) Har value se mean minus karo. (3) Differences ko square karo. (4) Average lo (population mein N se divide, sample mein n-1 se). n-1 kyu? Bessel's correction β sample se population estimate karte waqt bias correct karta hai. Variance ka unit squared hota hai (βΉΒ²) β interpret mushkil β isliye square root lete hain = Standard Deviation. Interview mein "Variance measures spread in squared units β I use standard deviation for interpretation and variance for mathematical calculations."
# Variance calculation
# Population variance (ddof=0)
print(f"Pop Variance: {np.var(salaries):,.0f}") # divide by N
# Sample variance (ddof=1) β use this!
print(f"Sample Variance: {np.var(salaries, ddof=1):,.0f}") # divide by n-1
# Manual calculation
mean = np.mean(salaries)
sq_diff = [(x - mean)**2 for x in salaries]
sample_var = sum(sq_diff) / (len(salaries) - 1)
print(f"Manual Sample Var: {sample_var:,.0f}")
Q15: What is Standard Deviation?
Answer: Standard Deviation (SD) is the square root of variance β it measures dispersion in the same units as the original data, making it interpretable. Population SD: Ο = β(ΟΒ²). Sample SD: s = β(sΒ²). A small SD means data points are clustered near the mean. A large SD means data points are spread far from the mean. The Empirical Rule (68-95-99.7) states: for normal distributions, ~68% data within Β±1 SD, ~95% within Β±2 SD, ~99.7% within Β±3 SD from the mean.
π― Explain: Standard Deviation = Variance ka square root β same unit mein answer. Salary ka SD = βΉ9,910 (approx) β matlab most salaries mean (63,750) se βΉ9,910 ke andar hain. Empirical Rule: 68% employees βΉ63,750 Β± βΉ9,910 ke beech (βΉ53,840 - βΉ73,660). 95% employees Β±2 SD ke beech. SD chhota = data tightly clustered (consistent). SD bada = data widely spread (variable). Interview mein "Standard Deviation tells me how much data varies from the mean β I use the 68-95-99.7 rule for quick range estimation in normal distributions."
# Standard Deviation
sd = np.std(salaries, ddof=1) # Sample SD
print(f"Standard Deviation: βΉ{sd:,.0f}")
# Empirical Rule β 68-95-99.7
mean = np.mean(salaries)
print(f"68% range: βΉ{mean-sd:,.0f} to βΉ{mean+sd:,.0f}")
print(f"95% range: βΉ{mean-2*sd:,.0f} to βΉ{mean+2*sd:,.0f}")
# Verify β what % falls within 1 SD?
within_1sd = [s for s in salaries if mean-sd <= s <= mean+sd]
print(f"Within 1 SD: {len(within_1sd)}/{len(salaries)} = {len(within_1sd)/len(salaries)*100:.0f}%")
Q16: What is Interquartile Range (IQR)?
Answer: IQR measures the spread of the middle 50% of data. IQR = Q3 - Q1, where Q1 is the 25th percentile and Q3 is the 75th percentile. IQR is robust to outliers because it ignores extreme values. It is used in the IQR method for outlier detection: values below Q1 - 1.5ΓIQR or above Q3 + 1.5ΓIQR are considered outliers. Box plots visually represent IQR β the box spans Q1 to Q3, the line inside is the median, and whiskers extend to 1.5ΓIQR.
π― Explain: IQR = Q3 - Q1 β middle 50% data ka spread. Sorted salaries: [48K, 55K, 58K, 62K, 65K, 70K, 72K, 80K]. Q1 = 57,250 (25th percentile), Q3 = 70,500 (75th percentile). IQR = 70,500 - 57,250 = 13,250. Outlier boundaries: Lower = Q1 - 1.5ΓIQR = 37,375. Upper = Q3 + 1.5ΓIQR = 90,375. Koi bhi value is range ke bahar β outlier. Box plot mein yahi dikhta hai. Interview mein "I use IQR for outlier detection because it's robust to extreme values β values beyond 1.5ΓIQR from Q1/Q3 are flagged."
# IQR Calculation
Q1 = np.percentile(salaries, 25)
Q3 = np.percentile(salaries, 75)
IQR = Q3 - Q1
print(f"Q1: βΉ{Q1:,.0f}") # Q1: βΉ57,250
print(f"Q3: βΉ{Q3:,.0f}") # Q3: βΉ70,500
print(f"IQR: βΉ{IQR:,.0f}") # IQR: βΉ13,250
# Outlier boundaries
lower = Q1 - 1.5 * IQR
upper = Q3 + 1.5 * IQR
print(f"Lower Bound: βΉ{lower:,.0f}") # βΉ37,375
print(f"Upper Bound: βΉ{upper:,.0f}") # βΉ90,375
# Find outliers
outliers = [s for s in salaries if s < lower or s > upper]
print(f"Outliers: {outliers}") # [] β no outliers
Q17: What is Coefficient of Variation (CV)?
Answer: Coefficient of Variation (CV) is the ratio of standard deviation to the mean, expressed as a percentage. Formula: CV = (SD / Mean) Γ 100. It measures relative variability β allowing comparison of spread between datasets with different units or scales. Example: comparing salary variability between two companies even if one pays in USD and another in INR. A higher CV indicates more relative variation. CV is unitless β making it a universal comparability metric.
π― Explain: CV = relative spread measure β SD ko mean se divide karke percentage banao. CV = (9910 / 63750) Γ 100 = 15.5%. Iska matlab data mean se approximately 15.5% vary karta hai. CV kyu useful? Do datasets compare karo β Salary (mean 63K, SD 10K) vs Sales (mean 72K, SD 15K) β SD compare karna unfair hai kyunki scales alag hain. CV compare karo β Salary CV = 15.5%, Sales CV = 20.8% β Sales mein zyada variability hai. Interview mein "CV enables comparing variability across different scales β I use it when comparing spread between metrics with different units."
# Coefficient of Variation
sal_cv = (np.std(salaries, ddof=1) / np.mean(salaries)) * 100
print(f"Salary CV: {sal_cv:.1f}%")
sales = [85000,92000,45000,78000,65000,52000,88000,71000]
sales_cv = (np.std(sales, ddof=1) / np.mean(sales)) * 100
print(f"Sales CV: {sales_cv:.1f}%")
print(f"\nSales has {'more' if sales_cv > sal_cv else 'less'} relative variability")
Q18: What is the Five-Number Summary and Box Plot?
Answer: The Five-Number Summary consists of: Minimum, Q1 (25th percentile), Median (Q2/50th percentile), Q3 (75th percentile), Maximum. A Box Plot (Box-and-Whisker plot) visually represents this summary β the box spans Q1 to Q3 (IQR), the line inside is the median, whiskers extend to the minimum and maximum (within 1.5ΓIQR), and points beyond whiskers are outliers shown as dots. Box plots are excellent for comparing distributions across groups and identifying outliers.
π― Explain: Five-Number Summary = data ke 5 key points β Min, Q1, Median, Q3, Max. Box Plot inka visual representation hai. Box = Q1 se Q3 (middle 50%). Line = Median. Whiskers = 1.5ΓIQR ke andar ki min/max values. Dots = outliers. Side-by-side box plots se departments compare karo β kaunsi dept mein salary zyada spread hai, kahan outliers hain. Interview mein "I use box plots for quick distribution comparison across groups β they show median, spread, skewness, and outliers in one view." Pandas se: df.boxplot(column='Salary', by='Dept').
# Five-Number Summary
print(f"Min: βΉ{np.min(salaries):,}") # βΉ48,000
print(f"Q1: βΉ{np.percentile(salaries,25):,.0f}") # βΉ57,250
print(f"Median: βΉ{np.median(salaries):,.0f}") # βΉ63,500
print(f"Q3: βΉ{np.percentile(salaries,75):,.0f}") # βΉ70,500
print(f"Max: βΉ{np.max(salaries):,}") # βΉ80,000
# Pandas describe gives five-number summary + more
import pandas as pd
print(pd.Series(salaries).describe())
# Box Plot
import matplotlib.pyplot as plt
plt.boxplot(salaries, vert=True)
plt.title('Salary Distribution')
plt.ylabel('Salary (βΉ)')
plt.show()
π’ Category 4: Distributions & Skewness (Q19βQ24)
Q19: What is a Normal Distribution?
Answer: Normal Distribution (Gaussian Distribution) is a symmetric, bell-shaped probability distribution where most values cluster around the mean. Key properties: (1) Mean = Median = Mode (perfectly symmetric). (2) Defined by two parameters β mean (ΞΌ) and standard deviation (Ο). (3) Follows the Empirical Rule β 68% within Β±1Ο, 95% within Β±2Ο, 99.7% within Β±3Ο. (4) Total area under the curve = 1. Many natural phenomena follow normal distribution β heights, IQ scores, measurement errors. It is the foundation of most statistical tests.
π― Explain: Normal Distribution = bell curve β sabse common distribution. Beech mein tall (most values), sides mein tapering (fewer extreme values). Mean = Median = Mode β perfectly balanced. 68-95-99.7 rule apply hota hai. Heights, weights, exam scores β approximate normal hote hain. Statistical tests jaise z-test, t-test β normal distribution assume karte hain. Data normal hai ya nahi check karo β histogram dikhao, skewness check karo. Interview mein "Normal distribution is the backbone of statistics β most parametric tests assume normality."
# Generate and visualize Normal Distribution
import numpy as np
import matplotlib.pyplot as plt
normal_data = np.random.normal(loc=63750, scale=10000, size=10000)
plt.hist(normal_data, bins=50, edgecolor='black', alpha=0.7)
plt.axvline(np.mean(normal_data), color='red', label='Mean')
plt.title('Normal Distribution β Salary Simulation')
plt.legend()
plt.show()
Q20: What is Skewness?
Answer: Skewness measures the asymmetry of a distribution. Zero skewness = perfectly symmetric (normal). Positive skewness (right-skewed) = tail extends to the right, most values concentrated on the left β Mean > Median. Common in: income, property prices. Negative skewness (left-skewed) = tail extends to the left, most values concentrated on the right β Mean < Median. Common in: exam scores, retirement age. Formula uses the third standardized moment. Pandas: df['col'].skew(). Skewness between -0.5 and 0.5 is approximately symmetric.
π― Explain: Skewness = distribution kitni tilted hai. Skew = 0 β symmetric (bell curve). Positive skew β right mein tail β income data (most earn moderate, few earn very high). Negative skew β left mein tail β exam scores (most score high, few score very low). pd.Series(salaries).skew() se check karo. |skew| < 0.5 β approximately symmetric. 0.5-1 β moderately skewed. >1 β highly skewed. Highly skewed data mein mean misleading hai β median use karo. Interview mein "I check skewness to decide between mean and median β positive skew means I report median for better representation."
# Skewness calculation
s = pd.Series(salaries)
print(f"Skewness: {s.skew():.3f}")
# Interpretation
skew = s.skew()
if abs(skew) < 0.5:
print("Approximately symmetric")
elif skew > 0:
print("Right-skewed (positive) β Mean > Median")
else:
print("Left-skewed (negative) β Mean < Median")
print(f"Mean: {s.mean():,.0f} | Median: {s.median():,.0f}")
Q21: What is Kurtosis?
Answer: Kurtosis measures the "tailedness" of a distribution β how heavy or light the tails are compared to a normal distribution. Three types: Mesokurtic (kurtosis β 0) β normal distribution tails. Leptokurtic (kurtosis > 0) β heavier tails, sharper peak β more extreme values than normal. Platykurtic (kurtosis < 0) β lighter tails, flatter peak β fewer extreme values. Pandas uses excess kurtosis (normal = 0). High kurtosis means more outliers. Used in risk analysis β financial returns with high kurtosis have higher tail risk.
π― Explain: Kurtosis = tails kitni heavy hain. Normal distribution ka kurtosis = 0 (excess). Positive kurtosis β sharp peak, heavy tails β extreme values zyada β stock market returns (crashes). Negative kurtosis β flat peak, light tails β extreme values kam β uniform-like. pd.Series(data).kurtosis() se check karo. Finance mein important β "fat tails" = unexpected crashes ka risk zyada. Interview mein "Kurtosis tells me about tail risk β high kurtosis means more extreme values than a normal distribution would predict."
# Kurtosis
print(f"Kurtosis: {pd.Series(salaries).kurtosis():.3f}")
# Interpretation
kurt = pd.Series(salaries).kurtosis()
if kurt > 0:
print("Leptokurtic β heavy tails, more outliers")
elif kurt < 0:
print("Platykurtic β light tails, fewer outliers")
else:
print("Mesokurtic β normal-like tails")
Q22: What is Frequency Distribution?
Answer: Frequency Distribution organizes data into groups (bins/classes) and shows how many values fall into each group. Types: (1) Absolute frequency β count in each bin. (2) Relative frequency β proportion (count/total). (3) Cumulative frequency β running total. Created using: Pandas value_counts() for categorical, pd.cut() for numerical bins, or histograms for visual representation. Frequency distributions reveal data patterns β where values concentrate, gaps in data, and overall shape of distribution.
π― Explain: Frequency Distribution = data ko groups mein divide karo aur count karo. Ratings: 3 appears 2 times, 4 appears 4 times, 5 appears 2 times. Salary bins: 40-50K: 1 employee, 50-60K: 2, 60-70K: 2, 70-80K: 2, 80-90K: 1. Histogram yahi visually dikhata hai. value_counts() se categorical frequency. pd.cut() se numerical bins banao. Interview mein "I use frequency distributions to understand data patterns β value_counts for categories and histograms for continuous data."
# Categorical frequency
ratings = pd.Series([4,5,3,4,4,3,5,4])
print(ratings.value_counts().sort_index())
# Numerical bins
sal_bins = pd.cut(pd.Series(salaries),
bins=[40000,50000,60000,70000,80000,90000])
print(sal_bins.value_counts().sort_index())
Q23: What is the difference between a Histogram and a Bar Chart?
Answer: Histograms display the distribution of continuous numerical data using adjacent bars (no gaps) β each bar represents a range (bin) of values. Bar Charts display comparisons of categorical data using separated bars (with gaps) β each bar represents a category. Histograms show frequency distribution and shape (skewness). Bar Charts show comparisons between discrete categories. Histogram x-axis is continuous, Bar Chart x-axis is categorical. Histogram bar order is fixed (numerical), Bar Chart bars can be reordered.
π― Explain: Histogram = continuous data ki distribution β salary ranges. Bars touching hain (no gaps) kyunki data continuous hai. Bar Chart = categories ki comparison β department-wise count. Bars alag alag hain (gaps) kyunki categories discrete hain. Histogram se shape pata chalta hai β symmetric, skewed. Bar chart se comparison hota hai β kaunsa department zyada hai. Interview mein "Histograms for distribution of continuous data with no gaps, Bar charts for categorical comparisons with gaps between bars."
Q24: What is the Z-Score (Standard Score)?
Answer: Z-Score measures how many standard deviations a value is from the mean. Formula: Z = (X - ΞΌ) / Ο. A Z-score of 0 means the value equals the mean. Positive Z means above mean, negative means below. Z > 2 or Z < -2 suggests the value is unusual. Z > 3 or Z < -3 is a potential outlier. Z-scores standardize data β enabling comparison of values from different distributions. Example: comparing a student's performance in two different subjects with different scoring scales.
π― Explain: Z-Score = value mean se kitne SD door hai. Rohan salary βΉ80,000. Mean = βΉ63,750, SD = βΉ9,910. Z = (80,000 - 63,750) / 9,910 = 1.64. Matlab Rohan mean se 1.64 SD upar hai β unusual nahi (2 se kam). Meera βΉ48,000. Z = (48,000 - 63,750) / 9,910 = -1.59. Mean se 1.59 SD neeche. Z-score se different scales compare karo β salary aur sales ko same scale pe lao. Interview mein "Z-score standardizes values for cross-comparison and outlier detection β values beyond Β±3 are potential outliers."
# Z-Score calculation
mean = np.mean(salaries)
sd = np.std(salaries, ddof=1)
z_scores = [(x - mean) / sd for x in salaries]
names = ["Aarav","Ishita","Kabir","Diya","Rohan","Meera","Arjun","Kavya"]
for name, sal, z in zip(names, salaries, z_scores):
flag = " β οΈ OUTLIER" if abs(z) > 2 else ""
print(f"{name}: βΉ{sal:,} β Z = {z:+.2f}{flag}")
# SciPy method
from scipy import stats
print(stats.zscore(salaries))
π’ Category 5: Probability Basics & Percentiles (Q25βQ30)
Q25: What is Probability?
Answer: Probability is a measure of the likelihood of an event occurring β a value between 0 (impossible) and 1 (certain). Formula: P(event) = Number of favorable outcomes / Total number of outcomes. Probability of 0.5 means 50% chance. Three approaches: (1) Classical β equal likelihood (coin toss = 0.5). (2) Relative Frequency β based on observed data (30 out of 100 customers churned = 0.30). (3) Subjective β based on judgment (70% chance of rain). Probability is the foundation of inferential statistics.
π― Explain: Probability = kuch hone ki chance kitni hai β 0 se 1 ke beech. 0 = impossible, 1 = certain. Coin toss: Heads ki probability = 1/2 = 0.5 = 50%. Data Analysis mein: 8 employees mein se salary > 60K wale kitne? 5 hain. P(salary > 60K) = 5/8 = 0.625 = 62.5%. Customer churn: 100 mein se 30 churn kiye β P(churn) = 0.30. Interview mein "Probability quantifies uncertainty β I use it for calculating churn rates, conversion rates, and event likelihoods from historical data."
# Basic probability
total = len(salaries) # 8
above_60k = sum(1 for s in salaries if s > 60000) # 5
prob = above_60k / total
print(f"P(Salary > 60K) = {above_60k}/{total} = {prob:.3f} = {prob*100:.1f}%")
# P(Salary > 60K) = 5/8 = 0.625 = 62.5%
# Complementary probability
print(f"P(Salary β€ 60K) = {1-prob:.3f} = {(1-prob)*100:.1f}%")
# P(Salary β€ 60K) = 0.375 = 37.5%
Q26: What are the basic rules of Probability?
Answer: Key probability rules: (1) Range: 0 β€ P(A) β€ 1. (2) Complement: P(not A) = 1 - P(A). (3) Addition Rule: P(A or B) = P(A) + P(B) - P(A and B) for non-mutually exclusive events. For mutually exclusive events: P(A or B) = P(A) + P(B). (4) Multiplication Rule: P(A and B) = P(A) Γ P(B|A). For independent events: P(A and B) = P(A) Γ P(B). (5) Conditional: P(A|B) = P(A and B) / P(B) β probability of A given B has occurred.
π― Explain: Complement: P(Not Raining) = 1 - P(Raining). Addition: P(IT or HR) β IT ke employees + HR ke employees β overlap hatao (agar koi dono mein hai). Multiplication: P(IT and North) β independent hain toh P(IT) Γ P(North). Conditional: P(High Salary | IT Dept) β "agar IT mein hai toh high salary ki chance kitni?" Interview mein "Addition for OR scenarios, Multiplication for AND scenarios, Conditional for given scenarios β these three rules cover most probability questions."
# Probability rules with our data
depts = ["IT","HR","Finance","IT","Marketing","HR","Finance","Marketing"]
n = len(depts)
# P(IT)
p_it = depts.count("IT") / n
print(f"P(IT) = {p_it:.3f}") # 0.250
# Complement: P(Not IT)
print(f"P(Not IT) = {1-p_it:.3f}") # 0.750
# P(IT or HR) β mutually exclusive (can't be in both)
p_hr = depts.count("HR") / n
print(f"P(IT or HR) = {p_it + p_hr:.3f}") # 0.500
# Conditional: P(Salary>60K | IT)
it_salaries = [s for s,d in zip(salaries,depts) if d=="IT"]
p_high_given_it = sum(1 for s in it_salaries if s>60000) / len(it_salaries)
print(f"P(Salary>60K | IT) = {p_high_given_it:.3f}") # 0.000
Q27: What are Percentiles and Quartiles?
Answer: Percentiles divide data into 100 equal parts β the Pth percentile is the value below which P% of data falls. The 50th percentile = Median. Quartiles divide data into 4 equal parts: Q1 (25th percentile) β 25% below, Q2 (50th/Median) β 50% below, Q3 (75th percentile) β 75% below. Deciles divide into 10 parts. Percentiles are used for ranking (exam scores β "you scored in the 95th percentile"), salary benchmarking, and performance evaluation (top 10% performers).
π― Explain: Percentile = kitne percent data isse neeche hai. 90th percentile salary βΉ78,000 β 90% employees isse kam kamate hain. Quartiles = data ko 4 equal parts mein divide karo. Q1 = 25th percentile, Q2 = Median = 50th, Q3 = 75th. Exam results mein: "95th percentile" = 95% students se better score kiya. Salary benchmarking: "P50 salary = median salary β market competitive." Interview mein "I use percentiles for benchmarking β P50 for median salary, P75 for competitive positioning, P90 for top performer identification."
# Percentiles and Quartiles
print(f"10th Percentile: βΉ{np.percentile(salaries,10):,.0f}")
print(f"25th (Q1): βΉ{np.percentile(salaries,25):,.0f}")
print(f"50th (Median): βΉ{np.percentile(salaries,50):,.0f}")
print(f"75th (Q3): βΉ{np.percentile(salaries,75):,.0f}")
print(f"90th Percentile: βΉ{np.percentile(salaries,90):,.0f}")
# Percentile rank of a specific value
from scipy import stats
print(f"\nRohan (βΉ80K) is at {stats.percentileofscore(salaries,80000):.0f}th percentile")
print(f"Meera (βΉ48K) is at {stats.percentileofscore(salaries,48000):.0f}th percentile")
Q28: What is the difference between Correlation and Causation?
Answer: Correlation measures the statistical relationship between two variables β how they move together. Positive correlation: both increase together. Negative: one increases, other decreases. Causation means one variable directly causes a change in another. Key principle: correlation does not imply causation. Example: ice cream sales and drowning deaths are correlated (both increase in summer) β but ice cream does not cause drowning (common cause: hot weather). Establishing causation requires controlled experiments or careful causal analysis.
π― Explain: Correlation = do variables saath mein move karti hain. Causation = ek variable dusri ko change KARTI hai. Ice cream sales β aur drowning β β correlated β lekin ice cream drowning cause nahi karti β dono ka cause garmi hai. Spurious correlations bahut hoti hain β Nicolas Cage movies aur swimming pool drownings correlated hain β lekin koi connection nahi! Causation prove karne ke liye: controlled experiment, time-order, no confounding variables. Interview mein "I always caution that correlation is not causation β I look for confounding variables and use A/B testing to establish causal relationships."
Q29: What is the Empirical Rule (68-95-99.7)?
Answer: The Empirical Rule applies to normal (bell-shaped) distributions: approximately 68% of data falls within Β±1 standard deviation from the mean. 95% falls within Β±2 standard deviations. 99.7% falls within Β±3 standard deviations. Also called the Three-Sigma Rule. Practical use: if Mean = 63,750 and SD = 9,910, then 68% of employees earn between βΉ53,840 - βΉ73,660. Values beyond 3Ο (Β±29,730 from mean) are extremely rare β only 0.3% of data.
π― Explain: 68-95-99.7 rule = normal distribution ka quick estimate tool. Mean Β± 1 SD β 68% data. Mean Β± 2 SD β 95% data. Mean Β± 3 SD β 99.7% data. Salary example: Mean βΉ63,750, SD βΉ9,910. 68%: βΉ53,840 to βΉ73,660. 95%: βΉ43,930 to βΉ83,570. 99.7%: βΉ34,020 to βΉ93,480. βΉ93,480 se zyada β bahut rare (<0.3%). Quality control mein use hota hai β 6 Sigma (Β±6 SD = almost zero defects). Interview mein "I use the Empirical Rule for quick range estimation β 95% of normal data falls within 2 standard deviations of the mean."
# Empirical Rule demonstration
mean = np.mean(salaries)
sd = np.std(salaries, ddof=1)
print(f"Mean: βΉ{mean:,.0f} | SD: βΉ{sd:,.0f}")
print(f"\n68% range: βΉ{mean-sd:,.0f} to βΉ{mean+sd:,.0f}")
print(f"95% range: βΉ{mean-2*sd:,.0f} to βΉ{mean+2*sd:,.0f}")
print(f"99.7% range: βΉ{mean-3*sd:,.0f} to βΉ{mean+3*sd:,.0f}")
# Verify with actual data
for k in [1, 2, 3]:
within = sum(1 for s in salaries if mean-k*sd <= s <= mean+k*sd)
print(f"Within Β±{k}Ο: {within}/{len(salaries)} = {within/len(salaries)*100:.0f}%")
Q30: What is the difference between Standard Error and Standard Deviation?
Answer: Standard Deviation (SD) measures the spread of individual data points around the mean of a single dataset. Standard Error (SE) measures the spread of sample means around the true population mean β it quantifies the precision of the sample mean as an estimator. Formula: SE = SD / βn. As sample size increases, SE decreases (more precise estimate), but SD stays approximately the same. SD describes data variability, SE describes estimation precision. SE is used in confidence intervals and hypothesis testing.
π― Explain: SD = individual values kitne spread hain β "salaries ka spread βΉ9,910 hai". SE = sample mean kitna accurate hai β "sample mean population mean se βΉ3,505 ke andar hoga approximately." SE = SD / βn = 9,910 / β8 = 3,505. Sample size badhao β SE decrease hoga β zyada precise estimate. 100 employees sample lo: SE = 9910/β100 = 991 β bahut precise. SD data describe karta hai, SE estimation quality describe karta hai. Interview mein "SD measures data spread, SE measures estimation precision β SE decreases with larger samples."
# Standard Deviation vs Standard Error
n = len(salaries)
sd = np.std(salaries, ddof=1)
se = sd / np.sqrt(n)
print(f"Standard Deviation: βΉ{sd:,.0f}") # Data spread
print(f"Standard Error: βΉ{se:,.0f}") # Estimation precision
# SE decreases with larger sample
for size in [8, 30, 100, 1000]:
print(f"n={size:>4} β SE = βΉ{sd/np.sqrt(size):,.0f}")
# n= 8 β SE = βΉ3,505
# n= 30 β SE = βΉ1,810
# n= 100 β SE = βΉ991
# n=1000 β SE = βΉ313
π Quick Revision Table β 30 Questions at a Glance
| Q# | Question | One-Line Answer |
|---|---|---|
| Q1 | What is Statistics? | Science of collecting, analyzing, interpreting data |
| Q2 | Descriptive vs Inferential? | Descriptive=summarize, Inferential=predict & generalize |
| Q3 | Population vs Sample? | Population=entire group, Sample=representative subset |
| Q4 | Parameter vs Statistic? | Parameter=population (ΞΌ), Statistic=sample (xΜ) |
| Q5 | Data types? | Nominal, Ordinal (categorical) | Discrete, Continuous (numerical) |
| Q6 | Scales of measurement? | NOIR β Nominal, Ordinal, Interval, Ratio |
| Q7 | Mean? | Sum/count β sensitive to outliers |
| Q8 | Median? | Middle value sorted β robust to outliers |
| Q9 | Mode? | Most frequent value β only for categorical data |
| Q10 | When to use which? | SymmetricβMean, SkewedβMedian, CategoricalβMode |
| Q11 | Weighted Mean? | Different weights for different values β GPA, portfolio |
| Q12 | Mean-Median-Mode relationship? | Symmetric: equal | Right skew: Mode<Median<Mean |
| Q13 | Range? | Max - Min β simple but outlier-sensitive |
| Q14 | Variance? | Average squared deviation β n-1 for sample (Bessel's) |
| Q15 | Standard Deviation? | βVariance β same unit as data β 68-95-99.7 rule |
| Q16 | IQR? | Q3-Q1 β middle 50% spread β outlier detection 1.5ΓIQR |
| Q17 | Coefficient of Variation? | (SD/Mean)Γ100 β relative variability comparison |
| Q18 | Five-Number Summary & Box Plot? | Min, Q1, Median, Q3, Max β visual distribution + outliers |
| Q19 | Normal Distribution? | Bell curve β Mean=Median=Mode β 68-95-99.7 rule |
| Q20 | Skewness? | Distribution asymmetry β positive=right tail, negative=left |
| Q21 | Kurtosis? | Tail heaviness β leptokurtic=heavy, platykurtic=light |
| Q22 | Frequency Distribution? | Data grouped into bins with counts β histograms |
| Q23 | Histogram vs Bar Chart? | Histogram=continuous no gaps, Bar=categorical with gaps |
| Q24 | Z-Score? | (X-Mean)/SD β how many SDs from mean β outlier >Β±3 |
| Q25 | Probability? | Likelihood 0-1 β favorable/total outcomes |
| Q26 | Probability rules? | Addition (OR), Multiplication (AND), Complement, Conditional |
| Q27 | Percentiles & Quartiles? | Pth percentile = P% below β Q1=25th, Q2=50th, Q3=75th |
| Q28 | Correlation vs Causation? | Correlation=together, Causation=causes β correlation β causation |
| Q29 | Empirical Rule? | 68% within Β±1Ο, 95% within Β±2Ο, 99.7% within Β±3Ο |
| Q30 | Standard Error vs SD? | SD=data spread, SE=estimation precision (SD/βn) |
Thanks for Reading! π
Thanks for reading! Data Insights par aur bhi Power BI, Excel, SQL, Python, Statistics topics available hain β explore karo aur apni analytics journey strong banao! Happy Learning & Keep Exploring! π
β JatinAnalytics