Distribution Charts in Seaborn
Distribution Charts in Seaborn: Complete Guide
Data analysis ka pehla step hai distribution samajhna — data kaise spread hai, skewed hai ya normal hai, outliers kahan hain. Seaborn ke 6 powerful distribution charts se yeh sab ek glance mein dikhta hai.
📑 Is Guide Mein 6 Distribution Charts Cover Honge:
- sns.histplot() — Enhanced Histogram with KDE overlay
- sns.kdeplot() — Smooth Kernel Density Estimation curve
- sns.rugplot() — Individual data point markers on axis
- sns.ecdfplot() — Empirical Cumulative Distribution
- sns.displot() — Figure-level unified distribution plot
- Combined Distribution Analysis — Hist + KDE + Rug together
⚙️ Setup — Har Example Ke Liye Common Code
import seaborn as sns
import matplotlib.pyplot as plt
import pandas as pd
import numpy as np
# Sample Dataset — Employee Data
np.random.seed(42)
df = pd.DataFrame({
"Salary": np.concatenate([
np.random.normal(55000, 10000, 400), # HR/Ops cluster
np.random.normal(82000, 15000, 350), # IT/Finance cluster
np.random.normal(120000, 20000, 150) # Senior Management
]),
"Department": np.random.choice(["IT","HR","Finance","Marketing"], 900),
"Experience": np.random.uniform(1, 20, 900),
"Age": np.random.randint(22, 60, 900),
"Gender": np.random.choice(["Male", "Female"], 900),
"Rating": np.random.choice([1,2,3,4,5], 900, p=[0.05,0.1,0.3,0.35,0.2])
})
# Global Seaborn theme
sns.set_theme(style="whitegrid", font_scale=1.1)
1. sns.histplot() — Enhanced Histogram
🔍 Kya Hai: Seaborn ka histplot() matplotlib ke hist() ka upgraded version hai — KDE overlay, hue-based split, multiple stat modes, aur DataFrame direct support ke saath. Distribution analysis ka primary tool.
💻 Basic Version:
fig, ax = plt.subplots(figsize=(10, 6))
sns.histplot(data=df, x="Salary", ax=ax)
ax.set_title("Salary Distribution")
plt.savefig("hist_basic.png", dpi=150, bbox_inches="tight")
plt.show()
💻 Professional Styled Version:
fig, ax = plt.subplots(figsize=(12, 6))
sns.histplot(data=df, x="Salary",
bins=30, # number of bins
kde=True, # overlay KDE curve
color="#3498db", # bar color
edgecolor="white", # bar border
linewidth=0.8, # border thickness
alpha=0.7, # transparency
stat="count", # "count","density","probability","percent","frequency"
element="bars", # "bars","step","poly"
fill=True, # fill bars with color
kde_kws={ # KDE curve styling
"color": "#e74c3c",
"linewidth": 2.5,
"label": "KDE Curve"
},
ax=ax
)
# Mean & Median lines
ax.axvline(df["Salary"].mean(), color="red", linewidth=2, linestyle="--",
label=f"Mean: ₹{df['Salary'].mean():,.0f}")
ax.axvline(df["Salary"].median(), color="green", linewidth=2, linestyle="-.",
label=f"Median: ₹{df['Salary'].median():,.0f}")
ax.set_title("Salary Distribution with KDE", fontsize=16, fontweight="bold", pad=15)
ax.set_xlabel("Salary (₹)", fontsize=12)
ax.set_ylabel("Count", fontsize=12)
ax.legend(fontsize=11)
sns.despine()
plt.savefig("hist_styled.png", dpi=150, bbox_inches="tight")
plt.show()
💻 Hue Split — Department-wise Distribution:
fig, ax = plt.subplots(figsize=(12, 6))
sns.histplot(data=df, x="Salary",
hue="Department", # split by category — separate colors!
multiple="stack", # "layer","dodge","stack","fill"
bins=25,
palette="Set2", # color palette
edgecolor="white",
alpha=0.8,
ax=ax
)
ax.set_title("Salary Distribution by Department (Stacked)", fontsize=15, fontweight="bold")
sns.despine()
plt.savefig("hist_hue.png", dpi=150, bbox_inches="tight")
plt.show()
📋 All Parameters:
| Parameter | Values | Description |
|---|---|---|
data | DataFrame | Source DataFrame |
x / y | column name | x=vertical hist, y=horizontal hist |
hue | column name | Category-wise color split |
bins | 10, 20, 30, "auto" | Number of bins |
kde | True / False | Overlay smooth KDE curve |
stat | "count","density","probability","percent" | Y-axis metric — count ya percentage |
multiple | "layer","dodge","stack","fill" | Hue groups kaise display hon |
element | "bars","step","poly" | Bar style — step = outline only |
palette | "Set1","Set2","viridis","husl" | Color palette for hue groups |
cumulative | True / False | Cumulative histogram |
2. sns.kdeplot() — Kernel Density Estimation
🔍 Kya Hai: KDE plot histogram ka smooth continuous version hai — bars ki jagah smooth curve se distribution dikhata hai. Multiple distributions overlap karna, bimodal patterns detect karna, aur probability density visualize karna — iske liye KDE best hai.
💻 Basic Version:
fig, ax = plt.subplots(figsize=(10, 6))
sns.kdeplot(data=df, x="Salary", ax=ax)
ax.set_title("Salary Density Curve")
plt.savefig("kde_basic.png", dpi=150, bbox_inches="tight")
plt.show()
💻 Professional — Multiple Groups + Filled:
fig, ax = plt.subplots(figsize=(12, 6))
sns.kdeplot(data=df, x="Salary",
hue="Department", # split by department
fill=True, # fill area under curve
alpha=0.4, # fill transparency
linewidth=2, # curve line thickness
palette="Set2", # color palette
bw_adjust=0.8, # bandwidth (lower=more detail, higher=smoother)
common_norm=False, # each group separately normalized
cut=0, # clip curve at data limits (no extending)
ax=ax
)
ax.set_title("Salary Density by Department", fontsize=16, fontweight="bold")
ax.set_xlabel("Salary (₹)", fontsize=12)
ax.set_ylabel("Density", fontsize=12)
sns.despine()
plt.savefig("kde_groups.png", dpi=150, bbox_inches="tight")
plt.show()
💻 2D KDE — Bivariate Density (Contour):
fig, ax = plt.subplots(figsize=(10, 7))
sns.kdeplot(data=df, x="Experience", y="Salary",
fill=True, # filled contours
cmap="YlOrRd", # colormap for density
levels=15, # number of contour levels
thresh=0.05, # minimum density threshold
alpha=0.8,
ax=ax
)
# Overlay scatter for actual data points
ax.scatter(df["Experience"], df["Salary"], color="black", alpha=0.1, s=5, zorder=2)
ax.set_title("Experience vs Salary — 2D Density", fontsize=15, fontweight="bold")
ax.set_xlabel("Experience (Years)", fontsize=12)
ax.set_ylabel("Salary (₹)", fontsize=12)
plt.savefig("kde_2d.png", dpi=150, bbox_inches="tight")
plt.show()
📋 KDE-Specific Parameters:
| Parameter | Values | Description |
|---|---|---|
fill | True / False | Curve ke neeche area fill karna |
bw_adjust | 0.5, 0.8, 1.0, 1.5 | Bandwidth — low=detail, high=smooth |
cut | 0, 1, 2, 3 | Curve extension beyond data (0=clip) |
common_norm | True / False | All groups same normalization ya separate |
levels | 5, 10, 15, 20 | 2D KDE mein contour levels count |
cmap | "YlOrRd","Blues","viridis" | 2D KDE mein colormap |
3. sns.rugplot() — Individual Data Point Markers
🔍 Kya Hai: Rugplot axis par chhote vertical tick marks lagata hai — har mark ek actual data point represent karta hai. Akela rarely use hota hai, lekin histogram ya KDE ke saath combine karne se data density clearly dikhti hai.
💻 Combined — KDE + Rug:
fig, ax = plt.subplots(figsize=(12, 6))
# KDE curve
sns.kdeplot(data=df, x="Salary", fill=True, color="#3498db", alpha=0.3, linewidth=2, ax=ax)
# Rug marks on x-axis
sns.rugplot(data=df, x="Salary",
height=0.05, # tick mark height (0-1 fraction of plot)
color="#e74c3c", # tick color
alpha=0.3, # transparency (important for overlapping!)
linewidth=1, # tick thickness
ax=ax
)
ax.set_title("Salary Distribution — KDE + Rug", fontsize=15, fontweight="bold")
ax.set_xlabel("Salary (₹)", fontsize=12)
sns.despine()
plt.savefig("kde_rug.png", dpi=150, bbox_inches="tight")
plt.show()
💻 Hue-Based Rug with Histogram:
fig, ax = plt.subplots(figsize=(12, 6))
sns.histplot(data=df, x="Salary", hue="Gender", kde=True,
multiple="layer", alpha=0.5, palette=["#3498db","#e74c3c"], ax=ax)
sns.rugplot(data=df, x="Salary", hue="Gender",
height=0.03,
palette=["#3498db","#e74c3c"],
alpha=0.2,
ax=ax
)
ax.set_title("Salary by Gender — Hist + KDE + Rug", fontsize=15, fontweight="bold")
sns.despine()
plt.savefig("hist_kde_rug.png", dpi=150, bbox_inches="tight")
plt.show()
📋 Rug Parameters:
| Parameter | Values | Description |
|---|---|---|
height | 0.03, 0.05, 0.1 | Tick mark ki height (plot fraction) |
alpha | 0.1, 0.2, 0.3 | Low rakhein — overlapping marks dense areas mein |
expand_margins | True / False | Rug ke liye extra space add karna |
4. sns.ecdfplot() — Empirical Cumulative Distribution
🔍 Kya Hai: ECDF plot data ki cumulative probability dikhata hai — x-axis par value aur y-axis par "kitne percent data is value se neeche hai". Percentile analysis, threshold decisions, aur distribution comparison ke liye histogram se zyada informative hai.
💻 Basic Version:
fig, ax = plt.subplots(figsize=(10, 6))
sns.ecdfplot(data=df, x="Salary",
stat="proportion", # "proportion" (0-1) or "count" (raw)
complementary=False, # True = survival function (1 - CDF)
linewidth=2.5,
color="#3498db",
ax=ax
)
ax.set_title("Salary ECDF — Cumulative Distribution", fontsize=15, fontweight="bold")
ax.set_xlabel("Salary (₹)")
ax.set_ylabel("Proportion")
sns.despine()
plt.savefig("ecdf_basic.png", dpi=150, bbox_inches="tight")
plt.show()
💻 Professional — Multi-Group + Percentile Lines:
fig, ax = plt.subplots(figsize=(12, 7))
sns.ecdfplot(data=df, x="Salary",
hue="Department",
linewidth=2.5,
palette="Set2",
ax=ax
)
# Percentile reference lines
for pct, style in [(0.25, ":"), (0.50, "--"), (0.75, "-.")]:
val = df["Salary"].quantile(pct)
ax.axhline(y=pct, color="gray", linestyle=style, alpha=0.5)
ax.axvline(x=val, color="gray", linestyle=style, alpha=0.5)
ax.text(val + 1000, pct + 0.02, f"P{int(pct*100)}: ₹{val:,.0f}",
fontsize=9, color="#7f8c8d")
ax.set_title("Salary ECDF by Department + Percentiles", fontsize=15, fontweight="bold")
ax.set_xlabel("Salary (₹)", fontsize=12)
ax.set_ylabel("Cumulative Proportion", fontsize=12)
sns.despine()
plt.savefig("ecdf_percentiles.png", dpi=150, bbox_inches="tight")
plt.show()
5. sns.displot() — Figure-Level Distribution Plot
🔍 Kya Hai: displot() Seaborn ka figure-level function hai jo histplot, kdeplot, aur ecdfplot teeno ko kind= parameter se control karta hai. Iska sabse powerful feature hai col= aur row= parameters jo automatic multi-panel faceted charts banate hain.
💻 Faceted Histogram — Department-wise Panels:
# Faceted histogram — separate panel for each department
g = sns.displot(data=df, x="Salary",
col="Department", # separate column for each department
kind="hist", # "hist", "kde", "ecdf"
kde=True, # overlay KDE on histogram
bins=20,
col_wrap=2, # 2 charts per row, wrap to next line
height=4, # each panel height (inches)
aspect=1.5, # width = height × aspect
color="#3498db",
edgecolor="white"
)
g.fig.suptitle("Salary Distribution — Each Department", fontsize=16, fontweight="bold", y=1.03)
plt.savefig("displot_faceted.png", dpi=150, bbox_inches="tight")
plt.show()
💻 Faceted KDE — Row + Col Grid:
# 2D facet grid — Department (columns) × Gender (rows)
g = sns.displot(data=df, x="Salary",
col="Department", # columns
row="Gender", # rows
kind="kde", # KDE curves
fill=True,
height=3.5,
aspect=1.3,
palette="Set2",
alpha=0.6
)
g.fig.suptitle("Salary KDE — Department × Gender Grid", fontsize=16, fontweight="bold", y=1.03)
plt.savefig("displot_grid.png", dpi=150, bbox_inches="tight")
plt.show()
💻 ECDF Mode:
# ECDF mode with hue
g = sns.displot(data=df, x="Salary",
hue="Gender",
kind="ecdf", # ECDF mode
height=5,
aspect=1.8,
palette=["#3498db", "#e74c3c"],
linewidth=2.5
)
g.fig.suptitle("Salary ECDF by Gender", fontsize=15, fontweight="bold", y=1.03)
plt.savefig("displot_ecdf.png", dpi=150, bbox_inches="tight")
plt.show()
📋 displot() Parameters:
| Parameter | Values | Description |
|---|---|---|
kind | "hist", "kde", "ecdf" | Chart type select karna |
col | column name | Category-wise separate columns |
row | column name | Category-wise separate rows |
col_wrap | 2, 3, 4 | Max columns per row, phir wrap |
height | 3, 4, 5 (inches) | Each panel ki height |
aspect | 1.0, 1.3, 1.5 | Width = height × aspect |
6. Combined Distribution Analysis — Complete EDA Template
🔍 Kya Hai: Production-ready distribution analysis — 4 panels mein complete picture: Histogram+KDE, KDE groups, ECDF, aur Box+Violin combined. Copy-paste ready template jo EDA mein directly use ho.
💻 Complete 4-Panel Distribution Dashboard:
fig, axes = plt.subplots(2, 2, figsize=(16, 12))
fig.patch.set_facecolor("#f8fafc")
# Panel 1: Histogram + KDE + Rug
sns.histplot(data=df, x="Salary", bins=30, kde=True, color="#3498db",
edgecolor="white", alpha=0.7, ax=axes[0][0])
sns.rugplot(data=df, x="Salary", height=0.03, color="#e74c3c", alpha=0.2, ax=axes[0][0])
axes[0][0].axvline(df["Salary"].mean(), color="red", linestyle="--", linewidth=1.5,
label=f"Mean: ₹{df['Salary'].mean():,.0f}")
axes[0][0].set_title("📊 Histogram + KDE + Rug", fontweight="bold")
axes[0][0].legend(fontsize=9)
# Panel 2: KDE by Department
sns.kdeplot(data=df, x="Salary", hue="Department", fill=True,
alpha=0.4, linewidth=2, palette="Set2", common_norm=False, ax=axes[0][1])
axes[0][1].set_title("📈 KDE by Department", fontweight="bold")
# Panel 3: ECDF with Percentiles
sns.ecdfplot(data=df, x="Salary", hue="Gender",
linewidth=2.5, palette=["#3498db","#e74c3c"], ax=axes[1][0])
axes[1][0].axhline(y=0.5, color="gray", linestyle="--", alpha=0.5)
axes[1][0].set_title("📉 ECDF by Gender", fontweight="bold")
# Panel 4: Box + Strip (distribution + individual points)
sns.boxplot(data=df, x="Department", y="Salary", palette="Set2",
width=0.5, fliersize=3, ax=axes[1][1])
sns.stripplot(data=df, x="Department", y="Salary", color="black",
alpha=0.1, size=2, jitter=True, ax=axes[1][1])
axes[1][1].set_title("📦 Box + Strip by Department", fontweight="bold")
# Clean up all panels
for ax in axes.flat:
sns.despine(ax=ax)
fig.suptitle("🏢 Salary Distribution — Complete EDA Dashboard",
fontsize=18, fontweight="bold", y=1.02)
plt.tight_layout()
plt.savefig("distribution_dashboard.png", dpi=150, bbox_inches="tight",
facecolor="#f8fafc")
plt.show()
💻 Quick Stats Print (Pair with Dashboard):
# Quick distribution stats — print alongside dashboard
col = "Salary"
print(f"📊 Distribution Stats for '{col}':")
print(f" Mean: ₹{df[col].mean():,.0f}")
print(f" Median: ₹{df[col].median():,.0f}")
print(f" Std Dev: ₹{df[col].std():,.0f}")
print(f" Skewness: {df[col].skew():.3f}")
print(f" Kurtosis: {df[col].kurtosis():.3f}")
print(f" Range: ₹{df[col].min():,.0f} — ₹{df[col].max():,.0f}")
print(f" IQR: ₹{df[col].quantile(0.75) - df[col].quantile(0.25):,.0f}")
Quick Reference: Distribution Charts Decision Guide
| Purpose | Chart | Function |
|---|---|---|
| Frequency distribution dekhna | Histogram | sns.histplot() |
| Smooth density curve | KDE | sns.kdeplot() |
| Individual data points | Rug | sns.rugplot() |
| Cumulative probability | ECDF | sns.ecdfplot() |
| Multi-panel faceted view | Facet Grid | sns.displot(col=, row=) |
| 2D density contour | 2D KDE | sns.kdeplot(x=, y=) |
| Complete EDA dashboard | Combined | subplots + multiple charts |
Next Post: Part 2
Next part mein hum cover karenge: Categorical Charts — barplot(), countplot(), boxplot(), violinplot(), swarmplot(), stripplot() aur catplot() ke saath professional categorical data visualization.
Happy Visualizing! 📊🚀
💬 Comments (0)
Loading comments...