Categorical Charts in Seaborn
Categorical Charts in Seaborn: Complete Guide
Categorical data analysis ka complete toolkit — Bar plots se lekar Swarm plots tak. Seaborn ke 7 powerful categorical chart types jo department-wise, gender-wise, aur har category-wise comparison professional level par dikhate hain.
📑 Is Guide Mein 7 Categorical Charts Cover Honge:
- sns.barplot() — Mean values + confidence intervals
- sns.countplot() — Category frequency counting
- sns.boxplot() — Distribution spread + outliers
- sns.violinplot() — Distribution shape visualization
- sns.swarmplot() — Individual data points (bee swarm)
- sns.stripplot() — Jittered individual points
- sns.catplot() — Figure-level unified categorical
1. sns.barplot() — Mean Values + Confidence Intervals
🔍 Kya Hai: Seaborn barplot() matplotlib bar() se alag hai — yeh automatically mean calculate karta hai aur confidence interval (error bar) dikhata hai. Statistical comparison ke liye designed hai, raw values ke liye nahi.
💻 Basic Version:
fig, ax = plt.subplots(figsize=(10, 6))
sns.barplot(data=df, x="Department", y="Salary", ax=ax)
ax.set_title("Average Salary by Department")
plt.savefig("barplot_basic.png", dpi=150, bbox_inches="tight")
plt.show()
💻 Professional Styled Version:
fig, ax = plt.subplots(figsize=(12, 6))
sns.barplot(data=df, x="Department", y="Salary",
hue="Gender", # split bars by gender
estimator="mean", # "mean","median","sum","std","min","max"
errorbar="ci", # "ci"(95% CI),"sd","se","pi", or None
ci=95, # confidence interval percentage
palette=["#3498db", "#e74c3c"], # custom colors
saturation=0.85, # color saturation (0-1)
width=0.7, # bar width
edgecolor="#2c3e50", # bar border
linewidth=1.2, # border thickness
capsize=0.1, # error bar cap width
errwidth=1.5, # error bar line width
dodge=True, # separate hue bars side-by-side
order=["IT","Finance","Marketing","HR"], # custom category order
ax=ax
)
ax.set_title("Average Salary by Department & Gender", fontsize=16, fontweight="bold", pad=15)
ax.set_xlabel("Department", fontsize=12)
ax.set_ylabel("Average Salary (₹)", fontsize=12)
ax.legend(title="Gender", fontsize=10)
sns.despine()
plt.savefig("barplot_styled.png", dpi=150, bbox_inches="tight")
plt.show()
💻 Horizontal + Custom Estimator:
fig, ax = plt.subplots(figsize=(10, 6))
sns.barplot(data=df, x="Salary", y="Department", # x,y swap = horizontal!
estimator="median", # median instead of mean
errorbar="sd", # standard deviation bars
palette="viridis",
orient="h", # horizontal orientation
ax=ax
)
# Data labels
for container in ax.containers:
ax.bar_label(container, fmt="₹%.0f", padding=5, fontsize=10)
ax.set_title("Median Salary (± Std Dev)", fontsize=15, fontweight="bold")
sns.despine()
plt.savefig("barplot_horizontal.png", dpi=150, bbox_inches="tight")
plt.show()
📋 barplot() Parameters:
| Parameter | Values | Description |
|---|---|---|
estimator | "mean","median","sum" | Bar height kaise calculate ho |
errorbar | "ci","sd","se","pi",None | Error bar type — None se remove |
hue | column name | Sub-category split |
dodge | True / False | Hue bars side-by-side ya overlapping |
order | list of categories | Custom category display order |
capsize | 0.05, 0.1, 0.15 | Error bar cap width |
2. sns.countplot() — Category Frequency Counter
🔍 Kya Hai: countplot() har category mein kitne records hain woh count karke bars dikhata hai — value_counts() ka visual version. barplot() mein y-column dena padta hai, countplot() mein sirf x ya y dena hai — count automatic hota hai.
💻 Professional Styled Version:
fig, ax = plt.subplots(figsize=(10, 6))
sns.countplot(data=df, x="Department",
hue="Gender", # split by gender
palette=["#3498db", "#e74c3c"],
edgecolor="#2c3e50",
linewidth=1,
saturation=0.85,
order=df["Department"].value_counts().index, # sorted by count
ax=ax
)
# Data labels on bars
for container in ax.containers:
ax.bar_label(container, fontsize=10, fontweight="bold", padding=3)
ax.set_title("Employee Count by Department & Gender", fontsize=16, fontweight="bold")
ax.set_ylabel("Count", fontsize=12)
ax.legend(title="Gender")
sns.despine()
plt.savefig("countplot.png", dpi=150, bbox_inches="tight")
plt.show()
💻 Percentage Countplot (Custom):
# Percentage labels instead of raw count
fig, ax = plt.subplots(figsize=(10, 6))
total = len(df)
sns.countplot(data=df, x="Rating", palette="RdYlGn",
order=sorted(df["Rating"].unique()), edgecolor="white", ax=ax)
for p in ax.patches:
pct = f"{p.get_height()/total*100:.1f}%"
ax.text(p.get_x() + p.get_width()/2, p.get_height() + 5,
pct, ha="center", fontweight="bold", fontsize=11)
ax.set_title("Rating Distribution (%)", fontsize=15, fontweight="bold")
sns.despine()
plt.savefig("countplot_pct.png", dpi=150, bbox_inches="tight")
plt.show()
3. sns.boxplot() — Distribution Spread + Outliers
🔍 Kya Hai: Seaborn boxplot() matplotlib version se zyada polished hai — automatic hue support, DataFrame integration, aur consistent styling. 5-number summary (min, Q1, median, Q3, max) + outliers ek chart mein.
💻 Professional Styled Version:
fig, ax = plt.subplots(figsize=(12, 6))
sns.boxplot(data=df, x="Department", y="Salary",
hue="Gender", # split by gender
palette=["#3498db", "#e74c3c"],
width=0.6, # box width
linewidth=1.5, # box border thickness
fliersize=4, # outlier dot size
flierprops={"marker": "o", "markerfacecolor": "#e74c3c", "alpha": 0.5},
medianprops={"color": "#2c3e50", "linewidth": 2},
whiskerprops={"linewidth": 1.5},
capprops={"linewidth": 1.5},
notch=True, # confidence interval notch
showmeans=True, # show mean marker
meanprops={"marker": "D", "markerfacecolor": "gold", "markersize": 6},
whis=1.5, # whisker length (IQR multiplier)
showfliers=True, # show outlier dots
ax=ax
)
ax.set_title("Salary Distribution by Dept & Gender", fontsize=16, fontweight="bold")
ax.set_ylabel("Salary (₹)", fontsize=12)
ax.legend(title="Gender")
ax.grid(axis="y", linestyle="--", alpha=0.3)
sns.despine()
plt.savefig("boxplot_seaborn.png", dpi=150, bbox_inches="tight")
plt.show()
📋 boxplot() Parameters:
| Parameter | Values | Description |
|---|---|---|
notch | True / False | Confidence interval notch dikhana |
showmeans | True / False | Mean marker (diamond) dikhana |
whis | 1.5, 2, 3 | Whisker length — IQR multiplier |
showfliers | True / False | Outlier dots dikhana |
fliersize | 3, 4, 5 | Outlier dot size |
4. sns.violinplot() — Distribution Shape Visualization
🔍 Kya Hai: Violin plot box plot + KDE combine karta hai — distribution ka actual shape dikhta hai. Bimodal distributions (2 peaks) violin mein clearly dikhti hain jo box plot mein invisible hoti hain.
💻 Professional Styled Version:
fig, ax = plt.subplots(figsize=(12, 7))
sns.violinplot(data=df, x="Department", y="Salary",
hue="Gender",
palette=["#3498db", "#e74c3c"],
split=True, # split violin — left=Male, right=Female
inner="quartile", # "box","quartile","point","stick",None
linewidth=1.5, # border thickness
bw_adjust=0.8, # KDE bandwidth
cut=0, # clip at data limits
density_norm="width", # "area","count","width"
saturation=0.85,
ax=ax
)
ax.set_title("Salary Distribution Shape — Split Violin", fontsize=16, fontweight="bold")
ax.set_ylabel("Salary (₹)", fontsize=12)
ax.legend(title="Gender")
ax.grid(axis="y", linestyle="--", alpha=0.3)
sns.despine()
plt.savefig("violin_split.png", dpi=150, bbox_inches="tight")
plt.show()
💻 Violin + Box Overlay (Best of Both):
fig, ax = plt.subplots(figsize=(12, 6))
# Violin (background shape)
sns.violinplot(data=df, x="Department", y="Salary",
inner=None, color="#d5e8f0", linewidth=0, alpha=0.6, ax=ax)
# Box (overlay details)
sns.boxplot(data=df, x="Department", y="Salary",
width=0.15, palette="Set2", linewidth=1.5,
showfliers=False, showmeans=True,
meanprops={"marker": "D", "markerfacecolor": "red", "markersize": 5},
ax=ax)
ax.set_title("Violin + Box Overlay (Shape + Summary)", fontsize=15, fontweight="bold")
sns.despine()
plt.savefig("violin_box_overlay.png", dpi=150, bbox_inches="tight")
plt.show()
📋 violinplot() Key Parameters:
| Parameter | Values | Description |
|---|---|---|
split | True / False | Hue groups ko ek violin mein split |
inner | "box","quartile","point","stick",None | Violin ke andar kya dikhna chahiye |
bw_adjust | 0.5, 0.8, 1.0 | KDE smoothness — lower=more detail |
density_norm | "area","count","width" | Violin width normalization method |
5. sns.swarmplot() — Bee Swarm (Individual Data Points)
🔍 Kya Hai: Swarmplot har individual data point ko dot ke roop mein dikhata hai — dots overlap nahi karte, bee swarm jaisa pattern banta hai. Small-medium datasets (<500 per category) ke liye perfect hai jahan har data point matter karta hai.
💻 Professional — Swarm + Box Overlay:
# Use smaller subset — swarm gets slow with 1000+ points
df_small = df.sample(200, random_state=42)
fig, ax = plt.subplots(figsize=(12, 6))
# Box (background)
sns.boxplot(data=df_small, x="Department", y="Salary",
color="#ecf0f1", width=0.5, showfliers=False, linewidth=1.5, ax=ax)
# Swarm (individual dots overlay)
sns.swarmplot(data=df_small, x="Department", y="Salary",
hue="Gender",
palette=["#3498db", "#e74c3c"],
size=5, # dot size
alpha=0.7, # transparency
edgecolor="#2c3e50", # dot border
linewidth=0.5, # dot border thickness
dodge=True, # separate hue groups
ax=ax
)
ax.set_title("Individual Salaries — Swarm + Box", fontsize=15, fontweight="bold")
ax.set_ylabel("Salary (₹)")
ax.legend(title="Gender")
sns.despine()
plt.savefig("swarm_box.png", dpi=150, bbox_inches="tight")
plt.show()
📋 swarmplot() Parameters:
| Parameter | Values | Description |
|---|---|---|
size | 3, 4, 5, 6 | Dot size — small rakhein dense data mein |
dodge | True / False | Hue groups alag alag position mein |
warn_thresh | 0.05 (default) | Overlap warning threshold |
6. sns.stripplot() — Jittered Individual Points
🔍 Kya Hai: Stripplot swarmplot jaisa hai lekin dots randomly jitter (spread) hote hain — overlap allowed hai lekin jitter se density roughly dikhti hai. Large datasets ke liye swarmplot se better hai kyunki swarm bahut slow ho jaata hai 500+ points par.
💻 Professional — Strip + Violin Overlay:
fig, ax = plt.subplots(figsize=(12, 6))
# Violin (background)
sns.violinplot(data=df, x="Department", y="Salary",
inner=None, color="#d5e8f0", alpha=0.5, linewidth=0, ax=ax)
# Strip (individual dots)
sns.stripplot(data=df, x="Department", y="Salary",
color="#3498db",
size=3, # dot size
alpha=0.4, # transparency (important!)
jitter=0.3, # random spread amount (0=no jitter, 0.5=max)
edgecolor="none", # no dot border — cleaner
ax=ax
)
ax.set_title("Salary — Violin + Strip (All 900 Points)", fontsize=15, fontweight="bold")
ax.set_ylabel("Salary (₹)")
sns.despine()
plt.savefig("strip_violin.png", dpi=150, bbox_inches="tight")
plt.show()
📋 stripplot() vs swarmplot():
| Feature | stripplot() | swarmplot() |
|---|---|---|
| Dots overlap? | Haan — alpha low rakhein | Nahi — algorithm se arrange |
| Speed | Fast — any size dataset | Slow — <500 per group best |
| Best for | Large datasets + overlays | Small datasets, exact view |
| Jitter control | Manual (jitter= param) | Automatic (algorithm) |
7. sns.catplot() — Figure-Level Unified Categorical
🔍 Kya Hai: catplot() Seaborn ka figure-level function hai jo sabhi categorical charts (bar, count, box, violin, swarm, strip, point) ko kind= parameter se control karta hai. col/row se automatic faceted panels bante hain — ek function se sab kuch.
💻 Faceted Box Plot — Department × Gender Grid:
g = sns.catplot(data=df, x="Rating", y="Salary",
col="Department", # separate panel per department
kind="box", # "bar","count","box","violin","swarm","strip","point"
palette="RdYlGn",
col_wrap=2, # 2 panels per row
height=4, # panel height
aspect=1.3, # width = height × aspect
showfliers=False,
linewidth=1.2
)
g.fig.suptitle("Salary by Rating — Per Department", fontsize=16, fontweight="bold", y=1.03)
plt.savefig("catplot_box.png", dpi=150, bbox_inches="tight")
plt.show()
💻 Faceted Bar Plot — Row + Col:
g = sns.catplot(data=df, x="Department", y="Salary",
row="Gender", # rows = Gender
kind="bar", # bar plot mode
palette="Set2",
height=4,
aspect=2.5,
errorbar="sd",
capsize=0.1
)
g.fig.suptitle("Salary by Department — Gender Rows", fontsize=16, fontweight="bold", y=1.02)
plt.savefig("catplot_bar.png", dpi=150, bbox_inches="tight")
plt.show()
💻 Point Plot (Mean + CI Line Chart):
g = sns.catplot(data=df, x="Rating", y="Salary",
hue="Department",
kind="point", # point = line + CI (like line chart for categories)
height=5,
aspect=1.8,
markers=["o", "s", "^", "D"], # different markers per hue
linestyles=["-", "--", "-.", ":"],
palette="Set2",
capsize=0.1,
errorbar="ci",
dodge=0.3 # spread hue groups slightly
)
g.fig.suptitle("Salary Trend by Rating & Department", fontsize=16, fontweight="bold", y=1.03)
plt.savefig("catplot_point.png", dpi=150, bbox_inches="tight")
plt.show()
📋 catplot() All kind= Options:
| kind= | Equivalent | Purpose |
|---|---|---|
"bar" | sns.barplot() | Mean + CI bars |
"count" | sns.countplot() | Frequency count bars |
"box" | sns.boxplot() | 5-number summary + outliers |
"violin" | sns.violinplot() | Distribution shape |
"swarm" | sns.swarmplot() | Non-overlapping dots |
"strip" | sns.stripplot() | Jittered dots |
"point" | sns.pointplot() | Mean + CI line chart |
Quick Reference: Categorical Chart Decision Guide
| Purpose | Chart | Best For |
|---|---|---|
| Mean comparison + CI | sns.barplot() | Statistical mean comparison |
| Count/frequency | sns.countplot() | Category size counting |
| Spread + outliers | sns.boxplot() | 5-number summary + outliers |
| Distribution shape | sns.violinplot() | Bimodal detection, KDE shape |
| Individual points (small data) | sns.swarmplot() | <500 points, no overlap |
| Individual points (large data) | sns.stripplot() | Any size, fast rendering |
| Faceted multi-panel | sns.catplot() | col=/row= automatic panels |
Next Post: Part 3
Next part mein hum cover karenge: Relationship & Regression Charts — scatterplot(), lineplot(), regplot(), lmplot(), relplot() aur statistical relationship visualization ke complete tools.
Happy Visualizing! 📊🚀
💬 Comments (0)
Loading comments...