Data Insights Case Studies — Part B
Case Study #2: Spotify Analysis (Python + Pandas) 🐍
Part B — Same Spotify case study, ab Python aur Pandas ke saath. Real-world analytics jab tumhare paas SQL nahi hai — sirf CSV files hain. Listener engagement, song duration analysis, YoY growth, revenue comparison, aur churn analysis — 5 detailed questions + final recommendations. Data Insights par.
🔧 Setup — Load Data in Pandas
Assume tumhare paas 5 CSV files hain (SQL tables ke corresponding). Pehle sab load karo:
import pandas as pd import numpy as np from datetime import datetime # Load all CSV files artists = pd.read_csv('artists.csv') songs = pd.read_csv('songs.csv') streams = pd.read_csv('streams.csv', parse_dates=['stream_date']) listeners = pd.read_csv('listeners.csv', parse_dates=['signup_date']) revenue = pd.read_csv('revenue.csv') # Convert songs release_date to datetime songs['release_date'] = pd.to_datetime(songs['release_date']) # Quick check print("Data Shape Overview:") print(f"Artists: {artists.shape}") print(f"Songs: {songs.shape}") print(f"Streams: {streams.shape}") print(f"Listeners: {listeners.shape}") print(f"Revenue: {revenue.shape}")parse_dates parameter directly load karte waqt datetime columns convert kar deta hai — later manual conversion se bachne ke liye. Production analysis mein hamesha use karo. Q6: Listener Engagement — Avg Streams per User
❓ Question: Analyze listener engagement — total streams per user, unique songs listened, favorite genre, avg session duration, and segment users into Power Users (100+ streams), Regular (30-99), Casual (10-29), Inactive (< 10).
💡 Step-by-step Approach:
1. Merge tables: streams + listeners + songs + artists (4-way merge)
2. GroupBy listener_id: Aggregate multiple metrics
3. Favorite genre: Har listener ka most-listened genre (mode)
4. Segmentation: pd.cut() se buckets banao based on stream count
5. Sort: By total_streams DESC
# Step 1: Merge all necessary tables df = streams.merge(listeners, on='listener_id') df = df.merge(songs, on='song_id') df = df.merge(artists[['artist_id', 'genre']], on='artist_id') # Step 2: Listener-level aggregations listener_engagement = df.groupby( ['listener_id', 'listener_name', 'country_x', 'subscription_type'] ).agg( total_streams=('stream_id', 'count'), unique_songs=('song_id', 'nunique'), avg_session_sec=('stream_duration_sec', 'mean'), total_listen_time=('stream_duration_sec', 'sum') ).reset_index() # Step 3: Find favorite genre for each listener fav_genre = df.groupby('listener_id')['genre'].agg( lambda x: x.value_counts().index[0] ).reset_index() fav_genre.columns = ['listener_id', 'favorite_genre'] # Merge favorite genre back listener_engagement = listener_engagement.merge(fav_genre, on='listener_id') # Step 4: Create user segments listener_engagement['segment'] = pd.cut( listener_engagement['total_streams'], bins=[0, 10, 30, 100, np.inf], labels=['Inactive', 'Casual', 'Regular', 'Power User'] ) # Step 5: Round avg_session_sec and total_listen_time (in minutes) listener_engagement['avg_session_sec'] = listener_engagement['avg_session_sec'].round(0) listener_engagement['total_listen_hours'] = ( listener_engagement['total_listen_time'] / 3600 ).round(1) # Sort by engagement result = listener_engagement.sort_values('total_streams', ascending=False).head(10) print(result)📈 Expected Output (Top 10):
| listener_id | name | country | subscription | streams | songs | avg_sec | hours | fav_genre | segment |
|---|---|---|---|---|---|---|---|---|---|
| L001 | Raj Sharma | India | Premium | 285 | 145 | 218 | 17.3 | Bollywood | Power User L048 |
| Priya Menon | India | Premium | 258 | 132 | 225 | 16.1 | Bollywood | Power User L002 | Emma Wilson |
| UK | Premium | 195 | 108 | 210 | 11.4 | Pop | Power User L005 | Jake Park | S.Korea |
| Free | 168 | 82 | 165 | 7.7 | K-Pop | Power User L015 | Sara Kim | S.Korea | Premium |
| 142 | 95 | 195 | 7.7 | K-Pop | Power User L022 | Alex Chen | USA | Free | 78 |
| 55 | 158 | 3.4 | Hip-Hop | Regular L003 | Mike Johnson | USA | Free | 65 | 48 |
| 145 | 2.6 | Hip-Hop | Regular L038 | Lisa Brown | USA | Premium | 52 | 38 | 205 |
| 3.0 | Pop | Regular L041 | Kim Sooyoung | S.Korea | Free | 28 | 22 | 175 | 1.4 |
| K-Pop | Casual L050 | Rahul Yadav | India | Free | 15 | 12 | 155 | 0.6 | Punjabi |
Casual
📊 Segment Summary:
# Segment-wise summary segment_summary = listener_engagement.groupby('segment', observed=True).agg( user_count=('listener_id', 'count'), avg_streams=('total_streams', 'mean'), total_streams=('total_streams', 'sum'), avg_listen_hours=('total_listen_hours', 'mean') ).round(1)
# Add percentage of total users
segment_summary['pct_of_users'] = (
segment_summary['user_count'] * 100
/ segment_summary['user_count'].sum()
).round(1)
print(segment_summary)
| segment | users | avg_streams | total_streams | avg_hours | pct_users |
|---|---|---|---|---|---|
| Power User | 45 | 168.5 | 7,58,250 | 12.8 | 8.1% Regular |
| 128 | 52.3 | 6,69,440 | 4.2 | 23.2% Casual | 245 |
| 18.5 | 4,53,250 | 1.5 | 44.4% Inactive | 133 | 4.8 |
6,38,40 0.4 24.1%
Insight 1 — Power Users Are Gold: 8.1% users (45) generate 39% of total streams (7.58L out of ~19.4L). Yeh users platform ke revenue backbone hain — treat them like VIPs. Personalized playlists, early access to new features, exclusive content.
Insight 2 — Inactive User Crisis: 24.1% users (133) sirf 3% streams generate karte hain (avg 4.8 streams total). Yeh silent churn indication hai — soon uninstall karenge. Re-engagement email campaigns launch karo pehle.
Insight 3 — Casual Users Are Growth Opportunity: 44.4% users Casual segment mein hain — biggest bucket. Agar inhe Regular tak convert karein, streams double ho sakte hain. Push notifications, personalized recommendations use karo.
Q7: Song Duration vs Popularity Analysis
❓ Question: Do shorter songs perform better? Analyze song popularity across duration buckets — Short (<180 sec), Medium (180-240 sec), Long (240-300 sec), Very Long (>300 sec). Show avg streams, completion rate, and unique listeners per bucket.
💡 Approach:
1. Songs table mein duration_bucket column create karo pd.cut() se
2. Streams table se aggregate karo per song
3. Merge kar ke songs ke saath
4. GroupBy duration_bucket, calculate metrics
5. Completion rate = avg_stream_duration / song_duration
# Step 1: Create duration buckets in songs table songs['duration_bucket'] = pd.cut( songs['duration_sec'], bins=[0, 180, 240, 300, np.inf], labels=['Short (<3min)', 'Medium (3-4min)', 'Long (4-5min)', 'Very Long (>5min)'] ) # Step 2: Aggregate stream metrics per song song_metrics = streams.groupby('song_id').agg( total_streams=('stream_id', 'count'), unique_listeners=('listener_id', 'nunique'), avg_stream_duration=('stream_duration_sec', 'mean') ).reset_index() # Step 3: Merge with songs songs_with_metrics = songs.merge(song_metrics, on='song_id') # Step 4: Calculate completion rate per song songs_with_metrics['completion_pct'] = ( songs_with_metrics['avg_stream_duration'] * 100 / songs_with_metrics['duration_sec'] ).round(1) # Step 5: Aggregate by duration bucket bucket_analysis = songs_with_metrics.groupby( 'duration_bucket', observed=True ).agg( songs_count=('song_id', 'count'), avg_duration=('duration_sec', 'mean'), total_streams=('total_streams', 'sum'), avg_streams_per_song=('total_streams', 'mean'), avg_unique_listeners=('unique_listeners', 'mean'), avg_completion=('completion_pct', 'mean') ).round(1) # Add percentage of total streams per bucket bucket_analysis['pct_total_streams'] = ( bucket_analysis['total_streams'] * 100 / bucket_analysis['total_streams'].sum() ).round(1) print(bucket_analysis)📈 Expected Output:
| duration_bucket | songs | avg_dur | total_streams | avg_str/song | avg_listen | comp% | pct% |
|---|---|---|---|---|---|---|---|
| Short (<3min) | 85 | 165 | 32,50,000 | 38,235 | 12,850 | 96.2 | 36.9 Medium (3-4min) |
| 142 | 212 | 38,80,000 | 27,323 | 9,420 | 92.8 | 44.1 Long (4-5min) | 62 |
| 268 | 12,50,000 | 20,161 | 6,800 | 87.5 | 14.2 Very Long (>5min) | 38 | 358 |
4,20,000 11,053 3,850 76.3 4.8
Insight 1 — Short Songs Win: Short songs (<3 min) sirf 26% songs hain (85/327) BUT 36.9% streams generate karte hain! Avg streams per song sabse highest (38,235). TikTok effect — short-form content trending hai.
Insight 2 — Sweet Spot 3-4 minutes: Medium bucket sabse zyada songs (142) aur 44.1% streams. This is the traditional pop song length — universally works. Artists ko advise karo — 3-4 min ka sweet spot hai.
Insight 3 — Long Songs Struggle: Very Long songs (>5 min) sirf 4.8% streams generate karte hain aur 76.3% completion — 24% listeners skip. Long songs sirf established artists (Taylor Swift, etc.) hi afford kar sakte hain. New artists ko avoid karna chahiye.
Q8: Year-over-Year Growth by Genre
❓ Question: Calculate YoY growth for each genre for last 2 years (2024 vs 2025 vs 2026). Show streams per year, YoY growth %, and identify fastest growing and declining genres.
💡 Approach:
1. Streams + songs + artists merge karo (need genre)
2. Year extract karo from stream_date
3. pivot_table use karo — rows=genre, columns=year, values=streams
4. YoY growth calculate karo — pct_change() function
5. Sort by 2025→2026 growth
# Step 1: Merge streams + songs + artists df_yoy = streams.merge(songs[['song_id', 'artist_id']], on='song_id') df_yoy = df_yoy.merge(artists[['artist_id', 'genre']], on='artist_id') # Step 2: Extract year df_yoy['year'] = df_yoy['stream_date'].dt.year # Step 3: Create pivot table yearly_streams = df_yoy.pivot_table( index='genre', columns='year', values='stream_id', aggfunc='count', fill_value=0 ) # Step 4: Calculate YoY growth yearly_streams['growth_2024_2025'] = ( (yearly_streams[2025] - yearly_streams[2024]) * 100 / yearly_streams[2024] ).round(1) yearly_streams['growth_2025_2026'] = ( (yearly_streams[2026] - yearly_streams[2025]) * 100 / yearly_streams[2025] ).round(1) # Step 5: Add trend indicator yearly_streams['trend_2026'] = yearly_streams['growth_2025_2026'].apply( lambda x: '📈 Growing' if x > 10 else ('➡️ Stable' if x > 0 else '📉 Declining') ) # Sort by 2026 growth yearly_streams = yearly_streams.sort_values( 'growth_2025_2026', ascending=False ) print(yearly_streams)📈 Expected Output:
| genre | 2024 | 2025 | 2026 | growth_24_25 | growth_25_26 | trend |
|---|---|---|---|---|---|---|
| Punjabi | 2,80,000 | 5,20,000 | 8,50,000 | +85.7% | +63.5% | 📈 Growing K-Pop |
| 4,50,000 | 7,80,000 | 11,00,000 | +73.3% | +41.0% | 📈 Growing Bollywood | 8,50,000 |
| 13,50,000 | 18,60,000 | +58.8% | +37.8% | 📈 Growing Pop | 18,00,000 | 23,50,000 |
| 28,50,000 | +30.6% | +21.3% | 📈 Growing Hip-Hop | 10,50,000 | 12,80,000 | 14,20,000 |
| +21.9% | +10.9% | 📈 Growing Latin | 3,20,000 | 3,85,000 | 4,20,000 | +20.3% |
| +9.1% | ➡️ Stable Classical | 1,60,000 | 1,75,000 | 1,80,000 | +9.4% | +2.9% |
| ➡️ Stable Others | 1,80,000 | 1,55,000 | 1,40,000 | -13.9% | -9.7% | 📉 Declining |
Insight 1 — Punjabi Explosion: +85.7% and +63.5% consecutive years — fastest growing genre! AP Dhillon effect + Punjabi music going mainstream globally. IMMEDIATELY sign more Punjabi artists — window of opportunity.
Insight 2 — K-Pop & Bollywood Momentum: Both showing 40%+ growth. Global appetite for non-English music increasing. Expand playlists like "Global Hits", "Non-English Charts".
Insight 3 — Classical Warning: Only 2.9% growth in 2026 — losing momentum. Focus on younger audience engagement — "Study Music", "Meditation" playlists that repackage classical for modern use.
Insight 4 — "Others" Declining: -9.7% — niche genres losing listeners. Consolidate — merge into main genres or discontinue small subgenres.
Q9: Premium vs Free Tier Revenue Analysis
❓ Question: Compare Premium vs Free tier — user counts, total revenue (subscription + ad), avg revenue per user (ARPU), avg streams per user, and identify conversion opportunities.
💡 Approach:
1. Revenue table + listeners merge
2. Streams count per listener calculate karo
3. Merge sab together
4. GroupBy subscription_type → aggregate
5. ARPU = total_revenue / user_count
# Step 1: Aggregate revenue per listener listener_revenue = revenue.groupby('listener_id').agg( total_sub_revenue=('subscription_revenue', 'sum'), total_ad_revenue=('ad_revenue', 'sum'), total_revenue=('total_revenue', 'sum'), active_months=('month', 'nunique') ).reset_index() # Step 2: Aggregate streams per listener listener_streams = streams.groupby('listener_id').agg( total_streams=('stream_id', 'count'), total_listen_time=('stream_duration_sec', 'sum') ).reset_index() # Step 3: Merge everything with listeners table full_data = listeners.merge(listener_revenue, on='listener_id', how='left') full_data = full_data.merge(listener_streams, on='listener_id', how='left') # Fill NaN with 0 (users with no data) full_data = full_data.fillna(0) # Step 4: Aggregate by subscription type tier_analysis = full_data.groupby('subscription_type').agg( user_count=('listener_id', 'count'), total_subscription_rev=('total_sub_revenue', 'sum'), total_ad_rev=('total_ad_revenue', 'sum'), total_revenue=('total_revenue', 'sum'), total_streams=('total_streams', 'sum'), avg_streams_per_user=('total_streams', 'mean'), avg_listen_hours=('total_listen_time', 'mean') ).round(1) # Step 5: Calculate ARPU tier_analysis['ARPU'] = ( tier_analysis['total_revenue'] / tier_analysis['user_count'] ).round(1) # Convert listen time to hours tier_analysis['avg_listen_hours'] = ( tier_analysis['avg_listen_hours'] / 3600 ).round(1) # Add revenue % of total tier_analysis['revenue_pct'] = ( tier_analysis['total_revenue'] * 100 / tier_analysis['total_revenue'].sum() ).round(1) print(tier_analysis)📈 Expected Output:
| subscription | users | sub_rev | ad_rev | total | streams | str/user | hours | ARPU | rev% |
|---|---|---|---|---|---|---|---|---|---|
| Premium | 218 | 25,94,200 | 0 | 25,94,200 | 12,50,000 | 5,733 | 42.5 | ₹11,900 | 82.5% Free |
333 0 5,50,800 5,50,800 6,90,000 2,072 15.2 ₹1,655 17.5%
📊 Conversion Opportunity Analysis:
# Identify Free users who behave like Premium (high engagement) # These are conversion opportunities!
free_high_engagement = full_data[
(full_data['subscription_type'] == 'Free') &
(full_data['total_streams'] >= 100)
].sort_values('total_streams', ascending=False)
print(f"Free users with 100+ streams (Conversion Targets): {len(free_high_engagement)}")
print(f"Potential monthly revenue if converted: ₹{len(free_high_engagement) * 119:,}")
# Show top conversion targets
print(free_high_engagement[[
'listener_name', 'country', 'total_streams',
'total_listen_time'
]].head(10))
33,200 Top Conversion Targets:
| listener_name | country | total_streams | total_listen_hours |
|---|---|---|---|
| Jake Park | S.Korea | 168 | 7.7 Alex Chen |
| USA | 78 | 3.4 Mike Johnson | USA |
65 2.6
Insight 1 — Premium Dominance: Premium users 39.6% (218/551) BUT contribute 82.5% revenue! Free users 60.4% but only 17.5% revenue. ARPU ratio: Premium ₹11,900 vs Free ₹1,655 — 7.2x more revenue per Premium user!
Insight 2 — Ad Revenue Is Weak: Free users generate ₹5.5L only from ads across 333 users. Ad density low hai — increase ad frequency or ad rates. Alternative: aggressive Premium conversion.
Insight 3 — Golden Conversion Opportunity: 28 Free users have 100+ streams — they LOVE the platform but haven't paid. If all convert to Premium = ₹3.33L extra monthly revenue = ₹40L annually! Send them personalized Premium trial offers, remove ads from their favorite playlists for a week.
Insight 4 — Engagement Gap: Premium users listen 42.5 hours/month, Free users only 15.2 hours. Premium users are 2.8x more engaged — showing that "no ads + offline listening" drives more usage. Marketing message: "Listen 3x more without interruptions."
Q10: Inactive Listeners — Churn Identification
❓ Question: Identify users at churn risk — those who haven't streamed in last 30 days. Segment them by churn severity (Warning, Risk, Churned), show their historical value, and prioritize re-engagement.
💡 Approach:
1. Har listener ka last stream date find karo
2. Days since last stream calculate karo
3. Churn severity segments banao based on inactivity days
4. Historical streams + subscription type merge karo
5. Priority scoring — high value + medium risk = highest priority for re-engagement
# Step 1: Find last stream date per listener current_date = pd.Timestamp('2026-02-01') # Analysis date last_activity = streams.groupby('listener_id').agg( last_stream_date=('stream_date', 'max'), total_lifetime_streams=('stream_id', 'count'), total_lifetime_hours=('stream_duration_sec', 'sum') ).reset_index() # Step 2: Calculate days since last stream last_activity['days_inactive'] = ( current_date - last_activity['last_stream_date'] ).dt.days # Convert hours last_activity['total_lifetime_hours'] = ( last_activity['total_lifetime_hours'] / 3600 ).round(1) # Step 3: Merge with listeners for context churn_analysis = last_activity.merge( listeners[['listener_id', 'listener_name', 'country', 'subscription_type', 'signup_date']], on='listener_id' ) # Step 4: Segment by churn severity churn_analysis['churn_status'] = pd.cut( churn_analysis['days_inactive'], bins=[-1, 7, 30, 60, 90, np.inf], labels=['Active', 'Warning', 'At Risk', 'Likely Churned', 'Churned'] ) # Step 5: Priority score (high value + medium inactivity = highest priority) churn_analysis['priority_score'] = ( churn_analysis['total_lifetime_streams'] * (1 / (1 + churn_analysis['days_inactive'] / 30)) ).round(1) # Filter only at-risk users (not Active) at_risk = churn_analysis[ churn_analysis['churn_status'] != 'Active' ].sort_values('priority_score', ascending=False) # Overall segment summary segment_summary = churn_analysis.groupby( 'churn_status', observed=True ).agg( user_count=('listener_id', 'count'), avg_lifetime_streams=('total_lifetime_streams', 'mean'), premium_users=('subscription_type', lambda x: (x == 'Premium').sum()) ).round(1) print("=== Churn Segment Summary ===") print(segment_summary) print("\n=== Top Priority Re-engagement Targets ===") print(at_risk[['listener_name', 'country', 'subscription_type', 'days_inactive', 'total_lifetime_streams', 'churn_status', 'priority_score']].head(10))📈 Expected Output — Segment Summary:
| churn_status | users | avg_lifetime_streams | premium_users |
|---|---|---|---|
| Active | 342 | 145.2 | 168 |
| Warning | 88 | 82.5 | 32 At Risk |
| 65 | 55.3 | 15 Likely Churned | 38 |
| 32.8 | 3 Churned | 18 | 18.5 |
0
🎯 Top Priority Re-engagement Targets:
| listener_name | country | subscription | days_inactive | lifetime_streams | status | priority |
|---|---|---|---|---|---|---|
| Amit Verma | India | Premium | 12 | 285 | Warning | 204.6 Sneha Iyer |
| India | Premium | 15 | 245 | Warning | 166.7 David Kim | S.Korea |
| Premium | 18 | 218 | Warning | 141.6 Ravi Kumar | India | Premium |
| 25 | 195 | Warning | 107.4 Emma Watson | UK | Premium | 35 |
| 178 | At Risk | 82.2 John Smith | USA | Free | 22 | 165 |
| Warning | 95.7 Aisha Khan | India | Free | 42 | 152 | At Risk |
62.7
Insight 1 — Churn Funnel: 342 Active → 88 Warning → 65 At Risk → 38 Likely Churned → 18 Churned. Progressive drop-off — early intervention CRITICAL at Warning stage.
Insight 2 — Premium Loss Alert: 32 Premium users in Warning (₹3,808 monthly), 15 in At Risk (₹1,785). If all churn = ₹5,593 monthly loss = ₹67K annually! Save even 50% = ₹33K annually.
Insight 3 — Highest Priority: Amit Verma (Premium, 285 lifetime streams, only 12 days inactive) — highest priority score 204.6. Highly engaged historically, recent drop. Immediate personalized outreach — "We noticed you're away, here's your favorite playlist waiting!"
Insight 4 — Segment-Specific Strategy:
• Warning (88 users): Push notification, personalized playlist
• At Risk (65 users): Email + special discount offer
• Likely Churned (38 users): Last-chance offer, 3 months at 50% off
• Churned (18 users): Win-back campaign, "We miss you" email
🎯 Final Business Recommendations to Spotify
Based on Part A + Part B complete analysis, yeh 8 actionable recommendations top management ko present karo:
1. 🎵 Punjabi Content Investment (URGENT): +85.7% and +63.5% growth consecutive years. Sign 5-8 new Punjabi artists in next quarter. Launch "Punjabi Unplugged" exclusive playlist series. Market to Punjabi diaspora globally.
2. 💰 Premium Conversion Campaign: 28 Free users have 100+ streams = ₹40L annual revenue opportunity. Personalized 30-day Premium trial with their favorite playlists ad-free. Expected 60-70% conversion.
3. 🚨 Churn Prevention Program: 88 users in Warning + 65 At Risk = ₹5,593 monthly Premium loss risk. Launch automated re-engagement — Day 7: push notification, Day 14: email with personalized playlist, Day 21: 30% off offer. Target: 40% recovery.
4. ⚡ Power User VIP Program: 45 Power Users = 39% of platform streams. Launch "Spotify Elite" — early access to new releases, exclusive artist AMAs, no ads even on Free tier features. Retention priority.
5. 🎤 Diversify K-Pop & Punjabi: BTS = 80% K-Pop, AP Dhillon = 72.9% Punjabi. Concentration risk. Sign 3-5 emerging artists in each genre. Reduce dependency on single artists.
6. 📏 Song Length Guidance: Short songs (<3 min) generate 38K avg streams vs Very Long (>5 min) 11K. Advise emerging artists — release 3-4 minute versions of long songs as "Radio Edit". Boost stream counts.
7. 🌏 November Content Push: November showed -2.54% (only declining month). Plan major releases in early November — Bollywood winter romance albums, "Wrapped Preview" campaigns, holiday playlists earlier.
8. 🎼 Classical Repackaging: Classical growth only 2.9%. Repackage as "Focus", "Meditation", "Study" playlists targeting Gen Z. Modernize discovery without changing content.
🎓 Key Learnings from This Case Study
- Pandas Skills: merge(), groupby() with agg(), pivot_table(), pd.cut() for segmentation, datetime handling, apply() for custom logic
- Business Segmentation: User segmentation (Power/Regular/Casual/Inactive) drives targeted marketing — one-size-fits-all doesn't work
- Churn Analysis: Days since last activity is the strongest churn indicator. Progressive intervention (Warning → At Risk → Churned) more effective than late reactions
- Duration Analysis: Understanding user behavior patterns (short vs long songs) informs content strategy — not just what to promote, but what to CREATE
- YoY Analysis: pct_change() and pivot_table() combination for time-series trend analysis — identifies rising and declining segments
- ARPU Calculation: Revenue-per-user metrics reveal true value of customer segments — Premium's 7.2x higher ARPU justifies aggressive conversion campaigns
- Priority Scoring: Combining multiple factors (lifetime value × recency) into single priority score helps focus limited resources on highest-impact opportunities
- SQL vs Python: Both work, but Python better for advanced statistics, custom logic, and integration with visualization/ML pipelines
Next: Case Study #3 — Amazon E-Commerce Analysis 🛒
Agle case study mein hum Amazon/Flipkart-style e-commerce data analyze karenge — customer segmentation, product performance, RFM analysis, seller ratings, return rates. SQL solutions ke saath. Data Insights par SQL Playground mein Amazon / Flipkart Style dataset available hai — queries directly run kar sakte ho! Spotify Case Study Part A (SQL) + Part B (Python) dono blog par available hain.
Happy Analyzing! Master Both SQL & Python! 🚀
💬 Comments (0)
Loading comments...