<DataInsights />
  • 🏠 Home
  • 📊 SQL
  • 🐍 Python
  • 📈 Power BI
  • 📗 Excel
  • 💼 Career
  • 🎯 Interview Q&A
  • 📁 Case Study
  • 📥 Downloads
  • 🚀 My Portfolio
<DataInsights />

Practical Data Analytics tutorials covering SQL, Python, Power BI, Excel and career guidance for aspiring analysts — 100% free.

Topics

  • SQL Tutorials
  • Python Guide
  • Power BI
  • Excel Tips
  • Career Guide

Quick Links

  • 🛠️ All Tools
  • 🗓️ Archive
  • 📬 Contact
  • 🔍 Search
  • Portfolio
  • Kaggle
  • GitHub

Legal & Info

  • About
  • Contact
  • Privacy Policy
  • Disclaimer
  • Terms & Conditions
  • DMCA
  • Sitemap
Copyright © 2026 Data Insights by Jatin Kumar. All Rights Reserved.Built with ❤️ for Data Analysts
Home/Python/Matrix Operations, Sorting, Stacking And NumPy + P...

Matrix Operations, Sorting, Stacking And NumPy + Pandas Integration

A
August 3, 2026 Jatin Kumar 24 min read Python
Data Insights NumPy Masterclass — Part 3

Matrix Operations, Sorting, Stacking & NumPy + Pandas Integration

NumPy ka final aur most powerful part — Matrix multiplication, array sorting, stacking/splitting, aur NumPy ko Pandas ke saath use karna. Real-world data pipeline examples ke saath Data Insights par.

📑 Is Part 3 Mein Aap Kya Sikhenge:

  • Topic 1: np.dot() & Matrix Multiplication
  • Topic 2: Sorting — np.sort(), np.argsort()
  • Topic 3: Stacking — np.vstack(), np.hstack(), np.concatenate()
  • Topic 4: Splitting — np.split(), np.hsplit(), np.vsplit()
  • Topic 5: np.unique(), np.intersect1d(), np.union1d()
  • Topic 6: NumPy + Pandas Integration — Real Data Pipeline

1. np.dot() & Matrix Multiplication

text

🔍 Definition: np.dot(A, B) performs dot product — for 1D arrays it returns scalar product, for 2D arrays it performs matrix multiplication (rows of A × columns of B). The @ operator is a cleaner alternative for matrix multiplication introduced in Python 3.5.

🎯 Samjho Simple Bhasha Mein: np.dot() aur * mein fark hai. * element-wise multiply karta hai — corresponding positions pe. np.dot() matrix multiplication karta hai — row × column style. Jaise ek class ke students ke marks hain aur har subject ka weightage hai — dot product se final weighted score nikalte hain. Machine learning mein weights aur features ka dot product prediction banata hai.

💡 Matrix Multiplication Rule:
A (m×n) dot B (n×p) = Result (m×p)
Inner dimensions must match: A's cols == B's rows

3 Ways to Matrix Multiply:
np.dot(A, B) → Classic way
A @ B → Modern Python 3.5+ operator
np.matmul(A, B) → Explicit matrix multiply

💻 Real-World Code Examples:

Example 1: 1D dot product aur 2D matrix multiplication.

import numpy as np
# 1D Dot Product (scalar result)
a = np.array([1, 2, 3])
b = np.array([4, 5, 6])
dot_1d = np.dot(a, b) # 14 + 25 + 3*6 = 32
print("1D dot product:", dot_1d)

# 2D Matrix Multiplication
A = np.array([[1, 2],
[3, 4]]) # Shape: (2,2)

B = np.array([[5, 6],
[7, 8]]) # Shape: (2,2)

# All three ways — same result
result1 = np.dot(A, B)
result2 = A @ B
result3 = np.matmul(A, B)

print("\nA @ B:\n", result1)
print("Same result?", np.array_equal(result1, result2))

# Element-wise vs Matrix multiply
print("\nElement-wise (*):\n", A * B)
print("Matrix multiply (@):\n", A @ B)

Example 2: Real-world — Weighted exam scores.

# 4 students, 3 subjects (Math, Science, English)
scores = np.array([[85, 90, 78],
[92, 88, 95],
[70, 75, 80],
[88, 92, 85]]) # Shape: (4,3)

# Weights: Math=40%, Science=35%, English=25%
weights = np.array([0.40, 0.35, 0.25]) # Shape: (3,)

# Weighted score for each student
# (4,3) dot (3,) = (4,) — each student's final score
final_scores = scores @ weights
print("Weighted final scores:", np.round(final_scores, 2))

# Rank students
ranks = np.argsort(-final_scores) + 1 # Descending
names = ['Rahul', 'Priya', 'Amit', 'Sneha']
for i, name in enumerate(names):
print(f"{name}: {final_scores[i]:.2f} (Rank {ranks[i]})")

📊 Expected Output:

1D dot product: 32
A @ B:
[[19 22]
[43 50]]
Same result? True

Element-wise (*):
[[ 5 12]
[21 32]]
Matrix multiply (@):
[[19 22]
[43 50]]

Weighted final scores: [85.45 91.45 74.25 89.05]
Rahul: 85.45 (Rank 4)
Priya: 91.45 (Rank 1)
Amit: 74.25 (Rank 4)
Sneha: 89.05 (Rank 2)

text

⚠️ Common Mistakes:

  • Mistake: Incompatible shapes ke saath dot product → A(2,3) @ B(2,3) → Error! Inner dimensions must match.
    Fix: A(2,3) @ B(3,2) → Result (2,2). Check: A cols == B rows.
  • Mistake: * aur @ confuse karna → * element-wise, @ matrix multiply — bilkul alag results!
    Fix: Machine learning mein predictions ke liye → @. Feature scaling ke liye → *.

💬 Interview Questions:

Q1: What is the difference between np.dot(), @ operator and np.multiply()?
Ans: np.dot() and @ (matmul) perform matrix multiplication — for 2D: rows of A times columns of B. np.multiply() and * perform element-wise multiplication — corresponding elements multiplied. For 1D arrays, np.dot() returns scalar dot product. @ is preferred modern syntax. np.multiply() is explicit element-wise. Key rule: @ requires inner dimensions to match (A cols == B rows).

Q2: In what shape must two matrices be for matrix multiplication?
Ans: For A @ B: A must be shape (m, n) and B must be shape (n, p). The inner dimensions must match — A's number of columns must equal B's number of rows. The result is shape (m, p). Example: (3,4) @ (4,5) = (3,5). If shapes are (3,4) @ (3,5) — error because 4 ≠ 3.

2. Sorting — np.sort() & np.argsort()

text

🔍 Definition: np.sort(arr) returns a sorted copy of the array. arr.sort() sorts in-place. np.argsort(arr) returns the indices that would sort the array — extremely useful for ranking and reordering other arrays based on sorted order.

🎯 Samjho Simple Bhasha Mein: np.sort() simply sorted array de deta hai. np.argsort() batata hai "kaunsa element kis position pe aata agar sort karte" — indices deta hai values nahi. Yeh bahut powerful hai — ek array ke sort order se dusre array bhi reorder kar sakte ho. Jaise salary ke hisaab se sort karke employee names bhi usi order mein dikh sakti hain.

💡 Sort Key Points:
np.sort(arr) → New sorted array (copy)
arr.sort() → In-place sort (modifies original)
np.sort(arr)[::-1] → Descending order
np.argsort(arr) → Indices for ascending
np.argsort(-arr) → Indices for descending
axis=0 → Column-wise sort (2D)
axis=1 → Row-wise sort (2D)

💻 Real-World Code Examples:

Example 1: Basic sorting operations.

import numpy as np
salaries = np.array([75000, 85000, 92000, 68000,
55000, 110000, 48000])

# Ascending sort (returns copy)
sorted_asc = np.sort(salaries)
print("Ascending:", sorted_asc)

# Descending sort
sorted_desc = np.sort(salaries)[::-1]
print("Descending:", sorted_desc)

# argsort — indices of sorted order
sort_idx = np.argsort(salaries)
print("Sort indices:", sort_idx)

# argsort descending
sort_idx_desc = np.argsort(-salaries)
print("Desc indices:", sort_idx_desc)

# Verify: salaries[sort_idx] == sorted
print("Verify:", salaries[sort_idx])

Example 2: Real-world — Rank employees by salary.

# Employee data
names = np.array(['Rahul', 'Priya', 'Amit',
'Sneha', 'Ravi', 'Anjali', 'Deepak'])
salaries = np.array([75000, 85000, 92000,
68000, 55000, 110000, 48000])

# Sort names by salary (descending)
desc_idx = np.argsort(-salaries)
sorted_names = names[desc_idx]
sorted_sal = salaries[desc_idx]

print("Salary Leaderboard:")
for rank, (name, sal) in enumerate(zip(sorted_names, sorted_sal), 1):
print(f" #{rank} {name}: ₹{sal:,}")

# 2D sort — sort each row
marks = np.array([[85, 70, 92],
[60, 95, 78],
[88, 55, 73]])
row_sorted = np.sort(marks, axis=1) # Sort each row
col_sorted = np.sort(marks, axis=0) # Sort each column
print("\nRow-sorted:\n", row_sorted)
print("Col-sorted:\n", col_sorted)

📊 Expected Output:

Ascending:  [ 48000  55000  68000  75000  85000  92000 110000]
Descending: [110000 92000 85000 75000 68000 55000 48000]
Sort indices: [6 4 3 0 1 2 5]
Desc indices: [5 2 1 0 3 4 6]
Verify: [ 48000 55000 68000 75000 85000 92000 110000]

Salary Leaderboard:
#1 Anjali: ₹1,10,000
#2 Amit: ₹92,000
#3 Priya: ₹85,000
#4 Rahul: ₹75,000
#5 Sneha: ₹68,000
#6 Ravi: ₹55,000
#7 Deepak: ₹48,000

Row-sorted:
[[70 85 92]
[60 78 95]
[55 73 88]]
Col-sorted:
[[60 55 73]
[85 70 78]
[88 95 92]]

⚠️ Common Mistakes:

  • Mistake: np.sort() aur arr.sort() ko same samajhna → np.sort() copy return karta hai. arr.sort() in-place karta hai aur None return karta hai.
    Fix: sorted_arr = np.sort(arr) ✅ — sorted_arr = arr.sort() ❌ (None milega)
  • Mistake: Descending sort ke liye arg ka use → np.sort(arr, reverse=True) — yeh parameter exist nahi karta NumPy mein!
    Fix: np.sort(arr)[::-1] ya np.argsort(-arr) use karo.

💬 Interview Questions:

Q1: What is the difference between np.sort() and np.argsort()?
Ans: np.sort(arr) returns the sorted values in a new array. np.argsort(arr) returns the indices that would sort the array — the actual values are not moved, only their positions are returned. argsort is powerful for multi-array sorting: get indices from one array and use them to reorder other arrays. Example: sort names by salary without losing name-salary pairing.

Q2: How to sort a 2D array row-wise vs column-wise?
Ans: Use axis parameter: np.sort(arr, axis=1) sorts each row independently (left to right within each row). np.sort(arr, axis=0) sorts each column independently (top to bottom within each column). Default axis=-1 sorts along the last axis. Note: these sort within rows/columns, not rearranging complete rows/columns.

3. Stacking — vstack, hstack, concatenate

text

🔍 Definition: Stacking combines multiple arrays into one. np.vstack() stacks vertically (row-wise — adds more rows). np.hstack() stacks horizontally (column-wise — adds more columns). np.concatenate() is the generalized version that works along any axis.

🎯 Samjho Simple Bhasha Mein: Stacking matlab arrays ko jodhna. vstack matlab "upar-neeche stack karo" — jaise do tables ke rows milana. hstack matlab "side-by-side rakhna" — jaise do tables ke columns milana. Data science mein yeh bahut use hota hai — train aur test data ko combine karna, features aur labels ko saath stack karna.

💡 Stacking Functions:
np.vstack([a, b]) → Vertical stack (row-wise), axis=0
np.hstack([a, b]) → Horizontal stack (col-wise), axis=1
np.concatenate([a,b], axis=0) → Same as vstack
np.concatenate([a,b], axis=1) → Same as hstack
np.stack([a,b], axis=0) → Creates NEW dimension
np.column_stack([a,b]) → 1D arrays as columns

💻 Real-World Code Examples:

Example 1: vstack, hstack aur concatenate.

import numpy as np
# Two batches of student marks
batch1 = np.array([[85, 90, 78],
[92, 88, 95]])

batch2 = np.array([[70, 75, 80],
[88, 92, 85]])

# vstack — add more students (rows)
all_students = np.vstack([batch1, batch2])
print("vstack (all students):\n", all_students)
print("Shape:", all_students.shape) # (4, 3)

# hstack — add more subjects (columns)
extra_subject = np.array([[88], [91]])
batch1_extended = np.hstack([batch1, extra_subject])
print("\nhstack (extra subject):\n", batch1_extended)
print("Shape:", batch1_extended.shape) # (2, 4)

# concatenate — generalized version
concat_v = np.concatenate([batch1, batch2], axis=0) # same as vstack
concat_h = np.concatenate([batch1, batch2], axis=1) # same as hstack
print("\nconcat axis=1:\n", concat_h)
print("Shape:", concat_h.shape) # (2, 6)

Example 2: column_stack — 1D arrays ko columns mein stack karo.

# Build feature matrix from separate arrays
age = np.array([25, 32, 28, 35, 29])
salary = np.array([50000, 75000, 60000, 90000, 55000])
exp_yrs = np.array([2, 7, 4, 10, 3])

# Column stack — each 1D array becomes a column
features = np.column_stack([age, salary, exp_yrs])
print("Feature Matrix:\n", features)
print("Shape:", features.shape) # (5, 3) — 5 samples, 3 features

# np.stack — creates NEW axis
a = np.array([1, 2, 3])
b = np.array([4, 5, 6])
stacked_0 = np.stack([a, b], axis=0) # (2,3)
stacked_1 = np.stack([a, b], axis=1) # (3,2)
print("\nstack axis=0:\n", stacked_0)
print("stack axis=1:\n", stacked_1)

📊 Expected Output:

vstack (all students):
[[85 90 78]
[92 88 95]
[70 75 80]
[88 92 85]]
Shape: (4, 3)

hstack (extra subject):
[[85 90 78 88]
[92 88 95 91]]
Shape: (2, 4)

concat axis=1:
[[85 90 78 70 75 80]
[92 88 95 88 92 85]]
Shape: (2, 6)

Feature Matrix:
[[ 25 50000 2]
[ 32 75000 7]
[ 28 60000 4]
[ 35 90000 10]
[ 29 55000 3]]
Shape: (5, 3)

stack axis=0: [[1 2 3] [4 5 6]]
stack axis=1: [[1 4] [2 5] [3 6]]

text

⚠️ Common Mistakes:

  • Mistake: vstack mein different number of columns → Error! Cols same honi chahiye for vstack.
    Fix: vstack → same columns. hstack → same rows. Check shapes pehle.
  • Mistake: np.stack() aur np.concatenate() confuse karna → stack NEW dimension banata hai. concatenate existing dimension mein jodhta hai.
    Fix: [1,2,3] aur [4,5,6] → stack → (2,3). concatenate → [1,2,3,4,5,6] (6,).

💬 Interview Questions:

Q1: What is the difference between np.vstack() and np.hstack()?
Ans: vstack (vertical stack) stacks arrays along axis=0 — adds more rows, requires same number of columns. hstack (horizontal stack) stacks along axis=1 — adds more columns, requires same number of rows. Think of vstack as placing arrays on top of each other, hstack as placing them side by side.

Q2: When would you use np.stack() vs np.concatenate()?
Ans: np.concatenate() joins arrays along an existing axis — no new dimension created. np.stack() always creates a NEW axis — input arrays must have the same shape. Use concatenate to extend existing data (add more rows/columns). Use stack to create a new batch dimension — like stacking image arrays to create a batch for deep learning: each image (H,W,C) stacked → (N,H,W,C) batch.

4. Splitting — np.split(), hsplit(), vsplit()

text

🔍 Definition: Splitting divides one array into multiple sub-arrays. np.split(arr, n) splits into n equal parts. np.vsplit() splits vertically (along rows). np.hsplit() splits horizontally (along columns). You can also specify exact indices where splits should occur.

🎯 Samjho Simple Bhasha Mein: Splitting stacking ka ulta hai — ek bade array ko chhote pieces mein todna. Machine learning mein train/validation/test split karna hota hai — np.split se hota hai. Batches mein data process karna ho toh bhi split use hota hai. Data ko different processing paths pe bhejna ho toh bhi split kaam aata hai.

💡 Split Functions:
np.split(arr, 3) → 3 equal parts
np.split(arr, [2, 5]) → Split at index 2 and 5
np.vsplit(arr, 2) → 2 equal vertical splits
np.hsplit(arr, 2) → 2 equal horizontal splits
np.array_split(arr, 3) → Allows unequal splits

💻 Real-World Code Examples:

Example 1: Basic splitting operations.

import numpy as np
arr = np.arange(12)
print("Original:", arr)

# Split into 3 equal parts
parts = np.split(arr, 3)
print("3 equal parts:", parts)

# Split at specific indices
splits = np.split(arr, [4, 8]) # Split at index 4 and 8
print("Split at [4,8]:", splits)

# Unequal splits (array_split — no error)
unequal = np.array_split(arr, 5)
print("5 unequal parts:")
for i, p in enumerate(unequal):
print(f" Part {i+1}: {p}")

Example 2: Train-Validation-Test split using NumPy.

import numpy as np
np.random.seed(42)

# Dataset: 1000 samples
X = np.random.randn(100, 5) # 100 samples, 5 features
y = np.random.randint(0, 2, 100) # Binary labels

# Shuffle indices
indices = np.random.permutation(100)
X_shuffled = X[indices]
y_shuffled = y[indices]

# 70% train, 15% val, 15% test
train_end = int(0.70 * 100) # 70
val_end = int(0.85 * 100) # 85

X_train, X_val, X_test = np.split(X_shuffled, [train_end, val_end])
y_train, y_val, y_test = np.split(y_shuffled, [train_end, val_end])

print(f"Train: {X_train.shape} | Val: {X_val.shape} | Test: {X_test.shape}")
print(f"Train labels: {y_train.shape}")

# 2D vsplit and hsplit
matrix = np.arange(24).reshape(4, 6)
top, bottom = np.vsplit(matrix, 2)
left, right = np.hsplit(matrix, 2)
print("\nTop half:\n", top)
print("Left half:\n", left)

📊 Expected Output:

Original: [ 0 1 2 3 4 5 6 7 8 9 10 11]
3 equal parts: [array([0,1,2,3]), array([4,5,6,7]), array([8,9,10,11])]
Split at [4,8]: [array([0,1,2,3]), array([4,5,6,7]), array([8,9,10,11])]

5 unequal parts:
Part 1: [0 1 2]
Part 2: [3 4 5]
Part 3: [6 7]
Part 4: [8 9]
Part 5: [10 11]

Train: (70, 5) | Val: (15, 5) | Test: (15, 5)
Train labels: (70,)

Top half:
[[ 0 1 2 3 4 5]
[ 6 7 8 9 10 11]]
Left half:
[[ 0 1 2]
[ 6 7 8]
[12 13 14]
[18 19 20]]

text

⚠️ Common Mistakes:

  • Mistake: np.split mein unequal division → np.split(arr, 5) for 12 elements → ValueError: array split does not result in equal division.
    Fix: np.array_split() use karo — unequal splits allow karta hai.
  • Mistake: Split result unpack bhool jaana → parts = np.split(arr, 3) list milti hai — a, b, c = np.split(arr, 3) directly unpack karo.

💬 Interview Questions:

Q1: What is the difference between np.split() and np.array_split()?
Ans: np.split(arr, n) requires equal division — if array size is not perfectly divisible by n, it raises ValueError. np.array_split(arr, n) allows unequal splits — some sub-arrays may have one more element than others. array_split is safer for production code where array size may not be perfectly divisible.

Q2: How do you perform train-test split using NumPy?
Ans: Shuffle indices with np.random.permutation(n), apply to X and y arrays. Then use np.split() with calculated indices: train_end = int(0.8 * n); X_train, X_test = np.split(X_shuffled, [train_end]). This is the manual equivalent of sklearn's train_test_split — useful when you need more control over the splitting process.

5. np.unique(), np.intersect1d(), np.union1d()

text

🔍 Definition: NumPy provides set operations for arrays. np.unique() returns unique sorted values. np.intersect1d() returns common elements. np.union1d() returns all unique elements from both. np.setdiff1d() returns elements in first but not second.

🎯 Samjho Simple Bhasha Mein: Yeh functions arrays ko sets ki tarah treat karte hain. np.unique() duplicates hata ke unique values deta hai. np.intersect1d() do arrays mein common values dhundhta hai. np.union1d() dono arrays ki saari unique values deta hai. Data analysis mein common customers dhundhna, unique categories nikalna — yeh sab set operations se hota hai.

💡 Set Operations Reference:
np.unique(arr) → Unique values (sorted)
np.unique(arr, return_counts=True) → With counts
np.unique(arr, return_index=True) → With first indices
np.intersect1d(a, b) → A ∩ B (common)
np.union1d(a, b) → A ∪ B (all unique)
np.setdiff1d(a, b) → A - B (in A not B)
np.in1d(a, b) → Boolean: is each a in b?

💻 Real-World Code Examples:

Example 1: np.unique() with extra information.

import numpy as np
depts = np.array(['IT', 'HR', 'IT', 'Sales', 'HR',
'IT', 'Finance', 'Sales', 'IT'])

# Basic unique
unique_depts = np.unique(depts)
print("Unique departments:", unique_depts)

# With counts (how many per dept)
unique_vals, counts = np.unique(depts, return_counts=True)
print("Department counts:")
for dept, cnt in zip(unique_vals, counts):
print(f" {dept}: {cnt} employees")

# With first occurrence index
unique_vals, first_idx = np.unique(depts, return_index=True)
print("First occurrence indices:", first_idx)

# Numeric unique
scores = np.array([85, 90, 85, 78, 90, 92, 78])
print("\nUnique scores:", np.unique(scores))

Example 2: Set operations — Customer analysis.

# January aur February customers
jan_customers = np.array([101, 102, 103, 104, 105])
feb_customers = np.array([103, 105, 106, 107, 108])

# Returning customers (both months)
returning = np.intersect1d(jan_customers, feb_customers)
print("Returning customers:", returning)

# All unique customers (either month)
all_cust = np.union1d(jan_customers, feb_customers)
print("All unique customers:", all_cust)

# Jan only customers (churned)
churned = np.setdiff1d(jan_customers, feb_customers)
print("Churned (Jan only):", churned)

# Feb only customers (new)
new_cust = np.setdiff1d(feb_customers, jan_customers)
print("New (Feb only):", new_cust)

# Check if each Jan customer returned
returned_mask = np.in1d(jan_customers, feb_customers)
print("Returned mask:", returned_mask)

📊 Expected Output:

Unique departments: ['Finance' 'HR' 'IT' 'Sales']
Department counts:
Finance: 1 employees
HR: 2 employees
IT: 4 employees
Sales: 2 employees
First occurrence indices: [6 1 0 3]

Unique scores: [78 85 90 92]

Returning customers: [103 105]
All unique customers: [101 102 103 104 105 106 107 108]
Churned (Jan only): [101 102 104]
New (Feb only): [106 107 108]
Returned mask: [False False True False True]

text

⚠️ Common Mistakes:

  • Mistake: np.unique() sorted output expect na karna → np.unique() hamesha sorted result deta hai.
    Fix: Original order chahiye toh return_index=True use karo aur indices se original order maintain karo.
  • Mistake: np.in1d() ko 2D arrays pe use karna → in1d 1D arrays ke liye hai.
    Fix: 2D ke liye np.isin() use karo — multi-dimensional support karta hai.

💬 Interview Questions:

Q1: How to find unique values and their frequencies in NumPy?
Ans: np.unique(arr, return_counts=True) returns two arrays: unique sorted values and their occurrence counts. Example: np.unique([1,2,2,3,3,3], return_counts=True) → values=[1,2,3], counts=[1,2,3]. This is equivalent to pandas value_counts() but for NumPy arrays.

Q2: Give a business use case for np.setdiff1d().
Ans: Customer churn analysis: compare last month's customers to this month's. np.setdiff1d(last_month, this_month) gives churned customers (were there before, not now). np.setdiff1d(this_month, last_month) gives new customers. np.intersect1d() gives retained customers. This pattern is used in e-commerce, subscription services, and marketing analytics.

6. NumPy + Pandas Integration — Real Data Pipeline

text

🔍 Definition: NumPy and Pandas work seamlessly together. Pandas DataFrames are built on top of NumPy arrays. You can convert between them, apply NumPy functions to Pandas columns, use NumPy arrays in Pandas operations, and build complete data processing pipelines using both libraries together.

🎯 Samjho Simple Bhasha Mein: NumPy aur Pandas ek team ki tarah kaam karte hain. Pandas ka kaam hai data organize karna (labeled rows aur columns), NumPy ka kaam hai fast mathematical operations. Jab bhi Pandas column pe math karna ho — internally NumPy chal raha hai. Dono ke beech easily data move ho sakta hai. Real data pipeline mein dono ka saath use karna normal hai.

💡 NumPy ↔ Pandas Conversion:
NumPy → Pandas: pd.DataFrame(np_array)
Pandas → NumPy: df.values or df.to_numpy()
Series → NumPy: series.values or series.to_numpy()
NumPy functions on Pandas: np.log(df['salary'])
Pandas column as NumPy: np.array(df['col'])

💻 Real-World Code Examples:

Example 1: Conversion between NumPy and Pandas.

import numpy as np
import pandas as pd

# NumPy Array → Pandas DataFrame
np_data = np.array([[1, 'Rahul', 75000, 'IT'],
[2, 'Priya', 85000, 'HR'],
[3, 'Amit', 92000, 'IT'],
[4, 'Sneha', 68000, 'Sales'],
[5, 'Deepak', 110000, 'Finance']])

df = pd.DataFrame(np_data, columns=['id', 'name', 'salary', 'dept'])
df['salary'] = df['salary'].astype(float)
print("DataFrame:\n", df)

# Pandas DataFrame → NumPy Array
np_from_df = df['salary'].to_numpy()
print("\nSalary as NumPy:", np_from_df)
print("Type:", type(np_from_df))

# Full DataFrame to numpy
numeric_cols = df[['id', 'salary']].to_numpy()
print("\nNumeric columns as array:\n", numeric_cols)

Example 2: NumPy functions on Pandas columns.

# Apply NumPy functions to Pandas columns
df['log_salary'] = np.log(df['salary'])
df['annual_salary'] = df['salary'] * 12
df['salary_grade'] = np.where(
df['salary'] > 80000,
'Senior', 'Junior'
)
df['normalized_sal'] = np.round(
(df['salary'] - df['salary'].mean()) / df['salary'].std(), 2
)
print(df[['name', 'salary', 'salary_grade', 'normalized_sal']])

Example 3: Complete Data Pipeline — NumPy + Pandas together.

import numpy as np
import pandas as pd
np.random.seed(42)

# Step 1: Generate synthetic dataset with NumPy
n = 200
age = np.random.randint(22, 55, n)
exp = np.random.randint(0, 20, n)
base_sal= 30000 + (exp * 3000) + np.random.normal(0, 5000, n)
dept_idx= np.random.randint(0, 4, n)
depts = np.array(['IT', 'HR', 'Sales', 'Finance'])[dept_idx]

# Step 2: Create Pandas DataFrame
df = pd.DataFrame({
'age': age,
'experience': exp,
'salary': np.clip(base_sal, 25000, 150000).astype(int),
'dept': depts
})

# Step 3: NumPy analysis on Pandas data
sal_arr = df['salary'].to_numpy()
print(f"Dataset: {df.shape}")
print(f"Mean: ₹{np.mean(sal_arr):,.0f}")
print(f"Median: ₹{np.median(sal_arr):,.0f}")
print(f"Std: ₹{np.std(sal_arr):,.0f}")
print(f"P25: ₹{np.percentile(sal_arr, 25):,.0f}")
print(f"P75: ₹{np.percentile(sal_arr, 75):,.0f}")

# Step 4: Department-wise analysis
for dept in np.unique(depts):
mask = df['dept'] == dept
dept_sal = sal_arr[mask]
print(f"{dept:8}: n={len(dept_sal):3}, mean=₹{np.mean(dept_sal):,.0f}")

# Step 5: Add NumPy-computed features back to DataFrame
df['z_score'] = np.round((sal_arr - sal_arr.mean()) / sal_arr.std(), 2)
df['is_senior'] = np.where(sal_arr > np.percentile(sal_arr, 75), 1, 0)
print("\nFirst 5 rows with new features:")
print(df.head())

📊 Expected Output:

DataFrame:
id name salary dept
0 1 Rahul 75000.0 IT
1 2 Priya 85000.0 HR
2 3 Amit 92000.0 IT
3 4 Sneha 68000.0 Sales
4 5 Deepak 110000.0 Finance

name salary salary_grade normalized_sal
Rahul 75000 Junior -0.55
Priya 85000 Senior -0.05
Amit 92000 Senior 0.30
Sneha 68000 Junior -0.90
Deepak 110000 Senior 1.20

Dataset: (200, 4)
Mean: ₹64,832
Median: ₹63,178
Std: ₹22,847
P25: ₹46,503
P75: ₹81,879

Finance : n= 51, mean=₹65,410
HR : n= 48, mean=₹63,891
IT : n= 54, mean=₹65,743
Sales : n= 47, mean=₹64,159

⚠️ Common Mistakes:

  • Mistake: df.values vs df.to_numpy() confusion → df.values returns NumPy array but may return object dtype for mixed columns. df.to_numpy() is more explicit.
    Fix: Prefer df.to_numpy() — clearer and recommended in modern Pandas.
  • Mistake: NumPy function result ko Pandas Series samajhna → np.log(df['col']) returns Pandas Series (preserves index) because pandas integrates numpy ufuncs.
    Fix: Yeh actually useful hai — result directly column assign ho sakta hai.
  • Mistake: Mixed dtype DataFrame ko numpy array mein convert karna → String + int columns → object dtype array, math operations fail.
    Fix: Sirf numeric columns select karo: df.select_dtypes(include=np.number).to_numpy()

💬 Interview Questions:

Q1: What is the relationship between NumPy and Pandas?
Ans: Pandas is built on top of NumPy — a Pandas DataFrame is essentially a labeled wrapper around NumPy arrays. Each Pandas column is a NumPy array with a label. Pandas adds: named rows/columns (index), mixed dtypes, missing value handling (NaN), and data manipulation methods. NumPy provides the fast computation engine underneath. They work seamlessly together.

Q2: What is the difference between df.values and df.to_numpy()?
Ans: Both convert DataFrame to NumPy array. df.values is the older approach — it returns the underlying NumPy array directly. df.to_numpy() was introduced in Pandas 0.24 as the official method — it has a dtype parameter for explicit type control. The key difference: for extension arrays (like Pandas Categorical, Int64NA), to_numpy() handles conversion better. to_numpy() is the recommended modern approach.

Q3: How do you apply a NumPy function to a Pandas column?
Ans: Directly apply: np.sqrt(df['salary']) works because NumPy ufuncs understand Pandas Series. Result is a Pandas Series with same index. Alternative: df['salary'].apply(np.sqrt) but this is slower. For complex operations: convert to NumPy first, compute, then assign back. np.where() and np.select() work directly on Pandas Series too.

NumPy Masterclass — Complete Reference

Category Key Functions Purpose
Creation array, zeros, ones, arange, linspace, eye Make arrays
Shape reshape, flatten, ravel, T, transpose Change structure
Indexing [idx], [start:stop:step], bool_mask, fancy Access elements
Math Ops +,-,*,/,**,@,dot, broadcasting Arithmetic
Conditions where, any, all, &, |, ~ Boolean logic
Aggregates sum, mean, std, min, max, median, percentile Statistics
Math Funcs sqrt, log, exp, abs, round, clip, cumsum Math operations
Sort sort, argsort Order arrays
Combine vstack, hstack, concatenate, stack Merge arrays
Split split, array_split, vsplit, hsplit Divide arrays
Set Ops unique, intersect1d, union1d, setdiff1d, in1d Set operations
Random seed, rand, randn, randint, choice, shuffle Random data

🎉 NumPy Masterclass — Complete!

Teeno parts complete ho gaye!

✅ Part 1 — Arrays, Creation, Indexing, Slicing, Shape

✅ Part 2 — Arithmetic, Broadcasting, Comparison, where(), Aggregates, Math, Random

✅ Part 3 — Matrix Ops, Sorting, Stacking, Splitting, Set Ops, Pandas Integration

Aage Kya — Data Analyst Roadmap

NumPy complete! Ab Data Analyst roadmap mein next topics: Power BI, Excel Advanced, ya Python + MySQL Integration — jo chahiye batao Data Insights par!

Happy Coding & Keep Learning! 🚀

👤
Jatin Kumar
Data Analyst & Educator

Python, SQL, Power BI aur Excel mein practical tutorials likhta hoon — taaki data analytics seekhna aasan ho. Portfolio: jatinanalytics.co.in

Portfolio LinkedIn GitHub Kaggle All Articles
Share:

💬 Comments (0)

Spam/links allowed nahi hain — respectful comments welcome!

Loading comments...

Was this article helpful?
Previous Article"Array Operations — Arithmetic, Broadcasting, Comparison AndNext Article Python Basics — Complete Foundation Guide

📚 More Articles Like This

"NumPy Fundamentals — Arrays, Creation And Indexing"

Read Article

Advanced Charts, Graph Objects And Export in Plotly

Read Article

Interactive Charts with Plotly Express

Read Article