<DataInsights />
  • 🏠 Home
  • 📊 SQL
  • 🐍 Python
  • 📈 Power BI
  • 📗 Excel
  • 💼 Career
  • 🎯 Interview Q&A
  • 📁 Case Study
  • 📥 Downloads
  • 🚀 My Portfolio
<DataInsights />

Practical Data Analytics tutorials covering SQL, Python, Power BI, Excel and career guidance for aspiring analysts — 100% free.

Topics

  • SQL Tutorials
  • Python Guide
  • Power BI
  • Excel Tips
  • Career Guide

Quick Links

  • 🛠️ All Tools
  • 🗓️ Archive
  • 📬 Contact
  • 🔍 Search
  • Portfolio
  • Kaggle
  • GitHub

Legal & Info

  • About
  • Contact
  • Privacy Policy
  • Disclaimer
  • Terms & Conditions
  • DMCA
  • Sitemap
Copyright © 2026 Data Insights by Jatin Kumar. All Rights Reserved.Built with ❤️ for Data Analysts
Home/Python/"Array Operations — Arithmetic, Broadcasting, Comp...

"Array Operations — Arithmetic, Broadcasting, Comparison And Math Functions"

A
August 3, 2026 Jatin Kumar 26 min read Python
Data Insights NumPy Masterclass — Part 2

Array Operations — Arithmetic, Broadcasting, Comparison & Math Functions

NumPy ka asli power — vectorized operations. Arithmetic se lekar broadcasting tak, np.where() se lekar mathematical functions tak — sab kuch real-world examples aur interview questions ke saath Data Insights par.

📑 Is Part 2 Mein Aap Kya Sikhenge:

  • Topic 1: Arithmetic Operations — +, -, *, /, //, %, **
  • Topic 2: Broadcasting — Different Shape Arrays
  • Topic 3: Comparison Operators — Boolean Arrays
  • Topic 4: np.where() — Conditional Selection
  • Topic 5: Aggregate Functions — sum, mean, std, min, max
  • Topic 6: Math Functions — sqrt, log, exp, round, abs
  • Topic 7: np.random — Random Number Generation

1. Arithmetic Operations — Vectorized Math

text

🔍 Definition: NumPy arithmetic operations are element-wise — they apply the operation to each corresponding pair of elements simultaneously without any loop. This is called vectorization. All standard operators (+, -, *, /, //, %, **) work element-wise on arrays of same shape.

🎯 Samjho Simple Bhasha Mein: Socho 5 employees hain aur sabko 10% salary hike deni hai. Python list mein loop lagana padta. NumPy mein sirf salaries * 1.10 — ek line mein saari salaries update! Yahi vectorization hai — har element par automatically operation apply hota hai bina explicit loop ke.

💡 Element-wise Operations:
arr1 + arr2 → Addition
arr1 - arr2 → Subtraction
arr1 * arr2 → Multiplication
arr1 / arr2 → Division (float result)
arr1 // arr2 → Floor Division (int result)
arr1 % arr2 → Modulo (remainder)
arr1 ** arr2 → Power/Exponent

💻 Real-World Code Examples:

Example 1: Basic arithmetic operations.

import numpy as np
a = np.array([10, 20, 30, 40, 50])
b = np.array([2, 4, 6, 8, 10])

print("a + b =", a + b)
print("a - b =", a - b)
print("a * b =", a * b)
print("a / b =", a / b)
print("a // b =", a // b)
print("a % b =", a % b)
print("a ** 2 =", a ** 2)

# Scalar operations (scalar broadcasts to all elements)
print("\na + 100 =", a + 100)
print("a * 1.5 =", a * 1.5)
print("a / 10 =", a / 10)

Example 2: Real-world — HR salary calculations.

# Employee data
salaries = np.array([75000, 85000, 92000,
68000, 55000, 110000])
years_exp = np.array([3, 5, 8, 2, 1, 12])

# 10% hike for all
new_salaries = salaries * 1.10
print("After 10% hike:", new_salaries)

# Annual salary
annual = salaries * 12
print("Annual salaries:", annual)

# Bonus based on experience (1000 per year)
bonus = years_exp * 1000
total_comp = salaries + bonus
print("Total compensation:", total_comp)

# Tax deduction (30%)
tax = salaries * 0.30
take_home = salaries - tax
print("Take home (after 30% tax):", take_home)

📊 Expected Output:

a + b = [12 24 36 48 60]
a - b = [ 8 16 24 32 40]
a * b = [ 20 80 180 320 500]
a / b = [5. 5. 5. 5. 5.]
a // b = [5 5 5 5 5]
a % b = [0 0 0 0 0]
a ** 2 = [ 100 400 900 1600 2500]

a + 100 = [110 120 130 140 150]
a * 1.5 = [15. 30. 45. 60. 75.]
a / 10 = [1. 2. 3. 4. 5.]

After 10% hike: [ 82500. 93500. 101200. 74800. 60500. 121000.]
Annual salaries: [ 900000 1020000 1104000 816000 660000 1320000]
Total compensation: [ 78000 90000 100000 70000 56000 122000]
Take home (after 30% tax): [ 52500. 59500. 64400. 47600. 38500. 77000.]

text

⚠️ Common Mistakes:

  • Mistake: Different shape arrays ka operation karna → np.array([1,2,3]) + np.array([1,2]) → ValueError!
    Fix: Same shape arrays use karo ya broadcasting rules samjho (Topic 2).
  • Mistake: Integer division result float expect karna → np.array([5,10]) / np.array([2,3]) → float result.
    Fix: Integer division ke liye // use karo.
  • Mistake: In-place operation bhool jaana → salaries * 1.10 original change nahi karta — new array return hoti hai.
    Fix: salaries = salaries * 1.10 ya salaries *= 1.10 use karo.

💬 Interview Questions:

Q1: What is vectorization in NumPy?
Ans: Vectorization means applying operations to entire arrays without explicit Python loops. NumPy operations are implemented in C and use SIMD (Single Instruction Multiple Data) CPU instructions to process multiple elements simultaneously. This eliminates the Python interpreter overhead for each element, making operations 50-100x faster than equivalent Python loops.

Q2: What is the difference between * and np.dot() for arrays?
Ans: * performs element-wise multiplication — each element multiplied by corresponding element, requires same shape. np.dot() performs matrix multiplication (dot product) — for 2D arrays it follows linear algebra rules (rows of first × columns of second). For 1D arrays, np.dot() returns the scalar dot product. In data science, use * for element-wise scaling and np.dot() for matrix operations.

2. Broadcasting — Different Shape Arrays

text

🔍 Definition: Broadcasting is NumPy's mechanism to perform operations between arrays of different shapes by automatically expanding the smaller array to match the larger one. NumPy follows specific rules to determine if two shapes are broadcast-compatible and how to expand them.

🎯 Samjho Simple Bhasha Mein: Broadcasting matlab "chota array automatically bada ho jaata hai" operation ke liye. Socho ek 3x3 matrix hai aur tumhe har row mein alag value add karni hai. Normally dono ka shape match hona chahiye. Broadcasting se ek 1D array (3 elements) ko 3x3 ke saath directly add kar sakte ho — NumPy internally row ko repeat karta hai. Yeh bahut common pattern hai data science mein.

💡 Broadcasting Rules:
Rule 1: Agar shapes alag hain toh choti shape ko left side pe 1s se pad karo
Rule 2: Size 1 wali dimension match karne wali dimension tak stretch hoti hai
Rule 3: Agar dono compatible nahi hain → ValueError

Compatible Examples:
(3,3) + (3,) → (3,3) ✅ (1D row broadcast)
(3,3) + (3,1) → (3,3) ✅ (column broadcast)
(3,3) + (2,) → Error ❌ (incompatible)

💻 Real-World Code Examples:

Example 1: Basic broadcasting patterns.

import numpy as np
# Scalar + Array (most basic broadcasting)
arr = np.array([1, 2, 3, 4, 5])
print("arr + 10:", arr + 10) # 10 broadcasts to [10,10,10,10,10]

# 2D + 1D (row-wise broadcast)
matrix = np.array([[1, 2, 3],
[4, 5, 6],
[7, 8, 9]])
row_add = np.array([10, 20, 30]) # Shape: (3,)
result = matrix + row_add # (3,3) + (3,) → (3,3)
print("\nMatrix + row:\n", result)

# 2D + column vector (column-wise broadcast)
col_add = np.array([[100], # Shape: (3,1)
[200],
[300]])
result_col = matrix + col_add # (3,3) + (3,1) → (3,3)
print("\nMatrix + column:\n", result_col)

Example 2: Real-world — Normalize scores department-wise.

# Student scores: 4 students, 3 subjects
scores = np.array([[85, 90, 78],
[92, 88, 95],
[70, 75, 80],
[88, 92, 85]])

# Subject-wise mean (mean of each column)
subject_mean = scores.mean(axis=0) # Shape: (3,)
print("Subject means:", subject_mean)

# Normalize: subtract subject mean from each student
# Broadcasting: (4,3) - (3,) → (4,3)
normalized = scores - subject_mean
print("\nNormalized scores:\n", normalized)

# Subject-wise standard deviation
subject_std = scores.std(axis=0)
# Z-score normalization
z_scores = (scores - subject_mean) / subject_std
print("\nZ-scores:\n", np.round(z_scores, 2))

📊 Expected Output:

arr + 10: [11 12 13 14 15]
Matrix + row:
[[11 22 33]
[14 25 36]
[17 28 39]]

Matrix + column:
[[101 102 103]
[204 205 206]
[307 308 309]]

Subject means: [83.75 86.25 84.5 ]

Normalized scores:
[[ 1.25 3.75 -6.5 ]
[ 8.25 1.75 10.5 ]
[-13.75 -11.25 -4.5]
[ 4.25 5.75 0.5 ]]

Z-scores:
[[ 0.14 0.49 -0.86]
[ 0.91 0.23 1.39]
[-1.52 -1.47 -0.6 ]
[ 0.47 0.75 0.07]]

text

⚠️ Common Mistakes:

  • Mistake: Incompatible shapes broadcast karna → (3,3) + (2,3) → Error!
    Fix: Shapes compare karo right-to-left — har dimension ya to equal honi chahiye ya ek mein 1 hona chahiye.
  • Mistake: Row vs column broadcast confuse karna → (3,) row broadcast karta hai. (3,1) column broadcast karta hai.
    Fix: Column broadcast ke liye reshape: arr.reshape(-1, 1) → shape (n, 1).

💬 Interview Questions:

Q1: What is Broadcasting in NumPy and what are its rules?
Ans: Broadcasting allows operations between arrays of different shapes by virtually expanding smaller arrays. Rules: (1) Compare shapes from right to left. (2) Dimensions are compatible if equal or one of them is 1. (3) Missing dimensions are treated as 1. Example: (4,3) and (3,) → compare 3==3 ✅, missing dim treated as 1 → broadcasts to (4,3). No actual data is copied — NumPy uses memory-efficient strides.

Q2: How do you subtract column-wise mean from a 2D array?
Ans: Use keepdims=True or reshape: mean = arr.mean(axis=0) gives shape (n_cols,). Then arr - mean broadcasts correctly (row-wise subtraction). For column-wise: mean = arr.mean(axis=1, keepdims=True) gives shape (n_rows, 1), then arr - mean broadcasts column-wise. keepdims=True preserves the dimension for correct broadcasting.

3. Comparison Operators — Boolean Arrays

text

🔍 Definition: Comparison operators in NumPy return a boolean array where each element is True or False based on the condition. These boolean arrays can be used for filtering (boolean indexing), counting (sum on bool array), and logical combinations using & (and), | (or), ~ (not).

🎯 Samjho Simple Bhasha Mein: Comparison operators se ek "mask" banta hai — True/False ka array. Socho tumhare paas 8 employees ki salaries hain aur dekhna chahte ho kitne 80K se zyada kamate hain. salaries > 80000 likhte hi ek True/False array milta hai. Is mask ko directly array pe apply karo — sirf True wale elements milenge. Yeh pandas ke boolean filtering jaisa hi kaam karta hai.

💡 Comparison + Logical Operators:
arr > val, arr < val, arr == val, arr != val
arr >= val, arr <= val
mask1 & mask2 → AND (dono True)
mask1 | mask2 → OR (koi ek True)
~mask → NOT (ulta karo)
np.any(mask) → Koi ek True hai?
np.all(mask) → Sab True hain?

💻 Real-World Code Examples:

Example 1: Basic comparison operations.

import numpy as np
salaries = np.array([75000, 85000, 92000, 68000,
55000, 110000, 48000, 95000])

# Basic comparisons — returns boolean array
print("> 80000:", salaries > 80000)
print("== 68000:", salaries == 68000)
print("!= 55000:", salaries != 55000)

# Count how many satisfy condition
high_count = np.sum(salaries > 80000)
print(f"\nEmployees earning >80K: {high_count}")

# Percentage
pct = np.mean(salaries > 80000) * 100
print(f"Percentage earning >80K: {pct:.1f}%")

Example 2: Logical combinations aur filtering.

# Logical AND: 60K se 90K ke beech
mid_mask = (salaries >= 60000) & (salaries 90000)
print("Mid range salaries:", salaries[mid_mask])

# Logical OR: 100K
extreme_mask = (salaries 55000) | (salaries > 100000)
print("Extreme salaries:", salaries[extreme_mask])

# NOT: 80K se zyada nahi kamane wale
not_high = salaries[~(salaries > 80000)]
print("Not high earners:", not_high)

# np.any() and np.all()
print("\nAny salary > 100K?", np.any(salaries > 100000))
print("All salaries > 40K?", np.all(salaries > 40000))
print("All salaries > 60K?", np.all(salaries > 60000))

# np.where for indices
high_indices = np.where(salaries > 80000)
print("Indices of high earners:", high_indices[0])

📊 Expected Output:

> 80000: [False True True False False True False True]
== 68000: [False False False True False False False False]
!= 55000: [ True True True True False True True True]

Employees earning >80K: 4
Percentage earning >80K: 50.0%

Mid range salaries: [75000 85000 68000]
Extreme salaries: [ 48000 110000]
Not high earners: [75000 68000 55000 48000]

Any salary > 100K? True
All salaries > 40K? True
All salaries > 60K? False
Indices of high earners: [1 2 5 7]

text

⚠️ Common Mistakes:

  • Mistake: and/or use karna instead of &/| → "ValueError: The truth value of an array is ambiguous."
    Fix: Always use &, |, ~ for NumPy boolean operations.
  • Mistake: Conditions ko parentheses mein wrap na karna → Operator precedence issue: arr > 5 & arr < 10 wrong hai.
    Fix: (arr > 5) & (arr < 10) — hamesha parentheses lagao.
  • Mistake: np.sum(bool_arr) ko np.count samajhna → np.sum(bool_arr) True values count karta hai (True=1, False=0).
    Fix: Yeh correct hai — NumPy mein True=1, False=0 hota hai.

💬 Interview Questions:

Q1: How do you count elements satisfying a condition in NumPy?
Ans: Use np.sum() on a boolean array: np.sum(arr > 80000) counts True values (True=1, False=0). Alternatively np.count_nonzero(arr > 80000) does the same. For percentage: np.mean(arr > 80000) * 100 gives the percentage of elements satisfying the condition.

Q2: What is the difference between np.any() and np.all()?
Ans: np.any(condition) returns True if at least ONE element satisfies the condition — like logical OR across all elements. np.all(condition) returns True only if ALL elements satisfy the condition — like logical AND. Use np.any() to check if any anomaly exists, np.all() to validate data quality (all values within range).

4. np.where() — Conditional Selection

text

🔍 Definition: np.where(condition, x, y) returns elements from array x where condition is True, and from array y where condition is False. It is the NumPy equivalent of a ternary operator or Excel IF function — applied element-wise to entire arrays.

🎯 Samjho Simple Bhasha Mein: np.where() Excel ke IF formula jaisa hai — "agar yeh condition True hai toh yeh value do, warna woh value do." Socho salary array hai — 80K se zyada wale ko "Senior", baaki ko "Junior" label dena hai. np.where se ek line mein poora array label ho jaata hai. Loop ki zaroorat nahi!

💡 np.where() — 2 Uses:
3 arguments: np.where(condition, true_val, false_val) → Replace values
1 argument: np.where(condition) → Returns INDICES where True

Nested np.where → Multiple conditions (like ELIF chain)

💻 Real-World Code Examples:

Example 1: Basic np.where() usage.

import numpy as np
salaries = np.array([75000, 85000, 92000, 68000,
55000, 110000, 48000, 95000])

# Label: Senior if >80K, else Junior
labels = np.where(salaries > 80000, 'Senior', 'Junior')
print("Labels:", labels)

# Apply 15% hike to high earners, 10% to others
new_sal = np.where(salaries > 80000,
salaries * 1.15, # True: 15% hike
salaries * 1.10) # False: 10% hike
print("New salaries:", new_sal)

# Get indices where condition is True
high_idx = np.where(salaries > 80000)[0]
print("High earner indices:", high_idx)

Example 2: Nested np.where() — multiple conditions.

# Grade system: A>=90, B>=80, C>=70, D otherwise
scores = np.array([95, 82, 71, 65, 88, 45, 93, 77])

grades = np.where(scores >= 90, 'A',
np.where(scores >= 80, 'B',
np.where(scores >= 70, 'C', 'D')))

print("Scores:", scores)
print("Grades:", grades)

# Real-world: Replace negative values with 0
data = np.array([5, -3, 8, -1, 0, 12, -5])
cleaned = np.where(data 0, 0, data) # ReLU function!
print("\nOriginal:", data)
print("Cleaned (negatives→0):", cleaned)

📊 Expected Output:

Labels: ['Junior' 'Senior' 'Senior' 'Junior' 'Junior' 'Senior' 'Junior' 'Senior']
New salaries: [ 82500. 97750. 105800. 74800. 60500. 126500. 52800. 109250.]
High earner indices: [1 2 5 7]

Scores: [95 82 71 65 88 45 93 77]
Grades: ['A' 'B' 'C' 'D' 'B' 'D' 'A' 'C']

Original: [ 5 -3 8 -1 0 12 -5]
Cleaned (negatives→0): [ 5 0 8 0 0 12 0]

text

⚠️ Common Mistakes:

  • Mistake: np.where(condition) ko 3-arg version samajhna → 1 arg version indices return karta hai as tuple.
    Fix: np.where(cond)[0] → 1D indices. np.where(cond, x, y) → Values.
  • Mistake: np.where mein string aur number mix karna → dtype issues aa sakte hain.
    Fix: Same type values use karo true_val aur false_val mein.

💬 Interview Questions:

Q1: What is np.where() and how does it differ from boolean indexing?
Ans: np.where(condition, x, y) selects values from x when True and from y when False — returns an array of same shape as condition. Boolean indexing arr[condition] returns only the matching elements (smaller array). np.where preserves the original array shape while replacing values; boolean indexing returns a subset. np.where is like Excel IF; boolean indexing is like filtering.

Q2: What is the ReLU function and how do you implement it with np.where?
Ans: ReLU (Rectified Linear Unit) is an activation function in neural networks: f(x) = max(0, x). In NumPy: np.where(x < 0, 0, x) or equivalently np.maximum(0, x). It replaces all negative values with 0 and keeps positive values unchanged. np.maximum(0, arr) is slightly more concise and preferred in practice.

5. Aggregate Functions — sum, mean, std, min, max

text

🔍 Definition: Aggregate functions reduce an array (or along an axis) to a single value or smaller array. Key functions: np.sum(), np.mean(), np.std(), np.var(), np.min(), np.max(), np.median(), np.percentile(). The axis parameter controls the direction of aggregation.

🎯 Samjho Simple Bhasha Mein: Aggregate functions data ko summarize karte hain. axis=0 matlab "column-wise" — har column ka separately calculation. axis=1 matlab "row-wise" — har row ka separately calculation. Bina axis ke — poora array ek number mein. Yeh Pandas ke .mean(), .sum() jaisa hi hai — NumPy ka version.

💡 Axis Parameter:
No axis: Poore array ka ek value
axis=0: Column-wise (collapse rows) → result shape = (n_cols,)
axis=1: Row-wise (collapse columns) → result shape = (n_rows,)

Trick: axis=0 → result mein rows khatam hoti hain
axis=1 → result mein columns khatam hoti hain

💻 Real-World Code Examples:

Example 1: All aggregate functions on salary data.

import numpy as np
salaries = np.array([75000, 85000, 92000, 68000,
55000, 110000, 48000, 95000],
dtype=np.float64)

print(f"Sum: {np.sum(salaries):,.0f}")
print(f"Mean: {np.mean(salaries):,.2f}")
print(f"Median: {np.median(salaries):,.2f}")
print(f"Std Dev: {np.std(salaries):,.2f}")
print(f"Variance: {np.var(salaries):,.2f}")
print(f"Min: {np.min(salaries):,.0f}")
print(f"Max: {np.max(salaries):,.0f}")
print(f"Range: {np.max(salaries) - np.min(salaries):,.0f}")

# Percentiles
print(f"\n25th percentile: {np.percentile(salaries, 25):,.0f}")
print(f"75th percentile: {np.percentile(salaries, 75):,.0f}")
print(f"IQR: {np.percentile(salaries, 75) - np.percentile(salaries, 25):,.0f}")

# Argmin, Argmax (index of min/max)
print(f"\nMin index: {np.argmin(salaries)}")
print(f"Max index: {np.argmax(salaries)}")

Example 2: Axis-wise aggregation — marks analysis.

# 4 students, 3 subjects
marks = np.array([[85, 90, 78],
[92, 88, 95],
[70, 75, 80],
[88, 92, 85]])

# No axis: Overall average
print("Overall avg:", np.mean(marks))

# axis=0: Subject-wise average (column-wise)
subject_avg = np.mean(marks, axis=0)
print("Subject avg (Math, Sci, Eng):", subject_avg)

# axis=1: Student-wise average (row-wise)
student_avg = np.mean(marks, axis=1)
print("Student avg:", student_avg)

# Student-wise total marks
student_total = np.sum(marks, axis=1)
print("Student totals:", student_total)

# keepdims — preserve dimensions
avg_keepdims = np.mean(marks, axis=1, keepdims=True)
print("keepdims shape:", avg_keepdims.shape) # (4,1) not (4,)

📊 Expected Output:

Sum: 628,000
Mean: 78,500.00
Median: 80,000.00
Std Dev: 20,263.94
Variance: 410,625,000.00
Min: 48,000
Max: 110,000
Range: 62,000

25th percentile: 62,750
75th percentile: 92,750
IQR: 30,000

Min index: 6
Max index: 5

Overall avg: 84.08
Subject avg (Math, Sci, Eng): [83.75 86.25 84.5]
Student avg: [84.33 91.67 75. 88.33]
Student totals: [253 275 225 265]
keepdims shape: (4, 1)

text

⚠️ Common Mistakes:

  • Mistake: axis direction ulta samajhna → axis=0 column-wise karta hai (rows collapse), axis=1 row-wise (columns collapse).
    Fix: Yaad rakho: axis=0 → result mein rows khatam. axis=1 → result mein columns khatam.
  • Mistake: mean() aur median() ko same samajhna → Outliers honge toh mean zyada shift hoga — median robust hota hai.
    Fix: Skewed data ke liye median prefer karo (e.g., salary data).
  • Mistake: argmax() ka result use karna bina [0] ke — waise toh theek hai, but 2D mein axis specify karo.
    Fix: np.argmax(arr, axis=1) for row-wise max index.

💬 Interview Questions:

Q1: What is the difference between np.mean() and np.median()?
Ans: np.mean() calculates arithmetic average — sensitive to outliers (one extreme value shifts it significantly). np.median() returns the middle value when sorted — robust to outliers. For salary data: [50K, 60K, 55K, 1M] → mean=291K (misleading), median=57.5K (representative). Use mean for symmetric distributions, median for skewed data with outliers.

Q2: What does keepdims=True do in aggregate functions?
Ans: Without keepdims, aggregation along an axis removes that dimension: shape (4,3) with axis=1 → shape (4,). With keepdims=True, the aggregated dimension is kept as size 1: shape (4,3) with axis=1 → shape (4,1). This is essential for broadcasting — keeping dimensions aligned allows direct subtraction: arr - arr.mean(axis=1, keepdims=True) normalizes each row correctly.

6. Math Functions — sqrt, log, exp, round, abs

text

🔍 Definition: NumPy provides universal mathematical functions (ufuncs) that operate element-wise on arrays. These include: np.sqrt(), np.log(), np.log2(), np.log10(), np.exp(), np.abs(), np.round(), np.floor(), np.ceil(), np.clip(), np.power(), and trigonometric functions.

🎯 Samjho Simple Bhasha Mein: NumPy ke math functions Python ke math module se alag hain — yeh arrays par directly kaam karte hain. math.sqrt(9) sirf ek number pe kaam karta hai. np.sqrt(arr) poore array ke har element ka square root ek saath nikalta hai. Data science mein log transformation, normalization, aur activation functions ke liye yeh bahut use hote hain.

💡 Common Math Functions:
np.sqrt(arr) → Square root
np.abs(arr) → Absolute value
np.round(arr, n) → Round to n decimals
np.floor(arr) → Round down
np.ceil(arr) → Round up
np.log(arr) → Natural log (ln)
np.log2(arr) → Log base 2
np.log10(arr) → Log base 10
np.exp(arr) → e^x
np.clip(arr, min, max) → Clamp values

💻 Real-World Code Examples:

Example 1: Common math operations.

import numpy as np
arr = np.array([1, 4, 9, 16, 25, 100])

print("sqrt:", np.sqrt(arr))
print("log:", np.round(np.log(arr), 2))
print("log10:", np.round(np.log10(arr), 2))
print("exp:", np.round(np.exp([0, 1, 2, 3]), 2))

# Rounding functions
decimals = np.array([3.14159, 2.71828, 1.41421, 9.87654])
print("\nround(2):", np.round(decimals, 2))
print("floor:", np.floor(decimals))
print("ceil:", np.ceil(decimals))

# Absolute value
data = np.array([-5, 3, -8, 2, -1])
print("\nabs:", np.abs(data))

Example 2: np.clip() aur real-world applications.

# np.clip() — clamp values between min and max
scores = np.array([105, 88, -5, 72, 110, 45, 95])
valid_scores = np.clip(scores, 0, 100) # Keep between 0-100
print("Original:", scores)
print("Clipped:", valid_scores)

# Log transformation (for skewed salary data)
salaries = np.array([25000, 50000, 75000, 200000, 500000])
log_salaries = np.log(salaries)
print("\nLog transformed salaries:", np.round(log_salaries, 2))

# Sigmoid function (neural networks)
z = np.array([-2, -1, 0, 1, 2])
sigmoid = 1 / (1 + np.exp(-z))
print("\nSigmoid:", np.round(sigmoid, 3))

# Cumulative sum (running total)
monthly_sales = np.array([100, 150, 120, 180, 200, 160])
cumulative = np.cumsum(monthly_sales)
print("\nCumulative sales:", cumulative)

📊 Expected Output:

sqrt: [ 1. 2. 3. 4. 5. 10.]
log: [0. 1.39 2.2 2.77 3.22 4.61]
log10: [0. 0.6 0.95 1.2 1.4 2. ]
exp: [ 1. 2.72 7.39 20.09]

round(2): [3.14 2.72 1.41 9.88]
floor: [3. 2. 1. 9.]
ceil: [4. 3. 2. 10.]

abs: [5 3 8 2 1]

Original: [105 88 -5 72 110 45 95]
Clipped: [100 88 0 72 100 45 95]

Log transformed salaries: [10.13 10.82 11.22 12.21 13.12]

Sigmoid: [0.119 0.269 0.5 0.731 0.881]

Cumulative sales: [100 250 370 550 750 910]

text

⚠️ Common Mistakes:

  • Mistake: np.log(0) ya negative number → -inf ya NaN return karega.
    Fix: Pehle check karo: np.log(np.where(arr > 0, arr, np.nan))
  • Mistake: Python's math.sqrt() use karna array pe → TypeError — sirf scalar ke liye.
    Fix: Arrays ke liye hamesha np.sqrt() use karo.
  • Mistake: np.round() ka rounding behavior bhool jaana → Banker's rounding (round half to even): np.round(2.5)=2, np.round(3.5)=4.
    Fix: Agar standard rounding chahiye → Python's round() or decimal module.

💬 Interview Questions:

Q1: Why do we use log transformation on salary/income data?
Ans: Income data is heavily right-skewed — most people earn low to medium salaries but a few earn very high amounts. Log transformation compresses the scale of large values while spreading small values, making the distribution more symmetric (closer to normal). This improves the performance of many machine learning algorithms that assume normally distributed features.

Q2: What is np.clip() and when is it used?
Ans: np.clip(arr, a_min, a_max) limits all values to be within [a_min, a_max] range — values below a_min become a_min, values above a_max become a_max, others stay unchanged. Used for: preventing out-of-range scores (0-100), preventing division by zero (clip to small positive value), gradient clipping in neural networks (prevent exploding gradients).

7. np.random — Random Number Generation

text

🔍 Definition: np.random module provides functions for generating random numbers from various probability distributions. Key functions: np.random.rand() (uniform), np.random.randn() (normal/Gaussian), np.random.randint() (integers), np.random.choice() (sampling), np.random.shuffle() (in-place shuffle), and np.random.seed() (reproducibility).

🎯 Samjho Simple Bhasha Mein: Data science mein random numbers bahut use hote hain — test data banana, model weight initialize karna, train-test split karna, bootstrapping karna. np.random module yeh sab karta hai. np.random.seed() set karo taaki results reproducible hon — same seed se hamesha same "random" numbers milenge. Yeh very important hai research aur testing mein.

💡 Key Random Functions:
np.random.seed(n) → Set seed for reproducibility
np.random.rand(shape) → Uniform [0,1)
np.random.randn(shape) → Standard Normal (mean=0, std=1)
np.random.randint(low, high, size) → Random integers
np.random.normal(mean, std, size) → Normal distribution
np.random.choice(arr, size, replace) → Random sampling
np.random.shuffle(arr) → In-place shuffle

💻 Real-World Code Examples:

Example 1: Basic random number generation.

import numpy as np
# Seed for reproducibility
np.random.seed(42)

# Uniform random [0,1)
uniform = np.random.rand(5)
print("Uniform [0,1):", np.round(uniform, 3))

# Normal distribution (mean=0, std=1)
normal = np.random.randn(5)
print("Normal (std):", np.round(normal, 3))

# Random integers
dice = np.random.randint(1, 7, size=10) # 1 to 6 dice rolls
print("Dice rolls:", dice)

# Custom normal distribution
salaries_sim = np.random.normal(
loc=75000, # mean
scale=15000, # std dev
size=1000 # 1000 samples
)
print(f"\nSimulated salaries — Mean: {salaries_sim.mean():.0f}, Std: {salaries_sim.std():.0f}")

Example 2: Sampling, shuffle aur train-test split.

import numpy as np
np.random.seed(42)

# Random sampling with replacement (Bootstrap)
population = np.array([10, 20, 30, 40, 50])
sample_with = np.random.choice(population, size=10, replace=True)
sample_without = np.random.choice(population, size=3, replace=False)
print("With replacement:", sample_with)
print("Without replacement:", sample_without)

# Shuffle array (in-place)
arr = np.arange(10)
np.random.shuffle(arr)
print("Shuffled:", arr)

# Manual Train-Test Split (80-20)
data = np.arange(100)
indices = np.random.permutation(100) # Shuffled indices
train_size = int(0.8 * 100)
train_idx = indices[:train_size]
test_idx = indices[train_size:]
train_data = data[train_idx]
test_data = data[test_idx]
print(f"\nTrain size: {len(train_data)}, Test size: {len(test_data)}")
print("First 5 train indices:", train_idx[:5])

📊 Expected Output:

Uniform [0,1): [0.374 0.951 0.732 0.599 0.156]
Normal (std): [ 0.497 -0.138 0.648 1.524 1.458]
Dice rolls: [1 4 5 3 2 6 2 4 1 3]

Simulated salaries — Mean: 74952, Std: 14987

With replacement: [40 20 50 10 30 40 30 20 10 50]
Without replacement: [20 40 10]
Shuffled: [3 7 1 0 5 9 2 4 6 8]

Train size: 80, Test size: 20
First 5 train indices: [51 92 14 71 60]

⚠️ Common Mistakes:

  • Mistake: np.random.seed() bhool jaana → Results reproducible nahi honge — baar baar alag numbers.
    Fix: Research aur testing mein hamesha seed set karo. Production mein seed mat lagao (actual randomness chahiye).
  • Mistake: np.random.rand() aur np.random.randn() confuse karna → rand = uniform [0,1). randn = normal (can be negative).
    Fix: Neural network weights → randn (normal). Probability sampling → rand (uniform).
  • Mistake: np.random.shuffle() ka return value use karna → shuffle in-place karta hai aur None return karta hai.
    Fix: np.random.shuffle(arr) — arr directly modify ho jaata hai. Return value use mat karo.

💬 Interview Questions:

Q1: What is np.random.seed() and why is it important?
Ans: np.random.seed(n) initializes the random number generator with a fixed starting point (seed). This makes "random" number generation deterministic — same seed always produces same sequence. Essential for: reproducible experiments, debugging (same data every run), sharing code where others need to replicate results. Without seed, results change every run making debugging difficult.

Q2: What is the difference between np.random.rand() and np.random.randn()?
Ans: np.random.rand(shape) generates from Uniform distribution [0, 1) — all values between 0 and 1 with equal probability. np.random.randn(shape) generates from Standard Normal distribution (mean=0, std=1) — values cluster around 0, can be negative, ~68% within [-1,1]. Use rand for probabilities/fractions, randn for initializing neural network weights.

Q3: What is the difference between np.random.shuffle() and np.random.permutation()?
Ans: np.random.shuffle(arr) shuffles arr IN-PLACE (modifies original, returns None). np.random.permutation(n) returns a NEW shuffled array (original unchanged) — can take an integer n to return shuffled range(n) or an array to return shuffled copy. Use shuffle when you want to modify original; use permutation when you need shuffled indices while keeping original intact.

Summary — NumPy Part 2 Quick Reference

Category Function Purpose Key Note
Arithmetic +, -, *, /, ** Element-wise math Same shape required
Broadcasting Auto shape expansion Different shapes Compare right-to-left
Comparison >, <, ==, &, |, ~ Boolean arrays Use & not 'and'
Conditional np.where() IF-ELSE on arrays 3 args=values, 1 arg=indices
Aggregate sum, mean, std, max Reduce to value axis=0 col, axis=1 row
Math sqrt, log, exp, clip Math operations Element-wise ufuncs
Random rand, randn, randint Random generation seed() for reproducibility

Next: Data Insights NumPy Masterclass — Part 3

Part 3 mein hum cover karenge: np.dot() Matrix Multiplication, Sorting, Stacking, Splitting, aur NumPy + Pandas Integration — real-world data pipeline examples ke saath Data Insights par.

Happy Coding & Keep Learning! 🚀

👤
Jatin Kumar
Data Analyst & Educator

Python, SQL, Power BI aur Excel mein practical tutorials likhta hoon — taaki data analytics seekhna aasan ho. Portfolio: jatinanalytics.co.in

Portfolio LinkedIn GitHub Kaggle All Articles
Share:

💬 Comments (0)

Spam/links allowed nahi hain — respectful comments welcome!

Loading comments...

Was this article helpful?
Previous Article"NumPy Fundamentals — Arrays, Creation And Indexing"Next Article Matrix Operations, Sorting, Stacking And NumPy + Pandas Inte

📚 More Articles Like This

Advanced Charts, Graph Objects And Export in Plotly

Read Article

Interactive Charts with Plotly Express

Read Article

Relationships, Matrix Charts And Styling in Seaborn

Read Article