Data Type Conversion And Mapping in Pandas
Data Type Conversion & Mapping in Pandas: Complete Guide
Galat data types systems memory aur performance ko dheema karte hain. Sikhiye columns ko map, rename aur convert karne ke saare professional techniques jatinanalytics par.
📑 Is Masterclass Guide Mein Aap Kya Sikhenge:
DataFrame structures ko modify, standardize aur scale karne ke 6 powerful methods:
- Explicit Type Casting: astype() and memory optimizations using 'category' types
- Robust Numeric Conversion: Handling dirty numeric strings with pd.to_numeric()
- Value Mapping & Translating: Element-wise dictionary mapping with map()
- DataFrame-wide Replacements: Bulk replacements using df.replace()
- Metadata Management: Clean schema column renaming with rename()
- Subset Selection: Filtering columns based on data types with select_dtypes()
1. astype() — Explicit Data Type Casting
🔍 Kya Hai: astype() Pandas ka sabse common casting operator hai. Iska use karke hum kisi bhi column ko manually float, integer, string, boolean ya categorical data type me force convert kar sakte hain.
🎯 Kyu Use Hota Hai: System databases se numbers kabhi-kabhi float formats me aate hain (e.g. Age: 25.0). Unhein standard integers me badalne ya object columns ko memory-saving category data type me badalne ke liye iska use hota hai.
💡 Kab Use Hota Hai: Data cleaning ke initial phase me, jab hume schema types ko standard formats me alignment dena hota hai.
💻 Real-World Code Examples:
Example 1: Age column ko float64 se clean integer type me convert karna.
import pandas as pd
employees = pd.read_csv("employees.csv")
# Convert Age from float/object to integer
employees["Age"] = employees["Age"].astype(int)
Example 2: High-cardinality text column ko "category" type me convert karke memory usage optimized karna.
employees["Department"] = employees["Department"].astype("category")
📊 Expected Output:
# Before: Age (float64) -> After: Age (int32)
# Before: Department (object) -> After: Department (category)
✅ Best Practices — Data Insights Rulebook:
- Garbage or empty values (NaN) wale numerical columns par direct
astype(int)mat lagayein, yeh float conversion par error dega kyunki Integer NaN support nahi karta (standard Pandas versions me). - Sirf unhi textual columns ko
categoryme convert karein jinki cardinality low ho (jaise Gender, City, Department) taaki memory optimize ho. - Multiple casting operations ek hi dict me run karein:
df.astype({'Age': int, 'Salary': float}).
💬 Crack the Interview:
Q1: NaN values wale numerical column ko integer me convert karne par kya error aata hai?
Ans: Yeh standard ValueError raise karega: Cannot convert non-finite values (NA or inf) to integer. Iska solution hai pehle missing values fill karna ya modern nullable integer type Int64 (capital I) use karna.
Q2: object data type ko category type me convert karne se memory kaise bachti hai?
Ans: category type strings ko uniquely code index (0, 1, 2) me convert karke background me save karta hai, jisse duplicate text string objects ka memory heap footprint reduce ho jata hai.
Q3: What is the difference between float64 and float32 types casted via astype?
Ans: float32 uses half the memory (32 bits) of float64, which is excellent for compressing massive numeric scale datasets in memory-constraint platforms.
2. pd.to_numeric() — Safe Numeric Imputation
🔍 Kya Hai: pd.to_numeric() ek robust system function hai jo string objects me likhe numerical metrics ko automatic standard numeric formats (float64 or int64) me cast karta hai, aur isme error handling ka advanced support milta hai.
🎯 Kyu Use Hota Hai: Kuch datasets me numeric fields ke beech me random alphanumeric garbage values (jaise "unknown", "abc", ya space) ghus aati hain. Standard astype() in garbage cases par direct crack crash ho jata hai, jabki to_numeric() ise safe handle karta hai.
💡 Kab Use Hota Hai: Web scraping data or messy user-generated numbers spreadsheets ko parse aur cleaning operations templates me transform karte waqt.
💻 Real-World Code Examples:
Example 1: Coercing errors to clean NaN values in messy salary metrics.
# Convert to numeric, bad values will gracefully become NaN
employees["Salary"] = pd.to_numeric(employees["Salary"], errors="coerce")
Example 2: Ignoring parsing errors, keeping original strings if transformation fails.
employees["Salary"] = pd.to_numeric(employees["Salary"], errors="ignore")
📊 Expected Output (Example 1):
# Input
values: ["45000", "55000", "not_disclosed", "60000"]
# Output
values: [45000.0, 55000.0, NaN, 60000.0] # Standard float type!
✅ Best Practices — Data Insights Rulebook:
errors='coerce'argument use karne se sara garbage dataNaNme convert ho jata hai. Is step ke immediate baad missing values imputation strategy lagayein.- Downcast numeric size dynamically to float32 or int32 using parameter:
downcast='float'ordowncast='signed'to compress memory limits. - Verify metrics results datatype after execution.
💬 Crack the Interview:
Q1: What are the three options for 'errors' parameter inside pd.to_numeric()?
Ans: 1. raise (default - crashes on error), 2. ignore (leaves invalid strings as is), 3. coerce (forces invalid strings to NaN).
Q2: astype('float') vs pd.to_numeric() me kya superiority difference hai?
Ans: astype() me dynamic error handling ka custom support nahi hota. Agar kisi row me text hoga toh astype crash ho jayega, jabki to_numeric() with coerce use smoothly handle kar lega.
Q3: How does 'downcast' parameter save memory?
Ans: It automatically shrinks the numeric limits down to smallest possible subtype (e.g. converting float64 elements to float32 if limits allow), preserving computational resources.
3. map() — Element-Wise Dictionary Translation
🔍 Kya Hai: map() ek translation engine hai. Yeh column ke har single cell value ko coordinate karke di gayi dictionary, standard key-value maps or dynamic mapping functions ke basis par target values me convert kar deta hai.
🎯 Kyu Use Hota Hai: Database normalization and standardizations me short keys ko full descriptions tags me swap out karne ke liye (e.g. translating "M" to "Male", "F" to "Female").
💡 Kab Use Hota Hai: Categorical variable transformation pipelines, code lookups mapping systems frameworks dashboards loops levels.
💻 Real-World Code Examples:
Example 1: Converting gender abbreviations key to standard descriptive text labels.
gender_dict = {"M": "Male", "F": "Female"}
employees["Gender"] = employees["Gender"].map(gender_dict)
Example 2: E-commerce mapping transaction card code descriptions via standard dictionary.
payment_dict = {"CC": "Credit Card", "UPI": "Unified Payments"}
ecommerce["PaymentMethod"] = ecommerce["PaymentMethod"].map(payment_dict)
📊 Expected Output:
# Before: ["M", "F", "M"]
# After: ["Male", "Female", "Male"] # Clean standardized
values!
✅ Best Practices — Data Insights Rulebook:
- Remember that any value NOT present inside your mapping dictionary will automatically be converted to
NaN. If you want to keep original values unchanged, usereplace()instead. - Validate mappings keys arrays matches.
- Keep lookups files structured clean separate modules.
💬 Crack the Interview:
Q1: Why does map() turn missing dictionary keys to NaN?
Ans: Because map() forces a complete translation of the entire domain. If a key is missing from dictionary definition, it assumes there is no mapping available and returns NaN.
Q2: How is Series.map() different from DataFrame.applymap()?
Ans: Series.map() works on a single column (1D Series) element-wise. applymap() works on the entire multi-column DataFrame grid.
Q3: Can we pass a standard Python custom function inside map()?
Ans: Yes, map() accepts functions as well: df['col'].map(lambda x: x.upper()) works similarly to apply.
4. df.replace() — DataFrame-Wide Replacements
🔍 Kya Hai: df.replace() target DataFrame level or column level par values ko replace karta hai. map() se alag, yeh un values ko safe chhod deta hai jo lookup list ya dict me nahi hain.
🎯 Kyu Use Hota Hai: Standard placeholders, outliers values, or custom system abbreviations elements swap out parameters configurations coordinates (without disturbing remaining valid columns structures).
💡 Kab Use Hota Hai: Clean targets values, multi-values modifications checks on huge files registries.
💻 Real-World Code Examples:
Example 1: Standard replacement, keeping undefined values safely intact.
employees["Department"] = employees["Department"].replace({"HR": "Human Resources"})
# Non matching
values e.g. "IT" remain untouched!
Example 2: Replacing system placeholders globally in the entire DataFrame.
import numpy as np
employees.replace("N/A", np.nan, inplace=True)
📊 Expected Output (Example 1):
# Before: ["HR", "IT", "Sales"]
# After: ["Human Resources", "IT", "Sales"] # "IT" & "Sales" are safely preserved!
✅ Best Practices — Data Insights Rulebook:
- Use
replace()instead ofmap()when you only want to update a subset of values and keep the rest intact. - To replace values in specific columns globally, use nested dictionary structures:
df.replace({'col1': {'old': 'new'}}). - Validate data counts before final commits.
💬 Crack the Interview:
Q1: What is the primary functional difference between map() and replace()?
Ans: map() will turn any value NOT listed in dictionary to NaN (complete domain translation). replace() will keep unlisted values unchanged.
Q2: Does replace() support wildcards or regex internally?
Ans: Yes, pass regex=True inside parameters list to support regex replacements: df.replace(to_replace=r'^temp_', value='clean_', regex=True).
Q3: Can we execute replace() globally across all string fields?
Ans: Yes, calling replace directly on DataFrame level (Example 2) will scan every cell across all columns and substitute matches.
5. rename() — Metadata Schema Management
🔍 Kya Hai: rename() ka use DataFrame ke indices labels ya column headers naming metadata schemas ko dictionary based mappings se safely rename karne ke liye kiya jata hai.
🎯 Kyu Use Hota Hai: Systems database standard schemas ke parameters naming requirements match indicators targets (e.g. converting dirty columns names "Emp_Sal" or "Joining_Dt" to clean professional labels "Salary" and "JoinDate").
💡 Kab Use Hota Hai: Data schema standardization phase, pre-analytics preparation pipelines setups me.
💻 Real-World Code Examples:
Example 1: Renaming dirty column headers cleanly using dictionary maps.
employees.rename(columns={"Emp_Salary": "Salary", "Dept_Id": "Department"}, inplace=True)
Example 2: Modifying indices labels values structures safely.
employees.rename(index={0: "First_Row"}, inplace=True)
📊 Expected Output:
# Column list transformed
from:
# ['Emp_Salary', 'Dept_Id', 'Name'] -> ['Salary', 'Department', 'Name']
✅ Best Practices — Data Insights Rulebook:
- Specify
columns=orindex=parameter explicitly. Directly passing dictionary without keywords will cause mapping errors. - To make changes permanent, hamesha
inplace=Trueuse karein ya original df variable me reassign karein. - Verify headers count list after rename step.
💬 Crack the Interview:
Q1: Can we rename columns using a function (like converting all headers to uppercase)?
Ans: Yes, pass string function reference inside columns: df.rename(columns=str.upper) will capitalize all column headers in 1 step.
Q2: If dictionary keys do not match any column name, will rename throw an error?
Ans: No, it will silently ignore non-matching keys and rename only the matched elements without raising errors.
Q3: How do we raise an error if a column to be renamed is missing?
Ans: Set parameter errors='raise' inside rename: df.rename(columns=dict, errors='raise').
6. select_dtypes() — Data Type Filtering
🔍 Kya Hai: select_dtypes() ek advanced filtering tool hai jo pure DataFrame me se sirf aapke specified criteria data types coordinates columns subset ko select or filter out karta hai.
🎯 Kyu Use Hota Hai: Batch calculations operations systems, separating numeric columns for math transforms or text columns for string cleanups (e.g. selecting all float columns to round them up in 1 step).
💡 Kab Use Hota Hai: Pipeline automation setups me, jahan data validation types checks parameters apply templates coordinates.
💻 Real-World Code Examples:
Example 1: Filtering out all numeric (float and int) columns for math analytics.
numeric_cols = employees.select_dtypes(include=["number"])
Example 2: Selecting everything except text object types columns.
non_text_cols = employees.select_dtypes(exclude=["object"])
📊 Expected Output (Example 1):
# Returns a filtered sub-dataframe containing ONLY: ['Age', 'Salary'] columns.
# Text columns like 'Name', 'Department' are excluded.
✅ Best Practices — Data Insights Rulebook:
- Use standard types tags coordinates:
include=['number']targets both integer and float columns, saving time. - To modify selected dtypes, safe reassign references cleanly using
.copy()structure logic. - Verify shape of the filtered output frame.
💬 Crack the Interview:
Q1: 'number' keyword include list checks target both float and integers?
Ans: Yes! 'number' keyword acts as a generic alias in Pandas, which includes int64, int32, float64, float32, etc. simultaneously.
Q2: Can we use both 'include' and 'exclude' parameters in a single select_dtypes() call?
Ans: Yes, but they must not overlap. If they overlap, it raises a ValueError. Best practice is to use one of them to keep logic clean.
Q3: What does select_dtypes(include=['category']) return if no category columns exist?
Ans: It returns an empty DataFrame with the original indices length but 0 columns, without raising errors.
Conclusion: Data Mapping Selection Matrix
Apne data conversion scenarios ke basis par right method chunye:
| Scenario / Situation | Recommended Tool | Key Implication |
|---|---|---|
| Explicit target casting | astype() |
Directly changes datatypes (float to int etc.). |
| Casting numeric strings with garbage | pd.to_numeric(errors='coerce') |
Forces invalid strings to standard NaN safely. |
| Complete element-wise translations | map() |
Maps key-values, converting unmatched items to NaN. |
| Partial replacements, keeping data intact | df.replace() |
Swaps target elements, leaving others unchanged. |
| Filtering columns for batch operations | select_dtypes() |
Selects numeric/object columns subsets cleanly. |
Next Post Preview: Masterclass Part 5
Next masterclass tutorial mein hum cover karenge: Numeric Cleaning & Mathematical Operations (round, abs, clip, and outliers capping) ko details layouts and interview questions ke sath jatinanalytics par.
Happy Coding & Keep Standardizing! 🚀
💬 Comments (0)
Loading comments...