Theory
Blind Slicing की दुविधा
आपके Semester 1 Python list labs में, आप matrix[0][1] जैसे basic integer brackets इस्तेमाल करके multi-dimensional matrices slice करते थे। पर जब Pandas में जटिल relational data tables process करते हैं, जहाँ rows को student roll numbers या date strings जैसे स्पष्ट custom labels से track किया जाता है, पूरी तरह hardcoded integer positions पर निर्भर रहना ख़तरनाक बन जाता है। अगर एक dataset shuffle, sort, या filter होता है, row position 3 पूरी तरह बदल जाती है! हम Python को कैसे बताते हैं कि एक ख़ास row को उसके official name badge बनाम उसकी असल vertical line position के आधार पर निकाले, और दोनों को उलझाना आपके production analytical pipelines को क्यों तोड़ सकता है?
Theory
Name-Badge Call बनाम Seat-Number Tap
अपने BCA classroom register की कल्पना कीजिए। हर student का एक unique संरचनात्मक identifier है, उसका Roll Number (Label)। वे आज एक ख़ास physical chair sequence भी घेरते हैं, आगे से पीछे तक (Integer Position: 0, 1, 2)। .loc इस्तेमाल करना classroom दरवाज़े पर खड़े होकर पुकारने जैसा है: 'मुझे Roll Number BCA102 के records लाओ।' उससे फ़र्क़ नहीं पड़ता वह student कहाँ बैठता है; label उन पर perfectly lock करता है। .iloc इस्तेमाल करना आँखें बंद करके classroom में चलने, 3री physical bench तक गिनने, और जो कोई भी उस chair में अभी बैठा हो उसे पकड़ने जैसा है, उनके नाम या roll number की परवाह किए बिना।
Theory
Label-Based बनाम Integer-Based Indexing औपचारिक रूप से
Pandas data selection को दो स्पष्ट methods में बाँटता है: .loc और .iloc। .loc property मुख्यतः label-based है; यह rows और columns को उनके string characters या स्पष्ट index values इस्तेमाल करके query करती है। इसके उलट, .iloc सख़्ती से integer position-based है; यह elements को zero-indexed integer offsets (0 से N-1) इस्तेमाल करके चुनती है। उनकी slicing boundaries में एक अहम mathematical अंतर है: standard Python और .iloc stop index boundary छोड़ते हैं (start:stop stop-1 पर रुकता है), जबकि .loc start और stop दोनों boundaries का पूरी तरह inclusive है।
At a glance
Table 1: Pandas .loc और .iloc extraction properties की गहरी functional mapping।
| Indexing Property | Target Query Mode | Slicing Boundary Rule | Syntax Example |
|---|---|---|---|
| df.loc[row_lbl, col_lbl] | स्पष्ट text/numeric names के ज़रिए label-based matching। | Inclusive: start और stop दोनों bounds रखे जाते हैं। | df.loc['t2', 'Amount'] |
| df.iloc[row_pos, col_pos] | numeric offsets के ज़रिए Integer position-based matching। | Exclusive: Stop bound छोड़ा जाता है (standard Python)। | df.iloc[1, 1] |
| df['ColumnName'] | सीधा bracket selection shortcut। | rows को पूरी तरह bypass करता है; individual columns सीधे निकालता है। | df['Amount'] |
Theory
Worked Example: PocketMoney Ledger Grid को विच्छेदित करना
आइए एक customized pocket-money ledger DataFrame पर precision extraction लागू करें। हम custom row strings (t1, t2, t3) बनाएँगे ताकि index names को physical positional indices से साफ़-साफ़ अलग करें, slicing operations के दौरान boundary अंतरों को उजागर करते हुए।
Practical
Precision Coordinate Extractor
import pandas as pd
# Step 1: Create a structured ledger with custom string index labels
raw_data = {
"Item": ["Samosa", "Bus", "Chai", "Notebook"],
"Cost": [35, 50, 15, 120]
}
ledger_df = pd.DataFrame(raw_data, index=["t1", "t2", "t3", "t4"])
# Step 2: Use label-based .loc to pull a specific coordinate value
bus_cost_loc = ledger_df.loc["t2", "Cost"]
# Step 3: Use integer position-based .iloc to pull the exact same cell
bus_cost_iloc = ledger_df.iloc[1, 1]
# Step 4: Compare label slicing vs position slicing boundaries
slice_loc = ledger_df.loc["t1":"t3"]
slice_iloc = ledger_df.iloc[0:2]
print(f"Extracted via .loc label: {bus_cost_loc}")
print(f"Extracted via .iloc integer position: {bus_cost_iloc}")
print("\n--- .loc Slicing ('t1':'t3') ---")
print(slice_loc)
print("\n--- .iloc Slicing (0:2) ---")
print(slice_iloc)This example runs in Gri-Learn on the web, where you can edit it and see the output.
Think first
Boundary Outputs का विश्लेषण करें
slicing lines को ध्यान से देखिए: slice_loc 't1':'t3' इस्तेमाल करती है, जबकि slice_iloc 0:2 इस्तेमाल करती है। हर statement कितनी rows output करेगी, और वे कौन से item records शामिल करेंगी?
Show the answer
script output करेगी:
Extracted via .loc label: 50
Extracted via .iloc integer position: 50
--- .loc Slicing ('t1':'t3') ---
Item Cost
t1 Samosa 35
t2 Bus 50
t3 Chai 15
--- .iloc Slicing (0:2) ---
Item Cost
t1 Samosa 35
t2 Bus 50
क्यों? .loc पूरी तरह inclusive है: 't1':'t3' तीनों rows ('t1', 't2', और 't3') निकालता है। इसके उलट, .iloc standard Python exclusion नियमों का पालन करता है: 0:2 rows 0 और 1 fetch करता है, index 2 से ठीक पहले रुकता है। यह दूसरे output block से 't3' को पूरी तरह छोड़ देता है!
Quiz
अगर एक DataFrame `df` के स्पष्ट index labels `['A', 'B', 'C', 'D']` हैं, statement `df.loc['A':'C']` execute करने से कौन से elements return होते हैं?
- केवल labels 'A' और 'B' से मेल खाती rows।
- labels 'A', 'B', और 'C' से मेल खाती rows।
- एक तुरंत ValueError क्योंकि string characters को logically slice नहीं किया जा सकता।
- index positions 0, 1, 2, और 3 से मेल खाती rows।
Show the answer
labels 'A', 'B', और 'C' से मेल खाती rows।
standard Python lists या .iloc indexing के उलट, Pandas .loc slicing endpoint का inclusive है। 'A' से 'C' तक slice करना rows 'A', 'B', और 'C' को पूरी तरह पकड़ता है।
Quiz
जब एक DataFrame के row indices को custom out-of-order numbers में बदला जाता है, जैसे index=[3, 1, 4, 2], और आप df.iloc[0] बनाम df.loc[0] execute करते हैं, कौन सा classic जाल सामने आता है?
- दोनों commands label 0 से मेल खाता एक जैसा row element return करती हैं।
- `df.iloc[0]` memory में असल पहली horizontal row (label 3) पकड़ता है, जबकि `df.loc[0]` एक KeyError के साथ crash होता है क्योंकि label 0 index map में मौजूद नहीं।
- दोनों commands एक तुरंत TypeError फेंकती हैं।
- Data selections DataFrame को memory space से पूरी तरह साफ़ कर देती हैं।
Show the answer
`df.iloc[0]` memory में असल पहली horizontal row (label 3) पकड़ता है, जबकि `df.loc[0]` एक KeyError के साथ crash होता है क्योंकि label 0 index map में मौजूद नहीं।
यह classic index जाल है! .iloc[0] position 0 खोजता है, screen पर बिल्कुल पहली line, उसके label की परवाह किए बिना। .loc[0] एक ऐसी row खोजता है जिसका official name label स्पष्ट रूप से integer 0 है। अगर वह label मौजूद नहीं, यह एक KeyError के साथ crash होता है।
Watch out
Classic जाल: Comma-Separated Subsetting की कमी
university lab exams में एक आम ग़लती single comma brackets df.loc['t1', 'Cost'] इस्तेमाल करने के बजाय df.loc['t1']['Cost'] लिखना है। जबकि double bracket shortcut अक्सर काम करता है, यह एक 'Chained Index' memory reference बनाता है। यह Pandas को पर्दे के पीछे temporary copy arrays बनाने पर मजबूर करता है, execution धीमा करते हुए और values update करने की कोशिश करते समय एक warning फेंकते हुए!
Theory
Subsetting को Semester 3 से जोड़ना
Precision data extraction advanced feature engineering pipelines का आधार बनाता है। Semester 3 Cloud Computing Architecture (BCA303) और Machine Learning models में, आप model training से पहले बड़े production tables को अलग feature matrices और target labels में बाँटने के लिए .loc conditional filters और .iloc coordinate slices इस्तेमाल करेंगे।
Summary
Key takeaways
- .loc property custom index name labels इस्तेमाल करके precision lookups सँभालती है।
- .iloc property 0 से N-1 तक numeric position index offsets इस्तेमाल करके lookups सँभालती है।
- .loc के ज़रिए slicing start और stop दोनों boundaries का पूरी तरह inclusive है।
- .iloc के ज़रिए slicing stop boundary का exclusive है, standard Python आदतों से मेल खाते हुए।
- df[row][col] जैसे chained selection syntax से साफ़ comma coordinates के पक्ष में बचना चाहिए।
- Memory Hook: .loc Label पर निर्भर करता है; .iloc Integer position map करता है!