Theory
આંધળા Slicing નો કોયડો
તમારા Semester 1 Python list labs માં, તમે matrix[0][1] જેવા સાદા integer brackets વાપરીને multi-dimensional matrices slice કરતા. પણ Pandas માં જટિલ relational data tables process કરતી વખતે, જ્યાં rows student roll numbers કે date strings જેવા સ્પષ્ટ custom labels થી track થાય છે, ત્યાં સંપૂર્ણપણે hardcoded integer positions પર આધાર રાખવો ખતરનાક બની જાય છે. જો એક dataset ને shuffle, sort, કે filter કરવામાં આવે, તો row position 3 તદ્દન બદલાઈ જાય છે! તો આપણે Python ને કેવી રીતે કહીએ કે એક ચોક્કસ row એના સત્તાવાર નામના badge પ્રમાણે કાઢે કે એની ખરેખરી ઊભી લીટીની જગ્યા પ્રમાણે, અને આ બંને વચ્ચે ગૂંચવાડો તમારી production analytical pipelines કેમ તોડી શકે છે?
Theory
નામનો પોકાર સામે Seat-Number પર ટકોરો
તમારા BCA classroom register ને કલ્પો. દરેક student પાસે એક અનન્ય સંરચનાત્મક ઓળખ છે, એમનો Roll Number (Label). એ આજે એક ચોક્કસ ભૌતિક ખુરશીના ક્રમમાં પણ બેસે છે, આગળથી પાછળ (Integer Position: 0, 1, 2). .loc વાપરવું એટલે classroom ના દરવાજે ઊભા રહીને પોકારવું: 'મને Roll Number BCA102 ના records લાવી આપો.' એ student ક્યાં બેઠો છે એનાથી ફરક પડતો નથી; label એને બરાબર પકડી લે છે. .iloc વાપરવું એટલે આંખો બંધ કરીને classroom માં જવું, 3જી ભૌતિક બેન્ચ સુધી ગણવું, અને અત્યારે એ ખુરશીમાં જે કોઈ બેઠું હોય એને પકડી લેવું, એમના નામ કે roll number ને ધ્યાનમાં લીધા વગર.
Theory
Label-આધારિત સામે Integer-આધારિત Indexing ઔપચારિક રીતે
Pandas data selection ને બે સ્પષ્ટ પદ્ધતિઓમાં વહેંચે છે: .loc અને .iloc. .loc property મુખ્યત્વે label-આધારિત છે; એ rows અને columns ને એમના string અક્ષરો કે સ્પષ્ટ index values વાપરીને query કરે છે. એથી ઊલટું, .iloc સખ્તાઈથી integer position-આધારિત છે; એ zero-indexed integer offsets ($0$ થી $N-1$) વાપરીને elements પસંદ કરે છે. એક નિર્ણાયક mathematical ભેદ એમની slicing સીમાઓમાં રહેલો છે: standard Python અને .iloc stop index સીમા છોડી દે છે (start:stop stop-1 પર અટકે છે), જ્યારે .loc start અને stop બંને સીમાઓનો સંપૂર્ણપણે સમાવેશ કરે છે.
At a glance
Table 1: Pandas .loc અને .iloc extraction properties નું ઊંડું functional mapping.
| Indexing Property | Target Query Mode | Slicing સીમાનો નિયમ | Syntax ઉદાહરણ |
|---|---|---|---|
| df.loc[row_lbl, col_lbl] | સ્પષ્ટ text/numeric નામો દ્વારા Label-આધારિત મેળ. | સમાવેશક: start અને stop બંને સીમાઓ રખાય છે. | df.loc['t2', 'Amount'] |
| df.iloc[row_pos, col_pos] | Numeric offsets દ્વારા Integer position-આધારિત મેળ. | બાકાત: stop સીમા છોડી દેવાય છે (standard Python). | df.iloc[1, 1] |
| df['ColumnName'] | સીધો bracket selection શોર્ટકટ. | Rows ને સંપૂર્ણપણે અવગણે છે; સીધા અલગ columns ખેંચે છે. | df['Amount'] |
Theory
Worked Example: PocketMoney Ledger Grid નું વિચ્છેદન
ચાલો એક customized pocket-money ledger DataFrame પર ચોકસાઈવાળું extraction લાગુ કરીએ. આપણે index નામોને ભૌતિક positional indices થી સ્પષ્ટ રીતે અલગ પાડવા custom row strings (t1, t2, t3) બનાવીશું, slicing operations દરમિયાનના સીમાના ભેદ ઉજાગર કરતાં.
Practical
Precision Coordinate Extractor
import pandas as pd
# Step 1: Create a structured ledger with custom string index labels
raw_data = {
"Item": ["Samosa", "Bus", "Chai", "Notebook"],
"Cost": [35, 50, 15, 120]
}
ledger_df = pd.DataFrame(raw_data, index=["t1", "t2", "t3", "t4"])
# Step 2: Use label-based .loc to pull a specific coordinate value
bus_cost_loc = ledger_df.loc["t2", "Cost"]
# Step 3: Use integer position-based .iloc to pull the exact same cell
bus_cost_iloc = ledger_df.iloc[1, 1]
# Step 4: Compare label slicing vs position slicing boundaries
slice_loc = ledger_df.loc["t1":"t3"]
slice_iloc = ledger_df.iloc[0:2]
print(f"Extracted via .loc label: {bus_cost_loc}")
print(f"Extracted via .iloc integer position: {bus_cost_iloc}")
print("\n--- .loc Slicing ('t1':'t3') ---")
print(slice_loc)
print("\n--- .iloc Slicing (0:2) ---")
print(slice_iloc)This example runs in Gri-Learn on the web, where you can edit it and see the output.
Think first
સીમાના Outputs નું વિશ્લેષણ કરો
slicing લીટીઓ ધ્યાનથી જુઓ: slice_loc 't1':'t3' વાપરે છે, જ્યારે slice_iloc 0:2 વાપરે છે. દરેક statement કેટલી rows output કરશે, અને એ કયા item records સામેલ કરશે?
Show the answer
script output કરશે:
Extracted via .loc label: 50
Extracted via .iloc integer position: 50
--- .loc Slicing ('t1':'t3') ---
Item Cost
t1 Samosa 35
t2 Bus 50
t3 Chai 15
--- .iloc Slicing (0:2) ---
Item Cost
t1 Samosa 35
t2 Bus 50
કેમ? .loc સંપૂર્ણપણે સમાવેશક છે: 't1':'t3' ત્રણેય rows ('t1', 't2', અને 't3') કાઢે છે. એથી ઊલટું, .iloc standard Python ના બાકાત નિયમો પાળે છે: 0:2 rows 0 અને 1 લાવે છે, index 2 ની બરાબર પહેલાં અટકતાં. આ બીજા output block માંથી 't3' ને સંપૂર્ણપણે કાઢી નાખે છે!
Quiz
જો એક DataFrame `df` પાસે સ્પષ્ટ index labels `['A', 'B', 'C', 'D']` હોય, તો statement `df.loc['A':'C']` ચલાવવાથી કયા elements પાછા મળે છે?
- માત્ર labels 'A' અને 'B' સાથે મળતી rows.
- labels 'A', 'B', અને 'C' સાથે મળતી rows.
- એક તરત ValueError કારણ કે string અક્ષરો તાર્કિક રીતે slice થઈ શકતા નથી.
- index positions 0, 1, 2, અને 3 સાથે મળતી rows.
Show the answer
labels 'A', 'B', અને 'C' સાથે મળતી rows.
Standard Python lists કે .iloc indexing થી વિપરીત, Pandas .loc slicing અંતિમ બિંદુનો સમાવેશ કરે છે. 'A' થી 'C' સુધી slice કરવાથી rows 'A', 'B', અને 'C' સંપૂર્ણપણે પકડાય છે.
Quiz
જ્યારે એક DataFrame ના row indices ને index=[3, 1, 4, 2] જેવા custom ક્રમ વગરના numbers માં બદલવામાં આવે, અને તમે df.iloc[0] સામે df.loc[0] ચલાવો, ત્યારે કયો classic ફાંદો સામે આવે છે?
- બંને commands label 0 સાથે મળતું એકસરખું row element પાછું આપે છે.
- `df.iloc[0]` memory માંની ખરેખરી પહેલી આડી row (label 3) પકડે છે, જ્યારે `df.loc[0]` એક KeyError સાથે crash થાય છે કારણ કે index map માં label 0 અસ્તિત્વમાં નથી.
- બંને commands તરત એક TypeError ફેંકે છે.
- Data selections DataFrame ને memory space માંથી સંપૂર્ણપણે સાફ કરી નાખે છે.
Show the answer
`df.iloc[0]` memory માંની ખરેખરી પહેલી આડી row (label 3) પકડે છે, જ્યારે `df.loc[0]` એક KeyError સાથે crash થાય છે કારણ કે index map માં label 0 અસ્તિત્વમાં નથી.
આ classic index ફાંદો છે! .iloc[0] position 0 શોધે છે, screen પરની સૌથી પહેલી લીટી, એના label ને ધ્યાનમાં લીધા વગર. .loc[0] એવી row શોધે છે જેનું સત્તાવાર નામ label સ્પષ્ટપણે integer 0 હોય. જો એ label અસ્તિત્વમાં ન હોય, એ એક KeyError સાથે crash થાય છે.
Watch out
Classic ફાંદો: Comma-Separated Subsetting ની ખોટ
university lab exams માં એક સામાન્ય ભૂલ એ છે એક comma વાળા brackets df.loc['t1', 'Cost'] વાપરવાને બદલે df.loc['t1']['Cost'] લખવું. જોકે બેવડા bracket નો શોર્ટકટ ઘણી વાર કામ કરે છે, એ એક 'Chained Index' memory reference બનાવે છે. આ Pandas ને પડદા પાછળ કામચલાઉ copy arrays બનાવવા મજબૂર કરે છે, execution ધીમું કરતાં અને values update કરવાનો પ્રયાસ કરતી વખતે એક warning ફેંકતાં!
Theory
Subsetting ને Semester 3 સાથે જોડવા
ચોકસાઈવાળું data extraction અદ્યતન feature engineering pipelines નો પાયો બનાવે છે. Semester 3 Cloud Computing Architecture (BCA303) અને Machine Learning models માં, તમે model training પહેલાં મોટા production tables ને અલગ feature matrices અને target labels માં વહેંચવા .loc conditional filters અને .iloc coordinate slices વાપરશો.
Summary
Key takeaways
- .loc property custom index name labels વાપરીને ચોકસાઈવાળા lookups સંભાળે છે.
- .iloc property 0 થી N-1 સુધીના numeric position index offsets વાપરીને lookups સંભાળે છે.
- .loc દ્વારા slicing start અને stop બંને સીમાઓનો સંપૂર્ણપણે સમાવેશ કરે છે.
- .iloc દ્વારા slicing stop સીમાને બાકાત રાખે છે, standard Python ની ટેવો સાથે મળતાં.
- df[row][col] જેવી chained selection syntax ટાળીને સાફ comma coordinates પસંદ કરવા જોઈએ.
- Memory Hook: .loc Label પર આધાર રાખે છે; .iloc Integer position map કરે છે!