Theory
साठ हज़ार rows, पाँच काम की झलकियाँ
ResultDesk अब university-wide CSV load करता है: साठ हज़ार rows। print(df) किसी film के credits की तरह scroll होता चला जाता है।
Analysts कभी पूरी tables नहीं पढ़ते। वे झलक लेते हैं: पहली rows (क्या सही load हुआ?), आख़िरी rows (क्या सही ख़त्म हुआ?), एक statistical summary (यह दिखता कैसा है?), और फिर वे exact cells की तरफ़ बढ़ते हैं।
पाँच functions यह सब cover करते हैं, और इनमें से दो, loc और iloc, एक distinction रखते हैं जिसे exams हर साल इस्तेमाल करते हैं।
Theory
Flat number बनाम बाएँ से तीसरा दरवाज़ा
एक building में एक घर ढूँढने के दो तरीक़े:
- Label से: "Flat 203": दरवाज़े पर लिखा number। यह loc है: यह index labels और column names पढ़ता है।
- Position से: "दूसरी मंज़िल, तीसरा दरवाज़ा": entrance से गिनते हुए। यह iloc (integer-location) है: यह 0, 1, 2 गिनता है चाहे दरवाज़ों पर कुछ भी लिखा हो।
एक ही building, दो addressing systems, और वे तभी agree करते हैं जब दरवाज़े संयोग से क्रम में numbered हों।
Theory
Glancing toolkit
- df.head(): पहली 5 rows (दस के लिए head(10)): क्या सही load हुआ यह check।
- df.tail(): आख़िरी 5: टूटी हुई final lines और import leftovers पकड़ता है।
- df.describe(): हर numeric column के लिए count, mean, std, min, quartiles (25%, 50%, 75%), max: पिछले lesson की statistics, एक call में।
- df.to_numpy(): नीचे का raw numpy array, labels उतारे हुए (पुरानी spelling
.valuesअभी भी किताबों में दिखती है: वही idea)।
चारों read-only झलकियाँ हैं: इनमें से कोई भी DataFrame को नहीं बदलता।
Practical
झलकियाँ, फिर exact addresses
import pandas as pd
df = pd.read_csv('marks.csv') # index: 0, 1, 2, ...
print(df.head()) # first 5 rows
print(df.tail(3)) # last 3
print(df.describe()) # the statistical X-ray
print(df.to_numpy()) # raw array, labels gone (aka .values)
# loc: BY LABEL (index label, column NAME)
print(df.loc[0, 'score']) # row labelled 0, column 'score'
print(df.loc[0:3, ['roll', 'score']]) # labels 0..3 INCLUSIVE: 4 rows!
# iloc: BY POSITION (counting from 0)
print(df.iloc[0, 2]) # first row, third column
print(df.iloc[0:3]) # positions 0,1,2: 3 rows (end EXCLUDED)
# loc happily takes last lesson's masks too:
print(df.loc[df['score'] > 60, ['roll', 'score']])
This example runs in Gri-Learn on the web, where you can edit it and see the output.
At a glance
loc बनाम iloc, exam table
| Aspect | loc | iloc |
|---|---|---|
| Select करता है | Label से (index value, column name) | Integer position से (0, 1, 2...) |
| Slice 0:3 देता है | 4 rows (दोनों ends INCLUDED) | 3 rows (end EXCLUDED) |
| Column form | df.loc[0, 'score'] | df.iloc[0, 2] |
| Boolean masks | हाँ | नहीं (सिर्फ़ positions) |
Think first
Four-versus-three trap
df का default index 0,1,2,...,9 है। tap करने से पहले दोनों का जवाब दीजिए: df.loc[0:3] कितनी rows return करता है, df.iloc[0:3] कितनी return करता है, और वे अलग क्यों हैं?
Show the answer
loc[0:3] → 4 rows (labels 0, 1, 2, 3): label slices दोनों ends INCLUDE करती हैं, क्योंकि arbitrary labels ('Riya':'Aman'?) के साथ pandas "one past the end" compute नहीं कर सकता।
iloc[0:3] → 3 rows (positions 0, 1, 2): integer slices Python list rules follow करती हैं, end excluded।
दिखने में एक जैसी expression, अलग row counts: यही THE loc/iloc exam question है, और अब आप इसे reason के साथ जवाब दे सकते हैं, सिर्फ़ numbers के साथ नहीं।
Quiz
Filtering के बाद, top = df[df['score'] > 80], बची हुई index labels 7, 23, 41 हैं। कौन सा call reliably पहली बची row return करता है?
- top.iloc[0]: position zero हमेशा exist करता है; top.loc[0] KeyError देता है क्योंकि कोई row LABEL 0 वाली नहीं है
- top.loc[0]: loc हमेशा पहली row return करता है
- top.head(): filtering के बाद head इस्तेमाल नहीं किया जा सकता
- दोनों काम करते हैं: filtered frames पर loc और iloc interchangeable हैं
Show the answer
top.iloc[0]: position zero हमेशा exist करता है; top.loc[0] KeyError देता है क्योंकि कोई row LABEL 0 वाली नहीं है
Filtering ORIGINAL labels रखता है (7, 23, 41), तो अब कोई label 0 नहीं है: loc[0] एक door number 0 ढूँढता है और fail होता है। iloc[0] physically गिनता है: पहली row, उसका label चाहे कुछ भी हो। यह divergence filtering के बाद वहीं है जहाँ दोनों systems असहमत होना शुरू करते हैं, और option A का reasoning professional habit है: 'first/last' के लिए positions, 'row known as X' के लिए labels। (head(1) भी काम करता: option C का claim ग़लत है।)
Watch out
तीन daily-driver slips
positions के साथ loc: तब तक luck से काम करता है जब तक index 0,1,2... है, किसी भी filter या sort के बाद टूट जाता है: TYPE करने से पहले label-या-position तय कीजिए।
describe() उम्मीद से कम columns दिखाता है: यह default से सिर्फ़ numeric columns summarise करता है; text columns के लिए describe(include='object') चाहिए।
पुरानी किताबों में .values: to_numpy() जैसा ही array; दोनों नाम पहचानिए, modern वाला लिखिए।
Theory
Unit 4 बंद होती है; तस्वीरें शुरू होती हैं
अपनी pandas kit गिनिए: load और save (read_csv/to_csv), slice (brackets और masks), summarise (mean से describe तक), और exactly address करना (loc/iloc)। यही एक working analyst का core है। Unit 5 इन numbers को तस्वीरों में बदलता है: matplotlib वही Series लेता है जो आप extract करते आए हैं और class की story खींचता है: lines, bars, scatters, histograms।
Summary
Key takeaways
- head(n)/tail(n): सिरों पर झलक; describe(): numeric columns का count-से-max तक statistical summary।
- to_numpy() (पहले .values) DataFrame को इसके raw numpy array तक strip करता है।
- loc LABEL से select करता है: slices दोनों ends include करती हैं; boolean masks accept करता है।
- iloc POSITION से select करता है: slices end exclude करती हैं, Python lists जैसा।
- Default index पर loc[0:3] = 4 rows, iloc[0:3] = 3 rows: classic exam pair।
- Filtering के बाद, labels में gaps रह जाते हैं: 'पहली row' के लिए iloc[0], 'row labelled X' के लिए loc।
- Memory hook: flat number बनाम बाएँ से तीसरा दरवाज़ा।