Theory
Sixty thousand rows, five useful glances
ResultDesk now loads the university-wide CSV: sixty thousand rows. print(df) scrolls past like credits in a film.
Analysts never read whole tables. They glance: the first rows (did it load right?), the last rows (did it end right?), one statistical summary (what does it look like?), and then they reach for exact cells.
Five functions cover all of it, and two of them, loc and iloc, carry a distinction that exams milk every single year.
Theory
Flat number vs third door on the left
Two ways to find a home in a building:
- By label: "Flat 203": the number on the door. That is loc: it reads the index labels and column names.
- By position: "second floor, third door": counting from the entrance. That is iloc (integer-location): it counts 0, 1, 2 regardless of what the doors say.
Same building, two addressing systems, and they only agree while doors happen to be numbered in order.
Theory
The glancing toolkit
- df.head(): first 5 rows (head(10) for ten): the did-it-load-right check.
- df.tail(): last 5: catches broken final lines and import leftovers.
- df.describe(): count, mean, std, min, quartiles (25%, 50%, 75%), max for every numeric column: last lesson's statistics, one call.
- df.to_numpy(): the raw numpy array underneath, labels stripped (the older spelling
.valuesstill appears in books: same idea).
All four are read-only glances: none of them changes the DataFrame.
Practical
Glances, then exact addresses
import pandas as pd
df = pd.read_csv('marks.csv') # index: 0, 1, 2, ...
print(df.head()) # first 5 rows
print(df.tail(3)) # last 3
print(df.describe()) # the statistical X-ray
print(df.to_numpy()) # raw array, labels gone (aka .values)
# loc: BY LABEL (index label, column NAME)
print(df.loc[0, 'score']) # row labelled 0, column 'score'
print(df.loc[0:3, ['roll', 'score']]) # labels 0..3 INCLUSIVE: 4 rows!
# iloc: BY POSITION (counting from 0)
print(df.iloc[0, 2]) # first row, third column
print(df.iloc[0:3]) # positions 0,1,2: 3 rows (end EXCLUDED)
# loc happily takes last lesson's masks too:
print(df.loc[df['score'] > 60, ['roll', 'score']])
This example runs in Gri-Learn on the web, where you can edit it and see the output.
At a glance
loc vs iloc, the exam table
| Aspect | loc | iloc |
|---|---|---|
| Selects by | Label (index value, column name) | Integer position (0, 1, 2...) |
| Slice 0:3 gives | 4 rows (both ends INCLUDED) | 3 rows (end EXCLUDED) |
| Column form | df.loc[0, 'score'] | df.iloc[0, 2] |
| Boolean masks | Yes | No (positions only) |
Think first
The four-versus-three trap
df has the default index 0,1,2,...,9. Before tapping, answer both: how many rows does df.loc[0:3] return, how many does df.iloc[0:3] return, and WHY do they differ?
Show the answer
loc[0:3] → 4 rows (labels 0, 1, 2, 3): label slices include BOTH ends, because with arbitrary labels ('Riya':'Aman'?) pandas cannot compute "one past the end".
iloc[0:3] → 3 rows (positions 0, 1, 2): integer slices follow Python list rules, end excluded.
Same-looking expression, different row counts: this is THE loc/iloc exam question, and now you can answer it with the reason, not just the numbers.
Quiz
After filtering, top = df[df['score'] > 80], the surviving index labels are 7, 23, 41. Which call reliably returns the FIRST surviving row?
- top.iloc[0]: position zero always exists; top.loc[0] raises KeyError since no row is LABELLED 0
- top.loc[0]: loc always returns the first row
- top.head(): head cannot be used after filtering
- Either works: loc and iloc are interchangeable on filtered frames
Show the answer
top.iloc[0]: position zero always exists; top.loc[0] raises KeyError since no row is LABELLED 0
Filtering keeps the ORIGINAL labels (7, 23, 41), so there is no label 0 anymore: loc[0] hunts a door numbered 0 and fails. iloc[0] counts physically: first row, whatever its label. This divergence after filtering is where the two systems stop agreeing, and option A's reasoning is the professional habit: positions for 'first/last', labels for 'the row known as X'. (head(1) would also work: option C's claim is false.)
Watch out
Three daily-driver slips
loc with positions: works by luck while the index is 0,1,2..., breaks after any filter or sort: decide label-or-position BEFORE typing.
describe() shows fewer columns than expected: it summarises numeric columns only by default; text columns need describe(include='object').
.values in old books: same array as to_numpy(); recognise both names, write the modern one.
Theory
Unit 4 closes; the pictures begin
Count your pandas kit: load and save (read_csv/to_csv), slice (brackets and masks), summarise (mean to describe), and address exactly (loc/iloc). That is a working analyst's core. Unit 5 turns these numbers into pictures: matplotlib takes the very Series you have been extracting and draws the class's story: lines, bars, scatters, histograms.
Summary
Key takeaways
- head(n)/tail(n): glance at the ends; describe(): count-to-max statistical summary of numeric columns.
- to_numpy() (formerly .values) strips the DataFrame to its raw numpy array.
- loc selects by LABEL: slices include BOTH ends; accepts boolean masks.
- iloc selects by POSITION: slices exclude the end, like Python lists.
- loc[0:3] = 4 rows, iloc[0:3] = 3 rows on a default index: the classic exam pair.
- After filtering, labels keep gaps: iloc[0] for 'first row', loc for 'row labelled X'.
- Memory hook: flat number vs third door on the left.