Theory
The comma that breaks split()
Armed with last lesson's open(), you parse marks.csv yourself: line.split(','). Works beautifully... until row 40:
104,"Patel, Riya",78
The name contains a comma. Your split() produces four pieces instead of three, and the score column now holds " Riya".
Quoted fields, embedded commas, stray newlines: the CSV format has edge cases, and hand-rolled parsers die on all of them. Python's csv module exists so you never write split(',') again.
Theory
A trained postman for tabular mail
Reading CSV by split() is grabbing envelopes and tearing them at every comma, even the commas inside the letter.
The csv module is a trained postman: csv.reader opens each envelope properly (quotes respected) and hands you the contents as a neat list. csv.writer is the same postman sealing your lists back into valid envelopes. The Dict pair goes further: envelopes labelled by column name instead of position.
Theory
reader and writer, formally
Both wrap an already-open file:
Reading: csv.reader(f) yields each row as a list of strings: ['101', 'DBMS', '78']. Note: every value is a string, even the numbers: convert with int()/float() before maths. The header line arrives as an ordinary first row: skip it with next(reader).
Writing: csv.writer(f) offers writerow(one_list) and writerows(list_of_lists).
One format-specific rule: open files for csv with `newline=''`, or Windows users get a blank line after every row.
Practical
The four workers on marks.csv
import csv
# READER: rows as lists (strings!)
with open('marks.csv', 'r', newline='') as f:
reader = csv.reader(f)
header = next(reader) # skip the header row
for row in reader:
print(row[0], int(row[2])) # roll, score as a NUMBER
# WRITER: lists back to disk
with open('toppers.csv', 'w', newline='') as f:
writer = csv.writer(f)
writer.writerow(['roll', 'subject', 'score']) # one row
writer.writerows([[101, 'DBMS', 88], [103, 'DBMS', 91]]) # many
# DICTREADER: rows as dictionaries, keys from the header
with open('marks.csv', newline='') as f:
for row in csv.DictReader(f):
print(row['roll'], row['score']) # by NAME, not position
# DICTWRITER: dictionaries to disk (fieldnames required)
with open('report.csv', 'w', newline='') as f:
w = csv.DictWriter(f, fieldnames=['roll', 'score'])
w.writeheader() # do not forget this line
w.writerows([{'roll': 101, 'score': 88}])
This example runs in Gri-Learn on the web, where you can edit it and see the output.
At a glance
The csv module at a glance
| Worker | Row looks like | Remember |
|---|---|---|
| csv.reader | ['101', 'DBMS', '78'] (list) | All strings; next() skips header |
| csv.writer | writerow / writerows | Open file with newline='' |
| csv.DictReader | {'roll': '101', 'score': '78'} | Header becomes the keys |
| csv.DictWriter | dicts in, needs fieldnames | Call writeheader() first |
Think first
Why did the maths crash?
A student computes a total: for row in csv.reader(f): total = total + row[2]. Python raises TypeError: unsupported operand type(s). The CSV is perfectly valid. Before tapping: what is row[2], actually, and what is the one-word fix?
Show the answer
row[2] is the string '78', not the number 78: csv.reader returns every field as text, because CSV files carry no type information. Adding a string to a number is the TypeError.
Fix: int(row[2]) (or float for decimals). This convert-before-maths habit is the same lesson SQLite's .import taught in Unit 2: text formats deliver text; numbers are your responsibility.
Quiz
When is DictReader clearly better than plain reader?
- When the CSV's column order might change: row['score'] still works, row[2] silently breaks
- When the file is very large: DictReader is faster
- When the values must arrive as numbers instead of strings
- When the file has no header row
Show the answer
When the CSV's column order might change: row['score'] still works, row[2] silently breaks
DictReader keys each row by the header names, so a reordered or extended CSV (faculty added a column) leaves row['score'] correct while positional row[2] starts reading the wrong field WITHOUT any error: the nastiest kind of bug. It is not faster (option B), values are still strings (option C), and with no header row DictReader is exactly the wrong tool (option D): it would eat the first data row as keys.
Watch out
The three csv-module slips
newline='' forgotten when writing: blank lines between rows (the classic Windows symptom, and an exam favourite "what is wrong with this code").
writeheader() skipped with DictWriter: a headerless file that breaks every DictReader downstream.
Header row processed as data with plain reader: 'roll' fails int(): skip it with next(reader) first.
Theory
One format, three tools now
You now hold three CSV instruments: the sqlite3 shell's .import/.export (Unit 2), Python's csv module (row-by-row control, this lesson), and next lesson pandas' read_csv (the whole file as one object, one line). Choosing among them is real engineering judgement: shell for one-off loads, csv module for streaming row logic, pandas for analysis. Exams love asking two ways to read a CSV in Python: you now have both.
Summary
Key takeaways
- csv.reader yields rows as lists of STRINGS; convert numbers with int()/float(); skip the header with next().
- csv.writer: writerow(list) and writerows(list of lists).
- DictReader keys rows by header names: survives column reordering; DictWriter needs fieldnames + writeheader().
- Open csv files with newline='' (writing especially) to avoid blank-line rows.
- Reader vs DictReader = position-based vs name-based access.
- Memory hook: a trained postman, envelopes by position or by label.