Important classes and functions of the CSV module: open(), reader(), writer(), writerows(), DictReader(), DictWriter()

csv module एक खुली file को एक reader (lists के रूप में rows) या writer (writerow/writerows) में wrap करता है, और Dict जोड़ी positions को column names से बदल देती है, newline='' एक setup rule के रूप में।

10 min read · 10 cards · 2 checks

Read in: English · हिन्दी · ગુજરાતી


Theory

वह comma जो split() तोड़ती है

पिछले lesson के open() से लैस होकर, आप marks.csv ख़ुद parse करते हैं: line.split(',')। ख़ूबसूरती से काम करता है... जब तक row 40 नहीं आता:

104,"Patel, Riya",78

नाम में एक comma है। आपका split() तीन के बजाय चार pieces produce करता है, और score column अब " Riya" रखता है।

Quoted fields, embedded commas, stray newlines: CSV format के edge cases हैं, और hand-rolled parsers इन सब पर मरते हैं। Python का csv module इसलिए मौजूद है ताकि आप फिर कभी split(',') न लिखें।

Theory

Tabular mail के लिए एक trained postman

split() से CSV पढ़ना envelopes पकड़ना और हर comma पर फाड़ना है, यहाँ तक कि letter के अंदर वाले commas भी।

csv module एक trained postman है: csv.reader हर envelope सही तरीक़े से खोलता है (quotes respect करते हुए) और contents आपको एक साफ़ list के रूप में देता है। csv.writer वही postman है जो आपकी lists को वापस valid envelopes में seal करता है। Dict जोड़ी और आगे जाती है: envelopes column name से labelled, position से नहीं।

Theory

reader और writer, formally

दोनों पहले से खुली file को wrap करते हैं:

Reading: csv.reader(f) हर row को एक strings की list के रूप में देता है: ['101', 'DBMS', '78']। नोटिस कीजिए: हर value एक string है, numbers भी: maths से पहले int()/float() से convert कीजिए। Header line एक साधारण पहली row के रूप में आती है: इसे next(reader) से skip कीजिए।

Writing: csv.writer(f) writerow(one_list) और writerows(list_of_lists) देता है।

एक format-specific rule: csv के लिए files `newline=''` के साथ खोलिए, वरना Windows users को हर row के बाद एक blank line मिलती है।

Practical

marks.csv पर चार workers

import csv

# READER: rows as lists (strings!)
with open('marks.csv', 'r', newline='') as f:
    reader = csv.reader(f)
    header = next(reader)            # skip the header row
    for row in reader:
        print(row[0], int(row[2]))   # roll, score as a NUMBER

# WRITER: lists back to disk
with open('toppers.csv', 'w', newline='') as f:
    writer = csv.writer(f)
    writer.writerow(['roll', 'subject', 'score'])          # one row
    writer.writerows([[101, 'DBMS', 88], [103, 'DBMS', 91]])  # many

# DICTREADER: rows as dictionaries, keys from the header
with open('marks.csv', newline='') as f:
    for row in csv.DictReader(f):
        print(row['roll'], row['score'])   # by NAME, not position

# DICTWRITER: dictionaries to disk (fieldnames required)
with open('report.csv', 'w', newline='') as f:
    w = csv.DictWriter(f, fieldnames=['roll', 'score'])
    w.writeheader()                        # do not forget this line
    w.writerows([{'roll': 101, 'score': 88}])

This example runs in Gri-Learn on the web, where you can edit it and see the output.

At a glance

एक नज़र में csv module

WorkerRow कैसी दिखती हैयाद रखिए
csv.reader['101', 'DBMS', '78'] (list)सभी strings; next() header skip करता है
csv.writerwriterow / writerowsnewline='' के साथ file खोलिए
csv.DictReader{'roll': '101', 'score': '78'}Header keys बन जाता है
csv.DictWriterdicts अंदर, fieldnames चाहिएपहले writeheader() call कीजिए

Think first

Maths क्यों crash हुई?

एक student एक total compute करता है: for row in csv.reader(f): total = total + row[2]। Python TypeError: unsupported operand type(s) उठाता है। CSV बिल्कुल valid है। tap करने से पहले: row[2] असल में क्या है, और one-word fix क्या है?

Show the answer

row[2] string '78' है, number 78 नहीं: csv.reader हर field text के रूप में return करता है, क्योंकि CSV files कोई type information नहीं ले जातीं। एक string को एक number में जोड़ना ही TypeError है।

Fix: int(row[2]) (या decimals के लिए float)। यह convert-before-maths आदत वही lesson है जो SQLite के .import ने Unit 2 में सिखाया: text formats text देते हैं; numbers आपकी responsibility हैं।

Quiz

DictReader plain reader से कब स्पष्ट रूप से बेहतर है?

  1. जब CSV का column order बदल सकता हो: row['score'] अभी भी काम करता है, row[2] चुपचाप टूट जाता है
  2. जब file बहुत बड़ी हो: DictReader तेज़ है
  3. जब values numbers के रूप में आनी चाहिए, strings के बजाय
  4. जब file में कोई header row न हो
Show the answer

जब CSV का column order बदल सकता हो: row['score'] अभी भी काम करता है, row[2] चुपचाप टूट जाता है

DictReader हर row को header names से key करता है, तो एक reordered या extended CSV (faculty ने एक column जोड़ा) row['score'] को सही छोड़ता है जबकि positional row[2] बिना किसी error के ग़लत field पढ़ना शुरू कर देता है: सबसे बुरी तरह का bug। यह तेज़ नहीं है (option B), values अभी भी strings हैं (option C), और बिना header row के DictReader बिल्कुल ग़लत tool है (option D): यह पहली data row को keys के रूप में खा लेगा।

Watch out

तीन csv-module slips

Writing करते समय newline='' भूलना: rows के बीच blank lines (classic Windows symptom, और एक exam favourite "इस code में क्या ग़लत है")।

DictWriter के साथ writeheader() skip करना: बिना header वाली file जो downstream हर DictReader तोड़ती है।

Header row को data के रूप में process करना plain reader के साथ: 'roll' int() से fail होता है: पहले इसे next(reader) से skip कीजिए।

Theory

एक format, अब तीन tools

अब आपके पास तीन CSV instruments हैं: sqlite3 shell का .import/.export (Unit 2), Python का csv module (row-by-row control, यह lesson), और अगला lesson pandas का read_csv (पूरी file एक object के रूप में, एक line)। इनके बीच चुनना असली engineering judgement है: one-off loads के लिए shell, streaming row logic के लिए csv module, analysis के लिए pandas। Exams Python में CSV पढ़ने के दो तरीक़े पूछना पसंद करते हैं: अब आपके पास दोनों हैं।

Summary

Key takeaways

  • csv.reader rows को STRINGS की lists के रूप में देता है; numbers को int()/float() से convert कीजिए; header को next() से skip कीजिए।
  • csv.writer: writerow(list) और writerows(list of lists)।
  • DictReader rows को header names से key करता है: column reordering से बचता है; DictWriter को fieldnames + writeheader() चाहिए।
  • csv files को newline='' के साथ खोलिए (ख़ासकर writing) blank-line rows से बचने के लिए।
  • Reader बनाम DictReader = position-based बनाम name-based access।
  • Memory hook: एक trained postman, envelopes position से या label से।

Study this properly

This page is the lesson to read. In Gri-Learn the same topic is a graded deck: the self-checks are scored and your weak topics are tracked. Free to start.

Start this topic

Already have an account? Sign in

More from Python interaction with text and CSV

Gri-Learn · syllabus-mapped B.C.A. lessons in English, Hindi and Gujarati