Importing data into R from different file formats (CSV, Excel, etc.)

read.csv() एक CSV को एक line में एक data frame में खींचता है (header handle होता है, types guess होते हैं), Excel files को readxl package के read_excel() की ज़रूरत होती है, और str() import के बाद का health check है।

9 min read · 9 cards · 2 checks

Read in: English · हिन्दी · ગુજરાતી


Theory

Survey फिर आती है

CampusPulse data survey.csv में बैठा है: 60 rows, चार columns। आपने यह exact shape पहले दो बार import की है: sqlite3 shell के .import के साथ, और pandas के read_csv के साथ।

R का version एक line है:

survey <- read.csv("survey.csv")

कोई import statement नहीं, कोई module नहीं: data पढ़ना language में ही built in है। इस lesson की skill line नहीं है: यह जानना है कि R ने इसके पीछे क्या किया, और कैसे check करें।

Theory

Customs officer

read.csv() R की सीमा पर एक customs officer है। जैसे file अंदर आती है, officer:

1. पहली line को labels के रूप में पढ़ता है (header = TRUE, default)।

2. इसके contents से हर column का type guess करता है: numbers numeric बन जाते हैं, text character बन जाता है।

3. shipment को एक data frame में stamp करता है: R की table।

अच्छे officers भी अभी audit होते हैं: हर import के बाद, आप stamped result inspect करते हैं। Inspection command str() है।

Practical

Import कीजिए, फिर audit कीजिए

# CSV: built into base R, one line
survey <- read.csv("survey.csv")     # header = TRUE is the default

# THE health check after every import:
str(survey)
# 'data.frame': 60 obs. of 4 variables:
#  $ roll   : int  101 102 103 ...
#  $ city   : chr  "Surat" "Navsari" ...
#  $ stay   : chr  "hostel" "day" ...
#  $ screen : num  120 90 240 ...

head(survey)        # first 6 rows, eyeball the values

# Excel needs a package: INSTALL once, LOAD every session
# install.packages("readxl")     # once per machine
library(readxl)                   # once per session
marks <- read_excel("marks.xlsx", sheet = 1)

# other delimiters: read.delim (tabs), read.csv2 (; with , decimals)

Theory

हर हिस्से का क्या मतलब है

header = TRUE (default): line 1 column names बन जाती है। header = FALSE के साथ, R नाम V1, V2, V3 बना लेता है और आपके labels एक data row बन जाते हैं: वही header decision जो हर import tool में आपने देखा।

Type guessing: digits का एक column int/num के रूप में आता है; text वाला कुछ भी character के रूप में आता है। (Modern R अब text को automatically factor में convert नहीं करता: आप जानबूझकर promote करते हैं, अगले lessons में।)

Excel: base R .xlsx नहीं पढ़ता। readxl package पढ़ता है: install.packages() प्रति machine एक बार, library() प्रति session एक बार: दो अलग verbs, हमेशा confuse होते हुए।

Quiz

एक student ने कल install.packages("readxl") चलाया। आज, एक नए RStudio session में, read_excel("marks.xlsx") error देता है: could not find function "read_excel"। क्यों?

  1. Installing package को disk पर एक बार रखता है; हर नए session को अभी भी इसे library(readxl) से load करना होगा
  2. Package को हर दिन दोबारा install करना होगा
  3. read_excel सिर्फ़ CSV files पर काम करता है
  4. File name को .csv extension चाहिए
Show the answer

Installing package को disk पर एक बार रखता है; हर नए session को अभी भी इसे library(readxl) से load करना होगा

दो verbs, दो lifetimes: install.packages() disk पर download करता है (प्रति machine एक बार); library() चल रहे session में load करता है (प्रति session एक बार)। एक नया session सिर्फ़ base R loaded के साथ शुरू होता है, तो library(readxl) चलने तक read_excel अनजान है। यह install-बनाम-load distinction R का सबसे ज़्यादा पूछा गया package सवाल है, labs और exams दोनों में।

Think first

Audit intruder पकड़ता है

str(survey) दिखाता है: $ screen: chr "120" "90" "absent" "240"। screen-time column CHARACTER के रूप में आया। tap करने से पहले: इसका कारण क्या था, यह क्यों मायने रखता है, और यह BCA303 के किस moment का twin है?

Show the answer

एक student की entry कहती है "absent": column में एक अकेली string, और R के homogeneous-column rule ने हर number को text में coerce कर दिया (syntax lesson का coercion trap, अब file scale पर)। mean(survey$screen) अब impossible है।

यह pandas/csv-module में numbers-as-strings और BCA303 के SQLite affinity surprise का twin है: text formats कोई type नहीं ले जाते, तो importers guess करते हैं, और एक गंदा value एक column ज़हरीला कर देता है। मरम्मत (NA handling, as.numeric) बिल्कुल वहीं है जहाँ अगली unit जाती है।

Watch out

Import checklist

1. File not found? पहले working directory: getwd(), फिर setwd() या पूरे path से fix कीजिए।

2. हर import के बाद str(): dimensions सही? Types सही?

3. Numeric column chr दिखता है: किसी statistics से पहले stray text value ढूँढिए।

4. V1, V2 column names: आपका header = FALSE (या file में सच में कोई header नहीं है)।

चार checks, दस सेकंड, और आगे का हर lesson trusted data पर काम करता है।

Theory

वही गाना, तीसरी language

अपनी fluency notice कीजिए: header row, type inference, working directory, verify-the-import: आपने ये ideas SQLite shell में सीखे, pandas में practice किए, और आज ये मिनटों में R में transfer हुए। Tool skills expire होती हैं; concepts compound होते हैं। अगला lesson: data frame ख़ुद: वह DataFrame जिसमें आप पहले से सोचते हैं उसका R का जवाब।

Summary

Key takeaways

  • read.csv("file.csv") एक CSV को data frame के रूप में import करता है; header = TRUE और comma separator defaults हैं।
  • Excel को readxl चाहिए: install.packages() प्रति machine एक बार, library() प्रति SESSION एक बार।
  • R column types guess करता है; एक stray string एक numeric column को character बना देती है।
  • str() अनिवार्य post-import audit है: dimensions + प्रति-column types; head() rows पर नज़र डालता है।
  • File-not-found लगभग हमेशा working directory का मतलब है, missing file का नहीं।
  • Tabs के लिए read.delim, semicolon files के लिए read.csv2; SPSS/JSON के लिए packages मौजूद हैं।
  • Memory hook: customs officer इसे stamp करता है; आप stamp audit करते हैं।

Study this properly

This page is the lesson to read. In Gri-Learn the same topic is a graded deck: the self-checks are scored and your weak topics are tracked. Free to start.

Start this topic

Already have an account? Sign in

More from Introduction to R and working with Data

Gri-Learn · syllabus-mapped B.C.A. lessons in English, Hindi and Gujarati

Importing data into R from different file formats (CSV, Excel, etc.) · Statistical Methods and Data Analysis (MDC-03) · Gri-Learn