Theory
The survey arrives, again
The CampusPulse data sits in survey.csv: 60 rows, four columns. You have imported this exact shape twice before: with the sqlite3 shell's .import, and with pandas' read_csv.
R's version is one line:
survey <- read.csv("survey.csv")
No import statement, no module: reading data is built into the language. The skill of this lesson is not the line: it is knowing what R did behind it, and how to check.
Theory
The customs officer
read.csv() is a customs officer at R's border. As the file enters, the officer:
1. Reads the first line as labels (header = TRUE, the default).
2. Guesses each column's type from its contents: numbers become numeric, text becomes character.
3. Stamps the shipment into a data frame: R's table.
Good officers still get audited: after every import, YOU inspect the stamped result. The inspection command is str().
Practical
Import, then audit
# CSV: built into base R, one line
survey <- read.csv("survey.csv") # header = TRUE is the default
# THE health check after every import:
str(survey)
# 'data.frame': 60 obs. of 4 variables:
# $ roll : int 101 102 103 ...
# $ city : chr "Surat" "Navsari" ...
# $ stay : chr "hostel" "day" ...
# $ screen : num 120 90 240 ...
head(survey) # first 6 rows, eyeball the values
# Excel needs a package: INSTALL once, LOAD every session
# install.packages("readxl") # once per machine
library(readxl) # once per session
marks <- read_excel("marks.xlsx", sheet = 1)
# other delimiters: read.delim (tabs), read.csv2 (; with , decimals)
Theory
What each piece means
header = TRUE (default): line 1 becomes column names. With header = FALSE, R invents names V1, V2, V3 and your labels become a data row: the same header decision as every import tool you have met.
Type guessing: a column of digits arrives as int/num; anything with text arrives as character. (Modern R no longer auto-converts text to factors: you promote deliberately, next lessons.)
Excel: base R does not read .xlsx. The readxl package does: install.packages() once per machine, library() once per session: two different verbs, endlessly confused.
Quiz
A student ran install.packages("readxl") yesterday. Today, in a fresh RStudio session, read_excel("marks.xlsx") errors: could not find function "read_excel". Why?
- Installing puts the package on disk once; each new session must still load it with library(readxl)
- The package must be reinstalled every day
- read_excel only works on CSV files
- The file name needs .csv extension
Show the answer
Installing puts the package on disk once; each new session must still load it with library(readxl)
Two verbs, two lifetimes: install.packages() downloads to disk (once per machine); library() loads into the running session (once per session). A fresh session starts with only base R loaded, so read_excel is unknown until library(readxl) runs. This install-vs-load distinction is R's most-asked package question, in labs and exams alike.
Think first
Audit catches the intruder
str(survey) shows: $ screen: chr "120" "90" "absent" "240". The screen-time column arrived as CHARACTER. Before tapping: what caused it, why does it matter, and which BCA303 moment is this the twin of?
Show the answer
One student's entry says "absent": a single string in the column, and R's homogeneous-column rule coerced every number to text (the coercion trap from the syntax lesson, now at file scale). mean(survey$screen) is now impossible.
It is the twin of pandas/csv-module numbers-arriving-as-strings and SQLite's affinity surprise from BCA303: text formats carry no types, so importers guess, and one dirty value poisons a column. The repair (NA handling, as.numeric) is exactly where the next unit goes.
Watch out
The import checklist
1. File not found? Working directory first: getwd(), then fix with setwd() or a full path.
2. str() after every import: dimensions right? Types right?
3. Numeric column shown as chr: hunt the stray text value before any statistics.
4. V1, V2 column names: your header = FALSE (or the file truly has no header).
Four checks, ten seconds, and every downstream lesson works on trusted data.
Theory
Same song, third language
Notice your own fluency: header row, type inference, working directory, verify-the-import: you learned these ideas in the SQLite shell, rehearsed them in pandas, and today they transferred to R in minutes. Tool skills expire; concepts compound. Next lesson: the data frame itself: R's answer to the DataFrame you already think in.
Summary
Key takeaways
- read.csv("file.csv") imports a CSV as a data frame; header = TRUE and comma separator are defaults.
- Excel needs readxl: install.packages() once per machine, library() once per SESSION.
- R guesses column types; one stray string turns a numeric column to character.
- str() is the mandatory post-import audit: dimensions + per-column types; head() eyeballs rows.
- File-not-found almost always means working directory, not a missing file.
- read.delim for tabs, read.csv2 for semicolon files; packages exist for SPSS/JSON.
- Memory hook: the customs officer stamps it; you audit the stamp.