Theory
The shape statistics actually lives in
Vectors carried single columns. But the CampusPulse survey is four columns wide: roll (numbers), city (text), stay (categories), screen time (numbers): different types, travelling together, one row per student.
R's container for that is the data frame: the type read.csv already handed you, the ancestor of pandas' DataFrame, and the shape every remaining lesson in this subject manipulates.
Today: create, inspect, measure, and save one: the daily verbs.
Theory
A register of stapled vectors
Picture the class register: one column of roll numbers, one of names, one of attendance: each column its own kind of data, all rows aligned so row 7 is the same student everywhere.
A data frame is exactly that: vectors of equal length stapled side by side: each column one type (a vector), the whole thing one table. Pull a column out ($) and you are back to a vector, with every vector skill still working.
Practical
Create, inspect, measure, save
# build one by hand (usually read.csv does this for you)
survey <- data.frame(
roll = c(101, 102, 103, 104),
city = c("Surat", "Navsari", "Surat", "Bardoli"),
screen = c(120, 90, 240, 150)
)
# inspect
head(survey, 2) # first 2 rows
str(survey) # structure: types per column
summary(survey) # per-column statistics (see below)
View(survey) # RStudio spreadsheet window: CAPITAL V
# measure
nrow(survey) # [1] 4
ncol(survey) # [1] 3
dim(survey) # [1] 4 3
names(survey) # "roll" "city" "screen"
# a column is just a vector again
mean(survey$screen) # [1] 150
# save: row.names = FALSE, always
write.csv(survey, "survey_out.csv", row.names = FALSE)
Theory
summary(): old friends in one command
For the screen column, summary(survey) prints:
Min. 1st Qu. Median Mean 3rd Qu. Max.
90 112.5 135 150 172.5 240
Read that row again slowly: minimum, Q1, median, Q3, maximum: the five-number summary from the box-plot lesson, plus the mean: computed for every numeric column at once.
Mean (150) above median (135)? Right skew, says your Unit 1 training. One command, and the whole descriptive-statistics toolkit reports in.
Quiz
A student saves with write.csv(survey, "out.csv") (no row.names argument). Reloading tomorrow shows a mystery first column X holding 1, 2, 3, 4. What happened?
- R wrote its row numbers into the file; write.csv(..., row.names = FALSE) prevents it
- The file was corrupted during saving
- read.csv always adds an X column
- The roll column was renamed to X
Show the answer
R wrote its row numbers into the file; write.csv(..., row.names = FALSE) prevents it
By default write.csv exports the data frame's ROW NAMES (1, 2, 3...) as a real column, which reloads under the invented name X: save again and they breed. The cure is the habit row.names = FALSE: the exact same disease and cure as pandas' to_csv index=False from BCA303 (there the mystery column was called 'Unnamed: 0'). Same lesson, second language: exporters must be told not to ship their internal numbering.
Think first
Predict four outputs
Using the survey data frame above (4 rows: screens 120, 90, 240, 150), work out what each prints, then tap:
1. dim(survey)
2. survey$screen > 100
3. mean(survey$screen > 100)
4. view(survey)
Show the answer
1. [1] 4 3: rows then columns.
2. TRUE FALSE TRUE TRUE: a logical vector (the $ pulled a vector; comparisons vectorise).
3. 0.75: proportion above 100: the mean-of-condition idiom returns.
4. Error: could not find function "view": the viewer is View(), capital V: case sensitivity never sleeps.
If you got all four, the vector lessons and the data frame just clicked together: columns ARE vectors.
Watch out
The daily slips
view() vs View(): lowercase fails: the one capital letter beginners forget most.
row.names = FALSE on every write.csv, or the X column haunts every reload.
$ is case-sensitive too: survey$Screen returns NULL (not an error!) when the column is screen: a silent NULL that crashes the NEXT line. Check names(df) when a column "disappears".
Theory
The bridge to real work
You can now round-trip data: read.csv in, str/summary to audit, $ to compute, write.csv out: the same loop you ran in pandas, one language over. From here the subject stops touring and starts working: next lesson slices this data frame (rows by condition, columns by name): R's answer to SQL's WHERE and pandas' masks: your third time meeting the same beautiful idea.
Summary
Key takeaways
- A data frame = equal-length vectors stapled as columns: one type per column, mixed types per table.
- Create with data.frame() or import via read.csv; df$col extracts a column as a vector.
- Inspect: head/tail, str (types), summary (five-number summary + mean per column), View (capital V).
- Measure: nrow, ncol, dim, names.
- Save with write.csv(df, file, row.names = FALSE): the index=False of R.
- A wrong-case column name returns silent NULL, not an error.
- Memory hook: a register of stapled vectors.