Theory
The survey needs a column it was never asked
The dean liked the shortlist, then asked the inevitable: "can I see screen time in hours, not minutes? And drop the roll numbers when you print it."
The CSV has no hours column. Nobody re-surveys 60 students to divide by 60.
A data frame is not the frozen photo of a file: it is clay in memory. Today's three verbs: sculpt new columns from old ones, cut columns away, and rename what is badly labelled.
Theory
The register grows in pencil
A class register printed in June still gains pencil columns all term: a teacher adds "attendance %", computed from the existing columns; crosses out a column nobody uses; relabels a cryptic heading.
The printed original (survey.csv on disk) never changes: the pencil work lives on YOUR copy (the data frame in memory), until you photocopy it into a new file with write.csv.
Practical
Add, remove, rename
survey <- read.csv("survey.csv") # roll, city, stay, screen
# ADD: derived from existing columns (vectorised: whole column at once)
survey$hours <- survey$screen / 60
# ADD: a logical flag column
survey$heavy <- survey$screen > 180
# ADD: a constant (recycled to every row)
survey$batch <- 2026
# REMOVE: immediate, no confirmation
survey$batch <- NULL
# RENAME by matching the old name (safe, position-independent)
names(survey)[names(survey) == "screen"] <- "screen_min"
# RENAME by position (fragile: breaks if order changes)
# names(survey)[4] <- "screen_min"
names(survey) # "roll" "city" "stay" "screen_min" "hours" "heavy"
# memory only! persist deliberately:
write.csv(survey, "survey_v2.csv", row.names = FALSE)
Theory
Why derived columns are the real skill
Adding a constant is clerical. The analyst's move is the derived column: a new variable computed row-wise from existing ones:
- unit conversions:
hours <- screen / 60 - flags:
heavy <- screen > 180 - combinations:
per_class <- marks / classes
Each is one vectorised line: no loop, every row at once. In statistics this is called transformation, and half of next lesson's "data cleaning" is exactly this: reshaping raw columns into analysable ones.
Quiz
A student runs survey$city <- NULL, meaning to remove a different column. What is the situation now?
- city is gone from the data frame immediately; recover it by re-importing from survey.csv (the file is untouched)
- R asked for confirmation before deleting, so nothing happened
- The column is in a recycle bin inside RStudio
- The CSV file on disk also lost its city column
Show the answer
city is gone from the data frame immediately; recover it by re-importing from survey.csv (the file is untouched)
NULL-deletion is instant and silent: no prompt, no undo within the session. The saving grace is the pencil-copy model: the data frame is a memory copy, so the file on disk still has every column: survey <- read.csv("survey.csv") restores the original (minus your other pencil work, redo it). Options B and C describe safety nets R does not have; D confuses memory with disk: only write.csv touches the file.
Think first
The typo that creates instead of edits
A student wants to convert the screen column to hours IN PLACE and types: survey$scren <- survey$screen / 60 (note the spelling). No error appears. Before tapping: what does the data frame look like now, and why is the silence dangerous?
Show the answer
R created a brand-new column named scren holding the hours, while screen sits unchanged: assignment to an unknown column name means "add it", never "did you mean...?".
The danger is the silence: no error now, confusion later: two similar columns, and any analysis using $screen still runs on minutes. Habit: after column surgery, glance at names(df) or str(df): the ten-second audit catches every typo of this species.
Watch out
Three column-surgery rules
Rename by NAME match, not position: names(df)[names(df) == "old"] <- "new" survives column reordering; [4] does not.
NULL is forever (in-session): prefer building a copy (df2 <- df[, keep]) when unsure.
Check names() after surgery: typos create columns silently: the one-glance audit.
Theory
Attributes, the exam word
Syllabi call columns variables or attributes (the database word from BCA303: same thing). So exam phrasing like "add an attribute to a data frame" is exactly today's df$new <- ... line. Next lesson zooms out from single columns to the whole cleaning workflow: duplicates, stray text, impossible values: turning a raw survey into data you would defend in front of the dean.
Summary
Key takeaways
- Add: df$new <- expression (vectorised, often derived from other columns); constants recycle.
- Remove: df$col <- NULL: instant, silent, session-permanent (the disk file is untouched).
- Rename: names(df)[names(df) == "old"] <- "new": name-matched beats position.
- Derived columns (conversions, flags) are the analyst's core transformation move.
- Typos in $assignment CREATE new columns silently: audit with names()/str() after surgery.
- All edits live in memory until write.csv persists them.
- Memory hook: the register grows in pencil; the printout on disk stays clean.