Adding, removing, and renaming variables/attributes

df$new <- values adds a column (often computed from existing ones), df$col <- NULL deletes one, and names(df)[i] <- "better" renames: the data frame is clay, not stone.

8 min read · 9 cards · 2 checks

Read in: English · हिन्दी · ગુજરાતી


Theory

The survey needs a column it was never asked

The dean liked the shortlist, then asked the inevitable: "can I see screen time in hours, not minutes? And drop the roll numbers when you print it."

The CSV has no hours column. Nobody re-surveys 60 students to divide by 60.

A data frame is not the frozen photo of a file: it is clay in memory. Today's three verbs: sculpt new columns from old ones, cut columns away, and rename what is badly labelled.

Theory

The register grows in pencil

A class register printed in June still gains pencil columns all term: a teacher adds "attendance %", computed from the existing columns; crosses out a column nobody uses; relabels a cryptic heading.

The printed original (survey.csv on disk) never changes: the pencil work lives on YOUR copy (the data frame in memory), until you photocopy it into a new file with write.csv.

Practical

Add, remove, rename

survey <- read.csv("survey.csv")   # roll, city, stay, screen

# ADD: derived from existing columns (vectorised: whole column at once)
survey$hours <- survey$screen / 60

# ADD: a logical flag column
survey$heavy <- survey$screen > 180

# ADD: a constant (recycled to every row)
survey$batch <- 2026

# REMOVE: immediate, no confirmation
survey$batch <- NULL

# RENAME by matching the old name (safe, position-independent)
names(survey)[names(survey) == "screen"] <- "screen_min"

# RENAME by position (fragile: breaks if order changes)
# names(survey)[4] <- "screen_min"

names(survey)   # "roll" "city" "stay" "screen_min" "hours" "heavy"

# memory only! persist deliberately:
write.csv(survey, "survey_v2.csv", row.names = FALSE)

Theory

Why derived columns are the real skill

Adding a constant is clerical. The analyst's move is the derived column: a new variable computed row-wise from existing ones:

  • unit conversions: hours <- screen / 60
  • flags: heavy <- screen > 180
  • combinations: per_class <- marks / classes

Each is one vectorised line: no loop, every row at once. In statistics this is called transformation, and half of next lesson's "data cleaning" is exactly this: reshaping raw columns into analysable ones.

Quiz

A student runs survey$city <- NULL, meaning to remove a different column. What is the situation now?

  1. city is gone from the data frame immediately; recover it by re-importing from survey.csv (the file is untouched)
  2. R asked for confirmation before deleting, so nothing happened
  3. The column is in a recycle bin inside RStudio
  4. The CSV file on disk also lost its city column
Show the answer

city is gone from the data frame immediately; recover it by re-importing from survey.csv (the file is untouched)

NULL-deletion is instant and silent: no prompt, no undo within the session. The saving grace is the pencil-copy model: the data frame is a memory copy, so the file on disk still has every column: survey <- read.csv("survey.csv") restores the original (minus your other pencil work, redo it). Options B and C describe safety nets R does not have; D confuses memory with disk: only write.csv touches the file.

Think first

The typo that creates instead of edits

A student wants to convert the screen column to hours IN PLACE and types: survey$scren <- survey$screen / 60 (note the spelling). No error appears. Before tapping: what does the data frame look like now, and why is the silence dangerous?

Show the answer

R created a brand-new column named scren holding the hours, while screen sits unchanged: assignment to an unknown column name means "add it", never "did you mean...?".

The danger is the silence: no error now, confusion later: two similar columns, and any analysis using $screen still runs on minutes. Habit: after column surgery, glance at names(df) or str(df): the ten-second audit catches every typo of this species.

Watch out

Three column-surgery rules

Rename by NAME match, not position: names(df)[names(df) == "old"] <- "new" survives column reordering; [4] does not.

NULL is forever (in-session): prefer building a copy (df2 <- df[, keep]) when unsure.

Check names() after surgery: typos create columns silently: the one-glance audit.

Theory

Attributes, the exam word

Syllabi call columns variables or attributes (the database word from BCA303: same thing). So exam phrasing like "add an attribute to a data frame" is exactly today's df$new <- ... line. Next lesson zooms out from single columns to the whole cleaning workflow: duplicates, stray text, impossible values: turning a raw survey into data you would defend in front of the dean.

Summary

Key takeaways

  • Add: df$new <- expression (vectorised, often derived from other columns); constants recycle.
  • Remove: df$col <- NULL: instant, silent, session-permanent (the disk file is untouched).
  • Rename: names(df)[names(df) == "old"] <- "new": name-matched beats position.
  • Derived columns (conversions, flags) are the analyst's core transformation move.
  • Typos in $assignment CREATE new columns silently: audit with names()/str() after surgery.
  • All edits live in memory until write.csv persists them.
  • Memory hook: the register grows in pencil; the printout on disk stays clean.

Study this properly

This page is the lesson to read. In Gri-Learn the same topic is a graded deck: the self-checks are scored and your weak topics are tracked. Free to start.

Start this topic

Already have an account? Sign in

More from Data Filtering and cleaning

Gri-Learn · syllabus-mapped B.C.A. lessons in English, Hindi and Gujarati

Adding, removing, and renaming variables/attributes · Statistical Methods and Data Analysis (MDC-03) · Gri-Learn