Theory
One card to close the loop
Unit 1 taught the measures by hand. Scattered lessons since taught their commands. The exam question that remains is consolidation: "Write R commands for the measures of central tendency and dispersion for a given dataset."
This lesson is that answer, assembled once: every measure, its command, its trap: the reference card you revise from. Nothing here is new; everything here is now one card.
Theory
The examiner's shopping list
Picture the examiner's checklist for a full-marks descriptive answer:
- Where is the centre? (mean, median, mode: pick with judgement)
- How wide is the spread? (range, IQR, sd: three widths, three meanings)
- Did you handle the holes? (the na.rm decision, stated)
- Sample or population? (name your divisor)
Four lines on the list: the card below buys all four.
At a glance
The reference card
| Measure | Command | Watch out |
|---|---|---|
| Mean | mean(x, na.rm = TRUE) | Outlier-dragged |
| Median | median(x, na.rm = TRUE) | Robust centre |
| Mode | names(which.max(table(x))) | No built-in; mode() is a prank |
| Range | range(x) then diff(range(x)) | range() returns min AND max |
| IQR | IQR(x, na.rm = TRUE) | Q3 - Q1, the box width |
| Variance / SD | var(x), sd(x) | Sample divisor (n-1) |
| Overview | summary(x), quantile(x) | fivenum(x) = five-number summary |
Practical
The full descriptive report, twelve lines
survey <- read.csv("survey_clean.csv")
x <- survey$screen
# centre
mean(x, na.rm = TRUE) # 171.4
median(x, na.rm = TRUE) # 135
get_mode <- function(v) names(which.max(table(v)))
get_mode(x) # "90" (your first user function!)
# spread
diff(range(x, na.rm = TRUE)) # max - min, ONE number
IQR(x, na.rm = TRUE) # the box plot's box width
sd(x, na.rm = TRUE) # sample SD (divides by n-1)
var(x, na.rm = TRUE) # sd squared
# overview + per-group spread
summary(x)
tapply(survey$screen, survey$stay, sd, na.rm = TRUE)
# day hostel <- who is more CONSISTENT?
# 38.2 71.5
Theory
Commands are easy; the sentence is the skill
The report block prints numbers; marks come from reading them aloud:
- mean 171 vs median 135: mean sits well above: right skew, report the median as typical.
- hostel sd 71.5 vs day 38.2: hostellers are twice as scattered: one average would misrepresent them.
- mode "90": the single most common habit.
Every number gets one interpreting sentence: that habit converts a computation answer into a statistics answer.
Quiz
A student wants "the range" as one number and writes range(x), getting [1] 30 600. What happened, and what gives the single number?
- range() returns the min AND max pair; the single spread number is diff(range(x)), here 570
- The data has two ranges because of NA values
- range() is broken for vectors longer than 100
- Use IQR(x): IQR and range are the same measure
Show the answer
range() returns the min AND max pair; the single spread number is diff(range(x)), here 570
R's range() hands you the two endpoints (30 and 600); the textbook "range = max - min" needs diff(range(x)) or max(x) - min(x): 570. Option D smuggles in the other confusion this card guards: IQR is Q3 - Q1 (the middle 50%'s width, outlier-resistant), a DIFFERENT and usually better spread measure than the outlier-fragile full range. Two names, two measures, one classic mix-up.
Think first
Write the exam answer cold
Exam task: "For marks stored in vector m (which may contain NA), write R commands for mean, median, mode, standard deviation and IQR, and state one precaution." Draft all five lines and the precaution before tapping.
Show the answer
mean(m, na.rm = TRUE)
median(m, na.rm = TRUE)
names(which.max(table(m))) # R has no built-in statistical mode
sd(m, na.rm = TRUE) # sample SD, divisor n-1
IQR(m, na.rm = TRUE)
Precaution: na.rm = TRUE on every call (one NA otherwise returns NA), and note the mode workaround since mode() returns the storage type. That answer, verbatim, is full marks: this card exists so you can produce it half-asleep.
Watch out
Type decides the toolkit
These commands serve numeric variables. Feed a factor to mean() and R returns NA with a warning: correctly, because Unit 1's law still rules: categorical data gets table()/prop.table() (last lesson), numerical data gets this card. The variable-type question you learned first is the router for every command you learned since.
Theory
You just wrote your first function
get_mode is three tokens of syntax around logic you already owned: function(v) wraps a recipe, and calling get_mode(x) replays it: the doorway to real programming in R. Two lessons remain: the normal distribution's R commands (pnorm and family), then the graphical finale where every number on today's card becomes a picture.
Summary
Key takeaways
- Centre: mean, median (na.rm = TRUE), mode via names(which.max(table(x))).
- Spread: diff(range(x)) for full width, IQR(x) for the robust middle-50% width, sd/var (sample divisor).
- Overview: summary(x), quantile(x), fivenum(x).
- range() returns the min-max PAIR, not one number.
- Interpret every number in one sentence: skew (mean vs median), consistency (sd), typical (mode).
- Numeric variables use this card; categoricals use table(): type routes the toolkit.
- Memory hook: the examiner's four-line shopping list.