Theory
Unit 1, replayed at machine speed
Weeks ago you computed the screen-time mean, median and SD by hand: deviations, squares, square roots, the full ritual.
In R, the ritual is four words:
mean(survey$screen, na.rm = TRUE)
Same number your paper produced. That is the deal this lesson formalises: the by-hand work taught you what the numbers MEAN; today's commands compute them instantly: with one famous omission that catches every beginner in the country.
Theory
The calculator with one missing button
R's statistics keyboard has a key for everything: mean, median, sd, var, quantile...
...except mode. The button LABELLED mode() is a prank: press it and R tells you the storage mode of your data ("numeric"): a programming answer to a statistics question.
The real most-frequent-value must be assembled from two other keys: table() to count, which.max() to crown the winner.
Practical
The whole Unit 1 toolkit, one line each
screen <- c(60, 90, 90, 120, 150, 180, 600, NA)
# central tendency
mean(screen, na.rm = TRUE) # [1] 184.3 (outlier-dragged, as Unit 1 warned)
median(screen, na.rm = TRUE) # [1] 120
# dispersion (SAMPLE formulas: divide by n-1)
sd(screen, na.rm = TRUE)
var(screen, na.rm = TRUE)
range(screen, na.rm = TRUE) # min and max together
quantile(screen, na.rm = TRUE) # 0% 25% 50% 75% 100%
# the five-number overview + mean, one call:
summary(screen)
# THE QUIRK: mode() is NOT the statistical mode
mode(screen) # [1] "numeric" (storage mode!)
# the real mode, assembled:
t <- table(screen)
names(which.max(t)) # [1] "90"
# group-wise statistics: R's GROUP BY
# tapply(survey$screen, survey$stay, mean, na.rm = TRUE)
# aggregate(screen ~ stay, data = survey, FUN = mean)
Theory
Three habits that make the numbers honest
na.rm = TRUE, always considered: one NA silences every one of these functions (the propagation rule): decide consciously each time.
Know your divisor: sd() and var() use the sample formula (n-1), matching pandas: say "sample SD" in reports, exactly as the dispersion lesson drilled.
Verify once by hand: run sd() on the tiny dataset you computed manually in Unit 1: same answer builds the trust that lets you delegate the arithmetic forever after.
Quiz
A student runs mode(survey$screen) hoping for the most frequent screen time and gets "numeric". What is going on, and what computes the real mode?
- mode() returns the STORAGE type in R; the statistical mode is assembled with names(which.max(table(x)))
- The data has no repeated values, so R printed its type instead
- mode() needs na.rm = TRUE to work
- Only factors have modes in R
Show the answer
mode() returns the STORAGE type in R; the statistical mode is assembled with names(which.max(table(x)))
R's mode() answers "how is this stored?": a naming collision with statistics that has trapped generations of students: the classic R gotcha, worth marks precisely because it is so surprising. The real mode: table(x) counts each value, which.max() finds the biggest count, names() reads off the winning value. Options B to D invent behaviours: the function NEVER computes frequency, regardless of data or arguments.
Think first
Group-wise: the dean's one-liner
The dean asks: "average screen time, hostellers vs day scholars, one command." Using tapply(survey$screen, survey$stay, mean, na.rm = TRUE): describe what each of the three arguments does, then tap.
Show the answer
tapply(values, groups, function): take survey$screen (what to summarise), split it by survey$stay (the piles: hostel and day), apply mean to each pile: output is one labelled average per group:
day hostel
142 201
It is the GROUP BY of R (aggregate(screen ~ stay, df, mean) says the same with a formula). You built these piles by hand in the GROUP BY lesson of BCA303: same idea, third syntax, ten seconds.
Watch out
The summary-statistics slips
mode() is never the mode: table + which.max, or your own two-line function.
NA silence: a single missing value returns NA from mean/sd: the na.rm decision is part of the answer, not an afterthought.
quantile() defaults to 0/25/50/75/100%: ask for others explicitly: quantile(x, 0.9) for the 90th percentile.
Theory
Where this plugs in next
These one-liners feed everything left in the subject: frequency tables count the categoricals these functions cannot summarise (next lesson), the commands lesson consolidates the whole surface, and the graphical finale draws what summary() describes. You are also now bilingual within the course: df['screen'].mean() and mean(df$screen) are the same thought: portability is the real syllabus.
Summary
Key takeaways
- mean, median, sd, var, range, quantile: one call each; summary() = five-number overview + mean.
- na.rm = TRUE is a conscious decision on every call: NA silences them otherwise.
- sd()/var() use the sample divisor (n-1): say 'sample SD', matching pandas.
- mode() returns the STORAGE type: the real mode is names(which.max(table(x))).
- Group-wise: tapply(values, groups, fun) or aggregate(y ~ group, df, fun): R's GROUP BY.
- Verify commands once against your Unit 1 hand calculations: then delegate forever.
- Memory hook: the calculator with one missing (mislabelled) button.