Theory
The mean that refused to answer
After last lesson's cleaning, roll 121's impossible screen time became NA. Now you ask for the average:
mean(survey$screen)
[1] NA
R refuses to answer. Not an error, not a guess: a flat "unknown".
Annoying? It is the most honest behaviour in this course: one value is unknown, so the exact mean is unknowable, and R says so. Today: how to detect the holes, and the three legitimate ways to work around them.
Theory
The blank in the attendance register
A register row shows: present, present, blank, present. What does the blank mean? Not "absent": absent has its own mark. The blank means nobody recorded it: unknown.
NA is that blank. It is not 0 (a real measured zero), not "" (an empty text), not NULL (no box at all): it is a box whose contents nobody knows. And any total computed OVER a blank is itself... blank. R takes that logic seriously.
Theory
NA propagates, and == cannot see it
Two rules govern NA:
Propagation: any arithmetic touching NA yields NA: mean, sum, sd, even 5 + NA. Missingness infects results, loudly, by design.
Comparison blindness: x == NA returns NA, never TRUE: asking "is this unknown equal to unknown?" has no answer. You met this exact logic in BCA303: SQL's = NULL failed the same way, for the same reason.
The only honest detector is a dedicated function: is.na().
Practical
Detect, then choose a strategy
screen <- c(120, 90, NA, 240, 150, NA)
# DETECT
is.na(screen) # FALSE FALSE TRUE FALSE FALSE TRUE
sum(is.na(screen)) # [1] 2 how many holes
mean(is.na(screen)) # [1] 0.333 what fraction is missing
# the wrong way (comparison blindness):
screen == NA # NA NA NA NA NA NA never TRUE!
# STRATEGY 1: compute around the holes
mean(screen, na.rm = TRUE) # [1] 150 NAs removed for this calc
# STRATEGY 2: keep only complete rows (whole data frame)
survey_complete <- na.omit(survey)
# per-column hole count for a data frame:
colSums(is.na(survey))
# STRATEGY 3 (advanced, use with care): impute
# screen[is.na(screen)] <- median(screen, na.rm = TRUE)
Quiz
A student "fixes" the NA problem with survey$screen[is.na(survey$screen)] <- 0 and proudly reports the new mean. What is wrong?
- Zero is a CLAIM (student uses no screen at all), not an absence: fake zeros drag the mean down and bias every statistic
- Nothing: 0 and NA mean the same thing
- The syntax is invalid: NAs cannot be assigned over
- It only fails if more than half the values are NA
Show the answer
Zero is a CLAIM (student uses no screen at all), not an absence: fake zeros drag the mean down and bias every statistic
NA says "we do not know"; 0 says "we know, and it is zero": replacing one with the other manufactures data, and every fake 0 pulls the mean toward zero (two fake zeros in our six values would report 100 instead of 150). The mechanically-correct fixes are na.rm = TRUE or na.omit(); imputation with a median is the defensible advanced option: with a note in the report. The syntax in the question is valid R: the crime is statistical, not syntactic.
Think first
Predict four outputs
With x <- c(10, NA, 30), work out each before tapping:
1. sum(x)
2. sum(x, na.rm = TRUE)
3. x == NA
4. sum(is.na(x))
Show the answer
1. NA: propagation: one unknown makes the total unknown.
2. 40: the hole is set aside for this calculation.
3. NA NA NA: comparisons with NA never resolve: the blindness rule.
4. 1: is.na() is the honest detector, and summing its TRUEs counts holes.
Four lines, all four NA rules: this exact quartet is how exams test the topic, and how you should self-check any dataset before trusting a statistic from it.
Watch out
Choosing the strategy honestly
na.rm = TRUE: quick statistics when holes are few and random.
na.omit(): analyses needing complete rows: but check how many rows you lose (nrow before vs after): dropping half the survey silently is its own crime.
Imputation: replacing NA with mean/median keeps the row but manufactures the value: acceptable in practice, ALWAYS disclosed.
Never: NA → 0. And always report the hole count next to any statistic.
Theory
The concept you have now met three times
SQLite called it NULL (IS NULL, the = NULL trap), pandas showed it as NaN, R spells it NA: three languages, one idea: unknown is not a value, and equality cannot see it. That cross-language recognition is worth more than any single syntax: interviewers probe exactly this. Next lesson: converting and recoding types: including the factor trap that turns categories into wrong numbers.
Summary
Key takeaways
- NA = unknown: distinct from 0 (a real zero), "" (empty text) and NULL (no object).
- NA propagates: mean/sum/sd over any NA return NA: loud by design.
- x == NA never works (returns NA): detect with is.na(); count with sum(is.na(x)).
- Compute around holes with na.rm = TRUE; drop incomplete rows with na.omit() (check the row loss).
- Imputation (median/mean) manufactures data: use sparingly, disclose always; NA → 0 never.
- Same concept as SQL NULL and pandas NaN.
- Memory hook: the blank in the register is not a zero.