Theory
The column is finally ready to be fixed
Remember the import lesson's wounded column: one "absent" entry turned every screen-time number into text. You diagnosed it with str(); today you repair it:
survey$screen <- as.numeric(survey$screen)
Warning: NAs introduced by coercion
The digits became numbers; "absent" became NA (exactly where last lesson's skills take over). One line, one honest warning. Conversion is that simple: except for one famous trap that turns categories into wrong numbers.
Theory
Currency exchange with a strict teller
The as.___() family is a currency exchange: as.numeric converts to number-currency, as.character to text-currency, as.factor to category-currency.
The teller is strict: "120" exchanges cleanly to 120, but "absent" has no number equivalent: the teller hands back NA and says so aloud (the coercion warning). Never ignore the announcement: it is your count of how many values failed the exchange.
Practical
Convert, repair, recode
# CONVERSION: the as.___() family
raw <- c("120", "90", "absent", "240")
screen <- as.numeric(raw) # Warning: NAs introduced by coercion
screen # [1] 120 90 NA 240 repaired!
class(screen) # "numeric": always verify after converting
# make a deliberate categorical (Unit 1's classification, executable):
stay <- as.factor(c("hostel", "day", "hostel"))
levels(stay) # "day" "hostel"
# RECODING: ifelse(condition, value_if_yes, value_if_no): vectorised
usage <- ifelse(screen > 180, "heavy", "normal")
usage # "normal" "normal" NA "heavy"
# three bands: nest the no-branch
band <- ifelse(screen > 180, "heavy",
ifelse(screen > 90, "medium", "light"))
# binning shortcut worth knowing: cut()
# cut(screen, breaks = c(0, 90, 180, 1440),
# labels = c("light", "medium", "heavy"))
Theory
THE trap: factors do not convert directly
A factor stores labels as internal level codes (1, 2, 3... in alphabetical order). Convert it straight to numeric and you get the codes, not the values:
f <- factor(c("120", "90"))
as.numeric(f)
[1] 2 1
Not 120 and 90: 2 and 1 ("120" sorts before "90" alphabetically!). The correct route is the two-step detour:
as.numeric(as.character(f))
[1] 120 90
Labels first, numbers second. This is R's most famous beginner bug and a guaranteed exam question.
Quiz
marks <- factor(c("85", "100", "92")). A student runs as.numeric(marks) and gets 2 1 3. Why, and what is the fix?
- Factors convert to their internal level codes (alphabetical: "100" < "85" < "92"); fix: as.numeric(as.character(marks))
- R cannot store marks above 99, so it renumbered them
- The values needed quotes removed in the CSV first
- as.numeric only works on integers
Show the answer
Factors convert to their internal level codes (alphabetical: "100" < "85" < "92"); fix: as.numeric(as.character(marks))
Alphabetically "100" comes first (character comparison, digit by digit), so the levels are 100=1, 85=2, 92=3, and as.numeric returns those codes. The two-step as.numeric(as.character(...)) recovers the real values: labels out first, numbers second. Options B and D invent limits R does not have; C would help prevention but does not explain the mechanism, which is what the exam asks.
Think first
Recode the survey yourself
screen holds 120, 90, NA, 240. Work out what usage <- ifelse(screen > 180, "heavy", "normal") produces for EACH element (careful with the third), then tap.
Show the answer
"normal" "normal" NA "heavy"
Element by element: 120 and 90 fail the condition → "normal"; 240 passes → "heavy"; and NA > 180 is... NA (last lesson's comparison blindness), so ifelse passes the NA through: unknown in, unknown out.
ifelse is vectorised (whole column, no loop) and NA-honest: both properties exams check. For three or more bands, nest ifelse or reach for cut().
Watch out
The conversion checklist
1. Heed the coercion warning: it counts your new NAs: follow up with sum(is.na(...)).
2. Factor to number = two steps: as.numeric(as.character(f)): never direct.
3. class() before and after: conversions that silently did nothing (or the wrong thing) are caught in one glance.
4. Recode consistently: "heavy"/"Heavy"/"HEAVY" would recreate the text mess you cleaned two lessons ago.
Theory
Unit 4 closes: the data is finally trustworthy
Count the pipeline you now run end to end: import, audit, subset, add/rename, deduplicate, standardise, range-check, NA-handle, convert, recode. That IS professional data preparation: the 60-80% of real work. Unit 5 is the reward: sorting, merging, summary commands, frequency tables, and the bell curve drawn in R: statistics on data you can finally defend.
Summary
Key takeaways
- as.numeric/as.character/as.factor convert types; unconvertible values become NA with a warning: count them.
- The coercion warning is the repair receipt for text-polluted numeric columns.
- FACTOR TRAP: as.numeric(factor) gives level codes; the fix is as.numeric(as.character(f)).
- ifelse(condition, yes, no) recodes vectorised; NAs pass through as NA; nest for bands or use cut().
- Promote categoricals deliberately with as.factor: levels() and table() come alive.
- class() before and after every conversion.
- Memory hook: the strict teller announces every failed exchange.