Data type conversion and recoding variables

as.numeric, as.character and as.factor convert between R's types (with NA for the unconvertible), ifelse() recodes values by condition, and factors must detour through as.character before becoming numbers.

9 min read · 9 cards · 2 checks

Read in: English · हिन्दी · ગુજરાતી


Theory

The column is finally ready to be fixed

Remember the import lesson's wounded column: one "absent" entry turned every screen-time number into text. You diagnosed it with str(); today you repair it:

survey$screen <- as.numeric(survey$screen)

Warning: NAs introduced by coercion

The digits became numbers; "absent" became NA (exactly where last lesson's skills take over). One line, one honest warning. Conversion is that simple: except for one famous trap that turns categories into wrong numbers.

Theory

Currency exchange with a strict teller

The as.___() family is a currency exchange: as.numeric converts to number-currency, as.character to text-currency, as.factor to category-currency.

The teller is strict: "120" exchanges cleanly to 120, but "absent" has no number equivalent: the teller hands back NA and says so aloud (the coercion warning). Never ignore the announcement: it is your count of how many values failed the exchange.

Practical

Convert, repair, recode

# CONVERSION: the as.___() family
raw <- c("120", "90", "absent", "240")
screen <- as.numeric(raw)     # Warning: NAs introduced by coercion
screen                        # [1] 120  90  NA 240   repaired!

class(screen)                 # "numeric": always verify after converting

# make a deliberate categorical (Unit 1's classification, executable):
stay <- as.factor(c("hostel", "day", "hostel"))
levels(stay)                  # "day" "hostel"

# RECODING: ifelse(condition, value_if_yes, value_if_no): vectorised
usage <- ifelse(screen > 180, "heavy", "normal")
usage                         # "normal" "normal" NA "heavy"

# three bands: nest the no-branch
band <- ifelse(screen > 180, "heavy",
        ifelse(screen > 90,  "medium", "light"))

# binning shortcut worth knowing: cut()
# cut(screen, breaks = c(0, 90, 180, 1440),
#     labels = c("light", "medium", "heavy"))

Theory

THE trap: factors do not convert directly

A factor stores labels as internal level codes (1, 2, 3... in alphabetical order). Convert it straight to numeric and you get the codes, not the values:

f <- factor(c("120", "90"))

as.numeric(f)

[1] 2 1

Not 120 and 90: 2 and 1 ("120" sorts before "90" alphabetically!). The correct route is the two-step detour:

as.numeric(as.character(f))

[1] 120 90

Labels first, numbers second. This is R's most famous beginner bug and a guaranteed exam question.

Quiz

marks <- factor(c("85", "100", "92")). A student runs as.numeric(marks) and gets 2 1 3. Why, and what is the fix?

  1. Factors convert to their internal level codes (alphabetical: "100" < "85" < "92"); fix: as.numeric(as.character(marks))
  2. R cannot store marks above 99, so it renumbered them
  3. The values needed quotes removed in the CSV first
  4. as.numeric only works on integers
Show the answer

Factors convert to their internal level codes (alphabetical: "100" < "85" < "92"); fix: as.numeric(as.character(marks))

Alphabetically "100" comes first (character comparison, digit by digit), so the levels are 100=1, 85=2, 92=3, and as.numeric returns those codes. The two-step as.numeric(as.character(...)) recovers the real values: labels out first, numbers second. Options B and D invent limits R does not have; C would help prevention but does not explain the mechanism, which is what the exam asks.

Think first

Recode the survey yourself

screen holds 120, 90, NA, 240. Work out what usage <- ifelse(screen > 180, "heavy", "normal") produces for EACH element (careful with the third), then tap.

Show the answer

"normal" "normal" NA "heavy"

Element by element: 120 and 90 fail the condition → "normal"; 240 passes → "heavy"; and NA > 180 is... NA (last lesson's comparison blindness), so ifelse passes the NA through: unknown in, unknown out.

ifelse is vectorised (whole column, no loop) and NA-honest: both properties exams check. For three or more bands, nest ifelse or reach for cut().

Watch out

The conversion checklist

1. Heed the coercion warning: it counts your new NAs: follow up with sum(is.na(...)).

2. Factor to number = two steps: as.numeric(as.character(f)): never direct.

3. class() before and after: conversions that silently did nothing (or the wrong thing) are caught in one glance.

4. Recode consistently: "heavy"/"Heavy"/"HEAVY" would recreate the text mess you cleaned two lessons ago.

Theory

Unit 4 closes: the data is finally trustworthy

Count the pipeline you now run end to end: import, audit, subset, add/rename, deduplicate, standardise, range-check, NA-handle, convert, recode. That IS professional data preparation: the 60-80% of real work. Unit 5 is the reward: sorting, merging, summary commands, frequency tables, and the bell curve drawn in R: statistics on data you can finally defend.

Summary

Key takeaways

  • as.numeric/as.character/as.factor convert types; unconvertible values become NA with a warning: count them.
  • The coercion warning is the repair receipt for text-polluted numeric columns.
  • FACTOR TRAP: as.numeric(factor) gives level codes; the fix is as.numeric(as.character(f)).
  • ifelse(condition, yes, no) recodes vectorised; NAs pass through as NA; nest for bands or use cut().
  • Promote categoricals deliberately with as.factor: levels() and table() come alive.
  • class() before and after every conversion.
  • Memory hook: the strict teller announces every failed exchange.

Study this properly

This page is the lesson to read. In Gri-Learn the same topic is a graded deck: the self-checks are scored and your weak topics are tracked. Free to start.

Start this topic

Already have an account? Sign in

More from Data Filtering and cleaning

Gri-Learn · syllabus-mapped B.C.A. lessons in English, Hindi and Gujarati

Data type conversion and recoding variables · Statistical Methods and Data Analysis (MDC-03) · Gri-Learn