Theory
A language where the data is plural
Type your first real R at the console:
heights <- c(160, 168, 175, 162, 171)
mean(heights)
Answer: [1] 167.2. Two lines: five students stored, average computed.
Notice what you did NOT write: no loop, no import, no square-bracket ceremony. R's founding design choice: the basic unit of data is a whole collection, not a single value: because statistics never studies one number.
Theory
The arrow puts things in boxes
heights <- c(160, 168, 175) reads exactly as it looks: the values flow along the arrow into the box named heights.
<- is R's assignment: the visual "put into". (Plain = mostly works too, but <- is the convention every textbook, paper and examiner expects: write the arrow.)
And c() is the combine function: R's way of saying "these values travel together as one vector".
Theory
Vectors and the vectorised superpower
A vector is an ordered collection of same-type values: R's fundamental unit. Even x <- 5 is a vector of length 1: which is why every output starts with [1]: R numbering elements.
The superpower: operations apply to every element at once:
heights + 2→ every height grown by 2.heights > 165→ a logical vector: FALSE TRUE TRUE FALSE TRUE.
Where Python's plain lists need a loop, R just... does it. One statistical thought, one line.
At a glance
The data types statistics runs on
| Type | Example | Notes |
|---|---|---|
| numeric | 168, 3.14 | The default for numbers |
| integer | 60L | The L suffix forces integer |
| character | "Surat" | Quotes required |
| logical | TRUE, FALSE | Uppercase; from comparisons |
| factor | factor(c("hostel","day")) | Categorical data with levels |
Practical
Ten lines of survey R
# assignment: the arrow puts values into names
heights <- c(160, 168, 175, 162, 171) # numeric vector
cities <- c("Surat", "Navsari", "Surat", "Bardoli", "Surat")
stay <- factor(c("hostel", "day", "hostel", "day", "day"))
mean(heights) # [1] 167.2 built in, no import
sd(heights) # [1] 6.058
length(heights) # [1] 5
heights > 165 # [1] FALSE TRUE TRUE FALSE TRUE (vectorised!)
class(cities) # [1] "character"
levels(stay) # [1] "day" "hostel" the factor's categories
Think first
Predict three answers
With heights <- c(160, 168, 175, 162, 171), work out what R prints for each, then tap:
1. heights * 2
2. heights > 170
3. mean(heights > 170)
Show the answer
1. 320 336 350 324 342: every element doubled (vectorised).
2. FALSE FALSE TRUE FALSE TRUE: a logical vector.
3. 0.4: the sly one: TRUE counts as 1, FALSE as 0, so the mean of the logicals = the PROPORTION above 170 (2 of 5).
That third trick (mean of a condition = proportion satisfying it) is a beloved R idiom and a beloved exam question: one line computing a percentage.
Quiz
A student types: marks <- c(78, 85, "absent", 91) and then mean(marks) errors. What happened?
- Vectors are homogeneous: the one string silently coerced ALL elements to character, so mean() received text
- c() cannot hold more than three values
- The word absent needed uppercase
- mean() requires the values to be sorted first
Show the answer
Vectors are homogeneous: the one string silently coerced ALL elements to character, so mean() received text
R vectors hold ONE type. Faced with numbers and a string, R silently converts everything to the most flexible type: character: "78" is now text, and mean() cannot average words. This silent coercion is R's sneakiest beginner trap precisely because the c() line itself shows no error: check with class(marks). The statistical fix for a missing value is NA, next unit's topic.
Watch out
The four syntax stumbles
Case: mean() works, Mean() does not: R never forgives capitals (TRUE, not true).
Quotes: Surat without quotes is treated as a variable name and errors.
Coercion: one string turns a numeric vector to text, silently.
The arrow: write <- in exams; examiners read = as not-quite-R even where it runs.
Theory
Factors: Unit 1 vocabulary becomes a type
Notice factor() closing the loop: Unit 1's categorical variables are literally a data type here: levels() lists the categories, and R refuses to average them (correctly: remember the nonsense-maths trap). Numeric vectors = Unit 1's numerical variables. The classification you learned by hand is enforced by the language: next lesson, whole survey files arrive via read.csv, typed and ready.
Summary
Key takeaways
- Assignment: name <- value (the arrow is the exam-expected form); # starts a comment.
- c() combines values into a VECTOR, R's fundamental unit; single values are length-1 vectors (the [1]).
- Operations are vectorised: heights + 2 and heights > 165 hit every element, no loops.
- Types: numeric, integer (L), character (quoted), logical (TRUE/FALSE), factor (categorical with levels).
- Vectors are homogeneous: a single string silently coerces everything to character.
- mean(condition) computes a proportion: TRUE = 1, FALSE = 0.
- Memory hook: the arrow puts the values in the box; c() makes them travel together.