Basic R syntax, variables, and data types

R assigns with the <- arrow, builds vectors with c(), and works element-wise on whole vectors at once, with numeric, character, logical and factor as the types statistics runs on.

10 min read · 10 cards · 2 checks

Read in: English · हिन्दी · ગુજરાતી


Theory

A language where the data is plural

Type your first real R at the console:

heights <- c(160, 168, 175, 162, 171)

mean(heights)

Answer: [1] 167.2. Two lines: five students stored, average computed.

Notice what you did NOT write: no loop, no import, no square-bracket ceremony. R's founding design choice: the basic unit of data is a whole collection, not a single value: because statistics never studies one number.

Theory

The arrow puts things in boxes

heights <- c(160, 168, 175) reads exactly as it looks: the values flow along the arrow into the box named heights.

<- is R's assignment: the visual "put into". (Plain = mostly works too, but <- is the convention every textbook, paper and examiner expects: write the arrow.)

And c() is the combine function: R's way of saying "these values travel together as one vector".

Theory

Vectors and the vectorised superpower

A vector is an ordered collection of same-type values: R's fundamental unit. Even x <- 5 is a vector of length 1: which is why every output starts with [1]: R numbering elements.

The superpower: operations apply to every element at once:

  • heights + 2 → every height grown by 2.
  • heights > 165 → a logical vector: FALSE TRUE TRUE FALSE TRUE.

Where Python's plain lists need a loop, R just... does it. One statistical thought, one line.

At a glance

The data types statistics runs on

TypeExampleNotes
numeric168, 3.14The default for numbers
integer60LThe L suffix forces integer
character"Surat"Quotes required
logicalTRUE, FALSEUppercase; from comparisons
factorfactor(c("hostel","day"))Categorical data with levels

Practical

Ten lines of survey R

# assignment: the arrow puts values into names
heights <- c(160, 168, 175, 162, 171)   # numeric vector
cities  <- c("Surat", "Navsari", "Surat", "Bardoli", "Surat")
stay    <- factor(c("hostel", "day", "hostel", "day", "day"))

mean(heights)        # [1] 167.2   built in, no import
sd(heights)          # [1] 6.058
length(heights)      # [1] 5

heights > 165        # [1] FALSE TRUE TRUE FALSE TRUE  (vectorised!)
class(cities)        # [1] "character"
levels(stay)         # [1] "day" "hostel"   the factor's categories

Think first

Predict three answers

With heights <- c(160, 168, 175, 162, 171), work out what R prints for each, then tap:

1. heights * 2

2. heights > 170

3. mean(heights > 170)

Show the answer

1. 320 336 350 324 342: every element doubled (vectorised).

2. FALSE FALSE TRUE FALSE TRUE: a logical vector.

3. 0.4: the sly one: TRUE counts as 1, FALSE as 0, so the mean of the logicals = the PROPORTION above 170 (2 of 5).

That third trick (mean of a condition = proportion satisfying it) is a beloved R idiom and a beloved exam question: one line computing a percentage.

Quiz

A student types: marks <- c(78, 85, "absent", 91) and then mean(marks) errors. What happened?

  1. Vectors are homogeneous: the one string silently coerced ALL elements to character, so mean() received text
  2. c() cannot hold more than three values
  3. The word absent needed uppercase
  4. mean() requires the values to be sorted first
Show the answer

Vectors are homogeneous: the one string silently coerced ALL elements to character, so mean() received text

R vectors hold ONE type. Faced with numbers and a string, R silently converts everything to the most flexible type: character: "78" is now text, and mean() cannot average words. This silent coercion is R's sneakiest beginner trap precisely because the c() line itself shows no error: check with class(marks). The statistical fix for a missing value is NA, next unit's topic.

Watch out

The four syntax stumbles

Case: mean() works, Mean() does not: R never forgives capitals (TRUE, not true).

Quotes: Surat without quotes is treated as a variable name and errors.

Coercion: one string turns a numeric vector to text, silently.

The arrow: write <- in exams; examiners read = as not-quite-R even where it runs.

Theory

Factors: Unit 1 vocabulary becomes a type

Notice factor() closing the loop: Unit 1's categorical variables are literally a data type here: levels() lists the categories, and R refuses to average them (correctly: remember the nonsense-maths trap). Numeric vectors = Unit 1's numerical variables. The classification you learned by hand is enforced by the language: next lesson, whole survey files arrive via read.csv, typed and ready.

Summary

Key takeaways

  • Assignment: name <- value (the arrow is the exam-expected form); # starts a comment.
  • c() combines values into a VECTOR, R's fundamental unit; single values are length-1 vectors (the [1]).
  • Operations are vectorised: heights + 2 and heights > 165 hit every element, no loops.
  • Types: numeric, integer (L), character (quoted), logical (TRUE/FALSE), factor (categorical with levels).
  • Vectors are homogeneous: a single string silently coerces everything to character.
  • mean(condition) computes a proportion: TRUE = 1, FALSE = 0.
  • Memory hook: the arrow puts the values in the box; c() makes them travel together.

Study this properly

This page is the lesson to read. In Gri-Learn the same topic is a graded deck: the self-checks are scored and your weak topics are tracked. Free to start.

Start this topic

Already have an account? Sign in

More from Introduction to R and working with Data

Gri-Learn · syllabus-mapped B.C.A. lessons in English, Hindi and Gujarati