Calculating summary statistics (mean, median, mode, standard deviation)

mean(), median(), sd() and var() compute Unit 1's statistics in one call each (with na.rm = TRUE as the reflex), summary() reports the five-number overview, and R famously has NO built-in mode function.

9 min read · 9 cards · 2 checks

Read in: English · हिन्दी · ગુજરાતી


Theory

Unit 1, replayed at machine speed

Weeks ago you computed the screen-time mean, median and SD by hand: deviations, squares, square roots, the full ritual.

In R, the ritual is four words:

mean(survey$screen, na.rm = TRUE)

Same number your paper produced. That is the deal this lesson formalises: the by-hand work taught you what the numbers MEAN; today's commands compute them instantly: with one famous omission that catches every beginner in the country.

Theory

The calculator with one missing button

R's statistics keyboard has a key for everything: mean, median, sd, var, quantile...

...except mode. The button LABELLED mode() is a prank: press it and R tells you the storage mode of your data ("numeric"): a programming answer to a statistics question.

The real most-frequent-value must be assembled from two other keys: table() to count, which.max() to crown the winner.

Practical

The whole Unit 1 toolkit, one line each

screen <- c(60, 90, 90, 120, 150, 180, 600, NA)

# central tendency
mean(screen, na.rm = TRUE)      # [1] 184.3   (outlier-dragged, as Unit 1 warned)
median(screen, na.rm = TRUE)    # [1] 120

# dispersion (SAMPLE formulas: divide by n-1)
sd(screen, na.rm = TRUE)
var(screen, na.rm = TRUE)
range(screen, na.rm = TRUE)     # min and max together
quantile(screen, na.rm = TRUE)  # 0% 25% 50% 75% 100%

# the five-number overview + mean, one call:
summary(screen)

# THE QUIRK: mode() is NOT the statistical mode
mode(screen)                    # [1] "numeric"   (storage mode!)

# the real mode, assembled:
t <- table(screen)
names(which.max(t))             # [1] "90"

# group-wise statistics: R's GROUP BY
# tapply(survey$screen, survey$stay, mean, na.rm = TRUE)
# aggregate(screen ~ stay, data = survey, FUN = mean)

Theory

Three habits that make the numbers honest

na.rm = TRUE, always considered: one NA silences every one of these functions (the propagation rule): decide consciously each time.

Know your divisor: sd() and var() use the sample formula (n-1), matching pandas: say "sample SD" in reports, exactly as the dispersion lesson drilled.

Verify once by hand: run sd() on the tiny dataset you computed manually in Unit 1: same answer builds the trust that lets you delegate the arithmetic forever after.

Quiz

A student runs mode(survey$screen) hoping for the most frequent screen time and gets "numeric". What is going on, and what computes the real mode?

  1. mode() returns the STORAGE type in R; the statistical mode is assembled with names(which.max(table(x)))
  2. The data has no repeated values, so R printed its type instead
  3. mode() needs na.rm = TRUE to work
  4. Only factors have modes in R
Show the answer

mode() returns the STORAGE type in R; the statistical mode is assembled with names(which.max(table(x)))

R's mode() answers "how is this stored?": a naming collision with statistics that has trapped generations of students: the classic R gotcha, worth marks precisely because it is so surprising. The real mode: table(x) counts each value, which.max() finds the biggest count, names() reads off the winning value. Options B to D invent behaviours: the function NEVER computes frequency, regardless of data or arguments.

Think first

Group-wise: the dean's one-liner

The dean asks: "average screen time, hostellers vs day scholars, one command." Using tapply(survey$screen, survey$stay, mean, na.rm = TRUE): describe what each of the three arguments does, then tap.

Show the answer

tapply(values, groups, function): take survey$screen (what to summarise), split it by survey$stay (the piles: hostel and day), apply mean to each pile: output is one labelled average per group:

day hostel

142 201

It is the GROUP BY of R (aggregate(screen ~ stay, df, mean) says the same with a formula). You built these piles by hand in the GROUP BY lesson of BCA303: same idea, third syntax, ten seconds.

Watch out

The summary-statistics slips

mode() is never the mode: table + which.max, or your own two-line function.

NA silence: a single missing value returns NA from mean/sd: the na.rm decision is part of the answer, not an afterthought.

quantile() defaults to 0/25/50/75/100%: ask for others explicitly: quantile(x, 0.9) for the 90th percentile.

Theory

Where this plugs in next

These one-liners feed everything left in the subject: frequency tables count the categoricals these functions cannot summarise (next lesson), the commands lesson consolidates the whole surface, and the graphical finale draws what summary() describes. You are also now bilingual within the course: df['screen'].mean() and mean(df$screen) are the same thought: portability is the real syllabus.

Summary

Key takeaways

  • mean, median, sd, var, range, quantile: one call each; summary() = five-number overview + mean.
  • na.rm = TRUE is a conscious decision on every call: NA silences them otherwise.
  • sd()/var() use the sample divisor (n-1): say 'sample SD', matching pandas.
  • mode() returns the STORAGE type: the real mode is names(which.max(table(x))).
  • Group-wise: tapply(values, groups, fun) or aggregate(y ~ group, df, fun): R's GROUP BY.
  • Verify commands once against your Unit 1 hand calculations: then delegate forever.
  • Memory hook: the calculator with one missing (mislabelled) button.

Study this properly

This page is the lesson to read. In Gri-Learn the same topic is a graded deck: the self-checks are scored and your weak topics are tracked. Free to start.

Start this topic

Already have an account? Sign in

More from Working with Data in R

Gri-Learn · syllabus-mapped B.C.A. lessons in English, Hindi and Gujarati

Calculating summary statistics (mean, median, mode, standard deviation) · Statistical Methods and Data Analysis (MDC-03) · Gri-Learn