Theory
Another language? You just learned pandas!
Fair objection. In BCA303 you loaded CSVs, computed means, drew histograms: all in Python. Now this subject hands you R to do... statistics?
Here is the honest sell: Python is a general-purpose language that learned statistics through libraries. R was born statistical: made by statisticians, for statisticians, where a t-test is as native as a for-loop.
When the world's statisticians publish methods, R gets them first. Your syllabus wants you fluent in the field's mother tongue.
Theory
The general hospital vs the speciality clinic
Python is a general hospital: it does everything: websites, games, databases, and (with pandas) data analysis in one wing.
R is the speciality clinic: it does one thing, statistics and its graphics, and the entire building is arranged around it: every instrument within arm's reach, every new technique arriving there first.
Professionals visit both. This subject takes you through the clinic.
Theory
What R is, formally
R is a free, open-source language and environment for statistical computing and graphics, created by Ross Ihaka and Robert Gentleman (University of Auckland, 1990s) as an open implementation of the earlier S language.
Its statistical DNA:
- mean(), median(), sd(), and full tests (t.test(), lm()) are built in, no imports.
- The data frame is a core language type (pandas borrowed the idea and the name).
- Publication-quality graphics ship in the box.
- It is interpreted and interactive: type a command, see the answer: made for exploration.
Theory
CRAN and the package universe
Base R is the clinic's ground floor. CRAN (the Comprehensive R Archive Network) is everything above: 20,000+ peer-contributed packages, one command away:
install.packages("ggplot2")
Stars worth recognising by name: ggplot2 (the famous grammar-of-graphics plots), dplyr (data manipulation), and the Bioconductor collection (genomics). The pattern mirrors Python's pip: base language + package ecosystem: you already know how this world works.
At a glance
Who runs on R
| Field | R at work |
|---|---|
| Academic research | The default language of statistics papers |
| Pharma / clinical trials | Drug-trial analysis and regulatory submissions |
| Epidemiology | Disease models, COVID dashboards |
| Data journalism | BBC and FT charts are R-made |
| Finance | Risk models, time-series forecasting |
| Bioinformatics | Gene-expression analysis (Bioconductor) |
Quiz
Which statement best captures the honest difference between R and Python-with-pandas?
- Python is general-purpose with analysis added via libraries; R is statistics-first with tests, data frames and graphics native to the language
- R is paid and proprietary while Python is free
- R replaces Python: pandas is obsolete
- R cannot read CSV files, so Python is required for data import
Show the answer
Python is general-purpose with analysis added via libraries; R is statistics-first with tests, data frames and graphics native to the language
Both are free and open source (option B is false), both read CSVs happily (D is false: read.csv is built into R), and neither replaces the other (C): real analysts use both, choosing by task. The true distinction is centre of gravity: Python orbits general programming, R orbits statistics: which is why your statistics syllabus teaches it, and why answer A is the line to write in exams.
Think first
Guess the one-liners
In BCA303 you wrote several pandas lines to load survey.csv and get the mean screen time. Before tapping: guess how R spells these two steps: loading a CSV, and taking a mean. (Hint: R names things plainly.)
Show the answer
data <- read.csv("survey.csv")
mean(data$screen_time)
Two lines, zero imports: reading files and averaging are built into the language itself. (That <- arrow is R's assignment operator, and $ picks a column: both explained properly two lessons ahead.) Your pandas instincts transfer almost one-to-one: R just spells them shorter for statistical work.
Watch out
Two mislabels to avoid
"R is just academic": pharma submissions, BBC graphics and bank risk models disagree: the applications table is real industry.
"R" vs "RStudio": R is the language (the engine); RStudio is the IDE you drive it from (next lesson installs both). Saying "I coded it in RStudio" is like saying you wrote Python "in VS Code": name the language, not the editor.
Theory
The road through this unit
The plan for Units 3-5: install R and RStudio, learn the syntax (that arrow, vectors, types), load the CampusPulse survey.csv, then redo everything from Units 1-2 in R: summaries, frequency tables, and finally the bell curve drawn by dnorm(). Every statistic you computed by hand becomes one command: the by-hand work is what lets you TRUST the command's answer.
Summary
Key takeaways
- R = free, open-source language for statistical computing and graphics (Ihaka & Gentleman, from S).
- Statistics is native: mean/sd/t.test/data frames/graphics built in, no imports.
- CRAN hosts 20,000+ packages via install.packages(); ggplot2 and dplyr are the stars.
- Used across research, pharma, epidemiology, journalism, finance, bioinformatics.
- Python = general-purpose + libraries; R = statistics-first: both free, professionals use both.
- R is the language; RStudio is the IDE.
- Memory hook: the speciality clinic next to the general hospital.