Theory
The subject ends with one picture
Twenty-five lessons of numbers deserve a closing image. The dean asks the final question: "you keep SAYING heights are normal and screen time is not: show me."
The proof is one chart: your real data as a histogram, and the theoretical bell curve drawn over it. Where the curve hugs the bars, the normal model holds and every pnorm answer is trustworthy. Where it does not, the model, not the data, must yield.
Today: the R commands that draw that verdict.
Theory
The tailor's fitting
The bell curve is a ready-made suit (two measurements: mean and sd). Your data is the customer.
Overlaying curve(dnorm(...)) on the histogram is the fitting: heights slip into the suit almost perfectly; screen time bulges left and drags a long right sleeve: no tailoring saves it.
Statistics is honest exactly here: fit the model TO the data, and show the fitting photo, never assume the suit fits.
Practical
The finale: histogram, groups, and the bell overlay
survey <- read.csv("survey_clean.csv")
# 1. the data's shape
hist(survey$height, breaks = 10,
main = "Heights of surveyed students",
xlab = "Height (cm)", ylab = "Frequency")
# 2. groups side by side: the formula reads 'screen BY stay'
boxplot(screen ~ stay, data = survey,
main = "Screen time by residence", ylab = "Minutes")
# 3. THE BELL OVERLAY: density scale first, then the curve
hist(survey$height, freq = FALSE, # y-axis = density now
main = "Heights vs the normal model", xlab = "Height (cm)")
curve(dnorm(x, mean = mean(survey$height, na.rm = TRUE),
sd = sd(survey$height, na.rm = TRUE)),
add = TRUE, col = "red", lwd = 2) # add = TRUE: draw ON TOP
abline(v = mean(survey$height, na.rm = TRUE), lty = 2)
# same overlay on screen time: watch the suit NOT fit
hist(survey$screen, freq = FALSE, main = "Screen time vs the bell")
curve(dnorm(x, mean(survey$screen, na.rm = TRUE),
sd(survey$screen, na.rm = TRUE)),
add = TRUE, col = "red", lwd = 2)
Theory
Reading the two fittings
Heights: bars rise and fall symmetrically; the red bell tracks them closely; the dashed mean line splits the picture evenly. Verdict: normal model fits: every pnorm answer about heights is on solid ground.
Screen time: bars pile up left with a long right tail (the binge-watchers); the bell floats over empty space on the right and undershoots the left pile. Verdict: not normal: report medians and IQRs, and treat any pnorm answer with suspicion.
Two pictures, and every claim this subject made becomes visible.
Quiz
A student runs hist(heights) then curve(dnorm(x, 168, 6), add = TRUE) and the red curve crawls flat along the x-axis, invisible. What went wrong?
- The histogram is on the FREQUENCY scale (counts up to ~15) while dnorm's density peaks near 0.07: hist needs freq = FALSE to share the density scale
- dnorm cannot be plotted with curve()
- add = TRUE erased the curve
- The sd of 6 is too small to draw
Show the answer
The histogram is on the FREQUENCY scale (counts up to ~15) while dnorm's density peaks near 0.07: hist needs freq = FALSE to share the density scale
Frequencies count students (0 to 15-ish); densities integrate to 1 (peaking around 0.07 here): drawn on one chart, the density curve is a flat whisper under frequency mountains. freq = FALSE rescales the histogram's y-axis to density so both speak the same units: the single most common overlay bug. add = TRUE does the opposite of erasing (it preserves the histogram); without it, curve() would REPLACE the plot: the sibling trap.
Think first
Deliver the verdict
You show the dean both overlay charts. Compose the two-sentence verdict: one sentence per variable, each naming the evidence in the picture AND the practical consequence for reporting. Then tap.
Show the answer
Heights: the histogram is symmetric and the fitted bell tracks it closely, so the normal model applies: means, SDs and pnorm-based percentages (e.g. share above 180 cm) are trustworthy.
Screen time: the distribution is right-skewed with a long tail the bell cannot follow, so normal-based answers would mislead: we report the median and IQR, and flag the heavy-user tail separately.
Evidence + consequence, twice: that is a statistics report in miniature, and the exam's dream answer.
Watch out
The closing traps
freq = FALSE before any dnorm overlay: frequency and density scales never mix.
add = TRUE on curve(): without it the bell replaces your histogram.
Labels are marks: main, xlab, ylab on every chart: an unlabelled chart is an unfinished answer.
The model serves the data: never force normal conclusions onto a shape that rejected the suit.
Theory
CampusPulse, closed
Look back along the whole arc: you planned a sample, computed its statistics by hand, learned the probability behind them, imported and cleaned real data in R, and ended by drawing the bell curve over your own survey. That loop (design, collect, clean, summarise, model, visualise, verdict) is not a syllabus: it is the actual daily shape of data work. BCA302 complete: qqnorm() and formal tests await whoever stays curious.
Summary
Key takeaways
- hist(x, breaks = n) shows the data's shape; main/xlab/ylab are mandatory manners.
- boxplot(y ~ group, data = df): side-by-side five-number comparisons: the formula reads 'y by group'.
- The overlay recipe: hist(x, freq = FALSE) then curve(dnorm(x, mean, sd), add = TRUE).
- freq = FALSE aligns the scales; add = TRUE preserves the histogram: both non-negotiable.
- Close fit = normal model trustworthy; visible skew = report median/IQR instead.
- abline() adds reference lines; qqnorm() is the formal normality tool beyond this course.
- Memory hook: the tailor's fitting: does the ready-made bell suit your data?