Commands to explore and view data distributions graphically (Bell curve)

hist() shows your data's real shape, boxplot(y ~ group) compares groups on one axis, and curve(dnorm(...), add = TRUE) lays the theoretical bell over the histogram: the subject's closing picture.

10 min read · 9 cards · 2 checks

Read in: English · हिन्दी · ગુજરાતી


Theory

The subject ends with one picture

Twenty-five lessons of numbers deserve a closing image. The dean asks the final question: "you keep SAYING heights are normal and screen time is not: show me."

The proof is one chart: your real data as a histogram, and the theoretical bell curve drawn over it. Where the curve hugs the bars, the normal model holds and every pnorm answer is trustworthy. Where it does not, the model, not the data, must yield.

Today: the R commands that draw that verdict.

Theory

The tailor's fitting

The bell curve is a ready-made suit (two measurements: mean and sd). Your data is the customer.

Overlaying curve(dnorm(...)) on the histogram is the fitting: heights slip into the suit almost perfectly; screen time bulges left and drags a long right sleeve: no tailoring saves it.

Statistics is honest exactly here: fit the model TO the data, and show the fitting photo, never assume the suit fits.

Practical

The finale: histogram, groups, and the bell overlay

survey <- read.csv("survey_clean.csv")

# 1. the data's shape
hist(survey$height, breaks = 10,
     main = "Heights of surveyed students",
     xlab = "Height (cm)", ylab = "Frequency")

# 2. groups side by side: the formula reads 'screen BY stay'
boxplot(screen ~ stay, data = survey,
        main = "Screen time by residence", ylab = "Minutes")

# 3. THE BELL OVERLAY: density scale first, then the curve
hist(survey$height, freq = FALSE,        # y-axis = density now
     main = "Heights vs the normal model", xlab = "Height (cm)")
curve(dnorm(x, mean = mean(survey$height, na.rm = TRUE),
               sd  = sd(survey$height,  na.rm = TRUE)),
      add = TRUE, col = "red", lwd = 2)  # add = TRUE: draw ON TOP
abline(v = mean(survey$height, na.rm = TRUE), lty = 2)

# same overlay on screen time: watch the suit NOT fit
hist(survey$screen, freq = FALSE, main = "Screen time vs the bell")
curve(dnorm(x, mean(survey$screen, na.rm = TRUE),
               sd(survey$screen, na.rm = TRUE)),
      add = TRUE, col = "red", lwd = 2)

Theory

Reading the two fittings

Heights: bars rise and fall symmetrically; the red bell tracks them closely; the dashed mean line splits the picture evenly. Verdict: normal model fits: every pnorm answer about heights is on solid ground.

Screen time: bars pile up left with a long right tail (the binge-watchers); the bell floats over empty space on the right and undershoots the left pile. Verdict: not normal: report medians and IQRs, and treat any pnorm answer with suspicion.

Two pictures, and every claim this subject made becomes visible.

Quiz

A student runs hist(heights) then curve(dnorm(x, 168, 6), add = TRUE) and the red curve crawls flat along the x-axis, invisible. What went wrong?

  1. The histogram is on the FREQUENCY scale (counts up to ~15) while dnorm's density peaks near 0.07: hist needs freq = FALSE to share the density scale
  2. dnorm cannot be plotted with curve()
  3. add = TRUE erased the curve
  4. The sd of 6 is too small to draw
Show the answer

The histogram is on the FREQUENCY scale (counts up to ~15) while dnorm's density peaks near 0.07: hist needs freq = FALSE to share the density scale

Frequencies count students (0 to 15-ish); densities integrate to 1 (peaking around 0.07 here): drawn on one chart, the density curve is a flat whisper under frequency mountains. freq = FALSE rescales the histogram's y-axis to density so both speak the same units: the single most common overlay bug. add = TRUE does the opposite of erasing (it preserves the histogram); without it, curve() would REPLACE the plot: the sibling trap.

Think first

Deliver the verdict

You show the dean both overlay charts. Compose the two-sentence verdict: one sentence per variable, each naming the evidence in the picture AND the practical consequence for reporting. Then tap.

Show the answer

Heights: the histogram is symmetric and the fitted bell tracks it closely, so the normal model applies: means, SDs and pnorm-based percentages (e.g. share above 180 cm) are trustworthy.

Screen time: the distribution is right-skewed with a long tail the bell cannot follow, so normal-based answers would mislead: we report the median and IQR, and flag the heavy-user tail separately.

Evidence + consequence, twice: that is a statistics report in miniature, and the exam's dream answer.

Watch out

The closing traps

freq = FALSE before any dnorm overlay: frequency and density scales never mix.

add = TRUE on curve(): without it the bell replaces your histogram.

Labels are marks: main, xlab, ylab on every chart: an unlabelled chart is an unfinished answer.

The model serves the data: never force normal conclusions onto a shape that rejected the suit.

Theory

CampusPulse, closed

Look back along the whole arc: you planned a sample, computed its statistics by hand, learned the probability behind them, imported and cleaned real data in R, and ended by drawing the bell curve over your own survey. That loop (design, collect, clean, summarise, model, visualise, verdict) is not a syllabus: it is the actual daily shape of data work. BCA302 complete: qqnorm() and formal tests await whoever stays curious.

Summary

Key takeaways

  • hist(x, breaks = n) shows the data's shape; main/xlab/ylab are mandatory manners.
  • boxplot(y ~ group, data = df): side-by-side five-number comparisons: the formula reads 'y by group'.
  • The overlay recipe: hist(x, freq = FALSE) then curve(dnorm(x, mean, sd), add = TRUE).
  • freq = FALSE aligns the scales; add = TRUE preserves the histogram: both non-negotiable.
  • Close fit = normal model trustworthy; visible skew = report median/IQR instead.
  • abline() adds reference lines; qqnorm() is the formal normality tool beyond this course.
  • Memory hook: the tailor's fitting: does the ready-made bell suit your data?

Study this properly

This page is the lesson to read. In Gri-Learn the same topic is a graded deck: the self-checks are scored and your weak topics are tracked. Free to start.

Start this topic

Already have an account? Sign in

More from Working with Data in R

Gri-Learn · syllabus-mapped B.C.A. lessons in English, Hindi and Gujarati

Commands to explore and view data distributions graphically (Bell curve) · Statistical Methods and Data Analysis (MDC-03) · Gri-Learn