Concepts of normal distribution

R speaks normal-distribution through four prefixed commands: dnorm gives the curve's height, pnorm the area (probability) left of a value, qnorm the value at a given percentile, and rnorm random samples.

10 min read · 9 cards · 2 checks

Read in: English · हिन्दी · ગુજરાતી


Theory

Replacing the z-table with four commands

In the bell-curve lesson you answered "what share of students are taller than 180 cm?" with the 68-95-99.7 rule: about 2.5%, because 180 sat exactly 2 SDs out.

But what about 178 cm? Or 183? The rule only knows whole rings; between them, statisticians once flipped through printed z-tables.

R retired those tables with a four-command family. Learn the prefixes d, p, q, r once, and every normal-curve question becomes one line.

Theory

Four questions to one bell

Stand before the bell curve and ask:

  • dnorm: how HIGH is the curve here? (density: the drawing question)
  • pnorm: how much area lies LEFT of here? (probability: the exam question)
  • qnorm: WHERE does 90% of the area end? (quantile: the cutoff question)
  • rnorm: hand me RANDOM students from this bell (simulation)

Same bell, four doors. pnorm and qnorm are inverses: one turns values into areas, the other turns areas back into values.

Practical

Heights: mu = 168, sd = 6, every question answered

# P(height <= 174): area LEFT of 174
pnorm(174, mean = 168, sd = 6)      # [1] 0.841

# P(taller than 180): right tail = 1 - left area
1 - pnorm(180, 168, 6)              # [1] 0.0228   (~the 2.5% ring!)
pnorm(180, 168, 6, lower.tail = FALSE)   # same, spelled directly

# P(between 162 and 174): big area minus small area
pnorm(174, 168, 6) - pnorm(162, 168, 6)  # [1] 0.683  (the 68% ring!)

# cutoff for the tallest 10%: qnorm is pnorm's inverse
qnorm(0.90, 168, 6)                 # [1] 175.7 cm

# defaults are the STANDARD normal (mean 0, sd 1):
pnorm(2)                            # [1] 0.977  = P(z <= 2)

# simulate 60 students; set.seed makes it reproducible
set.seed(42)
heights <- rnorm(60, 168, 6)
mean(heights)                       # close to 168, not exact: sampling!

# dnorm: curve HEIGHT (for drawing, next lesson), not probability
dnorm(168, 168, 6)                  # [1] 0.066

Theory

The three question patterns

Every exam question maps to one pattern:

  • Left area ("at most x"): pnorm(x, mean, sd).
  • Right area ("more than x"): 1 - pnorm(x, ...) (or lower.tail = FALSE): the 1-minus is the most-forgotten keystroke in R statistics.
  • Between a and b: pnorm(b, ...) - pnorm(a, ...).

And the reverse direction ("what value marks the top 10%?"): qnorm(0.90, ...): note it takes the area to the LEFT, so top 10% means 0.90, not 0.10.

Quiz

Scores are normal with mean 60, sd 8. Which command computes P(score > 70), and roughly what does it return?

  1. 1 - pnorm(70, 60, 8): about 0.106
  2. pnorm(70, 60, 8): about 0.894
  3. dnorm(70, 60, 8): the probability of exactly 70
  4. qnorm(70, 60, 8): the 70th percentile
Show the answer

1 - pnorm(70, 60, 8): about 0.106

pnorm gives the LEFT area, so P(above 70) needs the complement: 1 - pnorm(70, 60, 8) ≈ 1 - 0.894 = 0.106. Option B is the left area itself: the answer to the OPPOSITE question, and the most common wrong submission. dnorm is curve height (a continuous variable has zero probability at an exact point: the distributions lesson's law), and qnorm expects an AREA (0 to 1) as input, so qnorm(70, ...) is a type error in thought.

Think first

The scholarship cutoff, both directions

CGPA is roughly normal: mean 7.0, sd 0.8. Two tasks, one line each: (1) what share of students exceed CGPA 8.5? (2) The top 5% get a scholarship: what CGPA is the cutoff? Draft both before tapping.

Show the answer

1. 1 - pnorm(8.5, 7.0, 0.8) ≈ 0.030: about 3% exceed 8.5.

2. qnorm(0.95, 7.0, 0.8) ≈ 8.32: the scholarship line.

Note the mirror: task 1 goes value → area (pnorm), task 2 goes area → value (qnorm, fed 0.95 because 5% at the top means 95% to the left). Reading which direction a question travels is the entire skill of this command family.

Watch out

The three normal-command slips

dnorm is not probability: it is curve height (density): probabilities for continuous variables live in AREAS: pnorm territory.

The forgotten 1-minus: every "more than" question needs the complement (or lower.tail = FALSE).

Model-check first: pnorm answers are exact only if the data is (approximately) normal: histogram + mean≈median check before trusting them: screen time would fail, heights pass.

Theory

The rule vs the exact answer

Notice 1 - pnorm(180, 168, 6) gave 0.0228 where the empirical rule estimated 2.5%: the rule is the mental approximation, pnorm the precise instrument: quote the rule in reasoning, compute with pnorm in practice. One lesson remains: drawing everything: hist for your data, curve(dnorm(...)) to lay the theoretical bell over it, and the CampusPulse story closes.

Summary

Key takeaways

  • The d-p-q-r family: dnorm height, pnorm left area, qnorm area-to-value, rnorm random draws.
  • Defaults mean = 0, sd = 1: the standard normal; pnorm(2) ≈ 0.977.
  • Right tails: 1 - pnorm(x, ...) or lower.tail = FALSE; between: pnorm(b) - pnorm(a).
  • qnorm takes the LEFT area: top 10% cutoff = qnorm(0.90, ...).
  • pnorm is the exact instrument behind the 68-95-99.7 approximations.
  • Check normality (hist shape, mean ≈ median) before trusting exact answers.
  • Memory hook: four doors to one bell: height, area, cutoff, sample.

Study this properly

This page is the lesson to read. In Gri-Learn the same topic is a graded deck: the self-checks are scored and your weak topics are tracked. Free to start.

Start this topic

Already have an account? Sign in

More from Working with Data in R

Gri-Learn · syllabus-mapped B.C.A. lessons in English, Hindi and Gujarati

Concepts of normal distribution · Statistical Methods and Data Analysis (MDC-03) · Gri-Learn