Measures of dispersion (range, variance, standard deviation)

Range, variance and standard deviation measure how spread out data is: range uses only the extremes, variance averages squared deviations from the mean, and SD square-roots it back into the data's own units.

10 min read · 10 cards · 2 checks

Read in: English · हिन्दी · ગુજરાતી


Theory

Same average, different animals

Two batches sit the same test. Both average 60.

Batch P: 58, 59, 60, 61, 62. Batch Q: 20, 40, 60, 80, 100.

One is a tight, teachable group; the other spans failure to topper in one room. The mean, our star of last lesson, cannot tell them apart.

Every centre needs a second number: how far do values stray from it? That number is dispersion, and the exam wants three versions of it, by hand.

Theory

Archers around a bullseye

Two archers both average dead-centre. One groups every arrow within a coin's width; the other sprays the whole target face, misses cancelling misses.

Same average position, completely different archers. Dispersion measures the scatter of the arrows, not where their centre lies. A summary without it praises the sprayer and the sniper equally.

Theory

Range: the sixty-second answer

Range = maximum - minimum.

Batch P: 62 - 58 = 4. Batch Q: 100 - 20 = 80. Instant, and already tells the two batches apart.

Its weakness is structural: it consults only two values and ignores the other n-2. One eccentric outlier (a single 600-minute screen-timer) explodes the range while the crowd sits unchanged. Use it as a first glance, never as the verdict.

Theory

Variance: why we square

Better idea: measure every value's deviation from the mean and average them. One catch: raw deviations always sum to zero (the mean is their balance point): +2 and -2 cancel, reporting fake calm.

The fix: square each deviation first: negatives vanish, big misses count extra.

Population variance σ² = Σ(x - mean)² ÷ N.

Sample variance s² divides by (n - 1) instead (Bessel's correction: samples slightly understate spread; n-1 compensates: name it, use it whenever data is a sample).

Theory

Worked, start to finish

Marks: 4, 8, 6, 2 (a population, for simplicity).

1. Mean = 20 ÷ 4 = 5.

2. Deviations: -1, +3, +1, -3. (Sum = 0 ✓ the built-in error check.)

3. Squares: 1, 9, 1, 9. Sum = 20.

4. Variance σ² = 20 ÷ 4 = 5 (units: marks²).

5. SD σ = √5 ≈ 2.24 marks: back in real units, the reportable number.

As a sample: s² = 20 ÷ 3 ≈ 6.67, s ≈ 2.58. Same recipe, different divisor.

Think first

Your turn, full recipe

Screen-time hours of 5 students: 2, 4, 4, 5, 10. On paper: mean, deviations (check they sum to 0), squared deviations, POPULATION variance and SD. Then tap.

Show the answer

Mean = 25 ÷ 5 = 5.

Deviations: -3, -1, -1, 0, +5 (sum 0 ✓).

Squares: 9, 1, 1, 0, 25. Sum = 36.

Variance = 36 ÷ 5 = 7.2 hours². SD = √7.2 ≈ 2.68 hours.

Note how the single 10 contributed 25 of the 36: squaring makes outliers shout. If your deviations did not sum to zero, the mean was wrong: recompute before squaring anything.

Quiz

A student reports "the variance of marks is 5, so scores typically differ from the mean by 5 marks." What is wrong?

  1. Variance is in SQUARED units (marks²); the typical distance is the SD, √5 ≈ 2.24 marks
  2. Nothing: variance and SD are the same number
  3. Variance cannot be computed for marks
  4. The typical distance is the range, not the SD
Show the answer

Variance is in SQUARED units (marks²); the typical distance is the SD, √5 ≈ 2.24 marks

Variance lives in squared units, useful for algebra but unreadable as a distance: nobody strays "5 square marks". The standard deviation square-roots it back to real units, making √5 ≈ 2.24 the honest "typical distance from the mean". Confusing the two is the most common dispersion error in scripts; the range (option D) measures total width, not typical distance.

Watch out

The three dispersion traps

Skipping the squares: raw deviations sum to zero: every dataset would claim zero spread.

Divisor confusion: population ÷ N, sample ÷ (n-1). State which you used: R and pandas default to the SAMPLE formula, plain calculators often to population.

Reporting variance as a distance: convert to SD before interpreting. And remember the range consulted only two values: never call it robust.

Theory

Where SD quietly runs things

Result moderation scales marks using mean and SD. Quality control stops a production line when SD creeps up. Cricket commentators' "consistent batsman" is a low-SD claim. And two lessons ahead, the bell curve turns SD into a ruler: 68% of data within 1 SD, 95% within 2: today's hand-computed number becomes the unit the whole normal distribution is measured in. (Comparing spreads across units? Divide SD by mean: the coefficient of variation.)

Summary

Key takeaways

  • Dispersion answers "how scattered?": the mean alone cannot separate tight from wild data.
  • Range = max - min: instant, but two values only and outlier-fragile.
  • Deviations from the mean sum to zero: squaring kills cancellation and weights big misses.
  • Variance = average squared deviation (÷N population, ÷(n-1) sample: Bessel).
  • SD = √variance: back in real units, the number you report and interpret.
  • Deviation sum = 0 is the free arithmetic check before squaring.
  • Memory hook: archers around a bullseye: same centre, different scatter.

Study this properly

This page is the lesson to read. In Gri-Learn the same topic is a graded deck: the self-checks are scored and your weak topics are tracked. Free to start.

Start this topic

Already have an account? Sign in

More from Basic concepts of statistic

Gri-Learn · syllabus-mapped B.C.A. lessons in English, Hindi and Gujarati

Measures of dispersion (range, variance, standard deviation) · Statistical Methods and Data Analysis (MDC-03) · Gri-Learn