Sampling distributions

Take many samples and each gives a slightly different mean: those means form the sampling distribution, whose spread shrinks as 1/√n and whose shape turns normal by the Central Limit Theorem.

10 min read · 10 cards · 2 checks

Read in: English · हिन्दी · ગુજરાતી


Theory

Two honest surveys, two different answers

Your stratified CampusPulse survey reports average screen time 178 minutes. A rival team runs the SAME method on a fresh 60 students: 191 minutes.

Nobody cheated. Nobody erred. Two honest spoonfuls from one pot simply differ.

So which is "the" answer? Wrong question. The right question, the one this lesson answers: how much do honest sample means wobble around the truth? Statistics' most elegant idea lives here.

Theory

A thousand spoonfuls, plotted

Imagine a thousand teams, each surveying their own random 60 students and reporting a mean. Plot all thousand means as a histogram.

That histogram is the sampling distribution of the mean: not the distribution of students, but the distribution of survey results. Its centre sits on the true value; its narrowness tells you how much any single honest survey can be trusted. You will only ever run ONE survey, but knowing how the thousand behave prices your one.

Theory

The three distributions (keep them apart)

Exams love to blur these; refuse:

1. Population distribution: all 3000 students' screen times (possibly skewed, unknown).

2. Sample data distribution: YOUR 60 students' values (looks roughly like the population).

3. Sampling distribution of the mean: the values of x̄ across all possible samples of size n: a distribution of RESULTS, not of students.

The third is the new object, and it has two beautiful, exam-worthy properties.

Theory

Property 1: centred right. Property 2: predictable spread.

Centre: the mean of all possible sample means equals the population mean μ: the sample mean is an unbiased estimator: honest spoonfuls do not systematically over- or under-taste.

Spread: the standard deviation of the sample means, called the standard error (SE):

SE = σ / √n

Worked: σ = 60 min, n = 36 → SE = 60/√36 = 10 min. Sample means typically land within about 10 minutes of the truth. Note the √: quadrupling n only halves SE: precision is bought at a square-root price.

Theory

The Central Limit Theorem

Now the miracle. Screen time itself is right-skewed (a few binge-watchers). Surely the sample means inherit that skew?

No. The Central Limit Theorem (CLT): for sufficiently large n (rule of thumb n ≥ 30), the sampling distribution of the mean is approximately NORMAL, whatever shape the population has.

Averages of many values smooth individual quirks away. This is WHY the bell curve owns statistics: even when data is ugly, means behave normally, so next lesson's 68-95-99.7 ruler applies to survey results everywhere.

Quiz

Screen time: μ = 180, σ = 60. With samples of n = 36 (SE = 10), one team reports a sample mean of 210. How should you react?

  1. 210 is 3 standard errors above 180: honest samples rarely land that far out, so investigate the sample
  2. Perfectly ordinary: 210 is within one population SD (60) of 180
  3. Impossible: sample means always equal the population mean
  4. The team should just survey 36 more students to confirm
Show the answer

210 is 3 standard errors above 180: honest samples rarely land that far out, so investigate the sample

Sample MEANS are judged against the standard error, not the population SD: (210 - 180)/10 = 3 SEs out, an event honest sampling produces well under 1% of the time: suspect a biased frame or an error. Option B applies the wrong ruler (σ is for individual students; σ/√n is for means): the exact confusion this topic exists to cure. Means wobble (C is false), and more data cannot retroactively validate a flawed sample (D echoes last lesson).

Think first

Buy more precision, at the square-root price

SE with n = 36 was 10 minutes. The principal wants SE = 5 minutes. Before tapping: what sample size is needed, and why is the answer NOT 72?

Show the answer

SE = σ/√n → 5 = 60/√n → √n = 12 → n = 144.

Halving the error needs four times the students, not double, because n sits under a square root. This √n economics is a real planning fact: each extra digit of precision gets exponentially more expensive, which is why professional surveys stop at "good enough" rather than "perfect".

Watch out

The two classic confusions

SD vs SE: σ describes the spread of students; σ/√n describes the spread of sample means. Same units, different objects: mixing them is the most-penalised error in this chapter.

What the CLT normalises: the sampling distribution of the MEAN: never the raw data. Screen time stays skewed forever; only the survey averages turn bell-shaped. Writing "large samples make data normal" loses the mark instantly.

Theory

Why polls dare to predict crores from thousands

An exit poll asks ~10,000 voters and speaks about 90 crore, with a stated "margin of error": that margin IS a couple of standard errors. The CLT guarantees the poll's mean behaves normally; SE = σ/√n prices its wobble. Next lesson hands you the bell curve's ruler (68-95-99.7 and z-scores): the same ruler you will lay on sample means for the rest of the subject.

Summary

Key takeaways

  • A statistic varies sample to sample; the sampling distribution of the mean is the distribution of x̄ across all samples of size n.
  • It is centred on μ (the sample mean is unbiased).
  • Its spread is the STANDARD ERROR: SE = σ/√n: quadruple n to halve SE.
  • CLT: for n ≥ 30, sample means are approximately normal REGARDLESS of the population's shape.
  • Judge individual values with σ; judge sample means with SE.
  • CLT normalises the means, never the raw data.
  • Memory hook: a thousand spoonfuls, plotted.

Study this properly

This page is the lesson to read. In Gri-Learn the same topic is a graded deck: the self-checks are scored and your weak topics are tracked. Free to start.

Start this topic

Already have an account? Sign in

More from Data Representation and Sampling technique

Gri-Learn · syllabus-mapped B.C.A. lessons in English, Hindi and Gujarati

Sampling distributions · Statistical Methods and Data Analysis (MDC-03) · Gri-Learn