Theory
The survey that surveyed itself
First attempt at the CampusPulse survey: a volunteer stood outside the library at 8 am and surveyed the first 60 students who agreed.
Average reported study time: 5.2 hours a day. The principal beamed. The staff room laughed: 8 am library visitors are the college's most studious tribe.
The survey did not measure the college; it measured who was easy to catch. Sampling technique, the topic everyone skips, is the difference between statistics and self-congratulation.
Theory
Stir the pot, or spoon every dish
Back to the dal pot. Random sampling is stirring thoroughly, then tasting one spoon: every drop had an equal chance to be on it.
But a thali is not one pot: dal, sabzi, rice, salad. One stirred spoon of dal says nothing about the salad. For a thali you taste a little of each dish, in proportion: that is stratified sampling: divide first, then sample fairly inside every division.
Theory
Simple random sampling, formally
Simple random sampling (SRS): every population member has an equal chance of selection.
Recipe: build the sampling frame (a numbered list of all 3000 students), then pick 60 using the lottery method (chits in a drum) or random numbers (tables or a generator).
Strength: no human preference can sneak in: the method is unbiased by construction.
Weakness: pure chance can still hand you a lopsided sample: 60 random students might include only 9 hostellers when the college has 40%.
Theory
Stratified sampling, formally
Stratified sampling: divide the population into non-overlapping strata (groups that matter for your question), then draw a random sample inside each stratum, usually proportional to its size.
Worked allocation: 3000 students = 1800 day scholars + 1200 hostellers. For a sample of 60:
- Day scholars: (1800/3000) × 60 = 36
- Hostellers: (1200/3000) × 60 = 24
Both groups guaranteed present, in true proportion: no unlucky lottery can erase the hostellers.
At a glance
The techniques compared
| Technique | How | Strength / Weakness |
|---|---|---|
| Simple random | Equal chance for all (lottery, random numbers) | Unbiased / small groups may be missed by luck |
| Stratified | Divide into strata, random sample within each | Every stratum represented / needs strata known upfront |
| Systematic | Every kth from a list, random start | Easy in the field / hidden list patterns can bias |
| Convenience | Whoever is easy to reach | Cheap / BIASED, not a probability method |
Quiz
The library-queue survey (first 60 willing students at 8 am) is which technique, and what is its core defect?
- Convenience sampling: selection depended on being easy to reach, so studious students were over-represented (bias)
- Simple random sampling: the volunteer did not know the students personally
- Stratified sampling: the library queue is a stratum
- Systematic sampling: students were taken in order of arrival
Show the answer
Convenience sampling: selection depended on being easy to reach, so studious students were over-represented (bias)
Equal chance is the test, and a hosteller asleep in their room at 8 am had zero chance: this is convenience sampling, and its defect is bias: a systematic tilt toward accessible (here, studious) members. Not knowing names does not make it random (option B); one self-selected queue is not a designed stratum (C); and arrival order without a sampling frame is not systematic sampling (D).
Think first
Design the honest version
Redesign CampusPulse properly: the college has 3000 students: 1500 in Year 1, 900 in Year 2, 600 in Year 3, and year of study strongly affects screen time. Sample size stays 60. Before tapping: which technique, and how many from each year?
Show the answer
Stratified sampling by year (the strata differ on the variable being measured, the textbook trigger for stratification):
- Year 1: (1500/3000) × 60 = 30
- Year 2: (900/3000) × 60 = 18
- Year 3: (600/3000) × 60 = 12
...then simple random sampling WITHIN each year (chits or random numbers over each year's roll list). Proportional allocation + random selection inside strata: that pairing is the full-marks answer.
Watch out
The bias that size cannot cure
The deadliest misconception: "our sample is biased, so let us survey MORE people the same way." A bigger convenience sample is a more confident wrong answer: 600 library-queue students still exclude the sleepers.
Bias is a property of the method, not the count. Fix the selection process first; only then does size help (how much it helps is the next lesson's story).
Theory
Where you have seen this fail publicly
Election exit polls that missed rural voters, app-store ratings (only the delighted and the furious bother), "90% of dentists" surveys with 10 dentists: sampling sins power most statistics scandals. Every claim you will ever audit starts with one question you now know to ask: how exactly was the sample chosen?
Summary
Key takeaways
- The sampling METHOD decides whether a sample can speak for the population.
- Simple random sampling: equal chance for all, via a sampling frame + lottery/random numbers.
- Stratified: divide into strata, random-sample within each, usually proportional to stratum size.
- Stratify when subgroups differ on the measured variable; allocation = (stratum/population) × sample size.
- Convenience sampling is biased by construction; bigger biased samples stay biased.
- Systematic (every kth) exists; watch for list patterns.
- Memory hook: stir the pot, or spoon every dish of the thali.