Theory
Nobody asks the whole college
The principal wants to know the average daily screen time of the college's 3000 students. Option one: interview all 3000. Two weeks, four staff, exam season ruined.
Option two, what statisticians actually do: carefully pick 60 students, survey them in an afternoon, and use their answers to speak about all 3000.
That move, measuring a part to understand the whole, is the founding act of statistics, and this subject's first job is naming its pieces precisely.
Theory
One spoon judges the whole pot
A cook never drinks the whole pot of dal to check the salt. One stirred spoonful decides for the entire pot.
The pot is the population: everything you want to know about. The spoonful is the sample: the part you actually taste. And the stirring is the honesty condition: an unstirred spoon from the top tastes only the tadka, and lies about the pot.
Theory
The four exam words
- Population: the ENTIRE group under study (all 3000 students).
- Sample: the subset actually measured (the 60 surveyed).
- Parameter: a number describing the population (the TRUE average screen time of all 3000: usually unknown).
- Statistic: the same kind of number computed from the sample (the 60 students' average: what we can actually calculate).
Statistics as a discipline is the art of using statistics (sample numbers) to estimate parameters (population truths). A full measurement of everyone is a census: costly, slow, sometimes impossible (testing every bulb destroys the stock).
Theory
Variables: what the survey records
Each surveyed student answers four questions. Each answer is a variable, and every variable is one of two families:
Categorical (qualitative): labels.
- Nominal: no natural order: city, blood group.
- Ordinal: ordered labels: grade (A > B > C), satisfaction (high/medium/low).
Numerical (quantitative): numbers you can do maths on.
- Discrete: countable jumps: subjects failed (0, 1, 2...).
- Continuous: any precision: height (167.5 cm), screen time (143.7 min).
At a glance
The CampusPulse survey, classified
| Variable | Family | Sub-type |
|---|---|---|
| City | Categorical | Nominal (no order) |
| Grade (A/B/C) | Categorical | Ordinal (ordered) |
| Subjects failed | Numerical | Discrete (countable) |
| Height (cm) | Numerical | Continuous (measurable) |
| Screen time (min) | Numerical | Continuous |
Quiz
The 60 surveyed students report an average screen time of 4.2 hours. In precise exam language, what IS that 4.2?
- A statistic: computed from the sample, used to estimate the population parameter
- A parameter: any computed average is a parameter
- The population value: 60 students is enough to make it exact
- A census result
Show the answer
A statistic: computed from the sample, used to estimate the population parameter
Numbers from the SAMPLE are statistics; the unknown true value across all 3000 is the parameter the statistic estimates. The pairing (sample→statistic, population→parameter, both starting with the same letters) is the memory device examiners expect. A sample average never becomes exact truth (option C), and a census would mean all 3000 were measured (option D).
Think first
Classify these four
From a hostel survey: (1) mess rating: poor/okay/good; (2) number of chapatis eaten at dinner; (3) mother tongue; (4) commute time in minutes. Assign each its family AND sub-type before tapping.
Show the answer
(1) Categorical ordinal: the labels have an order.
(2) Numerical discrete: you count chapatis, no one eats 2.7.
(3) Categorical nominal: languages have no order.
(4) Numerical continuous: 23.5 minutes is meaningful.
The two-question test: is it a label or a number? Then: ordered or not / counted or measured? Every classification question in this subject falls to those two questions asked in order.
Watch out
The nonsense-maths trap
Code cities as 1 = Surat, 2 = Navsari, 3 = Bardoli and a spreadsheet will happily report "average city = 1.8". Meaningless: nominal labels wearing number costumes are still labels. Same for roll numbers and pin codes.
The variable's TYPE decides its legal maths: means and SDs belong to numerical variables; categorical variables get counts and percentages. Exams test this with "which measure is appropriate?" questions.
Theory
Why this vocabulary runs the whole subject
Every later topic stands on today's words: central tendency and dispersion summarise numerical variables; frequency tables count categorical ones; sampling techniques are about choosing the 60 honestly; the bell curve describes how sample statistics behave. And in Unit 3, R will make you DECLARE these types (numeric, factor): the computer refuses to guess what you have not classified.
Summary
Key takeaways
- Population = everyone of interest; sample = the part actually measured.
- Parameter describes the population (unknown); statistic is computed from the sample (estimates it).
- Census = measure everyone: costly, slow, sometimes impossible.
- Categorical variables are labels: nominal (unordered) or ordinal (ordered).
- Numerical variables are numbers: discrete (counted) or continuous (measured).
- Type decides legal maths: no averaging city codes or roll numbers.
- Memory hook: one stirred spoon judges the whole pot.