Population vs. sample, variables (categorical vs. numerical), datatypes

A population is every unit you care about, a sample is the part you actually measure, and every variable you record is either categorical (labels) or numerical (counts and measurements).

9 min read · 10 cards · 2 checks

Read in: English · हिन्दी · ગુજરાતી


Theory

Nobody asks the whole college

The principal wants to know the average daily screen time of the college's 3000 students. Option one: interview all 3000. Two weeks, four staff, exam season ruined.

Option two, what statisticians actually do: carefully pick 60 students, survey them in an afternoon, and use their answers to speak about all 3000.

That move, measuring a part to understand the whole, is the founding act of statistics, and this subject's first job is naming its pieces precisely.

Theory

One spoon judges the whole pot

A cook never drinks the whole pot of dal to check the salt. One stirred spoonful decides for the entire pot.

The pot is the population: everything you want to know about. The spoonful is the sample: the part you actually taste. And the stirring is the honesty condition: an unstirred spoon from the top tastes only the tadka, and lies about the pot.

Theory

The four exam words

  • Population: the ENTIRE group under study (all 3000 students).
  • Sample: the subset actually measured (the 60 surveyed).
  • Parameter: a number describing the population (the TRUE average screen time of all 3000: usually unknown).
  • Statistic: the same kind of number computed from the sample (the 60 students' average: what we can actually calculate).

Statistics as a discipline is the art of using statistics (sample numbers) to estimate parameters (population truths). A full measurement of everyone is a census: costly, slow, sometimes impossible (testing every bulb destroys the stock).

Theory

Variables: what the survey records

Each surveyed student answers four questions. Each answer is a variable, and every variable is one of two families:

Categorical (qualitative): labels.

  • Nominal: no natural order: city, blood group.
  • Ordinal: ordered labels: grade (A > B > C), satisfaction (high/medium/low).

Numerical (quantitative): numbers you can do maths on.

  • Discrete: countable jumps: subjects failed (0, 1, 2...).
  • Continuous: any precision: height (167.5 cm), screen time (143.7 min).

At a glance

The CampusPulse survey, classified

VariableFamilySub-type
CityCategoricalNominal (no order)
Grade (A/B/C)CategoricalOrdinal (ordered)
Subjects failedNumericalDiscrete (countable)
Height (cm)NumericalContinuous (measurable)
Screen time (min)NumericalContinuous

Quiz

The 60 surveyed students report an average screen time of 4.2 hours. In precise exam language, what IS that 4.2?

  1. A statistic: computed from the sample, used to estimate the population parameter
  2. A parameter: any computed average is a parameter
  3. The population value: 60 students is enough to make it exact
  4. A census result
Show the answer

A statistic: computed from the sample, used to estimate the population parameter

Numbers from the SAMPLE are statistics; the unknown true value across all 3000 is the parameter the statistic estimates. The pairing (sample→statistic, population→parameter, both starting with the same letters) is the memory device examiners expect. A sample average never becomes exact truth (option C), and a census would mean all 3000 were measured (option D).

Think first

Classify these four

From a hostel survey: (1) mess rating: poor/okay/good; (2) number of chapatis eaten at dinner; (3) mother tongue; (4) commute time in minutes. Assign each its family AND sub-type before tapping.

Show the answer

(1) Categorical ordinal: the labels have an order.

(2) Numerical discrete: you count chapatis, no one eats 2.7.

(3) Categorical nominal: languages have no order.

(4) Numerical continuous: 23.5 minutes is meaningful.

The two-question test: is it a label or a number? Then: ordered or not / counted or measured? Every classification question in this subject falls to those two questions asked in order.

Watch out

The nonsense-maths trap

Code cities as 1 = Surat, 2 = Navsari, 3 = Bardoli and a spreadsheet will happily report "average city = 1.8". Meaningless: nominal labels wearing number costumes are still labels. Same for roll numbers and pin codes.

The variable's TYPE decides its legal maths: means and SDs belong to numerical variables; categorical variables get counts and percentages. Exams test this with "which measure is appropriate?" questions.

Theory

Why this vocabulary runs the whole subject

Every later topic stands on today's words: central tendency and dispersion summarise numerical variables; frequency tables count categorical ones; sampling techniques are about choosing the 60 honestly; the bell curve describes how sample statistics behave. And in Unit 3, R will make you DECLARE these types (numeric, factor): the computer refuses to guess what you have not classified.

Summary

Key takeaways

  • Population = everyone of interest; sample = the part actually measured.
  • Parameter describes the population (unknown); statistic is computed from the sample (estimates it).
  • Census = measure everyone: costly, slow, sometimes impossible.
  • Categorical variables are labels: nominal (unordered) or ordinal (ordered).
  • Numerical variables are numbers: discrete (counted) or continuous (measured).
  • Type decides legal maths: no averaging city codes or roll numbers.
  • Memory hook: one stirred spoon judges the whole pot.

Study this properly

This page is the lesson to read. In Gri-Learn the same topic is a graded deck: the self-checks are scored and your weak topics are tracked. Free to start.

Start this topic

Already have an account? Sign in

More from Basic concepts of statistic

Gri-Learn · syllabus-mapped B.C.A. lessons in English, Hindi and Gujarati

Population vs. sample, variables (categorical vs. numerical), datatypes · Statistical Methods and Data Analysis (MDC-03) · Gri-Learn