Theory
Not all columns are the same kind
In the placement dataset, 'marks' is a number you can average, 'department' is a name you cannot, and 'grade' (A, B, C) is somewhere in between, ordered, but not a measurement. Treating them all the same way leads to nonsense, like computing the 'average department'.
So before analysing, you must know each column's data type. The types form a small family: quantitative (numeric) split into discrete and continuous, and qualitative (categorical) split into nominal and ordinal. Knowing which is which tells you what analysis is even valid.
At a glance
| Type | Meaning | Placement-data example |
|---|---|---|
| Quantitative - discrete | Countable whole numbers | Number of projects (0, 1, 2, 3) |
| Quantitative - continuous | Any value in a range (measured) | Marks percentage (e.g. 78.5) |
| Qualitative - nominal | Categories with no order | Department; placed (yes/no) |
| Qualitative - ordinal | Categories with a meaningful order | Grade (A, B, C); skill level (low/medium/high) |
Theory
Quantitative: discrete vs continuous
Quantitative data are numbers you can meaningfully measure and do arithmetic on. They split in two.
Discrete data are countable whole values, things you count: the number of projects a student did is 0, 1, 2, 3, never 2.5. Continuous data can take any value in a range, things you measure: a marks percentage could be 78, 78.5, 78.53, limited only by precision.
A quick test: can it sensibly have decimals and in-between values? If yes, it is continuous (marks, height, time); if it only jumps in whole steps you count, it is discrete (projects, number of attempts).
Theory
Qualitative: nominal vs ordinal
Qualitative (categorical) data are categories, not measurements. They also split in two, and the difference is order.
Nominal categories have no inherent order: department names (Computer, Commerce, Science) or a yes/no placement flag, none is 'higher' than another. Ordinal categories have a meaningful order but not measurable gaps: a grade of A is better than B is better than C, and satisfaction low/medium/high is ranked, yet you cannot say A minus B equals some fixed amount.
The test: do the categories have a natural ranking? Ranked means ordinal; unranked means nominal.
Quiz
In the placement dataset, the column 'grade' takes values A, B, and C, where A is better than B is better than C. What data type is this?
- Quantitative continuous, because grades are like marks
- Qualitative ordinal, because the categories have a meaningful order but are not measured numbers
- Qualitative nominal, because grades are just labels with no order
- Quantitative discrete, because there are three grades
Show the answer
Qualitative ordinal, because the categories have a meaningful order but are not measured numbers
Grade (A, B, C) is qualitative ordinal: the values are categories with a clear ranking (A better than B better than C), but they are not measured numbers with fixed gaps. Option A is wrong: grades are labels, not continuous measurements, you cannot say A equals 85.3; that would be the underlying marks, a different column. Option C misses the ordering: unlike department names, grades DO have a meaningful order, so they are ordinal, not nominal. Option D is wrong: discrete quantitative means countable numbers (like number of projects), whereas A/B/C are ranked categories, not counts. The deciding question: are they ordered categories (ordinal) or unordered ones (nominal), or actual numbers (quantitative)?
Think first
Why does the data type matter for analysis?
Why bother classifying each column? What goes wrong if you ignore data types? Then tap.
Show the answer
Because the data type decides which OPERATIONS, SUMMARIES, and CHARTS are valid, and applying the wrong one produces meaningless or misleading results. Consider the average. Taking the MEAN of a quantitative column like marks is perfectly sensible ('the average mark is 72'). But taking the 'mean' of a NOMINAL column like department is nonsense, there is no average of 'Computer, Commerce, Science'; the numbers you might assign to those labels are arbitrary, so their average means nothing. Ordinal data is subtler: you can find a median grade (the middle rank) but averaging A, B, C as if the gaps were equal is dubious. The type also dictates the right VISUALISATION: a histogram suits continuous data, a bar chart of counts suits categorical data, and a scatter plot suits two quantitative variables. It even guides how a machine-learning model should treat a column, numeric features and categorical features are handled differently (categories often need encoding). So classifying each column is not busywork; it is what keeps your analysis HONEST, ensuring every average, chart, and model input actually makes sense for the kind of data it describes. Know the type, and you know what you are allowed to do with it.
Summary
Key takeaways
- Each column has a data type that determines which analyses and charts are valid.
- Quantitative data are measurable numbers; qualitative data are categories.
- Quantitative splits into discrete (countable whole values, like number of projects) and continuous (any value in a range, like marks).
- Qualitative splits into nominal (unordered categories, like department) and ordinal (ordered categories, like grade A/B/C).
- Test quantitative: can it sensibly have in-between decimal values (continuous) or only whole counts (discrete)?
- Test qualitative: do the categories have a natural ranking (ordinal) or not (nominal)?
- Memory hook: quantitative counts or measures; qualitative labels, ranked (ordinal) or not (nominal).