Theory
Two kinds of prediction
Supervised learning predicts an output from labelled examples. But 'output' can mean two very different things. For the placement data you might ask: will this student be placed, yes or no? Or you might ask: what marks will this student score?
The first predicts a category; the second predicts a number. That distinction splits supervised learning into two families, classification and regression, and knowing which one your problem is decides everything that follows: the model, the metrics, the whole approach. This lesson makes the split crisp.
Theory
Classification vs regression
Classification predicts a category (a discrete class label). The output is one of a fixed set of classes: placed or not placed, spam or not spam, which of the digits 0 to 9. The answer is a label, not a measured quantity.
Regression predicts a continuous number. The output is a quantity on a scale: a student's marks, a house's price, tomorrow's temperature. The answer can take any value in a range.
Both are supervised (both learn from labelled training data). The only thing that differs is the type of output: a category means classification, a number means regression.
At a glance
| Aspect | Classification | Regression |
|---|---|---|
| Predicts | A category (class label) | A continuous number |
| Output example | Placed: yes or no | Marks: 78.5 |
| Other examples | Spam / not spam; disease / no disease | Price, temperature, sales |
| Question it answers | Which class does this belong to? | How much / how many? |
Formula
The one question that decides it
To tell which family a problem belongs to, ask one thing: is the target a category or a number?
If you are predicting which class something falls into, placed or not, pass or fail, cat or dog, it is classification. If you are predicting how much or how many, a marks score, a rupee amount, a temperature, it is regression. That single question, category versus quantity, reliably sorts any supervised problem into the right family, and stops the most common beginner mix-up.
Quiz
You want to predict the exact marks percentage a student will score (like 78.5) from their attendance and projects. Is this classification or regression?
- Classification, because students are in categories
- Regression, because the target is a continuous number (a marks value)
- Unsupervised, because it uses data
- Neither; predicting numbers is not machine learning
Show the answer
Regression, because the target is a continuous number (a marks value)
Predicting an exact marks percentage (a continuous number like 78.5) is regression, because the target is a quantity on a scale, not a category. Option A is wrong: you are not sorting students into fixed classes here; you are predicting a numeric value, which is regression, if instead you predicted 'pass/fail' or 'placed/not', that would be classification. Option C is wrong: the data is labelled with known marks, so this is supervised learning, not unsupervised. Option D is wrong: predicting a number from labelled examples is a core supervised ML task (regression). Apply the test: a number means regression, a category means classification.
Think first
Why does the classification-vs-regression distinction matter so much?
Both are supervised learning. Why is it important to know which one your problem is? Then tap.
Show the answer
Because almost everything downstream depends on it: the MODELS you can use, the way you measure SUCCESS, and how you interpret the output all differ between classification and regression. The evaluation metrics are different, and this is the most immediate consequence. For REGRESSION, you measure error as a distance between predicted and actual numbers, using metrics like mean squared error or mean absolute error (coming next lesson), because 'how far off is 78 from 80' is meaningful. For CLASSIFICATION, distance makes no sense (there is no 'distance' between 'placed' and 'not placed'), so you measure things like accuracy, precision, and recall, was the predicted class right? Using a regression metric on a classification problem, or vice versa, gives nonsense. The MODELS also differ: many algorithms have distinct classification and regression versions, and some suit one task better than the other. Even the OUTPUT is interpreted differently: a regression model gives you a number to act on, while a classification model gives you a label (often with a probability). So identifying the family first is not pedantry; it is the fork in the road that determines which tools, metrics, and interpretations are valid for the rest of the project. Get this wrong and everything after it is built on sand. Category or quantity, decide first, then proceed correctly.
Summary
Key takeaways
- Supervised learning splits into two families based on what you predict.
- Classification predicts a category (a discrete class label): placed or not, spam or not, digit 0 to 9.
- Regression predicts a continuous number: marks, price, temperature.
- Both are supervised (both learn from labelled data); only the output type differs.
- The deciding question: is the target a category (classification) or a number (regression)?
- The distinction matters because models, evaluation metrics, and interpretation all depend on it.
- Memory hook: category means classify, quantity means regress.