Theory
The words you need
Machine learning has a compact vocabulary, and using it precisely makes everything clearer. In the placement dataset, the columns you use to predict (marks, attendance, projects) have a name, the thing you predict (placed or not) has a name, and the way you split the data has names too.
This lesson pins down those terms, features, labels, training/test/validation data, and then the two failure modes every model must avoid: overfitting and underfitting. Get this vocabulary right and the rest of ML reads much more easily.
At a glance
| Term | Meaning | Placement example |
|---|---|---|
| Dataset | The whole collection of examples | All students' records |
| Features | The input variables used to predict | Marks, attendance, projects |
| Labels | The known correct output (target) | Placed: yes or no |
| Training data | The part the model learns from | 70 percent of students, say |
| Validation data | Used during development to tune the model | A slice to compare settings |
| Test data | Held out to judge final performance on unseen data | Students the model never saw in training |
Theory
Why split the data
You never judge a model on the same data it learned from, that would be like giving students the exam questions in advance. So you split the dataset.
The model learns from the training data. You keep back test data it never sees during training, and use it to measure how well the model does on new cases, the only fair test of whether it actually learned the pattern. Validation data is a further slice used during development to tune the model and choose settings, so that the test set stays truly untouched until the end. Train on one part, judge on another: that separation is fundamental.
Theory
Overfitting and underfitting
Two opposite failures threaten every model.
Overfitting: the model learns the training data too well, memorising its noise and quirks instead of the general pattern. It scores brilliantly on training data but poorly on new data, it did not generalise. Like a student who memorised past papers word for word and is lost when the questions change.
Underfitting: the model is too simple to capture the real pattern, so it does badly even on the training data, let alone new data. Like a student who studied too little to grasp the topic at all.
The goal sits between them: a model that captures the true pattern and so generalises well to unseen data.
Formula
Overfit memorises; underfit oversimplifies
The clean way to remember: an overfit model has learned too much detail (including noise), so it aces training data but fails on new data, it memorised rather than understood. An underfit model has learned too little, so it fails everywhere, even on training data.
The tell-tale sign of overfitting is a big gap: excellent on training data, poor on test data. Underfitting shows as poor on both. This is exactly why you hold out a test set, without it, you could not even detect overfitting.
Quiz
A model scores 99 percent on the training data but only 60 percent on the unseen test data. What has most likely happened?
- Underfitting, because 60 percent is low
- Overfitting: it learned the training data too well (including noise) and fails to generalise to new data
- The model is perfect, because training accuracy is high
- Nothing is wrong; training accuracy is all that matters
Show the answer
Overfitting: it learned the training data too well (including noise) and fails to generalise to new data
A large gap, excellent on training data (99 percent) but much worse on unseen test data (60 percent), is the classic signature of overfitting: the model memorised the training data's noise and quirks rather than the general pattern, so it fails to generalise. Option A is wrong: underfitting shows as poor performance on BOTH training and test data, but here training performance is very high, so it is not underfitting. Option C is a trap: high TRAINING accuracy alone is not success, what matters is performance on NEW data, which is poor here. Option D is the mistake overfitting punishes: judging only on training data hides the failure to generalise, which is exactly why you hold out a test set. Big train-test gap means overfitting.
Think first
Why must the test set be kept completely separate?
Why is it so important that the model never sees the test data during training? Then tap.
Show the answer
Because the test set's whole purpose is to simulate NEW, UNSEEN data, and that only works if the model has genuinely never encountered it, otherwise your performance estimate is dishonestly inflated. When a model is deployed in the real world, it faces cases it never saw while learning, and the one thing you want to know beforehand is: how well will it do on those? The test set answers that by standing in for the future: you train the model on the training data, then check it on the held-out test data that played no part in training. If, however, the test data leaked into training, even indirectly, the model may have effectively memorised those specific examples, so it would score well on them not because it learned the general pattern but because it saw the answers, giving you a falsely optimistic number that collapses the moment real new data arrives. This is also why the VALIDATION set exists as a separate slice: you use validation data to tune and choose your model during development, so that the test set stays pristine and untouched until the very end, preserving it as an honest final exam. The discipline is strict for a reason: the credibility of your performance estimate depends entirely on the test data being truly unseen. Keep it sealed, and your evaluation means something.
Summary
Key takeaways
- A dataset is the collection of examples; features are the input variables and labels are the known correct outputs.
- In the placement data, features are marks/attendance/projects and the label is placed (yes/no).
- Data is split: training data (the model learns from it), validation data (tune during development), test data (held out to judge unseen performance).
- You never judge a model on data it trained on; train on one part, test on another.
- Overfitting: the model learns training data too well (including noise) and fails on new data (great on train, poor on test).
- Underfitting: the model is too simple and does poorly even on training data (poor on both).
- Memory hook: features in, labels out; overfit memorises, underfit oversimplifies, the goal is to generalise.