Basic Terminologies: Dataset, Features, Labels; Training, Test and Validation Data; Overfitting, Underfitting

The core vocabulary of machine learning: features are the inputs and labels the known answers in your dataset, which you split into training, validation, and test sets, and a good model avoids both overfitting (memorising) and underfitting (oversimplifying).

11 min read · 8 cards · 2 checks

Read in: English · हिन्दी · ગુજરાતી


Theory

The words you need

Machine learning has a compact vocabulary, and using it precisely makes everything clearer. In the placement dataset, the columns you use to predict (marks, attendance, projects) have a name, the thing you predict (placed or not) has a name, and the way you split the data has names too.

This lesson pins down those terms, features, labels, training/test/validation data, and then the two failure modes every model must avoid: overfitting and underfitting. Get this vocabulary right and the rest of ML reads much more easily.

At a glance

TermMeaningPlacement example
DatasetThe whole collection of examplesAll students' records
FeaturesThe input variables used to predictMarks, attendance, projects
LabelsThe known correct output (target)Placed: yes or no
Training dataThe part the model learns from70 percent of students, say
Validation dataUsed during development to tune the modelA slice to compare settings
Test dataHeld out to judge final performance on unseen dataStudents the model never saw in training

Theory

Why split the data

You never judge a model on the same data it learned from, that would be like giving students the exam questions in advance. So you split the dataset.

The model learns from the training data. You keep back test data it never sees during training, and use it to measure how well the model does on new cases, the only fair test of whether it actually learned the pattern. Validation data is a further slice used during development to tune the model and choose settings, so that the test set stays truly untouched until the end. Train on one part, judge on another: that separation is fundamental.

Theory

Overfitting and underfitting

Two opposite failures threaten every model.

Overfitting: the model learns the training data too well, memorising its noise and quirks instead of the general pattern. It scores brilliantly on training data but poorly on new data, it did not generalise. Like a student who memorised past papers word for word and is lost when the questions change.

Underfitting: the model is too simple to capture the real pattern, so it does badly even on the training data, let alone new data. Like a student who studied too little to grasp the topic at all.

The goal sits between them: a model that captures the true pattern and so generalises well to unseen data.

Formula

Overfit memorises; underfit oversimplifies

The clean way to remember: an overfit model has learned too much detail (including noise), so it aces training data but fails on new data, it memorised rather than understood. An underfit model has learned too little, so it fails everywhere, even on training data.

The tell-tale sign of overfitting is a big gap: excellent on training data, poor on test data. Underfitting shows as poor on both. This is exactly why you hold out a test set, without it, you could not even detect overfitting.

Quiz

A model scores 99 percent on the training data but only 60 percent on the unseen test data. What has most likely happened?

  1. Underfitting, because 60 percent is low
  2. Overfitting: it learned the training data too well (including noise) and fails to generalise to new data
  3. The model is perfect, because training accuracy is high
  4. Nothing is wrong; training accuracy is all that matters
Show the answer

Overfitting: it learned the training data too well (including noise) and fails to generalise to new data

A large gap, excellent on training data (99 percent) but much worse on unseen test data (60 percent), is the classic signature of overfitting: the model memorised the training data's noise and quirks rather than the general pattern, so it fails to generalise. Option A is wrong: underfitting shows as poor performance on BOTH training and test data, but here training performance is very high, so it is not underfitting. Option C is a trap: high TRAINING accuracy alone is not success, what matters is performance on NEW data, which is poor here. Option D is the mistake overfitting punishes: judging only on training data hides the failure to generalise, which is exactly why you hold out a test set. Big train-test gap means overfitting.

Think first

Why must the test set be kept completely separate?

Why is it so important that the model never sees the test data during training? Then tap.

Show the answer

Because the test set's whole purpose is to simulate NEW, UNSEEN data, and that only works if the model has genuinely never encountered it, otherwise your performance estimate is dishonestly inflated. When a model is deployed in the real world, it faces cases it never saw while learning, and the one thing you want to know beforehand is: how well will it do on those? The test set answers that by standing in for the future: you train the model on the training data, then check it on the held-out test data that played no part in training. If, however, the test data leaked into training, even indirectly, the model may have effectively memorised those specific examples, so it would score well on them not because it learned the general pattern but because it saw the answers, giving you a falsely optimistic number that collapses the moment real new data arrives. This is also why the VALIDATION set exists as a separate slice: you use validation data to tune and choose your model during development, so that the test set stays pristine and untouched until the very end, preserving it as an honest final exam. The discipline is strict for a reason: the credibility of your performance estimate depends entirely on the test data being truly unseen. Keep it sealed, and your evaluation means something.

Summary

Key takeaways

  • A dataset is the collection of examples; features are the input variables and labels are the known correct outputs.
  • In the placement data, features are marks/attendance/projects and the label is placed (yes/no).
  • Data is split: training data (the model learns from it), validation data (tune during development), test data (held out to judge unseen performance).
  • You never judge a model on data it trained on; train on one part, test on another.
  • Overfitting: the model learns training data too well (including noise) and fails on new data (great on train, poor on test).
  • Underfitting: the model is too simple and does poorly even on training data (poor on both).
  • Memory hook: features in, labels out; overfit memorises, underfit oversimplifies, the goal is to generalise.

Study this properly

This page is the lesson to read. In Gri-Learn the same topic is a graded deck: the self-checks are scored and your weak topics are tracked. Free to start.

Start this topic

Already have an account? Sign in

More from Understanding Supervised Learning

Gri-Learn · syllabus-mapped B.C.A. lessons in English, Hindi and Gujarati