Theory
Learning instead of being told
Suppose you want to predict which students will be placed. You could try to write rules by hand: if marks above 70 and projects above 2, then placed. But real patterns are far too complex and subtle to capture in hand-written rules.
Machine learning (ML) takes a different path: instead of programming the rules, you let the computer learn them from data, from examples of past students and whether they were placed. This lesson introduces what ML is, its life cycle, and the crucial split between supervised and unsupervised learning.
Theory
What machine learning is
Machine learning is teaching a computer to learn patterns from data rather than being explicitly programmed with every rule. You feed it examples, and it works out the patterns that connect inputs to outcomes, then applies them to new cases.
The benefits follow from this: ML can tackle problems too complex to hand-code (recognising faces, filtering spam), it improves as it sees more data, and it automates predictions at scale. Where the rules are too intricate or hidden for a human to write out, but plenty of example data exists, ML shines. That describes a huge and growing range of real problems.
At a glance
| Stage | What happens |
|---|---|
| Collect data | Gather relevant data (e.g. past students' records) |
| Prepare data | Clean it: handle missing values, outliers, formatting (the EDA skills) |
| Train a model | Choose a model and let it learn patterns from the prepared data |
| Evaluate | Test how well it predicts on data it has not seen |
| Deploy and monitor | Put it to use on new cases, and keep checking it stays accurate |
Theory
Supervised vs unsupervised
The biggest split in ML is whether your data comes with labels, known correct answers.
Supervised learning uses labelled data: each example includes the right output. You show the model past students' features and whether each was placed, and it learns to predict placement for new students. The 'supervision' is the known answers it learns from.
Unsupervised learning uses unlabelled data: no correct output is given. Instead the model finds hidden structure on its own, for example grouping students into natural clusters by similarity, without being told what the groups are. Supervised learns to predict a known target; unsupervised discovers patterns you did not specify.
Quiz
You have past students' data where each record includes whether the student was placed (yes/no), and you train a model to predict placement for new students. Which type of machine learning is this?
- Unsupervised, because the computer figures it out
- Supervised, because the training data is labelled with the known outcome (placed yes/no)
- Neither; prediction is not machine learning
- Unsupervised, because there are only two outcomes
Show the answer
Supervised, because the training data is labelled with the known outcome (placed yes/no)
This is supervised learning: the training data is LABELLED, each past student's record includes the known correct outcome (placed: yes or no), and the model learns from those labelled examples to predict the outcome for new students. Option A is wrong: unsupervised learning works with UNLABELLED data and finds structure on its own; here the outcomes are provided, which is the definition of supervised. Option C is wrong: learning to predict from examples is a central kind of machine learning. Option D misunderstands the categories: the number of outcomes (two) does not make it unsupervised; what matters is that labelled answers guide the learning, making it supervised. The deciding question: does the training data include the known answers (supervised) or not (unsupervised)?
Think first
When is machine learning better than writing rules by hand?
Sometimes a few if-else rules solve a problem fine. When is ML the better tool, and when is it overkill? Then tap.
Show the answer
ML earns its place when the patterns are TOO COMPLEX or TOO HIDDEN to write out as rules, but plenty of EXAMPLE DATA exists; for simple, well-understood problems, hand-written rules are simpler and better. Consider spam filtering: the signals that make an email spam are countless, subtle, and constantly changing, no human could write and maintain rules covering every trick, but millions of labelled examples (spam / not spam) let an ML model learn the patterns and adapt as spam evolves. The same is true for recognising faces, recommending products, or predicting an outcome from dozens of interacting factors: the rules are effectively unknowable by hand, yet learnable from data. On the other hand, if a problem has clear, stable rules, converting Celsius to Fahrenheit, checking whether a number is even, applying a fixed tax slab, then ordinary code is simpler, faster, exact, and easier to trust than ML; using machine learning there would be needless complexity. So the decision hinges on two questions: are the rules too complex or hidden to hand-code, and do you have enough representative data to learn from? If both are yes, ML is powerful; if the logic is simple and known, write the rules directly. Right tool for the problem: rules when they are knowable, learning when they are not.
Summary
Key takeaways
- Machine learning teaches a computer to learn patterns from data instead of being explicitly programmed with rules.
- Benefits: it handles problems too complex to hand-code, improves with more data, and automates prediction at scale.
- A typical ML life cycle: collect data, prepare/clean it, train a model, evaluate it, then deploy and monitor.
- Supervised learning uses labelled data (known outputs) to learn to predict, e.g. predicting placement yes/no.
- Unsupervised learning uses unlabelled data to find hidden structure, e.g. clustering students into groups.
- ML suits complex, hidden patterns with lots of data; simple, known rules are better written by hand.
- Memory hook: ML learns from data; supervised has the answers (labels), unsupervised finds structure.