Theory
How wrong is the model?
You built a regression model to predict marks. It makes predictions; the real marks are known. So: how wrong is it? To improve a model you must first measure its error with a number, and that number is called a loss (or error) function.
This lesson covers the three you will use most for regression: mean absolute error (MAE), mean squared error (MSE), and R-squared. We will compute all three on one small, worked example so the definitions become concrete. Lower error is better; higher R-squared is better.
Theory
The worked example, and MAE
Take four predictions. The actual marks are 3, 5, 7, 9; the model predicts 4, 5, 7, 6. The error for each is actual minus predicted: -1, 0, 0, 3.
Mean Absolute Error (MAE) is the average of the error sizes (absolute values, ignoring sign):
MAE = (|-1| + |0| + |0| + |3|) / 4 = (1 + 0 + 0 + 3) / 4 = 4 / 4 = 1
So on average the model is off by 1 mark. MAE is in the same units as the target and is easy to interpret: it is simply the typical size of a miss.
Theory
MSE, and why it punishes big errors
Mean Squared Error (MSE) averages the squares of the errors:
MSE = ((-1)squared + 0squared + 0squared + 3squared) / 4 = (1 + 0 + 0 + 9) / 4 = 10 / 4 = 2.5
Notice MSE (2.5) is larger than MAE (1) for the same predictions. The reason is the error of 3: squaring it gives 9, which dominates the average. That is the key property of MSE, by squaring, it penalises large errors much more heavily than small ones. So MSE is the metric to use when big mistakes are especially bad and you want the model to avoid them. (Its downside: squaring also makes it more sensitive to outliers.)
Practical
Computing MAE and MSE (verified)
actual = [3, 5, 7, 9]
predicted = [4, 5, 7, 6]
n = len(actual)
errors = [a - p for a, p in zip(actual, predicted)] # [-1, 0, 0, 3]
mae = sum(abs(e) for e in errors) / n # (1+0+0+3)/4 = 1.0
mse = sum(e*e for e in errors) / n # (1+0+0+9)/4 = 2.5
print(mae) # 1.0
print(mse) # 2.5 -> larger, because the error of 3 squares to 9This example runs in Gri-Learn on the web, where you can edit it and see the output.
Theory
R-squared: how much is explained
R-squared answers a different question: what fraction of the variation in the target does the model explain? It compares the model's errors to the variation around a simple baseline (just predicting the mean).
R-squared = 1 - (SS_res / SS_tot), where SS_res is the sum of squared errors and SS_tot is the total squared variation from the mean. For our example: the mean of the actuals is 6, so SS_tot = (3-6)squared + (5-6)squared + (7-6)squared + (9-6)squared = 9 + 1 + 1 + 9 = 20, and SS_res = 1 + 0 + 0 + 9 = 10. So R-squared = 1 - 10/20 = 0.5.
An R-squared of 0.5 means the model explains half the variation. Near 1 is excellent; 0 means no better than just guessing the mean.
Quiz
For the errors [-1, 0, 0, 3], MAE is 1 but MSE is 2.5. Why is MSE larger than MAE here?
- Because MSE uses more data points than MAE
- Because MSE squares each error, so the error of 3 becomes 9, penalising large errors more heavily
- Because MAE ignores the error of 3 entirely
- It is a mistake; MSE should always equal MAE
Show the answer
Because MSE squares each error, so the error of 3 becomes 9, penalising large errors more heavily
MSE squares each error before averaging, so the largest error dominates: the error of 3 becomes 3 squared = 9, which pushes MSE (2.5) above MAE (1). Option A is wrong: both use the same four data points; the difference is squaring versus absolute value, not the count. Option C is wrong: MAE does include the error of 3, it contributes its absolute value 3 to the sum (1+0+0+3=4, divided by 4 is 1); it simply does not square it. Option D is wrong: MSE and MAE are generally different by design, MSE equals MAE only in special cases (like all errors being 0 or 1). The key property: squaring makes MSE penalise large errors far more than MAE does.
Think first
When would you prefer MAE over MSE, or the reverse?
Both measure regression error. When is squaring (MSE) the right choice, and when is MAE better? Then tap.
Show the answer
Choose based on how much you want to PUNISH large errors and how you feel about OUTLIERS. MSE squares each error, so a few big misses dominate the score; this is exactly what you want when large errors are disproportionately costly, if being off by 10 is much worse than twice as bad as being off by 5, MSE's heavy penalty pushes the model hard to avoid big mistakes. That sensitivity is also its weakness: a single wild outlier can blow up the MSE and drag the model toward chasing that one point. MAE treats every error in proportion to its size (no squaring), so it is more ROBUST to outliers, one freak error moves it only by that error's size, not its square. So MAE is preferable when your data has outliers you do not want to dominate the fit, or when you simply want an error in the target's own units that reads as 'the typical miss'. MSE is preferable when large errors are genuinely the ones to fear and your data is reasonably clean, and it also has nice mathematical properties that many algorithms exploit during training. In practice analysts often look at both: MAE for an interpretable typical-error figure, MSE (or its square root) when large errors matter most, and R-squared for the overall fraction of variance explained. Match the metric to whether big misses should hurt a little or a lot.
Summary
Key takeaways
- A loss function measures how wrong a model is with a single number; lower error is better.
- Error for each prediction is actual minus predicted; for actual [3,5,7,9] and predicted [4,5,7,6] the errors are [-1,0,0,3].
- Mean Absolute Error (MAE) averages the error sizes: (1+0+0+3)/4 = 1; it is in the target's units and easy to read.
- Mean Squared Error (MSE) averages the squared errors: (1+0+0+9)/4 = 2.5; squaring makes it penalise large errors more.
- MSE (2.5) exceeds MAE (1) here because the error of 3 squares to 9.
- R-squared = 1 - SS_res/SS_tot; here 1 - 10/20 = 0.5, so the model explains half the variation (near 1 is excellent, 0 is baseline).
- Memory hook: MAE typical miss, MSE punishes big errors, R-squared fraction of variance explained.