Theory
Does one thing predict another?
A natural question about the placement data: does higher attendance go with higher marks? And can you use attendance to predict marks? These are questions about the relationship between variables, and answering them is the job of regression and its supporting measures.
This lesson introduces regression (modelling how one variable depends on others), the vocabulary of dependent and independent variables, and the two tools for measuring how variables move together: covariance (direction) and correlation (strength on a fixed scale).
Theory
Regression, dependent, and independent
Regression models the relationship between variables in order to predict a numeric outcome. In the simplest case, you fit a line that best describes how one variable changes with another.
The variable you want to predict is the dependent variable (often called Y, the response or target), it depends on the others. The variables you use as inputs are the independent variables (X, the predictors or features). For example, to predict a student's marks (dependent) from their attendance (independent), marks is Y, attendance is X. Naming which is which is the first step of any regression.
Theory
Covariance and correlation
To measure whether two variables move together, you have two tools.
Covariance tells you the direction: a positive covariance means they tend to rise together (more attendance, more marks); a negative one means as one rises the other falls. But its size depends on the variables' units, so the raw number is hard to interpret.
Correlation fixes that by standardising to a fixed scale from minus 1 to plus 1: +1 is a perfect positive relationship, -1 a perfect negative one, and 0 means no linear relationship. The magnitude shows strength regardless of units, so a correlation of 0.99 is very strong, while 0.1 is weak. Correlation is covariance made comparable.
Practical
Correlation in pandas (verified)
import pandas as pd
df = pd.DataFrame({
"marks": [80, 55, 90, 70, 60],
"projects": [3, 1, 4, 2, 1],
})
print(df["marks"].corr(df["projects"])) # 0.991 -> strong positive
print(df["marks"].cov(df["projects"])) # positive number -> same direction
# Correlation is on a fixed scale: -1 ... 0 ... +1
# 0.991 is very close to +1, so marks and projects rise strongly together (here).This example runs in Gri-Learn on the web, where you can edit it and see the output.
Quiz
Two variables have a correlation of +0.99. What does this tell you?
- They are almost unrelated, because 0.99 is small
- They have a very strong positive linear relationship: as one rises, the other rises too
- One definitely causes the other
- The correlation is invalid, because correlation must be above 1
Show the answer
They have a very strong positive linear relationship: as one rises, the other rises too
Correlation runs from -1 to +1, so +0.99 is very close to the maximum, indicating a very strong POSITIVE linear relationship: as one variable rises, the other rises too. Option A is backwards: 0.99 is near +1, which is very strong, not small (weak would be near 0). Option C is the classic trap, correlation does NOT imply causation; a strong correlation shows the variables move together, but does not prove one causes the other (a hidden third factor could drive both). Option D is wrong: correlation is bounded within -1 to +1, so 0.99 is perfectly valid, values above 1 are impossible. Read correlation by sign (direction) and magnitude (strength), and never leap from correlation to cause.
Think first
Why does correlation not prove causation?
If two variables are strongly correlated, why can you not conclude that one causes the other? Then tap.
Show the answer
Because correlation only shows that two variables MOVE TOGETHER, not WHY, and there are several innocent reasons for that beyond one causing the other. The most common is a hidden THIRD variable (a confounder) that drives both. Suppose ice-cream sales and swimming-pool drownings are strongly correlated across the year; ice cream does not cause drowning, rather a third factor, hot weather, increases both independently. In the placement data, students with more projects might also have higher marks, but it could be that a third trait (motivation, or ability) raises both, rather than projects directly causing marks. There is also the problem of DIRECTION: even if two things are causally linked, correlation alone cannot tell you which way (does studying raise marks, or do high-marks students choose to study more?), and sometimes a correlation is pure COINCIDENCE, especially in small samples. To actually establish causation you need more than correlation: controlled experiments, or careful methods that rule out confounders and reverse causation. This is why 'correlation does not imply causation' is one of the most important cautions in all of data analysis: a strong correlation is a clue worth investigating, not proof of cause. Move together is not the same as one making the other happen.
Summary
Key takeaways
- Regression models the relationship between variables to predict a numeric outcome (often by fitting a line).
- The dependent variable (Y) is what you predict; the independent variable(s) (X) are the inputs used to predict it.
- Example: predict marks (dependent) from attendance (independent).
- Covariance shows the direction two variables move: positive (rise together) or negative (one rises as the other falls); its size is unit-dependent.
- Correlation standardises this to a fixed scale from -1 to +1: sign is direction, magnitude is strength (0 = none).
- Correlation does NOT imply causation: variables can move together due to a hidden third factor, reverse direction, or coincidence.
- Memory hook: dependent Y from independent X; covariance for direction, correlation (-1 to +1) for strength.