Theory
One class, three different "averages"
Five DBMS scores land in ResultDesk: 35, 40, 45, 50, 90.
The class representative announces: "class average 52!" Four of the five students scored below that "average". How can most of a class be below average?
One brilliant outlier (the 90) dragged the mean upward. Summarising data with a single number is a choice, and the three classic choices (mean, median, mode) can tell three different stories about the same marks.
Theory
Three reporters at one exam hall
Three reporters summarise the same results:
- The accountant (mean): adds everything, divides by the count: precise, but one crorepati in a village makes the "average villager" rich.
- The queue-watcher (median): lines everyone up, reports the middle person: outliers cannot budge the middle.
- The crowd-counter (mode): reports what occurs most often: the only one who can also summarise non-numbers (most common grade: 'B').
Theory
The three, formally (worked on our five scores)
Scores: 35, 40, 45, 50, 90.
Mean = sum ÷ count = 260 ÷ 5 = 52.
Median = middle value after sorting: 35, 40, 45, 50, 90 → 45. (Even count? Average the two middle values.)
Mode = most frequent value. Here every score appears once: no useful mode. In 55, 60, 60, 70: mode = 60.
The outlier rule: means chase outliers, medians ignore them: when mean and median disagree sharply, an outlier is hiding in the data.
Theory
Spread: variance and standard deviation
Two classes can share a mean of 60: one scoring 58 to 62, another 20 to 95. The missing number is spread.
Variance = the average of squared deviations from the mean (squaring stops + and - deviations cancelling).
Standard deviation (SD) = √variance, which returns the number to marks units: an SD of 4 means scores typically sit about 4 marks from the mean.
Small SD = consistent class. Large SD = a class of toppers and strugglers wearing one average.
Practical
All five, one line each
import pandas as pd
df = pd.read_csv('marks.csv')
dbms = df[df['subject'] == 'DBMS']['score'] # last lesson's slicing
print(dbms.mean()) # 52.0
print(dbms.median()) # 45.0
print(dbms.mode()) # a Series: there can be MULTIPLE modes
print(dbms.var()) # variance (sample formula: divides by n-1)
print(dbms.std()) # standard deviation, in marks
# the one-call summary: count, mean, std, min, quartiles, max
print(dbms.describe())
# numpy note: np.std(dbms) divides by n (population formula),
# so it gives a slightly SMALLER answer than dbms.std(). Name
# your formula in exams: sample (n-1) vs population (n).
This example runs in Gri-Learn on the web, where you can edit it and see the output.
Quiz
Scores 35, 40, 45, 50, 90: mean 52, median 45. A scholarship needs the "typical student's score". Which measure and why?
- Median 45: the outlier 90 drags the mean above 4 of the 5 students, the median resists it
- Mean 52: it uses every value, so it is always more accurate
- Mode: it is always the best measure for marks
- Variance: it shows the typical score directly
Show the answer
Median 45: the outlier 90 drags the mean above 4 of the 5 students, the median resists it
"Typical" asks for the centre the majority actually sits near: the median stays at 45 no matter how spectacular the topper is, while the mean got pulled to 52, above 80% of the class. Option B's 'uses every value' is exactly WHY the mean is vulnerable here. Mode is useless on all-distinct scores, and variance measures spread, not centre: mixing those two families is the classic exam slip.
Think first
Work one by hand (exams demand this)
Scores: 4, 8, 6, 2 (tiny numbers on purpose). On paper: compute the mean, then the four deviations, then the POPULATION variance (divide by n) and SD. Take your time, then tap.
Show the answer
Mean = (4+8+6+2) ÷ 4 = 5.
Deviations: -1, +3, +1, -3. Squared: 1, 9, 1, 9. Sum = 20.
Population variance = 20 ÷ 4 = 5. SD = √5 ≈ 2.24.
(Sample formula would divide by n-1 = 3: variance ≈ 6.67, SD ≈ 2.58: pandas' default.) If your deviations did not sum to zero before squaring (-1+3+1-3 = 0), recheck the mean: that zero-sum is the built-in error check.
Watch out
Where the marks leak
Median without sorting: the middle of the unsorted list is meaningless: sort first, always.
mode() returns a Series: several values can tie for most frequent; do not assume one number.
pandas vs numpy: .std() divides by n-1 (sample), np.std by n (population): two "correct" answers that differ. In exams, NAME the formula you used.
Theory
This is BCA302 knocking early
Next semester-slot subject BCA302 (Statistical Methods) builds everything on these five numbers: skewness, correlation, distributions. Meeting them here, running on your own marks data, is deliberate: statistics learned on data you built beats statistics learned on textbook tables. Next lesson: the DataFrame's inspection toolkit (head, tail, loc, iloc, describe): the daily-driver functions.
Summary
Key takeaways
- Mean = sum/count: uses all values, dragged by outliers; median = middle of the SORTED data: outlier-proof.
- Mode = most frequent value(s): works for categories; can be multiple.
- Mean far from median = an outlier is hiding.
- Variance = average squared deviation from the mean; SD = its square root, back in marks units.
- pandas: .mean() .median() .mode() .var() .std(), and .describe() for the full summary.
- pandas divides by n-1 (sample); numpy by n (population): name your formula.
- Memory hook: accountant, queue-watcher, crowd-counter.