Central tendency measures: mean, median, mode, variance, standard deviation

Mean, median and mode each summarise where the data sits (and disagree when outliers strike), while variance and standard deviation measure how spread out it is, and pandas computes all five in one call each.

11 min read · 10 cards · 2 checks

Read in: English · हिन्दी · ગુજરાતી


Theory

One class, three different "averages"

Five DBMS scores land in ResultDesk: 35, 40, 45, 50, 90.

The class representative announces: "class average 52!" Four of the five students scored below that "average". How can most of a class be below average?

One brilliant outlier (the 90) dragged the mean upward. Summarising data with a single number is a choice, and the three classic choices (mean, median, mode) can tell three different stories about the same marks.

Theory

Three reporters at one exam hall

Three reporters summarise the same results:

  • The accountant (mean): adds everything, divides by the count: precise, but one crorepati in a village makes the "average villager" rich.
  • The queue-watcher (median): lines everyone up, reports the middle person: outliers cannot budge the middle.
  • The crowd-counter (mode): reports what occurs most often: the only one who can also summarise non-numbers (most common grade: 'B').

Theory

The three, formally (worked on our five scores)

Scores: 35, 40, 45, 50, 90.

Mean = sum ÷ count = 260 ÷ 5 = 52.

Median = middle value after sorting: 35, 40, 45, 50, 90 → 45. (Even count? Average the two middle values.)

Mode = most frequent value. Here every score appears once: no useful mode. In 55, 60, 60, 70: mode = 60.

The outlier rule: means chase outliers, medians ignore them: when mean and median disagree sharply, an outlier is hiding in the data.

Theory

Spread: variance and standard deviation

Two classes can share a mean of 60: one scoring 58 to 62, another 20 to 95. The missing number is spread.

Variance = the average of squared deviations from the mean (squaring stops + and - deviations cancelling).

Standard deviation (SD) = √variance, which returns the number to marks units: an SD of 4 means scores typically sit about 4 marks from the mean.

Small SD = consistent class. Large SD = a class of toppers and strugglers wearing one average.

Practical

All five, one line each

import pandas as pd

df = pd.read_csv('marks.csv')
dbms = df[df['subject'] == 'DBMS']['score']   # last lesson's slicing

print(dbms.mean())     # 52.0
print(dbms.median())   # 45.0
print(dbms.mode())     # a Series: there can be MULTIPLE modes
print(dbms.var())      # variance  (sample formula: divides by n-1)
print(dbms.std())      # standard deviation, in marks

# the one-call summary: count, mean, std, min, quartiles, max
print(dbms.describe())

# numpy note: np.std(dbms) divides by n (population formula),
# so it gives a slightly SMALLER answer than dbms.std(). Name
# your formula in exams: sample (n-1) vs population (n).

This example runs in Gri-Learn on the web, where you can edit it and see the output.

Quiz

Scores 35, 40, 45, 50, 90: mean 52, median 45. A scholarship needs the "typical student's score". Which measure and why?

  1. Median 45: the outlier 90 drags the mean above 4 of the 5 students, the median resists it
  2. Mean 52: it uses every value, so it is always more accurate
  3. Mode: it is always the best measure for marks
  4. Variance: it shows the typical score directly
Show the answer

Median 45: the outlier 90 drags the mean above 4 of the 5 students, the median resists it

"Typical" asks for the centre the majority actually sits near: the median stays at 45 no matter how spectacular the topper is, while the mean got pulled to 52, above 80% of the class. Option B's 'uses every value' is exactly WHY the mean is vulnerable here. Mode is useless on all-distinct scores, and variance measures spread, not centre: mixing those two families is the classic exam slip.

Think first

Work one by hand (exams demand this)

Scores: 4, 8, 6, 2 (tiny numbers on purpose). On paper: compute the mean, then the four deviations, then the POPULATION variance (divide by n) and SD. Take your time, then tap.

Show the answer

Mean = (4+8+6+2) ÷ 4 = 5.

Deviations: -1, +3, +1, -3. Squared: 1, 9, 1, 9. Sum = 20.

Population variance = 20 ÷ 4 = 5. SD = √5 ≈ 2.24.

(Sample formula would divide by n-1 = 3: variance ≈ 6.67, SD ≈ 2.58: pandas' default.) If your deviations did not sum to zero before squaring (-1+3+1-3 = 0), recheck the mean: that zero-sum is the built-in error check.

Watch out

Where the marks leak

Median without sorting: the middle of the unsorted list is meaningless: sort first, always.

mode() returns a Series: several values can tie for most frequent; do not assume one number.

pandas vs numpy: .std() divides by n-1 (sample), np.std by n (population): two "correct" answers that differ. In exams, NAME the formula you used.

Theory

This is BCA302 knocking early

Next semester-slot subject BCA302 (Statistical Methods) builds everything on these five numbers: skewness, correlation, distributions. Meeting them here, running on your own marks data, is deliberate: statistics learned on data you built beats statistics learned on textbook tables. Next lesson: the DataFrame's inspection toolkit (head, tail, loc, iloc, describe): the daily-driver functions.

Summary

Key takeaways

  • Mean = sum/count: uses all values, dragged by outliers; median = middle of the SORTED data: outlier-proof.
  • Mode = most frequent value(s): works for categories; can be multiple.
  • Mean far from median = an outlier is hiding.
  • Variance = average squared deviation from the mean; SD = its square root, back in marks units.
  • pandas: .mean() .median() .mode() .var() .std(), and .describe() for the full summary.
  • pandas divides by n-1 (sample); numpy by n (population): name your formula.
  • Memory hook: accountant, queue-watcher, crowd-counter.

Study this properly

This page is the lesson to read. In Gri-Learn the same topic is a graded deck: the self-checks are scored and your weak topics are tracked. Free to start.

Start this topic

Already have an account? Sign in

More from Python interaction with text and CSV

Gri-Learn · syllabus-mapped B.C.A. lessons in English, Hindi and Gujarati

Central tendency measures: mean, median, mode, variance, standard deviation · Database Handling using Python · Gri-Learn