Statistical methods: Mean, Median, Mode, Standard Deviation, Variance

Statistical methods summarize messy distributions of data into crisp metrics that reveal the central balance, middle rank, and volumetric spread of numbers.

11 min read · 12 cards · 3 checks

Read in: English · हिन्दी · ગુજરાતી


Theory

The Blind Spots of Grand Totals

In your Semester 2 NumPy labs, you learned how to sum up thousands of records instantly using np.sum(). But if you are reviewing your campus canteen monthly food spends, does knowing the grand total amount spent tell you whether you have a consistent daily eating habit or if you splurged your entire month's pocket money on a single birthday party? Why do aggregate sums fail to tell the true story of how numbers fluctuate, and how do we capture the true essence of a dataset without staring at thousands of raw rows?

Theory

The Classroom Height Balance vs the Scatter Spray

Think of descriptive statistics like taking a group photograph of your BCA section. The Mean is like finding the exact center point of a physical balance scale where everyone can stand evenly. The Median is sorting the class from shortest to tallest and tapping the single student standing exactly in the middle of the queue. But to see if your classmates are all nearly the same height or a wild mix of basketball players and short students, you need Standard Deviation, which measures how far away individuals are sprayed or scattered out from that center balance line.

Theory

Descriptive Statistics Formally

Descriptive statistics split data analytics into two core categories: Central Tendency (where the data clusters) and Dispersion (how spread out the data is). The Mean represents the arithmetic average. The Median splits an ordered dataset into two equal halves, acting as a robust measure against extreme outliers. The Mode represents the most frequent data point. To quantify dispersion, Variance calculates the average squared difference from the mean, and Standard Deviation (SD) takes the square root of variance to return the spread metric back to the data's original scale unit.

At a glance

Table 1: Essential Python statistical tools and mathematical attributes.

Statistical MetricPython Syntax TemplateBehavioral Characteristic
Arithmetic Meannp.mean(dataset)Highly sensitive to extreme outliers or typo entries.
Middle Mediannp.median(dataset)Locks onto physical rank order; immune to skewing spikes.
Peak Modestats.mode(dataset)Tracks popularity; ideal for categorical parameters.
Squared Variancenp.var(dataset)Measures squared dispersion; heavily amplifies outer variations.
Standard Deviationnp.std(dataset)The definitive benchmark for standard distance spread.

Theory

Worked Example: Auditing the PocketMoney Variance

Let us compute descriptive metrics for a structured array of student expenses using NumPy and the SciPy stats module. We will look closely at how a single massive outlier impacts the mean versus the median, and determine the exact metric of spread.

Practical

PocketMoney Spread Auditor

import numpy as np
from scipy import stats

# Step 1: Expense logs with a major outlier spike (e.g., a 400 Rs book purchase)
spends = np.array([40, 50, 40, 60, 400, 50, 60])

# Step 2: Calculate central tendencies
mean_val   = np.mean(spends)
median_val = np.median(spends)
mode_res   = stats.mode(spends, keepdims=True)

# Step 3: Compute dispersion metrics
variance_val = np.var(spends)
std_dev_val  = np.std(spends)

print(f"Mean Spend: {mean_val:.2f}")
print(f"Median Spend: {median_val:.2f}")
print(f"Mode (Most Frequent): {mode_res.mode[0]} (Count: {mode_res.count[0]})")
print(f"Variance Score: {variance_val:.2f}")
print(f"Standard Deviation Spread: {std_dev_val:.2f}")

This example runs in Gri-Learn on the web, where you can edit it and see the output.

Think first

Analyze Outlier Sensitivity

Look at the spends dataset: five items are under 60 Rs, but one item jumps to 400 Rs. How will this single spike alter the balance between the mean and median inside your terminal output?

Show the answer

The script will output:

Mean Spend: 100.00

Median Spend: 50.00

Mode (Most Frequent): 40 (Count: 2)

Variance Score: 15057.14

Standard Deviation Spread: 122.71

Why? The mean is pulled up to 100 Rs because it accounts for the total sum divided by 7. However, the median sorts the array to [40, 40, 50, 50, 60, 60, 400] and picks 50, providing a much more accurate picture of a typical day's expense. The Standard Deviation of 122.71 highlights a high level of volatility in your spending data caused by that 400 Rs purchase.

Quiz

In a university exam paper question, if a distribution has a mean of 50 and a variance of 64, what is its exact Standard Deviation metric?

  1. 4096
  2. 8
  3. 32
  4. 6.4
Show the answer

8

Standard Deviation is mathematically defined as the positive square root of Variance. The square root of 64 is 8. Mean is a measure of center, so it does not affect this conversion calculation.

Quiz

Why do corporate database analytics systems rely heavily on the Median over the Mean when analyzing average consumer metrics or salary distributions?

  1. The median requires much less computing memory to run in Python loops.
  2. The mean is severely dragged out of proportion by extreme high value outliers, whereas the median remains stable.
  3. The mean cannot process arrays that contain floating decimal values.
  4. The median ignores duplicate entries in the distribution entirely.
Show the answer

The mean is severely dragged out of proportion by extreme high value outliers, whereas the median remains stable.

Because the mean factors in the exact value of every element, a few multi-millionaire salaries or massive purchase outliers can artificially inflate it. The median depends only on rank position, making it a more reliable metric for skewed datasets.

Watch out

The Classic Trap: The Default Population Variance Slip

A frequent mistake in university lab exams is assuming np.std() behaves exactly like standard textbook formulas for sample datasets. By default, NumPy calculates the population standard deviation ($N$ in the denominator, or ddof=0). If your question paper explicitly demands the sample standard deviation ($N-1$ degrees of freedom), you must supply the optional parameter np.std(dataset, ddof=1), or your final decimal value will lose marks!

Theory

Connecting Statistics to Semester 3

Descriptive statistics form the backbone of predictive intelligence. In Semester 3 Data Science (BCA302), when building regression pipelines or normalization engines, you will use mean and standard deviation matrices to clean and scale inputs, helping your machine learning models process diverse data fairly.

Summary

Key takeaways

  • Mean acts as the arithmetic balance point but is highly vulnerable to outliers.
  • Median identifies the true center of sorted data, offering resistance to extreme value spikes.
  • Mode isolates the highest-frequency recurring parameters inside a dataset.
  • Variance evaluates squared deviations, tracking the underlying variability of your records.
  • Standard Deviation returns the variance score to the original unit scale for clear interpretation.
  • Memory Hook: Mean sums all cells, Median tracks rank position, SD defines the deviation radius!

Study this properly

This page is the lesson to read. In Gri-Learn the same topic is a graded deck: the self-checks are scored and your weak topics are tracked. Free to start.

Start this topic

Already have an account? Sign in

More from Python Libraries

Gri-Learn · syllabus-mapped B.C.A. lessons in English, Hindi and Gujarati

Statistical methods: Mean, Median, Mode, Standard Deviation, Variance · Programming Skills · Gri-Learn