Spread of Data: Normal Distribution; Skewed Distribution; Skewness and Kurtosis

The shape of your data matters: a normal distribution is the symmetric bell curve, skewness measures how lopsided a distribution is (which way its tail leans), and kurtosis measures how heavy its tails and how sharp its peak are.

11 min read · 7 cards · 2 checks

Read in: English · हिन्दी · ગુજરાતી


Theory

The shape of the data

Averages alone can mislead. Two classes can have the same average marks but completely different shapes: one bunched tightly around the mean, another with a long tail of low scorers dragging things down. To really understand a variable, you look at how its values are distributed, their shape.

This lesson covers the key shapes: the symmetric normal distribution, skewed distributions that lean one way, and the two numbers that measure shape, skewness (lopsidedness) and kurtosis (tailedness). You met the bell curve in statistics; here you learn to read shape like an analyst.

Theory

The normal distribution

The normal distribution is the famous symmetric, bell-shaped curve. Most values cluster near the centre, and they thin out evenly on both sides. In a perfectly normal distribution, the mean, median, and mode all sit at the same central point.

Many natural measurements (heights, and often exam marks) are roughly normal. It is the reference shape against which we describe others: we ask how a real distribution departs from this ideal bell, and the two ways it can depart are captured by skewness and kurtosis.

Theory

Skewed distributions and skewness

A skewed distribution is asymmetric, one tail is longer than the other, and skewness is the number that measures this lopsidedness (0 means symmetric).

Right (positive) skew: a long tail stretches to the right (toward high values), pulling the mean above the median. Example: incomes, or marks where a few students score very high above a low cluster.

Left (negative) skew: a long tail stretches to the left (toward low values), pulling the mean below the median. Example: an easy exam where most score high but a few fail badly.

The handy rule: for right skew mean > median; for left skew mean < median; the mean is pulled toward the long tail.

Formula

Skewness leans; kurtosis peaks

Keep the two shape measures distinct. Skewness is about symmetry, which way (if any) the distribution leans, told by which tail is longer and by mean-versus-median.

Kurtosis is about tailedness and peakedness, how heavy the tails are and how sharp the peak is. High kurtosis means heavy tails and a sharp peak (more extreme outliers than a normal curve); low kurtosis means light tails and a flatter shape. So skewness answers 'is it lopsided, and which way?' while kurtosis answers 'how extreme are the tails?'.

Quiz

In a distribution of marks, the mean (3.5) is greater than the median (3.0), and a few very high scores stretch out to the right. What is this distribution?

  1. Left (negative) skew, because the mean is small
  2. Right (positive) skew, with a long right tail pulling the mean above the median
  3. Perfectly normal, because it has a mean and a median
  4. It has high kurtosis, not skew
Show the answer

Right (positive) skew, with a long right tail pulling the mean above the median

A long tail stretching to the RIGHT (toward high values) with the mean pulled ABOVE the median (3.5 > 3.0) is the signature of right (positive) skew. Option A is backwards: left skew has the tail on the LEFT and the mean BELOW the median (mean < median), the opposite of what is described. Option C is wrong: a perfectly normal distribution is symmetric with mean = median = mode, but here mean does not equal median, so it is not normal. Option D confuses the two measures: kurtosis describes tail heaviness/peakedness, whereas the lopsidedness and the mean-median gap here are about SKEWNESS. Remember: the mean is dragged toward the long tail, so mean > median means a right tail, right skew.

Think first

Why does the mean get pulled toward the long tail?

In a right-skewed distribution, why does the mean end up higher than the median? Then tap.

Show the answer

Because the MEAN is sensitive to extreme values while the MEDIAN is not, so a long tail of extreme values drags the mean toward it, but barely moves the median. The mean is the arithmetic average: every value contributes its full size to the total, so a handful of very large values (the long right tail) add a lot to the sum and pull the average UP. The median, by contrast, is just the middle value when the data is sorted, it only cares about POSITION, not magnitude. Adding a few enormous values on the right pushes the middle position only slightly, because the median counts data points, not their size. So in a right-skewed distribution, the extreme high values inflate the mean well above the median, giving mean > median; in a left-skewed distribution, extreme low values drag the mean below the median. This is exactly why, for skewed data like incomes, the median is often reported as the 'typical' value, it resists distortion by outliers, whereas the mean can be misleadingly high. It also explains why comparing mean and median is such a quick skew test: if they differ noticeably, the data is skewed toward whichever side the mean sits. The mean chases the tail; the median holds the middle.

Summary

Key takeaways

  • The distribution's shape, not just its average, reveals what a variable is really like.
  • The normal distribution is the symmetric bell curve, where mean = median = mode and values cluster centrally.
  • Skewness measures asymmetry: right (positive) skew has a long right tail with mean > median; left (negative) skew has a long left tail with mean < median.
  • The mean is pulled toward the long tail, so comparing mean and median is a quick skew test.
  • Kurtosis measures tailedness and peakedness: high = heavy tails and sharp peak (more extremes), low = light tails and flatter.
  • Skewness answers 'which way does it lean?'; kurtosis answers 'how extreme are the tails?'.
  • Memory hook: normal is the balanced bell; skew leans (mean chases the tail); kurtosis is about the tails' weight.

Study this properly

This page is the lesson to read. In Gri-Learn the same topic is a graded deck: the self-checks are scored and your weak topics are tracked. Free to start.

Start this topic

Already have an account? Sign in

More from Fundamentals of Data Analytics

Gri-Learn · syllabus-mapped B.C.A. lessons in English, Hindi and Gujarati

Spread of Data: Normal Distribution; Skewed Distribution; Skewness and Kurtosis · Data Analytics using Python (Major-14) · Gri-Learn