Theory
Same average, two different classes
Two DBMS divisions both average 60. Same mean, same syllabus. Then the papers are regraded for moderation:
- Division A: nearly everyone between 55 and 65.
- Division B: half the class near 40, the other half near 80, almost nobody at 60.
The mean saw no difference. The shape of the marks tells two totally different teaching stories. Seeing shape is a chart's job, and that chart is the histogram.
Theory
Sorting envelopes into pigeonholes
Take pigeonholes labelled 0-10, 10-20, ... 90-100. Drop each student's mark sheet into its slot, then look at the stack heights.
Tall stack in the middle: a clustered class. Two tall stacks apart: two distinct groups sharing one classroom. A lone envelope in 90-100: your topper.
A histogram IS those pigeonholes: it answers "how are the values distributed?": one variable, counted by range.
Theory
Histogram, formally
plt.hist(values, bins=n) splits the numeric range into n equal-width intervals (bins) and draws one bar per bin: height = how many values fell inside.
Note what it takes: one numeric column. Scatter needed x and y; hist() needs only the scores: the x axis (score ranges) is generated by the binning itself, the y axis is always a count.
bins= is the tuning knob: too few bins mash the story flat, too many shred it into noise. For a class of 60, start around 8-12.
Practical
The two divisions, exposed
import matplotlib.pyplot as plt
import pandas as pd
df = pd.read_csv('marks.csv')
dbms = df[df['subject'] == 'DBMS']['score'] # ONE numeric column
plt.hist(dbms, bins=10)
plt.title('Distribution of DBMS scores')
plt.xlabel('Score range')
plt.ylabel('Number of students')
plt.show()
# side by side: the two divisions on one figure
# plt.subplot(1, 2, 1); plt.hist(divA, bins=10); plt.title('Div A')
# plt.subplot(1, 2, 2); plt.hist(divB, bins=10); plt.title('Div B')
# Div A: one hump near 60. Div B: two humps: two groups, one mean.
This example runs in Gri-Learn on the web, where you can edit it and see the output.
Think first
Read the shape like an examiner
Division B's histogram shows two clear humps: one around 40, one around 80, valley at 60. Before tapping: what does this bimodal shape say about the class, and why did the mean of 60 hide it?
Show the answer
Two humps = two distinct populations in one room: one group struggling near 40, another comfortable near 80: perhaps different preparation, different batches, different backgrounds. Teaching "to the average student at 60" targets a student who barely exists.
The mean hid it because it balances the two groups arithmetically: 40s and 80s average to 60. This is the central-tendency lesson's warning made visible: a summary number always deserves a look at the shape behind it.
Quiz
Histogram vs bar chart: which statement is CORRECT?
- A histogram bins one continuous numeric variable (bars touch); a bar chart compares separate categories (bars apart)
- They are the same chart with different names
- A histogram compares subjects; a bar chart shows score ranges
- A bar chart requires bins= while a histogram does not
Show the answer
A histogram bins one continuous numeric variable (bars touch); a bar chart compares separate categories (bars apart)
The confusion pair, settled: a histogram chops a CONTINUOUS number line (scores 0-100) into bins, so its bars touch and their order is fixed by the number line. A bar chart (next lesson) compares SEPARATE categories (DBMS vs Maths), bars apart, order yours to choose. Option C swaps them exactly: the mistake examiners fish for. bins= belongs to hist() alone.
Watch out
The bins trap, both directions
bins=2 on the bimodal class: one fat block, both humps gone, crisis invisible. bins=60 on sixty students: every bar height 0 or 1, pure noise.
There is no single correct bins value: try a few (8, 10, 15) and keep the one where the shape stabilises. If an exam asks why bins matter, this hide-or-shred trade-off IS the answer.
Theory
Where distributions decide things
Result moderation (should this paper's marks be scaled?), normalization curves, income surveys, response-time monitoring in servers: all begin with a histogram, because all begin with "what does the distribution look like?". You now hold three of the four chart archetypes; the bar chart completes the set next, and with it, ResultDesk's dashboard.
Summary
Key takeaways
- plt.hist(values, bins=n): one numeric column, binned into n ranges, bar height = count.
- Histograms reveal SHAPE: humps, twin humps (bimodal = two groups), tails, outliers.
- Same mean can hide wildly different shapes: the two-divisions story.
- bins= too small hides structure, too large shreds it: try 8-15 for class-sized data.
- Histogram = continuous variable binned (bars touch); bar chart = separate categories (bars apart).
- Label: title, xlabel = score range, ylabel = count of students.
- Memory hook: envelopes in pigeonholes, read the stack heights.