Theory
एक जैसा average, दो अलग classes
दो DBMS divisions दोनों का average 60 है। वही mean, वही syllabus। फिर moderation के लिए papers दोबारा grade होते हैं:
- Division A: लगभग सब 55 से 65 के बीच।
- Division B: आधी class 40 के आस-पास, बाक़ी आधी 80 के आस-पास, 60 पर लगभग कोई नहीं।
Mean को कोई फ़र्क़ नहीं दिखा। Marks की shape दो बिल्कुल अलग teaching stories बताती है। Shape देखना एक chart का काम है, और वह chart histogram है।
Theory
Envelopes को pigeonholes में sort करना
0-10, 10-20, ... 90-100 label वाले pigeonholes लीजिए। हर student की mark sheet उसके slot में डालिए, फिर stack की heights देखिए।
बीच में लंबा stack: एक clustered class। दूर-दूर दो लंबे stacks: एक ही classroom share करते दो अलग groups। 90-100 में एक अकेला envelope: आपका topper।
एक histogram वही pigeonholes है: यह पूछता है "values कैसे distribute हुई हैं?": एक variable, range से गिना गया।
Theory
Histogram, formally
plt.hist(values, bins=n) numeric range को n equal-width intervals (bins) में split करता है और प्रति bin एक bar खींचता है: height = कितनी values अंदर गिरीं।
नोट कीजिए यह क्या लेता है: एक numeric column। Scatter को x और y चाहिए थे; hist() को सिर्फ़ scores चाहिए: x axis (score ranges) खुद binning से generate होता है, y axis हमेशा एक count होता है।
bins= tuning knob है: बहुत कम bins story को चपटा कर देते हैं, बहुत ज़्यादा इसे noise में तोड़ देते हैं। 60 students की class के लिए, 8-12 के आस-पास शुरू कीजिए।
Practical
दोनों divisions, उजागर
import matplotlib.pyplot as plt
import pandas as pd
df = pd.read_csv('marks.csv')
dbms = df[df['subject'] == 'DBMS']['score'] # ONE numeric column
plt.hist(dbms, bins=10)
plt.title('Distribution of DBMS scores')
plt.xlabel('Score range')
plt.ylabel('Number of students')
plt.show()
# side by side: the two divisions on one figure
# plt.subplot(1, 2, 1); plt.hist(divA, bins=10); plt.title('Div A')
# plt.subplot(1, 2, 2); plt.hist(divB, bins=10); plt.title('Div B')
# Div A: one hump near 60. Div B: two humps: two groups, one mean.
This example runs in Gri-Learn on the web, where you can edit it and see the output.
Think first
Shape को एक examiner की तरह पढ़िए
Division B के histogram में दो साफ़ humps दिखते हैं: एक 40 के आस-पास, एक 80 के आस-पास, 60 पर valley। tap करने से पहले: यह bimodal shape class के बारे में क्या कहती है, और 60 का mean इसे क्यों छुपा गया?
Show the answer
दो humps = एक कमरे में दो अलग populations: एक group 40 के आस-पास struggle करता, दूसरा 80 के आस-पास comfortable: शायद अलग preparation, अलग batches, अलग backgrounds। "60 पर average student के लिए" पढ़ाना एक ऐसे student को target करता है जो मुश्किल से मौजूद है।
Mean इसे इसलिए छुपा गया क्योंकि यह दोनों groups को arithmetically balance करता है: 40s और 80s का average 60 होता है। यह central-tendency lesson की चेतावनी को visible बना देता है: एक summary number को हमेशा उसके पीछे की shape देखनी चाहिए।
Quiz
Histogram बनाम bar chart: कौन सा statement CORRECT है?
- एक histogram एक continuous numeric variable को bin करता है (bars touch करते हैं); एक bar chart अलग categories compare करता है (bars अलग होते हैं)
- दोनों एक ही chart हैं, अलग नामों के साथ
- एक histogram subjects compare करता है; एक bar chart score ranges दिखाता है
- bar chart को bins= चाहिए जबकि histogram को नहीं
Show the answer
एक histogram एक continuous numeric variable को bin करता है (bars touch करते हैं); एक bar chart अलग categories compare करता है (bars अलग होते हैं)
Confusion pair, सुलझी: एक histogram एक CONTINUOUS number line (scores 0-100) को bins में काटता है, तो इसके bars touch करते हैं और उनका order number line से तय होता है। एक bar chart (अगला lesson) SEPARATE categories compare करता है (DBMS बनाम Maths), bars अलग-अलग, order आपकी choice। Option C इन्हें exactly swap करता है: वह mistake जो examiners fish करते हैं। bins= सिर्फ़ hist() का है।
Watch out
Bins trap, दोनों directions
bimodal class पर bins=2: एक मोटा block, दोनों humps ग़ायब, crisis invisible। साठ students पर bins=60: हर bar height 0 या 1, pure noise।
कोई एक सही bins value नहीं है: कुछ (8, 10, 15) try कीजिए और वह रखिए जहाँ shape stabilise हो। अगर कोई exam पूछे bins क्यों matter करता है, यही hide-or-shred trade-off जवाब है।
Theory
जहाँ distributions फ़ैसले तय करती हैं
Result moderation (क्या इस paper के marks scale होने चाहिए?), normalization curves, income surveys, servers में response-time monitoring: सब एक histogram से शुरू होते हैं, क्योंकि सब "distribution कैसी दिखती है?" से शुरू होते हैं। अब आपके पास चार chart archetypes में से तीन हैं; bar chart अगला set पूरा करता है, और उसके साथ, ResultDesk का dashboard।
Summary
Key takeaways
- plt.hist(values, bins=n): एक numeric column, n ranges में binned, bar height = count।
- Histograms SHAPE उजागर करते हैं: humps, twin humps (bimodal = दो groups), tails, outliers।
- एक जैसा mean बिल्कुल अलग shapes छुपा सकता है: दो-divisions वाली story।
- बहुत छोटा bins= structure छुपाता है, बहुत बड़ा इसे तोड़ देता है: class-sized data के लिए 8-15 try कीजिए।
- Histogram = binned continuous variable (bars touch); bar chart = separate categories (bars अलग)।
- Label: title, xlabel = score range, ylabel = students की count।
- Memory hook: pigeonholes में envelopes, stack heights पढ़िए।