Histogram chart: concepts of histogram hist(), set title, xlabel and ylabel

एक histogram एक numeric range को bins में काटता है और दिखाता है कितनी values हर bin में गिरती हैं, distribution की shape उजागर करते हुए: plt.hist(scores, bins=n) पूरी call है।

9 min read · 9 cards · 2 checks

Read in: English · हिन्दी · ગુજરાતી


Theory

एक जैसा average, दो अलग classes

दो DBMS divisions दोनों का average 60 है। वही mean, वही syllabus। फिर moderation के लिए papers दोबारा grade होते हैं:

  • Division A: लगभग सब 55 से 65 के बीच।
  • Division B: आधी class 40 के आस-पास, बाक़ी आधी 80 के आस-पास, 60 पर लगभग कोई नहीं।

Mean को कोई फ़र्क़ नहीं दिखा। Marks की shape दो बिल्कुल अलग teaching stories बताती है। Shape देखना एक chart का काम है, और वह chart histogram है।

Theory

Envelopes को pigeonholes में sort करना

0-10, 10-20, ... 90-100 label वाले pigeonholes लीजिए। हर student की mark sheet उसके slot में डालिए, फिर stack की heights देखिए।

बीच में लंबा stack: एक clustered class। दूर-दूर दो लंबे stacks: एक ही classroom share करते दो अलग groups। 90-100 में एक अकेला envelope: आपका topper।

एक histogram वही pigeonholes है: यह पूछता है "values कैसे distribute हुई हैं?": एक variable, range से गिना गया।

Theory

Histogram, formally

plt.hist(values, bins=n) numeric range को n equal-width intervals (bins) में split करता है और प्रति bin एक bar खींचता है: height = कितनी values अंदर गिरीं।

नोट कीजिए यह क्या लेता है: एक numeric column। Scatter को x और y चाहिए थे; hist() को सिर्फ़ scores चाहिए: x axis (score ranges) खुद binning से generate होता है, y axis हमेशा एक count होता है।

bins= tuning knob है: बहुत कम bins story को चपटा कर देते हैं, बहुत ज़्यादा इसे noise में तोड़ देते हैं। 60 students की class के लिए, 8-12 के आस-पास शुरू कीजिए।

Practical

दोनों divisions, उजागर

import matplotlib.pyplot as plt
import pandas as pd

df = pd.read_csv('marks.csv')
dbms = df[df['subject'] == 'DBMS']['score']   # ONE numeric column

plt.hist(dbms, bins=10)
plt.title('Distribution of DBMS scores')
plt.xlabel('Score range')
plt.ylabel('Number of students')
plt.show()

# side by side: the two divisions on one figure
# plt.subplot(1, 2, 1); plt.hist(divA, bins=10); plt.title('Div A')
# plt.subplot(1, 2, 2); plt.hist(divB, bins=10); plt.title('Div B')
# Div A: one hump near 60.  Div B: two humps: two groups, one mean.

This example runs in Gri-Learn on the web, where you can edit it and see the output.

Think first

Shape को एक examiner की तरह पढ़िए

Division B के histogram में दो साफ़ humps दिखते हैं: एक 40 के आस-पास, एक 80 के आस-पास, 60 पर valley। tap करने से पहले: यह bimodal shape class के बारे में क्या कहती है, और 60 का mean इसे क्यों छुपा गया?

Show the answer

दो humps = एक कमरे में दो अलग populations: एक group 40 के आस-पास struggle करता, दूसरा 80 के आस-पास comfortable: शायद अलग preparation, अलग batches, अलग backgrounds। "60 पर average student के लिए" पढ़ाना एक ऐसे student को target करता है जो मुश्किल से मौजूद है।

Mean इसे इसलिए छुपा गया क्योंकि यह दोनों groups को arithmetically balance करता है: 40s और 80s का average 60 होता है। यह central-tendency lesson की चेतावनी को visible बना देता है: एक summary number को हमेशा उसके पीछे की shape देखनी चाहिए।

Quiz

Histogram बनाम bar chart: कौन सा statement CORRECT है?

  1. एक histogram एक continuous numeric variable को bin करता है (bars touch करते हैं); एक bar chart अलग categories compare करता है (bars अलग होते हैं)
  2. दोनों एक ही chart हैं, अलग नामों के साथ
  3. एक histogram subjects compare करता है; एक bar chart score ranges दिखाता है
  4. bar chart को bins= चाहिए जबकि histogram को नहीं
Show the answer

एक histogram एक continuous numeric variable को bin करता है (bars touch करते हैं); एक bar chart अलग categories compare करता है (bars अलग होते हैं)

Confusion pair, सुलझी: एक histogram एक CONTINUOUS number line (scores 0-100) को bins में काटता है, तो इसके bars touch करते हैं और उनका order number line से तय होता है। एक bar chart (अगला lesson) SEPARATE categories compare करता है (DBMS बनाम Maths), bars अलग-अलग, order आपकी choice। Option C इन्हें exactly swap करता है: वह mistake जो examiners fish करते हैं। bins= सिर्फ़ hist() का है।

Watch out

Bins trap, दोनों directions

bimodal class पर bins=2: एक मोटा block, दोनों humps ग़ायब, crisis invisible। साठ students पर bins=60: हर bar height 0 या 1, pure noise।

कोई एक सही bins value नहीं है: कुछ (8, 10, 15) try कीजिए और वह रखिए जहाँ shape stabilise हो। अगर कोई exam पूछे bins क्यों matter करता है, यही hide-or-shred trade-off जवाब है।

Theory

जहाँ distributions फ़ैसले तय करती हैं

Result moderation (क्या इस paper के marks scale होने चाहिए?), normalization curves, income surveys, servers में response-time monitoring: सब एक histogram से शुरू होते हैं, क्योंकि सब "distribution कैसी दिखती है?" से शुरू होते हैं। अब आपके पास चार chart archetypes में से तीन हैं; bar chart अगला set पूरा करता है, और उसके साथ, ResultDesk का dashboard।

Summary

Key takeaways

  • plt.hist(values, bins=n): एक numeric column, n ranges में binned, bar height = count।
  • Histograms SHAPE उजागर करते हैं: humps, twin humps (bimodal = दो groups), tails, outliers।
  • एक जैसा mean बिल्कुल अलग shapes छुपा सकता है: दो-divisions वाली story।
  • बहुत छोटा bins= structure छुपाता है, बहुत बड़ा इसे तोड़ देता है: class-sized data के लिए 8-15 try कीजिए।
  • Histogram = binned continuous variable (bars touch); bar chart = separate categories (bars अलग)।
  • Label: title, xlabel = score range, ylabel = students की count।
  • Memory hook: pigeonholes में envelopes, stack heights पढ़िए।

Study this properly

This page is the lesson to read. In Gri-Learn the same topic is a graded deck: the self-checks are scored and your weak topics are tracked. Free to start.

Start this topic

Already have an account? Sign in

More from Data Visualization using dataframe

Gri-Learn · syllabus-mapped B.C.A. lessons in English, Hindi and Gujarati

Histogram chart: concepts of histogram hist(), set title, xlabel and ylabel · Database Handling using Python · Gri-Learn