Theory
Dashboard पर आख़िरी सवाल
ResultDesk ने relationships (scatter), trends (line) और shapes (histogram) का जवाब दे दिया। Principal की आख़िरी request सबसे simple लगती है: "हम एक नज़र में किन subjects में सबसे कमज़ोर हैं?"
पाँच subjects, पाँच averages: DBMS 51, Maths 68, Python 72, Stats 55, English 74।
एक नज़र का मतलब है कोई पढ़ना नहीं, digits compare करना नहीं: बस एक-दूसरे के बगल खड़ी heights। किताब का सबसे पुराना chart बिल्कुल यही करता है, और यही वह है जिसे सब पहले से पढ़ना जानते हैं।
Theory
Numbers के लिए एक police lineup
पाँचों subjects को एक lineup की तरह बगल-बगल खड़ा कीजिए, हर एक अपने average जितना लंबा।
आँख बाक़ी काम तुरंत कर देती है: सबसे लंबा = सबसे strong, सबसे छोटा = सबसे कमज़ोर, heights के बीच gaps = वे असल में कितनी दूर हैं। न digits छानने की ज़रूरत, न mental arithmetic। एक bar chart values के लिए एक lineup है: प्रति category एक bar, height गवाही है।
Theory
Bar chart, formally
plt.bar(categories, values) प्रति category एक separated bar खींचता है, height = value।
x axis discrete labels रखता है (subjects, cities, semesters-as-groups): उनका order आपकी choice है, और value से sort करना आमतौर पर सबसे बेहतर पढ़ा जाता है।
नाम देने लायक़ दो variants:
plt.barh(categories, values): horizontal bars: लंबे category names का fix।- Two-factor comparisons के लिए grouped/stacked bars मौजूद हैं: जानिए वे हैं, बाद में master कीजिए।
एक honesty rule: bar charts को अपना value axis zero से शुरू करना चाहिए: bar area value encode करता है, और एक truncated axis differences को inflate करता है।
Practical
सबसे कमज़ोर-subject lineup
import matplotlib.pyplot as plt
import pandas as pd
df = pd.read_csv('marks.csv')
# the aggregate: average score per subject (Unit 1's GROUP BY, pandas-style)
avg = df.groupby('subject')['score'].mean().sort_values()
plt.bar(avg.index, avg.values) # categories, heights
plt.title('Average score by subject')
plt.xlabel('Subject')
plt.ylabel('Average score')
plt.show()
# long subject names? go horizontal:
# plt.barh(avg.index, avg.values)
# the pandas shortcut, one line, same chart:
# df.groupby('subject')['score'].mean().plot(kind='bar')
This example runs in Gri-Learn on the web, where you can edit it and see the output.
Quiz
प्रति SUBJECT average score (पाँच subjects) को एक नज़र में compare करना है। Histogram के बजाय bar chart सही choice क्यों है?
- Subjects compare करने के लिए अलग categories हैं; एक histogram इसके बजाय एक continuous variable को bin करता है
- Histograms 50 से ऊपर averages नहीं दिखा सकते
- Bar charts हमेशा histograms से बेहतर होते हैं
- एक histogram को bins= चाहिए होता जो choose करना बहुत मुश्किल है
Show the answer
Subjects compare करने के लिए अलग categories हैं; एक histogram इसके बजाय एक continuous variable को bin करता है
यहाँ x axis पाँच discrete categories हैं, हर एक एक aggregate value ढो रहा है: bar chart की exact definition। एक histogram एक अलग सवाल का जवाब देता है (ONE numeric column की values कैसे distribute हैं?): इसका "DBMS बनाम Maths" का कोई notion नहीं। Options B और D limitations invent करते हैं; option C वह absolute है जो marks गँवाता है: charts सवालों के लिए सही होते हैं, general में कभी नहीं।
Think first
Chart चुनिए, चार बार
इस unit का final exam। हर एक के लिए scatter, line, histogram या bar चुनिए: (1) 200 students के पार study hours बनाम CGPA; (2) college admissions प्रति year, 2020 से 2026; (3) एक class में Python scores का spread; (4) 8 subjects में से हर एक का average score। tap करने से पहले चारों का जवाब दीजिए।
Show the answer
(1) Scatter: records के पार दो measures के बीच relationship।
(2) Line: एक ordered (time) axis के साथ एक trend।
(3) Histogram: एक numeric variable का distribution।
(4) Bar: अलग categories के पार comparison।
चारों सही = पूरी unit एक reflex में compress हुई: chart को QUESTION से match कीजिए, data की prettiness से नहीं। यह four-way choice एक guaranteed exam और interview सवाल है।
Watch out
Honesty rules
Bars को zero से शुरू कीजिए: 45 से 75 तक का y axis DBMS (51) को English (74) की ऊँचाई का एक-तिहाई दिखाता है: एक visual झूठ: असली ratio 0.69 है, 0.33 नहीं।
सिर्फ़ इसलिए एक time trend को bar-chart मत कीजिए क्योंकि bars solid दिखते हैं: line chart का slope trend story बेहतर बताता है।
और वह बार-बार आने वाला confusion: bars अलग-अलग = categories (bar chart); bars touching = binned number line (histogram)।
Theory
ResultDesk, complete
पीछे हटिए और देखिए आपने 26 topics में क्या बनाया: SQL जो सवालों के जवाब देता है, backups जो disasters से बचते हैं, Python जो इसे automate करता है, pandas जो इसे analyse करता है, और चार charts जो एक कमरे को समझा देते हैं। वह pipeline (store, query, process, visualize) सिर्फ़ इस subject का syllabus नहीं है: यह industry के हर data role की shape है। अगला semester-slot BCA302 statistics को गहरा करता है; toolkit पहले से आपकी है।
Summary
Key takeaways
- plt.bar(categories, values): प्रति category एक separated bar, height = value।
- Discrete categories के पार एक aggregate compare करने के लिए इस्तेमाल कीजिए; readability के लिए value से sort कीजिए।
- plt.barh() horizontal जाता है: लंबे labels का fix।
- Bar value-axes को zero से शुरू होना चाहिए: truncation differences को visually inflate करता है।
- Four-chart reflex: scatter = relationship, line = trend, histogram = distribution, bar = category comparison।
- pandas shortcut: df.groupby(...).mean().plot(kind='bar')।
- Memory hook: numbers के लिए एक police lineup।