Theory
क्या attendance असल में marks ख़रीदती है?
वह हमेशा से चलने वाली staff-room बहस ResultDesk तक पहुँचती है: "जो students ज़्यादा attend करते हैं, वे ज़्यादा score करते हैं... है ना?"
आपके पास दोनों columns हैं: attendance percentage और DBMS score, प्रति student एक pair। Averages इसे settle नहीं करेंगे: सवाल एक साथ साठ students में दो variables के बीच एक relationship के बारे में है।
ठीक इस सवाल के लिए बना एक chart है, और वह पूरी library का सबसे simple है: प्रति student एक dot।
Theory
Noticeboard पर साठ pins
हर student को एक pin दीजिए। ज़्यादा attendance के लिए इसे दाएँ खिसकाइए, ज़्यादा marks के लिए ऊपर, और इसे board में दबाइए।
पीछे हटिए। अगर pins दाएँ जाते हुए ऊपर की तरफ़ drift करते हैं, दोनों चीज़ें साथ बढ़ती हैं। अगर वे दाएँ जाते हुए गिरते हैं, एक बढ़ता है जब दूसरा घटता है। अगर वे एक बेढंगे cloud हैं, कोई relationship नहीं। Board साठ students के बारे में कभी झूठ नहीं बोलता जैसा एक anecdote बोल सकता है। वह board एक scatter plot है।
Theory
Scatter, formally
plt.scatter(x, y) प्रति (x, y) pair एक marker खींचता है, किसी से जुड़ा नहीं।
इसे तब इस्तेमाल कीजिए जब दोनों axes measurements हों और सवाल हो "वे साथ कैसे move करते हैं?"
Cloud पढ़ना:
- Dots ऊपर-दाएँ drift करते हुए: positive correlation (ज़्यादा attendance, ज़्यादा marks)।
- नीचे-दाएँ drift करते: negative correlation।
- बेढंगे: कोई visible relation नहीं।
- Cloud से दूर एक dot: एक outlier: वह topper जो कभी attend नहीं करता, roll number से investigate करने लायक़।
Practical
Attendance सवाल, छह lines में जवाब दिया गया
import matplotlib.pyplot as plt
import pandas as pd
df = pd.read_csv('students_full.csv') # roll, attendance, score
plt.scatter(df['attendance'], df['score'])
plt.title('Attendance vs DBMS score')
plt.xlabel('Attendance (%)')
plt.ylabel('DBMS score')
plt.show()
# optional polish: size and colour per point
# plt.scatter(df['attendance'], df['score'], s=20, c='teal')
# object API spelling (the syllabus's set_title):
# fig, ax = plt.subplots()
# ax.scatter(df['attendance'], df['score'])
# ax.set_title('Attendance vs DBMS score')
# ax.set_xlabel('Attendance (%)'); ax.set_ylabel('DBMS score')
This example runs in Gri-Learn on the web, where you can edit it and see the output.
Quiz
Scatter में dots साफ़ तौर पर ऊपर-दाएँ drift करते दिखते हैं, एक dot अकेला (35, 92) पर है। आप इसे कैसे पढ़ेंगे?
- कुल मिलाकर positive correlation, plus एक outlier: एक low-attendance high scorer जिसे individually check करना चाहिए
- कोई relationship नहीं: एक exception pattern को disprove करता है
- Negative correlation: वह अकेला dot direction तय करता है
- Chart invalid है क्योंकि एक point fit नहीं होता
Show the answer
कुल मिलाकर positive correlation, plus एक outlier: एक low-attendance high scorer जिसे individually check करना चाहिए
Cloud का drift story बताता है: ऊपर-दाएँ = positive correlation। एक अकेला दूर वाला dot एक outlier है: यह साठ students के pattern को पलटता नहीं, पर एक नज़र कमाता है (एक brilliant self-studier? एक data-entry slip?)। Patterns mass से पढ़े जाते हैं, exceptions flag किए जाते हैं, कोई भी दूसरे को मिटाता नहीं: यही dual reading exactly वह है जो examiners एक describe-this-plot answer में चाहते हैं।
Think first
Correlation एक verdict नहीं है
Scatter साबित करता है attendance और marks साथ बढ़ते हैं। Principal एक rule draft करता है: "attendance को 90% force कीजिए, marks बढ़ेंगे।" tap करने से पहले: उस jump में क्या flaw है, एक sentence में?
Show the answer
Correlation causation नहीं है: plot दिखाता है दोनों साथ MOVE करते हैं, यह नहीं कि एक दूसरे को CAUSE करता है। शायद sincere students दोनों ज़्यादा attend करते हैं और ज़्यादा पढ़ते हैं (दोनों को drive करता एक third factor); शरीरों को chairs में force करना symptom copy करता है, cause नहीं।
एक scatter investigate करना justify कर सकता है; अकेले यह rule justify नहीं कर सकता। किसी भी correlation के साथ यह एक sentence लिखना वही फ़र्क़ है analysis और superstition के बीच।
Watch out
ग़लत-chart trap
Students habit से plt.plot() पकड़ते हैं और साठ dots एक zigzag line से जुड़े पाते हैं: meaningless, क्योंकि students का x axis पर कोई order नहीं है: line एक ऐसे sequence का दावा करती है जो मौजूद नहीं।
Rule: records के पार दो measures के बीच relationship = scatter; एक ordered axis (tests, months) के साथ trend = line (अगला lesson)। और हमेशा की तरह: कोई title/xlabel/ylabel नहीं, कोई full marks नहीं।
Theory
जंगल में Scatter
Price बनाम demand, study hours बनाम CGPA, RAM बनाम app performance: business और research में हर "क्या ये दो चीज़ें साथ move करती हैं?" सवाल एक scatter plot से खुलता है। यह उस correlation coefficient का visual twin भी है जो आप BCA302 में compute करेंगे: number वह summarise करता है जो यह chart आपको DEKHNE देता है।
Summary
Key takeaways
- plt.scatter(x, y): प्रति record एक dot, कोई connecting line नहीं; two-measure relationship सवालों के लिए।
- Cloud पढ़िए: ऊपर-दाएँ drift = positive, नीचे-दाएँ = negative, बेढंगा = none।
- दूर वाले dots outliers हैं: flag कीजिए और investigate कीजिए, उन्हें mass पलटने मत दीजिए।
- Correlation causation नहीं है: plot co-movement दिखाता है, cause कभी नहीं।
- हमेशा title, xlabel, ylabel (object API में ax.set_title)।
- Records के पार relationships के लिए scatter; एक ordered axis के साथ trends के लिए line।
- Memory hook: noticeboard पर साठ pins।