Scatter plot: concept, set title, xlabel and ylabel

एक scatter plot प्रति record एक dot (x, y) पर रखता है ताकि दो variables के बीच relationship दिखाई दे, और plt.scatter() title/xlabel/ylabel के साथ पूरी recipe है।

8 min read · 9 cards · 2 checks

Read in: English · हिन्दी · ગુજરાતી


Theory

क्या attendance असल में marks ख़रीदती है?

वह हमेशा से चलने वाली staff-room बहस ResultDesk तक पहुँचती है: "जो students ज़्यादा attend करते हैं, वे ज़्यादा score करते हैं... है ना?"

आपके पास दोनों columns हैं: attendance percentage और DBMS score, प्रति student एक pair। Averages इसे settle नहीं करेंगे: सवाल एक साथ साठ students में दो variables के बीच एक relationship के बारे में है।

ठीक इस सवाल के लिए बना एक chart है, और वह पूरी library का सबसे simple है: प्रति student एक dot।

Theory

Noticeboard पर साठ pins

हर student को एक pin दीजिए। ज़्यादा attendance के लिए इसे दाएँ खिसकाइए, ज़्यादा marks के लिए ऊपर, और इसे board में दबाइए।

पीछे हटिए। अगर pins दाएँ जाते हुए ऊपर की तरफ़ drift करते हैं, दोनों चीज़ें साथ बढ़ती हैं। अगर वे दाएँ जाते हुए गिरते हैं, एक बढ़ता है जब दूसरा घटता है। अगर वे एक बेढंगे cloud हैं, कोई relationship नहीं। Board साठ students के बारे में कभी झूठ नहीं बोलता जैसा एक anecdote बोल सकता है। वह board एक scatter plot है।

Theory

Scatter, formally

plt.scatter(x, y) प्रति (x, y) pair एक marker खींचता है, किसी से जुड़ा नहीं।

इसे तब इस्तेमाल कीजिए जब दोनों axes measurements हों और सवाल हो "वे साथ कैसे move करते हैं?"

Cloud पढ़ना:

  • Dots ऊपर-दाएँ drift करते हुए: positive correlation (ज़्यादा attendance, ज़्यादा marks)।
  • नीचे-दाएँ drift करते: negative correlation।
  • बेढंगे: कोई visible relation नहीं।
  • Cloud से दूर एक dot: एक outlier: वह topper जो कभी attend नहीं करता, roll number से investigate करने लायक़।

Practical

Attendance सवाल, छह lines में जवाब दिया गया

import matplotlib.pyplot as plt
import pandas as pd

df = pd.read_csv('students_full.csv')   # roll, attendance, score

plt.scatter(df['attendance'], df['score'])
plt.title('Attendance vs DBMS score')
plt.xlabel('Attendance (%)')
plt.ylabel('DBMS score')
plt.show()

# optional polish: size and colour per point
# plt.scatter(df['attendance'], df['score'], s=20, c='teal')
# object API spelling (the syllabus's set_title):
# fig, ax = plt.subplots()
# ax.scatter(df['attendance'], df['score'])
# ax.set_title('Attendance vs DBMS score')
# ax.set_xlabel('Attendance (%)'); ax.set_ylabel('DBMS score')

This example runs in Gri-Learn on the web, where you can edit it and see the output.

Quiz

Scatter में dots साफ़ तौर पर ऊपर-दाएँ drift करते दिखते हैं, एक dot अकेला (35, 92) पर है। आप इसे कैसे पढ़ेंगे?

  1. कुल मिलाकर positive correlation, plus एक outlier: एक low-attendance high scorer जिसे individually check करना चाहिए
  2. कोई relationship नहीं: एक exception pattern को disprove करता है
  3. Negative correlation: वह अकेला dot direction तय करता है
  4. Chart invalid है क्योंकि एक point fit नहीं होता
Show the answer

कुल मिलाकर positive correlation, plus एक outlier: एक low-attendance high scorer जिसे individually check करना चाहिए

Cloud का drift story बताता है: ऊपर-दाएँ = positive correlation। एक अकेला दूर वाला dot एक outlier है: यह साठ students के pattern को पलटता नहीं, पर एक नज़र कमाता है (एक brilliant self-studier? एक data-entry slip?)। Patterns mass से पढ़े जाते हैं, exceptions flag किए जाते हैं, कोई भी दूसरे को मिटाता नहीं: यही dual reading exactly वह है जो examiners एक describe-this-plot answer में चाहते हैं।

Think first

Correlation एक verdict नहीं है

Scatter साबित करता है attendance और marks साथ बढ़ते हैं। Principal एक rule draft करता है: "attendance को 90% force कीजिए, marks बढ़ेंगे।" tap करने से पहले: उस jump में क्या flaw है, एक sentence में?

Show the answer

Correlation causation नहीं है: plot दिखाता है दोनों साथ MOVE करते हैं, यह नहीं कि एक दूसरे को CAUSE करता है। शायद sincere students दोनों ज़्यादा attend करते हैं और ज़्यादा पढ़ते हैं (दोनों को drive करता एक third factor); शरीरों को chairs में force करना symptom copy करता है, cause नहीं।

एक scatter investigate करना justify कर सकता है; अकेले यह rule justify नहीं कर सकता। किसी भी correlation के साथ यह एक sentence लिखना वही फ़र्क़ है analysis और superstition के बीच।

Watch out

ग़लत-chart trap

Students habit से plt.plot() पकड़ते हैं और साठ dots एक zigzag line से जुड़े पाते हैं: meaningless, क्योंकि students का x axis पर कोई order नहीं है: line एक ऐसे sequence का दावा करती है जो मौजूद नहीं।

Rule: records के पार दो measures के बीच relationship = scatter; एक ordered axis (tests, months) के साथ trend = line (अगला lesson)। और हमेशा की तरह: कोई title/xlabel/ylabel नहीं, कोई full marks नहीं।

Theory

जंगल में Scatter

Price बनाम demand, study hours बनाम CGPA, RAM बनाम app performance: business और research में हर "क्या ये दो चीज़ें साथ move करती हैं?" सवाल एक scatter plot से खुलता है। यह उस correlation coefficient का visual twin भी है जो आप BCA302 में compute करेंगे: number वह summarise करता है जो यह chart आपको DEKHNE देता है।

Summary

Key takeaways

  • plt.scatter(x, y): प्रति record एक dot, कोई connecting line नहीं; two-measure relationship सवालों के लिए।
  • Cloud पढ़िए: ऊपर-दाएँ drift = positive, नीचे-दाएँ = negative, बेढंगा = none।
  • दूर वाले dots outliers हैं: flag कीजिए और investigate कीजिए, उन्हें mass पलटने मत दीजिए।
  • Correlation causation नहीं है: plot co-movement दिखाता है, cause कभी नहीं।
  • हमेशा title, xlabel, ylabel (object API में ax.set_title)।
  • Records के पार relationships के लिए scatter; एक ordered axis के साथ trends के लिए line।
  • Memory hook: noticeboard पर साठ pins।

Study this properly

This page is the lesson to read. In Gri-Learn the same topic is a graded deck: the self-checks are scored and your weak topics are tracked. Free to start.

Start this topic

Already have an account? Sign in

More from Data Visualization using dataframe

Gri-Learn · syllabus-mapped B.C.A. lessons in English, Hindi and Gujarati