Theory
Does attendance actually buy marks?
The eternal staff-room debate reaches ResultDesk: "students who attend more, score more... right?"
You have both columns: attendance percentage and DBMS score, one pair per student. Averages will not settle this: the question is about a relationship between two variables, across sixty students at once.
There is a chart built precisely for that question, and it is the simplest one in the whole library: one dot per student.
Theory
Sixty pins on a noticeboard
Give every student a pin. Slide it right for more attendance, up for more marks, and press it into the board.
Step back. If the pins drift upward as they go right, the two things rise together. If they fall going right, one rises as the other drops. If they are a shapeless cloud, no relationship. The board never lies about sixty students the way one anecdote can. That board is a scatter plot.
Theory
Scatter, formally
plt.scatter(x, y) draws one marker per (x, y) pair, connected by nothing.
Use it when both axes are measurements and the question is "how do they move together?"
Reading the cloud:
- Dots drifting up-right: positive correlation (more attendance, more marks).
- Drifting down-right: negative correlation.
- Shapeless: no visible relation.
- A dot far from the cloud: an outlier: the topper who never attends, worth investigating by roll number.
Practical
The attendance question, answered in six lines
import matplotlib.pyplot as plt
import pandas as pd
df = pd.read_csv('students_full.csv') # roll, attendance, score
plt.scatter(df['attendance'], df['score'])
plt.title('Attendance vs DBMS score')
plt.xlabel('Attendance (%)')
plt.ylabel('DBMS score')
plt.show()
# optional polish: size and colour per point
# plt.scatter(df['attendance'], df['score'], s=20, c='teal')
# object API spelling (the syllabus's set_title):
# fig, ax = plt.subplots()
# ax.scatter(df['attendance'], df['score'])
# ax.set_title('Attendance vs DBMS score')
# ax.set_xlabel('Attendance (%)'); ax.set_ylabel('DBMS score')
This example runs in Gri-Learn on the web, where you can edit it and see the output.
Quiz
The scatter shows dots clearly drifting up and to the right, with one dot alone at (35, 92). How do you read it?
- Positive correlation overall, plus one outlier: a low-attendance high scorer worth checking individually
- No relationship: one exception disproves the pattern
- Negative correlation: the lone dot sets the direction
- The chart is invalid because one point does not fit
Show the answer
Positive correlation overall, plus one outlier: a low-attendance high scorer worth checking individually
The cloud's drift carries the story: up-right = positive correlation. A single distant dot is an outlier: it does not overturn sixty students' pattern, but it earns a look (a brilliant self-studier? a data-entry slip?). Patterns are read from the mass, exceptions are flagged, neither erases the other: that dual reading is exactly what examiners want in a describe-this-plot answer.
Think first
Correlation is not a verdict
The scatter proves attendance and marks rise together. The principal drafts a rule: "force attendance to 90%, marks will rise." Before tapping: what is the flaw in that jump, in one sentence?
Show the answer
Correlation is not causation: the plot shows the two MOVE together, not that one CAUSES the other. Maybe sincere students both attend more and study more (a third factor driving both); forcing bodies into chairs copies the symptom, not the cause.
One scatter can justify investigating; it cannot alone justify the rule. Writing this one sentence next to any correlation you report is the difference between analysis and superstition.
Watch out
The wrong-chart trap
Students reach for plt.plot() out of habit and get sixty dots connected by a zigzag line: meaningless, because students have no order along the x axis: the line implies a sequence that does not exist.
Rule: relationship between two measures across records = scatter; trend along an ordered axis (tests, months) = line (next lesson). And as always: no title/xlabel/ylabel, no full marks.
Theory
Scatter in the wild
Price vs demand, study hours vs CGPA, RAM vs app performance: every "do these two things move together?" question in business and research opens with a scatter plot. It is also the visual twin of the correlation coefficient you will compute in BCA302: the number summarises what this chart lets you SEE.
Summary
Key takeaways
- plt.scatter(x, y): one dot per record, no connecting line; for two-measure relationship questions.
- Read the cloud: up-right drift = positive, down-right = negative, shapeless = none.
- Distant dots are outliers: flag and investigate, do not let them overturn the mass.
- Correlation is not causation: the plot shows co-movement, never the cause.
- Always title, xlabel, ylabel (ax.set_title in the object API).
- Scatter for relationships across records; line for trends along an ordered axis.
- Memory hook: sixty pins on a noticeboard.