Commands to explore and view data distributions graphically (Bell curve)

hist() आपके data का असली shape दिखाता है, boxplot(y ~ group) एक axis पर groups compare करता है, और curve(dnorm(...), add = TRUE) theoretical bell को histogram पर बिठाता है: subject की closing picture।

10 min read · 9 cards · 2 checks

Read in: English · हिन्दी · ગુજરાતી


Theory

Subject एक picture के साथ ख़त्म होता है

पच्चीस lessons के numbers एक closing image लायक़ हैं। Dean आख़िरी सवाल पूछते हैं: "आप कहते रहते हैं heights normal हैं और screen time नहीं: दिखाइए।"

Proof एक chart है: आपका असली data एक histogram के रूप में, और theoretical bell curve इसके ऊपर खींचा गया। जहाँ curve bars को गले लगाता है, normal model टिकता है और हर pnorm जवाब trustworthy है। जहाँ नहीं, model को झुकना होगा, data को नहीं।

आज: वे R commands जो वह verdict खींचते हैं।

Theory

Tailor की fitting

Bell curve एक ready-made suit है (दो measurements: mean और sd)। आपका data customer है।

हistogram पर curve(dnorm(...)) बिठाना fitting है: heights लगभग perfectly suit में फिसल जाती हैं; screen time बायें उभरता है और एक लंबी right sleeve खींचता है: कोई tailoring इसे नहीं बचाती।

Statistics बिल्कुल यहाँ honest है: model को DATA पर fit कीजिए, और fitting photo दिखाइए, कभी मान मत लीजिए suit fit करता है।

Practical

Finale: histogram, groups, और bell overlay

survey <- read.csv("survey_clean.csv")

# 1. the data's shape
hist(survey$height, breaks = 10,
     main = "Heights of surveyed students",
     xlab = "Height (cm)", ylab = "Frequency")

# 2. groups side by side: the formula reads 'screen BY stay'
boxplot(screen ~ stay, data = survey,
        main = "Screen time by residence", ylab = "Minutes")

# 3. THE BELL OVERLAY: density scale first, then the curve
hist(survey$height, freq = FALSE,        # y-axis = density now
     main = "Heights vs the normal model", xlab = "Height (cm)")
curve(dnorm(x, mean = mean(survey$height, na.rm = TRUE),
               sd  = sd(survey$height,  na.rm = TRUE)),
      add = TRUE, col = "red", lwd = 2)  # add = TRUE: draw ON TOP
abline(v = mean(survey$height, na.rm = TRUE), lty = 2)

# same overlay on screen time: watch the suit NOT fit
hist(survey$screen, freq = FALSE, main = "Screen time vs the bell")
curve(dnorm(x, mean(survey$screen, na.rm = TRUE),
               sd(survey$screen, na.rm = TRUE)),
      add = TRUE, col = "red", lwd = 2)

Theory

दो fittings पढ़ना

Heights: bars symmetrically ऊपर-नीचे होते हैं; red bell उन्हें क़रीब से track करता है; dashed mean line picture को evenly बाँटती है। Verdict: normal model fits करता है: heights के बारे में हर pnorm जवाब solid ground पर है।

Screen time: bars बायें ढेर होते हैं एक लंबी right tail के साथ (binge-watchers); bell दायें empty space के ऊपर तैरता है और बायें ढेर को undershoot करता है। Verdict: normal नहीं: medians और IQRs report कीजिए, और किसी भी pnorm जवाब को suspicion से देखिए।

दो pictures, और इस subject का हर claim visible बन जाता है।

Quiz

एक student hist(heights) चलाता है फिर curve(dnorm(x, 168, 6), add = TRUE), और red curve x-axis के साथ flat, invisible रेंगता है। क्या ग़लत हुआ?

  1. Histogram FREQUENCY scale पर है (counts ~15 तक) जबकि dnorm की density लगभग 0.07 पर peak करती है: hist को density scale share करने के लिए freq = FALSE चाहिए
  2. curve() से dnorm plot नहीं किया जा सकता
  3. add = TRUE ने curve मिटा दिया
  4. 6 का sd खींचने के लिए बहुत छोटा है
Show the answer

Histogram FREQUENCY scale पर है (counts ~15 तक) जबकि dnorm की density लगभग 0.07 पर peak करती है: hist को density scale share करने के लिए freq = FALSE चाहिए

Frequencies students गिनते हैं (0 से लगभग 15); densities 1 तक integrate होती हैं (यहाँ लगभग 0.07 पर peak): एक chart पर खींची, density curve frequency mountains के नीचे एक flat फुसफुसाहट है। freq = FALSE histogram के y-axis को density में rescale करता है ताकि दोनों वही units बोलें: सबसे common overlay bug। add = TRUE मिटाने का उल्टा करता है (यह histogram preserve करता है); इसके बिना, curve() plot को REPLACE कर देता: sibling trap।

Think first

Verdict deliver कीजिए

आप dean को दोनों overlay charts दिखाते हैं। दो-sentence verdict compose कीजिए: प्रति variable एक sentence, हर एक picture में evidence AND reporting के लिए practical consequence नाम देते हुए। फिर tap कीजिए।

Show the answer

Heights: histogram symmetric है और fitted bell इसे क़रीब से track करता है, तो normal model लागू होता है: means, SDs और pnorm-based percentages (जैसे 180 cm से ऊपर का हिस्सा) trustworthy हैं।

Screen time: distribution right-skewed है एक लंबी tail के साथ जिसे bell follow नहीं कर सकता, तो normal-based जवाब गुमराह करेंगे: हम median और IQR report करते हैं, और heavy-user tail को अलग से flag करते हैं।

Evidence + consequence, दो बार: यही एक statistics report miniature में है, और exam का dream answer।

Watch out

Closing traps

किसी भी dnorm overlay से पहले freq = FALSE: frequency और density scales कभी नहीं मिलते।

curve() पर add = TRUE: इसके बिना bell आपका histogram replace कर देता है।

Labels marks हैं: हर chart पर main, xlab, ylab: बिना label वाला chart एक अधूरा जवाब है।

Model data की service करता है: कभी उस shape पर normal conclusions मत थोपिए जिसने suit reject कर दिया।

Theory

CampusPulse, बंद

पूरे arc पर पीछे मुड़कर देखिए: आपने एक sample plan किया, इसकी statistics हाथ से compute कीं, उनके पीछे की probability सीखी, R में असली data import और clean किया, और अपनी ही survey पर bell curve खींचकर ख़त्म किया। वह loop (design, collect, clean, summarise, model, visualise, verdict) कोई syllabus नहीं है: यह data work का असली रोज़ाना shape है। BCA302 complete: qqnorm() और formal tests उनका इंतज़ार करते हैं जो curious बने रहते हैं।

Summary

Key takeaways

  • hist(x, breaks = n) data का shape दिखाता है; main/xlab/ylab अनिवार्य manners हैं।
  • boxplot(y ~ group, data = df): side-by-side five-number comparisons: formula 'y by group' पढ़ता है।
  • Overlay recipe: hist(x, freq = FALSE) फिर curve(dnorm(x, mean, sd), add = TRUE)।
  • freq = FALSE scales align करता है; add = TRUE histogram preserve करता है: दोनों non-negotiable।
  • Close fit = normal model trustworthy; visible skew = इसके बजाय median/IQR report कीजिए।
  • abline() reference lines जोड़ता है; qqnorm() इस course से आगे का formal normality tool है।
  • Memory hook: tailor की fitting: क्या ready-made bell आपके data को suit करता है?

Study this properly

This page is the lesson to read. In Gri-Learn the same topic is a graded deck: the self-checks are scored and your weak topics are tracked. Free to start.

Start this topic

Already have an account? Sign in

More from Working with Data in R

Gri-Learn · syllabus-mapped B.C.A. lessons in English, Hindi and Gujarati

Commands to explore and view data distributions graphically (Bell curve) · Statistical Methods and Data Analysis (MDC-03) · Gri-Learn