Concepts of normal distribution

R चार prefixed commands से normal-distribution बोलता है: dnorm curve की height देता है, pnorm एक value के बायें area (probability) देता है, qnorm दिए percentile पर value देता है, और rnorm random samples देता है।

10 min read · 9 cards · 2 checks

Read in: English · हिन्दी · ગુજરાતી


Theory

z-table को चार commands से replace करना

Bell-curve lesson में आपने "180 cm से taller students का कितना हिस्सा है?" का जवाब 68-95-99.7 rule से दिया: लगभग 2.5%, क्योंकि 180 बिल्कुल 2 SDs बाहर बैठा।

पर 178 cm का क्या? या 183? Rule सिर्फ़ पूरी rings जानता है; उनके बीच, statisticians कभी printed z-tables पलटते थे।

R ने उन tables को एक चार-command family से retire कर दिया। Prefixes d, p, q, r एक बार सीखिए, और हर normal-curve सवाल एक line बन जाता है।

Theory

एक bell के चार सवाल

Bell curve के सामने खड़े होकर पूछिए:

  • dnorm: यहाँ curve कितनी HIGH है? (density: drawing सवाल)
  • pnorm: यहाँ के LEFT कितना area है? (probability: exam सवाल)
  • qnorm: 90% area कहाँ ख़त्म होता है? (quantile: cutoff सवाल)
  • rnorm: मुझे इस bell से RANDOM students दीजिए (simulation)

वही bell, चार दरवाज़े। pnorm और qnorm inverses हैं: एक values को areas में बदलता है, दूसरा areas को वापस values में।

Practical

Heights: mu = 168, sd = 6, हर सवाल का जवाब

# P(height <= 174): area LEFT of 174
pnorm(174, mean = 168, sd = 6)      # [1] 0.841

# P(taller than 180): right tail = 1 - left area
1 - pnorm(180, 168, 6)              # [1] 0.0228   (~the 2.5% ring!)
pnorm(180, 168, 6, lower.tail = FALSE)   # same, spelled directly

# P(between 162 and 174): big area minus small area
pnorm(174, 168, 6) - pnorm(162, 168, 6)  # [1] 0.683  (the 68% ring!)

# cutoff for the tallest 10%: qnorm is pnorm's inverse
qnorm(0.90, 168, 6)                 # [1] 175.7 cm

# defaults are the STANDARD normal (mean 0, sd 1):
pnorm(2)                            # [1] 0.977  = P(z <= 2)

# simulate 60 students; set.seed makes it reproducible
set.seed(42)
heights <- rnorm(60, 168, 6)
mean(heights)                       # close to 168, not exact: sampling!

# dnorm: curve HEIGHT (for drawing, next lesson), not probability
dnorm(168, 168, 6)                  # [1] 0.066

Theory

तीन question patterns

हर exam सवाल एक pattern में map होता है:

  • Left area ("at most x"): pnorm(x, mean, sd)।
  • Right area ("more than x"): 1 - pnorm(x, ...) (या lower.tail = FALSE): 1-minus R statistics में सबसे भूला keystroke है।
  • a और b के बीच: pnorm(b, ...) - pnorm(a, ...)।

और reverse direction ("कौन सा value top 10% mark करता है?"): qnorm(0.90, ...): नोटिस कीजिए यह LEFT की area लेता है, तो top 10% मतलब 0.90, 0.10 नहीं।

Quiz

Scores normal हैं mean 60, sd 8 के साथ। कौन सा command P(score > 70) compute करता है, और यह लगभग क्या return करता है?

  1. 1 - pnorm(70, 60, 8): लगभग 0.106
  2. pnorm(70, 60, 8): लगभग 0.894
  3. dnorm(70, 60, 8): exactly 70 की probability
  4. qnorm(70, 60, 8): 70th percentile
Show the answer

1 - pnorm(70, 60, 8): लगभग 0.106

pnorm LEFT area देता है, तो P(70 से ऊपर) को complement चाहिए: 1 - pnorm(70, 60, 8) ≈ 1 - 0.894 = 0.106। Option B ख़ुद left area है: OPPOSITE सवाल का जवाब, और सबसे common ग़लत submission। dnorm curve height है (एक continuous variable की एक exact point पर zero probability होती है: distributions lesson का law), और qnorm input के रूप में एक AREA (0 से 1) चाहता है, तो qnorm(70, ...) thought में एक type error है।

Think first

Scholarship cutoff, दोनों directions

CGPA roughly normal है: mean 7.0, sd 0.8। दो tasks, प्रति एक line: (1) कितना हिस्सा CGPA 8.5 से exceed करता है? (2) Top 5% को scholarship मिलती है: cutoff CGPA क्या है? tap करने से पहले दोनों draft कीजिए।

Show the answer

1. 1 - pnorm(8.5, 7.0, 0.8) ≈ 0.030: लगभग 3% 8.5 से exceed करते हैं।

2. qnorm(0.95, 7.0, 0.8) ≈ 8.32: scholarship line।

Mirror नोटिस कीजिए: task 1 value → area जाता है (pnorm), task 2 area → value जाता है (qnorm, 0.95 fed क्योंकि top पर 5% मतलब left पर 95%)। कौन सी direction एक सवाल travel करता है यह पढ़ना ही इस command family की पूरी skill है।

Watch out

तीन normal-command slips

dnorm probability नहीं है: यह curve height है (density): continuous variables की probabilities AREAS में रहती हैं: pnorm territory।

Forgotten 1-minus: हर "more than" सवाल को complement चाहिए (या lower.tail = FALSE)।

पहले model-check कीजिए: pnorm जवाब सिर्फ़ तभी exact हैं जब data (approximately) normal हो: भरोसा करने से पहले histogram + mean≈median check कीजिए: screen time fail होगा, heights pass।

Theory

Rule बनाम exact जवाब

नोटिस कीजिए 1 - pnorm(180, 168, 6) ने 0.0228 दिया जहाँ empirical rule ने 2.5% estimate किया: rule mental approximation है, pnorm precise instrument: reasoning में rule quote कीजिए, practice में pnorm से compute कीजिए। एक lesson बचा है: सब कुछ खींचना: आपके data के लिए hist, इसके ऊपर theoretical bell बिठाने के लिए curve(dnorm(...)), और CampusPulse story बंद होती है।

Summary

Key takeaways

  • d-p-q-r family: dnorm height, pnorm left area, qnorm area-to-value, rnorm random draws।
  • Defaults mean = 0, sd = 1: standard normal; pnorm(2) ≈ 0.977।
  • Right tails: 1 - pnorm(x, ...) या lower.tail = FALSE; between: pnorm(b) - pnorm(a)।
  • qnorm LEFT area लेता है: top 10% cutoff = qnorm(0.90, ...)।
  • pnorm 68-95-99.7 approximations के पीछे exact instrument है।
  • Exact answers पर भरोसा करने से पहले normality check कीजिए (hist shape, mean ≈ median)।
  • Memory hook: एक bell के चार दरवाज़े: height, area, cutoff, sample।

Study this properly

This page is the lesson to read. In Gri-Learn the same topic is a graded deck: the self-checks are scored and your weak topics are tracked. Free to start.

Start this topic

Already have an account? Sign in

More from Working with Data in R

Gri-Learn · syllabus-mapped B.C.A. lessons in English, Hindi and Gujarati

Concepts of normal distribution · Statistical Methods and Data Analysis (MDC-03) · Gri-Learn