Theory
z-table को चार commands से replace करना
Bell-curve lesson में आपने "180 cm से taller students का कितना हिस्सा है?" का जवाब 68-95-99.7 rule से दिया: लगभग 2.5%, क्योंकि 180 बिल्कुल 2 SDs बाहर बैठा।
पर 178 cm का क्या? या 183? Rule सिर्फ़ पूरी rings जानता है; उनके बीच, statisticians कभी printed z-tables पलटते थे।
R ने उन tables को एक चार-command family से retire कर दिया। Prefixes d, p, q, r एक बार सीखिए, और हर normal-curve सवाल एक line बन जाता है।
Theory
एक bell के चार सवाल
Bell curve के सामने खड़े होकर पूछिए:
- dnorm: यहाँ curve कितनी HIGH है? (density: drawing सवाल)
- pnorm: यहाँ के LEFT कितना area है? (probability: exam सवाल)
- qnorm: 90% area कहाँ ख़त्म होता है? (quantile: cutoff सवाल)
- rnorm: मुझे इस bell से RANDOM students दीजिए (simulation)
वही bell, चार दरवाज़े। pnorm और qnorm inverses हैं: एक values को areas में बदलता है, दूसरा areas को वापस values में।
Practical
Heights: mu = 168, sd = 6, हर सवाल का जवाब
# P(height <= 174): area LEFT of 174
pnorm(174, mean = 168, sd = 6) # [1] 0.841
# P(taller than 180): right tail = 1 - left area
1 - pnorm(180, 168, 6) # [1] 0.0228 (~the 2.5% ring!)
pnorm(180, 168, 6, lower.tail = FALSE) # same, spelled directly
# P(between 162 and 174): big area minus small area
pnorm(174, 168, 6) - pnorm(162, 168, 6) # [1] 0.683 (the 68% ring!)
# cutoff for the tallest 10%: qnorm is pnorm's inverse
qnorm(0.90, 168, 6) # [1] 175.7 cm
# defaults are the STANDARD normal (mean 0, sd 1):
pnorm(2) # [1] 0.977 = P(z <= 2)
# simulate 60 students; set.seed makes it reproducible
set.seed(42)
heights <- rnorm(60, 168, 6)
mean(heights) # close to 168, not exact: sampling!
# dnorm: curve HEIGHT (for drawing, next lesson), not probability
dnorm(168, 168, 6) # [1] 0.066
Theory
तीन question patterns
हर exam सवाल एक pattern में map होता है:
- Left area ("at most x"):
pnorm(x, mean, sd)। - Right area ("more than x"):
1 - pnorm(x, ...)(या lower.tail = FALSE): 1-minus R statistics में सबसे भूला keystroke है। - a और b के बीच:
pnorm(b, ...) - pnorm(a, ...)।
और reverse direction ("कौन सा value top 10% mark करता है?"): qnorm(0.90, ...): नोटिस कीजिए यह LEFT की area लेता है, तो top 10% मतलब 0.90, 0.10 नहीं।
Quiz
Scores normal हैं mean 60, sd 8 के साथ। कौन सा command P(score > 70) compute करता है, और यह लगभग क्या return करता है?
- 1 - pnorm(70, 60, 8): लगभग 0.106
- pnorm(70, 60, 8): लगभग 0.894
- dnorm(70, 60, 8): exactly 70 की probability
- qnorm(70, 60, 8): 70th percentile
Show the answer
1 - pnorm(70, 60, 8): लगभग 0.106
pnorm LEFT area देता है, तो P(70 से ऊपर) को complement चाहिए: 1 - pnorm(70, 60, 8) ≈ 1 - 0.894 = 0.106। Option B ख़ुद left area है: OPPOSITE सवाल का जवाब, और सबसे common ग़लत submission। dnorm curve height है (एक continuous variable की एक exact point पर zero probability होती है: distributions lesson का law), और qnorm input के रूप में एक AREA (0 से 1) चाहता है, तो qnorm(70, ...) thought में एक type error है।
Think first
Scholarship cutoff, दोनों directions
CGPA roughly normal है: mean 7.0, sd 0.8। दो tasks, प्रति एक line: (1) कितना हिस्सा CGPA 8.5 से exceed करता है? (2) Top 5% को scholarship मिलती है: cutoff CGPA क्या है? tap करने से पहले दोनों draft कीजिए।
Show the answer
1. 1 - pnorm(8.5, 7.0, 0.8) ≈ 0.030: लगभग 3% 8.5 से exceed करते हैं।
2. qnorm(0.95, 7.0, 0.8) ≈ 8.32: scholarship line।
Mirror नोटिस कीजिए: task 1 value → area जाता है (pnorm), task 2 area → value जाता है (qnorm, 0.95 fed क्योंकि top पर 5% मतलब left पर 95%)। कौन सी direction एक सवाल travel करता है यह पढ़ना ही इस command family की पूरी skill है।
Watch out
तीन normal-command slips
dnorm probability नहीं है: यह curve height है (density): continuous variables की probabilities AREAS में रहती हैं: pnorm territory।
Forgotten 1-minus: हर "more than" सवाल को complement चाहिए (या lower.tail = FALSE)।
पहले model-check कीजिए: pnorm जवाब सिर्फ़ तभी exact हैं जब data (approximately) normal हो: भरोसा करने से पहले histogram + mean≈median check कीजिए: screen time fail होगा, heights pass।
Theory
Rule बनाम exact जवाब
नोटिस कीजिए 1 - pnorm(180, 168, 6) ने 0.0228 दिया जहाँ empirical rule ने 2.5% estimate किया: rule mental approximation है, pnorm precise instrument: reasoning में rule quote कीजिए, practice में pnorm से compute कीजिए। एक lesson बचा है: सब कुछ खींचना: आपके data के लिए hist, इसके ऊपर theoretical bell बिठाने के लिए curve(dnorm(...)), और CampusPulse story बंद होती है।
Summary
Key takeaways
- d-p-q-r family: dnorm height, pnorm left area, qnorm area-to-value, rnorm random draws।
- Defaults mean = 0, sd = 1: standard normal; pnorm(2) ≈ 0.977।
- Right tails: 1 - pnorm(x, ...) या lower.tail = FALSE; between: pnorm(b) - pnorm(a)।
- qnorm LEFT area लेता है: top 10% cutoff = qnorm(0.90, ...)।
- pnorm 68-95-99.7 approximations के पीछे exact instrument है।
- Exact answers पर भरोसा करने से पहले normality check कीजिए (hist shape, mean ≈ median)।
- Memory hook: एक bell के चार दरवाज़े: height, area, cutoff, sample।