Theory
Loop बंद करने के लिए एक card
Unit 1 ने measures हाथ से सिखाए। तब से बिखरे lessons ने उनके commands सिखाए। जो exam सवाल बचा है वह consolidation है: "किसी दिए dataset के लिए central tendency और dispersion के measures के लिए R commands लिखिए।"
यह lesson वह जवाब है, एक बार assembled: हर measure, इसका command, इसका trap: वह reference card जिससे आप revise करते हैं। यहाँ कुछ नया नहीं है; यहाँ सब कुछ अब एक card है।
Theory
Examiner की shopping list
एक full-marks descriptive answer के लिए examiner की checklist की कल्पना कीजिए:
- Centre कहाँ है? (mean, median, mode: judgement से चुनिए)
- Spread कितनी चौड़ी है? (range, IQR, sd: तीन widths, तीन meanings)
- क्या आपने holes handle किए? (na.rm decision, stated)
- Sample या population? (अपना divisor नाम दीजिए)
List पर चार lines: नीचे वाला card चारों ख़रीदता है।
At a glance
Reference card
| Measure | Command | सावधान |
|---|---|---|
| Mean | mean(x, na.rm = TRUE) | Outlier-dragged |
| Median | median(x, na.rm = TRUE) | Robust centre |
| Mode | names(which.max(table(x))) | कोई built-in नहीं; mode() एक prank है |
| Range | range(x) फिर diff(range(x)) | range() min AND max return करता है |
| IQR | IQR(x, na.rm = TRUE) | Q3 - Q1, box width |
| Variance / SD | var(x), sd(x) | Sample divisor (n-1) |
| Overview | summary(x), quantile(x) | fivenum(x) = five-number summary |
Practical
पूरी descriptive report, बारह lines
survey <- read.csv("survey_clean.csv")
x <- survey$screen
# centre
mean(x, na.rm = TRUE) # 171.4
median(x, na.rm = TRUE) # 135
get_mode <- function(v) names(which.max(table(v)))
get_mode(x) # "90" (your first user function!)
# spread
diff(range(x, na.rm = TRUE)) # max - min, ONE number
IQR(x, na.rm = TRUE) # the box plot's box width
sd(x, na.rm = TRUE) # sample SD (divides by n-1)
var(x, na.rm = TRUE) # sd squared
# overview + per-group spread
summary(x)
tapply(survey$screen, survey$stay, sd, na.rm = TRUE)
# day hostel <- who is more CONSISTENT?
# 38.2 71.5
Theory
Commands आसान हैं; sentence skill है
Report block numbers print करता है; marks उन्हें ज़ोर से पढ़ने से आते हैं:
- mean 171 बनाम median 135: mean काफ़ी ऊपर बैठा है: right skew, typical के रूप में median report कीजिए।
- hostel sd 71.5 बनाम day 38.2: hostellers दोगुने scattered हैं: एक average उन्हें misrepresent करेगा।
- mode "90": सबसे common अकेली habit।
हर number एक interpreting sentence पाता है: यह आदत एक computation answer को statistics answer में बदल देती है।
Quiz
एक student "the range" को एक number चाहता है और range(x) लिखता है, [1] 30 600 पाता है। क्या हुआ, और अकेला number क्या देता है?
- range() min AND max pair return करता है; अकेला spread number diff(range(x)) है, यहाँ 570
- NA values के कारण data के दो ranges हैं
- 100 से लंबे vectors के लिए range() broken है
- IQR(x) इस्तेमाल कीजिए: IQR और range एक ही measure हैं
Show the answer
range() min AND max pair return करता है; अकेला spread number diff(range(x)) है, यहाँ 570
R का range() आपको दो endpoints देता है (30 और 600); textbook का "range = max - min" चाहता है diff(range(x)) या max(x) - min(x): 570। Option D एक और confusion smuggle करता है जो यह card guard करता है: IQR है Q3 - Q1 (middle 50% की width, outlier-resistant), full range से अलग और आमतौर पर बेहतर spread measure। दो नाम, दो measures, एक classic mix-up।
Think first
Exam answer cold लिखिए
Exam task: "vector m में stored marks के लिए (जिसमें NA हो सकता है), mean, median, mode, standard deviation और IQR के लिए R commands लिखिए, और एक precaution बताइए।" tap करने से पहले पाँचों lines और precaution draft कीजिए।
Show the answer
mean(m, na.rm = TRUE)
median(m, na.rm = TRUE)
names(which.max(table(m))) # R has no built-in statistical mode
sd(m, na.rm = TRUE) # sample SD, divisor n-1
IQR(m, na.rm = TRUE)
Precaution: हर call पर na.rm = TRUE (वरना एक NA NA return करता है), और mode workaround नोट कीजिए क्योंकि mode() storage type return करता है। वह answer, verbatim, पूरे marks है: यह card इसलिए मौजूद है ताकि आप इसे आधी नींद में produce कर सकें।
Watch out
Type toolkit decide करता है
ये commands numeric variables को serve करते हैं। mean() को एक factor दीजिए और R एक warning के साथ NA return करता है: सही तरीक़े से, क्योंकि Unit 1 का law अभी भी लागू है: categorical data को table()/prop.table() मिलता है (पिछला lesson), numerical data को यह card। जो variable-type सवाल आपने पहले सीखा वह उन हर command का router है जो आपने बाद में सीखा।
Theory
आपने अभी अपना पहला function लिखा
get_mode उस logic के चारों तरफ़ syntax के तीन tokens हैं जो आप पहले से रखते थे: function(v) एक recipe wrap करता है, और get_mode(x) call करना इसे replay करता है: R में असली programming का दरवाज़ा। दो lessons बचे हैं: normal distribution के R commands (pnorm और family), फिर graphical finale जहाँ आज के card का हर number एक picture बनता है।
Summary
Key takeaways
- Centre: mean, median (na.rm = TRUE), mode names(which.max(table(x))) के ज़रिए।
- Spread: पूरी width के लिए diff(range(x)), robust middle-50% width के लिए IQR(x), sd/var (sample divisor)।
- Overview: summary(x), quantile(x), fivenum(x)।
- range() min-max PAIR return करता है, एक number नहीं।
- हर number को एक sentence में interpret कीजिए: skew (mean बनाम median), consistency (sd), typical (mode)।
- Numeric variables यह card इस्तेमाल करते हैं; categoricals table() इस्तेमाल करते हैं: type toolkit route करता है।
- Memory hook: examiner की चार-line shopping list।