Theory
Unit 1, machine speed पर replayed
हफ़्तों पहले आपने screen-time mean, median और SD हाथ से compute किए: deviations, squares, square roots, पूरा ritual।
R में, ritual चार शब्द है:
mean(survey$screen, na.rm = TRUE)
वही number जो आपके paper ने produce किया। यही deal है जो यह lesson formalise करता है: हाथ के काम ने आपको सिखाया numbers का MEANING क्या है; आज के commands उन्हें instantly compute करते हैं: एक famous omission के साथ जो देश के हर beginner को पकड़ता है।
Theory
एक missing button वाला calculator
R के statistics keyboard में हर चीज़ के लिए एक key है: mean, median, sd, var, quantile...
...सिवाय mode के। mode() LABELLED वाला button एक prank है: इसे दबाइए और R आपको आपके data का storage mode बताता है ("numeric"): एक statistics सवाल का एक programming जवाब।
असली most-frequent-value को दो अन्य keys से assemble करना होगा: गिनने के लिए table(), विजेता चुनने के लिए which.max()।
Practical
पूरी Unit 1 toolkit, प्रति line एक
screen <- c(60, 90, 90, 120, 150, 180, 600, NA)
# central tendency
mean(screen, na.rm = TRUE) # [1] 184.3 (outlier-dragged, as Unit 1 warned)
median(screen, na.rm = TRUE) # [1] 120
# dispersion (SAMPLE formulas: divide by n-1)
sd(screen, na.rm = TRUE)
var(screen, na.rm = TRUE)
range(screen, na.rm = TRUE) # min and max together
quantile(screen, na.rm = TRUE) # 0% 25% 50% 75% 100%
# the five-number overview + mean, one call:
summary(screen)
# THE QUIRK: mode() is NOT the statistical mode
mode(screen) # [1] "numeric" (storage mode!)
# the real mode, assembled:
t <- table(screen)
names(which.max(t)) # [1] "90"
# group-wise statistics: R's GROUP BY
# tapply(survey$screen, survey$stay, mean, na.rm = TRUE)
# aggregate(screen ~ stay, data = survey, FUN = mean)
Theory
तीन आदतें जो numbers को honest बनाती हैं
na.rm = TRUE, हमेशा सोचा गया: एक NA इनमें से हर function को चुप कर देता है (propagation rule): हर बार consciously decide कीजिए।
अपना divisor जानिए: sd() और var() sample formula (n-1) इस्तेमाल करते हैं, pandas से मेल खाते हुए: reports में "sample SD" कहिए, बिल्कुल जैसे dispersion lesson ने drill किया।
एक बार हाथ से verify कीजिए: उस छोटे dataset पर sd() चलाइए जो आपने Unit 1 में manually compute किया: वही answer वह trust बनाता है जो आपको हमेशा के लिए arithmetic delegate करने देता है।
Quiz
एक student mode(survey$screen) चलाता है सबसे frequent screen time की उम्मीद में और "numeric" पाता है। क्या हो रहा है, और असली mode क्या compute करता है?
- mode() R में STORAGE type return करता है; statistical mode names(which.max(table(x))) से assemble होता है
- Data में कोई repeated values नहीं हैं, तो R ने इसके बजाय अपना type print किया
- mode() को काम करने के लिए na.rm = TRUE चाहिए
- R में सिर्फ़ factors के modes होते हैं
Show the answer
mode() R में STORAGE type return करता है; statistical mode names(which.max(table(x))) से assemble होता है
R का mode() जवाब देता है "यह कैसे stored है?": statistics के साथ एक naming collision जिसने पीढ़ियों students को फँसाया: classic R gotcha, marks लायक़ बिल्कुल इसलिए क्योंकि यह इतना surprising है। असली mode: table(x) हर value गिनता है, which.max() सबसे बड़ी count ढूँढता है, names() जीतने वाली value पढ़ता है। Options B से D behaviours invent करते हैं: function data या arguments चाहे जो हों, कभी frequency compute नहीं करता।
Think first
Group-wise: dean का one-liner
Dean पूछते हैं: "average screen time, hostellers बनाम day scholars, एक command।" tapply(survey$screen, survey$stay, mean, na.rm = TRUE) इस्तेमाल करते हुए: describe कीजिए तीनों arguments में से हर एक क्या करता है, फिर tap कीजिए।
Show the answer
tapply(values, groups, function): survey$screen लीजिए (क्या summarise करना है), इसे survey$stay से split कीजिए (ढेर: hostel और day), हर ढेर पर mean apply कीजिए: output प्रति group एक labelled average है:
day hostel
142 201
यह R का GROUP BY है (aggregate(screen ~ stay, df, mean) वही बात एक formula से कहता है)। आपने ये ढेर BCA303 के GROUP BY lesson में हाथ से बनाए: वही idea, तीसरी syntax, दस सेकंड।
Watch out
Summary-statistics की slips
mode() कभी mode नहीं है: table + which.max, या आपका अपना two-line function।
NA silence: एक अकेला missing value mean/sd से NA return करता है: na.rm decision जवाब का हिस्सा है, afterthought नहीं।
quantile() defaults 0/25/50/75/100% पर: दूसरों के लिए explicitly पूछिए: 90th percentile के लिए quantile(x, 0.9)।
Theory
यह आगे कहाँ plug होता है
ये one-liners subject में बाक़ी हर चीज़ feed करते हैं: frequency tables categoricals गिनते हैं जिन्हें ये functions summarise नहीं कर सकते (अगला lesson), commands lesson पूरी surface consolidate करता है, और graphical finale वही draw करता है जो summary() describe करता है। अब आप course के अंदर bilingual भी हैं: df['screen'].mean() और mean(df$screen) वही thought हैं: portability असली syllabus है।
Summary
Key takeaways
- mean, median, sd, var, range, quantile: प्रति एक call; summary() = five-number overview + mean।
- na.rm = TRUE हर call पर एक conscious decision है: वरना NA इन्हें चुप कर देता है।
- sd()/var() sample divisor (n-1) इस्तेमाल करते हैं: 'sample SD' कहिए, pandas से मेल खाते हुए।
- mode() STORAGE type return करता है: असली mode names(which.max(table(x))) है।
- Group-wise: tapply(values, groups, fun) या aggregate(y ~ group, df, fun): R का GROUP BY।
- Commands को अपने Unit 1 हाथ के calculations के ख़िलाफ़ एक बार verify कीजिए: फिर हमेशा के लिए delegate कीजिए।
- Memory hook: एक missing (mislabelled) button वाला calculator।