Calculating summary statistics (mean, median, mode, standard deviation)

mean(), median(), sd() और var() Unit 1 की statistics को एक-एक call में compute करते हैं (na.rm = TRUE reflex के साथ), summary() five-number overview report करता है, और R में famously कोई built-in mode function NAHI है।

9 min read · 9 cards · 2 checks

Read in: English · हिन्दी · ગુજરાતી


Theory

Unit 1, machine speed पर replayed

हफ़्तों पहले आपने screen-time mean, median और SD हाथ से compute किए: deviations, squares, square roots, पूरा ritual।

R में, ritual चार शब्द है:

mean(survey$screen, na.rm = TRUE)

वही number जो आपके paper ने produce किया। यही deal है जो यह lesson formalise करता है: हाथ के काम ने आपको सिखाया numbers का MEANING क्या है; आज के commands उन्हें instantly compute करते हैं: एक famous omission के साथ जो देश के हर beginner को पकड़ता है।

Theory

एक missing button वाला calculator

R के statistics keyboard में हर चीज़ के लिए एक key है: mean, median, sd, var, quantile...

...सिवाय mode के। mode() LABELLED वाला button एक prank है: इसे दबाइए और R आपको आपके data का storage mode बताता है ("numeric"): एक statistics सवाल का एक programming जवाब।

असली most-frequent-value को दो अन्य keys से assemble करना होगा: गिनने के लिए table(), विजेता चुनने के लिए which.max()।

Practical

पूरी Unit 1 toolkit, प्रति line एक

screen <- c(60, 90, 90, 120, 150, 180, 600, NA)

# central tendency
mean(screen, na.rm = TRUE)      # [1] 184.3   (outlier-dragged, as Unit 1 warned)
median(screen, na.rm = TRUE)    # [1] 120

# dispersion (SAMPLE formulas: divide by n-1)
sd(screen, na.rm = TRUE)
var(screen, na.rm = TRUE)
range(screen, na.rm = TRUE)     # min and max together
quantile(screen, na.rm = TRUE)  # 0% 25% 50% 75% 100%

# the five-number overview + mean, one call:
summary(screen)

# THE QUIRK: mode() is NOT the statistical mode
mode(screen)                    # [1] "numeric"   (storage mode!)

# the real mode, assembled:
t <- table(screen)
names(which.max(t))             # [1] "90"

# group-wise statistics: R's GROUP BY
# tapply(survey$screen, survey$stay, mean, na.rm = TRUE)
# aggregate(screen ~ stay, data = survey, FUN = mean)

Theory

तीन आदतें जो numbers को honest बनाती हैं

na.rm = TRUE, हमेशा सोचा गया: एक NA इनमें से हर function को चुप कर देता है (propagation rule): हर बार consciously decide कीजिए।

अपना divisor जानिए: sd() और var() sample formula (n-1) इस्तेमाल करते हैं, pandas से मेल खाते हुए: reports में "sample SD" कहिए, बिल्कुल जैसे dispersion lesson ने drill किया।

एक बार हाथ से verify कीजिए: उस छोटे dataset पर sd() चलाइए जो आपने Unit 1 में manually compute किया: वही answer वह trust बनाता है जो आपको हमेशा के लिए arithmetic delegate करने देता है।

Quiz

एक student mode(survey$screen) चलाता है सबसे frequent screen time की उम्मीद में और "numeric" पाता है। क्या हो रहा है, और असली mode क्या compute करता है?

  1. mode() R में STORAGE type return करता है; statistical mode names(which.max(table(x))) से assemble होता है
  2. Data में कोई repeated values नहीं हैं, तो R ने इसके बजाय अपना type print किया
  3. mode() को काम करने के लिए na.rm = TRUE चाहिए
  4. R में सिर्फ़ factors के modes होते हैं
Show the answer

mode() R में STORAGE type return करता है; statistical mode names(which.max(table(x))) से assemble होता है

R का mode() जवाब देता है "यह कैसे stored है?": statistics के साथ एक naming collision जिसने पीढ़ियों students को फँसाया: classic R gotcha, marks लायक़ बिल्कुल इसलिए क्योंकि यह इतना surprising है। असली mode: table(x) हर value गिनता है, which.max() सबसे बड़ी count ढूँढता है, names() जीतने वाली value पढ़ता है। Options B से D behaviours invent करते हैं: function data या arguments चाहे जो हों, कभी frequency compute नहीं करता।

Think first

Group-wise: dean का one-liner

Dean पूछते हैं: "average screen time, hostellers बनाम day scholars, एक command।" tapply(survey$screen, survey$stay, mean, na.rm = TRUE) इस्तेमाल करते हुए: describe कीजिए तीनों arguments में से हर एक क्या करता है, फिर tap कीजिए।

Show the answer

tapply(values, groups, function): survey$screen लीजिए (क्या summarise करना है), इसे survey$stay से split कीजिए (ढेर: hostel और day), हर ढेर पर mean apply कीजिए: output प्रति group एक labelled average है:

day hostel

142 201

यह R का GROUP BY है (aggregate(screen ~ stay, df, mean) वही बात एक formula से कहता है)। आपने ये ढेर BCA303 के GROUP BY lesson में हाथ से बनाए: वही idea, तीसरी syntax, दस सेकंड।

Watch out

Summary-statistics की slips

mode() कभी mode नहीं है: table + which.max, या आपका अपना two-line function।

NA silence: एक अकेला missing value mean/sd से NA return करता है: na.rm decision जवाब का हिस्सा है, afterthought नहीं।

quantile() defaults 0/25/50/75/100% पर: दूसरों के लिए explicitly पूछिए: 90th percentile के लिए quantile(x, 0.9)।

Theory

यह आगे कहाँ plug होता है

ये one-liners subject में बाक़ी हर चीज़ feed करते हैं: frequency tables categoricals गिनते हैं जिन्हें ये functions summarise नहीं कर सकते (अगला lesson), commands lesson पूरी surface consolidate करता है, और graphical finale वही draw करता है जो summary() describe करता है। अब आप course के अंदर bilingual भी हैं: df['screen'].mean() और mean(df$screen) वही thought हैं: portability असली syllabus है।

Summary

Key takeaways

  • mean, median, sd, var, range, quantile: प्रति एक call; summary() = five-number overview + mean।
  • na.rm = TRUE हर call पर एक conscious decision है: वरना NA इन्हें चुप कर देता है।
  • sd()/var() sample divisor (n-1) इस्तेमाल करते हैं: 'sample SD' कहिए, pandas से मेल खाते हुए।
  • mode() STORAGE type return करता है: असली mode names(which.max(table(x))) है।
  • Group-wise: tapply(values, groups, fun) या aggregate(y ~ group, df, fun): R का GROUP BY।
  • Commands को अपने Unit 1 हाथ के calculations के ख़िलाफ़ एक बार verify कीजिए: फिर हमेशा के लिए delegate कीजिए।
  • Memory hook: एक missing (mislabelled) button वाला calculator।

Study this properly

This page is the lesson to read. In Gri-Learn the same topic is a graded deck: the self-checks are scored and your weak topics are tracked. Free to start.

Start this topic

Already have an account? Sign in

More from Working with Data in R

Gri-Learn · syllabus-mapped B.C.A. lessons in English, Hindi and Gujarati

Calculating summary statistics (mean, median, mode, standard deviation) · Statistical Methods and Data Analysis (MDC-03) · Gri-Learn