Commands for measures of central tendency and dispersion

एक reference card R की पूरी descriptive surface रखता है: centre के लिए mean/median और assembled mode, spread के लिए range/IQR/var/sd, overview के लिए quantile और summary, हर एक में na.rm consciously decided।

9 min read · 10 cards · 2 checks

Read in: English · हिन्दी · ગુજરાતી


Theory

Loop बंद करने के लिए एक card

Unit 1 ने measures हाथ से सिखाए। तब से बिखरे lessons ने उनके commands सिखाए। जो exam सवाल बचा है वह consolidation है: "किसी दिए dataset के लिए central tendency और dispersion के measures के लिए R commands लिखिए।"

यह lesson वह जवाब है, एक बार assembled: हर measure, इसका command, इसका trap: वह reference card जिससे आप revise करते हैं। यहाँ कुछ नया नहीं है; यहाँ सब कुछ अब एक card है।

Theory

Examiner की shopping list

एक full-marks descriptive answer के लिए examiner की checklist की कल्पना कीजिए:

  • Centre कहाँ है? (mean, median, mode: judgement से चुनिए)
  • Spread कितनी चौड़ी है? (range, IQR, sd: तीन widths, तीन meanings)
  • क्या आपने holes handle किए? (na.rm decision, stated)
  • Sample या population? (अपना divisor नाम दीजिए)

List पर चार lines: नीचे वाला card चारों ख़रीदता है।

At a glance

Reference card

MeasureCommandसावधान
Meanmean(x, na.rm = TRUE)Outlier-dragged
Medianmedian(x, na.rm = TRUE)Robust centre
Modenames(which.max(table(x)))कोई built-in नहीं; mode() एक prank है
Rangerange(x) फिर diff(range(x))range() min AND max return करता है
IQRIQR(x, na.rm = TRUE)Q3 - Q1, box width
Variance / SDvar(x), sd(x)Sample divisor (n-1)
Overviewsummary(x), quantile(x)fivenum(x) = five-number summary

Practical

पूरी descriptive report, बारह lines

survey <- read.csv("survey_clean.csv")
x <- survey$screen

# centre
mean(x, na.rm = TRUE)          # 171.4
median(x, na.rm = TRUE)        # 135
get_mode <- function(v) names(which.max(table(v)))
get_mode(x)                    # "90"    (your first user function!)

# spread
diff(range(x, na.rm = TRUE))   # max - min, ONE number
IQR(x, na.rm = TRUE)           # the box plot's box width
sd(x, na.rm = TRUE)            # sample SD (divides by n-1)
var(x, na.rm = TRUE)           # sd squared

# overview + per-group spread
summary(x)
tapply(survey$screen, survey$stay, sd, na.rm = TRUE)
#   day  hostel     <- who is more CONSISTENT?
#  38.2    71.5

Theory

Commands आसान हैं; sentence skill है

Report block numbers print करता है; marks उन्हें ज़ोर से पढ़ने से आते हैं:

  • mean 171 बनाम median 135: mean काफ़ी ऊपर बैठा है: right skew, typical के रूप में median report कीजिए।
  • hostel sd 71.5 बनाम day 38.2: hostellers दोगुने scattered हैं: एक average उन्हें misrepresent करेगा।
  • mode "90": सबसे common अकेली habit।

हर number एक interpreting sentence पाता है: यह आदत एक computation answer को statistics answer में बदल देती है।

Quiz

एक student "the range" को एक number चाहता है और range(x) लिखता है, [1] 30 600 पाता है। क्या हुआ, और अकेला number क्या देता है?

  1. range() min AND max pair return करता है; अकेला spread number diff(range(x)) है, यहाँ 570
  2. NA values के कारण data के दो ranges हैं
  3. 100 से लंबे vectors के लिए range() broken है
  4. IQR(x) इस्तेमाल कीजिए: IQR और range एक ही measure हैं
Show the answer

range() min AND max pair return करता है; अकेला spread number diff(range(x)) है, यहाँ 570

R का range() आपको दो endpoints देता है (30 और 600); textbook का "range = max - min" चाहता है diff(range(x)) या max(x) - min(x): 570। Option D एक और confusion smuggle करता है जो यह card guard करता है: IQR है Q3 - Q1 (middle 50% की width, outlier-resistant), full range से अलग और आमतौर पर बेहतर spread measure। दो नाम, दो measures, एक classic mix-up।

Think first

Exam answer cold लिखिए

Exam task: "vector m में stored marks के लिए (जिसमें NA हो सकता है), mean, median, mode, standard deviation और IQR के लिए R commands लिखिए, और एक precaution बताइए।" tap करने से पहले पाँचों lines और precaution draft कीजिए।

Show the answer

mean(m, na.rm = TRUE)

median(m, na.rm = TRUE)

names(which.max(table(m))) # R has no built-in statistical mode

sd(m, na.rm = TRUE) # sample SD, divisor n-1

IQR(m, na.rm = TRUE)

Precaution: हर call पर na.rm = TRUE (वरना एक NA NA return करता है), और mode workaround नोट कीजिए क्योंकि mode() storage type return करता है। वह answer, verbatim, पूरे marks है: यह card इसलिए मौजूद है ताकि आप इसे आधी नींद में produce कर सकें।

Watch out

Type toolkit decide करता है

ये commands numeric variables को serve करते हैं। mean() को एक factor दीजिए और R एक warning के साथ NA return करता है: सही तरीक़े से, क्योंकि Unit 1 का law अभी भी लागू है: categorical data को table()/prop.table() मिलता है (पिछला lesson), numerical data को यह card। जो variable-type सवाल आपने पहले सीखा वह उन हर command का router है जो आपने बाद में सीखा।

Theory

आपने अभी अपना पहला function लिखा

get_mode उस logic के चारों तरफ़ syntax के तीन tokens हैं जो आप पहले से रखते थे: function(v) एक recipe wrap करता है, और get_mode(x) call करना इसे replay करता है: R में असली programming का दरवाज़ा। दो lessons बचे हैं: normal distribution के R commands (pnorm और family), फिर graphical finale जहाँ आज के card का हर number एक picture बनता है।

Summary

Key takeaways

  • Centre: mean, median (na.rm = TRUE), mode names(which.max(table(x))) के ज़रिए।
  • Spread: पूरी width के लिए diff(range(x)), robust middle-50% width के लिए IQR(x), sd/var (sample divisor)।
  • Overview: summary(x), quantile(x), fivenum(x)।
  • range() min-max PAIR return करता है, एक number नहीं।
  • हर number को एक sentence में interpret कीजिए: skew (mean बनाम median), consistency (sd), typical (mode)।
  • Numeric variables यह card इस्तेमाल करते हैं; categoricals table() इस्तेमाल करते हैं: type toolkit route करता है।
  • Memory hook: examiner की चार-line shopping list।

Study this properly

This page is the lesson to read. In Gri-Learn the same topic is a graded deck: the self-checks are scored and your weak topics are tracked. Free to start.

Start this topic

Already have an account? Sign in

More from Working with Data in R

Gri-Learn · syllabus-mapped B.C.A. lessons in English, Hindi and Gujarati