Descriptive statistics: measures of central tendency (mean, median, mode)

Mean, median और mode "इस data का केंद्र कहाँ है?" के तीन प्रतिस्पर्धी जवाब हैं, और exam skill है हाथ से हर एक compute करना AND दिए गए dataset के लिए किस पर भरोसा करना है यह judge करना।

10 min read · 9 cards · 2 checks

Read in: English · हिन्दी · ગુજરાતી


Theory

तीन honest जवाब, एक सवाल

CampusPulse survey का सात hostellers का screen-time column, मिनटों में:

60, 90, 90, 120, 150, 180, 600

(600 असली है: एक student ने पूरा cricket match देखा।)

"typical screen time क्या है?" Mean कहता है 184.3। Median कहता है 120। Mode कहता है 90। एक सवाल के तीन defensible जवाब, और इस topic के marks हर एक को सही compute करने AND कौन सा report करना है यह जानने से आते हैं।

Theory

तीन committee members

एक committee से पूछिए class का "centre" कहाँ है:

  • Accountant सब कुछ जोड़ता और divide करता है: precise, पर एक extreme member report को skew कर देता है।
  • Middle-finder सबको sort करता है और central person पर point करता है: extremes उन्हें हिला नहीं सकते।
  • Popularity-counter जो भी सबसे ज़्यादा होता है वह नाम देता है: वह अकेला member जो labels भी handle कर सकता है (सबसे common city)।

आप इस committee से BCA303 में pandas के arithmetic करते समय मिले। आज exam इसे हाथ से चाहता है, reasons के साथ।

Theory

हर एक precisely compute करना

Data: 60, 90, 90, 120, 150, 180, 600 (n = 7, पहले से sorted)।

Mean = sum ÷ n = 1290 ÷ 7 ≈ 184.3 min।

Median = position (n+1)/2 पर value = 4th value = 120। (हमेशा पहले SORT कीजिए। Even n? दोनों middle values का average लीजिए।)

Mode = सबसे frequent = 90 (दो बार आता है)।

जवाबों का फैलाव नोटिस कीजिए: 600 ने mean को सात में से छह students से ऊपर खींच दिया। Median ने इसे कभी महसूस नहीं किया।

At a glance

Merits और demerits (5-mark table)

MeasureStrengthsWeaknesses
Meanहर value इस्तेमाल करता है; algebra-friendlyOutliers से बिगड़ता है; सिर्फ़ numeric
MedianOutlier-proof; ordinal data के लिए काम करता हैExact magnitudes नज़रअंदाज़ करता है
ModeNominal data के लिए एकमात्र choice; एक असली observed valueNone हो सकती, या कई; unrepresentative हो सकती है

Think first

Even n के साथ एक हल कीजिए

6 students के commute minutes: 25, 10, 40, 20, 30, 15। कागज़ पर: sort कीजिए, फिर median ढूँढिए (सावधान: n even है), mean, और mode। फिर tap कीजिए।

Show the answer

Sorted: 10, 15, 20, 25, 30, 40।

Median: n = 6 even है, तो 3rd और 4th values का average लीजिए: (20 + 25) / 2 = 22.5।

Mean: 140 ÷ 6 ≈ 23.3।

Mode: हर value एक बार आती है: कोई mode नहीं।

दो traps बच गए: even-n averaging step, और honest जवाब "कोई mode नहीं" (data आपको एक देने का क़र्ज़दार नहीं है)। यहाँ Mean ≈ median: कोई outlier नहीं, symmetric-ish data, दोनों centres सहमत हैं।

Quiz

Hostel screen time पर एक report को एक typical value बताना है। 600-minute outlier को देखते हुए, कौन सा measure और कौन सी justification पूरे marks कमाती है?

  1. Median 120, क्योंकि median उन outliers का सामना करता है जो skewed data में mean को बिगाड़ते हैं
  2. Mean 184.3, क्योंकि mean हर observation इस्तेमाल करता है और इसलिए हमेशा best है
  3. Mode 90, क्योंकि mode हमेशा सबसे safe choice है
  4. इनमें से कोई भी, क्योंकि तीनों central tendency के measures हैं
Show the answer

Median 120, क्योंकि median उन outliers का सामना करता है जो skewed data में mean को बिगाड़ते हैं

एक outlier वाला skewed data बिल्कुल median का home ground है: यह बताता है bulk students असल में कहाँ बैठते हैं, जबकि mean सात में से छह observations से ऊपर एक value तक खिंच गया। Option B की justification असली है (mean सच में हर value इस्तेमाल करता है) पर यही property यहाँ इसकी कमज़ोरी है। "Always" और "any" जवाब (C, D) उस judgement को छोड़ देते हैं जो सवाल test कर रहा है।

Watch out

Marks कहाँ लीक होते हैं

Unsorted median: raw data के बीच में एक random number है। Sort कीजिए, फिर position (n+1)/2 तक गिनिए।

Even n: दो middle values का average लेना भूलना।

Skew signature: अगर mean > median noticeably है, data right-skewed है (बड़ी values की एक लंबी tail): अपने numbers के बगल में यह one-line diagnosis बताना एक जवाब को computed से understood तक upgrade करता है।

Theory

हर measure असली दुनिया में कहाँ चलता है

Median: घर की क़ीमतें और salaries (एक करोड़पति को "typical" define नहीं करना चाहिए)। Mean: exam marks, cricket averages, कुछ भी roughly symmetric। Mode: एक shop stock करती shoe sizes, blood bank में सबसे common blood group। BCA303 में आपने ये pandas में compute किए; इस subject की Unit 5 इसे R में करती है (एक surprise के साथ: R में कोई built-in mode function नहीं है)। तीनों में concepts identical रहते हैं।

Summary

Key takeaways

  • Mean = sum/n: सब कुछ इस्तेमाल करता है, outliers की तरफ़ झुकता है।
  • Median = SORTED data का middle, position (n+1)/2; even n दो middles का average लेता है; outlier-proof।
  • Mode = सबसे frequent: none, एक, या कई हो सकती है; nominal data के लिए एकमात्र measure।
  • Skewed data या outliers → median report कीजिए; symmetric numeric data → mean; labels → mode।
  • Mean noticeably median से ऊपर = right skew (बड़ी values की लंबी tail)।
  • Memory hook: accountant, middle-finder, popularity-counter: सही committee member चुनिए।

Study this properly

This page is the lesson to read. In Gri-Learn the same topic is a graded deck: the self-checks are scored and your weak topics are tracked. Free to start.

Start this topic

Already have an account? Sign in

More from Basic concepts of statistic

Gri-Learn · syllabus-mapped B.C.A. lessons in English, Hindi and Gujarati

Descriptive statistics: measures of central tendency (mean, median, mode) · Statistical Methods and Data Analysis (MDC-03) · Gri-Learn