Sampling distributions

कई samples लीजिए और हर एक थोड़ा अलग mean देता है: वे means sampling distribution बनाते हैं, जिसका spread 1/√n से सिकुड़ता है और जिसका shape Central Limit Theorem से normal बन जाता है।

10 min read · 10 cards · 2 checks

Read in: English · हिन्दी · ગુજરાતી


Theory

दो honest surveys, दो अलग जवाब

आपकी stratified CampusPulse survey average screen time 178 मिनट report करती है। एक rival team वही method 60 नए students पर चलाती है: 191 मिनट।

किसी ने cheat नहीं किया। किसी ने error नहीं किया। एक पॉट से दो honest spoonfuls बस अलग होते हैं।

तो कौन सा "असली" जवाब है? ग़लत सवाल। सही सवाल, जिसका यह lesson जवाब देता है: honest sample means truth के चारों तरफ़ कितना wobble करते हैं? Statistics का सबसे elegant idea यहीं रहता है।

Theory

हज़ार spoonfuls, plotted

कल्पना कीजिए एक हज़ार teams हैं, हर एक अपने random 60 students survey करती है और एक mean report करती है। सभी हज़ार means को एक histogram के रूप में plot कीजिए।

वह histogram sampling distribution of the mean है: students की distribution नहीं, बल्कि survey results की distribution। इसका centre असली value पर बैठता है; इसकी narrowness बताती है कोई भी अकेली honest survey कितनी भरोसे लायक़ है। आप कभी सिर्फ़ ONE survey चलाएँगे, पर हज़ार कैसे behave करती हैं यह जानना आपके एक को price देता है।

Theory

तीन distributions (उन्हें अलग रखिए)

Exams इन्हें blur करना पसंद करते हैं; मना कीजिए:

1. Population distribution: सभी 3000 students के screen times (शायद skewed, unknown)।

2. Sample data distribution: आपके 60 students की values (population जैसी दिखती है roughly)।

3. Sampling distribution of the mean: n size के सभी possible samples भर x̄ की values: RESULTS की distribution, students की नहीं।

तीसरी नई चीज़ है, और इसके दो सुंदर, exam-worthy properties हैं।

Theory

Property 1: सही centred। Property 2: predictable spread।

Centre: सभी possible sample means का mean population mean μ के बराबर है: sample mean एक unbiased estimator है: honest spoonfuls systematically over- या under-taste नहीं करते।

Spread: sample means का standard deviation, जिसे standard error (SE) कहते हैं:

SE = σ / √n

हल किया गया: σ = 60 min, n = 36 → SE = 60/√36 = 10 min। Sample means आमतौर पर truth के लगभग 10 मिनट के अंदर उतरते हैं। √ नोटिस कीजिए: n को quadruple करना SE को सिर्फ़ आधा करता है: precision एक square-root price पर ख़रीदी जाती है।

Theory

Central Limit Theorem

अब miracle। Screen time ख़ुद right-skewed है (कुछ binge-watchers)। ज़रूर sample means वह skew inherit करेंगे?

नहीं। Central Limit Theorem (CLT): sufficiently बड़े n के लिए (rule of thumb n ≥ 30), mean की sampling distribution approximately NORMAL है, population का shape चाहे जो हो।

कई values के averages individual quirks को smooth कर देते हैं। यही कारण है bell curve statistics पर राज करता है: data ugly होने पर भी, means normally behave करते हैं, तो अगले lesson का 68-95-99.7 ruler हर जगह survey results पर लागू होता है।

Quiz

Screen time: μ = 180, σ = 60। n = 36 वाले samples के साथ (SE = 10), एक team 210 का sample mean report करती है। आपको कैसे react करना चाहिए?

  1. 210, 180 से 3 standard errors ऊपर है: honest samples बहुत कम ही इतनी दूर उतरते हैं, तो sample investigate कीजिए
  2. बिल्कुल normal: 210, 180 के एक population SD (60) के अंदर है
  3. Impossible: sample means हमेशा population mean के बराबर होते हैं
  4. Team को confirm करने के लिए बस 36 students और survey करने चाहिए
Show the answer

210, 180 से 3 standard errors ऊपर है: honest samples बहुत कम ही इतनी दूर उतरते हैं, तो sample investigate कीजिए

Sample MEANS को standard error के ख़िलाफ़ judge किया जाता है, population SD के नहीं: (210 - 180)/10 = 3 SEs बाहर, एक event जो honest sampling 1% से भी कम समय produce करती है: एक biased frame या error का शक कीजिए। Option B ग़लत ruler लगाता है (σ individual students के लिए है; σ/√n means के लिए है): बिल्कुल वही confusion जिसे ठीक करने के लिए यह topic मौजूद है। Means wobble करते हैं (C ग़लत है), और ज़्यादा data एक flawed sample को retroactively validate नहीं कर सकता (D पिछले lesson को echo करता है)।

Think first

Square-root price पर ज़्यादा precision ख़रीदिए

n = 36 के साथ SE 10 मिनट था। Principal चाहते हैं SE = 5 मिनट। tap करने से पहले: कौन सा sample size चाहिए, और जवाब 72 क्यों NAHI है?

Show the answer

SE = σ/√n → 5 = 60/√n → √n = 12 → n = 144।

Error को आधा करने के लिए चार गुना students चाहिए, double नहीं, क्योंकि n एक square root के नीचे बैठता है। यह √n economics एक असली planning fact है: precision की हर extra digit exponentially महंगी हो जाती है, यही कारण है professional surveys "perfect" के बजाय "good enough" पर रुक जाते हैं।

Watch out

दो classic confusions

SD बनाम SE: σ students के spread को describe करता है; σ/√n sample means के spread को describe करता है। वही units, अलग objects: उन्हें mix करना इस chapter का सबसे penalised error है।

CLT क्या normalise करता है: MEAN की sampling distribution: raw data कभी नहीं। Screen time हमेशा के लिए skewed रहता है; सिर्फ़ survey averages bell-shaped बनते हैं। "बड़े samples data को normal बनाते हैं" लिखना तुरंत mark खो देता है।

Theory

Polls हज़ारों से करोड़ों predict करने की हिम्मत क्यों करते हैं

एक exit poll ~10,000 voters से पूछता है, और 90 करोड़ के बारे में बोलता है, एक stated "margin of error" के साथ: वह margin कुछ standard errors HI है। CLT guarantee करता है poll का mean normally behave करे; SE = σ/√n इसके wobble को price करता है। अगला lesson आपको bell curve का ruler देता है (68-95-99.7 और z-scores): वही ruler जो आप subject के बाक़ी हिस्से में sample means पर लगाएँगे।

Summary

Key takeaways

  • एक statistic sample से sample अलग होता है; mean की sampling distribution n size के सभी samples भर x̄ की distribution है।
  • यह μ पर centred है (sample mean unbiased है)।
  • इसका spread STANDARD ERROR है: SE = σ/√n: SE आधा करने के लिए n quadruple कीजिए।
  • CLT: n ≥ 30 के लिए, sample means population के shape के बावजूद approximately normal हैं।
  • Individual values को σ से judge कीजिए; sample means को SE से।
  • CLT means को normalise करता है, raw data को कभी नहीं।
  • Memory hook: हज़ार spoonfuls, plotted।

Study this properly

This page is the lesson to read. In Gri-Learn the same topic is a graded deck: the self-checks are scored and your weak topics are tracked. Free to start.

Start this topic

Already have an account? Sign in

More from Data Representation and Sampling technique

Gri-Learn · syllabus-mapped B.C.A. lessons in English, Hindi and Gujarati

Sampling distributions · Statistical Methods and Data Analysis (MDC-03) · Gri-Learn