Theory
वही average, अलग जानवर
दो batches वही test देते हैं। दोनों का average 60 है।
Batch P: 58, 59, 60, 61, 62। Batch Q: 20, 40, 60, 80, 100।
एक tight, teachable group है; दूसरा एक ही room में failure से topper तक फैला है। Mean, पिछले lesson का हमारा star, इन्हें अलग नहीं बता सकता।
हर centre को दूसरा number चाहिए: values इससे कितनी दूर भटकती हैं? वह number dispersion है, और exam इसकी तीन versions हाथ से चाहता है।
Theory
एक bullseye के चारों तरफ़ archers
दो archers दोनों dead-centre average करते हैं। एक हर arrow को एक coin की चौड़ाई के अंदर group करता है; दूसरा पूरे target face पर छिड़कता है, misses misses को cancel करते हुए।
वही average position, बिल्कुल अलग archers। Dispersion arrows के scatter को measure करता है, यह नहीं कि उनका centre कहाँ है। इसके बिना एक summary sprayer और sniper दोनों की समान तारीफ़ करता है।
Theory
Range: साठ-सेकंड का जवाब
Range = maximum - minimum।
Batch P: 62 - 58 = 4। Batch Q: 100 - 20 = 80। Instant, और पहले से दोनों batches को अलग बताता है।
इसकी कमज़ोरी structural है: यह सिर्फ़ दो values से पूछता है और बाक़ी n-2 को नज़रअंदाज़ करता है। एक विचित्र outlier (एक अकेला 600-minute screen-timer) range को explode कर देता है जबकि crowd unchanged बैठी रहती है। इसे पहली नज़र के रूप में इस्तेमाल कीजिए, कभी verdict के रूप में नहीं।
Theory
Variance: हम square क्यों करते हैं
बेहतर idea: हर value की mean से deviation measure कीजिए और उन्हें average कीजिए। एक catch: raw deviations हमेशा zero तक sum होते हैं (mean उनका balance point है): +2 और -2 cancel होते हैं, fake calm report करते हुए।
Fix: पहले हर deviation को square कीजिए: negatives ग़ायब हो जाते हैं, बड़ी misses अतिरिक्त गिनती हैं।
Population variance σ² = Σ(x - mean)² ÷ N।
Sample variance s² इसके बजाय (n - 1) से divide करता है (Bessel's correction: samples spread थोड़ा understate करते हैं; n-1 compensate करता है: इसे नाम दीजिए, इसे इस्तेमाल कीजिए जब भी data एक sample हो)।
Theory
शुरू से अंत तक, हल किया गया
Marks: 4, 8, 6, 2 (सादगी के लिए, एक population)।
1. Mean = 20 ÷ 4 = 5।
2. Deviations: -1, +3, +1, -3। (Sum = 0 ✓ built-in error check।)
3. Squares: 1, 9, 1, 9। Sum = 20।
4. Variance σ² = 20 ÷ 4 = 5 (units: marks²)।
5. SD σ = √5 ≈ 2.24 marks: असली units में वापस, वह reportable number।
एक sample के रूप में: s² = 20 ÷ 3 ≈ 6.67, s ≈ 2.58। वही recipe, अलग divisor।
Think first
आपकी बारी, पूरी recipe
5 students के screen-time घंटे: 2, 4, 4, 5, 10। कागज़ पर: mean, deviations (check कीजिए वे 0 तक sum हों), squared deviations, POPULATION variance और SD। फिर tap कीजिए।
Show the answer
Mean = 25 ÷ 5 = 5।
Deviations: -3, -1, -1, 0, +5 (sum 0 ✓)।
Squares: 9, 1, 1, 0, 25। Sum = 36।
Variance = 36 ÷ 5 = 7.2 घंटे²। SD = √7.2 ≈ 2.68 घंटे।
नोटिस कीजिए अकेले 10 ने 36 में से 25 contribute किए: squaring outliers को चीख़ने देती है। अगर आपकी deviations zero तक sum नहीं हुईं, mean ग़लत था: कुछ भी square करने से पहले recompute कीजिए।
Quiz
एक student report करता है "marks का variance 5 है, तो scores आमतौर पर mean से 5 marks अलग होते हैं।" क्या ग़लत है?
- Variance SQUARED units (marks²) में है; typical distance SD है, √5 ≈ 2.24 marks
- कुछ नहीं: variance और SD एक ही number हैं
- Marks के लिए variance compute नहीं किया जा सकता
- Typical distance range है, SD नहीं
Show the answer
Variance SQUARED units (marks²) में है; typical distance SD है, √5 ≈ 2.24 marks
Variance squared units में रहता है, algebra के लिए useful पर एक distance के रूप में unreadable: कोई "5 square marks" नहीं भटकता। Standard deviation इसे वापस असली units में square-root करता है, √5 ≈ 2.24 को honest "mean से typical distance" बनाते हुए। दोनों को confuse करना scripts में सबसे common dispersion error है; range (option D) total width measure करता है, typical distance नहीं।
Watch out
तीन dispersion traps
Squares skip करना: raw deviations zero तक sum होते हैं: हर dataset zero spread claim करेगा।
Divisor confusion: population ÷ N, sample ÷ (n-1)। बताइए आपने कौन सा इस्तेमाल किया: R और pandas default रूप से SAMPLE formula इस्तेमाल करते हैं, plain calculators अक्सर population।
Variance को एक distance के रूप में report करना: interpret करने से पहले SD में convert कीजिए। और याद रखिए range सिर्फ़ दो values से पूछता है: इसे कभी robust मत कहिए।
Theory
SD चुपचाप कहाँ चीज़ें चलाता है
Result moderation mean और SD का इस्तेमाल करके marks scale करता है। Quality control एक production line रोकता है जब SD ऊपर चढ़ता है। Cricket commentators का "consistent batsman" एक low-SD claim है। और दो lessons आगे, bell curve SD को एक ruler में बदलता है: 68% data 1 SD के अंदर, 95% 2 के अंदर: आज हाथ से compute किया number वह unit बन जाता है जिसमें पूरा normal distribution measure होता है। (Units भर spreads compare कर रहे हैं? SD को mean से divide कीजिए: coefficient of variation।)
Summary
Key takeaways
- Dispersion "कितना scattered?" का जवाब देता है: अकेला mean tight को wild data से अलग नहीं बता सकता।
- Range = max - min: instant, पर सिर्फ़ दो values और outlier-fragile।
- Mean से deviations zero तक sum होते हैं: squaring cancellation ख़त्म करता है और बड़ी misses को ज़्यादा weight देता है।
- Variance = average squared deviation (÷N population, ÷(n-1) sample: Bessel)।
- SD = √variance: असली units में वापस, वह number जो आप report और interpret करते हैं।
- Deviation sum = 0 squaring से पहले free arithmetic check है।
- Memory hook: एक bullseye के चारों तरफ़ archers: वही centre, अलग scatter।