Theory
पूरी college से कोई नहीं पूछता
Principal college के 3000 students की average daily screen time जानना चाहते हैं। Option one: सभी 3000 का interview कीजिए। दो हफ़्ते, चार staff, exam season बर्बाद।
Option two, जो statisticians असल में करते हैं: सावधानी से 60 students चुनिए, उन्हें एक दोपहर में survey कीजिए, और उनके जवाबों से सभी 3000 के बारे में बताइए।
यह move, पूरे को समझने के लिए एक हिस्सा measure करना, statistics का founding act है, और इस subject का पहला काम इसके pieces को precisely नाम देना है।
Theory
एक चम्मच पूरी हांडी judge करता है
एक cook salt check करने के लिए कभी पूरी हांडी दाल नहीं पीता। एक हिलाया हुआ चम्मच भर पूरी हांडी के लिए decide करता है।
हांडी population है: हर वह चीज़ जो आप जानना चाहते हैं। चम्मच भर sample है: वह हिस्सा जो आप असल में चखते हैं। और हिलाना honesty condition है: ऊपर से बिना हिलाया चम्मच सिर्फ़ tadka चखता है, और हांडी के बारे में झूठ बोलता है।
Theory
चार exam शब्द
- Population: study के तहत ENTIRE group (सभी 3000 students)।
- Sample: असल में measure किया subset (survey हुए 60)।
- Parameter: population describe करता नंबर (सभी 3000 का TRUE average screen time: आमतौर पर unknown)।
- Statistic: sample से compute किया वैसा ही नंबर (60 students का average: जो हम असल में calculate कर सकते हैं)।
एक discipline के रूप में Statistics parameters (population truths) estimate करने के लिए statistics (sample numbers) इस्तेमाल करने की कला है। सबकी पूरी measurement एक census है: costly, slow, कभी-कभी impossible (हर bulb test करने से stock ख़त्म हो जाता है)।
Theory
Variables: survey क्या record करता है
हर survey हुआ student चार सवालों के जवाब देता है। हर जवाब एक variable है, और हर variable दो families में से एक है:
Categorical (qualitative): labels।
- Nominal: कोई natural order नहीं: city, blood group।
- Ordinal: ordered labels: grade (A > B > C), satisfaction (high/medium/low)।
Numerical (quantitative): numbers जिन पर आप maths कर सकते हैं।
- Discrete: countable jumps: fail हुए subjects (0, 1, 2...)।
- Continuous: किसी भी precision पर: height (167.5 cm), screen time (143.7 min)।
At a glance
CampusPulse survey, classified
| Variable | Family | Sub-type |
|---|---|---|
| City | Categorical | Nominal (कोई order नहीं) |
| Grade (A/B/C) | Categorical | Ordinal (ordered) |
| Subjects failed | Numerical | Discrete (countable) |
| Height (cm) | Numerical | Continuous (measurable) |
| Screen time (min) | Numerical | Continuous |
Quiz
Survey हुए 60 students औसतन 4.2 घंटे screen time report करते हैं। सटीक exam भाषा में, वह 4.2 क्या HAI?
- एक statistic: sample से compute किया, population parameter estimate करने के लिए इस्तेमाल किया
- एक parameter: कोई भी compute किया average एक parameter है
- Population value: 4.2 को exact बनाने के लिए 60 students काफ़ी हैं
- एक census result
Show the answer
एक statistic: sample से compute किया, population parameter estimate करने के लिए इस्तेमाल किया
SAMPLE से numbers statistics हैं; सभी 3000 भर unknown true value वह parameter है जिसे statistic estimate करता है। यह pairing (sample→statistic, population→parameter, दोनों एक ही letters से शुरू) वह memory device है जो examiners चाहते हैं। एक sample average कभी exact truth नहीं बनता (option C), और एक census का मतलब होता सभी 3000 measure हुए (option D)।
Think first
इन चार को classify कीजिए
एक hostel survey से: (1) mess rating: poor/okay/good; (2) dinner में खाए chapatis की संख्या; (3) mother tongue; (4) commute time मिनटों में। tap करने से पहले हर एक को इसकी family AND sub-type दीजिए।
Show the answer
(1) Categorical ordinal: labels में एक order है।
(2) Numerical discrete: आप chapatis गिनते हैं, कोई 2.7 नहीं खाता।
(3) Categorical nominal: languages में कोई order नहीं।
(4) Numerical continuous: 23.5 मिनट meaningful है।
दो-सवाल test: क्या यह एक label है या एक number? फिर: ordered है या नहीं / counted है या measured? इस subject का हर classification सवाल order में पूछे गए उन दो सवालों पर गिरता है।
Watch out
Nonsense-maths trap
Cities को 1 = Surat, 2 = Navsari, 3 = Bardoli code कीजिए और एक spreadsheet ख़ुशी से report करेगा "average city = 1.8"। Meaningless: number costumes पहने nominal labels अभी भी labels हैं। यही roll numbers और pin codes के लिए भी सच है।
Variable का TYPE इसकी legal maths decide करता है: means और SDs numerical variables की हैं; categorical variables को counts और percentages मिलते हैं। Exams इसे "कौन सा measure appropriate है?" सवालों से test करते हैं।
Theory
यह vocabulary पूरे subject को क्यों चलाती है
बाद का हर topic आज के शब्दों पर खड़ा है: central tendency और dispersion numerical variables को summarise करते हैं; frequency tables categorical वालों को गिनती हैं; sampling techniques उन 60 को ईमानदारी से चुनने के बारे में हैं; bell curve बताता है sample statistics कैसे behave करते हैं। और Unit 3 में, R आपसे ये types DECLARE करवाएगा (numeric, factor): computer जो आपने classify नहीं किया उसे guess करने से मना करता है।
Summary
Key takeaways
- Population = वे सब जिनकी परवाह है; sample = असल में measure किया हिस्सा।
- Parameter population describe करता है (unknown); statistic sample से compute होता है (इसे estimate करता है)।
- Census = सबको measure करना: costly, slow, कभी-कभी impossible।
- Categorical variables labels हैं: nominal (unordered) या ordinal (ordered)।
- Numerical variables numbers हैं: discrete (counted) या continuous (measured)।
- Type legal maths decide करता है: city codes या roll numbers को average मत कीजिए।
- Memory hook: एक हिलाया चम्मच पूरी हांडी judge करता है।