Theory
वे columns जिन्हें mean() नहीं छू सकता
CampusPulse survey का आधा हिस्सा labels है: city, stay, वे usage bands जो आपने recode किए। mean(survey$city) nonsense है (Unit 1 का सबसे पुराना rule)।
फिर भी labels के बारे में dean के सवाल पूरी तरह numeric हैं: हर city से कितने? क्या hostellers heavy usage की तरफ़ skew करते हैं?
Labels गिनने का toolkit एक function है दो gears के साथ: table() अकेली एक variable गिनता है; इसे दो दीजिए और यह cross-tabulation बनाता है: वह grid जहाँ से हर social statistic शुरू होता है।
Theory
Tally marks, फिर एक tally grid
One-way table: classic tally sheet: forms पढ़े जाने पर हर city के नीचे एक stroke: Surat |||| ||, Navsari |||।
Cross-tab: tally cells की एक grid: rows stay हैं (hostel/day), columns usage हैं (heavy/normal), और हर form बिल्कुल ONE cell में एक stroke जोड़ता है। तैयार grid relationship सवालों का जवाब देता है जो कोई भी अकेली tally sheet नहीं दे सकती: इसलिए इसका अपना नाम है।
Practical
Counts, grids, proportions
survey <- read.csv("survey_clean.csv") # city, stay, usage (recoded)
# ONE-WAY: frequency table of a categorical
table(survey$city)
# bardoli navsari surat
# 8 16 36
# sorted, most common first
sort(table(survey$city), decreasing = TRUE)
# TWO-WAY: the cross-tabulation (rows = 1st arg, cols = 2nd)
t <- table(survey$stay, survey$usage)
t
# heavy normal
# day 6 30
# hostel 14 10
# proportions: three DIFFERENT questions
prop.table(t) # share of ALL students per cell
prop.table(t, margin = 1) # within each ROW (each row sums to 1)
prop.table(t, margin = 2) # within each COLUMN
# formula spelling + totals
xtabs(~ stay + usage, data = survey)
addmargins(t) # adds row/column sums
Theory
margin: वह argument जो सवाल बदल देता है
वही grid, तीन proportion सवाल:
- prop.table(t): सभी surveyed students का कितना हिस्सा हर cell है? (cells overall 1 तक sum होते हैं)
- margin = 1 (rows): hostellers के अंदर, कितना हिस्सा heavy है? हर row 1 तक sum होती है: stay groups भर usage compare करती है।
- margin = 2 (columns): heavy users के अंदर, कितना हिस्सा hostellers है? हर column 1 तक sum होता है।
Row 'hostel': 14/(14+10) = margin = 1 के साथ 58% heavy। Dean का "क्या hostellers heavy की तरफ़ skew करते हैं?" एक margin = 1 सवाल है: margin चुनना ही analysis है।
Quiz
ऊपर वाली grid से (hostel: 14 heavy, 10 normal; day: 6 heavy, 30 normal), dean पूछते हैं: "HOSTELLERS का कितना percentage heavy users हैं?" कौन सा command और कौन सा number?
- prop.table(t, margin = 1): hostel row देता है 14/24 ≈ 58%
- prop.table(t, margin = 2): heavy column देता है 14/20 = 70%
- prop.table(t): 14/60 ≈ 23%
- table(t): raw 14 पहले से एक percentage है
Show the answer
prop.table(t, margin = 1): hostel row देता है 14/24 ≈ 58%
"Hostellers का" row को पूरे के रूप में fix करता है: row-wise proportions, margin = 1: 24 hostellers में से 14 ≈ 58%। Option B REVERSE सवाल का जवाब देता है (heavy users में से, कितने hostellers हैं: 70%): एक अलग, भी-valid सवाल जिसे careless reading swap कर देती है। Option C सबके सामने share करता है। margin argument बिल्कुल वहाँ है जहाँ ये तीन सवाल अलग होते हैं: exams और असली reports दोनों इस पर टिकते हैं।
Think first
Grid को एक analyst की तरह पढ़िए
margin = 1 table दिखाती है: hostel → heavy 0.58, normal 0.42; day → heavy 0.17, normal 0.83। tap करने से पहले: honest one-sentence finding क्या है, और कौन सा CAUTION साथ आना चाहिए (Unit 2 ने इसे सिखाया)?
Show the answer
Finding: hostellers day scholars से heavy users होने की कहीं ज़्यादा संभावना रखते हैं (58% बनाम 17%): stay और usage के बीच एक strong association।
Caution: association causation नहीं है: hostels ज़रूरी नहीं heavy usage CAUSE करते हों (free evenings, wifi access, या कौन hostels चुनता है दोनों drive कर सकता है)। Scatter-plot lesson का rule grids पर भी लागू होता है: cross-tabs दिखाते हैं कि दो labels साथ चलते हैं, कभी क्यों नहीं।
Watch out
Tabulation की slips
Counts जब proportions पूछे गए (और उल्टा): सवाल का "कितने" बनाम "कितना हिस्सा" पढ़िए।
margin mix-ups: margin = 1 rows, margin = 2 columns: within-row सवाल common exam phrasing है।
NA चुपचाप ग़ायब होता है: table() default रूप से missing values drop करता है: table(x, useNA = "ifany") उन्हें दिखाता है: हमेशा की तरह holes report कीजिए।
Theory
Cross-tabs आगे कहाँ जाते हैं
हर opinion poll ("age group से support"), हर medical study ("treatment से outcome"), हर churn report ("plan से cancelled") एक cross-tab है business clothes पहने: और chi-square test जो आप बाद के semesters में मिल सकते हैं वह बस formal तरीक़ा है यह पूछने का कि क्या grid का association असली है या luck। pandas इन्हें value_counts() और pd.crosstab() spell करता है: चौथा tool, वही tally grid।
Summary
Key takeaways
- table(x): एक categorical के लिए frequency counts: label data का mean()।
- table(a, b): cross-tabulation grid: rows = पहला argument, columns = दूसरा।
- prop.table(t): overall shares; margin = 1 row-wise, margin = 2 column-wise: margin ही सवाल है।
- xtabs(~ a + b, df) formula spelling है; addmargins() totals जोड़ता है।
- table() चुपचाप NA drop करता है: उन्हें दिखाने के लिए useNA = 'ifany'।
- Grid associations causation नहीं हैं।
- Memory hook: tally sheet, फिर tally grid।