Generating frequency tables and cross-tabulations

table(x) गिनता है हर category कितनी बार होती है, table(a, b) दो categoricals को एक grid में cross करता है, और prop.table() किसी को भी proportions में बदलता है: label data का summary toolkit।

9 min read · 9 cards · 2 checks

Read in: English · हिन्दी · ગુજરાતી


Theory

वे columns जिन्हें mean() नहीं छू सकता

CampusPulse survey का आधा हिस्सा labels है: city, stay, वे usage bands जो आपने recode किए। mean(survey$city) nonsense है (Unit 1 का सबसे पुराना rule)।

फिर भी labels के बारे में dean के सवाल पूरी तरह numeric हैं: हर city से कितने? क्या hostellers heavy usage की तरफ़ skew करते हैं?

Labels गिनने का toolkit एक function है दो gears के साथ: table() अकेली एक variable गिनता है; इसे दो दीजिए और यह cross-tabulation बनाता है: वह grid जहाँ से हर social statistic शुरू होता है।

Theory

Tally marks, फिर एक tally grid

One-way table: classic tally sheet: forms पढ़े जाने पर हर city के नीचे एक stroke: Surat |||| ||, Navsari |||।

Cross-tab: tally cells की एक grid: rows stay हैं (hostel/day), columns usage हैं (heavy/normal), और हर form बिल्कुल ONE cell में एक stroke जोड़ता है। तैयार grid relationship सवालों का जवाब देता है जो कोई भी अकेली tally sheet नहीं दे सकती: इसलिए इसका अपना नाम है।

Practical

Counts, grids, proportions

survey <- read.csv("survey_clean.csv")  # city, stay, usage (recoded)

# ONE-WAY: frequency table of a categorical
table(survey$city)
#  bardoli navsari   surat
#        8      16      36

# sorted, most common first
sort(table(survey$city), decreasing = TRUE)

# TWO-WAY: the cross-tabulation (rows = 1st arg, cols = 2nd)
t <- table(survey$stay, survey$usage)
t
#          heavy normal
#   day       6     30
#   hostel   14     10

# proportions: three DIFFERENT questions
prop.table(t)              # share of ALL students per cell
prop.table(t, margin = 1)  # within each ROW (each row sums to 1)
prop.table(t, margin = 2)  # within each COLUMN

# formula spelling + totals
xtabs(~ stay + usage, data = survey)
addmargins(t)              # adds row/column sums

Theory

margin: वह argument जो सवाल बदल देता है

वही grid, तीन proportion सवाल:

  • prop.table(t): सभी surveyed students का कितना हिस्सा हर cell है? (cells overall 1 तक sum होते हैं)
  • margin = 1 (rows): hostellers के अंदर, कितना हिस्सा heavy है? हर row 1 तक sum होती है: stay groups भर usage compare करती है।
  • margin = 2 (columns): heavy users के अंदर, कितना हिस्सा hostellers है? हर column 1 तक sum होता है।

Row 'hostel': 14/(14+10) = margin = 1 के साथ 58% heavy। Dean का "क्या hostellers heavy की तरफ़ skew करते हैं?" एक margin = 1 सवाल है: margin चुनना ही analysis है।

Quiz

ऊपर वाली grid से (hostel: 14 heavy, 10 normal; day: 6 heavy, 30 normal), dean पूछते हैं: "HOSTELLERS का कितना percentage heavy users हैं?" कौन सा command और कौन सा number?

  1. prop.table(t, margin = 1): hostel row देता है 14/24 ≈ 58%
  2. prop.table(t, margin = 2): heavy column देता है 14/20 = 70%
  3. prop.table(t): 14/60 ≈ 23%
  4. table(t): raw 14 पहले से एक percentage है
Show the answer

prop.table(t, margin = 1): hostel row देता है 14/24 ≈ 58%

"Hostellers का" row को पूरे के रूप में fix करता है: row-wise proportions, margin = 1: 24 hostellers में से 14 ≈ 58%। Option B REVERSE सवाल का जवाब देता है (heavy users में से, कितने hostellers हैं: 70%): एक अलग, भी-valid सवाल जिसे careless reading swap कर देती है। Option C सबके सामने share करता है। margin argument बिल्कुल वहाँ है जहाँ ये तीन सवाल अलग होते हैं: exams और असली reports दोनों इस पर टिकते हैं।

Think first

Grid को एक analyst की तरह पढ़िए

margin = 1 table दिखाती है: hostel → heavy 0.58, normal 0.42; day → heavy 0.17, normal 0.83। tap करने से पहले: honest one-sentence finding क्या है, और कौन सा CAUTION साथ आना चाहिए (Unit 2 ने इसे सिखाया)?

Show the answer

Finding: hostellers day scholars से heavy users होने की कहीं ज़्यादा संभावना रखते हैं (58% बनाम 17%): stay और usage के बीच एक strong association।

Caution: association causation नहीं है: hostels ज़रूरी नहीं heavy usage CAUSE करते हों (free evenings, wifi access, या कौन hostels चुनता है दोनों drive कर सकता है)। Scatter-plot lesson का rule grids पर भी लागू होता है: cross-tabs दिखाते हैं कि दो labels साथ चलते हैं, कभी क्यों नहीं।

Watch out

Tabulation की slips

Counts जब proportions पूछे गए (और उल्टा): सवाल का "कितने" बनाम "कितना हिस्सा" पढ़िए।

margin mix-ups: margin = 1 rows, margin = 2 columns: within-row सवाल common exam phrasing है।

NA चुपचाप ग़ायब होता है: table() default रूप से missing values drop करता है: table(x, useNA = "ifany") उन्हें दिखाता है: हमेशा की तरह holes report कीजिए।

Theory

Cross-tabs आगे कहाँ जाते हैं

हर opinion poll ("age group से support"), हर medical study ("treatment से outcome"), हर churn report ("plan से cancelled") एक cross-tab है business clothes पहने: और chi-square test जो आप बाद के semesters में मिल सकते हैं वह बस formal तरीक़ा है यह पूछने का कि क्या grid का association असली है या luck। pandas इन्हें value_counts() और pd.crosstab() spell करता है: चौथा tool, वही tally grid।

Summary

Key takeaways

  • table(x): एक categorical के लिए frequency counts: label data का mean()।
  • table(a, b): cross-tabulation grid: rows = पहला argument, columns = दूसरा।
  • prop.table(t): overall shares; margin = 1 row-wise, margin = 2 column-wise: margin ही सवाल है।
  • xtabs(~ a + b, df) formula spelling है; addmargins() totals जोड़ता है।
  • table() चुपचाप NA drop करता है: उन्हें दिखाने के लिए useNA = 'ifany'।
  • Grid associations causation नहीं हैं।
  • Memory hook: tally sheet, फिर tally grid।

Study this properly

This page is the lesson to read. In Gri-Learn the same topic is a graded deck: the self-checks are scored and your weak topics are tracked. Free to start.

Start this topic

Already have an account? Sign in

More from Working with Data in R

Gri-Learn · syllabus-mapped B.C.A. lessons in English, Hindi and Gujarati