Theory
The columns mean() cannot touch
Half the CampusPulse survey is labels: city, stay, the usage bands you recoded. mean(survey$city) is nonsense (Unit 1's oldest rule).
Yet the dean's questions about labels are perfectly numeric: how many from each city? Do hostellers skew toward heavy usage?
The toolkit for counting labels is one function with two gears: table() alone counts one variable; give it two and it builds the cross-tabulation: the grid every social statistic starts from.
Theory
Tally marks, then a tally grid
One-way table: the classic tally sheet: a stroke under each city as forms are read out: Surat |||| ||, Navsari |||.
Cross-tab: a grid of tally cells: rows are stay (hostel/day), columns are usage (heavy/normal), and each form adds one stroke to exactly ONE cell. The finished grid answers relationship questions no single tally sheet can: that is why it gets its own name.
Practical
Counts, grids, proportions
survey <- read.csv("survey_clean.csv") # city, stay, usage (recoded)
# ONE-WAY: frequency table of a categorical
table(survey$city)
# bardoli navsari surat
# 8 16 36
# sorted, most common first
sort(table(survey$city), decreasing = TRUE)
# TWO-WAY: the cross-tabulation (rows = 1st arg, cols = 2nd)
t <- table(survey$stay, survey$usage)
t
# heavy normal
# day 6 30
# hostel 14 10
# proportions: three DIFFERENT questions
prop.table(t) # share of ALL students per cell
prop.table(t, margin = 1) # within each ROW (each row sums to 1)
prop.table(t, margin = 2) # within each COLUMN
# formula spelling + totals
xtabs(~ stay + usage, data = survey)
addmargins(t) # adds row/column sums
Theory
margin: the argument that changes the question
The same grid, three proportion questions:
- prop.table(t): what share of ALL surveyed students is each cell? (cells sum to 1 overall)
- margin = 1 (rows): within hostellers, what share is heavy? Each row sums to 1: compares usage across stay groups.
- margin = 2 (columns): within heavy users, what share is hostellers? Each column sums to 1.
Row 'hostel': 14/(14+10) = 58% heavy with margin = 1. The dean's "do hostellers skew heavy?" is a margin = 1 question: choosing the margin IS the analysis.
Quiz
From the grid above (hostel: 14 heavy, 10 normal; day: 6 heavy, 30 normal), the dean asks: "what percentage of HOSTELLERS are heavy users?" Which command and which number?
- prop.table(t, margin = 1): the hostel row gives 14/24 ≈ 58%
- prop.table(t, margin = 2): the heavy column gives 14/20 = 70%
- prop.table(t): 14/60 ≈ 23%
- table(t): the raw 14 is already a percentage
Show the answer
prop.table(t, margin = 1): the hostel row gives 14/24 ≈ 58%
"Of hostellers" fixes the row as the whole: row-wise proportions, margin = 1: 14 of the 24 hostellers ≈ 58%. Option B answers the REVERSE question (of heavy users, how many are hostellers: 70%): a different, also-valid question that careless reading swaps. Option C shares against everyone. The margin argument is exactly where these three questions part ways: exams and real reports both hinge on it.
Think first
Read the grid like an analyst
The margin = 1 table shows: hostel → heavy 0.58, normal 0.42; day → heavy 0.17, normal 0.83. Before tapping: what is the honest one-sentence finding, and what CAUTION must accompany it (Unit 2 taught it)?
Show the answer
Finding: hostellers are far likelier to be heavy users than day scholars (58% vs 17%): a strong association between stay and usage.
Caution: association is not causation: hostels do not necessarily CAUSE heavy usage (free evenings, wifi access, or who chooses hostels could drive both). The scatter-plot lesson's rule applies to grids too: cross-tabs reveal that two labels move together, never why.
Watch out
The tabulation slips
Counts when proportions were asked (and vice versa): read the question's "how many" vs "what share".
margin mix-ups: margin = 1 rows, margin = 2 columns: within-row questions are the common exam phrasing.
NA vanishes silently: table() drops missing values by default: table(x, useNA = "ifany") shows them: report holes, as always.
Theory
Where cross-tabs go next
Every opinion poll ("support by age group"), every medical study ("outcome by treatment"), every churn report ("cancelled by plan") is a cross-tab wearing business clothes: and the chi-square test you may meet later semesters is simply the formal way to ask whether a grid's association is real or luck. pandas spells these value_counts() and pd.crosstab(): fourth tool, same tally grid.
Summary
Key takeaways
- table(x): frequency counts for one categorical: the mean() of label data.
- table(a, b): the cross-tabulation grid: rows = first argument, columns = second.
- prop.table(t): overall shares; margin = 1 row-wise, margin = 2 column-wise: the margin IS the question.
- xtabs(~ a + b, df) is the formula spelling; addmargins() adds totals.
- table() drops NA silently: useNA = 'ifany' to show them.
- Grid associations are not causation.
- Memory hook: tally sheet, then tally grid.