Generating frequency tables and cross-tabulations

table(x) counts how often each category occurs, table(a, b) crosses two categoricals into a grid, and prop.table() turns either into proportions: the summary toolkit for label data.

9 min read · 9 cards · 2 checks

Read in: English · हिन्दी · ગુજરાતી


Theory

The columns mean() cannot touch

Half the CampusPulse survey is labels: city, stay, the usage bands you recoded. mean(survey$city) is nonsense (Unit 1's oldest rule).

Yet the dean's questions about labels are perfectly numeric: how many from each city? Do hostellers skew toward heavy usage?

The toolkit for counting labels is one function with two gears: table() alone counts one variable; give it two and it builds the cross-tabulation: the grid every social statistic starts from.

Theory

Tally marks, then a tally grid

One-way table: the classic tally sheet: a stroke under each city as forms are read out: Surat |||| ||, Navsari |||.

Cross-tab: a grid of tally cells: rows are stay (hostel/day), columns are usage (heavy/normal), and each form adds one stroke to exactly ONE cell. The finished grid answers relationship questions no single tally sheet can: that is why it gets its own name.

Practical

Counts, grids, proportions

survey <- read.csv("survey_clean.csv")  # city, stay, usage (recoded)

# ONE-WAY: frequency table of a categorical
table(survey$city)
#  bardoli navsari   surat
#        8      16      36

# sorted, most common first
sort(table(survey$city), decreasing = TRUE)

# TWO-WAY: the cross-tabulation (rows = 1st arg, cols = 2nd)
t <- table(survey$stay, survey$usage)
t
#          heavy normal
#   day       6     30
#   hostel   14     10

# proportions: three DIFFERENT questions
prop.table(t)              # share of ALL students per cell
prop.table(t, margin = 1)  # within each ROW (each row sums to 1)
prop.table(t, margin = 2)  # within each COLUMN

# formula spelling + totals
xtabs(~ stay + usage, data = survey)
addmargins(t)              # adds row/column sums

Theory

margin: the argument that changes the question

The same grid, three proportion questions:

  • prop.table(t): what share of ALL surveyed students is each cell? (cells sum to 1 overall)
  • margin = 1 (rows): within hostellers, what share is heavy? Each row sums to 1: compares usage across stay groups.
  • margin = 2 (columns): within heavy users, what share is hostellers? Each column sums to 1.

Row 'hostel': 14/(14+10) = 58% heavy with margin = 1. The dean's "do hostellers skew heavy?" is a margin = 1 question: choosing the margin IS the analysis.

Quiz

From the grid above (hostel: 14 heavy, 10 normal; day: 6 heavy, 30 normal), the dean asks: "what percentage of HOSTELLERS are heavy users?" Which command and which number?

  1. prop.table(t, margin = 1): the hostel row gives 14/24 ≈ 58%
  2. prop.table(t, margin = 2): the heavy column gives 14/20 = 70%
  3. prop.table(t): 14/60 ≈ 23%
  4. table(t): the raw 14 is already a percentage
Show the answer

prop.table(t, margin = 1): the hostel row gives 14/24 ≈ 58%

"Of hostellers" fixes the row as the whole: row-wise proportions, margin = 1: 14 of the 24 hostellers ≈ 58%. Option B answers the REVERSE question (of heavy users, how many are hostellers: 70%): a different, also-valid question that careless reading swaps. Option C shares against everyone. The margin argument is exactly where these three questions part ways: exams and real reports both hinge on it.

Think first

Read the grid like an analyst

The margin = 1 table shows: hostel → heavy 0.58, normal 0.42; day → heavy 0.17, normal 0.83. Before tapping: what is the honest one-sentence finding, and what CAUTION must accompany it (Unit 2 taught it)?

Show the answer

Finding: hostellers are far likelier to be heavy users than day scholars (58% vs 17%): a strong association between stay and usage.

Caution: association is not causation: hostels do not necessarily CAUSE heavy usage (free evenings, wifi access, or who chooses hostels could drive both). The scatter-plot lesson's rule applies to grids too: cross-tabs reveal that two labels move together, never why.

Watch out

The tabulation slips

Counts when proportions were asked (and vice versa): read the question's "how many" vs "what share".

margin mix-ups: margin = 1 rows, margin = 2 columns: within-row questions are the common exam phrasing.

NA vanishes silently: table() drops missing values by default: table(x, useNA = "ifany") shows them: report holes, as always.

Theory

Where cross-tabs go next

Every opinion poll ("support by age group"), every medical study ("outcome by treatment"), every churn report ("cancelled by plan") is a cross-tab wearing business clothes: and the chi-square test you may meet later semesters is simply the formal way to ask whether a grid's association is real or luck. pandas spells these value_counts() and pd.crosstab(): fourth tool, same tally grid.

Summary

Key takeaways

  • table(x): frequency counts for one categorical: the mean() of label data.
  • table(a, b): the cross-tabulation grid: rows = first argument, columns = second.
  • prop.table(t): overall shares; margin = 1 row-wise, margin = 2 column-wise: the margin IS the question.
  • xtabs(~ a + b, df) is the formula spelling; addmargins() adds totals.
  • table() drops NA silently: useNA = 'ifany' to show them.
  • Grid associations are not causation.
  • Memory hook: tally sheet, then tally grid.

Study this properly

This page is the lesson to read. In Gri-Learn the same topic is a graded deck: the self-checks are scored and your weak topics are tracked. Free to start.

Start this topic

Already have an account? Sign in

More from Working with Data in R

Gri-Learn · syllabus-mapped B.C.A. lessons in English, Hindi and Gujarati

Generating frequency tables and cross-tabulations · Statistical Methods and Data Analysis (MDC-03) · Gri-Learn