Theory
The dean wants it worst-first
The cleaned survey is ready, and the dean's next request is the oldest one in reporting: "same table, but sorted: highest screen time first."
In SQL this was ORDER BY. In pandas, sort_values. R's spelling looks stranger at first:
survey[order(-survey$screen), ]
Why "order" and not "sort"? Because R splits the job into two different functions, and confusing them is the lesson's famous trap. One returns values; the other returns positions.
Theory
Sorting people vs calling out seat numbers
Two ways to arrange a class by height:
- sort() physically lines up the HEIGHTS: 150, 158, 163... just the numbers, detached from their owners.
- order() calls out seat positions: "seat 7 first, then seat 2, then seat 9...": a rearrangement plan that the whole class (every column!) can follow.
To sort a data frame you need the PLAN, not the bare values: rows must move together, or scores drift away from their roll numbers.
Practical
order() sorts tables; sort() sorts vectors
survey <- read.csv("survey_clean.csv") # roll, city, stay, screen
# sort a LONE VECTOR: sort() returns the values, ascending
sort(survey$screen) # 60 90 120 150 240 ...
sort(survey$screen, decreasing = TRUE)
# sort the DATA FRAME: order() returns row positions, brackets apply them
survey[order(survey$screen), ] # ascending
survey[order(survey$screen, decreasing = TRUE), ] # descending
survey[order(-survey$screen), ] # minus: same, numerics only
# multiple keys: city A-Z, then screen high-to-low within city
survey[order(survey$city, -survey$screen), ]
# what order() actually returns: a permutation of row numbers
order(c(300, 100, 200)) # [1] 2 3 1 (2nd smallest first...)
# reorder COLUMNS: plain subsetting
survey <- survey[, c("roll", "screen", "city", "stay")]
Theory
Reshaping: wide vs long
Sorting rearranges rows; reshaping rearranges the table's very layout.
Wide: one row per student, one column per test:
roll test1 test2 test3
Long: one row per student-test pair:
roll test score
Wide reads nicely; long computes nicely (group by test, plot by test: every tool prefers it). Base R's reshape() converts between them but is famously clunky; the modern tools are tidyr's pivot_longer()/pivot_wider() (names worth knowing). For exams: know the two shapes, their trade-off, and that conversion tools exist.
Quiz
A student sorts the survey with survey$screen <- sort(survey$screen) and moves on. What silently broke?
- Only the screen column was sorted: every value now sits beside the WRONG roll number: row alignment is destroyed
- Nothing: sorting one column sorts the whole data frame
- R raised an error, so no harm done
- The column was deleted
Show the answer
Only the screen column was sorted: every value now sits beside the WRONG roll number: row alignment is destroyed
sort() returned the bare values and the assignment overwrote the column IN PLACE: screen times are now ascending while roll, city and stay kept their old order: student 101 wears someone else's screen time, with no error to warn you. This silent misalignment is why data frames are sorted with df[order(df$col), ]: the order() plan moves every column together. Arguably the most dangerous single line a beginner can run.
Think first
Read the permutation
screen <- c(300, 100, 200) for rolls 101, 102, 103. Work out: (1) what order(screen) returns, (2) which roll ends up FIRST in survey[order(screen), ], (3) what sort(screen) would return. Then tap.
Show the answer
1. order(screen) = 2 3 1: "row 2 is smallest, then row 3, then row 1".
2. Row 2 leads, so roll 102 (screen 100) comes first: ascending by default.
3. sort(screen) = 100 200 300: values only, owners left behind.
If you can read "2 3 1" as a seating plan rather than as values, order() has clicked: the bracket notation df[plan, ] simply seats every column by that plan.
Watch out
The three sorting slips
sort() inside df[...]: breaks alignment or errors: data frames take order().
The missing comma: df[order(df$screen), ]: the row-slot comma from the subsetting lesson still rules.
Minus-for-descending works only on numeric keys: for text or mixed keys use decreasing = TRUE (applies to all keys) or wrap keys accordingly.
Theory
Third language, same ORDER BY
SQL: ORDER BY screen DESC. pandas: sort_values('screen', ascending=False). R: df[order(-df$screen), ]. One idea, three accents: and in all three, multi-key sorting (city, then screen) is just listing the keys in priority order. Next lesson the survey meets a second table: merge(), R's JOIN, and your BCA303 join instincts come straight back to work.
Summary
Key takeaways
- order(x) returns the row-position PLAN that would sort x; df[order(x), ] applies it to every column.
- sort(x) returns sorted VALUES of a lone vector: never use it to sort a data frame column in place.
- Descending: decreasing = TRUE, or a minus sign on numeric keys; multi-key: order(key1, key2).
- Reorder columns by subsetting with a name vector.
- Wide (one column per measure) vs long (one row per observation): long computes and plots better.
- reshape() exists in base R; tidyr's pivot_longer/pivot_wider are the modern names.
- Memory hook: sort lines up heights; order calls out seat numbers.