Reordering and reshaping data frames

order() returns the row positions that would sort a data frame (df[order(df$screen), ] sorts it), sort() only sorts a lone vector, and reshaping moves data between wide and long layouts.

9 min read · 9 cards · 2 checks

Read in: English · हिन्दी · ગુજરાતી


Theory

The dean wants it worst-first

The cleaned survey is ready, and the dean's next request is the oldest one in reporting: "same table, but sorted: highest screen time first."

In SQL this was ORDER BY. In pandas, sort_values. R's spelling looks stranger at first:

survey[order(-survey$screen), ]

Why "order" and not "sort"? Because R splits the job into two different functions, and confusing them is the lesson's famous trap. One returns values; the other returns positions.

Theory

Sorting people vs calling out seat numbers

Two ways to arrange a class by height:

  • sort() physically lines up the HEIGHTS: 150, 158, 163... just the numbers, detached from their owners.
  • order() calls out seat positions: "seat 7 first, then seat 2, then seat 9...": a rearrangement plan that the whole class (every column!) can follow.

To sort a data frame you need the PLAN, not the bare values: rows must move together, or scores drift away from their roll numbers.

Practical

order() sorts tables; sort() sorts vectors

survey <- read.csv("survey_clean.csv")   # roll, city, stay, screen

# sort a LONE VECTOR: sort() returns the values, ascending
sort(survey$screen)                  # 60 90 120 150 240 ...
sort(survey$screen, decreasing = TRUE)

# sort the DATA FRAME: order() returns row positions, brackets apply them
survey[order(survey$screen), ]                   # ascending
survey[order(survey$screen, decreasing = TRUE), ] # descending
survey[order(-survey$screen), ]                  # minus: same, numerics only

# multiple keys: city A-Z, then screen high-to-low within city
survey[order(survey$city, -survey$screen), ]

# what order() actually returns: a permutation of row numbers
order(c(300, 100, 200))     # [1] 2 3 1  (2nd smallest first...)

# reorder COLUMNS: plain subsetting
survey <- survey[, c("roll", "screen", "city", "stay")]

Theory

Reshaping: wide vs long

Sorting rearranges rows; reshaping rearranges the table's very layout.

Wide: one row per student, one column per test:

roll test1 test2 test3

Long: one row per student-test pair:

roll test score

Wide reads nicely; long computes nicely (group by test, plot by test: every tool prefers it). Base R's reshape() converts between them but is famously clunky; the modern tools are tidyr's pivot_longer()/pivot_wider() (names worth knowing). For exams: know the two shapes, their trade-off, and that conversion tools exist.

Quiz

A student sorts the survey with survey$screen <- sort(survey$screen) and moves on. What silently broke?

  1. Only the screen column was sorted: every value now sits beside the WRONG roll number: row alignment is destroyed
  2. Nothing: sorting one column sorts the whole data frame
  3. R raised an error, so no harm done
  4. The column was deleted
Show the answer

Only the screen column was sorted: every value now sits beside the WRONG roll number: row alignment is destroyed

sort() returned the bare values and the assignment overwrote the column IN PLACE: screen times are now ascending while roll, city and stay kept their old order: student 101 wears someone else's screen time, with no error to warn you. This silent misalignment is why data frames are sorted with df[order(df$col), ]: the order() plan moves every column together. Arguably the most dangerous single line a beginner can run.

Think first

Read the permutation

screen <- c(300, 100, 200) for rolls 101, 102, 103. Work out: (1) what order(screen) returns, (2) which roll ends up FIRST in survey[order(screen), ], (3) what sort(screen) would return. Then tap.

Show the answer

1. order(screen) = 2 3 1: "row 2 is smallest, then row 3, then row 1".

2. Row 2 leads, so roll 102 (screen 100) comes first: ascending by default.

3. sort(screen) = 100 200 300: values only, owners left behind.

If you can read "2 3 1" as a seating plan rather than as values, order() has clicked: the bracket notation df[plan, ] simply seats every column by that plan.

Watch out

The three sorting slips

sort() inside df[...]: breaks alignment or errors: data frames take order().

The missing comma: df[order(df$screen), ]: the row-slot comma from the subsetting lesson still rules.

Minus-for-descending works only on numeric keys: for text or mixed keys use decreasing = TRUE (applies to all keys) or wrap keys accordingly.

Theory

Third language, same ORDER BY

SQL: ORDER BY screen DESC. pandas: sort_values('screen', ascending=False). R: df[order(-df$screen), ]. One idea, three accents: and in all three, multi-key sorting (city, then screen) is just listing the keys in priority order. Next lesson the survey meets a second table: merge(), R's JOIN, and your BCA303 join instincts come straight back to work.

Summary

Key takeaways

  • order(x) returns the row-position PLAN that would sort x; df[order(x), ] applies it to every column.
  • sort(x) returns sorted VALUES of a lone vector: never use it to sort a data frame column in place.
  • Descending: decreasing = TRUE, or a minus sign on numeric keys; multi-key: order(key1, key2).
  • Reorder columns by subsetting with a name vector.
  • Wide (one column per measure) vs long (one row per observation): long computes and plots better.
  • reshape() exists in base R; tidyr's pivot_longer/pivot_wider are the modern names.
  • Memory hook: sort lines up heights; order calls out seat numbers.

Study this properly

This page is the lesson to read. In Gri-Learn the same topic is a graded deck: the self-checks are scored and your weak topics are tracked. Free to start.

Start this topic

Already have an account? Sign in

More from Working with Data in R

Gri-Learn · syllabus-mapped B.C.A. lessons in English, Hindi and Gujarati

Reordering and reshaping data frames · Statistical Methods and Data Analysis (MDC-03) · Gri-Learn