Theory
Small tools, big questions
A log file on the server has 10000 lines. You want just the last 20. Or a count of how many times each error appears. Or the third column of a data file. Linux does not give you one giant program for this; it gives you many small text tools, each doing a single job well.
The family includes head, tail, cut, sort, uniq, wc, tr, cmp, and tee. Each reads text and writes text, and, as you will see, they are designed to be chained together. This lesson introduces them and previews that chaining.
At a glance
| Command | What it does |
|---|---|
| head -n N | Show the first N lines |
| tail -n N | Show the last N lines |
| cut -d D -f N | Extract field N, using D as the delimiter |
| sort | Sort lines into order |
| uniq | Collapse adjacent duplicate lines (uniq -c counts them) |
| wc | Count lines (-l), words (-w), or bytes (-c) |
| tr | Translate or delete characters (e.g. tr a-z A-Z) |
| cmp / tee | Compare two files / write to a file and pass output through |
Practical
Worked examples (verified output)
$ cat fruits.txt
banana
apple
cherry
apple
$ wc -l fruits.txt # count lines
4 fruits.txt
$ sort fruits.txt | uniq -c # sort, then count each unique line
2 apple
1 banana
1 cherry
$ head -2 fruits.txt # first 2 lines
banana
apple
$ echo "hello" | tr a-z A-Z # translate lowercase to uppercase
HELLOFormula
These are filters, built to be piped
Each of these tools reads from standard input and writes to standard output. That means you can feed the output of one straight into the next using the pipe symbol |, which you will study properly in Unit 3.
That is the real power: sort fruits.txt | uniq -c sends sorted lines into uniq to count them. A few tiny tools, combined, answer questions no single one could. This 'do one thing well, then combine' idea is the heart of the Linux command line.
Quiz
Your file has the lines: banana, apple, cherry, apple (in that order). Why do you run sort before uniq, as in sort file | uniq?
- sort is just decoration; uniq removes duplicates anywhere in the file
- uniq only collapses ADJACENT duplicate lines, so sorting first brings identical lines together for uniq to combine
- sort deletes duplicates, so uniq is unnecessary
- uniq requires the file to be uppercase first
Show the answer
uniq only collapses ADJACENT duplicate lines, so sorting first brings identical lines together for uniq to combine
uniq only removes duplicates that are ADJACENT (right next to each other). In the original order (banana, apple, cherry, apple), the two 'apple' lines are not adjacent, so uniq alone would not combine them. Sorting first groups all identical lines together (apple, apple, banana, cherry), and then uniq can collapse them, which is why sort | uniq is the standard idiom. Option A is wrong: without sorting, uniq misses non-adjacent duplicates. Option C is wrong: sort orders lines, it does not remove duplicates (that is uniq's job, or sort -u). Option D is invented. Remember the pairing: sort THEN uniq, because uniq only sees neighbours.
Think first
Why so many tiny tools instead of one big one?
Why does Linux give you a dozen small text commands rather than a single powerful text program? Then tap.
Show the answer
Because small, single-purpose tools that COMBINE are more flexible and reusable than one large fixed program, this is the famous Unix philosophy: 'do one thing well, and work together'. Each tool, sort, uniq, cut, wc, has one clear job, which makes it simple to learn, easy to trust, and easy to maintain. The power comes from CONNECTING them with pipes: because every tool reads standard input and writes standard output, you can snap them together in endless combinations to solve problems nobody anticipated. Want the three most common errors in a log? cut the error column, sort it, uniq -c to count, sort -rn to rank, head -3 to take the top, five tiny tools you already know, assembled on the spot. A single giant 'text program' could never anticipate every such need, but small composable tools let YOU build the exact pipeline you want. Fewer features per tool, but limitless combinations: that is why the toolbox beats the monolith, and why learning these small commands pays off far beyond their individual simplicity.
Summary
Key takeaways
- Linux provides many small text tools, each doing one job: head, tail, cut, sort, uniq, wc, tr, cmp, tee.
- head -n shows the first N lines; tail -n the last N; cut extracts a field by delimiter; wc counts lines/words/bytes.
- sort orders lines; uniq collapses ADJACENT duplicates (uniq -c counts), so you usually sort before uniq.
- tr translates or deletes characters (tr a-z A-Z uppercases); cmp compares files; tee writes to a file and passes output on.
- These are filters: they read standard input and write standard output, so they combine with the pipe | (Unit 3).
- Combining tiny tools with pipes answers questions no single command could: the core Linux idea.
- Memory hook: do one thing well, then pipe them together.