Theory
Describing text by its shape
Sometimes you do not know the exact text you want, only its shape. Every line that starts with ERROR. Every string that contains a digit. Every word ending in .txt. You cannot search for these with a fixed string, because they are patterns, not specific words.
A regular expression (regex) is exactly this: a compact pattern that describes a set of possible texts. It is one of the most powerful ideas on the command line, and it powers the grep and sed tools you will meet next. This lesson introduces the building blocks.
Theory
The building blocks
A regex is built from ordinary characters (which match themselves) plus special ones (which have powers). The essentials:
- . matches any single character.
- [ ] matches any one character from a set:
[aeiou]a vowel,[0-9]a digit. - ^ anchors to the start of a line; $ anchors to the end.
- ***** means 'zero or more of the preceding item'.
So ^ERROR matches lines beginning with ERROR, [0-9] matches any line containing a digit, and ^$ matches an empty line (start immediately followed by end).
At a glance
| Pattern | Matches |
|---|---|
| . | Any single character |
| [0-9] | Any one digit (a range); [a-z] any lowercase letter |
| [^0-9] | Any one character that is NOT a digit (negated set) |
| ^ | The start of a line (anchor) |
| $ | The end of a line (anchor) |
| * | Zero or more of the preceding item |
| + ? | ( ) | Extended regex: one-or-more, zero-or-one, OR, grouping |
Practical
Patterns in action (verified with grep)
# [0-9] matches any line containing a digit
$ printf "abc\na1b\nxyz\n" | grep "[0-9]"
a1b
# ^ERROR matches lines that START with ERROR
$ grep "^ERROR" log.txt
ERROR disk full
ERROR timeout
# Extended regex: ERROR OR WARN (needs grep -E)
$ grep -E "ERROR|WARN" log.txt
ERROR disk full
ERROR timeout
WARN low memoryFormula
Anchors and classes are the workhorses
Two ideas do most of the work in everyday patterns. Anchors pin a match to a position: ^ for the start of a line, $ for the end, so ^ERROR finds only lines that begin with ERROR, not ones that merely contain it. Character classes in [ ] match one character from a set, with ranges like [0-9] and [a-z], and [^...] to negate.
Combine them and you can describe surprisingly precise shapes: ^[0-9] for lines starting with a digit, for example. Learn anchors and classes first; they appear in nearly every real pattern.
Quiz
What does the regular expression ^ERROR match?
- Any line that contains the word ERROR anywhere
- Only lines that begin with ERROR, because ^ anchors the match to the start of the line
- Lines that end with ERROR
- The literal text caret-E-R-R-O-R, including the ^ symbol
Show the answer
Only lines that begin with ERROR, because ^ anchors the match to the start of the line
The ^ is a start-of-line anchor, so ^ERROR matches only lines that BEGIN with ERROR. Option A drops the anchor's meaning: a plain ERROR (no ^) matches ERROR anywhere on a line, but ^ restricts it to the start. Option C describes the OTHER anchor, $, which matches the end of a line (ERROR$ would match lines ending in ERROR). Option D is wrong: ^ is special in a regex (an anchor), not a literal character; the pattern does not include a caret in the matched text. Anchors position your match: ^ at the start, $ at the end, which is essential for precise searching.
Think first
Why are there two flavours, basic and extended?
Regex comes in basic (BRE) and extended (ERE) forms. Why the split, and when does it matter? Then tap.
Show the answer
It is largely a historical inheritance, and the practical difference is which special characters work WITHOUT backslashes. The older BASIC regular expressions (BRE), used by grep and sed by default, treat some powerful operators, +, ?, |, and grouping ( ), as ordinary characters unless you escape them with a backslash (\+, \?, \|, \( \)). The newer EXTENDED regular expressions (ERE), used by grep -E (egrep) and sed -E, treat those same operators as special AUTOMATICALLY, so you write them plainly: ERROR|WARN for 'ERROR or WARN', or (ab)+ for 'one or more ab'. The matching power is essentially the same; ERE is just cleaner to write when you need alternation, grouping, or one-or-more. That is why, for anything with |, +, or ( ), people reach for grep -E: the pattern reads naturally instead of being littered with backslashes. For simple patterns (a word, an anchor, a character class), basic regex is perfectly fine and both flavours behave identically. So the rule of thumb is: simple pattern, either works; needs OR, grouping, or plus, prefer extended (grep -E) to avoid escaping. Knowing which flavour a tool uses saves you from mysterious 'why won't my | work?' moments.
Summary
Key takeaways
- A regular expression is a pattern that describes text by its shape, not by exact spelling.
- Core elements: . (any character), [ ] (a set, like [0-9]), [^...] (negated set), * (zero or more).
- Anchors position a match: ^ ties it to the start of a line, $ to the end.
- So ^ERROR matches lines beginning with ERROR; [0-9] matches lines containing a digit.
- Basic regex (BRE, used by grep/sed) needs backslashes for + ? | ( ); extended regex (ERE, grep -E) does not.
- Reach for extended regex (grep -E) when you need alternation |, grouping ( ), or one-or-more +.
- Memory hook: . any, [ ] a set, ^ start, $ end; -E for OR and grouping.