Introduction to Regular Expressions (Basic and Extended)

एक regular expression एक pattern है जो text को describe करता है इसे spell out करने की बजाय: एक dot का मतलब है कोई भी character, square brackets एक set, एक anchor match को line के start या end से बाँधता है, और ये building blocks आपको text के shapes खोजने देते हैं, सिर्फ़ exact words नहीं।

11 min read · 8 cards · 2 checks

Read in: English · हिन्दी · ગુજરાતી


Theory

Text को इसके Shape से Describe करना

कभी-कभी आप वह exact text नहीं जानते जो आप चाहते हैं, सिर्फ़ इसका shape। हर line जो ERROR से शुरू होती है। हर string जिसमें एक digit है। हर word जो .txt से ख़त्म होती है। आप इन्हें एक fixed string से search नहीं कर सकते, क्योंकि ये patterns हैं, specific words नहीं।

एक regular expression (regex) exactly यही है: एक compact pattern जो possible texts के एक set को describe करता है। यह command line के सबसे powerful ideas में से एक है, और यह उस grep और sed tools को power करता है जो आप अगले मिलेंगे। यह lesson building blocks introduce करता है।

Theory

Building Blocks

एक regex ordinary characters (जो खुद को match करते हैं) plus special वालों (जिनमें powers होती हैं) से built होता है। Essentials:

  • . कोई भी single character match करता है।
  • [ ] एक set में से कोई एक character match करता है: [aeiou] एक vowel, [0-9] एक digit।
  • ^ line के start से anchor करता है; $ end से anchor करता है।
  • ***** का मतलब है 'preceding item के zero या ज़्यादा'।

तो ^ERROR ERROR से शुरू होने वाली lines match करता है, [0-9] कोई भी digit वाली line match करता है, और ^$ एक empty line match करता है (start तुरंत end के बाद)।

At a glance

Patternक्या Match करता है
.कोई भी single character
[0-9]कोई एक digit (एक range); [a-z] कोई lowercase letter
[^0-9]कोई एक character जो digit NAHI है (negated set)
^एक line का start (anchor)
$एक line का end (anchor)
*Preceding item के zero या ज़्यादा
+ ? | ( )Extended regex: one-or-more, zero-or-one, OR, grouping

Practical

Action में Patterns (grep से verified)

# [0-9] matches any line containing a digit
$ printf "abc\na1b\nxyz\n" | grep "[0-9]"
a1b

# ^ERROR matches lines that START with ERROR
$ grep "^ERROR" log.txt
ERROR disk full
ERROR timeout

# Extended regex: ERROR OR WARN (needs grep -E)
$ grep -E "ERROR|WARN" log.txt
ERROR disk full
ERROR timeout
WARN low memory

Formula

Anchors और Classes Workhorses हैं

दो ideas everyday patterns में ज़्यादातर काम करते हैं। Anchors एक match को एक position से pin करते हैं: ^ line के start के लिए, $ end के लिए, तो ^ERROR सिर्फ़ वे lines ढूँढता है जो ERROR से शुरू होती हैं, ऐसी नहीं जो सिर्फ़ इसे contain करती हैं। [ ] में Character classes एक set में से एक character match करते हैं, [0-9] और [a-z] जैसे ranges के साथ, और negate करने के लिए [^...]।

इन्हें combine कीजिए और आप surprisingly precise shapes describe कर सकते हैं: उदाहरण के लिए digit से शुरू होने वाली lines के लिए ^[0-9]। पहले anchors और classes सीखिए; ये लगभग हर real pattern में appear होते हैं।

Quiz

Regular expression ^ERROR क्या match करता है?

  1. कोई भी line जिसमें कहीं भी word ERROR है
  2. सिर्फ़ वे lines जो ERROR से शुरू होती हैं, क्योंकि ^ match को line के start से anchor करता है
  3. वे lines जो ERROR से ख़त्म होती हैं
  4. Literal text caret-E-R-R-O-R, ^ symbol समेत
Show the answer

सिर्फ़ वे lines जो ERROR से शुरू होती हैं, क्योंकि ^ match को line के start से anchor करता है

^ एक start-of-line anchor है, तो ^ERROR सिर्फ़ वे lines match करता है जो ERROR से BEGIN होती हैं। Option A anchor के meaning को drop करता है: एक plain ERROR (^ के बिना) line पर कहीं भी ERROR match करता है, पर ^ इसे start तक restrict करता है। Option C दूसरे anchor को describe करता है, $, जो line के end से match करता है (ERROR$ ERROR से ख़त्म होने वाली lines match करेगा)। Option D ग़लत है: regex में ^ special है (एक anchor), literal character नहीं; pattern matched text में एक caret शामिल नहीं करता। Anchors आपका match position करते हैं: ^ start पर, $ end पर, जो precise searching के लिए essential है।

Think first

दो Flavours, Basic और Extended, क्यों हैं?

Regex basic (BRE) और extended (ERE) forms में आता है। Split क्यों, और यह कब matter करता है? फिर tap कीजिए।

Show the answer

यह ज़्यादातर एक historical inheritance है, और practical difference है कौन से special characters WITHOUT backslashes काम करते हैं। पुराने BASIC regular expressions (BRE), जो grep और sed default रूप से इस्तेमाल करते हैं, कुछ powerful operators, +, ?, |, और grouping ( ), को ordinary characters की तरह treat करते हैं जब तक आप इन्हें एक backslash से escape न करें (\+, \?, \|, \( \))। नए EXTENDED regular expressions (ERE), जो grep -E (egrep) और sed -E इस्तेमाल करते हैं, उन्हीं same operators को AUTOMATICALLY special treat करते हैं, तो आप इन्हें plainly लिखते हैं: 'ERROR या WARN' के लिए ERROR|WARN, या 'एक या ज़्यादा ab' के लिए (ab)+। Matching power essentially same है; ERE बस लिखने में cleaner है जब आपको alternation, grouping, या one-or-more चाहिए। यही वजह है, |, +, या ( ) वाली किसी भी चीज़ के लिए, लोग grep -E की तरफ़ reach करते हैं: pattern naturally पढ़ता है instead of backslashes से littered होने की बजाय। Simple patterns के लिए (एक word, एक anchor, एक character class), basic regex बिल्कुल fine है और दोनों flavours identically behave करते हैं। तो rule of thumb है: simple pattern, दोनों काम करते हैं; OR, grouping, या plus चाहिए, extended (grep -E) prefer कीजिए escaping avoid करने के लिए। कौन सा tool कौन सा flavour इस्तेमाल करता है यह जानना आपको mysterious 'मेरा | क्यों काम नहीं कर रहा?' moments से बचाता है।

Summary

Key takeaways

  • एक regular expression एक pattern है जो text को इसके shape से describe करता है, exact spelling से नहीं।
  • Core elements: . (कोई भी character), [ ] (एक set, [0-9] जैसा), [^...] (negated set), * (zero या ज़्यादा)।
  • Anchors एक match position करते हैं: ^ इसे line के start से बाँधता है, $ end से।
  • तो ^ERROR ERROR से शुरू होने वाली lines match करता है; [0-9] digit वाली lines match करता है।
  • Basic regex (BRE, grep/sed इस्तेमाल करते हैं) को + ? | ( ) के लिए backslashes चाहिए; extended regex (ERE, grep -E) को नहीं।
  • Extended regex (grep -E) की तरफ़ reach कीजिए जब आपको alternation |, grouping ( ), या one-or-more + चाहिए।
  • Memory hook: . कोई भी, [ ] एक set, ^ start, $ end; OR और grouping के लिए -E।

Study this properly

This page is the lesson to read. In Gri-Learn the same topic is a graded deck: the self-checks are scored and your weak topics are tracked. Free to start.

Start this topic

Already have an account? Sign in

More from Advanced Text Processing Tools

Gri-Learn · syllabus-mapped B.C.A. lessons in English, Hindi and Gujarati

Introduction to Regular Expressions (Basic and Extended) · Linux Operating System (LOS) (Minor-04) · Gri-Learn