Theory
छोटे Tools, बड़े Questions
Server पर एक log file में 10000 lines हैं। आपको बस last 20 चाहिए। या हर error कितनी बार appear होता है इसकी एक count। या एक data file का तीसरा column। Linux आपको इसके लिए एक giant program नहीं देता; यह आपको कई छोटे text tools देता है, हर एक एक single job अच्छी तरह करते हुए।
Family में शामिल हैं head, tail, cut, sort, uniq, wc, tr, cmp, और tee। हर एक text पढ़ता है और text लिखता है, और, जैसा आप देखेंगे, इन्हें साथ chained होने के लिए design किया गया है। यह lesson इन्हें introduce करता है और वह chaining preview करता है।
At a glance
| Command | यह क्या करता है |
|---|---|
| head -n N | पहली N lines दिखाता है |
| tail -n N | आख़िरी N lines दिखाता है |
| cut -d D -f N | field N extract करता है, D को delimiter की तरह इस्तेमाल करते हुए |
| sort | Lines को order में sort करता है |
| uniq | Adjacent duplicate lines collapse करता है (uniq -c इन्हें count करता है) |
| wc | Lines count करता है (-l), words (-w), या bytes (-c) |
| tr | Characters translate या delete करता है (जैसे tr a-z A-Z) |
| cmp / tee | दो files compare करता है / एक file में लिखता है और output pass through करता है |
Practical
Worked Examples (verified output)
$ cat fruits.txt
banana
apple
cherry
apple
$ wc -l fruits.txt # count lines
4 fruits.txt
$ sort fruits.txt | uniq -c # sort, then count each unique line
2 apple
1 banana
1 cherry
$ head -2 fruits.txt # first 2 lines
banana
apple
$ echo "hello" | tr a-z A-Z # translate lowercase to uppercase
HELLOFormula
ये Filters हैं, Piped होने के लिए Built
इनमें से हर tool standard input से पढ़ता है और standard output पर लिखता है। इसका मतलब है आप एक का output दूसरे में सीधे pipe symbol | इस्तेमाल करके feed कर सकते हैं, जिसे आप Unit 3 में properly पढ़ेंगे।
यही real power है: sort fruits.txt | uniq -c sorted lines को uniq में भेजता है इन्हें count करने के लिए। कुछ tiny tools, combined, ऐसे सवालों के answer देते हैं जो कोई single tool नहीं दे सकता। यह 'एक चीज़ अच्छी तरह करो, फिर combine करो' idea Linux command line का heart है।
Quiz
आपकी file में lines हैं: banana, apple, cherry, apple (उसी order में)। आप uniq से पहले sort क्यों चलाते हैं, sort file | uniq की तरह?
- sort बस decoration है; uniq file में कहीं भी duplicates हटाता है
- uniq सिर्फ़ ADJACENT duplicate lines collapse करता है, तो पहले sorting identical lines को साथ लाता है uniq के combine करने के लिए
- sort duplicates delete करता है, तो uniq unnecessary है
- uniq को पहले file uppercase में चाहिए
Show the answer
uniq सिर्फ़ ADJACENT duplicate lines collapse करता है, तो पहले sorting identical lines को साथ लाता है uniq के combine करने के लिए
uniq सिर्फ़ वे duplicates हटाता है जो ADJACENT (एक-दूसरे के बिल्कुल बगल में) हैं। Original order में (banana, apple, cherry, apple), दोनों 'apple' lines adjacent नहीं हैं, तो अकेला uniq इन्हें combine नहीं करेगा। पहले sorting सारी identical lines को साथ group करता है (apple, apple, banana, cherry), और फिर uniq इन्हें collapse कर सकता है, यही वजह है sort | uniq standard idiom है। Option A ग़लत है: sorting के बिना, uniq non-adjacent duplicates miss करता है। Option C ग़लत है: sort lines order करता है, यह duplicates नहीं हटाता (यह uniq का job है, या sort -u)। Option D invented है। Pairing याद रखिए: sort फिर uniq, क्योंकि uniq सिर्फ़ neighbours देखता है।
Think first
एक बड़े वाले की बजाय इतने सारे Tiny Tools क्यों?
Linux आपको एक single powerful text program की बजाय एक दर्जन छोटे text commands क्यों देता है? फिर tap कीजिए।
Show the answer
क्योंकि छोटे, single-purpose tools जो COMBINE होते हैं एक बड़े fixed program से ज़्यादा flexible और reusable होते हैं, यह famous Unix philosophy है: 'एक चीज़ अच्छी तरह करो, और साथ काम करो'। हर tool, sort, uniq, cut, wc, का एक clear job है, जो इसे सीखने में simple, trust करने में easy, और maintain करने में easy बनाता है। Power pipes से इन्हें CONNECT करने से आती है: क्योंकि हर tool standard input पढ़ता है और standard output लिखता है, आप इन्हें endless combinations में साथ snap कर सकते हैं उन problems को solve करने के लिए जिन्हें किसी ने anticipate नहीं किया। एक log में तीन सबसे common errors चाहिए? error column cut कीजिए, इसे sort कीजिए, count के लिए uniq -c, rank के लिए sort -rn, top 3 लेने के लिए head -3, पाँच tiny tools जो आप पहले से जानते हैं, spot पर assembled। एक single giant 'text program' कभी ऐसी हर need anticipate नहीं कर सकता, पर छोटे composable tools आपको वह exact pipeline build करने देते हैं जो आप चाहते हैं। हर tool में कम features, पर limitless combinations: यही वजह है toolbox monolith को beat करता है, और यही वजह है ये छोटे commands सीखना इनकी individual simplicity से कहीं ज़्यादा pay off करता है।
Summary
Key takeaways
- Linux कई छोटे text tools provide करता है, हर एक एक job करते हुए: head, tail, cut, sort, uniq, wc, tr, cmp, tee।
- head -n पहली N lines दिखाता है; tail -n आख़िरी N; cut delimiter से एक field extract करता है; wc lines/words/bytes count करता है।
- sort lines order करता है; uniq ADJACENT duplicates collapse करता है (uniq -c count करता है), तो आप usually uniq से पहले sort करते हैं।
- tr characters translate या delete करता है (tr a-z A-Z uppercase करता है); cmp files compare करता है; tee एक file में लिखता है और output आगे pass करता है।
- ये filters हैं: ये standard input पढ़ते हैं और standard output लिखते हैं, तो ये pipe | (Unit 3) से combine होते हैं।
- Pipes से tiny tools combine करना ऐसे सवालों के answer देता है जो कोई single command नहीं दे सकता: core Linux idea।
- Memory hook: एक चीज़ अच्छी तरह करो, फिर इन्हें साथ pipe करो।