Theory
एक language जहाँ data plural है
Console पर अपनी पहली असली R type कीजिए:
heights <- c(160, 168, 175, 162, 171)
mean(heights)
जवाब: [1] 167.2। दो lines: पाँच students store हुए, average compute हुआ।
नोटिस कीजिए आपने क्या नहीं लिखा: कोई loop नहीं, कोई import नहीं, कोई square-bracket ceremony नहीं। R का founding design choice: data की basic unit एक पूरा collection है, अकेला value नहीं: क्योंकि statistics कभी एक नंबर पर study नहीं करती।
Theory
Arrow चीज़ों को boxes में डालता है
heights <- c(160, 168, 175) बिल्कुल वैसे ही पढ़ी जाती है जैसे दिखती है: values arrow के साथ बहते हुए उस box में जाते हैं जिसका नाम heights है।
<- R का assignment है: visual "put into"। (सादा = भी ज़्यादातर काम करता है, पर <- वह convention है जो हर textbook, paper और examiner चाहता है: arrow लिखिए।)
और c() combine function है: R का यह कहने का तरीक़ा "ये values एक साथ एक vector के रूप में यात्रा करते हैं"।
Theory
Vectors और vectorised superpower
एक vector समान-type values का एक ordered collection है: R की fundamental unit। यहाँ तक कि x <- 5 भी length 1 का vector है: यही कारण है हर output [1] से शुरू होता है: R elements को numbering कर रहा है।
Superpower: operations हर element पर एक साथ apply होते हैं:
heights + 2→ हर height 2 से बढ़ी।heights > 165→ एक logical vector: FALSE TRUE TRUE FALSE TRUE।
जहाँ Python की सादी lists को loop चाहिए, R बस... कर देता है। एक statistical thought, एक line।
At a glance
वे data types जिन पर statistics चलती है
| Type | Example | Notes |
|---|---|---|
| numeric | 168, 3.14 | Numbers के लिए default |
| integer | 60L | L suffix integer force करता है |
| character | "Surat" | Quotes required |
| logical | TRUE, FALSE | Uppercase; comparisons से |
| factor | factor(c("hostel","day")) | Levels वाला categorical data |
Practical
Survey R की दस lines
# assignment: the arrow puts values into names
heights <- c(160, 168, 175, 162, 171) # numeric vector
cities <- c("Surat", "Navsari", "Surat", "Bardoli", "Surat")
stay <- factor(c("hostel", "day", "hostel", "day", "day"))
mean(heights) # [1] 167.2 built in, no import
sd(heights) # [1] 6.058
length(heights) # [1] 5
heights > 165 # [1] FALSE TRUE TRUE FALSE TRUE (vectorised!)
class(cities) # [1] "character"
levels(stay) # [1] "day" "hostel" the factor's categories
Think first
तीन जवाब अंदाज़ा लगाइए
heights <- c(160, 168, 175, 162, 171) के साथ, अंदाज़ा लगाइए R हर एक के लिए क्या print करता है, फिर tap कीजिए:
1. heights * 2
2. heights > 170
3. mean(heights > 170)
Show the answer
1. 320 336 350 324 342: हर element double हुआ (vectorised)।
2. FALSE FALSE TRUE FALSE TRUE: एक logical vector।
3. 0.4: sly वाला: TRUE 1 गिनता है, FALSE 0, तो logicals का mean = 170 से ऊपर का PROPORTION (5 में से 2)।
वह तीसरा trick (एक condition का mean = इसे satisfy करने का proportion) एक loved R idiom है और एक loved exam question: एक line एक percentage compute करती है।
Quiz
एक student type करता है: marks <- c(78, 85, "absent", 91) और फिर mean(marks) error देता है। क्या हुआ?
- Vectors homogeneous हैं: उस एक string ने चुपचाप सारे elements को character में coerce कर दिया, तो mean() को text मिला
- c() तीन से ज़्यादा values नहीं रख सकता
- absent शब्द को uppercase चाहिए था
- mean() को values पहले sorted चाहिए
Show the answer
Vectors homogeneous हैं: उस एक string ने चुपचाप सारे elements को character में coerce कर दिया, तो mean() को text मिला
R vectors एक TYPE रखते हैं। numbers और एक string के सामने, R चुपचाप सब कुछ सबसे flexible type में बदल देता है: character: "78" अब text है, और mean() शब्दों का average नहीं निकाल सकता। यह silent coercion R का सबसे चालाक beginner trap है ठीक इसलिए क्योंकि c() line ख़ुद कोई error नहीं दिखाती: class(marks) से check कीजिए। missing value का statistical fix NA है, अगली unit का topic।
Watch out
चार syntax stumbles
Case: mean() काम करता है, Mean() नहीं: R capitals कभी माफ़ नहीं करता (TRUE, true नहीं)।
Quotes: बिना quotes के Surat को एक variable name माना जाता है और error देता है।
Coercion: एक string एक numeric vector को चुपचाप text में बदल देती है।
Arrow: exams में <- लिखिए; examiners = को not-quite-R पढ़ते हैं वहाँ भी जहाँ यह चलता है।
Theory
Factors: Unit 1 की vocabulary एक type बनती है
नोटिस कीजिए factor() loop बंद कर रहा है: Unit 1 के categorical variables यहाँ literally एक data type हैं: levels() categories list करता है, और R इन्हें average करने से मना करता है (सही तरीक़े से: वह nonsense-maths trap याद कीजिए)। Numeric vectors = Unit 1 के numerical variables। जो classification आपने हाथ से सीखी वह language द्वारा enforce होती है: अगला lesson, पूरी survey files read.csv के ज़रिए आती हैं, typed और तैयार।
Summary
Key takeaways
- Assignment: name <- value (arrow ही exam-expected form है); # comment शुरू करता है।
- c() values को एक VECTOR में combine करता है, R की fundamental unit; अकेली values length-1 vectors हैं ([1])।
- Operations vectorised हैं: heights + 2 और heights > 165 हर element पर लगते हैं, कोई loops नहीं।
- Types: numeric, integer (L), character (quoted), logical (TRUE/FALSE), factor (levels वाला categorical)।
- Vectors homogeneous हैं: एक अकेली string चुपचाप सब कुछ character में coerce कर देती है।
- mean(condition) एक proportion compute करता है: TRUE = 1, FALSE = 0।
- Memory hook: arrow values को box में डालता है; c() उन्हें साथ यात्रा कराता है।