Basic R syntax, variables, and data types

R `<-` arrow से assign करता है, c() से vectors बनाता है, और पूरे vectors पर एक साथ element-wise काम करता है, numeric, character, logical और factor उन types के साथ जिन पर statistics चलती है।

10 min read · 10 cards · 2 checks

Read in: English · हिन्दी · ગુજરાતી


Theory

एक language जहाँ data plural है

Console पर अपनी पहली असली R type कीजिए:

heights <- c(160, 168, 175, 162, 171)

mean(heights)

जवाब: [1] 167.2। दो lines: पाँच students store हुए, average compute हुआ।

नोटिस कीजिए आपने क्या नहीं लिखा: कोई loop नहीं, कोई import नहीं, कोई square-bracket ceremony नहीं। R का founding design choice: data की basic unit एक पूरा collection है, अकेला value नहीं: क्योंकि statistics कभी एक नंबर पर study नहीं करती।

Theory

Arrow चीज़ों को boxes में डालता है

heights <- c(160, 168, 175) बिल्कुल वैसे ही पढ़ी जाती है जैसे दिखती है: values arrow के साथ बहते हुए उस box में जाते हैं जिसका नाम heights है।

<- R का assignment है: visual "put into"। (सादा = भी ज़्यादातर काम करता है, पर <- वह convention है जो हर textbook, paper और examiner चाहता है: arrow लिखिए।)

और c() combine function है: R का यह कहने का तरीक़ा "ये values एक साथ एक vector के रूप में यात्रा करते हैं"।

Theory

Vectors और vectorised superpower

एक vector समान-type values का एक ordered collection है: R की fundamental unit। यहाँ तक कि x <- 5 भी length 1 का vector है: यही कारण है हर output [1] से शुरू होता है: R elements को numbering कर रहा है।

Superpower: operations हर element पर एक साथ apply होते हैं:

  • heights + 2 → हर height 2 से बढ़ी।
  • heights > 165 → एक logical vector: FALSE TRUE TRUE FALSE TRUE।

जहाँ Python की सादी lists को loop चाहिए, R बस... कर देता है। एक statistical thought, एक line।

At a glance

वे data types जिन पर statistics चलती है

TypeExampleNotes
numeric168, 3.14Numbers के लिए default
integer60LL suffix integer force करता है
character"Surat"Quotes required
logicalTRUE, FALSEUppercase; comparisons से
factorfactor(c("hostel","day"))Levels वाला categorical data

Practical

Survey R की दस lines

# assignment: the arrow puts values into names
heights <- c(160, 168, 175, 162, 171)   # numeric vector
cities  <- c("Surat", "Navsari", "Surat", "Bardoli", "Surat")
stay    <- factor(c("hostel", "day", "hostel", "day", "day"))

mean(heights)        # [1] 167.2   built in, no import
sd(heights)          # [1] 6.058
length(heights)      # [1] 5

heights > 165        # [1] FALSE TRUE TRUE FALSE TRUE  (vectorised!)
class(cities)        # [1] "character"
levels(stay)         # [1] "day" "hostel"   the factor's categories

Think first

तीन जवाब अंदाज़ा लगाइए

heights <- c(160, 168, 175, 162, 171) के साथ, अंदाज़ा लगाइए R हर एक के लिए क्या print करता है, फिर tap कीजिए:

1. heights * 2

2. heights > 170

3. mean(heights > 170)

Show the answer

1. 320 336 350 324 342: हर element double हुआ (vectorised)।

2. FALSE FALSE TRUE FALSE TRUE: एक logical vector।

3. 0.4: sly वाला: TRUE 1 गिनता है, FALSE 0, तो logicals का mean = 170 से ऊपर का PROPORTION (5 में से 2)।

वह तीसरा trick (एक condition का mean = इसे satisfy करने का proportion) एक loved R idiom है और एक loved exam question: एक line एक percentage compute करती है।

Quiz

एक student type करता है: marks <- c(78, 85, "absent", 91) और फिर mean(marks) error देता है। क्या हुआ?

  1. Vectors homogeneous हैं: उस एक string ने चुपचाप सारे elements को character में coerce कर दिया, तो mean() को text मिला
  2. c() तीन से ज़्यादा values नहीं रख सकता
  3. absent शब्द को uppercase चाहिए था
  4. mean() को values पहले sorted चाहिए
Show the answer

Vectors homogeneous हैं: उस एक string ने चुपचाप सारे elements को character में coerce कर दिया, तो mean() को text मिला

R vectors एक TYPE रखते हैं। numbers और एक string के सामने, R चुपचाप सब कुछ सबसे flexible type में बदल देता है: character: "78" अब text है, और mean() शब्दों का average नहीं निकाल सकता। यह silent coercion R का सबसे चालाक beginner trap है ठीक इसलिए क्योंकि c() line ख़ुद कोई error नहीं दिखाती: class(marks) से check कीजिए। missing value का statistical fix NA है, अगली unit का topic।

Watch out

चार syntax stumbles

Case: mean() काम करता है, Mean() नहीं: R capitals कभी माफ़ नहीं करता (TRUE, true नहीं)।

Quotes: बिना quotes के Surat को एक variable name माना जाता है और error देता है।

Coercion: एक string एक numeric vector को चुपचाप text में बदल देती है।

Arrow: exams में <- लिखिए; examiners = को not-quite-R पढ़ते हैं वहाँ भी जहाँ यह चलता है।

Theory

Factors: Unit 1 की vocabulary एक type बनती है

नोटिस कीजिए factor() loop बंद कर रहा है: Unit 1 के categorical variables यहाँ literally एक data type हैं: levels() categories list करता है, और R इन्हें average करने से मना करता है (सही तरीक़े से: वह nonsense-maths trap याद कीजिए)। Numeric vectors = Unit 1 के numerical variables। जो classification आपने हाथ से सीखी वह language द्वारा enforce होती है: अगला lesson, पूरी survey files read.csv के ज़रिए आती हैं, typed और तैयार।

Summary

Key takeaways

  • Assignment: name <- value (arrow ही exam-expected form है); # comment शुरू करता है।
  • c() values को एक VECTOR में combine करता है, R की fundamental unit; अकेली values length-1 vectors हैं ([1])।
  • Operations vectorised हैं: heights + 2 और heights > 165 हर element पर लगते हैं, कोई loops नहीं।
  • Types: numeric, integer (L), character (quoted), logical (TRUE/FALSE), factor (levels वाला categorical)।
  • Vectors homogeneous हैं: एक अकेली string चुपचाप सब कुछ character में coerce कर देती है।
  • mean(condition) एक proportion compute करता है: TRUE = 1, FALSE = 0।
  • Memory hook: arrow values को box में डालता है; c() उन्हें साथ यात्रा कराता है।

Study this properly

This page is the lesson to read. In Gri-Learn the same topic is a graded deck: the self-checks are scored and your weak topics are tracked. Free to start.

Start this topic

Already have an account? Sign in

More from Introduction to R and working with Data

Gri-Learn · syllabus-mapped B.C.A. lessons in English, Hindi and Gujarati

Basic R syntax, variables, and data types · Statistical Methods and Data Analysis (MDC-03) · Gri-Learn