Theory
एक और language? आपने अभी pandas सीखी!
निष्पक्ष आपत्ति। BCA303 में आपने CSVs load किए, means calculate किए, histograms खींचे: सब Python में। अब यह subject आपको R थमाता है यह करने के लिए... statistics?
ईमानदार बात यह है: Python एक general-purpose language है जिसने libraries के ज़रिए statistics सीखी। R statistical पैदा हुई थी: statisticians द्वारा, statisticians के लिए बनी, जहाँ एक t-test उतना ही native है जितना एक for-loop।
जब दुनिया के statisticians methods publish करते हैं, R को वे पहले मिलते हैं। आपका syllabus चाहता है आप field की मातृभाषा में fluent हों।
Theory
General hospital बनाम speciality clinic
Python एक general hospital है: यह सब कुछ करता है: websites, games, databases, और (pandas के साथ) एक wing में data analysis।
R speciality clinic है: यह एक चीज़ करता है, statistics और इसके graphics, और पूरी इमारत इसके चारों तरफ़ व्यवस्थित है: हर instrument हाथ की पहुँच में, हर नई technique वहाँ पहले पहुँचती है।
Professionals दोनों जाते हैं। यह subject आपको clinic से होकर ले जाता है।
Theory
R क्या है, औपचारिक रूप से
R एक free, open-source statistical computing और graphics के लिए language और environment है, Ross Ihaka और Robert Gentleman (University of Auckland, 1990s) द्वारा बनाई गई पहले की S language के एक open implementation के रूप में।
इसका statistical DNA:
- mean(), median(), sd(), और पूरे tests (t.test(), lm()) built in हैं, कोई imports नहीं।
- data frame एक core language type है (pandas ने idea और नाम दोनों उधार लिए)।
- Publication-quality graphics box में आते हैं।
- यह interpreted और interactive है: एक command type कीजिए, जवाब देखिए: exploration के लिए बनी।
Theory
CRAN और Package Universe
Base R clinic का ground floor है। CRAN (Comprehensive R Archive Network) ऊपर की हर चीज़ है: 20,000+ peer-contributed packages, एक command दूर:
install.packages("ggplot2")
नाम से पहचानने लायक़ Stars: ggplot2 (famous grammar-of-graphics plots), dplyr (data manipulation), और Bioconductor collection (genomics)। Pattern Python के pip को mirror करता है: base language + package ecosystem: आप पहले से जानते हैं यह दुनिया कैसे काम करती है।
At a glance
R पर कौन चलता है
| Field | काम पर R |
|---|---|
| Academic research | Statistics papers की default language |
| Pharma / Clinical Trials | Drug-trial analysis और regulatory submissions |
| Epidemiology | Disease models, COVID dashboards |
| Data Journalism | BBC और FT charts R-made हैं |
| Finance | Risk models, time-series forecasting |
| Bioinformatics | Gene-expression analysis (Bioconductor) |
Quiz
कौन सा statement R और Python-with-pandas के बीच honest अंतर सबसे अच्छे से capture करता है?
- Python general-purpose है analysis libraries के ज़रिए जोड़ी गई; R statistics-first है tests, data frames, और graphics language के लिए native के साथ
- R paid और proprietary है जबकि Python free है
- R Python को replace करता है: pandas obsolete है
- R CSV files पढ़ नहीं सकता, इसलिए data import के लिए Python चाहिए
Show the answer
Python general-purpose है analysis libraries के ज़रिए जोड़ी गई; R statistics-first है tests, data frames, और graphics language के लिए native के साथ
दोनों free और open source हैं (option B ग़लत है), दोनों ख़ुशी से CSVs पढ़ते हैं (D ग़लत है: read.csv R में built-in है), और कोई भी दूसरे को replace नहीं करता (C): असली analysts दोनों इस्तेमाल करते हैं, task के हिसाब से चुनते हुए। असली अंतर center of gravity है: Python general programming के चारों तरफ़ घूमता है, R statistics के चारों तरफ़: यही कारण है आपका statistics syllabus इसे पढ़ाता है, और यही कारण है answer A exams में लिखने लायक़ line है।
Think first
One-liners अंदाज़ा लगाइए
BCA303 में आपने survey.csv load करने और mean screen time पाने के लिए कई pandas lines लिखीं। tap करने से पहले: अंदाज़ा लगाइए R इन दो steps को कैसे spell करता है: एक CSV load करना, और एक mean लेना। (Hint: R चीज़ों को सादगी से नाम देता है।)
Show the answer
data <- read.csv("survey.csv")
mean(data$screen_time)
दो lines, शून्य imports: files पढ़ना और average निकालना language में ही built in हैं। (वह <- arrow R का assignment operator है, और $ एक column चुनता है: दोनों दो lessons आगे ठीक से explain किए गए।) आपके pandas instincts लगभग एक-से-एक transfer होते हैं: R बस उन्हें statistical काम के लिए छोटे तरीक़े से spell करता है।
Watch out
दो mislabels से बचना
"R बस academic है": pharma submissions, BBC graphics, और bank risk models असहमत हैं: applications table असली industry है।
"R" बनाम "RStudio": R language है (engine); RStudio वह IDE है जिससे आप इसे चलाते हैं (अगला lesson दोनों install करता है)। "मैंने इसे RStudio में code किया" कहना ऐसा है जैसे कहना आपने Python "VS Code में" लिखा: language का नाम लीजिए, editor का नहीं।
Theory
इस unit के आगे की सड़क
Units 3-5 का plan: R और RStudio install कीजिए, syntax सीखिए (वह arrow, vectors, types), CampusPulse survey.csv load कीजिए, फिर Units 1-2 का सब कुछ R में दोबारा कीजिए: summaries, frequency tables, और अंत में dnorm() से खींचा bell curve। जो हर statistic आपने हाथ से calculate किया वह एक command बन जाता है: hand-by-hand काम ही है जो आपको command के जवाब पर TRUST करने देता है।
Summary
Key takeaways
- R = statistical computing और graphics के लिए free, open-source language (Ihaka और Gentleman, S से)।
- Statistics native है: mean/sd/t.test/data frames/graphics built in, कोई imports नहीं।
- CRAN install.packages() के ज़रिए 20,000+ packages host करता है; ggplot2 और dplyr stars हैं।
- Research, pharma, epidemiology, journalism, finance, bioinformatics भर इस्तेमाल होता है।
- Python = general-purpose + libraries; R = statistics-first: दोनों free, professionals दोनों इस्तेमाल करते हैं।
- R language है; RStudio IDE है।
- Memory hook: general hospital के बगल में speciality clinic।