Theory
Survey को एक ऐसा column चाहिए जो कभी माँगा नहीं गया
Dean को shortlist पसंद आई, फिर उन्होंने वह अनिवार्य सवाल पूछा: "क्या मैं screen time hours में देख सकता हूँ, minutes में नहीं? और print करते समय roll numbers हटा दीजिए।"
CSV में hours column नहीं है। कोई 60 से divide करने के लिए 60 students को दोबारा survey नहीं करता।
एक data frame किसी file की जमी हुई photo नहीं है: यह memory में मिट्टी है। आज के तीन verbs: पुराने columns से नए sculpt कीजिए, columns काट दीजिए, और जो बुरी तरह labelled है उसे rename कीजिए।
Theory
Register pencil में बढ़ता है
June में print हुआ एक class register अभी भी पूरे term pencil columns पाता रहता है: एक teacher "attendance %" जोड़ते हैं, मौजूदा columns से compute किया गया; एक column काट देते हैं जिसे कोई इस्तेमाल नहीं करता; एक cryptic heading को relabel करते हैं।
Print हुआ original (disk पर survey.csv) कभी नहीं बदलता: pencil का काम आपकी COPY पर रहता है (memory में data frame), जब तक आप इसे write.csv से एक नई file में photocopy नहीं करते।
Practical
जोड़ना, हटाना, rename करना
survey <- read.csv("survey.csv") # roll, city, stay, screen
# ADD: derived from existing columns (vectorised: whole column at once)
survey$hours <- survey$screen / 60
# ADD: a logical flag column
survey$heavy <- survey$screen > 180
# ADD: a constant (recycled to every row)
survey$batch <- 2026
# REMOVE: immediate, no confirmation
survey$batch <- NULL
# RENAME by matching the old name (safe, position-independent)
names(survey)[names(survey) == "screen"] <- "screen_min"
# RENAME by position (fragile: breaks if order changes)
# names(survey)[4] <- "screen_min"
names(survey) # "roll" "city" "stay" "screen_min" "hours" "heavy"
# memory only! persist deliberately:
write.csv(survey, "survey_v2.csv", row.names = FALSE)
Theory
Derived columns असली skill क्यों हैं
एक constant जोड़ना clerical है। Analyst का move derived column है: मौजूदा columns से row-wise compute किया एक नया variable:
- unit conversions:
hours <- screen / 60 - flags:
heavy <- screen > 180 - combinations:
per_class <- marks / classes
हर एक vectorised line है: कोई loop नहीं, हर row एक साथ। Statistics में इसे transformation कहते हैं, और अगले lesson की "data cleaning" का आधा हिस्सा बिल्कुल यही है: raw columns को analysable columns में reshape करना।
Quiz
एक student survey$city <- NULL चलाता है, किसी दूसरे column को हटाने के इरादे से। अब स्थिति क्या है?
- city data frame से तुरंत चला जाता है; survey.csv से दोबारा import करके इसे recover कीजिए (file untouched है)
- R ने delete करने से पहले confirmation माँगा, तो कुछ नहीं हुआ
- Column RStudio के अंदर एक recycle bin में है
- Disk पर CSV file ने भी अपना city column खोया
Show the answer
city data frame से तुरंत चला जाता है; survey.csv से दोबारा import करके इसे recover कीजिए (file untouched है)
NULL-deletion instant और silent है: कोई prompt नहीं, session के अंदर कोई undo नहीं। बचाव pencil-copy model है: data frame एक memory copy है, तो disk पर file का हर column अभी भी है: survey <- read.csv("survey.csv") original restore करता है (आपके बाक़ी pencil काम को छोड़कर, उसे दोबारा कीजिए)। Options B और C ऐसे safety nets describe करते हैं जो R के पास नहीं हैं; D memory को disk से confuse करता है: सिर्फ़ write.csv file को touch करता है।
Think first
वह typo जो edit की जगह create करता है
एक student screen column को IN PLACE hours में convert करना चाहता है और type करता है: survey$scren <- survey$screen / 60 (spelling नोट कीजिए)। कोई error नहीं आती। tap करने से पहले: data frame अब कैसा दिखता है, और यह silence ख़तरनाक क्यों है?
Show the answer
R ने hours रखता एक बिल्कुल नया column scren नाम का बनाया, जबकि screen unchanged रहता है: किसी unknown column name को assignment का मतलब है "इसे जोड़ो", कभी "क्या आपका मतलब..." नहीं।
ख़तरा silence है: अभी कोई error नहीं, बाद में confusion: दो similar columns, और $screen इस्तेमाल करता कोई भी analysis अभी भी minutes पर चलता है। आदत: column surgery के बाद, names(df) या str(df) पर एक नज़र डालिए: दस-सेकंड का audit इस तरह का हर typo पकड़ता है।
Watch out
Column-surgery के तीन rules
NAME match से rename कीजिए, position से नहीं: names(df)[names(df) == "old"] <- "new" column reordering में भी टिकता है; [4] नहीं।
NULL forever है (in-session): अनिश्चित हों तो एक copy बनाना पसंद कीजिए (df2 <- df[, keep])।
Surgery के बाद names() check कीजिए: typos चुपचाप columns बनाते हैं: one-glance audit।
Theory
Attributes, exam word
Syllabi columns को variables या attributes कहते हैं (BCA303 का database word: वही चीज़)। तो "data frame में एक attribute जोड़िए" जैसा exam phrasing बिल्कुल आज की df$new <- ... line है। अगला lesson single columns से पूरी cleaning workflow तक zoom out करता है: duplicates, stray text, impossible values: एक raw survey को dean के सामने defend करने लायक़ data में बदलना।
Summary
Key takeaways
- जोड़िए: df$new <- expression (vectorised, अक्सर दूसरे columns से derived); constants recycle होते हैं।
- हटाइए: df$col <- NULL: instant, silent, session-permanent (disk file untouched है)।
- Rename कीजिए: names(df)[names(df) == "old"] <- "new": name-matched position को हराता है।
- Derived columns (conversions, flags) analyst का core transformation move हैं।
- $assignment में typos चुपचाप नए columns CREATE करते हैं: surgery के बाद names()/str() से audit कीजिए।
- सभी edits memory में रहते हैं जब तक write.csv उन्हें persist न करे।
- Memory hook: register pencil में बढ़ता है; disk पर printout साफ़ रहता है।