Theory
Column आख़िरकार fix होने के लिए तैयार है
Import lesson का घायल column याद कीजिए: एक "absent" entry ने हर screen-time number को text बना दिया था। आपने इसे str() से diagnose किया; आज इसे repair कीजिए:
survey$screen <- as.numeric(survey$screen)
Warning: NAs introduced by coercion
Digits numbers बन गए; "absent" NA बन गया (बिल्कुल जहाँ पिछले lesson की skills संभाल लेती हैं)। एक line, एक honest warning। Conversion इतना ही आसान है: सिवाय एक famous trap के जो categories को ग़लत numbers में बदल देता है।
Theory
एक strict teller के साथ currency exchange
as.___() family एक currency exchange है: as.numeric number-currency में convert करता है, as.character text-currency में, as.factor category-currency में।
Teller strict है: "120" साफ़ 120 में exchange होता है, पर "absent" का कोई number equivalent नहीं है: teller NA वापस देता है और इसे ज़ोर से कहता है (coercion warning)। announcement को कभी ignore मत कीजिए: यह exchange में fail हुई values की आपकी count है।
Practical
Convert कीजिए, repair कीजिए, recode कीजिए
# CONVERSION: the as.___() family
raw <- c("120", "90", "absent", "240")
screen <- as.numeric(raw) # Warning: NAs introduced by coercion
screen # [1] 120 90 NA 240 repaired!
class(screen) # "numeric": always verify after converting
# make a deliberate categorical (Unit 1's classification, executable):
stay <- as.factor(c("hostel", "day", "hostel"))
levels(stay) # "day" "hostel"
# RECODING: ifelse(condition, value_if_yes, value_if_no): vectorised
usage <- ifelse(screen > 180, "heavy", "normal")
usage # "normal" "normal" NA "heavy"
# three bands: nest the no-branch
band <- ifelse(screen > 180, "heavy",
ifelse(screen > 90, "medium", "light"))
# binning shortcut worth knowing: cut()
# cut(screen, breaks = c(0, 90, 180, 1440),
# labels = c("light", "medium", "heavy"))
Theory
THE trap: factors सीधे convert नहीं होते
एक factor labels को internal level codes (1, 2, 3... alphabetical order में) के रूप में store करता है। इसे सीधे numeric में convert कीजिए और आपको codes मिलते हैं, values नहीं:
f <- factor(c("120", "90"))
as.numeric(f)
[1] 2 1
120 और 90 नहीं: 2 और 1 ("120" alphabetically "90" से पहले sort होता है!)। सही route two-step detour है:
as.numeric(as.character(f))
[1] 120 90
पहले labels, फिर numbers। यह R का सबसे famous beginner bug है और एक guaranteed exam question।
Quiz
marks <- factor(c("85", "100", "92"))। एक student as.numeric(marks) चलाता है और 2 1 3 पाता है। क्यों, और fix क्या है?
- Factors अपने internal level codes में convert होते हैं (alphabetical: "100" < "85" < "92"); fix: as.numeric(as.character(marks))
- R 99 से ऊपर marks store नहीं कर सकता, तो इसने उन्हें renumber किया
- CSV में values से पहले quotes हटाने की ज़रूरत थी
- as.numeric सिर्फ़ integers पर काम करता है
Show the answer
Factors अपने internal level codes में convert होते हैं (alphabetical: "100" < "85" < "92"); fix: as.numeric(as.character(marks))
Alphabetically "100" पहले आता है (character comparison, digit by digit), तो levels हैं 100=1, 85=2, 92=3, और as.numeric वे codes return करता है। Two-step as.numeric(as.character(...)) असली values recover करता है: पहले labels बाहर, फिर numbers। Options B और D limits invent करते हैं जो R के पास नहीं हैं; C prevention में मदद करेगा पर mechanism explain नहीं करता, जो exam पूछता है।
Think first
Survey को ख़ुद recode कीजिए
screen में 120, 90, NA, 240 है। अंदाज़ा लगाइए usage <- ifelse(screen > 180, "heavy", "normal") हर ELEMENT के लिए क्या produce करता है (तीसरे से सावधान रहें), फिर tap कीजिए।
Show the answer
"normal" "normal" NA "heavy"
Element by element: 120 और 90 condition fail करते हैं → "normal"; 240 pass करता है → "heavy"; और NA > 180 है... NA (पिछले lesson की comparison blindness), तो ifelse NA को through pass कर देता है: unknown अंदर, unknown बाहर।
ifelse vectorised है (पूरा column, कोई loop नहीं) और NA-honest है: exams दोनों properties check करते हैं। तीन या ज़्यादा bands के लिए, ifelse nest कीजिए या cut() इस्तेमाल कीजिए।
Watch out
Conversion checklist
1. Coercion warning सुनिए: यह आपके नए NAs गिनती है: sum(is.na(...)) से follow up कीजिए।
2. Factor to number = दो steps: as.numeric(as.character(f)): कभी direct नहीं।
3. पहले और बाद में class(): वे conversions जिन्होंने चुपचाप कुछ नहीं किया (या ग़लत किया) एक नज़र में पकड़े जाते हैं।
4. Consistently recode कीजिए: "heavy"/"Heavy"/"HEAVY" वही text mess फिर बनाएगा जो आपने दो lessons पहले साफ़ किया था।
Theory
Unit 4 बंद होती है: data आख़िरकार trustworthy है
उस pipeline को गिनिए जिसे आप अब end to end चलाते हैं: import, audit, subset, add/rename, deduplicate, standardise, range-check, NA-handle, convert, recode। यही असली professional data preparation है: असली काम का 60-80%। Unit 5 इनाम है: sorting, merging, summary commands, frequency tables, और R में खींचा bell curve: उस data पर statistics जिसे आप आख़िरकार defend कर सकते हैं।
Summary
Key takeaways
- as.numeric/as.character/as.factor types convert करते हैं; unconvertible values एक warning के साथ NA बनते हैं: उन्हें गिनिए।
- Coercion warning text-polluted numeric columns के लिए repair receipt है।
- FACTOR TRAP: as.numeric(factor) level codes देता है; fix है as.numeric(as.character(f))।
- ifelse(condition, yes, no) vectorised recode करता है; NAs NA के रूप में through pass होते हैं; bands के लिए nest कीजिए या cut() इस्तेमाल कीजिए।
- as.factor से categoricals को जानबूझकर promote कीजिए: levels() और table() जीवंत हो जाते हैं।
- हर conversion से पहले और बाद में class()।
- Memory hook: strict teller हर fail हुई exchange को ज़ोर से announce करता है।