Subsetting and filtering data

df[rows, columns] एक data frame को दोनों तरफ़ से एक साथ slice करता है: df$screen > 180 जैसी conditions rows चुनती हैं, name vectors columns चुनते हैं, और subset() वही चीज़ पढ़ने लायक़ English में कहता है।

9 min read · 9 cards · 2 checks

Read in: English · हिन्दी · ગુજરાતી


Theory

Dean को एक shortlist चाहिए

Dean CampusPulse summary पढ़ते हैं और narrow करते हैं: "मुझे बस 180 मिनट से ऊपर screen time वाले hostellers दिखाइए: roll numbers और उनके minutes, और कुछ नहीं।"

एक row condition AND एक column selection, एक ही request में। आपने यह shape दो बार जवाब दी है: SQL का WHERE + SELECT, pandas के masks + double brackets।

R की spelling एक जोड़ी square brackets है जिसके बीच में एक comma है, और वह comma पूरी grammar ढोता है।

Theory

दो-हिस्से की order slip

df[ ___ , ___ ] को एक दो blanks वाली order slip समझिए:

  • Comma के LEFT: कौन सी rows (कौन से students)।
  • Comma के RIGHT: कौन से columns (उनके बारे में कौन से facts)।
  • एक blank empty छोड़िए और इसका मतलब है उस dimension का ALL।

df[rows, ]: चुने students, हर fact। df[, cols]: हर student, चुने facts। Comma slip की dividing line है: इसे भूलिए और kitchen order ग़लत पढ़ता है।

Practical

Bracket grammar, फिर subset()

survey <- read.csv("survey.csv")   # roll, city, stay, screen

# rows by NUMBER, all columns (empty after comma)
survey[1:5, ]

# rows by CONDITION: the workhorse
survey[survey$screen > 180, ]

# columns by NAME, all rows (empty before comma)
survey[, c("roll", "screen")]

# both at once: the dean's request
survey[survey$stay == "hostel" & survey$screen > 180,
       c("roll", "screen")]

# membership test with %in%
survey[survey$city %in% c("Surat", "Navsari"), ]

# subset(): same result, reads like English (no survey$ prefixes)
subset(survey, stay == "hostel" & screen > 180,
       select = c(roll, screen))

# keep a result by assigning it
heavy <- subset(survey, screen > 180)
nrow(heavy)

Theory

Conditions: वही logic, R में spelled

Row conditions logical vectors हैं (syntax lesson का survey$screen > 180) row-pickers के रूप में इस्तेमाल किए गए:

  • == equality test करता है (अकेला = assign करता है: वह classic slip)।
  • & and, | or, ! not: element-wise, हर side अपनी पूरी comparison में: survey$screen > 180 & survey$stay == "hostel"।
  • %in% एक set में membership test करता है: R का IN।

Translation table: rows-condition = SQL WHERE = pandas mask; columns blank / select = SQL की SELECT list। तीसरी language, वही दो moves।

Quiz

एक student survey[survey$screen > 180] type करता है (कोई comma नहीं) और heavy screen-time rows expect करता है। R ने क्या समझा?

  1. Comma के बिना, R इसे COLUMN selection मानता है, row filtering नहीं: दो-blank slip को इसका comma चाहिए: survey[survey$screen > 180, ]
  2. यह बिल्कुल वैसा ही काम करता है: comma optional style है
  3. यह matching rows delete कर देता है
  4. R data frame को screen time से sort कर देता है
Show the answer

Comma के बिना, R इसे COLUMN selection मानता है, row filtering नहीं: दो-blank slip को इसका comma चाहिए: survey[survey$screen > 180, ]

एक data frame पर single-argument brackets columns select करते हैं, तो R logical vector को एक column-picker के रूप में इस्तेमाल करने की कोशिश करता है: ग़लत columns या एक error, इच्छित rows कभी नहीं। Row/column grammar पूरी तरह उस comma में रहती है: left = rows, right = columns, empty = all। df[condition, ] trailing comma-space के साथ लिखना वह आदत है जो यह bug असंभव बनाती है।

Think first

Dean को तीन तरीक़ों से translate कीजिए

Dean की request (screen > 180 वाले hostellers; सिर्फ़ roll और screen) को लिखिए: (1) R brackets में, (2) R subset() में, (3) वह SQL जो आपने BCA303 में सीखा। tap करने से पहले तीनों sketch कीजिए।

Show the answer

1. survey[survey$stay == "hostel" & survey$screen > 180, c("roll", "screen")]

2. subset(survey, stay == "hostel" & screen > 180, select = c(roll, screen))

3. SELECT roll, screen FROM survey WHERE stay = 'hostel' AND screen > 180;

तीन spellings, एक thought: rows filter कीजिए, columns चुनिए। अगर तीनों अलग accents में एक ही sentence जैसे लगें, concept उतर चुका है: interviews वही recognition test करते हैं।

Watch out

Filtering की चार slips

Missing comma: df[condition, ]: left blank rows, right blank columns, हमेशा।

= बनाम ==: conditions == से compare करते हैं; brackets के अंदर अकेला = एक error या accident है।

'and' मौजूद नहीं है: R को & और | चाहिए, एक ampersand (&& single values के लिए है, vectors के लिए नहीं)।

Unassigned results ग़ायब हो जाते हैं: filtering print करता है और भूल जाता है: इसे heavy <- subset(...) से रखिए।

Theory

Filtering आगे सब कुछ feed करता है

अब से आप जो भी statistic compute करते हैं वह असल में "statistic OF a subset" है: hostellers OF mean screen time, year 3 OF SD, हर city OF boxplot। Slicing statistics का preposition है। अगले lessons: columns जोड़ना और rename करना (survey derived variables पाता है), फिर cleaning: जहाँ आपकी conditions ने चुपचाप छोड़े NA values आख़िरकार justice का सामना करते हैं।

Summary

Key takeaways

  • df[rows, cols]: comma के left rows चुनता है, right columns चुनता है, EMPTY मतलब all।
  • Rows number से (1:5), condition से (df$screen > 180), columns name vector से।
  • Conditions को & | ! से combine कीजिए, == से compare कीजिए, %in% से membership।
  • subset(df, condition, select = cols) वही बात बिना df$ prefixes के कहता है।
  • Pattern = SQL WHERE + SELECT = pandas mask + column list: तीसरी language, वही idea।
  • Filtered results assign कीजिए वरना वे print होकर ग़ायब हो जाते हैं।
  • Memory hook: दो-blank order slip, अपने comma से divided।

Study this properly

This page is the lesson to read. In Gri-Learn the same topic is a graded deck: the self-checks are scored and your weak topics are tracked. Free to start.

Start this topic

Already have an account? Sign in

More from Data Filtering and cleaning

Gri-Learn · syllabus-mapped B.C.A. lessons in English, Hindi and Gujarati

Subsetting and filtering data · Statistical Methods and Data Analysis (MDC-03) · Gri-Learn