Theory
Leap करने से पहले Look कीजिए
मान लीजिए college आपको एक spreadsheet देती है: per student एक row, marks, attendance, projects की number, और वे placed हुए या नहीं इसके लिए columns। Placement predict करने के लिए कोई भी model बनाने से पहले, आपको actually इस data को समझना है, इसका shape, इसके quirks, इसके gaps।
वह first step exploratory data analysis (EDA) है: model करने से पहले summaries और visualisations के through एक dataset को जानना। यह पूरा subject उस placement dataset को अपने running example की तरह इस्तेमाल करता है, और यह सब यहीं शुरू होता है, आपके पास जो है उसे carefully देखने से।
Theory
EDA क्या है
Exploratory data analysis formal models apply करने से पहले एक dataset को इसकी main characteristics समझने के लिए examine करने की practice है। आप summaries compute करते हैं (averages, ranges, counts) और simple charts draw करते हैं इन questions के answer देने के लिए: हर column कौन सी values लेता है? ये कैसे distributed हैं? कौन से columns एक-दूसरे से relate करते हैं? क्या gaps या strange values हैं?
EDA एक formality नहीं है; यहीं आप problems catch करते हैं और hunches form करते हैं। इसे skip करने का मतलब है उस data को model करना जिसे आप नहीं समझते, wrong conclusions का एक recipe। Good analysts यहाँ पहले serious time spend करते हैं।
At a glance
| Type | Examine किए गए Variables | Placement-Data Example |
|---|---|---|
| Univariate | एक time पर एक | अकेले marks की distribution |
| Bivariate | दो, और इनका relationship | Marks versus student placed था या नहीं |
| Multivariate | तीन या ज़्यादा साथ में | Marks, attendance, और projects साथ में vs placement |
Theory
Missing Data और Outliers
Real datasets messy होते हैं, और दो problems constantly आती हैं।
Missing data: कुछ values बस absent हैं (एक student की attendance record नहीं हुई)। आपको decide करना पड़ता है हर gap कैसे handle करें: affected rows या columns drop कीजिए, या इन्हें impute कीजिए, इन्हें उस column के mean, median, या mode जैसे एक sensible substitute से fill कीजिए।
Outliers: कुछ values normal range से बहुत बाहर बैठती हैं (0 से 100 वाले column में 500 की एक marks entry, या एक genuine पर extreme case)। आप इन्हें detect करते हैं और decide करते हैं: क्या यह fix या remove करने के लिए एक error है, या रखने और समझने के लिए एक real extreme value है? दोनों को honestly handle करना EDA का एक core part है।
Quiz
आप students के marks और वे placed हुए या नहीं इनके relationship को examine करते हैं, exactly उन दो variables को साथ देखते हुए। यह EDA का कौन सा type है?
- Univariate, क्योंकि placement एक outcome है
- Bivariate, क्योंकि आप दो variables और इनका relationship examine कर रहे हैं
- Multivariate, क्योंकि placement सब कुछ पर depend करता है
- यह EDA नहीं है; यह modelling है
Show the answer
Bivariate, क्योंकि आप दो variables और इनका relationship examine कर रहे हैं
Exactly दो variables (marks और placement) और ये कैसे relate करते हैं यह examine करना bivariate analysis है। Option A wrong है: univariate सिर्फ़ एक variable अकेले देखता है (मान लीजिए, अकेले marks की spread), पर यहाँ आप दो को relate कर रहे हैं। Option C, multivariate, तीन या ज़्यादा variables साथ में involve करता है (उदाहरण के लिए marks, attendance, AND projects एक साथ); इस example में सिर्फ़ दो हैं। Option D wrong है: summaries और charts के through variables के बीच relationships explore करना precisely EDA है, formal modelling नहीं। आप साथ में जो variables examine करते हैं उन्हें count कीजिए: एक univariate है, दो bivariate है, तीन या ज़्यादा multivariate है।
Think first
Missing Value वाली हर Row Delete क्यों न करें?
Missing data वाली rows drop करना easy है। Imputing (values fill करना) अक्सर better क्यों है? फिर tap कीजिए।
Show the answer
क्योंकि rows drop करना बहुत सारा valuable data THROW AWAY कर सकता है और आपके results को BIAS भी कर सकता है, तो imputation अक्सर wiser choice है। Imagine कीजिए placement dataset में 500 students हैं, और इनमें से 120 के लिए attendance missing है, शायद क्योंकि एक department ने इसे अलग तरह से record किया। अगर आप किसी भी missing value वाली हर row delete करते हैं, आप शायद उन 120 students को पूरी तरह खो दें, अपना data लगभग एक चौथाई shrink करते हुए और उन rows में मौजूद बाकी सारी good information (इनके marks, projects, placement outcome) खोते हुए। और worse, अगर missingness random नहीं है, मान लीजिए attendance mostly एक department के लिए missing है, तो उन rows को drop करना quietly आपके analysis को उन departments की तरफ BIAS करता है जिन्होंने इसे actually record किया, आपके conclusions को distort करते हुए। Imputation, gaps को column के mean, median, या mode (या एक smarter estimate) जैसी एक reasonable value से fill करना, आपको हर row की बाकी useful information रखने देता है जबकि missing piece के लिए एक principled guess बनाता है। यह भी हमेशा सही नहीं होता, imputation अपनी खुद की assumptions introduce करता है, तो आप choose करते हैं कितना missing है और क्यों इसके आधार पर। पर blindly delete करना rarely best होता है। Professional habit यह समझना है data WHY missing है, फिर row-by-column decide करना drop करें या impute करें, reflexively discard करने की बजाय। जो information रख सकते हैं रखिए; gaps thoughtfully fill कीजिए।
Summary
Key takeaways
- Exploratory data analysis (EDA) first step है: model करने से पहले summaries और charts के through एक dataset को समझना।
- यह answer देता है हर column कौन सी values लेता है, ये कैसे distributed हैं, ये कैसे relate करते हैं, और problems कहाँ हैं।
- Univariate EDA एक variable examine करता है; bivariate दो और इनका relationship examine करता है; multivariate तीन या ज़्यादा साथ में examine करता है।
- Missing data (absent values) को affected rows/columns drop करके या imputing करके (mean, median, या mode से fill करके) handle किया जाता है।
- Outliers (बाकी से बहुत दूर extreme values) को detect किया जाना चाहिए और fix करने के लिए errors या रखने के लिए genuine extremes की तरह judge किया जाना चाहिए।
- Rows drop करना data खो सकता है और results bias कर सकता है, तो imputation अक्सर preferable है; पहले समझिए data क्यों missing है।
- Memory hook: EDA पहले; variables count कीजिए (uni/bi/multi), और gaps और extremes honestly handle कीजिए।