Theory
सारे Columns एक जैसी Kind के नहीं होते
Placement dataset में, 'marks' एक number है जिसे आप average कर सकते हैं, 'department' एक नाम है जिसे आप नहीं कर सकते, और 'grade' (A, B, C) बीच में कहीं है, ordered, पर एक measurement नहीं। इन सबको same तरह treat करना nonsense की तरफ ले जाता है, जैसे 'average department' compute करना।
तो analyse करने से पहले, आपको हर column की data type जाननी पड़ती है। Types एक small family बनाती हैं: quantitative (numeric) discrete और continuous में split होती है, और qualitative (categorical) nominal और ordinal में split होती है। कौन सी कौन है यह जानना आपको बताता है कौन सा analysis बिल्कुल भी valid है।
At a glance
| Type | Meaning | Placement-Data Example |
|---|---|---|
| Quantitative - discrete | Countable whole numbers | Projects की number (0, 1, 2, 3) |
| Quantitative - continuous | एक range में कोई भी value (measured) | Marks percentage (जैसे 78.5) |
| Qualitative - nominal | बिना order की categories | Department; placed (yes/no) |
| Qualitative - ordinal | एक meaningful order वाली categories | Grade (A, B, C); skill level (low/medium/high) |
Theory
Quantitative: Discrete vs Continuous
Quantitative data numbers हैं जिन्हें आप meaningfully measure कर सकते हैं और इन पर arithmetic कर सकते हैं। ये दो में split होते हैं।
Discrete data countable whole values हैं, वे चीज़ें जो आप count करते हैं: एक student ने कितने projects किए इसकी number 0, 1, 2, 3 है, कभी 2.5 नहीं। Continuous data एक range में कोई भी value ले सकता है, वे चीज़ें जो आप measure करते हैं: एक marks percentage 78, 78.5, 78.53 हो सकता है, सिर्फ़ precision से limited।
एक quick test: क्या इसमें sensibly decimals और in-between values हो सकते हैं? अगर हाँ, यह continuous है (marks, height, time); अगर यह सिर्फ़ whole steps में jump करता है जिन्हें आप count करते हैं, यह discrete है (projects, attempts की number)।
Theory
Qualitative: Nominal vs Ordinal
Qualitative (categorical) data categories हैं, measurements नहीं। ये भी दो में split होते हैं, और difference order है।
Nominal categories में कोई inherent order नहीं होता: department names (Computer, Commerce, Science) या एक yes/no placement flag, इनमें से कोई दूसरे से 'higher' नहीं है। Ordinal categories में एक meaningful order होता है पर measurable gaps नहीं: A का grade B से better है B, C से better है, और satisfaction low/medium/high ranked है, फिर भी आप नहीं कह सकते A minus B एक fixed amount के बराबर है।
Test: क्या categories में एक natural ranking है? Ranked का मतलब ordinal है; unranked का मतलब nominal है।
Quiz
Placement dataset में, column 'grade' values A, B, और C लेता है, जहाँ A, B से better है B, C से better है। यह कौन सी data type है?
- Quantitative continuous, क्योंकि grades marks जैसे हैं
- Qualitative ordinal, क्योंकि categories में एक meaningful order है पर ये measured numbers नहीं हैं
- Qualitative nominal, क्योंकि grades सिर्फ़ बिना order के labels हैं
- Quantitative discrete, क्योंकि तीन grades हैं
Show the answer
Qualitative ordinal, क्योंकि categories में एक meaningful order है पर ये measured numbers नहीं हैं
Grade (A, B, C) qualitative ordinal है: values categories हैं जिनकी clear ranking है (A, B से better है B, C से better है), पर ये fixed gaps वाले measured numbers नहीं हैं। Option A wrong है: grades labels हैं, continuous measurements नहीं, आप नहीं कह सकते A 85.3 के बराबर है; वह underlying marks होगा, एक अलग column। Option C ordering को miss करता है: department names के unlike, grades में DOES एक meaningful order होता है, तो ये ordinal हैं, nominal नहीं। Option D wrong है: discrete quantitative का मतलब countable numbers है (जैसे projects की number), जबकि A/B/C ranked categories हैं, counts नहीं। Deciding question: क्या ये ordered categories हैं (ordinal) या unordered (nominal), या actual numbers (quantitative)?
Think first
Analysis के लिए Data Type क्यों Matter करती है?
हर column classify करने की तकलीफ़ क्यों उठाएँ? Data types ignore करने पर क्या गलत होता है? फिर tap कीजिए।
Show the answer
क्योंकि data type decide करती है कौन से OPERATIONS, SUMMARIES, और CHARTS valid हैं, और wrong एक apply करना meaningless या misleading results produce करता है। Average सोचिए। Marks जैसे एक quantitative column का MEAN लेना बिल्कुल sensible है ('average mark 72 है')। पर department जैसे एक NOMINAL column का 'mean' लेना nonsense है, 'Computer, Commerce, Science' का कोई average नहीं होता; जो numbers आप उन labels को assign कर सकते हैं वे arbitrary हैं, तो इनका average कुछ मतलब नहीं रखता। Ordinal data subtler है: आप एक median grade (middle rank) पा सकते हैं पर A, B, C को इस तरह average करना जैसे gaps equal हों dubious है। Type सही VISUALISATION भी dictate करती है: एक histogram continuous data को suit करता है, counts का एक bar chart categorical data को suit करता है, और एक scatter plot दो quantitative variables को suit करता है। यह यहाँ तक guide करता है कि एक machine-learning model को एक column को कैसे treat करना चाहिए, numeric features और categorical features अलग तरह से handle होते हैं (categories को अक्सर encoding चाहिए)। तो हर column classify करना busywork नहीं है; यही वह है जो आपके analysis को HONEST रखता है, ensure करते हुए हर average, chart, और model input actually उस kind के data के लिए sense बनाता है जिसे यह describe करता है। Type जानिए, और आप जानते हैं इसके साथ आपको क्या करने की permission है।
Summary
Key takeaways
- हर column की एक data type होती है जो decide करती है कौन से analyses और charts valid हैं।
- Quantitative data measurable numbers हैं; qualitative data categories हैं।
- Quantitative discrete (countable whole values, जैसे projects की number) और continuous (एक range में कोई भी value, जैसे marks) में split होता है।
- Qualitative nominal (unordered categories, जैसे department) और ordinal (ordered categories, जैसे grade A/B/C) में split होता है।
- Quantitative test कीजिए: क्या इसमें sensibly in-between decimal values हो सकते हैं (continuous) या सिर्फ़ whole counts (discrete)?
- Qualitative test कीजिए: क्या categories में एक natural ranking है (ordinal) या नहीं (nominal)?
- Memory hook: quantitative counts या measures करता है; qualitative labels करता है, ranked (ordinal) या नहीं (nominal)।