Theory
Prediction की दो Kinds
Supervised learning labelled examples से एक output predict करता है। पर 'output' का मतलब दो बहुत अलग चीज़ें हो सकती हैं। Placement data के लिए आप पूछ सकते हैं: क्या यह student placed होगा, yes या no? या आप पूछ सकते हैं: यह student कितने marks score करेगा?
पहला एक category predict करता है; दूसरा एक number predict करता है। वह distinction supervised learning को दो families में split करता है, classification और regression, और जानना आपकी problem कौन सी है बाकी सब कुछ decide करता है जो follow करता है: model, metrics, पूरा approach। यह lesson split को crisp बनाता है।
Theory
Classification vs Regression
Classification एक category predict करता है (एक discrete class label)। Output classes के एक fixed set में से एक है: placed या not placed, spam या not spam, digits 0 से 9 में से कौन सा। Answer एक label है, एक measured quantity नहीं।
Regression एक continuous number predict करता है। Output एक scale पर एक quantity है: एक student के marks, एक house की price, कल का temperature। Answer एक range में कोई भी value ले सकता है।
दोनों supervised हैं (दोनों labelled training data से सीखते हैं)। सिर्फ़ output के type में फर्क है: एक category का मतलब classification है, एक number का मतलब regression है।
At a glance
| Aspect | Classification | Regression |
|---|---|---|
| Predict करता है | एक category (class label) | एक continuous number |
| Output Example | Placed: yes या no | Marks: 78.5 |
| दूसरे Examples | Spam / not spam; disease / no disease | Price, temperature, sales |
| जो Question Answer करता है | यह किस class में belong करता है? | कितना / कितने? |
Formula
एक Question जो इसे Decide करता है
एक problem कौन सी family से belong करती है यह बताने के लिए, एक चीज़ पूछिए: क्या target एक category है या एक number?
अगर आप predict कर रहे हैं कोई चीज़ किस class में गिरती है, placed या not, pass या fail, cat या dog, यह classification है। अगर आप predict कर रहे हैं कितना या कितने, एक marks score, एक rupee amount, एक temperature, यह regression है। वह एक question, category versus quantity, किसी भी supervised problem को reliably सही family में sort कर देता है, और सबसे common beginner mix-up रोकता है।
Quiz
आप attendance और projects से एक student कितने exact marks percentage score करेगा (जैसे 78.5) यह predict करना चाहते हैं। क्या यह classification है या regression?
- Classification, क्योंकि students categories में हैं
- Regression, क्योंकि target एक continuous number है (एक marks value)
- Unsupervised, क्योंकि यह data इस्तेमाल करता है
- कोई नहीं; numbers predict करना machine learning नहीं है
Show the answer
Regression, क्योंकि target एक continuous number है (एक marks value)
एक exact marks percentage predict करना (78.5 जैसा एक continuous number) regression है, क्योंकि target एक scale पर एक quantity है, category नहीं। Option A wrong है: आप यहाँ students को fixed classes में sort नहीं कर रहे; आप एक numeric value predict कर रहे हैं, जो regression है, अगर इसके बजाय आपने 'pass/fail' या 'placed/not' predict किया होता, वह classification होता। Option C wrong है: data known marks से labelled है, तो यह supervised learning है, unsupervised नहीं। Option D wrong है: labelled examples से एक number predict करना एक core supervised ML task है (regression)। Test apply कीजिए: एक number का मतलब regression है, एक category का मतलब classification है।
Think first
Classification-vs-Regression Distinction इतना Matter क्यों करता है?
दोनों supervised learning हैं। आपकी problem कौन सी है यह जानना important क्यों है? फिर tap कीजिए।
Show the answer
क्योंकि लगभग हर downstream चीज़ इस पर depend करती है: आप जो MODELS इस्तेमाल कर सकते हैं, SUCCESS measure करने का तरीका, और आप output कैसे interpret करते हैं ये सब classification और regression के बीच अलग होते हैं। Evaluation metrics अलग होते हैं, और यह सबसे immediate consequence है। REGRESSION के लिए, आप error को predicted और actual numbers के बीच एक distance की तरह measure करते हैं, mean squared error या mean absolute error (अगला lesson) जैसे metrics इस्तेमाल करते हुए, क्योंकि '78, 80 से कितना दूर है' meaningful है। CLASSIFICATION के लिए, distance कोई sense नहीं बनाता ('placed' और 'not placed' के बीच कोई 'distance' नहीं है), तो आप accuracy, precision, और recall जैसी चीज़ें measure करते हैं, क्या predicted class सही था? एक classification problem पर regression metric इस्तेमाल करना, या इसका उलटा, nonsense देता है। MODELS भी अलग होते हैं: कई algorithms के distinct classification और regression versions होते हैं, और कुछ एक task को दूसरे से better suit करते हैं। यहाँ तक कि OUTPUT भी अलग तरह interpreted होता है: एक regression model आपको act करने के लिए एक number देता है, जबकि एक classification model आपको एक label देता है (अक्सर एक probability के साथ)। तो पहले family identify करना pedantry नहीं है; यह वह fork in the road है जो decide करता है project के बाकी हिस्से के लिए कौन से tools, metrics, और interpretations valid हैं। यह गलत करिए और इसके बाद सब कुछ sand पर built है। Category या quantity, पहले decide कीजिए, फिर correctly proceed कीजिए।
Summary
Key takeaways
- आप क्या predict करते हैं इसके आधार पर supervised learning दो families में split होता है।
- Classification एक category predict करता है (एक discrete class label): placed या not, spam या not, digit 0 से 9।
- Regression एक continuous number predict करता है: marks, price, temperature।
- दोनों supervised हैं (दोनों labelled data से सीखते हैं); सिर्फ़ output type में फर्क है।
- Deciding question: क्या target एक category है (classification) या एक number (regression)?
- Distinction matter करता है क्योंकि models, evaluation metrics, और interpretation सब इस पर depend करते हैं।
- Memory hook: category का मतलब classify करना है, quantity का मतलब regress करना है।