Loss functions: Mean Squared Error (MSE) - definition, computing and properties; Mean Absolute Error (MAE) - definition, computing and properties; understanding Regression and R-squared

एक regression model कितना गलत है यह measure करने के लिए, आप इसके errors compute करते हैं: mean absolute error इनके sizes average करता है, mean squared error इनके squares average करता है तो बड़ी mistakes ज़्यादा hurt करें, और R-squared बताता है model variation का कितना fraction explain करता है।

12 min read · 8 cards · 2 checks

Read in: English · हिन्दी · ગુજરાતી


Theory

Model कितना Wrong है?

आपने marks predict करने के लिए एक regression model बनाया। यह predictions बनाता है; real marks known हैं। तो: यह कितना wrong है? एक model improve करने के लिए आपको पहले इसकी error को एक number से measure करना पड़ता है, और उस number को एक loss (या error) function कहते हैं।

यह lesson वे तीन cover करता है जो आप regression के लिए सबसे ज़्यादा इस्तेमाल करेंगे: mean absolute error (MAE), mean squared error (MSE), और R-squared। हम एक छोटे, worked example पर तीनों compute करेंगे तो definitions concrete बन जाएँ। Lower error better है; higher R-squared better है।

Theory

Worked Example, और MAE

चार predictions लीजिए। Actual marks हैं 3, 5, 7, 9; model 4, 5, 7, 6 predict करता है। हर एक का error actual minus predicted है: -1, 0, 0, 3।

Mean Absolute Error (MAE) error sizes (absolute values, sign ignore करते हुए) का average है:

MAE = (|-1| + |0| + |0| + |3|) / 4 = (1 + 0 + 0 + 3) / 4 = 4 / 4 = 1

तो average में model 1 mark से off है। MAE target की same units में है और interpret करना easy है: यह बस miss का typical size है।

Theory

MSE, और यह बड़ी Errors को क्यों Punish करता है

Mean Squared Error (MSE) errors के squares को average करता है:

MSE = ((-1)squared + 0squared + 0squared + 3squared) / 4 = (1 + 0 + 0 + 9) / 4 = 10 / 4 = 2.5

नोटिस कीजिए MSE (2.5) same predictions के लिए MAE (1) से बड़ा है। इसकी वजह 3 का error है: इसे square करने पर 9 मिलता है, जो average को dominate करता है। यही MSE की key property है, squaring करके, यह छोटी errors से बड़ी errors को कहीं ज़्यादा heavily penalise करता है। तो MSE वह metric है जो तब इस्तेमाल करना चाहिए जब बड़ी mistakes especially bad हों और आप चाहते हैं model इन्हें avoid करे। (इसका downside: squaring इसे outliers के लिए भी ज़्यादा sensitive बनाता है।)

Practical

MAE और MSE Compute करना (verified)

actual    = [3, 5, 7, 9]
predicted = [4, 5, 7, 6]
n = len(actual)

errors = [a - p for a, p in zip(actual, predicted)]   # [-1, 0, 0, 3]

mae = sum(abs(e) for e in errors) / n     # (1+0+0+3)/4 = 1.0
mse = sum(e*e   for e in errors) / n      # (1+0+0+9)/4 = 2.5

print(mae)   # 1.0
print(mse)   # 2.5  -> larger, because the error of 3 squares to 9

This example runs in Gri-Learn on the web, where you can edit it and see the output.

Theory

R-squared: कितना Explained है

R-squared एक अलग question answer देता है: target में variation का कौन सा fraction model explain करता है? यह model की errors को एक simple baseline (बस mean predict करना) के around variation से compare करता है।

R-squared = 1 - (SS_res / SS_tot), जहाँ SS_res squared errors का sum है और SS_tot mean से total squared variation है। हमारे example के लिए: actuals का mean 6 है, तो SS_tot = (3-6)squared + (5-6)squared + (7-6)squared + (9-6)squared = 9 + 1 + 1 + 9 = 20, और SS_res = 1 + 0 + 0 + 9 = 10। तो R-squared = 1 - 10/20 = 0.5।

0.5 का एक R-squared का मतलब है model half variation explain करता है। 1 के near excellent है; 0 का मतलब है बस mean guess करने से better नहीं है।

Quiz

Errors [-1, 0, 0, 3] के लिए, MAE 1 है पर MSE 2.5 है। यहाँ MSE, MAE से बड़ा क्यों है?

  1. क्योंकि MSE, MAE से ज़्यादा data points इस्तेमाल करता है
  2. क्योंकि MSE हर error square करता है, तो 3 का error 9 बन जाता है, बड़ी errors को ज़्यादा heavily penalise करते हुए
  3. क्योंकि MAE 3 के error को पूरी तरह ignore करता है
  4. यह एक mistake है; MSE हमेशा MAE के बराबर होना चाहिए
Show the answer

क्योंकि MSE हर error square करता है, तो 3 का error 9 बन जाता है, बड़ी errors को ज़्यादा heavily penalise करते हुए

MSE average करने से पहले हर error square करता है, तो largest error dominate करता है: 3 का error 3 squared = 9 बन जाता है, जो MSE (2.5) को MAE (1) से ऊपर push करता है। Option A wrong है: दोनों same चार data points इस्तेमाल करते हैं; difference squaring versus absolute value का है, count का नहीं। Option C wrong है: MAE में 3 का error include होता है, यह sum में अपना absolute value 3 contribute करता है (1+0+0+3=4, 4 से divide करने पर 1 है); यह बस इसे square नहीं करता। Option D wrong है: MSE और MAE design से generally अलग होते हैं, MSE MAE के बराबर सिर्फ़ special cases में होता है (जैसे सारी errors 0 या 1 होना)। Key property: squaring MSE को MAE से कहीं ज़्यादा बड़ी errors penalise कराता है।

Think first

MSE की बजाय MAE कब Prefer करेंगे, या Reverse?

दोनों regression error measure करते हैं। Squaring (MSE) right choice कब है, और MAE कब better है? फिर tap कीजिए।

Show the answer

आप बड़ी errors को कितना PUNISH करना चाहते हैं और OUTLIERS के बारे में आप कैसा feel करते हैं इसके आधार पर choose कीजिए। MSE हर error square करता है, तो कुछ बड़ी misses score को dominate करती हैं; यह exactly तब चाहिए जब बड़ी errors disproportionately costly हों, अगर 10 से off होना 5 से off होने से सिर्फ़ दुगुना नहीं बल्कि कहीं ज़्यादा worse है, MSE की heavy penalty model को बड़ी mistakes avoid करने के लिए hard push करती है। वह sensitivity इसकी weakness भी है: एक single wild outlier MSE को blow up कर सकता है और model को उस एक point के पीछे भाग खींच सकता है। MAE हर error को इसके size के proportion में treat करता है (squaring नहीं), तो यह outliers के लिए ज़्यादा ROBUST है, एक freak error इसे सिर्फ़ उस error के size से move करता है, इसके square से नहीं। तो MAE preferable है जब आपके data में outliers हैं जिन्हें आप fit dominate नहीं करने देना चाहते, या जब आप बस target की अपनी units में एक error चाहते हैं जो 'typical miss' की तरह पढ़ी जाए। MSE preferable है जब बड़ी errors genuinely डरने लायक हों और आपका data reasonably clean हो, और इसकी nice mathematical properties भी हैं जिनका training के दौरान कई algorithms exploit करते हैं। Practice में analysts अक्सर दोनों देखते हैं: एक interpretable typical-error figure के लिए MAE, जब बड़ी errors सबसे ज़्यादा matter करती हैं तब MSE (या इसका square root), और explained variance के overall fraction के लिए R-squared। Metric को match कीजिए इस बात से कि बड़ी misses थोड़ा hurt करनी चाहिए या बहुत।

Summary

Key takeaways

  • एक loss function एक single number से measure करता है एक model कितना wrong है; lower error better है।
  • हर prediction का error actual minus predicted है; actual [3,5,7,9] और predicted [4,5,7,6] के लिए errors [-1,0,0,3] हैं।
  • Mean Absolute Error (MAE) error sizes average करता है: (1+0+0+3)/4 = 1; यह target की units में है और पढ़ना easy है।
  • Mean Squared Error (MSE) squared errors average करता है: (1+0+0+9)/4 = 2.5; squaring इसे बड़ी errors ज़्यादा penalise कराता है।
  • यहाँ MSE (2.5), MAE (1) से ज़्यादा है क्योंकि 3 का error square होकर 9 बनता है।
  • R-squared = 1 - SS_res/SS_tot; यहाँ 1 - 10/20 = 0.5, तो model half variation explain करता है (1 के near excellent है, 0 baseline है)।
  • Memory hook: MAE typical miss है, MSE बड़ी errors punish करता है, R-squared explained variance का fraction है।

Study this properly

This page is the lesson to read. In Gri-Learn the same topic is a graded deck: the self-checks are scored and your weak topics are tracked. Free to start.

Start this topic

Already have an account? Sign in

More from Understanding Supervised Learning

Gri-Learn · syllabus-mapped B.C.A. lessons in English, Hindi and Gujarati

Loss functions: Mean Squared Error (MSE) - definition, computing and properties; Mean Absolute Error (MAE) - definition, computing and properties; understanding Regression and R-squared · Data Analytics using Python (Major-14) · Gri-Learn