Statistical methods: Mean, Median, Mode, Standard Deviation, Variance

Statistical methods data के गंदे distributions को crisp metrics में समेटती हैं जो numbers का central balance, middle rank, और volumetric spread उजागर करती हैं।

11 min read · 12 cards · 3 checks

Read in: English · हिन्दी · ગુજરાતી


Theory

Grand Totals के Blind Spots

आपके Semester 2 NumPy labs में, आपने सीखा कि np.sum() इस्तेमाल करके हज़ारों records को तुरंत कैसे जोड़ें। पर अगर आप अपने campus canteen के मासिक food spends की समीक्षा कर रहे हैं, क्या ख़र्च किया कुल grand total amount जानना आपको बताता है कि आपकी एक consistent daily eating habit है या आपने अपने पूरे महीने की pocket money एक अकेली birthday party पर उड़ा दी? aggregate sums numbers कैसे उतार-चढ़ाव करते हैं इसकी असली कहानी बताने में क्यों विफल होते हैं, और हम हज़ारों raw rows को घूरे बिना एक dataset का असली सार कैसे पकड़ते हैं?

Theory

Classroom Height Balance बनाम Scatter Spray

descriptive statistics को अपनी BCA section की एक group photograph लेने की तरह सोचिए। Mean एक physical balance scale का बिल्कुल center point खोजने जैसा है जहाँ हर कोई समान रूप से खड़ा हो सके। Median class को सबसे छोटे से सबसे लंबे तक sort करना और queue के बिल्कुल बीच में खड़े अकेले student पर tap करना है। पर यह देखने के लिए कि क्या आपके classmates सब लगभग एक ही height के हैं या basketball players और छोटे students का एक जंगली मिश्रण, आपको Standard Deviation चाहिए, जो मापता है कि व्यक्ति उस center balance line से कितनी दूर छिड़के या बिखरे हुए हैं।

Theory

Descriptive Statistics औपचारिक रूप से

Descriptive statistics data analytics को दो मूल श्रेणियों में बाँटती हैं: Central Tendency (जहाँ data गुच्छित होता है) और Dispersion (data कितना फैला है)। Mean arithmetic average दर्शाता है। Median एक ordered dataset को दो बराबर हिस्सों में बाँटता है, extreme outliers के ख़िलाफ़ एक मज़बूत माप के रूप में काम करते हुए। Mode सबसे बार-बार आने वाला data point दर्शाता है। dispersion मापने के लिए, Variance mean से average squared अंतर calculate करता है, और Standard Deviation (SD) spread metric को data के original scale unit पर वापस लाने के लिए variance का square root लेता है।

At a glance

Table 1: ज़रूरी Python statistical tools और mathematical attributes।

Statistical MetricPython Syntax TemplateBehavioral Characteristic
Arithmetic Meannp.mean(dataset)extreme outliers या typo entries के प्रति अत्यधिक संवेदनशील।
Middle Mediannp.median(dataset)physical rank order पर lock करता है; skewing spikes से अछूता।
Peak Modestats.mode(dataset)popularity track करता है; categorical parameters के लिए आदर्श।
Squared Variancenp.var(dataset)squared dispersion मापता है; outer variations को भारी बढ़ाता है।
Standard Deviationnp.std(dataset)standard distance spread के लिए निश्चित benchmark।

Theory

Worked Example: PocketMoney Variance का Audit

आइए NumPy और SciPy stats module इस्तेमाल करके student expenses की एक structured array के लिए descriptive metrics compute करें। हम ध्यान से देखेंगे कि एक अकेला विशाल outlier mean बनाम median को कैसे प्रभावित करता है, और spread का बिल्कुल metric तय करेंगे।

Practical

PocketMoney Spread Auditor

import numpy as np
from scipy import stats

# Step 1: Expense logs with a major outlier spike (e.g., a 400 Rs book purchase)
spends = np.array([40, 50, 40, 60, 400, 50, 60])

# Step 2: Calculate central tendencies
mean_val   = np.mean(spends)
median_val = np.median(spends)
mode_res   = stats.mode(spends, keepdims=True)

# Step 3: Compute dispersion metrics
variance_val = np.var(spends)
std_dev_val  = np.std(spends)

print(f"Mean Spend: {mean_val:.2f}")
print(f"Median Spend: {median_val:.2f}")
print(f"Mode (Most Frequent): {mode_res.mode[0]} (Count: {mode_res.count[0]})")
print(f"Variance Score: {variance_val:.2f}")
print(f"Standard Deviation Spread: {std_dev_val:.2f}")

This example runs in Gri-Learn on the web, where you can edit it and see the output.

Think first

Outlier Sensitivity का विश्लेषण करें

spends dataset देखिए: पाँच items 60 Rs से कम हैं, पर एक item 400 Rs तक कूदता है। यह अकेला spike आपके terminal output के अंदर mean और median के बीच balance को कैसे बदलेगा?

Show the answer

script output करेगी:

Mean Spend: 100.00

Median Spend: 50.00

Mode (Most Frequent): 40 (Count: 2)

Variance Score: 15057.14

Standard Deviation Spread: 122.71

क्यों? mean 100 Rs तक ऊपर खिंचता है क्योंकि यह कुल sum को 7 से भाग देकर हिसाब रखता है। पर, median array को [40, 40, 50, 50, 60, 60, 400] में sort करता है और 50 चुनता है, एक typical दिन के expense की बहुत ज़्यादा सटीक तस्वीर देते हुए। 122.71 का Standard Deviation उस 400 Rs purchase के कारण आपके spending data में उच्च स्तर की volatility उजागर करता है।

Quiz

एक university exam paper सवाल में, अगर एक distribution का mean 50 और variance 64 है, इसका बिल्कुल Standard Deviation metric क्या है?

  1. 4096
  2. 8
  3. 32
  4. 6.4
Show the answer

8

Standard Deviation गणितीय रूप से Variance के positive square root के रूप में परिभाषित है। 64 का square root 8 है। Mean center का एक माप है, तो यह इस conversion calculation को प्रभावित नहीं करता।

Quiz

corporate database analytics systems average consumer metrics या salary distributions analyze करते समय Mean के मुक़ाबले Median पर भारी क्यों निर्भर करते हैं?

  1. median को Python loops में चलाने के लिए बहुत कम computing memory चाहिए।
  2. mean extreme high value outliers से गंभीर रूप से अनुपात से बाहर खिंचता है, जबकि median स्थिर रहता है।
  3. mean floating decimal values रखती arrays process नहीं कर सकता।
  4. median distribution में duplicate entries को पूरी तरह अनदेखा करता है।
Show the answer

mean extreme high value outliers से गंभीर रूप से अनुपात से बाहर खिंचता है, जबकि median स्थिर रहता है।

क्योंकि mean हर element के बिल्कुल value का हिसाब रखता है, कुछ multi-millionaire salaries या भारी purchase outliers इसे कृत्रिम रूप से फुला सकते हैं। median केवल rank position पर निर्भर करता है, इसे skewed datasets के लिए एक ज़्यादा भरोसेमंद metric बनाते हुए।

Watch out

Classic जाल: Default Population Variance की चूक

university lab exams में एक बार-बार होने वाली ग़लती यह मानना है कि np.std() sample datasets के लिए standard textbook formulas की तरह बिल्कुल बर्ताव करता है। default रूप से, NumPy population standard deviation (denominator में N, या ddof=0) calculate करता है। अगर आपका question paper स्पष्ट रूप से sample standard deviation (N-1 degrees of freedom) माँगता है, आपको optional parameter np.std(dataset, ddof=1) देना होगा, वरना आपका अंतिम decimal value marks खो देगा!

Theory

Statistics को Semester 3 से जोड़ना

Descriptive statistics predictive intelligence की रीढ़ बनाती हैं। Semester 3 Data Science (BCA302) में, जब regression pipelines या normalization engines बनाते हैं, आप inputs साफ़ और scale करने के लिए mean और standard deviation matrices इस्तेमाल करेंगे, अपने machine learning models को विविध data को निष्पक्ष रूप से process करने में मदद करते हुए।

Summary

Key takeaways

  • Mean arithmetic balance point के रूप में काम करता है पर outliers के प्रति अत्यधिक कमज़ोर है।
  • Median sorted data का असली center पहचानता है, extreme value spikes के प्रति प्रतिरोध देते हुए।
  • Mode एक dataset के अंदर highest-frequency दोहराए parameters अलग करता है।
  • Variance squared deviations का मूल्यांकन करता है, आपके records की अंतर्निहित variability track करते हुए।
  • Standard Deviation variance score को साफ़ interpretation के लिए original unit scale पर वापस लाता है।
  • Memory Hook: Mean सारे cells जोड़ता है, Median rank position track करता है, SD deviation radius परिभाषित करता है!

Study this properly

This page is the lesson to read. In Gri-Learn the same topic is a graded deck: the self-checks are scored and your weak topics are tracked. Free to start.

Start this topic

Already have an account? Sign in

More from Python Libraries

Gri-Learn · syllabus-mapped B.C.A. lessons in English, Hindi and Gujarati

Statistical methods: Mean, Median, Mode, Standard Deviation, Variance · Programming Skills · Gri-Learn