Python Libraries to Automate EDA: Pandas and NumPy

दो Python libraries data analysis की heavy lifting करती हैं: pandas आपको DataFrame देता है, एक smart table जिसमें आप एक CSV load कर सकते हैं और एक line में summarise कर सकते हैं, और numpy इसके नीचे fast numeric work power करता है।

10 min read · 7 cards · 2 checks

Read in: English · हिन्दी · ગુજરાતી


Theory

Boring Part को Automate करना

Placement dataset के 500 rows पर हाथ से EDA करना, missing values count करना, हर column को average करना, manually करना impossible है और scratch से script करना tedious है। Python की data libraries इसे automate करती हैं, pages के काम को दो lines में बदलते हुए।

दो libraries यह काम करती हैं: pandas, जो आपको DataFrame नाम की एक smart table देता है, और numpy, जो इसके नीचे fast numeric maths power करता है। आप earlier थोड़ा pandas से मिले थे; यहाँ आप exploratory analysis automate करने के लिए इसे इस्तेमाल करते हैं। यह lesson वे दो commands दिखाता है जो एक पूरे dataset को instantly summarise करते हैं।

Theory

pandas और numpy

pandas Python में data analysis का workhorse है। इसका central object DataFrame है: एक labelled two-dimensional table, rows और named columns, code में exactly एक spreadsheet जैसा। आप pd.read_csv() से एक CSV file directly इसमें load करते हैं।

numpy नीचे बैठता है, fast numeric arrays और mathematical operations provide करते हुए जिन पर pandas rely करता है। आप अक्सर calculations के लिए directly numpy इस्तेमाल करेंगे, पर ज़्यादातर समय pandas quietly आपके लिए इसे इस्तेमाल करता है। साथ में ये आपको बहुत कम code के साथ data load, clean, summarise, और analyse करने देते हैं, यही वजह है लगभग सारा Python data work इन्हें इस्तेमाल करता है।

Practical

Placement Data को Load और Summarise करना

import pandas as pd
import numpy as np

# Load the CSV straight into a DataFrame (a table in code)
df = pd.read_csv("placement.csv")

df.info()        # columns, their data types, and non-null (missing) counts
df.describe()    # count, mean, std, min, quartiles, max for numeric columns

# describe() output (numeric columns) looks like:
#        marks  projects
# mean    71.0       2.2
# min     55.0       1.0
# max     90.0       4.0

This example runs in Gri-Learn on the web, where you can edit it and see the output.

Formula

info() और describe(): Instant EDA

दो methods आपका ज़्यादातर first-pass EDA करते हैं। df.info() हर column, इसकी data type, और इसमें कितने non-null values हैं यह list करता है, तो missing data तुरंत उभर आता है (rows से कम non-nulls वाले एक column में gaps हैं)।

df.describe() summary statistics compute करता है, count, mean, standard deviation, min, quartiles, max, हर numeric column के लिए एक साथ। यही पूरे dataset के लिए एक single call में univariate analysis है। दोनों मिलकर, info() आपको structure और gaps बताता है, और describe() आपको numeric shape बताता है, seconds में।

Quiz

आपको अपने DataFrame के हर numeric column के mean, min, max, और quartiles का एक one-line summary चाहिए। कौन सा pandas method यह देता है?

  1. df.info(), क्योंकि यह सब कुछ print करता है
  2. df.describe(), जो numeric columns के लिए summary statistics compute करता है
  3. pd.read_csv(), क्योंकि यह data load करता है
  4. numpy, क्योंकि यह maths करता है
Show the answer

df.describe(), जो numeric columns के लिए summary statistics compute करता है

df.describe() summary statistics return करता है, count, mean, standard deviation, min, quartiles, और max, हर numeric column के लिए एक call में, exactly requested summary। Option A, df.info(), इसके बजाय structure दिखाता है: columns, इनकी data types, और non-null counts (missing data spot करने के लिए great), पर mean/min/max/quartile statistics नहीं। Option C, pd.read_csv(), एक CSV को DataFrame में load करता है; यह summarise नहीं करता। Option D, numpy, low-level numeric arrays और maths provide करता है पर एक DataFrame के लिए one-line summary method नहीं है। Practice में इन्हें pair कीजिए: structure और gaps के लिए info(), numeric summary के लिए describe()।

Think first

Plain Python Loops की बजाय pandas और numpy क्यों इस्तेमाल करें?

आप ordinary Python के साथ rows पर loop कर सकते थे। Data work के लिए pandas और numpy इतने better क्यों हैं? फिर tap कीजिए।

Show the answer

क्योंकि ये FAST, CONCISE हैं, और specifically data के लिए बने हैं, जबकि plain Python loops large datasets के लिए slow और verbose हैं। Under the hood, numpy numbers को compact arrays में store करता है और पूरे array पर एक साथ operations perform करता है (vectorised operations optimised low-level code में implemented), तो दस लाख numbers के एक column को add करना, या इनका mean compute करना, एक quick step में होता है, दस लाख बार iterate करने वाले एक slow Python loop की बजाय। pandas इस पर build करके आपको DataFrame देता है, जहाँ एक पूरा analysis, rows filter करना, group करना, summarise करना, tables join करना, clear code की एक या दो lines में express होता है जो fast भी run करता है। हाथ से बनाए Python loops के साथ same चीज़ लिखने में कहीं ज़्यादा lines लगतीं, यह कहीं slower होता, और कहीं ज़्यादा error-prone होता। एक rich ecosystem भी है: CSVs और Excel पढ़ना, missing values handle करना, plotting, और machine-learning libraries में data feed करना सब pandas के साथ smoothly integrate करते हैं। तो essentially हर Python data analyst pandas और numpy की तरफ क्यों जाता है इसकी वजह यह है कि ये slow, fiddly, low-level work को fast, readable, high-level operations में बदल देते हैं, आपको plumbing की बजाय analysis पर focus करने देते हैं। Data के लिए right tools: expressive और fast।

Summary

Key takeaways

  • pandas और numpy data analysis के लिए core Python libraries हैं, EDA automate करते हुए।
  • pandas DataFrame provide करता है: एक labelled 2D table (code में एक spreadsheet जैसा); pd.read_csv() से एक CSV load कीजिए।
  • numpy fast numeric arrays और maths provide करता है जिन पर pandas built है।
  • df.info() हर column की data type और non-null count दिखाता है, missing data को visible बनाते हुए।
  • df.describe() सारे numeric columns के लिए एक call में count, mean, std, min, quartiles, और max देता है, univariate EDA automate करते हुए।
  • ये fast और concise हैं क्योंकि numpy पूरे arrays पर एक साथ operate करता है (vectorised), slow Python loops के unlike।
  • Memory hook: pandas DataFrame table है; structure और gaps के लिए info(), numeric summary के लिए describe()।

Study this properly

This page is the lesson to read. In Gri-Learn the same topic is a graded deck: the self-checks are scored and your weak topics are tracked. Free to start.

Start this topic

Already have an account? Sign in

More from Automate EDA (Exploratory Data Analysis)

Gri-Learn · syllabus-mapped B.C.A. lessons in English, Hindi and Gujarati

Python Libraries to Automate EDA: Pandas and NumPy · Data Analytics using Python (Major-14) · Gri-Learn