Theory
Matrix Identity संकट
आपके Semester 2 NumPy labs में, आपने सीखा कि homogeneous arrays इस्तेमाल करके हज़ारों numeric transaction amounts को बिजली की गति पर कैसे process करें। पर क्या होता है जब आपके campus pocket-money tracker को एक साथ पूरी तरह अलग data types track करने होते हैं, जैसे categories के लिए text labels ('Food', 'Chai'), item counts के लिए integers, floating decimal percentages, और true/false validation flags? क्योंकि एक NumPy array हर cell को एक single data type में मजबूर करता है, mixed strings और numbers देना सब कुछ raw text strings में बदल देता है, आपके math tools को पूरी तरह बर्बाद करते हुए। हम एक high-performance tabular layout कैसे बनाते हैं जो एक professional Excel sheet की तरह rows, columns, और mixed data types सँभालता है?
Theory
Single-Column Ruler बनाम Multi-Department Ledger
एक Pandas Series को एक student attendance register के अंदर print एक अकेले vertical column की तरह सोचिए। यह data points को एक line में नीचे सूचीबद्ध करता है, पर इसके हर line के लिए एक custom painted label होता है (जैसे raw index positions 0, 1, 2 के बजाय student roll numbers)। एक Pandas DataFrame पूरी multi-page physical ledger book है। यह कई columns को साथ-साथ जोड़ता है। हर column पूरी तरह अलग चीज़ें रख सकता है, एक में text names हैं, अगले में numeric balances, और तीसरे में pass/fail marks। फिर भी, वे सब बाईं margin पर एक ही unique horizontal row index labels तक perfectly बँधे रहते हैं।
Theory
Pandas Architecture औपचारिक रूप से
Pandas library व्यावहारिक data analysis के लिए design की गई high-performance, इस्तेमाल में आसान data structures देती है। पहली मूल structure एक Series है, जो एक one-dimensional labeled array है जो कोई भी data type रख सकता है। दूसरी और सबसे व्यापक रूप से इस्तेमाल होने वाली structure DataFrame है, एक two-dimensional, size-mutable, और heterogeneous tabular data structure जिसमें labeled axes (rows और columns) हैं। NumPy के उलट, एक DataFrame में columns अलग heterogeneous data types (integers, floats, objects) रखते हैं, जबकि rows में सख़्त index alignment बनाए रखते हैं।
At a glance
Table 1: Pandas Series और DataFrame architectures की मूल संरचनात्मक तुलना।
| Pandas Data Layer | Dimensional Layout | Type Homogeneity Constraint |
|---|---|---|
| pd.Series(data, index) | 1D: custom label mappings वाला एक अकेला column। | Homogeneous: एक single Series के अंदर हर cell को types से मेल खाना होगा। |
| pd.DataFrame(data) | 2D: स्वतंत्र columns वाला एक पूर्ण tabular matrix। | Heterogeneous: Columns पूरी तरह अलग types store कर सकते हैं। |
| df.columns | Metadata: horizontal field headings उजागर करता है। | Immutable text labels या coordinate sequence descriptors। |
| df.index | Metadata: vertical row alignment सीमाएँ track करता है। | integers, text strings, timestamps, या roll numbers हो सकते हैं। |
Theory
Worked Example: Tabular Ledger का Blueprint बनाना
आइए explore करें कि हमारा चलता project, PocketMoney, mixed transaction logs को model करने के लिए Pandas components कैसे इस्तेमाल करता है। हम पहले एक अलग 1D Series बनाएँगे, और फिर कई विविध sequences को एक एकीकृत 2D DataFrame grid में जोड़ेंगे।
Practical
PocketMoney Pandas Blueprint
import pandas as pd
# Step 1: Construct an isolated 1D Series tracking category types
categories = pd.Series(["Samosa", "Bus", "Chai"], index=["t1", "t2", "t3"])
# Step 2: Assemble structured dictionary metrics with mixed data types
raw_data = {
"Category": ["Samosa", "Bus", "Chai", "Stationery"],
"Amount": [35, 50, 15, 120],
"Essential": [False, True, False, True]
}
# Step 3: Compile into a heterogeneous 2D DataFrame table
ledger_df = pd.DataFrame(raw_data)
print("--- 1D Series Output ---")
print(categories)
print("\n--- 2D DataFrame Table ---")
print(ledger_df)
print("\nData Types Per Column:\n", ledger_df.dtypes)This example runs in Gri-Learn on the web, where you can edit it and see the output.
Think first
Structured Alignment को ट्रेस करें
ledger_df के printed output layout का विश्लेषण कीजिए। अगर DataFrame constructor को कोई स्पष्ट index array न दिया जाए तो Pandas rows को अपने-आप कैसे format करता है? कौन से data types inferred होते हैं?
Show the answer
script output करेगी:
--- 1D Series Output ---
t1 Samosa
t2 Bus
t3 Chai
dtype: object
--- 2D DataFrame Table ---
Category Amount Essential
0 Samosa 35 False
1 Bus 50 True
2 Chai 15 False
3 Stationery 120 True
Data Types Per Column:
Category object
Amount int64
Essential bool
dtype: object
क्यों? जब DataFrame के लिए कोई स्पष्ट row index निर्दिष्ट नहीं होता, Pandas अपने-आप 0 से शुरू होता एक standard numeric sequence generate करता है। हर column अपना native data type perfectly रखता है, Category एक 'object' (text string) पर map होता है, Amount 'int64' पर, और Essential 'bool' पर, किसी भी safe mathematical properties को degrade होने से रोकते हुए।
Quiz
एक NumPy ndarray और एक Pandas DataFrame structure के बीच प्राथमिक संरचनात्मक अंतर क्या है?
- एक DataFrame केवल floating-point numbers रख सकता है, जबकि एक ndarray characters रखता है।
- एक ndarray सख़्ती से homogeneous है (एक साझा type), जबकि एक DataFrame heterogeneous है (columns के पूरी तरह स्वतंत्र data types हो सकते हैं)।
- एक ndarray columns पर automated text labels support करता है, जबकि एक DataFrame पूरी तरह numerical indices पर निर्भर करता है।
- DataFrames memory में initialize होने के बाद update या resize नहीं किए जा सकते।
Show the answer
एक ndarray सख़्ती से homogeneous है (एक साझा type), जबकि एक DataFrame heterogeneous है (columns के पूरी तरह स्वतंत्र data types हो सकते हैं)।
NumPy arrays execution speed हासिल करने के लिए memory में हर एक element में पूर्ण data type एकरूपता लागू करते हैं। Pandas DataFrames tabular databases की तरह काम करते हैं, आपको adjacent columns में text, integers, और booleans को उनकी अलग पहचान खोए बिना साफ़-सुथरे mix करने का लचीलापन देते हुए।
Quiz
अगर आप एक multi-column Pandas DataFrame से निकाले एक single column के internal type को देखते हैं, Python कौन सा data object type return करेगा?
- एक standard primitive Python list
- एक अकेला अलग NumPy scalar integer
- एक Pandas Series object
- एक flat text string file string
Show the answer
एक Pandas Series object
एक Pandas DataFrame संरचनात्मक रूप से एक साझा horizontal index row map के साथ साथ-साथ aligned कई individual columns से बना है। एक DataFrame से निकाला हर individual column एक 1D labeled sequence के रूप में काम करता है, जो एक Pandas Series object है।
Watch out
Classic जाल: Equal-Length Column Collapse
एक dictionary से एक DataFrame बनाते समय university lab exams में सबसे बार-बार ग़लती length में मेल न खाती lists देना है (जैसे Column A में 4 values हैं, पर Column B में केवल 3)। Python blanks को zeros या None से नहीं भरेगा; बजाय इसके, constructor एक ValueError: All arrays must be of the same length के साथ तुरंत crash होगा। initialize करने से पहले अपनी list counts दोबारा जाँचिए!
Theory
Tables को Semester 3 से जोड़ना
Tabular index tracking automated industrial analytical pipelines का मूल framework बनाता है। Semester 3 Data Science (BCA302) और Web Architecture (BCA303) में, आप raw transactional databases साफ़ करने, null anomalies filter करने, और multi-type inputs को सीधे predictive model engines में feed करने के लिए DataFrames को memory tables के रूप में इस्तेमाल करेंगे।
Summary
Key takeaways
- Pandas heterogeneous mixed types को आसानी से सँभालने के लिए संरचनात्मक data objects पेश करता है।
- एक Series स्पष्ट user-defined tracking labels से बँधे एक 1D sequence को दर्शाता है।
- एक DataFrame एक 2D tabular layout है जिसमें साझा index alignments वाले अलग columns हैं।
- अगर कोई custom labels न दिए जाएँ तो DataFrames अपने-आप 0 से शुरू होता एक integer-based index generate करते हैं।
- एक व्यापक DataFrame table से निकाला हर एक स्वतंत्र column एक Series के रूप में काम करता है।
- Memory Hook: एक Series एक अकेला labeled column है; एक DataFrame पूरा heterogeneous multi-column ledger है!