Theory
Raw Memory Streams की चुनौती
आपके Semester 1 BCA104 labs में, आप parallel 1D primitive arrays या multi-dimensional nested arrays इस्तेमाल करके record tables store करते थे। पर, values की एक matrix parse करने के लिए hardcoded coordinate integer tracking चाहिए था, और descriptive headers जोड़ना या orientations बदलना शुरू से custom layout manipulation algorithms लिखने की माँग करता था। Pandas इस संरचनात्मक setup को क्यों सरल बनाता है, आपको nested lists या standard data dictionaries को सीधे एक constructor में देकर साफ़, tabular database tables generate करने देते हुए?
Theory
Horizontal Blueprint बनाम Vertical Filing Cabinet
raw cards से एक physical spreadsheet table assemble करने की कल्पना कीजिए। एक List of Lists से एक DataFrame बनाना लंबे paper receipt tapes को horizontally, row-दर-row बिछाने जैसा है। आप row 1 stack करते हैं, फिर row 2 उसके नीचे रखते हैं, और आख़िर में column names बनाने के लिए हर vertical lane के ऊपर sticky labels लगाते हैं। दूसरी ओर, एक Dictionary of Lists से एक DataFrame बनाना एक ख़ाली vertical organizer तक चलने जैसा है। हर अलग dictionary folder एक pre-labeled vertical column दर्शाता है, और आप data slips की एक जैसी rows को हर slot में सीधे नीचे गिराते हैं।
Theory
Data Ingestion औपचारिक रूप से
pd.DataFrame() constructor एक 2D grid बनाने के लिए विविध structured Python sequencers स्वीकारता है। एक 2D List of Lists इस्तेमाल करते समय, हर sub-list को स्पष्ट रूप से एक अलग horizontal row block के रूप में treat किया जाता है; आप optional columns list parameter इस्तेमाल करके column name labels देते हैं। इसके उलट, एक Dictionary of Lists इस्तेमाल करते समय, हर dictionary key को एक vertical column header के रूप में treat किया जाता है, और मेल खाती associated list उस field के elements पूरे page में नीचे रखती है। यह method माँग करती है कि हर list value एक सख़्त equal-length sequence constraint साझा करे।
At a glance
Table 1: Pandas में row-wise list layouts और column-wise dictionary layouts के बीच संरचनात्मक अंतर।
| Data Structure Source | Constructor Parsing Rule | Required Shape Parameters |
|---|---|---|
| [[Row1], [Row2]] | nested components को horizontally rows के रूप में parse करता है। | optional columns=['A', 'B'] parameter स्वीकारता है। |
| {'Col1': [Val1, Val2]} | values को vertically स्वतंत्र columns के रूप में parse करता है। | Keys अपने-आप tabular column headings परिभाषित करती हैं। |
| Length Mismatch | runtime के दौरान संरचनात्मक data errors trigger करता है। | list sizes मेल न खाने पर एक ValueError फेंकता है। |
Theory
Worked Example: Semester Budget assemble करना
आइए देखें कि हमारा PocketMoney tracking system दोनों संरचनात्मक entry methods इस्तेमाल करके एक जैसे tabular data containers कैसे बनाता है। हम पहले एक nested raw list structure parse करेंगे, और फिर एक key-mapped dictionary structure इस्तेमाल करके बिल्कुल वही layout बनाएँगे।
Practical
PocketMoney DataFrame Constructor
import pandas as pd
# Method A: Constructing row-by-row using a 2D List of Lists
row_matrix = [
["Samosa", 35, "Food"],
["Bus", 50, "Travel"],
["Chai", 15, "Food"]
]
col_headers = ["Item", "Cost", "Type"]
df_from_lists = pd.DataFrame(row_matrix, columns=col_headers)
# Method B: Constructing column-by-column using a Dictionary of equal-length lists
col_dictionary = {
"Item": ["Samosa", "Bus", "Chai"],
"Cost": [35, 50, 15],
"Type": ["Food", "Travel", "Food"]
}
df_from_dict = pd.DataFrame(col_dictionary)
print("--- DataFrame from List of Lists ---")
print(df_from_lists)
print("\n--- DataFrame from Dictionary ---")
print(df_from_dict)This example runs in Gri-Learn on the web, where you can edit it and see the output.
Think first
Structural Outlines की तुलना करें
दोनों compilation outputs का मन में विश्लेषण कीजिए। क्या terminal पर print होने पर संरचनात्मक shapes, columns, और internal coordinate alignments अलग दिखेंगे?
Show the answer
script output करेगी:
--- DataFrame from List of Lists ---
Item Cost Type
0 Samosa 35 Food
1 Bus 50 Travel
2 Chai 15 Food
--- DataFrame from Dictionary ---
Item Cost Type
0 Samosa 35 Food
1 Bus 50 Travel
2 Chai 15 Food
क्यों? दोनों creation दृष्टिकोण एक जैसे 2D relational matrices देते हैं। Method A horizontal coordinates को स्पष्ट रूप से rows के रूप में map करता है और headers को एक अलग parameter के ज़रिए सौंपता है। Method B dictionary keys को vertical संरचनात्मक labels के रूप में इस्तेमाल करता है और columns में मेल खाती list indices मिलाकर rows बनाता है।
Quiz
क्या होता है अगर आप दो keys रखता एक dictionary इस्तेमाल करके एक DataFrame initialize करते हैं, जहाँ key 'A' length 5 की एक list पर map होती है और key 'B' length 4 की एक list पर?
- Pandas ग़ायब cell को अपने-आप 0 पर map करता है।
- Pandas ग़ायब cell को एक NaN null object flag पर map करता है।
- यह एक तुरंत ValueError फेंकता है यह बताते हुए कि arrays की length बराबर होनी चाहिए।
- यह एक 2D table के बजाय एक 1D Series array container बनाता है।
Show the answer
यह एक तुरंत ValueError फेंकता है यह बताते हुए कि arrays की length बराबर होनी चाहिए।
data को सुरक्षित रूप से continuous rows और columns में align करने के लिए, pd.DataFrame() में दी गई एक dictionary में एक जैसी length की lists होनी चाहिए। कोई भी dimension असंतुलन data corruption रोकने के लिए एक तुरंत ValueError: All arrays must be of the same length फेंकता है।
Quiz
जब आप एक nested list of lists, जैसे data = [[10, 20], [30, 40]], से सीधे एक DataFrame बनाते हैं बिना एक स्पष्ट columns argument दिए, output कौन से headers दिखाएगा?
- यह अपने-आप labels Column1 और Column2 इस्तेमाल करता है।
- यह column headers के रूप में integer index markers 0 और 1 इस्तेमाल करता है।
- यह एक MissingArgumentException फेंकता है।
- यह पहली inner sub-list [10, 20] को headers block के रूप में treat करता है।
Show the answer
यह column headers के रूप में integer index markers 0 और 1 इस्तेमाल करता है।
अगर कोई custom column labels न दिए जाएँ, Pandas column labels को बिल्कुल row indexes की तरह सँभालता है: यह 0 से शुरू होते एक integer sequence पर default करता है। columns page में नीचे 0 और 1 labeled होंगे।
Watch out
Classic जाल: Row-Wise Dictionary Assumption की गिरावट
university examinations में सबसे आम marks-गँवाने वाली ग़लती constructor में lists की एक dictionary देना और keys के rows बनाने की उम्मीद करना है। याद रखिए: dictionary keys हमेशा vertical column names बनाती हैं। अगर आपको dictionary keys को इसके बजाय rows के रूप में चाहिए, आपको अपना data वैकल्पिक ingestion methods जैसे DataFrame.from_dict(data, orient='index') से देना होगा!
Theory
Ingestion को Semester 3 से जोड़ना
memory primitives से datasets संरचित करना आधुनिक software integrations के लिए एक ज़रूरी कौशल है। Semester 3 (BCA303/BCA304) में, जब API connections से web data payloads (JSON logs) parse करते हैं, आप नियमित रूप से raw dictionaries को Pandas DataFrames में decode करेंगे। यह आपको web data को back-end databases में save करने से पहले साफ़, index, और organize करने देता है।
Summary
Key takeaways
- Pandas native Python lists, matrices, या dictionaries से DataFrames साफ़-सुथरे बनाता है।
- एक nested 2D List of Lists देना data arrays को horizontally individual rows के रूप में parse करता है।
- एक Dictionary of Lists देना records को vertically संरचित करता है, keys को column names के रूप में इस्तेमाल करते हुए।
- एक input dictionary में सारे element sequences एक ValueError से बचने के लिए एक जैसी length के होने चाहिए।
- Missing index या column arguments 0 से शुरू होते एक integer sequence पर default होते हैं।
- Memory Hook: Nested lists horizontally rows के रूप में stack होती हैं; dictionary keys vertically columns के रूप में गिरती हैं!