Creating dataframe using list / dict of equal length lists

Pandas DataFrames को raw data को standard 2D lists of lists या key-aligned equal-length rows के dictionaries में organize करके तुरंत बनाया जा सकता है।

11 min read · 12 cards · 3 checks

Read in: English · हिन्दी · ગુજરાતી


Theory

Raw Memory Streams की चुनौती

आपके Semester 1 BCA104 labs में, आप parallel 1D primitive arrays या multi-dimensional nested arrays इस्तेमाल करके record tables store करते थे। पर, values की एक matrix parse करने के लिए hardcoded coordinate integer tracking चाहिए था, और descriptive headers जोड़ना या orientations बदलना शुरू से custom layout manipulation algorithms लिखने की माँग करता था। Pandas इस संरचनात्मक setup को क्यों सरल बनाता है, आपको nested lists या standard data dictionaries को सीधे एक constructor में देकर साफ़, tabular database tables generate करने देते हुए?

Theory

Horizontal Blueprint बनाम Vertical Filing Cabinet

raw cards से एक physical spreadsheet table assemble करने की कल्पना कीजिए। एक List of Lists से एक DataFrame बनाना लंबे paper receipt tapes को horizontally, row-दर-row बिछाने जैसा है। आप row 1 stack करते हैं, फिर row 2 उसके नीचे रखते हैं, और आख़िर में column names बनाने के लिए हर vertical lane के ऊपर sticky labels लगाते हैं। दूसरी ओर, एक Dictionary of Lists से एक DataFrame बनाना एक ख़ाली vertical organizer तक चलने जैसा है। हर अलग dictionary folder एक pre-labeled vertical column दर्शाता है, और आप data slips की एक जैसी rows को हर slot में सीधे नीचे गिराते हैं।

Theory

Data Ingestion औपचारिक रूप से

pd.DataFrame() constructor एक 2D grid बनाने के लिए विविध structured Python sequencers स्वीकारता है। एक 2D List of Lists इस्तेमाल करते समय, हर sub-list को स्पष्ट रूप से एक अलग horizontal row block के रूप में treat किया जाता है; आप optional columns list parameter इस्तेमाल करके column name labels देते हैं। इसके उलट, एक Dictionary of Lists इस्तेमाल करते समय, हर dictionary key को एक vertical column header के रूप में treat किया जाता है, और मेल खाती associated list उस field के elements पूरे page में नीचे रखती है। यह method माँग करती है कि हर list value एक सख़्त equal-length sequence constraint साझा करे।

At a glance

Table 1: Pandas में row-wise list layouts और column-wise dictionary layouts के बीच संरचनात्मक अंतर।

Data Structure SourceConstructor Parsing RuleRequired Shape Parameters
[[Row1], [Row2]]nested components को horizontally rows के रूप में parse करता है।optional columns=['A', 'B'] parameter स्वीकारता है।
{'Col1': [Val1, Val2]}values को vertically स्वतंत्र columns के रूप में parse करता है।Keys अपने-आप tabular column headings परिभाषित करती हैं।
Length Mismatchruntime के दौरान संरचनात्मक data errors trigger करता है।list sizes मेल न खाने पर एक ValueError फेंकता है।

Theory

Worked Example: Semester Budget assemble करना

आइए देखें कि हमारा PocketMoney tracking system दोनों संरचनात्मक entry methods इस्तेमाल करके एक जैसे tabular data containers कैसे बनाता है। हम पहले एक nested raw list structure parse करेंगे, और फिर एक key-mapped dictionary structure इस्तेमाल करके बिल्कुल वही layout बनाएँगे।

Practical

PocketMoney DataFrame Constructor

import pandas as pd

# Method A: Constructing row-by-row using a 2D List of Lists
row_matrix = [
    ["Samosa", 35, "Food"],
    ["Bus", 50, "Travel"],
    ["Chai", 15, "Food"]
]
col_headers = ["Item", "Cost", "Type"]
df_from_lists = pd.DataFrame(row_matrix, columns=col_headers)

# Method B: Constructing column-by-column using a Dictionary of equal-length lists
col_dictionary = {
    "Item": ["Samosa", "Bus", "Chai"],
    "Cost": [35, 50, 15],
    "Type": ["Food", "Travel", "Food"]
}
df_from_dict = pd.DataFrame(col_dictionary)

print("--- DataFrame from List of Lists ---")
print(df_from_lists)
print("\n--- DataFrame from Dictionary ---")
print(df_from_dict)

This example runs in Gri-Learn on the web, where you can edit it and see the output.

Think first

Structural Outlines की तुलना करें

दोनों compilation outputs का मन में विश्लेषण कीजिए। क्या terminal पर print होने पर संरचनात्मक shapes, columns, और internal coordinate alignments अलग दिखेंगे?

Show the answer

script output करेगी:

--- DataFrame from List of Lists ---

Item Cost Type

0 Samosa 35 Food

1 Bus 50 Travel

2 Chai 15 Food

--- DataFrame from Dictionary ---

Item Cost Type

0 Samosa 35 Food

1 Bus 50 Travel

2 Chai 15 Food

क्यों? दोनों creation दृष्टिकोण एक जैसे 2D relational matrices देते हैं। Method A horizontal coordinates को स्पष्ट रूप से rows के रूप में map करता है और headers को एक अलग parameter के ज़रिए सौंपता है। Method B dictionary keys को vertical संरचनात्मक labels के रूप में इस्तेमाल करता है और columns में मेल खाती list indices मिलाकर rows बनाता है।

Quiz

क्या होता है अगर आप दो keys रखता एक dictionary इस्तेमाल करके एक DataFrame initialize करते हैं, जहाँ key 'A' length 5 की एक list पर map होती है और key 'B' length 4 की एक list पर?

  1. Pandas ग़ायब cell को अपने-आप 0 पर map करता है।
  2. Pandas ग़ायब cell को एक NaN null object flag पर map करता है।
  3. यह एक तुरंत ValueError फेंकता है यह बताते हुए कि arrays की length बराबर होनी चाहिए।
  4. यह एक 2D table के बजाय एक 1D Series array container बनाता है।
Show the answer

यह एक तुरंत ValueError फेंकता है यह बताते हुए कि arrays की length बराबर होनी चाहिए।

data को सुरक्षित रूप से continuous rows और columns में align करने के लिए, pd.DataFrame() में दी गई एक dictionary में एक जैसी length की lists होनी चाहिए। कोई भी dimension असंतुलन data corruption रोकने के लिए एक तुरंत ValueError: All arrays must be of the same length फेंकता है।

Quiz

जब आप एक nested list of lists, जैसे data = [[10, 20], [30, 40]], से सीधे एक DataFrame बनाते हैं बिना एक स्पष्ट columns argument दिए, output कौन से headers दिखाएगा?

  1. यह अपने-आप labels Column1 और Column2 इस्तेमाल करता है।
  2. यह column headers के रूप में integer index markers 0 और 1 इस्तेमाल करता है।
  3. यह एक MissingArgumentException फेंकता है।
  4. यह पहली inner sub-list [10, 20] को headers block के रूप में treat करता है।
Show the answer

यह column headers के रूप में integer index markers 0 और 1 इस्तेमाल करता है।

अगर कोई custom column labels न दिए जाएँ, Pandas column labels को बिल्कुल row indexes की तरह सँभालता है: यह 0 से शुरू होते एक integer sequence पर default करता है। columns page में नीचे 0 और 1 labeled होंगे।

Watch out

Classic जाल: Row-Wise Dictionary Assumption की गिरावट

university examinations में सबसे आम marks-गँवाने वाली ग़लती constructor में lists की एक dictionary देना और keys के rows बनाने की उम्मीद करना है। याद रखिए: dictionary keys हमेशा vertical column names बनाती हैं। अगर आपको dictionary keys को इसके बजाय rows के रूप में चाहिए, आपको अपना data वैकल्पिक ingestion methods जैसे DataFrame.from_dict(data, orient='index') से देना होगा!

Theory

Ingestion को Semester 3 से जोड़ना

memory primitives से datasets संरचित करना आधुनिक software integrations के लिए एक ज़रूरी कौशल है। Semester 3 (BCA303/BCA304) में, जब API connections से web data payloads (JSON logs) parse करते हैं, आप नियमित रूप से raw dictionaries को Pandas DataFrames में decode करेंगे। यह आपको web data को back-end databases में save करने से पहले साफ़, index, और organize करने देता है।

Summary

Key takeaways

  • Pandas native Python lists, matrices, या dictionaries से DataFrames साफ़-सुथरे बनाता है।
  • एक nested 2D List of Lists देना data arrays को horizontally individual rows के रूप में parse करता है।
  • एक Dictionary of Lists देना records को vertically संरचित करता है, keys को column names के रूप में इस्तेमाल करते हुए।
  • एक input dictionary में सारे element sequences एक ValueError से बचने के लिए एक जैसी length के होने चाहिए।
  • Missing index या column arguments 0 से शुरू होते एक integer sequence पर default होते हैं।
  • Memory Hook: Nested lists horizontally rows के रूप में stack होती हैं; dictionary keys vertically columns के रूप में गिरती हैं!

Study this properly

This page is the lesson to read. In Gri-Learn the same topic is a graded deck: the self-checks are scored and your weak topics are tracked. Free to start.

Start this topic

Already have an account? Sign in

More from Python Libraries

Gri-Learn · syllabus-mapped B.C.A. lessons in English, Hindi and Gujarati

Creating dataframe using list / dict of equal length lists · Programming Skills · Gri-Learn