Creating dataframe using list / dict of equal length lists

Pandas DataFrames raw data ને standard 2D lists of lists કે સરખી લંબાઈની rows વાળી key-aligned dictionaries માં ગોઠવીને તરત બનાવી શકાય છે.

11 min read · 12 cards · 3 checks

Read in: English · हिन्दी · ગુજરાતી


Theory

Raw Memory Streams નો પડકાર

તમારા Semester 1 BCA104 labs માં, તમે parallel 1D primitive arrays કે multi-dimensional nested arrays વાપરીને record tables સંઘરતા. પણ, values ના એક matrix ને parse કરવા hardcoded coordinate integer tracking જોઈતું, અને descriptive headers ઉમેરવા કે orientations બદલવા શરૂઆતથી custom layout manipulation algorithms લખવા પડતા. Pandas આ સંરચનાત્મક setup ને કેમ સરળ બનાવે છે, તમને સાફ, tabular database tables બનાવવા nested lists કે standard data dictionaries સીધા એક constructor માં આપવા દેતાં?

Theory

આડું Blueprint સામે ઊભું Filing Cabinet

raw cards માંથી એક ભૌતિક spreadsheet table જોડવાનું કલ્પો. એક List of Lists માંથી DataFrame બનાવવું એટલે લાંબા કાગળના receipt tapes ને આડા, હરોળ-દર-હરોળ પાથરવા. તમે row 1 ગોઠવો, પછી row 2 એની નીચે મૂકો, અને છેલ્લે દરેક ઊભી લેન ની ટોચે column names બનાવવા sticky labels ચોંટાડો. બીજી બાજુ, એક Dictionary of Lists માંથી DataFrame બનાવવું એટલે એક ખાલી ઊભા organizer પાસે જવું. દરેક અલગ dictionary folder એક પહેલેથી labeled ઊભા column ને દર્શાવે છે, અને તમે data slips ની એકસરખી rows સીધી દરેક ખાનામાં નીચે નાખો છો.

Theory

Data Ingestion ઔપચારિક રીતે

pd.DataFrame() constructor એક 2D grid બનાવવા વૈવિધ્યસભર સંરચિત Python sequencers સ્વીકારે છે. જ્યારે એક 2D List of Lists વાપરો, દરેક sub-list ને સ્પષ્ટપણે એક અલગ આડા row block તરીકે ગણવામાં આવે છે; તમે optional columns list parameter વાપરીને column name labels આપો છો. એથી ઊલટું, જ્યારે એક Dictionary of Lists વાપરો, દરેક dictionary key ને એક ઊભા column header તરીકે ગણવામાં આવે છે, અને એની સાથે જોડાયેલી list એ field ના elements પાનાની નીચે તરફ રાખે છે. આ પદ્ધતિ માંગે છે કે દરેક list value એક કડક equal-length sequence constraint પાળે.

At a glance

Table 1: Pandas માં row-wise list layouts અને column-wise dictionary layouts વચ્ચેના સંરચનાત્મક ભેદ.

Data Structure નો સ્રોતConstructor નો Parsing નિયમજરૂરી Shape Parameters
[[Row1], [Row2]]Nested components ને આડા rows તરીકે parse કરે છે.Optional columns=['A', 'B'] parameter સ્વીકારે છે.
{'Col1': [Val1, Val2]}Values ને ઊભા સ્વતંત્ર columns તરીકે parse કરે છે.Keys આપોઆપ tabular column headings વ્યાખ્યાયિત કરે છે.
Length MismatchRuntime દરમિયાન સંરચનાત્મક data errors trigger કરે છે.જો list sizes ન મળે તો એક ValueError ઊભી કરે છે.

Theory

Worked Example: Semester Budget જોડવું

ચાલો જોઈએ કે આપણી PocketMoney tracking system બંને સંરચનાત્મક entry પદ્ધતિઓ વાપરીને એકસરખા tabular data containers કેવી રીતે બનાવે છે. આપણે પહેલાં એક nested raw list structure parse કરીશું, અને પછી એક key-mapped dictionary structure વાપરીને બિલકુલ એ જ layout બનાવીશું.

Practical

PocketMoney DataFrame Constructor

import pandas as pd

# Method A: Constructing row-by-row using a 2D List of Lists
row_matrix = [
    ["Samosa", 35, "Food"],
    ["Bus", 50, "Travel"],
    ["Chai", 15, "Food"]
]
col_headers = ["Item", "Cost", "Type"]
df_from_lists = pd.DataFrame(row_matrix, columns=col_headers)

# Method B: Constructing column-by-column using a Dictionary of equal-length lists
col_dictionary = {
    "Item": ["Samosa", "Bus", "Chai"],
    "Cost": [35, 50, 15],
    "Type": ["Food", "Travel", "Food"]
}
df_from_dict = pd.DataFrame(col_dictionary)

print("--- DataFrame from List of Lists ---")
print(df_from_lists)
print("\n--- DataFrame from Dictionary ---")
print(df_from_dict)

This example runs in Gri-Learn on the web, where you can edit it and see the output.

Think first

સંરચનાત્મક રૂપરેખાઓ સરખાવો

બંને compilation outputs મનમાં તપાસો. terminal પર છપાય ત્યારે સંરચનાત્મક shapes, columns, અને આંતરિક coordinate alignments અલગ દેખાશે?

Show the answer

script output કરશે:

--- DataFrame from List of Lists ---

Item Cost Type

0 Samosa 35 Food

1 Bus 50 Travel

2 Chai 15 Food

--- DataFrame from Dictionary ---

Item Cost Type

0 Samosa 35 Food

1 Bus 50 Travel

2 Chai 15 Food

કેમ? બંને creation અભિગમો એકસરખા 2D relational matrices આપે છે. Method A આડા coordinates ને સ્પષ્ટપણે rows તરીકે map કરે છે અને એક અલગ parameter દ્વારા headers સોંપે છે. Method B dictionary keys ને ઊભા સંરચનાત્મક labels તરીકે વાપરે છે અને columns પર મળતા list indices જોડીને rows બનાવે છે.

Quiz

જો તમે બે keys ધરાવતી એક dictionary વાપરીને એક DataFrame initialize કરો, જ્યાં key 'A' લંબાઈ 5 ની list પર map થાય અને key 'B' લંબાઈ 4 ની list પર map થાય, તો શું થાય છે?

  1. Pandas ખૂટતા cell ને આપોઆપ 0 પર map કરે છે.
  2. Pandas ખૂટતા cell ને એક NaN null object flag પર map કરે છે.
  3. એ તરત એક ValueError ફેંકે છે જે જણાવે છે કે arrays સરખી લંબાઈની હોવી જોઈએ.
  4. એ એક 2D table ને બદલે એક 1D Series array container બનાવે છે.
Show the answer

એ તરત એક ValueError ફેંકે છે જે જણાવે છે કે arrays સરખી લંબાઈની હોવી જોઈએ.

Data ને સળંગ rows અને columns માં સલામત રીતે ગોઠવવા, pd.DataFrame() માં અપાયેલી dictionary માં એકસરખી લંબાઈની lists હોવી જ જોઈએ. કોઈ પણ પરિમાણનું અસંતુલન data બગડતું અટકાવવા તરત એક ValueError: All arrays must be of the same length ફેંકે છે.

Quiz

જ્યારે એક સ્પષ્ટ columns argument આપ્યા વગર, data = [[10, 20], [30, 40]] જેવી એક nested list of lists માંથી સીધી એક DataFrame બનાવો, ત્યારે output કયા headers દેખાડશે?

  1. એ આપોઆપ labels Column1 અને Column2 વાપરે છે.
  2. એ column headers તરીકે integer index markers 0 અને 1 વાપરે છે.
  3. એ એક MissingArgumentException ફેંકે છે.
  4. એ પહેલી અંદરની sub-list [10, 20] ને headers block તરીકે ગણે છે.
Show the answer

એ column headers તરીકે integer index markers 0 અને 1 વાપરે છે.

જો કોઈ custom column labels ન અપાય, Pandas column labels ને બિલકુલ row indexes ની જેમ સંભાળે છે: એ 0 થી શરૂ થતા એક integer sequence પર default થાય છે. Columns પાનાની નીચે તરફ 0 અને 1 તરીકે labeled થશે.

Watch out

Classic ફાંદો: Row-Wise Dictionary ની ધારણા

university examinations માં સૌથી સામાન્ય marks-ગુમાવતી ભૂલ એ છે constructor માં lists ની એક dictionary આપવી અને keys rows બનાવશે એવી અપેક્ષા રાખવી. યાદ રાખો: dictionary keys હંમેશા ઊભા column names બનાવે છે. જો તમારે dictionary keys ને એને બદલે rows તરીકે જોઈએ, તો તમારે તમારો data DataFrame.from_dict(data, orient='index') જેવી વૈકલ્પિક ingestion પદ્ધતિઓ દ્વારા આપવો પડશે!

Theory

Ingestion ને Semester 3 સાથે જોડવા

Memory primitives માંથી datasets સંરચિત કરવું આધુનિક software integrations માટે એક આવશ્યક કૌશલ્ય છે. Semester 3 (BCA303/BCA304) માં, જ્યારે API connections માંથી web data payloads (JSON logs) parse કરશો, ત્યારે તમે નિયમિતપણે raw dictionaries ને Pandas DataFrames માં decode કરશો. આ તમને web data ને back-end databases માં સાચવતાં પહેલાં સાફ, index, અને વ્યવસ્થિત કરવા દે છે.

Summary

Key takeaways

  • Pandas native Python lists, matrices, કે dictionaries માંથી સાફ રીતે DataFrames બનાવે છે.
  • એક nested 2D List of Lists આપવાથી data arrays આડા અલગ rows તરીકે parse થાય છે.
  • એક Dictionary of Lists આપવાથી records ઊભા સંરચિત થાય છે, keys ને column names તરીકે વાપરતાં.
  • એક ValueError ટાળવા input dictionary માંના બધા element sequences એકસરખી લંબાઈના હોવા જોઈએ.
  • ખૂટતા index કે column arguments 0 થી શરૂ થતા એક integer sequence પર default થાય છે.
  • Memory Hook: Nested lists આડી rows તરીકે ખડકાય છે; dictionary keys ઊભા columns તરીકે ઊતરે છે!

Study this properly

This page is the lesson to read. In Gri-Learn the same topic is a graded deck: the self-checks are scored and your weak topics are tracked. Free to start.

Start this topic

Already have an account? Sign in

More from Python Libraries

Gri-Learn · syllabus-mapped B.C.A. lessons in English, Hindi and Gujarati

Creating dataframe using list / dict of equal length lists · Programming Skills · Gri-Learn