Theory
Raw Memory Streams નો પડકાર
તમારા Semester 1 BCA104 labs માં, તમે parallel 1D primitive arrays કે multi-dimensional nested arrays વાપરીને record tables સંઘરતા. પણ, values ના એક matrix ને parse કરવા hardcoded coordinate integer tracking જોઈતું, અને descriptive headers ઉમેરવા કે orientations બદલવા શરૂઆતથી custom layout manipulation algorithms લખવા પડતા. Pandas આ સંરચનાત્મક setup ને કેમ સરળ બનાવે છે, તમને સાફ, tabular database tables બનાવવા nested lists કે standard data dictionaries સીધા એક constructor માં આપવા દેતાં?
Theory
આડું Blueprint સામે ઊભું Filing Cabinet
raw cards માંથી એક ભૌતિક spreadsheet table જોડવાનું કલ્પો. એક List of Lists માંથી DataFrame બનાવવું એટલે લાંબા કાગળના receipt tapes ને આડા, હરોળ-દર-હરોળ પાથરવા. તમે row 1 ગોઠવો, પછી row 2 એની નીચે મૂકો, અને છેલ્લે દરેક ઊભી લેન ની ટોચે column names બનાવવા sticky labels ચોંટાડો. બીજી બાજુ, એક Dictionary of Lists માંથી DataFrame બનાવવું એટલે એક ખાલી ઊભા organizer પાસે જવું. દરેક અલગ dictionary folder એક પહેલેથી labeled ઊભા column ને દર્શાવે છે, અને તમે data slips ની એકસરખી rows સીધી દરેક ખાનામાં નીચે નાખો છો.
Theory
Data Ingestion ઔપચારિક રીતે
pd.DataFrame() constructor એક 2D grid બનાવવા વૈવિધ્યસભર સંરચિત Python sequencers સ્વીકારે છે. જ્યારે એક 2D List of Lists વાપરો, દરેક sub-list ને સ્પષ્ટપણે એક અલગ આડા row block તરીકે ગણવામાં આવે છે; તમે optional columns list parameter વાપરીને column name labels આપો છો. એથી ઊલટું, જ્યારે એક Dictionary of Lists વાપરો, દરેક dictionary key ને એક ઊભા column header તરીકે ગણવામાં આવે છે, અને એની સાથે જોડાયેલી list એ field ના elements પાનાની નીચે તરફ રાખે છે. આ પદ્ધતિ માંગે છે કે દરેક list value એક કડક equal-length sequence constraint પાળે.
At a glance
Table 1: Pandas માં row-wise list layouts અને column-wise dictionary layouts વચ્ચેના સંરચનાત્મક ભેદ.
| Data Structure નો સ્રોત | Constructor નો Parsing નિયમ | જરૂરી Shape Parameters |
|---|---|---|
| [[Row1], [Row2]] | Nested components ને આડા rows તરીકે parse કરે છે. | Optional columns=['A', 'B'] parameter સ્વીકારે છે. |
| {'Col1': [Val1, Val2]} | Values ને ઊભા સ્વતંત્ર columns તરીકે parse કરે છે. | Keys આપોઆપ tabular column headings વ્યાખ્યાયિત કરે છે. |
| Length Mismatch | Runtime દરમિયાન સંરચનાત્મક data errors trigger કરે છે. | જો list sizes ન મળે તો એક ValueError ઊભી કરે છે. |
Theory
Worked Example: Semester Budget જોડવું
ચાલો જોઈએ કે આપણી PocketMoney tracking system બંને સંરચનાત્મક entry પદ્ધતિઓ વાપરીને એકસરખા tabular data containers કેવી રીતે બનાવે છે. આપણે પહેલાં એક nested raw list structure parse કરીશું, અને પછી એક key-mapped dictionary structure વાપરીને બિલકુલ એ જ layout બનાવીશું.
Practical
PocketMoney DataFrame Constructor
import pandas as pd
# Method A: Constructing row-by-row using a 2D List of Lists
row_matrix = [
["Samosa", 35, "Food"],
["Bus", 50, "Travel"],
["Chai", 15, "Food"]
]
col_headers = ["Item", "Cost", "Type"]
df_from_lists = pd.DataFrame(row_matrix, columns=col_headers)
# Method B: Constructing column-by-column using a Dictionary of equal-length lists
col_dictionary = {
"Item": ["Samosa", "Bus", "Chai"],
"Cost": [35, 50, 15],
"Type": ["Food", "Travel", "Food"]
}
df_from_dict = pd.DataFrame(col_dictionary)
print("--- DataFrame from List of Lists ---")
print(df_from_lists)
print("\n--- DataFrame from Dictionary ---")
print(df_from_dict)This example runs in Gri-Learn on the web, where you can edit it and see the output.
Think first
સંરચનાત્મક રૂપરેખાઓ સરખાવો
બંને compilation outputs મનમાં તપાસો. terminal પર છપાય ત્યારે સંરચનાત્મક shapes, columns, અને આંતરિક coordinate alignments અલગ દેખાશે?
Show the answer
script output કરશે:
--- DataFrame from List of Lists ---
Item Cost Type
0 Samosa 35 Food
1 Bus 50 Travel
2 Chai 15 Food
--- DataFrame from Dictionary ---
Item Cost Type
0 Samosa 35 Food
1 Bus 50 Travel
2 Chai 15 Food
કેમ? બંને creation અભિગમો એકસરખા 2D relational matrices આપે છે. Method A આડા coordinates ને સ્પષ્ટપણે rows તરીકે map કરે છે અને એક અલગ parameter દ્વારા headers સોંપે છે. Method B dictionary keys ને ઊભા સંરચનાત્મક labels તરીકે વાપરે છે અને columns પર મળતા list indices જોડીને rows બનાવે છે.
Quiz
જો તમે બે keys ધરાવતી એક dictionary વાપરીને એક DataFrame initialize કરો, જ્યાં key 'A' લંબાઈ 5 ની list પર map થાય અને key 'B' લંબાઈ 4 ની list પર map થાય, તો શું થાય છે?
- Pandas ખૂટતા cell ને આપોઆપ 0 પર map કરે છે.
- Pandas ખૂટતા cell ને એક NaN null object flag પર map કરે છે.
- એ તરત એક ValueError ફેંકે છે જે જણાવે છે કે arrays સરખી લંબાઈની હોવી જોઈએ.
- એ એક 2D table ને બદલે એક 1D Series array container બનાવે છે.
Show the answer
એ તરત એક ValueError ફેંકે છે જે જણાવે છે કે arrays સરખી લંબાઈની હોવી જોઈએ.
Data ને સળંગ rows અને columns માં સલામત રીતે ગોઠવવા, pd.DataFrame() માં અપાયેલી dictionary માં એકસરખી લંબાઈની lists હોવી જ જોઈએ. કોઈ પણ પરિમાણનું અસંતુલન data બગડતું અટકાવવા તરત એક ValueError: All arrays must be of the same length ફેંકે છે.
Quiz
જ્યારે એક સ્પષ્ટ columns argument આપ્યા વગર, data = [[10, 20], [30, 40]] જેવી એક nested list of lists માંથી સીધી એક DataFrame બનાવો, ત્યારે output કયા headers દેખાડશે?
- એ આપોઆપ labels Column1 અને Column2 વાપરે છે.
- એ column headers તરીકે integer index markers 0 અને 1 વાપરે છે.
- એ એક MissingArgumentException ફેંકે છે.
- એ પહેલી અંદરની sub-list [10, 20] ને headers block તરીકે ગણે છે.
Show the answer
એ column headers તરીકે integer index markers 0 અને 1 વાપરે છે.
જો કોઈ custom column labels ન અપાય, Pandas column labels ને બિલકુલ row indexes ની જેમ સંભાળે છે: એ 0 થી શરૂ થતા એક integer sequence પર default થાય છે. Columns પાનાની નીચે તરફ 0 અને 1 તરીકે labeled થશે.
Watch out
Classic ફાંદો: Row-Wise Dictionary ની ધારણા
university examinations માં સૌથી સામાન્ય marks-ગુમાવતી ભૂલ એ છે constructor માં lists ની એક dictionary આપવી અને keys rows બનાવશે એવી અપેક્ષા રાખવી. યાદ રાખો: dictionary keys હંમેશા ઊભા column names બનાવે છે. જો તમારે dictionary keys ને એને બદલે rows તરીકે જોઈએ, તો તમારે તમારો data DataFrame.from_dict(data, orient='index') જેવી વૈકલ્પિક ingestion પદ્ધતિઓ દ્વારા આપવો પડશે!
Theory
Ingestion ને Semester 3 સાથે જોડવા
Memory primitives માંથી datasets સંરચિત કરવું આધુનિક software integrations માટે એક આવશ્યક કૌશલ્ય છે. Semester 3 (BCA303/BCA304) માં, જ્યારે API connections માંથી web data payloads (JSON logs) parse કરશો, ત્યારે તમે નિયમિતપણે raw dictionaries ને Pandas DataFrames માં decode કરશો. આ તમને web data ને back-end databases માં સાચવતાં પહેલાં સાફ, index, અને વ્યવસ્થિત કરવા દે છે.
Summary
Key takeaways
- Pandas native Python lists, matrices, કે dictionaries માંથી સાફ રીતે DataFrames બનાવે છે.
- એક nested 2D List of Lists આપવાથી data arrays આડા અલગ rows તરીકે parse થાય છે.
- એક Dictionary of Lists આપવાથી records ઊભા સંરચિત થાય છે, keys ને column names તરીકે વાપરતાં.
- એક ValueError ટાળવા input dictionary માંના બધા element sequences એકસરખી લંબાઈના હોવા જોઈએ.
- ખૂટતા index કે column arguments 0 થી શરૂ થતા એક integer sequence પર default થાય છે.
- Memory Hook: Nested lists આડી rows તરીકે ખડકાય છે; dictionary keys ઊભા columns તરીકે ઊતરે છે!