Theory
Matrix ની ઓળખની કટોકટી
તમારા Semester 2 NumPy labs માં, તમે શીખ્યા કે homogeneous arrays વાપરીને હજારો numeric transaction amounts ને વીજળીની ઝડપે કેવી રીતે process કરવા. પણ શું થાય જ્યારે તમારા campus pocket-money tracker ને એકસાથે તદ્દન અલગ અલગ data types સાચવવા હોય, જેમ કે categories માટે text labels ('Food', 'Chai'), item counts માટે integers, floating decimal percentages, અને true/false validation flags? કારણ કે એક NumPy array દરેક cell ને એક જ data type માં જકડે છે, mixed strings અને numbers આપવાથી બધું જ raw text strings બની જાય છે, તમારાં math tools સંપૂર્ણપણે બગાડી નાખતાં. તો આપણે એક high-performance tabular layout કેવી રીતે બનાવીએ જે એક professional Excel sheet ની જેમ rows, columns, અને mixed data types સંભાળે?
Theory
એક-Column Ruler સામે Multi-Department Ledger
એક Pandas Series ને એક student attendance register ની અંદર છપાયેલા એક ઊભા column જેવી વિચારો. એ એક લીટીમાં નીચે તરફ data points યાદી કરે છે, પણ દરેક લીટી માટે એની પાસે એક custom રંગેલું label હોય છે (જેમ કે raw index positions 0, 1, 2 ને બદલે student roll numbers). એક Pandas DataFrame આખું multi-page ભૌતિક ledger પુસ્તક છે. એ અનેક columns ને પડખોપડખ જોડે છે. દરેક column તદ્દન અલગ વસ્તુઓ રાખી શકે છે, એકમાં text names હોય, બીજામાં numeric balances, અને ત્રીજામાં pass/fail marks. છતાં, એ બધા ડાબા હાંસિયામાં નીચે તરફના એ જ અનન્ય આડા row index labels સાથે બરાબર જોડાયેલા રહે છે.
Theory
Pandas Architecture ઔપચારિક રીતે
Pandas library વ્યવહારુ data analysis માટે રચાયેલી high-performance, વાપરવામાં સરળ data structures આપે છે. પહેલી મુખ્ય structure એક Series છે, જે એક એક-પરિમાણીય labeled array છે જે કોઈ પણ data type રાખી શકે છે. બીજી અને સૌથી વધુ વપરાતી structure DataFrame છે, એક દ્વિ-પરિમાણીય, કદમાં બદલી શકાય એવી, અને heterogeneous tabular data structure જેમાં labeled axes (rows અને columns) હોય છે. NumPy થી વિપરીત, એક DataFrame માં columns અલગ અલગ heterogeneous data types (integers, floats, objects) રાખે છે, જ્યારે rows પર કડક index alignment જાળવે છે.
At a glance
Table 1: Pandas Series અને DataFrame architectures ની મુખ્ય સંરચનાત્મક સરખામણી.
| Pandas Data Layer | પરિમાણીય Layout | Type Homogeneity ની મર્યાદા |
|---|---|---|
| pd.Series(data, index) | 1D: custom label mappings સાથે એક જ column. | Homogeneous: એક Series ની અંદરના દરેક cell નો type મળતો હોવો જોઈએ. |
| pd.DataFrame(data) | 2D: સ્વતંત્ર columns સાથે એક પૂરું tabular matrix. | Heterogeneous: Columns તદ્દન અલગ types સંઘરી શકે છે. |
| df.columns | Metadata: આડા field headings ખુલ્લા કરે છે. | અપરિવર્તનીય text labels કે coordinate sequence descriptors. |
| df.index | Metadata: ઊભી row alignment સીમાઓ track કરે છે. | integers, text strings, timestamps, કે roll numbers હોઈ શકે છે. |
Theory
Worked Example: Tabular Ledger નું Blueprint
ચાલો જોઈએ કે આપણો ચાલુ project, PocketMoney, mixed transaction logs ને model કરવા Pandas components કેવી રીતે વાપરે છે. આપણે પહેલાં એક અલગ 1D Series બનાવીશું, અને પછી અનેક વૈવિધ્યસભર sequences ને એક એકીકૃત 2D DataFrame grid માં જોડીશું.
Practical
PocketMoney Pandas Blueprint
import pandas as pd
# Step 1: Construct an isolated 1D Series tracking category types
categories = pd.Series(["Samosa", "Bus", "Chai"], index=["t1", "t2", "t3"])
# Step 2: Assemble structured dictionary metrics with mixed data types
raw_data = {
"Category": ["Samosa", "Bus", "Chai", "Stationery"],
"Amount": [35, 50, 15, 120],
"Essential": [False, True, False, True]
}
# Step 3: Compile into a heterogeneous 2D DataFrame table
ledger_df = pd.DataFrame(raw_data)
print("--- 1D Series Output ---")
print(categories)
print("\n--- 2D DataFrame Table ---")
print(ledger_df)
print("\nData Types Per Column:\n", ledger_df.dtypes)This example runs in Gri-Learn on the web, where you can edit it and see the output.
Think first
સંરચિત Alignment ટ્રેસ કરો
ledger_df નું છપાયેલું output layout તપાસો. જો DataFrame constructor ને કોઈ સ્પષ્ટ index array ન અપાય તો Pandas rows ને આપોઆપ કેવી રીતે format કરે છે? કયા data types તારવાય છે?
Show the answer
script output કરશે:
--- 1D Series Output ---
t1 Samosa
t2 Bus
t3 Chai
dtype: object
--- 2D DataFrame Table ---
Category Amount Essential
0 Samosa 35 False
1 Bus 50 True
2 Chai 15 False
3 Stationery 120 True
Data Types Per Column:
Category object
Amount int64
Essential bool
dtype: object
કેમ? જ્યારે DataFrame માટે કોઈ સ્પષ્ટ row index નિર્દિષ્ટ ન હોય, Pandas આપોઆપ 0 થી શરૂ થતો એક standard numeric sequence generate કરે છે. દરેક column પોતાનો મૂળ data type બરાબર જાળવે છે, Category એક 'object' (text string) પર map થાય છે, Amount 'int64' પર map થાય છે, અને Essential 'bool' પર map થાય છે, કોઈ પણ સલામત mathematical properties ને બગડતી અટકાવતાં.
Quiz
એક NumPy ndarray અને એક Pandas DataFrame structure વચ્ચેનો મુખ્ય સંરચનાત્મક ભેદ શું છે?
- એક DataFrame માત્ર floating-point numbers રાખી શકે છે, જ્યારે એક ndarray characters રાખે છે.
- એક ndarray સખ્તાઈથી homogeneous છે (એક જ સહિયારો type), જ્યારે એક DataFrame heterogeneous છે (columns પાસે તદ્દન સ્વતંત્ર data types હોઈ શકે છે).
- એક ndarray columns પર automated text labels ને support કરે છે, જ્યારે એક DataFrame સંપૂર્ણપણે numerical indices પર આધાર રાખે છે.
- DataFrames memory માં initialize થયા પછી update કે resize થઈ શકતી નથી.
Show the answer
એક ndarray સખ્તાઈથી homogeneous છે (એક જ સહિયારો type), જ્યારે એક DataFrame heterogeneous છે (columns પાસે તદ્દન સ્વતંત્ર data types હોઈ શકે છે).
NumPy arrays execution speed મેળવવા memory માં દરેકે દરેક element પર સંપૂર્ણ data type એકરૂપતા લાદે છે. Pandas DataFrames tabular databases ની જેમ કામ કરે છે, તમને પડખેપડખના columns માં text, integers, અને booleans ને એમની અલગ ઓળખ ગુમાવ્યા વગર સાફ રીતે ભેળવવાની છૂટ આપતાં.
Quiz
જો તમે એક multi-column Pandas DataFrame માંથી કાઢેલા એક જ column નો આંતરિક type જુઓ, તો Python કયો data object type પાછો આપશે?
- એક standard primitive Python list
- એક જ અલગ NumPy scalar integer
- એક Pandas Series object
- એક સપાટ text string file string
Show the answer
એક Pandas Series object
એક Pandas DataFrame સંરચનાત્મક રીતે એક સહિયારા આડા index row map પર ગોઠવાયેલા અનેક અલગ columns ની બનેલી છે. એક DataFrame માંથી કાઢેલો દરેક અલગ column એક 1D labeled sequence તરીકે કામ કરે છે, જે એક Pandas Series object છે.
Watch out
Classic ફાંદો: સરખી લંબાઈના Columns નું પતન
એક dictionary માંથી DataFrame બનાવતી વખતે university lab exams માં સૌથી વારંવારની ભૂલ એ છે એવી lists આપવી જેમની લંબાઈ મળતી ન હોય (દા.ત. Column A પાસે 4 values છે, પણ Column B પાસે માત્ર 3). Python ખાલી જગ્યાઓ zeros કે None થી ભરશે નહીં; એને બદલે, constructor તરત ValueError: All arrays must be of the same length સાથે crash થશે. Initialize કરતાં પહેલાં તમારી list counts બે વાર તપાસો!
Theory
Tables ને Semester 3 સાથે જોડવા
Tabular index tracking automated industrial analytical pipelines નું મુખ્ય માળખું બનાવે છે. Semester 3 Data Science (BCA302) અને Web Architecture (BCA303) માં, તમે raw transactional databases સાફ કરવા, null anomalies filter કરવા, અને predictive model engines માં સીધા multi-type inputs આપવા DataFrames ને memory tables તરીકે વાપરશો.
Summary
Key takeaways
- Pandas heterogeneous mixed types સહેલાઈથી સંભાળવા સંરચનાત્મક data objects રજૂ કરે છે.
- એક Series સ્પષ્ટ user-defined tracking labels સાથે જોડાયેલો એક 1D sequence દર્શાવે છે.
- એક DataFrame એક 2D tabular layout છે જેમાં સહિયારા index alignments સાથે અલગ columns હોય છે.
- જો કોઈ custom labels ન અપાય તો DataFrames આપોઆપ 0 થી શરૂ થતો એક integer-આધારિત index generate કરે છે.
- એક પહોળા DataFrame table માંથી કાઢેલો દરેકે દરેક સ્વતંત્ર column એક Series તરીકે કામ કરે છે.
- Memory Hook: એક Series એક જ labeled column છે; એક DataFrame આખું heterogeneous multi-column ledger છે!