Pandas overview: Dataframe and Series

Pandas અવ્યવસ્થિત data logs ને સાફ, excel-જેવા સંરચિત grids માં ફેરવે છે જે તમને multidimensional datasets તરત analyze કરવા દે છે.

10 min read · 12 cards · 3 checks

Read in: English · हिन्दी · ગુજરાતી


Theory

Matrix ની ઓળખની કટોકટી

તમારા Semester 2 NumPy labs માં, તમે શીખ્યા કે homogeneous arrays વાપરીને હજારો numeric transaction amounts ને વીજળીની ઝડપે કેવી રીતે process કરવા. પણ શું થાય જ્યારે તમારા campus pocket-money tracker ને એકસાથે તદ્દન અલગ અલગ data types સાચવવા હોય, જેમ કે categories માટે text labels ('Food', 'Chai'), item counts માટે integers, floating decimal percentages, અને true/false validation flags? કારણ કે એક NumPy array દરેક cell ને એક જ data type માં જકડે છે, mixed strings અને numbers આપવાથી બધું જ raw text strings બની જાય છે, તમારાં math tools સંપૂર્ણપણે બગાડી નાખતાં. તો આપણે એક high-performance tabular layout કેવી રીતે બનાવીએ જે એક professional Excel sheet ની જેમ rows, columns, અને mixed data types સંભાળે?

Theory

એક-Column Ruler સામે Multi-Department Ledger

એક Pandas Series ને એક student attendance register ની અંદર છપાયેલા એક ઊભા column જેવી વિચારો. એ એક લીટીમાં નીચે તરફ data points યાદી કરે છે, પણ દરેક લીટી માટે એની પાસે એક custom રંગેલું label હોય છે (જેમ કે raw index positions 0, 1, 2 ને બદલે student roll numbers). એક Pandas DataFrame આખું multi-page ભૌતિક ledger પુસ્તક છે. એ અનેક columns ને પડખોપડખ જોડે છે. દરેક column તદ્દન અલગ વસ્તુઓ રાખી શકે છે, એકમાં text names હોય, બીજામાં numeric balances, અને ત્રીજામાં pass/fail marks. છતાં, એ બધા ડાબા હાંસિયામાં નીચે તરફના એ જ અનન્ય આડા row index labels સાથે બરાબર જોડાયેલા રહે છે.

Theory

Pandas Architecture ઔપચારિક રીતે

Pandas library વ્યવહારુ data analysis માટે રચાયેલી high-performance, વાપરવામાં સરળ data structures આપે છે. પહેલી મુખ્ય structure એક Series છે, જે એક એક-પરિમાણીય labeled array છે જે કોઈ પણ data type રાખી શકે છે. બીજી અને સૌથી વધુ વપરાતી structure DataFrame છે, એક દ્વિ-પરિમાણીય, કદમાં બદલી શકાય એવી, અને heterogeneous tabular data structure જેમાં labeled axes (rows અને columns) હોય છે. NumPy થી વિપરીત, એક DataFrame માં columns અલગ અલગ heterogeneous data types (integers, floats, objects) રાખે છે, જ્યારે rows પર કડક index alignment જાળવે છે.

At a glance

Table 1: Pandas Series અને DataFrame architectures ની મુખ્ય સંરચનાત્મક સરખામણી.

Pandas Data Layerપરિમાણીય LayoutType Homogeneity ની મર્યાદા
pd.Series(data, index)1D: custom label mappings સાથે એક જ column.Homogeneous: એક Series ની અંદરના દરેક cell નો type મળતો હોવો જોઈએ.
pd.DataFrame(data)2D: સ્વતંત્ર columns સાથે એક પૂરું tabular matrix.Heterogeneous: Columns તદ્દન અલગ types સંઘરી શકે છે.
df.columnsMetadata: આડા field headings ખુલ્લા કરે છે.અપરિવર્તનીય text labels કે coordinate sequence descriptors.
df.indexMetadata: ઊભી row alignment સીમાઓ track કરે છે.integers, text strings, timestamps, કે roll numbers હોઈ શકે છે.

Theory

Worked Example: Tabular Ledger નું Blueprint

ચાલો જોઈએ કે આપણો ચાલુ project, PocketMoney, mixed transaction logs ને model કરવા Pandas components કેવી રીતે વાપરે છે. આપણે પહેલાં એક અલગ 1D Series બનાવીશું, અને પછી અનેક વૈવિધ્યસભર sequences ને એક એકીકૃત 2D DataFrame grid માં જોડીશું.

Practical

PocketMoney Pandas Blueprint

import pandas as pd

# Step 1: Construct an isolated 1D Series tracking category types
categories = pd.Series(["Samosa", "Bus", "Chai"], index=["t1", "t2", "t3"])

# Step 2: Assemble structured dictionary metrics with mixed data types
raw_data = {
    "Category": ["Samosa", "Bus", "Chai", "Stationery"],
    "Amount": [35, 50, 15, 120],
    "Essential": [False, True, False, True]
}

# Step 3: Compile into a heterogeneous 2D DataFrame table
ledger_df = pd.DataFrame(raw_data)

print("--- 1D Series Output ---")
print(categories)
print("\n--- 2D DataFrame Table ---")
print(ledger_df)
print("\nData Types Per Column:\n", ledger_df.dtypes)

This example runs in Gri-Learn on the web, where you can edit it and see the output.

Think first

સંરચિત Alignment ટ્રેસ કરો

ledger_df નું છપાયેલું output layout તપાસો. જો DataFrame constructor ને કોઈ સ્પષ્ટ index array ન અપાય તો Pandas rows ને આપોઆપ કેવી રીતે format કરે છે? કયા data types તારવાય છે?

Show the answer

script output કરશે:

--- 1D Series Output ---

t1 Samosa

t2 Bus

t3 Chai

dtype: object

--- 2D DataFrame Table ---

Category Amount Essential

0 Samosa 35 False

1 Bus 50 True

2 Chai 15 False

3 Stationery 120 True

Data Types Per Column:

Category object

Amount int64

Essential bool

dtype: object

કેમ? જ્યારે DataFrame માટે કોઈ સ્પષ્ટ row index નિર્દિષ્ટ ન હોય, Pandas આપોઆપ 0 થી શરૂ થતો એક standard numeric sequence generate કરે છે. દરેક column પોતાનો મૂળ data type બરાબર જાળવે છે, Category એક 'object' (text string) પર map થાય છે, Amount 'int64' પર map થાય છે, અને Essential 'bool' પર map થાય છે, કોઈ પણ સલામત mathematical properties ને બગડતી અટકાવતાં.

Quiz

એક NumPy ndarray અને એક Pandas DataFrame structure વચ્ચેનો મુખ્ય સંરચનાત્મક ભેદ શું છે?

  1. એક DataFrame માત્ર floating-point numbers રાખી શકે છે, જ્યારે એક ndarray characters રાખે છે.
  2. એક ndarray સખ્તાઈથી homogeneous છે (એક જ સહિયારો type), જ્યારે એક DataFrame heterogeneous છે (columns પાસે તદ્દન સ્વતંત્ર data types હોઈ શકે છે).
  3. એક ndarray columns પર automated text labels ને support કરે છે, જ્યારે એક DataFrame સંપૂર્ણપણે numerical indices પર આધાર રાખે છે.
  4. DataFrames memory માં initialize થયા પછી update કે resize થઈ શકતી નથી.
Show the answer

એક ndarray સખ્તાઈથી homogeneous છે (એક જ સહિયારો type), જ્યારે એક DataFrame heterogeneous છે (columns પાસે તદ્દન સ્વતંત્ર data types હોઈ શકે છે).

NumPy arrays execution speed મેળવવા memory માં દરેકે દરેક element પર સંપૂર્ણ data type એકરૂપતા લાદે છે. Pandas DataFrames tabular databases ની જેમ કામ કરે છે, તમને પડખેપડખના columns માં text, integers, અને booleans ને એમની અલગ ઓળખ ગુમાવ્યા વગર સાફ રીતે ભેળવવાની છૂટ આપતાં.

Quiz

જો તમે એક multi-column Pandas DataFrame માંથી કાઢેલા એક જ column નો આંતરિક type જુઓ, તો Python કયો data object type પાછો આપશે?

  1. એક standard primitive Python list
  2. એક જ અલગ NumPy scalar integer
  3. એક Pandas Series object
  4. એક સપાટ text string file string
Show the answer

એક Pandas Series object

એક Pandas DataFrame સંરચનાત્મક રીતે એક સહિયારા આડા index row map પર ગોઠવાયેલા અનેક અલગ columns ની બનેલી છે. એક DataFrame માંથી કાઢેલો દરેક અલગ column એક 1D labeled sequence તરીકે કામ કરે છે, જે એક Pandas Series object છે.

Watch out

Classic ફાંદો: સરખી લંબાઈના Columns નું પતન

એક dictionary માંથી DataFrame બનાવતી વખતે university lab exams માં સૌથી વારંવારની ભૂલ એ છે એવી lists આપવી જેમની લંબાઈ મળતી ન હોય (દા.ત. Column A પાસે 4 values છે, પણ Column B પાસે માત્ર 3). Python ખાલી જગ્યાઓ zeros કે None થી ભરશે નહીં; એને બદલે, constructor તરત ValueError: All arrays must be of the same length સાથે crash થશે. Initialize કરતાં પહેલાં તમારી list counts બે વાર તપાસો!

Theory

Tables ને Semester 3 સાથે જોડવા

Tabular index tracking automated industrial analytical pipelines નું મુખ્ય માળખું બનાવે છે. Semester 3 Data Science (BCA302) અને Web Architecture (BCA303) માં, તમે raw transactional databases સાફ કરવા, null anomalies filter કરવા, અને predictive model engines માં સીધા multi-type inputs આપવા DataFrames ને memory tables તરીકે વાપરશો.

Summary

Key takeaways

  • Pandas heterogeneous mixed types સહેલાઈથી સંભાળવા સંરચનાત્મક data objects રજૂ કરે છે.
  • એક Series સ્પષ્ટ user-defined tracking labels સાથે જોડાયેલો એક 1D sequence દર્શાવે છે.
  • એક DataFrame એક 2D tabular layout છે જેમાં સહિયારા index alignments સાથે અલગ columns હોય છે.
  • જો કોઈ custom labels ન અપાય તો DataFrames આપોઆપ 0 થી શરૂ થતો એક integer-આધારિત index generate કરે છે.
  • એક પહોળા DataFrame table માંથી કાઢેલો દરેકે દરેક સ્વતંત્ર column એક Series તરીકે કામ કરે છે.
  • Memory Hook: એક Series એક જ labeled column છે; એક DataFrame આખું heterogeneous multi-column ledger છે!

Study this properly

This page is the lesson to read. In Gri-Learn the same topic is a graded deck: the self-checks are scored and your weak topics are tracked. Free to start.

Start this topic

Already have an account? Sign in

More from Python Libraries

Gri-Learn · syllabus-mapped B.C.A. lessons in English, Hindi and Gujarati

Pandas overview: Dataframe and Series · Programming Skills · Gri-Learn