DataFrame functions: head, tail, loc, iloc, value, to_numpy(), describe()

head/tail एक DataFrame के दोनों सिरों पर झलक देते हैं, describe() इसे statistically summarise करता है, to_numpy() इसे raw array तक strip करता है, और loc/iloc जोड़ी LABEL से बनाम POSITION से select करती है: वह distinction जिस पर बाक़ी सब कुछ टिका है।

10 min read · 10 cards · 2 checks

Read in: English · हिन्दी · ગુજરાતી


Theory

साठ हज़ार rows, पाँच काम की झलकियाँ

ResultDesk अब university-wide CSV load करता है: साठ हज़ार rows। print(df) किसी film के credits की तरह scroll होता चला जाता है।

Analysts कभी पूरी tables नहीं पढ़ते। वे झलक लेते हैं: पहली rows (क्या सही load हुआ?), आख़िरी rows (क्या सही ख़त्म हुआ?), एक statistical summary (यह दिखता कैसा है?), और फिर वे exact cells की तरफ़ बढ़ते हैं।

पाँच functions यह सब cover करते हैं, और इनमें से दो, loc और iloc, एक distinction रखते हैं जिसे exams हर साल इस्तेमाल करते हैं।

Theory

Flat number बनाम बाएँ से तीसरा दरवाज़ा

एक building में एक घर ढूँढने के दो तरीक़े:

  • Label से: "Flat 203": दरवाज़े पर लिखा number। यह loc है: यह index labels और column names पढ़ता है।
  • Position से: "दूसरी मंज़िल, तीसरा दरवाज़ा": entrance से गिनते हुए। यह iloc (integer-location) है: यह 0, 1, 2 गिनता है चाहे दरवाज़ों पर कुछ भी लिखा हो।

एक ही building, दो addressing systems, और वे तभी agree करते हैं जब दरवाज़े संयोग से क्रम में numbered हों।

Theory

Glancing toolkit

  • df.head(): पहली 5 rows (दस के लिए head(10)): क्या सही load हुआ यह check।
  • df.tail(): आख़िरी 5: टूटी हुई final lines और import leftovers पकड़ता है।
  • df.describe(): हर numeric column के लिए count, mean, std, min, quartiles (25%, 50%, 75%), max: पिछले lesson की statistics, एक call में।
  • df.to_numpy(): नीचे का raw numpy array, labels उतारे हुए (पुरानी spelling .values अभी भी किताबों में दिखती है: वही idea)।

चारों read-only झलकियाँ हैं: इनमें से कोई भी DataFrame को नहीं बदलता।

Practical

झलकियाँ, फिर exact addresses

import pandas as pd
df = pd.read_csv('marks.csv')      # index: 0, 1, 2, ...

print(df.head())        # first 5 rows
print(df.tail(3))       # last 3
print(df.describe())    # the statistical X-ray
print(df.to_numpy())    # raw array, labels gone (aka .values)

# loc: BY LABEL (index label, column NAME)
print(df.loc[0, 'score'])            # row labelled 0, column 'score'
print(df.loc[0:3, ['roll', 'score']])  # labels 0..3 INCLUSIVE: 4 rows!

# iloc: BY POSITION (counting from 0)
print(df.iloc[0, 2])                 # first row, third column
print(df.iloc[0:3])                  # positions 0,1,2: 3 rows (end EXCLUDED)

# loc happily takes last lesson's masks too:
print(df.loc[df['score'] > 60, ['roll', 'score']])

This example runs in Gri-Learn on the web, where you can edit it and see the output.

At a glance

loc बनाम iloc, exam table

Aspectlociloc
Select करता हैLabel से (index value, column name)Integer position से (0, 1, 2...)
Slice 0:3 देता है4 rows (दोनों ends INCLUDED)3 rows (end EXCLUDED)
Column formdf.loc[0, 'score']df.iloc[0, 2]
Boolean masksहाँनहीं (सिर्फ़ positions)

Think first

Four-versus-three trap

df का default index 0,1,2,...,9 है। tap करने से पहले दोनों का जवाब दीजिए: df.loc[0:3] कितनी rows return करता है, df.iloc[0:3] कितनी return करता है, और वे अलग क्यों हैं?

Show the answer

loc[0:3] → 4 rows (labels 0, 1, 2, 3): label slices दोनों ends INCLUDE करती हैं, क्योंकि arbitrary labels ('Riya':'Aman'?) के साथ pandas "one past the end" compute नहीं कर सकता।

iloc[0:3] → 3 rows (positions 0, 1, 2): integer slices Python list rules follow करती हैं, end excluded।

दिखने में एक जैसी expression, अलग row counts: यही THE loc/iloc exam question है, और अब आप इसे reason के साथ जवाब दे सकते हैं, सिर्फ़ numbers के साथ नहीं।

Quiz

Filtering के बाद, top = df[df['score'] > 80], बची हुई index labels 7, 23, 41 हैं। कौन सा call reliably पहली बची row return करता है?

  1. top.iloc[0]: position zero हमेशा exist करता है; top.loc[0] KeyError देता है क्योंकि कोई row LABEL 0 वाली नहीं है
  2. top.loc[0]: loc हमेशा पहली row return करता है
  3. top.head(): filtering के बाद head इस्तेमाल नहीं किया जा सकता
  4. दोनों काम करते हैं: filtered frames पर loc और iloc interchangeable हैं
Show the answer

top.iloc[0]: position zero हमेशा exist करता है; top.loc[0] KeyError देता है क्योंकि कोई row LABEL 0 वाली नहीं है

Filtering ORIGINAL labels रखता है (7, 23, 41), तो अब कोई label 0 नहीं है: loc[0] एक door number 0 ढूँढता है और fail होता है। iloc[0] physically गिनता है: पहली row, उसका label चाहे कुछ भी हो। यह divergence filtering के बाद वहीं है जहाँ दोनों systems असहमत होना शुरू करते हैं, और option A का reasoning professional habit है: 'first/last' के लिए positions, 'row known as X' के लिए labels। (head(1) भी काम करता: option C का claim ग़लत है।)

Watch out

तीन daily-driver slips

positions के साथ loc: तब तक luck से काम करता है जब तक index 0,1,2... है, किसी भी filter या sort के बाद टूट जाता है: TYPE करने से पहले label-या-position तय कीजिए।

describe() उम्मीद से कम columns दिखाता है: यह default से सिर्फ़ numeric columns summarise करता है; text columns के लिए describe(include='object') चाहिए।

पुरानी किताबों में .values: to_numpy() जैसा ही array; दोनों नाम पहचानिए, modern वाला लिखिए।

Theory

Unit 4 बंद होती है; तस्वीरें शुरू होती हैं

अपनी pandas kit गिनिए: load और save (read_csv/to_csv), slice (brackets और masks), summarise (mean से describe तक), और exactly address करना (loc/iloc)। यही एक working analyst का core है। Unit 5 इन numbers को तस्वीरों में बदलता है: matplotlib वही Series लेता है जो आप extract करते आए हैं और class की story खींचता है: lines, bars, scatters, histograms।

Summary

Key takeaways

  • head(n)/tail(n): सिरों पर झलक; describe(): numeric columns का count-से-max तक statistical summary।
  • to_numpy() (पहले .values) DataFrame को इसके raw numpy array तक strip करता है।
  • loc LABEL से select करता है: slices दोनों ends include करती हैं; boolean masks accept करता है।
  • iloc POSITION से select करता है: slices end exclude करती हैं, Python lists जैसा।
  • Default index पर loc[0:3] = 4 rows, iloc[0:3] = 3 rows: classic exam pair।
  • Filtering के बाद, labels में gaps रह जाते हैं: 'पहली row' के लिए iloc[0], 'row labelled X' के लिए loc।
  • Memory hook: flat number बनाम बाएँ से तीसरा दरवाज़ा।

Study this properly

This page is the lesson to read. In Gri-Learn the same topic is a graded deck: the self-checks are scored and your weak topics are tracked. Free to start.

Start this topic

Already have an account? Sign in

More from Python interaction with text and CSV

Gri-Learn · syllabus-mapped B.C.A. lessons in English, Hindi and Gujarati