Theory
Automating the boring part
Doing EDA by hand on the placement dataset's 500 rows, counting missing values, averaging each column, is impossible to do manually and tedious to script from scratch. Python's data libraries automate it, turning pages of work into a couple of lines.
Two libraries do the job: pandas, which gives you a smart table called the DataFrame, and numpy, which powers the fast numeric maths beneath it. You met pandas briefly earlier; here you use it to automate exploratory analysis. This lesson shows the two commands that summarise a whole dataset instantly.
Theory
pandas and numpy
pandas is the workhorse of data analysis in Python. Its central object is the DataFrame: a labelled two-dimensional table, rows and named columns, exactly like a spreadsheet in code. You load a CSV file straight into one with pd.read_csv().
numpy sits underneath, providing fast numeric arrays and the mathematical operations that pandas relies on. You will often use numpy directly for calculations, but much of the time pandas quietly uses it for you. Together they let you load, clean, summarise, and analyse data with very little code, which is why nearly all Python data work uses them.
Practical
Loading and summarising the placement data
import pandas as pd
import numpy as np
# Load the CSV straight into a DataFrame (a table in code)
df = pd.read_csv("placement.csv")
df.info() # columns, their data types, and non-null (missing) counts
df.describe() # count, mean, std, min, quartiles, max for numeric columns
# describe() output (numeric columns) looks like:
# marks projects
# mean 71.0 2.2
# min 55.0 1.0
# max 90.0 4.0This example runs in Gri-Learn on the web, where you can edit it and see the output.
Formula
info() and describe(): instant EDA
Two methods do most of your first-pass EDA. df.info() lists every column, its data type, and how many non-null values it has, so missing data jumps out immediately (a column with fewer non-nulls than rows has gaps).
df.describe() computes summary statistics, count, mean, standard deviation, min, quartiles, max, for every numeric column at once. That is univariate analysis for the whole dataset in a single call. Between them, info() tells you the structure and gaps, and describe() tells you the numeric shape, in seconds.
Quiz
You want a one-line summary of the mean, min, max, and quartiles of every numeric column in your DataFrame. Which pandas method gives it?
- df.info(), because it prints everything
- df.describe(), which computes summary statistics for numeric columns
- pd.read_csv(), because it loads the data
- numpy, because it does maths
Show the answer
df.describe(), which computes summary statistics for numeric columns
df.describe() returns summary statistics, count, mean, standard deviation, min, quartiles, and max, for each numeric column in one call, exactly the requested summary. Option A, df.info(), shows structure instead: the columns, their data types, and non-null counts (great for spotting missing data), but not the mean/min/max/quartile statistics. Option C, pd.read_csv(), loads a CSV into a DataFrame; it does not summarise. Option D, numpy, provides low-level numeric arrays and maths but is not the one-line summary method for a DataFrame. Pair them in practice: info() for structure and gaps, describe() for the numeric summary.
Think first
Why use pandas and numpy instead of plain Python loops?
You could loop over rows with ordinary Python. Why are pandas and numpy so much better for data work? Then tap.
Show the answer
Because they are FAST, CONCISE, and built specifically for data, whereas plain Python loops are slow and verbose for large datasets. Under the hood, numpy stores numbers in compact arrays and performs operations on the whole array at once (vectorised operations implemented in optimised low-level code), so adding a column of a million numbers, or computing their mean, happens in one quick step instead of a slow Python loop iterating a million times. pandas builds on this to give you the DataFrame, where a whole analysis, filtering rows, grouping, summarising, joining tables, is expressed in a line or two of clear code that also runs fast. Writing the same thing with hand-rolled Python loops would be many more lines, far slower, and much more error-prone. There is also a rich ecosystem: reading CSVs and Excel, handling missing values, plotting, and feeding data into machine-learning libraries all integrate smoothly with pandas. So the reason essentially every Python data analyst reaches for pandas and numpy is that they turn slow, fiddly, low-level work into fast, readable, high-level operations, letting you focus on the analysis rather than the plumbing. Right tools for data: expressive and fast.
Summary
Key takeaways
- pandas and numpy are the core Python libraries for data analysis, automating EDA.
- pandas provides the DataFrame: a labelled 2D table (like a spreadsheet in code); load a CSV with pd.read_csv().
- numpy provides fast numeric arrays and maths that pandas is built on.
- df.info() shows each column's data type and non-null count, making missing data visible.
- df.describe() gives count, mean, std, min, quartiles, and max for all numeric columns in one call, automating univariate EDA.
- They are fast and concise because numpy operates on whole arrays at once (vectorised), unlike slow Python loops.
- Memory hook: pandas DataFrame is the table; info() for structure and gaps, describe() for the numeric summary.