Python Libraries to Automate EDA: Pandas and NumPy

Two Python libraries do the heavy lifting of data analysis: pandas gives you the DataFrame, a smart table you can load a CSV into and summarise in one line, and numpy powers fast numeric work underneath it.

10 min read · 7 cards · 2 checks

Read in: English · हिन्दी · ગુજરાતી


Theory

Automating the boring part

Doing EDA by hand on the placement dataset's 500 rows, counting missing values, averaging each column, is impossible to do manually and tedious to script from scratch. Python's data libraries automate it, turning pages of work into a couple of lines.

Two libraries do the job: pandas, which gives you a smart table called the DataFrame, and numpy, which powers the fast numeric maths beneath it. You met pandas briefly earlier; here you use it to automate exploratory analysis. This lesson shows the two commands that summarise a whole dataset instantly.

Theory

pandas and numpy

pandas is the workhorse of data analysis in Python. Its central object is the DataFrame: a labelled two-dimensional table, rows and named columns, exactly like a spreadsheet in code. You load a CSV file straight into one with pd.read_csv().

numpy sits underneath, providing fast numeric arrays and the mathematical operations that pandas relies on. You will often use numpy directly for calculations, but much of the time pandas quietly uses it for you. Together they let you load, clean, summarise, and analyse data with very little code, which is why nearly all Python data work uses them.

Practical

Loading and summarising the placement data

import pandas as pd
import numpy as np

# Load the CSV straight into a DataFrame (a table in code)
df = pd.read_csv("placement.csv")

df.info()        # columns, their data types, and non-null (missing) counts
df.describe()    # count, mean, std, min, quartiles, max for numeric columns

# describe() output (numeric columns) looks like:
#        marks  projects
# mean    71.0       2.2
# min     55.0       1.0
# max     90.0       4.0

This example runs in Gri-Learn on the web, where you can edit it and see the output.

Formula

info() and describe(): instant EDA

Two methods do most of your first-pass EDA. df.info() lists every column, its data type, and how many non-null values it has, so missing data jumps out immediately (a column with fewer non-nulls than rows has gaps).

df.describe() computes summary statistics, count, mean, standard deviation, min, quartiles, max, for every numeric column at once. That is univariate analysis for the whole dataset in a single call. Between them, info() tells you the structure and gaps, and describe() tells you the numeric shape, in seconds.

Quiz

You want a one-line summary of the mean, min, max, and quartiles of every numeric column in your DataFrame. Which pandas method gives it?

  1. df.info(), because it prints everything
  2. df.describe(), which computes summary statistics for numeric columns
  3. pd.read_csv(), because it loads the data
  4. numpy, because it does maths
Show the answer

df.describe(), which computes summary statistics for numeric columns

df.describe() returns summary statistics, count, mean, standard deviation, min, quartiles, and max, for each numeric column in one call, exactly the requested summary. Option A, df.info(), shows structure instead: the columns, their data types, and non-null counts (great for spotting missing data), but not the mean/min/max/quartile statistics. Option C, pd.read_csv(), loads a CSV into a DataFrame; it does not summarise. Option D, numpy, provides low-level numeric arrays and maths but is not the one-line summary method for a DataFrame. Pair them in practice: info() for structure and gaps, describe() for the numeric summary.

Think first

Why use pandas and numpy instead of plain Python loops?

You could loop over rows with ordinary Python. Why are pandas and numpy so much better for data work? Then tap.

Show the answer

Because they are FAST, CONCISE, and built specifically for data, whereas plain Python loops are slow and verbose for large datasets. Under the hood, numpy stores numbers in compact arrays and performs operations on the whole array at once (vectorised operations implemented in optimised low-level code), so adding a column of a million numbers, or computing their mean, happens in one quick step instead of a slow Python loop iterating a million times. pandas builds on this to give you the DataFrame, where a whole analysis, filtering rows, grouping, summarising, joining tables, is expressed in a line or two of clear code that also runs fast. Writing the same thing with hand-rolled Python loops would be many more lines, far slower, and much more error-prone. There is also a rich ecosystem: reading CSVs and Excel, handling missing values, plotting, and feeding data into machine-learning libraries all integrate smoothly with pandas. So the reason essentially every Python data analyst reaches for pandas and numpy is that they turn slow, fiddly, low-level work into fast, readable, high-level operations, letting you focus on the analysis rather than the plumbing. Right tools for data: expressive and fast.

Summary

Key takeaways

  • pandas and numpy are the core Python libraries for data analysis, automating EDA.
  • pandas provides the DataFrame: a labelled 2D table (like a spreadsheet in code); load a CSV with pd.read_csv().
  • numpy provides fast numeric arrays and maths that pandas is built on.
  • df.info() shows each column's data type and non-null count, making missing data visible.
  • df.describe() gives count, mean, std, min, quartiles, and max for all numeric columns in one call, automating univariate EDA.
  • They are fast and concise because numpy operates on whole arrays at once (vectorised), unlike slow Python loops.
  • Memory hook: pandas DataFrame is the table; info() for structure and gaps, describe() for the numeric summary.

Study this properly

This page is the lesson to read. In Gri-Learn the same topic is a graded deck: the self-checks are scored and your weak topics are tracked. Free to start.

Start this topic

Already have an account? Sign in

More from Automate EDA (Exploratory Data Analysis)

Gri-Learn · syllabus-mapped B.C.A. lessons in English, Hindi and Gujarati

Python Libraries to Automate EDA: Pandas and NumPy · Data Analytics using Python (Major-14) · Gri-Learn