Exploratory Data Analysis (EDA): types of EDA; Univariate, Bivariate and Multivariate Analysis; handling missing data and outliers

Before you model anything, you explore: exploratory data analysis means getting to know your dataset by looking at one variable at a time, then pairs, then many together, and dealing honestly with missing values and outliers.

11 min read · 7 cards · 2 checks

Read in: English · हिन्दी · ગુજરાતી


Theory

Look before you leap

Suppose the college hands you a spreadsheet: one row per student, columns for marks, attendance, number of projects, and whether they got placed. Before you build any model to predict placement, you must actually understand this data, its shape, its quirks, its gaps.

That first step is exploratory data analysis (EDA): getting to know a dataset through summaries and visualisations before modelling it. This whole subject uses that placement dataset as its running example, and it all starts here, with looking carefully at what you have.

Theory

What EDA is

Exploratory data analysis is the practice of examining a dataset to understand its main characteristics before you apply formal models. You compute summaries (averages, ranges, counts) and draw simple charts to answer questions like: what values does each column take? How are they distributed? Which columns relate to each other? Are there gaps or strange values?

EDA is not a formality; it is where you catch problems and form hunches. Skipping it means modelling data you do not understand, a recipe for wrong conclusions. Good analysts spend serious time here first.

At a glance

TypeVariables examinedPlacement-data example
UnivariateOne at a timeThe distribution of marks on its own
BivariateTwo, and their relationshipMarks versus whether the student was placed
MultivariateThree or more togetherMarks, attendance, and projects together vs placement

Theory

Missing data and outliers

Real datasets are messy, and two problems come up constantly.

Missing data: some values are simply absent (a student's attendance not recorded). You must decide how to handle each gap: drop the affected rows or columns, or impute them, fill them in with a sensible substitute like the mean, median, or mode of that column.

Outliers: some values sit far outside the normal range (a marks entry of 500 in a 0 to 100 column, or a genuine but extreme case). You detect them and decide: is it an error to fix or remove, or a real extreme value to keep and understand? Handling both honestly is a core part of EDA.

Quiz

You examine the relationship between students' marks and whether they were placed, looking at exactly those two variables together. What type of EDA is this?

  1. Univariate, because placement is one outcome
  2. Bivariate, because you are examining two variables and their relationship
  3. Multivariate, because placement depends on everything
  4. It is not EDA; it is modelling
Show the answer

Bivariate, because you are examining two variables and their relationship

Examining exactly TWO variables (marks and placement) and how they relate is bivariate analysis. Option A is wrong: univariate looks at just ONE variable in isolation (say, the spread of marks alone), but here you are relating two. Option C, multivariate, involves THREE or more variables together (for example marks, attendance, AND projects at once); this example has only two. Option D is wrong: exploring relationships between variables through summaries and charts is precisely EDA, not formal modelling. Count the variables you examine together: one is univariate, two is bivariate, three or more is multivariate.

Think first

Why not just delete every row with a missing value?

Dropping rows with missing data is easy. Why is imputing (filling values in) often better? Then tap.

Show the answer

Because dropping rows can THROW AWAY a lot of valuable data and even BIAS your results, so imputation is often the wiser choice. Imagine the placement dataset has 500 students, and attendance is missing for 120 of them, perhaps because one department recorded it differently. If you delete every row with any missing value, you might lose those 120 students entirely, shrinking your data by nearly a quarter and losing all the other good information those rows contained (their marks, projects, placement outcome). Worse, if the missingness is not random, say attendance is missing mostly for one department, then dropping those rows quietly BIASES your analysis toward the departments that did record it, distorting your conclusions. Imputation, filling the gaps with a reasonable value like the column's mean, median, or mode (or a smarter estimate), lets you keep the rest of each row's useful information while making a principled guess for the missing piece. It is not always right either, imputation introduces its own assumptions, so you choose based on how much is missing and why. But blindly deleting is rarely best. The professional habit is to understand WHY data is missing, then decide row-by-column whether to drop or impute, rather than reflexively discarding. Keep the information you can; fill gaps thoughtfully.

Summary

Key takeaways

  • Exploratory data analysis (EDA) is the first step: understanding a dataset through summaries and charts before modelling.
  • It answers what values each column takes, how they are distributed, how they relate, and where the problems are.
  • Univariate EDA examines one variable; bivariate examines two and their relationship; multivariate examines three or more together.
  • Missing data (absent values) is handled by dropping the affected rows/columns or imputing (filling with mean, median, or mode).
  • Outliers (extreme values far from the rest) must be detected and judged as errors to fix or genuine extremes to keep.
  • Dropping rows can lose data and bias results, so imputation is often preferable; understand why data is missing first.
  • Memory hook: EDA first; count variables (uni/bi/multi), and handle gaps and extremes honestly.

Study this properly

This page is the lesson to read. In Gri-Learn the same topic is a graded deck: the self-checks are scored and your weak topics are tracked. Free to start.

Start this topic

Already have an account? Sign in

More from Fundamentals of Data Analytics

Gri-Learn · syllabus-mapped B.C.A. lessons in English, Hindi and Gujarati

Exploratory Data Analysis (EDA): types of EDA; Univariate, Bivariate and Multivariate Analysis; handling missing data and outliers · Data Analytics using Python (Major-14) · Gri-Learn