Importing data into R from different file formats (CSV, Excel, etc.)

read.csv() pulls a CSV into a data frame in one line (header handled, types guessed), Excel files need the readxl package's read_excel(), and str() is the after-import health check.

9 min read · 9 cards · 2 checks

Read in: English · हिन्दी · ગુજરાતી


Theory

The survey arrives, again

The CampusPulse data sits in survey.csv: 60 rows, four columns. You have imported this exact shape twice before: with the sqlite3 shell's .import, and with pandas' read_csv.

R's version is one line:

survey <- read.csv("survey.csv")

No import statement, no module: reading data is built into the language. The skill of this lesson is not the line: it is knowing what R did behind it, and how to check.

Theory

The customs officer

read.csv() is a customs officer at R's border. As the file enters, the officer:

1. Reads the first line as labels (header = TRUE, the default).

2. Guesses each column's type from its contents: numbers become numeric, text becomes character.

3. Stamps the shipment into a data frame: R's table.

Good officers still get audited: after every import, YOU inspect the stamped result. The inspection command is str().

Practical

Import, then audit

# CSV: built into base R, one line
survey <- read.csv("survey.csv")     # header = TRUE is the default

# THE health check after every import:
str(survey)
# 'data.frame': 60 obs. of 4 variables:
#  $ roll   : int  101 102 103 ...
#  $ city   : chr  "Surat" "Navsari" ...
#  $ stay   : chr  "hostel" "day" ...
#  $ screen : num  120 90 240 ...

head(survey)        # first 6 rows, eyeball the values

# Excel needs a package: INSTALL once, LOAD every session
# install.packages("readxl")     # once per machine
library(readxl)                   # once per session
marks <- read_excel("marks.xlsx", sheet = 1)

# other delimiters: read.delim (tabs), read.csv2 (; with , decimals)

Theory

What each piece means

header = TRUE (default): line 1 becomes column names. With header = FALSE, R invents names V1, V2, V3 and your labels become a data row: the same header decision as every import tool you have met.

Type guessing: a column of digits arrives as int/num; anything with text arrives as character. (Modern R no longer auto-converts text to factors: you promote deliberately, next lessons.)

Excel: base R does not read .xlsx. The readxl package does: install.packages() once per machine, library() once per session: two different verbs, endlessly confused.

Quiz

A student ran install.packages("readxl") yesterday. Today, in a fresh RStudio session, read_excel("marks.xlsx") errors: could not find function "read_excel". Why?

  1. Installing puts the package on disk once; each new session must still load it with library(readxl)
  2. The package must be reinstalled every day
  3. read_excel only works on CSV files
  4. The file name needs .csv extension
Show the answer

Installing puts the package on disk once; each new session must still load it with library(readxl)

Two verbs, two lifetimes: install.packages() downloads to disk (once per machine); library() loads into the running session (once per session). A fresh session starts with only base R loaded, so read_excel is unknown until library(readxl) runs. This install-vs-load distinction is R's most-asked package question, in labs and exams alike.

Think first

Audit catches the intruder

str(survey) shows: $ screen: chr "120" "90" "absent" "240". The screen-time column arrived as CHARACTER. Before tapping: what caused it, why does it matter, and which BCA303 moment is this the twin of?

Show the answer

One student's entry says "absent": a single string in the column, and R's homogeneous-column rule coerced every number to text (the coercion trap from the syntax lesson, now at file scale). mean(survey$screen) is now impossible.

It is the twin of pandas/csv-module numbers-arriving-as-strings and SQLite's affinity surprise from BCA303: text formats carry no types, so importers guess, and one dirty value poisons a column. The repair (NA handling, as.numeric) is exactly where the next unit goes.

Watch out

The import checklist

1. File not found? Working directory first: getwd(), then fix with setwd() or a full path.

2. str() after every import: dimensions right? Types right?

3. Numeric column shown as chr: hunt the stray text value before any statistics.

4. V1, V2 column names: your header = FALSE (or the file truly has no header).

Four checks, ten seconds, and every downstream lesson works on trusted data.

Theory

Same song, third language

Notice your own fluency: header row, type inference, working directory, verify-the-import: you learned these ideas in the SQLite shell, rehearsed them in pandas, and today they transferred to R in minutes. Tool skills expire; concepts compound. Next lesson: the data frame itself: R's answer to the DataFrame you already think in.

Summary

Key takeaways

  • read.csv("file.csv") imports a CSV as a data frame; header = TRUE and comma separator are defaults.
  • Excel needs readxl: install.packages() once per machine, library() once per SESSION.
  • R guesses column types; one stray string turns a numeric column to character.
  • str() is the mandatory post-import audit: dimensions + per-column types; head() eyeballs rows.
  • File-not-found almost always means working directory, not a missing file.
  • read.delim for tabs, read.csv2 for semicolon files; packages exist for SPSS/JSON.
  • Memory hook: the customs officer stamps it; you audit the stamp.

Study this properly

This page is the lesson to read. In Gri-Learn the same topic is a graded deck: the self-checks are scored and your weak topics are tracked. Free to start.

Start this topic

Already have an account? Sign in

More from Introduction to R and working with Data

Gri-Learn · syllabus-mapped B.C.A. lessons in English, Hindi and Gujarati

Importing data into R from different file formats (CSV, Excel, etc.) · Statistical Methods and Data Analysis (MDC-03) · Gri-Learn