readrreadxlR always reads and writes files relative to a working directory unless you give a full path.
getwd() # see current working directory
setwd("C:/Users/YourName/Data") # change it (use your own path)Best Practice: In RStudio, create an RStudio Project (
File > New Project) for each course/assignment. This automatically manages your working directory and keeps files organized — far safer thansetwd().
CSV (Comma-Separated Values) is the most common format for tabular data.
readr (faster, more consistent, tidyverse-style)
readr::read_csv()is generally preferred in modern R because it is faster, gives a helpful column-type summary, and never auto-converts text to factors.
Since we can’t guarantee a file exists on your machine, let’s create one to practice with:
sample_data <- data.frame(
name = c("Amina", "Hodan", "Yusuf"),
score = c(88, 92, 65)
)
write.csv(sample_data, "sample_students.csv", row.names = FALSE)
# Now read it back in
loaded_data <- read.csv("sample_students.csv")
loaded_data## name score
## 1 Amina 88
## 2 Hodan 92
## 3 Yusuf 65
Base R cannot read .xlsx files directly — we need the
readxl package.
RStudio also offers a menu-driven importer:
File > Import Dataset > From Text (base) /
From Text (readr) / From Excel
This is a great way to see the code RStudio generates for
you (it shows the exact read_csv() call in the console) — a
useful way to learn the syntax.
Always run these checks immediately after loading new data — this habit will save you hours of debugging later:
## 'data.frame': 3 obs. of 2 variables:
## $ name : chr "Amina" "Hodan" "Yusuf"
## $ score: int 88 92 65
## name score
## 1 Amina 88
## 2 Hodan 92
## 3 Yusuf 65
## [1] 3 2
## name score
## Length:3 Min. :65.00
## Class :character 1st Qu.:76.50
## Mode :character Median :88.00
## Mean :81.67
## 3rd Qu.:90.00
## Max. :92.00
## name score
## 0 0
| Problem | Likely Cause | Fix |
|---|---|---|
| All columns read as text | Wrong delimiter or encoding | Check sep = ";" or sep = "\t" |
| Numbers read as character | Commas used as decimal points | read.csv(..., dec = ",") |
| Strange symbols in text | Wrong file encoding | read.csv(..., fileEncoding = "UTF-8") |
| Missing values not recognized | Uses a custom code like "N/A", "-" |
read.csv(..., na.strings = c("N/A", "-")) |
| Extra blank rows/columns | Excel file has formatting artifacts | Use readxl::read_excel(), then na.omit()
or manual cleaning |
Scenario: You export a small gradebook, then re-import and validate it, simulating a real workflow.
gradebook <- data.frame(
student_id = 1:4,
name = c("Ali", "Sara", "Deka", "Omar"),
midterm = c(70, 85, 60, 92),
final = c(75, 88, 55, 95)
)
write.csv(gradebook, "gradebook.csv", row.names = FALSE)
gradebook_check <- read.csv("gradebook.csv")
identical(gradebook, gradebook_check) # sanity check: did the round trip preserve the data?## [1] FALSE
## 'data.frame': 4 obs. of 4 variables:
## $ student_id: int 1 2 3 4
## $ name : chr "Ali" "Sara" "Deka" "Omar"
## $ midterm : int 70 85 60 92
## $ final : int 75 88 55 95
write.csv().head().colSums(is.na(...)) to detect it after
re-importing.readxl and writexl, then export
your data frame to an .xlsx file and read it back.read.csv() on a file that doesn’t
exist. What error message does R give you?getwd() and setwd().row.names = FALSE recommended when writing
CSVs?Q1. Which package provides read_excel()
for importing .xlsx files?
readrreadxldplyrQ2. What does getwd() do? a) Sets a new
working directory b) Gets the current working directory c) Gets the
Windows Desktop path d) Deletes the working directory
Q3. Why might you specify
na.strings = c("N/A", "-") when reading a CSV?
Q4. What function quickly shows you the number of
missing values in each column of a data frame df?
sum(df)colSums(is.na(df))nrow(df)str(df)Q5. What argument prevents write.csv()
from adding an unwanted row-number column?
header = FALSEsep = ","row.names = FALSEquote = FALSEread.csv() / readr::read_csv() import
CSVs; readxl::read_excel() imports Excel files.write.csv() and writexl::write_xlsx()
export data back out.str(),
head(), summary(), and a missing-value
check.Next Lesson: Data Cleaning and Manipulation Basics.