R Markdown

Question 1

Each row represents an observation, and each column represents a variable. Different types of observational units are organized into separate tables.

Question 2

Tidy datasets organize data in a clear and consistent way, making it easier to work with. Packages like ggplot2 and dplyr are designed to work with this format for data visualization and manipulation. By keeping data in a standardized structure, these packages can work together smoothly, making it easier to clean, visualize, and analyze data without constantly reorganizing it.

Question 3

The airline_safety_smaller dataset can be converted into tidy format using the pivot_longer() function. This combines the two fatalities columns into one column called fatalities_years, which identifies the time period, and another column called count, which shows the number of fatalities. The airline names remain in their own column.

library(fivethirtyeight)
## Some larger datasets need to be installed separately, like senators and
## house_district_forecast. To install these, we recommend you install the
## fivethirtyeightdata package by running:
## install.packages('fivethirtyeightdata', repos =
## 'https://fivethirtyeightdata.github.io/drat/', type = 'source')
library(tidyverse)
## ── Attaching core tidyverse packages ──────────────────────── tidyverse 2.0.0 ──
## ✔ dplyr     1.2.1     ✔ readr     2.2.0
## ✔ forcats   1.0.1     ✔ stringr   1.6.0
## ✔ ggplot2   4.0.3     ✔ tibble    3.3.1
## ✔ lubridate 1.9.5     ✔ tidyr     1.3.2
## ✔ purrr     1.2.2
## ── Conflicts ────────────────────────────────────────── tidyverse_conflicts() ──
## ✖ dplyr::filter() masks stats::filter()
## ✖ dplyr::lag()    masks stats::lag()
## ℹ Use the conflicted package (<http://conflicted.r-lib.org/>) to force all conflicts to become errors
airline_safety_smaller <- airline_safety |>
  select(airline, starts_with("fatalities"))

airline_safety_smaller_tidy <- airline_safety_smaller |>
  pivot_longer(
    cols = -airline,
    names_to = "fatalities_years",
    values_to = "count"
  )

airline_safety_smaller_tidy
## # A tibble: 112 × 3
##    airline               fatalities_years count
##    <chr>                 <chr>            <int>
##  1 Aer Lingus            fatalities_85_99     0
##  2 Aer Lingus            fatalities_00_14     0
##  3 Aeroflot              fatalities_85_99   128
##  4 Aeroflot              fatalities_00_14    88
##  5 Aerolineas Argentinas fatalities_85_99     0
##  6 Aerolineas Argentinas fatalities_00_14     0
##  7 Aeromexico            fatalities_85_99    64
##  8 Aeromexico            fatalities_00_14     0
##  9 Air Canada            fatalities_85_99     0
## 10 Air Canada            fatalities_00_14     0
## # ℹ 102 more rows