Assignment4-Part2

Quarto

Quarto enables you to weave together content and executable code into a finished document. To learn more about Quarto see https://quarto.org.

Running Code

When you click the Render button a document will be generated that includes both content and the output of embedded code. You can embed code like this:

library(tidyverse)
── Attaching core tidyverse packages ──────────────────────── tidyverse 2.0.0 ──
✔ dplyr     1.2.1     ✔ readr     2.2.0
✔ forcats   1.0.1     ✔ stringr   1.6.0
✔ ggplot2   4.0.3     ✔ tibble    3.3.1
✔ lubridate 1.9.5     ✔ tidyr     1.3.2
✔ purrr     1.2.2     
── Conflicts ────────────────────────────────────────── tidyverse_conflicts() ──
✖ dplyr::filter() masks stats::filter()
✖ dplyr::lag()    masks stats::lag()
ℹ Use the conflicted package (<http://conflicted.r-lib.org/>) to force all conflicts to become errors
library(nycflights23)

data(flights)
dim(flights)
[1] 435352     19
head(flights)
# A tibble: 6 × 19
   year month   day dep_time sched_dep_time dep_delay arr_time sched_arr_time
  <int> <int> <int>    <int>          <int>     <dbl>    <int>          <int>
1  2023     1     1        1           2038       203      328              3
2  2023     1     1       18           2300        78      228            135
3  2023     1     1       31           2344        47      500            426
4  2023     1     1       33           2140       173      238           2352
5  2023     1     1       36           2048       228      223           2252
6  2023     1     1      503            500         3      808            815
# ℹ 11 more variables: arr_delay <dbl>, carrier <chr>, flight <int>,
#   tailnum <chr>, origin <chr>, dest <chr>, air_time <dbl>, distance <dbl>,
#   hour <dbl>, minute <dbl>, time_hour <dttm>
late_flights <- flights %>%
  filter(arr_delay > 0 | dep_delay > 0)
dim(late_flights)
[1] 189854     19
late_by_carrier <- late_flights %>% group_by(carrier)
late_by_carrier
# A tibble: 189,854 × 19
# Groups:   carrier [14]
    year month   day dep_time sched_dep_time dep_delay arr_time sched_arr_time
   <int> <int> <int>    <int>          <int>     <dbl>    <int>          <int>
 1  2023     1     1        1           2038       203      328              3
 2  2023     1     1       18           2300        78      228            135
 3  2023     1     1       31           2344        47      500            426
 4  2023     1     1       33           2140       173      238           2352
 5  2023     1     1       36           2048       228      223           2252
 6  2023     1     1      503            500         3      808            815
 7  2023     1     1      520            510        10      948            949
 8  2023     1     1      537            520        17      926            818
 9  2023     1     1      547            545         2      845            852
10  2023     1     1      549            559       -10      905            901
# ℹ 189,844 more rows
# ℹ 11 more variables: arr_delay <dbl>, carrier <chr>, flight <int>,
#   tailnum <chr>, origin <chr>, dest <chr>, air_time <dbl>, distance <dbl>,
#   hour <dbl>, minute <dbl>, time_hour <dttm>
late_totals <- late_by_carrier %>%
  summarise(
    total_count = n()                                
  )
print(late_totals)
# A tibble: 14 × 2
   carrier total_count
   <chr>         <int>
 1 9E            17894
 2 AA            17117
 3 AS             3727
 4 B6            36434
 5 DL            28494
 6 F9              825
 7 G4              208
 8 HA              274
 9 MQ              199
10 NK             7369
11 OO             2941
12 UA            40767
13 WN             7840
14 YX            25765
library(treemap)
treemap(
  late_totals,
  index = c("carrier"),
  vSize = "total_count",
  type = "index",
  title = "No. of flights that were late for each airlines"
  
)

The visualization that I created shows exactly how many observations and variables are within the flight’s dataset. It then shows just the arrival and departure delay columns in a way where only the late numbers are shown, and the number of observations and variables for those columns. After that, each dataset of delayed flights is grouped with their assigned airline carrier. In addition to that, it even shows the total amount for each of the carriers that are listed in the carrier column. Lastly, there is a treemap visualization that shows the number of late flights for each of the airlines. One aspect of this plot that I would like to highlight is that it accurately shows how many observations and variables there are in this dataset. Another aspect I would like to highlight is that the late totals show the exact amount for each of the carriers that are listed.