Haleluya Tesfamariam - NYC Flights Project

Project Description:

For this project, I analyzed flight data from New York City to determine which airlines operate the most flights from John F. Kennedy International Airport (JFK) and LaGuardia Airport (LGA).

I cleaned and transformed the data through five steps in R. In step1, I merged the flights dataset with airports to convert three-letter airport codes into full airport names. In step2, I joined step1 with the airlines dataset to replace carrier codes with full airline names. In step3, I filtered out flights originating from Newark Liberty International Airport (EWR), leaving only JFK and LGA.

In step4, I grouped step3 by airline and airport and used summarise() to calculate the total number of flights for each combination. In step5, I filtered step4 to focus on the five busiest airlines at each airport, reducing clutter in the final visualization.

Finally, I used step5 in ggplot2 to create side-by-side charts comparing the top airlines at JFK and LGA. The results show that Delta Air Lines is a major carrier at both airports, while JetBlue Airways has a particularly large presence at JFK and American Airlines is more prominent at LGA.

library(tidyverse)
── Attaching core tidyverse packages ──────────────────────── tidyverse 2.0.0 ──
✔ dplyr     1.2.1     ✔ readr     2.2.0
✔ forcats   1.0.1     ✔ stringr   1.6.0
✔ ggplot2   4.0.3     ✔ tibble    3.3.1
✔ lubridate 1.9.5     ✔ tidyr     1.3.2
✔ purrr     1.2.2     
── Conflicts ────────────────────────────────────────── tidyverse_conflicts() ──
✖ dplyr::filter() masks stats::filter()
✖ dplyr::lag()    masks stats::lag()
ℹ Use the conflicted package (<http://conflicted.r-lib.org/>) to force all conflicts to become errors
library(nycflights23)
library(dplyr)
library(ggplot2)
#Converting 3 Letter Airports to Full Names
data("airports")
step1 <- flights %>%
left_join(airports, by = c("origin" = "faa")) %>%
  rename(origin_name = name)
head(step1)
# A tibble: 6 × 26
   year month   day dep_time sched_dep_time dep_delay arr_time sched_arr_time
  <int> <int> <int>    <int>          <int>     <dbl>    <int>          <int>
1  2023     1     1        1           2038       203      328              3
2  2023     1     1       18           2300        78      228            135
3  2023     1     1       31           2344        47      500            426
4  2023     1     1       33           2140       173      238           2352
5  2023     1     1       36           2048       228      223           2252
6  2023     1     1      503            500         3      808            815
# ℹ 18 more variables: arr_delay <dbl>, carrier <chr>, flight <int>,
#   tailnum <chr>, origin <chr>, dest <chr>, air_time <dbl>, distance <dbl>,
#   hour <dbl>, minute <dbl>, time_hour <dttm>, origin_name <chr>, lat <dbl>,
#   lon <dbl>, alt <dbl>, tz <dbl>, dst <chr>, tzone <chr>
#Converting Airlines to Full Names
data("airlines")
step2 <- step1 %>%
  left_join(airlines, by = "carrier")

#Filtering Flights Originating from Newark (EWR)
step3 <- step2 %>%
  filter(!origin_name %in% c("Newark Liberty International Airport"))

#Grouping Number of Flights Originating from LGA vs JFK by Airline

step4 <- step3 %>%
  group_by(name, origin) %>%
  summarise(total_flights = n())
`summarise()` has regrouped the output.
ℹ Summaries were computed grouped by name and origin.
ℹ Output is grouped by name.
ℹ Use `summarise(.groups = "drop_last")` to silence this message.
ℹ Use `summarise(.by = c(name, origin))` for per-operation grouping
  (`?dplyr::dplyr_by`) instead.
#filtering for top 5 airlines only because graph was crowded
step5 <- step4 %>%
  filter(row_number() <= 5)

#Plotting graphs facet wrapped to show JFK and LGA in seperate boxes. Flipped vertically to show airline names, scaled total_flights by 1,000 to make the graph easier to read.

ggplot(data = step5) +
  geom_col(aes(x = name, y = total_flights/1000, fill = name), position = "dodge") +
  facet_wrap(~origin) +
  coord_flip() +

#Cosmetic Edits
labs(
  title = "Top Airlines Flying out of New York City Airports",
  x = "Airline",
  y = "Total Number of Flights",
  fill = "Airlines",
  caption = "Data Source: nycflights23 package (U.S. Bureau of Transportation)"
)