For this project, I analyzed flight data from New York City to determine which airlines operate the most flights from John F. Kennedy International Airport (JFK) and LaGuardia Airport (LGA).
I cleaned and transformed the data through five steps in R. In step1, I merged the flights dataset with airports to convert three-letter airport codes into full airport names. In step2, I joined step1 with the airlines dataset to replace carrier codes with full airline names. In step3, I filtered out flights originating from Newark Liberty International Airport (EWR), leaving only JFK and LGA.
In step4, I grouped step3 by airline and airport and used summarise() to calculate the total number of flights for each combination. In step5, I filtered step4 to focus on the five busiest airlines at each airport, reducing clutter in the final visualization.
Finally, I used step5 in ggplot2 to create side-by-side charts comparing the top airlines at JFK and LGA. The results show that Delta Air Lines is a major carrier at both airports, while JetBlue Airways has a particularly large presence at JFK and American Airlines is more prominent at LGA.
library(tidyverse)
── Attaching core tidyverse packages ──────────────────────── tidyverse 2.0.0 ──
✔ dplyr 1.2.1 ✔ readr 2.2.0
✔ forcats 1.0.1 ✔ stringr 1.6.0
✔ ggplot2 4.0.3 ✔ tibble 3.3.1
✔ lubridate 1.9.5 ✔ tidyr 1.3.2
✔ purrr 1.2.2
── Conflicts ────────────────────────────────────────── tidyverse_conflicts() ──
✖ dplyr::filter() masks stats::filter()
✖ dplyr::lag() masks stats::lag()
ℹ Use the conflicted package (<http://conflicted.r-lib.org/>) to force all conflicts to become errors
#Converting 3 Letter Airports to Full Namesdata("airports")step1 <- flights %>%left_join(airports, by =c("origin"="faa")) %>%rename(origin_name = name)head(step1)
#Converting Airlines to Full Namesdata("airlines")step2 <- step1 %>%left_join(airlines, by ="carrier")#Filtering Flights Originating from Newark (EWR)step3 <- step2 %>%filter(!origin_name %in%c("Newark Liberty International Airport"))#Grouping Number of Flights Originating from LGA vs JFK by Airlinestep4 <- step3 %>%group_by(name, origin) %>%summarise(total_flights =n())
`summarise()` has regrouped the output.
ℹ Summaries were computed grouped by name and origin.
ℹ Output is grouped by name.
ℹ Use `summarise(.groups = "drop_last")` to silence this message.
ℹ Use `summarise(.by = c(name, origin))` for per-operation grouping
(`?dplyr::dplyr_by`) instead.
#filtering for top 5 airlines only because graph was crowdedstep5 <- step4 %>%filter(row_number() <=5)#Plotting graphs facet wrapped to show JFK and LGA in seperate boxes. Flipped vertically to show airline names, scaled total_flights by 1,000 to make the graph easier to read.ggplot(data = step5) +geom_col(aes(x = name, y = total_flights/1000, fill = name), position ="dodge") +facet_wrap(~origin) +coord_flip() +#Cosmetic Editslabs(title ="Top Airlines Flying out of New York City Airports",x ="Airline",y ="Total Number of Flights",fill ="Airlines",caption ="Data Source: nycflights23 package (U.S. Bureau of Transportation)")