It is the goal of this report to analyze and explain the trend of cancellations and delays of flights in 2013 using data tables and graphs to inform best practices when flying. Most of the data used in this report will be from the “Flights” data set found in the library “nycflights13” that contains data from the Bureau of Transportation Statistics. The variables found in this data set are as follows:
Year, Month, and Day
sched_dep_time, sched_arr_time | The scheduled Departure and Arrival time, respectively
dep_time, arr_time | The actual departure and arrival time, respectively
dep_delay, arr_delay | The delay between the scheduled departure or arrival time and the actual departure or arrival time
carrier | A two letter abbreviation of the company that handled the flight
flight | flight number
tailnum | the identification number of the aircraft used for the flight
origin, dest | The airport of origin and destination abbreviated using three letters
air_time | The time spent in the air, in minutes
distance | The total distance traveled, in miles
hour, minute | time of departure broken into hours and minutes
time_hour | Scheduled time of departure as a standard date
This data set contains data on over 300,000 flights. Below are the libraries used in this report. Unfortunately due to the extreme quantity of data in this table it is not possible to display the data table in this report.
library(nycflights13)
library(tidyverse)
## ── Attaching core tidyverse packages ──────────────────────── tidyverse 2.0.0 ──
## ✔ dplyr 1.2.1 ✔ readr 2.2.0
## ✔ forcats 1.0.1 ✔ stringr 1.6.0
## ✔ ggplot2 4.0.3 ✔ tibble 3.3.1
## ✔ lubridate 1.9.5 ✔ tidyr 1.3.2
## ✔ purrr 1.2.2
## ── Conflicts ────────────────────────────────────────── tidyverse_conflicts() ──
## ✖ dplyr::filter() masks stats::filter()
## ✖ dplyr::lag() masks stats::lag()
## ℹ Use the conflicted package (<http://conflicted.r-lib.org/>) to force all conflicts to become errors
library(DT)
flights %>%
group_by(month, day) %>%
summarize(large_delay = sum(dep_delay >= 60, na.rm = TRUE),
small_delay = sum(dep_delay < 60, na.rm = TRUE),
cancelled = sum(is.na(dep_time)),
large_delay_percentage = ((large_delay + cancelled)/(large_delay + small_delay + cancelled))*100) %>%
filter(large_delay_percentage >= 35) %>%
arrange(month, day)
Here is a table that contains all days that had over 35% of flights delayed or cancelled in chronological order. One hypothesis about the cause of these delays and cancellations that is easy to test is if weather was the foremost cause of the delays. A quick search reveals that:
On February 2nd and 3rd New York was struck with a historically catastrophic blizzard that slowed all city functions to a halt
Similarly, on March 8th heavy fog and snow clouded vision and delayed air travel
On May 23rd it was no longer cold weather phenomena that disrupted travel but spring thunderstorms that hit New York and grounded or delayed aircraft
June 24th and 28th are very similar as large amounts of rainfall hit the city and caused scheduling disturbances
July 1st and 10th follow a similar pattern as heavy rainfall and thunderstorms caused delays along the entire east coast
September 2nd and 12th follow a similar pattern as severe thunderstorms caused delays in the greater New York area
Finally, on December 5th a large amount of dense fog covered the ground at multiple airports and reduced visibility enough that it was unsafe to take off or land.
Now that there is a catalog of days with an abnormally large number of delays or cancellations, it would be useful to see what flights managed to leave on time or early on those days.
flights %>%
filter((month == 2 & day == 8 | month == 2 & day == 9) | (month == 3 & day == 8) | (month == 5 & day == 23) | (month == 6 & day == 24 | month == 6 & day == 28) | (month == 7 & day == 1 | month == 7 & day == 10 | month == 7 & day == 23) | (month == 9 & day == 2 | month == 9 & day == 12) | (month == 12 & day == 5)) %>%
filter(dep_delay <= 0) %>%
arrange(month, day)
A more useful statistic yet would be to know the average time these flights left to avoid delays and cancellations.
flights %>%
filter((month == 2 & day == 8 | month == 2 & day == 9) | (month == 3 & day == 8) | (month == 5 & day == 23) | (month == 6 & day == 24 | month == 6 & day == 28) | (month == 7 & day == 1 | month == 7 & day == 10 | month == 7 & day == 23) | (month == 9 & day == 2 | month == 9 & day == 12) | (month == 12 & day == 5)) %>%
filter(dep_delay <= 0) %>%
group_by(month, day) %>%
summarise(average_dep_time = mean(dep_time),
count = n()) %>%
arrange(month, day)
It can be reasoned then that the best time to leave to avoid delays is early morning, though it can sometimes be effective to leave in the late afternoon. Looking at the February 8th and 9th average departure time elicits some curiosity. There is a large gap between the average departure times. As noted above, there was a historically large snowstorm that hit New York on these two days. One likely explanation for this discrepancy then could be that many flights were able to leave early enough to avoid the snowstorm on the 8th. The snowstorm then hit and stopped all flights until it passed in the evening and airport crews were able to clear runways to get flights running again on the 9th.
While the table above implied the best time to travel on days with significant delays it would also be important to know the best time to travel in general.
flights %>%
group_by(hour) %>%
summarize(large_delay = sum(dep_delay >= 60, na.rm = TRUE),
small_delay = sum(dep_delay < 60, na.rm = TRUE),
cancelled = sum(is.na(dep_time)),
large_delay_percentage = ((large_delay + cancelled)/(large_delay + small_delay + cancelled))*100) %>%
ggplot() +
geom_point(mapping = aes(x = hour, y = large_delay_percentage)) +
labs(x = "Hour",
y = "% Flights Delayed or Cancelled",
title = "Relation of Flight Delays to Time of Day")
One plausible reason that there is a general increase in delays as the day progresses could be an idea of “compounding tardiness.” If an aircraft’s first flight of the day is delayed by ten minutes it would not appear on this graph and be barely noticeable to all passengers. However, it may then arrive ten minutes late as well. Then, if there is another delay of ten minutes, suddenly this flight is leaving 20 minutes late. If this effect compounds over a significant number of flights each day then it would reasonably explain the trend in this graph.
It is interesting to note a large outlier near 1AM. Surely it is not possible that all flights at 1AM were delayed by more than an hour. What is actually being displayed is the fights scheduled at late night that were delayed until a new day. No flights were scheduled for 1AM but all flights that took off then were delayed by more than an hour. Thus, 100% of the flights that took off at 1AM were significantly delayed. This graph suggests that the riskiest time to fly is late night between 9PM-11PM.
The original aim of this project was to analyze the trend of delays and cancellations of flights to understand why delays happen and how to plan around them. There was a large correlation between the number of delays on a given day and weather. An observable increase in delays as a day progresses was also noted. With this information it is recommended to travel early in the morning on a day with little adverse weather.