Introduction

New York is densely populated and described as a city that is constantly moving. People commute into, across, or within the city lines every day for work, tourism, or simply to enjoy themselves. Understandably as a result of this constant motion, the city is notorious for its traffic. With this traffic comes accidents. This report seeks to interpret a data set regarding information about accidents that occurred in New York City over the span of 14 years. Through this report we can hope to recognize trends in what time most accidents occur, what areas seem to be the most accident-prone, what years had the highest recorded accidents, and more. By observing this data we can hope to identify problems in the current transportation infrastructure of the city and potentially propose solutions to mitigate the risk of a crash in these areas.

Data set

This data set contains 1,048,538 records pertaining to accidents that occurred in New York City. Data has been collected from 2012-2026 however the majority of observable recordings take place from 2017-2024. There are 29 fields in this data set including information on the crash date and time, area of occurrence (zip code, street, borough), type of cars involved, and data on injuries/deaths.

library(lubridate)
library(plyr)
library(ggplot2)
library(ggthemes)
library(RColorBrewer)
library(scales)
library(ggrepel)
library(plotly)
library(data.table)
library(dplyr)
library(cowplot)
library(ggpubr)

#----------------------------------------------------------#
#data cleaning / variable set up

setwd("~/ds736 directory/R/R Project/R Module Deliverable")

filename <- "C:\\Users\\evere\\OneDrive\\Documents\\ds736 directory\\R\\R Project\\R Module Deliverable\\NYC_Crashes.csv"

df <- fread(filename, na.strings = c(NA, ""))
#---------------------for chart 1:-----------------------
times <- lubridate::hm(df$`CRASH TIME`)
hour_df <- data.frame(hour = lubridate::hour(times)) %>%
dplyr::count(hour)

x_axis = min(hour_df$hour):max(hour_df$hour)

#---------------------for chart 2:------------------------
Crash_date <- lubridate::mdy(df$`CRASH DATE`)
crash_year <- data.frame(year = lubridate::year(Crash_date))
crash_year <-crash_year %>% count()

Crash_info <- df %>%
  mutate(crash_year = year(mdy(`CRASH DATE`))) %>%
  dplyr::count(.data$BOROUGH, crash_year)
Crash_info <- Crash_info %>%
  filter(!crash_year %in% c(2012,2013,2014,2015,2016,2025,2026),
         !is.na(BOROUGH)) 

max_y <-round_any(max(Crash_info$n),150000,ceiling)

#---------------------for chart 3:------------------------
vehicle_type <- c(df$`VEHICLE TYPE CODE 1`,
                  df$`VEHICLE TYPE CODE 2`,
                  df$ `VEHICLE TYPE CODE 3`,
                  df$`VEHICLE TYPE CODE 4`,
                  df$`VEHICLE TYPE CODE 5`)
vehicle_type <- vehicle_type[!is.na(vehicle_type)] %>% data.frame(type = .)

type_count <- vehicle_type %>%
  dplyr::count(type, name = "freq")

top_10_vehicle_types <- type_count[order(type_count$freq, decreasing = TRUE), ]
top_10_vehicle_types <- top_10_vehicle_types %>%
  dplyr::mutate(type = ifelse(type == "Station Wagon/Sport Utility Vehicle", "SUV", type))

#----------------------for chart 4:------------------------
injuries_or_deaths <- rbind(df$`NUMBER OF PERSONS INJURED`,
  df$`NUMBER OF PERSONS KILLED`)

injuries_or_deaths <- df %>%
  group_by(`NUMBER OF PERSONS INJURED`, `NUMBER OF PERSONS KILLED`) %>%
  dplyr::summarise(freq = n(), .groups = "drop") %>%
  arrange(desc(freq))

injuries_or_deaths <- df %>%
  mutate(category = ifelse(
    `NUMBER OF PERSONS INJURED` == 0 & `NUMBER OF PERSONS KILLED` == 0, "No Deaths, No Injuries",
    ifelse(
      `NUMBER OF PERSONS INJURED` > 0 & `NUMBER OF PERSONS KILLED` == 0, "Injuries, No Deaths",
      ifelse(
        `NUMBER OF PERSONS INJURED` == 0 & `NUMBER OF PERSONS KILLED` > 0, "Deaths, No Injuries",
        "Both Deaths and Injuries"
      )
    )
  )) %>%
  dplyr::count(category, sort = TRUE, name = "total_occurrences")

injuries_or_deaths <- na.omit(injuries_or_deaths)

#-------------------for chart 5:------------------------

top_zips <- df %>%
  dplyr::count(`ZIP CODE`, `BOROUGH`, name = "freq")

top_zips <- top_zips[order(top_zips$freq, decreasing = TRUE),]
top_zips <-na.omit(top_zips)

top5_bk <- top_zips%>%
  filter(BOROUGH == "BROOKLYN")
top5_bk <- top5_bk[1:5,]

top5_manhattan <- top_zips%>%
  filter(BOROUGH == "MANHATTAN")
top5_manhattan <- top5_manhattan[1:5,]

top5_bx <- top_zips%>%
  filter(BOROUGH == "BRONX")
top5_bx <- top5_bx[1:5,]

top5_staten <- top_zips%>%
  filter(BOROUGH == "STATEN ISLAND")
top5_staten <- top5_staten[1:5,]

top5_queens <- top_zips%>%
  filter(BOROUGH == "QUEENS")
top5_queens <- top5_queens[1:5,]

top5_NY <- rbind(top5_bk,
                 top5_bx,
                 top5_manhattan,
                 top5_queens,
                 top5_staten
)

boroughs_ranked <- top5_NY %>%
  dplyr::group_by(BOROUGH) %>%
  dplyr::arrange(BOROUGH, freq) %>%         
  dplyr::mutate(rank = row_number()) %>%            
  dplyr::ungroup()

Visualizations

Crash Occurrences by Time of Day

#-----------------------chart 1---------------------
ggplot(hour_df, aes(x = hour, y = n))+
  geom_line(color = 'blue', size=1) +
  geom_point(shape=21, size =3, color = 'black',
             fill = 'white') +
  labs(title = "Crash Occurances by Time of Day", x= 'Hour',
       y='Total Crashes')+
  scale_y_continuous(labels=comma)+
  theme_light()+
  theme(plot.title = element_text(hjust = .5))+
  scale_x_continuous(labels = x_axis, breaks = x_axis, 
                     minor_breaks = NULL)+
  geom_label_repel(aes(label = ifelse(hour == min(hour)|
                                        hour == max(hour) |
                                        n == min(n) |
                                        n == max(n),
                                      scales::comma(n),
                                      "")), size =3,
                   box.padding = 1,
                   segment.colour = 'black')

Crashes By Year By Borough

#-------------------------chart 2------------------------
ggplot(Crash_info, 
       aes(x = crash_year, y = n , fill = BOROUGH)) + 
  geom_bar(stat = "identity")+
  coord_flip() +
  labs(title = "Crashes by Year by Borough 2017-2024", 
       y = "Crashes", x = "") + 
  theme_economist() +
  theme(plot.title = element_text(hjust = .5)) +
  scale_fill_brewer(palette = "Spectral", name ="",
                    guide = guide_legend(reverse = TRUE)) +
  scale_y_continuous(labels = comma, limits = c(0,max_y))

Most Common Vehicle Types in NYC Crashes

#-------------------------chart 3------------------------
ggplot(top_10_vehicle_types[1:10,], aes(x = reorder(type,freq), y=freq)) +
  geom_bar(colour = "black", fill="blue", stat="identity") + 
  coord_flip() +
  labs(title = "Top 10 Recurring Vehicle Types in New York Accidents",
       x = "Vehicle Type", y= "Occurences") +
  theme(plot.title = element_text(hjust = 0.5)) +
  scale_y_continuous(labels=comma)+
  geom_text(aes(x= type, y = freq, label = scales::comma(freq)),
            hjust = -.1,
            size = 3.5)

Proportions of Injuries and/or Deaths

#------------------------chart 4--------------------------
plot_ly(injuries_or_deaths, labels=~category, values = ~total_occurrences, type = "pie",
        textposition = "outside", textinfo = "label + percent")%>%
  layout(title = "Proportion of injuries and/or deaths that occured in NY Accidents")

Accidents By Zip Within Each Borough

#------------------------chart 5---------------------------
ggplot(boroughs_ranked, 
                 aes(x = factor(rank), 
                  y = reorder(BOROUGH, freq, FUN = sum), 
                            fill = freq)) +
  geom_tile(color = "black") +
  geom_text(aes(label = paste0("Zip: ", `ZIP CODE`, "\n", comma(freq))), 
            size = 3.5) +
  scale_fill_gradient(low = "white", high = "turquoise", labels = comma) +
  labs(
    title = "Accidents by Zip Code Within Each Borough",
    x = "Accident frequency",
    y = "Borough",
    fill = "Accident Count"
  ) +
  theme_light() +
  theme(
    plot.title = element_text(hjust = 0.5)
  )

Conclusion

Through these visualizations we were able to observe that 4-5pm was the time of day where most accidents in this database occurred. This is not surprising as this can be considered rush hour, so the roads would be busier than usual. 2018 and 2019 were the two years with the most recorded accidents, and within this 2 year span, Brooklyn was the borough with the most recorded accidents. The most common vehicle type within the data set is a sedan. This is also unsurprising as sedans and compact cars tend to be a more popular choice in a city. 72.3% of all recorded accidents had no deaths or injuries tied to them, 27.5% were tied to an injury but not a death, .13% were tied to one or more deaths but no injuries, and .0435% were tied to at least one death and at least one injury. The low death rates of these accidents makes sense when you consider that many of New York’s roads are often lined with traffic lights or tend to get crowded. Being in this environment would lead to many accidents occurring at lower speeds, resulting in fewer injuries and deaths. The zip code that contained the highest amount of accident recordings was 11207 in Brooklyn with 14,800 observations. The neighboring zip, 11208 was also in the top 5 most accident prone zip codes with 9,709 occurrences. With this data, we can target the problem areas and try to lower risk of collision. This could be done by the means of creating alternative routes or changing the traffic infrastructure in these areas.