Introduction

Every day, thousands of car accidents occur in the United States. Traffic accidents remain a significant public safety issue around the world. This dataset contains information on car accidents across 49 states using APIs to access traffic accident data. Numerous variables, including weather, location, and road conditions, could have an impact on the amount of accidents that occur. By analyzing accident data, we can discover patterns in the data and get a better understanding of how traffic accidents happen. This analysis will examine the relationship between environmental and road conditions with car accidents as well as the states that have the most collisions reported.

Dataset

This dataset includes accident data gathered from February 2016 to March 2023 in the United States. When viewing the visualizations, it is important to recognize that the number of accidents reported in the year 2023 is smaller than the previous years because the data only covers the first three months of the year. The dataset contains approximately 7.7 million accident records. For this report, a sampled version will be used with 500,000 observations and 46 different variables. Live road data was collected from APIs using information from traffic cameras, departments of transportation, and law enforcement.

library(data.table)
library(ggplot2)
library(dplyr)
library(RColorBrewer)
library(scales)
library(lubridate)
library(plotly)
filename <- "/Users/isabella/Desktop/DS736//US_Accidents_March23_sampled_500k.csv"
df <- fread(filename, na.strings = c(NA, ""))

Findings

The analysis is focused on variables including state, accident severity, year, weather condition, and traffic signals. Several trends and patterns were shown after examining the car accident data. The number of accidents varied greatly between the states and the top 10 states with the most traffic accidents were selected. These states were then analyzed to show the differences in the severity of the accidents and resulted in severity level 2 being the most common among the states. The number of accidents were examined over the years and showed an increase in the data. Weather conditions and traffic signals were also analyzed to discover patterns in the data.

Top 10 States by the Number of Accidents

The bar chart shows the top 10 states with the highest number of accidents out of the 49 states in the dataset. California is the leading state with 113,274 reported traffic accidents. The counts vary greatly between the states with the tenth state being Oregon with 11,559 accidents.

# Count the number of accidents in each state and select the top 10 states
state_count <- data.frame(count(df, State))
top_states <- state_count[order(-state_count$n), ][1:10, ]

# Create a bar chart for the top 10 states by the number of accidents
ggplot(top_states, aes(x = reorder(State, -n), y = n)) +
  geom_bar(colour = "black", fill = "lightblue", stat = "identity") +
  geom_text(aes(label = scales::comma(n)), vjust = -0.5) +
  labs(title = "Top 10 States by Number of Accidents", x = "State", y = "Accident Count") +
  scale_y_continuous(labels = scales::comma) +
  theme_light() +
  theme(plot.title = element_text(hjust = 0.5))

Accident Severity in the Top 10 States

The stacked bar chart shows the accident severity across the top 10 states with the highest number of traffic accidents. The severity of the accident is shown on a scale of 1 to 4 where 4 represents the most impact on traffic. The first chart shows the count of accidents for each severity level of the top 10 states. The second chart was included to show the proportion of each severity level in the states. This chart allows you to get a better visual comparison of the severity distributions.

# Count accidents by state and severity
state_severity <- data.frame(count(df, State, Severity))
state_severity <- state_severity[state_severity$State %in% top_states$State, ]

# Calculate the total number of accidents to label each state
agg_tot <- state_severity %>%
  select(State, n) %>%
  group_by(State) %>%
  summarise(tot = sum(n), .groups = 'keep') %>%
  data.frame()

# Create a stacked bar chart for the accident severity in the top 10 states
ggplot(state_severity, aes(x = State, y = n, fill = factor(Severity))) +
  geom_bar(stat = "identity") +
  labs(title = "Accident Severity in the Top 10 States", x = "State", y = "Accident Count", fill = "Severity") +
  scale_y_continuous(labels = scales::comma) +
  theme_light() +
  theme(plot.title = element_text(hjust = 0.5)) +
  scale_fill_brewer(palette = "RdYlBu", direction = -1) +
  geom_text(data = agg_tot, aes(x = State, y = tot, label = scales::comma(tot), fill = NULL), vjust = -0.5)

# Show proportion of accidents
ggplot(state_severity, aes(x = State, y = n, fill = factor(Severity))) +
  geom_bar(stat = "identity", position = "fill") +
  labs(title = "Proportion of Accident Severity in the Top 10 States", x = "State", y = "Proportion of Accidents", fill = "Severity") +
  theme_light() +
  theme(plot.title = element_text(hjust = 0.5)) +
  scale_fill_brewer(palette = "RdYlBu", direction = -1)

Number of Accidents Over the Years

The line graph shows the number of accidents over the years 2016 to 2023. The accident count increased throughout the years 2016 to 2022. The year 2023 shows a sharp decrease in the data. It is important to note that data was only collected up to March in 2023.

# Create a data frame for years
year_df <- df %>%
  select(Start_Time) %>%
  mutate(year = year(ymd_hms(Start_Time))) %>%
  group_by(year) %>%
  summarise(n = length(Start_Time), .groups = 'keep') %>%
  data.frame()
year_df <- year_df[!is.na(year_df$year), ]

# Create a line plot for the number of accidents every year
ggplot(year_df, aes(x = year, y = n)) +
  geom_line(color = "lightblue", linewidth = 1) +
  geom_point(color = "red", size = 2) +
  geom_text(aes(label = scales::comma(n)), vjust = -0.5, size = 3.5) +
  labs(title = "Number of Accidents Over the Years", x = "Year", y = "Accident Count") +
  theme_light() +
  theme(plot.title = element_text(hjust = 0.5)) +
  scale_y_continuous(labels = comma) +
  scale_x_continuous(breaks = 2016:2023)

Proportion of Accidents by Traffic Signal

The pie chart shows the proportion of accidents that occurred at locations with or without traffic signals. There is a significant difference in the accident count as 85.2% of collisions took place without a traffic signal and 14.8% took place with a traffic signal.

# Count the traffic signal values
traffic_signal_count <- data.frame(count(df, Traffic_Signal))

# Change true/false to yes/no
traffic_signal_count$Traffic_Signal <- ifelse(traffic_signal_count$Traffic_Signal == TRUE, "Yes", "No")

# Use plotly to create a pie chart for the proportion of accidents by traffic signal
plot_ly(traffic_signal_count, labels = ~Traffic_Signal, values = ~n, type = "pie", 
        textposition = "outside", textinfo = "label + percent", 
        marker = list(colors = c("lightblue", "navyblue"))) %>%
  layout(title = "Accidents by Traffic Signal")

Accident Severity by Weather Condition

The trellis chart shows the number of accidents at the different severity levels across the top 6 weather conditions. Out of the total 109 weather conditions, the six with the highest accident counts were selected for analysis. Note that scales=”free_y” was used to adjust the y-axis for each weather condition to make it easier to visualize the differences in accident severity within each condition. “Fair” weather contains the highest number of traffic accidents. For all six of the conditions, the accident count is the largest at severity level 2.

# Count the number of accidents by weather condition and select the top 6 conditions
weather_count <- data.frame(count(df, Weather_Condition))
top_weather <- weather_count[order(-weather_count$n), ][1:6, ]

# Create trellis chart for accident severity by weather condition
weather_df <- df[df$Weather_Condition %in% top_weather$Weather_Condition, ]
ggplot(weather_df, aes(x = factor(Severity), fill = factor(Severity))) +
  geom_bar() +
  facet_wrap(~Weather_Condition, scales = "free_y") +
  labs(title = "Accident Severity by Weather Condition", x = "Severity", y = "Number of Accidents") +
  scale_y_continuous(labels = scales::comma) +
  scale_fill_brewer(palette = "RdYlBu", direction = -1) +
  theme_light() +
  theme(plot.title = element_text(hjust = 0.5))

Conclusion

This report analyzed multiple variables in the United States accidents dataset to find patterns in the data. The visualizations revealed that California had the highest accident count among the 49 states and severity level 2 was the most common across the states and weather conditions. The number of accidents increased from 2016 to 2022 and more accidents occurred at locations without a traffic signal. Further analysis can be done on other variables to get a deeper understanding of the trends in road accidents.