Setting up my Data base

I loaded tidyverse and ggplot2 so that I can manipulate my data and then create any graphs with the data.

library(tidyverse)
library(ggplot2)
library(tinytex)
library(kableExtra)

Loading Data Set

I used SQL to clean the data set “SDOT Collisions - All Years.” I found this data set on Data.gov.

accidents_table <- read.csv("C:\\Users\\charl\\OneDrive\\Documents\\accidents_table.csv")
accidents_table_percent <- read.csv("C:\\Users\\charl\\OneDrive\\Documents\\accidents_table_percents.csv")

Seattle Number of Accidents from 2004 to 2022

Here, I inserted my table to present my two data sets: Total number of automobile accidents from 2004 to 2022. I wanted to focus on how many collisions involved cyclysists and pedestrians.

accidents_table %>% 
  kbl(caption = "Seattle Number of Accidents from 2004 to 2022") %>% 
  kable_styling() 
Seattle Number of Accidents from 2004 to 2022
year num_of_accidents num_of_fatalities num_of_pedestrian_collision num_of_cyclist_collision
2004 12117 30 386 232
2005 15349 26 474 279
2006 15513 28 559 354
2007 14715 14 475 346
2008 13873 20 459 351
2009 11958 24 443 360
2010 11005 17 485 358
2011 11072 9 385 359
2012 10440 17 473 369
2013 10159 22 399 394
2014 11820 16 478 422
2015 12976 16 502 464
2016 11131 21 517 412
2017 10682 23 510 362
2018 10106 13 534 379
2019 8993 22 523 393
2020 5515 19 276 190
2021 5968 32 318 223
2022 3027 14 145 111

Percent Data Set

Here, I took the data from the first table and found the percent of each collision type based on the total number of accidents.

accidents_table_percent %>% 
  kbl(caption = "Percentage of Collision types based on Total number of accidents") %>% 
  kable_styling()
Percentage of Collision types based on Total number of accidents
year percent_of_fatalities percent_of_pedestrian_collisions percent_of_cyclists_collisions
2004 0.2475860 3.185607 1.914665
2005 0.1693921 3.088149 1.817708
2006 0.1804938 3.603429 2.281957
2007 0.0951410 3.227999 2.351342
2008 0.1441649 3.308585 2.530094
2009 0.2007025 3.704633 3.010537
2010 0.1544752 4.407088 3.253067
2011 0.0812861 3.477240 3.242413
2012 0.1628352 4.530651 3.534483
2013 0.2165567 3.927552 3.878334
2014 0.1353638 4.043993 3.570220
2015 0.1233046 3.868681 3.575832
2016 0.1886623 4.644686 3.701375
2017 0.2153155 4.774387 3.388879
2018 0.1286365 5.283990 3.750247
2019 0.2446347 5.815634 4.370066
2020 0.3445150 5.004533 3.445150
2021 0.5361930 5.328418 3.736595
2022 0.4625041 4.790221 3.666997

Bar Graph

I wanted to show the number of accidents per year in Seattle in order to see any trends.

ggplot(accidents_table, aes(x = year))+
  geom_line(aes(y = num_of_accidents))+
  labs(title = "Number of Total Accidents per Year", x = "Year", y = "Number of Accidents")

accidents <- ggplot(accidents_table, aes(x = year))+  
  geom_line(aes(y = num_of_pedestrian_collision, color = "Number of Pedestrians Collisions"))+
  geom_line(aes(y = num_of_cyclist_collision, color = "Number of Cyclist Collisions"))
print(accidents + labs(title = "Cyclists Collisions vs Pedestrians Collisions", x = "Year", y = "Number of Collisions", colour = "Collision Type"))

Summary

I wanted to showcased the trend of automobile collisions over the last 18 years in Seattle, WA. The general trend is that the total number of collisions are going down with the peak of 15,513 in 2006. What is important to note is that there is a trend with more pedestrian accidents than cyclists. This could indicate that there needs to be more safety implications implemented for pedestrians in the city of Seattle. For future studies, I would want to look at data from other cities to see if it shows the same trends and to see if there is potentially lower pedestrian and bike collisions in other cities. If so, I would look to see if there is infrastructure in place that lowers pedestrian and bike collisions.