I loaded tidyverse and ggplot2 so that I can manipulate my data and then create any graphs with the data.
library(tidyverse)
library(ggplot2)
library(tinytex)
library(kableExtra)
I used SQL to clean the data set “SDOT Collisions - All Years.” I found this data set on Data.gov.
accidents_table <- read.csv("C:\\Users\\charl\\OneDrive\\Documents\\accidents_table.csv")
accidents_table_percent <- read.csv("C:\\Users\\charl\\OneDrive\\Documents\\accidents_table_percents.csv")
Here, I inserted my table to present my two data sets: Total number of automobile accidents from 2004 to 2022. I wanted to focus on how many collisions involved cyclysists and pedestrians.
accidents_table %>%
kbl(caption = "Seattle Number of Accidents from 2004 to 2022") %>%
kable_styling()
| year | num_of_accidents | num_of_fatalities | num_of_pedestrian_collision | num_of_cyclist_collision |
|---|---|---|---|---|
| 2004 | 12117 | 30 | 386 | 232 |
| 2005 | 15349 | 26 | 474 | 279 |
| 2006 | 15513 | 28 | 559 | 354 |
| 2007 | 14715 | 14 | 475 | 346 |
| 2008 | 13873 | 20 | 459 | 351 |
| 2009 | 11958 | 24 | 443 | 360 |
| 2010 | 11005 | 17 | 485 | 358 |
| 2011 | 11072 | 9 | 385 | 359 |
| 2012 | 10440 | 17 | 473 | 369 |
| 2013 | 10159 | 22 | 399 | 394 |
| 2014 | 11820 | 16 | 478 | 422 |
| 2015 | 12976 | 16 | 502 | 464 |
| 2016 | 11131 | 21 | 517 | 412 |
| 2017 | 10682 | 23 | 510 | 362 |
| 2018 | 10106 | 13 | 534 | 379 |
| 2019 | 8993 | 22 | 523 | 393 |
| 2020 | 5515 | 19 | 276 | 190 |
| 2021 | 5968 | 32 | 318 | 223 |
| 2022 | 3027 | 14 | 145 | 111 |
Here, I took the data from the first table and found the percent of each collision type based on the total number of accidents.
accidents_table_percent %>%
kbl(caption = "Percentage of Collision types based on Total number of accidents") %>%
kable_styling()
| year | percent_of_fatalities | percent_of_pedestrian_collisions | percent_of_cyclists_collisions |
|---|---|---|---|
| 2004 | 0.2475860 | 3.185607 | 1.914665 |
| 2005 | 0.1693921 | 3.088149 | 1.817708 |
| 2006 | 0.1804938 | 3.603429 | 2.281957 |
| 2007 | 0.0951410 | 3.227999 | 2.351342 |
| 2008 | 0.1441649 | 3.308585 | 2.530094 |
| 2009 | 0.2007025 | 3.704633 | 3.010537 |
| 2010 | 0.1544752 | 4.407088 | 3.253067 |
| 2011 | 0.0812861 | 3.477240 | 3.242413 |
| 2012 | 0.1628352 | 4.530651 | 3.534483 |
| 2013 | 0.2165567 | 3.927552 | 3.878334 |
| 2014 | 0.1353638 | 4.043993 | 3.570220 |
| 2015 | 0.1233046 | 3.868681 | 3.575832 |
| 2016 | 0.1886623 | 4.644686 | 3.701375 |
| 2017 | 0.2153155 | 4.774387 | 3.388879 |
| 2018 | 0.1286365 | 5.283990 | 3.750247 |
| 2019 | 0.2446347 | 5.815634 | 4.370066 |
| 2020 | 0.3445150 | 5.004533 | 3.445150 |
| 2021 | 0.5361930 | 5.328418 | 3.736595 |
| 2022 | 0.4625041 | 4.790221 | 3.666997 |
I wanted to show the number of accidents per year in Seattle in order to see any trends.
ggplot(accidents_table, aes(x = year))+
geom_line(aes(y = num_of_accidents))+
labs(title = "Number of Total Accidents per Year", x = "Year", y = "Number of Accidents")
accidents <- ggplot(accidents_table, aes(x = year))+
geom_line(aes(y = num_of_pedestrian_collision, color = "Number of Pedestrians Collisions"))+
geom_line(aes(y = num_of_cyclist_collision, color = "Number of Cyclist Collisions"))
print(accidents + labs(title = "Cyclists Collisions vs Pedestrians Collisions", x = "Year", y = "Number of Collisions", colour = "Collision Type"))
I wanted to showcased the trend of automobile collisions over the last 18 years in Seattle, WA. The general trend is that the total number of collisions are going down with the peak of 15,513 in 2006. What is important to note is that there is a trend with more pedestrian accidents than cyclists. This could indicate that there needs to be more safety implications implemented for pedestrians in the city of Seattle. For future studies, I would want to look at data from other cities to see if it shows the same trends and to see if there is potentially lower pedestrian and bike collisions in other cities. If so, I would look to see if there is infrastructure in place that lowers pedestrian and bike collisions.