── Attaching core tidyverse packages ──────────────────────── tidyverse 2.0.0 ──
✔ dplyr 1.2.1 ✔ readr 2.2.0
✔ forcats 1.0.1 ✔ stringr 1.6.0
✔ ggplot2 4.0.3 ✔ tibble 3.3.1
✔ lubridate 1.9.5 ✔ tidyr 1.3.2
✔ purrr 1.2.2
── Conflicts ────────────────────────────────────────── tidyverse_conflicts() ──
✖ dplyr::filter() masks stats::filter()
✖ dplyr::lag() masks stats::lag()
ℹ Use the conflicted package (<http://conflicted.r-lib.org/>) to force all conflicts to become errors
library(knitr)setwd("C:/Users/maria/OneDrive/School/MC - Business Analyst Cert/DATA110/RStudio Stuff")hatecrimes<-read_csv("NYPD_Hate_Crimes_19-26.csv")
Rows: 4029 Columns: 14
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (9): Record Create Date, Patrol Borough Name, County, Law Code Category ...
dbl (4): Full Complaint ID, Complaint Year Number, Month Number, Complaint P...
lgl (1): Arrest Date
ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
bias_count |>head(10) |>ggplot(aes(x=reorder(biasmotivedescription,n),y=n))+geom_col()+coord_flip()+labs(x="",y="Counts of hatecrime types based on motive",title ="Bar Graph of Hate Crimes from 2019-2026",subtitle="Counts based on the hatecrime motive",caption ="Source: NY State Division of Criminal Justice Services")
bias_count |>head(10) |>ggplot(aes(x=reorder(biasmotivedescription,n),y=n))+geom_col(fill="salmon")+coord_flip()+labs(x="",y="Counts of hatecrime types based on motive",title ="Bar Graph of Hate Crimes from 2019-2026",subtitle="Counts based on the hatecrime motive",caption ="Source: NY State Division of Criminal Justice Services")+theme_minimal()
bias_count |>head(10) |>ggplot(aes(x=reorder(biasmotivedescription,n),y=n))+geom_col(fill="salmon")+coord_flip()+labs(x="",y="Counts of hatecrime types based on motive",title ="Bar Graph of Hate Crimes from 2019-2026",subtitle="Counts based on the hatecrime motive",caption ="Source: NY State Division of Criminal Justice Services")+theme_minimal()+geom_text(aes(label=n),hjust=-0.5,size=3)+theme(axis.text.x =element_blank())
# A tibble: 127 × 4
# Groups: complaintyearnumber, county [35]
complaintyearnumber county biasmotivedescription n
<dbl> <chr> <chr> <int>
1 2024 KINGS ANTI-JEWISH 152
2 2024 NEW YORK ANTI-JEWISH 136
3 2025 KINGS ANTI-JEWISH 136
4 2019 KINGS ANTI-JEWISH 128
5 2023 KINGS ANTI-JEWISH 126
6 2022 KINGS ANTI-JEWISH 125
7 2023 NEW YORK ANTI-JEWISH 124
8 2025 NEW YORK ANTI-JEWISH 110
9 2022 NEW YORK ANTI-JEWISH 104
10 2021 NEW YORK ANTI-ASIAN 84
# ℹ 117 more rows
ggplot(data=hate2)+geom_bar(aes(x=complaintyearnumber,y=n,fill=biasmotivedescription),position="dodge",stat="identity")+labs(fill="Hate Crime Type",y="Number of Hate Crime Incidents",title="Hate Crime Type in NY Counties Between 2010-2016",caption="Source: NY State Division of Criminal Justice Services")
ggplot(data=hate2)+geom_bar(aes(x=county,y=n,fill=biasmotivedescription),position="dodge",stat="identity")+labs(fill="Hate Crime Type",y="Number of Hate Crime Incidents",title="Hate Crime Type in NY Counties Between 2010-2016",caption="Source: NY State Division of Criminal Justice Services")
ggplot(data=hate2)+geom_bar(aes(x=complaintyearnumber,y=n,fill=biasmotivedescription),position="dodge",stat="identity")+facet_wrap(~county)+labs(fill="Hate Crime Type",y="Number of Hate Crime Incidents",title="Hate Crime Type in NY Counties Between 2010-2016",caption="Source: NY State Division of Criminal Justice Services")
Hate Crimes in Counties Per Year By Population Densities
setwd("C:/Users/maria/OneDrive/School/MC - Business Analyst Cert/DATA110/RStudio Stuff")nypop<-read_csv("nyc_census_pop_2020.csv")
Rows: 62 Columns: 4
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Area Name, Population Percent Change
num (2): 2020 Census Population, Population Change
ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
# A tibble: 127 × 5
# Groups: complaintyearnumber, county [35]
complaintyearnumber county biasmotivedescription n 2020 Census Populati…¹
<dbl> <chr> <chr> <int> <dbl>
1 2024 KINGS ANTI-JEWISH 152 NA
2 2024 NEW Y… ANTI-JEWISH 136 NA
3 2025 KINGS ANTI-JEWISH 136 NA
4 2019 KINGS ANTI-JEWISH 128 NA
5 2023 KINGS ANTI-JEWISH 126 NA
6 2022 KINGS ANTI-JEWISH 125 NA
7 2023 NEW Y… ANTI-JEWISH 124 NA
8 2025 NEW Y… ANTI-JEWISH 110 NA
9 2022 NEW Y… ANTI-JEWISH 104 NA
10 2021 NEW Y… ANTI-ASIAN 84 NA
# ℹ 117 more rows
# ℹ abbreviated name: ¹`2020 Census Population`
# A tibble: 127 × 5
# Groups: complaintyearnumber, county [35]
complaintyearnumber county biasmotivedescription n 2020 Census Populati…¹
<dbl> <fct> <chr> <int> <dbl>
1 2024 kings ANTI-JEWISH 152 2736074
2 2024 new y… ANTI-JEWISH 136 1694251
3 2025 kings ANTI-JEWISH 136 2736074
4 2019 kings ANTI-JEWISH 128 2736074
5 2023 kings ANTI-JEWISH 126 2736074
6 2022 kings ANTI-JEWISH 125 2736074
7 2023 new y… ANTI-JEWISH 124 1694251
8 2025 new y… ANTI-JEWISH 110 1694251
9 2022 new y… ANTI-JEWISH 104 1694251
10 2021 new y… ANTI-ASIAN 84 1694251
# ℹ 117 more rows
# ℹ abbreviated name: ¹`2020 Census Population`
datajoinrate<-datajoin |>mutate(rate=n/`2020 Census Population`*100000) |>arrange(desc(rate))datajoinrate
# A tibble: 127 × 6
# Groups: complaintyearnumber, county [35]
complaintyearnumber county biasmotivedescription n 2020 Census Populati…¹
<dbl> <fct> <chr> <int> <dbl>
1 2024 new y… ANTI-JEWISH 136 1694251
2 2023 new y… ANTI-JEWISH 124 1694251
3 2025 new y… ANTI-JEWISH 110 1694251
4 2022 new y… ANTI-JEWISH 104 1694251
5 2024 kings ANTI-JEWISH 152 2736074
6 2025 kings ANTI-JEWISH 136 2736074
7 2021 new y… ANTI-ASIAN 84 1694251
8 2021 new y… ANTI-JEWISH 84 1694251
9 2019 kings ANTI-JEWISH 128 2736074
10 2023 kings ANTI-JEWISH 126 2736074
# ℹ 117 more rows
# ℹ abbreviated name: ¹`2020 Census Population`
# ℹ 1 more variable: rate <dbl>
In the hatecrimes dataset, the positive aspects are that the variables are descriptive, especially the “bias motive description” which made it easy to create categories and do a count for the bar chart. A negative aspect would be from the variable “PD Code Description” because there are multiple descriptions that would make it hard to get data that could be useful for visualization.
One path I could hypothetically like to study about this dataset in the future is the time of the year (year, months) these crimes mostly happened and how it could correlate with social movements or political events at the time. For example, Anti-Asian hate increased in 2020 and 2021 due to the leaked news that the COVID-19 virus allegedly originated from laboratory in China, and this could possibly be corroborated with the presentation of this data. Another path I could hypothetically study is through the “Offense Category” that categorizes the crime as either religious, race, sexual orientation, gender, etc. and compare the counts for those variables.