Rows: 4029 Columns: 14
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (9): Record Create Date, Patrol Borough Name, County, Law Code Category ...
dbl (4): Full Complaint ID, Complaint Year Number, Month Number, Complaint P...
lgl (1): Arrest Date
ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
bias_count |>head(10) |>ggplot(aes(x = biasmotivedescription, y = n)) +geom_col()
bias_count |>head(10) |>ggplot(aes(x =reorder(biasmotivedescription, n), y = n)) +geom_col() +coord_flip() +labs(x =" ", y ="Counts of hatecrime types based on motive", title ="Bar Graphs of Hate Crimes from 2019 - 2026", subtitle ="Counts based on the hatecrime motive", caption ="Source: NY State Division of Criminal Justice Services") +theme_minimal() +geom_text(aes(label = n), hjust =-.05, size =3) +theme(axis.text.x =element_blank())
# A tibble: 127 × 4
# Groups: complaintyearnumber, county [35]
complaintyearnumber county biasmotivedescription n
<dbl> <chr> <chr> <int>
1 2024 KINGS ANTI-JEWISH 152
2 2024 NEW YORK ANTI-JEWISH 136
3 2025 KINGS ANTI-JEWISH 136
4 2019 KINGS ANTI-JEWISH 128
5 2023 KINGS ANTI-JEWISH 126
6 2022 KINGS ANTI-JEWISH 125
7 2023 NEW YORK ANTI-JEWISH 124
8 2025 NEW YORK ANTI-JEWISH 110
9 2022 NEW YORK ANTI-JEWISH 104
10 2021 NEW YORK ANTI-ASIAN 84
# ℹ 117 more rows
ggplot(data = hate2) +geom_bar(aes(x=complaintyearnumber, y = n, fill = biasmotivedescription), position ="dodge", stat ="identity") +labs(fill ="Hate Crime Type", y ="Number of Hate Crimes Incidents", title ="Hate Crime Type in NY Counties Between 2010-2016", caption ="Source: NY State Division of Criminal Justice Services")
ggplot(data = hate2) +geom_bar(aes(x=county, y=n, fill = biasmotivedescription), position ="dodge", stat ="identity") +labs(fill ="Hate Crime Type", y ="Number of Hate Crimes Incidents", title ="Hate Crime Type in NY Counties Between 2010-2016", caption ="Source: NY State Division of Criminal Justice Services")
ggplot(data = hate2) +geom_bar(aes(x=complaintyearnumber, y = n, fill = biasmotivedescription), position ="dodge", stat ="identity") +facet_wrap(~county) +labs(fill ="Hate Crime Type", y ="Number of Hate Crimes Incidents", title ="Hate Crime Type in NY Counties Between 2010-2016", caption ="Source: NY State Division of Criminal Justice Services")
nypop =read_csv("nyc_census_pop_2020.csv")
Rows: 62 Columns: 4
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Area Name, Population Percent Change
num (2): 2020 Census Population, Population Change
ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
# A tibble: 127 × 5
# Groups: complaintyearnumber, county [35]
complaintyearnumber county biasmotivedescription n 2020 Census Populati…¹
<dbl> <chr> <chr> <int> <dbl>
1 2024 KINGS ANTI-JEWISH 152 NA
2 2024 NEW Y… ANTI-JEWISH 136 NA
3 2025 KINGS ANTI-JEWISH 136 NA
4 2019 KINGS ANTI-JEWISH 128 NA
5 2023 KINGS ANTI-JEWISH 126 NA
6 2022 KINGS ANTI-JEWISH 125 NA
7 2023 NEW Y… ANTI-JEWISH 124 NA
8 2025 NEW Y… ANTI-JEWISH 110 NA
9 2022 NEW Y… ANTI-JEWISH 104 NA
10 2021 NEW Y… ANTI-ASIAN 84 NA
# ℹ 117 more rows
# ℹ abbreviated name: ¹`2020 Census Population`
# A tibble: 127 × 5
# Groups: complaintyearnumber, county [35]
complaintyearnumber county biasmotivedescription n 2020 Census Populati…¹
<dbl> <chr> <chr> <int> <dbl>
1 2024 kings ANTI-JEWISH 152 2736074
2 2024 new y… ANTI-JEWISH 136 1694251
3 2025 kings ANTI-JEWISH 136 2736074
4 2019 kings ANTI-JEWISH 128 2736074
5 2023 kings ANTI-JEWISH 126 2736074
6 2022 kings ANTI-JEWISH 125 2736074
7 2023 new y… ANTI-JEWISH 124 1694251
8 2025 new y… ANTI-JEWISH 110 1694251
9 2022 new y… ANTI-JEWISH 104 1694251
10 2021 new y… ANTI-ASIAN 84 1694251
# ℹ 117 more rows
# ℹ abbreviated name: ¹`2020 Census Population`
Assignment #2 — Ending Essay
Data on hate crimes must be collected to this extent, as it can bring awareness to the scale to which different communities are discriminated against. The USA is commonly referred to as a “melting pot,” especially in major cities like New York, so understanding and analyzing this data is eye-opening about how much we can improve. Furthermore, this data set is massive and provides a variety of different variables to analyze. Dissecting the data allows us to better understand the whole and ensure no major bias. An obvious negative point of this data set is the fact that it exists in the first place. Although this data is important to bring awareness, it is saddening to see how many marginalized groups are persecuted simply for being different. However, the data is inconsistent, as error messages when visualizing it. For example, the counties in the data sets had capitalization differences, making it difficult to combine the necessary data.
The first path that would be interesting to study is the dates when each hate crime was reported. Looking at the dates, as well as the respective marginalized group the hate crime was against, allows us to find patterns. We can find patterns of when certain groups were targeted and use events happening at that time to better understand. A second path is looking for outliers within the data. Major outliers lead to skewed data, and with a data set as complex as this one, it can be difficult to grasp outliers immediately. Investigating outliers and working to comprehend them allows us to understand the data on a deeper level.