Week 3 Homework Assignment Summary

Author

Konci Lawrence

Week 3 Homework Assignment Summary

Name: Konci Lawrence

Due date: 9/20/2026

Flawed hate crime data collection

library(tidyverse)
── Attaching core tidyverse packages ──────────────────────── tidyverse 2.0.0 ──
✔ dplyr     1.2.1     ✔ readr     2.2.0
✔ forcats   1.0.1     ✔ stringr   1.6.0
✔ ggplot2   4.0.3     ✔ tibble    3.3.1
✔ lubridate 1.9.5     ✔ tidyr     1.3.2
✔ purrr     1.2.2     
── Conflicts ────────────────────────────────────────── tidyverse_conflicts() ──
✖ dplyr::filter() masks stats::filter()
✖ dplyr::lag()    masks stats::lag()
ℹ Use the conflicted package (<http://conflicted.r-lib.org/>) to force all conflicts to become errors
library(knitr)
url <- "https://data.cityofnewyork.us/api/v3/views/bqiq-cu78/query.csv"
hatecrimes <- read_csv(url)
Rows: 4408 Columns: 18
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr  (10): :id, :version, patrol_borough_name, county, law_code_category_des...
dbl   (4): full_complaint_id, complaint_year_number, month_number, complaint...
lgl   (1): arrest_date
dttm  (3): :created_at, :updated_at, record_create_date

ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.

Clean up the data

names(hatecrimes) <- tolower(names(hatecrimes))
names(hatecrimes) <- gsub(" ","",names(hatecrimes))
head(hatecrimes)
# A tibble: 6 × 18
  `:id`     `:version` `:created_at`       `:updated_at`       full_complaint_id
  <chr>     <chr>      <dttm>              <dttm>                          <dbl>
1 row-quav… rv-kfkc~8… 2026-07-27 15:18:04 2026-07-27 15:18:04           2.02e14
2 row-mwkr… rv-fsqr-p… 2026-07-27 15:18:04 2026-07-27 15:18:04           2.02e14
3 row-mn3m… rv-d9cd-2… 2026-07-27 15:18:04 2026-07-27 15:18:04           2.02e14
4 row-4wzc… rv-yibn~9… 2026-07-27 15:18:04 2026-07-27 15:18:04           2.02e14
5 row-emmq… rv-ag2g.2… 2026-07-27 15:18:04 2026-07-27 15:18:04           2.02e14
6 row-ibxi… rv-tcdx.6… 2026-07-27 15:18:04 2026-07-27 15:18:04           2.02e14
# ℹ 13 more variables: complaint_year_number <dbl>, month_number <dbl>,
#   record_create_date <dttm>, complaint_precinct_code <dbl>,
#   patrol_borough_name <chr>, county <chr>,
#   law_code_category_description <chr>, offense_description <chr>,
#   pd_code_description <chr>, bias_motive_description <chr>,
#   offense_category <chr>, arrest_date <lgl>, arrest_id <chr>

Explore the bias motive description variable

bias_count <- hatecrimes |>
  select(bias_motive_description) |>
  group_by(bias_motive_description) |>
  count() |>
  arrange(desc(n))
head(bias_count)
# A tibble: 6 × 2
# Groups:   bias_motive_description [6]
  bias_motive_description        n
  <chr>                      <int>
1 ANTI-JEWISH                 2113
2 ANTI-MALE HOMOSEXUAL (GAY)   521
3 ANTI-ASIAN                   414
4 ANTI-BLACK                   341
5 ANTI-MUSLIM                  181
6 ANTI-OTHER ETHNICITY         174

Visualize counts as bar graph

ggplot(hatecrimes, aes(x = bias_motive_description))+
  geom_bar()

Use inclusion/exclusion criteria as filter

bias_count |>
  head(10) |>
  ggplot(aes(x=bias_motive_description, y = n)) +
  geom_col()

Arrange the bars according to height and rotate

bias_count |>
  head(10) |>
  ggplot(aes(x=reorder(bias_motive_description, n), y = n)) +
  geom_col() +
  coord_flip()

Add title, caption for data source, and x-axis label

bias_count |>
  head(10) |>
  ggplot(aes(x=reorder(bias_motive_description, n), y = n)) +
  geom_col() +
  coord_flip()+
  labs(x = "",
       y = "Counts of hatecrime types based on motive",
       title = "Bar Graph of Hate Crimes from 2019-2026",
       subtitle = "Counts based on the hatecrime motive",
       caption = "Source: NY State Division of Criminal Justice Services")

Add color and change theme

bias_count |>
  head(10) |>
  ggplot(aes(x=reorder(bias_motive_description, n), y = n)) +
  geom_col(fill = "salmon") +
  coord_flip()+
  labs(x = "",
       y = "Counts of hatecrime types based on motive",
       title = "Bar Graph of Hate Crimes from 2019-2026",
       subtitle = "Counts based on the hatecrime motive",
       caption = "Source: NY State Division of Criminal Justice Services") +
  theme_minimal()

Add annotations for counts and remove x-axis values

bias_count |>
  head(10) |>
  ggplot(aes(x=reorder(bias_motive_description, n), y = n)) +
  geom_col(fill = "salmon") +
  coord_flip()+
  labs(x = "",
       y = "Counts of hatecrime types based on motive",
       title = "Bar Graph of Hate Crimes from 2019-2026",
       subtitle = "Counts based on the hatecrime motive",
       caption = "Source: NY State Division of Criminal Justice Services") +
  theme_minimal()+
  geom_text(aes(label = n), hjust = -.05, size = 3) +
  theme(axis.text.x = element_blank())

Look deeper into crimes by ethnicity

hate_year <- hatecrimes |>
  filter(bias_motive_description %in% c("ANTI-JEWISH", "ANTI-MALE HOMOSEXUAL (GAY)", "ANTI-ASIAN", "ANTI-BLACK"))|>
  group_by(complaint_year_number) |>
  count(bias_motive_description)|>
  arrange(desc(n))
hate_year
# A tibble: 32 × 3
# Groups:   complaint_year_number [8]
   complaint_year_number bias_motive_description        n
                   <dbl> <chr>                      <int>
 1                  2024 ANTI-JEWISH                  371
 2                  2023 ANTI-JEWISH                  342
 3                  2025 ANTI-JEWISH                  340
 4                  2022 ANTI-JEWISH                  280
 5                  2019 ANTI-JEWISH                  252
 6                  2021 ANTI-JEWISH                  215
 7                  2026 ANTI-JEWISH                  187
 8                  2021 ANTI-ASIAN                   150
 9                  2020 ANTI-JEWISH                  126
10                  2023 ANTI-MALE HOMOSEXUAL (GAY)   116
# ℹ 22 more rows

Check county totals

hate_county <- hatecrimes |>
  filter(bias_motive_description %in% c("ANTI-JEWISH", "ANTI-MALE HOMOSEXUAL (GAY)", "ANTI-ASIAN", "ANTI-BLACK"))|>
  group_by(county) |>
  count(bias_motive_description)|>
  arrange(desc(n))
hate_county
# A tibble: 20 × 3
# Groups:   county [5]
   county   bias_motive_description        n
   <chr>    <chr>                      <int>
 1 KINGS    ANTI-JEWISH                  873
 2 NEW YORK ANTI-JEWISH                  711
 3 QUEENS   ANTI-JEWISH                  338
 4 NEW YORK ANTI-MALE HOMOSEXUAL (GAY)   252
 5 NEW YORK ANTI-ASIAN                   235
 6 KINGS    ANTI-MALE HOMOSEXUAL (GAY)   125
 7 BRONX    ANTI-JEWISH                  103
 8 KINGS    ANTI-BLACK                   101
 9 NEW YORK ANTI-BLACK                   101
10 QUEENS   ANTI-MALE HOMOSEXUAL (GAY)    96
11 RICHMOND ANTI-JEWISH                   88
12 KINGS    ANTI-ASIAN                    85
13 QUEENS   ANTI-ASIAN                    79
14 QUEENS   ANTI-BLACK                    75
15 BRONX    ANTI-MALE HOMOSEXUAL (GAY)    42
16 RICHMOND ANTI-BLACK                    35
17 BRONX    ANTI-BLACK                    29
18 BRONX    ANTI-ASIAN                    10
19 RICHMOND ANTI-MALE HOMOSEXUAL (GAY)     6
20 RICHMOND ANTI-ASIAN                     5

Combine county totals and years

hate2 <- hatecrimes |>
  filter(bias_motive_description %in% c("ANTI-JEWISH", "ANTI-MALE HOMOSEXUAL (GAY)", "ANTI-ASIAN", "ANTI-BLACK"))|>
  group_by(complaint_year_number, county) |>
  count(bias_motive_description)|>
  arrange(desc(n))
hate2
# A tibble: 143 × 4
# Groups:   complaint_year_number, county [40]
   complaint_year_number county   bias_motive_description     n
                   <dbl> <chr>    <chr>                   <int>
 1                  2024 KINGS    ANTI-JEWISH               151
 2                  2025 KINGS    ANTI-JEWISH               142
 3                  2024 NEW YORK ANTI-JEWISH               137
 4                  2019 KINGS    ANTI-JEWISH               128
 5                  2023 KINGS    ANTI-JEWISH               126
 6                  2022 KINGS    ANTI-JEWISH               125
 7                  2025 NEW YORK ANTI-JEWISH               124
 8                  2023 NEW YORK ANTI-JEWISH               123
 9                  2022 NEW YORK ANTI-JEWISH               104
10                  2021 NEW YORK ANTI-ASIAN                 84
# ℹ 133 more rows

Plot hate crimes together

ggplot(data = hate2) +
  geom_bar(aes(x=complaint_year_number, y=n, fill = bias_motive_description),
           position = "dodge", stat = "identity") +
  labs(fill = "Hate Crime Type",
       y = "Number of Hate Crime Incidents",
       title = "Hate Crime Type in NY Counties Between 2010-2016",
       caption = "Source: NY State Division of Criminal Justice Services")

Make bar graphs by counties

ggplot(data = hate2) +
  geom_bar(aes(x=county, y=n, fill = bias_motive_description),
           position = "dodge", stat = "identity") +
  labs(fill = "Hate Crime Type",
       y = "Number of Hate Crime Incidents",
       title = "Hate Crime Type in NY Counties Between 2010-2016",
       caption = "Source: NY State Division of Criminal Justice Services")

Put it all together using “facet”

ggplot(data = hate2) +
  geom_bar(aes(x=complaint_year_number, y=n, fill = bias_motive_description),
           position = "dodge", stat = "identity") +
  facet_wrap(~county) +
  labs(fill = "Hate Crime Type",
       y = "Number of Hate Crime Incidents",
       title = "Hate Crime Type in NY Counties Between 2010-2016",
       caption = "Source: NY State Division of Criminal Justice Services")

Counties per year by population density

url <- "https://data.ny.gov/api/v3/views/8wqk-8put/query.csv"
nypop <- read.csv(url)
head(nypop)
  area_type          area_name X_2010_census_population
1    County      Albany County                   304204
2    County    Allegany County                    48946
3    County       Bronx County                  1385108
4    County      Broome County                   200600
5    County Cattaraugus County                    80317
6    County      Cayuga County                    80026
  X_2020_census_population population_change population_percent_change
1                   314848             10644                    0.0350
2                    46456             -2490                   -0.0509
3                  1472654             87546                    0.0632
4                   198683             -1917                   -0.0096
5                    77042             -3275                   -0.0408
6                    76248             -3778                   -0.0472
                X.id         X.version             X.created_at
1 row-jf6e-9e4x_ggc8 rv-k2gu.2emy_k8eu 2022-01-03T17:19:22.938Z
2 row-k2v8_n9y5-fkxq rv-d4yw-9u74.4rsf 2022-01-03T17:19:22.938Z
3 row-7q86_jjjz.7j7h rv-xu6v.p8mn~8wkt 2022-01-03T17:19:22.938Z
4 row-pf2m~cznk.jq85 rv-9c9w_irz9_pdud 2022-01-03T17:19:22.938Z
5 row-x7gs-xrnj~a2iu rv-zrr3.8aks_pzix 2022-01-03T17:19:22.938Z
6 row-fqea.h2g5~h4p8 rv-ktaf-ir73_87m7 2022-01-03T17:19:22.938Z
              X.updated_at
1 2022-01-03T17:19:22.938Z
2 2022-01-03T17:19:22.938Z
3 2022-01-03T17:19:22.938Z
4 2022-01-03T17:19:22.938Z
5 2022-01-03T17:19:22.938Z
6 2022-01-03T17:19:22.938Z

Clean county name to match datasets

names(nypop) <- gsub(" County", "", names(nypop))
head(nypop)
  area_type          area_name X_2010_census_population
1    County      Albany County                   304204
2    County    Allegany County                    48946
3    County       Bronx County                  1385108
4    County      Broome County                   200600
5    County Cattaraugus County                    80317
6    County      Cayuga County                    80026
  X_2020_census_population population_change population_percent_change
1                   314848             10644                    0.0350
2                    46456             -2490                   -0.0509
3                  1472654             87546                    0.0632
4                   198683             -1917                   -0.0096
5                    77042             -3275                   -0.0408
6                    76248             -3778                   -0.0472
                X.id         X.version             X.created_at
1 row-jf6e-9e4x_ggc8 rv-k2gu.2emy_k8eu 2022-01-03T17:19:22.938Z
2 row-k2v8_n9y5-fkxq rv-d4yw-9u74.4rsf 2022-01-03T17:19:22.938Z
3 row-7q86_jjjz.7j7h rv-xu6v.p8mn~8wkt 2022-01-03T17:19:22.938Z
4 row-pf2m~cznk.jq85 rv-9c9w_irz9_pdud 2022-01-03T17:19:22.938Z
5 row-x7gs-xrnj~a2iu rv-zrr3.8aks_pzix 2022-01-03T17:19:22.938Z
6 row-fqea.h2g5~h4p8 rv-ktaf-ir73_87m7 2022-01-03T17:19:22.938Z
              X.updated_at
1 2022-01-03T17:19:22.938Z
2 2022-01-03T17:19:22.938Z
3 2022-01-03T17:19:22.938Z
4 2022-01-03T17:19:22.938Z
5 2022-01-03T17:19:22.938Z
6 2022-01-03T17:19:22.938Z
nypop2 <- nypop |>
  rename(county = `area_name`)|>
  select(county, `X_2020_census_population`)
head(nypop2)
              county X_2020_census_population
1      Albany County                   314848
2    Allegany County                    46456
3       Bronx County                  1472654
4      Broome County                   198683
5 Cattaraugus County                    77042
6      Cayuga County                    76248

Join hate2 data with nypop

hate_new <- hate2 |>
  mutate(county = as_factor(str_to_lower(as.character(county))))
nypop_new <- nypop2 |>
  mutate(county = as_factor(str_to_lower(as.character(county)))) #enure that counties are in lowercase

datajoin <- left_join(hate_new, nypop_new, by=c("county"))
datajoin
# A tibble: 143 × 5
# Groups:   complaint_year_number, county [40]
   complaint_year_number county   bias_motive_description     n
                   <dbl> <fct>    <chr>                   <int>
 1                  2024 kings    ANTI-JEWISH               151
 2                  2025 kings    ANTI-JEWISH               142
 3                  2024 new york ANTI-JEWISH               137
 4                  2019 kings    ANTI-JEWISH               128
 5                  2023 kings    ANTI-JEWISH               126
 6                  2022 kings    ANTI-JEWISH               125
 7                  2025 new york ANTI-JEWISH               124
 8                  2023 new york ANTI-JEWISH               123
 9                  2022 new york ANTI-JEWISH               104
10                  2021 new york ANTI-ASIAN                 84
# ℹ 133 more rows
# ℹ 1 more variable: X_2020_census_population <int>

Calculate rate of incidents per 100,000 and arrange in descending order

datajoinrate <- datajoin |>
  mutate(rate = n/`X_2020_census_population`* 100000) |>
  arrange(desc(rate))
datajoinrate
# A tibble: 143 × 6
# Groups:   complaint_year_number, county [40]
   complaint_year_number county   bias_motive_description     n
                   <dbl> <fct>    <chr>                   <int>
 1                  2024 kings    ANTI-JEWISH               151
 2                  2025 kings    ANTI-JEWISH               142
 3                  2024 new york ANTI-JEWISH               137
 4                  2019 kings    ANTI-JEWISH               128
 5                  2023 kings    ANTI-JEWISH               126
 6                  2022 kings    ANTI-JEWISH               125
 7                  2025 new york ANTI-JEWISH               124
 8                  2023 new york ANTI-JEWISH               123
 9                  2022 new york ANTI-JEWISH               104
10                  2021 new york ANTI-ASIAN                 84
# ℹ 133 more rows
# ℹ 2 more variables: X_2020_census_population <int>, rate <dbl>

Summary

The hate crimes dataset, “NYC hate crimes 2019-2026”, by R Saidi acknowledges that hate crime data collection is flawed as local law enforcement agencies cannot be compelled to submit data to the Federal Bureau of Information (FBI). In consideration of this flaw in data, the positive aspects of this dataset are the ways in which the flaws of hate crime data collection is shown. An example of this is the disproportion in the reported bias motive descriptions being majorly anti-Jewish. This gap is even more evident in its visualization as a bar graph as Jewish hate crimes exceed 1500 but other hate crimes against gay males, Asian people, and black people do not exceed 500. The acknowledgement of flawed data collection is sufficient to assume a bias in reporting of hate crimes. On the other hand, the negative aspects of this data set is the mistake of heading bar graphs with the years between 2010-2016 when the data covers 2019 to 2026. Two different paths I would like to hypothetically study about this data set are the demographics of the population and the rate of hate crimes outside of New York and Kings counties.