Assignment Week 3

Author

Surafel Haile

Assignment Week 3:

library(tidyverse)
Warning: package 'tidyverse' was built under R version 4.5.3
Warning: package 'readr' was built under R version 4.5.3
Warning: package 'dplyr' was built under R version 4.5.3
── Attaching core tidyverse packages ──────────────────────── tidyverse 2.0.0 ──
✔ dplyr     1.2.1     ✔ readr     2.2.0
✔ forcats   1.0.1     ✔ stringr   1.6.0
✔ ggplot2   4.0.2     ✔ tibble    3.3.1
✔ lubridate 1.9.5     ✔ tidyr     1.3.2
✔ purrr     1.2.1     
── Conflicts ────────────────────────────────────────── tidyverse_conflicts() ──
✖ dplyr::filter() masks stats::filter()
✖ dplyr::lag()    masks stats::lag()
ℹ Use the conflicted package (<http://conflicted.r-lib.org/>) to force all conflicts to become errors
library(knitr)
setwd("C:/Users/suraf/OneDrive/New folder")
hatecrimes <- read.csv("NYPD_Hate_Crimes_19-26.csv")

Make all headers lowercase and remove spaces:

names(hatecrimes) <- tolower(names(hatecrimes))
names(hatecrimes) <- gsub(" ","",names(hatecrimes))
head(hatecrimes)
  full.complaint.id complaint.year.number month.number record.create.date
1       2.01906e+14                  2019            1          1/23/2019
2       2.01906e+14                  2019            2          2/25/2019
3       2.01906e+14                  2019            2          2/27/2019
4       2.01906e+14                  2019            4          4/16/2019
5       2.01906e+14                  2019            6          6/20/2019
6       2.01906e+14                  2019            7          7/31/2019
  complaint.precinct.code     patrol.borough.name county
1                      60 PATROL BORO BKLYN SOUTH  KINGS
2                      60 PATROL BORO BKLYN SOUTH  KINGS
3                      60 PATROL BORO BKLYN SOUTH  KINGS
4                      60 PATROL BORO BKLYN SOUTH  KINGS
5                      60 PATROL BORO BKLYN SOUTH  KINGS
6                      60 PATROL BORO BKLYN SOUTH  KINGS
  law.code.category.description     offense.description     pd.code.description
1                        FELONY MISCELLANEOUS PENAL LAW AGGRAVATED HARASSMENT 1
2                        FELONY MISCELLANEOUS PENAL LAW AGGRAVATED HARASSMENT 1
3                        FELONY MISCELLANEOUS PENAL LAW AGGRAVATED HARASSMENT 1
4                        FELONY MISCELLANEOUS PENAL LAW AGGRAVATED HARASSMENT 1
5                        FELONY MISCELLANEOUS PENAL LAW AGGRAVATED HARASSMENT 1
6                        FELONY MISCELLANEOUS PENAL LAW AGGRAVATED HARASSMENT 1
  bias.motive.description            offense.category arrest.date arrest.id
1             ANTI-JEWISH Religion/Religious Practice          NA          
2             ANTI-JEWISH Religion/Religious Practice          NA          
3             ANTI-JEWISH Religion/Religious Practice          NA          
4             ANTI-JEWISH Religion/Religious Practice          NA          
5             ANTI-JEWISH Religion/Religious Practice          NA          
6             ANTI-JEWISH Religion/Religious Practice          NA          

Exploring the data:

glimpse(hatecrimes)
Rows: 4,029
Columns: 14
$ full.complaint.id             <dbl> 2.01906e+14, 2.01906e+14, 2.01906e+14, 2…
$ complaint.year.number         <int> 2019, 2019, 2019, 2019, 2019, 2019, 2019…
$ month.number                  <int> 1, 2, 2, 4, 6, 7, 9, 11, 5, 6, 6, 8, 12,…
$ record.create.date            <chr> "1/23/2019", "2/25/2019", "2/27/2019", "…
$ complaint.precinct.code       <int> 60, 60, 60, 60, 60, 60, 60, 60, 61, 61, …
$ patrol.borough.name           <chr> "PATROL BORO BKLYN SOUTH", "PATROL BORO …
$ county                        <chr> "KINGS", "KINGS", "KINGS", "KINGS", "KIN…
$ law.code.category.description <chr> "FELONY", "FELONY", "FELONY", "FELONY", …
$ offense.description           <chr> "MISCELLANEOUS PENAL LAW", "MISCELLANEOU…
$ pd.code.description           <chr> "AGGRAVATED HARASSMENT 1", "AGGRAVATED H…
$ bias.motive.description       <chr> "ANTI-JEWISH", "ANTI-JEWISH", "ANTI-JEWI…
$ offense.category              <chr> "Religion/Religious Practice", "Religion…
$ arrest.date                   <lgl> NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, …
$ arrest.id                     <chr> "", "", "", "", "", "", "", "", "", "", …
summary(hatecrimes)
 full.complaint.id   complaint.year.number  month.number    record.create.date
 Min.   :2.019e+14   Min.   :2019          Min.   : 1.000   Length:4029       
 1st Qu.:2.021e+14   1st Qu.:2021          1st Qu.: 4.000   Class :character  
 Median :2.023e+14   Median :2023          Median : 6.000   Mode  :character  
 Mean   :2.022e+14   Mean   :2022          Mean   : 6.456                     
 3rd Qu.:2.024e+14   3rd Qu.:2024          3rd Qu.: 9.000                     
 Max.   :2.025e+14   Max.   :2025          Max.   :12.000                     
 complaint.precinct.code patrol.borough.name    county         
 Min.   :  1.00          Length:4029         Length:4029       
 1st Qu.: 19.00          Class :character    Class :character  
 Median : 66.00          Mode  :character    Mode  :character  
 Mean   : 59.41                                                
 3rd Qu.: 90.00                                                
 Max.   :123.00                                                
 law.code.category.description offense.description pd.code.description
 Length:4029                   Length:4029         Length:4029        
 Class :character              Class :character    Class :character   
 Mode  :character              Mode  :character    Mode  :character   
                                                                      
                                                                      
                                                                      
 bias.motive.description offense.category   arrest.date     arrest.id        
 Length:4029             Length:4029        Mode:logical   Length:4029       
 Class :character        Class :character   NA's:4029      Class :character  
 Mode  :character        Mode  :character                  Mode  :character  
                                                                             
                                                                             
                                                                             

Explore the bias motive(bias.motive.description):

bias_count <- hatecrimes |>
  select(bias.motive.description) |>
  group_by(bias.motive.description) |>
  count() |>
  arrange(desc(n))
head(bias_count)
# A tibble: 6 × 2
# Groups:   bias.motive.description [6]
  bias.motive.description        n
  <chr>                      <int>
1 ANTI-JEWISH                 1906
2 ANTI-MALE HOMOSEXUAL (GAY)   489
3 ANTI-ASIAN                   401
4 ANTI-BLACK                   315
5 ANTI-OTHER ETHNICITY         168
6 ANTI-MUSLIM                  156

Counts as a bar graph:

ggplot(hatecrimes, aes(x = bias.motive.description))+
  geom_bar()

Use inclusion/exclusion criteria to filter:

bias_count |>
  head(10) |>
  ggplot(aes(x=bias.motive.description, y = n)) +
  geom_col()

Rotate and arrange the bars according to height:

bias_count |>
  head(10) |>
  ggplot(aes(x=reorder(bias.motive.description,n),y=n))+
  geom_col()+
  coord_flip()

Labels, caption, and title:

bias_count |>
  head(10) |>
  ggplot(aes(x=reorder(bias.motive.description, n), y = n)) +
  geom_col() +
  coord_flip()+
  labs(x = "",
       y = "Counts of hatecrime types based on motive",
       title = "Bar Graph of Hate Crimes from 2019-2026",
       subtitle = "Counts based on the hatecrime motive",
       caption = "Source: NY State Division of Criminal Justice Services")

Finally add color and change the theme:

bias_count |>
  head(10) |>
  ggplot(aes(x=reorder(bias.motive.description, n), y = n)) +
  geom_col(fill = "salmon") +
  coord_flip()+
  labs(x = "",
       y = "Counts of hatecrime types based on motive",
       title = "Bar Graph of Hate Crimes from 2019-2026",
       subtitle = "Counts based on the hatecrime motive",
       caption = "Source: NY State Division of Criminal Justice Services") +
  theme_minimal()

Add annotations for counts and remove the x-axis values:

bias_count |>
  head(10) |>
  ggplot(aes(x=reorder(bias.motive.description, n), y = n)) +
  geom_col(fill = "salmon") +
  coord_flip()+
  labs(x = "",
       y = "Counts of hatecrime types based on motive",
       title = "Bar Graph of Hate Crimes from 2019-2026",
       subtitle = "Counts based on the hatecrime motive",
       caption = "Source: NY State Division of Criminal Justice Services") +
  theme_minimal()+
  geom_text(aes(label = n), hjust = -.05, size = 3) +
  theme(axis.text.x = element_blank())

First check the year totals:

hate_year <- hatecrimes |>
  filter(bias.motive.description %in% c("ANTI-JEWISH", "ANTI-MALE HOMOSEXUAL (GAY)", "ANTI-ASIAN", "ANTI-BLACK"))|>
  group_by(complaint.year.number) |>
  count(bias.motive.description)|>
  arrange(desc(n))
hate_year
# A tibble: 28 × 3
# Groups:   complaint.year.number [7]
   complaint.year.number bias.motive.description        n
                   <int> <chr>                      <int>
 1                  2024 ANTI-JEWISH                  371
 2                  2023 ANTI-JEWISH                  343
 3                  2025 ANTI-JEWISH                  320
 4                  2022 ANTI-JEWISH                  279
 5                  2019 ANTI-JEWISH                  252
 6                  2021 ANTI-JEWISH                  215
 7                  2021 ANTI-ASIAN                   150
 8                  2020 ANTI-JEWISH                  126
 9                  2023 ANTI-MALE HOMOSEXUAL (GAY)   116
10                  2022 ANTI-ASIAN                    91
# ℹ 18 more rows

Then check the county totals:

hate_county <- hatecrimes |>
  filter(bias.motive.description %in% c("ANTI-JEWISH", "ANTI-MALE HOMOSEXUAL (GAY)", "ANTI-ASIAN", "ANTI-BLACK"))|>
  group_by(county) |>
  count(bias.motive.description)|>
  arrange(desc(n))
hate_county
# A tibble: 20 × 3
# Groups:   county [5]
   county   bias.motive.description        n
   <chr>    <chr>                      <int>
 1 KINGS    ANTI-JEWISH                  798
 2 NEW YORK ANTI-JEWISH                  651
 3 QUEENS   ANTI-JEWISH                  289
 4 NEW YORK ANTI-MALE HOMOSEXUAL (GAY)   237
 5 NEW YORK ANTI-ASIAN                   228
 6 KINGS    ANTI-MALE HOMOSEXUAL (GAY)   120
 7 KINGS    ANTI-BLACK                    99
 8 BRONX    ANTI-JEWISH                   92
 9 QUEENS   ANTI-MALE HOMOSEXUAL (GAY)    91
10 KINGS    ANTI-ASIAN                    80
11 NEW YORK ANTI-BLACK                    79
12 QUEENS   ANTI-ASIAN                    78
13 RICHMOND ANTI-JEWISH                   76
14 QUEENS   ANTI-BLACK                    75
15 BRONX    ANTI-MALE HOMOSEXUAL (GAY)    35
16 RICHMOND ANTI-BLACK                    35
17 BRONX    ANTI-BLACK                    27
18 BRONX    ANTI-ASIAN                    10
19 RICHMOND ANTI-MALE HOMOSEXUAL (GAY)     6
20 RICHMOND ANTI-ASIAN                     5

Check information combining totals from counties and years :

hate2 <- hatecrimes |>
  filter(bias.motive.description %in% c("ANTI-JEWISH", "ANTI-MALE HOMOSEXUAL (GAY)", "ANTI-ASIAN", "ANTI-BLACK"))|>
  group_by(complaint.year.number, county) |>
  count(bias.motive.description)|>
  arrange(desc(n))
hate2
# A tibble: 127 × 4
# Groups:   complaint.year.number, county [35]
   complaint.year.number county   bias.motive.description     n
                   <int> <chr>    <chr>                   <int>
 1                  2024 KINGS    ANTI-JEWISH               152
 2                  2024 NEW YORK ANTI-JEWISH               136
 3                  2025 KINGS    ANTI-JEWISH               136
 4                  2019 KINGS    ANTI-JEWISH               128
 5                  2023 KINGS    ANTI-JEWISH               126
 6                  2022 KINGS    ANTI-JEWISH               125
 7                  2023 NEW YORK ANTI-JEWISH               124
 8                  2025 NEW YORK ANTI-JEWISH               110
 9                  2022 NEW YORK ANTI-JEWISH               104
10                  2021 NEW YORK ANTI-ASIAN                 84
# ℹ 117 more rows

Plot these three types of hate crimes together:

ggplot(data = hate2) +
  geom_bar(aes(x=complaint.year.number, y=n, fill = bias.motive.description),
      position = "dodge", stat = "identity") +
  labs(fill = "Hate Crime Type",
       y = "Number of Hate Crime Incidents",
       title = "Hate Crime Type in NY Counties Between 2010-2016",
       caption = "Source: NY State Division of Criminal Justice Services")

ggplot(data = hate2) +
  geom_bar(aes(x=county, y=n, fill = bias.motive.description),
      position = "dodge", stat = "identity") +
  labs(fill = "Hate Crime Type",
       y = "Number of Hate Crime Incidents",
       title = "Hate Crime Type in NY Counties Between 2010-2016",
       caption = "Source: NY State Division of Criminal Justice Services")

Put it all together with years and counties using “facet”:

ggplot(data = hate2) +
  geom_bar(aes(x=complaint.year.number, y=n, fill = bias.motive.description),
      position = "dodge", stat = "identity") +
  facet_wrap(~county) +
  labs(fill = "Hate Crime Type",
       y = "Number of Hate Crime Incidents",
       title = "Hate Crime Type in NY Counties Between 2010-2016",
       caption = "Source: NY State Division of Criminal Justice Services")

setwd("C:/Users/suraf/OneDrive/New folder")
nypop <- read_csv("nyc_census_pop_2020.csv")
Rows: 62 Columns: 4
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Area Name, Population Percent Change
num (2): 2020 Census Population, Population Change

ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.

Clean the county name to match the other dataset:

nypop$`Area Name` <- gsub(" County", "", nypop$`Area Name`)
nypop2 <- nypop |>
  rename(county = `Area Name`)|>
  select(county, `2020 Census Population`)
head(nypop2)
# A tibble: 6 × 2
  county      `2020 Census Population`
  <chr>                          <dbl>
1 Albany                        314848
2 Allegany                       46456
3 Bronx                        1472654
4 Broome                        198683
5 Cattaraugus                    77042
6 Cayuga                         76248

Join the hate2 data with nypop:

datajoin <- left_join(hate2, nypop2, by=c("county"))
datajoin
# A tibble: 127 × 5
# Groups:   complaint.year.number, county [35]
   complaint.year.number county   bias.motive.description     n
                   <int> <chr>    <chr>                   <int>
 1                  2024 KINGS    ANTI-JEWISH               152
 2                  2024 NEW YORK ANTI-JEWISH               136
 3                  2025 KINGS    ANTI-JEWISH               136
 4                  2019 KINGS    ANTI-JEWISH               128
 5                  2023 KINGS    ANTI-JEWISH               126
 6                  2022 KINGS    ANTI-JEWISH               125
 7                  2023 NEW YORK ANTI-JEWISH               124
 8                  2025 NEW YORK ANTI-JEWISH               110
 9                  2022 NEW YORK ANTI-JEWISH               104
10                  2021 NEW YORK ANTI-ASIAN                 84
# ℹ 117 more rows
# ℹ 1 more variable: `2020 Census Population` <dbl>

It didn’t work -the new column has NA values:

hate_new <- hate2 |>
  mutate(county = as_factor(str_to_lower(as.character(county))))
nypop_new <- nypop2 |>
  mutate(county = as_factor(str_to_lower(as.character(county))))

Try joining again:

datajoin <- left_join(hate_new, nypop_new, by=c("county"))
datajoin
# A tibble: 127 × 5
# Groups:   complaint.year.number, county [35]
   complaint.year.number county   bias.motive.description     n
                   <int> <fct>    <chr>                   <int>
 1                  2024 kings    ANTI-JEWISH               152
 2                  2024 new york ANTI-JEWISH               136
 3                  2025 kings    ANTI-JEWISH               136
 4                  2019 kings    ANTI-JEWISH               128
 5                  2023 kings    ANTI-JEWISH               126
 6                  2022 kings    ANTI-JEWISH               125
 7                  2023 new york ANTI-JEWISH               124
 8                  2025 new york ANTI-JEWISH               110
 9                  2022 new york ANTI-JEWISH               104
10                  2021 new york ANTI-ASIAN                 84
# ℹ 117 more rows
# ℹ 1 more variable: `2020 Census Population` <dbl>

Calculate the rate of incidents per 100,000. Then arrange in descending order:

datajoinrate <- datajoin |>
  mutate(rate = n/`2020 Census Population`* 100000) |>
  arrange(desc(rate))
datajoinrate
# A tibble: 127 × 6
# Groups:   complaint.year.number, county [35]
   complaint.year.number county   bias.motive.description     n
                   <int> <fct>    <chr>                   <int>
 1                  2024 new york ANTI-JEWISH               136
 2                  2023 new york ANTI-JEWISH               124
 3                  2025 new york ANTI-JEWISH               110
 4                  2022 new york ANTI-JEWISH               104
 5                  2024 kings    ANTI-JEWISH               152
 6                  2025 kings    ANTI-JEWISH               136
 7                  2021 new york ANTI-ASIAN                 84
 8                  2021 new york ANTI-JEWISH                84
 9                  2019 kings    ANTI-JEWISH               128
10                  2023 kings    ANTI-JEWISH               126
# ℹ 117 more rows
# ℹ 2 more variables: `2020 Census Population` <dbl>, rate <dbl>

Questions:

Once you complete this tutorial, include an essay of about 150-200 words which that answers the following questions:

  1. Write about the positive and negative aspects of this hatecrimes data set.

    This data set does indicate whether the crime is anti-black or anti-Asian, but we don’t get any info on the victims or the perpetrators. This means that we know who was targeted, but we don’t know to much else about them. Same for the suspects, we don’t get much on them either. Are they a repeat offender ? or not.

    Positives about this data include the specifics that it does show. Geographic indicators help us compare levels based on boroughs, states, and countries. Bias motive descriptions are helpful for social trends analysis, and arrest data helps us visualize clearance rates.

  2. List 2 different paths you could hypothetically like to study about this data set at some future point.

    An interesting plot would be one that shows enforcement and bias motives . Some interesting possible questions that could be asked include: Which bias motive has the highest rate of clearance ? Is there some inherent bias within enforcement agencies themselves ? We could also look at the relationship between bias motives and months. Are certain bias motives more prevalent during specific time of the month.