This dataset looks at all types of hate crimes in New York counties by the type of hate crime from 2010 to 2016.
Before any analysis, it matters how the data was collected. As Nathan Yau of FlowingData noted (December 5, 2017), data can give you important information, but when the collection process is flawed there is not much you can do about it. Ken Schwencke, reporting for ProPublica, described the tiered system the FBI relies on:
“Under a federal law passed in 1990, the FBI is required to track and tabulate crimes in which there was ‘manifest evidence of prejudice’ against a host of protected groups, regardless of differences in how state laws define who’s protected. The FBI, in turn, relies on local law enforcement agencies to collect and submit this data, but can’t compel them to do so.”
ProPublica article: https://www.propublica.org/article/why-america-fails-at-gathering-hate-crime-statistics
Map of where hate crimes do not get reported: https://projects.propublica.org/graphics/hatecrime-map
So we know there is possible bias in the dataset. What can we do with it?
library(tidyverse)
hatecrimes <- read_csv("hateCrimes2010.csv")
Make all headers lowercase and remove spaces.
names(hatecrimes) <- tolower(names(hatecrimes))
names(hatecrimes) <- gsub(" ", "", names(hatecrimes))
head(hatecrimes)
## # A tibble: 6 × 44
## county year crimetype `anti-male` `anti-female` `anti-transgender`
## <chr> <dbl> <chr> <dbl> <dbl> <dbl>
## 1 Albany 2016 Crimes Against Pe… 0 0 0
## 2 Albany 2016 Property Crimes 0 0 0
## 3 Allegany 2016 Property Crimes 0 0 0
## 4 Bronx 2016 Crimes Against Pe… 0 0 0
## 5 Bronx 2016 Property Crimes 0 0 0
## 6 Broome 2016 Crimes Against Pe… 0 0 0
## # ℹ 38 more variables: `anti-genderidentityexpression` <dbl>,
## # `anti-age*` <dbl>, `anti-white` <dbl>, `anti-black` <dbl>,
## # `anti-americanindian/alaskannative` <dbl>, `anti-asian` <dbl>,
## # `anti-nativehawaiian/pacificislander` <dbl>,
## # `anti-multi-racialgroups` <dbl>, `anti-otherrace` <dbl>,
## # `anti-jewish` <dbl>, `anti-catholic` <dbl>, `anti-protestant` <dbl>,
## # `anti-islamic(muslim)` <dbl>, `anti-multi-religiousgroups` <dbl>, …
With 44 variables, summary helps decide which hate
crimes to focus on. Looking at the min/max values, some categories have
a maximum of only 1.
summary(hatecrimes)
## county year crimetype anti-male
## Length :423 Min. :2010 Length :423 Min. :0.000000
## N.unique : 60 1st Qu.:2011 N.unique : 3 1st Qu.:0.000000
## N.blank : 0 Median :2013 N.blank : 0 Median :0.000000
## Min.nchar: 4 Mean :2013 Min.nchar: 9 Mean :0.007092
## Max.nchar: 12 3rd Qu.:2015 Max.nchar: 22 3rd Qu.:0.000000
## Max. :2016 Max. :1.000000
## anti-female anti-transgender anti-genderidentityexpression
## Min. :0.00000 Min. :0.00000 Min. :0.00000
## 1st Qu.:0.00000 1st Qu.:0.00000 1st Qu.:0.00000
## Median :0.00000 Median :0.00000 Median :0.00000
## Mean :0.01655 Mean :0.05674 Mean :0.04728
## 3rd Qu.:0.00000 3rd Qu.:0.00000 3rd Qu.:0.00000
## Max. :1.00000 Max. :3.00000 Max. :5.00000
## anti-age* anti-white anti-black
## Min. :0.00000 Min. : 0.0000 Min. : 0.000
## 1st Qu.:0.00000 1st Qu.: 0.0000 1st Qu.: 0.000
## Median :0.00000 Median : 0.0000 Median : 1.000
## Mean :0.05201 Mean : 0.3357 Mean : 1.761
## 3rd Qu.:0.00000 3rd Qu.: 0.0000 3rd Qu.: 2.000
## Max. :9.00000 Max. :11.0000 Max. :18.000
## anti-americanindian/alaskannative anti-asian
## Min. :0.000000 Min. :0.0000
## 1st Qu.:0.000000 1st Qu.:0.0000
## Median :0.000000 Median :0.0000
## Mean :0.007092 Mean :0.1773
## 3rd Qu.:0.000000 3rd Qu.:0.0000
## Max. :1.000000 Max. :8.0000
## anti-nativehawaiian/pacificislander anti-multi-racialgroups anti-otherrace
## Min. :0 Min. :0.00000 Min. :0
## 1st Qu.:0 1st Qu.:0.00000 1st Qu.:0
## Median :0 Median :0.00000 Median :0
## Mean :0 Mean :0.08511 Mean :0
## 3rd Qu.:0 3rd Qu.:0.00000 3rd Qu.:0
## Max. :0 Max. :3.00000 Max. :0
## anti-jewish anti-catholic anti-protestant anti-islamic(muslim)
## Min. : 0.000 Min. : 0.0000 Min. :0.00000 Min. : 0.0000
## 1st Qu.: 0.000 1st Qu.: 0.0000 1st Qu.:0.00000 1st Qu.: 0.0000
## Median : 0.000 Median : 0.0000 Median :0.00000 Median : 0.0000
## Mean : 3.981 Mean : 0.2695 Mean :0.02364 Mean : 0.4704
## 3rd Qu.: 3.000 3rd Qu.: 0.0000 3rd Qu.:0.00000 3rd Qu.: 0.0000
## Max. :82.000 Max. :12.0000 Max. :1.00000 Max. :10.0000
## anti-multi-religiousgroups anti-atheism/agnosticism
## Min. : 0.00000 Min. :0
## 1st Qu.: 0.00000 1st Qu.:0
## Median : 0.00000 Median :0
## Mean : 0.07565 Mean :0
## 3rd Qu.: 0.00000 3rd Qu.:0
## Max. :10.00000 Max. :0
## anti-religiouspracticegenerally anti-otherreligion anti-buddhist
## Min. :0.000000 Min. :0.000 Min. :0
## 1st Qu.:0.000000 1st Qu.:0.000 1st Qu.:0
## Median :0.000000 Median :0.000 Median :0
## Mean :0.007092 Mean :0.104 Mean :0
## 3rd Qu.:0.000000 3rd Qu.:0.000 3rd Qu.:0
## Max. :2.000000 Max. :4.000 Max. :0
## anti-easternorthodox(greek,russian,etc.) anti-hindu
## Min. :0.000000 Min. :0.000000
## 1st Qu.:0.000000 1st Qu.:0.000000
## Median :0.000000 Median :0.000000
## Mean :0.002364 Mean :0.002364
## 3rd Qu.:0.000000 3rd Qu.:0.000000
## Max. :1.000000 Max. :1.000000
## anti-jehovahswitness anti-mormon anti-otherchristian anti-sikh
## Min. :0 Min. :0 Min. :0.00000 Min. :0
## 1st Qu.:0 1st Qu.:0 1st Qu.:0.00000 1st Qu.:0
## Median :0 Median :0 Median :0.00000 Median :0
## Mean :0 Mean :0 Mean :0.01891 Mean :0
## 3rd Qu.:0 3rd Qu.:0 3rd Qu.:0.00000 3rd Qu.:0
## Max. :0 Max. :0 Max. :3.00000 Max. :0
## anti-hispanic anti-arab anti-otherethnicity/nationalorigin
## Min. : 0.0000 Min. :0.00000 Min. : 0.0000
## 1st Qu.: 0.0000 1st Qu.:0.00000 1st Qu.: 0.0000
## Median : 0.0000 Median :0.00000 Median : 0.0000
## Mean : 0.3735 Mean :0.06619 Mean : 0.2837
## 3rd Qu.: 0.0000 3rd Qu.:0.00000 3rd Qu.: 0.0000
## Max. :17.0000 Max. :2.00000 Max. :19.0000
## anti-non-hispanic* anti-gaymale anti-gayfemale anti-gay(maleandfemale)
## Min. :0 Min. : 0.000 Min. :0.0000 Min. :0.0000
## 1st Qu.:0 1st Qu.: 0.000 1st Qu.:0.0000 1st Qu.:0.0000
## Median :0 Median : 0.000 Median :0.0000 Median :0.0000
## Mean :0 Mean : 1.499 Mean :0.2411 Mean :0.1017
## 3rd Qu.:0 3rd Qu.: 1.000 3rd Qu.:0.0000 3rd Qu.:0.0000
## Max. :0 Max. :36.000 Max. :8.0000 Max. :4.0000
## anti-heterosexual anti-bisexual anti-physicaldisability
## Min. :0.000000 Min. :0.000000 Min. :0.00000
## 1st Qu.:0.000000 1st Qu.:0.000000 1st Qu.:0.00000
## Median :0.000000 Median :0.000000 Median :0.00000
## Mean :0.002364 Mean :0.004728 Mean :0.01182
## 3rd Qu.:0.000000 3rd Qu.:0.000000 3rd Qu.:0.00000
## Max. :1.000000 Max. :1.000000 Max. :1.00000
## anti-mentaldisability totalincidents totalvictims totaloffenders
## Min. :0.000000 Min. : 1.00 Min. : 1.00 Min. : 1.00
## 1st Qu.:0.000000 1st Qu.: 1.00 1st Qu.: 1.00 1st Qu.: 1.00
## Median :0.000000 Median : 3.00 Median : 3.00 Median : 3.00
## Mean :0.009456 Mean : 10.09 Mean : 10.48 Mean : 11.78
## 3rd Qu.:0.000000 3rd Qu.: 10.00 3rd Qu.: 10.00 3rd Qu.: 11.00
## Max. :1.000000 Max. :101.00 Max. :106.00 Max. :113.00
Only the hate-crime types with a maximum count of 9 or more are kept, so the analysis focuses on the most prominent types.
hatecrimes2 <- hatecrimes |>
select(county, year, `anti-black`, `anti-white`, `anti-jewish`, `anti-catholic`,
`anti-age*`, `anti-islamic(muslim)`, `anti-multi-religiousgroups`,
`anti-gaymale`, `anti-hispanic`, `anti-otherethnicity/nationalorigin`) |>
group_by(county, year)
head(hatecrimes2)
## # A tibble: 6 × 12
## # Groups: county, year [4]
## county year `anti-black` `anti-white` `anti-jewish` `anti-catholic`
## <chr> <dbl> <dbl> <dbl> <dbl> <dbl>
## 1 Albany 2016 1 0 0 0
## 2 Albany 2016 2 0 0 0
## 3 Allegany 2016 1 0 0 0
## 4 Bronx 2016 0 1 0 0
## 5 Bronx 2016 0 1 1 0
## 6 Broome 2016 1 0 0 0
## # ℹ 6 more variables: `anti-age*` <dbl>, `anti-islamic(muslim)` <dbl>,
## # `anti-multi-religiousgroups` <dbl>, `anti-gaymale` <dbl>,
## # `anti-hispanic` <dbl>, `anti-otherethnicity/nationalorigin` <dbl>
dim(hatecrimes2)
## [1] 423 12
There are 12 variables with 423 rows.
summary(hatecrimes2)
## county year anti-black anti-white
## Length :423 Min. :2010 Min. : 0.000 Min. : 0.0000
## N.unique : 60 1st Qu.:2011 1st Qu.: 0.000 1st Qu.: 0.0000
## N.blank : 0 Median :2013 Median : 1.000 Median : 0.0000
## Min.nchar: 4 Mean :2013 Mean : 1.761 Mean : 0.3357
## Max.nchar: 12 3rd Qu.:2015 3rd Qu.: 2.000 3rd Qu.: 0.0000
## Max. :2016 Max. :18.000 Max. :11.0000
## anti-jewish anti-catholic anti-age* anti-islamic(muslim)
## Min. : 0.000 Min. : 0.0000 Min. :0.00000 Min. : 0.0000
## 1st Qu.: 0.000 1st Qu.: 0.0000 1st Qu.:0.00000 1st Qu.: 0.0000
## Median : 0.000 Median : 0.0000 Median :0.00000 Median : 0.0000
## Mean : 3.981 Mean : 0.2695 Mean :0.05201 Mean : 0.4704
## 3rd Qu.: 3.000 3rd Qu.: 0.0000 3rd Qu.:0.00000 3rd Qu.: 0.0000
## Max. :82.000 Max. :12.0000 Max. :9.00000 Max. :10.0000
## anti-multi-religiousgroups anti-gaymale anti-hispanic
## Min. : 0.00000 Min. : 0.000 Min. : 0.0000
## 1st Qu.: 0.00000 1st Qu.: 0.000 1st Qu.: 0.0000
## Median : 0.00000 Median : 0.000 Median : 0.0000
## Mean : 0.07565 Mean : 1.499 Mean : 0.3735
## 3rd Qu.: 0.00000 3rd Qu.: 1.000 3rd Qu.: 0.0000
## Max. :10.00000 Max. :36.000 Max. :17.0000
## anti-otherethnicity/nationalorigin
## Min. : 0.0000
## 1st Qu.: 0.0000
## Median : 0.0000
## Mean : 0.2837
## 3rd Qu.: 0.0000
## Max. :19.0000
pivot_longer collapses each hate-crime column into a
single victim_cat column, with the counts going into
crimecount. This is what makes facet_wrap
possible.
hatelong <- hatecrimes2 |>
pivot_longer(
cols = 3:12,
names_to = "victim_cat",
values_to = "crimecount")
head(hatelong)
## # A tibble: 6 × 4
## # Groups: county, year [1]
## county year victim_cat crimecount
## <chr> <dbl> <chr> <dbl>
## 1 Albany 2016 anti-black 1
## 2 Albany 2016 anti-white 0
## 3 Albany 2016 anti-jewish 0
## 4 Albany 2016 anti-catholic 0
## 5 Albany 2016 anti-age* 0
## 6 Albany 2016 anti-islamic(muslim) 1
hatecrimplot <- hatelong |>
ggplot(aes(year, crimecount)) +
geom_point() +
aes(color = victim_cat) +
facet_wrap(~victim_cat)
hatecrimplot
From the facet plot, the anti-Black, anti-gay-male, and anti-Jewish categories have the highest reported counts. Filter for just those three.
hatenew <- hatelong |>
filter(victim_cat %in% c("anti-black", "anti-jewish", "anti-gaymale")) |>
group_by(year, county) |>
arrange(desc(crimecount))
hatenew
## # A tibble: 1,269 × 4
## # Groups: year, county [277]
## county year victim_cat crimecount
## <chr> <dbl> <chr> <dbl>
## 1 Kings 2012 anti-jewish 82
## 2 Kings 2016 anti-jewish 51
## 3 Suffolk 2014 anti-jewish 48
## 4 Suffolk 2012 anti-jewish 48
## 5 Kings 2011 anti-jewish 44
## 6 Kings 2013 anti-jewish 41
## 7 Kings 2010 anti-jewish 39
## 8 Nassau 2011 anti-jewish 38
## 9 Suffolk 2013 anti-jewish 37
## 10 Nassau 2016 anti-jewish 36
## # ℹ 1,259 more rows
position = "dodge" gives side-by-side rather than
stacked bars, and stat = "identity" plots the counts as
they are.
plot2 <- hatenew |>
ggplot() +
geom_bar(aes(x = year, y = crimecount, fill = victim_cat),
position = "dodge", stat = "identity") +
labs(fill = "Hate Crime Type",
y = "Number of Hate Crime Incidents",
title = "Hate Crime Type in NY Counties Between 2010-2016",
caption = "Source: NY State Division of Criminal Justice Services")
plot2
Hate crimes against Jews spiked in 2012. Other years were relatively consistent with a slight upward trend. There was also an upward trend in hate crimes against gay males, and an apparent downward trend in hate crimes against Blacks over this period.
plot3 <- hatenew |>
ggplot() +
geom_bar(aes(x = county, y = crimecount, fill = victim_cat),
position = "dodge", stat = "identity") +
labs(fill = "Hate Crime Type",
y = "Number of Hate Crime Incidents",
title = "Hate Crime Type in NY Counties Between 2010-2016",
caption = "Source: NY State Division of Criminal Justice Services")
plot3
There are too many counties for that plot to be readable, so group and sum to find the highest.
counties <- hatenew |>
group_by(year, county) |>
summarize(sum = sum(crimecount)) |>
arrange(desc(sum))
counties
## # A tibble: 277 × 3
## # Groups: year [7]
## year county sum
## <dbl> <chr> <dbl>
## 1 2012 Kings 136
## 2 2010 Kings 110
## 3 2016 Kings 101
## 4 2013 Kings 96
## 5 2014 Kings 94
## 6 2015 Kings 90
## 7 2011 Kings 86
## 8 2016 New York 86
## 9 2012 Suffolk 83
## 10 2013 New York 75
## # ℹ 267 more rows
counties2 <- hatenew |>
group_by(county) |>
summarize(sum = sum(crimecount)) |>
slice_max(order_by = sum, n = 5)
counties2
## # A tibble: 5 × 2
## county sum
## <chr> <dbl>
## 1 Kings 713
## 2 New York 459
## 3 Suffolk 360
## 4 Nassau 298
## 5 Queens 235
plot4 <- hatenew |>
filter(county %in% c("Kings", "New York", "Suffolk", "Nassau", "Queens")) |>
ggplot() +
geom_bar(aes(x = county, y = crimecount, fill = victim_cat),
position = "dodge", stat = "identity") +
labs(y = "Number of Hate Crime Incidents",
title = "5 Counties in NY with Highest Incidents of Hate Crimes",
subtitle = "Between 2010-2016",
fill = "Hate Crime Type",
caption = "Source: NY State Division of Criminal Justice Services")
plot4
Raw counts mostly tell us where the people are. Bringing in census population estimates for New York counties lets us compute rates instead.
nypop <- read_csv("newyorkpopulation.csv")
head(nypop)
## # A tibble: 6 × 8
## Geography `2010` `2011` `2012` `2013` `2014` `2015` `2016`
## <chr> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl>
## 1 Albany County, New York 304078 3.05e5 3.06e5 3.07e5 3.08e5 3.08e5 3.09e5
## 2 Allegany County, New York 48949 4.88e4 4.82e4 4.80e4 4.78e4 4.74e4 4.71e4
## 3 Bronx County, New York 1388240 1.40e6 1.41e6 1.43e6 1.44e6 1.45e6 1.46e6
## 4 Broome County, New York 200469 1.99e5 1.99e5 1.98e5 1.98e5 1.97e5 1.95e5
## 5 Cattaraugus County, New York 80249 7.98e4 7.94e4 7.90e4 7.86e4 7.79e4 7.77e4
## 6 Cayuga County, New York 79844 7.98e4 7.96e4 7.92e4 7.89e4 7.83e4 7.79e4
The Geography column reads like “Albany County, New
York” while the hate crimes data uses just “Albany”. Rename
Geography to county and strip the extra text
so the two datasets can be joined.
nypop$Geography <- gsub(" , New York", "", nypop$Geography)
nypop$Geography <- gsub("County", "", nypop$Geography)
nypoplong <- nypop |>
rename(county = Geography) |>
gather("year", "population", 2:8)
nypoplong$year <- as.double(nypoplong$year)
head(nypoplong)
## # A tibble: 6 × 3
## county year population
## <chr> <dbl> <dbl>
## 1 Albany , New York 2010 304078
## 2 Allegany , New York 2010 48949
## 3 Bronx , New York 2010 1388240
## 4 Broome , New York 2010 200469
## 5 Cattaraugus , New York 2010 80249
## 6 Cayuga , New York 2010 79844
Since 2012 had the highest counts, look at county populations in 2012.
nypoplong12 <- nypoplong |>
filter(year == 2012) |>
arrange(desc(population)) |>
head(10)
nypoplong12$county <- gsub(" , New York", "", nypoplong12$county)
nypoplong12
## # A tibble: 10 × 3
## county year population
## <chr> <dbl> <dbl>
## 1 Kings 2012 2572282
## 2 Queens 2012 2278024
## 3 New York 2012 1625121
## 4 Suffolk 2012 1499382
## 5 Bronx 2012 1414774
## 6 Nassau 2012 1350748
## 7 Westchester 2012 961073
## 8 Erie 2012 920792
## 9 Monroe 2012 748947
## 10 Richmond 2012 470978
Four of the five counties with the highest populations are also among those with the highest hate crime counts. Only the Bronx, fifth in population, is missing from the high-count list.
Recall the total hate crime counts: Kings 713, New York 459, Suffolk 360, Nassau 298, Queens 235.
counties12 <- counties |>
filter(year == 2012) |>
arrange(desc(sum))
counties12
## # A tibble: 41 × 3
## # Groups: year [1]
## year county sum
## <dbl> <chr> <dbl>
## 1 2012 Kings 136
## 2 2012 Suffolk 83
## 3 2012 New York 71
## 4 2012 Nassau 48
## 5 2012 Queens 48
## 6 2012 Erie 28
## 7 2012 Bronx 23
## 8 2012 Richmond 18
## 9 2012 Multiple 14
## 10 2012 Westchester 13
## # ℹ 31 more rows
This is the key step. Both tables now have a county
column and a numeric year column, which is what makes them
joinable. The county names had to be cleaned first, and
year had to be converted with as.double after
gather turned the year column headings into character
strings. Without both fixes the join would silently match nothing.
datajoin <- counties12 |>
full_join(nypoplong12, by = c("county", "year"))
datajoin
## # A tibble: 41 × 4
## # Groups: year [1]
## year county sum population
## <dbl> <chr> <dbl> <dbl>
## 1 2012 Kings 136 2572282
## 2 2012 Suffolk 83 1499382
## 3 2012 New York 71 1625121
## 4 2012 Nassau 48 1350748
## 5 2012 Queens 48 2278024
## 6 2012 Erie 28 920792
## 7 2012 Bronx 23 1414774
## 8 2012 Richmond 18 470978
## 9 2012 Multiple 14 NA
## 10 2012 Westchester 13 961073
## # ℹ 31 more rows
datajoinrate <- datajoin |>
mutate(rate = sum/population*100000) |>
arrange(desc(rate))
datajoinrate
## # A tibble: 41 × 5
## # Groups: year [1]
## year county sum population rate
## <dbl> <chr> <dbl> <dbl> <dbl>
## 1 2012 Suffolk 83 1499382 5.54
## 2 2012 Kings 136 2572282 5.29
## 3 2012 New York 71 1625121 4.37
## 4 2012 Richmond 18 470978 3.82
## 5 2012 Nassau 48 1350748 3.55
## 6 2012 Erie 28 920792 3.04
## 7 2012 Queens 48 2278024 2.11
## 8 2012 Bronx 23 1414774 1.63
## 9 2012 Westchester 13 961073 1.35
## 10 2012 Monroe 5 748947 0.668
## # ℹ 31 more rows
dt <- datajoinrate[, c("county", "rate")]
dt
## # A tibble: 41 × 2
## county rate
## <chr> <dbl>
## 1 Suffolk 5.54
## 2 Kings 5.29
## 3 New York 4.37
## 4 Richmond 3.82
## 5 Nassau 3.55
## 6 Erie 3.04
## 7 Queens 2.11
## 8 Bronx 1.63
## 9 Westchester 1.35
## 10 Monroe 0.668
## # ℹ 31 more rows
The counties with the highest 2012 rates are Suffolk, Kings, New York, Richmond, and Nassau. The most populous counties were Kings, Queens, New York, Suffolk, Bronx, and Nassau. Those lists are similar but they do not correspond directly, which is the whole point of computing a rate. Queens drops from second in population to seventh in rate, and the Bronx has a much lower rate than its population would predict.
The individual categories are thin, so collapsing them into broader
groups gives more usable counts. Note that columns 42 through 44 of the
raw data are totalincidents, totalvictims, and
totaloffenders, so the pivot uses columns 4 through 41 to
include only actual victim categories.
aggregategroups <- hatecrimes |>
pivot_longer(
cols = 4:41,
names_to = "victim_cat",
values_to = "crimecount")
unique(aggregategroups$victim_cat)
## [1] "anti-male"
## [2] "anti-female"
## [3] "anti-transgender"
## [4] "anti-genderidentityexpression"
## [5] "anti-age*"
## [6] "anti-white"
## [7] "anti-black"
## [8] "anti-americanindian/alaskannative"
## [9] "anti-asian"
## [10] "anti-nativehawaiian/pacificislander"
## [11] "anti-multi-racialgroups"
## [12] "anti-otherrace"
## [13] "anti-jewish"
## [14] "anti-catholic"
## [15] "anti-protestant"
## [16] "anti-islamic(muslim)"
## [17] "anti-multi-religiousgroups"
## [18] "anti-atheism/agnosticism"
## [19] "anti-religiouspracticegenerally"
## [20] "anti-otherreligion"
## [21] "anti-buddhist"
## [22] "anti-easternorthodox(greek,russian,etc.)"
## [23] "anti-hindu"
## [24] "anti-jehovahswitness"
## [25] "anti-mormon"
## [26] "anti-otherchristian"
## [27] "anti-sikh"
## [28] "anti-hispanic"
## [29] "anti-arab"
## [30] "anti-otherethnicity/nationalorigin"
## [31] "anti-non-hispanic*"
## [32] "anti-gaymale"
## [33] "anti-gayfemale"
## [34] "anti-gay(maleandfemale)"
## [35] "anti-heterosexual"
## [36] "anti-bisexual"
## [37] "anti-physicaldisability"
## [38] "anti-mentaldisability"
aggregategroups <- aggregategroups |>
mutate(group = case_when(
victim_cat %in% c("anti-transgender", "anti-gayfemale", "anti-genderidentityexpression",
"anti-gaymale", "anti-gay(maleandfemale)", "anti-bisexual") ~ "anti-lgbtq",
victim_cat %in% c("anti-multi-racialgroups", "anti-jewish", "anti-protestant",
"anti-multi-religiousgroups", "anti-religiouspracticegenerally",
"anti-buddhist", "anti-hindu", "anti-mormon", "anti-sikh",
"anti-catholic", "anti-islamic(muslim)", "anti-atheism/agnosticism",
"anti-otherreligion", "anti-easternorthodox(greek,russian,etc.)",
"anti-jehovahswitness", "anti-otherchristian") ~ "anti-religion",
victim_cat %in% c("anti-asian", "anti-arab", "anti-non-hispanic*", "anti-white",
"anti-americanindian/alaskannative",
"anti-nativehawaiian/pacificislander", "anti-otherrace",
"anti-hispanic", "anti-otherethnicity/nationalorigin") ~ "anti-ethnicity",
victim_cat %in% c("anti-physicaldisability", "anti-mentaldisability") ~ "anti-disability",
victim_cat %in% c("anti-female", "anti-male") ~ "anti-gender",
TRUE ~ "others"))
head(aggregategroups)
## # A tibble: 6 × 9
## county year crimetype totalincidents totalvictims totaloffenders victim_cat
## <chr> <dbl> <chr> <dbl> <dbl> <dbl> <chr>
## 1 Albany 2016 Crimes Aga… 3 4 3 anti-male
## 2 Albany 2016 Crimes Aga… 3 4 3 anti-fema…
## 3 Albany 2016 Crimes Aga… 3 4 3 anti-tran…
## 4 Albany 2016 Crimes Aga… 3 4 3 anti-gend…
## 5 Albany 2016 Crimes Aga… 3 4 3 anti-age*
## 6 Albany 2016 Crimes Aga… 3 4 3 anti-white
## # ℹ 2 more variables: crimecount <dbl>, group <chr>
aggregategroups |>
group_by(group) |>
summarize(total = sum(crimecount)) |>
arrange(desc(total))
## # A tibble: 6 × 2
## group total
## <chr> <dbl>
## 1 anti-religion 2132
## 2 anti-lgbtq 825
## 3 others 768
## 4 anti-ethnicity 526
## 5 anti-gender 10
## 6 anti-disability 9
lgbtq <- hatecrimes |>
pivot_longer(
cols = 4:41,
names_to = "victim_cat",
values_to = "crimecount") |>
filter(victim_cat %in% c("anti-transgender", "anti-gayfemale",
"anti-genderidentityexpression", "anti-gaymale",
"anti-gay(maleandfemale)", "anti-bisexual"))
head(lgbtq)
## # A tibble: 6 × 8
## county year crimetype totalincidents totalvictims totaloffenders victim_cat
## <chr> <dbl> <chr> <dbl> <dbl> <dbl> <chr>
## 1 Albany 2016 Crimes Aga… 3 4 3 anti-tran…
## 2 Albany 2016 Crimes Aga… 3 4 3 anti-gend…
## 3 Albany 2016 Crimes Aga… 3 4 3 anti-gaym…
## 4 Albany 2016 Crimes Aga… 3 4 3 anti-gayf…
## 5 Albany 2016 Crimes Aga… 3 4 3 anti-gay(…
## 6 Albany 2016 Crimes Aga… 3 4 3 anti-bise…
## # ℹ 1 more variable: crimecount <dbl>
One strength. The best move here is the join with census data and the shift from raw counts to a rate per 100,000. Before it, every plot really just ranks counties by population: Kings, Queens, and New York lead because they are the largest, not because they are the most hostile. Dividing by population changes the story. Suffolk passes Kings, Queens falls from second in population to seventh in rate, and the Bronx lands far below what its size would predict. It is the same lesson as the Stop and Frisk table in Unit 1, where raw counts across unequal groups misled and rates did not.
One limitation. The data measures reported and
classified hate crimes, not hate crimes. The FBI cannot compel
local agencies to submit, and agencies differ in training, willingness
to apply the label, and community trust. That makes the Bronx finding
ambiguous: it may have fewer hate crimes, or it may report fewer, and
nothing here separates those. The join has a narrower flaw too. Because
the population table was trimmed to ten counties before joining,
everyone else gets NA and drops out of the rate ranking,
including a county listed literally as “Multiple” with 14 incidents in
2012.
Two research questions. First, is variation in reported rates driven more by reporting infrastructure than by actual incidence? Joining FBI agency-participation records and county police resources would test whether reporting capacity predicts recorded rates better than demographics do. Second, was the 2012 anti-Jewish spike in Kings County a countywide shift or a single cluster? Monthly, precinct-level data would show whether those 136 incidents reflect a broad change in climate or a short burst tied to specific neighborhoods, which matters for what kind of response fits.