Toplines

This analysis makes use of the migrant death data reported on the Humane Borders website (https://www.humaneborders.org/). These are freely available data and the R code generating this analysis is shown in the file (I do not claim to be a pristine coder, but the code works). The following are the major toplines:

  1. The migrant death crisis emerges in the early 2000s and has been persistent since then, with obvious yearly variation.

  2. The vast majority of remains recovered are males, although the number of female remains has increased in recent years.

  3. The typical age range of migrant deaths is 20 to 34 years of age. Over 1/3rd of remains recovered upon which age can be determined fall within this range.

  4. There are substantially large numbers of remains for which gender, age, and cause of death cannot be determined.

  5. About 34% of deaths are due to exposure. About 35% of remains recovered are skeletal.

  6. About 23% of remains recovered are of migrants who likely died more than 6-8 months before they were recovered (which could imply the remains may be many years old). I call these “noncontemporaneous deaths.” Most remains recovered over the time frame are of migrants estimated to have died within 6-8 months of recovery. I call these “contemporaneous deaths.” The number of noncontemporaneous deaths has increased overtime, implying that in more recent years, many remains recovered are of deaths that occurred at least 6-8 months earlier.

  7. However, in every year, except for 2025, the number of contemporaneous deaths exceeds noncontemporaneous deaths, implying the death crisis is ongoing.

Understanding the migrant death issue

The data on migrant deaths comes from the Arizona OpenGIS Project, cosponsored by Humane Borders and the Pima County Office of Medical Examiner. I’ve edited the data set to include all recorded migrant deaths between 2000 up through July of 2026. My goal with any analysis of these data is to always honor and humanize those who have died in the desert. These data permit a better understanding of the death crisis.

A note about the data

The data on migrant deaths comes from the Arizona OpenGIS Project, cosponsored by Humane Borders and the Pima County Office of Medical Examiner. Recorded remains are between the years 1981 to 2026. However systematic data on migrant deaths only are found from 2000 and later. This is primarily because the death crisis really doesn’t emerge until around the year 2000.

The reason the crisis was nonexistent before 2000 is that very, very few people crossed through Arizona. This changed due to changes made in the Clinton Administration that led to the funnel effect pushing migrants into the Arizona corridor. So if one downloads the death map data, one will find that 4,505 migrant remains have been recovered (as of 7/31/2026); however, 4,371 have been recovered since the year 2000. This means that in the death map, only 134 remains are recorded as being found between 1981 and 1999 (an 18 year period). I don’t mean “only” in a demeaning way; rather, relative to the mass death crisis, the total number of deaths reported in the 18 year period between 1981 to 1999 is far lower than the average yearly number of deaths after the year 2000.

The data file is a csv file saved to my GitHub site. I dynamically update the data as new information is added. This file is current through 7/31/26.

md="https://raw.githubusercontent.com/mightyjoemoon/POL51/main/ogis_migrant_deaths-14.csv"

md<-read_csv(url(md))
## Rows: 4505 Columns: 21
## ── Column specification ────────────────────────────────────────────────────────
## Delimiter: ","
## chr  (13): ML Number, Name, Sex, Surface Management, Location, Location Prec...
## dbl   (7): Age, Corridor Code, Condition Code, Latitude, Longitude, UTM X, U...
## date  (1): Reporting Date
## 
## ℹ Use `spec()` to retrieve the full column specification for this data.
## ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.

Migrant deaths by year and decade

Below, I create a table of the total number of recovered remains by calendar year. The percentage in the table corresponds to the “contribution” in percentage terms each year makes to the total. In other words, consider the year 2002. In that year, the table shows that 151 remains were recovered.

Of the total number of remains reported between 1981 to July 31, 2026, 2002 contributed 3.35% of the total number of deaths (i.e. total reported deaths is 4,505 and the total deaths in 2002 was 151; the percentage contribution is thus \(\frac{151}{4505}\times 100 \approx 3.35\%\)).

Observe that years before 2000, there were very few recovered remains; after 2000, the death crisis emerges. Note that the years 2020 and 2021—years enmeshed in the COVID pandemic—mark two of the three highest recovery years in the times series. I will have more to say on this below.

tabyl(md$yeardecade) %>%  adorn_pct_formatting(rounding = "half up", digits = 2) %>%
   adorn_totals() %>%
  knitr::kable()
## Warning: 'xfun::attr()' is deprecated.
## Use 'xfun::attr2()' instead.
## See help("Deprecated")

## Warning: 'xfun::attr()' is deprecated.
## Use 'xfun::attr2()' instead.
## See help("Deprecated")
md$yeardecade n percent
1981 1 0.02%
1982 1 0.02%
1985 3 0.07%
1987 1 0.02%
1990 9 0.20%
1991 6 0.13%
1992 7 0.16%
1993 17 0.38%
1994 4 0.09%
1995 12 0.27%
1996 13 0.29%
1997 22 0.49%
1998 15 0.33%
1999 23 0.51%
2000 75 1.66%
2001 79 1.75%
2002 151 3.35%
2003 164 3.64%
2004 186 4.13%
2005 203 4.51%
2006 174 3.86%
2007 221 4.91%
2008 166 3.68%
2009 197 4.37%
2010 224 4.97%
2011 182 4.04%
2012 163 3.62%
2013 184 4.08%
2014 140 3.11%
2015 147 3.26%
2016 164 3.64%
2017 127 2.82%
2018 129 2.86%
2019 146 3.24%
2020 212 4.71%
2021 229 5.08%
2022 174 3.86%
2023 200 4.44%
2024 158 3.51%
2025 103 2.29%
2026 73 1.62%
Total 4505 -

Visualizing migrant deaths by year

Below, I produce a bar plot of migrant deaths by year. Each bar corresponds to the number of migrant remains recovered in each year. As is clear from the plot, the death crisis emerges in 2000 and is persistent over time.

yearplot<-ggplot(md, aes(yeardecade)) +
  geom_bar(fill = "lightskyblue4") +
    labs(title="Migrant remains recovered by year along the Arizona/Mexico border, \n1981- July, 2026", 
subtitle="Data from the Arizon OpenGIS Project (https://humaneborders.info/app/map.asp)",
          y="Number recovered", 
          x="Year") +
  theme_classic()

yearplot<- yearplot + theme(axis.text.x = element_text(angle = 90, vjust = 0.5, hjust=1))

yearplot

Understanding gender and migrant deaths

The overwhelming number of migrant deaths are male. The analysis below shows this clearly; about 81% of the total remains recovered are male. Since gender is not always determined, there is a category called “undetermined.”

md$gender<-factor(md$Sex,
                  levels=c("male", "female", "undetermined"),
                  labels=c("male", "female", "undetermined"))
tabyl(md$gender)
##     md$gender    n      percent valid_percent
##          male 3630 0.8057713651    0.80630831
##        female  643 0.1427302997    0.14282541
##  undetermined  229 0.0508324084    0.05086628
##          <NA>    3 0.0006659267            NA

There are no discernible trends in gender differences. There has been some speculation that if asylum seeker deaths begin to rise, then more females will die given a large share of asylum seekers are women. Here is a link to a fairly recent article on this: https://19thnews.org/2024/07/women-migrants-deaths-us-mexico-border/

Below is a frequency table of remains recovered by year and gender.

gendertable<-tabyl(md, yeardecade, gender)

gendertable<-gendertable %>% adorn_percentages(denominator = "row")%>%
  adorn_totals(c()) %>%
  adorn_percentages("row") %>% 
  adorn_pct_formatting(rounding = "half up", digits = 0) %>%
  adorn_ns() %>%
  adorn_title("combined") %>%
  knitr::kable()
## Warning: 'xfun::attr()' is deprecated.
## Use 'xfun::attr2()' instead.
## See help("Deprecated")
gendertable
## Warning: 'xfun::attr()' is deprecated.
## Use 'xfun::attr2()' instead.
## See help("Deprecated")
yeardecade/gender male female undetermined NA_
1981 0% (0) 0% (0) 100% (1) 0% (0)
1982 0% (0) 100% (1) 0% (0) 0% (0)
1985 100% (3) 0% (0) 0% (0) 0% (0)
1987 100% (1) 0% (0) 0% (0) 0% (0)
1990 56% (5) 0% (0) 44% (4) 0% (0)
1991 67% (4) 0% (0) 33% (2) 0% (0)
1992 71% (5) 0% (0) 29% (2) 0% (0)
1993 59% (10) 0% (0) 41% (7) 0% (0)
1994 50% (2) 0% (0) 50% (2) 0% (0)
1995 42% (5) 0% (0) 58% (7) 0% (0)
1996 23% (3) 0% (0) 77% (10) 0% (0)
1997 18% (4) 5% (1) 77% (17) 0% (0)
1998 20% (3) 0% (0) 80% (12) 0% (0)
1999 22% (5) 4% (1) 74% (17) 0% (0)
2000 76% (57) 24% (18) 0% (0) 0% (0)
2001 75% (59) 24% (19) 1% (1) 0% (0)
2002 76% (115) 24% (36) 0% (0) 0% (0)
2003 80% (131) 20% (33) 0% (0) 0% (0)
2004 81% (150) 19% (35) 1% (1) 0% (0)
2005 80% (163) 20% (40) 0% (0) 0% (0)
2006 82% (143) 18% (31) 0% (0) 0% (0)
2007 77% (171) 23% (50) 0% (0) 0% (0)
2008 80% (133) 20% (33) 0% (0) 0% (0)
2009 85% (167) 14% (27) 2% (3) 0% (0)
2010 87% (195) 13% (29) 0% (0) 0% (0)
2011 88% (160) 11% (20) 1% (2) 0% (0)
2012 87% (141) 12% (20) 1% (2) 0% (0)
2013 90% (166) 10% (18) 0% (0) 0% (0)
2014 91% (127) 8% (11) 1% (2) 0% (0)
2015 88% (130) 8% (12) 3% (5) 0% (0)
2016 93% (152) 6% (10) 1% (2) 0% (0)
2017 98% (124) 2% (3) 0% (0) 0% (0)
2018 90% (116) 9% (12) 1% (1) 0% (0)
2019 88% (129) 8% (12) 3% (5) 0% (0)
2020 82% (174) 9% (19) 9% (19) 0% (0)
2021 76% (174) 14% (31) 9% (21) 1% (3)
2022 78% (135) 13% (23) 9% (16) 0% (0)
2023 74% (147) 19% (38) 8% (15) 0% (0)
2024 68% (108) 20% (32) 11% (18) 0% (0)
2025 70% (72) 13% (13) 17% (18) 0% (0)
2026 56% (41) 21% (15) 23% (17) 0% (0)

Visualizing deaths by gender

I created a summary data set of remains recoverd for the years 2000 to 2026 (noting that 2026 is incomplete). This summary data set makes it easy to visualize migrant deaths by year and gender.

mdg="https://raw.githubusercontent.com/mightyjoemoon/POL51/main/mdg.csv"

mdg<-read_csv(url(mdg))
## Rows: 27 Columns: 4
## ── Column specification ────────────────────────────────────────────────────────
## Delimiter: ","
## dbl (4): year, male, female, undetermined
## 
## ℹ Use `spec()` to retrieve the full column specification for this data.
## ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
ggplot(mdg, aes(x = year)) +
  geom_line(aes(y = male, color="Male"), size=1) +
  geom_line(aes(y = female, color="Female"), size=1) +
  geom_line(aes(y = undetermined, color="Undetermined"), size=1) +
  scale_color_manual(values = c("coral2", "lightskyblue3", "gray")) +
  scale_x_continuous(n.breaks = 20) +
labs(title="Percentage of migrant deaths accounted for by gender, \n2000-July, 2026",
     y="Percent of total", x="Year",
     color="") +
  theme_bw()
## Warning: Using `size` aesthetic for lines was deprecated in ggplot2 3.4.0.
## ℹ Please use `linewidth` instead.
## This warning is displayed once every 8 hours.
## Call `lifecycle::last_lifecycle_warnings()` to see where this warning was
## generated.

Below are two barplots. The first shows the total number of remains recovered by gender; the second is a stacked barplot of deaths by year and gender.

ggplot(md, aes(x=gender)) +
  geom_bar(fill = "lightskyblue4") +
           labs(title="Migrant remains recovered by year on the Arizona/Mexico border, \n2000-July, 2026", 
                subtitle="Data from the Arizona OpenGIS Project",
          y="Number recovered", 
          x="Year") +
  theme_classic()

In the stacked barplot, the total height of the bar corresponds to the total number of remains recovered by year. The color codes represent the number associated with males, females, and those whose gender is undetermined.

stack_gender<-ggplot(md, mapping = aes(x = yeardecade)) +
         geom_bar(aes(fill = gender)) + 
         labs(title="Migrant remains by gender recovered on the Arizona/Mexico border, \n1981-July, 2026", subtitle="Data from the Arizon OpenGIS Project",
          y="Number recovered", 
          x="Year") + 
  theme_classic()  
  
stack_gender <- stack_gender + theme(axis.text.x = element_text(angle = 90, vjust = 0.5, hjust=1))

stack_gender 

Age of deceased migrants

The OpenGIS data records the age of the migrant if this determination is possible. To understand the relationship between age and migrant deaths, I’ve created a new variable recording agegroups. These groups are in 5-year increments with the exception of the first group (0-9 years) and the last group (60 and over). We see that about 38% of the migrant deaths have an indeterminant age. We can see that most migrant deaths (about 32% of the total) are in the age range of 20 to 34 years of age. This percentage is based on including all of the individuals whose age cannot be determined. This group constitutes 40% of the data.

md$agegroup <- case_when(md$Age <=9 ~ '0-9',
                             md$Age > 9  & md$Age <= 14 ~ '10-14',
                             md$Age > 14  & md$Age <= 19 ~ '15-19',
                         md$Age > 19  & md$Age <= 24 ~ '20-24',
                         md$Age > 24  & md$Age <= 29 ~ '25-29',
                         md$Age > 29  & md$Age <= 34 ~ '30-34',
                         md$Age > 34  & md$Age <= 39 ~ '35-39',
                         md$Age > 39  & md$Age <= 44 ~ '40-44',
                         md$Age > 44  & md$Age <= 49 ~ '45-49',
                         md$Age > 49  & md$Age <= 54 ~ '49-54',
                         md$Age > 54  & md$Age <= 59 ~ '55-59',
                         md$Age > 59 ~ '60 and over',
                             FALSE ~ 'NA' )

tabyl(md$agegroup) %>%  adorn_pct_formatting(rounding = "half up", digits = 0) %>%
  knitr::kable()
## Warning: 'xfun::attr()' is deprecated.
## Use 'xfun::attr2()' instead.
## See help("Deprecated")

## Warning: 'xfun::attr()' is deprecated.
## Use 'xfun::attr2()' instead.
## See help("Deprecated")
md$agegroup n percent valid_percent
0-9 7 0% 0%
10-14 22 0% 1%
15-19 255 6% 9%
20-24 498 11% 18%
25-29 490 11% 18%
30-34 470 10% 17%
35-39 391 9% 14%
40-44 281 6% 10%
45-49 168 4% 6%
49-54 85 2% 3%
55-59 40 1% 1%
60 and over 24 1% 1%
NA 1774 39% -

Below, I tabulate migrant deaths by year and age group. To understand the resulting table, consider the year 2006 and the age group 20-24 years of age. This group accounted for about 13% of that year’s migrant deaths. For most years, those whose age is undetermined (i.e. the “NAs”) accounts for a large share of migrant deaths. Note that the number of “non-determined” seems to be increasing with time.

tabyl(md, yeardecade, agegroup) %>%  adorn_percentages(denominator = "row")%>%
  adorn_totals(c("row", "col")) %>%
  adorn_percentages("row") %>% 
  adorn_pct_formatting(rounding = "half up", digits = 0) %>%
  adorn_ns() %>%
  adorn_title("combined") %>%
  knitr::kable()
## Warning: 'xfun::attr()' is deprecated.
## Use 'xfun::attr2()' instead.
## See help("Deprecated")

## Warning: 'xfun::attr()' is deprecated.
## Use 'xfun::attr2()' instead.
## See help("Deprecated")
yeardecade/agegroup 0-9 10-14 15-19 20-24 25-29 30-34 35-39 40-44 45-49 49-54 55-59 60 and over NA_ Total
1981 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 100% (1) 100% (1)
1982 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 100% (1) 100% (1)
1985 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 100% (3) 100% (3)
1987 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 100% (1) 100% (1)
1990 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 100% (9) 100% (9)
1991 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 100% (6) 100% (6)
1992 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 100% (7) 100% (7)
1993 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 100% (17) 100% (17)
1994 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 100% (4) 100% (4)
1995 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 100% (12) 100% (12)
1996 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 100% (13) 100% (13)
1997 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 100% (22) 100% (22)
1998 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 100% (15) 100% (15)
1999 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 100% (23) 100% (23)
2000 0% (0) 4% (3) 11% (8) 17% (13) 19% (14) 12% (9) 4% (3) 8% (6) 5% (4) 1% (1) 0% (0) 0% (0) 19% (14) 100% (75)
2001 0% (0) 0% (0) 6% (5) 19% (15) 13% (10) 13% (10) 6% (5) 9% (7) 6% (5) 3% (2) 0% (0) 0% (0) 25% (20) 100% (79)
2002 0% (0) 3% (4) 11% (17) 15% (22) 13% (20) 11% (17) 11% (17) 7% (10) 5% (8) 3% (4) 1% (2) 0% (0) 20% (30) 100% (151)
2003 1% (2) 0% (0) 10% (16) 15% (25) 14% (23) 10% (17) 9% (15) 8% (13) 2% (3) 1% (2) 2% (3) 1% (2) 26% (43) 100% (164)
2004 0% (0) 1% (1) 7% (13) 13% (24) 17% (31) 9% (16) 10% (18) 10% (19) 3% (5) 2% (3) 2% (3) 1% (1) 28% (52) 100% (186)
2005 0% (1) 1% (3) 10% (20) 15% (30) 11% (23) 12% (25) 10% (21) 5% (10) 5% (11) 3% (7) 0% (1) 0% (0) 25% (51) 100% (203)
2006 1% (1) 3% (5) 8% (14) 13% (23) 10% (17) 11% (19) 7% (13) 7% (12) 2% (4) 2% (3) 1% (2) 1% (2) 34% (59) 100% (174)
2007 0% (0) 1% (2) 8% (18) 9% (20) 13% (28) 14% (30) 13% (28) 8% (17) 2% (5) 3% (6) 1% (2) 1% (3) 28% (62) 100% (221)
2008 1% (1) 1% (1) 10% (17) 13% (22) 11% (18) 13% (21) 6% (10) 8% (13) 4% (7) 4% (7) 1% (1) 1% (1) 28% (47) 100% (166)
2009 0% (0) 1% (1) 8% (16) 10% (20) 14% (27) 15% (29) 11% (21) 3% (6) 2% (4) 2% (4) 2% (3) 1% (1) 33% (65) 100% (197)
2010 0% (0) 0% (0) 6% (13) 12% (27) 16% (36) 11% (25) 13% (28) 8% (18) 5% (12) 1% (2) 0% (0) 0% (0) 28% (63) 100% (224)
2011 0% (0) 0% (0) 5% (9) 9% (16) 12% (21) 12% (21) 11% (20) 7% (12) 3% (5) 1% (2) 1% (1) 1% (1) 41% (74) 100% (182)
2012 0% (0) 1% (1) 4% (7) 6% (9) 13% (22) 12% (20) 9% (15) 8% (13) 4% (7) 2% (3) 1% (1) 0% (0) 40% (65) 100% (163)
2013 0% (0) 0% (0) 5% (9) 9% (16) 11% (20) 10% (18) 12% (22) 5% (10) 3% (5) 2% (3) 1% (2) 1% (1) 42% (78) 100% (184)
2014 0% (0) 0% (0) 4% (6) 6% (9) 8% (11) 6% (9) 5% (7) 7% (10) 4% (6) 3% (4) 1% (2) 0% (0) 54% (76) 100% (140)
2015 0% (0) 0% (0) 6% (9) 12% (18) 8% (12) 12% (17) 12% (17) 5% (7) 2% (3) 3% (4) 1% (2) 1% (1) 39% (57) 100% (147)
2016 0% (0) 0% (0) 2% (3) 13% (21) 9% (14) 12% (19) 10% (16) 3% (5) 4% (6) 1% (1) 2% (3) 1% (2) 45% (74) 100% (164)
2017 0% (0) 0% (0) 2% (3) 9% (11) 9% (12) 11% (14) 5% (6) 3% (4) 5% (6) 1% (1) 1% (1) 0% (0) 54% (69) 100% (127)
2018 0% (0) 0% (0) 4% (5) 5% (7) 9% (12) 9% (11) 9% (12) 9% (11) 7% (9) 2% (2) 2% (3) 1% (1) 43% (56) 100% (129)
2019 1% (1) 0% (0) 3% (4) 12% (18) 9% (13) 10% (15) 10% (14) 3% (5) 3% (5) 3% (5) 0% (0) 1% (1) 45% (65) 100% (146)
2020 0% (0) 0% (0) 5% (11) 13% (27) 8% (17) 9% (20) 8% (16) 6% (12) 3% (7) 0% (1) 0% (1) 0% (1) 47% (99) 100% (212)
2021 0% (0) 0% (0) 4% (10) 12% (27) 10% (23) 9% (21) 8% (19) 5% (12) 4% (10) 1% (2) 1% (2) 0% (0) 45% (103) 100% (229)
2022 0% (0) 0% (0) 4% (7) 12% (21) 11% (19) 13% (22) 5% (8) 10% (17) 3% (6) 1% (1) 1% (1) 1% (1) 41% (71) 100% (174)
2023 1% (1) 1% (1) 3% (6) 15% (29) 11% (21) 12% (24) 11% (21) 5% (9) 7% (13) 4% (8) 1% (1) 1% (2) 32% (64) 100% (200)
2024 0% (0) 0% (0) 3% (5) 11% (18) 9% (15) 9% (15) 7% (11) 9% (15) 5% (8) 2% (3) 1% (1) 2% (3) 41% (64) 100% (158)
2025 0% (0) 0% (0) 2% (2) 5% (5) 4% (4) 4% (4) 5% (5) 6% (6) 3% (3) 3% (3) 1% (1) 0% (0) 68% (70) 100% (103)
2026 0% (0) 0% (0) 3% (2) 7% (5) 10% (7) 3% (2) 4% (3) 3% (2) 1% (1) 1% (1) 1% (1) 0% (0) 67% (49) 100% (73)
Total 0% (7) 0% (22) 4% (255) 7% (498) 7% (490) 7% (470) 6% (391) 4% (281) 3% (168) 1% (85) 1% (40) 0% (24) 59% (1,774) 100% (4,505)

Below I produce a barplot of migrant deaths by age group (including those whose age cannot be determined).

ggplot(md, aes(agegroup)) +
  geom_bar(fill = "lightskyblue4") +
    labs(title="Migrant remains by age group, 1981-July, 2026", 
subtitle="Data from the Arizon OpenGIS Project (https://humaneborders.info/app/map.asp)",
          y="Number recovered", 
          x="Age group") +
  theme_classic()

Below, I reproduce the age plot but omit the cases where age is indeterminate. This thus gives the number of remains recovered accounted for each group among those whose age is determined.

md_ssage<-md[which(md$agegroup!="NA"),]

ggplot(md_ssage, aes(agegroup)) +
  geom_bar(fill = "lightskyblue4") +
    labs(title="Migrant remains by age group (excluding NAs), \n2000-July, 2026", 
subtitle="Data from the Arizon OpenGIS Project (https://humaneborders.info/app/map.asp)",
          y="Number recovered", 
          x="Age group") +
  theme_classic() +
 theme(axis.text.x = element_text(angle = 45, vjust = 0.5, hjust=.5))

Suppose we are interested in see if there are any trends in average ages of deceased migrants? One way to do this is through a simple regression model. This model will give us an estimate of the average age conditional on year. The key thing to look at is the dots plotted for each year. This dot is our estimate of the average age of death for migrants whose remains are recovered. So consider the year 2003. This dot is right on the line corresponding to 30 years of age. This implies that among those whose remains were recovered in 2003, the average age at death was about 30 years of age. The vertical lines above and below the dot correspond to the error bars around the estimate. In general, the plot reveals that there has been no appreciable change over time in the average age at death. The year 2018 seems to be an outlier, but all other years give estimates in the range of late 20s to early 30s.

md$YEAR<-as.factor(md$yeardecade)
  
agemod<-lm(Age~YEAR,data=md)

summary(agemod)
## 
## Call:
## lm(formula = Age ~ YEAR, data = md)
## 
## Residuals:
##     Min      1Q  Median      3Q     Max 
## -32.735  -7.824  -1.132   6.726  66.289 
## 
## Coefficients:
##             Estimate Std. Error t value             Pr(>|t|)    
## (Intercept)   28.230      1.307  21.591 < 0.0000000000000002 ***
## YEAR2001       2.686      1.865   1.440             0.149888    
## YEAR2002       1.903      1.604   1.187             0.235498    
## YEAR2003       1.770      1.604   1.104             0.269641    
## YEAR2004       2.957      1.577   1.875             0.060926 .  
## YEAR2005       1.909      1.548   1.233             0.217620    
## YEAR2006       1.797      1.617   1.111             0.266790    
## YEAR2007       3.815      1.538   2.480             0.013190 *  
## YEAR2008       2.594      1.608   1.613             0.106829    
## YEAR2009       2.324      1.581   1.470             0.141768    
## YEAR2010       3.174      1.535   2.067             0.038786 *  
## YEAR2011       3.891      1.636   2.379             0.017433 *  
## YEAR2012       4.434      1.665   2.662             0.007808 ** 
## YEAR2013       3.959      1.641   2.412             0.015911 *  
## YEAR2014       5.739      1.827   3.141             0.001703 ** 
## YEAR2015       3.459      1.694   2.043             0.041183 *  
## YEAR2016       4.482      1.694   2.646             0.008186 ** 
## YEAR2017       3.098      1.873   1.654             0.098196 .  
## YEAR2018       7.044      1.771   3.977            0.0000717 ***
## YEAR2019       3.845      1.731   2.221             0.026447 *  
## YEAR2020       2.549      1.622   1.571             0.116245    
## YEAR2021       3.112      1.593   1.954             0.050852 .  
## YEAR2022       3.411      1.650   2.068             0.038768 *  
## YEAR2023       4.506      1.574   2.863             0.004224 ** 
## YEAR2024       5.579      1.679   3.323             0.000903 ***
## YEAR2025       6.922      2.207   3.137             0.001726 ** 
## YEAR2026       2.729      2.461   1.109             0.267524    
## ---
## Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
## 
## Residual standard error: 10.21 on 2704 degrees of freedom
##   (1774 observations deleted due to missingness)
## Multiple R-squared:  0.01732,    Adjusted R-squared:  0.00787 
## F-statistic: 1.833 on 26 and 2704 DF,  p-value: 0.006251
regage<-plot_model(agemod, type = "pred", terms = c("YEAR"), ci.lvl = .95, title="Estimated average age of deceased migrants, \n2000-July, 2026", axis.title=c("Year", "Age"), colors=c("skyblue4"), 
  legend.title="Time Period") + 
  geom_line() + 
  theme_bw() 

regage <- regage + theme(axis.text.x = element_text(angle = 90, vjust = 0.5, hjust=1))

regage

Understanding cause of death

The OpenGIS data records the likely cause of death. Several codes are given for migrant deaths. Two of these codes suggest cause of death is not determined. In one case, it is simply impossible to determine how the migrant dies and in the second case, only skeletal remains are found (making it extraordinarily difficult to discern cause of death). Below, I create a table of the causes of death reported in the OpenGIS data.

tabyl(md$`Cause of Death`)%>%  adorn_pct_formatting(rounding = "half up", digits = 0) %>%
  knitr::kable()
## Warning: 'xfun::attr()' is deprecated.
## Use 'xfun::attr2()' instead.
## See help("Deprecated")

## Warning: 'xfun::attr()' is deprecated.
## Use 'xfun::attr2()' instead.
## See help("Deprecated")
md$Cause of Death n percent valid_percent
Asphyxia 8 0% 0%
Blunt Force Injury 222 5% 5%
Diabetes 5 0% 0%
Drowning 45 1% 1%
Drug Overdose 7 0% 0%
Exposure 1547 34% 35%
Exsanguination 1 0% 0%
Gunshot Wound 98 2% 2%
Heart Disease 26 1% 1%
Lightning Strike 1 0% 0%
Motor Vehicle Accident 25 1% 1%
Nonviable Fetus 2 0% 0%
Other Disease 24 1% 1%
Other Injury 13 0% 0%
Other Injury / Homicide 26 1% 1%
Other injury 2 0% 0%
Pending 29 1% 1%
Pregnancy Complication 1 0% 0%
Skeletal Remains 1576 35% 36%
Undetermined 720 16% 16%
NA 127 3% -

Because “exposure,” “skeletal remains,” and “undetermined” are the three dominant codes given, let’s consider how these kinds of cases vary across the years. Observe that the number of “skeletal remains” seem to be increasing with time. I suspect this is due to the number of search and recovery groups heading out into the desert in the past few years. I think we are likely finding the remains of people who may have died long before the calendar year in which their remains were recovered. There is more evidence of this below.

md$cod <- case_when(md$`Cause of Death` == "Exposure" ~ 'Exposure',
                             md$`Cause of Death`=="Skeletal Remains" ~ 'Skeletal remains',
                             md$`Cause of Death`=="Undetermined" ~ 'Undetermined',
                             TRUE ~ 'Other' )


my_tabyl<-tabyl(md, yeardecade, cod) %>%  adorn_percentages(denominator = "row")%>%
  #adorn_totals(c("row", "col")) %>%
  adorn_percentages("row") %>% 
  adorn_pct_formatting(rounding = "half up", digits = 0) %>%
  #adorn_ns() %>%
  #adorn_title("combined") %>%
  knitr::kable()
## Warning: 'xfun::attr()' is deprecated.
## Use 'xfun::attr2()' instead.
## See help("Deprecated")
#clean_df <- as.data.frame(my_tabyl)

# 3. Export directly to a clean CSV file
#write.csv(clean_df, "cod2000_2026.csv", row.names = FALSE)

I created a summary data set using the information in the table we just considered and plot the number of individuals whose cause-of-death is exposure and all other causes combined (i.e. the category “other”), along with those whose cause-of-death is not determined or is classified as skeletal remains. Two points are very clear. First, relative to all known causes of death, exposure deaths are dominant. Relative to the category “other,” there is a consistent and often large gap. Second, the number of skeletal remains has increased over time. Again, I suspect this is due to greater vigilance in trying to find remains of those who have been lost.

mdcod="https://raw.githubusercontent.com/mightyjoemoon/POL51/main/cod2000_2026.csv"

mdcod<-read_csv(url(mdcod))
## Rows: 27 Columns: 5
## ── Column specification ────────────────────────────────────────────────────────
## Delimiter: ","
## dbl (5): year20, Exposure, Other, Skeletal remains, Undetermined
## 
## ℹ Use `spec()` to retrieve the full column specification for this data.
## ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
summary(mdcod)
##      year20        Exposure          Other        Skeletal remains
##  Min.   :2000   Min.   :0.1000   Min.   :0.0200   Min.   :0.0200  
##  1st Qu.:2006   1st Qu.:0.2150   1st Qu.:0.0600   1st Qu.:0.1200  
##  Median :2013   Median :0.3300   Median :0.1000   Median :0.4400  
##  Mean   :2013   Mean   :0.3496   Mean   :0.1219   Mean   :0.3663  
##  3rd Qu.:2020   3rd Qu.:0.4300   3rd Qu.:0.1600   3rd Qu.:0.5650  
##  Max.   :2026   Max.   :0.6800   Max.   :0.3200   Max.   :0.7500  
##   Undetermined   
##  Min.   :0.0000  
##  1st Qu.:0.1100  
##  Median :0.1400  
##  Mean   :0.1619  
##  3rd Qu.:0.2100  
##  Max.   :0.4100
mdcod$sr<-mdcod$`Skeletal remains`*100
mdcod$Undetermined<-mdcod$Undetermined*100
mdcod$Exposure<-mdcod$Exposure*100
mdcod$Other<-mdcod$Other*100


ggplot(mdcod, aes(x = year20)) +
  geom_line(aes(y = Exposure, color="Exposure"), size=.7) +
  geom_line(aes(y = Other, color="Other"), size=.7) +
  geom_line(aes(y = sr, color="Skeletal remains"), size=.7) +
  geom_line(aes(y = Undetermined, color="Undetermined"), size=.7) +

  scale_color_manual(values = c("coral2", "gray","lightskyblue3",  "black")) + 
  scale_x_continuous(n.breaks = 20) +
labs(title="Percentage of migrant deaths by cause of death, \n2000-July, 2026",
     y="Percent of total", x="Year",
     color="") +
  theme_bw()

Making sense of the PMI

In the OpenGIS data, the Medical Examiner records the likely “post-mortem interval.” This is essentially the time frame in which a migrant likely died. Thus, the code “< 1 day” would imply the migrant died within the past 24 hours, while the code “> 6-8 months” means that the migrant likely died, potentially, much further back in time. Below is a table of the data. One way I like to think of the death data is in terms of what I call “contemporaneous deaths.” By this I mean deaths that likely occurred within a 12-month window. To explain, suppose a migrant remains are found in February? Suppose PCOME assesses the PMI as death occuring less than 6-8 months later. This would seem to imply that the migrant didn’t die in the year his/her death is recorded, but did die within a 12-month window. For those whose PMI is assessed as being greater than 6-8 months, it’s impossible to determine what the time window is in which they died. It could be one year or 15 years.

tabyl(md$`Post Mortem Interval`) %>% adorn_pct_formatting(rounding = "half up", digits = 0) %>%
  knitr::kable()
## Warning: 'xfun::attr()' is deprecated.
## Use 'xfun::attr2()' instead.
## See help("Deprecated")

## Warning: 'xfun::attr()' is deprecated.
## Use 'xfun::attr2()' instead.
## See help("Deprecated")
md$Post Mortem Interval n percent valid_percent
< 1 day 1207 27% 29%
< 1 week 479 11% 12%
< 3 months 359 8% 9%
< 3 weeks 277 6% 7%
< 5 weeks 383 9% 9%
< 6-8 months 398 9% 10%
> 6-8 months 1051 23% 25%
NA 351 8% -

One thing that may be of interest might be to consider how the PMI codes vary with respect to time. Below I create a table to inspect this.

md$pmi<-md$`Post Mortem Interval`

tabyl(md, yeardecade, pmi)
##  yeardecade < 1 day < 1 week < 3 months < 3 weeks < 5 weeks < 6-8 months
##        1981       0        0          0         0         0            0
##        1982       0        0          0         0         0            0
##        1985       0        0          0         0         0            0
##        1987       0        0          0         0         0            0
##        1990       0        0          0         0         0            0
##        1991       0        0          0         0         0            0
##        1992       0        0          0         0         0            0
##        1993       0        0          0         0         0            0
##        1994       0        0          0         0         0            0
##        1995       0        0          0         0         0            0
##        1996       0        0          0         0         0            0
##        1997       0        0          0         0         0            0
##        1998       0        0          0         0         0            0
##        1999       0        0          0         0         0            0
##        2000      45       18          2         1         2            3
##        2001      35       15          2         0        14            4
##        2002      60       29          4        15        19            1
##        2003      56       24          7        13        16            8
##        2004      95       15          5         8        10           10
##        2005      83       23          6        31        26            9
##        2006      71       21          5        10         7           22
##        2007      86       25         12        14        21           17
##        2008      60       32          3        12         7           18
##        2009      83       16         13        12        19            6
##        2010      76       32         25        12        18           27
##        2011      39       15         24        11        14           20
##        2012      28       17         18        15        17           27
##        2013      37       11         25        16        16           30
##        2014      14        6         21        12        11           33
##        2015      19       11         12        15        22           24
##        2016      25       13         28        11        18           15
##        2017      11        5         12         3        13           26
##        2018      18        7         14        10        14           22
##        2019      11       15         24        16        14            7
##        2020      39       21         20         5        20           12
##        2021      49       32         16        13        15           17
##        2022      36       21         20         8        14            9
##        2023      62       23         20         7        18           19
##        2024      41       20          9         4        12            7
##        2025      12        2         11         2         3            3
##        2026      16       10          1         1         3            2
##  > 6-8 months NA_
##             0   1
##             0   1
##             0   3
##             0   1
##             0   9
##             0   6
##             0   7
##             0  17
##             0   4
##             0  12
##             0  13
##             0  22
##             0  15
##             0  23
##             3   1
##             7   2
##            19   4
##            31   9
##            26  17
##            18   7
##            32   6
##            39   7
##            30   4
##            41   7
##            32   2
##            55   4
##            31  10
##            30  19
##            29  14
##            31  13
##            37  17
##            48   9
##            35   9
##            49  10
##            87   8
##            73  14
##            63   3
##            48   3
##            55  10
##            65   5
##            37   3
tabyl(md, yeardecade, pmi) %>%  adorn_percentages(denominator = "row")%>%
  adorn_totals(c("row", "col")) %>%
  adorn_percentages("row") %>% 
  adorn_pct_formatting(rounding = "half up", digits = 0) %>%
  adorn_ns() %>%
  adorn_title("combined") %>%
  knitr::kable()
## Warning: 'xfun::attr()' is deprecated.
## Use 'xfun::attr2()' instead.
## See help("Deprecated")

## Warning: 'xfun::attr()' is deprecated.
## Use 'xfun::attr2()' instead.
## See help("Deprecated")
yeardecade/pmi < 1 day < 1 week < 3 months < 3 weeks < 5 weeks < 6-8 months > 6-8 months NA_ Total
1981 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 100% (1) 100% (1)
1982 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 100% (1) 100% (1)
1985 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 100% (3) 100% (3)
1987 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 100% (1) 100% (1)
1990 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 100% (9) 100% (9)
1991 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 100% (6) 100% (6)
1992 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 100% (7) 100% (7)
1993 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 100% (17) 100% (17)
1994 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 100% (4) 100% (4)
1995 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 100% (12) 100% (12)
1996 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 100% (13) 100% (13)
1997 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 100% (22) 100% (22)
1998 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 100% (15) 100% (15)
1999 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 0% (0) 100% (23) 100% (23)
2000 60% (45) 24% (18) 3% (2) 1% (1) 3% (2) 4% (3) 4% (3) 1% (1) 100% (75)
2001 44% (35) 19% (15) 3% (2) 0% (0) 18% (14) 5% (4) 9% (7) 3% (2) 100% (79)
2002 40% (60) 19% (29) 3% (4) 10% (15) 13% (19) 1% (1) 13% (19) 3% (4) 100% (151)
2003 34% (56) 15% (24) 4% (7) 8% (13) 10% (16) 5% (8) 19% (31) 5% (9) 100% (164)
2004 51% (95) 8% (15) 3% (5) 4% (8) 5% (10) 5% (10) 14% (26) 9% (17) 100% (186)
2005 41% (83) 11% (23) 3% (6) 15% (31) 13% (26) 4% (9) 9% (18) 3% (7) 100% (203)
2006 41% (71) 12% (21) 3% (5) 6% (10) 4% (7) 13% (22) 18% (32) 3% (6) 100% (174)
2007 39% (86) 11% (25) 5% (12) 6% (14) 10% (21) 8% (17) 18% (39) 3% (7) 100% (221)
2008 36% (60) 19% (32) 2% (3) 7% (12) 4% (7) 11% (18) 18% (30) 2% (4) 100% (166)
2009 42% (83) 8% (16) 7% (13) 6% (12) 10% (19) 3% (6) 21% (41) 4% (7) 100% (197)
2010 34% (76) 14% (32) 11% (25) 5% (12) 8% (18) 12% (27) 14% (32) 1% (2) 100% (224)
2011 21% (39) 8% (15) 13% (24) 6% (11) 8% (14) 11% (20) 30% (55) 2% (4) 100% (182)
2012 17% (28) 10% (17) 11% (18) 9% (15) 10% (17) 17% (27) 19% (31) 6% (10) 100% (163)
2013 20% (37) 6% (11) 14% (25) 9% (16) 9% (16) 16% (30) 16% (30) 10% (19) 100% (184)
2014 10% (14) 4% (6) 15% (21) 9% (12) 8% (11) 24% (33) 21% (29) 10% (14) 100% (140)
2015 13% (19) 7% (11) 8% (12) 10% (15) 15% (22) 16% (24) 21% (31) 9% (13) 100% (147)
2016 15% (25) 8% (13) 17% (28) 7% (11) 11% (18) 9% (15) 23% (37) 10% (17) 100% (164)
2017 9% (11) 4% (5) 9% (12) 2% (3) 10% (13) 20% (26) 38% (48) 7% (9) 100% (127)
2018 14% (18) 5% (7) 11% (14) 8% (10) 11% (14) 17% (22) 27% (35) 7% (9) 100% (129)
2019 8% (11) 10% (15) 16% (24) 11% (16) 10% (14) 5% (7) 34% (49) 7% (10) 100% (146)
2020 18% (39) 10% (21) 9% (20) 2% (5) 9% (20) 6% (12) 41% (87) 4% (8) 100% (212)
2021 21% (49) 14% (32) 7% (16) 6% (13) 7% (15) 7% (17) 32% (73) 6% (14) 100% (229)
2022 21% (36) 12% (21) 11% (20) 5% (8) 8% (14) 5% (9) 36% (63) 2% (3) 100% (174)
2023 31% (62) 12% (23) 10% (20) 4% (7) 9% (18) 10% (19) 24% (48) 2% (3) 100% (200)
2024 26% (41) 13% (20) 6% (9) 3% (4) 8% (12) 4% (7) 35% (55) 6% (10) 100% (158)
2025 12% (12) 2% (2) 11% (11) 2% (2) 3% (3) 3% (3) 63% (65) 5% (5) 100% (103)
2026 22% (16) 14% (10) 1% (1) 1% (1) 4% (3) 3% (2) 51% (37) 4% (3) 100% (73)
Total 18% (1,207) 7% (479) 5% (359) 4% (277) 6% (383) 6% (398) 16% (1,051) 37% (351) 100% (4,505)
#Simple table of frequencies 
#tabyl(md, yeardecade, pmi) %>%  
 # adorn_title("combined") %>%
#  knitr::kable()

To simplify our understanding of the data, consider the table that’s created below. Here I compare cases where the PMI is 5 weeks or less to cases that are: 3 months or less; 6-8 months or less; and cases greater than 6-8 months.

md$pmi2 <- case_when(md$`Post Mortem Interval` == "< 1 day" ~ 'Within 5 weeks',
                             md$`Post Mortem Interval`=="< 1 week" ~ 'Within 5 weeks',
                             md$`Post Mortem Interval`=="< 3 weeks" ~ 'Within 5 weeks',
                             md$`Post Mortem Interval`=="< 5 weeks" ~ 'Within 5 weeks',
                             md$`Post Mortem Interval`=="< 3 months" ~ 'Within 3 months',
                             md$`Post Mortem Interval`=="< 6-8 months" ~ 'Within 6-8 months',
                        md$`Post Mortem Interval`=="> 6-8 months" ~ 'Greater than 6-8 months',
                             TRUE ~ 'Not reported' )


tabyl(md, yeardecade, pmi2) %>%  adorn_percentages(denominator = "row")%>%
  adorn_totals(c("row", "col")) %>%
  adorn_percentages("row") %>% 
  adorn_pct_formatting(rounding = "half up", digits = 0) %>%
  adorn_ns() %>%
  adorn_title("combined") %>%
  knitr::kable()
## Warning: 'xfun::attr()' is deprecated.
## Use 'xfun::attr2()' instead.
## See help("Deprecated")

## Warning: 'xfun::attr()' is deprecated.
## Use 'xfun::attr2()' instead.
## See help("Deprecated")
yeardecade/pmi2 Greater than 6-8 months Not reported Within 3 months Within 5 weeks Within 6-8 months Total
1981 0% (0) 100% (1) 0% (0) 0% (0) 0% (0) 100% (1)
1982 0% (0) 100% (1) 0% (0) 0% (0) 0% (0) 100% (1)
1985 0% (0) 100% (3) 0% (0) 0% (0) 0% (0) 100% (3)
1987 0% (0) 100% (1) 0% (0) 0% (0) 0% (0) 100% (1)
1990 0% (0) 100% (9) 0% (0) 0% (0) 0% (0) 100% (9)
1991 0% (0) 100% (6) 0% (0) 0% (0) 0% (0) 100% (6)
1992 0% (0) 100% (7) 0% (0) 0% (0) 0% (0) 100% (7)
1993 0% (0) 100% (17) 0% (0) 0% (0) 0% (0) 100% (17)
1994 0% (0) 100% (4) 0% (0) 0% (0) 0% (0) 100% (4)
1995 0% (0) 100% (12) 0% (0) 0% (0) 0% (0) 100% (12)
1996 0% (0) 100% (13) 0% (0) 0% (0) 0% (0) 100% (13)
1997 0% (0) 100% (22) 0% (0) 0% (0) 0% (0) 100% (22)
1998 0% (0) 100% (15) 0% (0) 0% (0) 0% (0) 100% (15)
1999 0% (0) 100% (23) 0% (0) 0% (0) 0% (0) 100% (23)
2000 4% (3) 1% (1) 3% (2) 88% (66) 4% (3) 100% (75)
2001 9% (7) 3% (2) 3% (2) 81% (64) 5% (4) 100% (79)
2002 13% (19) 3% (4) 3% (4) 81% (123) 1% (1) 100% (151)
2003 19% (31) 5% (9) 4% (7) 66% (109) 5% (8) 100% (164)
2004 14% (26) 9% (17) 3% (5) 69% (128) 5% (10) 100% (186)
2005 9% (18) 3% (7) 3% (6) 80% (163) 4% (9) 100% (203)
2006 18% (32) 3% (6) 3% (5) 63% (109) 13% (22) 100% (174)
2007 18% (39) 3% (7) 5% (12) 66% (146) 8% (17) 100% (221)
2008 18% (30) 2% (4) 2% (3) 67% (111) 11% (18) 100% (166)
2009 21% (41) 4% (7) 7% (13) 66% (130) 3% (6) 100% (197)
2010 14% (32) 1% (2) 11% (25) 62% (138) 12% (27) 100% (224)
2011 30% (55) 2% (4) 13% (24) 43% (79) 11% (20) 100% (182)
2012 19% (31) 6% (10) 11% (18) 47% (77) 17% (27) 100% (163)
2013 16% (30) 10% (19) 14% (25) 43% (80) 16% (30) 100% (184)
2014 21% (29) 10% (14) 15% (21) 31% (43) 24% (33) 100% (140)
2015 21% (31) 9% (13) 8% (12) 46% (67) 16% (24) 100% (147)
2016 23% (37) 10% (17) 17% (28) 41% (67) 9% (15) 100% (164)
2017 38% (48) 7% (9) 9% (12) 25% (32) 20% (26) 100% (127)
2018 27% (35) 7% (9) 11% (14) 38% (49) 17% (22) 100% (129)
2019 34% (49) 7% (10) 16% (24) 38% (56) 5% (7) 100% (146)
2020 41% (87) 4% (8) 9% (20) 40% (85) 6% (12) 100% (212)
2021 32% (73) 6% (14) 7% (16) 48% (109) 7% (17) 100% (229)
2022 36% (63) 2% (3) 11% (20) 45% (79) 5% (9) 100% (174)
2023 24% (48) 2% (3) 10% (20) 55% (110) 10% (19) 100% (200)
2024 35% (55) 6% (10) 6% (9) 49% (77) 4% (7) 100% (158)
2025 63% (65) 5% (5) 11% (11) 18% (19) 3% (3) 100% (103)
2026 51% (37) 4% (3) 1% (1) 41% (30) 3% (2) 100% (73)
Total 16% (1,051) 37% (351) 5% (359) 35% (2,346) 6% (398) 100% (4,505)

Below I explore the idea of contemporaneous deaths. I create a variable that records migrant remains recovered within 6 to 8 months (this would include all migrant recoveries on any time-scale up to and within 6-8 months). I take this as a measure of contemporaneous deaths, or alternatively, deaths that likely occurred within a 12-month time window.

To understand the composition of contemporaneous deaths with older remains recovered, I simply take a ratio of contemporaneous deaths with older deaths. To understand the logic, consider the hypothetical example: suppose in year \(x\), there were 75 contemporaneous deaths and 75 older deaths. This would give a ratio of \(\frac{75}{75}=1\), implying that for every 1 recent death recorded, there is 1 older death recorded. Thus, as this ratio exceeds 1, there are more recent deaths recovered than older remains recovered. Below, I plot this ratio by calendar year. For this exercise, I do not include PMI codes listed as undetermined.

The plot shows some interesting characteristics of the death crisis. In the early years, recoveries were overwhelmingly contemporaneous. However, there are several possible reasons for this. First, migrant crossings through Arizona were likely relatively low in the early period implying most deaths would be recent deaths. Second, as the crisis evolves, more and more cross but more and more die without a trace. As time moves forward, and individuals and groups are more vigilant in looking for remains, we begin to find more and more older remains. In my view, the fact that older remains are now being found at a higher rate is alarming for it tells me that the scope and magnitude of the death crisis was far bigger than the reported data in the early years suggested.

mdpmi="https://raw.githubusercontent.com/mightyjoemoon/POL51/main/pmi2000_2026.csv"

mdpmi<-read_csv(url(mdpmi))
## Rows: 27 Columns: 10
## ── Column specification ────────────────────────────────────────────────────────
## Delimiter: ","
## dbl (10): Year, lt1day, lt1week, lt3months, lt3weeks, lt5weeks, lt6_8months,...
## 
## ℹ Use `spec()` to retrieve the full column specification for this data.
## ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
summary(mdpmi)
##       Year          lt1day        lt1week        lt3months        lt3weeks    
##  Min.   :2000   Min.   :11.0   Min.   : 2.00   Min.   : 1.00   Min.   : 0.00  
##  1st Qu.:2006   1st Qu.:22.0   1st Qu.:12.00   1st Qu.: 5.50   1st Qu.: 6.00  
##  Median :2013   Median :39.0   Median :17.00   Median :12.00   Median :11.00  
##  Mean   :2013   Mean   :44.7   Mean   :17.74   Mean   :13.33   Mean   :10.26  
##  3rd Qu.:2020   3rd Qu.:61.0   3rd Qu.:23.00   3rd Qu.:20.50   3rd Qu.:13.50  
##  Max.   :2026   Max.   :95.0   Max.   :32.00   Max.   :28.00   Max.   :31.00  
##     lt5weeks      lt6_8months     gt6_8months     undetermined   
##  Min.   : 2.00   Min.   : 1.00   Min.   : 3.00   Min.   : 0.000  
##  1st Qu.:11.50   1st Qu.: 7.00   1st Qu.:30.00   1st Qu.: 4.000  
##  Median :14.00   Median :15.00   Median :35.00   Median : 7.000  
##  Mean   :14.15   Mean   :14.74   Mean   :39.26   Mean   : 7.556  
##  3rd Qu.:18.00   3rd Qu.:22.00   3rd Qu.:48.50   3rd Qu.:10.000  
##  Max.   :26.00   Max.   :33.00   Max.   :97.00   Max.   :19.000  
##      Total      
##  Min.   : 73.0  
##  1st Qu.:142.0  
##  Median :164.0  
##  Mean   :161.7  
##  3rd Qu.:191.5  
##  Max.   :225.0
mdpmi$in12window <- mdpmi$lt1day + mdpmi$lt1week  + mdpmi$lt3months + mdpmi$lt3weeks + 
                    mdpmi$lt5weeks + mdpmi$lt6_8months
mdpmi$outside12window <- mdpmi$gt6_8months
mdpmi$ratio_inout12<- mdpmi$in12window/mdpmi$outside12window
ggplot(mdpmi, aes(x = Year)) +
  geom_line(aes(y = ratio_inout12)) +
  scale_x_continuous(n.breaks = 20) +
  
labs(title="Ratio of recent remains recovered to older remains recovered, \n2000-July, 2026",
      subtitle="2026 only includes data up through July 31, 2026",
     y="Ratio", x="Year") +
  geom_hline(yintercept = 1) +
  theme_bw()

To get a sense of what recent remains recovered looks like compared to older remains recovered, I plot both time series. There is one crucial remark to make. Earlier in this report, I wrote: “Note that the years 2020 and 2021—years enmeshed in the COVID pandemic—mark two of the three highest recovery years in the times series. I will have more to say on this below.” This statement is a bit misleading (and news reports at this time would have been misleading) and underscores the need to think about the PMI information.

In 2020, while there were indeed many deaths that occurred within in the year (especially June), there were many older remains recovered. Since we cannot unambiguously determine the time frame of death for these older remains, adding them to the contemporaneous deaths may be misleading. Saying this, 2021 and 2023 would rank among the most lethal years in the time-series in terms of contemporaneous deaths.

ggplot(mdpmi, aes(x = Year)) +
  geom_line(aes(y = in12window, color="Likely within 12 months"), size=.7) +
  geom_line(aes(y = outside12window, color="Likely outside 12 months"), size=.7) +
  scale_color_manual(values = c("coral2", "lightskyblue3")) +
  scale_x_continuous(n.breaks = 20) +
labs(title="Number of migrant remains recovered for recent vs. older PMI, \n2000-July, 2026",
      subtitle="2026 only includes data up through July 31, 2026",
     y="Total", x="Year",
     color="") +
  theme_bw()