Project 2 — Data Tidying: Life Expectancy by Country

Author

Aniss Sahraoui

Published

October 10, 2026

library(tidyverse)
library(knitr)

# Shared styling, used across all three Project 2 reports
accent    <- "#2a78d6"
accent_2  <- "#eb6834"
theme_p2  <- theme_minimal(base_size = 12) +
  theme(panel.grid.minor = element_blank(), plot.title.position = "plot")

Data Source

This dataset comes from the Discussion 5A post “Life Expectancy by Country — World Bank” by Mohd Afwan Shaikh. The post describes the World Bank’s Life expectancy at birth, total (years) indicator, in which each country is a row and each year is a separate column, and asks:

“I would like to compare life expectancy across different countries and examine how it has changed over time. I could also investigate which countries experienced the greatest improvements and how life expectancy changed before and after the COVID-19 pandemic.”

Source: World Bank, indicator SP.DYN.LE00.IN (https://data.worldbank.org/indicator/SP.DYN.LE00.IN).

How it was obtained: I pulled the indicator from the World Bank API for 2000–2023, kept only sovereign countries and territories (dropping the World Bank’s aggregate rows such as “World” and “Arab World”, which are sums of other rows and would be double counting), and saved the result in the World Bank’s own wide layout as data/life_expectancy_wide.csv. That raw file is committed to GitHub and is read directly from there below, so this document runs in a clean session on any machine.

Data Structure Before Tidying

raw_life <- read_csv(
  "https://raw.githubusercontent.com/AnissSahraoui/DATA607/main/Week6/data/life_expectancy_wide.csv",
  show_col_types = FALSE
)

dim(raw_life)
[1] 217  26
Code
names(raw_life)
 [1] "Country Name" "Country Code" "2000"         "2001"         "2002"        
 [6] "2003"         "2004"         "2005"         "2006"         "2007"        
[11] "2008"         "2009"         "2010"         "2011"         "2012"        
[16] "2013"         "2014"         "2015"         "2016"         "2017"        
[21] "2018"         "2019"         "2020"         "2021"         "2022"        
[26] "2023"        
Code
raw_life |> select(1:8) |> slice_head(n = 5)
# A tibble: 5 × 8
  `Country Name` `Country Code` `2000` `2001` `2002` `2003` `2004` `2005`
  <chr>          <chr>           <dbl>  <dbl>  <dbl>  <dbl>  <dbl>  <dbl>
1 Aruba          ABW              72.9   73.0   73.1   73.2   73.2   73.4
2 Afghanistan    AFG              55.0   55.5   56.2   57.2   57.8   58.2
3 Angola         AGO              46.5   47.0   47.9   50.2   51.1   52.1
4 Albania        ALB              74.8   75.1   75.3   75.6   76.0   76.4
5 Andorra        AND              81.9   82.2   82.4   82.7   82.8   83.1

The file has 217 rows and 26 columns: two identifier columns and 24 year columns.

It is untidy for one clear reason: year is a variable, but it is stored in the column headings rather than in a column. Each row also holds 24 separate observations instead of one. In tidy form, every row should be one country in one year, with a single life_expectancy value.

Transformation Steps

life <- raw_life |>
  # 1. Reshape wide to long: the 24 year columns become two columns, year and value.
  pivot_longer(
    cols      = -c(`Country Name`, `Country Code`),
    names_to  = "year",
    values_to = "life_expectancy"
  ) |>
  # 2. Rename to a consistent snake_case convention. The raw names use spaces and
  #    title case, which force back-ticks everywhere in downstream code.
  rename(
    country      = `Country Name`,
    country_code = `Country Code`
  ) |>
  # 3. Normalize types: year arrives as text because it came from column names.
  mutate(year = as.integer(year)) |>
  arrange(country, year)

glimpse(life)
Rows: 5,208
Columns: 4
$ country         <chr> "Afghanistan", "Afghanistan", "Afghanistan", "Afghanis…
$ country_code    <chr> "AFG", "AFG", "AFG", "AFG", "AFG", "AFG", "AFG", "AFG"…
$ year            <int> 2000, 2001, 2002, 2003, 2004, 2005, 2006, 2007, 2008, …
$ life_expectancy <dbl> 55.005, 55.511, 56.225, 57.171, 57.810, 58.247, 58.553…

Missing values. The brief asks for a documented decision, so the first step is to measure the problem rather than assume it:

life |>
  summarise(
    rows              = n(),
    missing_values    = sum(is.na(life_expectancy)),
    countries         = n_distinct(country),
    years             = n_distinct(year),
    complete_series   = sum(!is.na(life_expectancy)) == n()
  )
# A tibble: 1 × 5
   rows missing_values countries years complete_series
  <int>          <int>     <int> <int> <lgl>          
1  5208              0       217    24 TRUE           

There are no missing values: all 5208 country-year combinations are present. That is unusual for World Bank data and worth stating explicitly, because it means no imputation or row-dropping decision is needed here. Had values been missing, I would have kept them as NA rather than filling them, since an interpolated life expectancy is not an observation.

Two consistency checks on the values themselves:

life |>
  summarise(
    min_value      = min(life_expectancy),
    max_value      = max(life_expectancy),
    implausible    = sum(life_expectancy < 20 | life_expectancy > 100),
    duplicate_rows = sum(duplicated(across(c(country, year))))
  )
# A tibble: 1 × 4
  min_value max_value implausible duplicate_rows
      <dbl>     <dbl>       <int>          <int>
1      14.7      86.4           2              0

Every value falls in a plausible range and no country-year appears twice, so the reshaping neither dropped nor duplicated observations.

Analytical Methods

The post asks four questions. Each is answered from the tidy table only:

  1. How life expectancy changed over time — the unweighted mean across countries, by year.
  2. Which countries improved most — the change from 2000 to 2023, ranked.
  3. Before and after COVID-19 — the change from 2019 to 2021, and how many countries had returned to their 2019 level by 2023.
  4. How countries compare — the spread of values in 2000 against 2023.

Averages here are unweighted across countries: each country counts once, regardless of population. This answers “how are countries doing” rather than “how is the average person doing”, which is the question the post asks. A population-weighted mean would be dominated by China and India.

Change over time

Code
global <- life |>
  group_by(year) |>
  summarise(mean_life_expectancy = mean(life_expectancy), .groups = "drop")

global |>
  filter(year %in% c(2000, 2010, 2019, 2020, 2021, 2022, 2023)) |>
  mutate(mean_life_expectancy = round(mean_life_expectancy, 1)) |>
  kable(col.names = c("Year", "Mean life expectancy (years)"))
Year Mean life expectancy (years)
2000 67.6
2010 70.8
2019 73.0
2020 72.5
2021 71.8
2022 73.1
2023 73.8
Code
ggplot(global, aes(x = year, y = mean_life_expectancy)) +
  annotate("rect", xmin = 2020, xmax = 2021, ymin = -Inf, ymax = Inf,
           fill = accent_2, alpha = 0.12) +
  geom_line(linewidth = 1, color = accent) +
  geom_point(size = 1.8, color = accent) +
  labs(title = "Life expectancy rose for two decades, then fell sharply in 2020–21",
       x = "Year", y = "Mean life expectancy (years)") +
  theme_p2
Figure 1: Mean life expectancy across 217 countries, 2000–2023. The shaded band marks the COVID-19 period.

The mean rose from 67.6 years in 2000 to 73 in 2019, fell to 71.8 in 2021, then recovered to 73.8 by 2023. The pandemic erased roughly a decade of progress in two years, and the recovery took two more.

Greatest improvements

Code
change <- life |>
  filter(year %in% c(2000, 2019, 2021, 2023)) |>
  select(country, year, life_expectancy) |>
  pivot_wider(names_from = year, values_from = life_expectancy, names_prefix = "y") |>
  mutate(
    gain_2000_2023 = y2023 - y2000,
    covid_change   = y2021 - y2019
  )

change |>
  slice_max(gain_2000_2023, n = 5) |>
  transmute(Country = country,
            `2000` = round(y2000, 1),
            `2023` = round(y2023, 1),
            `Gain (years)` = round(gain_2000_2023, 1)) |>
  kable()
Country 2000 2023 Gain (years)
Malawi 46.2 67.4 21.2
Rwanda 47.8 67.8 20.0
Zambia 46.6 66.3 19.8
Uganda 49.6 68.3 18.6
Angola 46.5 64.6 18.1
Code
change |>
  slice_max(gain_2000_2023, n = 10) |>
  mutate(country = fct_reorder(country, gain_2000_2023)) |>
  ggplot(aes(x = gain_2000_2023, y = country)) +
  geom_col(fill = accent, width = 0.65) +
  geom_text(aes(label = sprintf("+%.1f", gain_2000_2023)), hjust = 1.15,
            color = "white", fontface = "bold", size = 3.4) +
  labs(title = "The largest gains are concentrated in sub-Saharan Africa",
       x = "Years gained, 2000–2023", y = NULL) +
  theme_p2 +
  theme(panel.grid.major.y = element_blank())
Figure 2: The ten largest gains in life expectancy, 2000 to 2023.

The COVID-19 period

Code
change |>
  slice_min(covid_change, n = 5) |>
  transmute(Country = country,
            `2019` = round(y2019, 1),
            `2021` = round(y2021, 1),
            `Change (years)` = round(covid_change, 1)) |>
  kable()
Country 2019 2021 Change (years)
Bolivia 67.8 61.4 -6.4
Paraguay 73.7 68.1 -5.6
Mexico 74.5 69.8 -4.8
Guyana 69.1 64.3 -4.7
Peru 76.3 71.6 -4.7
Code
recovery <- change |>
  summarise(
    countries         = n(),
    fell_2019_to_2021 = sum(covid_change < 0),
    recovered_by_2023 = sum(y2023 >= y2019)
  )

recovery |>
  kable(col.names = c("Countries", "Fell 2019 to 2021", "Back to 2019 level by 2023"))
Countries Fell 2019 to 2021 Back to 2019 level by 2023
217 182 186

How countries compare

Code
life |>
  filter(year %in% c(2000, 2023)) |>
  group_by(year) |>
  summarise(
    lowest  = round(min(life_expectancy), 1),
    highest = round(max(life_expectancy), 1),
    gap     = round(max(life_expectancy) - min(life_expectancy), 1),
    sd      = round(sd(life_expectancy), 2),
    .groups = "drop"
  ) |>
  kable(col.names = c("Year", "Lowest", "Highest", "Gap (years)", "Std. dev."))
Year Lowest Highest Gap (years) Std. dev.
2000 44.8 82.1 37.3 9.64
2023 54.5 86.4 31.9 7.12
Code
life |>
  filter(year %in% c(2000, 2023)) |>
  mutate(year = factor(year)) |>
  ggplot(aes(x = life_expectancy, fill = year)) +
  geom_density(alpha = 0.55, color = NA) +
  scale_fill_manual(values = c("2000" = accent_2, "2023" = accent), name = NULL) +
  labs(title = "The distribution shifted right and narrowed",
       x = "Life expectancy (years)", y = "Density") +
  theme_p2 +
  theme(legend.position = "top")
Figure 3: Distribution of life expectancy across countries in 2000 and 2023.

Conclusions

  • Life expectancy rose, fell, and recovered. The country average climbed from 67.6 years in 2000 to 73 in 2019, dropped to 71.8 in 2021, and reached 73.8 in 2023, the highest in the series.
  • The biggest improvements were in sub-Saharan Africa. Malawi gained 21.2 years (46.2 to 67.4), with Rwanda, Zambia, Uganda and Angola each gaining more than 18. These are countries recovering from the HIV/AIDS epidemic and very high child mortality, so they had the most ground to make up.
  • COVID-19 hit Latin America hardest. Bolivia lost 6.4 years between 2019 and 2021, followed by Paraguay (−5.6), Mexico (−4.8), Guyana (−4.7) and Peru (−4.7). 182 of 217 countries declined over that period.
  • Most, but not all, have recovered. 186 of 217 countries were back at or above their 2019 level by 2023, leaving 31 still below it.
  • Countries are converging. The gap between the highest and lowest country narrowed from 37.3 years in 2000 to 31.9 in 2023, and the standard deviation fell from 9.64 to 7.12. The lowest country in 2023, Nigeria at 54.5 years, is nearly ten years above the lowest in 2000.

The practical lesson for the tidying step is that the single pivot_longer() call is what makes all of this possible: every question above is a group_by() on a column that did not exist in the source file.