Project 2: World Bank Life Expectancy

Author

Ummay Rukiya

Published

October 8, 2026

Overview

For this part of Project 2, I will tidy and analyze the World Bank Life Expectancy at Birth dataset. The original dataset stores every year in a separate column, making it a wide-format dataset. I will transform it into a tidy format with one row for each country and year.

My main analysis will compare life expectancy in Bangladesh and the United States over time. I will also examine whether the gap between the two countries has increased or decreased.

Data Source

The dataset comes from the World Bank Life Expectancy at Birth indicator. It contains annual life-expectancy measurements for countries and regions beginning in 1960.

The original wide-format CSV file was uploaded to my public GitHub repository so the analysis can be reproduced directly from an internet source.

Loading Required Packages

Code
library(tidyverse)
library(knitr)

Importing the Raw Wide-Format Data

The original World Bank CSV contains several descriptive lines above the actual column names. I used skip = 4 so R begins reading from the row containing the dataset’s column names.

Code
raw_data_url <- paste0(
  "https://cdn.jsdelivr.net/gh/UR-71/",
  "DATA607-Project2@main/",
  "world_bank_life_expectancy_wide.csv"
)

life_expectancy_wide <- readr::read_csv(
  raw_data_url,
  skip = 4,
  show_col_types = FALSE
)

dim(life_expectancy_wide)
[1] 265  71
Code
head(life_expectancy_wide[, 1:10])
# A tibble: 6 × 10
  `Country Name`  `Country Code` `Indicator Name` `Indicator Code` `1960` `1961`
  <chr>           <chr>          <chr>            <chr>             <dbl>  <dbl>
1 Aruba           ABW            Life expectancy… SP.DYN.LE00.IN     64.0   64.2
2 Africa Eastern… AFE            Life expectancy… SP.DYN.LE00.IN     44.2   44.5
3 Afghanistan     AFG            Life expectancy… SP.DYN.LE00.IN     32.8   33.3
4 Africa Western… AFW            Life expectancy… SP.DYN.LE00.IN     37.8   38.1
5 Angola          AGO            Life expectancy… SP.DYN.LE00.IN     37.9   36.9
6 Albania         ALB            Life expectancy… SP.DYN.LE00.IN     56.4   57.5
# ℹ 4 more variables: `1962` <dbl>, `1963` <dbl>, `1964` <dbl>, `1965` <dbl>

Data Structure Before Tidying

The original data set contains 265 rows and 71 columns. Each row represents a country or region. The first four columns contain identifying information, while the remaining columns contain life-expectancy values for separate years. This structure is wide because the years are being used as column names.

Code
data.frame(
  rows = nrow(life_expectancy_wide),
  columns = ncol(life_expectancy_wide)
)
  rows columns
1  265      71
Code
names(life_expectancy_wide)
 [1] "Country Name"   "Country Code"   "Indicator Name" "Indicator Code"
 [5] "1960"           "1961"           "1962"           "1963"          
 [9] "1964"           "1965"           "1966"           "1967"          
[13] "1968"           "1969"           "1970"           "1971"          
[17] "1972"           "1973"           "1974"           "1975"          
[21] "1976"           "1977"           "1978"           "1979"          
[25] "1980"           "1981"           "1982"           "1983"          
[29] "1984"           "1985"           "1986"           "1987"          
[33] "1988"           "1989"           "1990"           "1991"          
[37] "1992"           "1993"           "1994"           "1995"          
[41] "1996"           "1997"           "1998"           "1999"          
[45] "2000"           "2001"           "2002"           "2003"          
[49] "2004"           "2005"           "2006"           "2007"          
[53] "2008"           "2009"           "2010"           "2011"          
[57] "2012"           "2013"           "2014"           "2015"          
[61] "2016"           "2017"           "2018"           "2019"          
[65] "2020"           "2021"           "2022"           "2023"          
[69] "2024"           "2025"           "...71"         
Code
kable(
  head(life_expectancy_wide[, 1:10]),
  caption = "Sample of the Original Wide-Format Dataset"
)
Sample of the Original Wide-Format Dataset
Country Name Country Code Indicator Name Indicator Code 1960 1961 1962 1963 1964 1965
Aruba ABW Life expectancy at birth, total (years) SP.DYN.LE00.IN 64.04900 64.21500 64.60200 64.94400 65.30300 65.61500
Africa Eastern and Southern AFE Life expectancy at birth, total (years) SP.DYN.LE00.IN 44.16966 44.46884 44.87789 45.16058 45.53570 45.77072
Afghanistan AFG Life expectancy at birth, total (years) SP.DYN.LE00.IN 32.79900 33.29100 33.75700 34.20100 34.67300 35.12400
Africa Western and Central AFW Life expectancy at birth, total (years) SP.DYN.LE00.IN 37.77964 38.05897 38.68179 38.93692 39.19438 39.47888
Angola AGO Life expectancy at birth, total (years) SP.DYN.LE00.IN 37.93300 36.90200 37.16800 37.41900 37.70400 37.96800
Albania ALB Life expectancy at birth, total (years) SP.DYN.LE00.IN 56.41300 57.48800 58.49400 59.47900 60.40400 61.27300

Transforming the Data from Wide to Long Format

I used pivot_longer() to move the year columns into rows. The tidy dataset contains one column for the country, one for the year, and one for the life-expectancy value. I kept missing measurements as NA during the transformation so that no information from the original dataset was incorrectly replaced.

Code
life_expectancy_tidy <- life_expectancy_wide |>
  select(-starts_with("...")) |>
  rename(
    country_name = `Country Name`,
    country_code = `Country Code`,
    indicator_name = `Indicator Name`,
    indicator_code = `Indicator Code`
  ) |>
  pivot_longer(
    cols = matches("^[0-9]{4}$"),
    names_to = "year",
    values_to = "life_expectancy"
  ) |>
  mutate(
    year = as.integer(year),
    life_expectancy = as.numeric(life_expectancy)
  ) |>
  arrange(country_name, year)

dim(life_expectancy_tidy)
[1] 17490     6
Code
head(life_expectancy_tidy)
# A tibble: 6 × 6
  country_name country_code indicator_name  indicator_code  year life_expectancy
  <chr>        <chr>        <chr>           <chr>          <int>           <dbl>
1 Afghanistan  AFG          Life expectanc… SP.DYN.LE00.IN  1960            32.8
2 Afghanistan  AFG          Life expectanc… SP.DYN.LE00.IN  1961            33.3
3 Afghanistan  AFG          Life expectanc… SP.DYN.LE00.IN  1962            33.8
4 Afghanistan  AFG          Life expectanc… SP.DYN.LE00.IN  1963            34.2
5 Afghanistan  AFG          Life expectanc… SP.DYN.LE00.IN  1964            34.7
6 Afghanistan  AFG          Life expectanc… SP.DYN.LE00.IN  1965            35.1

Handling Missing Values

Some country-and-year combinations do not have a reported life-expectancy value. I kept these values as NA while tidying the data set so the original missing information would remain visible. For the analysis, I created a separate dataframe that excludes rows with missing life-expectancy values. I did not replace them with zero because zero would not be a meaningful life-expectancy measurement.

Code
missing_value_summary <- life_expectancy_tidy |>
  summarize(
    total_rows = n(),
    missing_values = sum(is.na(life_expectancy)),
    percent_missing = round(
      mean(is.na(life_expectancy)) * 100,
      2
    )
  )

missing_value_summary
# A tibble: 1 × 3
  total_rows missing_values percent_missing
       <int>          <int>           <dbl>
1      17490            364            2.08
Code
life_expectancy_analysis <- life_expectancy_tidy |>
  filter(!is.na(life_expectancy))

dim(life_expectancy_analysis)
[1] 17126     6
Code
range(life_expectancy_analysis$year)
[1] 1960 2024

Comparing Bangladesh and the United States

I selected Bangladesh and the United States for the main comparison. The following summary shows each country’s earliest and latest available life-expectancy measurement and the total change between those years.

Code
country_comparison <- life_expectancy_analysis |>
  filter(country_code %in% c("BGD", "USA")) |>
  select(
    country_name,
    country_code,
    year,
    life_expectancy
  )

country_summary <- country_comparison |>
  group_by(country_name, country_code) |>
  arrange(year, .by_group = TRUE) |>
  summarize(
    starting_year = first(year),
    starting_life_expectancy = first(life_expectancy),
    ending_year = last(year),
    ending_life_expectancy = last(life_expectancy),
    total_increase = last(life_expectancy) -
      first(life_expectancy),
    .groups = "drop"
  ) |>
  mutate(
    across(
      c(
        starting_life_expectancy,
        ending_life_expectancy,
        total_increase
      ),
      ~ round(.x, 2)
    )
  )

kable(
  country_summary,
  caption = paste(
    "Life Expectancy Changes in Bangladesh",
    "and the United States"
  )
)
Life Expectancy Changes in Bangladesh and the United States
country_name country_code starting_year starting_life_expectancy ending_year ending_life_expectancy total_increase
Bangladesh BGD 1960 43.98 2024 74.93 30.95
United States USA 1960 69.77 2024 78.89 9.12

Life Expectancy Over Time

The following line chart shows how life expectancy changed in Bangladesh and the United States between 1960 and 2024.

Code
ggplot(
  country_comparison,
  aes(
    x = year,
    y = life_expectancy,
    color = country_name
  )
) +
  geom_line(linewidth = 1) +
  labs(
    title = paste(
      "Life Expectancy in Bangladesh",
      "and the United States"
    ),
    x = "Year",
    y = "Life Expectancy at Birth (Years)",
    color = "Country"
  ) +
  scale_color_manual(
    values = c(
      "Bangladesh" = "#2C7FB8",
      "United States" = "#D95F0E"
    )
  ) +
  theme_minimal()

Measuring the Life-Expectancy Gap

To measure the gap directly, I reshaped the two-country data so Bangladesh and the United States had separate columns. I then subtracted Bangladesh’s life expectancy from the United States’ life expectancy for each year.

Code
life_expectancy_gap <- country_comparison |>
  select(
    country_code,
    year,
    life_expectancy
  ) |>
  pivot_wider(
    names_from = country_code,
    values_from = life_expectancy
  ) |>
  drop_na(BGD, USA) |>
  mutate(
    life_expectancy_gap = USA - BGD
  ) |>
  arrange(year)

gap_summary <- life_expectancy_gap |>
  summarize(
    starting_year = first(year),
    starting_gap = first(life_expectancy_gap),
    ending_year = last(year),
    ending_gap = last(life_expectancy_gap),
    change_in_gap = last(life_expectancy_gap) -
      first(life_expectancy_gap)
  ) |>
  mutate(
    across(
      c(
        starting_gap,
        ending_gap,
        change_in_gap
      ),
      ~ round(.x, 2)
    )
  )

kable(
  gap_summary,
  caption = paste(
    "Life-Expectancy Gap Between the",
    "United States and Bangladesh"
  )
)
Life-Expectancy Gap Between the United States and Bangladesh
starting_year starting_gap ending_year ending_gap change_in_gap
1960 25.79 2024 3.96 -21.83

Change in the Gap Over Time

Code
ggplot(
  life_expectancy_gap,
  aes(
    x = year,
    y = life_expectancy_gap
  )
) +
  geom_line(
    color = "#6A3D9A",
    linewidth = 1
  ) +
  labs(
    title = paste(
      "Life-Expectancy Gap Between the",
      "United States and Bangladesh"
    ),
    subtitle = paste(
      "Positive values show how many years",
      "higher U.S. life expectancy was"
    ),
    x = "Year",
    y = "Difference in Life Expectancy (Years)"
  ) +
  theme_minimal()

Interpretation of the Life-Expectancy Gap

The results confirm that the life-expectancy gap between the two countries became much smaller. In 1960, life expectancy in the United States was 25.79 years higher than in Bangladesh. By 2024, the difference had decreased to only 3.96 years. This means that the gap narrowed by 21.83 years during the period.

The graph also shows a sharp temporary increase in the gap around 1971. According to the World Bank data, Bangladesh’s life expectancy fell to 26.52 years during that year. This period included the Bangladesh Liberation War, political conflict, natural disaster, and severe food insecurity. The value recovered after 1971, and the long-term gap continued to decrease.

Conclusion

The original World Bank data set was successfully transformed from a wide format with separate year columns into a tidy format containing one row for each country and year. The missing values were kept as NA during the transformation and removed only from the data used for analysis.

Both Bangladesh and the United States experienced improvements in life expectancy between 1960 and 2024. However, Bangladesh experienced a much larger increase of 30.95 years, compared with 9.12 years in the United States. As a result, the life-expectancy gap between the two countries decreased from 25.79 years to 3.96 years. These results show that Bangladesh moved substantially closer to the United States in this measurement over time, even though a smaller gap still remained in 2024.