Code
library(tidyverse)
library(knitr)For this part of Project 2, I will tidy and analyze the World Bank Life Expectancy at Birth dataset. The original dataset stores every year in a separate column, making it a wide-format dataset. I will transform it into a tidy format with one row for each country and year.
My main analysis will compare life expectancy in Bangladesh and the United States over time. I will also examine whether the gap between the two countries has increased or decreased.
The dataset comes from the World Bank Life Expectancy at Birth indicator. It contains annual life-expectancy measurements for countries and regions beginning in 1960.
The original wide-format CSV file was uploaded to my public GitHub repository so the analysis can be reproduced directly from an internet source.
library(tidyverse)
library(knitr)The original World Bank CSV contains several descriptive lines above the actual column names. I used skip = 4 so R begins reading from the row containing the dataset’s column names.
raw_data_url <- paste0(
"https://cdn.jsdelivr.net/gh/UR-71/",
"DATA607-Project2@main/",
"world_bank_life_expectancy_wide.csv"
)
life_expectancy_wide <- readr::read_csv(
raw_data_url,
skip = 4,
show_col_types = FALSE
)
dim(life_expectancy_wide)[1] 265 71
head(life_expectancy_wide[, 1:10])# A tibble: 6 × 10
`Country Name` `Country Code` `Indicator Name` `Indicator Code` `1960` `1961`
<chr> <chr> <chr> <chr> <dbl> <dbl>
1 Aruba ABW Life expectancy… SP.DYN.LE00.IN 64.0 64.2
2 Africa Eastern… AFE Life expectancy… SP.DYN.LE00.IN 44.2 44.5
3 Afghanistan AFG Life expectancy… SP.DYN.LE00.IN 32.8 33.3
4 Africa Western… AFW Life expectancy… SP.DYN.LE00.IN 37.8 38.1
5 Angola AGO Life expectancy… SP.DYN.LE00.IN 37.9 36.9
6 Albania ALB Life expectancy… SP.DYN.LE00.IN 56.4 57.5
# ℹ 4 more variables: `1962` <dbl>, `1963` <dbl>, `1964` <dbl>, `1965` <dbl>
The original data set contains 265 rows and 71 columns. Each row represents a country or region. The first four columns contain identifying information, while the remaining columns contain life-expectancy values for separate years. This structure is wide because the years are being used as column names.
data.frame(
rows = nrow(life_expectancy_wide),
columns = ncol(life_expectancy_wide)
) rows columns
1 265 71
names(life_expectancy_wide) [1] "Country Name" "Country Code" "Indicator Name" "Indicator Code"
[5] "1960" "1961" "1962" "1963"
[9] "1964" "1965" "1966" "1967"
[13] "1968" "1969" "1970" "1971"
[17] "1972" "1973" "1974" "1975"
[21] "1976" "1977" "1978" "1979"
[25] "1980" "1981" "1982" "1983"
[29] "1984" "1985" "1986" "1987"
[33] "1988" "1989" "1990" "1991"
[37] "1992" "1993" "1994" "1995"
[41] "1996" "1997" "1998" "1999"
[45] "2000" "2001" "2002" "2003"
[49] "2004" "2005" "2006" "2007"
[53] "2008" "2009" "2010" "2011"
[57] "2012" "2013" "2014" "2015"
[61] "2016" "2017" "2018" "2019"
[65] "2020" "2021" "2022" "2023"
[69] "2024" "2025" "...71"
kable(
head(life_expectancy_wide[, 1:10]),
caption = "Sample of the Original Wide-Format Dataset"
)| Country Name | Country Code | Indicator Name | Indicator Code | 1960 | 1961 | 1962 | 1963 | 1964 | 1965 |
|---|---|---|---|---|---|---|---|---|---|
| Aruba | ABW | Life expectancy at birth, total (years) | SP.DYN.LE00.IN | 64.04900 | 64.21500 | 64.60200 | 64.94400 | 65.30300 | 65.61500 |
| Africa Eastern and Southern | AFE | Life expectancy at birth, total (years) | SP.DYN.LE00.IN | 44.16966 | 44.46884 | 44.87789 | 45.16058 | 45.53570 | 45.77072 |
| Afghanistan | AFG | Life expectancy at birth, total (years) | SP.DYN.LE00.IN | 32.79900 | 33.29100 | 33.75700 | 34.20100 | 34.67300 | 35.12400 |
| Africa Western and Central | AFW | Life expectancy at birth, total (years) | SP.DYN.LE00.IN | 37.77964 | 38.05897 | 38.68179 | 38.93692 | 39.19438 | 39.47888 |
| Angola | AGO | Life expectancy at birth, total (years) | SP.DYN.LE00.IN | 37.93300 | 36.90200 | 37.16800 | 37.41900 | 37.70400 | 37.96800 |
| Albania | ALB | Life expectancy at birth, total (years) | SP.DYN.LE00.IN | 56.41300 | 57.48800 | 58.49400 | 59.47900 | 60.40400 | 61.27300 |
I used pivot_longer() to move the year columns into rows. The tidy dataset contains one column for the country, one for the year, and one for the life-expectancy value. I kept missing measurements as NA during the transformation so that no information from the original dataset was incorrectly replaced.
life_expectancy_tidy <- life_expectancy_wide |>
select(-starts_with("...")) |>
rename(
country_name = `Country Name`,
country_code = `Country Code`,
indicator_name = `Indicator Name`,
indicator_code = `Indicator Code`
) |>
pivot_longer(
cols = matches("^[0-9]{4}$"),
names_to = "year",
values_to = "life_expectancy"
) |>
mutate(
year = as.integer(year),
life_expectancy = as.numeric(life_expectancy)
) |>
arrange(country_name, year)
dim(life_expectancy_tidy)[1] 17490 6
head(life_expectancy_tidy)# A tibble: 6 × 6
country_name country_code indicator_name indicator_code year life_expectancy
<chr> <chr> <chr> <chr> <int> <dbl>
1 Afghanistan AFG Life expectanc… SP.DYN.LE00.IN 1960 32.8
2 Afghanistan AFG Life expectanc… SP.DYN.LE00.IN 1961 33.3
3 Afghanistan AFG Life expectanc… SP.DYN.LE00.IN 1962 33.8
4 Afghanistan AFG Life expectanc… SP.DYN.LE00.IN 1963 34.2
5 Afghanistan AFG Life expectanc… SP.DYN.LE00.IN 1964 34.7
6 Afghanistan AFG Life expectanc… SP.DYN.LE00.IN 1965 35.1
Some country-and-year combinations do not have a reported life-expectancy value. I kept these values as NA while tidying the data set so the original missing information would remain visible. For the analysis, I created a separate dataframe that excludes rows with missing life-expectancy values. I did not replace them with zero because zero would not be a meaningful life-expectancy measurement.
missing_value_summary <- life_expectancy_tidy |>
summarize(
total_rows = n(),
missing_values = sum(is.na(life_expectancy)),
percent_missing = round(
mean(is.na(life_expectancy)) * 100,
2
)
)
missing_value_summary# A tibble: 1 × 3
total_rows missing_values percent_missing
<int> <int> <dbl>
1 17490 364 2.08
life_expectancy_analysis <- life_expectancy_tidy |>
filter(!is.na(life_expectancy))
dim(life_expectancy_analysis)[1] 17126 6
range(life_expectancy_analysis$year)[1] 1960 2024
I selected Bangladesh and the United States for the main comparison. The following summary shows each country’s earliest and latest available life-expectancy measurement and the total change between those years.
country_comparison <- life_expectancy_analysis |>
filter(country_code %in% c("BGD", "USA")) |>
select(
country_name,
country_code,
year,
life_expectancy
)
country_summary <- country_comparison |>
group_by(country_name, country_code) |>
arrange(year, .by_group = TRUE) |>
summarize(
starting_year = first(year),
starting_life_expectancy = first(life_expectancy),
ending_year = last(year),
ending_life_expectancy = last(life_expectancy),
total_increase = last(life_expectancy) -
first(life_expectancy),
.groups = "drop"
) |>
mutate(
across(
c(
starting_life_expectancy,
ending_life_expectancy,
total_increase
),
~ round(.x, 2)
)
)
kable(
country_summary,
caption = paste(
"Life Expectancy Changes in Bangladesh",
"and the United States"
)
)| country_name | country_code | starting_year | starting_life_expectancy | ending_year | ending_life_expectancy | total_increase |
|---|---|---|---|---|---|---|
| Bangladesh | BGD | 1960 | 43.98 | 2024 | 74.93 | 30.95 |
| United States | USA | 1960 | 69.77 | 2024 | 78.89 | 9.12 |
The following line chart shows how life expectancy changed in Bangladesh and the United States between 1960 and 2024.
ggplot(
country_comparison,
aes(
x = year,
y = life_expectancy,
color = country_name
)
) +
geom_line(linewidth = 1) +
labs(
title = paste(
"Life Expectancy in Bangladesh",
"and the United States"
),
x = "Year",
y = "Life Expectancy at Birth (Years)",
color = "Country"
) +
scale_color_manual(
values = c(
"Bangladesh" = "#2C7FB8",
"United States" = "#D95F0E"
)
) +
theme_minimal()Both countries experienced an increase in life expectancy between 1960 and 2024, but the improvement was much larger in Bangladesh. Bangladesh’s life expectancy increased from 43.98 years to 74.93 years, which was a gain of 30.95 years. During the same period, life expectancy in the United States increased from 69.77 years to 78.89 years, a gain of 9.12 years. The two lines became much closer over time, suggesting that the life-expectancy gap between the countries decreased.
To measure the gap directly, I reshaped the two-country data so Bangladesh and the United States had separate columns. I then subtracted Bangladesh’s life expectancy from the United States’ life expectancy for each year.
life_expectancy_gap <- country_comparison |>
select(
country_code,
year,
life_expectancy
) |>
pivot_wider(
names_from = country_code,
values_from = life_expectancy
) |>
drop_na(BGD, USA) |>
mutate(
life_expectancy_gap = USA - BGD
) |>
arrange(year)
gap_summary <- life_expectancy_gap |>
summarize(
starting_year = first(year),
starting_gap = first(life_expectancy_gap),
ending_year = last(year),
ending_gap = last(life_expectancy_gap),
change_in_gap = last(life_expectancy_gap) -
first(life_expectancy_gap)
) |>
mutate(
across(
c(
starting_gap,
ending_gap,
change_in_gap
),
~ round(.x, 2)
)
)
kable(
gap_summary,
caption = paste(
"Life-Expectancy Gap Between the",
"United States and Bangladesh"
)
)| starting_year | starting_gap | ending_year | ending_gap | change_in_gap |
|---|---|---|---|---|
| 1960 | 25.79 | 2024 | 3.96 | -21.83 |
ggplot(
life_expectancy_gap,
aes(
x = year,
y = life_expectancy_gap
)
) +
geom_line(
color = "#6A3D9A",
linewidth = 1
) +
labs(
title = paste(
"Life-Expectancy Gap Between the",
"United States and Bangladesh"
),
subtitle = paste(
"Positive values show how many years",
"higher U.S. life expectancy was"
),
x = "Year",
y = "Difference in Life Expectancy (Years)"
) +
theme_minimal()The results confirm that the life-expectancy gap between the two countries became much smaller. In 1960, life expectancy in the United States was 25.79 years higher than in Bangladesh. By 2024, the difference had decreased to only 3.96 years. This means that the gap narrowed by 21.83 years during the period.
The graph also shows a sharp temporary increase in the gap around 1971. According to the World Bank data, Bangladesh’s life expectancy fell to 26.52 years during that year. This period included the Bangladesh Liberation War, political conflict, natural disaster, and severe food insecurity. The value recovered after 1971, and the long-term gap continued to decrease.
The original World Bank data set was successfully transformed from a wide format with separate year columns into a tidy format containing one row for each country and year. The missing values were kept as NA during the transformation and removed only from the data used for analysis.
Both Bangladesh and the United States experienced improvements in life expectancy between 1960 and 2024. However, Bangladesh experienced a much larger increase of 30.95 years, compared with 9.12 years in the United States. As a result, the life-expectancy gap between the two countries decreased from 25.79 years to 3.96 years. These results show that Bangladesh moved substantially closer to the United States in this measurement over time, even though a smaller gap still remained in 2024.