Human Development Index vs Gross Domestic Product


Marc Siciliano
MBA Student | Data Analytics
ECON-6635-01 Business Forecasting | A.E. Rodriguez
Pompea College of Business | University of New Haven

Introduction

The United Nations Development Programme created the Human Development Index (HDI) “to emphasize that people and their capabilities should be the ultimate criteria for assessing the development of a country, not economic growth alone.” But does it really measure something different than economic growth or is it just a country’s Gross Domestic Product (GDP) in disguise?

The HDI is a composite index of three key dimensions:

  1. A Long and healthy life, as measured by Life Expectancy (LE) at birth.

  2. Knowledge, as measured by a combination of Expected Years of Schooling (EYS) and Mean Years of Schooling (MYS).

  3. A decent standard of living, as measured by Gross National Income per capita (GNI).

Let’s take a closer look at what really drives the HDI by analyzing these individual metrics in relation to the HDI as a whole.

# Start with a clean R Envrionment
options(digits = 3, scipen = 9, stringsAsFactors = FALSE, warn = -1)
remove(list = ls())
graphics.off()

# Set location for importing data
setwd("~/Library/CloudStorage/OneDrive-UniversityofNewHaven/2024-08 Fall/ECON 6635/Final Project")

#load the libraries required for data analysis
suppressWarnings({
  suppressPackageStartupMessages({
    library(tidyverse)
    library(ggplot2)
    library(rio)
    library(tidyr)
    library(Boruta)
    library(randomForest)
  })
})

Data Import and Preprocessing

Data for the individual Human Development Index (HDI) dimensions per country as well as the HDI itself is available at https://hdr.undp.org/data-center/documentation-and-downloads. The latest complete set of country data (from year 2022) was downloaded and imported for analysis. Per-capita GDP data was obtained from World Bank Group at https://data.worldbank.org/. All data was reduced into one dataframe by country for easier analysis.

# Import data for Expected Years of Schooling (EYS)
eys_raw <- rio::import("EYS.xlsx") %>% 
  dplyr::select('country','value')

# Import data for Mean Years of Schooling (EYS)
mys_raw <- rio::import("MYS.xlsx") %>% 
  dplyr::select('country','value')

# Import data for Life Expectancy (LE)
le_raw <- rio::import("LE.xlsx") %>%
  dplyr::select('country','value')

# Import data for per-capita Gross National Income (GNI)
gni_raw <- rio::import("GNIPC.xlsx") %>%
  dplyr::select('country','value')

# Import data for Human Development Index (HDI)
hdi_raw <- rio::import("HDI.xlsx") %>%
  dplyr::select('country','value')

# Import data for per-capita Gross Domestic Product (GDP)
gdp_raw <- rio::import("GDPPC.xls") %>%
  dplyr::select(country = 'Country Name', value = '2022')

# Combine individual dimension datasets into one dataframe by country
df <- list(eys_raw, mys_raw, le_raw, gni_raw, hdi_raw, gdp_raw) %>%
  reduce(inner_join, by = "country")

# Reset column names in new dataframe
colnames(df) <- c("Country", "EYS", "MYS", "LE", "GNI", "HDI", "GDP")

# Transform the "Country" column to the dataframe's row names
rownames(df) <- NULL
df <- df |> column_to_rownames("Country")

head(df); tail(df)
# Check for NAs
sum(is.na(df))
## [1] 7
# Very small number of NAs present, remove.
df <- na.omit(df)

Variability

So, how much variation is there between countries for each of the HDI’s metrics? The Coefficient of Variation was calculated for each metric and is shown in the table below. Right away we see that economic forces do play a big role in the HDI as the variability of GNI is much greater than any of the other dimensions.

# Compute CV (Coefficient of Variation) for each column
cv <- sapply(df, function(x) sd(x, na.rm = TRUE) / mean(x, na.rm = TRUE) * 100)
print(cv[1:4])
##  EYS  MYS   LE  GNI 
## 21.5 36.3 10.9 99.0

Feature Importance

But what about feature importance? The Boruta function can help us see which of the four individual metrics (EYS, MYS, LE, GNI) are most important to determining the HDI.

myboruta <- Boruta(HDI ~ EYS + MYS + LE + GNI, data = df)
plot(myboruta)

Years of Schooling and Life Expectancy are certainly important and not at all irrelevant to the HDI, but clearly economics (i.e. the GNI) is most important when determining the HDI for a country.

Correlation Matrix

We can look at this another way with a correlation matrix after scaling the data to make sure all are on a level playing field. Nominal values for GNI are naturally going to be significantly higher than values for years of school and life expectancy. Scaling the data ensures we don’t give undue weight to GNI if not warranted.

# Scale the dataset
df_s <- scale(df) |> as.data.frame()
head(df_s); tail(df_s)
Correlation Matrix
cor(df_s[,-6])
##       EYS   MYS    LE   GNI   HDI
## EYS 1.000 0.808 0.786 0.694 0.897
## MYS 0.808 1.000 0.723 0.687 0.916
## LE  0.786 0.723 1.000 0.758 0.905
## GNI 0.694 0.687 0.758 1.000 0.823
## HDI 0.897 0.916 0.905 0.823 1.000

The correlation matrix reveals something a little different than what we’ve seen so far. With the data scaled, GNI does not show the strongest correlation with HDI. Instead, the strongest correlation with HDI is MYS (Mean Years of Schooling) so maybe it’s a country’s collective knowledge of its population that is really driving the HDI? Let’s keep exploring…

Random Forest Model

Rather than a simple correlation matrix, let’s run a random forest simulation with the scaled data for HDI and see which of the four variables show the greatest importance.

# Create random forest model with scaled HDI data
rf <- randomForest(HDI~ EYS + MYS + LE + GNI, data = df_s, 
                   ntree=100, keep.forest=TRUE, importance=TRUE)

# Store importance scores
iscores = randomForest::varImpPlot(rf, sort=FALSE, scale = FALSE, main="Variable Importance")

Visualizing the Random Forest importance scores…

# Create new dataframe of importance scores for ggplot
importance_df <- data.frame(
  Feature = rownames(iscores),
  Importance = iscores[, 1]
)

# Plot Random Forest Importance Scores
ggplot(importance_df, aes(x = reorder(Feature, Importance), y = Importance, fill = Feature)) +
  geom_bar(stat = "identity", color = "black") +
  labs(
    title = "Feature Importance from Random Forest",
    x = "Features",
    y = "Importance Score"
  ) +
  scale_fill_brewer(palette = "Dark2") +
  theme_minimal(base_size = 14) +
  theme(
    plot.title = element_text(face = "bold", size = 16, hjust = 0.5),
    axis.title = element_text(face = "bold", size = 12),
    axis.text = element_text(size = 10),
    legend.position = "none",
    plot.caption = element_text(size = 10, hjust = 0),
    panel.grid.major.x = element_blank(),
    panel.grid.minor = element_blank()
  )

With the Random Forest model, GNI comes out on top again as the most important to the HDI.

So, we see that GNI plays perhaps the most significant role in what the HDI looks like for a country, but we started off this exploration asking what it looks like compared to GDP.

HDI & GDP Ranks

Let’s rank all the countries in our dataset by their HDI, rank them again by their GDP, and plot the results against each other. Any country that falls directly on the 45-degree line has the same rank regardless of whether we’re talking about HDI or GDP.

# Create country ranks for HDI and GDP
df$HDI_Rank <- rank(df$HDI)
df$GDP_Rank <- rank(df$GDP)

# Plot the 2 ranks against each other
ggplot(df, aes(x = HDI_Rank, y = GDP_Rank, label = rownames(df))) +
  geom_point(color = "lightblue", size = 3, alpha = 0.8) +
  geom_text(nudge_y = 2, size = 3.5, check_overlap = TRUE) +
  geom_smooth(method = "lm", se = FALSE, linetype = "dashed", color = "darkred") +
  labs(
    title = "HDI vs per-capita GDP - Rank by Country",
    x = "Country Rank: Human Development Index",
    y = "Country Rank: Per-Capita GDP",
    caption = "Data Sources: World Bank Group (https://data.worldbank.com);
    United Nations Development Programme (https://hdr.undp.org)"
  ) +
  theme_minimal(base_size = 14) +
  theme(
    plot.title = element_text(color = "darkblue", face = "bold", size = 16, hjust = 0.5),
    axis.title = element_text(face = "bold", size = 12),
    plot.caption = element_text(size = 10, hjust = 0),
    legend.position = "none"
  )
## `geom_smooth()` using formula = 'y ~ x'

There are a few outliers here for sure, but the trend along the 45-degree line is very clear.

Conclusion

The United Nations Development Programme doesn’t want to rely solely on economic data to assess the development of a country. And the Human Development Index with its combination of life expectancy, schooling, and economic data does that… to a point. It’s not all smoke and mirrors; education and life expectancy data do play a role in shaping the HDI. But if we rank a country’s development based on their Human Development Index or their per-capita Gross Domestic Product, there’s hardly much of a difference.