In this exercise, I try to use the ELO formula of the probability of winning in the context of the work I am tying to do regarding disparities among children in Latin America. As in previous exercises, I am using information about birth registration for just two countries: Argentina and Jamaica.
Hopefully, comparing these probabilities could function as a way to normalize the differences and to compare disparities across axes of disparities and countries.
As usual, first, some libraries and the data are downloaded.
library(tidyverse)
## ── Attaching core tidyverse packages ──────────────────────── tidyverse 2.0.0 ──
## ✔ dplyr 1.2.1 ✔ readr 2.2.0
## ✔ forcats 1.0.1 ✔ stringr 1.6.0
## ✔ ggplot2 4.0.3 ✔ tibble 3.3.1
## ✔ lubridate 1.9.5 ✔ tidyr 1.3.2
## ✔ purrr 1.2.2
## ── Conflicts ────────────────────────────────────────── tidyverse_conflicts() ──
## ✖ dplyr::filter() masks stats::filter()
## ✖ dplyr::lag() masks stats::lag()
## ℹ Use the conflicted package (<http://conflicted.r-lib.org/>) to force all conflicts to become errors
library(dplyr)
library(ggplot2)
url<-"https://raw.githubusercontent.com/Enrique01234/607-Fall-2026/refs/heads/main/Arg%20Jam%20BReg%20simplified.csv"
The raw data from GitHub are converted into a data frame.
BReg <- read_csv(file = url, show_col_types = FALSE, progress = FALSE)
BReg
## # A tibble: 30 × 4
## Ctry Axis Disp Rate
## <chr> <chr> <chr> <dbl>
## 1 Arg Total Total 1.8
## 2 Arg Sex Boy 1.7
## 3 Arg Sex Girl 1.8
## 4 Arg Educ Lower_secondary 2
## 5 Arg Educ Upper_secondary 1.8
## 6 Arg Educ Post_secondary_plus 0.8
## 7 Arg Region One 2.5
## 8 Arg Region Two 2
## 9 Arg Region Three 100
## 10 Arg Region Four 29.5
## # ℹ 20 more rows
A word on each column may be useful. “Ctry” refers to the name of the country (using my own abbreviations, not the UN Statistics Division standard three letter code). “Axis” refers to the four axes of disparity I have included in the data set: sex, level of education of the mother, sub-national region, and wealth quintiles. “Disp” includes the categories in each of these axes of disparities, for instance boys and girls for sex. In the geographic axis of disparities (Region), I have labeled the regions as “One”, “Two”, “Three”, and “Four” to avoid using the actual names which are different in each country (I also dropped one region in Argentina in order to have the same number of regions in each country). Finally “Rate” is the actual variable, the percentage of children under five years of age whose birth was registered and whose birth certificate is available.
For each country and for each axis of disparity, the category with the highest and lowest probabilities are found. These values could be interpreted as the “rankings” in an ELO system.
BReg<- BReg|>
group_by(Ctry, Axis)|>
mutate(MaxRate = max(Rate))
BReg<- BReg|>
group_by(Ctry, Axis)|>
mutate(minRate = min(Rate))
Then, for each county and for each axis of disparity, the difference between the maximum and minimum rates are calculated.
BReg<- BReg|>
group_by(Ctry, Axis)|>
mutate(Diff = MaxRate - minRate)
These rates are then included in the ELO formula. This is done only for the first of the formulae. I cannot apply the second one as there is no change, there is only one observation in time. This is done in four steps.
BReg<- BReg|>
group_by(Ctry, Axis)|>
mutate(Diffby400 = (MaxRate - minRate)/400)
BReg<- BReg|>
group_by(Ctry, Axis)|>
mutate(Power10 = 10^((MaxRate - minRate)/400))
BReg<- BReg|>
group_by(Ctry, Axis)|>
mutate(Oneplus = 1 + (10^((MaxRate - minRate)/400)))
BReg<- BReg|>
group_by(Ctry, Axis)|>
mutate(OneOver = 1/ (1 + (10^((MaxRate - minRate)/400))))
As I tried to explore the calculated probabilities, I tried to look at the data frame and to summarize it.
BReg
## # A tibble: 30 × 11
## # Groups: Ctry, Axis [10]
## Ctry Axis Disp Rate MaxRate minRate Diff Diffby400 Power10 Oneplus
## <chr> <chr> <chr> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl>
## 1 Arg Total Total 1.8 1.8 1.8 0 0 1 2
## 2 Arg Sex Boy 1.7 1.8 1.7 0.100 0.000250 1.00 2.00
## 3 Arg Sex Girl 1.8 1.8 1.7 0.100 0.000250 1.00 2.00
## 4 Arg Educ Lower_se… 2 2 0.8 1.2 0.003 1.01 2.01
## 5 Arg Educ Upper_se… 1.8 2 0.8 1.2 0.003 1.01 2.01
## 6 Arg Educ Post_sec… 0.8 2 0.8 1.2 0.003 1.01 2.01
## 7 Arg Region One 2.5 100 2 98 0.245 1.76 2.76
## 8 Arg Region Two 2 100 2 98 0.245 1.76 2.76
## 9 Arg Region Three 100 100 2 98 0.245 1.76 2.76
## 10 Arg Region Four 29.5 100 2 98 0.245 1.76 2.76
## # ℹ 20 more rows
## # ℹ 1 more variable: OneOver <dbl>
BReg|>
summary()
## Ctry Axis Disp Rate MaxRate
## Length :30 Length :30 Length :30 Min. : 0.20 Min. : 1.80
## N.unique : 2 N.unique : 5 N.unique :15 1st Qu.: 1.80 1st Qu.: 2.40
## N.blank : 0 N.blank : 0 N.blank : 0 Median : 13.85 Median : 18.40
## Min.nchar: 3 Min.nchar: 3 Min.nchar: 3 Mean : 13.36 Mean : 24.09
## Max.nchar: 3 Max.nchar:14 Max.nchar:19 3rd Qu.: 17.05 3rd Qu.: 20.90
## Max. :100.00 Max. :100.00
## minRate Diff Diffby400 Power10
## Min. : 0.200 Min. : 0.00 Min. :0.00000 Min. :1.000
## 1st Qu.: 1.025 1st Qu.: 2.20 1st Qu.:0.00550 1st Qu.:1.013
## Median : 6.200 Median : 2.70 Median :0.00675 Median :1.016
## Mean : 7.223 Mean :16.86 Mean :0.04216 Mean :1.123
## 3rd Qu.:13.500 3rd Qu.:10.50 3rd Qu.:0.02625 3rd Qu.:1.062
## Max. :16.400 Max. :98.00 Max. :0.24500 Max. :1.758
## Oneplus OneOver
## Min. :2.000 Min. :0.3626
## 1st Qu.:2.013 1st Qu.:0.4849
## Median :2.016 Median :0.4961
## Mean :2.123 Mean :0.4762
## 3rd Qu.:2.062 3rd Qu.:0.4968
## Max. :2.758 Max. :0.5000
However, there were too many variables. Consequently, I reduced the data frame by keeping only the variable with the calculated probabilities (OneOver). I also kept the countries (Ctry) and the axes of disparities (Axis).
BReg2 <- BReg|>
select(Ctry, Axis, OneOver)
As it was still difficult to clearly see the results, I decided to convert the data frame back to a more traditional matrix look. However I could not do it because the different categories in each axis of disparity have the same probability (as it should be). Then, I eliminated the duplicates.
BReg2<- BReg2|>
distinct(OneOver)
Finally, I was able to reshape (untidy?) the data frame.
BReg2<- BReg2|>
pivot_wider(names_from = Axis, values_from = OneOver)
Thus, I was able to see clearly the probabilities in each category for each country.
BReg2
## # A tibble: 2 × 6
## # Groups: Ctry [2]
## Ctry Total Sex Educ Region WealthQuintile
## <chr> <dbl> <dbl> <dbl> <dbl> <dbl>
## 1 Arg 0.5 0.500 0.498 0.363 0.497
## 2 Jam 0.5 0.494 0.487 0.496 0.485
Although the total for the country has no specific meaning (because there is only one category), it helps to see the value of the probability (of “winning” in the ELO context) when there is no variation: the probability is 0.5. For the sex category, in both countries, the probability is almost 0.5 which can be interpreted as very little variation in birth registration (and having the birth certificate). The levels of disparities according to the wealth index or mother’s education are similar in both countries. It is a bit lower in Jamaica than in Argentina which means in Jamaica the gaps in birth registration (and having the birth certificate) between poorer and richer parents and between mothers with little of higher formal education are higher. Also, the gaps are higher than between boys and girls in Jamaica.
However, the most interesting (and promising) result is the low value in Argentina for the geographic axis of disparity (Region). In the first tibble above it can be seen that there is a region in Argentina where the rate of birth registration (and having the actual certificate) is 100 percent. This could be an error or an outlier. For the purposes of the exercise, it does not matter.
What matters, and makes the result promising, is that using the probabilities works to compare disparities across axes of disparities. This is a point I would like to keep on exploring.