Introduction

In this exercise, I try to use the ELO formula of the probability of winning in the context of the work I am tying to do regarding disparities among children in Latin America. As in previous exercises, I am using information about birth registration for just two countries: Argentina and Jamaica.

Hopefully, comparing these probabilities could function as a way to normalize the differences and to compare disparities across axes of disparities and countries.

As usual, first, some libraries and the data are downloaded.

library(tidyverse)
## ── Attaching core tidyverse packages ──────────────────────── tidyverse 2.0.0 ──
## ✔ dplyr     1.2.1     ✔ readr     2.2.0
## ✔ forcats   1.0.1     ✔ stringr   1.6.0
## ✔ ggplot2   4.0.3     ✔ tibble    3.3.1
## ✔ lubridate 1.9.5     ✔ tidyr     1.3.2
## ✔ purrr     1.2.2     
## ── Conflicts ────────────────────────────────────────── tidyverse_conflicts() ──
## ✖ dplyr::filter() masks stats::filter()
## ✖ dplyr::lag()    masks stats::lag()
## ℹ Use the conflicted package (<http://conflicted.r-lib.org/>) to force all conflicts to become errors
library(dplyr)
library(ggplot2)
url<-"https://raw.githubusercontent.com/Enrique01234/607-Fall-2026/refs/heads/main/Arg%20Jam%20BReg%20simplified.csv"

The raw data from GitHub are converted into a data frame.

BReg <- read_csv(file = url, show_col_types = FALSE, progress = FALSE)
BReg
## # A tibble: 30 × 4
##    Ctry  Axis   Disp                 Rate
##    <chr> <chr>  <chr>               <dbl>
##  1 Arg   Total  Total                 1.8
##  2 Arg   Sex    Boy                   1.7
##  3 Arg   Sex    Girl                  1.8
##  4 Arg   Educ   Lower_secondary       2  
##  5 Arg   Educ   Upper_secondary       1.8
##  6 Arg   Educ   Post_secondary_plus   0.8
##  7 Arg   Region One                   2.5
##  8 Arg   Region Two                   2  
##  9 Arg   Region Three               100  
## 10 Arg   Region Four                 29.5
## # ℹ 20 more rows

A word on each column may be useful. “Ctry” refers to the name of the country (using my own abbreviations, not the UN Statistics Division standard three letter code). “Axis” refers to the four axes of disparity I have included in the data set: sex, level of education of the mother, sub-national region, and wealth quintiles. “Disp” includes the categories in each of these axes of disparities, for instance boys and girls for sex. In the geographic axis of disparities (Region), I have labeled the regions as “One”, “Two”, “Three”, and “Four” to avoid using the actual names which are different in each country (I also dropped one region in Argentina in order to have the same number of regions in each country). Finally “Rate” is the actual variable, the percentage of children under five years of age whose birth was registered and whose birth certificate is available.

Calculations

For each country and for each axis of disparity, the category with the highest and lowest probabilities are found. These values could be interpreted as the “rankings” in an ELO system.

BReg<- BReg|>
 group_by(Ctry, Axis)|>
  mutate(MaxRate = max(Rate))
BReg<- BReg|>
 group_by(Ctry, Axis)|>
  mutate(minRate = min(Rate))

Then, for each county and for each axis of disparity, the difference between the maximum and minimum rates are calculated.

BReg<- BReg|>
 group_by(Ctry, Axis)|>
  mutate(Diff = MaxRate - minRate)

These rates are then included in the ELO formula. This is done only for the first of the formulae. I cannot apply the second one as there is no change, there is only one observation in time. This is done in four steps.

BReg<- BReg|>
 group_by(Ctry, Axis)|>
  mutate(Diffby400 = (MaxRate - minRate)/400)
BReg<- BReg|>
 group_by(Ctry, Axis)|>
  mutate(Power10 = 10^((MaxRate - minRate)/400))
BReg<- BReg|>
 group_by(Ctry, Axis)|>
  mutate(Oneplus = 1 + (10^((MaxRate - minRate)/400)))
BReg<- BReg|>
 group_by(Ctry, Axis)|>
  mutate(OneOver = 1/ (1 + (10^((MaxRate - minRate)/400))))

Summary results and interpretation

As I tried to explore the calculated probabilities, I tried to look at the data frame and to summarize it.

BReg
## # A tibble: 30 × 11
## # Groups:   Ctry, Axis [10]
##    Ctry  Axis   Disp       Rate MaxRate minRate   Diff Diffby400 Power10 Oneplus
##    <chr> <chr>  <chr>     <dbl>   <dbl>   <dbl>  <dbl>     <dbl>   <dbl>   <dbl>
##  1 Arg   Total  Total       1.8     1.8     1.8  0      0           1       2   
##  2 Arg   Sex    Boy         1.7     1.8     1.7  0.100  0.000250    1.00    2.00
##  3 Arg   Sex    Girl        1.8     1.8     1.7  0.100  0.000250    1.00    2.00
##  4 Arg   Educ   Lower_se…   2       2       0.8  1.2    0.003       1.01    2.01
##  5 Arg   Educ   Upper_se…   1.8     2       0.8  1.2    0.003       1.01    2.01
##  6 Arg   Educ   Post_sec…   0.8     2       0.8  1.2    0.003       1.01    2.01
##  7 Arg   Region One         2.5   100       2   98      0.245       1.76    2.76
##  8 Arg   Region Two         2     100       2   98      0.245       1.76    2.76
##  9 Arg   Region Three     100     100       2   98      0.245       1.76    2.76
## 10 Arg   Region Four       29.5   100       2   98      0.245       1.76    2.76
## # ℹ 20 more rows
## # ℹ 1 more variable: OneOver <dbl>
BReg|>
  summary()
##         Ctry           Axis           Disp         Rate           MaxRate      
##  Length   :30   Length   :30   Length   :30   Min.   :  0.20   Min.   :  1.80  
##  N.unique : 2   N.unique : 5   N.unique :15   1st Qu.:  1.80   1st Qu.:  2.40  
##  N.blank  : 0   N.blank  : 0   N.blank  : 0   Median : 13.85   Median : 18.40  
##  Min.nchar: 3   Min.nchar: 3   Min.nchar: 3   Mean   : 13.36   Mean   : 24.09  
##  Max.nchar: 3   Max.nchar:14   Max.nchar:19   3rd Qu.: 17.05   3rd Qu.: 20.90  
##                                               Max.   :100.00   Max.   :100.00  
##     minRate            Diff         Diffby400          Power10     
##  Min.   : 0.200   Min.   : 0.00   Min.   :0.00000   Min.   :1.000  
##  1st Qu.: 1.025   1st Qu.: 2.20   1st Qu.:0.00550   1st Qu.:1.013  
##  Median : 6.200   Median : 2.70   Median :0.00675   Median :1.016  
##  Mean   : 7.223   Mean   :16.86   Mean   :0.04216   Mean   :1.123  
##  3rd Qu.:13.500   3rd Qu.:10.50   3rd Qu.:0.02625   3rd Qu.:1.062  
##  Max.   :16.400   Max.   :98.00   Max.   :0.24500   Max.   :1.758  
##     Oneplus         OneOver      
##  Min.   :2.000   Min.   :0.3626  
##  1st Qu.:2.013   1st Qu.:0.4849  
##  Median :2.016   Median :0.4961  
##  Mean   :2.123   Mean   :0.4762  
##  3rd Qu.:2.062   3rd Qu.:0.4968  
##  Max.   :2.758   Max.   :0.5000

However, there were too many variables. Consequently, I reduced the data frame by keeping only the variable with the calculated probabilities (OneOver). I also kept the countries (Ctry) and the axes of disparities (Axis).

BReg2 <- BReg|>
 select(Ctry, Axis, OneOver)

As it was still difficult to clearly see the results, I decided to convert the data frame back to a more traditional matrix look. However I could not do it because the different categories in each axis of disparity have the same probability (as it should be). Then, I eliminated the duplicates.

BReg2<- BReg2|>
 distinct(OneOver)

Finally, I was able to reshape (untidy?) the data frame.

BReg2<- BReg2|>
 pivot_wider(names_from = Axis, values_from = OneOver)

Thus, I was able to see clearly the probabilities in each category for each country.

BReg2
## # A tibble: 2 × 6
## # Groups:   Ctry [2]
##   Ctry  Total   Sex  Educ Region WealthQuintile
##   <chr> <dbl> <dbl> <dbl>  <dbl>          <dbl>
## 1 Arg     0.5 0.500 0.498  0.363          0.497
## 2 Jam     0.5 0.494 0.487  0.496          0.485

Although the total for the country has no specific meaning (because there is only one category), it helps to see the value of the probability (of “winning” in the ELO context) when there is no variation: the probability is 0.5. For the sex category, in both countries, the probability is almost 0.5 which can be interpreted as very little variation in birth registration (and having the birth certificate). The levels of disparities according to the wealth index or mother’s education are similar in both countries. It is a bit lower in Jamaica than in Argentina which means in Jamaica the gaps in birth registration (and having the birth certificate) between poorer and richer parents and between mothers with little of higher formal education are higher. Also, the gaps are higher than between boys and girls in Jamaica.

However, the most interesting (and promising) result is the low value in Argentina for the geographic axis of disparity (Region). In the first tibble above it can be seen that there is a region in Argentina where the rate of birth registration (and having the actual certificate) is 100 percent. This could be an error or an outlier. For the purposes of the exercise, it does not matter.

What matters, and makes the result promising, is that using the probabilities works to compare disparities across axes of disparities. This is a point I would like to keep on exploring.