Intoduction

Admittedly the level of complexity of what I managed to do in this exercise, as described below, is inferior to the ELO exercise. I apologize for that.

Change in course

On Friday morning, I was stuck with the exercise about the chess players and the ELO rating system. I had managed (after many trials and errors, watching some videos, and reading from https://www.rdocumentation.org/) to instruct R to skip the first lines of the text file when they were messing the setting up and counting of the columns, to write down in the syntax how to let R know how columns are separated in the text file and how decimals are written. The syntax allowed me to read the text file.

dfr<- read.table("https://brightspace.cuny.edu/content/enforced/1344368-SPS01_DATA_607_1269_1_34037/tournamentinfo.txt?ou=1344368", sep = "|",  skip = 2, row.names = NULL)

However, there were still something amiss. I was getting a data frame of one variable with 384 observations, I had not realized this when I submitted the approach answer on Thursday.

Alternative data set

As I had to put the project aside, I started to do some of my own work. I started collecting the data for the analysis of disparities among children in Latin America. In assignment one I explored how to select information about birth registration (disaggregated between boys and girls) for Latin American countries out of the global UNICEF.

However, other types of disaggregation (e.g. urban/rural on mother’s education) are not available online. They need to be retrieved from the actual country reports They are available in PDF by country here: https://mics.unicef.org/surveys).

Then I realized, I could try reading them with R, instead of having to type the values one by one in an Excel file. So, I used the reports from Argentina (https://mics.unicef.org/sites/mics/files/Argentina%202019-20%20MICS%20Survey%20Findings%20Report_Spanish.pdf) and Jamaica (https://mics.unicef.org/sites/mics/files/2024-07/Jamaica%202022%20MICS%20Survey%20Findings_English.pdf). I extracted the pages with the tables (uploaded in GitHub: https://github.com/Enrique01234/607-Fall-2026/blob/main/Argentina%202019-20%20MICS%20BReg.pdf ; https://github.com/Enrique01234/607-Fall-2026/blob/main/MICS%20Jamaica%20BReg.pdf ) and converted them to text (uploaded in GitHub: https://github.com/Enrique01234/607-Fall-2026/blob/main/Argentina%202019-20%20MICS%20BReg.txt ; https://github.com/Enrique01234/607-Fall-2026/blob/main/MICS%20Jamaica%20BReg.txt ).

Then, I started to face similar challenges as with the ELO rating system. For instance, I was getting errors about rows not having the required number of columns. I tried to fix this problem by skipping lines as with the ELO ranking exercise. Nevertheless, it was not working . I figured there was so much text, it was difficult for me to actually know how many rows to skip. I thought by just deleting the text this problem would not arise.

At that point, I decided to reverse engineer the text file to deal with the various issues that were coming up as errors. In other words, there were some elements I was able to fix with the syntax (as with the ELO ranking exercise) but I was still getting stuck. By reverse engineering, using the error messages, I could get a text file that could be read. Also, I discovered issues I could not have figured out otherwise For example, I noticed that R was getting confused by the spaces between words in the rows describing the categories (e.g. “lower secondary” when referring to the level of education). I used underscores to avoid this problem. Also, there were rows with no data, they were just subtitles to indicate the type of disaggregation (e.g. Mother’s education). I eliminated these lines.

Additionally, there was an instance when a line did not have the expected number of columns. Somehow the value had not been picked up when converting from PDF to text. In other words, there was an actual missing element (it was not recorded as not available, it was just a blank space). I manually typed the value from the PDF into the text file.

I used these files (https://github.com/Enrique01234/607-Fall-2026/blob/main/Argentina%202019-20%20MICS%20BReg%20without%20intro%20paragraph%20columns%20ONLY%20relevant%20without%20Table%20Title.txt ; https://github.com/Enrique01234/607-Fall-2026/blob/main/MICS%20Jamaica%20BReg%20without%20intro%20paragrap%20Titles%20subtitles%20and%20with%20underscore%20row%20names.txt ) to be read by R.

BReg_Arg <- read.table("https://github.com/Enrique01234/607-Fall-2026/blob/main/Argentina%202019-20%20MICS%20BReg%20without%20intro%20paragraph%20columns%20ONLY%20relevant%20without%20Table%20Title.txt")
## Error in `scan()`:
## ! line 3 did not have 2 elements
BReg_Jam <- read.table("https://github.com/Enrique01234/607-Fall-2026/blob/main/MICS%20Jamaica%20BReg%20without%20intro%20paragrap%20Titles%20subtitles%20and%20with%20underscore%20row%20names.txt", dec = ".")
## Error in `scan()`:
## ! line 3 did not have 2 elements

I did not get these errors when running the syntax on the same files from the C: drive. Please, see separate RMarkdown files.

I also tried to use the syntax to name the columns, given I had deleted the row with the column names. However, it did not work

Then, I converted to CSV files again, even without the proper column names. Although I could not do so in a reproducible way, I managed it in the C: drive. Please see attached RMarkdown files

Once I had done this for each country, I tried to read them both together. As I would need to have all of this information in one file for the comparative analysis, it would be better to read all countries in one command.

ctry_files <-c("Argentina 2019-20 MICS BReg without intro paragraph columns ONLY relevant without Table Title.txt", "MICS Jamaica BReg without intro paragrap Titles subtitles and with underscore row names.txt")
Combined <- read.table(ctry_files, id = "file")
## Error in `read.table()`:
## ! unused argument (id = "file")

However, it did not work and I could not figure out why.

Conclusion

Nevertheless, what I have found out most intriguing is that the syntax would work when running it with the file in the C: drive but not from GitHub, although it was exactly the same file.