url <- "https://raw.githubusercontent.com/stevenmacaluso/DATA607-Project1/refs/heads/main/tournamentinfo.txt"
lines <- readLines(url, warn = FALSE)Project 1
Introduction
In this project, I will be using a dataset of chess tournament results, and generating a .CSV file with important information regarding each player. The file will include the player’s name, state, total number of points, pre-rating, and average pre-rating of their opponents. This endeavor will require using R to read the information from the .TXT file containing the data, extracting the relevant data and performing necessary calculations, and exporting these findings as a .CSV file. I anticipate the data extraction and calculations will be the most difficult part of this project. These tasks will require the most amount of coding and attention to detail to ensure the data is being organized properly, and the correct data is being used for the calculations.
Body
To start, I load the data that I have stored in a GitHub repository dedicated to this project. The following code will save the URL as a variable and then read the lines from the file:
“warn = FALSE” isn’t necessary, but the TXT file does not end with a final newline character, which generates a warning. The data is still imported correctly, so the warning can be ignored, and so I’ve set “warn” to FALSE to avoid viewing the warning.
Now, I will do some data cleaning. The TXT file organizes the data in a table-like format, which helps with readability. That being said, in order for the data to be properly prepared to be stored as a dataframe, is it essential to remove unnecessary characters. The following lines of code will remove whitespace, remove blank lines, remove dashed separator lines, and remove the two header lines.
lines <- trimws(lines)
lines <- lines[nzchar(lines)]
lines <- lines[!grepl("^-+$", lines)]
lines <- lines[!grepl("^(Pair|Num)\\b", lines)]The next thing I do is write two lines of code to separate the first and second line for each player. This will make the data more accessible and greatly assist with creating the dataframe.
player_lines <- lines[seq(1, length(lines), by = 2)]
detail_lines <- lines[seq(2, length(lines), by = 2)]After that, I write a function to split each line using the symbol, “|”
split_line <- function(x) {
pieces <- trimws(strsplit(x, "|", fixed = TRUE)[[1]])
pieces[pieces != ""]
}I then implement this function on both “player_lines” and “detail_lines.”
player_data <- do.call(rbind, lapply(player_lines, split_line))
detail_data <- do.call(rbind, lapply(detail_lines, split_line))The last step necessary before I create the dataframe is to extract the pre-ratings and remove provisional rating information. This is done with the following code,
pre_rating_text <- sub(".*R:\\s*", "", detail_data[, 2])
pre_rating_text <- sub("->.*", "", pre_rating_text)
pre_rating <- as.numeric(sub("P.*", "", trimws(pre_rating_text)))Now I can create my dataframe. The intial dataframe will contain the columns Pair, Player_Name, State, Total_Points, and Pre-Rating. Later, we will add Average_Opponent_Pre_Rating, and remove Pair. The dataframe is created with the following code,
chess <- data.frame(
Pair = as.integer(player_data[, 1]),
Player_Name = player_data[, 2],
State = detail_data[, 1],
Total_Points = as.numeric(player_data[, 3]),
Pre_Rating = pre_rating
)library(tidyverse)glimpse(chess)Rows: 64
Columns: 5
$ Pair <int> 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17…
$ Player_Name <chr> "GARY HUA", "DAKSHESH DARURI", "ADITYA BAJAJ", "PATRICK H…
$ State <chr> "ON", "MI", "MI", "MI", "MI", "OH", "MI", "MI", "ON", "MI…
$ Total_Points <dbl> 6.0, 6.0, 6.0, 5.5, 5.5, 5.0, 5.0, 5.0, 5.0, 5.0, 4.5, 4.…
$ Pre_Rating <dbl> 1794, 1553, 1384, 1716, 1655, 1686, 1649, 1641, 1411, 136…
Now that the dataframe is created, it is time to calculate the average opponent pre-ratings and store them in a new column. First, I need to extract the result columns from player_data.
round_results <- player_data[, 4:10]The next step is to create a function to extract the opponent’s pair number. This function will remove everything except numbers, record any value without a number as “NA,” and then ensure the pair number is in integer form.
extract_opponent <- function(x) {
opponent <- gsub("\\D", "", x)
opponent[opponent == ""] <- NA
as.integer(opponent)
}This function must be applied to all seven rounds, which is done with the following code:
opponent_numbers <- apply(
round_results,
2,
extract_opponent
)The next step is to look up each opponent’s pre-rating using their pair number
opponent_ratings <- matrix(
chess$Pre_Rating[
match(opponent_numbers, chess$Pair)
],
nrow = nrow(opponent_numbers),
ncol = ncol(opponent_numbers)
)At this point, I am now ready to make the calculation. This code will calculate the player’s average opponent pre-rating and ignore rounds where no opponent was played.
chess$Avg_Opponent_Pre_Rating <- rowMeans(
opponent_ratings,
na.rm = TRUE
)After this, I have to round to the nearest whole number.
chess$Avg_Opponent_Pre_Rating <- round(
chess$Avg_Opponent_Pre_Rating
)Here is a glimpse of the updated dataframe:
glimpse(chess)Rows: 64
Columns: 6
$ Pair <int> 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14,…
$ Player_Name <chr> "GARY HUA", "DAKSHESH DARURI", "ADITYA BAJAJ",…
$ State <chr> "ON", "MI", "MI", "MI", "MI", "OH", "MI", "MI"…
$ Total_Points <dbl> 6.0, 6.0, 6.0, 5.5, 5.5, 5.0, 5.0, 5.0, 5.0, 5…
$ Pre_Rating <dbl> 1794, 1553, 1384, 1716, 1655, 1686, 1649, 1641…
$ Avg_Opponent_Pre_Rating <dbl> 1605, 1469, 1564, 1574, 1501, 1519, 1372, 1468…
After this, I must prepare the data to be exported. First, I need to remove the “Pair” column. I will create a new dataframe without the pair column using this code,
final_chess <- chess[, c(
"Player_Name",
"State",
"Total_Points",
"Pre_Rating",
"Avg_Opponent_Pre_Rating"
)]glimpse(final_chess)Rows: 64
Columns: 5
$ Player_Name <chr> "GARY HUA", "DAKSHESH DARURI", "ADITYA BAJAJ",…
$ State <chr> "ON", "MI", "MI", "MI", "MI", "OH", "MI", "MI"…
$ Total_Points <dbl> 6.0, 6.0, 6.0, 5.5, 5.5, 5.0, 5.0, 5.0, 5.0, 5…
$ Pre_Rating <dbl> 1794, 1553, 1384, 1716, 1655, 1686, 1649, 1641…
$ Avg_Opponent_Pre_Rating <dbl> 1605, 1469, 1564, 1574, 1501, 1519, 1372, 1468…
The last step is to export this dataframe as a CSV file.
write.csv(final_chess, "chess.csv", row.names = FALSE)“row.names” is set to FALSE to avoid row numbers that are added by default.
Conclusion
To summarize, given a TXT file full of chess tournament information, I used R to read and clean the data, extract key values, perform necessary calculations, and finally export my findings as a CSV file. To build upon this work, the CSV file could be added to a SQL database with many similar findings from different tournaments. Potentially, I could have a database of all the major chess tournaments in this region, and then perform analysis on this larger set of data. If I obtain data from other regions, I could then compare the skill level of the different regions.