The objective of this assignment is to work through the wrangling and format of data using different packages. In order to transform the txt file of Chess tournament results into a tidy CSV file with our desired fields will take several steps, which is an interesting challenge. Leveraging packages such as dplyr for data wrangling/manipulation and tidyr for organizing the structure will become useful.
First, I’d have to load in the txt file into RStudio. From there, I can begin to work on cleaning out the initial format of the file. Once I remove any extra noise, like whitespace, from the file will assist on only filtering out the fields I want to use, such as the Player’s Name, Player’s State, Total Number of Points, Player’s Pre-Rating, and then calculate the new field, Average Pre Chess Rating of Opponents. Those are fields I want to be used in the generated CSV. I expect some challenges with the formatting piece, so I will have to explore how to handle wrangling data without a clear machine-readable structure.
# Loading needed libraries
library(tidyverse)
## ── Attaching core tidyverse packages ──────────────────────── tidyverse 2.0.0 ──
## ✔ dplyr 1.2.1 ✔ readr 2.2.0
## ✔ forcats 1.0.1 ✔ stringr 1.6.0
## ✔ ggplot2 4.0.3 ✔ tibble 3.3.1
## ✔ lubridate 1.9.5 ✔ tidyr 1.3.2
## ✔ purrr 1.2.2
## ── Conflicts ────────────────────────────────────────── tidyverse_conflicts() ──
## ✖ dplyr::filter() masks stats::filter()
## ✖ dplyr::lag() masks stats::lag()
## ℹ Use the conflicted package (<http://conflicted.r-lib.org/>) to force all conflicts to become errors
library(readr)
# Importing data from GitHub Raw Link
chess_file_import <- "https://raw.githubusercontent.com/shanicesmith98/data-607-assignments/refs/heads/main/week-4-assignment/tournamentinfo.txt"
chess_file <- readLines(chess_file_import, warn = FALSE) # added the warn arg to hide the final line warning from appearing
head(chess_file, 10)
## [1] "-----------------------------------------------------------------------------------------"
## [2] " Pair | Player Name |Total|Round|Round|Round|Round|Round|Round|Round| "
## [3] " Num | USCF ID / Rtg (Pre->Post) | Pts | 1 | 2 | 3 | 4 | 5 | 6 | 7 | "
## [4] "-----------------------------------------------------------------------------------------"
## [5] " 1 | GARY HUA |6.0 |W 39|W 21|W 18|W 14|W 7|D 12|D 4|"
## [6] " ON | 15445895 / R: 1794 ->1817 |N:2 |W |B |W |B |W |B |W |"
## [7] "-----------------------------------------------------------------------------------------"
## [8] " 2 | DAKSHESH DARURI |6.0 |W 63|W 58|L 4|W 17|W 16|W 20|W 7|"
## [9] " MI | 14598900 / R: 1553 ->1663 |N:2 |B |W |B |W |B |W |B |"
## [10] "-----------------------------------------------------------------------------------------"
# Cleaning the separators and header rows
file_lines <- chess_file[!str_detect(chess_file, "^-+$")]
file_lines <- file_lines[-(1:2)]
line1 <- file_lines[seq(1, length(file_lines), by = 2)]
line2 <- file_lines[seq(2, length(file_lines), by = 2)]
p1 <- str_split_fixed(line1, "\\|", 11)
p2 <- str_split_fixed(line2, "\\|", 11)
glimpse(file_lines)
## chr [1:128] " 1 | GARY HUA |6.0 |W 39|W 21|W 18|W 14|W 7|D 12|D 4|" ...
# Create player df to get existing fields
players <- tibble(
id = as.integer(str_trim(p1[, 1])),
name = str_to_title(str_trim(p1[, 2])),
points = as.numeric(str_trim(p1[, 3])),
state = str_trim(p2[, 1]),
pre_rating = as.integer(str_match(p2[, 2], "R:\\s*(\\d+)")[, 2])
)
opp_ids <- apply(p1[, 4:10], 2, function(x) as.integer(str_extract(x, "\\d+")))
# Calculate field
opp_ratings <- matrix(players$pre_rating[match(opp_ids, players$id)],
nrow = nrow(opp_ids))
players$games_played <- rowSums(!is.na(opp_ids))
players$avg_opp_rating <- round(rowMeans(opp_ratings, na.rm = TRUE))
# Validate
players %>% filter(id %in% c(1, 41, 62))
## # A tibble: 3 × 7
## id name points state pre_rating games_played avg_opp_rating
## <int> <chr> <dbl> <chr> <int> <dbl> <dbl>
## 1 1 Gary Hua 6 ON 1794 7 1605
## 2 41 Kyle William Murphy 3 MI 1403 4 1248
## 3 62 Ashwin Balaji 1 MI 1530 1 1186
# Gary Hua: 7 games, 1605 | Kyle William Murphy: 4 games, 1248 | Ashwin Balaji: 1 game, 1186
nrow(players) # 64
## [1] 64
sum(is.na(players$pre_rating)) # 0
## [1] 0
players %>%
mutate(points = sprintf("%.1f", points)) %>% # keeps 6.0 instead of 6
select(name, state, points, pre_rating, avg_opp_rating) %>%
write_csv("chess_ratings.csv")