Approach

The objective of this assignment is to work through the wrangling and format of data using different packages. In order to transform the txt file of Chess tournament results into a tidy CSV file with our desired fields will take several steps, which is an interesting challenge. Leveraging packages such as dplyr for data wrangling/manipulation and tidyr for organizing the structure will become useful.

First, I’d have to load in the txt file into RStudio. From there, I can begin to work on cleaning out the initial format of the file. Once I remove any extra noise, like whitespace, from the file will assist on only filtering out the fields I want to use, such as the Player’s Name, Player’s State, Total Number of Points, Player’s Pre-Rating, and then calculate the new field, Average Pre Chess Rating of Opponents. Those are fields I want to be used in the generated CSV. I expect some challenges with the formatting piece, so I will have to explore how to handle wrangling data without a clear machine-readable structure.

Codebase (Prep)

# Loading needed libraries
library(tidyverse)
## ── Attaching core tidyverse packages ──────────────────────── tidyverse 2.0.0 ──
## ✔ dplyr     1.2.1     ✔ readr     2.2.0
## ✔ forcats   1.0.1     ✔ stringr   1.6.0
## ✔ ggplot2   4.0.3     ✔ tibble    3.3.1
## ✔ lubridate 1.9.5     ✔ tidyr     1.3.2
## ✔ purrr     1.2.2     
## ── Conflicts ────────────────────────────────────────── tidyverse_conflicts() ──
## ✖ dplyr::filter() masks stats::filter()
## ✖ dplyr::lag()    masks stats::lag()
## ℹ Use the conflicted package (<http://conflicted.r-lib.org/>) to force all conflicts to become errors
library(readr)

# Importing data from GitHub Raw Link

chess_file_import <- "https://raw.githubusercontent.com/shanicesmith98/data-607-assignments/refs/heads/main/week-4-assignment/tournamentinfo.txt"
chess_file <- readLines(chess_file_import, warn = FALSE) # added the warn arg to hide the final line warning from appearing
head(chess_file, 10)
##  [1] "-----------------------------------------------------------------------------------------" 
##  [2] " Pair | Player Name                     |Total|Round|Round|Round|Round|Round|Round|Round| "
##  [3] " Num  | USCF ID / Rtg (Pre->Post)       | Pts |  1  |  2  |  3  |  4  |  5  |  6  |  7  | "
##  [4] "-----------------------------------------------------------------------------------------" 
##  [5] "    1 | GARY HUA                        |6.0  |W  39|W  21|W  18|W  14|W   7|D  12|D   4|" 
##  [6] "   ON | 15445895 / R: 1794   ->1817     |N:2  |W    |B    |W    |B    |W    |B    |W    |" 
##  [7] "-----------------------------------------------------------------------------------------" 
##  [8] "    2 | DAKSHESH DARURI                 |6.0  |W  63|W  58|L   4|W  17|W  16|W  20|W   7|" 
##  [9] "   MI | 14598900 / R: 1553   ->1663     |N:2  |B    |W    |B    |W    |B    |W    |B    |" 
## [10] "-----------------------------------------------------------------------------------------"

Codebase (Clean + Execution)

# Cleaning the separators and header rows
file_lines <- chess_file[!str_detect(chess_file, "^-+$")]
file_lines <- file_lines[-(1:2)]

line1 <- file_lines[seq(1, length(file_lines), by = 2)] 
line2 <- file_lines[seq(2, length(file_lines), by = 2)]
p1 <- str_split_fixed(line1, "\\|", 11)
p2 <- str_split_fixed(line2, "\\|", 11)

glimpse(file_lines)
##  chr [1:128] "    1 | GARY HUA                        |6.0  |W  39|W  21|W  18|W  14|W   7|D  12|D   4|" ...
# Create player df to get existing fields

players <- tibble(
  id         = as.integer(str_trim(p1[, 1])),
  name       = str_to_title(str_trim(p1[, 2])),
  points     = as.numeric(str_trim(p1[, 3])),
  state      = str_trim(p2[, 1]),
  pre_rating = as.integer(str_match(p2[, 2], "R:\\s*(\\d+)")[, 2])
)

opp_ids <- apply(p1[, 4:10], 2, function(x) as.integer(str_extract(x, "\\d+")))

# Calculate field

opp_ratings <- matrix(players$pre_rating[match(opp_ids, players$id)],
                      nrow = nrow(opp_ids))
players$games_played   <- rowSums(!is.na(opp_ids))
players$avg_opp_rating <- round(rowMeans(opp_ratings, na.rm = TRUE))

Codebase (Test)

# Validate

players %>% filter(id %in% c(1, 41, 62))
## # A tibble: 3 × 7
##      id name                points state pre_rating games_played avg_opp_rating
##   <int> <chr>                <dbl> <chr>      <int>        <dbl>          <dbl>
## 1     1 Gary Hua                 6 ON          1794            7           1605
## 2    41 Kyle William Murphy      3 MI          1403            4           1248
## 3    62 Ashwin Balaji            1 MI          1530            1           1186
# Gary Hua: 7 games, 1605 | Kyle William Murphy: 4 games, 1248 | Ashwin Balaji: 1 game, 1186
nrow(players)                  # 64
## [1] 64
sum(is.na(players$pre_rating)) # 0
## [1] 0

Codebase (Export)

players %>%
  mutate(points = sprintf("%.1f", points)) %>%   # keeps 6.0 instead of 6
  select(name, state, points, pre_rating, avg_opp_rating) %>%
  write_csv("chess_ratings.csv")