##Setup

Approach

For this project, I plan to use R to clean and organize the provided chess tournament text file into a structured dataset. I will extract each player’s name, state, total points, pre-tournament rating, and opponent information from the text. I will then use the opponent numbers to match each opponent with their pre-tournament rating and calculate the average opponent rating for every player. Finally, I will organize the results into a data frame and export the completed data as a CSV file. I will use online resources and LLM tools as needed to help understand the text-processing and data-cleaning methods required for the project.

Data Import

tournament <- readLines("tournamentinfo.txt")
## Warning in readLines("tournamentinfo.txt"): incomplete final line found on
## 'tournamentinfo.txt'
head(tournament, 10)
##  [1] "-----------------------------------------------------------------------------------------" 
##  [2] " Pair | Player Name                     |Total|Round|Round|Round|Round|Round|Round|Round| "
##  [3] " Num  | USCF ID / Rtg (Pre->Post)       | Pts |  1  |  2  |  3  |  4  |  5  |  6  |  7  | "
##  [4] "-----------------------------------------------------------------------------------------" 
##  [5] "    1 | GARY HUA                        |6.0  |W  39|W  21|W  18|W  14|W   7|D  12|D   4|" 
##  [6] "   ON | 15445895 / R: 1794   ->1817     |N:2  |W    |B    |W    |B    |W    |B    |W    |" 
##  [7] "-----------------------------------------------------------------------------------------" 
##  [8] "    2 | DAKSHESH DARURI                 |6.0  |W  63|W  58|L   4|W  17|W  16|W  20|W   7|" 
##  [9] "   MI | 14598900 / R: 1553   ->1663     |N:2  |B    |W    |B    |W    |B    |W    |B    |" 
## [10] "-----------------------------------------------------------------------------------------"

Data Parsing

player_lines <- tournament[seq(5, length(tournament), by = 3)]

head(player_lines)
## [1] "    1 | GARY HUA                        |6.0  |W  39|W  21|W  18|W  14|W   7|D  12|D   4|"
## [2] "    2 | DAKSHESH DARURI                 |6.0  |W  63|W  58|L   4|W  17|W  16|W  20|W   7|"
## [3] "    3 | ADITYA BAJAJ                    |6.0  |L   8|W  61|W  25|W  21|W  11|W  13|W  12|"
## [4] "    4 | PATRICK H SCHILLING             |5.5  |W  23|D  28|W   2|W  26|D   5|W  19|D   1|"
## [5] "    5 | HANSHI ZUO                      |5.5  |W  45|W  37|D  12|D  13|D   4|W  14|W  17|"
## [6] "    6 | HANSEN SONG                     |5.0  |W  34|D  29|L  11|W  35|D  10|W  27|W  21|"
length(player_lines)
## [1] 64
player_number <- as.numeric(
  trimws(substr(player_lines, 1, 5))
)

player_name <- trimws(
  substr(player_lines, 7, 40)
)

total_points <- as.numeric(
  trimws(substr(player_lines, 42, 45))
)

players <- data.frame(
  player_number,
  player_name,
  total_points
)

head(players)
##   player_number           player_name total_points
## 1             1            | GARY HUA          6.0
## 2             2     | DAKSHESH DARURI          6.0
## 3             3        | ADITYA BAJAJ          6.0
## 4             4 | PATRICK H SCHILLING          5.5
## 5             5          | HANSHI ZUO          5.5
## 6             6         | HANSEN SONG          5.0
rating_lines <- tournament[seq(6, length(tournament), by = 3)]

players$state <- trimws(substr(rating_lines, 1, 5))

players$pre_rating <- as.numeric(
  sub(".*R:\\s*([0-9]+).*", "\\1", rating_lines)
)

head(players)
##   player_number           player_name total_points state pre_rating
## 1             1            | GARY HUA          6.0    ON       1794
## 2             2     | DAKSHESH DARURI          6.0    MI       1553
## 3             3        | ADITYA BAJAJ          6.0    MI       1384
## 4             4 | PATRICK H SCHILLING          5.5    MI       1716
## 5             5          | HANSHI ZUO          5.5    MI       1655
## 6             6         | HANSEN SONG          5.0    OH       1686
round_results <- strsplit(player_lines, "|", fixed = TRUE)

opponent_numbers <- lapply(round_results, function(fields) {
  as.numeric(stringr::str_extract(fields[4:10], "\\d+"))
})

opponent_numbers[[1]]
## [1] 39 21 18 14  7 12  4
players$avg_opponent_rating <- sapply(opponent_numbers, function(opponents) {
  opponent_ratings <- players$pre_rating[
    match(opponents, players$player_number)
  ]

  round(mean(opponent_ratings, na.rm = TRUE))
})

head(players)
##   player_number           player_name total_points state pre_rating
## 1             1            | GARY HUA          6.0    ON       1794
## 2             2     | DAKSHESH DARURI          6.0    MI       1553
## 3             3        | ADITYA BAJAJ          6.0    MI       1384
## 4             4 | PATRICK H SCHILLING          5.5    MI       1716
## 5             5          | HANSHI ZUO          5.5    MI       1655
## 6             6         | HANSEN SONG          5.0    OH       1686
##   avg_opponent_rating
## 1                1605
## 2                1469
## 3                1564
## 4                1574
## 5                1501
## 6                1519

Creating CSV file

final_results <- data.frame(
  "Player Name" = players$player_name,
  "State" = players$state,
  "Total Points" = players$total_points,
  "Pre-Rating" = players$pre_rating,
  "Average Opponent Pre-Rating" = players$avg_opponent_rating,
  check.names = FALSE
)

knitr::kable(head(final_results))
Player Name State Total Points Pre-Rating Average Opponent Pre-Rating
| GARY HUA ON 6.0 1794 1605
| DAKSHESH DARURI MI 6.0 1553 1469
| ADITYA BAJAJ MI 6.0 1384 1564
| PATRICK H SCHILLING MI 5.5 1716 1574
| HANSHI ZUO MI 5.5 1655 1501
| HANSEN SONG OH 5.0 1686 1519
write.csv(final_results, "chess_results.csv", row.names = FALSE)

Findings and Suggestions

The tournament text file was successfully transformed into a structured dataset containing 64 players. The final CSV includes each player’s name, state, total points, pre-tournament rating, and average opponent pre-tournament rating. Rounds without an opponent were excluded from the average. As a validation check, Gary Hua’s average opponent rating matched the expected value of 1605.

Average opponent rating provides additional context when comparing players’ tournament results because players faced opponents of different strengths. Future analysis could explore the relationship between pre-tournament ratings and total points, or examine how opponent strength relates to performance. The extraction process could also be improved by identifying player records through text patterns rather than relying on fixed line positions, making it more adaptable to changes in the source file.