This project turns a messy chess tournament text file into a clean CSV with one row per player: name, state, total points, pre-tournament rating, and the average pre-tournament rating of their opponents. Each player’s data is split across two lines, and the round columns list opponent pair numbers instead of ratings, so the opponent ratings have to be looked up separately.
Read the File
library(tidyverse)
── Attaching core tidyverse packages ──────────────────────── tidyverse 2.0.0 ──
✔ dplyr 1.2.1 ✔ readr 2.2.0
✔ forcats 1.0.1 ✔ stringr 1.6.0
✔ ggplot2 4.0.3 ✔ tibble 3.3.1
✔ lubridate 1.9.5 ✔ tidyr 1.3.2
✔ purrr 1.2.2
── Conflicts ────────────────────────────────────────── tidyverse_conflicts() ──
✖ dplyr::filter() masks stats::filter()
✖ dplyr::lag() masks stats::lag()
ℹ Use the conflicted package (<http://conflicted.r-lib.org/>) to force all conflicts to become errors
content <- raw_lines[!str_detect(str_trim(raw_lines), "^-+$")]content <- content[-c(1, 2)]line1 <- content[seq(1, length(content), by =2)]line2 <- content[seq(2, length(content), by =2)]
Dash-only separator lines and the two header rows are removed. What remains is two lines per player, so odd lines hold name, points, and rounds, and even lines hold state and rating.
Each line is split on |. The pre-rating is pulled from text like R: 1794 ->1817, and each round like W 39 gives the opponent’s pair number. Byes and unplayed rounds have no number, so they become NA. The result has one row per player per round.
The tournament had 64 players, mostly from Michigan (55), with 8 from Ontario and 1 from Ohio. Gary Hua, Dakshesh Daruri, and Aditya Bajaj tied for the top score with 6.0 points. Aditya Bajaj stands out because his pre-rating of 1384 was the lowest of the top finishers, yet his opponents averaged 1564, so he beat stronger players than his rating predicted. Gary Hua faced the toughest field overall, with the highest average opponent rating in the tournament at 1605.
23 players played fewer than 7 games because of byes or withdrawals. Their averages only include games actually played, which matches how the assignment calculated Gary Hua’s example.
Conclusions
The hardest part was realizing the round columns store opponent pair numbers, not ratings, so the full player table had to be built before any averages could be calculated. Testing the output against Gary Hua’s known result was a quick way to confirm the parsing worked. A next step would be comparing each player’s actual points to their expected score from the Elo formula, which would show who over or underperformed their rating, like Aditya Bajaj.
AI Citation
Anthropic. (2026). Claude Sonnet 5 [Large language model]. https://claude.ai. Accessed September 2026.