For this project, I worked with a chess tournament cross-table that was not in a tidy format. My goal was to convert the original data into a clean dataset that could be used for analysis.
The final dataset has one row for each player and includes these columns:
• Player Name
• Player State
• Total Points
•
Pre-Tournament Rating
• Average Pre-Tournament Rating of Opponents
I used the tournament crosstable provided for the assignment. I think this dataset was complex enough to demonstrate the data cleaning and parsing skills required for this project. Each player is represented across multiple lines, and the round information includes opponent numbers and game results.
The main goal was to extract the player information from the original crosstable and determine the average pre-tournament rating of each player’s opponents.
The most difficult part of the project was not calculating an average. The main challenge was correctly identifying each opponent and connecting that opponent to the correct pre-tournament rating.
I also needed to handle situations where a player did not play a particular round, such as a bye or an unplayed game. Finally, I wanted the entire process to be reproducible so that the same code could be run again and produce the same results.
I loaded the tournament information directly from my GitHub repository using the raw text file URL. I used readLines() to read the file into R and then checked the number of lines and the first 10 rows to make sure the data was loaded correctly.
Identifying Player Information Lines
I identified the two types of lines containing player information in the raw tournament data. I then separated these lines so that the player details could be extracted and processed correctly.
I extracted the player’s pair number, name, and total points from the first player information line. I stored these values in a tidy data frame for use in the next steps of the analysis.
1. Data Ingestion
I read the tournament text file from a public URL so that the project did not depend on a file stored only on my computer.
I read the file line by line and preserved the original formatting. This allowed me to work with the structure of the original cross-table.
2. Reconstructing Player Records
Each player is represented by two lines in the cross-table.
The first line contains information such as:
• Pair number
• Player name
• Total points
• Results for
each round
The second line contains information such as:
• State
• USCF ID
• Pre-tournament and post-tournament
ratings
I identified the two types of player information separately and combined them in their original order after verifying that 64 records of each type were found.
Regular expressions and string manipulation were used to extract the information I needed.
3. Extracting the Variables
For each player, I extracted:
• pair_number
• player_name
• state
• total_points
•
pre_rating
• Opponent pair numbers for each round
The round results may look like:
• W 39
• L 21
• D 12
The letter shows the result of the game, while the number identifies the opponent. I only needed the opponent number for calculating the average opponent rating.
Calculating the Average Opponent Rating
To calculate the average opponent rating, I created a lookup table that connected each player’s pair number with their pre-tournament rating.
For example:
pair_number → pre_rating
Then, for each player, I collected the opponent pair numbers from all rounds and used the lookup table to obtain the corresponding pre-tournament ratings.
I counted a round as a game when it contained a result indicator
(W, L, or D) followed by an opponent number. I did not count entries
such as:
• H – half-point bye
• B – bye
• U – unplayed
• X –
forfeit or win by absence
These entries do not provide an actual opponent whose pre-tournament
rating should be included in the average.
After identifying the
valid opponents, I used their pair numbers to look up their
pre-tournament ratings and calculate the mean.
This means that a player who played all seven rounds had the average
calculated from all of their opponents, while a player who played fewer
rounds had the average calculated only from the opponents they actually
played.
I performed manual checks to make sure the program produced the correct results.
Test Case 1: Player Who Played All Games
I selected a player who played all rounds and manually identified their opponents. I identified the opponents, found the pre-tournament rating of each opponent, and calculated the average by hand. I then compared this result with the value produced by my program.
Test Case 2: Player Who Played Fewer Games
I selected a player who did not play every round. I identified the rounds that were actually played and excluded byes or other non-game entries. I then manually calculated the average opponent rating and compared it with the program’s result.
These checks helped confirm that the opponent numbers were parsed correctly and that the correct denominator was used when calculating the average.
I stored all of the code in the Quarto file so that the complete process can be reproduced from beginning to end. The data was read from a public URL rather than from a local file. There was no manual editing of the data.
I also included checks to make sure that:
• Exactly 64 players
were processed.
• Opponent pair numbers were between 1 and 64.
• Opponent pair numbers successfully matched a player in the rating
lookup table.
• No missing pre-tournament ratings were created
during the lookup process.
• The final dataset contained one row
per player.
The main challenges in processing the data included:
• Correctly combining the two lines for each player.
• Dealing
with inconsistent spacing in the original text.
• Correctly
extracting opponent numbers from the round results.
•
Distinguishing actual games from byes and unplayed rounds.
• Making
sure that opponent numbers are connected to the correct player ratings.
• Avoiding parsing errors that could shift information from one
player to another.
For Player 1 (Pair #1), the opponents played were:
39, 21, 18, 14, 7, 12, 4
Their corresponding pre-ratings are:
1436, 1563, 1600, 1610, 1649, 1663, 1716
Hand calculation:
Sum = 1436 + 1563 + 1600 + 1610 + 1649 + 1663 + 1716
Sum = 11237
Number of games played = 7
Average = 11237 / 7 = 1605.2857
Rounded to nearest integer = 1605
The program output for Player 1 is 1605, confirming correct extraction, mapping, and denominator logic.
For Player 16 (Pair #16), the valid games were against:
10, 15, 39, 2, 36
The entries “H” and “U” are excluded because they do not represent valid opponents.
Their corresponding pre-ratings are:
1365, 1220, 1436, 1553, 1355
Hand calculation:
Sum = 1365 + 1220 + 1436 + 1553 + 1355
Sum = 6929
Number of valid games played = 5
Average = 6929 / 5 = 1385.8
Rounded to nearest integer = 1386
The program output (1386) matches the hand calculation, confirming correct handling of partial participation and exclusion of non-games.
The final dataset contains exactly one row for each player. The columns are:
• Player_Name
• State
• Total_Points
• Pre_Rating
•
Average_Opponent_Pre_Rating
The complete final dataset was exported as a CSV file.
The objective of this assignment is to compute each player’s expected tournament score using the Elo rating model and compare it with their actual score from the Project 1 tournament.
We will then identify:
This analysis builds directly on the structured dataset created in Project 1.
The dataset contains exactly one row per player with the following variables:
This dataset was exported as a CSV file and uploaded to GitHub to ensure reproducibility. The file will be loaded into R using the GitHub raw link.
The Elo rating system models the probability that Player A defeats Player B using a logistic function based on rating differences.
The expected score formula used in this assignment is:
\[ E_A = \frac{1}{1 + 10^{(R_B - R_A)/400}} \]
Where:
This formulation implies that a 400-point rating difference corresponds to approximately a 10-to-1 expected win ratio.
Sources:
- Glickman, M. E. A Comprehensive Guide to Chess Ratings: https://www.glicko.net/research/acjpaper.pdf
- Elo rating system (Wikipedia): https://en.wikipedia.org/wiki/Elo_rating_system
- The Elo Rating System for Chess and Beyond (YouTube, 2019):
https://www.youtube.com/watch?v=AsYfbmp0To0
In the standard Elo system, expected score is computed separately for each game using each opponent’s rating, and then summed across games.
However, the Project 1 dataset retains only the average pre-tournament rating of opponents, rather than the full match level opponent list. Therefore, this analysis approximates expected tournament score by applying the Elo expected score formula using the player’s rating versus their average opponent rating, and then scaling by the number of rounds.
The expected per-game score is computed as:
\[ E = \frac{1}{1 + 10^{(\overline{R_{opp}} - R_{player})/400}} \]
The expected tournament score is then:
\[ \text{Expected Tournament Score} = n \times E \]
Where:
This provides a transparent approximation consistent with the Elo logistic framework given the available Project 1 output.
The analysis will proceed as follows:
Project1_Chess_Summary.csv directly from the
GitHub raw link using read.csv() to ensure
reproducibility.\[ \text{Performance Difference} = \text{Actual Score} - \text{Expected Score} \]
The final output will include:
Players will be ranked to identify:
This approach builds directly upon the structured output from Project 1 and applies the Elo rating model to evaluate tournament performance relative to expectation.
This analysis uses the cleaned Project 1 summary dataset (one row per player) containing:
Player Name
Player State
Total Points (Actual Score)
Pre-Tournament Rating
Average Pre-Tournament Rating of Opponents
chess <- read.csv(
"https://raw.githubusercontent.com/BIKASHBHOWMIK15/Data-607/main/Project-1/Project1_Chess_Data_Summary.csv"
)
chess <- as_tibble(chess)
glimpse(chess)Rows: 64
Columns: 5
$ Player_Name <chr> "Gary Hua", "Dakshesh Daruri", "Aditya Baj…
$ State <chr> "ON", "MI", "MI", "MI", "MI", "OH", "MI", …
$ Total_Points <dbl> 6.0, 6.0, 6.0, 5.5, 5.5, 5.0, 5.0, 5.0, 5.…
$ Pre_Rating <int> 1794, 1553, 1384, 1716, 1655, 1686, 1649, …
$ Average_Opponent_Pre_Rating <int> 1605, 1469, 1564, 1574, 1501, 1519, 1372, …
[1] "Player_Name" "State"
[3] "Total_Points" "Pre_Rating"
[5] "Average_Opponent_Pre_Rating"
We use the standard Elo expected score formula:
\[ E_A = \frac{1}{1 + 10^{(R_B - R_A)/400}} \]
This follows the standard Elo logistic model (see Approach section for detailed explanation and sources).
Because the Project 1 output retains only the average opponent pre-rating (not the full opponent list), I approximated the expected tournament score by computing expected per-game score versus the average opponent rating, then multiplying by the number of rounds played.
In this tournament there are 7 rounds, so we use
n_rounds = 7.
n_rounds <- 7
elo_expected <- function(r_player, r_opp_avg) {
1 / (1 + 10^((r_opp_avg - r_player) / 400))
}
results <- chess %>%
mutate(
expected_per_game = elo_expected(Pre_Rating, Average_Opponent_Pre_Rating),
expected_score = n_rounds * expected_per_game,
actual_score = Total_Points,
performance_diff = actual_score - expected_score
)
results %>%
select(Player_Name, State, Pre_Rating, Average_Opponent_Pre_Rating, actual_score, expected_score, performance_diff) %>%
arrange(desc(performance_diff)) %>%
head(10) %>%
kable(digits = 3, caption = "Top 10 players by performance difference (Actual - Expected)")| Player_Name | State | Pre_Rating | Average_Opponent_Pre_Rating | actual_score | expected_score | performance_diff |
|---|---|---|---|---|---|---|
| Aditya Bajaj | MI | 1384 | 1564 | 6.0 | 1.833 | 4.167 |
| Zachary James Houghton | MI | 1220 | 1484 | 4.5 | 1.257 | 3.243 |
| Anvit Rao | MI | 1365 | 1554 | 5.0 | 1.764 | 3.236 |
| Jacob Alexander Lavalley | MI | 377 | 1358 | 3.0 | 0.025 | 2.975 |
| Amiyatosh Pwnanandam | MI | 980 | 1385 | 3.5 | 0.620 | 2.880 |
| Stefano Lee | ON | 1411 | 1523 | 5.0 | 2.409 | 2.591 |
| Ethan Guo | MI | 935 | 1495 | 2.5 | 0.268 | 2.232 |
| Michael R Aldrich | MI | 1229 | 1357 | 4.0 | 2.266 | 1.734 |
| Dakshesh Daruri | MI | 1553 | 1469 | 6.0 | 4.330 | 1.670 |
| Tejas Ayyagari | MI | 1011 | 1356 | 2.5 | 0.845 | 1.655 |
We rank players by: \[
Performance Difference=Actual Score−Expected Score
\]
Positive values indicate overperformance
Negative values indicate underperformance
top_5_over <- results %>%
arrange(desc(performance_diff)) %>%
select(Player_Name, State, Pre_Rating, actual_score, expected_score, performance_diff) %>%
slice(1:5)
kable(top_5_over, digits = 3, caption = "Top 5 Overperformers (Actual - Expected)")| Player_Name | State | Pre_Rating | actual_score | expected_score | performance_diff |
|---|---|---|---|---|---|
| Aditya Bajaj | MI | 1384 | 6.0 | 1.833 | 4.167 |
| Zachary James Houghton | MI | 1220 | 4.5 | 1.257 | 3.243 |
| Anvit Rao | MI | 1365 | 5.0 | 1.764 | 3.236 |
| Jacob Alexander Lavalley | MI | 377 | 3.0 | 0.025 | 2.975 |
| Amiyatosh Pwnanandam | MI | 980 | 3.5 | 0.620 | 2.880 |
top_5_under <- results %>%
arrange(performance_diff) %>%
select(Player_Name, State, Pre_Rating, actual_score, expected_score, performance_diff) %>%
slice(1:5)
kable(top_5_under, digits = 3, caption = "Top 5 Underperformers (Actual - Expected)")| Player_Name | State | Pre_Rating | actual_score | expected_score | performance_diff |
|---|---|---|---|---|---|
| Ashwin Balaji | MI | 1530 | 1.0 | 6.151 | -5.151 |
| Loren Schwiebert | MI | 1745 | 3.5 | 6.301 | -2.801 |
| George Avery Jones | ON | 1522 | 3.5 | 6.286 | -2.786 |
| Gaurav Gidwani | MI | 1552 | 3.5 | 6.089 | -2.589 |
| Chiedozie Okorie | MI | 1602 | 3.5 | 5.880 | -2.380 |
ggplot(results, aes(x = performance_diff)) +
geom_histogram(bins = 15) +
labs(
title = "Distribution of Performance Difference (Actual - Expected)",
x = "Performance Difference",
y = "Number of Players"
)Players with the largest positive performance differences significantly exceeded their expected scores based on pre-tournament ratings and average opponent strength. In several cases, lower-rated players substantially outperformed rating-based expectations.
Conversely, the largest negative differences indicate higher-rated players who scored fewer points than predicted by the Elo model.
Because expected scores were calculated using average opponent ratings (rather than game-by-game calculations), these results represent an approximation of true Elo expectations. However, the ranking clearly highlights players whose tournament performance diverged most from rating-based predictions.