Zaina Hassan
GitHub
Repository
For this assignment, I will use the chess tournament data from Project 1 to compare each player’s actual tournament performance with their expected performance based on the Elo rating system. The dataset already contains each player’s pre-tournament rating, actual total points, and the opponents they faced during the tournament.
I will use an Elo expected-score formula to calculate the expected result of each individual matchup based on the difference between the player’s pre-tournament rating and their opponent’s pre-tournament rating. Each player’s expected scores from their individual games will then be added together to calculate their total expected tournament score.
Next, I will compare each player’s expected score with their actual total points by calculating the difference between the two. A positive difference will indicate that a player scored more points than expected, while a negative difference will indicate that a player scored fewer points than expected. Finally, I will sort these differences to identify the five players who most overperformed and the five players who most underperformed relative to their expected scores.
One anticipated complication is handling rounds in which a player did not face an opponent, such as byes or unplayed rounds. These rounds will need to be excluded from the Elo matchup calculations so that expected scores are based only on games against actual opponents. Another consideration is ensuring that every opponent number is correctly matched to that opponent’s pre-tournament rating before calculating the expected scores.
For this assignment, I will use the same tournament text file from Project 1. The following code extracts each player’s name, total points, pre-tournament rating, and opponent numbers from the original tournament data.
chess_raw <- readLines("tournamentinfo.txt")
## Warning in readLines("tournamentinfo.txt"): incomplete final line found on
## 'tournamentinfo.txt'
# Extract player and rating records
player_lines <- chess_raw[seq(5, length(chess_raw), by = 3)]
rating_lines <- chess_raw[seq(6, length(chess_raw), by = 3)]
# Extract player names and total points
player_split <- strsplit(player_lines, "\\|")
player_name <- trimws(sapply(player_split, `[`, 2))
total_points <- as.numeric(trimws(sapply(player_split, `[`, 3)))
# Extract states and pre-tournament ratings
rating_split <- strsplit(rating_lines, "\\|")
player_state <- trimws(sapply(rating_split, `[`, 1))
rating_text <- sapply(rating_split, `[`, 2)
pre_rating <- as.numeric(
sub(".*R:\\s*([0-9]+).*", "\\1", rating_text)
)
# Create player data frame
players <- data.frame(
player_number = 1:length(player_name),
player_name = player_name,
state = player_state,
total_points = total_points,
pre_rating = pre_rating
)
head(players)
## player_number player_name state total_points pre_rating
## 1 1 GARY HUA ON 6.0 1794
## 2 2 DAKSHESH DARURI MI 6.0 1553
## 3 3 ADITYA BAJAJ MI 6.0 1384
## 4 4 PATRICK H SCHILLING MI 5.5 1716
## 5 5 HANSHI ZUO MI 5.5 1655
## 6 6 HANSEN SONG OH 5.0 1686
Next, I will extract the opponent numbers from each player’s seven tournament rounds. Rounds that do not contain an opponent number will be stored as NA
# Extract the seven tournament rounds
rounds <- t(sapply(player_split, function(x) x[4:10]))
# Keep only opponent numbers
opponents <- apply(rounds, c(1, 2), function(x) {
number <- gsub("[^0-9]", "", x)
ifelse(number == "", NA, as.numeric(number))
})
head(opponents)
## [,1] [,2] [,3] [,4] [,5] [,6] [,7]
## [1,] 39 21 18 14 7 12 4
## [2,] 63 58 4 17 16 20 7
## [3,] 8 61 25 21 11 13 12
## [4,] 23 28 2 26 5 19 1
## [5,] 45 37 12 13 4 14 17
## [6,] 34 29 11 35 10 27 21
I will calculate the expected score for every game based on the difference between the player’s rating and their opponent’s rating.
# Create a matrix to store expected scores
expected_scores <- matrix(
NA,
nrow = nrow(opponents),
ncol = ncol(opponents)
)
# Calculate expected score for each matchup
for (i in 1:nrow(opponents)) {
for (j in 1:ncol(opponents)) {
opponent_number <- opponents[i, j]
if (!is.na(opponent_number)) {
player_rating <- players$pre_rating[i]
opponent_rating <- players$pre_rating[opponent_number]
expected_scores[i, j] <- 1 / (
1 + 10^((opponent_rating - player_rating) / 400)
)
}
}
}
head(expected_scores)
## [,1] [,2] [,3] [,4] [,5] [,6] [,7]
## [1,] 0.8870357 0.7907981 0.7533861 0.7425356 0.6973451 0.6800707 0.6104024
## [2,] 0.8980683 0.9749402 0.2812432 0.3923389 0.4271277 0.4398499 0.3652567
## [3,] 0.1855164 0.9219774 0.1112454 0.2630052 0.1314590 0.1647472 0.1671373
## [4,] 0.8841194 0.7690759 0.7187568 0.6875382 0.5868950 0.7057814 0.3895976
## [5,] 0.9150891 0.9798780 0.4884891 0.4841750 0.4131050 0.5644005 0.5373473
## [6,] 0.8391753 0.6185841 0.4626527 0.8065275 0.8638715 0.6838163 0.6699690
expected_total <- rowSums(
expected_scores,
na.rm = TRUE
)
head(expected_total)
## [1] 5.161574 3.778825 1.945088 4.741764 4.382484 4.944596
performance <- data.frame(
Player_Name = players$player_name,
Pre_Rating = players$pre_rating,
Actual_Score = players$total_points,
Expected_Score = expected_total
)
performance$Difference <-
performance$Actual_Score - performance$Expected_Score
# Round values for easier interpretation
performance$Expected_Score <- round(
performance$Expected_Score, 2
)
performance$Difference <- round(
performance$Difference, 2
)
head(performance)
## Player_Name Pre_Rating Actual_Score Expected_Score Difference
## 1 GARY HUA 1794 6.0 5.16 0.84
## 2 DAKSHESH DARURI 1553 6.0 3.78 2.22
## 3 ADITYA BAJAJ 1384 6.0 1.95 4.05
## 4 PATRICK H SCHILLING 1716 5.5 4.74 0.76
## 5 HANSHI ZUO 1655 5.5 4.38 1.12
## 6 HANSEN SONG 1686 5.0 4.94 0.06
underperformed <- performance[
order(performance$Difference),
]
top_5_underperformed <- head(underperformed, 5)
top_5_underperformed
## Player_Name Pre_Rating Actual_Score Expected_Score Difference
## 25 LOREN SCHWIEBERT 1745 3.5 6.28 -2.78
## 30 GEORGE AVERY JONES 1522 3.5 6.02 -2.52
## 42 JARED GE 1332 3.0 5.01 -2.01
## 31 RISHI SHETTY 1494 3.5 5.09 -1.59
## 35 JOSHUA DAVID LEE 1438 3.5 4.96 -1.46
overperformed <- performance[
order(performance$Difference, decreasing = TRUE),
]
top_5_overperformed <- head(overperformed, 5)
top_5_overperformed
## Player_Name Pre_Rating Actual_Score Expected_Score Difference
## 3 ADITYA BAJAJ 1384 6.0 1.95 4.05
## 15 ZACHARY JAMES HOUGHTON 1220 4.5 1.37 3.13
## 10 ANVIT RAO 1365 5.0 1.94 3.06
## 46 JACOB ALEXANDER LAVALLEY 377 3.0 0.04 2.96
## 37 AMIYATOSH PWNANANDAM 980 3.5 0.77 2.73
# Check that all 64 players received an expected score
nrow(performance)
## [1] 64
sum(is.na(performance$Expected_Score))
## [1] 0
Using the Elo expected score formula, I calculated an expected score for each individual matchup based on the difference between the player’s pre-tournament rating and their opponent’s pre-tournament rating. These individual expected scores were then added together to determine each player’s total expected tournament score.
The five players who most overperformed were Aditya Bajaj, Zachary James Houghton, Anvit Rao, Jacob Alexander Lavalley, and Amiyatosh Pwnanandam. Aditya Bajaj had the largest positive difference, earning 6.0 actual points compared with an expected score of 1.95, resulting in a difference of +4.05 points.
The five players who most underperformed were Loren Schwiebert, George Avery Jones, Jared Ge, Rishi Shetty, and Joshua David Lee. Loren Schwiebert had the largest negative difference, earning 3.5 actual points compared with an expected score of 6.28, resulting in a difference of -2.78 points.
The results demonstrate how Elo ratings can be used to estimate a player’s expected performance based on the ratings of the opponents they faced. By comparing each player’s actual tournament score with their expected score, I was able to identify players whose performance was substantially above or below expectations.
The analysis showed that pre-tournament rating differences can provide an expected score for each matchup, while the difference between expected and actual scores provides a simple way to evaluate tournament performance relative to those expectations.