Approach

For this assignment, I will use the chess tournament data from Project 1 to compare each player’s actual tournament performance with their expected performance based on the Elo rating system. The dataset already contains each player’s pre-tournament rating, actual total points, and the opponents they faced during the tournament.

I will use an Elo expected-score formula to calculate the expected result of each individual matchup based on the difference between the player’s pre-tournament rating and their opponent’s pre-tournament rating. Each player’s expected scores from their individual games will then be added together to calculate their total expected tournament score.

Next, I will compare each player’s expected score with their actual total points by calculating the difference between the two. A positive difference will indicate that a player scored more points than expected, while a negative difference will indicate that a player scored fewer points than expected. Finally, I will sort these differences to identify the five players who most overperformed and the five players who most underperformed relative to their expected scores.

One anticipated complication is handling rounds in which a player did not face an opponent, such as byes or unplayed rounds. These rounds will need to be excluded from the Elo matchup calculations so that expected scores are based only on games against actual opponents. Another consideration is ensuring that every opponent number is correctly matched to that opponent’s pre-tournament rating before calculating the expected scores.

Code Base

Import and Prepare the Tournament Data

For this assignment, I will use the same tournament text file from Project 1. The following code extracts each player’s name, total points, pre-tournament rating, and opponent numbers from the original tournament data.

chess_raw <- readLines("tournamentinfo.txt")
## Warning in readLines("tournamentinfo.txt"): incomplete final line found on
## 'tournamentinfo.txt'
# Extract player and rating records
player_lines <- chess_raw[seq(5, length(chess_raw), by = 3)]
rating_lines <- chess_raw[seq(6, length(chess_raw), by = 3)]

# Extract player names and total points
player_split <- strsplit(player_lines, "\\|")
player_name <- trimws(sapply(player_split, `[`, 2))
total_points <- as.numeric(trimws(sapply(player_split, `[`, 3)))

# Extract states and pre-tournament ratings
rating_split <- strsplit(rating_lines, "\\|")
player_state <- trimws(sapply(rating_split, `[`, 1))
rating_text <- sapply(rating_split, `[`, 2)

pre_rating <- as.numeric(
  sub(".*R:\\s*([0-9]+).*", "\\1", rating_text)
)

# Create player data frame
players <- data.frame(
  player_number = 1:length(player_name),
  player_name = player_name,
  state = player_state,
  total_points = total_points,
  pre_rating = pre_rating
)

head(players)
##   player_number         player_name state total_points pre_rating
## 1             1            GARY HUA    ON          6.0       1794
## 2             2     DAKSHESH DARURI    MI          6.0       1553
## 3             3        ADITYA BAJAJ    MI          6.0       1384
## 4             4 PATRICK H SCHILLING    MI          5.5       1716
## 5             5          HANSHI ZUO    MI          5.5       1655
## 6             6         HANSEN SONG    OH          5.0       1686

Extract Opponent Numbers

Next, I will extract the opponent numbers from each player’s seven tournament rounds. Rounds that do not contain an opponent number will be stored as NA

# Extract the seven tournament rounds
rounds <- t(sapply(player_split, function(x) x[4:10]))

# Keep only opponent numbers
opponents <- apply(rounds, c(1, 2), function(x) {
  number <- gsub("[^0-9]", "", x)
  ifelse(number == "", NA, as.numeric(number))
})

head(opponents)
##      [,1] [,2] [,3] [,4] [,5] [,6] [,7]
## [1,]   39   21   18   14    7   12    4
## [2,]   63   58    4   17   16   20    7
## [3,]    8   61   25   21   11   13   12
## [4,]   23   28    2   26    5   19    1
## [5,]   45   37   12   13    4   14   17
## [6,]   34   29   11   35   10   27   21

Calculate Expected Scores

I will calculate the expected score for every game based on the difference between the player’s rating and their opponent’s rating.

# Create a matrix to store expected scores
expected_scores <- matrix(
  NA,
  nrow = nrow(opponents),
  ncol = ncol(opponents)
)

# Calculate expected score for each matchup
for (i in 1:nrow(opponents)) {
  for (j in 1:ncol(opponents)) {
    
    opponent_number <- opponents[i, j]
    
    if (!is.na(opponent_number)) {
      
      player_rating <- players$pre_rating[i]
      opponent_rating <- players$pre_rating[opponent_number]
      
      expected_scores[i, j] <- 1 / (
        1 + 10^((opponent_rating - player_rating) / 400)
      )
    }
  }
}

head(expected_scores)
##           [,1]      [,2]      [,3]      [,4]      [,5]      [,6]      [,7]
## [1,] 0.8870357 0.7907981 0.7533861 0.7425356 0.6973451 0.6800707 0.6104024
## [2,] 0.8980683 0.9749402 0.2812432 0.3923389 0.4271277 0.4398499 0.3652567
## [3,] 0.1855164 0.9219774 0.1112454 0.2630052 0.1314590 0.1647472 0.1671373
## [4,] 0.8841194 0.7690759 0.7187568 0.6875382 0.5868950 0.7057814 0.3895976
## [5,] 0.9150891 0.9798780 0.4884891 0.4841750 0.4131050 0.5644005 0.5373473
## [6,] 0.8391753 0.6185841 0.4626527 0.8065275 0.8638715 0.6838163 0.6699690

Calculate Total Expected Score

expected_total <- rowSums(
  expected_scores,
  na.rm = TRUE
)

head(expected_total)
## [1] 5.161574 3.778825 1.945088 4.741764 4.382484 4.944596

Compare Expected and Actual Scores

performance <- data.frame(
  Player_Name = players$player_name,
  Pre_Rating = players$pre_rating,
  Actual_Score = players$total_points,
  Expected_Score = expected_total
)

performance$Difference <- 
  performance$Actual_Score - performance$Expected_Score

# Round values for easier interpretation
performance$Expected_Score <- round(
  performance$Expected_Score, 2
)

performance$Difference <- round(
  performance$Difference, 2
)

head(performance)
##           Player_Name Pre_Rating Actual_Score Expected_Score Difference
## 1            GARY HUA       1794          6.0           5.16       0.84
## 2     DAKSHESH DARURI       1553          6.0           3.78       2.22
## 3        ADITYA BAJAJ       1384          6.0           1.95       4.05
## 4 PATRICK H SCHILLING       1716          5.5           4.74       0.76
## 5          HANSHI ZUO       1655          5.5           4.38       1.12
## 6         HANSEN SONG       1686          5.0           4.94       0.06

Underpreformers

underperformed <- performance[
  order(performance$Difference),
]

top_5_underperformed <- head(underperformed, 5)

top_5_underperformed
##           Player_Name Pre_Rating Actual_Score Expected_Score Difference
## 25   LOREN SCHWIEBERT       1745          3.5           6.28      -2.78
## 30 GEORGE AVERY JONES       1522          3.5           6.02      -2.52
## 42           JARED GE       1332          3.0           5.01      -2.01
## 31       RISHI SHETTY       1494          3.5           5.09      -1.59
## 35   JOSHUA DAVID LEE       1438          3.5           4.96      -1.46

Overpreformers

overperformed <- performance[
  order(performance$Difference, decreasing = TRUE),
]

top_5_overperformed <- head(overperformed, 5)

top_5_overperformed
##                 Player_Name Pre_Rating Actual_Score Expected_Score Difference
## 3              ADITYA BAJAJ       1384          6.0           1.95       4.05
## 15   ZACHARY JAMES HOUGHTON       1220          4.5           1.37       3.13
## 10                ANVIT RAO       1365          5.0           1.94       3.06
## 46 JACOB ALEXANDER LAVALLEY        377          3.0           0.04       2.96
## 37     AMIYATOSH PWNANANDAM        980          3.5           0.77       2.73
# Check that all 64 players received an expected score
nrow(performance)
## [1] 64
sum(is.na(performance$Expected_Score))
## [1] 0

Results

Using the Elo expected score formula, I calculated an expected score for each individual matchup based on the difference between the player’s pre-tournament rating and their opponent’s pre-tournament rating. These individual expected scores were then added together to determine each player’s total expected tournament score.

The five players who most overperformed were Aditya Bajaj, Zachary James Houghton, Anvit Rao, Jacob Alexander Lavalley, and Amiyatosh Pwnanandam. Aditya Bajaj had the largest positive difference, earning 6.0 actual points compared with an expected score of 1.95, resulting in a difference of +4.05 points.

The five players who most underperformed were Loren Schwiebert, George Avery Jones, Jared Ge, Rishi Shetty, and Joshua David Lee. Loren Schwiebert had the largest negative difference, earning 3.5 actual points compared with an expected score of 6.28, resulting in a difference of -2.78 points.

Conclusion

The results demonstrate how Elo ratings can be used to estimate a player’s expected performance based on the ratings of the opponents they faced. By comparing each player’s actual tournament score with their expected score, I was able to identify players whose performance was substantially above or below expectations.

The analysis showed that pre-tournament rating differences can provide an expected score for each matchup, while the difference between expected and actual scores provides a simple way to evaluate tournament performance relative to those expectations.