library(tidyverse)Elo Calculations: Chess Tournament
Introduction
In Project 1, I turned a chess tournament cross-table into a clean dataset with each player’s name, state, total points, pre-rating, and average opponent rating. In this project, I use the Elo rating system to estimate how many points each player was expected to score against their opponents, and compare that to what they actually scored.
The Elo expected score for player A against player B is:
\[E_A = \frac{1}{1 + 10^{(R_B - R_A)/400}}\]
Here, \(E_A\) is player A’s expected score, \(R_A\) is player A’s rating, and \(R_B\) is player B’s rating. This formula comes from Arpad Elo’s The Rating of Chessplayers, Past and Present (1978) and is summarized in the Wikipedia article “Elo rating system” (https://en.wikipedia.org/wiki/Elo_rating_system).
It gives a number between 0 and 1 that shows how much of a point the player is expected to earn. For example, two equally rated players each have an expected score of 0.5, and a player rated 200 points higher than their opponent has an expected score of about 0.76. Adding up a player’s expected scores across their games gives their expected tournament score. I will then find the five players who most overperformed and the five who most underperformed.
Planned Approach
First, I will start from the data parsed in Project 1, keeping each player’s round-by-round results and opponent IDs. I will reshape this into one row per game and attach both the player’s and the opponent’s pre-ratings.
Then, for each game actually played, I will calculate the expected score using the Elo formula and convert the result into actual points: win = 1, draw = 0.5, loss = 0. Summing these for each player gives their expected and actual scores, and the difference (actual − expected) shows whether they overperformed or underperformed.
Finally, I will sort players by this difference and report the five largest overperformers and five largest underperformers, showing each player’s name, actual score, expected score, and difference.
Anticipated Challenges
Some rounds are byes, forfeits, or unplayed games with no opponent, so I will compare expected and actual scores using only games that were actually played. Since Elo formulas vary slightly between sources, I will use the standard formula and cite it. Finally, reshaping and joining the data can create duplicate or missing rows, so I will check row counts at each step.
url <- "https://raw.githubusercontent.com/esradogan3/data607-project1/refs/heads/main/tournamentinfo.txt"
chess_raw <- readLines(url, warn = FALSE) # hide the warning while reading lines
chess_raw <- chess_raw[!grepl("^-+\\s*$", chess_raw)] # drop the dashed separator lines
chess_raw <- chess_raw[-(1:2)] # drop the two header lines
row1 <- chess_raw[seq(1, length(chess_raw), by = 2)] # pair num, name, points, rounds
row2 <- chess_raw[seq(2, length(chess_raw), by = 2)] # state, USCF ID / ratings
p1 <- read.table(text = row1, sep = "|", strip.white = TRUE,
quote = "", stringsAsFactors = FALSE)
p2 <- read.table(text = row2, sep = "|", strip.white = TRUE,
quote = "", stringsAsFactors = FALSE)
chess <- data.frame(
id = p1$V1,
name = p1$V2,
state = p2$V1,
points = as.numeric(p1$V3),
pre_rating = as.numeric(sub(".*R:\\s*(\\d+).*", "\\1", p2$V2))
)
nrow(chess)[1] 64
head(chess) id name state points pre_rating
1 1 GARY HUA ON 6.0 1794
2 2 DAKSHESH DARURI MI 6.0 1553
3 3 ADITYA BAJAJ MI 6.0 1384
4 4 PATRICK H SCHILLING MI 5.5 1716
5 5 HANSHI ZUO MI 5.5 1655
6 6 HANSEN SONG OH 5.0 1686
Getting the Results and Opponents
Each round entry, like “W 39”, holds two values: the result letter and the opponent’s ID. In Project 1, I only took the opponent’s ID. This time I also need the result letter, so I take both from columns V4 to V10.
letters_df <- p1[, 4:10] # result letters
opp_df <- p1[, 4:10] # opponent IDs
for (r in 1:7) {
letters_df[[r]] <- substr(trimws(p1[[r + 3]]), 1, 1) # first letter only
opp_df[[r]] <- as.integer(gsub("\\D", "", p1[[r + 3]])) # numbers only
}
head(letters_df) V4 V5 V6 V7 V8 V9 V10
1 W W W W W D D
2 W W L W W W W
3 L W W W W W W
4 W D W W D W D
5 W W D D D W W
6 W D L W D W W
head(opp_df) V4 V5 V6 V7 V8 V9 V10
1 39 21 18 14 7 12 4
2 63 58 4 17 16 20 7
3 8 61 25 21 11 13 12
4 23 28 2 26 5 19 1
5 45 37 12 13 4 14 17
6 34 29 11 35 10 27 21
Calculating Expected and Actual Scores
For each player, I go through their seven rounds. I only use rounds with W (win), L (loss), or D (draw), because byes (B, H), unplayed rounds (U), and forfeits (X) have no real game. For each game, I turn the result into points: win = 1, draw = 0.5, loss = 0, use the Elo formula to find the expected score, and add both to the player’s totals.
actual <- numeric(64)
expected <- numeric(64)
games <- numeric(64)
for (i in 1:64) {
for (r in 1:7) {
letter <- letters_df[i, r]
opp <- opp_df[i, r]
if (letter %in% c("W", "L", "D")) {
# actual points
if (letter == "W") actual[i] <- actual[i] + 1
if (letter == "D") actual[i] <- actual[i] + 0.5
# expected points (Elo formula)
my_rating <- chess$pre_rating[i]
opp_rating <- chess$pre_rating[opp]
expected[i] <- expected[i] + 1 / (1 + 10^((opp_rating - my_rating) / 400))
games[i] <- games[i] + 1
}
}
}Now I put everything in one table and calculate the difference between the actual and expected scores.
results <- data.frame(
name = chess$name,
pre_rating = chess$pre_rating,
games_played = games,
actual = actual,
expected = round(expected, 2),
difference = round(actual - expected, 2)
)
head(results) name pre_rating games_played actual expected difference
1 GARY HUA 1794 7 6.0 5.16 0.84
2 DAKSHESH DARURI 1553 7 6.0 3.78 2.22
3 ADITYA BAJAJ 1384 7 6.0 1.95 4.05
4 PATRICK H SCHILLING 1716 7 5.5 4.74 0.76
5 HANSHI ZUO 1655 7 5.5 4.38 1.12
6 HANSEN SONG 1686 7 5.0 4.94 0.06
Checking if everything looks good: In every game, the two players’ points add up to 1, and so do their expected scores. So the total actual score and the total expected score should both equal the number of games played in the tournament.
sum(games) / 2 # number of games (each game is counted twice)[1] 204
sum(actual)[1] 204
sum(expected)[1] 204
All three numbers are the same, so the calculations look correct.
Results
Top 5 Overperformers
results %>%
arrange(desc(difference)) %>%
head(5) name pre_rating games_played actual expected difference
1 ADITYA BAJAJ 1384 7 6.0 1.95 4.05
2 ZACHARY JAMES HOUGHTON 1220 7 4.5 1.37 3.13
3 ANVIT RAO 1365 7 5.0 1.94 3.06
4 JACOB ALEXANDER LAVALLEY 377 7 3.0 0.04 2.96
5 STEFANO LEE 1411 7 5.0 2.29 2.71
Top 5 Underperformers
results %>%
arrange(difference) %>%
head(5) name pre_rating games_played actual expected difference
1 LOREN SCHWIEBERT 1745 7 3.5 6.28 -2.78
2 GEORGE AVERY JONES 1522 7 3.5 6.02 -2.52
3 LARRY HODGE 1270 6 1.0 3.40 -2.40
4 JARED GE 1332 7 3.0 5.01 -2.01
5 RISHI SHETTY 1494 7 3.5 5.09 -1.59
Charts
The first chart shows the ten players from the tables above. Green bars are overperformers and red bars are underperformers.
top_bottom <- rbind(head(arrange(results, desc(difference)), 5),
head(arrange(results, difference), 5))
ggplot(top_bottom, aes(x = reorder(name, difference), y = difference,
fill = difference > 0)) +
geom_col(show.legend = FALSE) +
scale_fill_manual(values = c("firebrick", "seagreen")) +
coord_flip() +
labs(title = "Biggest Over- and Underperformers",
x = NULL, y = "Actual score - expected score")The second chart shows all 64 players. Each dot is one player. The dashed line shows where the actual score equals the expected score. Players above the line did better than expected, and players below the line did worse.
ggplot(results, aes(x = expected, y = actual)) +
geom_point() +
geom_abline(slope = 1, intercept = 0, linetype = "dashed") +
labs(title = "Expected vs. Actual Score",
x = "Expected score", y = "Actual score")Conclusion
All five overperformers had pre-ratings close to or below the middle of the field (the median pre-rating is 1407). This makes sense, because Elo expects lower-rated players to lose against stronger opponents, so every win or draw against a stronger player raises their difference a lot. Aditya Bajaj stands out the most: he was expected to score about 2 points but scored 6. Jacob Alexander Lavalley, rated only 377, had an expected score close to 0 and still scored 3 points. His rating, like Zachary James Houghton’s, is provisional, which means it is based on only a few games. A provisional rating may not show how strong a player really is, especially for young players who improve quickly.
The underperformers were more mixed. Loren Schwiebert, one of the higher-rated players, was expected to score about 6.3 points but scored only 3.5. On the other hand, Larry Hodge and Jared Ge were rated below the median and still scored about 2 points less than expected. Overall, the results show that pre-tournament ratings are a useful guide, but they do not always predict how a player will perform in a single tournament.
Using LLM
I used Claude (Anthropic, 2026) for this project. Building on my Project 1 code, Claude wrote most of the code for calculating the Elo scores and helped me edit my explanations. I went through each step to understand how it works and ran it myself.
References
Anthropic. (2026). Claude (Opus 5.5 version) [Large language model]. https://claude.ai
Elo, A. E. (1978). The rating of chessplayers, past and present. Arco.
Wikipedia contributors. (n.d.). Elo rating system. In Wikipedia. https://en.wikipedia.org/wiki/Elo_rating_system