Elo Calculations: Chess Tournament

Author

Esra Dogan

Introduction

In Project 1, I turned a chess tournament cross-table into a clean dataset with each player’s name, state, total points, pre-rating, and average opponent rating. In this project, I use the Elo rating system to estimate how many points each player was expected to score against their opponents, and compare that to what they actually scored.

The Elo expected score for player A against player B is:

\[E_A = \frac{1}{1 + 10^{(R_B - R_A)/400}}\]

Here, \(E_A\) is player A’s expected score, \(R_A\) is player A’s rating, and \(R_B\) is player B’s rating. This formula comes from Arpad Elo’s The Rating of Chessplayers, Past and Present (1978) and is summarized in the Wikipedia article “Elo rating system” (https://en.wikipedia.org/wiki/Elo_rating_system).

It gives a number between 0 and 1 that shows how much of a point the player is expected to earn. For example, two equally rated players each have an expected score of 0.5, and a player rated 200 points higher than their opponent has an expected score of about 0.76. Adding up a player’s expected scores across their games gives their expected tournament score. I will then find the five players who most overperformed and the five who most underperformed.

Planned Approach

First, I will start from the data parsed in Project 1, keeping each player’s round-by-round results and opponent IDs. I will reshape this into one row per game and attach both the player’s and the opponent’s pre-ratings.

Then, for each game actually played, I will calculate the expected score using the Elo formula and convert the result into actual points: win = 1, draw = 0.5, loss = 0. Summing these for each player gives their expected and actual scores, and the difference (actual − expected) shows whether they overperformed or underperformed.

Finally, I will sort players by this difference and report the five largest overperformers and five largest underperformers, showing each player’s name, actual score, expected score, and difference.

Anticipated Challenges

Some rounds are byes, forfeits, or unplayed games with no opponent, so I will compare expected and actual scores using only games that were actually played. Since Elo formulas vary slightly between sources, I will use the standard formula and cite it. Finally, reshaping and joining the data can create duplicate or missing rows, so I will check row counts at each step.

library(tidyverse)
url <- "https://raw.githubusercontent.com/esradogan3/data607-project1/refs/heads/main/tournamentinfo.txt"

chess_raw <- readLines(url, warn = FALSE) # hide the warning while reading lines
chess_raw <- chess_raw[!grepl("^-+\\s*$", chess_raw)] # drop the dashed separator lines
chess_raw <- chess_raw[-(1:2)] # drop the two header lines

row1 <- chess_raw[seq(1, length(chess_raw), by = 2)] # pair num, name, points, rounds
row2 <- chess_raw[seq(2, length(chess_raw), by = 2)] # state, USCF ID / ratings

p1 <- read.table(text = row1, sep = "|", strip.white = TRUE,
                 quote = "", stringsAsFactors = FALSE)
p2 <- read.table(text = row2, sep = "|", strip.white = TRUE,
                 quote = "", stringsAsFactors = FALSE)

chess <- data.frame(
  id         = p1$V1,
  name       = p1$V2,
  state      = p2$V1,
  points     = as.numeric(p1$V3),
  pre_rating = as.numeric(sub(".*R:\\s*(\\d+).*", "\\1", p2$V2))
)

nrow(chess)
[1] 64
head(chess)
  id                name state points pre_rating
1  1            GARY HUA    ON    6.0       1794
2  2     DAKSHESH DARURI    MI    6.0       1553
3  3        ADITYA BAJAJ    MI    6.0       1384
4  4 PATRICK H SCHILLING    MI    5.5       1716
5  5          HANSHI ZUO    MI    5.5       1655
6  6         HANSEN SONG    OH    5.0       1686

Getting the Results and Opponents

Each round entry, like “W 39”, holds two values: the result letter and the opponent’s ID. In Project 1, I only took the opponent’s ID. This time I also need the result letter, so I take both from columns V4 to V10.

letters_df <- p1[, 4:10] # result letters
opp_df     <- p1[, 4:10] # opponent IDs

for (r in 1:7) {
  letters_df[[r]] <- substr(trimws(p1[[r + 3]]), 1, 1)       # first letter only
  opp_df[[r]]     <- as.integer(gsub("\\D", "", p1[[r + 3]])) # numbers only
}

head(letters_df)
  V4 V5 V6 V7 V8 V9 V10
1  W  W  W  W  W  D   D
2  W  W  L  W  W  W   W
3  L  W  W  W  W  W   W
4  W  D  W  W  D  W   D
5  W  W  D  D  D  W   W
6  W  D  L  W  D  W   W
head(opp_df)
  V4 V5 V6 V7 V8 V9 V10
1 39 21 18 14  7 12   4
2 63 58  4 17 16 20   7
3  8 61 25 21 11 13  12
4 23 28  2 26  5 19   1
5 45 37 12 13  4 14  17
6 34 29 11 35 10 27  21

Calculating Expected and Actual Scores

For each player, I go through their seven rounds. I only use rounds with W (win), L (loss), or D (draw), because byes (B, H), unplayed rounds (U), and forfeits (X) have no real game. For each game, I turn the result into points: win = 1, draw = 0.5, loss = 0, use the Elo formula to find the expected score, and add both to the player’s totals.

actual   <- numeric(64)
expected <- numeric(64)
games    <- numeric(64)

for (i in 1:64) {
  for (r in 1:7) {
    letter <- letters_df[i, r]
    opp    <- opp_df[i, r]

    if (letter %in% c("W", "L", "D")) {
      # actual points
      if (letter == "W") actual[i] <- actual[i] + 1
      if (letter == "D") actual[i] <- actual[i] + 0.5

      # expected points (Elo formula)
      my_rating  <- chess$pre_rating[i]
      opp_rating <- chess$pre_rating[opp]
      expected[i] <- expected[i] + 1 / (1 + 10^((opp_rating - my_rating) / 400))

      games[i] <- games[i] + 1
    }
  }
}

Now I put everything in one table and calculate the difference between the actual and expected scores.

results <- data.frame(
  name         = chess$name,
  pre_rating   = chess$pre_rating,
  games_played = games,
  actual       = actual,
  expected     = round(expected, 2),
  difference   = round(actual - expected, 2)
)

head(results)
                 name pre_rating games_played actual expected difference
1            GARY HUA       1794            7    6.0     5.16       0.84
2     DAKSHESH DARURI       1553            7    6.0     3.78       2.22
3        ADITYA BAJAJ       1384            7    6.0     1.95       4.05
4 PATRICK H SCHILLING       1716            7    5.5     4.74       0.76
5          HANSHI ZUO       1655            7    5.5     4.38       1.12
6         HANSEN SONG       1686            7    5.0     4.94       0.06

Checking if everything looks good: In every game, the two players’ points add up to 1, and so do their expected scores. So the total actual score and the total expected score should both equal the number of games played in the tournament.

sum(games) / 2   # number of games (each game is counted twice)
[1] 204
sum(actual)
[1] 204
sum(expected)
[1] 204

All three numbers are the same, so the calculations look correct.

Results

Top 5 Overperformers

results %>%
  arrange(desc(difference)) %>%
  head(5)
                      name pre_rating games_played actual expected difference
1             ADITYA BAJAJ       1384            7    6.0     1.95       4.05
2   ZACHARY JAMES HOUGHTON       1220            7    4.5     1.37       3.13
3                ANVIT RAO       1365            7    5.0     1.94       3.06
4 JACOB ALEXANDER LAVALLEY        377            7    3.0     0.04       2.96
5              STEFANO LEE       1411            7    5.0     2.29       2.71

Top 5 Underperformers

results %>%
  arrange(difference) %>%
  head(5)
                name pre_rating games_played actual expected difference
1   LOREN SCHWIEBERT       1745            7    3.5     6.28      -2.78
2 GEORGE AVERY JONES       1522            7    3.5     6.02      -2.52
3        LARRY HODGE       1270            6    1.0     3.40      -2.40
4           JARED GE       1332            7    3.0     5.01      -2.01
5       RISHI SHETTY       1494            7    3.5     5.09      -1.59

Charts

The first chart shows the ten players from the tables above. Green bars are overperformers and red bars are underperformers.

top_bottom <- rbind(head(arrange(results, desc(difference)), 5),
                    head(arrange(results, difference), 5))

ggplot(top_bottom, aes(x = reorder(name, difference), y = difference,
                       fill = difference > 0)) +
  geom_col(show.legend = FALSE) +
  scale_fill_manual(values = c("firebrick", "seagreen")) +
  coord_flip() +
  labs(title = "Biggest Over- and Underperformers",
       x = NULL, y = "Actual score - expected score")

The second chart shows all 64 players. Each dot is one player. The dashed line shows where the actual score equals the expected score. Players above the line did better than expected, and players below the line did worse.

ggplot(results, aes(x = expected, y = actual)) +
  geom_point() +
  geom_abline(slope = 1, intercept = 0, linetype = "dashed") +
  labs(title = "Expected vs. Actual Score",
       x = "Expected score", y = "Actual score")

Conclusion

All five overperformers had pre-ratings close to or below the middle of the field (the median pre-rating is 1407). This makes sense, because Elo expects lower-rated players to lose against stronger opponents, so every win or draw against a stronger player raises their difference a lot. Aditya Bajaj stands out the most: he was expected to score about 2 points but scored 6. Jacob Alexander Lavalley, rated only 377, had an expected score close to 0 and still scored 3 points. His rating, like Zachary James Houghton’s, is provisional, which means it is based on only a few games. A provisional rating may not show how strong a player really is, especially for young players who improve quickly.

The underperformers were more mixed. Loren Schwiebert, one of the higher-rated players, was expected to score about 6.3 points but scored only 3.5. On the other hand, Larry Hodge and Jared Ge were rated below the median and still scored about 2 points less than expected. Overall, the results show that pre-tournament ratings are a useful guide, but they do not always predict how a player will perform in a single tournament.

Using LLM

I used Claude (Anthropic, 2026) for this project. Building on my Project 1 code, Claude wrote most of the code for calculating the Elo scores and helped me edit my explanations. I went through each step to understand how it works and ran it myself.

References

Anthropic. (2026). Claude (Opus 5.5 version) [Large language model]. https://claude.ai

Elo, A. E. (1978). The rating of chessplayers, past and present. Arco.

Wikipedia contributors. (n.d.). Elo rating system. In Wikipedia. https://en.wikipedia.org/wiki/Elo_rating_system