For this assignment, I used the chess tournament data from Project 1 to calculate each player’s expected score with the Elo formula and compared it to their actual score. The difference shows which players performed better or worse than their pre-tournament rating predicted.
Load and Parse the Data
library(tidyverse)
── Attaching core tidyverse packages ──────────────────────── tidyverse 2.0.0 ──
✔ dplyr 1.2.1 ✔ readr 2.2.0
✔ forcats 1.0.1 ✔ stringr 1.6.0
✔ ggplot2 4.0.3 ✔ tibble 3.3.1
✔ lubridate 1.9.5 ✔ tidyr 1.3.2
✔ purrr 1.2.2
── Conflicts ────────────────────────────────────────── tidyverse_conflicts() ──
✖ dplyr::filter() masks stats::filter()
✖ dplyr::lag() masks stats::lag()
ℹ Use the conflicted package (<http://conflicted.r-lib.org/>) to force all conflicts to become errors
This uses the same parsing approach as Project 1, with one addition. Each round’s result letter (W, L, D, or a bye or forfeit code) is now saved along with the opponent’s pair number.
Source: “The Elo Rating System for Chess and Beyond” [video], February 15, 2019. https://www.youtube.com/watch?v=AsYfbmp0To0
ratings <- players_long %>%distinct(pair_num, opp_rating = pre_rating)games <- players_long %>%filter(result %in%c("W", "D", "L")) %>%left_join(ratings, by =c("opponent"="pair_num")) %>%mutate(score =case_when(result =="W"~1, result =="D"~0.5, result =="L"~0),expected =1/ (1+10^((opp_rating - pre_rating) /400)) )
Only games that were actually played (wins, draws, and losses) are included. Some players received points from byes or forfeits, but those rounds have no opponent rating to calculate an expected result from. Counting those points in the actual score would make those players look like they overperformed. For example, Amiyatosh Pwnanandam earned 3.5 total points, but only 2.0 of those came from games he played. Including his bye points would have moved him into the top five overperformers.
ggplot(results, aes(x = pre_rating, y = difference)) +geom_point() +geom_hline(yintercept =0, linetype ="dashed") +labs(title ="Actual Minus Expected Score by Pre-Tournament Rating",x ="Pre-Tournament Rating", y ="Actual - Expected" ) +theme_minimal()
Interpreting the Results
Aditya Bajaj was the biggest overperformer by a wide margin. With a rating of 1384, he was expected to score about 1.95 points against his opponents but finished with 6.0, tying for first place. Zachary James Houghton, Anvit Rao, and Stefano Lee also scored roughly three points more than expected. Jacob Alexander Lavalley is an unusual case. His pre-rating of 377 is provisional, based on only three previous games, which made his expected score close to zero. His three wins look like a huge overperformance, but his rating was likely just inaccurate going in, which his post-tournament rating of 1076 reflects.
On the other end, Loren Schwiebert was the biggest underperformer. With a rating of 1745, he was expected to score about 6.3 points but finished with 3.5. George Avery Jones, Larry Hodge, Jared Ge, and Rishi Shetty also finished well below their expected scores.
Conclusions
The Elo expected score gives a useful way to judge performance relative to the strength of each player’s opponents, rather than looking at total points alone. The bye issue showed that how the data is handled can change the results, so it was important to compare only games that were actually played. Provisional ratings like Lavalley’s are another limitation, since the formula assumes each rating is an accurate measure of ability. A next step would be to apply the Elo update formula to calculate each player’s new rating after every round and compare it to the post-tournament ratings listed in the file.
AI Citation
Anthropic. (2026). Claude Sonnet 5 [Large language model]. https://claude.ai. Accessed October 2026.