In the Project 1 tournament, each player’s score depends on how strong their opponents were. A player who scores 4 out of 7 against weak opposition has done worse than one who scores 3 against strong opposition. The Elo system makes that comparison concrete: from the rating difference in each game, it calculates the score a player was expected to achieve.
This report calculates every player’s expected score, compares it with what they actually scored, and lists the five players who most overperformed and the five who most underperformed.
2 The Elo Expected Score
For a single game, a player’s expected score against one opponent is:
\[
E = \frac{1}{1 + 10^{(R_{\text{opponent}} - R_{\text{player}})/400}}
\]
\(E\) is a number between 0 and 1, read as the share of a point the player is expected to take. A player’s expected score for the tournament is the sum of \(E\) across the games they played.
Source:Elo rating system, Wikipedia, which gives this logistic form of Arpad Elo’s expected-score formula. The same formula appears in the FIDE Handbook rating regulations, where the expectancy is published as a lookup table of rating differences.
The table shows the shape of the formula: equal ratings give 0.5, a 100-point advantage gives 0.64, and a 200-point advantage gives 0.76. These match the figures quoted in the source.
3 The Data
The tournament file is the same one I parsed for Project 1, read from my GitHub repository. The parsing is unchanged, except that I now keep the result letter in each round as well as the opponent’s number, because the result is what the expected score gets compared against.
Only W, L and D have an opponent number. The other letters are byes and unplayed rounds, which have no opponent and therefore no rating difference.
rounds |>count(result, has_opponent =!is.na(opponent)) |>mutate(meaning =case_when( result =="W"~"Win", result =="L"~"Loss", result =="D"~"Draw", result =="B"~"Full-point bye", result =="H"~"Half-point bye", result =="U"~"Unplayed", result =="X"~"Forfeit win" ))
result
has_opponent
n
meaning
B
FALSE
7
Full-point bye
D
TRUE
58
Draw
H
FALSE
16
Half-point bye
L
TRUE
175
Loss
U
FALSE
16
Unplayed
W
TRUE
175
Win
X
FALSE
1
Forfeit win
Byes are excluded from both sides of the comparison. A bye awards points without a game, so including it in the actual score while the expected score has nothing to match it against would make every player with a bye look like an overperformer.
games <- rounds |>filter(!is.na(opponent)) |>left_join(players |>select(pair, pre_rating), by ="pair") |>left_join(players |>select(opponent = pair, opponent_rating = pre_rating), by ="opponent") |>mutate(actual =case_when(result =="W"~1, result =="D"~0.5, result =="L"~0),expected =expected_score(pre_rating, opponent_rating) )games |>select(pair, round, opponent, pre_rating, opponent_rating, actual, expected) |>mutate(expected =round(expected, 3)) |>head(7)
In any single game, the two players’ expected scores add up to 1, because one player’s advantage is the other’s disadvantage. Across the whole tournament, then, the total expected score must equal the total actual score, which is simply the number of games played.
These are players whose official total includes bye points. Their actual_score here counts only real games, which is what the expected score can be compared against.
highlight <-c(top_over$name, top_under$name)performance |>mutate(name =fct_reorder(name, difference),group =case_when(name %in% top_over$name ~"Overperformed", name %in% top_under$name ~"Underperformed",TRUE~"Other") ) |>ggplot(aes(x = difference, y = name, color = group)) +geom_vline(xintercept =0, color ="grey60") +geom_segment(aes(x =0, xend = difference, yend = name), linewidth =0.6) +geom_point(size =2) +scale_color_manual(values =c(Overperformed = col_over, Underperformed = col_under,Other ="grey75"), name =NULL) +# 64 names will not fit legibly on this axis; the ten highlighted players are# named in the tables above, and the two extremes are labelled directlyannotate("text", x =4.0, y =55, label ="Aditya Bajaj, +4.05",hjust =1, size =3.4, color = col_over) +annotate("text", x =-2.7, y =9, label ="Loren Schwiebert, -2.78",hjust =0, size =3.4, color = col_under) +labs(x ="Actual score minus expected score (points)", y ="Players, worst to best") + theme_report +theme(axis.text.y =element_blank(), legend.position ="top",panel.grid.major.y =element_blank())
Figure 1: Actual score minus expected score for all 64 players, sorted. The highlighted points are the five biggest over- and underperformers listed above.
performance |>mutate(group =case_when(name %in% top_over$name ~"Overperformed", name %in% top_under$name ~"Underperformed",TRUE~"Other")) |>ggplot(aes(x = expected_score, y = actual_score, color = group)) +geom_abline(linetype ="dashed", color ="grey60") +geom_point(size =2.6, alpha =0.9) +scale_color_manual(values =c(Overperformed = col_over, Underperformed = col_under,Other ="grey70"), name =NULL) +coord_equal(xlim =c(0, 7), ylim =c(0, 7)) +labs(x ="Expected score", y ="Actual score") + theme_report +theme(legend.position ="top")
Figure 2: Expected score against actual score. Players above the line scored more than their ratings predicted.
5.1 What the lists show
Aditya Bajaj is the clearest overperformance in the tournament: rated 1384, he was expected to score about 1.95 points from his seven games and scored 6, a difference of +4.05. At the other end, Loren Schwiebert, the second-highest rated player in the field at 1745, was expected to score 6.28 and managed 3.5, a difference of −2.78.
The two lists are not symmetric, and that is a feature of the formula rather than an accident. A strong player can only lose points they were expected to win, so the worst possible underperformance is bounded by their expected score. A weak player has almost nothing to lose and a great deal to gain, so the overperformance list is dominated by lower-rated players having a good week.
5.2 A sanity check from the file itself
The tournament file also records each player’s post-tournament rating, which was calculated independently by the rating system. If my expected scores are reasonable, players who overperformed should have gained rating and players who underperformed should have lost it.
The correlation is 0.78, and every one of the ten players moves in the expected direction: all five overperformers gained rating, and all five underperformers lost it. This is a genuinely independent check, because those post-tournament ratings were in the file before I calculated anything.
5.3 One player worth a caveat
games |>filter(pair == performance$pair[performance$name =="Jacob Alexander Lavalley"]) |>select(round, opponent, opponent_rating, result, actual, expected) |>mutate(expected =round(expected, 4))
round
opponent
opponent_rating
result
actual
expected
1
35
1438
W
1
0.0022
2
7
1649
L
0
0.0007
3
27
1552
L
0
0.0012
4
50
1056
L
0
0.0197
5
64
1163
W
1
0.0107
6
43
1283
W
1
0.0054
7
23
1363
L
0
0.0034
Jacob Alexander Lavalley entered rated 377 and won three games against opponents rated 1438, 1163 and 1283. The formula gives him an expected score of 0.04 points from seven games, so his 3 points look like an extraordinary overperformance. The more likely explanation is that his 377 rating was simply wrong, or based on too few games to mean anything. The rating system reached the same conclusion: his rating jumped 699 points to 1076. Elo measures performance against a rating, so it is only as trustworthy as the rating it starts from.
6 Conclusions
Across 204 games, the total expected score equals the total actual score exactly, which confirms the calculation is internally consistent.
The five biggest overperformers were Aditya Bajaj (+4.05), Zachary James Houghton (+3.13), Anvit Rao (+3.06), Jacob Alexander Lavalley (+2.96) and Stefano Lee (+2.71).
The five biggest underperformers were Loren Schwiebert (−2.78), George Avery Jones (−2.52), Larry Hodge (−2.40), Jared Ge (−2.01) and Rishi Shetty (−1.59).
The official rating changes in the file agree with these results, with a correlation of 0.78 and all ten players moving in the expected direction.
Byes had to be excluded from both scores. Had they been left in the actual score, the players who received them would have appeared to overperform by exactly the value of the bye.
7 References
Elo rating system, Wikipedia — source of the expected-score formula used here.
FIDE Handbook, rating regulations — the same expectancy published as a table.