This analysis builds on the chess tournament data from Project 1 to compare each player’s actual tournament score with the score that would be expected from the rating differences between that player and each opponent. The business question is: Which players performed most above or below what their pre-tournament ratings predicted? I use the standard Elo expected-score formula and compare the summed expected score with each player’s official tournament score.
For the expected-score calculation, I use the logistic Elo formula \(E_A = 1/(1 + 10^{(R_B-R_A)/400})\), where \(R_A\) is the player’s rating and \(R_B\) is the opponent’s rating. FIDE’s rating regulations similarly determine a scoring probability from the rating difference for each rated game: FIDE Rating Regulations. The logistic formula used here is also described in this Elo rating system reference.
I use tidyverse for importing, parsing, joining,
reshaping, and summarizing the tournament data, and knitr
for formatted tables.
library(tidyverse)
library(knitr)
options(dplyr.summarise.inform = FALSE)
To make the analysis reproducible, the tournament file is read directly from the public GitHub repository used for Project 1 rather than from a file stored on my computer.
data_url <- "https://raw.githubusercontent.com/chanicemcken/Data-607-Project-1/main/tournamentinfo.txt"
tournament_raw <- readLines(data_url, warn = FALSE)
head(tournament_raw, 10)
## [1] "-----------------------------------------------------------------------------------------"
## [2] " Pair | Player Name |Total|Round|Round|Round|Round|Round|Round|Round| "
## [3] " Num | USCF ID / Rtg (Pre->Post) | Pts | 1 | 2 | 3 | 4 | 5 | 6 | 7 | "
## [4] "-----------------------------------------------------------------------------------------"
## [5] " 1 | GARY HUA |6.0 |W 39|W 21|W 18|W 14|W 7|D 12|D 4|"
## [6] " ON | 15445895 / R: 1794 ->1817 |N:2 |W |B |W |B |W |B |W |"
## [7] "-----------------------------------------------------------------------------------------"
## [8] " 2 | DAKSHESH DARURI |6.0 |W 63|W 58|L 4|W 17|W 16|W 20|W 7|"
## [9] " MI | 14598900 / R: 1553 ->1663 |N:2 |B |W |B |W |B |W |B |"
## [10] "-----------------------------------------------------------------------------------------"
The tournament file stores each player across two lines. The first line contains pair number, player name, total score, and seven round results. The second line contains state and rating information. I identify the 64 player rows, match each one to the following detail row, and extract the fields needed for the Elo analysis.
player_idx <- which(str_detect(
tournament_raw,
"^\\s*\\d+\\s*\\|"
))
stopifnot(length(player_idx) == 64)
player_lines <- tournament_raw[player_idx]
detail_lines <- tournament_raw[player_idx + 1]
player_fields <- str_split_fixed(player_lines, "\\|", 11)
detail_fields <- str_split_fixed(detail_lines, "\\|", 11)
players <- tibble(
Pair_Number = as.integer(str_trim(player_fields[, 1])),
Player_Name = str_squish(player_fields[, 2]),
Actual_Score = as.numeric(str_trim(player_fields[, 3])),
State = str_trim(detail_fields[, 1]),
Pre_Rating = as.integer(
str_extract(detail_fields[, 2], "(?<=R:)\\s*\\d+")
),
Round_1 = str_trim(player_fields[, 4]),
Round_2 = str_trim(player_fields[, 5]),
Round_3 = str_trim(player_fields[, 6]),
Round_4 = str_trim(player_fields[, 7]),
Round_5 = str_trim(player_fields[, 8]),
Round_6 = str_trim(player_fields[, 9]),
Round_7 = str_trim(player_fields[, 10])
)
kable(
head(players, 10),
caption = "First 10 Parsed Player Records"
)
| Pair_Number | Player_Name | Actual_Score | State | Pre_Rating | Round_1 | Round_2 | Round_3 | Round_4 | Round_5 | Round_6 | Round_7 |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | GARY HUA | 6.0 | ON | 1794 | W 39 | W 21 | W 18 | W 14 | W 7 | D 12 | D 4 |
| 2 | DAKSHESH DARURI | 6.0 | MI | 1553 | W 63 | W 58 | L 4 | W 17 | W 16 | W 20 | W 7 |
| 3 | ADITYA BAJAJ | 6.0 | MI | 1384 | L 8 | W 61 | W 25 | W 21 | W 11 | W 13 | W 12 |
| 4 | PATRICK H SCHILLING | 5.5 | MI | 1716 | W 23 | D 28 | W 2 | W 26 | D 5 | W 19 | D 1 |
| 5 | HANSHI ZUO | 5.5 | MI | 1655 | W 45 | W 37 | D 12 | D 13 | D 4 | W 14 | W 17 |
| 6 | HANSEN SONG | 5.0 | OH | 1686 | W 34 | D 29 | L 11 | W 35 | D 10 | W 27 | W 21 |
| 7 | GARY DEE SWATHELL | 5.0 | MI | 1649 | W 57 | W 46 | W 13 | W 11 | L 1 | W 9 | L 2 |
| 8 | EZEKIEL HOUGHTON | 5.0 | MI | 1641 | W 3 | W 32 | L 14 | L 9 | W 47 | W 28 | W 19 |
| 9 | STEFANO LEE | 5.0 | ON | 1411 | W 25 | L 18 | W 59 | W 8 | W 26 | L 7 | W 20 |
| 10 | ANVIT RAO | 5.0 | MI | 1365 | D 16 | L 19 | W 55 | W 31 | D 6 | W 25 | W 18 |
The numeric extraction of Pre_Rating intentionally keeps
the rating itself while dropping provisional-rating text that may follow
it. This prevents values such as provisional ratings from being treated
as nonnumeric.
The seven rounds are stored in separate columns, so I transform them
into long format. Each row below represents one player-round
observation. The result code (W, D,
L, H, or U) is separated from the
opponent pair number.
games <- players %>%
pivot_longer(
cols = starts_with("Round_"),
names_to = "Round",
values_to = "Round_Result"
) %>%
mutate(
Round = as.integer(str_remove(Round, "Round_")),
Result = str_extract(Round_Result, "^[WDLHU]"),
Opponent_Pair = as.integer(str_extract(Round_Result, "\\d+")),
Game_Score = case_when(
Result == "W" ~ 1,
Result == "D" ~ 0.5,
Result == "L" ~ 0,
TRUE ~ NA_real_
)
)
kable(
head(games, 14),
caption = "Tournament Results in Long Format"
)
| Pair_Number | Player_Name | Actual_Score | State | Pre_Rating | Round | Round_Result | Result | Opponent_Pair | Game_Score |
|---|---|---|---|---|---|---|---|---|---|
| 1 | GARY HUA | 6 | ON | 1794 | 1 | W 39 | W | 39 | 1.0 |
| 1 | GARY HUA | 6 | ON | 1794 | 2 | W 21 | W | 21 | 1.0 |
| 1 | GARY HUA | 6 | ON | 1794 | 3 | W 18 | W | 18 | 1.0 |
| 1 | GARY HUA | 6 | ON | 1794 | 4 | W 14 | W | 14 | 1.0 |
| 1 | GARY HUA | 6 | ON | 1794 | 5 | W 7 | W | 7 | 1.0 |
| 1 | GARY HUA | 6 | ON | 1794 | 6 | D 12 | D | 12 | 0.5 |
| 1 | GARY HUA | 6 | ON | 1794 | 7 | D 4 | D | 4 | 0.5 |
| 2 | DAKSHESH DARURI | 6 | MI | 1553 | 1 | W 63 | W | 63 | 1.0 |
| 2 | DAKSHESH DARURI | 6 | MI | 1553 | 2 | W 58 | W | 58 | 1.0 |
| 2 | DAKSHESH DARURI | 6 | MI | 1553 | 3 | L 4 | L | 4 | 0.0 |
| 2 | DAKSHESH DARURI | 6 | MI | 1553 | 4 | W 17 | W | 17 | 1.0 |
| 2 | DAKSHESH DARURI | 6 | MI | 1553 | 5 | W 16 | W | 16 | 1.0 |
| 2 | DAKSHESH DARURI | 6 | MI | 1553 | 6 | W 20 | W | 20 | 1.0 |
| 2 | DAKSHESH DARURI | 6 | MI | 1553 | 7 | W 7 | W | 7 | 1.0 |
Special entries such as H and U do not
contain a rated opponent pair number, so they are not treated as normal
opponent matchups. I retain them in the data for validation, but only
rounds with a valid opponent number are used to calculate Elo expected
score.
The expected score for a game depends on both players’ ratings. I therefore create a lookup table using pair number and pre-rating, then join it back to each player’s opponent number.
opponent_lookup <- players %>%
select(
Opponent_Pair = Pair_Number,
Opponent_Name = Player_Name,
Opponent_Rating = Pre_Rating
)
rated_games <- games %>%
filter(!is.na(Opponent_Pair)) %>%
left_join(opponent_lookup, by = "Opponent_Pair")
opponent_validation <- rated_games %>%
summarise(
Rated_Game_Records = n(),
Missing_Opponent_Matches = sum(is.na(Opponent_Rating))
)
kable(
opponent_validation,
caption = "Opponent Matching Validation"
)
| Rated_Game_Records | Missing_Opponent_Matches |
|---|---|
| 408 | 0 |
stopifnot(opponent_validation$Missing_Opponent_Matches == 0)
Before calculating expected scores, I validate the opponent matching using Gary Hua from Project 1. His listed opponents are pair numbers 39, 21, 18, 14, 7, 12, and 4. Their average pre-rating should be approximately 1605.
gary_check <- rated_games %>%
filter(Player_Name == "GARY HUA") %>%
select(
Round,
Player_Name,
Pre_Rating,
Opponent_Pair,
Opponent_Name,
Opponent_Rating
)
kable(
gary_check,
caption = "Gary Hua Opponent Matching Check"
)
| Round | Player_Name | Pre_Rating | Opponent_Pair | Opponent_Name | Opponent_Rating |
|---|---|---|---|---|---|
| 1 | GARY HUA | 1794 | 39 | JOEL R HENDON | 1436 |
| 2 | GARY HUA | 1794 | 21 | DINH DANG BUI | 1563 |
| 3 | GARY HUA | 1794 | 18 | DAVID SUNDEEN | 1600 |
| 4 | GARY HUA | 1794 | 14 | BRADLEY SHAW | 1610 |
| 5 | GARY HUA | 1794 | 7 | GARY DEE SWATHELL | 1649 |
| 6 | GARY HUA | 1794 | 12 | KENNETH J TACK | 1663 |
| 7 | GARY HUA | 1794 | 4 | PATRICK H SCHILLING | 1716 |
gary_average <- mean(gary_check$Opponent_Rating)
tibble(
Validation = "Gary Hua average opponent pre-rating",
Calculated_Value = round(gary_average, 2),
Expected_Approximate_Value = 1605
) %>%
kable(caption = "Project 1 Validation Check")
| Validation | Calculated_Value | Expected_Approximate_Value |
|---|---|---|
| Gary Hua average opponent pre-rating | 1605.29 | 1605 |
This check is important because an incorrect opponent match would affect the expected probability for a game and therefore the player’s final expected tournament score.
For each rated matchup, I calculate the player’s expected score using:
\[ E_A = \frac{1}{1 + 10^{(R_B-R_A)/400}} \]
An expected value near 0.50 indicates an approximately even matchup. A value above 0.50 means the player was favored based on pre-tournament rating, while a value below 0.50 means the opponent was favored.
elo_expected <- function(player_rating, opponent_rating) {
1 / (1 + 10 ^ ((opponent_rating - player_rating) / 400))
}
rated_games <- rated_games %>%
mutate(
Rating_Difference = Pre_Rating - Opponent_Rating,
Expected_Score_Game = elo_expected(
Pre_Rating,
Opponent_Rating
)
)
kable(
rated_games %>%
select(
Player_Name,
Round,
Pre_Rating,
Opponent_Name,
Opponent_Rating,
Rating_Difference,
Expected_Score_Game
) %>%
head(12),
digits = 3,
caption = "Example Elo Expected-Score Calculations"
)
| Player_Name | Round | Pre_Rating | Opponent_Name | Opponent_Rating | Rating_Difference | Expected_Score_Game |
|---|---|---|---|---|---|---|
| GARY HUA | 1 | 1794 | JOEL R HENDON | 1436 | 358 | 0.887 |
| GARY HUA | 2 | 1794 | DINH DANG BUI | 1563 | 231 | 0.791 |
| GARY HUA | 3 | 1794 | DAVID SUNDEEN | 1600 | 194 | 0.753 |
| GARY HUA | 4 | 1794 | BRADLEY SHAW | 1610 | 184 | 0.743 |
| GARY HUA | 5 | 1794 | GARY DEE SWATHELL | 1649 | 145 | 0.697 |
| GARY HUA | 6 | 1794 | KENNETH J TACK | 1663 | 131 | 0.680 |
| GARY HUA | 7 | 1794 | PATRICK H SCHILLING | 1716 | 78 | 0.610 |
| DAKSHESH DARURI | 1 | 1553 | THOMAS JOSEPH HOSMER | 1175 | 378 | 0.898 |
| DAKSHESH DARURI | 2 | 1553 | VIRAJ MOHILE | 917 | 636 | 0.975 |
| DAKSHESH DARURI | 3 | 1553 | PATRICK H SCHILLING | 1716 | -163 | 0.281 |
| DAKSHESH DARURI | 4 | 1553 | RONALD GRZEGORCZYK | 1629 | -76 | 0.392 |
| DAKSHESH DARURI | 5 | 1553 | MIKE NIKITIN | 1604 | -51 | 0.427 |
A player’s expected tournament score is the sum of the expected scores from all rounds with a rated opponent. I then subtract expected score from the official tournament score. Positive values indicate overperformance and negative values indicate underperformance.
Because H and U entries do not identify a
rated opponent, they do not receive an Elo expected-score value. The
official tournament score is retained as the assignment’s actual-score
measure, and the number of rated games is shown so these special cases
remain visible when interpreting the results.
expected_totals <- rated_games %>%
group_by(Pair_Number, Player_Name, Pre_Rating) %>%
summarise(
Rated_Games = n(),
Expected_Score = sum(Expected_Score_Game)
)
player_performance <- players %>%
select(
Pair_Number,
Player_Name,
State,
Pre_Rating,
Actual_Score
) %>%
left_join(
expected_totals,
by = c("Pair_Number", "Player_Name", "Pre_Rating")
) %>%
mutate(
Expected_Score = replace_na(Expected_Score, 0),
Rated_Games = replace_na(Rated_Games, 0L),
Performance_Difference = Actual_Score - Expected_Score
) %>%
arrange(desc(Performance_Difference))
kable(
player_performance %>%
mutate(
Expected_Score = round(Expected_Score, 2),
Performance_Difference = round(Performance_Difference, 2)
),
caption = "Actual and Expected Tournament Scores for All Players"
)
| Pair_Number | Player_Name | State | Pre_Rating | Actual_Score | Rated_Games | Expected_Score | Performance_Difference |
|---|---|---|---|---|---|---|---|
| 3 | ADITYA BAJAJ | MI | 1384 | 6.0 | 7 | 1.95 | 4.05 |
| 15 | ZACHARY JAMES HOUGHTON | MI | 1220 | 4.5 | 7 | 1.37 | 3.13 |
| 10 | ANVIT RAO | MI | 1365 | 5.0 | 7 | 1.94 | 3.06 |
| 46 | JACOB ALEXANDER LAVALLEY | MI | 377 | 3.0 | 7 | 0.04 | 2.96 |
| 37 | AMIYATOSH PWNANANDAM | MI | 980 | 3.5 | 5 | 0.77 | 2.73 |
| 9 | STEFANO LEE | ON | 1411 | 5.0 | 7 | 2.29 | 2.71 |
| 2 | DAKSHESH DARURI | MI | 1553 | 6.0 | 7 | 3.78 | 2.22 |
| 52 | ETHAN GUO | MI | 935 | 2.5 | 7 | 0.30 | 2.20 |
| 59 | SEAN M MC CORMICK | MI | 853 | 2.0 | 6 | 0.41 | 1.59 |
| 58 | VIRAJ MOHILE | MI | 917 | 2.0 | 6 | 0.43 | 1.57 |
| 51 | TEJAS AYYAGARI | MI | 1011 | 2.5 | 7 | 1.03 | 1.47 |
| 24 | MICHAEL R ALDRICH | MI | 1229 | 4.0 | 7 | 2.55 | 1.45 |
| 5 | HANSHI ZUO | MI | 1655 | 5.5 | 7 | 4.38 | 1.12 |
| 50 | SHIVAM JHA | MI | 1056 | 2.5 | 6 | 1.42 | 1.08 |
| 44 | JUSTIN D SCHILLING | MI | 1199 | 3.0 | 6 | 2.07 | 0.93 |
| 56 | MARISA RICCI | MI | 1153 | 2.0 | 5 | 1.08 | 0.92 |
| 60 | JULIA SHEN | MI | 967 | 1.5 | 5 | 0.60 | 0.90 |
| 38 | BRIAN LIU | MI | 1423 | 3.0 | 6 | 2.13 | 0.87 |
| 1 | GARY HUA | ON | 1794 | 6.0 | 7 | 5.16 | 0.84 |
| 36 | SIDDHARTH JHA | MI | 1355 | 3.5 | 6 | 2.70 | 0.80 |
| 4 | PATRICK H SCHILLING | MI | 1716 | 5.5 | 7 | 4.74 | 0.76 |
| 57 | MICHAEL LU | MI | 1092 | 2.0 | 6 | 1.30 | 0.70 |
| 41 | KYLE WILLIAM MURPHY | MI | 1403 | 3.0 | 4 | 2.36 | 0.64 |
| 55 | ALEX KONG | MI | 1186 | 2.0 | 6 | 1.44 | 0.56 |
| 61 | JEZZEL FARKAS | ON | 955 | 1.5 | 7 | 0.97 | 0.53 |
| 7 | GARY DEE SWATHELL | MI | 1649 | 5.0 | 7 | 4.58 | 0.42 |
| 12 | KENNETH J TACK | MI | 1663 | 4.5 | 6 | 4.11 | 0.39 |
| 14 | BRADLEY SHAW | MI | 1610 | 4.5 | 7 | 4.18 | 0.32 |
| 53 | JOSE C YBARRA | MI | 1393 | 2.0 | 3 | 1.72 | 0.28 |
| 16 | MIKE NIKITIN | MI | 1604 | 4.0 | 5 | 3.80 | 0.20 |
| 28 | SOFIA ADINA STANESCU-BELLU | MI | 1507 | 3.5 | 7 | 3.31 | 0.19 |
| 62 | ASHWIN BALAJI | MI | 1530 | 1.0 | 1 | 0.88 | 0.12 |
| 34 | MICHAEL JEFFERY THOMAS | MI | 1399 | 3.5 | 7 | 3.44 | 0.06 |
| 40 | FOREST ZHANG | MI | 1348 | 3.0 | 7 | 2.94 | 0.06 |
| 23 | ALAN BUI | ON | 1363 | 4.0 | 7 | 3.94 | 0.06 |
| 6 | HANSEN SONG | OH | 1686 | 5.0 | 7 | 4.94 | 0.06 |
| 48 | DANIEL KHAIN | MI | 1382 | 2.5 | 5 | 2.53 | -0.03 |
| 8 | EZEKIEL HOUGHTON | MI | 1641 | 5.0 | 7 | 5.03 | -0.03 |
| 49 | MICHAEL J MARTIN | MI | 1291 | 2.5 | 5 | 2.54 | -0.04 |
| 32 | JOSHUA PHILIP MATHEWS | ON | 1441 | 3.5 | 7 | 3.72 | -0.22 |
| 21 | DINH DANG BUI | ON | 1563 | 4.0 | 7 | 4.32 | -0.32 |
| 19 | DIPANKAR ROY | MI | 1564 | 4.0 | 7 | 4.33 | -0.33 |
| 63 | THOMAS JOSEPH HOSMER | MI | 1175 | 1.0 | 5 | 1.43 | -0.43 |
| 13 | TORRANCE HENRY JR | MI | 1666 | 4.5 | 7 | 4.95 | -0.45 |
| 22 | EUGENE L MCCLURE | MI | 1555 | 4.0 | 6 | 4.48 | -0.48 |
| 27 | GAURAV GIDWANI | MI | 1552 | 3.5 | 6 | 4.00 | -0.50 |
| 18 | DAVID SUNDEEN | MI | 1600 | 4.0 | 7 | 4.59 | -0.59 |
| 26 | MAX ZHU | ON | 1579 | 3.5 | 7 | 4.10 | -0.60 |
| 39 | JOEL R HENDON | MI | 1436 | 3.0 | 7 | 3.62 | -0.62 |
| 17 | RONALD GRZEGORCZYK | MI | 1629 | 4.0 | 7 | 4.66 | -0.66 |
| 47 | ERIC WRIGHT | MI | 1362 | 2.5 | 7 | 3.19 | -0.69 |
| 11 | CAMERON WILLIAM MC LEMAN | MI | 1712 | 4.5 | 7 | 5.34 | -0.84 |
| 29 | CHIEDOZIE OKORIE | MI | 1602 | 3.5 | 6 | 4.60 | -1.10 |
| 20 | JASON ZHENG | MI | 1595 | 4.0 | 7 | 5.13 | -1.13 |
| 33 | JADE GE | MI | 1449 | 3.5 | 7 | 4.64 | -1.14 |
| 64 | BEN LI | MI | 1163 | 1.0 | 7 | 2.27 | -1.27 |
| 43 | ROBERT GLEN VASEY | MI | 1283 | 3.0 | 7 | 4.33 | -1.33 |
| 45 | DEREK YAN | MI | 1242 | 3.0 | 7 | 4.37 | -1.37 |
| 54 | LARRY HODGE | MI | 1270 | 2.0 | 6 | 3.40 | -1.40 |
| 35 | JOSHUA DAVID LEE | MI | 1438 | 3.5 | 7 | 4.96 | -1.46 |
| 31 | RISHI SHETTY | MI | 1494 | 3.5 | 7 | 5.09 | -1.59 |
| 42 | JARED GE | MI | 1332 | 3.0 | 7 | 5.01 | -2.01 |
| 30 | GEORGE AVERY JONES | ON | 1522 | 3.5 | 7 | 6.02 | -2.52 |
| 25 | LOREN SCHWIEBERT | MI | 1745 | 3.5 | 7 | 6.28 | -2.78 |
Each single-game expected score must fall between 0 and 1. A player’s total expected score must also fall between 0 and the number of rated games. I check both conditions before ranking the players.
validation_checks <- tibble(
Check = c(
"All game expected scores are between 0 and 1",
"All player expected totals are between 0 and rated games",
"All 64 tournament players are included"
),
Passed = c(
all(between(rated_games$Expected_Score_Game, 0, 1)),
all(
player_performance$Expected_Score >= 0 &
player_performance$Expected_Score <= player_performance$Rated_Games
),
nrow(player_performance) == 64
)
)
kable(
validation_checks,
caption = "Expected-Score Validation Checks"
)
| Check | Passed |
|---|---|
| All game expected scores are between 0 and 1 | TRUE |
| All player expected totals are between 0 and rated games | TRUE |
| All 64 tournament players are included | TRUE |
stopifnot(all(validation_checks$Passed))
The five largest positive values of
Performance_Difference identify the players who scored the
most points above their Elo-based expectation.
top_overperformers <- player_performance %>%
slice_max(
order_by = Performance_Difference,
n = 5,
with_ties = FALSE
) %>%
mutate(
Expected_Score = round(Expected_Score, 2),
Performance_Difference = round(Performance_Difference, 2)
) %>%
select(
Player_Name,
Pre_Rating,
Actual_Score,
Expected_Score,
Performance_Difference
)
kable(
top_overperformers,
caption = "Five Players Who Most Overperformed Their Expected Score"
)
| Player_Name | Pre_Rating | Actual_Score | Expected_Score | Performance_Difference |
|---|---|---|---|---|
| ADITYA BAJAJ | 1384 | 6.0 | 1.95 | 4.05 |
| ZACHARY JAMES HOUGHTON | 1220 | 4.5 | 1.37 | 3.13 |
| ANVIT RAO | 1365 | 5.0 | 1.94 | 3.06 |
| JACOB ALEXANDER LAVALLEY | 377 | 3.0 | 0.04 | 2.96 |
| AMIYATOSH PWNANANDAM | 980 | 3.5 | 0.77 | 2.73 |
The five most negative values identify the players who scored the most points below their Elo-based expectation.
top_underperformers <- player_performance %>%
slice_min(
order_by = Performance_Difference,
n = 5,
with_ties = FALSE
) %>%
mutate(
Expected_Score = round(Expected_Score, 2),
Performance_Difference = round(Performance_Difference, 2)
) %>%
select(
Player_Name,
Pre_Rating,
Actual_Score,
Expected_Score,
Performance_Difference
)
kable(
top_underperformers,
caption = "Five Players Who Most Underperformed Their Expected Score"
)
| Player_Name | Pre_Rating | Actual_Score | Expected_Score | Performance_Difference |
|---|---|---|---|---|
| LOREN SCHWIEBERT | 1745 | 3.5 | 6.28 | -2.78 |
| GEORGE AVERY JONES | 1522 | 3.5 | 6.02 | -2.52 |
| JARED GE | 1332 | 3.0 | 5.01 | -2.01 |
| RISHI SHETTY | 1494 | 3.5 | 5.09 | -1.59 |
| JOSHUA DAVID LEE | 1438 | 3.5 | 4.96 | -1.46 |
The chart provides an additional way to see how actual tournament performance differs from Elo expectation. Players above zero scored more points than expected, while players below zero scored fewer points than expected.
player_performance %>%
mutate(
Player_Name = fct_reorder(
Player_Name,
Performance_Difference
)
) %>%
ggplot(
aes(
x = Performance_Difference,
y = Player_Name,
fill = Performance_Difference > 0
)
) +
geom_col(show.legend = FALSE) +
geom_vline(xintercept = 0, linetype = "dashed") +
labs(
title = "Tournament Performance Relative to Elo Expectation",
subtitle = "Actual score minus expected score",
x = "Performance Difference (Points)",
y = NULL
) +
theme_minimal() +
theme(
axis.text.y = element_text(size = 6)
)
This analysis compares tournament results with the scores predicted by each player’s pre-tournament rating and the ratings of the opponents they actually faced. The final overperformance and underperformance tables identify the five players with the largest positive and negative differences between actual and expected score.
The results should be interpreted as performance relative to rating-based expectation rather than as a ranking of the strongest players. A lower-rated player can overperform by scoring more points than expected against stronger opponents, while a highly rated player can underperform even with a relatively high tournament score if that score falls below expectation.
A useful extension would be to compare the logistic Elo probabilities used here with FIDE’s published scoring-probability table to see whether the choice of implementation changes the top-five rankings. Another extension would be to repeat the analysis across multiple tournaments to determine whether the largest overperformances persist or are specific to this event.
The analysis reads the tournament text file directly from a public GitHub URL and does not depend on local file paths. The session information below documents the R environment and package versions used when the report is rendered.
sessionInfo()
## R version 4.5.2 (2025-10-31)
## Platform: aarch64-apple-darwin20
## Running under: macOS Sequoia 15.7.3
##
## Matrix products: default
## BLAS: /System/Library/Frameworks/Accelerate.framework/Versions/A/Frameworks/vecLib.framework/Versions/A/libBLAS.dylib
## LAPACK: /Library/Frameworks/R.framework/Versions/4.5-arm64/Resources/lib/libRlapack.dylib; LAPACK version 3.12.1
##
## locale:
## [1] en_US.UTF-8/en_US.UTF-8/en_US.UTF-8/C/en_US.UTF-8/en_US.UTF-8
##
## time zone: America/New_York
## tzcode source: internal
##
## attached base packages:
## [1] stats graphics grDevices utils datasets methods base
##
## other attached packages:
## [1] knitr_1.51 lubridate_1.9.4 forcats_1.0.1 stringr_1.6.0
## [5] dplyr_1.2.1 purrr_1.2.1 readr_2.2.0 tidyr_1.3.2
## [9] tibble_3.3.1 ggplot2_4.0.2 tidyverse_2.0.0
##
## loaded via a namespace (and not attached):
## [1] gtable_0.3.6 jsonlite_2.0.0 compiler_4.5.2 tidyselect_1.2.1
## [5] dichromat_2.0-0.1 jquerylib_0.1.4 scales_1.4.0 yaml_2.3.12
## [9] fastmap_1.2.0 R6_2.6.1 labeling_0.4.3 generics_0.1.4
## [13] tzdb_0.5.0 bslib_0.10.0 pillar_1.11.1 RColorBrewer_1.1-3
## [17] rlang_1.1.7 stringi_1.8.7 cachem_1.1.0 xfun_0.56
## [21] sass_0.4.10 S7_0.2.1 otel_0.2.0 timechange_0.4.0
## [25] cli_3.6.5 withr_3.0.2 magrittr_2.0.4 digest_0.6.39
## [29] grid_4.5.2 rstudioapi_0.18.0 hms_1.1.4 lifecycle_1.0.5
## [33] vctrs_0.7.1 evaluate_1.0.5 glue_1.8.0 farver_2.1.2
## [37] rmarkdown_2.30 tools_4.5.2 pkgconfig_2.0.3 htmltools_0.5.9
OpenAI. (2026). ChatGPT (Version 5.6) [Large language model]. https://chat.openai.com. Accessed October 3, 2026.