For this project, I will use the Elo rating system (learn the Elo rating system ) to compare each player’s actual tournament performance with their expected performance. The goal is to calculate an expected score for every player based on the ratings of the opponents they faced during the tournament. After finding each player’s expected score, I will compare it to their actual score to determine whether they overperformed or underperformed relative to expectations. Finally, I will identify the five players who exceeded expectations the most and the five who performed below expectations the most.
To complete this assignment, I will begin by using the player ratings and opponent information from the chess tournament dataset created in Project 1. For each match, I will apply the Elo expected score formula to calculate the probability that a player will score against a specific opponent. I will then sum the expected scores across all opponents to obtain each player’s total expected tournament score. Next, I will compare the expected score to the player’s actual score and calculate the difference. Positive differences will indicate overperformance, while negative differences will indicate underperformance. Lastly, I will sort the results to identify the top five overperformers and underperformers.
One challenge I anticipate is matching each player with the ratings of all their opponents. While the Elo formula itself is straightforward, obtaining the correct opponent ratings and linking them to each player’s tournament record may require additional data manipulation and joins. Ensuring that all opponent ratings are correctly associated with the appropriate player will be important for producing accurate expected score calculations.
library(readr)
players_df <- read_csv("https://raw.githubusercontent.com/lioneljr17/LDATA607/refs/heads/main/607Project01/players_table.csv")
## Rows: 64 Columns: 12
## ── Column specification ────────────────────────────────────────────────────────
## Delimiter: ","
## chr (2): player_name, state
## dbl (10): player_num, points, pre_rating, r1, r2, r3, r4, r5, r6, r7
##
## ℹ Use `spec()` to retrieve the full column specification for this data.
## ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
head(players_df)
## # A tibble: 6 × 12
## player_num player_name state points pre_rating r1 r2 r3 r4 r5
## <dbl> <chr> <chr> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl>
## 1 1 GARY HUA ON 6 1794 39 21 18 14 7
## 2 2 DAKSHESH DAR… MI 6 1553 63 58 4 17 16
## 3 3 ADITYA BAJAJ MI 6 1384 8 61 25 21 11
## 4 4 PATRICK H SC… MI 5.5 1716 23 28 2 26 5
## 5 5 HANSHI ZUO MI 5.5 1655 45 37 12 13 4
## 6 6 HANSEN SONG OH 5 1686 34 29 11 35 10
## # ℹ 2 more variables: r6 <dbl>, r7 <dbl>
gameplayed_table <- read_csv("https://raw.githubusercontent.com/lioneljr17/LDATA607/refs/heads/main/607Project01/gameplayed_table.csv")
## Rows: 408 Columns: 3
## ── Column specification ────────────────────────────────────────────────────────
## Delimiter: ","
## dbl (3): player_num, round_num, opponent_num
##
## ℹ Use `spec()` to retrieve the full column specification for this data.
## ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
head(gameplayed_table,10)
## # A tibble: 10 × 3
## player_num round_num opponent_num
## <dbl> <dbl> <dbl>
## 1 1 1 39
## 2 1 2 21
## 3 1 3 18
## 4 1 4 14
## 5 1 5 7
## 6 1 6 12
## 7 1 7 4
## 8 2 1 63
## 9 2 2 58
## 10 2 3 4
\(E_O = \frac{1}{1 + 10^{(R_oop - R_player /400)}}\)
expected_score <- function (player_rating,
opponent_rating){1/(1+10^((opponent_rating - player_rating)/400))}
games_elo <- gameplayed_table %>%
left_join(
players_df %>%
select(player_num,
pre_rating),
by = "player_num"
) %>%
rename(player_rating = pre_rating)
view(games_elo)
games_elo <- games_elo %>%
left_join(
players_df %>%
select(player_num,
pre_rating),
by = c("opponent_num" = "player_num")
) %>%
rename(opponent_rating = pre_rating)
games_elo <- games_elo %>%
mutate(
expected_game =
expected_score(
player_rating,
opponent_rating
)
)
expected_totals <- games_elo %>%
group_by(player_num) %>%
summarise(
expected_score =
sum(expected_game,
na.rm = TRUE)
)
head(expected_totals)
## # A tibble: 6 × 2
## player_num expected_score
## <dbl> <dbl>
## 1 1 5.16
## 2 2 3.78
## 3 3 1.95
## 4 4 4.74
## 5 5 4.38
## 6 6 4.94
results <- players_df %>%
left_join(
expected_totals,
by = "player_num"
) %>%
mutate(
difference =
points - expected_score
)
final_results <- results %>%
select(
player_name,
points,
pre_rating,
expected_score,
difference
)
head(final_results)
## # A tibble: 6 × 5
## player_name points pre_rating expected_score difference
## <chr> <dbl> <dbl> <dbl> <dbl>
## 1 GARY HUA 6 1794 5.16 0.838
## 2 DAKSHESH DARURI 6 1553 3.78 2.22
## 3 ADITYA BAJAJ 6 1384 1.95 4.05
## 4 PATRICK H SCHILLING 5.5 1716 4.74 0.758
## 5 HANSHI ZUO 5.5 1655 4.38 1.12
## 6 HANSEN SONG 5 1686 4.94 0.0554
top_overperformers <- final_results %>%
arrange(desc(difference)) %>%
slice(1:5)
head(top_overperformers,10)
## # A tibble: 5 × 5
## player_name points pre_rating expected_score difference
## <chr> <dbl> <dbl> <dbl> <dbl>
## 1 ADITYA BAJAJ 6 1384 1.95 4.05
## 2 ZACHARY JAMES HOUGHTON 4.5 1220 1.37 3.13
## 3 ANVIT RAO 5 1365 1.94 3.06
## 4 JACOB ALEXANDER LAVALLEY 3 377 0.0432 2.96
## 5 AMIYATOSH PWNANANDAM 3.5 980 0.773 2.73
top_underperformers <- final_results %>%
arrange(difference) %>%
slice(1:5)
head(top_underperformers,10)
## # A tibble: 5 × 5
## player_name points pre_rating expected_score difference
## <chr> <dbl> <dbl> <dbl> <dbl>
## 1 LOREN SCHWIEBERT 3.5 1745 6.28 -2.78
## 2 GEORGE AVERY JONES 3.5 1522 6.02 -2.52
## 3 JARED GE 3 1332 5.01 -2.01
## 4 RISHI SHETTY 3.5 1494 5.09 -1.59
## 5 JOSHUA DAVID LEE 3.5 1438 4.96 -1.46
ggplot(
final_results,
aes(
x = reorder(player_name,
difference),
y = difference
)
) +
geom_col() +
coord_flip() +
labs(
title = "Difference Between Actual and Expected Scores",
x = "Player",
y = "Actual - Expected"
)
top_overperformers %>%
ggplot(
aes(
x = reorder(player_name, difference),
y = difference
)
) +
geom_col(fill = "steelblue") +
coord_flip() +
labs(
title = "Top 5 Overperformers",
x = "Player",
y = "Actual - Expected"
) +
theme_minimal()
#Top 5 Underperformers
top_underperformers %>%
ggplot(
aes(
x = reorder(player_name, difference),
y = difference
)
) +
geom_col(fill = "firebrick") +
coord_flip() +
labs(
title = "Top 5 Underperformers",
x = "Player",
y = "Actual - Expected"
) +
theme_minimal()
##Conclusion
The Elo rating system provides a way to estimate how many points a player should score based on the ratings of their opponents. By comparing expected scores to actual tournament results, it is possible to identify players who performed significantly above or below expectations. The analysis highlights the top five overperformers and underperformers, providing insight into which players achieved results that differed most from what the Elo system predicted.