This assignment is a continuation of the last assignment of chess tournament data. The goal of this assignment however is to use the ELO ranking formula to generate these different player’s algorithmically projected scores, and for us to see who did the best and who did the worst. It is a further exercise in cleaning and organizing data, as well as building on top of previous exercises. There were some obstacles that have come up from this that I will go into a little more detail in the body of this assignment.
I stole my original code from the first assignment and used it heavily to discern the data.
First we will load in the .txt as is and see what we can do to start chipping away at it. after struggling through read.csv, read_csv, and other similar functions, I found that read.delim was the option that produced the best data load-in, and using “|” as the delimiter allowed us to get a rough graph to start.
Chess_Data <- read.delim("tournamentinfo.txt", header = FALSE, sep = "|")
glimpse(Chess_Data)
## Rows: 196
## Columns: 11
## $ V1 <chr> "-----------------------------------------------------------------…
## $ V2 <chr> "", " Player Name ", " USCF ID / Rtg (Pre->Pos…
## $ V3 <chr> "", "Total", " Pts ", "", "6.0 ", "N:2 ", "", "6.0 ", "N:2 ", …
## $ V4 <chr> "", "Round", " 1 ", "", "W 39", "W ", "", "W 63", "B ", …
## $ V5 <chr> "", "Round", " 2 ", "", "W 21", "B ", "", "W 58", "W ", …
## $ V6 <chr> "", "Round", " 3 ", "", "W 18", "W ", "", "L 4", "B ", …
## $ V7 <chr> "", "Round", " 4 ", "", "W 14", "B ", "", "W 17", "W ", …
## $ V8 <chr> "", "Round", " 5 ", "", "W 7", "W ", "", "W 16", "B ", …
## $ V9 <chr> "", "Round", " 6 ", "", "D 12", "B ", "", "W 20", "W ", …
## $ V10 <chr> "", "Round", " 7 ", "", "D 4", "W ", "", "W 7", "B ", …
## $ V11 <lgl> NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA…
More cleaning from the previous assignment.
## Rows: 64
## Columns: 3
## $ Area <chr> " ON ", " MI ", " MI ", " MI ", " MI ", " OH ", " …
## $ Rating <chr> " 15445895 / R: 1794 ->1817 ", " 14598900 / R: 1553 ->1…
## $ index <int> 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, …
## Rows: 64
## Columns: 11
## $ Rank <chr> " 1 ", " 2 ", " 3 ", " 4 ", " 5 ", " 6 ", " …
## $ Name <chr> " GARY HUA ", " DAKSHESH DARURI …
## $ Points <chr> "6.0 ", "6.0 ", "6.0 ", "5.5 ", "5.5 ", "5.0 ", "5.0 ", …
## $ R1 <chr> "W 39", "W 63", "L 8", "W 23", "W 45", "W 34", "W 57", …
## $ R2 <chr> "W 21", "W 58", "W 61", "D 28", "W 37", "D 29", "W 46", …
## $ R3 <chr> "W 18", "L 4", "W 25", "W 2", "D 12", "L 11", "W 13", …
## $ R4 <chr> "W 14", "W 17", "W 21", "W 26", "D 13", "W 35", "W 11", …
## $ R5 <chr> "W 7", "W 16", "W 11", "D 5", "D 4", "D 10", "L 1", …
## $ R6 <chr> "D 12", "W 20", "W 13", "W 19", "W 14", "W 27", "W 9", …
## $ R7 <chr> "D 4", "W 7", "W 12", "D 1", "W 17", "W 21", "L 2", …
## $ index <int> 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, …
Last chunk of data from project 1 before I added more for this assignment. Getting data joined, values cleaned, etc.
Full_Clean_Data <- full_join(Chess_Data_Even,Chess_Data_Odd)
## Joining with `by = join_by(index)`
glimpse(Full_Clean_Data)
## Rows: 64
## Columns: 13
## $ Area <chr> " ON ", " MI ", " MI ", " MI ", " MI ", " OH ", " …
## $ Rating <chr> " 15445895 / R: 1794 ->1817 ", " 14598900 / R: 1553 ->1…
## $ index <int> 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, …
## $ Rank <chr> " 1 ", " 2 ", " 3 ", " 4 ", " 5 ", " 6 ", " …
## $ Name <chr> " GARY HUA ", " DAKSHESH DARURI …
## $ Points <chr> "6.0 ", "6.0 ", "6.0 ", "5.5 ", "5.5 ", "5.0 ", "5.0 ", …
## $ R1 <chr> "W 39", "W 63", "L 8", "W 23", "W 45", "W 34", "W 57", …
## $ R2 <chr> "W 21", "W 58", "W 61", "D 28", "W 37", "D 29", "W 46", …
## $ R3 <chr> "W 18", "L 4", "W 25", "W 2", "D 12", "L 11", "W 13", …
## $ R4 <chr> "W 14", "W 17", "W 21", "W 26", "D 13", "W 35", "W 11", …
## $ R5 <chr> "W 7", "W 16", "W 11", "D 5", "D 4", "D 10", "L 1", …
## $ R6 <chr> "D 12", "W 20", "W 13", "W 19", "W 14", "W 27", "W 9", …
## $ R7 <chr> "D 4", "W 7", "W 12", "D 1", "W 17", "W 21", "L 2", …
Full_Clean_Data2 <- Full_Clean_Data |>
mutate(
across(
c(R1, R2, R3, R4, R5, R6, R7),
~str_remove_all(.x, "W|L|D")
)
) |>
mutate(
across(
c(R1, R2, R3, R4, R5, R6, R7, Rank),
~as.integer(.x)
)
) |>
mutate(
Rating = str_remove(Rating, ".*:"),
Rating = str_remove(Rating, "-.*"),
Rating = str_remove(Rating, "P.*")
)
## Warning: There were 7 warnings in `mutate()`.
## The first warning was:
## ℹ In argument: `across(c(R1, R2, R3, R4, R5, R6, R7, Rank), ~as.integer(.x))`.
## Caused by warning:
## ! NAs introduced by coercion
## ℹ Run `dplyr::last_dplyr_warnings()` to see the 6 remaining warnings.
glimpse(Full_Clean_Data2)
## Rows: 64
## Columns: 13
## $ Area <chr> " ON ", " MI ", " MI ", " MI ", " MI ", " OH ", " …
## $ Rating <chr> " 1794 ", " 1553 ", " 1384 ", " 1716 ", " 1655 ", " 1…
## $ index <int> 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, …
## $ Rank <int> 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, …
## $ Name <chr> " GARY HUA ", " DAKSHESH DARURI …
## $ Points <chr> "6.0 ", "6.0 ", "6.0 ", "5.5 ", "5.5 ", "5.0 ", "5.0 ", …
## $ R1 <int> 39, 63, 8, 23, 45, 34, 57, 3, 25, 16, 38, 42, 36, 54, 19, 10, 4…
## $ R2 <int> 21, 58, 61, 28, 37, 29, 46, 32, 18, 19, 56, 33, 27, 44, 16, 15,…
## $ R3 <int> 18, 4, 25, 2, 12, 11, 13, 14, 59, 55, 6, 5, 7, 8, 30, NA, 26, 1…
## $ R4 <int> 14, 17, 21, 26, 13, 35, 11, 9, 8, 31, 7, 38, 5, 1, 22, 39, 2, 3…
## $ R5 <int> 7, 16, 11, 5, 4, 10, 1, 47, 26, 6, 3, NA, 33, 27, 54, 2, 23, 19…
## $ R6 <int> 12, 20, 13, 19, 14, 27, 9, 28, 7, 25, 34, 1, 3, 5, 33, 36, 22, …
## $ R7 <int> 4, 7, 12, 1, 17, 21, 2, 19, 20, 18, 26, 3, 32, 31, 38, NA, 5, 1…
Chess_Data_Joins <- Full_Clean_Data2 |>
select(!index) |>
mutate(R1elo = Rating[match(R1, Rank)],
R2elo = Rating[match(R2, Rank)],
R3elo = Rating[match(R3, Rank)],
R4elo = Rating[match(R4, Rank)],
R5elo = Rating[match(R5, Rank)],
R6elo = Rating[match(R6, Rank)],
R7elo = Rating[match(R7, Rank)]) |>
mutate(
across(
c(R1elo, R2elo, R3elo, R4elo, R5elo, R6elo, R7elo),
~as.integer(.x)
)
) |>
rowwise() |>
mutate(Opponent_Avg_Elo = mean(
c(R1elo, R2elo, R3elo, R4elo, R5elo, R6elo, R7elo
)
, na.rm = TRUE)) |>
relocate(Name, Area, Points, Rating, Opponent_Avg_Elo) |>
arrange(Rank) |>
select(Name, Area, Points, Rating, Opponent_Avg_Elo) |>
mutate(Opponent_Avg_Elo = round(Opponent_Avg_Elo, digits = 0))
Now for the new calculations below. I first got everything in the right format as integers, then I created new fields for probability, predicted score, and difference. The ELO formula I used was from “The Elo Rating System for Chess and Beyond” by singingbanana on Youtube. the big issue I encountered with this is that my original data did not decode the different chess tournament match outcomes such as B, U, H, and X. These returned as NA’s and I chose to keep them as NA for the sake of this exercise. I will go over more in the conclusion of what could be done to remediate these.
With the newly calculated ratings, I took the differences of these with the player’s actual scores to show who over and under performed.
Chess_Data_Elo <- Full_Clean_Data2 |>
select(!index) |>
mutate(R1elo = Rating[match(R1, Rank)],
R2elo = Rating[match(R2, Rank)],
R3elo = Rating[match(R3, Rank)],
R4elo = Rating[match(R4, Rank)],
R5elo = Rating[match(R5, Rank)],
R6elo = Rating[match(R6, Rank)],
R7elo = Rating[match(R7, Rank)]) |>
mutate(
across(
c(R1elo, R2elo, R3elo, R4elo, R5elo, R6elo, R7elo, Rating, Points),
~as.integer(.x)
)) |>
rowwise() |>
mutate(R1prob = 1/(1 + 10^((R1elo - Rating)/400)),
R2prob = 1/(1 + 10^((R2elo - Rating)/400)),
R3prob = 1/(1 + 10^((R3elo - Rating)/400)),
R4prob = 1/(1 + 10^((R4elo - Rating)/400)),
R5prob = 1/(1 + 10^((R5elo - Rating)/400)),
R6prob = 1/(1 + 10^((R6elo - Rating)/400)),
R7prob = 1/(1 + 10^((R7elo - Rating)/400)),
Predicted_Score = sum(across(c(R1prob,R2prob,R3prob,R4prob,R5prob,R6prob,R7prob)), na.rm = TRUE),
Difference = round((Points - Predicted_Score), digits = 2)
)|>
arrange(desc(Difference)) |>
select(!R1:R7prob)
For a bit more of a zoom in:
head(Chess_Data_Elo)
## # A tibble: 6 × 7
## # Rowwise:
## Area Rating Rank Name Points Predicted_Score Difference
## <chr> <int> <int> <chr> <int> <dbl> <dbl>
## 1 " MI " 1384 3 " ADITYA BAJAJ … 6 1.95 4.05
## 2 " MI " 1365 10 " ANVIT RAO … 5 1.94 3.06
## 3 " MI " 377 46 " JACOB ALEXANDER LAV… 3 0.0432 2.96
## 4 " ON " 1411 9 " STEFANO LEE … 5 2.29 2.71
## 5 " MI " 1220 15 " ZACHARY JAMES HOUGH… 4 1.37 2.63
## 6 " MI " 980 37 " AMIYATOSH PWNANANDA… 3 0.773 2.23
tail(Chess_Data_Elo)
## # A tibble: 6 × 7
## # Rowwise:
## Area Rating Rank Name Points Predicted_Score Difference
## <chr> <int> <int> <chr> <int> <dbl> <dbl>
## 1 " MI " 1449 33 " JADE GE … 3 4.64 -1.64
## 2 " MI " 1438 35 " JOSHUA DAVID LEE … 3 4.96 -1.96
## 3 " MI " 1332 42 " JARED GE … 3 5.01 -2.01
## 4 " MI " 1494 31 " RISHI SHETTY … 3 5.09 -2.09
## 5 " ON " 1522 30 " GEORGE AVERY JONES … 3 6.02 -3.02
## 6 " MI " 1745 25 " LOREN SCHWIEBERT … 3 6.28 -3.28
The ELO calculation is more intuitive than you would originally think, given that the probability of winning then becomes the points that you can use for theoretical placings of different players. There is no roundabout way to have to extrapolate the score past that point and I think that is a testament to how effective and popular of a rating style the Elo method is.
For how this assignment played out, it was nice to work more with this set of data as I thought it was inherently very interesting. I enjoy chess myself, so seeing how tournaments count scores and diving into the mechanics of those rankings was very neat for me. As I mentioned earlier, the one gap left open for me with this assignment is those NA values. What I would have to do to fix these is go back and assign a set value for the different U, X, B, H, for which some are assumed to be .5 and some are assumed to be 0. This could have an impact on the top and bottom 5 performers, but upon visual inspection it didn’t effect the top players as much as it did the lower ranks.
From my data, players like Aditya Bajaj, Anvit Rao, Jacob Lavalley, Stefano Lee, and Zachary Houghton has outstanding performances, while Loren Schwiebert, George Jones , Rishi Shetty, Jared Ge, and Joshua Lee were upset in the tournament.