Introduction

This assignment is a continuation of the last assignment of chess tournament data. The goal of this assignment however is to use the ELO ranking formula to generate these different player’s algorithmically projected scores, and for us to see who did the best and who did the worst. It is a further exercise in cleaning and organizing data, as well as building on top of previous exercises. There were some obstacles that have come up from this that I will go into a little more detail in the body of this assignment.

The Code

I stole my original code from the first assignment and used it heavily to discern the data.

First we will load in the .txt as is and see what we can do to start chipping away at it. after struggling through read.csv, read_csv, and other similar functions, I found that read.delim was the option that produced the best data load-in, and using “|” as the delimiter allowed us to get a rough graph to start.

Chess_Data <- read.delim("tournamentinfo.txt", header = FALSE, sep = "|")

glimpse(Chess_Data)
## Rows: 196
## Columns: 11
## $ V1  <chr> "-----------------------------------------------------------------…
## $ V2  <chr> "", " Player Name                     ", " USCF ID / Rtg (Pre->Pos…
## $ V3  <chr> "", "Total", " Pts ", "", "6.0  ", "N:2  ", "", "6.0  ", "N:2  ", …
## $ V4  <chr> "", "Round", "  1  ", "", "W  39", "W    ", "", "W  63", "B    ", …
## $ V5  <chr> "", "Round", "  2  ", "", "W  21", "B    ", "", "W  58", "W    ", …
## $ V6  <chr> "", "Round", "  3  ", "", "W  18", "W    ", "", "L   4", "B    ", …
## $ V7  <chr> "", "Round", "  4  ", "", "W  14", "B    ", "", "W  17", "W    ", …
## $ V8  <chr> "", "Round", "  5  ", "", "W   7", "W    ", "", "W  16", "B    ", …
## $ V9  <chr> "", "Round", "  6  ", "", "D  12", "B    ", "", "W  20", "W    ", …
## $ V10 <chr> "", "Round", "  7  ", "", "D   4", "W    ", "", "W   7", "B    ", …
## $ V11 <lgl> NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA…

More cleaning from the previous assignment.

## Rows: 64
## Columns: 3
## $ Area   <chr> "   ON ", "   MI ", "   MI ", "   MI ", "   MI ", "   OH ", "  …
## $ Rating <chr> " 15445895 / R: 1794   ->1817     ", " 14598900 / R: 1553   ->1…
## $ index  <int> 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, …
## Rows: 64
## Columns: 11
## $ Rank   <chr> "    1 ", "    2 ", "    3 ", "    4 ", "    5 ", "    6 ", "  …
## $ Name   <chr> " GARY HUA                        ", " DAKSHESH DARURI         …
## $ Points <chr> "6.0  ", "6.0  ", "6.0  ", "5.5  ", "5.5  ", "5.0  ", "5.0  ", …
## $ R1     <chr> "W  39", "W  63", "L   8", "W  23", "W  45", "W  34", "W  57", …
## $ R2     <chr> "W  21", "W  58", "W  61", "D  28", "W  37", "D  29", "W  46", …
## $ R3     <chr> "W  18", "L   4", "W  25", "W   2", "D  12", "L  11", "W  13", …
## $ R4     <chr> "W  14", "W  17", "W  21", "W  26", "D  13", "W  35", "W  11", …
## $ R5     <chr> "W   7", "W  16", "W  11", "D   5", "D   4", "D  10", "L   1", …
## $ R6     <chr> "D  12", "W  20", "W  13", "W  19", "W  14", "W  27", "W   9", …
## $ R7     <chr> "D   4", "W   7", "W  12", "D   1", "W  17", "W  21", "L   2", …
## $ index  <int> 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, …

Last chunk of data from project 1 before I added more for this assignment. Getting data joined, values cleaned, etc.

Full_Clean_Data <- full_join(Chess_Data_Even,Chess_Data_Odd) 
## Joining with `by = join_by(index)`
glimpse(Full_Clean_Data)
## Rows: 64
## Columns: 13
## $ Area   <chr> "   ON ", "   MI ", "   MI ", "   MI ", "   MI ", "   OH ", "  …
## $ Rating <chr> " 15445895 / R: 1794   ->1817     ", " 14598900 / R: 1553   ->1…
## $ index  <int> 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, …
## $ Rank   <chr> "    1 ", "    2 ", "    3 ", "    4 ", "    5 ", "    6 ", "  …
## $ Name   <chr> " GARY HUA                        ", " DAKSHESH DARURI         …
## $ Points <chr> "6.0  ", "6.0  ", "6.0  ", "5.5  ", "5.5  ", "5.0  ", "5.0  ", …
## $ R1     <chr> "W  39", "W  63", "L   8", "W  23", "W  45", "W  34", "W  57", …
## $ R2     <chr> "W  21", "W  58", "W  61", "D  28", "W  37", "D  29", "W  46", …
## $ R3     <chr> "W  18", "L   4", "W  25", "W   2", "D  12", "L  11", "W  13", …
## $ R4     <chr> "W  14", "W  17", "W  21", "W  26", "D  13", "W  35", "W  11", …
## $ R5     <chr> "W   7", "W  16", "W  11", "D   5", "D   4", "D  10", "L   1", …
## $ R6     <chr> "D  12", "W  20", "W  13", "W  19", "W  14", "W  27", "W   9", …
## $ R7     <chr> "D   4", "W   7", "W  12", "D   1", "W  17", "W  21", "L   2", …
Full_Clean_Data2 <- Full_Clean_Data |>
  mutate(
    across(
      c(R1, R2, R3, R4, R5, R6, R7),
         ~str_remove_all(.x, "W|L|D")
      )
  ) |>
    mutate(
    across(
      c(R1, R2, R3, R4, R5, R6, R7, Rank),
      ~as.integer(.x)
    )
  ) |>
  mutate(
    Rating = str_remove(Rating, ".*:"),
    Rating = str_remove(Rating, "-.*"),
    Rating = str_remove(Rating, "P.*")
  )
## Warning: There were 7 warnings in `mutate()`.
## The first warning was:
## ℹ In argument: `across(c(R1, R2, R3, R4, R5, R6, R7, Rank), ~as.integer(.x))`.
## Caused by warning:
## ! NAs introduced by coercion
## ℹ Run `dplyr::last_dplyr_warnings()` to see the 6 remaining warnings.
glimpse(Full_Clean_Data2)
## Rows: 64
## Columns: 13
## $ Area   <chr> "   ON ", "   MI ", "   MI ", "   MI ", "   MI ", "   OH ", "  …
## $ Rating <chr> " 1794   ", " 1553   ", " 1384   ", " 1716   ", " 1655   ", " 1…
## $ index  <int> 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, …
## $ Rank   <int> 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, …
## $ Name   <chr> " GARY HUA                        ", " DAKSHESH DARURI         …
## $ Points <chr> "6.0  ", "6.0  ", "6.0  ", "5.5  ", "5.5  ", "5.0  ", "5.0  ", …
## $ R1     <int> 39, 63, 8, 23, 45, 34, 57, 3, 25, 16, 38, 42, 36, 54, 19, 10, 4…
## $ R2     <int> 21, 58, 61, 28, 37, 29, 46, 32, 18, 19, 56, 33, 27, 44, 16, 15,…
## $ R3     <int> 18, 4, 25, 2, 12, 11, 13, 14, 59, 55, 6, 5, 7, 8, 30, NA, 26, 1…
## $ R4     <int> 14, 17, 21, 26, 13, 35, 11, 9, 8, 31, 7, 38, 5, 1, 22, 39, 2, 3…
## $ R5     <int> 7, 16, 11, 5, 4, 10, 1, 47, 26, 6, 3, NA, 33, 27, 54, 2, 23, 19…
## $ R6     <int> 12, 20, 13, 19, 14, 27, 9, 28, 7, 25, 34, 1, 3, 5, 33, 36, 22, …
## $ R7     <int> 4, 7, 12, 1, 17, 21, 2, 19, 20, 18, 26, 3, 32, 31, 38, NA, 5, 1…
Chess_Data_Joins <- Full_Clean_Data2 |>
  select(!index) |>
  mutate(R1elo = Rating[match(R1, Rank)],
         R2elo = Rating[match(R2, Rank)],
         R3elo = Rating[match(R3, Rank)],
         R4elo = Rating[match(R4, Rank)],
         R5elo = Rating[match(R5, Rank)],
         R6elo = Rating[match(R6, Rank)],
         R7elo = Rating[match(R7, Rank)]) |>
      mutate(
    across(
      c(R1elo, R2elo, R3elo, R4elo, R5elo, R6elo, R7elo),
      ~as.integer(.x)
    )
  ) |>
  rowwise() |>
  mutate(Opponent_Avg_Elo = mean(
      c(R1elo, R2elo, R3elo, R4elo, R5elo, R6elo, R7elo
      )
      , na.rm = TRUE)) |>
  relocate(Name, Area, Points, Rating, Opponent_Avg_Elo) |>
  arrange(Rank) |>
  select(Name, Area, Points, Rating, Opponent_Avg_Elo) |>
  mutate(Opponent_Avg_Elo = round(Opponent_Avg_Elo, digits = 0))

Now for the new calculations below. I first got everything in the right format as integers, then I created new fields for probability, predicted score, and difference. The ELO formula I used was from “The Elo Rating System for Chess and Beyond” by singingbanana on Youtube. the big issue I encountered with this is that my original data did not decode the different chess tournament match outcomes such as B, U, H, and X. These returned as NA’s and I chose to keep them as NA for the sake of this exercise. I will go over more in the conclusion of what could be done to remediate these.

With the newly calculated ratings, I took the differences of these with the player’s actual scores to show who over and under performed.

Chess_Data_Elo <- Full_Clean_Data2 |>
  select(!index) |>
  mutate(R1elo = Rating[match(R1, Rank)],
         R2elo = Rating[match(R2, Rank)],
         R3elo = Rating[match(R3, Rank)],
         R4elo = Rating[match(R4, Rank)],
         R5elo = Rating[match(R5, Rank)],
         R6elo = Rating[match(R6, Rank)],
         R7elo = Rating[match(R7, Rank)]) |>
      mutate(
    across(
      c(R1elo, R2elo, R3elo, R4elo, R5elo, R6elo, R7elo, Rating, Points),
      ~as.integer(.x)
    )) |>
  rowwise() |>
  mutate(R1prob = 1/(1 + 10^((R1elo - Rating)/400)),
         R2prob = 1/(1 + 10^((R2elo - Rating)/400)),
         R3prob = 1/(1 + 10^((R3elo - Rating)/400)),
         R4prob = 1/(1 + 10^((R4elo - Rating)/400)),
         R5prob = 1/(1 + 10^((R5elo - Rating)/400)),
         R6prob = 1/(1 + 10^((R6elo - Rating)/400)),
         R7prob = 1/(1 + 10^((R7elo - Rating)/400)),
         Predicted_Score = sum(across(c(R1prob,R2prob,R3prob,R4prob,R5prob,R6prob,R7prob)), na.rm = TRUE),
         Difference = round((Points - Predicted_Score), digits = 2)
)|>
  arrange(desc(Difference)) |>
  select(!R1:R7prob)

For a bit more of a zoom in:

head(Chess_Data_Elo)
## # A tibble: 6 × 7
## # Rowwise: 
##   Area     Rating  Rank Name                   Points Predicted_Score Difference
##   <chr>     <int> <int> <chr>                   <int>           <dbl>      <dbl>
## 1 "   MI "   1384     3 " ADITYA BAJAJ       …      6          1.95         4.05
## 2 "   MI "   1365    10 " ANVIT RAO          …      5          1.94         3.06
## 3 "   MI "    377    46 " JACOB ALEXANDER LAV…      3          0.0432       2.96
## 4 "   ON "   1411     9 " STEFANO LEE        …      5          2.29         2.71
## 5 "   MI "   1220    15 " ZACHARY JAMES HOUGH…      4          1.37         2.63
## 6 "   MI "    980    37 " AMIYATOSH PWNANANDA…      3          0.773        2.23
tail(Chess_Data_Elo)
## # A tibble: 6 × 7
## # Rowwise: 
##   Area     Rating  Rank Name                   Points Predicted_Score Difference
##   <chr>     <int> <int> <chr>                   <int>           <dbl>      <dbl>
## 1 "   MI "   1449    33 " JADE GE            …      3            4.64      -1.64
## 2 "   MI "   1438    35 " JOSHUA DAVID LEE   …      3            4.96      -1.96
## 3 "   MI "   1332    42 " JARED GE           …      3            5.01      -2.01
## 4 "   MI "   1494    31 " RISHI SHETTY       …      3            5.09      -2.09
## 5 "   ON "   1522    30 " GEORGE AVERY JONES …      3            6.02      -3.02
## 6 "   MI "   1745    25 " LOREN SCHWIEBERT   …      3            6.28      -3.28

Conclusion

The ELO calculation is more intuitive than you would originally think, given that the probability of winning then becomes the points that you can use for theoretical placings of different players. There is no roundabout way to have to extrapolate the score past that point and I think that is a testament to how effective and popular of a rating style the Elo method is.

For how this assignment played out, it was nice to work more with this set of data as I thought it was inherently very interesting. I enjoy chess myself, so seeing how tournaments count scores and diving into the mechanics of those rankings was very neat for me. As I mentioned earlier, the one gap left open for me with this assignment is those NA values. What I would have to do to fix these is go back and assign a set value for the different U, X, B, H, for which some are assumed to be .5 and some are assumed to be 0. This could have an impact on the top and bottom 5 performers, but upon visual inspection it didn’t effect the top players as much as it did the lower ranks.

From my data, players like Aditya Bajaj, Anvit Rao, Jacob Lavalley, Stefano Lee, and Zachary Houghton has outstanding performances, while Loren Schwiebert, George Jones , Rishi Shetty, Jared Ge, and Joshua Lee were upset in the tournament.