We will be calculating the elo scores of the players from our first project and compare the expected score to the actual elo change that each player experienced. To do this, we will first get the original data from github used in project 1 and perform the same transformations as the original project.
Warning in
readLines("https://raw.githubusercontent.com/Keenan-Roe/DATA-607/main/Project%201/ELO.txt"):
incomplete final line found on
'https://raw.githubusercontent.com/Keenan-Roe/DATA-607/main/Project%201/ELO.txt'
In order to perform our calculations we will create two new data sets that we will use to keep track of the opponents as well as the results of matches separately. We will need to replace the win, lose, and draw indicators with 1, 0, and .5 respectively. This will allow us to us the results data set directly in our calculations later on.
In order to perform elo changes we will be using two formulas derived from this video. The basics of the algorithm states that a player facing an opponent with 400 more elo than them means that the opponent has a 10 times higher chance to win. A K-value of 32 is used which will be how much the elo changes with each match. a higher K-value introduces more volatility and change to elo while a lower K-value will result in lower elo changes overall. The previously produced data set with our win/loss/draw numbers will be used in place of score.
E_value <-function(p2, p1) { E <-1/(1+10^((p2-p1)/400))}point_change <-function( elo, score, E ){ newElo <- elo +32*( score - E )}
Analysis
Now that we have all the pieces we need we can put it all together and even keep track of elo changes in a dataframe. We will create a new data set called elo_results and then update the column starting_elo in our players new starting elo for each match so we will be using their results in the tournament in the calculation. We will need the players elo and their opponent’s elo to calculate the expected score. We will them use that as well as their starting elo and whether they won, lost, or drew to determine the player’s new score. Any NA value can be ignored.
With these results we see that our final score is different from the final elo score stated in the original file, but this can be down to a different k-value or a different calculation altogether. No provisional rating was used in our calculations either. We can start comparing the expected score vs the actual score each player achieved. We can see there was a large spread of performances, but looking at the top ten performing players compared to their expected score we can see that Aditya Bajaj outperformed themselves by almost 4 whole points with each point mapping onto an unexpected win based on the elo differences.
top10 <- score_compare |>arrange(desc(performance)) |>slice_head(n=10)ggplot( top10, aes( x =reorder ( players, -performance ), y = performance)) +geom_col( fill ="firebrick")+theme( axis.text.x =element_text( angle =45, hjust =1 )) +labs( x ="Players")
bottom10 <- score_compare |>arrange(performance) |>slice_head(n=10)ggplot( bottom10, aes( x =reorder ( players, -performance ), y = performance)) +geom_col( fill ="orangered")+theme( axis.text.x =element_text( angle =45, hjust =1 )) +labs( x ="Players")
Conclusion
the results are a fairly even spread when looking at a 5 to 6 game average. The total difference between the top performing player and the lowest performing player is 6.4323848. A max differential in this format would be a number approaching 12 with one player winning all 6 games when they are expected to lose all 6 and another player losing all 6 games when they are expected to win all six games.