Using the player’s pre-ratings and expected scores for each round match up (Project 1), we calculate the difference between expected score and actual score to evaluate under and over performance. The expected score \(E_A\) is calcualted using the following formula:
The tournament data is imported using the same code as used in Project 1. New code begins after the dashed divider.
Data Import
Import the crosstable as plain text are remove lines that are just the separator rows and headers, just leaving lines with real content.
Code
url <-"https://raw.githubusercontent.com/jocslater-code/DATA607/refs/heads/main/Project1/tournamentinfo.txt"raw_txt <-suppressWarnings(readLines(url))# Remove the separator lines and headercontent_lines <- raw_txt[!str_detect(raw_txt, "^\\s*---")] content_lines <- content_lines[3:length(content_lines)]
Parse Data
Each player has data that spans two lines, so we need to combine the information from those lines into one combined string, and then parse that string by the | delimiter and white space. Then we place the matrix of parsed strings into a dataframe.
Code
# Create a grouping variable for line combinationplayer_groups <-rep(1:(length(content_lines) /2), each =2)# Combine the two lines for each player into a single stringcombined_records <-tapply(content_lines, player_groups, paste0, collapse =" | ")# Split by '|' whitespace split_matrix <-str_split(combined_records, "\\|")cleaned_data <-lapply(split_matrix, function(row) str_trim(row))# Convert to a data framedf_all <-as.data.frame(do.call(rbind, cleaned_data), stringsAsFactors =FALSE)
Tidy Dataframe
To only work with necessary data, we create a tidy data frame with just subselected columns, relabeled for clarity.
We create a new column with each player’s pre-tournament rating,
Code
# Parse the rating string to just get the pre-tournament rating# Extract the digits inside the capture group after 'R:'df_tidy$pre_rating <-str_match(df_tidy$rating_str, "R:\\s*(\\d+)")[, 2]# \s is for spaces# (\\d+) grabs all of the numbers up until the next non-digit character like a P, space, or ->df_tidy$pre_rating <-as.numeric(df_tidy$pre_rating)
Then we extracted the opponent number for each round and saved them as seven new columns. Opponent numbers were subbed for NA in cases where the game was anything else then a win, lose, or draw since, for the purposes of our exercize, those were the only games that contributed.
Code
# If the game is not won, lost, or draw, fill with NAdf_tidy <- df_tidy %>%mutate(across(starts_with("R", ignore.case =FALSE), ~ifelse(str_detect(., "[WLD]"), ., NA)))# Create new column that just has the opponent number for each round played# There is a better way to do this with * or \\d df_tidy <- df_tidy %>%mutate(across(starts_with("R1", ignore.case =FALSE), ~as.numeric(str_extract(., "\\d+")), .names ="{.col}_opp_name"))df_tidy <- df_tidy %>%mutate(across(starts_with("R2", ignore.case =FALSE), ~as.numeric(str_extract(., "\\d+")), .names ="{.col}_opp_name"))df_tidy <- df_tidy %>%mutate(across(starts_with("R3", ignore.case =FALSE), ~as.numeric(str_extract(., "\\d+")), .names ="{.col}_opp_name"))df_tidy <- df_tidy %>%mutate(across(starts_with("R4", ignore.case =FALSE), ~as.numeric(str_extract(., "\\d+")), .names ="{.col}_opp_name"))df_tidy <- df_tidy %>%mutate(across(starts_with("R5", ignore.case =FALSE), ~as.numeric(str_extract(., "\\d+")), .names ="{.col}_opp_name"))df_tidy <- df_tidy %>%mutate(across(starts_with("R6", ignore.case =FALSE), ~as.numeric(str_extract(., "\\d+")), .names ="{.col}_opp_name"))df_tidy <- df_tidy %>%mutate(across(starts_with("R7", ignore.case =FALSE), ~as.numeric(str_extract(., "\\d+")), .names ="{.col}_opp_name"))
Once we had an opponent’s name for each round, we then pulled the corresponding pre-rating and saved it as a numeric in another column
Code
# Make new columns for opponent pre-ratingsopp_cols <-paste0("R", 1:7, "_opp_name")# Loop through each opponent column and pull the matching pre_ratingfor (col in opp_cols) { new_col <-paste0(col, "_score") df_tidy[[new_col]] <- df_tidy$pre_rating[df_tidy[[col]]]}
Elo Calculation
Code
# Make new dataframe with only pertinent Elo columns and rename them because this getting long...df_Elo <-select(df_tidy, num, player_name, pre_rating, total_pts, R1_opp_name_score, R2_opp_name_score, R3_opp_name_score, R4_opp_name_score, R5_opp_name_score, R6_opp_name_score, R7_opp_name_score)colnames(df_Elo) <-c("num", "player_name", "pre_rating", "total_pts", "R1_opp_rating", "R2_opp_rating", "R3_opp_rating", "R4_opp_rating", "R5_opp_rating", "R6_opp_rating", "R7_opp_rating")# Calculate expected scoredf_Elo <- df_Elo %>%mutate(across(starts_with("R") &ends_with("_opp_rating"),~1/ (1+10^((.x - pre_rating) /400)),.names ="{.col}_expected" ) )# Rename columns againcolnames(df_Elo) <-c("num", "player_name", "pre_rating", "total_pts", "R1_opp_rating", "R2_opp_rating", "R3_opp_rating", "R4_opp_rating", "R5_opp_rating", "R6_opp_rating", "R7_opp_rating", "R1_EA","R2_EA", "R3_EA", "R4_EA", "R5_EA", "R6_EA", "R7_EA")
Once we have the expected score for each round, we can sum these up to get the total expected score, and compare this with the actual score in a new column.
Code
# Add new column with the sum of all of the expected scores so we can compare this to the actual scoreexpected_cols <-grep("_EA", names(df_Elo), value =TRUE)# Compute total expected score from the rounds for each playerdf_Elo$total_EA <-rowSums(df_Elo[, expected_cols], na.rm =TRUE)df_Elo$diff_EA <-as.numeric(df_Elo$total_pts)-df_Elo$total_EA
Expected Score Analysis
Now that we have the expected scores, actual, and the differences, we can analyze player performance. Specifically, we are looking for the five players who underperformed (largest negative difference) and who overperformed (largest positive difference).
Code
# Count how many players under/overperformeddf_Elo %>%filter(!is.na(diff_EA)) %>%summarise(Positive =sum(diff_EA >0),Negative =sum(diff_EA <0),Zero =sum(diff_EA ==0) )
# Get top 5 largest and top 5 smallest in new dataframetop_bottom_df <- df_Elo %>%filter(!is.na(diff_EA)) %>%slice_max(order_by = diff_EA, n =5) %>%bind_rows( df_Elo %>%filter(!is.na(diff_EA)) %>%slice_min(order_by = diff_EA, n =5) ) top_bottom_df <-select(top_bottom_df, player_name, diff_EA) top_bottom_df <-arrange(top_bottom_df, diff_EA)# Plot this in a sideways bar chartggplot(top_bottom_df, aes(x =reorder(player_name, diff_EA), y = diff_EA, fill = diff_EA >0)) +geom_col() +geom_text(aes(label =round(diff_EA, 2)),hjust =ifelse(top_bottom_df$diff_EA >=0, -0.2, 1.2),size =3.5 ) +coord_flip() +scale_fill_manual(values =c("TRUE"="#2e7d32", "FALSE"="#c62828"),labels =c("TRUE"="Overperformed", "FALSE"="Underperformed"),name ="Performance" ) +labs(title ="Top 5 and Bottom 5 Players",x ="Player Name",y ="Performance Difference" ) +theme_minimal() +theme(panel.grid.minor =element_blank())
Discussion and Next Steps
It seems that for this chess tournament, the majority of players had a good day - 56% of them overperformed, exceeding their expected scores. Aditya Baiaj had the best day of all though, overperforming by a full 4 points!
Further work could be to repeat this analysis with other tournament data to observe trends among other player groups. We could explore if certain demographics for players or factors of the day (time games are played, weather, location) impacts performance. I would also be currious to explore the cumulative effect of each game; if a player wins/loses the prior game, how does that statistically effect the following games?
Citations
Elo Chess Rating Calculator. (2026). Expected score in Elo chess ratings. https://elochessratingcalculator.com/learn/elo/expected-score/
Google DeepMind. (2026). Gemini 3.6 Flash [Large language model]. https://gemini.google.com. Accessed Sept 29, 2026. Transcript available in Github as Assignment5B_Gemini_Transcript.pdf