For this project, I worked with a text file containing chess tournament results. The file contains information about each player’s name, state, total points, pre-tournament rating, and the opponents they played.
I used R to clean and organize the text data. I extracted the important information for each player and used the opponent numbers to find each opponent’s pre-tournament rating. I then calculated the average opponent rating for every player.
Next, I created a clean data frame containing each player’s name, state, total points, pre-rating, and average opponent pre-rating, and exported the results as a CSV file.
I also performed exploratory data analysis to examine the relationships between player ratings, opponent strength, and tournament performance using visualizations.
Finally, I went a step further and created an interactive dashboard called Beyond the Chessboard using R Shiny and Plotly. The dashboard allows users to explore player statistics, compare ratings, analyze tournament performance, and view player rankings by state.
In this project, I analyzed chess tournament results for 64 players using R. My goal was to convert the original text file into a structured dataset containing player information and the average pre-tournament rating of each player’s opponents.
I also checked the accuracy of the extracted data and explored the relationships between player ratings, opponent strength, and tournament performance using visualizations.
##
## Attaching package: 'dplyr'
## The following objects are masked from 'package:stats':
##
## filter, lag
## The following objects are masked from 'package:base':
##
## intersect, setdiff, setequal, union
library(stringr)
url <- "https://raw.githubusercontent.com/Afwans/DATA607-Project1/main/tournamentinfo.txt"
chess_data <- readLines(url)## Warning in readLines(url): incomplete final line found on
## 'https://raw.githubusercontent.com/Afwans/DATA607-Project1/main/tournamentinfo.txt'
## [1] "-----------------------------------------------------------------------------------------"
## [2] " Pair | Player Name |Total|Round|Round|Round|Round|Round|Round|Round| "
## [3] " Num | USCF ID / Rtg (Pre->Post) | Pts | 1 | 2 | 3 | 4 | 5 | 6 | 7 | "
## [4] "-----------------------------------------------------------------------------------------"
## [5] " 1 | GARY HUA |6.0 |W 39|W 21|W 18|W 14|W 7|D 12|D 4|"
## [6] " ON | 15445895 / R: 1794 ->1817 |N:2 |W |B |W |B |W |B |W |"
## [7] "-----------------------------------------------------------------------------------------"
## [8] " 2 | DAKSHESH DARURI |6.0 |W 63|W 58|L 4|W 17|W 16|W 20|W 7|"
## [9] " MI | 14598900 / R: 1553 ->1663 |N:2 |B |W |B |W |B |W |B |"
## [10] "-----------------------------------------------------------------------------------------"
The original tournament file contains player information, column headers, and separator lines. Each player’s information is stored across two lines.
I will first remove the unnecessary headers and separators. This will make it easier to extract the player information and organize it into a structured dataset.
# Remove the first four header lines
player_data <- chess_data[-(1:4)]
# Remove separator lines
player_data <- player_data[
!str_detect(player_data, "^\\s*-+\\s*$")
]
# Check the number of remaining lines
length(player_data)## [1] 128
## [1] " 1 | GARY HUA |6.0 |W 39|W 21|W 18|W 14|W 7|D 12|D 4|"
## [2] " ON | 15445895 / R: 1794 ->1817 |N:2 |W |B |W |B |W |B |W |"
## [3] " 2 | DAKSHESH DARURI |6.0 |W 63|W 58|L 4|W 17|W 16|W 20|W 7|"
## [4] " MI | 14598900 / R: 1553 ->1663 |N:2 |B |W |B |W |B |W |B |"
## [5] " 3 | ADITYA BAJAJ |6.0 |L 8|W 61|W 25|W 21|W 11|W 13|W 12|"
## [6] " MI | 14959604 / R: 1384 ->1640 |N:2 |W |B |W |B |W |B |W |"
After removing the unnecessary lines, I have 128 lines representing 64 players. Each player has two lines of information.
I will separate these lines and extract the player number, name, state, total points, and pre-tournament rating.
# Separate the two lines for each player
player_lines <- player_data[seq(1, length(player_data), by = 2)]
rating_lines <- player_data[seq(2, length(player_data), by = 2)]
# Extract player information
players <- data.frame(
Player_Number = as.integer(
str_trim(str_extract(player_lines, "^\\s*\\d+"))
),
Player_Name = str_trim(
str_match(player_lines, "^\\s*\\d+\\s*\\|([^|]+)")[, 2]
),
State = str_trim(
str_match(rating_lines, "^\\s*([^|]+)\\|")[, 2]
),
Total_Points = as.numeric(
str_trim(str_match(player_lines, "^\\s*\\d+\\s*\\|[^|]+\\|([^|]+)")[, 2])
),
Pre_Rating = as.integer(
str_match(rating_lines, "R:\\s*(\\d+)")[, 2]
)
)
head(players)## Player_Number Player_Name State Total_Points Pre_Rating
## 1 1 GARY HUA ON 6.0 1794
## 2 2 DAKSHESH DARURI MI 6.0 1553
## 3 3 ADITYA BAJAJ MI 6.0 1384
## 4 4 PATRICK H SCHILLING MI 5.5 1716
## 5 5 HANSHI ZUO MI 5.5 1655
## 6 6 HANSEN SONG OH 5.0 1686
## [1] 64
## Player_Number Player_Name State Total_Points Pre_Rating
## 0 0 0 0 0
Before calculating the average opponent ratings, I will check that all 64 players were extracted and that none of the required fields contain missing values.
## [1] 64
## Player_Number Player_Name State Total_Points Pre_Rating
## 0 0 0 0 0
## [1] 0
## [1] 1 64
Each player has seven rounds of tournament results. The results include wins, losses, draws, and sometimes byes or unplayed rounds.
I will extract the opponent numbers from each round. Only rounds containing an actual opponent number will be used to calculate the average opponent rating.
# Extract the seven round columns
round_data <- str_split_fixed(player_lines, "\\|", 11)
# Keep only the seven round results
round_data <- round_data[, 4:10]
# Function to extract opponent numbers
get_opponents <- function(rounds) {
# Extract numbers from rounds with W, L, or D
opponent_numbers <- str_match(
rounds,
"^\\s*[WLD]\\s+(\\d+)\\s*$"
)[, 2]
# Remove rounds without opponents
opponent_numbers <- opponent_numbers[!is.na(opponent_numbers)]
# Convert opponent numbers to integers
as.integer(opponent_numbers)
}
# Extract opponents for every player
opponents <- apply(round_data, 1, get_opponents)
stopifnot(length(opponents) == 64)
stopifnot(
all(unlist(opponents) %in% players$Player_Number)
)
# Display Gary Hua's opponents
opponents[[1]]## [1] 39 21 18 14 7 12 4
##
## 1 3 4 5 6 7
## 1 1 1 7 13 41
## [1] 7
After extracting the opponent numbers, I will match each opponent with their pre-tournament rating.
I will then calculate the average opponent rating for every player. Rounds without an actual opponent will not be included in the calculation.
# Function to calculate average opponent rating
get_average_rating <- function(opponent_ids) {
# Find the pre-ratings of the opponents
ratings <- players$Pre_Rating[
match(opponent_ids, players$Player_Number)
]
# Calculate the average
if (length(ratings) == 0 || any(is.na(ratings))) {
return(NA_real_)
}
round(mean(ratings))
}
# Calculate average opponent ratings for all players
players$Avg_Opponent_Rating <- sapply(
opponents,
get_average_rating
)
# Display the first six players
head(players)## Player_Number Player_Name State Total_Points Pre_Rating
## 1 1 GARY HUA ON 6.0 1794
## 2 2 DAKSHESH DARURI MI 6.0 1553
## 3 3 ADITYA BAJAJ MI 6.0 1384
## 4 4 PATRICK H SCHILLING MI 5.5 1716
## 5 5 HANSHI ZUO MI 5.5 1655
## 6 6 HANSEN SONG OH 5.0 1686
## Avg_Opponent_Rating
## 1 1605
## 2 1469
## 3 1564
## 4 1574
## 5 1501
## 6 1519
# Display Gary Hua's opponent ratings
players$Pre_Rating[
match(opponents[[1]], players$Player_Number)
]## [1] 1436 1563 1600 1610 1649 1663 1716
## [1] 1605
## [1] 0
After extracting and validating the player information, I will create the final dataset containing the five required columns.
I will export this dataset as a CSV file so that it can be used for further analysis.
# Create the final dataset
final_data <- players %>%
select(
Player_Name,
State,
Total_Points,
Pre_Rating,
Avg_Opponent_Rating
)
# Display the first six players
head(final_data)## Player_Name State Total_Points Pre_Rating Avg_Opponent_Rating
## 1 GARY HUA ON 6.0 1794 1605
## 2 DAKSHESH DARURI MI 6.0 1553 1469
## 3 ADITYA BAJAJ MI 6.0 1384 1564
## 4 PATRICK H SCHILLING MI 5.5 1716 1574
## 5 HANSHI ZUO MI 5.5 1655 1501
## 6 HANSEN SONG OH 5.0 1686 1519
## [1] 64 5
# Read the exported CSV
check_data <- read.csv("chess_tournament_clean.csv")
# Confirm the dimensions
dim(check_data)## [1] 64 5
## Player_Name State Total_Points Pre_Rating
## 0 0 0 0
## Avg_Opponent_Rating
## 0
## [1] TRUE
## [1] TRUE
The exported CSV contains 64 rows and five columns, with no missing values. I compared the exported file with the original data frame to verify that the data was saved correctly.
Now that the dataset has been cleaned and validated, I will explore the relationship between player ratings, opponent strength, and tournament performance.
This visualization compares each player’s pre-tournament rating with the average pre-tournament rating of their opponents.
library(ggplot2)
ggplot(final_data, aes(
x = Pre_Rating,
y = Avg_Opponent_Rating
)) +
geom_point(
color = "steelblue",
size = 3,
alpha = 0.7
) +
geom_smooth(
method = "lm",
se = FALSE,
color = "darkred"
) +
labs(
title = "Player Rating vs. Opponent Strength",
subtitle = "Chess Tournament Performance Analysis",
x = "Player Pre-Tournament Rating",
y = "Average Opponent Pre-Tournament Rating",
caption = "Source: Chess Tournament Dataset"
) +
theme_minimal()## `geom_smooth()` using formula = 'y ~ x'
## [1] 0.2839375
The correlation between player pre-tournament ratings and average opponent ratings was approximately 0.284.
This indicates a weak positive relationship. Higher-rated players tended to face slightly stronger opponents, but there was considerable variation.
The scatter plot also shows that players with similar pre-tournament ratings sometimes faced opponents with very different average ratings.
This suggests that a player’s initial rating alone does not fully explain the strength of their opponents in this tournament.
To further explore tournament performance, I will calculate the difference between each player’s pre-tournament rating and the average rating of their opponents.
I will then examine the relationship between this rating difference and total tournament points.
# Calculate the rating difference
analysis_data <- final_data %>%
mutate(
Rating_Difference = Avg_Opponent_Rating - Pre_Rating
)
# Display the first six players
head(analysis_data)## Player_Name State Total_Points Pre_Rating Avg_Opponent_Rating
## 1 GARY HUA ON 6.0 1794 1605
## 2 DAKSHESH DARURI MI 6.0 1553 1469
## 3 ADITYA BAJAJ MI 6.0 1384 1564
## 4 PATRICK H SCHILLING MI 5.5 1716 1574
## 5 HANSHI ZUO MI 5.5 1655 1501
## 6 HANSEN SONG OH 5.0 1686 1519
## Rating_Difference
## 1 -189
## 2 -84
## 3 180
## 4 -142
## 5 -154
## 6 -167
This visualization explores the relationship between opponent strength relative to a player’s rating and the total points earned during the tournament.
ggplot(analysis_data, aes(
x = Rating_Difference,
y = Total_Points
)) +
geom_point(
color = "steelblue",
size = 3,
alpha = 0.7
) +
geom_smooth(
method = "lm",
se = FALSE,
color = "darkred"
) +
geom_vline(
xintercept = 0,
linetype = "dashed",
color = "gray50"
) +
labs(
title = "Opponent Strength vs. Tournament Performance",
subtitle = "Rating Difference and Total Tournament Points",
x = "Rating Difference (Opponent Rating - Player Rating)",
y = "Total Tournament Points",
caption = "Source: Chess Tournament Dataset"
) +
theme_minimal()## `geom_smooth()` using formula = 'y ~ x'
## [1] -0.369523
The correlation between rating difference and total tournament points was approximately -0.37, indicating a moderate negative relationship.
In the graph, the red line represents the linear regression trend. Since the line slopes downward, it shows that players who faced stronger opponents relative to their own ratings generally earned fewer tournament points.
However, the scatter plot shows considerable variation. Some players earned high scores despite facing stronger opponents, while others earned fewer points against comparatively weaker opponents.
These findings describe an association within this tournament and do not establish causation.
This visualization shows how tournament points are distributed among all 64 players. It helps identify the most common scores and understand the overall distribution of tournament results.
# Count the number of players for each score
points_data <- final_data %>%
count(Total_Points)
# Display the results
points_data## Total_Points n
## 1 1.0 3
## 2 1.5 2
## 3 2.0 7
## 4 2.5 6
## 5 3.0 9
## 6 3.5 13
## 7 4.0 9
## 8 4.5 5
## 9 5.0 5
## 10 5.5 2
## 11 6.0 3
# Create a bar chart
ggplot(points_data, aes(
x = factor(Total_Points),
y = n
)) +
geom_col(
fill = "steelblue",
width = 0.7
) +
geom_text(
aes(label = n),
vjust = -0.5,
size = 4
) +
labs(
title = "Distribution of Tournament Points",
subtitle = "Scores Earned by 64 Chess Players",
x = "Total Tournament Points",
y = "Number of Players",
caption = "Source: Chess Tournament Dataset"
) +
theme_minimal()The bar chart shows that 3.5 points was the most common tournament score, earned by 13 players.
Most players earned between 2 and 4 points, with 44 out of 64 players (68.75%) falling within this range.
The results show that tournament scores were concentrated around the middle, with fewer players achieving the highest or lowest scores.
In this project, I used R to clean and analyze chess tournament data for 64 players. I extracted player information from an unstructured text file and calculated the average pre-tournament rating of each player’s opponents.
I validated the dataset by checking for missing values and duplicate player numbers, and by verifying Gary Hua’s average opponent rating.. I then exported the results into a CSV file containing the five required columns.
My analysis revealed three findings:
Player ratings and average opponent ratings had a weak positive correlation of 0.284.
Rating difference and tournament points had a moderate negative correlation of -0.370.
The most common tournament score was 3.5 points, earned by 13 players.
These findings demonstrate how R can transform unstructured text into useful information and support exploratory data analysis.
I also developed Beyond the Chessboard, an interactive Shiny dashboard that allows users to explore tournament results, compare player ratings, and view player rankings by state.
However, these relationships describe only this tournament and should not be interpreted as evidence of causation.