The objective of Project1 is to to extract data from the given text file https://brightspace.cuny.edu/content/enforced/1344367-SPS01_DATA_607_1269_1_34036/tournamentinfo.txt?ou=1344367 and generates a .csv file. The required information includes the Player’s Name, Player’s State, Total Number of Points, Player’s Pre-Rating, and the Average Pre-Chess Rating of Opponents.
Following is my approach to accomplish the above objective.
Step 1: First, I will read the given text file into R.
Step 2: I will examine the structure of the text file and determine which regular expressions (regex) are needed to extract the required information.
Step 3: I will use str_extract() and regular expressions to extract information such as Player’s Name, Player’s State, Total Number of Points, and Player’s Pre-Rating. I will then create appropriate variable/column names for the extracted data.
Step 4: Calculate Average Pre-Chess Rating of Opponents.
Step4.1: I will identify each player’s opponents using the opponent numbers provided for each round and match those opponents with their Pre-Ratings.
Step 4.2: I will calculate each player’s Average Pre-Chess Rating of Opponents by dividing the total Pre-Rating of the opponents by the number of games played.
Step 5: There may be NA values in the dataset. I will identify and handle the missing values appropriately during the data-cleaning and calculation process.
Step 6: Finally, I will generate the .csv file using the write_csv() function in R.
Anticipated Data challenges
Extracting and organizing data from an unstructured text file.
Matching each opponent to the correct player in order to calculate the Average Pre-Chess Rating of Opponents.
I converted text into a tibble with one column called text. I remove separator lines containing dashes and remove the first two rows because they contain information that is not needed for the analysis. I grouped the data by player and combined the two rows for each player into one row.
library(tidyverse)
## ── Attaching core tidyverse packages ──────────────────────── tidyverse 2.0.0 ──
## ✔ dplyr 1.2.0 ✔ readr 2.2.0
## ✔ forcats 1.0.1 ✔ stringr 1.6.0
## ✔ ggplot2 4.0.3 ✔ tibble 3.3.1
## ✔ lubridate 1.9.5 ✔ tidyr 1.3.2
## ✔ purrr 1.2.1
## ── Conflicts ────────────────────────────────────────── tidyverse_conflicts() ──
## ✖ dplyr::filter() masks stats::filter()
## ✖ dplyr::lag() masks stats::lag()
## ℹ Use the conflicted package (<http://conflicted.r-lib.org/>) to force all conflicts to become errors
text<-read_lines("https://raw.githubusercontent.com/lhamo07/Data-607-Assignment/refs/heads/Project1/tournamentinfo.txt")
tournamentInfo_df <- tibble(text = text)%>%
filter(!str_detect(text, "^\\s*-+\\s*$")) %>%
slice(-(1:2))%>%
mutate(
person = ceiling(row_number() / 2)
) %>%
group_by(person) %>%
summarise(
text = paste(text, collapse = " "),
.groups = "drop"
) %>%
# I use regular expressions with str_extract() to extract the player number, player name, state, total points, and pre-tournament rating from the combined text. I use as.numeric() to convert the player number, total points, and pre-rating from character values to numeric values. After extracting the required information, I remove the temporary text and person columns because they are no longer needed.
mutate(
Player_num=as.numeric(str_extract(text,"^\\s*(\\d+)")),
Player_name=str_extract(text, "(?<=\\|\\s)[A-Z]+(?:[-\\s]+[A-Z]+)+(?=\\s*\\|)"
),
Player_state = str_extract(text, "[A-Z]{2}(?=\\s*\\|\\s*\\d{8})"),
Total_Points=as.numeric(str_extract(text,"(?<=\\|)\\s*\\d+\\.\\d+")),
Player_pre_Rating=as.numeric(str_extract(text,"(?<=R:)\\s*\\d+")))%>%
select(-text)%>%
select(-person)
tournamentInfo_df
## # A tibble: 64 × 5
## Player_num Player_name Player_state Total_Points Player_pre_Rating
## <dbl> <chr> <chr> <dbl> <dbl>
## 1 1 GARY HUA ON 6 1794
## 2 2 DAKSHESH DARURI MI 6 1553
## 3 3 ADITYA BAJAJ MI 6 1384
## 4 4 PATRICK H SCHILLING MI 5.5 1716
## 5 5 HANSHI ZUO MI 5.5 1655
## 6 6 HANSEN SONG OH 5 1686
## 7 7 GARY DEE SWATHELL MI 5 1649
## 8 8 EZEKIEL HOUGHTON MI 5 1641
## 9 9 STEFANO LEE ON 5 1411
## 10 10 ANVIT RAO MI 5 1365
## # ℹ 54 more rows
glimpse(tournamentInfo_df)
## Rows: 64
## Columns: 5
## $ Player_num <dbl> 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 1…
## $ Player_name <chr> "GARY HUA", "DAKSHESH DARURI", "ADITYA BAJAJ", "PATR…
## $ Player_state <chr> "ON", "MI", "MI", "MI", "MI", "OH", "MI", "MI", "ON"…
## $ Total_Points <dbl> 6.0, 6.0, 6.0, 5.5, 5.5, 5.0, 5.0, 5.0, 5.0, 5.0, 4.…
## $ Player_pre_Rating <dbl> 1794, 1553, 1384, 1716, 1655, 1686, 1649, 1641, 1411…
Extracting Players opponent number was quite challenging so to make it simple, I created another data frame called opponents_df to extract Opponent_Player_num along with their game results. The str_extract_all() function is used because each player can have multiple opponents. unnest_longer() takes those multiple values and puts each opponent on its own row.
opponents_df <- tibble(text = text) %>%
filter(str_detect(text, "^\\s*\\d+\\s*\\|")) %>%
mutate(
Player_num = str_extract(text, "^\\s*\\d+"),
Opponent_Player_num = str_extract_all(text, "[WDL]\\s+\\d+")
) %>%
unnest_longer(Opponent_Player_num) %>%
mutate(
Player_num = as.numeric(Player_num),
Opponent_Player_num = as.numeric(
str_extract(Opponent_Player_num, "\\d+")
)
)
opponents_df <- opponents_df %>%
left_join(
tournamentInfo_df %>%
select(Player_num, Player_pre_Rating),
by = c("Opponent_Player_num" = "Player_num")
)
# I calculated the average pre-rating of each player's opponents and then joined the opponent_mean data frame to tournamentInfo_df using Player_num. The final result was stored in the result data frame.
opponent_mean <- opponents_df %>%
group_by(Player_num) %>%
summarise(
Avg_Opponent_Pre_Rating = round(mean(Player_pre_Rating, na.rm = TRUE),digit=0)
)
opponent_mean
## # A tibble: 64 × 2
## Player_num Avg_Opponent_Pre_Rating
## <dbl> <dbl>
## 1 1 1605
## 2 2 1469
## 3 3 1564
## 4 4 1574
## 5 5 1501
## 6 6 1519
## 7 7 1372
## 8 8 1468
## 9 9 1523
## 10 10 1554
## # ℹ 54 more rows
result <- tournamentInfo_df %>%
left_join(opponent_mean %>% select(Player_num,Avg_Opponent_Pre_Rating), by = "Player_num")
result
## # A tibble: 64 × 6
## Player_num Player_name Player_state Total_Points Player_pre_Rating
## <dbl> <chr> <chr> <dbl> <dbl>
## 1 1 GARY HUA ON 6 1794
## 2 2 DAKSHESH DARURI MI 6 1553
## 3 3 ADITYA BAJAJ MI 6 1384
## 4 4 PATRICK H SCHILLING MI 5.5 1716
## 5 5 HANSHI ZUO MI 5.5 1655
## 6 6 HANSEN SONG OH 5 1686
## 7 7 GARY DEE SWATHELL MI 5 1649
## 8 8 EZEKIEL HOUGHTON MI 5 1641
## 9 9 STEFANO LEE ON 5 1411
## 10 10 ANVIT RAO MI 5 1365
## # ℹ 54 more rows
## # ℹ 1 more variable: Avg_Opponent_Pre_Rating <dbl>
write_csv(result, "chess_tournament.csv")
In this project, I successfully extracted and organized the required chess tournament data, calculated the average pre-rating of each player’s opponents, and exported the results as chess_tournament.csv. This project helped me practice regular expressions, data cleaning, and joining data frames in R. The code could be improved by using functions to make it more organized, reusable, and efficient.
I used ChatGPT to help with this project, particularly with understanding and developing R code for extracting and organizing data from the chess tournament text file. I reviewed, tested, and modified the suggested code to complete the analysis and meet the project requirements. OpenAI. (2026). ChatGPT [Large language model]. https://chatgpt.com. Accessed September 27, 2026.
Chat conversation: https://chatgpt.com/share/6ab99ab2-fc58-83ea-8d69-0e254d8e6227