library(tidyverse)
library(stringr)
library(ggplot2)
library(readr)
library(kableExtra)Chess Tournament (Project 1 final)
Chess Tournament
Objective
We are given a text file containing a chess tournament cross-table where the information has some structure with which we are to transform in order to facilitate data manipulation and calculate each player’s “stats” including total number of points, their pre-rating and their opponent’s pre-rating.
Approach
Prior to doing any coding, I watched the Youtube video, “Reading a Chess Tournament Cross-Table” by Andy Catlin, that was included in the assignment requirements to understand what information a chess tournament cross-table conveys. Once I had an understanding of what the data meant, I was able to conceptualize my approach.
To begin, I will import it into R, study it and then use appropriate appropriate functions from the tidyverse package, including but not limited to stringr and tidyr to manipulate the data and parse out extraneous data/characters. Other functions such as skimr and ggrepel may be used to transform the data into something more suitable for exporting into a .csv file with one row per player with columns that show :
- Player’s Name, Player’s State, Total number of Points, Player’s Pre-Rating, Average Pre Chess Rating of Opponents
For example, for the information for first player would be:
Gary Hua, ON, 6.0, 1794, 1605
where 1794 represents his rating and his opponent’s rating of 1605 which was calculated by using their pre-tournament ratings of 1436, 1563, 1600, 1610, 1649, 1663, 1716, and dividing by the total number of games played.
Load libraries used
Read file from Github
Using “readlines”, I imported the flat file, “tournamentinfo.txt”, into RStudio. I wrapped the entire import statement in with “suppressWarnings” to eliminate warning regarding that about last line being all dashes.
chess_tbl = suppressWarnings(readLines("https://raw.githubusercontent.com/carolc57/Data607/main/tournamentinfo.txt"))
head(chess_tbl, 10) [1] "-----------------------------------------------------------------------------------------"
[2] " Pair | Player Name |Total|Round|Round|Round|Round|Round|Round|Round| "
[3] " Num | USCF ID / Rtg (Pre->Post) | Pts | 1 | 2 | 3 | 4 | 5 | 6 | 7 | "
[4] "-----------------------------------------------------------------------------------------"
[5] " 1 | GARY HUA |6.0 |W 39|W 21|W 18|W 14|W 7|D 12|D 4|"
[6] " ON | 15445895 / R: 1794 ->1817 |N:2 |W |B |W |B |W |B |W |"
[7] "-----------------------------------------------------------------------------------------"
[8] " 2 | DAKSHESH DARURI |6.0 |W 63|W 58|L 4|W 17|W 16|W 20|W 7|"
[9] " MI | 14598900 / R: 1553 ->1663 |N:2 |B |W |B |W |B |W |B |"
[10] "-----------------------------------------------------------------------------------------"
I then removed all lines that contained dashes using.
chess_tbl_mod <- str_replace_all(string = chess_tbl, pattern = "^-+$", "")
head (chess_tbl_mod)[1] ""
[2] " Pair | Player Name |Total|Round|Round|Round|Round|Round|Round|Round| "
[3] " Num | USCF ID / Rtg (Pre->Post) | Pts | 1 | 2 | 3 | 4 | 5 | 6 | 7 | "
[4] ""
[5] " 1 | GARY HUA |6.0 |W 39|W 21|W 18|W 14|W 7|D 12|D 4|"
[6] " ON | 15445895 / R: 1794 ->1817 |N:2 |W |B |W |B |W |B |W |"
Our table is looking better but still needs some refinement. Let’s remove the empty rows.
chess_tbl_mod <- chess_tbl[sapply(chess_tbl_mod, nchar) > 0]
head(chess_tbl_mod)[1] " Pair | Player Name |Total|Round|Round|Round|Round|Round|Round|Round| "
[2] " Num | USCF ID / Rtg (Pre->Post) | Pts | 1 | 2 | 3 | 4 | 5 | 6 | 7 | "
[3] " 1 | GARY HUA |6.0 |W 39|W 21|W 18|W 14|W 7|D 12|D 4|"
[4] " ON | 15445895 / R: 1794 ->1817 |N:2 |W |B |W |B |W |B |W |"
[5] " 2 | DAKSHESH DARURI |6.0 |W 63|W 58|L 4|W 17|W 16|W 20|W 7|"
[6] " MI | 14598900 / R: 1553 ->1663 |N:2 |B |W |B |W |B |W |B |"
I decided to remove the header rows to facilitate data manipulation.
chess_tbl_mod <- chess_tbl_mod[-(1:2)] ##remove column headings because interfere w/ later operations
head(chess_tbl_mod, 8)[1] " 1 | GARY HUA |6.0 |W 39|W 21|W 18|W 14|W 7|D 12|D 4|"
[2] " ON | 15445895 / R: 1794 ->1817 |N:2 |W |B |W |B |W |B |W |"
[3] " 2 | DAKSHESH DARURI |6.0 |W 63|W 58|L 4|W 17|W 16|W 20|W 7|"
[4] " MI | 14598900 / R: 1553 ->1663 |N:2 |B |W |B |W |B |W |B |"
[5] " 3 | ADITYA BAJAJ |6.0 |L 8|W 61|W 25|W 21|W 11|W 13|W 12|"
[6] " MI | 14959604 / R: 1384 ->1640 |N:2 |W |B |W |B |W |B |W |"
[7] " 4 | PATRICK H SCHILLING |5.5 |W 23|D 28|W 2|W 26|D 5|W 19|D 1|"
[8] " MI | 12616049 / R: 1716 ->1744 |N:2 |W |B |W |B |W |B |B |"
Since the data for one competitor is spread between two rows, I created two separate vectors, one for even rows and another for odd rows, then paste them together
from https://stackoverflow.com/questions/24440258/selecting-multiple-odd-or-even-columns-rows-for-dataframe
chess_tbl_mod_odd = chess_tbl_mod[seq(1,128,2)]
chess_tbl_mod_even = chess_tbl_mod[seq(2,128,2)]This code will combine the odd and even vectors to create one unified vector, thus each player’s data is now on a single row.
chess_tbl_mod_combined <- paste(chess_tbl_mod_odd, chess_tbl_mod_even)
head(chess_tbl_mod_combined, 5)[1] " 1 | GARY HUA |6.0 |W 39|W 21|W 18|W 14|W 7|D 12|D 4| ON | 15445895 / R: 1794 ->1817 |N:2 |W |B |W |B |W |B |W |"
[2] " 2 | DAKSHESH DARURI |6.0 |W 63|W 58|L 4|W 17|W 16|W 20|W 7| MI | 14598900 / R: 1553 ->1663 |N:2 |B |W |B |W |B |W |B |"
[3] " 3 | ADITYA BAJAJ |6.0 |L 8|W 61|W 25|W 21|W 11|W 13|W 12| MI | 14959604 / R: 1384 ->1640 |N:2 |W |B |W |B |W |B |W |"
[4] " 4 | PATRICK H SCHILLING |5.5 |W 23|D 28|W 2|W 26|D 5|W 19|D 1| MI | 12616049 / R: 1716 ->1744 |N:2 |W |B |W |B |W |B |B |"
[5] " 5 | HANSHI ZUO |5.5 |W 45|W 37|D 12|D 13|D 4|W 14|W 17| MI | 14601533 / R: 1655 ->1690 |N:2 |B |W |B |W |B |W |B |"
Now that we have the data for each player, let’s transform it using various stringr and regex functions to parse out the required fields for the final results.
#change file name to preserve data prior to manipulation
chess_final <- chess_tbl_mod_combined First let’s get each player’s name
Name <- str_extract(string = chess_final, pattern = "\\s([[:alpha:] ]{5,})\\b\\s")
head(Name, 5)[1] " GARY HUA " " DAKSHESH DARURI " " ADITYA BAJAJ "
[4] " PATRICK H SCHILLING " " HANSHI ZUO "
Obtain total points for each player
Player_Points <- str_extract(string = chess_final, pattern = "[0-9]\\.[0-9]")
head (Player_Points)[1] "6.0" "6.0" "6.0" "5.5" "5.5" "5.0"
Get State code
#need to remove | and surrounding space around two letter State code
State <- unlist(str_extract_all(chess_final, "\\|[[:space:]]{1,}[[A-Z]]{2} \\|"))
State <- str_replace_all(State, pattern = "(\\|[[:space:]]{1,})|([[:space:]]{1,}\\|)", replacement = "")
head(State)[1] "ON" "MI" "MI" "MI" "MI" "OH"
Player_rating <- str_extract(string = chess_final, pattern = "\\s\\d{3,4}[^\\d]")
#extract player numerical rating (some have a P in them)
Player_rating <- as.integer(str_extract(Player_rating, "\\d+"))
head(Player_rating)[1] 1794 1553 1384 1716 1655 1686
Get Player number
#extract player number
PlayerNumTemp <- as.integer(str_extract(chess_final, "\\d+"))
PlayerNum <- subset(c(PlayerNumTemp), c(PlayerNumTemp)!="0",64)
head(PlayerNum)[1] 1 2 3 4 5 6
Get USCF_id
USCF_id <- str_extract(string = chess_final, pattern ="[[:digit:]]{8}")
head(USCF_id)[1] "15445895" "14598900" "14959604" "12616049" "14601533" "15055204"
Opponent Data - parse player’s opponents rating per round
I used Copilot for part of this section. My original code,
“Opp_Id <-str_extract_all(str_extract_all(chess_final,”\d+\|“),”\d+“) Opp_Id <-subset(c(Opp_Id), c(Opp_Id)!=”0”) head(Opp_Id)”
gave me this warning,
“Warning in stri_extract_all_regex(string, pattern, simplify = simplify, : argument is not an atomic vector; coercing”
which I didn’t understand. I prompted Copilot to “Explain this warning?
I used the code below as suggested to obtain the opponent’s data.
# Extract only digits that are followed by a pipe (|)
Opp_Id <- str_extract_all(chess_final, "\\d+(?=\\|)")
#Flatten the list into a single vector and remove "0"
Opp_Id_vector <- unlist(Opp_Id)
Opp_Id_clean <- Opp_Id_vector[Opp_Id_vector != "0"]
head(Opp_Id_clean)[1] "39" "21" "18" "14" "7" "12"
Calculate opponent average ratings using a for loop
#calculate opponent average ratings
x<-length(chess_tbl_mod)
OppAvgRtg <-numeric(x/2)
for (i in 1:(x/2)) {
OppAvgRtg[i] <- mean(Player_rating[as.numeric(unlist(Opp_Id[PlayerNum[i]]))])
}
OppAvgRtg <- round((OppAvgRtg),0)
head(OppAvgRtg)[1] 1605 1469 1564 1574 1501 1519
Create a chess_ratings data frame
Now that I have all the required elements, I’ll create a ratings data frame
chess_ratings <- data.frame(Name,State,Player_Points,Player_rating, OppAvgRtg)
head(chess_ratings) Name State Player_Points Player_rating OppAvgRtg
1 GARY HUA ON 6.0 1794 1605
2 DAKSHESH DARURI MI 6.0 1553 1469
3 ADITYA BAJAJ MI 6.0 1384 1564
4 PATRICK H SCHILLING MI 5.5 1716 1574
5 HANSHI ZUO MI 5.5 1655 1501
6 HANSEN SONG OH 5.0 1686 1519
Write output to .csv file
Saved data frame as .csv file in my working directory.
write.csv(chess_ratings, file = "Carols Chess Ratings.csv");Final Chess Ratings Table
kable(chess_ratings) %>%
kable_styling(bootstrap_options = "striped", full_width = F, position = "center")| Name | State | Player_Points | Player_rating | OppAvgRtg |
|---|---|---|---|---|
| GARY HUA | ON | 6.0 | 1794 | 1605 |
| DAKSHESH DARURI | MI | 6.0 | 1553 | 1469 |
| ADITYA BAJAJ | MI | 6.0 | 1384 | 1564 |
| PATRICK H SCHILLING | MI | 5.5 | 1716 | 1574 |
| HANSHI ZUO | MI | 5.5 | 1655 | 1501 |
| HANSEN SONG | OH | 5.0 | 1686 | 1519 |
| GARY DEE SWATHELL | MI | 5.0 | 1649 | 1372 |
| EZEKIEL HOUGHTON | MI | 5.0 | 1641 | 1468 |
| STEFANO LEE | ON | 5.0 | 1411 | 1523 |
| ANVIT RAO | MI | 5.0 | 1365 | 1554 |
| CAMERON WILLIAM MC LEMAN | MI | 4.5 | 1712 | 1468 |
| KENNETH J TACK | MI | 4.5 | 1663 | 1506 |
| TORRANCE HENRY JR | MI | 4.5 | 1666 | 1498 |
| BRADLEY SHAW | MI | 4.5 | 1610 | 1515 |
| ZACHARY JAMES HOUGHTON | MI | 4.5 | 1220 | 1484 |
| MIKE NIKITIN | MI | 4.0 | 1604 | 1386 |
| RONALD GRZEGORCZYK | MI | 4.0 | 1629 | 1499 |
| DAVID SUNDEEN | MI | 4.0 | 1600 | 1480 |
| DIPANKAR ROY | MI | 4.0 | 1564 | 1426 |
| JASON ZHENG | MI | 4.0 | 1595 | 1411 |
| DINH DANG BUI | ON | 4.0 | 1563 | 1470 |
| EUGENE L MCCLURE | MI | 4.0 | 1555 | 1300 |
| ALAN BUI | ON | 4.0 | 1363 | 1214 |
| MICHAEL R ALDRICH | MI | 4.0 | 1229 | 1357 |
| LOREN SCHWIEBERT | MI | 3.5 | 1745 | 1363 |
| MAX ZHU | ON | 3.5 | 1579 | 1507 |
| GAURAV GIDWANI | MI | 3.5 | 1552 | 1222 |
| SOFIA ADINA | MI | 3.5 | 1507 | 1522 |
| CHIEDOZIE OKORIE | MI | 3.5 | 1602 | 1314 |
| GEORGE AVERY JONES | ON | 3.5 | 1522 | 1144 |
| RISHI SHETTY | MI | 3.5 | 1494 | 1260 |
| JOSHUA PHILIP MATHEWS | ON | 3.5 | 1441 | 1379 |
| JADE GE | MI | 3.5 | 1449 | 1277 |
| MICHAEL JEFFERY THOMAS | MI | 3.5 | 1399 | 1375 |
| JOSHUA DAVID LEE | MI | 3.5 | 1438 | 1150 |
| SIDDHARTH JHA | MI | 3.5 | 1355 | 1388 |
| AMIYATOSH PWNANANDAM | MI | 3.5 | 980 | 1385 |
| BRIAN LIU | MI | 3.0 | 1423 | 1539 |
| JOEL R HENDON | MI | 3.0 | 1436 | 1430 |
| FOREST ZHANG | MI | 3.0 | 1348 | 1391 |
| KYLE WILLIAM MURPHY | MI | 3.0 | 1403 | 1248 |
| JARED GE | MI | 3.0 | 1332 | 1150 |
| ROBERT GLEN VASEY | MI | 3.0 | 1283 | 1107 |
| JUSTIN D SCHILLING | MI | 3.0 | 1199 | 1327 |
| DEREK YAN | MI | 3.0 | 1242 | 1152 |
| JACOB ALEXANDER LAVALLEY | MI | 3.0 | 377 | 1358 |
| ERIC WRIGHT | MI | 2.5 | 1362 | 1392 |
| DANIEL KHAIN | MI | 2.5 | 1382 | 1356 |
| MICHAEL J MARTIN | MI | 2.5 | 1291 | 1286 |
| SHIVAM JHA | MI | 2.5 | 1056 | 1296 |
| TEJAS AYYAGARI | MI | 2.5 | 1011 | 1356 |
| ETHAN GUO | MI | 2.5 | 935 | 1495 |
| JOSE C YBARRA | MI | 2.0 | 1393 | 1345 |
| LARRY HODGE | MI | 2.0 | 1270 | 1206 |
| ALEX KONG | MI | 2.0 | 1186 | 1406 |
| MARISA RICCI | MI | 2.0 | 1153 | 1414 |
| MICHAEL LU | MI | 2.0 | 1092 | 1363 |
| VIRAJ MOHILE | MI | 2.0 | 917 | 1391 |
| SEAN M MC CORMICK | MI | 2.0 | 853 | 1319 |
| JULIA SHEN | MI | 1.5 | 967 | 1330 |
| JEZZEL FARKAS | ON | 1.5 | 955 | 1327 |
| ASHWIN BALAJI | MI | 1.0 | 1530 | 1186 |
| THOMAS JOSEPH HOSMER | MI | 1.0 | 1175 | 1350 |
| BEN LI | MI | 1.0 | 1163 | 1263 |