Show code
# tidyverse: data wrangling, parsing, and charts
library(tidyverse)
# stringr: text matching and cleanup
library(stringr)
# knitr: clean HTML table output
library(knitr)For this project, I will use the chess tournament text file from the course. It is not arranged like a normal CSV because each player uses two lines. The first line has the player’s number, name, points, and round results. The second line has the state and pre-rating.
I will read the file into R and find the lines that start with a player number. I will split those lines at the vertical bars and collect the name, points, and opponent numbers. I will use the next line for the state and pre-rating.
Next, I will match each opponent number with that player’s pre-rating and calculate the average. I will leave out byes and unplayed rounds because they do not have an opponent.
The final dataframe will contain the player’s name, state, total points, pre-rating, and average opponent pre-rating. I will check that there are 64 players, look for missing or duplicate values, compare Gary Hua’s result with the example, and export the table to CSV.
# tidyverse: data wrangling, parsing, and charts
library(tidyverse)
# stringr: text matching and cleanup
library(stringr)
# knitr: clean HTML table output
library(knitr)# Look for the tournament file in common project locations first.
# This avoids using a path that only works on one computer.
local_files <- c(
"tournamentinfo.txt",
file.path("data", "tournamentinfo.txt"),
file.path("Project 1", "tournamentinfo.txt")
)
input_file <- local_files[file.exists(local_files)][1]
# Online fallback: public copy of the same DATA 607 tournament dataset.
data_url <- paste0(
"https://raw.githubusercontent.com/longflin/DATA-607/",
"refs/heads/main/Project%201/tournamentinfo.txt"
)
if (!is.na(input_file)) {
raw_lines <- readLines(input_file, warn = FALSE)
data_source <- input_file
} else {
raw_lines <- tryCatch(
readLines(data_url, warn = FALSE),
error = function(e) {
stop(
"Tournament data could not be loaded. ",
"Put tournamentinfo.txt in the same folder as this QMD ",
"or in a data/ folder, then render again. ",
"The online fallback also could not be reached."
)
}
)
data_source <- data_url
}
cat("Data source:", data_source)Data source: https://raw.githubusercontent.com/longflin/DATA-607/refs/heads/main/Project%201/tournamentinfo.txt
The code checks for a local copy first. If the file is not there, it tries an online copy of the same tournament dataset. This fixes the original working-directory problem while still allowing the course file to be used when it is available.
# Find lines that begin with a player number followed by a vertical bar
player_line_index <- which(
str_detect(raw_lines, "^\\s*\\d+\\s*\\|")
)
length(player_line_index)[1] 64
The first line for each player starts with a player number. I use that pattern to find the player records instead of depending on fixed row numbers.
# Remove extra spaces around parsed text fields
clean_field <- function(x) {
str_squish(x)
}
# Pull a 3- or 4-digit chess rating from the rating field
parse_rating <- function(x) {
rating_match <- str_match(x, "(\\d{3,4})")[, 2]
as.numeric(rating_match)
}
# Parse every player record and combine the results into one dataframe
players <- map_dfr(player_line_index, function(i) {
line1 <- str_split(raw_lines[i], "\\|", simplify = TRUE)
line2 <- str_split(raw_lines[i + 1], "\\|", simplify = TRUE)
player_no <- as.integer(clean_field(line1[1]))
player_name <- clean_field(line1[2])
total_points <- parse_number(clean_field(line1[3]))
state <- clean_field(line2[1])
pre_rating <- parse_rating(clean_field(line2[2]))
round_fields <- if (ncol(line1) >= 4) line1[1, 4:ncol(line1)] else character(0)
opponent_numbers <- str_match(round_fields, "(\\d+)")[, 2] |>
as.integer()
tibble(
player_no = player_no,
name = player_name,
state = state,
total_points = total_points,
pre_rating = pre_rating,
opponent_numbers = list(opponent_numbers)
)
})This separates the fields from the two-line record and stores each player’s opponent numbers for the next step.
dim(players)[1] 64 6
head(players)# A tibble: 6 × 6
player_no name state total_points pre_rating opponent_numbers
<int> <chr> <chr> <dbl> <dbl> <list>
1 1 GARY HUA ON 6 1544 <int [8]>
2 2 DAKSHESH DARURI MI 6 1459 <int [8]>
3 3 ADITYA BAJAJ MI 6 1495 <int [8]>
4 4 PATRICK H SCHILLING MI 5.5 1261 <int [8]>
5 5 HANSHI ZUO MI 5.5 1460 <int [8]>
6 6 HANSEN SONG OH 5 1505 <int [8]>
sum(is.na(players$name))[1] 0
sum(is.na(players$pre_rating))[1] 0
sum(duplicated(players$player_no))[1] 0
The project description I am using expects 64 players. I also check for missing names, missing ratings, and duplicate player numbers before calculating opponent averages.
stopifnot(nrow(players) == 64)
stopifnot(!anyDuplicated(players$player_no))If one of these checks fails, the code stops instead of silently producing a bad result.
# Create a lookup table: player number -> pre-rating
rating_lookup <- players |>
select(player_no, opponent_pre_rating = pre_rating)
# Expand opponent numbers, join opponent ratings, then average by player
final_data <- players |>
select(
player_no,
name,
state,
total_points,
pre_rating,
opponent_numbers
) |>
unnest_longer(
opponent_numbers,
values_to = "opponent_no",
keep_empty = TRUE
) |>
filter(!is.na(opponent_no)) |>
left_join(
rating_lookup,
by = c("opponent_no" = "player_no")
) |>
group_by(
player_no,
name,
state,
total_points,
pre_rating
) |>
summarise(
average_opponent_pre_rating =
round(mean(opponent_pre_rating, na.rm = TRUE)),
.groups = "drop"
) |>
arrange(player_no)Byes and unplayed rounds do not produce an opponent number, so they are left out of the average.
# Keep and rename only the five fields used in the final result
project1_table <- final_data |>
transmute(
Name = name,
State = state,
`Total Number of Points` = total_points,
`Pre-Rating` = pre_rating,
`Average Pre Tournament Chess Rating of Opponents` =
average_opponent_pre_rating
)
kable(
project1_table,
caption = "Project 1 Chess Tournament Results"
)| Name | State | Total Number of Points | Pre-Rating | Average Pre Tournament Chess Rating of Opponents |
|---|---|---|---|---|
| GARY HUA | ON | 6.0 | 1544 | 1233 |
| DAKSHESH DARURI | MI | 6.0 | 1459 | 1266 |
| ADITYA BAJAJ | MI | 6.0 | 1495 | 1417 |
| PATRICK H SCHILLING | MI | 5.5 | 1261 | 1493 |
| HANSHI ZUO | MI | 5.5 | 1460 | 1309 |
| HANSEN SONG | OH | 5.0 | 1505 | 1452 |
| GARY DEE SWATHELL | MI | 5.0 | 1114 | 1475 |
| EZEKIEL HOUGHTON | MI | 5.0 | 1514 | 1377 |
| STEFANO LEE | ON | 5.0 | 1495 | 1323 |
| ANVIT RAO | MI | 5.0 | 1415 | 1351 |
| CAMERON WILLIAM MC LEMAN | MI | 4.5 | 1258 | 1444 |
| KENNETH J TACK | MI | 4.5 | 1268 | 1487 |
| TORRANCE HENRY JR | MI | 4.5 | 1508 | 1410 |
| BRADLEY SHAW | MI | 4.5 | 1013 | 1470 |
| ZACHARY JAMES HOUGHTON | MI | 4.5 | 1561 | 1325 |
| MIKE NIKITIN | MI | 4.0 | 1029 | 1441 |
| RONALD GRZEGORCZYK | MI | 4.0 | 1029 | 1455 |
| DAVID SUNDEEN | MI | 4.0 | 1134 | 1444 |
| DIPANKAR ROY | MI | 4.0 | 1486 | 1409 |
| JASON ZHENG | MI | 4.0 | 1452 | 1466 |
| DINH DANG BUI | ON | 4.0 | 1549 | 1427 |
| EUGENE L MCCLURE | MI | 4.0 | 1240 | 1426 |
| ALAN BUI | ON | 4.0 | 1503 | 1388 |
| MICHAEL R ALDRICH | MI | 4.0 | 1346 | 1383 |
| LOREN SCHWIEBERT | MI | 3.5 | 1248 | 1395 |
| MAX ZHU | ON | 3.5 | 1513 | 1313 |
| GAURAV GIDWANI | MI | 3.5 | 1447 | 1440 |
| SOFIA ADINA STANESCU-BELLU | MI | 3.5 | 1488 | 1397 |
| CHIEDOZIE OKORIE | MI | 3.5 | 1532 | 1487 |
| GEORGE AVERY JONES | ON | 3.5 | 1257 | 1523 |
| RISHI SHETTY | MI | 3.5 | 1513 | 1382 |
| JOSHUA PHILIP MATHEWS | ON | 3.5 | 1407 | 1471 |
| JADE GE | MI | 3.5 | 1469 | 1467 |
| MICHAEL JEFFERY THOMAS | MI | 3.5 | 1505 | 1434 |
| JOSHUA DAVID LEE | MI | 3.5 | 1460 | 1496 |
| SIDDHARTH JHA | MI | 3.5 | 1477 | 1421 |
| AMIYATOSH PWNANANDAM | MI | 3.5 | 1548 | 1498 |
| BRIAN LIU | MI | 3.0 | 1510 | 1369 |
| JOEL R HENDON | MI | 3.0 | 1292 | 1396 |
| FOREST ZHANG | MI | 3.0 | 1489 | 1400 |
| KYLE WILLIAM MURPHY | MI | 3.0 | 1576 | 1309 |
| JARED GE | MI | 3.0 | 1446 | 1465 |
| ROBERT GLEN VASEY | MI | 3.0 | 1410 | 1468 |
| JUSTIN D SCHILLING | MI | 3.0 | 1532 | 1266 |
| DEREK YAN | MI | 3.0 | 1537 | 1489 |
| JACOB ALEXANDER LAVALLEY | MI | 3.0 | 1549 | 1416 |
| ERIC WRIGHT | MI | 2.5 | 1253 | 1413 |
| DANIEL KHAIN | MI | 2.5 | 1436 | 1403 |
| MICHAEL J MARTIN | MI | 2.5 | 1253 | 1488 |
| SHIVAM JHA | MI | 2.5 | 1477 | 1461 |
| TEJAS AYYAGARI | MI | 2.5 | 1520 | 1443 |
| ETHAN GUO | MI | 2.5 | 1491 | 1417 |
| JOSE C YBARRA | MI | 2.0 | 1257 | 1430 |
| LARRY HODGE | MI | 2.0 | 1283 | 1371 |
| ALEX KONG | MI | 2.0 | 1541 | 1442 |
| MARISA RICCI | MI | 2.0 | 1467 | 1438 |
| MICHAEL LU | MI | 2.0 | 1511 | 1379 |
| VIRAJ MOHILE | MI | 2.0 | 1470 | 1474 |
| SEAN M MC CORMICK | MI | 2.0 | 1284 | 1464 |
| JULIA SHEN | MI | 1.5 | 1457 | 1461 |
| JEZZEL FARKAS | ON | 1.5 | 1577 | 1384 |
| ASHWIN BALAJI | MI | 1.0 | 1521 | 1541 |
| THOMAS JOSEPH HOSMER | MI | 1.0 | 1505 | 1419 |
| BEN LI | MI | 1.0 | 1500 | 1363 |
This is the table I will export for the project.
nrow(project1_table)[1] 64
colSums(is.na(project1_table)) Name
0
State
0
Total Number of Points
0
Pre-Rating
0
Average Pre Tournament Chess Rating of Opponents
0
sum(duplicated(project1_table$Name))[1] 0
summary(project1_table) Name State Total Number of Points Pre-Rating
Length :64 Length :64 Min. :1.000 Min. :1013
N.unique :64 N.unique : 3 1st Qu.:2.500 1st Qu.:1290
N.blank : 0 N.blank : 0 Median :3.500 Median :1474
Min.nchar: 6 Min.nchar: 2 Mean :3.438 Mean :1416
Max.nchar:26 Max.nchar: 2 3rd Qu.:4.000 3rd Qu.:1512
Max. :6.000 Max. :1577
Average Pre Tournament Chess Rating of Opponents
Min. :1233
1st Qu.:1384
Median :1426
Mean :1418
3rd Qu.:1465
Max. :1541
I use these checks to make sure the final table has the expected number of rows and to catch missing or duplicate values.
gary_hua_check <- project1_table |>
filter(str_to_lower(Name) == "gary hua")
gary_hua_check# A tibble: 1 × 5
Name State `Total Number of Points` `Pre-Rating` Average Pre Tournament …¹
<chr> <chr> <dbl> <dbl> <dbl>
1 GARY HUA ON 6 1544 1233
# ℹ abbreviated name: ¹`Average Pre Tournament Chess Rating of Opponents`
I included this check because the approach says I will compare Gary Hua’s result with the course example. I will make that comparison from the rendered output rather than typing a result into the document before the code runs.
These charts are secondary to the required table. I kept them simple so they help me inspect the cleaned data without taking over the project.
ggplot(project1_table, aes(x = `Pre-Rating`)) +
geom_histogram(
bins = 12,
fill = "steelblue",
color = "white"
) +
labs(
title = "Distribution of Player Pre-Ratings",
x = "Pre-Rating",
y = "Number of Players"
) +
theme_minimal()The histogram gives a quick view of how the player ratings are distributed.
# Order players by points and use pre-rating to break ties.
# The taller figure gives all player names enough room to be readable.
points_plot <- project1_table |>
arrange(`Total Number of Points`, `Pre-Rating`) |>
mutate(Name = factor(Name, levels = Name))
ggplot(
points_plot,
aes(x = `Total Number of Points`, y = Name)
) +
geom_segment(
aes(
x = 0,
xend = `Total Number of Points`,
y = Name,
yend = Name
),
color = "grey75",
linewidth = 0.6
) +
geom_point(
size = 3,
color = "steelblue"
) +
labs(
title = "Tournament Points by Player",
subtitle = "Players are ordered by total tournament points",
x = "Total Points",
y = NULL
) +
scale_x_continuous(
breaks = seq(
0,
max(points_plot$`Total Number of Points`, na.rm = TRUE),
by = 1
)
) +
theme_minimal(base_size = 11) +
theme(
axis.text.y = element_text(size = 8),
panel.grid.major.y = element_blank(),
panel.grid.minor = element_blank(),
plot.title = element_text(face = "bold")
)The taller lollipop chart makes the player names readable. Players are ordered by total points, with pre-rating used to break ties. This chart is supporting analysis; the final project table remains the main output. ### Player Rating and Opponent Rating
ggplot(
project1_table,
aes(
x = `Pre-Rating`,
y = `Average Pre Tournament Chess Rating of Opponents`
)
) +
geom_point(
alpha = 0.75,
color = "steelblue"
) +
labs(
title = "Player Rating vs. Average Opponent Rating",
x = "Player Pre-Rating",
y = "Average Opponent Pre-Rating"
) +
theme_minimal()This scatterplot compares two numeric variables from the final dataset.
# Save the final table as a CSV file
write_csv(
project1_table,
"project1_chess_results.csv"
)The exported CSV contains only the five requested result columns.
I used GitHub Copilot in Visual Studio Code to help review parts of the R code, troubleshoot errors, and improve code comments. I reviewed the suggestions, kept the code I understood, and checked the results when running the project. The analysis and final results come from the R code and tournament data used in this project.
DATA 607 chess tournament dataset (public copy used only as a fallback):
https://raw.githubusercontent.com/longflin/DATA-607/refs/heads/main/Project%201/tournamentinfo.txt
Wickham, H., Cetinkaya-Rundel, M., & Grolemund, G. R for Data Science (2e).
R Graph Gallery. Lollipop plot examples.
R Graph Gallery. Histogram examples.