Author

Andre Thomson

Approach

For this project, I will use the chess tournament text file from the course. It is not arranged like a normal CSV because each player uses two lines. The first line has the player’s number, name, points, and round results. The second line has the state and pre-rating.

I will read the file into R and find the lines that start with a player number. I will split those lines at the vertical bars and collect the name, points, and opponent numbers. I will use the next line for the state and pre-rating.

Next, I will match each opponent number with that player’s pre-rating and calculate the average. I will leave out byes and unplayed rounds because they do not have an opponent.

The final dataframe will contain the player’s name, state, total points, pre-rating, and average opponent pre-rating. I will check that there are 64 players, look for missing or duplicate values, compare Gary Hua’s result with the example, and export the table to CSV.

Setup

Show code
# tidyverse: data wrangling, parsing, and charts
library(tidyverse)
# stringr: text matching and cleanup
library(stringr)
# knitr: clean HTML table output
library(knitr)

Load the Tournament File

Show code
# Look for the tournament file in common project locations first.
# This avoids using a path that only works on one computer.
local_files <- c(
  "tournamentinfo.txt",
  file.path("data", "tournamentinfo.txt"),
  file.path("Project 1", "tournamentinfo.txt")
)

input_file <- local_files[file.exists(local_files)][1]

# Online fallback: public copy of the same DATA 607 tournament dataset.
data_url <- paste0(
  "https://raw.githubusercontent.com/longflin/DATA-607/",
  "refs/heads/main/Project%201/tournamentinfo.txt"
)

if (!is.na(input_file)) {
  raw_lines <- readLines(input_file, warn = FALSE)
  data_source <- input_file
} else {
  raw_lines <- tryCatch(
    readLines(data_url, warn = FALSE),
    error = function(e) {
      stop(
        "Tournament data could not be loaded. ",
        "Put tournamentinfo.txt in the same folder as this QMD ",
        "or in a data/ folder, then render again. ",
        "The online fallback also could not be reached."
      )
    }
  )
  data_source <- data_url
}

cat("Data source:", data_source)
Data source: https://raw.githubusercontent.com/longflin/DATA-607/refs/heads/main/Project%201/tournamentinfo.txt

The code checks for a local copy first. If the file is not there, it tries an online copy of the same tournament dataset. This fixes the original working-directory problem while still allowing the course file to be used when it is available.

Find the Player Records

Show code
# Find lines that begin with a player number followed by a vertical bar
player_line_index <- which(
  str_detect(raw_lines, "^\\s*\\d+\\s*\\|")
)

length(player_line_index)
[1] 64

The first line for each player starts with a player number. I use that pattern to find the player records instead of depending on fixed row numbers.

Parse the Player Information

Show code
# Remove extra spaces around parsed text fields
clean_field <- function(x) {
  str_squish(x)
}

# Pull a 3- or 4-digit chess rating from the rating field
parse_rating <- function(x) {
  rating_match <- str_match(x, "(\\d{3,4})")[, 2]
  as.numeric(rating_match)
}

# Parse every player record and combine the results into one dataframe
players <- map_dfr(player_line_index, function(i) {

  line1 <- str_split(raw_lines[i], "\\|", simplify = TRUE)
  line2 <- str_split(raw_lines[i + 1], "\\|", simplify = TRUE)

  player_no <- as.integer(clean_field(line1[1]))
  player_name <- clean_field(line1[2])
  total_points <- parse_number(clean_field(line1[3]))

  state <- clean_field(line2[1])
  pre_rating <- parse_rating(clean_field(line2[2]))

  round_fields <- if (ncol(line1) >= 4) line1[1, 4:ncol(line1)] else character(0)

  opponent_numbers <- str_match(round_fields, "(\\d+)")[, 2] |>
    as.integer()

  tibble(
    player_no = player_no,
    name = player_name,
    state = state,
    total_points = total_points,
    pre_rating = pre_rating,
    opponent_numbers = list(opponent_numbers)
  )
})

This separates the fields from the two-line record and stores each player’s opponent numbers for the next step.

Check the Parsed Data

Show code
dim(players)
[1] 64  6
Show code
head(players)
# A tibble: 6 × 6
  player_no name                state total_points pre_rating opponent_numbers
      <int> <chr>               <chr>        <dbl>      <dbl> <list>          
1         1 GARY HUA            ON             6         1544 <int [8]>       
2         2 DAKSHESH DARURI     MI             6         1459 <int [8]>       
3         3 ADITYA BAJAJ        MI             6         1495 <int [8]>       
4         4 PATRICK H SCHILLING MI             5.5       1261 <int [8]>       
5         5 HANSHI ZUO          MI             5.5       1460 <int [8]>       
6         6 HANSEN SONG         OH             5         1505 <int [8]>       
Show code
sum(is.na(players$name))
[1] 0
Show code
sum(is.na(players$pre_rating))
[1] 0
Show code
sum(duplicated(players$player_no))
[1] 0

The project description I am using expects 64 players. I also check for missing names, missing ratings, and duplicate player numbers before calculating opponent averages.

Show code
stopifnot(nrow(players) == 64)
stopifnot(!anyDuplicated(players$player_no))

If one of these checks fails, the code stops instead of silently producing a bad result.

Calculate Average Opponent Pre-Rating

Show code
# Create a lookup table: player number -> pre-rating
rating_lookup <- players |>
  select(player_no, opponent_pre_rating = pre_rating)

# Expand opponent numbers, join opponent ratings, then average by player
final_data <- players |>
  select(
    player_no,
    name,
    state,
    total_points,
    pre_rating,
    opponent_numbers
  ) |>
  unnest_longer(
    opponent_numbers,
    values_to = "opponent_no",
    keep_empty = TRUE
  ) |>
  filter(!is.na(opponent_no)) |>
  left_join(
    rating_lookup,
    by = c("opponent_no" = "player_no")
  ) |>
  group_by(
    player_no,
    name,
    state,
    total_points,
    pre_rating
  ) |>
  summarise(
    average_opponent_pre_rating =
      round(mean(opponent_pre_rating, na.rm = TRUE)),
    .groups = "drop"
  ) |>
  arrange(player_no)

Byes and unplayed rounds do not produce an opponent number, so they are left out of the average.

Final Table

Show code
# Keep and rename only the five fields used in the final result
project1_table <- final_data |>
  transmute(
    Name = name,
    State = state,
    `Total Number of Points` = total_points,
    `Pre-Rating` = pre_rating,
    `Average Pre Tournament Chess Rating of Opponents` =
      average_opponent_pre_rating
  )

kable(
  project1_table,
  caption = "Project 1 Chess Tournament Results"
)
Project 1 Chess Tournament Results
Name State Total Number of Points Pre-Rating Average Pre Tournament Chess Rating of Opponents
GARY HUA ON 6.0 1544 1233
DAKSHESH DARURI MI 6.0 1459 1266
ADITYA BAJAJ MI 6.0 1495 1417
PATRICK H SCHILLING MI 5.5 1261 1493
HANSHI ZUO MI 5.5 1460 1309
HANSEN SONG OH 5.0 1505 1452
GARY DEE SWATHELL MI 5.0 1114 1475
EZEKIEL HOUGHTON MI 5.0 1514 1377
STEFANO LEE ON 5.0 1495 1323
ANVIT RAO MI 5.0 1415 1351
CAMERON WILLIAM MC LEMAN MI 4.5 1258 1444
KENNETH J TACK MI 4.5 1268 1487
TORRANCE HENRY JR MI 4.5 1508 1410
BRADLEY SHAW MI 4.5 1013 1470
ZACHARY JAMES HOUGHTON MI 4.5 1561 1325
MIKE NIKITIN MI 4.0 1029 1441
RONALD GRZEGORCZYK MI 4.0 1029 1455
DAVID SUNDEEN MI 4.0 1134 1444
DIPANKAR ROY MI 4.0 1486 1409
JASON ZHENG MI 4.0 1452 1466
DINH DANG BUI ON 4.0 1549 1427
EUGENE L MCCLURE MI 4.0 1240 1426
ALAN BUI ON 4.0 1503 1388
MICHAEL R ALDRICH MI 4.0 1346 1383
LOREN SCHWIEBERT MI 3.5 1248 1395
MAX ZHU ON 3.5 1513 1313
GAURAV GIDWANI MI 3.5 1447 1440
SOFIA ADINA STANESCU-BELLU MI 3.5 1488 1397
CHIEDOZIE OKORIE MI 3.5 1532 1487
GEORGE AVERY JONES ON 3.5 1257 1523
RISHI SHETTY MI 3.5 1513 1382
JOSHUA PHILIP MATHEWS ON 3.5 1407 1471
JADE GE MI 3.5 1469 1467
MICHAEL JEFFERY THOMAS MI 3.5 1505 1434
JOSHUA DAVID LEE MI 3.5 1460 1496
SIDDHARTH JHA MI 3.5 1477 1421
AMIYATOSH PWNANANDAM MI 3.5 1548 1498
BRIAN LIU MI 3.0 1510 1369
JOEL R HENDON MI 3.0 1292 1396
FOREST ZHANG MI 3.0 1489 1400
KYLE WILLIAM MURPHY MI 3.0 1576 1309
JARED GE MI 3.0 1446 1465
ROBERT GLEN VASEY MI 3.0 1410 1468
JUSTIN D SCHILLING MI 3.0 1532 1266
DEREK YAN MI 3.0 1537 1489
JACOB ALEXANDER LAVALLEY MI 3.0 1549 1416
ERIC WRIGHT MI 2.5 1253 1413
DANIEL KHAIN MI 2.5 1436 1403
MICHAEL J MARTIN MI 2.5 1253 1488
SHIVAM JHA MI 2.5 1477 1461
TEJAS AYYAGARI MI 2.5 1520 1443
ETHAN GUO MI 2.5 1491 1417
JOSE C YBARRA MI 2.0 1257 1430
LARRY HODGE MI 2.0 1283 1371
ALEX KONG MI 2.0 1541 1442
MARISA RICCI MI 2.0 1467 1438
MICHAEL LU MI 2.0 1511 1379
VIRAJ MOHILE MI 2.0 1470 1474
SEAN M MC CORMICK MI 2.0 1284 1464
JULIA SHEN MI 1.5 1457 1461
JEZZEL FARKAS ON 1.5 1577 1384
ASHWIN BALAJI MI 1.0 1521 1541
THOMAS JOSEPH HOSMER MI 1.0 1505 1419
BEN LI MI 1.0 1500 1363

This is the table I will export for the project.

Quality Checks

Show code
nrow(project1_table)
[1] 64
Show code
colSums(is.na(project1_table))
                                            Name 
                                               0 
                                           State 
                                               0 
                          Total Number of Points 
                                               0 
                                      Pre-Rating 
                                               0 
Average Pre Tournament Chess Rating of Opponents 
                                               0 
Show code
sum(duplicated(project1_table$Name))
[1] 0
Show code
summary(project1_table)
        Name          State    Total Number of Points   Pre-Rating  
 Length   :64   Length   :64   Min.   :1.000          Min.   :1013  
 N.unique :64   N.unique : 3   1st Qu.:2.500          1st Qu.:1290  
 N.blank  : 0   N.blank  : 0   Median :3.500          Median :1474  
 Min.nchar: 6   Min.nchar: 2   Mean   :3.438          Mean   :1416  
 Max.nchar:26   Max.nchar: 2   3rd Qu.:4.000          3rd Qu.:1512  
                               Max.   :6.000          Max.   :1577  
 Average Pre Tournament Chess Rating of Opponents
 Min.   :1233                                    
 1st Qu.:1384                                    
 Median :1426                                    
 Mean   :1418                                    
 3rd Qu.:1465                                    
 Max.   :1541                                    

I use these checks to make sure the final table has the expected number of rows and to catch missing or duplicate values.

Gary Hua Check

Show code
gary_hua_check <- project1_table |>
  filter(str_to_lower(Name) == "gary hua")

gary_hua_check
# A tibble: 1 × 5
  Name     State `Total Number of Points` `Pre-Rating` Average Pre Tournament …¹
  <chr>    <chr>                    <dbl>        <dbl>                     <dbl>
1 GARY HUA ON                           6         1544                      1233
# ℹ abbreviated name: ¹​`Average Pre Tournament Chess Rating of Opponents`

I included this check because the approach says I will compare Gary Hua’s result with the course example. I will make that comparison from the rendered output rather than typing a result into the document before the code runs.

Exploratory Charts

These charts are secondary to the required table. I kept them simple so they help me inspect the cleaned data without taking over the project.

Pre-Rating Distribution

Show code
ggplot(project1_table, aes(x = `Pre-Rating`)) +
  geom_histogram(
    bins = 12,
    fill = "steelblue",
    color = "white"
  ) +
  labs(
    title = "Distribution of Player Pre-Ratings",
    x = "Pre-Rating",
    y = "Number of Players"
  ) +
  theme_minimal()

The histogram gives a quick view of how the player ratings are distributed.

Points by Player

Show code
# Order players by points and use pre-rating to break ties.
# The taller figure gives all player names enough room to be readable.
points_plot <- project1_table |>
  arrange(`Total Number of Points`, `Pre-Rating`) |>
  mutate(Name = factor(Name, levels = Name))

ggplot(
  points_plot,
  aes(x = `Total Number of Points`, y = Name)
) +
  geom_segment(
    aes(
      x = 0,
      xend = `Total Number of Points`,
      y = Name,
      yend = Name
    ),
    color = "grey75",
    linewidth = 0.6
  ) +
  geom_point(
    size = 3,
    color = "steelblue"
  ) +
  labs(
    title = "Tournament Points by Player",
    subtitle = "Players are ordered by total tournament points",
    x = "Total Points",
    y = NULL
  ) +
  scale_x_continuous(
    breaks = seq(
      0,
      max(points_plot$`Total Number of Points`, na.rm = TRUE),
      by = 1
    )
  ) +
  theme_minimal(base_size = 11) +
  theme(
    axis.text.y = element_text(size = 8),
    panel.grid.major.y = element_blank(),
    panel.grid.minor = element_blank(),
    plot.title = element_text(face = "bold")
  )

The taller lollipop chart makes the player names readable. Players are ordered by total points, with pre-rating used to break ties. This chart is supporting analysis; the final project table remains the main output. ### Player Rating and Opponent Rating

Show code
ggplot(
  project1_table,
  aes(
    x = `Pre-Rating`,
    y = `Average Pre Tournament Chess Rating of Opponents`
  )
) +
  geom_point(
    alpha = 0.75,
    color = "steelblue"
  ) +
  labs(
    title = "Player Rating vs. Average Opponent Rating",
    x = "Player Pre-Rating",
    y = "Average Opponent Pre-Rating"
  ) +
  theme_minimal()

This scatterplot compares two numeric variables from the final dataset.

Export the CSV

Show code
# Save the final table as a CSV file
write_csv(
  project1_table,
  "project1_chess_results.csv"
)

The exported CSV contains only the five requested result columns.

AI Use

I used GitHub Copilot in Visual Studio Code to help review parts of the R code, troubleshoot errors, and improve code comments. I reviewed the suggestions, kept the code I understood, and checked the results when running the project. The analysis and final results come from the R code and tournament data used in this project.

References

DATA 607 chess tournament dataset (public copy used only as a fallback):
https://raw.githubusercontent.com/longflin/DATA-607/refs/heads/main/Project%201/tournamentinfo.txt

Wickham, H., Cetinkaya-Rundel, M., & Grolemund, G. R for Data Science (2e).

R Graph Gallery. Lollipop plot examples.

R Graph Gallery. Histogram examples.