This is a short practice project to keep my R skills current. It uses open data on how preschoolers rate their classmates as friends.
Data source: [https://osf.io/9rw6s/files/fa5mg Horn, L. (2025, January 29). Data. Retrieved from osf.io/9rw6s]
The file is separated by semicolons, not commas, so it needs
read_delim() instead of read_csv().
library(tidyverse)
library(janitor)
dat <- read_delim(
"data_friendship assessment_children_anonymized.CSV",
delim = ";"
) |>
clean_names()
glimpse(dat)
## Rows: 1,130
## Columns: 13
## $ id_assessor <chr> "A_F01", "A_F01", "A_F01", "A_F01", "A_F01", "A_F0…
## $ group_assessor <chr> "A", "A", "A", "A", "A", "A", "A", "A", "A", "A", …
## $ group_size <dbl> 11, 11, 11, 11, 11, 11, 11, 11, 11, 11, 11, 11, 11…
## $ age_assessor_months <dbl> 52, 52, 52, 52, 52, 52, 52, 52, 52, 52, 52, 61, 61…
## $ gender_assessor <chr> "F", "F", "F", "F", "F", "F", "F", "F", "F", "F", …
## $ round <dbl> 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1,…
## $ id_assessed <chr> "A_F01", "A_F02", "A_F03", "A_F04", "A_F05", "A_M0…
## $ group_assessed <chr> "A", "A", "A", "A", "A", "A", "A", "A", "A", "A", …
## $ age_assessed_months <dbl> 52, 61, 63, 69, 76, 41, 39, 62, 62, 68, 68, 52, 61…
## $ gender_assessed <chr> "F", "F", "F", "F", "F", "M", "M", "M", "M", "M", …
## $ assessment1 <chr> NA, "green", "green", "yellow", "yellow", "red", "…
## $ assessment2 <chr> NA, "green", "green", "green", "yellow", "red", "y…
## $ friends <dbl> NA, 1, 1, 0, 0, 0, 0, 0, 1, 0, 0, 1, NA, 1, 0, 0, …
Each row is one child (the assessor) rating one classmate (the assessed), so every child appears many times.
data.frame(missing = sort(colSums(is.na(dat)), decreasing = TRUE))
## missing
## assessment1 197
## assessment2 193
## friends 182
## id_assessed 20
## age_assessed_months 20
## gender_assessed 20
## id_assessor 0
## group_assessor 0
## group_size 0
## age_assessor_months 0
## gender_assessor 0
## round 0
## group_assessed 0
The ratings are missing most often. To find out why, I labelled the rows where a child is paired with themselves.
dat <- dat |>
mutate(self = id_assessor == id_assessed)
count(dat, self)
## # A tibble: 3 × 2
## self n
## <lgl> <int>
## 1 FALSE 1012
## 2 TRUE 98
## 3 NA 20
The 98 TRUE rows are children paired with their own
name, which are blank by design. The 20 NA rows have no
classmate recorded at all, so I checked who they belong to.
dat |>
filter(is.na(self)) |>
count(id_assessor)
## # A tibble: 2 × 2
## id_assessor n
## <chr> <int>
## 1 C_F06 10
## 2 C_F07 10
Those 20 rows come from two children who have no ratings recorded, most likely because they did not take part.
dat |>
filter(self == FALSE) |>
summarise(missing_a1 = sum(is.na(assessment1)))
## # A tibble: 1 × 1
## missing_a1
## <int>
## 1 79
So the 197 blank first ratings have three causes: 98 own-name rows, 20 from two children with no ratings, and 79 real pairs where a rating is missing.
friends is calculatedcount(dat, assessment1, assessment2, friends)
## # A tibble: 15 × 4
## assessment1 assessment2 friends n
## <chr> <chr> <dbl> <int>
## 1 green green 1 379
## 2 green red 0 21
## 3 green yellow 0 34
## 4 red green 0 19
## 5 red red 0 153
## 6 red yellow 0 48
## 7 red <NA> 0 4
## 8 yellow green 0 34
## 9 yellow red 0 48
## 10 yellow yellow 0 190
## 11 yellow <NA> 0 3
## 12 <NA> red 0 3
## 13 <NA> yellow 0 8
## 14 <NA> <NA> 0 4
## 15 <NA> <NA> NA 182
friends is 1 only when both ratings are green. Any other
combination, even one green, is coded 0.
Prediction: same-gender pairs are more likely to be rated as friends.
pairs <- dat |>
filter(self == FALSE, !is.na(friends)) |>
mutate(same_gender = if_else(
gender_assessor == gender_assessed,
"Same gender", "Different gender"
))
pairs |>
group_by(same_gender) |>
summarise(prop_friends = mean(friends), n = n())
## # A tibble: 2 × 3
## same_gender prop_friends n
## <chr> <dbl> <int>
## 1 Different gender 0.347 453
## 2 Same gender 0.448 495
pairs |>
group_by(same_gender) |>
summarise(prop_friends = mean(friends)) |>
ggplot(aes(same_gender, prop_friends)) +
geom_col() +
labs(x = NULL, y = "Proportion rated as a friend")
[Within these results I found that same gendered pairs are more likely to rate each other as friends by about .05 than different gendered pairs.]
friends only goes
one way