About this project

This is a short practice project to keep my R skills current. It uses open data on how preschoolers rate their classmates as friends.

Data source: [https://osf.io/9rw6s/files/fa5mg Horn, L. (2025, January 29). Data. Retrieved from osf.io/9rw6s]

Loading the data

The file is separated by semicolons, not commas, so it needs read_delim() instead of read_csv().

library(tidyverse)
library(janitor)

dat <- read_delim(
  "data_friendship assessment_children_anonymized.CSV",
  delim = ";"
) |>
  clean_names()

glimpse(dat)
## Rows: 1,130
## Columns: 13
## $ id_assessor         <chr> "A_F01", "A_F01", "A_F01", "A_F01", "A_F01", "A_F0…
## $ group_assessor      <chr> "A", "A", "A", "A", "A", "A", "A", "A", "A", "A", …
## $ group_size          <dbl> 11, 11, 11, 11, 11, 11, 11, 11, 11, 11, 11, 11, 11…
## $ age_assessor_months <dbl> 52, 52, 52, 52, 52, 52, 52, 52, 52, 52, 52, 61, 61…
## $ gender_assessor     <chr> "F", "F", "F", "F", "F", "F", "F", "F", "F", "F", …
## $ round               <dbl> 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1,…
## $ id_assessed         <chr> "A_F01", "A_F02", "A_F03", "A_F04", "A_F05", "A_M0…
## $ group_assessed      <chr> "A", "A", "A", "A", "A", "A", "A", "A", "A", "A", …
## $ age_assessed_months <dbl> 52, 61, 63, 69, 76, 41, 39, 62, 62, 68, 68, 52, 61…
## $ gender_assessed     <chr> "F", "F", "F", "F", "F", "M", "M", "M", "M", "M", …
## $ assessment1         <chr> NA, "green", "green", "yellow", "yellow", "red", "…
## $ assessment2         <chr> NA, "green", "green", "green", "yellow", "red", "y…
## $ friends             <dbl> NA, 1, 1, 0, 0, 0, 0, 0, 1, 0, 0, 1, NA, 1, 0, 0, …

Each row is one child (the assessor) rating one classmate (the assessed), so every child appears many times.

Missing values

data.frame(missing = sort(colSums(is.na(dat)), decreasing = TRUE))
##                     missing
## assessment1             197
## assessment2             193
## friends                 182
## id_assessed              20
## age_assessed_months      20
## gender_assessed          20
## id_assessor               0
## group_assessor            0
## group_size                0
## age_assessor_months       0
## gender_assessor           0
## round                     0
## group_assessed            0

The ratings are missing most often. To find out why, I labelled the rows where a child is paired with themselves.

dat <- dat |>
  mutate(self = id_assessor == id_assessed)

count(dat, self)
## # A tibble: 3 × 2
##   self      n
##   <lgl> <int>
## 1 FALSE  1012
## 2 TRUE     98
## 3 NA       20

The 98 TRUE rows are children paired with their own name, which are blank by design. The 20 NA rows have no classmate recorded at all, so I checked who they belong to.

dat |>
  filter(is.na(self)) |>
  count(id_assessor)
## # A tibble: 2 × 2
##   id_assessor     n
##   <chr>       <int>
## 1 C_F06          10
## 2 C_F07          10

Those 20 rows come from two children who have no ratings recorded, most likely because they did not take part.

dat |>
  filter(self == FALSE) |>
  summarise(missing_a1 = sum(is.na(assessment1)))
## # A tibble: 1 × 1
##   missing_a1
##        <int>
## 1         79

So the 197 blank first ratings have three causes: 98 own-name rows, 20 from two children with no ratings, and 79 real pairs where a rating is missing.

How friends is calculated

count(dat, assessment1, assessment2, friends)
## # A tibble: 15 × 4
##    assessment1 assessment2 friends     n
##    <chr>       <chr>         <dbl> <int>
##  1 green       green             1   379
##  2 green       red               0    21
##  3 green       yellow            0    34
##  4 red         green             0    19
##  5 red         red               0   153
##  6 red         yellow            0    48
##  7 red         <NA>              0     4
##  8 yellow      green             0    34
##  9 yellow      red               0    48
## 10 yellow      yellow            0   190
## 11 yellow      <NA>              0     3
## 12 <NA>        red               0     3
## 13 <NA>        yellow            0     8
## 14 <NA>        <NA>              0     4
## 15 <NA>        <NA>             NA   182

friends is 1 only when both ratings are green. Any other combination, even one green, is coded 0.

Same-gender friendships

Prediction: same-gender pairs are more likely to be rated as friends.

pairs <- dat |>
  filter(self == FALSE, !is.na(friends)) |>
  mutate(same_gender = if_else(
    gender_assessor == gender_assessed,
    "Same gender", "Different gender"
  ))

pairs |>
  group_by(same_gender) |>
  summarise(prop_friends = mean(friends), n = n())
## # A tibble: 2 × 3
##   same_gender      prop_friends     n
##   <chr>                   <dbl> <int>
## 1 Different gender        0.347   453
## 2 Same gender             0.448   495
pairs |>
  group_by(same_gender) |>
  summarise(prop_friends = mean(friends)) |>
  ggplot(aes(same_gender, prop_friends)) +
  geom_col() +
  labs(x = NULL, y = "Proportion rated as a friend")

[Within these results I found that same gendered pairs are more likely to rate each other as friends by about .05 than different gendered pairs.]

What I would explore next