Overview

Project goals

The goal of this project is to establish if children and adults can adjust their generalizations about a social group to account for sampling skew.

Previously on..

In Studies 1a-c (boat studies), we found that 4yo fail to adjust generalization against probabilistic structural skew (generalizing anyway), while adults adjust but just barely (tiny effect size). This result makes an interesting contrast compared to infant literature, where infants discount a sample of marbles deterministically skewed by an agent’s preferences.

In Study 2a, we use a paradigm involving agent preference to test if the challenge has to do with structural vs agentic sources of skew. If this is the case, participants should now pass. In contrast, if the challenge has to do with adjusting inference from probabilistically skewed samples vs discounting deterministically skewed samples entirely, participants should continue to fail here.

Results

Adult participants passed all the checks, indicating that they knew what soccer and basketball balls were, that they knew Alex either had no preference (not skewed) or had a preference for one sport (skewed) among kids, and among Gorps.

Nevertheless, adult participants were insensitive to sampling skew. Specifically:

  • Adult participants in both conditions generalized equally from the sample, predicting that most Gorps would like the sample-majority sport (i.e., the sport that the majority of Alex’s Gorp friends preferred).

  • There was no condition difference in participants’ explicit comparison of Alex’s Gorp friends to Gorps at large - participants were equally likely to say that they liked the sample-majority sport “the same”, or the Alex’s Gorp friends liked the sample-majority sport more.

As a result, it appears that adults fail to account for probabilistically skewed sampling.

Methods

Participants

Data was collected from 200 adults recruited via Prolific on Weds 8/26/2026 as a standard sample. Participants were required to be in the United States and have not participated in any previous studies in this project.

Participants were paid $2.25 for an estimated 9 minute task. In fact, the study generally took about 10.5 minutes for participants.

condition participants
not_skewed 93
skewed 98

The final sample included 191 adults (n = 93-98 in each of the 2 conditions).

Exclusion criteria

Participants were excluded if they failed the sound check or task check.

Exclusion reasons
n_collect sound_check check_task check_AI n n_excl excl_rate
200 0 7 2 191 9 4.50%

Demographics

Age

age
mean sd n
40.36 12.71 191

Gender

gender n prop
Female 92 48.2%
Male 92 48.2%
Prefer not to specify 4 2.1%
Non-binary 2 1.0%
nonconforming 1 0.5%

Race

race n prop
White, Caucasian, or European American 101 52.9%
Black or African American 26 13.6%
South or Southeast Asian 13 6.8%
Hispanic or Latino/a 12 6.3%
East Asian 10 5.2%
Prefer not to specify 6 3.1%
White, Caucasian, or European American,Black or African American 4 2.1%
White, Caucasian, or European American,Hispanic or Latino/a 3 1.6%
White, Caucasian, or European American,Native American, American Indian, or Alaska Native 3 1.6%
Hispanic or Latino/a,Black or African American 2 1.0%
White, Caucasian, or European American,Middle Eastern or North African 2 1.0%
Hebrew 1 0.5%
Hispanic or Latino/a,Black or African American,South or Southeast Asian 1 0.5%
Hispanic or Latino/a,Native Hawaiian or other Pacific Islander 1 0.5%
Hispanic or Latino/a,South or Southeast Asian 1 0.5%
Middle Eastern or North African 1 0.5%
South or Southeast Asian,East Asian 1 0.5%
White, Caucasian, or European American,East Asian 1 0.5%
White, Caucasian, or European American,Hispanic or Latino/a,Native American, American Indian, or Alaska Native 1 0.5%
White, Caucasian, or European American,South or Southeast Asian 1 0.5%

Education

education n prop
High school/GED 17 8.9%
Some college 60 31.4%
Bachelor's (B.A., B.S.) 86 45.0%
Master's (M.A., M.S.) 24 12.6%
Doctoral (Ph.D., J.D., M.D.) 1 0.5%
Prefer not to specify 3 1.6%

Procedure

This study was administered as a Qualtrics survey, and approved by the NYU IRB (IRB-FY2024-9169).

After providing their consent, participants completed a captcha and sound check, and were asked to watch videos sound on. Participants then watched the following videos in order:

  1. In the warmup phase, to confirm participants’ understanding of balls and sports, participants heard the narrator label a soccer ball and a basketball, and were asked to click on each.

    In the alternate counterbalance version, the left/right position of soccer and basketball buttons was switched on this and all questions in the study.

  2. In the familiarization phase, participants were introduced to an agent called Alex (depicted using a photograph of a white female child), and learned how Alex chooses friends by watching her make friends at a playground.

Each trial showed pictures of two children (pictures matched on race and gender; races and genders varied across trials), one holding a soccer ball and another holding a basketball. Children were unique to each trial.

In the skewed condition, Alex approached the child holding the soccer ball on 5 out of 6 trials, and approached the child holding the basketball on the remaining oddball trial. Trial order was randomized, such that the oddball trial always appeared in 2nd, 3rd, 4th, or 5th position.

In the not skewed condition, Alex approached children of each sport on 3 out of 6 trials. Sport selections alternated, with the first selection being randomized.

In the alternate counterbalance version, the position of children was fixed, while Alex’s selections were switched, such that the skewed condition saw Alex approach mostly basketball.

  1. As familiarization phase checks, participants were asked to (in the following fixed order):
  1. Familiarization: friends check: Predict which child Alex might befriend between a novel soccer kid and a novel basketball kid, to confirm their understanding of the agent’s preference. The images used were fixed images of a Black girl holding a basketball (fixed), and another Black girl holding a soccer ball (fixed), their positions counterbalanced on screen.

After responding, participants in the skewed condition were told that Alex will probably choose the kid holding the soccer ball, because Alex likes soccer (or basketball, in the alternate counterbalance version). Participants in the not skewed condition were told, “it might be hard for Alex to choose, because Alex likes soccer and basketball”.

  1. Familiarization: sport base rate check: Confirm their understanding of whether kids at the playground liked basketball, soccer, or both the same. After responding, all participants were told that kids at the playground liked both the same.
  1. In the sample observation phase, participants observed the same sample of Gorps that Alex befriends on Gorp Planet. In both conditions, Alex befriends 8 Gorps, 6 of which (fixed positions and colors) like soccer. The remaining 2 Gorps (fixed positions and colors: light pink, light yellow) like basketball. (In the alternate counterbalance version, the 6 like basketball, and the remaining 2 like soccer.)

  1. In the inference phase, participants had to make inferences about Gorps in general, after Alex left. They completed the below measures in fixed order:
  1. Inference: prediction trials: As one of our dependent measures, participants were asked to predict the sport preferences of 4 novel Gorps (a brown, pink, purple, and teal Gorp in fixed order). These four trials were averaged into a prediction proportion for each participant.

    For piloting purposes, participants were also asked how they decided their responses in the prediction task.

  1. Inference: comparison forced-choice: As another dependent measure, adults were shown Alex’s Gorp friends again, and were asked to make a forced-choice comparison of who likes the sample-majority sport more: Alex’s Gorp friends, the Gorps on Gorp Planet, or if they like it the same.

  1. Finally, participants were asked for any problems or confusion they had, what they thought the task was about, and demographic information.

Checks

Ball training

After labeling, all participants correctly identified a soccer ball and a basketball.

Familiarization: friends check

Both conditions largely passed the friends check.

Participants in the skewed condition understood Alex would befriend kids who prefered one sport (aligned with their counterbalance condition), while participants in the not skewed condition appeared more mixed.

All participants received information about the expected response after their response.

Familiarization: sport base rate check

Participants were asked to recall that Alex met many kids at the park today, and to recall which sport more of the kids on the playground liked: basketball, soccer, or did they like them the same.

The correct answer to this question is “the same”, as every trial showed a basketball kid and a soccer kid.

Participants mostly answered this question correctly; participants who gave incorrect answers were still included. All participants received the correct answer after their response.

Agent friends check (Gorps)

Both conditions largely passed the agent friends check.

Participants in the skewed condition understood Alex would befriend the Gorp who preferred one sport (aligned with their counterbalance condition), while participants in the not skewed condition appeared more mixed.

Results

Inference: prediction trials

Participants were asked to predict the sport preferences (soccer or basketball) of 4 novel group members (a brown, pink, purple, and teal Gorp in fixed order).

First prediction trial:

glmer_infer_sport <-
  glmer(infer_sport ~ condition + (1 | participant), 
        data = data_tidy, 
        family = binomial)

# condition difference?
glmer_infer_sport %>% 
  summary()

There was no statistically significant effect of condition on predictions, in a logistic regression with condition as the sole predictor and random intercepts per participant (b = -0.2, SE = 0.16, z = -1.29, p = 0.196).)

# set priors
priors <- c(
  prior(normal(0, 1.5), class = "Intercept"), # centered prior for baseline log-odds
  prior(normal(0, 1),   class = "b"), # weakly informative prior
  prior(student_t(3, 0, 1), class = "sd") # half-Student-t
)

# bayesian model
brm_infer_sport <-
  brm(infer_sport ~ condition + (1 | participant), 
      data = data_tidy, 
      family = bernoulli(link = "logit"),
      prior = priors,
      save_pars = save_pars(all = TRUE))
# prior predictive check
brm_prior_check <-
  brm(infer_sport ~ condition + (1 | participant),
      data = data_tidy,
      family = bernoulli(link = "logit"),
      prior = priors,
      sample_prior = "only")

# get draws of the linear predictor, transformed to probability scale
prior_draws <- data_tidy %>%
  distinct(condition) %>%
  add_epred_draws(brm_prior_check, re_formula = NA)  # NA = population-level only, ignore participant RE

# check that each condition's density is largely similar, no spikes around 0, 0.5, 1
ggplot(prior_draws, aes(x = .epred, fill = condition)) +
  geom_density(alpha = 0.5) +
  labs(x = "Prior predicted P(infer_sport = 1)", y = "Density") 

# set priors for null model
priors_null <- c(
  prior(normal(0, 1.5), class = "Intercept"), # centered prior for baseline log-odds
  prior(student_t(3, 0, 1), class = "sd") # half-Student-t
)

# null model
brm_infer_sport_null <- 
  brm(infer_sport ~ 1 + (1 | participant), 
      data = data_tidy, family = bernoulli(),
      prior = priors_null, 
      save_pars = save_pars(all = TRUE))

# get bf
bf <- bayes_factor(brm_infer_sport, brm_infer_sport_null)$bf

A Bayesian analysis suggests moderate evidence against a condition difference (BF =0.21).

Inference: comparison forced-choice

Participants were shown Alex’s Gorp friends again, and were asked to infer whether Gorps on Gorp Planet liked the sample-majority sport “less”, “the same”, or “more” than Alex’s Gorp friends.

# multinomial regression: do responses differ by condition?
infer_comp_multinom <- data %>% 
  multinom(infer_comp ~ condition, data = .) %>% 
  Anova()

infer_comp_multinom

There was no statistically significant effect of condition on responses to the comparison forced-choice question (LR Chisq(2) = 0.42, p = 0.812).

Supplementary results

Inference: comparison forced-choice vs prediction

Participants’ responses to the comparison forced-choice question were somewhat related to their earlier responses to the prediction trials.

The comparison forced-choice asked whether Gorps on Gorp Planet liked the sample-majority sport “less”, “the same”, or “more” than Alex’s Gorp friends.

Since 6 out of 8 of Alex’s Gorp friends liked the sample-majority sport, matching the sample proportion would be 3/4. As a result:

  • predictions of 0, 0.25, and 0.5 should correspond with responding “less”

  • predictions of 0.75 should correspond with responding “the same”

  • predictions of 1 should correspond with responding “more”

Session info

## R version 4.5.2 (2025-10-31)
## Platform: aarch64-apple-darwin20
## Running under: macOS Tahoe 26.5.2
## 
## Matrix products: default
## BLAS:   /System/Library/Frameworks/Accelerate.framework/Versions/A/Frameworks/vecLib.framework/Versions/A/libBLAS.dylib 
## LAPACK: /Library/Frameworks/R.framework/Versions/4.5-arm64/Resources/lib/libRlapack.dylib;  LAPACK version 3.12.1
## 
## locale:
## [1] en_US.UTF-8/en_US.UTF-8/en_US.UTF-8/C/en_US.UTF-8/en_US.UTF-8
## 
## time zone: America/New_York
## tzcode source: internal
## 
## attached base packages:
## [1] stats     graphics  grDevices utils     datasets  methods   base     
## 
## other attached packages:
##  [1] car_3.1-3           carData_3.0-5       nnet_7.3-20        
##  [4] tidybayes_3.0.7     broom.mixed_0.2.9.6 brms_2.23.0        
##  [7] Rcpp_1.1.2          lmerTest_3.2-0      lme4_2.0-6         
## [10] Matrix_1.7-6        viridis_0.6.5       viridisLite_0.4.2  
## [13] ggtext_0.1.2        lubridate_1.9.4     forcats_1.0.1      
## [16] stringr_1.6.0       dplyr_1.1.4         purrr_1.2.1        
## [19] readr_2.1.6         tidyr_1.3.2         tibble_3.3.1       
## [22] ggplot2_4.0.3       tidyverse_2.0.0     gt_1.3.0           
## [25] scales_1.4.0        janitor_2.2.1       here_1.0.2         
## [28] knitr_1.51         
## 
## loaded via a namespace (and not attached):
##   [1] RColorBrewer_1.1-3    tensorA_0.36.2.1      rstudioapi_0.18.0    
##   [4] jsonlite_2.0.0        magrittr_2.0.4        TH.data_1.1-5        
##   [7] estimability_1.5.1    farver_2.1.2          nloptr_2.2.1         
##  [10] rmarkdown_2.30        fs_1.6.6              ragg_1.5.0           
##  [13] vctrs_0.7.1           minqa_1.2.8           base64enc_0.1-3      
##  [16] htmltools_0.5.9       curl_7.0.0            distributional_0.6.0 
##  [19] broom_1.0.12          Formula_1.2-5         StanHeaders_2.32.10  
##  [22] sass_0.4.10           parallelly_1.46.1     bslib_0.10.0         
##  [25] htmlwidgets_1.6.4     sandwich_3.1-1        emmeans_2.0.1        
##  [28] zoo_1.8-15            cachem_1.1.0          lifecycle_1.0.5      
##  [31] pkgconfig_2.0.3       R6_2.6.1              fastmap_1.2.0        
##  [34] rbibutils_2.4.1       future_1.69.0         snakecase_0.11.1     
##  [37] digest_0.6.39         numDeriv_2016.8-1.1   colorspace_2.1-2     
##  [40] furrr_0.3.1           ps_1.9.1              rprojroot_2.1.1      
##  [43] textshaping_1.0.4     Hmisc_5.2-5           labeling_0.4.3       
##  [46] timechange_0.3.0      abind_1.4-8           compiler_4.5.2       
##  [49] bit64_4.6.0-1         withr_3.0.2           inline_0.3.21        
##  [52] htmlTable_2.4.3       S7_0.2.1              backports_1.5.0      
##  [55] QuickJSR_1.9.0        pkgbuild_1.4.8        MASS_7.3-65          
##  [58] loo_2.9.0             tools_4.5.2           foreign_0.8-90       
##  [61] otel_0.2.0            glue_1.8.0            callr_3.7.6          
##  [64] nlme_3.1-168          gridtext_0.1.5        grid_4.5.2           
##  [67] checkmate_2.3.3       cluster_2.1.8.1       generics_0.1.4       
##  [70] gtable_0.3.6          tzdb_0.5.0            data.table_1.18.0    
##  [73] hms_1.1.4             xml2_1.5.2            pillar_1.11.1        
##  [76] ggdist_3.3.3          vroom_1.6.7           posterior_1.6.1      
##  [79] splines_4.5.2         lattice_0.22-7        survival_3.8-6       
##  [82] bit_4.6.0             tidyselect_1.2.1      reformulas_0.4.3.1   
##  [85] arrayhelpers_1.1-2    gridExtra_2.3         V8_8.0.1             
##  [88] stats4_4.5.2          xfun_0.56             bridgesampling_1.2-1 
##  [91] matrixStats_1.5.0     rstan_2.32.7          stringi_1.8.7        
##  [94] yaml_2.3.12           boot_1.3-32           evaluate_1.0.5       
##  [97] codetools_0.2-20      cli_3.6.5             rpart_4.1.24         
## [100] RcppParallel_5.1.11-1 xtable_1.8-4          systemfonts_1.3.1    
## [103] Rdpack_2.6.5          processx_3.8.6        jquerylib_0.1.4      
## [106] globals_0.18.0        coda_0.19-4.1         svUnit_1.0.8         
## [109] parallel_4.5.2        rstantools_2.6.0      bayesplot_1.15.0     
## [112] Brobdingnag_1.2-9     listenv_0.10.0        ggthemes_5.2.0       
## [115] mvtnorm_1.3-3         crayon_1.5.3          rlang_1.1.7          
## [118] multcomp_1.4-29