The goal of this project is to establish if children and adults can adjust their generalizations about a social group to account for sampling skew.
In Studies 1a-c (boat studies), we found that 4yo fail to adjust generalization against probabilistic structural skew (generalizing anyway), while adults adjust but just barely (tiny effect size). This result makes an interesting contrast compared to infant literature, where infants discount a sample of marbles deterministically skewed by an agent’s preferences.
In Study 2a, we use a paradigm involving agent preference to test if the challenge has to do with structural vs agentic sources of skew. If this is the case, participants should now pass. In contrast, if the challenge has to do with adjusting inference from probabilistically skewed samples vs discounting deterministically skewed samples entirely, participants should continue to fail here.
Adult participants passed all the checks, indicating that they knew what soccer and basketball balls were, that they knew Alex either had no preference (not skewed) or had a preference for one sport (skewed) among kids, and among Gorps.
Nevertheless, adult participants were insensitive to sampling skew. Specifically:
Adult participants in both conditions generalized equally from the sample, predicting that most Gorps would like the sample-majority sport (i.e., the sport that the majority of Alex’s Gorp friends preferred).
There was no condition difference in participants’ explicit comparison of Alex’s Gorp friends to Gorps at large - participants were equally likely to say that they liked the sample-majority sport “the same”, or the Alex’s Gorp friends liked the sample-majority sport more.
As a result, it appears that adults fail to account for probabilistically skewed sampling.
Data was collected from 200 adults recruited via Prolific on Weds 8/26/2026 as a standard sample. Participants were required to be in the United States and have not participated in any previous studies in this project.
Participants were paid $2.25 for an estimated 9 minute task. In fact, the study generally took about 10.5 minutes for participants.
| condition | participants |
|---|---|
| not_skewed | 93 |
| skewed | 98 |
The final sample included 191 adults (n = 93-98 in each of the 2 conditions).
Participants were excluded if they failed the sound check or task check.
| Exclusion reasons | ||||||
| n_collect | sound_check | check_task | check_AI | n | n_excl | excl_rate |
|---|---|---|---|---|---|---|
| 200 | 0 | 7 | 2 | 191 | 9 | 4.50% |
| age | ||
| mean | sd | n |
|---|---|---|
| 40.36 | 12.71 | 191 |
| gender | n | prop |
|---|---|---|
| Female | 92 | 48.2% |
| Male | 92 | 48.2% |
| Prefer not to specify | 4 | 2.1% |
| Non-binary | 2 | 1.0% |
| nonconforming | 1 | 0.5% |
| race | n | prop |
|---|---|---|
| White, Caucasian, or European American | 101 | 52.9% |
| Black or African American | 26 | 13.6% |
| South or Southeast Asian | 13 | 6.8% |
| Hispanic or Latino/a | 12 | 6.3% |
| East Asian | 10 | 5.2% |
| Prefer not to specify | 6 | 3.1% |
| White, Caucasian, or European American,Black or African American | 4 | 2.1% |
| White, Caucasian, or European American,Hispanic or Latino/a | 3 | 1.6% |
| White, Caucasian, or European American,Native American, American Indian, or Alaska Native | 3 | 1.6% |
| Hispanic or Latino/a,Black or African American | 2 | 1.0% |
| White, Caucasian, or European American,Middle Eastern or North African | 2 | 1.0% |
| Hebrew | 1 | 0.5% |
| Hispanic or Latino/a,Black or African American,South or Southeast Asian | 1 | 0.5% |
| Hispanic or Latino/a,Native Hawaiian or other Pacific Islander | 1 | 0.5% |
| Hispanic or Latino/a,South or Southeast Asian | 1 | 0.5% |
| Middle Eastern or North African | 1 | 0.5% |
| South or Southeast Asian,East Asian | 1 | 0.5% |
| White, Caucasian, or European American,East Asian | 1 | 0.5% |
| White, Caucasian, or European American,Hispanic or Latino/a,Native American, American Indian, or Alaska Native | 1 | 0.5% |
| White, Caucasian, or European American,South or Southeast Asian | 1 | 0.5% |
| education | n | prop |
|---|---|---|
| High school/GED | 17 | 8.9% |
| Some college | 60 | 31.4% |
| Bachelor's (B.A., B.S.) | 86 | 45.0% |
| Master's (M.A., M.S.) | 24 | 12.6% |
| Doctoral (Ph.D., J.D., M.D.) | 1 | 0.5% |
| Prefer not to specify | 3 | 1.6% |
This study was administered as a Qualtrics survey, and approved by the NYU IRB (IRB-FY2024-9169).
After providing their consent, participants completed a captcha and sound check, and were asked to watch videos sound on. Participants then watched the following videos in order:
In the warmup phase, to confirm participants’ understanding of balls and sports, participants heard the narrator label a soccer ball and a basketball, and were asked to click on each.
In the alternate counterbalance version, the left/right position of soccer and basketball buttons was switched on this and all questions in the study.
In the familiarization phase, participants were introduced to an agent called Alex (depicted using a photograph of a white female child), and learned how Alex chooses friends by watching her make friends at a playground.
Each trial showed pictures of two children (pictures matched on race and gender; races and genders varied across trials), one holding a soccer ball and another holding a basketball. Children were unique to each trial.
In the skewed condition, Alex approached the child holding the soccer ball on 5 out of 6 trials, and approached the child holding the basketball on the remaining oddball trial. Trial order was randomized, such that the oddball trial always appeared in 2nd, 3rd, 4th, or 5th position.
In the not skewed condition, Alex approached children of each sport on 3 out of 6 trials. Sport selections alternated, with the first selection being randomized.
In the alternate counterbalance version, the position of children was fixed, while Alex’s selections were switched, such that the skewed condition saw Alex approach mostly basketball.
After responding, participants in the skewed condition were told that Alex will probably choose the kid holding the soccer ball, because Alex likes soccer (or basketball, in the alternate counterbalance version). Participants in the not skewed condition were told, “it might be hard for Alex to choose, because Alex likes soccer and basketball”.
Inference: prediction trials: As one of our dependent measures, participants were asked to predict the sport preferences of 4 novel Gorps (a brown, pink, purple, and teal Gorp in fixed order). These four trials were averaged into a prediction proportion for each participant.
For piloting purposes, participants were also asked how they decided their responses in the prediction task.
After labeling, all participants correctly identified a soccer ball and a basketball.
Both conditions largely passed the friends check.
Participants in the skewed condition understood Alex would befriend kids who prefered one sport (aligned with their counterbalance condition), while participants in the not skewed condition appeared more mixed.
All participants received information about the expected response after their response.
Participants were asked to recall that Alex met many kids at the park today, and to recall which sport more of the kids on the playground liked: basketball, soccer, or did they like them the same.
The correct answer to this question is “the same”, as every trial showed a basketball kid and a soccer kid.
Participants mostly answered this question correctly; participants who gave incorrect answers were still included. All participants received the correct answer after their response.
Both conditions largely passed the agent friends check.
Participants in the skewed condition understood Alex would befriend the Gorp who preferred one sport (aligned with their counterbalance condition), while participants in the not skewed condition appeared more mixed.
Participants were asked to predict the sport preferences (soccer or basketball) of 4 novel group members (a brown, pink, purple, and teal Gorp in fixed order).
First prediction trial:
glmer_infer_sport <-
glmer(infer_sport ~ condition + (1 | participant),
data = data_tidy,
family = binomial)
# condition difference?
glmer_infer_sport %>%
summary()
There was no statistically significant effect of condition on predictions, in a logistic regression with condition as the sole predictor and random intercepts per participant (b = -0.2, SE = 0.16, z = -1.29, p = 0.196).)
# set priors
priors <- c(
prior(normal(0, 1.5), class = "Intercept"), # centered prior for baseline log-odds
prior(normal(0, 1), class = "b"), # weakly informative prior
prior(student_t(3, 0, 1), class = "sd") # half-Student-t
)
# bayesian model
brm_infer_sport <-
brm(infer_sport ~ condition + (1 | participant),
data = data_tidy,
family = bernoulli(link = "logit"),
prior = priors,
save_pars = save_pars(all = TRUE))
# prior predictive check
brm_prior_check <-
brm(infer_sport ~ condition + (1 | participant),
data = data_tidy,
family = bernoulli(link = "logit"),
prior = priors,
sample_prior = "only")
# get draws of the linear predictor, transformed to probability scale
prior_draws <- data_tidy %>%
distinct(condition) %>%
add_epred_draws(brm_prior_check, re_formula = NA) # NA = population-level only, ignore participant RE
# check that each condition's density is largely similar, no spikes around 0, 0.5, 1
ggplot(prior_draws, aes(x = .epred, fill = condition)) +
geom_density(alpha = 0.5) +
labs(x = "Prior predicted P(infer_sport = 1)", y = "Density")
# set priors for null model
priors_null <- c(
prior(normal(0, 1.5), class = "Intercept"), # centered prior for baseline log-odds
prior(student_t(3, 0, 1), class = "sd") # half-Student-t
)
# null model
brm_infer_sport_null <-
brm(infer_sport ~ 1 + (1 | participant),
data = data_tidy, family = bernoulli(),
prior = priors_null,
save_pars = save_pars(all = TRUE))
# get bf
bf <- bayes_factor(brm_infer_sport, brm_infer_sport_null)$bf
A Bayesian analysis suggests moderate evidence against a condition difference (BF =0.21).
Participants were shown Alex’s Gorp friends again, and were asked to infer whether Gorps on Gorp Planet liked the sample-majority sport “less”, “the same”, or “more” than Alex’s Gorp friends.
# multinomial regression: do responses differ by condition?
infer_comp_multinom <- data %>%
multinom(infer_comp ~ condition, data = .) %>%
Anova()
infer_comp_multinom
There was no statistically significant effect of condition on responses to the comparison forced-choice question (LR Chisq(2) = 0.42, p = 0.812).
Participants’ responses to the comparison forced-choice question were somewhat related to their earlier responses to the prediction trials.
The comparison forced-choice asked whether Gorps on Gorp Planet liked the sample-majority sport “less”, “the same”, or “more” than Alex’s Gorp friends.
Since 6 out of 8 of Alex’s Gorp friends liked the sample-majority sport, matching the sample proportion would be 3/4. As a result:
predictions of 0, 0.25, and 0.5 should correspond with responding “less”
predictions of 0.75 should correspond with responding “the same”
predictions of 1 should correspond with responding “more”
## R version 4.5.2 (2025-10-31)
## Platform: aarch64-apple-darwin20
## Running under: macOS Tahoe 26.5.2
##
## Matrix products: default
## BLAS: /System/Library/Frameworks/Accelerate.framework/Versions/A/Frameworks/vecLib.framework/Versions/A/libBLAS.dylib
## LAPACK: /Library/Frameworks/R.framework/Versions/4.5-arm64/Resources/lib/libRlapack.dylib; LAPACK version 3.12.1
##
## locale:
## [1] en_US.UTF-8/en_US.UTF-8/en_US.UTF-8/C/en_US.UTF-8/en_US.UTF-8
##
## time zone: America/New_York
## tzcode source: internal
##
## attached base packages:
## [1] stats graphics grDevices utils datasets methods base
##
## other attached packages:
## [1] car_3.1-3 carData_3.0-5 nnet_7.3-20
## [4] tidybayes_3.0.7 broom.mixed_0.2.9.6 brms_2.23.0
## [7] Rcpp_1.1.2 lmerTest_3.2-0 lme4_2.0-6
## [10] Matrix_1.7-6 viridis_0.6.5 viridisLite_0.4.2
## [13] ggtext_0.1.2 lubridate_1.9.4 forcats_1.0.1
## [16] stringr_1.6.0 dplyr_1.1.4 purrr_1.2.1
## [19] readr_2.1.6 tidyr_1.3.2 tibble_3.3.1
## [22] ggplot2_4.0.3 tidyverse_2.0.0 gt_1.3.0
## [25] scales_1.4.0 janitor_2.2.1 here_1.0.2
## [28] knitr_1.51
##
## loaded via a namespace (and not attached):
## [1] RColorBrewer_1.1-3 tensorA_0.36.2.1 rstudioapi_0.18.0
## [4] jsonlite_2.0.0 magrittr_2.0.4 TH.data_1.1-5
## [7] estimability_1.5.1 farver_2.1.2 nloptr_2.2.1
## [10] rmarkdown_2.30 fs_1.6.6 ragg_1.5.0
## [13] vctrs_0.7.1 minqa_1.2.8 base64enc_0.1-3
## [16] htmltools_0.5.9 curl_7.0.0 distributional_0.6.0
## [19] broom_1.0.12 Formula_1.2-5 StanHeaders_2.32.10
## [22] sass_0.4.10 parallelly_1.46.1 bslib_0.10.0
## [25] htmlwidgets_1.6.4 sandwich_3.1-1 emmeans_2.0.1
## [28] zoo_1.8-15 cachem_1.1.0 lifecycle_1.0.5
## [31] pkgconfig_2.0.3 R6_2.6.1 fastmap_1.2.0
## [34] rbibutils_2.4.1 future_1.69.0 snakecase_0.11.1
## [37] digest_0.6.39 numDeriv_2016.8-1.1 colorspace_2.1-2
## [40] furrr_0.3.1 ps_1.9.1 rprojroot_2.1.1
## [43] textshaping_1.0.4 Hmisc_5.2-5 labeling_0.4.3
## [46] timechange_0.3.0 abind_1.4-8 compiler_4.5.2
## [49] bit64_4.6.0-1 withr_3.0.2 inline_0.3.21
## [52] htmlTable_2.4.3 S7_0.2.1 backports_1.5.0
## [55] QuickJSR_1.9.0 pkgbuild_1.4.8 MASS_7.3-65
## [58] loo_2.9.0 tools_4.5.2 foreign_0.8-90
## [61] otel_0.2.0 glue_1.8.0 callr_3.7.6
## [64] nlme_3.1-168 gridtext_0.1.5 grid_4.5.2
## [67] checkmate_2.3.3 cluster_2.1.8.1 generics_0.1.4
## [70] gtable_0.3.6 tzdb_0.5.0 data.table_1.18.0
## [73] hms_1.1.4 xml2_1.5.2 pillar_1.11.1
## [76] ggdist_3.3.3 vroom_1.6.7 posterior_1.6.1
## [79] splines_4.5.2 lattice_0.22-7 survival_3.8-6
## [82] bit_4.6.0 tidyselect_1.2.1 reformulas_0.4.3.1
## [85] arrayhelpers_1.1-2 gridExtra_2.3 V8_8.0.1
## [88] stats4_4.5.2 xfun_0.56 bridgesampling_1.2-1
## [91] matrixStats_1.5.0 rstan_2.32.7 stringi_1.8.7
## [94] yaml_2.3.12 boot_1.3-32 evaluate_1.0.5
## [97] codetools_0.2-20 cli_3.6.5 rpart_4.1.24
## [100] RcppParallel_5.1.11-1 xtable_1.8-4 systemfonts_1.3.1
## [103] Rdpack_2.6.5 processx_3.8.6 jquerylib_0.1.4
## [106] globals_0.18.0 coda_0.19-4.1 svUnit_1.0.8
## [109] parallel_4.5.2 rstantools_2.6.0 bayesplot_1.15.0
## [112] Brobdingnag_1.2-9 listenv_0.10.0 ggthemes_5.2.0
## [115] mvtnorm_1.3-3 crayon_1.5.3 rlang_1.1.7
## [118] multcomp_1.4-29