The goal of this project is to establish if children and adults can adjust their generalizations about a social group to account for sampling skew.
Adults failed to account for probabilistic skew that resulted in a skewed but still mixed sample (study 2a). However, adults showed evidence of accounting for deterministic skew that resulted in a uniform sample (study 2b). Here, we try to identify whether adults require skew to be deterministic and/or the sample to be all uniform to account for sampling skew.
Here, the agent’s preference was probabilistic, while the sample they chose ended up being entirely uniform (100% aligned with the agent’s preference). If adults require skew to be deterministic, they should fail. If adults require the sample to be uniform, they should succeed.
Adult participants passed all the checks, indicating that they knew what soccer and basketball balls were, that they knew Alex either had no preference (not skewed) or had a preference for one sport (skewed) among kids, and among Gorps.
In contrast to Study 2b, where adults successfully account for deterministic skew, adult participants here showed sensitivity to probabilistic skew that resulted in a uniform sample. Specifically:
Adults showed less generalization of the sample’s sport in the skewed condition than in the not skewed condition.
When comparing Alex’s Gorp friends and Gorps on Gorp Planet, there was no condition difference in forced-choice responses of who likes soccer more, although qualitatively it leans towards a condition effect in the predicted direction. (Note that an effect on this measure was not preregistered.) Adults’ responses to this measure do seem modestly related to their prediction responses.
The results of this study are an interesting contrast to Study 2b, as the only difference is a single familiarization trial, making Alex’s preferences probabilistic (befriending soccer on 5 out of 6 trials) rather than deterministic (6 out of 6).
As a result, it appears that adults do account for (probabilistic) skewed sampling when it results in a uniform sample, such that the uniform sample acts as a cue that skewed sampling has contaminated the sample.
Data was collected from 400 adults recruited via Prolific on Mon 7/27/2026 as a standard sample. Participants were required to be in the United States, fluent in English, and have not participated in any previous studies in this project.
Participants were paid $2.25 for an estimated 9 minute task. In fact, the study generally took about 10.5 minutes for participants.
| condition | participants |
|---|---|
| not_skewed | 191 |
| skewed | 192 |
The final sample included 383 adults (n = 191-192 in each of the 2 conditions).
Participants were excluded if they failed the sound check or task check.
| Exclusion reasons | |||||
| n_collect | sound_check | check_task | n | n_excl | excl_rate |
|---|---|---|---|---|---|
| 400 | 3 | 14 | 383 | 17 | 4.25% |
| age | ||
| mean | sd | n |
|---|---|---|
| 42.37 | 14.18 | 383 |
| gender | n | prop |
|---|---|---|
| Female | 214 | 55.9% |
| Male | 165 | 43.1% |
| Non-binary | 2 | 0.5% |
| Prefer not to specify | 1 | 0.3% |
| Woman | 1 | 0.3% |
| race | n | prop |
|---|---|---|
| White, Caucasian, or European American | 238 | 62.1% |
| Black or African American | 67 | 17.5% |
| Hispanic or Latino/a | 20 | 5.2% |
| South or Southeast Asian | 16 | 4.2% |
| East Asian | 12 | 3.1% |
| White, Caucasian, or European American,Hispanic or Latino/a | 8 | 2.1% |
| White, Caucasian, or European American,Black or African American | 4 | 1.0% |
| Middle Eastern or North African | 3 | 0.8% |
| Prefer not to specify | 3 | 0.8% |
| White, Caucasian, or European American,Native American, American Indian, or Alaska Native | 3 | 0.8% |
| Black or African American,South or Southeast Asian | 1 | 0.3% |
| HUMAN RACE | 1 | 0.3% |
| Hispanic or Latino/a,Black or African American | 1 | 0.3% |
| Hispanic or Latino/a,Native American, American Indian, or Alaska Native | 1 | 0.3% |
| South or Southeast Asian,East Asian | 1 | 0.3% |
| White, Caucasian, or European American,East Asian | 1 | 0.3% |
| White, Caucasian, or European American,Middle Eastern or North African | 1 | 0.3% |
| White, Caucasian, or European American,Native American, American Indian, or Alaska Native,South or Southeast Asian | 1 | 0.3% |
| mixed | 1 | 0.3% |
| education | n | prop |
|---|---|---|
| Less than high school | 4 | 1.0% |
| High school/GED | 43 | 11.2% |
| Some college | 100 | 26.1% |
| Bachelor's (B.A., B.S.) | 151 | 39.4% |
| Master's (M.A., M.S.) | 68 | 17.8% |
| Doctoral (Ph.D., J.D., M.D.) | 15 | 3.9% |
| Prefer not to specify | 2 | 0.5% |
This study was administered as a Qualtrics survey, and approved by the NYU IRB (IRB-FY2024-9169).
After providing their consent, participants completed a captcha and sound check, and were asked to watch videos sound on. Participants then watched the following videos in order:
In the warmup phase, to confirm participants’ understanding of balls and sports, participants heard the narrator label a soccer ball and a basketball, and were asked to click on each.
In the alternate counterbalance version, the left/right position of soccer and basketball buttons was switched on this and all questions in the study.
In the familiarization phase, participants were introduced to an agent called Alex (depicted using a photograph of a white female child), and learned how Alex chooses friends by watching her make friends at a playground.
Each trial showed pictures of two children (pictures matched on race and gender; races and genders varied across trials), one holding a soccer ball and another holding a basketball. Children were unique to each trial.
In the **skewed condition**, Alex approached the child holding the soccer ball on 5 out of 6 trials, and approached the child holding the basketball on the remaining oddball trial. Trial order was randomized, such that the oddball trial always appeared in 2nd, 3rd, 4th, or 5th position.
In the **not skewed condition**, Alex approached children of each sport on 3 out of 6 trials. Sport selections alternated, with the first selection being randomized.
In the alternate counterbalance version, the position of children was fixed, while Alex's selections were switched, such that the skewed condition saw Alex approach mostly basketball.
After responding, participants in the skewed condition were told that Alex will probably choose the kid holding the soccer ball, because Alex likes soccer (or basketball, in the alternate counterbalance version). Participants in the not skewed condition were told, "it might be hard for Alex to choose, because Alex likes soccer *and* basketball".
Inference: prediction trials: As one of our dependent measures, participants were asked to predict the sport preferences of 4 novel Gorps (a brown, pink, purple, and teal Gorp in fixed order). These four trials were averaged into a prediction proportion for each participant.
For piloting purposes, participants were also asked how they decided their responses in the prediction task.
Inference: comparison forced-choice: As another dependent measure, adults were shown Alex’s Gorp friends again, and were asked to make a forced-choice comparison: Who liked soccer more? Alex’s Gorp friends, Gorps on Gorp Planet, or do they like soccer the same. In the counterbalanced version, this question asked about basketball.
For piloting purposes, participants were also asked how they decided their response to this question.
After labeling, all participants correctly identified a soccer ball and a basketball.
Both conditions largely passed the friends check.
Participants in the skewed condition understood Alex would befriend kids who prefered one sport (aligned with their counterbalance condition), while participants in the not skewed condition appeared more mixed.
All participants received information about the expected response after their response.
Participants were asked to recall that Alex met many kids at the park today, and to recall which sport more of the kids on the playground liked: basketball, soccer, or did they like them the same.
The correct answer to this question is “the same”, as every trial showed a basketball kid and a soccer kid.
Participants mostly answered this question correctly; participants who gave incorrect answers were still included. All participants received the correct answer after their response.
Both conditions largely passed the agent friends check.
Participants in the skewed condition understood Alex would befriend the Gorp who preferred one sport (aligned with their counterbalance condition), while participants in the not skewed condition appeared more mixed.
Participants were asked to predict the sport preferences (soccer or basketball) of 4 novel group members (a brown, pink, purple, and teal Gorp in fixed order).
First prediction trial:
glmer_infer_sport <-
glmer(infer_sport ~ condition + (1 | participant),
data = data_tidy,
family = binomial)
# condition difference?
glmer_infer_sport %>%
summary()
# set priors
priors <- c(
prior(normal(0, 1.5), class = "Intercept"), # centered prior for baseline log-odds
prior(normal(0, 1), class = "b"), # weakly informative prior
prior(student_t(3, 0, 1), class = "sd") # half-Student-t
)
# bayesian model
brm_infer_sport <-
brm(infer_sport ~ condition + (1 | participant),
data = data_tidy,
family = bernoulli(link = "logit"),
prior = priors,
save_pars = save_pars(all = TRUE))
## Running /Library/Frameworks/R.framework/Resources/bin/R CMD SHLIB foo.c
## using C compiler: ‘Apple clang version 17.0.0 (clang-1700.6.4.2)’
## using SDK: ‘MacOSX26.2.sdk’
## clang -arch arm64 -std=gnu2x -I"/Library/Frameworks/R.framework/Resources/include" -DNDEBUG -I"/Library/Frameworks/R.framework/Versions/4.5-arm64/Resources/library/Rcpp/include/" -I"/Library/Frameworks/R.framework/Versions/4.5-arm64/Resources/library/RcppEigen/include/" -I"/Library/Frameworks/R.framework/Versions/4.5-arm64/Resources/library/RcppEigen/include/unsupported" -I"/Library/Frameworks/R.framework/Versions/4.5-arm64/Resources/library/BH/include" -I"/Library/Frameworks/R.framework/Versions/4.5-arm64/Resources/library/StanHeaders/include/src/" -I"/Library/Frameworks/R.framework/Versions/4.5-arm64/Resources/library/StanHeaders/include/" -I"/Library/Frameworks/R.framework/Versions/4.5-arm64/Resources/library/RcppParallel/include/" -I"/Library/Frameworks/R.framework/Versions/4.5-arm64/Resources/library/rstan/include" -DEIGEN_NO_DEBUG -DBOOST_DISABLE_ASSERTS -DBOOST_PENDING_INTEGER_LOG2_HPP -DSTAN_THREADS -DUSE_STANC3 -DSTRICT_R_HEADERS -DBOOST_PHOENIX_NO_VARIADIC_EXPRESSION -D_HAS_AUTO_PTR_ETC=0 -include '/Library/Frameworks/R.framework/Versions/4.5-arm64/Resources/library/StanHeaders/include/stan/math/prim/fun/Eigen.hpp' -D_REENTRANT -DRCPP_PARALLEL_USE_TBB=1 -I/opt/R/arm64/include -fPIC -falign-functions=64 -Wall -g -O2 -c foo.c -o foo.o
## In file included from <built-in>:1:
## In file included from /Library/Frameworks/R.framework/Versions/4.5-arm64/Resources/library/StanHeaders/include/stan/math/prim/fun/Eigen.hpp:22:
## In file included from /Library/Frameworks/R.framework/Versions/4.5-arm64/Resources/library/RcppEigen/include/Eigen/Dense:1:
## In file included from /Library/Frameworks/R.framework/Versions/4.5-arm64/Resources/library/RcppEigen/include/Eigen/Core:19:
## /Library/Frameworks/R.framework/Versions/4.5-arm64/Resources/library/RcppEigen/include/Eigen/src/Core/util/Macros.h:679:10: fatal error: 'cmath' file not found
## 679 | #include <cmath>
## | ^~~~~~~
## 1 error generated.
## make: *** [foo.o] Error 1
##
## SAMPLING FOR MODEL 'anon_model' NOW (CHAIN 1).
## Chain 1:
## Chain 1: Gradient evaluation took 4.5e-05 seconds
## Chain 1: 1000 transitions using 10 leapfrog steps per transition would take 0.45 seconds.
## Chain 1: Adjust your expectations accordingly!
## Chain 1:
## Chain 1:
## Chain 1: Iteration: 1 / 2000 [ 0%] (Warmup)
## Chain 1: Iteration: 200 / 2000 [ 10%] (Warmup)
## Chain 1: Iteration: 400 / 2000 [ 20%] (Warmup)
## Chain 1: Iteration: 600 / 2000 [ 30%] (Warmup)
## Chain 1: Iteration: 800 / 2000 [ 40%] (Warmup)
## Chain 1: Iteration: 1000 / 2000 [ 50%] (Warmup)
## Chain 1: Iteration: 1001 / 2000 [ 50%] (Sampling)
## Chain 1: Iteration: 1200 / 2000 [ 60%] (Sampling)
## Chain 1: Iteration: 1400 / 2000 [ 70%] (Sampling)
## Chain 1: Iteration: 1600 / 2000 [ 80%] (Sampling)
## Chain 1: Iteration: 1800 / 2000 [ 90%] (Sampling)
## Chain 1: Iteration: 2000 / 2000 [100%] (Sampling)
## Chain 1:
## Chain 1: Elapsed Time: 0.14 seconds (Warm-up)
## Chain 1: 0.142 seconds (Sampling)
## Chain 1: 0.282 seconds (Total)
## Chain 1:
##
## SAMPLING FOR MODEL 'anon_model' NOW (CHAIN 2).
## Chain 2:
## Chain 2: Gradient evaluation took 5e-06 seconds
## Chain 2: 1000 transitions using 10 leapfrog steps per transition would take 0.05 seconds.
## Chain 2: Adjust your expectations accordingly!
## Chain 2:
## Chain 2:
## Chain 2: Iteration: 1 / 2000 [ 0%] (Warmup)
## Chain 2: Iteration: 200 / 2000 [ 10%] (Warmup)
## Chain 2: Iteration: 400 / 2000 [ 20%] (Warmup)
## Chain 2: Iteration: 600 / 2000 [ 30%] (Warmup)
## Chain 2: Iteration: 800 / 2000 [ 40%] (Warmup)
## Chain 2: Iteration: 1000 / 2000 [ 50%] (Warmup)
## Chain 2: Iteration: 1001 / 2000 [ 50%] (Sampling)
## Chain 2: Iteration: 1200 / 2000 [ 60%] (Sampling)
## Chain 2: Iteration: 1400 / 2000 [ 70%] (Sampling)
## Chain 2: Iteration: 1600 / 2000 [ 80%] (Sampling)
## Chain 2: Iteration: 1800 / 2000 [ 90%] (Sampling)
## Chain 2: Iteration: 2000 / 2000 [100%] (Sampling)
## Chain 2:
## Chain 2: Elapsed Time: 0.146 seconds (Warm-up)
## Chain 2: 0.146 seconds (Sampling)
## Chain 2: 0.292 seconds (Total)
## Chain 2:
##
## SAMPLING FOR MODEL 'anon_model' NOW (CHAIN 3).
## Chain 3:
## Chain 3: Gradient evaluation took 1e-05 seconds
## Chain 3: 1000 transitions using 10 leapfrog steps per transition would take 0.1 seconds.
## Chain 3: Adjust your expectations accordingly!
## Chain 3:
## Chain 3:
## Chain 3: Iteration: 1 / 2000 [ 0%] (Warmup)
## Chain 3: Iteration: 200 / 2000 [ 10%] (Warmup)
## Chain 3: Iteration: 400 / 2000 [ 20%] (Warmup)
## Chain 3: Iteration: 600 / 2000 [ 30%] (Warmup)
## Chain 3: Iteration: 800 / 2000 [ 40%] (Warmup)
## Chain 3: Iteration: 1000 / 2000 [ 50%] (Warmup)
## Chain 3: Iteration: 1001 / 2000 [ 50%] (Sampling)
## Chain 3: Iteration: 1200 / 2000 [ 60%] (Sampling)
## Chain 3: Iteration: 1400 / 2000 [ 70%] (Sampling)
## Chain 3: Iteration: 1600 / 2000 [ 80%] (Sampling)
## Chain 3: Iteration: 1800 / 2000 [ 90%] (Sampling)
## Chain 3: Iteration: 2000 / 2000 [100%] (Sampling)
## Chain 3:
## Chain 3: Elapsed Time: 0.143 seconds (Warm-up)
## Chain 3: 0.143 seconds (Sampling)
## Chain 3: 0.286 seconds (Total)
## Chain 3:
##
## SAMPLING FOR MODEL 'anon_model' NOW (CHAIN 4).
## Chain 4:
## Chain 4: Gradient evaluation took 6e-06 seconds
## Chain 4: 1000 transitions using 10 leapfrog steps per transition would take 0.06 seconds.
## Chain 4: Adjust your expectations accordingly!
## Chain 4:
## Chain 4:
## Chain 4: Iteration: 1 / 2000 [ 0%] (Warmup)
## Chain 4: Iteration: 200 / 2000 [ 10%] (Warmup)
## Chain 4: Iteration: 400 / 2000 [ 20%] (Warmup)
## Chain 4: Iteration: 600 / 2000 [ 30%] (Warmup)
## Chain 4: Iteration: 800 / 2000 [ 40%] (Warmup)
## Chain 4: Iteration: 1000 / 2000 [ 50%] (Warmup)
## Chain 4: Iteration: 1001 / 2000 [ 50%] (Sampling)
## Chain 4: Iteration: 1200 / 2000 [ 60%] (Sampling)
## Chain 4: Iteration: 1400 / 2000 [ 70%] (Sampling)
## Chain 4: Iteration: 1600 / 2000 [ 80%] (Sampling)
## Chain 4: Iteration: 1800 / 2000 [ 90%] (Sampling)
## Chain 4: Iteration: 2000 / 2000 [100%] (Sampling)
## Chain 4:
## Chain 4: Elapsed Time: 0.141 seconds (Warm-up)
## Chain 4: 0.146 seconds (Sampling)
## Chain 4: 0.287 seconds (Total)
## Chain 4:
# set priors for null model
priors_null <- c(
prior(normal(0, 1.5), class = "Intercept"), # centered prior for baseline log-odds
prior(student_t(3, 0, 1), class = "sd") # half-Student-t
)
# null model
brm_infer_sport_null <-
brm(infer_sport ~ 1 + (1 | participant),
data = data_tidy, family = bernoulli(),
prior = priors_null,
save_pars = save_pars(all = TRUE))
# get bf
bf <- bayes_factor(brm_infer_sport, brm_infer_sport_null)$bf
bf
A Bayesian analysis shows very strong evidence against a condition effect (BF = 0).
Participants were shown Alex’s Gorp friends again, and were asked to infer whether Gorps on Gorp Planet liked the sample-majority sport “less”, “the same”, or “more” than Alex’s Gorp friends.
# multinomial regression: do responses differ by condition?
infer_comp_multinom <- data %>%
multinom(infer_comp ~ condition, data = .) %>%
Anova()
infer_comp_multinom
There was no evidence participants made different responses to the comparison forced-choice question depending on condition (LR Chisq(2) = 1.05, p = 0.591).
In the prediction trials figure, predictions appear to be bimodal in the skewed condition: although the condition mean appears similar to .75 (as predicted by genuine adjustment), participants’ individual means appear to cluster at either .5 (discounting generalization, or guessing) or 1 (direct generalization from sample).
To test for bimodality in the skewed condition, we treat each participant’s prediction trials as draws from a Binomial(4, p) process, and ask whether participants’ responses in the skewed condition are better described by a single shared p (a unimodal population, with ordinary binomial sampling noise across the 4 trials) or by a finite mixture of subpopulations with different p’s (a truly bimodal population).
library(flexmix) # bimodality analysis, loads modeltools, load after brms models so modeltools doesn't override brms::prior
d_mix <- data_infer_summary %>%
ungroup() %>%
mutate(fail = infer_non_na_count - infer_sport_count)
d_mix_skewed <- d_mix %>% filter(condition == "skewed")
# fit binomial mixtures with 1, 2, and 3 latent classes
set.seed(42)
fm_skewed <- map(1:3, ~ flexmix(
cbind(infer_sport_count, fail) ~ 1,
data = d_mix_skewed,
k = .x,
model = FLXMRglm(family = "binomial")
))
bic_skewed <- tibble(
k = 1:3,
AIC = map_dbl(fm_skewed, AIC),
BIC = map_dbl(fm_skewed, BIC)
)
# class-specific probabilities & sizes for the 2-class model, ordered low to high
comp_p_raw <- plogis(parameters(fm_skewed[[2]])) %>% as.numeric()
comp_n_raw <- table(factor(clusters(fm_skewed[[2]]), levels = 1:2))
comp_stats <- tibble(p = comp_p_raw, n = as.numeric(comp_n_raw)) %>% arrange(p)
| Binomial mixture model comparison (skewed condition) | ||
| k | AIC | BIC |
|---|---|---|
| 1 | 505.4 | 508.7 |
| 2 | 453.9 | 463.6 |
| 3 | 457.9 | 474.1 |
The data in the skewed condition is best fit by a mixture of two binomial classes (BIC: 508.7 for a single shared p vs. 463.6 for 2 classes, vs. 474.1 for 3 classes), with estimated classes at p = 0.6 (n = 93) and p = 1 (n = 99).
To test the specific bimodality of exactly .5 (guessing/discounting) and exactly 1 (generalizing) — we compare a single free-p binomial model against a 2-class mixture with class probabilities fixed at .5 and 1 where only the mixing weight is estimated.
Because the regularity conditions for a standard chi-square likelihood-ratio test do not hold when comparing mixture models with different numbers of components, we obtain a null distribution via parametric bootstrap (simulating from the fitted single-p null model).
loglik_binom <- function(y, n, p) sum(dbinom(y, n, p, log = TRUE))
# M0: single free p shared by everyone
p_hat <- sum(d_mix_skewed$infer_sport_count) / sum(d_mix_skewed$infer_non_na_count)
ll_M0 <- loglik_binom(d_mix_skewed$infer_sport_count, 4, p_hat)
# M1: 2-class mixture, p fixed at .5 and 1, only mixing weight (pi) estimated
negloglik_mix_fixed <- function(logit_pi, y, n) {
pi_ <- plogis(logit_pi)
lik <- pi_ * dbinom(y, n, 0.5) + (1 - pi_) * dbinom(y, n, 1)
-sum(log(pmax(lik, 1e-300)))
}
opt_M1 <- optimize(negloglik_mix_fixed, interval = c(-10, 10),
y = d_mix_skewed$infer_sport_count, n = 4)
pi_hat <- plogis(opt_M1$minimum)
ll_M1 <- -opt_M1$objective
lrt_obs <- 2 * (ll_M1 - ll_M0)
# parametric bootstrap null distribution: simulate under M0, refit both models each time
set.seed(42)
n_boot <- 1000
n_obs <- nrow(d_mix_skewed)
lrt_boot <- map_dbl(1:n_boot, function(i) {
y_sim <- rbinom(n_obs, 4, p_hat)
ll0 <- loglik_binom(y_sim, 4, sum(y_sim) / (4 * n_obs))
opt1 <- optimize(negloglik_mix_fixed, interval = c(-10, 10), y = y_sim, n = 4)
2 * (-opt1$objective - ll0)
})
p_boot <- (sum(lrt_boot >= lrt_obs) + 1) / (n_boot + 1)
The fixed .5/1 mixture (mixing weight 0.52 at .5, 0.48 at 1) fits substantially better than a single shared probability (p = 0.78, LRT = 44.8, parametric-bootstrap p < .001). This supports treating the skewed condition as a mixture of guessing/discounting (0.5) and generalizing (1) participants, rather than a single population.
Finally, we test whether class membership (0.5 vs 1) depends on condition, by fitting a single 2-class mixture across both conditions with condition entered as a predictor of class membership (a “concomitant variable” mixture model), and comparing it via likelihood-ratio test to a version where condition does not predict class membership. Because both models fix the number of classes at 2 and differ only in the (ordinary, regular) concomitant logistic model, a standard chi-square LRT is valid here.
set.seed(42)
fm_null_mix <- flexmix(cbind(infer_sport_count, fail) ~ 1,
data = d_mix, k = 2,
model = FLXMRglm(family = "binomial"),
concomitant = FLXPmultinom(~ 1))
set.seed(42)
fm_cond_mix <- flexmix(cbind(infer_sport_count, fail) ~ 1,
data = d_mix, k = 2,
model = FLXMRglm(family = "binomial"),
concomitant = FLXPmultinom(~ condition))
ll0 <- logLik(fm_null_mix)
ll1 <- logLik(fm_cond_mix)
lrt_cond <- as.numeric(2 * (ll1 - ll0))
df_cond <- attr(ll1, "df") - attr(ll0, "df")
p_cond <- pchisq(lrt_cond, df = df_cond, lower.tail = FALSE)
# identify which cluster index is the "chance-like" (lower p) one
cond_p <- plogis(parameters(fm_cond_mix)) %>% as.numeric()
chance_cluster <- which.min(cond_p)
d_mix$cluster <- clusters(fm_cond_mix)
cluster_by_condition <- d_mix %>%
count(condition, cluster) %>%
group_by(condition) %>%
mutate(prop = n / sum(n))
skewed_chance_prop <- cluster_by_condition %>%
filter(condition == "skewed", cluster == chance_cluster) %>% pull(prop)
notskewed_chance_prop <- cluster_by_condition %>%
filter(condition == "not_skewed", cluster == chance_cluster) %>% pull(prop)
| Mixture class membership by condition | ||
| cluster | n | prop |
|---|---|---|
| not_skewed | ||
| 1 | 145 | 75.9% |
| 2 | 46 | 24.1% |
| skewed | ||
| 1 | 99 | 51.6% |
| 2 | 93 | 48.4% |
Condition significantly predicts response pattern (LRT(1) = 24.9, p < .001): 48.4% of skewed condition participants fall into the 0.5 class, compared to 24.1% of not skewed condition participants. This confirms the bimodal split is substantially more pronounced in the skewed condition than in the not skewed condition.
Participants’ responses to the comparison forced-choice question appear to be modestly related to their earlier responses to the prediction trials.
The comparison forced-choice asked who likes the sample-majority sport more: Alex’s Gorp friends, Gorps on Gorp Planet, or if they like it the same.
Since 8 out of 8 of Alex’s Gorp friends liked the sample-majority sport, matching the sample proportion would correspond with predictions of 1 (100%). As a result:
predictions of 0, 0.25, and 0.5, 0.75 should correspond with responding “Alex’s Gorp friends”
predictions of 1 should correspond with responding “the same”
Indeed, we see that participants were more likely to say “the same” rather than “Alex’s Gorp friends” after making 100% predictions compared to after other predictions.
Likewise, if participants were responding consistently across the two measures, we would expect that:
participants who responded that Alex’s Gorp friends like [sample-majority sport] “the same” as Gorps on Gorp Planet should show higher levels of generalization on the prediction trials, since Alex’s Gorp friends all liked soccer
participants who responded “Alex’s Gorp friends” like soccer more than Gorps on Gorp Planet should show lower levels of generalization on the prediction trials.
Indeed, participants who responded “the same” do show higher generalization in prediction trials than participants who responded “Alex’s Gorp friends”.
## R version 4.5.2 (2025-10-31)
## Platform: aarch64-apple-darwin20
## Running under: macOS Tahoe 26.7
##
## Matrix products: default
## BLAS: /System/Library/Frameworks/Accelerate.framework/Versions/A/Frameworks/vecLib.framework/Versions/A/libBLAS.dylib
## LAPACK: /Library/Frameworks/R.framework/Versions/4.5-arm64/Resources/lib/libRlapack.dylib; LAPACK version 3.12.1
##
## locale:
## [1] en_US.UTF-8/en_US.UTF-8/en_US.UTF-8/C/en_US.UTF-8/en_US.UTF-8
##
## time zone: America/New_York
## tzcode source: internal
##
## attached base packages:
## [1] stats graphics grDevices utils datasets methods base
##
## other attached packages:
## [1] flexmix_2.3-20 lattice_0.22-7 car_3.1-3
## [4] carData_3.0-5 nnet_7.3-20 tidybayes_3.0.7
## [7] broom.mixed_0.2.9.6 brms_2.23.0 Rcpp_1.1.2
## [10] lmerTest_3.2-0 lme4_2.0-6 Matrix_1.7-6
## [13] viridis_0.6.5 viridisLite_0.4.2 ggtext_0.1.2
## [16] lubridate_1.9.4 forcats_1.0.1 stringr_1.6.0
## [19] dplyr_1.1.4 purrr_1.2.1 readr_2.1.6
## [22] tidyr_1.3.2 tibble_3.3.1 ggplot2_4.0.3
## [25] tidyverse_2.0.0 gt_1.3.0 scales_1.4.0
## [28] janitor_2.2.1 here_1.0.2 knitr_1.51
##
## loaded via a namespace (and not attached):
## [1] RColorBrewer_1.1-3 tensorA_0.36.2.1 rstudioapi_0.18.0
## [4] jsonlite_2.0.0 magrittr_2.0.4 modeltools_0.2-24
## [7] TH.data_1.1-5 estimability_1.5.1 farver_2.1.2
## [10] nloptr_2.2.1 rmarkdown_2.30 fs_1.6.6
## [13] ragg_1.5.0 vctrs_0.7.1 minqa_1.2.8
## [16] base64enc_0.1-3 htmltools_0.5.9 curl_7.0.0
## [19] distributional_0.6.0 broom_1.0.12 Formula_1.2-5
## [22] StanHeaders_2.32.10 sass_0.4.10 parallelly_1.46.1
## [25] bslib_0.10.0 htmlwidgets_1.6.4 sandwich_3.1-1
## [28] emmeans_2.0.1 zoo_1.8-15 cachem_1.1.0
## [31] lifecycle_1.0.5 pkgconfig_2.0.3 R6_2.6.1
## [34] fastmap_1.2.0 rbibutils_2.4.1 future_1.69.0
## [37] snakecase_0.11.1 digest_0.6.39 numDeriv_2016.8-1.1
## [40] colorspace_2.1-2 furrr_0.3.1 ps_1.9.1
## [43] rprojroot_2.1.1 textshaping_1.0.4 Hmisc_5.2-5
## [46] labeling_0.4.3 timechange_0.3.0 abind_1.4-8
## [49] compiler_4.5.2 bit64_4.6.0-1 withr_3.0.2
## [52] inline_0.3.21 htmlTable_2.4.3 S7_0.2.1
## [55] backports_1.5.0 QuickJSR_1.9.0 pkgbuild_1.4.8
## [58] MASS_7.3-65 loo_2.9.0 tools_4.5.2
## [61] foreign_0.8-90 otel_0.2.0 glue_1.8.0
## [64] callr_3.7.6 nlme_3.1-168 gridtext_0.1.5
## [67] grid_4.5.2 checkmate_2.3.3 cluster_2.1.8.1
## [70] generics_0.1.4 gtable_0.3.6 tzdb_0.5.0
## [73] data.table_1.18.0 hms_1.1.4 xml2_1.5.2
## [76] pillar_1.11.1 ggdist_3.3.3 vroom_1.6.7
## [79] posterior_1.6.1 splines_4.5.2 survival_3.8-6
## [82] bit_4.6.0 tidyselect_1.2.1 reformulas_0.4.3.1
## [85] arrayhelpers_1.1-2 gridExtra_2.3 V8_8.0.1
## [88] stats4_4.5.2 xfun_0.56 bridgesampling_1.2-1
## [91] matrixStats_1.5.0 rstan_2.32.7 stringi_1.8.7
## [94] yaml_2.3.12 boot_1.3-32 evaluate_1.0.5
## [97] codetools_0.2-20 cli_3.6.5 rpart_4.1.24
## [100] RcppParallel_5.1.11-1 xtable_1.8-4 systemfonts_1.3.1
## [103] Rdpack_2.6.5 processx_3.8.6 jquerylib_0.1.4
## [106] globals_0.18.0 coda_0.19-4.1 svUnit_1.0.8
## [109] parallel_4.5.2 rstantools_2.6.0 bayesplot_1.15.0
## [112] Brobdingnag_1.2-9 listenv_0.10.0 ggthemes_5.2.0
## [115] mvtnorm_1.3-3 crayon_1.5.3 rlang_1.1.7
## [118] multcomp_1.4-29