Overview

Project goals

The goal of this project is to establish if children and adults can adjust their generalizations about a social group to account for sampling skew.

Previously on..

In Study 1b (boat study), 4yo failed to show sensitivity to probabilistic structural skew in social group generalization, while 7yo succeed (although the latter’s success is complicated by potential revealed population information).

In contrast, even infants succeed in discounting samples skewed by an agent’s deterministic preference in physical generalization (Xu & Denison, 2009).

Do 4-7yo succeed in discounting samples skewed by an agent’s deterministic preference in social group generalization?

Results

Children largely passed all the checks, indicating that they knew what soccer and basketball balls were, that they knew Alex either had no preference (not skewed) or had a preference for one sport (skewed) among kids, and among Gorps.

In contrast to adults, children were not sensitive to the agent’s skewed sampling in either condition.

  • Children made similar predictions in the skewed condition than in the not skewed condition.

  • On the comparison forced-choice, children made similar responses in both conditions, either responding that Alex’s Gorp friends liked soccer “the same” or “more” than Gorps on Gorp Planet. Unlike adults, children’s predictions did not seem related to their responses to the comparison forced-choice (see prediction vs comparison forced-choice).

  • There was no evidence of an interaction between condition and age (exact) on children’s predictions. (see age results)

  • Children’s predictions were similar across trial number, regardless of condition. Even on the first trial, there was no evidence of condition differences. (see prediction by trial number)

Children’s predictions indicated that they generalized similarly in both conditions.

  • Children’s predictions in each condition were above chance, indicating generalization from the sample to Gorps in general.

Surprisingly, compared to adults, generalization was relatively weak in the not skewed condition.

  • In not skewed condition, as in the skewed condition, about half of children picked the sample property at chance levels (2/4 trials) and about half of children above chance (3/4 or 4/4 trials). Analysis of bimodality suggests similar mixtures of two response distributions in each condition.

  • There is a significant interaction between children and adults and condition on predictions, but surprisingly, the interaction largely lies in predictions in the not skewed condition, rather than in the skewed condition.

Discussion

Why are children making such different predictions in the not skewed condition compared to adults? Some ideas:

  • maybe children are more likely than adults to produce chance responding in general (almost certainly true), and/or alternating responses on repeated measures (which would result in a chance profile)

  • maybe children compared to adults extract a stronger prior from familiarization that soccer/basketball proportion is going to be 0.5 (because same number of kids at playground liked soccer and basketball, and we ask a question to check participants understand this)

    • seems odd though that the prior would be so strong to mostly overcome seeing 8 out of 8 soccer-playing Gorps though, which seems like a pretty strong sample
  • maybe children compared to adults have weaker overhypotheses that groups share (sport) preferences so are more hesitant to generalize from the sample of 8 soccer Gorps to Gorps in general

All of these would draw responses generally towards chance, although I don’t think they help explain why children are making such different predictions from adults in the not skewed condition specifically (their predictions are no different from adults in the skewed condition).

If children are truly insensitive to the agent’s skewed sampling in this task, then why did infants seem to be sensitive to skewed sampling in Xu & Denison (2009)?

  • Evidence from non-random condition in Xu & Denison (2009) is a frequentist null effect, which is weak, indirect: no difference in infants’ looking time to scene based on likelihood of the sample from box contents (all red or all white sample from mostly red or mostly white box), combining across trials where sample is consistent vs inconsistent with agent’s preferences (agent prefers red or white)

  • Reasoning about social vs non-social population - priors about social populations?

  • Visibility of sampling process - sampling process visible in Xu & Denison (2009), not in our study, which may make it easier to take it into account

  • Size of hypothesis space - constrained in Xu & Denison (2009) (infants are familiarized to box being mostly white or mostly red), vs unconstrained in our study (prior set to 0.5 but that’s it)

  • Implicit vs explicit measures - looking time to outcomes vs prediction about individual members

Methods

The study was preregistered on OSF.

Participants

219 children participated by submitting a valid PANDA video Thurs 9/25 - Fri 9/26/2026.

Participants’ families were paid $10 for an estimated 10-15 minute task. Children’s parent/guardian provided consent and children provided assent.

After applying exclusion criteria, the final sample included 209 children (n = 19-31 in each of the 8 age*condition bins).

Exclusion criteria

A total of 10 participants (4.6% of all participants) were excluded for meeting at least 1 of the following exclusion criteria:

  • not completing the survey (n = 3)

Demographics

Of included participants.

Age

age
mean sd n
6.08 1.17 209
age groups
age_cat n
4 45
5 58
6 44
7 62
age groups by condition
not_skewed skewed
4
20 25
5
29 29
6
25 19
7
31 31

Gender

gender n prop
female 116 55.5%
male 93 44.5%

Race

race n prop
Caucasian 88 42.1%
Asian 35 16.7%
African American 22 10.5%
Asian, Caucasian 13 6.2%
Hispanic 13 6.2%
Caucasian, Hispanic 11 5.3%
African American, Caucasian 7 3.3%
NA 4 1.9%
Asian, Caucasian, Hispanic 3 1.4%
Hawaiian Pacific 3 1.4%
African American, Hispanic 2 1.0%
Asian, Hispanic 2 1.0%
White & Asian 2 1.0%
African American, Caucasian, Hispanic 1 0.5%
African American, Caucasian, Native American 1 0.5%
Asian, Caucasian, Native American 1 0.5%
Middle Eastern 1 0.5%

Parental education

education n prop
High school/GED 7 3.3%
Some college 26 12.4%
Bachelor's (B.A., B.S.) 83 39.7%
Master's (M.A., M.S.) 59 28.2%
Doctoral (Ph.D., J.D., M.D.) 27 12.9%
Prefer not to specify 3 1.4%
NA 4 1.9%

Geographic location

country n
US 209
state n
AL 3
AR 1
AZ 2
CA 28
CO 6
CT 2
DC 4
FL 15
GA 10
HI 1
IA 4
IL 2
IN 1
KY 2
LA 1
MA 6
MD 7
MI 4
MN 4
MO 2
NC 7
ND 2
NJ 9
NV 1
NY 21
OH 6
OK 1
OR 2
PA 9
RI 2
TN 3
TX 16
UT 5
VA 11
VT 1
WA 2
WI 3
WV 2
NA 1

Almost all participants (n = 209) were based in the United States, from across 38 states.

Procedure

This study was administered as a Qualtrics survey, and approved by the NYU IRB (IRB-FY2024-9169).

After providing their consent, participants completed a captcha and sound check, and were asked to watch videos sound on. Participants then watched the following videos in order:

  1. In the warmup phase, to confirm participants’ understanding of balls and sports, participants heard the narrator label a soccer ball and a basketball, and were asked to click on each.

    In the alternate counterbalance version, the left/right position of soccer and basketball buttons was switched on this and all questions in the study.

  2. In the familiarization phase, participants were introduced to an agent called Alex (depicted using a photograph of a white female child), and learned how Alex chooses friends by watching her make friends at a playground.

    Each trial showed pictures of two children (pictures matched on race and gender; races and genders varied across trials), one holding a soccer ball and another holding a basketball. Children were unique to each trial.

In the skewed condition, Alex approached the child holding the soccer ball on 6 out of 6 trials. Trial order was randomized.

In the not skewed condition, Alex approached children of each sport on 3 out of 6 trials. Sport selections alternated, with the first selection being randomized.

In the alternate counterbalance version, the position of children was fixed, while Alex’s selections were switched, such that the skewed condition saw Alex approach mostly basketball.

  1. As familiarization phase checks, participants were asked to (in the following fixed order):
  1. Familiarization: friends check: Predict which child Alex might befriend between a novel soccer kid and a novel basketball kid, to confirm their understanding of the agent’s preference. The images used were fixed images of a Black girl holding a basketball (fixed), and another Black girl holding a soccer ball (fixed), their positions counterbalanced on screen.

After responding, participants in the skewed condition were told that Alex will probably choose the kid holding the soccer ball, because Alex likes soccer (or basketball, in the alternate counterbalance version). Participants in the not skewed condition were told, “it might be hard for Alex to choose, because Alex likes soccer and basketball”.

  1. Familiarization: sport base rate check: Confirm their understanding of whether kids at the playground liked basketball, soccer, or both the same. After responding, all participants were told that kids at the playground liked both the same.
  1. In the sample observation phase, participants observed the same sample of Gorps that Alex befriends on Gorp Planet. In both conditions, Alex befriends 8 Gorps, all of which (fixed positions and colors) like soccer. (In the alternate counterbalance version, the 8 like basketball.)

  1. In the inference phase, participants had to make inferences about Gorps in general, after Alex left. They completed the below measures in fixed order:
  1. Inference: prediction trials: As one of our dependent measures, participants were asked to predict the sport preferences of 4 novel Gorps (a brown, pink, purple, and teal Gorp in fixed order). These four trials were averaged into a prediction proportion for each participant.

  1. Inference: comparison forced-choice: As another dependent measure, adults were shown Alex’s Gorp friends again, and were asked to make a forced-choice comparison: Who liked soccer more? Alex’s Gorp friends, Gorps on Gorp Planet, or do they like soccer the same. In the counterbalanced version, this question asked about basketball.

  1. As a final friends check, participants were asked which Gorp Alex might befriend between a soccer Gorp and a basketball Gorp, to check whether they understood the agent’s preference for kids extended to Gorps. The images used were fixed images of a pink Gorp and a purple Gorp from the prediction phase (which had soccer ball and which had basketball was counterbalanced).

  1. Finally, participants’ parent/guardian were asked for any problems or confusion they had, and demographic information.

Checks

Ball training

After labeling, most participants correctly identified a soccer ball and a basketball.

Familiarization: friends check (kids)

Both conditions largely passed the friends check.

Participants in the skewed condition understood Alex would befriend kids who prefered one sport (aligned with their counterbalance condition), while participants in the not skewed condition appeared more mixed.

All participants received information about the expected response after their response.

Familiarization: sport base rate check

Participants were asked to recall that Alex met many kids at the park today, and to recall which sport more of the kids on the playground liked: basketball, soccer, or did they like them the same.

The correct answer to this question is “the same”, as every trial showed a basketball kid and a soccer kid.

Participants mostly answered this question correctly; participants who gave incorrect answers were still included. All participants received the correct answer after their response.

Agent friends check (Gorps)

Both conditions largely passed the agent friends check.

Participants in the skewed condition understood Alex would befriend the Gorp who preferred one sport (aligned with their counterbalance condition), while participants in the not skewed condition appeared more mixed.

Main results

Predictions

Participants were asked to predict the sport preferences (soccer or basketball) of 4 novel group members (a brown, pink, purple, and teal Gorp in fixed order).

First prediction trial:

infer_sport_prop not_skewed skewed
0.00 4% 3%
0.25 8% 8%
0.50 51% 41%
0.75 15% 21%
1.00 22% 27%
condition infer_sport_avg infer_sport_sd
not_skewed 61% 26%
skewed 65% 26%
glmer_infer_sport <-
  glmer(infer_sport ~ condition + (1 | participant), 
        data = data_tidy, 
        family = binomial)

# condition difference?
glmer_infer_sport %>% 
  summary()

There was no evidence children made different predictions in the skewed vs not skewed conditions about the sport preferences of novel Gorps (b = 0.2, z = 1.25, p = 0.212). This result suggests children were not sensitive to the agent’s skew in this study.

# set priors
priors <- c(
  prior(normal(0, 1.5), class = "Intercept"), # centered prior for baseline log-odds
  prior(normal(0, 1),   class = "b"), # weakly informative prior
  prior(student_t(3, 0, 1), class = "sd") # half-Student-t
)

# bayesian model
brm_infer_sport <-
  brm(infer_sport ~ condition + (1 | participant), 
      data = data_tidy, 
      family = bernoulli(link = "logit"),
      prior = priors,
      save_pars = save_pars(all = TRUE))

# set priors for null model
priors_null <- c(
  prior(normal(0, 1.5), class = "Intercept"), # centered prior for baseline log-odds
  prior(student_t(3, 0, 1), class = "sd") # half-Student-t
)

# null model
brm_infer_sport_null <- 
  brm(infer_sport ~ 1 + (1 | participant), 
      data = data_tidy, family = bernoulli(),
      prior = priors_null, 
      save_pars = save_pars(all = TRUE))

# get bf
bf <- bayes_factor(brm_infer_sport, brm_infer_sport_null)

A Bayesian analysis indicates anecdotal evidence against a condition difference in predictions, indicating that children made similar predictions in the skewed vs not skewed conditions about the sport preferences of novel Gorps (BF = 0.5).

# compare each condition to chance (50%)
# = test if intercept differs from 0 (=logit of 0.5)

# not skewed condition
glmer_infer_sport_notskewed_chance <-
  glmer(infer_sport ~ 1 + (1 | participant),
        data = data_tidy %>% 
          filter(condition == "not_skewed"),
        family = binomial)

glmer_infer_sport_notskewed_chance %>%
  summary()

# skewed condition
glmer_infer_sport_skewed_chance <-
  glmer(infer_sport ~ 1 + (1 | participant),
        data = data_tidy %>% 
          filter(condition == "skewed"),
        family = binomial)

glmer_infer_sport_skewed_chance %>%
  summary()

In both conditions, participants’ predictions differed from chance. In the not skewed condition, participants predicted the sample sport more than chance (b = 0.46, z = 4.19, p < .001). In the skewed condition, participants predicted the sample sport more than chance as well (b = 0.68, z = 5.5, p < .001),

Comparison forced-choice

Participants were shown Alex’s Gorp friends again, and were asked to infer whether Gorps on Gorp Planet liked the sample sport “less”, “the same”, or “more” than Alex’s Gorp friends.

# multinomial regression: do responses differ by condition?
infer_comp_multinom <- data %>% 
  multinom(infer_comp ~ condition, data = .) %>% 
  Anova()

infer_comp_multinom

Participants did not make different responses to the comparison forced-choice question depending on condition (LR Chisq(2) = 0.9, p = 0.637).

# compare "population" responses
glm_pop <- data %>% 
  mutate(infer_comp_pop = ifelse(infer_comp == "Gorps on Gorp Planet", 1, 0)) %>% 
  glm(infer_comp_pop ~ condition, 
      data = .,
      family = binomial)

glm_pop %>% 
  summary()

# compare "the same" responses
glm_same <- data %>% 
  mutate(infer_comp_same = ifelse(infer_comp == "the same", 1, 0)) %>% 
  glm(infer_comp_same ~ condition, 
      data = .,
      family = binomial)

glm_same %>% 
  summary()

# compare "sample" responses
glm_sample <- data %>% 
  mutate(dv_comp_sample = ifelse(infer_comp == "Alex's Gorp friends", 1, 0)) %>% 
  glm(dv_comp_sample ~ condition, 
      data = .,
      family = binomial)

glm_sample %>% 
  summary()

Specifically, there was no evidence of a condition difference in responding with “Alex’s Gorp friends” (b = 0.26, z = 0.91, p = 0.362), “the same” (b = -0.17, z = -0.62, p = 0.536), or “Gorps on Gorp Planet” (b = -0.23, z = -0.47, p = 0.638).

Age results

Prediction by age

First plot made by Shiloh (RA).

glmer_infer_sport_by_age <-
  glmer(infer_sport ~ condition * age_exact + (1 | participant), 
        data = data_tidy, 
        family = binomial)

# condition * age_exact interaction?
glmer_infer_sport_by_age %>% 
  summary()

There was no evidence of an interaction between condition and age (exact) on prediction profiles (b = -0.01, z = -0.05, p = 0.958).

Supplementary results

Prediction by trial number

Conditions appear largely consistent in predictions across trial number/across novel Gorps (novel Gorps were presented in fixed order in prediction trials).

Even in the first trial, there does not appear to be a difference between conditions.

glmer_infer_sport_trial_1 <-
  glmer(infer_sport ~ condition + (1 | participant), 
        data = data_tidy %>% filter(infer_trial_num == 1),
        family = binomial)

# condition difference?
glmer_infer_sport_trial_1 %>% 
  summary()

On trial 1, there was no evidence children made different predictions in the skewed vs not skewed conditions about the sport preferences of novel Gorps (b = 0.03, z = 0.1, p = 0.922).

Predictions: bimodality in responses

To test for bimodality, we treat each participant’s prediction trials as draws from a Binomial(4, p) process, and ask whether participants’ responses in the skewed condition are better described by a single shared p (a unimodal population, with ordinary binomial sampling noise across the 4 trials) or by a finite mixture of subpopulations with different p’s (a truly bimodal population).

Unimodality vs multimodality

library(flexmix)    # bimodality analysis, loads modeltools, load after brms models so modeltools doesn't override brms::prior
set.seed(42)

d_mix <- data_infer_summary %>%
  ungroup() %>%
  mutate(fail = infer_non_na_count - infer_sport_count)

d_mix_skewed <- d_mix %>% 
  filter(condition == "skewed")

d_mix_not_skewed <- d_mix %>% 
  filter(condition == "not_skewed")
# fit binomial mixtures with 1, 2, and 3 latent classes
fm_not_skewed <- map(1:3, ~ flexmix(
  cbind(infer_sport_count, fail) ~ 1,
  data = d_mix_not_skewed,
  k = .x,
  model = FLXMRglm(family = "binomial")
))

bic_not_skewed <- tibble(
  k = 1:3,
  AIC = map_dbl(fm_not_skewed, AIC),
  BIC = map_dbl(fm_not_skewed, BIC)
)

# class-specific probabilities & sizes for the 2-class model, ordered low to high
not_skewed_comp_p_raw <- plogis(parameters(fm_not_skewed[[2]])) %>% as.numeric()
not_skewed_comp_n_raw <- table(factor(clusters(fm_not_skewed[[2]]), levels = 1:2))
not_skewed_comp_stats <- tibble(p = not_skewed_comp_p_raw, n = as.numeric(not_skewed_comp_n_raw)) %>% arrange(p)
Binomial mixture model comparison (not skewed condition)
k AIC BIC
1 303.9 306.5
2 298.6 306.5
3 302.6 315.8

The data in the not skewed condition is best fit by a mixture of two binomial classes (BIC: 306.5 for a single shared p vs. 306.5 for 2 classes, vs. 315.8 for 3 classes), with estimated classes at p = 0.54 (n = 82) and p = 1 (n = 23).

# fit binomial mixtures with 1, 2, and 3 latent classes
fm_skewed <- map(1:3, ~ flexmix(
  cbind(infer_sport_count, fail) ~ 1,
  data = d_mix_skewed,
  k = .x,
  model = FLXMRglm(family = "binomial")
))

bic_skewed <- tibble(
  k = 1:3,
  AIC = map_dbl(fm_skewed, AIC),
  BIC = map_dbl(fm_skewed, BIC)
)

# class-specific probabilities & sizes for the 2-class model, ordered low to high
skewed_comp_p_raw <- plogis(parameters(fm_skewed[[2]])) %>% as.numeric()
skewed_comp_n_raw <- table(factor(clusters(fm_skewed[[2]]), levels = 1:2))
skewed_comp_stats <- tibble(p = skewed_comp_p_raw, n = as.numeric(skewed_comp_n_raw)) %>% arrange(p)
Binomial mixture model comparison (skewed condition)
k AIC BIC
1 301.4 304.0
2 295.8 303.7
3 299.8 313.0

The data in the skewed condition is best fit by a mixture of two binomial classes (BIC: 304 for a single shared p vs. 303.7 for 2 classes, vs. 313 for 3 classes), with estimated classes at p = 0.58 (n = 76) and p = 1 (n = 28).

Best unimodal model vs best bimodal mixture of .5 and 1

To test the specific bimodality of exactly .5 (guessing/discounting) and exactly 1 (generalizing) — we compare a single free-p binomial model against a 2-class mixture with class probabilities fixed at .5 and 1 where only the mixing weight is estimated.

Because the regularity conditions for a standard chi-square likelihood-ratio test do not hold when comparing mixture models with different numbers of components, we obtain a null distribution via parametric bootstrap (simulating from the fitted single-p null model).

loglik_binom <- function(y, n, p) sum(dbinom(y, n, p, log = TRUE))

# M0: single free p shared by everyone
p_hat <- sum(d_mix_not_skewed$infer_sport_count) / sum(d_mix_not_skewed$infer_non_na_count)
ll_M0 <- loglik_binom(d_mix_not_skewed$infer_sport_count, 4, p_hat)

# M1: 2-class mixture, p fixed at .5 and 1, only mixing weight (pi) estimated
negloglik_mix_fixed <- function(logit_pi, y, n) {
  pi_ <- plogis(logit_pi)
  lik <- pi_ * dbinom(y, n, 0.5) + (1 - pi_) * dbinom(y, n, 1)
  -sum(log(pmax(lik, 1e-300)))
}
opt_M1 <- optimize(negloglik_mix_fixed, interval = c(-10, 10),
                    y = d_mix_not_skewed$infer_sport_count, n = 4)
pi_hat <- plogis(opt_M1$minimum)
ll_M1 <- -opt_M1$objective

lrt_obs <- 2 * (ll_M1 - ll_M0)

# parametric bootstrap null distribution: simulate under M0, refit both models each time
set.seed(42)
n_boot <- 1000
n_obs <- nrow(d_mix_not_skewed)
lrt_boot <- map_dbl(1:n_boot, function(i) {
  y_sim <- rbinom(n_obs, 4, p_hat)
  ll0 <- loglik_binom(y_sim, 4, sum(y_sim) / (4 * n_obs))
  opt1 <- optimize(negloglik_mix_fixed, interval = c(-10, 10), y = y_sim, n = 4)
  2 * (-opt1$objective - ll0)
})
p_boot <- (sum(lrt_boot >= lrt_obs) + 1) / (n_boot + 1)

The fixed .5/1 mixture (mixing weight 0.83 at .5, 0.17 at 1) fits substantially better than a single shared probability (p = 0.61, LRT = 7.4, parametric-bootstrap p = 0.003). This supports treating the not skewed condition as a mixture of guessing/discounting (0.5) and generalizing (1) participants, rather than a single population.

loglik_binom <- function(y, n, p) sum(dbinom(y, n, p, log = TRUE))

# M0: single free p shared by everyone
p_hat <- sum(d_mix_skewed$infer_sport_count) / sum(d_mix_skewed$infer_non_na_count)
ll_M0 <- loglik_binom(d_mix_skewed$infer_sport_count, 4, p_hat)

# M1: 2-class mixture, p fixed at .5 and 1, only mixing weight (pi) estimated
negloglik_mix_fixed <- function(logit_pi, y, n) {
  pi_ <- plogis(logit_pi)
  lik <- pi_ * dbinom(y, n, 0.5) + (1 - pi_) * dbinom(y, n, 1)
  -sum(log(pmax(lik, 1e-300)))
}
opt_M1 <- optimize(negloglik_mix_fixed, interval = c(-10, 10),
                    y = d_mix_skewed$infer_sport_count, n = 4)
pi_hat <- plogis(opt_M1$minimum)
ll_M1 <- -opt_M1$objective

lrt_obs <- 2 * (ll_M1 - ll_M0)

# parametric bootstrap null distribution: simulate under M0, refit both models each time
set.seed(42)
n_boot <- 1000
n_obs <- nrow(d_mix_skewed)
lrt_boot <- map_dbl(1:n_boot, function(i) {
  y_sim <- rbinom(n_obs, 4, p_hat)
  ll0 <- loglik_binom(y_sim, 4, sum(y_sim) / (4 * n_obs))
  opt1 <- optimize(negloglik_mix_fixed, interval = c(-10, 10), y = y_sim, n = 4)
  2 * (-opt1$objective - ll0)
})
p_boot <- (sum(lrt_boot >= lrt_obs) + 1) / (n_boot + 1)

The fixed .5/1 mixture (mixing weight 0.78 at .5, 0.22 at 1) fits substantially better than a single shared probability (p = 0.65, LRT = 3.9, parametric-bootstrap p = 0.002). This supports treating the skewed condition as a mixture of guessing/discounting (0.5) and generalizing (1) participants, rather than a single population.

Condition differences in mixture

Finally, we test whether class membership (0.5 vs 1) depends on condition, by fitting a single 2-class mixture across both conditions with condition entered as a predictor of class membership (a “concomitant variable” mixture model), and comparing it via likelihood-ratio test to a version where condition does not predict class membership. Because both models fix the number of classes at 2 and differ only in the (ordinary, regular) concomitant logistic model, a standard chi-square LRT is valid here.

set.seed(42)
fm_null_mix <- flexmix(cbind(infer_sport_count, fail) ~ 1,
                        data = d_mix, k = 2,
                        model = FLXMRglm(family = "binomial"),
                        concomitant = FLXPmultinom(~ 1))
set.seed(42)
fm_cond_mix <- flexmix(cbind(infer_sport_count, fail) ~ 1,
                        data = d_mix, k = 2,
                        model = FLXMRglm(family = "binomial"),
                        concomitant = FLXPmultinom(~ condition))

ll0 <- logLik(fm_null_mix)
ll1 <- logLik(fm_cond_mix)
lrt_cond <- as.numeric(2 * (ll1 - ll0))
df_cond <- attr(ll1, "df") - attr(ll0, "df")
p_cond <- pchisq(lrt_cond, df = df_cond, lower.tail = FALSE)

# identify which cluster index is the "chance-like" (lower p) one
cond_p <- plogis(parameters(fm_cond_mix)) %>% as.numeric()
chance_cluster <- which.min(cond_p)

d_mix$cluster <- clusters(fm_cond_mix)
cluster_by_condition <- d_mix %>%
  count(condition, cluster) %>%
  group_by(condition) %>%
  mutate(prop = n / sum(n))

skewed_chance_prop <- cluster_by_condition %>%
  filter(condition == "skewed", cluster == chance_cluster) %>% pull(prop)

notskewed_chance_prop <- cluster_by_condition %>%
  filter(condition == "not_skewed", cluster == chance_cluster) %>% pull(prop)
Mixture class membership by condition
cluster n prop
not_skewed
1 82 78.1%
2 23 21.9%
skewed
1 76 73.1%
2 28 26.9%

Condition does not significantly predict response pattern (LRT(1) = 0.7, p = 0.398): 73.1% of skewed condition participants fall into the 0.5 class, compared to 78.1% of not skewed condition participants.

Prediction vs comparison forced-choice

Participants’ responses to the comparison forced-choice question appear to be unrelated to their earlier responses to the prediction trials.

The comparison forced-choice asked who likes the sample sport more: Alex’s Gorp friends, Gorps on Gorp Planet, or if they like it the same.

Since 8 out of 8 of Alex’s Gorp friends liked the sample sport, matching the sample proportion would correspond with predictions of 1 (100%). As a result:

  • predictions of 0%, 25%, 50%, 75% should correspond with responding “Alex’s Gorp friends” like the sample sport more

  • predictions of 100% should correspond with responding that Alex’s Gorp friends and Gorps on Gorp Planet like the sample sport “the same”

However, that’s not what we see: if anything, responses of “Alex’s Gorp friends” liking the sample sport more appear to be more common, not less common, after predictions of 100% than in other predictions.

Likewise, if participants were responding consistently across the two measures, we would expect that:

  • participants who responded that Alex’s Gorp friends like [sample sport] “the same” as Gorps on Gorp Planet should show higher levels of generalization on the prediction trials, since Alex’s Gorp friends all liked soccer

  • participants who responded “Alex’s Gorp friends” like soccer more than Gorps on Gorp Planet should show lower levels of generalization on the prediction trials.

However, that’s not what we see; generalization appears to be largely similar whether participants say “the same” or “Alex’s Gorp friends”.

Session info

## R version 4.5.2 (2025-10-31)
## Platform: aarch64-apple-darwin20
## Running under: macOS Tahoe 26.7
## 
## Matrix products: default
## BLAS:   /System/Library/Frameworks/Accelerate.framework/Versions/A/Frameworks/vecLib.framework/Versions/A/libBLAS.dylib 
## LAPACK: /Library/Frameworks/R.framework/Versions/4.5-arm64/Resources/lib/libRlapack.dylib;  LAPACK version 3.12.1
## 
## locale:
## [1] en_US.UTF-8/en_US.UTF-8/en_US.UTF-8/C/en_US.UTF-8/en_US.UTF-8
## 
## time zone: America/New_York
## tzcode source: internal
## 
## attached base packages:
## [1] stats     graphics  grDevices utils     datasets  methods   base     
## 
## other attached packages:
##  [1] flexmix_2.3-20      lattice_0.22-7      car_3.1-3          
##  [4] carData_3.0-5       nnet_7.3-20         tidybayes_3.0.7    
##  [7] broom.mixed_0.2.9.6 brms_2.23.0         Rcpp_1.1.2         
## [10] lmerTest_3.2-0      lme4_2.0-6          Matrix_1.7-6       
## [13] tidycensus_1.7.3    zipcodeR_0.3.5      viridis_0.6.5      
## [16] viridisLite_0.4.2   ggtext_0.1.2        lubridate_1.9.4    
## [19] forcats_1.0.1       stringr_1.6.0       dplyr_1.1.4        
## [22] purrr_1.2.1         readr_2.1.6         tidyr_1.3.2        
## [25] tibble_3.3.1        ggplot2_4.0.3       tidyverse_2.0.0    
## [28] gt_1.3.0            scales_1.4.0        janitor_2.2.1      
## [31] here_1.0.2          knitr_1.51         
## 
## loaded via a namespace (and not attached):
##   [1] svUnit_1.0.8          splines_4.5.2         rpart_4.1.24         
##   [4] lifecycle_1.0.5       Rdpack_2.6.5          sf_1.0-24            
##   [7] StanHeaders_2.32.10   rprojroot_2.1.1       processx_3.8.6       
##  [10] globals_0.18.0        vroom_1.6.7           MASS_7.3-65          
##  [13] ggdist_3.3.3          backports_1.5.0       magrittr_2.0.4       
##  [16] Hmisc_5.2-5           sass_0.4.10           rmarkdown_2.30       
##  [19] jquerylib_0.1.4       yaml_2.3.12           otel_0.2.0           
##  [22] pkgbuild_1.4.8        sp_2.2-0              DBI_1.2.3            
##  [25] minqa_1.2.8           RColorBrewer_1.1-3    multcomp_1.4-29      
##  [28] abind_1.4-8           rvest_1.0.5           TH.data_1.1-5        
##  [31] tensorA_0.36.2.1      rappdirs_0.3.4        sandwich_3.1-1       
##  [34] inline_0.3.21         listenv_0.10.0        terra_1.8-93         
##  [37] units_1.0-0           bridgesampling_1.2-1  parallelly_1.46.1    
##  [40] codetools_0.2-20      xml2_1.5.2            tidyselect_1.2.1     
##  [43] raster_3.6-32         bayesplot_1.15.0      farver_2.1.2         
##  [46] matrixStats_1.5.0     stats4_4.5.2          base64enc_0.1-3      
##  [49] jsonlite_2.0.0        e1071_1.7-17          Formula_1.2-5        
##  [52] survival_3.8-6        emmeans_2.0.1         systemfonts_1.3.1    
##  [55] tools_4.5.2           ragg_1.5.0            glue_1.8.0           
##  [58] gridExtra_2.3         mgcv_1.9-4            xfun_0.56            
##  [61] distributional_0.6.0  ggthemes_5.2.0        loo_2.9.0            
##  [64] withr_3.0.2           numDeriv_2016.8-1.1   fastmap_1.2.0        
##  [67] tigris_2.2.1          boot_1.3-32           callr_3.7.6          
##  [70] digest_0.6.39         timechange_0.3.0      R6_2.6.1             
##  [73] estimability_1.5.1    textshaping_1.0.4     colorspace_2.1-2     
##  [76] RSQLite_2.4.5         generics_0.1.4        data.table_1.18.0    
##  [79] class_7.3-23          httr_1.4.7            htmlwidgets_1.6.4    
##  [82] pkgconfig_2.0.3       gtable_0.3.6          modeltools_0.2-24    
##  [85] blob_1.3.0            S7_0.2.1              furrr_0.3.1          
##  [88] htmltools_0.5.9       posterior_1.6.1       reformulas_0.4.3.1   
##  [91] snakecase_0.11.1      rstudioapi_0.18.0     tzdb_0.5.0           
##  [94] uuid_1.2-2            coda_0.19-4.1         checkmate_2.3.3      
##  [97] nlme_3.1-168          curl_7.0.0            nloptr_2.2.1         
## [100] proxy_0.4-29          cachem_1.1.0          zoo_1.8-15           
## [103] KernSmooth_2.23-26    parallel_4.5.2        foreign_0.8-90       
## [106] pillar_1.11.1         grid_4.5.2            vctrs_0.7.1          
## [109] arrayhelpers_1.1-2    xtable_1.8-4          cluster_2.1.8.1      
## [112] htmlTable_2.4.3       evaluate_1.0.5        mvtnorm_1.3-3        
## [115] cli_3.6.5             compiler_4.5.2        rlang_1.1.7          
## [118] crayon_1.5.3          rstantools_2.6.0      labeling_0.4.3       
## [121] classInt_0.4-11       ps_1.9.1              fs_1.6.6             
## [124] stringi_1.8.7         rstan_2.32.7          QuickJSR_1.9.0       
## [127] V8_8.0.1              Brobdingnag_1.2-9     hms_1.1.4            
## [130] bit64_4.6.0-1         future_1.69.0         rbibutils_2.4.1      
## [133] gridtext_0.1.5        broom_1.0.12          memoise_2.0.1        
## [136] RcppParallel_5.1.11-1 bslib_0.10.0          bit_4.6.0