The main result is a script-by-lexicality pattern in reaction times among fluent Chinese readers. Real words were identified faster in Chinese characters than in Pinyin, whereas the corresponding script difference for nonwords was small and uncertain. Accuracy was high overall and showed little evidence for a corresponding main script effect in the primary analysis. The exploratory language-dominance analysis is hypothesis-generating rather than confirmatory, but it raises the possibility that lower Chinese-family exposure is associated with more errors on Chinese-script trials.
Eligibility required participants to be fluent Chinese readers. Chinese-reader eligibility was checked from LEAP-Q language-background responses and, where available, the experiment’s Chinese-reader gate. The final analysed sample included 75 Chinese-English bilinguals after excluding 9 participants with accuracy below 70%. Participants were aged 18-42 years (\(M = 24.55\), \(SD = 4.73\)). First-language responses were predominantly Mandarin/Chinese (59 participants), followed by English (10), Cantonese (3), and Fuzhou dialect (1). For second language, most participants reported English (59), with a smaller number reporting Mandarin/Chinese (11). 33 participants reported an additional language or dialect beyond their main Chinese-English background.
Language exposure was coded from free-text LEAP-Q responses. Mandarin, Chinese, Cantonese, Fuzhou dialect, Hokkien, and other Chinese-family varieties were grouped as Chinese-family exposure. 5 responses whose percentages summed to more than 100 were rescaled proportionally. Exposure could be coded for 73 participants. Mean English exposure was 50.5% (\(SD = 21.0\), range = 2.0-95.0), and mean Chinese-family exposure was 48.6% (\(SD = 21.1\), range = 5.0-97.0). 30 participants were English-dominant, 27 were Chinese-family-dominant, and 16 were approximately balanced (within 10 percentage points).
Reaction-time analyses are based on correct responses only. Very slow responses are unlikely to reflect the intended speeded lexical decision process. The analysis therefore applied the language-background and participant-accuracy exclusions, kept correct responses, and then applied a 3000 ms RT cutoff. The RT cutoff removed 284 correct trials, corresponding to 10.2% of correct trials and 9.5% of all trials in the analysed participant sample.
Data were analysed using Bayesian mixed-effects models (Gelman et al., 2014; McElreath, 2016; Nicenboim & Vasishth, 2016). Reaction times were modelled with a shifted-lognormal distribution (Rouder, 2005), which is appropriate for positively skewed response-time data because it estimates a non-decision-time shift in addition to the lognormal location and scale parameters. This approach follows recommendations to model response-time distributions directly rather than relying only on transformed Gaussian models (Lo & Andrews, 2015). Accuracy was analysed with a Bernoulli model with a logit link. Both primary models included fixed effects of lexicality, script, and their interaction, with correlated random intercepts and slopes for the lexicality-by-script structure by participants and by items. For the shifted-lognormal RT models, the lognormal location parameter included participant and item random intercepts and slopes, and the non-decision-time parameter included a participant random intercept.
Predictors were coded with scaled sum contrasts (-0.5, +0.5), following recommendations for contrast coding in mixed models (Brehm & Alday, 2022; Schad et al., 2020). With this coding, two-level main effects are estimated on the full pairwise-difference scale and interaction terms are estimated on the corresponding difference-in-differences scale. This keeps main effects and interactions on comparable scales for ROPE evaluation. Models were fitted with brms (Bürkner, 2017), which uses Stan and Hamiltonian Monte Carlo for Bayesian estimation (Carpenter et al., 2017; Hoffman & Gelman, 2014).
Weakly informative priors were used to regularise estimates without strongly constraining the direction of the effects (McElreath, 2016). For the RT model, fixed effects were assigned \(\mathcal{N}(0, 0.30)\) priors on the shifted-lognormal location scale, with a \(\mathcal{N}(log(1000), 1)\) prior on the intercept and a \(\mathcal{N}(log(250), 0.35)\) prior on the non-decision-time intercept. For the accuracy model, fixed effects were assigned \(\mathcal{N}(0, 1)\) priors on the log-odds scale, with a \(\mathcal{N}(2.5, 1.5)\) prior on the intercept. Random-effect standard deviations were assigned exponential priors, and random-effect correlation matrices were assigned LKJ(6) priors. Model convergence was assessed using \(\hat{R}\), effective sample sizes, divergent transitions, and visual inspection of trace plots.
Posterior summaries are reported using posterior medians, 95% probability intervals (PIs), the posterior probability that the effect is negative [P(neg.)], and the posterior probability that the effect falls inside a region of practical equivalence (ROPE). PIs directly summarise uncertainty in the posterior distribution of the parameter (Sorensen et al., 2016). ROPE values quantify the posterior probability that an effect is practically equivalent to zero (Kruschke, 2018; Kruschke & Liddell, 2018). The ROPE was set to \(\pm .05\) log units for RT coefficients and \(\pm .10\) log-odds for accuracy coefficients. For pairwise comparisons in the tables, the ROPE was set to \(\pm 20\) ms for RT differences, following similar applications to psycholinguistic response-time data (Vasishth et al., 2018), and \(\pm 2\) percentage points for accuracy differences.
Reaction times were analysed using a Bayesian shifted-lognormal mixed-effects model. The model included fixed effects of lexicality, script, and their interaction, with correlated random intercepts and slopes for lexicality, script, and their interaction by participants and by items. The non-decision-time parameter included a participant random intercept. For the RT model, the posterior summary for lexicality was \(\hat\beta\) = -0.27, 95% PI [-0.38, -0.15], P(neg.) > .99, ROPE < .01; for script, \(\hat\beta\) = 0.12, 95% PI [-0.03, 0.26], P(neg.) = 0.05, ROPE = 0.16; and for the lexicality by script interaction, \(\hat\beta\) = 0.21, 95% PI [0.04, 0.38], P(neg.) < .01, ROPE = 0.03.
Posterior means and pairwise script comparisons are shown in Table 1, and the RT estimates are plotted in Figure 1. Pairwise comparisons within lexicality indicated that responses to real words were slower for Pinyin than for Chinese. For nonwords, the estimated script difference was small and uncertain.
| Outcome | Lexicality | Pinyin | Chinese | Difference | Ratio / OR | P(neg.) | ROPE |
|---|---|---|---|---|---|---|---|
| Reaction time | nonword | 1110 ms [1025, 1208] | 1103 ms [1010, 1209] | 8 ms [-85, 100] | 1.01 [0.93, 1.09] | 0.43 | 0.33 |
| Reaction time | word | 1019 ms [942, 1106] | 916 ms [850, 991] | 102 ms [21, 185] | 1.11 [1.02, 1.21] | < .01 | 0.02 |
| Accuracy | nonword | 96.8% [94.4, 98.3] | 97.3% [95.2, 98.7] | -0.5 pp [-2.9, 1.8] | 0.83 [0.36, 1.78] | 0.68 | 0.88 |
| Accuracy | word | 97.2% [95.1, 98.6] | 98.0% [96.4, 99.0] | -0.8 pp [-2.9, 1.0] | 0.71 [0.31, 1.58] | 0.80 | 0.88 |
Accuracy was analysed using a Bayesian Bernoulli mixed-effects model with fixed effects of lexicality, script, and their interaction, with correlated random intercepts and slopes for lexicality, script, and their interaction by participants and by items. For the accuracy model, the posterior summary for lexicality was \(\hat\beta\) = 0.23, 95% PI [-0.39, 0.81], P(neg.) = 0.23, ROPE = 0.20; for script, \(\hat\beta\) = -0.27, 95% PI [-0.89, 0.35], P(neg.) = 0.81, ROPE = 0.17; and for the lexicality by script interaction, \(\hat\beta\) = -0.15, 95% PI [-1.16, 0.88], P(neg.) = 0.61, ROPE = 0.15. Pairwise comparisons revealed no evidence for script effects. Posterior means and evidence for script effects can be found in Table 1.
Figure 1: Estimated marginal mean reaction times by lexicality and written form. Error bars show 95% probability intervals.
As a posthoc, hypothesis-generating analysis, the RT model was extended by adding exposure-dominance group and all interactions with lexicality and script. This analysis included only the 55 participants classified as English-dominant or Chinese-family-dominant (English-dominant: n = 28; Chinese-family-dominant: n = 27). The 16 approximately balanced participants and 2 participants whose exposure responses could not be coded were not included in this exploratory subgroup analysis. The exploratory RT model included correlated participant random intercepts and slopes for lexicality, script, and their interaction, a participant random intercept for non-decision time, and correlated item random intercepts and slopes for the lexicality-by-script-by-exposure-dominance structure. The subgroup estimated marginal means and pairwise script comparisons are shown in Table 2. In addition to the three-way interaction (\(\hat\beta\) = -0.07, 95% PI [-0.33, 0.17], P(neg.) = 0.72, ROPE = 0.26), the exploratory model estimated the exposure-group main effect (\(\hat\beta\) = 0.10, 95% PI [-0.13, 0.32], P(neg.) = 0.20, ROPE = 0.24), the lexicality by exposure-group interaction (\(\hat\beta\) = -0.22, 95% PI [-0.39, -0.05], P(neg.) > .99, ROPE = 0.03), and the script by exposure-group interaction (\(\hat\beta\) = -0.19, 95% PI [-0.38, 0.00], P(neg.) = 0.97, ROPE = 0.07). These subgroup estimates and pairwise comparisons should be treated as descriptive and hypothesis-generating.
| Exposure group | Lexicality | Pinyin | Chinese | Difference | Ratio | P(neg.) | ROPE |
|---|---|---|---|---|---|---|---|
| Chinese-family-dominant | nonword | 1061 ms [951, 1192] | 1002 ms [896, 1136] | 59 ms [-48, 165] | 1.06 [0.96, 1.17] | 0.13 | 0.17 |
| Chinese-family-dominant | word | 1040 ms [934, 1169] | 893 ms [809, 998] | 148 ms [46, 259] | 1.17 [1.05, 1.30] | < .01 | < .01 |
| English-dominant | nonword | 1140 ms [1015, 1288] | 1166 ms [1030, 1338] | -25 ms [-170, 104] | 0.98 [0.87, 1.09] | 0.65 | 0.22 |
| English-dominant | word | 976 ms [876, 1089] | 935 ms [846, 1044] | 41 ms [-61, 137] | 1.04 [0.94, 1.15] | 0.20 | 0.22 |
For the corresponding exploratory, hypothesis-generating error analysis, accuracy was modelled with the same two exposure-dominance groups and the same fixed-effect structure as the exploratory RT model. The model included correlated participant random intercepts and slopes for lexicality, script, and their interaction, and correlated item random intercepts and slopes for the lexicality-by-script-by-exposure-dominance structure. Descriptively, English-dominant participants made more errors overall than Chinese-family-dominant participants, especially on Chinese-script trials. In the model, the exposure-group main effect was \(\hat\beta\) = -0.81, 95% PI [-1.56, -0.08], P(neg.) = 0.98, ROPE = 0.02, the lexicality by exposure-group interaction was \(\hat\beta\) = 0.37, 95% PI [-0.56, 1.29], P(neg.) = 0.22, ROPE = 0.13, the script by exposure-group interaction was \(\hat\beta\) = 0.86, 95% PI [-0.12, 1.83], P(neg.) = 0.04, ROPE = 0.04, and the three-way interaction was \(\hat\beta\) = -0.67, 95% PI [-2.12, 0.76], P(neg.) = 0.82, ROPE = 0.07. Observed accuracies, estimated accuracies, and pairwise script comparisons by exposure group are shown in Table 3.
| Exposure group | Lexicality | Observed Pinyin | Observed Chinese | Pinyin | Chinese | Difference | Odds ratio | P(neg.) | ROPE |
|---|---|---|---|---|---|---|---|---|---|
| Chinese-family-dominant | nonword | 92.2% | 96.3% | 96.9% [93.5, 98.8] | 98.9% [97.1, 99.6] | -1.9 pp [-5.1, 0.1] | 0.35 [0.11, 1.05] | 0.97 | 0.53 |
| Chinese-family-dominant | word | 94.4% | 96.3% | 97.5% [94.5, 99.0] | 98.6% [96.5, 99.5] | -1.0 pp [-3.9, 0.9] | 0.56 [0.19, 1.62] | 0.85 | 0.79 |
| English-dominant | nonword | 90.3% | 86.7% | 95.5% [91.1, 98.1] | 94.9% [89.6, 97.8] | 0.6 pp [-4.0, 5.7] | 1.15 [0.43, 3.01] | 0.40 | 0.62 |
| English-dominant | word | 90.7% | 92.0% | 96.5% [92.8, 98.5] | 96.7% [93.1, 98.6] | -0.2 pp [-3.9, 3.3] | 0.94 [0.36, 2.50] | 0.55 | 0.77 |
The following exploratory appendix analyses only the participants who did not report knowing any Chinese-family language in the LEAP-Q/background screening. No participant-level accuracy exclusion was applied for this subset.
The no-Chinese subset contained 7 participants and 280 lexical-decision response trials. The accuracy analysis used all response trials from these participants. The RT model used the 134 correct responses at or below 3000 ms, matching the upper RT cutoff used in the main analysis.
Accuracy was modelled with a Bayesian Bernoulli mixed-effects model with fixed effects of lexicality, script, and their interaction, with correlated random intercepts and slopes for lexicality, script, and their interaction by participants and by items. This matches the fixed-effect and random-effect structure used in the main accuracy analysis. No participant-level accuracy exclusion was applied.
| Lexicality | Pinyin | Chinese | Difference | Odds ratio | Odds-ratio 95% PI | P(Chinese > Pinyin) | ROPE |
|---|---|---|---|---|---|---|---|
| nonword | 55.3% [33.4, 75.8] | 69.0% [47.7, 85.1] | 13.2 pp [-11.0, 37.8] | 1.80 | [0.62, 5.50] | 0.86 | 0.07 |
| word | 53.9% [32.3, 75.0] | 66.5% [45.7, 82.3] | 12.3 pp [-13.9, 37.1] | 1.71 | [0.54, 5.11] | 0.83 | 0.08 |
Reaction times were modelled with a Bayesian shifted-lognormal mixed-effects model for correct responses at or below 3000 ms, without excluding participants for poor accuracy. This matches the fixed-effect, random-effect, and distributional structure used in the main RT analysis, including correlated random intercepts and slopes for the lognormal location parameter and a participant random intercept for non-decision time. The model retained 2 post-warmup divergent transitions after tightening the sampler controls, most likely reflecting the small and noisy appendix subset under the intentionally matched random-effects structure rather than a problem that would be solved by further increasing the target acceptance rate or tree depth. The table reports posterior-predicted RTs in milliseconds and response-scale pairwise differences.
| Lexicality | Pinyin | Chinese | Difference | P(Chinese > Pinyin) | ROPE |
|---|---|---|---|---|---|
| nonword | 1317 ms [1013, 1745] | 1330 ms [1033, 1769] | 12 ms [-316, 333] | 0.53 | 0.11 |
| word | 1180 ms [928, 1564] | 1217 ms [957, 1608] | 34 ms [-238, 323] | 0.60 | 0.12 |
Table 6 reports \(\hat{R}\) values for the main reported parameters from each fitted model, together with the number of post-warmup divergent transitions. Trace plots are then shown for the same fixed-effect parameters. The chains showed adequate mixing for the parameters reported in the main text. The only remaining divergent transitions were in the no-Chinese appendix RT model.
| Model | Parameter | R-hat | Divergences |
|---|---|---|---|
| Reaction time | Intercept | 1.00 | 0 |
| Reaction time | Lexicality | 1.00 | 0 |
| Reaction time | Script | 1.00 | 0 |
| Reaction time | Lexicality x script | 1.00 | 0 |
| Reaction time | Non-decision time intercept | 1.00 | 0 |
| Accuracy | Intercept | 1.00 | 0 |
| Accuracy | Lexicality | 1.00 | 0 |
| Accuracy | Script | 1.00 | 0 |
| Accuracy | Lexicality x script | 1.00 | 0 |
| Exploratory reaction time | Intercept | 1.00 | 0 |
| Exploratory reaction time | Exposure group | 1.00 | 0 |
| Exploratory reaction time | Lexicality x exposure group | 1.00 | 0 |
| Exploratory reaction time | Script x exposure group | 1.00 | 0 |
| Exploratory reaction time | Lexicality x script x exposure group | 1.00 | 0 |
| Exploratory accuracy | Intercept | 1.00 | 0 |
| Exploratory accuracy | Exposure group | 1.00 | 0 |
| Exploratory accuracy | Lexicality x exposure group | 1.00 | 0 |
| Exploratory accuracy | Script x exposure group | 1.00 | 0 |
| Exploratory accuracy | Lexicality x script x exposure group | 1.00 | 0 |
| No-Chinese reaction time | Intercept | 1.00 | 2 |
| No-Chinese reaction time | Lexicality | 1.00 | 2 |
| No-Chinese reaction time | Script | 1.00 | 2 |
| No-Chinese reaction time | Lexicality x script | 1.00 | 2 |
| No-Chinese reaction time | Non-decision time intercept | 1.00 | 2 |
| No-Chinese accuracy | Intercept | 1.00 | 0 |
| No-Chinese accuracy | Lexicality | 1.00 | 0 |
| No-Chinese accuracy | Script | 1.00 | 0 |
| No-Chinese accuracy | Lexicality x script | 1.00 | 0 |
Figure 2: Trace plots for the main reaction-time model fixed-effect parameters.
Figure 3: Trace plots for the accuracy model fixed-effect parameters.
Figure 4: Trace plots for the exploratory reaction-time model fixed-effect parameters involving language-exposure group.
Figure 5: Trace plots for the exploratory accuracy model fixed-effect parameters involving language-exposure group.
Posterior predictive checks compare the observed data with draws from the posterior predictive distribution. For reaction times, the checks compare the observed RT distribution with replicated RT distributions generated by the fitted shifted-lognormal models. For accuracy, the check compares the observed distribution of binary responses with replicated binary-response data generated by the Bernoulli model.
Figure 6: Posterior predictive check for the main reaction-time model. The dark line shows the observed data; lighter lines show replicated data sets generated from the posterior predictive distribution.
Figure 7: Posterior predictive check for the exploratory accuracy model. Bars compare the observed binary-response distribution with replicated binary-response data generated from the posterior predictive distribution.
Figure 8: Posterior predictive check for the accuracy model. Bars compare the observed binary-response distribution with replicated data sets generated from the posterior predictive distribution.
Figure 9: Posterior predictive check for the exploratory reaction-time model. The dark line shows the observed data; lighter lines show replicated data sets generated from the posterior predictive distribution.