Brief Summary

The main result is a script-by-lexicality pattern in reaction times among fluent Chinese readers. Real words were identified faster in Chinese characters than in Pinyin, whereas the corresponding script difference for nonwords was small and uncertain. Accuracy was high overall and showed little evidence for a corresponding main script effect in the primary analysis. The exploratory language-dominance analysis is hypothesis-generating rather than confirmatory, but it raises the possibility that lower Chinese-family exposure is associated with more errors on Chinese-script trials.

Sample

Eligibility required participants to be fluent Chinese readers. Chinese-reader eligibility was checked from LEAP-Q language-background responses and, where available, the experiment’s Chinese-reader gate. The final analysed sample included 75 Chinese-English bilinguals after excluding 9 participants with accuracy below 70%. Participants were aged 18-42 years (\(M = 24.55\), \(SD = 4.73\)). First-language responses were predominantly Mandarin/Chinese (59 participants), followed by English (10), Cantonese (3), and Fuzhou dialect (1). For second language, most participants reported English (59), with a smaller number reporting Mandarin/Chinese (11). 33 participants reported an additional language or dialect beyond their main Chinese-English background.

Language exposure was coded from free-text LEAP-Q responses. Mandarin, Chinese, Cantonese, Fuzhou dialect, Hokkien, and other Chinese-family varieties were grouped as Chinese-family exposure. 5 responses whose percentages summed to more than 100 were rescaled proportionally. Exposure could be coded for 73 participants. Mean English exposure was 50.5% (\(SD = 21.0\), range = 2.0-95.0), and mean Chinese-family exposure was 48.6% (\(SD = 21.1\), range = 5.0-97.0). 30 participants were English-dominant, 27 were Chinese-family-dominant, and 16 were approximately balanced (within 10 percentage points).

Data Analysis

Reaction-time analyses are based on correct responses only. Very slow responses are unlikely to reflect the intended speeded lexical decision process. The analysis therefore applied the language-background and participant-accuracy exclusions, kept correct responses, and then applied a 3000 ms RT cutoff. The RT cutoff removed 284 correct trials, corresponding to 10.2% of correct trials and 9.5% of all trials in the analysed participant sample.

Data were analysed using Bayesian mixed-effects models (Gelman et al., 2014; McElreath, 2016; Nicenboim & Vasishth, 2016). Reaction times were modelled with a shifted-lognormal distribution (Rouder, 2005), which is appropriate for positively skewed response-time data because it estimates a non-decision-time shift in addition to the lognormal location and scale parameters. This approach follows recommendations to model response-time distributions directly rather than relying only on transformed Gaussian models (Lo & Andrews, 2015). Accuracy was analysed with a Bernoulli model with a logit link. Both primary models included fixed effects of lexicality, script, and their interaction, with correlated random intercepts and slopes for the lexicality-by-script structure by participants and by items. For the shifted-lognormal RT models, the lognormal location parameter included participant and item random intercepts and slopes, and the non-decision-time parameter included a participant random intercept.

Predictors were coded with scaled sum contrasts (-0.5, +0.5), following recommendations for contrast coding in mixed models (Brehm & Alday, 2022; Schad et al., 2020). With this coding, two-level main effects are estimated on the full pairwise-difference scale and interaction terms are estimated on the corresponding difference-in-differences scale. This keeps main effects and interactions on comparable scales for ROPE evaluation. Models were fitted with brms (Bürkner, 2017), which uses Stan and Hamiltonian Monte Carlo for Bayesian estimation (Carpenter et al., 2017; Hoffman & Gelman, 2014).

Weakly informative priors were used to regularise estimates without strongly constraining the direction of the effects (McElreath, 2016). For the RT model, fixed effects were assigned \(\mathcal{N}(0, 0.30)\) priors on the shifted-lognormal location scale, with a \(\mathcal{N}(log(1000), 1)\) prior on the intercept and a \(\mathcal{N}(log(250), 0.35)\) prior on the non-decision-time intercept. For the accuracy model, fixed effects were assigned \(\mathcal{N}(0, 1)\) priors on the log-odds scale, with a \(\mathcal{N}(2.5, 1.5)\) prior on the intercept. Random-effect standard deviations were assigned exponential priors, and random-effect correlation matrices were assigned LKJ(6) priors. Model convergence was assessed using \(\hat{R}\), effective sample sizes, divergent transitions, and visual inspection of trace plots.

Posterior summaries are reported using posterior medians, 95% probability intervals (PIs), the posterior probability that the effect is negative [P(neg.)], and the posterior probability that the effect falls inside a region of practical equivalence (ROPE). PIs directly summarise uncertainty in the posterior distribution of the parameter (Sorensen et al., 2016). ROPE values quantify the posterior probability that an effect is practically equivalent to zero (Kruschke, 2018; Kruschke & Liddell, 2018). The ROPE was set to \(\pm .05\) log units for RT coefficients and \(\pm .10\) log-odds for accuracy coefficients. For pairwise comparisons in the tables, the ROPE was set to \(\pm 20\) ms for RT differences, following similar applications to psycholinguistic response-time data (Vasishth et al., 2018), and \(\pm 2\) percentage points for accuracy differences.

Results

Reaction times were analysed using a Bayesian shifted-lognormal mixed-effects model. The model included fixed effects of lexicality, script, and their interaction, with correlated random intercepts and slopes for lexicality, script, and their interaction by participants and by items. The non-decision-time parameter included a participant random intercept. For the RT model, the posterior summary for lexicality was \(\hat\beta\) = -0.27, 95% PI [-0.38, -0.15], P(neg.) > .99, ROPE < .01; for script, \(\hat\beta\) = 0.12, 95% PI [-0.03, 0.26], P(neg.) = 0.05, ROPE = 0.16; and for the lexicality by script interaction, \(\hat\beta\) = 0.21, 95% PI [0.04, 0.38], P(neg.) < .01, ROPE = 0.03.

Posterior means and pairwise script comparisons are shown in Table 1, and the RT estimates are plotted in Figure 1. Pairwise comparisons within lexicality indicated that responses to real words were slower for Pinyin than for Chinese. For nonwords, the estimated script difference was small and uncertain.

Table 1: Posterior estimated marginal means and pairwise script comparisons for reaction time and accuracy. Reaction-time differences are in milliseconds; accuracy differences are in percentage points. Intervals are 95% posterior intervals.
Outcome Lexicality Pinyin Chinese Difference Ratio / OR P(neg.) ROPE
Reaction time nonword 1110 ms [1025, 1208] 1103 ms [1010, 1209] 8 ms [-85, 100] 1.01 [0.93, 1.09] 0.43 0.33
Reaction time word 1019 ms [942, 1106] 916 ms [850, 991] 102 ms [21, 185] 1.11 [1.02, 1.21] < .01 0.02
Accuracy nonword 96.8% [94.4, 98.3] 97.3% [95.2, 98.7] -0.5 pp [-2.9, 1.8] 0.83 [0.36, 1.78] 0.68 0.88
Accuracy word 97.2% [95.1, 98.6] 98.0% [96.4, 99.0] -0.8 pp [-2.9, 1.0] 0.71 [0.31, 1.58] 0.80 0.88

Accuracy was analysed using a Bayesian Bernoulli mixed-effects model with fixed effects of lexicality, script, and their interaction, with correlated random intercepts and slopes for lexicality, script, and their interaction by participants and by items. For the accuracy model, the posterior summary for lexicality was \(\hat\beta\) = 0.23, 95% PI [-0.39, 0.81], P(neg.) = 0.23, ROPE = 0.20; for script, \(\hat\beta\) = -0.27, 95% PI [-0.89, 0.35], P(neg.) = 0.81, ROPE = 0.17; and for the lexicality by script interaction, \(\hat\beta\) = -0.15, 95% PI [-1.16, 0.88], P(neg.) = 0.61, ROPE = 0.15. Pairwise comparisons revealed no evidence for script effects. Posterior means and evidence for script effects can be found in Table 1.

Estimated marginal mean reaction times by lexicality and written form. Error bars show 95% probability intervals.

Figure 1: Estimated marginal mean reaction times by lexicality and written form. Error bars show 95% probability intervals.

Exploratory Analyses by Language Exposure

As a posthoc, hypothesis-generating analysis, the RT model was extended by adding exposure-dominance group and all interactions with lexicality and script. This analysis included only the 55 participants classified as English-dominant or Chinese-family-dominant (English-dominant: n = 28; Chinese-family-dominant: n = 27). The 16 approximately balanced participants and 2 participants whose exposure responses could not be coded were not included in this exploratory subgroup analysis. The exploratory RT model included correlated participant random intercepts and slopes for lexicality, script, and their interaction, a participant random intercept for non-decision time, and correlated item random intercepts and slopes for the lexicality-by-script-by-exposure-dominance structure. The subgroup estimated marginal means and pairwise script comparisons are shown in Table 2. In addition to the three-way interaction (\(\hat\beta\) = -0.07, 95% PI [-0.33, 0.17], P(neg.) = 0.72, ROPE = 0.26), the exploratory model estimated the exposure-group main effect (\(\hat\beta\) = 0.10, 95% PI [-0.13, 0.32], P(neg.) = 0.20, ROPE = 0.24), the lexicality by exposure-group interaction (\(\hat\beta\) = -0.22, 95% PI [-0.39, -0.05], P(neg.) > .99, ROPE = 0.03), and the script by exposure-group interaction (\(\hat\beta\) = -0.19, 95% PI [-0.38, 0.00], P(neg.) = 0.97, ROPE = 0.07). These subgroup estimates and pairwise comparisons should be treated as descriptive and hypothesis-generating.

Table 2: Posterior estimated marginal mean reaction times and pairwise script comparisons by language-exposure group. Differences are in milliseconds and intervals are 95% posterior intervals.
Exposure group Lexicality Pinyin Chinese Difference Ratio P(neg.) ROPE
Chinese-family-dominant nonword 1061 ms [951, 1192] 1002 ms [896, 1136] 59 ms [-48, 165] 1.06 [0.96, 1.17] 0.13 0.17
Chinese-family-dominant word 1040 ms [934, 1169] 893 ms [809, 998] 148 ms [46, 259] 1.17 [1.05, 1.30] < .01 < .01
English-dominant nonword 1140 ms [1015, 1288] 1166 ms [1030, 1338] -25 ms [-170, 104] 0.98 [0.87, 1.09] 0.65 0.22
English-dominant word 976 ms [876, 1089] 935 ms [846, 1044] 41 ms [-61, 137] 1.04 [0.94, 1.15] 0.20 0.22

For the corresponding exploratory, hypothesis-generating error analysis, accuracy was modelled with the same two exposure-dominance groups and the same fixed-effect structure as the exploratory RT model. The model included correlated participant random intercepts and slopes for lexicality, script, and their interaction, and correlated item random intercepts and slopes for the lexicality-by-script-by-exposure-dominance structure. Descriptively, English-dominant participants made more errors overall than Chinese-family-dominant participants, especially on Chinese-script trials. In the model, the exposure-group main effect was \(\hat\beta\) = -0.81, 95% PI [-1.56, -0.08], P(neg.) = 0.98, ROPE = 0.02, the lexicality by exposure-group interaction was \(\hat\beta\) = 0.37, 95% PI [-0.56, 1.29], P(neg.) = 0.22, ROPE = 0.13, the script by exposure-group interaction was \(\hat\beta\) = 0.86, 95% PI [-0.12, 1.83], P(neg.) = 0.04, ROPE = 0.04, and the three-way interaction was \(\hat\beta\) = -0.67, 95% PI [-2.12, 0.76], P(neg.) = 0.82, ROPE = 0.07. Observed accuracies, estimated accuracies, and pairwise script comparisons by exposure group are shown in Table 3.

Table 3: Observed and posterior estimated accuracy by language-exposure group. Model-based differences are Pinyin minus Chinese in percentage points; intervals are 95% posterior intervals.
Exposure group Lexicality Observed Pinyin Observed Chinese Pinyin Chinese Difference Odds ratio P(neg.) ROPE
Chinese-family-dominant nonword 92.2% 96.3% 96.9% [93.5, 98.8] 98.9% [97.1, 99.6] -1.9 pp [-5.1, 0.1] 0.35 [0.11, 1.05] 0.97 0.53
Chinese-family-dominant word 94.4% 96.3% 97.5% [94.5, 99.0] 98.6% [96.5, 99.5] -1.0 pp [-3.9, 0.9] 0.56 [0.19, 1.62] 0.85 0.79
English-dominant nonword 90.3% 86.7% 95.5% [91.1, 98.1] 94.9% [89.6, 97.8] 0.6 pp [-4.0, 5.7] 1.15 [0.43, 3.01] 0.40 0.62
English-dominant word 90.7% 92.0% 96.5% [92.8, 98.5] 96.7% [93.1, 98.6] -0.2 pp [-3.9, 3.3] 0.94 [0.36, 2.50] 0.55 0.77

References

Brehm, L., & Alday, P. M. (2022). Contrast coding choices in a decade of mixed models. Journal of Memory and Language, 125, 104334. https://doi.org/10.1016/j.jml.2022.104334
Bürkner, P.-C. (2017). Brms: An r package for bayesian multilevel models using stan. Journal of Statistical Software, 80(1), 1–28. https://doi.org/10.18637/jss.v080.i01
Carpenter, B., Gelman, A., Hoffman, M. D., Lee, D., Goodrich, B., Betancourt, M., Brubaker, M., Guo, J., Li, P., & Riddell, A. (2017). Stan: A probabilistic programming language. Journal of Statistical Software, 76(1), 1–32. https://doi.org/10.18637/jss.v076.i01
Gelman, A., Carlin, J. B., Stern, H. S., Dunson, D. B., Vehtari, A., & Rubin, D. B. (2014). Bayesian data analysis (3rd ed.). Chapman; Hall/CRC.
Hoffman, M. D., & Gelman, A. (2014). The no-u-turn sampler: Adaptively setting path lengths in hamiltonian monte carlo. Journal of Machine Learning Research, 15(1), 1593–1623.
Kruschke, J. K. (2018). Rejecting or accepting parameter values in bayesian estimation. Advances in Methods and Practices in Psychological Science, 1(2), 270–280. https://doi.org/10.1177/2515245918771304
Kruschke, J. K., & Liddell, T. M. (2018). The bayesian new statistics: Hypothesis testing, estimation, meta-analysis, and power analysis from a bayesian perspective. Psychonomic Bulletin & Review, 25(1), 178–206. https://doi.org/10.3758/s13423-016-1221-4
Lo, S., & Andrews, S. (2015). To transform or not to transform: Using generalized linear mixed models to analyse reaction time data. Frontiers in Psychology, 6, 1171. https://doi.org/10.3389/fpsyg.2015.01171
McElreath, R. (2016). Statistical rethinking: A bayesian course with examples in r and stan. CRC Press.
Nicenboim, B., & Vasishth, S. (2016). Statistical methods for linguistic research: Foundational ideas - part II. Language and Linguistics Compass, 10(11), 591–613. https://doi.org/10.1111/lnc3.12207
Rouder, J. N. (2005). Are unshifted distributional models appropriate for response time? Psychometrika, 70(2), 377–381. https://doi.org/10.1007/s11336-005-1297-7
Schad, D. J., Vasishth, S., Hohenstein, S., & Kliegl, R. (2020). How to capitalize on a priori contrasts in linear (mixed) models: A tutorial. Journal of Memory and Language, 110, 104038. https://doi.org/10.1016/j.jml.2019.104038
Sorensen, T., Hohenstein, S., & Vasishth, S. (2016). Bayesian linear mixed models using stan: A tutorial for psychologists, linguists, and cognitive scientists. The Quantitative Methods for Psychology, 12(3), 175–200. https://doi.org/10.20982/tqmp.12.3.p175
Vasishth, S., Mertzen, D., Jäger, L. A., & Gelman, A. (2018). The statistical significance filter leads to overoptimistic expectations of replicability. Journal of Memory and Language, 103, 151–175. https://doi.org/10.1016/j.jml.2018.07.004

Appendix

No-Chinese Participant Subset

The following exploratory appendix analyses only the participants who did not report knowing any Chinese-family language in the LEAP-Q/background screening. No participant-level accuracy exclusion was applied for this subset.

The no-Chinese subset contained 7 participants and 280 lexical-decision response trials. The accuracy analysis used all response trials from these participants. The RT model used the 134 correct responses at or below 3000 ms, matching the upper RT cutoff used in the main analysis.

Accuracy was modelled with a Bayesian Bernoulli mixed-effects model with fixed effects of lexicality, script, and their interaction, with correlated random intercepts and slopes for lexicality, script, and their interaction by participants and by items. This matches the fixed-effect and random-effect structure used in the main accuracy analysis. No participant-level accuracy exclusion was applied.

Table 4: Model-estimated accuracy and Bayesian script contrasts for participants who did not report knowing Chinese. Differences compare Chinese-script trials with Pinyin trials in percentage points.
Lexicality Pinyin Chinese Difference Odds ratio Odds-ratio 95% PI P(Chinese > Pinyin) ROPE
nonword 55.3% [33.4, 75.8] 69.0% [47.7, 85.1] 13.2 pp [-11.0, 37.8] 1.80 [0.62, 5.50] 0.86 0.07
word 53.9% [32.3, 75.0] 66.5% [45.7, 82.3] 12.3 pp [-13.9, 37.1] 1.71 [0.54, 5.11] 0.83 0.08

Reaction times were modelled with a Bayesian shifted-lognormal mixed-effects model for correct responses at or below 3000 ms, without excluding participants for poor accuracy. This matches the fixed-effect, random-effect, and distributional structure used in the main RT analysis, including correlated random intercepts and slopes for the lognormal location parameter and a participant random intercept for non-decision time. The model retained 2 post-warmup divergent transitions after tightening the sampler controls, most likely reflecting the small and noisy appendix subset under the intentionally matched random-effects structure rather than a problem that would be solved by further increasing the target acceptance rate or tree depth. The table reports posterior-predicted RTs in milliseconds and response-scale pairwise differences.

Table 5: Model-estimated reaction times and Bayesian script contrasts for participants who did not report knowing Chinese. Differences compare Chinese-script trials with Pinyin trials in milliseconds.
Lexicality Pinyin Chinese Difference P(Chinese > Pinyin) ROPE
nonword 1317 ms [1013, 1745] 1330 ms [1033, 1769] 12 ms [-316, 333] 0.53 0.11
word 1180 ms [928, 1564] 1217 ms [957, 1608] 34 ms [-238, 323] 0.60 0.12

Posterior Model Checks

Table 6 reports \(\hat{R}\) values for the main reported parameters from each fitted model, together with the number of post-warmup divergent transitions. Trace plots are then shown for the same fixed-effect parameters. The chains showed adequate mixing for the parameters reported in the main text. The only remaining divergent transitions were in the no-Chinese appendix RT model.

Table 6: R-hat values for the main reported parameters from each Bayesian model. Divergences are post-warmup divergent transitions for the full fitted model.
Model Parameter R-hat Divergences
Reaction time Intercept 1.00 0
Reaction time Lexicality 1.00 0
Reaction time Script 1.00 0
Reaction time Lexicality x script 1.00 0
Reaction time Non-decision time intercept 1.00 0
Accuracy Intercept 1.00 0
Accuracy Lexicality 1.00 0
Accuracy Script 1.00 0
Accuracy Lexicality x script 1.00 0
Exploratory reaction time Intercept 1.00 0
Exploratory reaction time Exposure group 1.00 0
Exploratory reaction time Lexicality x exposure group 1.00 0
Exploratory reaction time Script x exposure group 1.00 0
Exploratory reaction time Lexicality x script x exposure group 1.00 0
Exploratory accuracy Intercept 1.00 0
Exploratory accuracy Exposure group 1.00 0
Exploratory accuracy Lexicality x exposure group 1.00 0
Exploratory accuracy Script x exposure group 1.00 0
Exploratory accuracy Lexicality x script x exposure group 1.00 0
No-Chinese reaction time Intercept 1.00 2
No-Chinese reaction time Lexicality 1.00 2
No-Chinese reaction time Script 1.00 2
No-Chinese reaction time Lexicality x script 1.00 2
No-Chinese reaction time Non-decision time intercept 1.00 2
No-Chinese accuracy Intercept 1.00 0
No-Chinese accuracy Lexicality 1.00 0
No-Chinese accuracy Script 1.00 0
No-Chinese accuracy Lexicality x script 1.00 0
Trace plots for the main reaction-time model fixed-effect parameters.

Figure 2: Trace plots for the main reaction-time model fixed-effect parameters.

Trace plots for the accuracy model fixed-effect parameters.

Figure 3: Trace plots for the accuracy model fixed-effect parameters.

Trace plots for the exploratory reaction-time model fixed-effect parameters involving language-exposure group.

Figure 4: Trace plots for the exploratory reaction-time model fixed-effect parameters involving language-exposure group.

Trace plots for the exploratory accuracy model fixed-effect parameters involving language-exposure group.

Figure 5: Trace plots for the exploratory accuracy model fixed-effect parameters involving language-exposure group.

Posterior predictive checks compare the observed data with draws from the posterior predictive distribution. For reaction times, the checks compare the observed RT distribution with replicated RT distributions generated by the fitted shifted-lognormal models. For accuracy, the check compares the observed distribution of binary responses with replicated binary-response data generated by the Bernoulli model.

Posterior predictive check for the main reaction-time model. The dark line shows the observed data; lighter lines show replicated data sets generated from the posterior predictive distribution.

Figure 6: Posterior predictive check for the main reaction-time model. The dark line shows the observed data; lighter lines show replicated data sets generated from the posterior predictive distribution.

Posterior predictive check for the exploratory accuracy model. Bars compare the observed binary-response distribution with replicated binary-response data generated from the posterior predictive distribution.

Figure 7: Posterior predictive check for the exploratory accuracy model. Bars compare the observed binary-response distribution with replicated binary-response data generated from the posterior predictive distribution.

Posterior predictive check for the accuracy model. Bars compare the observed binary-response distribution with replicated data sets generated from the posterior predictive distribution.

Figure 8: Posterior predictive check for the accuracy model. Bars compare the observed binary-response distribution with replicated data sets generated from the posterior predictive distribution.

Posterior predictive check for the exploratory reaction-time model. The dark line shows the observed data; lighter lines show replicated data sets generated from the posterior predictive distribution.

Figure 9: Posterior predictive check for the exploratory reaction-time model. The dark line shows the observed data; lighter lines show replicated data sets generated from the posterior predictive distribution.