This experiment compares the tossing accuracy of three participants. Each participant made ten recorded tosses toward a target. The response is distance in centimeters from the tossed object to the target; smaller distances indicate greater accuracy. Person is the single factor, with three levels. The analysis assumes that the tossing and measurement conditions were kept consistent. Independence depends on how the tosses were carried out; it cannot be established from a statistical test alone.
In this recorded dataset, mean distances were 9.96, 13.19, and 7.23 cm for Persons 1, 2, and 3, respectively. The residual diagnostics did not show strong evidence against normality or equal variances. One-way ANOVA found a difference among mean distances, \(F(2,27)=53.155\), \(p=4.343\times10^{-10}\). An equivalent regression with the coding used in class produced the same overall result. Bonferroni, Tukey, and Scheffe 95% family confidence intervals all excluded zero for each pair. Person 3 had the smallest recorded mean distance, followed by Person 1 and then Person 2. These conclusions describe the recorded tosses under the stated model assumptions.
The 30 recorded distances are listed by participant and toss number. Each column has ten observations, in centimeters.
| Toss | Person 1 | Person 2 | Person 3 |
|---|---|---|---|
| 1 | 8.4 | 13.5 | 6.7 |
| 2 | 11.2 | 11.8 | 8.1 |
| 3 | 9.7 | 14.2 | 7.4 |
| 4 | 7.9 | 12.6 | 5.9 |
| 5 | 10.5 | 15.1 | 9.0 |
| 6 | 12.1 | 10.9 | 6.2 |
| 7 | 8.8 | 13.1 | 7.8 |
| 8 | 9.3 | 14.7 | 8.5 |
| 9 | 11.6 | 12.2 | 5.6 |
| 10 | 10.1 | 13.8 | 7.1 |
| Person | n | Mean | SD | Median |
|---|---|---|---|---|
| Person 1 | 10 | 9.96 | 1.400 | 9.90 |
| Person 2 | 10 | 13.19 | 1.330 | 13.30 |
| Person 3 | 10 | 7.23 | 1.137 | 7.25 |
The group means suggest that Person 3’s tosses landed closest to the target on average, while Person 2’s landed farthest away. The standard deviations range from 1.137 to 1.400 cm.
Distribution of recorded distances by participant.
The one-way ANOVA model assumes independent errors, approximately normal errors within each participant group, and equal error variances. Independence should be justified from how the tosses were carried out. To assess normality and variance, we examined residual plots and used Shapiro-Wilk and median-centered Levene tests at the 5% significance level.
For Shapiro-Wilk, \(H_0\) says the errors are normally distributed; \(H_a\) says they are not. For Levene, \(H_0:\sigma_1^2=\sigma_2^2=\sigma_3^2\); \(H_a\) says at least one group variance differs. A p-value above 0.05 means we fail to reject the relevant null hypothesis; it does not prove that the assumption holds exactly.
Residual diagnostics for the recorded one-way ANOVA.
| Test | Statistic | p-value |
|---|---|---|
| Shapiro-Wilk residual normality | 0.9709 | 0.5653 |
| Levene equality of variances | 0.2514 | 0.7795 |
The Q-Q plot is nearly linear through the middle with modest tail departures. Residual spreads across the three fitted group means appear reasonably similar; three vertical columns are expected because the fitted value is constant within each participant. Shapiro-Wilk gave \(W=0.9709\), \(p=0.5653\), and Levene gave \(F(2,27)=0.2514\), \(p=0.7795\). Thus, neither test found evidence of a serious violation at \(\alpha=0.05\). The actual throwing procedure must still be reviewed for independence.
Let \(\mu_j\) denote the population mean distance for Person \(j\). We tested
\[H_0:\mu_1=\mu_2=\mu_3 \quad\text{versus}\quad H_a:\text{at least one mean differs}. \]
We reject \(H_0\) when the ANOVA p-value is at most 0.05, equivalently when \(F>F_{0.95;2,27}=3.3541\).
| Source | df | SS | MS | F |
|---|---|---|---|---|
| Person | 2 | 178.0247 | 89.0123 | 53.1546 |
| Error | 27 | 45.2140 | 1.6746 | NA |
| Total | 29 | 223.2387 | NA | NA |
The test statistic is \(F(2,27)=53.1546\), with \(p=4.343\times10^{-10}\). Because the p-value is below 0.05, we reject \(H_0\): the recorded data provide strong evidence that not all three population mean distances are equal. The omnibus test does not, by itself, identify which pairs differ.
Using the effect coding provided in class, the equivalent model is
\[Y_i=\mu_{\cdot}+\tau_1X_{1i}+\tau_2X_{2i}+\varepsilon_i,\]
where
\[X_{1i}=\begin{cases}1,&\text{Person 1},\\-1,&\text{Person 3},\\0,&\text{otherwise},\end{cases} \qquad X_{2i}=\begin{cases}1,&\text{Person 2},\\-1,&\text{Person 3},\\0,&\text{otherwise}.\end{cases}\]
Here \(\mu_{\cdot}\) is the unweighted average of the three group means. The model implies \(\mu_1=\mu_{\cdot}+\tau_1\), \(\mu_2=\mu_{\cdot}+\tau_2\), and \(\mu_3=\mu_{\cdot}-\tau_1-\tau_2\). Thus, the equivalent omnibus hypothesis is \(H_0:\tau_1=\tau_2=0\) against the alternative that at least one is nonzero.
| Df | Sum Sq | Mean Sq | F value | Pr(>F) | |
|---|---|---|---|---|---|
| X1 | 1 | 37.2645 | 37.2645 | 22.2529 | 1e-04 |
| X2 | 1 | 140.7602 | 140.7602 | 84.0564 | 0e+00 |
| Residuals | 27 | 45.2140 | 1.6746 | NA | NA |
The regression assigns 37.2645 cm\(^2\) of sequential sum of squares to \(X_1\) and a further 140.7602 cm\(^2\) to \(X_2\), giving 178.0247 cm\(^2\) in total. Individual sequential sums of squares depend on predictor order. The joint regression test is \(F(2,27)=53.1546\), \(p=4.343\times10^{-10}\), so we reject \(H_0\) at the 5% level, matching the one-way ANOVA. The fitted coefficients are \(\hat\mu_{\cdot}=10.1267\), \(\hat\tau_1=-0.1667\), and \(\hat\tau_2=3.0633\) cm.
We formed simultaneous 95% family confidence intervals for all three pairwise mean differences using Bonferroni, Tukey, and Scheffe procedures. A difference is defined as the mean for the first named participant minus the mean for the second. An interval excluding zero indicates a difference for that pair under the procedure.
For participants \(i\) and \(j\), define the estimated difference and its standard error by
\[d_{ij}=\bar Y_i-\bar Y_j,\qquad SE(d_{ij})=\sqrt{MSE\left(\frac{1}{n_i}+\frac{1}{n_j}\right)}.\]
Each interval has the form \(d_{ij}\pm c\,SE(d_{ij})\), with the following critical values:
\[\begin{aligned} \text{Bonferroni:}\quad &c=t_{1-\alpha/(2m),\,\nu},\\ \text{Tukey:}\quad &c=\frac{q_{1-\alpha;\,k,\nu}}{\sqrt{2}},\\ \text{Scheffe:}\quad &c=\sqrt{(k-1)F_{1-\alpha;\,k-1,\nu}}. \end{aligned}\]
Here \(\alpha=0.05\), \(k=3\) participants, \(m=\binom{3}{2}=3\) comparisons, \(\nu=30-3=27\) error degrees of freedom, \(n_i=n_j=10\), and \(MSE\) is the ANOVA residual mean square. Tukey’s \(q\) is the studentized-range critical value; its formula applies directly because the group sizes are equal.
| Pair | Method | Difference | Lower | Upper | Width | Finding |
|---|---|---|---|---|---|---|
| 1 - 2 | Bonferroni | -3.23 | -4.707 | -1.753 | 2.954 | Significant |
| 1 - 2 | Tukey | -3.23 | -4.665 | -1.795 | 2.870 | Significant |
| 1 - 2 | Scheffe | -3.23 | -4.729 | -1.731 | 2.998 | Significant |
| 1 - 3 | Bonferroni | 2.73 | 1.253 | 4.207 | 2.954 | Significant |
| 1 - 3 | Tukey | 2.73 | 1.295 | 4.165 | 2.870 | Significant |
| 1 - 3 | Scheffe | 2.73 | 1.231 | 4.229 | 2.998 | Significant |
| 2 - 3 | Bonferroni | 5.96 | 4.483 | 7.437 | 2.954 | Significant |
| 2 - 3 | Tukey | 5.96 | 4.525 | 7.395 | 2.870 | Significant |
| 2 - 3 | Scheffe | 5.96 | 4.461 | 7.459 | 2.998 | Significant |
For example, Tukey gives a 95% family interval of \((-4.665,-1.795)\) cm for Person 1 minus Person 2; Person 1 therefore has a lower mean distance. The Tukey intervals for Person 1 minus Person 3 and Person 2 minus Person 3 are \((1.295,4.165)\) and \((4.525,7.395)\) cm, respectively. None includes zero, so Tukey distinguishes all three pairs in the recorded data. Bonferroni and Scheffe reach the same pairwise conclusions.
The Tukey intervals are the narrowest here (width 2.870 cm), compared with Bonferroni (2.954 cm) and Scheffe (2.998 cm). Because the objective is all pairwise comparisons, Tukey is the most precise of these three methods for these data. Scheffe also protects more general contrasts and may therefore give wider pairwise intervals.
Simultaneous 95% family confidence intervals. A dashed line marks zero difference.
For the recorded data, mean distances differ among the three participants. Both the one-way ANOVA and its regression formulation reject equality of mean distances. The normality and equal-variance checks do not suggest severe problems, although independence must be assessed from the actual tossing procedure. All three simultaneous-interval methods distinguish every pair, and Tukey gives the narrowest intervals for this pairwise objective. Person 3 appears most accurate in these recorded tosses because their mean distance is smallest. The findings depend on the experimental procedure and need not generalize to other tossing conditions.
The following code generated the tables, tests, and figures. The tables and figures use the recorded measurements in the setup code.
knitr::opts_chunk$set(echo = FALSE, warning = FALSE, message = FALSE,
fig.width = 6.4, fig.height = 3.8,
out.width = "85%", fig.align = "center", fig.pos = "H")
alpha <- 0.05
person_1 <- c(8.4, 11.2, 9.7, 7.9, 10.5, 12.1, 8.8, 9.3, 11.6, 10.1)
person_2 <- c(13.5, 11.8, 14.2, 12.6, 15.1, 10.9, 13.1, 14.7, 12.2, 13.8)
person_3 <- c(6.7, 8.1, 7.4, 5.9, 9.0, 6.2, 7.8, 8.5, 5.6, 7.1)
toss_data <- data.frame(
Person = factor(rep(paste("Person", 1:3), each = 10)),
Toss = rep(1:10, 3), Distance = c(person_1, person_2, person_3)
)
fit <- aov(Distance ~ Person, toss_data)
atab <- anova(fit)
res <- residuals(fit)
shap <- shapiro.test(res)
if (!requireNamespace("car", quietly = TRUE)) {
stop("Install the car package once with install.packages('car')")
}
lev <- car::leveneTest(Distance ~ Person, data = toss_data)
knitr::kable(data.frame(Toss = 1:10, `Person 1` = person_1,
`Person 2` = person_2, `Person 3` = person_3,
check.names = FALSE),
caption = "Recorded distances to the target (cm).")
summ <- do.call(rbind, lapply(split(toss_data$Distance, toss_data$Person),
function(x) data.frame(n = length(x), Mean = mean(x), SD = sd(x),
Median = median(x))))
knitr::kable(data.frame(Person = rownames(summ), summ, row.names = NULL),
digits = 3, caption = "Descriptive statistics by participant (cm).")
boxplot(Distance ~ Person, data = toss_data,
col = c("#7CA5C8", "#85BBA5", "#D9A477"),
xlab = "Participant", ylab = "Distance from target (cm)")
stripchart(Distance ~ Person, data = toss_data, method = "jitter",
vertical = TRUE, add = TRUE, pch = 16, col = "#303030", jitter = 0.1)
op <- par(mfrow = c(2, 1), mar = c(4, 4.5, 2, 1), cex.main = 0.95)
qqnorm(res, main = "Q-Q plot of residuals")
qqline(res, col = "firebrick", lwd = 2)
plot(fitted(fit), res, pch = 16, xlab = "Fitted group mean (cm)",
ylab = "Residual (cm)", main = "Residuals versus fitted")
abline(h = 0, col = "firebrick", lty = 2)
par(op)
knitr::kable(data.frame(Test = c("Shapiro-Wilk residual normality",
"Levene equality of variances"),
Statistic = c(unname(shap$statistic), lev[1, "F value"]),
`p-value` = c(shap$p.value, lev[1, "Pr(>F)"]),
check.names = FALSE), digits = 4,
caption = "Formal assumption checks.")
tab <- data.frame(Source = c("Person", "Error", "Total"),
df = c(2, 27, 29),
SS = c(atab[1, "Sum Sq"], atab[2, "Sum Sq"], sum(atab[, "Sum Sq"])),
MS = c(atab[1, "Mean Sq"], atab[2, "Mean Sq"], NA),
F = c(atab[1, "F value"], NA, NA))
knitr::kable(tab, digits = 4, caption = "One-way ANOVA for distance to target.")
toss_data$X1 <- c(rep(1, 10), rep(0, 10), rep(-1, 10))
toss_data$X2 <- c(rep(0, 10), rep(1, 10), rep(-1, 10))
reg_fit <- lm(Distance ~ X1 + X2, toss_data)
reg_aov <- anova(reg_fit)
knitr::kable(reg_aov, digits = 4,
caption = "Sequential regression ANOVA; X1 enters before X2.")
stopifnot(isTRUE(all.equal(unname(summary(reg_fit)$fstatistic[1]),
unname(atab[1, "F value"]), tolerance = 1e-8)))
means <- tapply(toss_data$Distance, toss_data$Person, mean)
MSE <- atab[2, "Mean Sq"]
nu <- df.residual(fit)
k <- length(means)
m <- choose(k, 2)
mult <- c(Bonferroni = qt(1 - alpha/(2*m), nu),
Tukey = qtukey(1 - alpha, k, nu)/sqrt(2),
Scheffe = sqrt((k - 1)*qf(1 - alpha, k - 1, nu)))
ij <- combn(seq_along(means), 2)
pair_results <- do.call(rbind, lapply(seq_len(ncol(ij)), function(a) {
i <- ij[1, a]; j <- ij[2, a]
est <- unname(means[i] - means[j])
se <- sqrt(MSE * (1/10 + 1/10))
data.frame(Pair = paste(i, "-", j), Method = names(mult),
Difference = est, Lower = est - mult*se, Upper = est + mult*se,
Width = 2*mult*se, row.names = NULL)
}))
pair_results$Finding <- ifelse(pair_results$Lower > 0 |
pair_results$Upper < 0, "Significant", "Not significant")
knitr::kable(pair_results, digits = 3,
caption = "95% family confidence intervals for mean differences (cm).")
cols <- c(Bonferroni = "#4477AA", Tukey = "#228833", Scheffe = "#CC6677")
oldpar <- par(mar = c(4, 6, 1, 1))
limits <- range(c(pair_results$Lower, pair_results$Upper)) + c(-0.5, 0.5)
y <- rev(seq_len(nrow(pair_results)))
plot(pair_results$Difference, y, xlim = limits, ylim = c(0.5, 9.5),
yaxt = "n", xlab = "Difference in mean distance (cm)", ylab = "",
pch = 19, col = cols[pair_results$Method])
segments(pair_results$Lower, y, pair_results$Upper, y,
col = cols[pair_results$Method], lwd = 2)
axis(2, at = y, labels = paste(pair_results$Pair, pair_results$Method,
sep = " "), las = 1, tick = FALSE, cex.axis = 0.77)
abline(v = 0, lty = 2, col = "gray40")
for (pos in c(3.5, 6.5)) {
abline(h = pos, col = "gray85")
}
par(oldpar)