A curious student wants to know if there’s a difference in typing speed (WPM) between Keyboard A and Keyboard B, based on \(n = 15\) people typing on each brand.
typing <- read.csv("typing.csv")
str(typing)
## 'data.frame': 30 obs. of 2 variables:
## $ Brand: chr "A" "A" "A" "A" ...
## $ Speed: int 72 73 69 75 77 74 74 71 72 78 ...
ggplot(typing, aes(x = Brand, y = Speed, fill = Brand)) +
geom_boxplot() +
scale_fill_manual(values = c("mediumpurple2", "goldenrod2")) +
labs(title = "Typing Speed by Keyboard Brand",
x = "Keyboard Brand",
y = "Typing Speed (WPM)") +
theme_minimal(base_size = 13) +
theme(plot.title = element_text(face = "bold", hjust = 0.5),
legend.position = "none")
Interpretation: The boxplot shows Keyboard B’s typing speeds sit noticeably higher than Keyboard A’s, with very little overlap between the two groups’ interquartile ranges. Keyboard A’s distribution also appears more spread out (its box and whiskers cover a wider range), while Keyboard B’s results are tightly clustered near the top of the scale. This is a strong preliminary visual indication that typing speed differs between the two brands, with Keyboard B appearing to allow for faster typing.
First, we check whether equal variances can be assumed using
var.test():
var.test(Speed ~ Brand, data = typing)
##
## F test to compare two variances
##
## data: Speed by Brand
## F = 7.2406, num df = 14, denom df = 14, p-value = 0.0006805
## alternative hypothesis: true ratio of variances is not equal to 1
## 95 percent confidence interval:
## 2.430884 21.566765
## sample estimates:
## ratio of variances
## 7.240602
The F-test gives \(F = 7.241\) on (14, 14) degrees of freedom, with a p-value of about 0.00068. Since this p-value is far less than \(\alpha = 0.05\), we reject \(H_0: \sigma_A^2 = \sigma_B^2\) and conclude the variances are not equal, so we cannot assume equal variances. This matches what the boxplot already suggested — Keyboard A’s spread is visibly larger than Keyboard B’s.
Because equal variances cannot be assumed, we use Welch’s t-test (the
R default, var.equal = FALSE):
t.test(Speed ~ Brand, data = typing, alternative = "two.sided", var.equal = FALSE)
##
## Welch Two Sample t-test
##
## data: Speed by Brand
## t = -7.5922, df = 17.795, p-value = 5.513e-07
## alternative hypothesis: true difference in means between group A and group B is not equal to 0
## 95 percent confidence interval:
## -8.087352 -4.579315
## sample estimates:
## mean in group A mean in group B
## 73.80000 80.13333
Interpretation: The Welch two-sample t-test gives \(t = -7.592\) on \(df \approx 17.79\), with a p-value of about \(5.51 \times 10^{-7}\). Keyboard A’s mean typing speed is about 73.8 WPM, compared to about 80.13 WPM for Keyboard B. The 95% confidence interval for the difference in means (A \(-\) B) is approximately (-8.09, -4.58) WPM, which does not contain 0.
Since the p-value (\(\approx 5.51 \times 10^{-7}\)) is far less than \(\alpha = 0.05\), the student rejects \(H_0: \mu_A = \mu_B\) and concludes there is a statistically significant difference in typing speed between the two keyboard brands. Because the confidence interval for \(\mu_A - \mu_B\) is entirely negative, we can further conclude that Keyboard B is associated with significantly faster typing speeds than Keyboard A.
This matches the boxplot from part (a): the two groups’ boxes barely overlap, with Keyboard B’s entire distribution sitting above most of Keyboard A’s distribution, visually reinforcing that the ~6.3 WPM difference seen in the hypothesis test is a real, meaningful gap between the two brands rather than something attributable to random sampling variation.
Researchers measured a physical trait in \(n = 10\) sets of monozygotic twins, relative to birth order (First-born vs. Second-born twin).
twins <- read.csv("twins.csv")
str(twins)
## 'data.frame': 10 obs. of 3 variables:
## $ Twins : int 1 2 3 4 5 6 7 8 9 10
## $ First : num 6.08 6.22 7.99 7.44 6.48 7.99 6.32 7.6 6.03 7.52
## $ Second: num 5.73 5.8 8.42 6.84 6.43 8.76 6.32 7.62 6.59 7.67
Because each row represents the same set of twins, measured
once as the First-born and once as the Second-born, the First and Second
measurements are dependent (paired) rather than
independent — each pair shares the same genetic background and
upbringing. This calls for a paired t-test rather than
an independent two-sample test, and because the two samples are not
independent, checking for equal variances with var.test()
does not apply here; that check is only meaningful for independent
samples.
t.test(twins$First, twins$Second, paired = TRUE, alternative = "two.sided", conf.level = 0.95)
##
## Paired t-test
##
## data: twins$First and twins$Second
## t = -0.36577, df = 9, p-value = 0.723
## alternative hypothesis: true mean difference is not equal to 0
## 95 percent confidence interval:
## -0.3664148 0.2644148
## sample estimates:
## mean difference
## -0.051
Interpretation: The paired t-test gives a mean difference (First \(-\) Second) of about \(-0.051\), with \(t = -0.366\) on \(df = 9\) and a p-value of about 0.723. The 95% confidence interval for the mean difference is approximately (-0.366, 0.264), which comfortably contains 0.
Since the p-value (0.723) is far greater than \(\alpha = 0.05\), we fail to reject \(H_0: \mu_d = 0\). The researchers do not have sufficient evidence to claim there is a difference in the measured trait between identical twins based on birth order. The small observed difference between First- and Second-born measurements is easily explained by random variation rather than a genuine effect of birth order.
twins_long <- data.frame(
order = rep(c("First", "Second"), each = nrow(twins)),
trait = c(twins$First, twins$Second)
)
ggplot(twins_long, aes(x = order, y = trait, fill = order)) +
geom_boxplot() +
scale_fill_manual(values = c("lightskyblue3", "salmon2")) +
labs(title = "Twin Trait Measurement by Birth Order",
x = "Birth Order",
y = "Trait Measurement") +
theme_minimal(base_size = 13) +
theme(plot.title = element_text(face = "bold", hjust = 0.5),
legend.position = "none")
Interpretation: The boxplots for the First-born and Second-born twins are very similar in both center and spread, with heavily overlapping interquartile ranges and nearly identical medians. This visual similarity supports the conclusion from the paired t-test in part (a): there is no meaningful difference in the measured trait between twins based on birth order.