A snack food company uses a machine to package 454 oz bags of peanuts. As part of an in-line inspection, operators measured the weights (in oz) of \(n = 25\) randomly selected bags.
peanuts <- c(456.1, 454.9, 463.4, 454.4, 439.9, 439.4, 433.6, 454.4, 441.2, 451.7,
451.1, 454.1, 449.7, 450.1, 449.6, 449.8, 448.2, 451.5, 447.9, 449.2,
455.1, 454.5, 459.2, 453.7, 456.5)
peanuts_df <- data.frame(weight = peanuts)
ggplot(peanuts_df, aes(sample = weight)) +
stat_qq(color = "darkorange", size = 2) +
stat_qq_line(color = "steelblue", linetype = "dashed", linewidth = 1) +
labs(title = "Normal QQ Plot of Peanut Bag Weights",
x = "Theoretical Quantiles",
y = "Sample Quantiles (oz)") +
theme_minimal(base_size = 13) +
theme(plot.title = element_text(face = "bold", hjust = 0.5))
We can also run a Shapiro-Wilk test to formally check the normality assumption:
shapiro.test(peanuts)
##
## Shapiro-Wilk normality test
##
## data: peanuts
## W = 0.92568, p-value = 0.06913
Interpretation: The majority of the points in the QQ plot fall reasonably close to the reference line, but a few of the lowest observations (433.6, 439.4, and 439.9 oz) fall noticeably below the line, pulling away from what we’d expect under normality. This suggests a mild left-skew / heavier lower tail rather than a perfectly normal distribution. The Shapiro-Wilk test supports this reading: the p-value (≈ 0.069) is greater than \(\alpha = 0.05\), so at that significance level we technically fail to reject the assumption of normality, but the result is close enough to the cutoff that the departure from normality is worth noting rather than dismissing outright.
Even so, the one-sample \(t\)-test results will still be valid. With \(n = 25\), the sample size is large enough that, by the Central Limit Theorem, the sampling distribution of the mean will be approximately normal even though the underlying data show a mild departure from normality. Since the deviation we see is a modest number of lower-tail points rather than an extreme skew or severe outliers, the \(t\)-test is considered robust to this kind of minor violation, so we can proceed with confidence in the hypothesis test conclusions below.
We test whether the machine is packaging bags at the expected weight of 454 oz, using a two-sided alternative hypothesis \(H_1: \mu \neq 454\), against a significance level of \(\alpha = 0.05\).
t.test(x = peanuts, mu = 454, alternative = "two.sided", conf.level = 0.95)
##
## One Sample t-test
##
## data: peanuts
## t = -2.45, df = 24, p-value = 0.02196
## alternative hypothesis: true mean is not equal to 454
## 95 percent confidence interval:
## 448.0454 453.4906
## sample estimates:
## mean of x
## 450.768
Interpretation: The sample mean weight is about 450.77 oz, and the one-sample \(t\)-test yields \(t = -2.450\) on \(df = 24\), giving a p-value of about 0.022. Since this p-value is less than \(\alpha = 0.05\), we reject the null hypothesis \(H_0: \mu = 454\). This means there is sufficient evidence, at the 5% significance level, to conclude the true mean weight of the packaged bags differs from the target of 454 oz.
The 95% confidence interval is approximately (448.05, 453.49) oz. Because this entire interval falls below 454 oz, it tells us not only that the mean weight differs from the target, but specifically that it is significantly lower than 454 oz. In other words, we do not expect the machine to be packaging bags at the correct weight — on average, it appears to be underfilling the bags relative to the stated 454 oz target, and the machine likely needs to be recalibrated.
According to guidebooks, an eruption of Old Faithful lasts about 3
minutes on average. A disgruntled visitor claimed the eruption they
witnessed was shorter than expected. Using the built-in
faithful data set, we test \(H_0:
\mu = 3\) against \(H_1: \mu <
3\) at a significance level of \(\alpha
= 0.1\).
data("faithful")
t.test(x = faithful$eruptions, mu = 3, alternative = "less", conf.level = 0.9)
##
## One Sample t-test
##
## data: faithful$eruptions
## t = 7.0483, df = 271, p-value = 1
## alternative hypothesis: true mean is less than 3
## 90 percent confidence interval:
## -Inf 3.576691
## sample estimates:
## mean of x
## 3.487783
Interpretation: The sample mean eruption length in
the faithful data set is about 3.49 minutes, which is
actually higher than the 3-minute benchmark from the
guidebooks, not lower. Because the sample mean lies above the
hypothesized value of 3, the one-sided test statistic for \(H_1: \mu < 3\) comes out positive rather
than negative, which drives the p-value very close to 1 — far greater
than \(\alpha = 0.1\). Therefore, we
fail to reject \(H_0\).
Based on this result, the “disgruntled visitor” was not
correct in their claim: there is no evidence that Old
Faithful’s eruptions are, on average, shorter than the advertised 3
minutes. If anything, the overall average eruption length in the data
set slightly exceeds 3 minutes. (It’s worth noting the
faithful eruption times are actually bimodal — there is a
cluster of shorter eruptions around 2 minutes and a cluster of longer
eruptions around 4.5 minutes — so it’s entirely possible the visitor
really did observe one of the shorter eruptions. However, that does not
change the conclusion about the population mean as a whole,
which is the quantity this hypothesis test addresses.)
Through-hikers on the Appalachian Trail are assumed, based on AMC data, to complete the trail at a rate of \(p = 0.25\). A journalist interviewed \(n = 72\) randomly selected through-hikers in Georgia and later found that \(x = 12\) of them completed the trail. We test \(H_0: p = 0.25\) against a two-sided alternative \(H_1: p \neq 0.25\) at \(\alpha = 0.05\).
prop.test(x = 12, n = 72, p = 0.25, alternative = "two.sided", conf.level = 0.95)
##
## 1-sample proportions test with continuity correction
##
## data: 12 out of 72, null probability 0.25
## X-squared = 2.2407, df = 1, p-value = 0.1344
## alternative hypothesis: true p is not equal to 0.25
## 95 percent confidence interval:
## 0.09272582 0.27697767
## sample estimates:
## p
## 0.1666667
Interpretation: The sample proportion of hikers who completed the trail is \(\hat{p} = 12/72 \approx 0.167\). The test statistic is \(X^2 = 2.241\) on \(df = 1\), giving a p-value ofabout 0.134. Because this p-value is greater than \(\alpha = 0.05\), we fail to reject the null hypothesis.
The 95% confidence interval for the proportion is approximately (0.093, 0.277). Since this interval contains the AMC’s assumed value of \(p = 0.25\), this is consistent with failing to reject \(H_0\).
Based on this sample, the journalist cannot conclude that the true completion proportion of through-hikers differs from the 0.25 rate assumed by the AMC. Although the observed sample proportion(≈ 16.7%) is numerically lower than 25%, with only \(n = 72\) hikers interviewed, the sample size is not large enough to provide statistically significant evidence that the true completion rate differs from what the AMC has established.