The effect of five different ingredients (A, B, C, D, E) on the reaction time of a chemical process is being studied. Each batch of new material is only large enough to permit five runs to be made. Furthermore, each run requires approximately 1.5 hours, so only five runs can be made in one day. The experimenter decides to run the experiment as a Latin square so that day and batch effects may be systematically controlled. She obtains the data that follows.
To analyze the data, it will need to be integrated together into a data frame. The code chunk below shows how the data is put into a data frame.
batch<-c(rep(1, 5), rep(2, 5), rep(3, 5), rep(4, 5), rep(5, 5))
day<-c(seq(1, 5), seq(1, 5), seq(1, 5), seq(1, 5), seq(1, 5))
treatment<-c("A", "B", "D", "C", "E",
"C", "E", "A", "D", "B",
"B", "A", "C", "E", "D",
"D", "C", "E", "B", "A",
"E", "D", "B", "A", "C")
obs<-c(8, 7, 1, 7, 3,
11, 2, 7, 3, 8,
4, 9, 10, 1, 5,
6, 8, 6, 6, 10,
4, 2, 3, 8, 8)
batch<-as.factor(batch)
day<-as.factor(day)
treatment<-as.factor(treatment)
dat<-data.frame(batch, day, treatment, obs)
Is this a valid Latin Square? (explain).
This data set does represent a valid Latin Square. All of the treatments (ingredients) are organized such that each one occurs exactly once per day and exactly once per batch. In other words, each ingredient appears exactly once in every row and column.
Write the model equation.
The model equation for a Latin Square design has five variables that have an impact on the response. These five variables are the grand mean, the effect of the individual treatments (or populations), the effect of the first blocking factor applied to the columns, the effect of the second blocking factor applied to the rows, and the random error associated with the measurement. A more concise way to express this is in the equation below.
\[ X_{ijk} = \mu + \tau_k + \alpha_i + \beta_j + \epsilon_{ijk} \]
Analyze the data from this experiment (use α=0.05) and draw conclusions about the factor of interest. (Note: Use aov() instead of gad() for Latin Square Designs, be sure all blocks are recognized as factors).
An analysis of variance will now be performed on the data set to determine whether or not the means of all the reaction times of the different ingredients are equal given a significance level of \(\alpha=0.05\). This will be done using the aov() function and can be seen in the code chunk below.
anova<-aov(obs~treatment+batch+day, data = dat)
summary(anova)
## Df Sum Sq Mean Sq F value Pr(>F)
## treatment 4 141.44 35.36 11.309 0.000488 ***
## batch 4 15.44 3.86 1.235 0.347618
## day 4 12.24 3.06 0.979 0.455014
## Residuals 12 37.52 3.13
## ---
## Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
The analysis yielded a p-value of 0.000488 for the “treatment” factor that we are interested in. Using a significance level of \(\alpha=0.05\), we reject the null hypothesis since the p-value is less than the significance level. This analysis shows that there is enough statistical evidence to suggest that at least one of the reaction time means of the ingredients differs from the others.
Now that we’ve concluded that not all the means are equal, we can perform Tukey’s test to determine which pairs of ingredients significantly differ from each other. This will be done using the TukeyHSD() function and can be seen in the code chunk below.
TukeyHSD(anova, "treatment")
## Tukey multiple comparisons of means
## 95% family-wise confidence level
##
## Fit: aov(formula = obs ~ treatment + batch + day, data = dat)
##
## $treatment
## diff lwr upr p adj
## B-A -2.8 -6.3646078 0.7646078 0.1539433
## C-A 0.4 -3.1646078 3.9646078 0.9960012
## D-A -5.0 -8.5646078 -1.4353922 0.0055862
## E-A -5.2 -8.7646078 -1.6353922 0.0041431
## C-B 3.2 -0.3646078 6.7646078 0.0864353
## D-B -2.2 -5.7646078 1.3646078 0.3365811
## E-B -2.4 -5.9646078 1.1646078 0.2631551
## D-C -5.4 -8.9646078 -1.8353922 0.0030822
## E-C -5.6 -9.1646078 -2.0353922 0.0023007
## E-D -0.2 -3.7646078 3.3646078 0.9997349
plot(TukeyHSD(anova, "treatment"), las = 1)
The Tukey’s test results show that there are four pairs of ingredients that significantly differ from each other. These pairs are A-D, A-E, C-D, and C-E.
# Data Input
batch<-c(rep(1, 5), rep(2, 5), rep(3, 5), rep(4, 5), rep(5, 5))
day<-c(seq(1, 5), seq(1, 5), seq(1, 5), seq(1, 5), seq(1, 5))
treatment<-c("A", "B", "D", "C", "E",
"C", "E", "A", "D", "B",
"B", "A", "C", "E", "D",
"D", "C", "E", "B", "A",
"E", "D", "B", "A", "C")
obs<-c(8, 7, 1, 7, 3,
11, 2, 7, 3, 8,
4, 9, 10, 1, 5,
6, 8, 6, 6, 10,
4, 2, 3, 8, 8)
batch<-as.factor(batch)
day<-as.factor(day)
treatment<-as.factor(treatment)
dat<-data.frame(batch, day, treatment, obs)
# Problem 3
anova<-aov(obs~treatment+batch+day, data = dat)
summary(anova)
TukeyHSD(anova, "treatment")
plot(TukeyHSD(anova, "treatment"), las = 1)