1 Exercise 1

Suppose we wish to design a new experiment that tests for a significant difference between the mean effective life of four insulating fluids at an accelerated load of 35kV. The variance of fluid life is estimated to be 3.5 hrs based on preliminary data. We would like this test to have a type 1 error probability of 0.05, and for this test to have an 80% probability of rejecting the assumption that the mean life of all the fluids are the same if there is a difference greater than 2 hours between the mean lives of the fluids, with a min of 18 hrs and max of 20 hrs. How many samples of each fluid will need to be collected to achieve this design criterion in the case of Min, Intermediate, and Max variability?

# Known parameters
# H0: m1 = m2 = m3 = m4
k <- 4 # number of fluid types
sigma2 <- 3.5 # error variance between group 
alpha <- 0.05 # Type 1 error probability
target_power <- 0.80 # power

1.1 Design criterion in the case of minimum variability

means_min <- c(18, 19, 19, 20)
var_min <-var(means_min)

# Sample size per group
power_min <- power.anova.test(groups = k,
                              between.var = var_min,
                              within.var = sigma2,
                              sig.level = alpha,
                              power = target_power)
print(power_min)
## 
##      Balanced one-way analysis of variance power calculation 
## 
##          groups = 4
##               n = 20.08368
##     between.var = 0.6666667
##      within.var = 3.5
##       sig.level = 0.05
##           power = 0.8
## 
## NOTE: n is number in each group

For the minimum variability we will need of 21 samples in each group, therefore 84 samples.

1.2 Design criterion in the case of intermediate variability

means_int <- c(18, 18.667, 19.333, 20)
var_int <-var(means_int)

# Sample size per group
power_int <- power.anova.test(groups = k,
                              between.var = var_int,
                              within.var = sigma2,
                              sig.level = alpha,
                              power = target_power)
print(power_int)
## 
##      Balanced one-way analysis of variance power calculation 
## 
##          groups = 4
##               n = 18.18209
##     between.var = 0.7405927
##      within.var = 3.5
##       sig.level = 0.05
##           power = 0.8
## 
## NOTE: n is number in each group

For the intermediate variability we will need of 19 samples in each group, therefore 76 samples.

1.3 Design criterion in the case of maximun variability

means_max <- c(18, 18, 20, 20)
var_max <-var(means_max)

# Sample size per group
power_max <- power.anova.test(groups = k,
                              between.var = var_max,
                              within.var = sigma2,
                              sig.level = alpha,
                              power = target_power)
print(power_max)
## 
##      Balanced one-way analysis of variance power calculation 
## 
##          groups = 4
##               n = 10.56952
##     between.var = 1.333333
##      within.var = 3.5
##       sig.level = 0.05
##           power = 0.8
## 
## NOTE: n is number in each group

In the maximum variability case, we will need a 11 samples in each group, therefore 44 samples.

1.4 Conclusion

The minimum variability case requires the largest sample size (21 per group), whereas the maximum variability case requires the smallest sample size (11 per group).

2 Exercise 2

The experimenter decided to collect six observations for each type of fluid, resulting in the following test data.

# Data Frame
wide_data <- data.frame(
Fluid_1 = c(17.6, 18.9, 16.3, 17.4, 20.1, 21.6),
Fluid_2 = c(16.9, 15.3, 18.6, 17.1, 19.5, 20.3),
Fluid_3 = c(21.4, 23.6, 19.4, 18.5, 20.5, 22.3),
Fluid_4 = c(19.3, 21.1, 16.9, 17.5, 18.3, 19.8)
)

library(knitr)

# Displaying the table
kable(wide_data,
      col.names = c("Fluid 1", "Fluid 2", "Fluid 3", "Fluid 4"),
      caption = "Table: Test Data")
Table: Test Data
Fluid 1 Fluid 2 Fluid 3 Fluid 4
17.6 16.9 21.4 19.3
18.9 15.3 23.6 21.1
16.3 18.6 19.4 16.9
17.4 17.1 18.5 17.5
20.1 19.5 20.5 18.3
21.6 20.3 22.3 19.8
  1. Test the hypothesis that the mean life of fluids is the same against the alternative that they differ at an α=0.10 level of significance.
library(tidyverse)

# Convert to tidy format - pivot_longer
tidy_data <- wide_data %>%
  pivot_longer(cols = everything(),
               names_to = "FluidType",
               values_to = "Life") %>%
  mutate(FluidType = as.factor(FluidType))

# One-way Anova
anova_model <- aov(Life ~ FluidType, data = tidy_data)
summary(anova_model)
##             Df Sum Sq Mean Sq F value Pr(>F)  
## FluidType    3  30.17   10.05   3.047 0.0525 .
## Residuals   20  65.99    3.30                 
## ---
## Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

Conclusion: We tested the follow hypotheses: \(H_0\): \(\mu_1 = \mu_2 = \mu_3 = \mu_4\) and
\(H_1\): at least one mean differs from the other.

Because p-value (0.0525) is < than alpha (0.10) we can reject \(H_0\).

  1. Is the model adequate?
# Graphics
par(mfrow = c(2,2)) # divide the display
plot(anova_model)

par(mfrow = c(1,1)) # restore the display

Conclusion: The first plot (Residuals vs Fitted) with the red line flat, means that the variance is stable. The residuals are divided in vertical columns and uniform between groups. The Q-Q Residuals indicates the residuals follows a normal distribution. Scale-location demonstrates the plots distributed homogeneous and the red line is almost flat, as the variance is acceptable. The last graphic (Constant Leverage) shows the fluid group with a vertical dispersion of residues equal and near of zero, therefore the fluids are not indicate error or extremely variance. Furthermore, the diagnostic plots do not indicate violations of the ANOVA assumptions, confirming that the model is adequate.

  1. Assuming the null hypothesis in question 1 is rejected, which fluids significantly differ using a family wise error rate of α=0.10 (use Tukey’s test). Include the plot of confidence intervals.
# Tukey's Test
tukey_result <- TukeyHSD(anova_model, which = "FluidType", conf.level = 0.90)

print(tukey_result)
##   Tukey multiple comparisons of means
##     90% family-wise confidence level
## 
## Fit: aov(formula = Life ~ FluidType, data = tidy_data)
## 
## $FluidType
##                       diff        lwr       upr     p adj
## Fluid_2-Fluid_1 -0.7000000 -3.2670196 1.8670196 0.9080815
## Fluid_3-Fluid_1  2.3000000 -0.2670196 4.8670196 0.1593262
## Fluid_4-Fluid_1  0.1666667 -2.4003529 2.7336862 0.9985213
## Fluid_3-Fluid_2  3.0000000  0.4329804 5.5670196 0.0440578
## Fluid_4-Fluid_2  0.8666667 -1.7003529 3.4336862 0.8413288
## Fluid_4-Fluid_3 -2.1333333 -4.7003529 0.4336862 0.2090635
par(mar = c(5, 9, 4, 2)) 
plot(tukey_result, las = 1, col = "blue")

par(mar = c(5, 4, 4, 2))

Conclusion: The Tukey’s test revels that only the comparison between Fluid 3 and Fluid 2 differs significantly, as padj = 0.0441 < alpha = 0.10.

3 Complete R Code

knitr::opts_chunk$set(echo = TRUE, warning=FALSE, message = FALSE)
# Known parameters
# H0: m1 = m2 = m3 = m4
k <- 4 # number of fluid types
sigma2 <- 3.5 # error variance between group 
alpha <- 0.05 # Type 1 error probability
target_power <- 0.80 # power

means_min <- c(18, 19, 19, 20)
var_min <-var(means_min)

# Sample size per group
power_min <- power.anova.test(groups = k,
                              between.var = var_min,
                              within.var = sigma2,
                              sig.level = alpha,
                              power = target_power)
print(power_min)

means_int <- c(18, 18.667, 19.333, 20)
var_int <-var(means_int)

# Sample size per group
power_int <- power.anova.test(groups = k,
                              between.var = var_int,
                              within.var = sigma2,
                              sig.level = alpha,
                              power = target_power)
print(power_int)

means_max <- c(18, 18, 20, 20)
var_max <-var(means_max)

# Sample size per group
power_max <- power.anova.test(groups = k,
                              between.var = var_max,
                              within.var = sigma2,
                              sig.level = alpha,
                              power = target_power)
print(power_max)

# Data Frame
wide_data <- data.frame(
Fluid_1 = c(17.6, 18.9, 16.3, 17.4, 20.1, 21.6),
Fluid_2 = c(16.9, 15.3, 18.6, 17.1, 19.5, 20.3),
Fluid_3 = c(21.4, 23.6, 19.4, 18.5, 20.5, 22.3),
Fluid_4 = c(19.3, 21.1, 16.9, 17.5, 18.3, 19.8)
)

library(knitr)

# Displaying the table
kable(wide_data,
      col.names = c("Fluid 1", "Fluid 2", "Fluid 3", "Fluid 4"),
      caption = "Table: Test Data")

library(tidyverse)

# Convert to tidy format - pivot_longer
tidy_data <- wide_data %>%
  pivot_longer(cols = everything(),
               names_to = "FluidType",
               values_to = "Life") %>%
  mutate(FluidType = as.factor(FluidType))

# One-way Anova
anova_model <- aov(Life ~ FluidType, data = tidy_data)
summary(anova_model)

# Graphics
par(mfrow = c(2,2)) # divide the display
plot(anova_model)
par(mfrow = c(1,1)) # restore the display

# Tukey's Test
tukey_result <- TukeyHSD(anova_model, which = "FluidType", conf.level = 0.90)

print(tukey_result)