Suppose we wish to design a new experiment that tests for a significant difference between the mean effective life of four insulating fluids at an accelerated load of 35kV. The variance of fluid life is estimated to be 3.5 hrs based on preliminary data. We would like this test to have a type 1 error probability of 0.05, and for this test to have an 80% probability of rejecting the assumption that the mean life of all the fluids are the same if there is a difference greater than 2 hours between the mean lives of the fluids, with a min of 18 hrs and max of 20 hrs. How many samples of each fluid will need to be collected to achieve this design criterion in the case of Min, Intermediate, and Max variability?
# Known parameters
# H0: m1 = m2 = m3 = m4
k <- 4 # number of fluid types
sigma2 <- 3.5 # error variance between group
alpha <- 0.05 # Type 1 error probability
target_power <- 0.80 # power
means_min <- c(18, 19, 19, 20)
var_min <-var(means_min)
# Sample size per group
power_min <- power.anova.test(groups = k,
between.var = var_min,
within.var = sigma2,
sig.level = alpha,
power = target_power)
print(power_min)
##
## Balanced one-way analysis of variance power calculation
##
## groups = 4
## n = 20.08368
## between.var = 0.6666667
## within.var = 3.5
## sig.level = 0.05
## power = 0.8
##
## NOTE: n is number in each group
For the minimum variability we will need of 21 samples in each group, therefore 84 samples.
means_int <- c(18, 18.667, 19.333, 20)
var_int <-var(means_int)
# Sample size per group
power_int <- power.anova.test(groups = k,
between.var = var_int,
within.var = sigma2,
sig.level = alpha,
power = target_power)
print(power_int)
##
## Balanced one-way analysis of variance power calculation
##
## groups = 4
## n = 18.18209
## between.var = 0.7405927
## within.var = 3.5
## sig.level = 0.05
## power = 0.8
##
## NOTE: n is number in each group
For the intermediate variability we will need of 19 samples in each group, therefore 76 samples.
means_max <- c(18, 18, 20, 20)
var_max <-var(means_max)
# Sample size per group
power_max <- power.anova.test(groups = k,
between.var = var_max,
within.var = sigma2,
sig.level = alpha,
power = target_power)
print(power_max)
##
## Balanced one-way analysis of variance power calculation
##
## groups = 4
## n = 10.56952
## between.var = 1.333333
## within.var = 3.5
## sig.level = 0.05
## power = 0.8
##
## NOTE: n is number in each group
In the maximum variability case, we will need a 11 samples in each group, therefore 44 samples.
The minimum variability case requires the largest sample size (21 per group), whereas the maximum variability case requires the smallest sample size (11 per group).
The experimenter decided to collect six observations for each type of fluid, resulting in the following test data.
# Data Frame
wide_data <- data.frame(
Fluid_1 = c(17.6, 18.9, 16.3, 17.4, 20.1, 21.6),
Fluid_2 = c(16.9, 15.3, 18.6, 17.1, 19.5, 20.3),
Fluid_3 = c(21.4, 23.6, 19.4, 18.5, 20.5, 22.3),
Fluid_4 = c(19.3, 21.1, 16.9, 17.5, 18.3, 19.8)
)
library(knitr)
# Displaying the table
kable(wide_data,
col.names = c("Fluid 1", "Fluid 2", "Fluid 3", "Fluid 4"),
caption = "Table: Test Data")
| Fluid 1 | Fluid 2 | Fluid 3 | Fluid 4 |
|---|---|---|---|
| 17.6 | 16.9 | 21.4 | 19.3 |
| 18.9 | 15.3 | 23.6 | 21.1 |
| 16.3 | 18.6 | 19.4 | 16.9 |
| 17.4 | 17.1 | 18.5 | 17.5 |
| 20.1 | 19.5 | 20.5 | 18.3 |
| 21.6 | 20.3 | 22.3 | 19.8 |
library(tidyverse)
# Convert to tidy format - pivot_longer
tidy_data <- wide_data %>%
pivot_longer(cols = everything(),
names_to = "FluidType",
values_to = "Life") %>%
mutate(FluidType = as.factor(FluidType))
# One-way Anova
anova_model <- aov(Life ~ FluidType, data = tidy_data)
summary(anova_model)
## Df Sum Sq Mean Sq F value Pr(>F)
## FluidType 3 30.17 10.05 3.047 0.0525 .
## Residuals 20 65.99 3.30
## ---
## Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
Conclusion: We tested the follow hypotheses: \(H_0\): \(\mu_1 =
\mu_2 = \mu_3 = \mu_4\) and
\(H_1\): at least one mean differs from
the other.
Because p-value (0.0525) is < than alpha (0.10) we can reject \(H_0\).
# Graphics
par(mfrow = c(2,2)) # divide the display
plot(anova_model)
par(mfrow = c(1,1)) # restore the display
Conclusion: The first plot (Residuals vs Fitted) with the red line flat, means that the variance is stable. The residuals are divided in vertical columns and uniform between groups. The Q-Q Residuals indicates the residuals follows a normal distribution. Scale-location demonstrates the plots distributed homogeneous and the red line is almost flat, as the variance is acceptable. The last graphic (Constant Leverage) shows the fluid group with a vertical dispersion of residues equal and near of zero, therefore the fluids are not indicate error or extremely variance. Furthermore, the diagnostic plots do not indicate violations of the ANOVA assumptions, confirming that the model is adequate.
# Tukey's Test
tukey_result <- TukeyHSD(anova_model, which = "FluidType", conf.level = 0.90)
print(tukey_result)
## Tukey multiple comparisons of means
## 90% family-wise confidence level
##
## Fit: aov(formula = Life ~ FluidType, data = tidy_data)
##
## $FluidType
## diff lwr upr p adj
## Fluid_2-Fluid_1 -0.7000000 -3.2670196 1.8670196 0.9080815
## Fluid_3-Fluid_1 2.3000000 -0.2670196 4.8670196 0.1593262
## Fluid_4-Fluid_1 0.1666667 -2.4003529 2.7336862 0.9985213
## Fluid_3-Fluid_2 3.0000000 0.4329804 5.5670196 0.0440578
## Fluid_4-Fluid_2 0.8666667 -1.7003529 3.4336862 0.8413288
## Fluid_4-Fluid_3 -2.1333333 -4.7003529 0.4336862 0.2090635
par(mar = c(5, 9, 4, 2))
plot(tukey_result, las = 1, col = "blue")
par(mar = c(5, 4, 4, 2))
Conclusion: The Tukey’s test revels that only the comparison between Fluid 3 and Fluid 2 differs significantly, as padj = 0.0441 < alpha = 0.10.
knitr::opts_chunk$set(echo = TRUE, warning=FALSE, message = FALSE)
# Known parameters
# H0: m1 = m2 = m3 = m4
k <- 4 # number of fluid types
sigma2 <- 3.5 # error variance between group
alpha <- 0.05 # Type 1 error probability
target_power <- 0.80 # power
means_min <- c(18, 19, 19, 20)
var_min <-var(means_min)
# Sample size per group
power_min <- power.anova.test(groups = k,
between.var = var_min,
within.var = sigma2,
sig.level = alpha,
power = target_power)
print(power_min)
means_int <- c(18, 18.667, 19.333, 20)
var_int <-var(means_int)
# Sample size per group
power_int <- power.anova.test(groups = k,
between.var = var_int,
within.var = sigma2,
sig.level = alpha,
power = target_power)
print(power_int)
means_max <- c(18, 18, 20, 20)
var_max <-var(means_max)
# Sample size per group
power_max <- power.anova.test(groups = k,
between.var = var_max,
within.var = sigma2,
sig.level = alpha,
power = target_power)
print(power_max)
# Data Frame
wide_data <- data.frame(
Fluid_1 = c(17.6, 18.9, 16.3, 17.4, 20.1, 21.6),
Fluid_2 = c(16.9, 15.3, 18.6, 17.1, 19.5, 20.3),
Fluid_3 = c(21.4, 23.6, 19.4, 18.5, 20.5, 22.3),
Fluid_4 = c(19.3, 21.1, 16.9, 17.5, 18.3, 19.8)
)
library(knitr)
# Displaying the table
kable(wide_data,
col.names = c("Fluid 1", "Fluid 2", "Fluid 3", "Fluid 4"),
caption = "Table: Test Data")
library(tidyverse)
# Convert to tidy format - pivot_longer
tidy_data <- wide_data %>%
pivot_longer(cols = everything(),
names_to = "FluidType",
values_to = "Life") %>%
mutate(FluidType = as.factor(FluidType))
# One-way Anova
anova_model <- aov(Life ~ FluidType, data = tidy_data)
summary(anova_model)
# Graphics
par(mfrow = c(2,2)) # divide the display
plot(anova_model)
par(mfrow = c(1,1)) # restore the display
# Tukey's Test
tukey_result <- TukeyHSD(anova_model, which = "FluidType", conf.level = 0.90)
print(tukey_result)