#install.packages("tidyverse")
library(tidyverse)
#install.packages("car")
library(car)
Suppose we wish to design a new experiment that tests for a significant difference between the mean effective life of four insulating fluids at an accelerated load of 35kV. The variance of fluid life is estimated to be 3.5 hrs based on preliminary data. We would like this test to have a type 1 error probability of 0.05, and for this test to have an 80% probability of rejecting the assumption that the mean life of all the fluids are the same if there is a difference greater than 2 hours between the mean lives of the fluids, with a min of 18 hrs and max of 20 hrs. How many samples of each fluid will need to be collected to achieve this design criterion in the case of Min, Intermediate, and Max variability?
Using the minimum variability criterion, we assume one fluid at the minimum lifetime of 18 hours, one fluid at the maximum lifetime of 20 hours, and two fluids at the midpoint of 19 hours with four groups. The variance within the populations is 3.5 hours according to the preliminary data. The significance level is 0.05, and a power of 0.80 is desired. With all of these parameters established, the power.anova.test function can be used to determine the number of samples needed to detect a difference in the means at a power of 0.80. The code chunk for the test can be seen below.
power.anova.test(groups = 4, n = NULL, between.var = var(c(18, 19, 19, 20)), within.var = 3.5, sig.level = 0.05, power = 0.8)
##
## Balanced one-way analysis of variance power calculation
##
## groups = 4
## n = 20.08368
## between.var = 0.6666667
## within.var = 3.5
## sig.level = 0.05
## power = 0.8
##
## NOTE: n is number in each group
The ANOVA power calculation shows that 21 samples need to be taken for each group.
Using the intermediate variability criterion, we assume one fluid at the minimum lifetime of 18 hours, one fluid at a lifetime of 18.67 hours, one fluid at a lifetime of 19.33 hours and one fluid at the maximum lifetime of 20 hours with four groups. The other parameters from the minimum variability will remain the same. With all of these parameters established, the power.anova.test function can be used to determine the number of samples needed to detect a difference in the means at a power of 0.80. The code chunk for the test can be seen below.
power.anova.test(groups = 4, n = NULL, between.var = var(c(18, 18.67, 19.33, 20)), within.var = 3.5, sig.level = 0.05, power = 0.8)
##
## Balanced one-way analysis of variance power calculation
##
## groups = 4
## n = 18.21285
## between.var = 0.7392667
## within.var = 3.5
## sig.level = 0.05
## power = 0.8
##
## NOTE: n is number in each group
The ANOVA power calculation shows that 19 samples need to be taken for each group.
Using the maximum variability criterion, we assume two fluids at the minimum lifetime of 18 hours and two fluids at the maximum lifetime of 20 hours with four groups. The other parameters from the minimum and intermediate variability tests will remain the same. With all of these parameters established, the power.anova.test function can be used to determine the number of samples needed to detect a difference in the means at a power of 0.80. The code chunk for the test can be seen below.
power.anova.test(groups = 4, n = NULL, between.var = var(c(18, 18, 20, 20)), within.var = 3.5, sig.level = 0.05, power = 0.8)
##
## Balanced one-way analysis of variance power calculation
##
## groups = 4
## n = 10.56952
## between.var = 1.333333
## within.var = 3.5
## sig.level = 0.05
## power = 0.8
##
## NOTE: n is number in each group
The ANOVA power calculation shows that 11 samples need to be taken for each group.
The experimenter decided to collect six observations for each type of fluid, resulting in the following test data.
f1<-c(17.6, 18.9, 16.3, 17.4, 20.1, 21.6)
f2<-c(16.9, 15.3, 18.6, 17.1, 19.5, 20.3)
f3<-c(21.4, 23.6, 19.4, 18.5, 20.5, 22.3)
f4<-c(19.3, 21.1, 16.9, 17.5, 18.3, 19.8)
dat<-data.frame(f1, f2, f3, f4)
dat1<-pivot_longer(dat, cols = c("f1", "f2", "f3", "f4"))
Test the hypothesis that the mean life of fluids is the same against the alternative that they differ at an α=0.10 level of significance (Remember to enter the data in a tidy format when using R, or to pivot_longer to a tidy format using tidyr)
The hypothesis will be tested using the aov() function to determine whether or not the mean life of the fluids differ. The code chunk can be seen below.
anova<-aov(value~name, data = dat1)
summary(anova)
## Df Sum Sq Mean Sq F value Pr(>F)
## name 3 30.17 10.05 3.047 0.0525 .
## Residuals 20 65.99 3.30
## ---
## Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
Since we are using a type-one error of 0.10 and the one-way ANOVA returned a p-value (0.0525) less than the type-one error, we reject the null hypothesis that the mean life of the fluids are equal. There is enough statistical evidence to suggest that at least one of the mean life of fluids differs from the others.
Is the model adequate? (show plots and comment)
The plots for the one-way ANOVA can be seen in the code chunk below.
plot(anova)
We will use the plots above to determine if this data set follows the two main assumptions to perform a valid one-way ANOVA analysis. These two assumptions are normality of the data and constant variance between the populations. Looking at the “Residuals vs Fitted” plot, the four different populations all have a similar spread of the data. This indicates a constant variance between the four populations. Additionally, the “Q-Q Residuals” plot shows a linear trend, indicating normality between the four populations. The model appears to be adequate since it meets the two main assumptions needed to perform a one-way ANOVA.
Assuming the null hypothesis in question 1 is rejected, which fluids significantly differ using a familywise error rate of α=0.10 (use Tukey’s test). Include the plot of confidence intervals.
We can see which fluids significantly differ using Tukey’s test. This test is performed by using the function TukeyHSD(), which can be accessed from the car library. The code chunk can be seen below.
TukeyHSD(anova, conf.level = 0.9)
## Tukey multiple comparisons of means
## 90% family-wise confidence level
##
## Fit: aov(formula = value ~ name, data = dat1)
##
## $name
## diff lwr upr p adj
## f2-f1 -0.7000000 -3.2670196 1.8670196 0.9080815
## f3-f1 2.3000000 -0.2670196 4.8670196 0.1593262
## f4-f1 0.1666667 -2.4003529 2.7336862 0.9985213
## f3-f2 3.0000000 0.4329804 5.5670196 0.0440578
## f4-f2 0.8666667 -1.7003529 3.4336862 0.8413288
## f4-f3 -2.1333333 -4.7003529 0.4336862 0.2090635
plot(TukeyHSD(anova, conf.level = 0.9))
The “90% family-wise confidence level” plot shows that there is only one pair of fluids that significantly differ from each other. Fluid 2 & Fluid 3 is the only pair of fluids that does not contain zero in its lower to upper range. This provides evidence that Fluids 2 & 3 significantly differ at the 90% confidence level.
# Package Install
#install.packages("tidyverse")
library(tidyverse)
#install.packages("car")
library(car)
## Minimum Variability Samples
power.anova.test(groups = 4, n = NULL, between.var = var(c(18, 19, 19, 20)), within.var = 3.5, sig.level = 0.05, power = 0.8)
## Intermediate Variability Samples
power.anova.test(groups = 4, n = NULL, between.var = var(c(18, 18.67, 19.33, 20)), within.var = 3.5, sig.level = 0.05, power = 0.8)
## Maximum Variability Samples
power.anova.test(groups = 4, n = NULL, between.var = var(c(18, 18, 20, 20)), within.var = 3.5, sig.level = 0.05, power = 0.8)
# Problem 2
f1<-c(17.6, 18.9, 16.3, 17.4, 20.1, 21.6)
f2<-c(16.9, 15.3, 18.6, 17.1, 19.5, 20.3)
f3<-c(21.4, 23.6, 19.4, 18.5, 20.5, 22.3)
f4<-c(19.3, 21.1, 16.9, 17.5, 18.3, 19.8)
dat<-data.frame(f1, f2, f3, f4)
dat1<-pivot_longer(dat, cols = c("f1", "f2", "f3", "f4"))
## Problem 2.A
anova<-aov(value~name, data = dat1)
summary(anova)
## Problem 2.B
plot(anova)
## Problem 2.C
TukeyHSD(anova, conf.level = 0.9)
plot(TukeyHSD(anova, conf.level = 0.9))