Basic Statistical Inference

In this lab, we will be completing basic statistical inference testing in order to understand more about our datasets and the relationships between the variables.

T Test

First, we will create a T test, comparing the Legolas actors heights’ to the Aragorns’, and then the Gimlis’.

aragorn = rnorm(50, mean = 180, sd = 10)
gimli = rnorm(50, mean = 132, sd = 15)
legolas = rnorm(50,195,15)

t.test(legolas, aragorn, alternative = "two.sided")
## 
##  Welch Two Sample t-test
## 
## data:  legolas and aragorn
## t = 6.7622, df = 85.26, p-value = 1.604e-09
## alternative hypothesis: true difference in means is not equal to 0
## 95 percent confidence interval:
##  12.97131 23.77535
## sample estimates:
## mean of x mean of y 
##  195.5124  177.1390
t.test(legolas, gimli, alternative = "two.sided")
## 
##  Welch Two Sample t-test
## 
## data:  legolas and gimli
## t = 19.572, df = 97.866, p-value < 2.2e-16
## alternative hypothesis: true difference in means is not equal to 0
## 95 percent confidence interval:
##  55.25613 67.72563
## sample estimates:
## mean of x mean of y 
##  195.5124  134.0215

What we find is that there IS a significant difference in height between Legolas and Aragorn actors, as well as a significant difference between Legolas and Gimli actors.

F Test

We will now run an F test to compare the difference between Gimli and Legolas actors and look for statistical significance.

var.test(gimli, legolas)
## 
##  F test to compare two variances
## 
## data:  gimli and legolas
## F = 0.92853, num df = 49, denom df = 49, p-value = 0.7963
## alternative hypothesis: true ratio of variances is not equal to 1
## 95 percent confidence interval:
##  0.5269161 1.6362369
## sample estimates:
## ratio of variances 
##          0.9285255

Due to a P-value of 0.5447, we can not say that the Gimlis and Legolas have significantly different variances.

Correlation Tests

We will now run a correlation test between sepal length and sepal width for the Iris dataset for the three individual species.

cor(iris$Sepal.Length[iris$Species=="setosa"], iris$Sepal.Width[iris$Species=="setosa"])
## [1] 0.7425467

With a correlation of 0.74, the Setosa species has a very high correlation between sepal length and sepal width.

cor(iris$Sepal.Length[iris$Species=="versicolor"], iris$Sepal.Width[iris$Species=="versicolor"])
## [1] 0.5259107

With a correlation of 0.525, the Versicolor species has a high correlation between sepal length and sepal width.

cor(iris$Sepal.Length[iris$Species=="virginica"], iris$Sepal.Width[iris$Species=="virginica"])
## [1] 0.4572278

With a correlation of 0.457, the Virginica species has a moderate correlation between sepal length and sepal width.

Chi Squared Test

We will now test if there are significant differences in the number of deer caught per month, as well as look for significant differences between the cases of tuberculosis across all farms.

deer = read.csv("Deer.csv")
chisq.test(table(deer$Month))
## 
##  Chi-squared test for given probabilities
## 
## data:  table(deer$Month)
## X-squared = 997.07, df = 11, p-value < 2.2e-16

Due to an extremely low P-value, we can say with confidence that the distribution of deer caught across months is not uniform.

chisq.test(table(deer$Farm, deer$Tb))
## Warning in chisq.test(table(deer$Farm, deer$Tb)): Chi-squared approximation may
## be incorrect
## 
##  Pearson's Chi-squared test
## 
## data:  table(deer$Farm, deer$Tb)
## X-squared = 129.09, df = 26, p-value = 1.243e-15

Additionally, due to a p value < 0.05, we can also say that there is a significant difference in the distribution of deers with tuberculosis across different farms.

Thanks!