Week 5 Discussion

Part 1 - Normal and t distributions

# x values
x <- seq(-4,4,by=0.01)

# plot normal distribution first
plot(x,dnorm(x),
     type = "l",
     lwd=2,
     ylim=c(0,0.4),
     xlab="x",
     ylab = "Density",
     main = "Normal and t distributions"
     )

# add a line for each t distribution

lines(x,dt(x,df=2),col="blue")
lines(x,dt(x,df=5),col="red")
lines(x,dt(x,df=15),col="purple")
lines(x,dt(x,df=30),col="green")
lines(x,dt(x,df=120),col="brown")

# add a legend

legend("topright",
       legend=c("Normal","df=2","df=5","df=15","df=30","df=120"),
       col=c("black","blue","red","purple","green","brown"),
       lty=1)

Part 2

set.seed(123)  # Set seed for reproducibility

mu      <-  108

sigma <-  7.2

data_values <- rnorm(n = 1000,   mean = mu,  sd = sigma)

# calculate z scores
z_score <- (data_values-mu)/sigma

#display charts next to each other
par(mfrow=c(1,2))

#both charts will be histograms
# for each graph - adding a density line to emphasis distribution shape

# graph 1 - original data
hist(data_values,probability = TRUE,ylim=c(0,0.08),main = "Normally distributed data",xlab = "Data values")

lines(density(data_values),lwd=2)

#graph 2 - z scores distribution
hist(z_score,probability=TRUE,ylim=c(0,0.8),main="Z score distribution",xlab="Z scores")

lines(density(z_score),lwd=2)

The normally distributed data and the z score distribution graphs has the same shape. They both follow a normal distribution. Converting the data into z scores only changes the center and scale of the distribution, it doesn’t change the shape of the distribution. The z score distribution shows how many standard deviations the data values below or above the mean, so it is basically shows the original data in a different scale, following the same shape.

*Part 3 - P value**

A P-value is a way to quantify the strength of the evidence against the null hypothesis, meaning in favor of the alternative hypothesis. It tells you the probability of getting results as extreme, or more extreme than the observed results, assuming the null hypothesis is true. After the p value is computed, you compare it to the desired significance level of the hypothesis test. If the p value is greater than the significance level, you fail to reject the null hypothesis, stating that there isn’t enough evidence to support the alternative hypothesis. If the p value is less than the significance level, you reject the null hypothesis, meaning that there is enough evidence to support the alternative hypothesis.