1. Normal & T Distribution Plots

set.seed(930)

I added a density curve in red to the histogram curve with curve() and dnorm(). Needed to add the probability = T argument so the plot would show desnity rather than value counts.

par(mfrow = c(1, 2))
Normal_Dist <- rnorm(n = 500, mean = 0, sd = 1)
hist(Normal_Dist, probability = T, main = "Normal Dist. (random)", ylim = c(0, 0.5))
curve(dnorm(x, mean = 0, sd = 1), add = T, col = "red", lwd = 5)

dfs <- c(2,5,15,30,120)
for (i in dfs){
  curve(dt(x, df = i), from = -4, to = 4, add = T, col = "blue", lwd = 1)
}

hist(Normal_Dist, probability = T, main = "Normal Dist. (random)", ylim = c(0, 0.5))
curve(dnorm(x, mean = 0, sd = 1), add = T, col = "red", lwd = 5)
curve(dt(x, df = 2), from = -4, to = 4, add = T, col = "blue", lwd = 2)
curve(dt(x, df = 120), from = -4, to = 4, add = T, col = "blue", lwd = 2)

I had the t distributions, with the varying degrees of freedom plotted in blue. I added a second plot with just the df = 2 and df = 120 just to see the t-dist’s journey to become approximately normal distributions.

2 P-Value

set.seed(931)

Normal Distribution

par(mfrow = c(1, 1))
mu <- 108
sigma <- 7.2
data_values <- rnorm(n = 1000, mean = mu, sd = sigma)
hist(data_values, probability = T, main = "data_values Normal Dist")

Z-Score Distribution

data_values_Zscore <- (data_values - mean(data_values)) / sd(data_values)
hist(data_values_Zscore, probability = T, main = "data_values Z Scores")

The two plotted distributions share the same distributional shape. This is not unexpected since changing from the original data set values to z-scores doesn’t change the shape of the distribution itself. Rather, a z-score is a method of standardizing the distribution’s values. When we have distributions with means and standard deviations that are not 0 and 1, it can be a little abstract when trying to describe how far away certain values are from the center. When we standardize the values, the distribution has a mean of 0 and a standard deviation of 1, giving us a clear distance from the mean in terms of standard deviations

3. P-Value

A p-value is a measurement of how likely we’re to have witnessed our observed result (or something even more unusual), under the pretense that the null hypothesis (H0) is actually true. After calculating the appropriate test stat, and comparing it to the data set distribution we would expect to see, we can determine a p-value for the observed result.

Now large p-values mean this outcome isn’t particularly out of the norm, under the guise of the H0, so we would Fail to Reject the H0 in these cases. Now, if the corresponding p-value is extremely small, then we may need to rethink our belief in the H0. But how small does it need to be? Beforehand, a cutoff point, or alpha, is decided upon. If the p-value is larger than alpha, the result isn’t considered statistically significant and we fail to reject H0. But if it’s smaller than alpha, then we do reject H0.

We can look at the appropriate distribution to help visualize how large or small the p-value is. The distribution used depends on the type of hypothesis test being performed, such as normal, t, chi-square or F distribution. The p-value can be depicted by the area under the curve, corresponding to our observed test statistic and outcomes at least as unusual as it, again all assuming that H0 is true.