2024-06-05
In statistics, p-values are a method of showing statistical significance and hypothesis testing. The P-Value is used to quantify how likely the outcome obtained is due to chance versus due to an actual causal relationship in the data.
To show the process for calculating a P Value, we will utilize the “Fatalities” data set in R to calculate if a lower drinking age leads to a higher per 100,000 crash fatality rate.
To apply a P Value, you first have to calculate some form of Test Statistic value, which is dependent upon the type of data you are using and the form of your test. For this example, we’ll use the 1-sample T Test. Other test statistic methods are available. The equation is \[ t = \frac{\overline{x}-\mu_0}{s/\sqrt{n}} \]
A brief explaination of the T-Test components:
\[\overline{x}\] - The value being tested for significance, the sample mean \[\mu_0\] - The Null Hypothesis value, or the value being test against for significant difference \[s\] - The Standard Deviation of the dataset \[\sqrt{n}\] - The squareroot of the sample size of the dataset
For our dataset, the data can be used to calculate average driving fatalities per 100,000 people. The average for each state is shown on the histogram below for all the years of the dataset (1982-1988)
To calculate the values for our T value equation, we can use the following R code to easily plug in our values.
statesWithNon21 <- Fatalities %>% filter(drinkage < 21) %>%
reframe(state,fatal,pop,drinkage)
#Filter the Data set to only include states with a drinking age below 21
stdDev <- sd( (statesWithNon21$fatal / statesWithNon21$pop) * 100000)
#Calculate the standard deviation
sampleSize <- nrow(statesWithNon21)
#Calculate the sample size of the filtered data
sampleMean <- mean((statesWithNon21$fatal / statesWithNon21$pop) * 100000)
#Calculate the filtered data mean
popMean <- mean((Fatalities$fatal / Fatalities$pop) * 100000)
#Calculate the population mean
With the means we just calculated, we can visualize what the difference in means before calculating our T value. The box plot below shows how the fatalities per 100,000 are distributed for each state, split into 2 groups based on drinking age.
With the values from our R code, the T Value equation comes out to \[ t = \frac{20.683-20.404}{6.102/\sqrt{104}} \] Which gives us a T Value of \[ t = .466 \]
#Doing a P-Value test
To determine if our value is significant, we use a test on a T-Distribution graph. On the graph below, the area to reject the null hypothesis is shown in red, and our T Value is shown on the blue line. With the given data, the difference in fatalities per 100000 is not significant enough to pass our selected p value of 0.025
While the selected hypothesis, that increasing the legal drinking age to 21 will reduce the amount of traffic fatalities per 100,000 people yearly did not show a high enough change to be statistically significant, this is just one of the many tests that can be ran on the given data to try and gleam information from it. This particular data set also lacked any direct changes, i.e. a state upping the drinking age from year to year that may be a more significant and representitive test.