2024-06-09

What is a Hypothesis?

A hypothesis is a statement that can be tested.

Is this a hypothesis statement?

  • This is a question, not a statement; therefore it is NOT a hypothesis.

If I change the previous question into a statement, then it will be a hypothesis.

  • Whether this statement is true or not, it is a statement, and it is testable; therefore, it is a hypothesis.

What is a Hypothesis? (cont.)

Hypotheses are often conditional if-then statements, but they don’t have to be.

  • Our ‘If’ is the condition that changes, or our independent variable.

  • Our ‘Then’ is the result of that conditional change, or our dependent variable.

Our hypotheses don’t always need an If and Then, but our hypotheses do always need an independent variable and dependent variable.

Null Hypotheses

After generating our hypothesis, we must also define what our null hypothesis is before we can perform our hypothesis test.

  • We take the same independent variable, but instead state our dependent variable will have no change.

NOTE: The null of our hypothesis and opposite of our hypothesis are two different statements.

Hypothesis and Null Hypothesis: Example

I typically finish my assignments one day before they are due.

Hypothesis:

If I work an extra hour a day, I will finish this assignment faster than normal. H1: \(\mu\) > 1_day

Null Hypothesis: If I work an extra hour a day, I will not finish this assignment faster than normal. H0: \(\mu\) \(\le\) 1_day

NOTE: My null hypothesis was “I will not finish this assignment faster than normal”, NOT “I will finish this assignment slower than normal”

  • “Slower than normal” omits \(\mu\) = 1_day as a result.

Our Data

For the rest of our examples, we’ll be drawing from the following data frame:

movies_PresentationDF = select(movies, title, year, budget, rating, 
                               votes, Action, Animation, Comedy, 
                               Drama, Documentary, Romance, Short)

The original dataframe, “movies”, is included in the library “ggplot2movies” and draws its data from IMDB, an internet movie database.

Our Data, Continued

head(movies_PresentationDF)
## # A tibble: 6 × 12
##   title       year budget rating votes Action Animation Comedy Drama Documentary
##   <chr>      <int>  <int>  <dbl> <int>  <int>     <int>  <int> <int>       <int>
## 1 $           1971     NA    6.4   348      0         0      1     1           0
## 2 $1000 a T…  1939     NA    6      20      0         0      1     0           0
## 3 $21 a Day…  1941     NA    8.2     5      0         1      0     0           0
## 4 $40,000     1996     NA    8.2     6      0         0      1     0           0
## 5 $50,000 C…  1975     NA    3.4    17      0         0      0     0           0
## 6 $pent       2000     NA    4.3    45      0         0      0     1           0
## # ℹ 2 more variables: Romance <int>, Short <int>

First Hypothesis Test:

Hypothesis: If a comedy movie was made in 2005, then its budget was higher than a comedy movie made in 1990.

  • H1: \(\mu\) > 1990_Budget

Null Hypothesis: If a comedy movie was made in 2005, then its budget was not higher than a comedy movie made in 1990.

  • H0: \(\mu\) \(\le\) 1990_Budget

First Hypothesis Test, First Code:

FirstMin_MaxDF = movies %>% 
  filter(Short == 0, budget > 0, year %in% c(1990:2005)) %>%
  select(year,budget,Comedy) %>%
  group_by(year) %>%
  summarize(FirstMin_Max_Budget = min(budget))

First Hypothesis Test, First Graph:

Our null hypothesis is proven true: a comedy movie made in 1990 had a higher budget than one made in 2005

First Hypothesis Test, Second Code:

SecondMin_MaxDF = movies %>% 
  filter(Short == 0, budget > 0, year %in% c(1990:2005)) %>%
  select(year,budget,Comedy) %>%
  group_by(year) %>%
  summarize(SecondMin_Max_Budget = max(budget))

First Hypothesis Test, Second Graph:

Our hypothesis is proven true: a comedy movie made in 2005 does have a higher budget than one made in 1990.

First Hypothesis Test, Conclusion:

This is a bad hypothesis. It is too vague and allows too much room for interpretation.

Because we didn’t specify exactly what we wanted to measure, our hypothesis test allowed for a variety of different results. Depending on the movies we choose to measure, our hypothesis could be proven right, or it could be wrong.

The only way to prove our hypothesis would be to directly compare every comedy movie made in 1990 to every comedy movie made in 2005 one-by-one, which is unrealistic.

Proving our null hypothesis is much easier: we simply find one comedy movie made in 1990 with a higher budget than a comedy movie made in 2005.

Second Hypothesis Test:

Let’s write a more specific hypothesis this time:

Hypothesis: The average comedy movie budget in 2005 was higher than the average comedy movie budget in 1990.

  • H1: \(\mu\) > 1990_AverageBudget

Null Hypothesis: The average comedy movie budget in 2005 was not higher than the average comedy movie budget in 1990.

  • H0: \(\mu\) \(\le\) 1990_AverageBudget

Second Hypothesis Test, Graph:

Our hypothesis is proven correct: the average comedy movie budget in 2005 was higher than in 1990.

Conclusion:

These examples were simple to better demonstrate how to structure a hypothesis.

The best hypotheses teach us something new by stating answers to questions we don’t know the answer to. They give us structure to perform advanced statistical analysis, whether that’s through ANOVA tests, regression tests, or any of the many other statistical tests at our fingertips.

We have to know what question we’re answering before we can answer it.

Hypothesis testing defines our question, and our hypothesis guides our search for its answer.