Customize this page so that it becomes a useful Guide to Statistical Analysis for your current and future self.
If you want to learn more about customizing this table of contents, see the “R Help -> R Markdown Hints” page of your Math 325 Notebook.
Note: to edit this page click here and then
open the index.Rmd file that is highlighted.
1 Quantitative Variable
Types of tests to use:
\[ H_0: \mu = \text{some number} \] \[ H_a: \mu \ \left\{\underset{<}{\stackrel{>}{\neq}}\right\} \ \text{some number} \]
\[ H_0: \mu_d = \text{some number, but typically 0} \] \[ H_a: \mu_d \ \left\{\underset{<}{\stackrel{>}{\neq}}\right\} \ \text{some number, but typically 0} \]
\[ H_0: \text{median of differences} = 0 \] \[ H_a: \text{median of differences} \ \left\{\underset{<}{\stackrel{>}{\neq}}\right\} \ 0 \]
1 Quantitative Variable | 2 Groups
Both samples are representative of the population. The differences of pairs of observed. There are two groups being compared to one type of quantitative data type. There are two different medians.
Ex: Measuring height in men and women Men and women are the two groups.
Wilcoxon Rank Sum (Mann-Whitney) Test
\[ H_0: \text{difference in medians} = 0 \] \[ H_a: \text{difference in medians} \neq 0 \]
\[ H_0: \mu_1 - \mu_2 = \text{some number, but typically 0} \] \[ H_a: \mu_1 - \mu_2 \ \left\{\underset{<}{\stackrel{>}{\neq}}\right\} \ \text{some number, but typically 0} \]
Wilcoxon Rank Sum (Mann-Whitney) Test
\[ H_0: \text{difference in medians} = 0 \]
\[ H_a: \text{difference in medians } \neq 0 \] &
\[ H_0: \text{the distributions are stochastically equal} \]
\[ H_a: \text{one distribution is stochastically greater than the other} \]
1 Quantitative Variable | 3+ Groups
Three groups of quantitative data representating a population to being compared. There are many means.
\[ H_0: \alpha_1 = \alpha_2 = \ldots = 0 \]
\[ H_a: \alpha_i \neq 0 \ \text{for at least one} \ i \]
Hypotheses:
First Factor:
\[ H_0: \alpha_1 = \alpha_2 = \ldots = 0 \] \[ H_a: \alpha_i \neq 0 \ \text{for at least one} \ i \]
\[ H_0: \beta_1 = \beta_2 = \ldots = 0 \]
\[ H_a: \beta_j \neq 0 \ \text{for at least one} \ j \]
\[ H_0: \text{The effect of} \ \alpha \ \text{is the same for all levels of} \ \beta \]
\[ H_a: \text{The effect of} \ \alpha \ \text{is different for at least one level of} \ \beta \]
\[ H_0: \text{All samples are from the same distribution.} \] \[ H_a: \text{At least one sample's distribution is stochastically different.} \] OR do it stating the hypotheses in the style of an ANOVA Test:
\[ H_0: \mu_1 = \mu_2 = \ldots = \mu \]
\[ H_0: \mu_1 = \mu_2 = \ldots = \mu \] \[ H_a: \mu_i \neq \mu \ \text{for at least one} \ i \] Permutation Tests
Ex: The heights of children, teenagers, and adults being measured.
2 Quantitative Variables
A table where two quantitative samples are being compared against one another. Use a scatterplot to graph the data.
\[ \left.\begin{array}{ll} H_0: \beta_1 = 0 \\ H_a: \beta_1 \neq 0 \end{array} \right\} \ \text{Slope Hypotheses} \] \[ \left.\begin{array}{ll} H_0: \beta_0 = 0 \\ H_a: \beta_0 \neq 0 \end{array} \right\} \ \text{Intercept Hypotheses}^{\quad\text{(sometimes useful)}} \]
\[ Y_i = \beta_0 + \beta_1 X_{1} + \epsilon_i \]
Ex: Measuring height and foot lengths
1 Quantitative Response | Multiple Explanatory Variables
With multiple options, multiple quantitative data samples may be used to compare against each other.
** Follow the same hypotheses as a simple linear regression, but follow the different slopes looking at the model: **
\[ Y_i = \beta_0 + \beta_1 X_{1i} + \beta_2 X_{2i} + \cdots + \beta_p X_{pi} + \epsilon_i \]
ex: Height and foot length being compared against gender, age, etc.
Binomial Response | 1 Explanatory Variable
Based off of quantitive data, the qualitative data can be predicted
\[ H_0: \beta_1 = 0 \\ H_a: \beta_1 \neq 0 \]
ex: Using shoe size to predict male or female.
ex:
0 = failed/false
1 = success/true
ex: Y column may represent the success/fail of someone saying “bless you”” after sneezing.
Binomial Response | Multiple Explanatory Variables
Similar to above, but sometimes it is necessary to use more information to make a prediction.
0 = failed/false
1 = success/true
For example: Y column is fail of success while x1 column is someone saying thank you after holding the door open for them. x2 is qualitative data like if the person had brown hair or not.
2 Qualitative Variables
Comparing 2 qualities against one another. For example, names and birth months
\(H_0\) The row variable and column variable are independent.
\(H_a\): The row variable and column variable are associated.
Use a bar chart with this data table.
fav_stats(1:10) fav_stats(faithful$eruptions) favstats(Sepal.Length ~ Species, data=iris) # Note: this is favstats() rather than fav_stats()
Type 2 Error
If you want to put graphs on the sides of each other, then use
the command:
_______________________ | | |par (mfrow = c(1,2)))|
_______________________
as.factor for ANOVA test
Histogram, Boxplot, or Dotplot
Mean, median, five-number summary, standard deviation
gsub function to get rid of or swap out symbols/letters ex: choc\(Cocao <- as.numeric(gsub("%", "", choc\)Cocao)) Looking at the choc dataset (View(choc)).
Example of performing simple linear regression in R.
plot(Height ~ Volume, data=trees) trees.lm <- lm(Height ~ Volume, data=trees) abline(trees.lm)
par(mfrow=c(1,2)) plot(trees.lm, which=1:2)
par(mfrow = c(1,1)) #This resets your plotting window for future plots.
Organize by row: rbind(Apples=c(Good = 80, Bruised= 20, Rotten = 15))
Organize by column: cbind(Apples=c(Good = 80, Bruised= 20, Rotten = 15))
Independence: Everything acts equally
For example, you might:
write a quick note to tell yourself that the image on the left represents “height” or “weight” data.
You may include a link to dot plots, histograms, or standard deviation.
Or, you might highlight some text to show its importance, or just change the text color or text size or even all three.