Part 1

I decided on the Swiss data set.

(Swiss Fertility and Socioeconomic Indicators (1888) Data)

There are 6 variables, all are either numeric or integers.

data(swiss)
swiss_data <- swiss
str(swiss_data)
## 'data.frame':    47 obs. of  6 variables:
##  $ Fertility       : num  80.2 83.1 92.5 85.8 76.9 76.1 83.8 92.4 82.4 82.9 ...
##  $ Agriculture     : num  17 45.1 39.7 36.5 43.5 35.3 70.2 67.8 53.3 45.2 ...
##  $ Examination     : int  15 6 5 12 17 9 16 14 12 16 ...
##  $ Education       : int  12 9 5 7 15 7 7 8 7 13 ...
##  $ Catholic        : num  9.96 84.84 93.4 33.77 5.16 ...
##  $ Infant.Mortality: num  22.2 22.2 20.2 20.3 20.6 26.6 23.6 24.9 21 24.4 ...

I used the cor() function to get a matrix of all the variable correlations. Then made a variable to house the matrix so I could call on the largest value that isn’t 1.

cor(swiss_data)
##                   Fertility Agriculture Examination   Education   Catholic
## Fertility         1.0000000  0.35307918  -0.6458827 -0.66378886  0.4636847
## Agriculture       0.3530792  1.00000000  -0.6865422 -0.63952252  0.4010951
## Examination      -0.6458827 -0.68654221   1.0000000  0.69841530 -0.5727418
## Education        -0.6637889 -0.63952252   0.6984153  1.00000000 -0.1538589
## Catholic          0.4636847  0.40109505  -0.5727418 -0.15385892  1.0000000
## Infant.Mortality  0.4165560 -0.06085861  -0.1140216 -0.09932185  0.1754959
##                  Infant.Mortality
## Fertility              0.41655603
## Agriculture           -0.06085861
## Examination           -0.11402160
## Education             -0.09932185
## Catholic               0.17549591
## Infant.Mortality       1.00000000
max(cor(swiss_data))
## [1] 1
swiss_data_cor <- cor(swiss_data)
max(swiss_data_cor[swiss_data_cor < 1])
## [1] 0.6984153
max(abs(swiss_data_cor[swiss_data_cor < 1]))
## [1] 0.6984153

Independent Variable (X): Education: “% education beyond primary school for draftees”

Dependent Variable (Y): Examination. “% draftees receiving highest mark on army examination”

\[ \widehat{{Examination_i}} = {\beta_0\ + \beta_1(Education_i)}\ + \epsilon_i \]

Intercept 10.1275: Tells us the % of draftees we should expect to still recieve high marks on any exam, even if the X varaible of Education was at a value of 0. This is our starting point.

Education 0.5795: This value tells us the estimated rate of change the Y variable of Examination would experience if the X variable were to increase by 1 unit. Since the units of measurement is %, if we were to increase education percentage by 1%, we would expect to see Examination percentage to also go up about 0.58%.

The fact that it’s a positive value is positive relation between the 2 variables. And the fact that it is closer to 1 than 0 means that it can be considered to have a strong amount of correlation.

lm(Examination ~ Education, data = swiss_data)
## 
## Call:
## lm(formula = Examination ~ Education, data = swiss_data)
## 
## Coefficients:
## (Intercept)    Education  
##     10.1275       0.5795

To get the slope and intercept values, I needed to use the covariance, variance and mean functions and input them into the supplied formulas. Both the results are able to recreate the lm() function’s return.

$$ {_1} = \ {_0} = {y} - {_1}{x}

$$

swiss_edu_exam_slope <- cov(swiss_data$Education, swiss_data$Examination) / var(swiss_data$Education)
swiss_edu_exam_slope
## [1] 0.5794737
swiss_edu_exam_intercept <- mean(swiss_data$Examination) - (swiss_edu_exam_slope * mean(swiss_data$Education))
swiss_edu_exam_intercept  
## [1] 10.12748

Part 2

OLS, or Ordinary Least Squares, is a way to determine which linear regression line best fits the data at hand. Its overall purpose is to find the line that gets as close as possible to all the data points by minimizing the sum of the squared errors.

The linear regression line contains the intercept B0, slope B1 and the explanatory variable (X). Each point on the line represents the expected value of Y for a given value of X. What we expect and what we actually observe will usually have some discrepancy, called a residual.

\[ e_i = y_i - \hat{y}_i \]

Residuals are the vertical distance from observations to the regression line and can be positive or negative. OLS squares these residuals so they don’t cancel each other out and finds the line that minimizes their total.

OLS also relies on several assumptions.

Linearity assumes the relation between the variables X and Y can be represented by a linear model.

Independence assumes the observations and their errors aren’t influenced by one another.

Homoscedasticty assumes the residuals have roughly constant variance across all the values of X.

Lastly, Normality assumes the errors are approximately normally distributed.

Now Residual plots can be used to help in the recognition of potential patterns and ID’ing extreme outliers.

\[ R^2 = 1 - \frac{\sum_{i=1}^{n}(y_i-\hat{y}_i)^2} {\sum_{i=1}^{n}(y_i-\bar{y})^2} \]