Part 1: Data Exploration (to be ignored due to previous assignment yet load data for future steps).

Most numbers in this assignment have been rounded somewhat.

library(ggpubr)
## Loading required package: ggplot2
library(broom)
library(ggplot2)
library(corrplot)
## corrplot 0.95 loaded
car <- mtcars
datasets::mtcars
##                      mpg cyl  disp  hp drat    wt  qsec vs am gear carb
## Mazda RX4           21.0   6 160.0 110 3.90 2.620 16.46  0  1    4    4
## Mazda RX4 Wag       21.0   6 160.0 110 3.90 2.875 17.02  0  1    4    4
## Datsun 710          22.8   4 108.0  93 3.85 2.320 18.61  1  1    4    1
## Hornet 4 Drive      21.4   6 258.0 110 3.08 3.215 19.44  1  0    3    1
## Hornet Sportabout   18.7   8 360.0 175 3.15 3.440 17.02  0  0    3    2
## Valiant             18.1   6 225.0 105 2.76 3.460 20.22  1  0    3    1
## Duster 360          14.3   8 360.0 245 3.21 3.570 15.84  0  0    3    4
## Merc 240D           24.4   4 146.7  62 3.69 3.190 20.00  1  0    4    2
## Merc 230            22.8   4 140.8  95 3.92 3.150 22.90  1  0    4    2
## Merc 280            19.2   6 167.6 123 3.92 3.440 18.30  1  0    4    4
## Merc 280C           17.8   6 167.6 123 3.92 3.440 18.90  1  0    4    4
## Merc 450SE          16.4   8 275.8 180 3.07 4.070 17.40  0  0    3    3
## Merc 450SL          17.3   8 275.8 180 3.07 3.730 17.60  0  0    3    3
## Merc 450SLC         15.2   8 275.8 180 3.07 3.780 18.00  0  0    3    3
## Cadillac Fleetwood  10.4   8 472.0 205 2.93 5.250 17.98  0  0    3    4
## Lincoln Continental 10.4   8 460.0 215 3.00 5.424 17.82  0  0    3    4
## Chrysler Imperial   14.7   8 440.0 230 3.23 5.345 17.42  0  0    3    4
## Fiat 128            32.4   4  78.7  66 4.08 2.200 19.47  1  1    4    1
## Honda Civic         30.4   4  75.7  52 4.93 1.615 18.52  1  1    4    2
## Toyota Corolla      33.9   4  71.1  65 4.22 1.835 19.90  1  1    4    1
## Toyota Corona       21.5   4 120.1  97 3.70 2.465 20.01  1  0    3    1
## Dodge Challenger    15.5   8 318.0 150 2.76 3.520 16.87  0  0    3    2
## AMC Javelin         15.2   8 304.0 150 3.15 3.435 17.30  0  0    3    2
## Camaro Z28          13.3   8 350.0 245 3.73 3.840 15.41  0  0    3    4
## Pontiac Firebird    19.2   8 400.0 175 3.08 3.845 17.05  0  0    3    2
## Fiat X1-9           27.3   4  79.0  66 4.08 1.935 18.90  1  1    4    1
## Porsche 914-2       26.0   4 120.3  91 4.43 2.140 16.70  0  1    5    2
## Lotus Europa        30.4   4  95.1 113 3.77 1.513 16.90  1  1    5    2
## Ford Pantera L      15.8   8 351.0 264 4.22 3.170 14.50  0  1    5    4
## Ferrari Dino        19.7   6 145.0 175 3.62 2.770 15.50  0  1    5    6
## Maserati Bora       15.0   8 301.0 335 3.54 3.570 14.60  0  1    5    8
## Volvo 142E          21.4   4 121.0 109 4.11 2.780 18.60  1  1    4    2
colSums(is.na(mtcars))
##  mpg  cyl disp   hp drat   wt qsec   vs   am gear carb 
##    0    0    0    0    0    0    0    0    0    0    0

Part 2: Using lm through Linear Regression (note)

1. Build a Linear Regression model through predicting mpg solely based from wt and hp.

LR <- lm(mpg ~ wt + hp, data = mtcars)
summary(LR)
## 
## Call:
## lm(formula = mpg ~ wt + hp, data = mtcars)
## 
## Residuals:
##    Min     1Q Median     3Q    Max 
## -3.941 -1.600 -0.182  1.050  5.854 
## 
## Coefficients:
##             Estimate Std. Error t value Pr(>|t|)    
## (Intercept) 37.22727    1.59879  23.285  < 2e-16 ***
## wt          -3.87783    0.63273  -6.129 1.12e-06 ***
## hp          -0.03177    0.00903  -3.519  0.00145 ** 
## ---
## Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
## 
## Residual standard error: 2.593 on 29 degrees of freedom
## Multiple R-squared:  0.8268, Adjusted R-squared:  0.8148 
## F-statistic: 69.21 on 2 and 29 DF,  p-value: 9.109e-12

2.Interpret these coefficients.

The intercept estimate or beta 0 will be 37.22 which is the average value of mpg when both wt and hp are 0. The estimates for both wt and hp will be -3.88 and -.03 respectively. The previously mentioned estimates are known as beta 1 x1 and beta 2 x2. The numbers in terms of columns can connect to it’s ability to showcase the mean difference. The standard error for the intercept estimate is 1.60 while the wt has .63 coupled with hp being .01. These numbers state how they showcase the standard deviation of the sample. These numbers in terms of standard deviation are also ideal. The intercept estimate contains a t value of 23.29 which allows for a distance of the standard error from the mean. The other numbers for the columns of hp and wt include for t value are -3.52 to -6.13. The p values for all of the terms are all below .05 which defines significance, they are all great predictors to use.

The columns also contain no negative values as well.

3. Find the assumptions when using linear regression and if they have been met here. This can be found by using diagnostic plots and discuss the findings. (note)

par(mfrow = c(2, 2))
plot(LR)

par(mfrow = c(1, 1))

When looking at the standard analysis of the 4 beginning graphs at first for linearity of the data, I can see how most of the fitted values for the residual graph are centered around 15 to 25. The line here is curved and most are spread out, thus confirming that the data is not normal. In the case for Q-Q Residuals or normality of residuals, most of the points do follow the line which is satisfactory due to it’s ability to be normal, note some points though are outliers and are apart from the trend. When looking at Scale-Location data through homogeneity of variance, the line is curved similar to the first graph yet the points will be scattered from the line, proof of high variance. In the graph of Residuals vs Leverage, the Cook’s distance is large through most of the points being focused around the .05 area. There are some points though that are outliers starting at .1 to almost .4.

4. Report and understand the MSE.

lm_mse <- mean(residuals(LR)^2)

print(paste("Mean Squared Error for Linear Model:", round(lm_mse, 2)))
## [1] "Mean Squared Error for Linear Model: 6.1"

For the Mean Square Error for this assignment, it will be set at 6.1. This number tells us that it is quite high and as such it defines over-fitting.

5. Report and understand the R^2

summary(LR)$r.squared
## [1] 0.8267855

The information can be found inside of the beginning of the linear regression model such that the R^2 will be .8268, which is a proof of concept of being a larger number and a good sight.

6. Add an interaction that is inside both wt and hp for the linear regression model in part 2. (note)

LRint <- lm(mpg ~ wt * hp, data = mtcars)
summary(LRint)
## 
## Call:
## lm(formula = mpg ~ wt * hp, data = mtcars)
## 
## Residuals:
##     Min      1Q  Median      3Q     Max 
## -3.0632 -1.6491 -0.7362  1.4211  4.5513 
## 
## Coefficients:
##             Estimate Std. Error t value Pr(>|t|)    
## (Intercept) 49.80842    3.60516  13.816 5.01e-14 ***
## wt          -8.21662    1.26971  -6.471 5.20e-07 ***
## hp          -0.12010    0.02470  -4.863 4.04e-05 ***
## wt:hp        0.02785    0.00742   3.753 0.000811 ***
## ---
## Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
## 
## Residual standard error: 2.153 on 28 degrees of freedom
## Multiple R-squared:  0.8848, Adjusted R-squared:  0.8724 
## F-statistic: 71.66 on 3 and 28 DF,  p-value: 2.981e-13

Report the interaction significance for .05. (note)

The estimate for the interaction is based through the value being .008. That is definitely less than .05 with being a sign to use it as a predictor.

Confirm the interaction term can have an impact the model’s ability to handle for R^2 and how can you discuss the matter. (ask for final part)

summary(LRint)$r.squared
## [1] 0.8847637

The interaction term increased the traditional R^2 by .20 which is a positive indicator of including an interaction will result in more desirable results.

7. Suppose that hp has some outliers and our goal is to set winsorization to 5% and 95%.

lower_bound_hp <- quantile(mtcars$hp, 0.05, na.rm = TRUE)
upper_bound_hp <- quantile(mtcars$hp, 0.95, na.rm = TRUE)


mtcars$hp[mtcars$hp < lower_bound_hp] <- lower_bound_hp
mtcars$hp[mtcars$hp > upper_bound_hp] <- upper_bound_hp

summary(mtcars[, "hp"])
##    Min. 1st Qu.  Median    Mean 3rd Qu.    Max. 
##   63.65   96.50  123.00  144.23  180.00  253.55

7. Add a new linear regression model for the winsorized hp along with the baseline wt. Ignore any interaction terms and compare by the newly fitted model. Find the differences for R^2 with the coefficients.

LRlhp <- lm(mpg ~ wt + hp, data = mtcars)
summary(LRlhp)
## 
## Call:
## lm(formula = mpg ~ wt + hp, data = mtcars)
## 
## Residuals:
##     Min      1Q  Median      3Q     Max 
## -3.8825 -1.6545 -0.0968  0.8367  5.7259 
## 
## Coefficients:
##             Estimate Std. Error t value Pr(>|t|)    
## (Intercept) 37.31722    1.56964  23.774  < 2e-16 ***
## wt          -3.58279    0.66427  -5.394  8.5e-06 ***
## hp          -0.03952    0.01059  -3.732 0.000824 ***
## ---
## Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
## 
## Residual standard error: 2.546 on 29 degrees of freedom
## Multiple R-squared:  0.833,  Adjusted R-squared:  0.8215 
## F-statistic: 72.34 on 2 and 29 DF,  p-value: 5.348e-12
summary(LRlhp)$r.squared
## [1] 0.8330309

You assign a singular winsorized hp variable based upon it’s adaptation to the percents of the above step. When this is done, your beta 0 of the intercept estimate will now be 37.32. While both the beta 1 x1 and beta 2 x2 are -3.58 to -.04. The values stated here were for wt and hp. Intercept increased by .10 while beta 1 x1 decreased compared to beta 2 x2 increasing. In addition to this information, the R^2 will be .833, also by a .01 increase. For the formula in terms of the bounds, each bound will be set to the bound either greater than or less than the bound in question.

8. Look for multicolinearity through vif() for the model for part 2, what is the analysis?

car::vif(LR)
##       wt       hp 
## 1.766625 1.766625

For the beta 1 x1 and beta 2 x2, they have the same colinearity. Along with this fact, they also have low values which makes them usable.

9. How or how not with an edited R^2 increase predictability in the model?

An R^2 will likely not turn a decision based on other factors that could impact predictability such as coefficients, p values, and residual graphs.