Most numbers in this assignment have been rounded somewhat.
library(ggpubr)
## Loading required package: ggplot2
library(broom)
library(ggplot2)
library(corrplot)
## corrplot 0.95 loaded
car <- mtcars
datasets::mtcars
## mpg cyl disp hp drat wt qsec vs am gear carb
## Mazda RX4 21.0 6 160.0 110 3.90 2.620 16.46 0 1 4 4
## Mazda RX4 Wag 21.0 6 160.0 110 3.90 2.875 17.02 0 1 4 4
## Datsun 710 22.8 4 108.0 93 3.85 2.320 18.61 1 1 4 1
## Hornet 4 Drive 21.4 6 258.0 110 3.08 3.215 19.44 1 0 3 1
## Hornet Sportabout 18.7 8 360.0 175 3.15 3.440 17.02 0 0 3 2
## Valiant 18.1 6 225.0 105 2.76 3.460 20.22 1 0 3 1
## Duster 360 14.3 8 360.0 245 3.21 3.570 15.84 0 0 3 4
## Merc 240D 24.4 4 146.7 62 3.69 3.190 20.00 1 0 4 2
## Merc 230 22.8 4 140.8 95 3.92 3.150 22.90 1 0 4 2
## Merc 280 19.2 6 167.6 123 3.92 3.440 18.30 1 0 4 4
## Merc 280C 17.8 6 167.6 123 3.92 3.440 18.90 1 0 4 4
## Merc 450SE 16.4 8 275.8 180 3.07 4.070 17.40 0 0 3 3
## Merc 450SL 17.3 8 275.8 180 3.07 3.730 17.60 0 0 3 3
## Merc 450SLC 15.2 8 275.8 180 3.07 3.780 18.00 0 0 3 3
## Cadillac Fleetwood 10.4 8 472.0 205 2.93 5.250 17.98 0 0 3 4
## Lincoln Continental 10.4 8 460.0 215 3.00 5.424 17.82 0 0 3 4
## Chrysler Imperial 14.7 8 440.0 230 3.23 5.345 17.42 0 0 3 4
## Fiat 128 32.4 4 78.7 66 4.08 2.200 19.47 1 1 4 1
## Honda Civic 30.4 4 75.7 52 4.93 1.615 18.52 1 1 4 2
## Toyota Corolla 33.9 4 71.1 65 4.22 1.835 19.90 1 1 4 1
## Toyota Corona 21.5 4 120.1 97 3.70 2.465 20.01 1 0 3 1
## Dodge Challenger 15.5 8 318.0 150 2.76 3.520 16.87 0 0 3 2
## AMC Javelin 15.2 8 304.0 150 3.15 3.435 17.30 0 0 3 2
## Camaro Z28 13.3 8 350.0 245 3.73 3.840 15.41 0 0 3 4
## Pontiac Firebird 19.2 8 400.0 175 3.08 3.845 17.05 0 0 3 2
## Fiat X1-9 27.3 4 79.0 66 4.08 1.935 18.90 1 1 4 1
## Porsche 914-2 26.0 4 120.3 91 4.43 2.140 16.70 0 1 5 2
## Lotus Europa 30.4 4 95.1 113 3.77 1.513 16.90 1 1 5 2
## Ford Pantera L 15.8 8 351.0 264 4.22 3.170 14.50 0 1 5 4
## Ferrari Dino 19.7 6 145.0 175 3.62 2.770 15.50 0 1 5 6
## Maserati Bora 15.0 8 301.0 335 3.54 3.570 14.60 0 1 5 8
## Volvo 142E 21.4 4 121.0 109 4.11 2.780 18.60 1 1 4 2
colSums(is.na(mtcars))
## mpg cyl disp hp drat wt qsec vs am gear carb
## 0 0 0 0 0 0 0 0 0 0 0
LR <- lm(mpg ~ wt + hp, data = mtcars)
summary(LR)
##
## Call:
## lm(formula = mpg ~ wt + hp, data = mtcars)
##
## Residuals:
## Min 1Q Median 3Q Max
## -3.941 -1.600 -0.182 1.050 5.854
##
## Coefficients:
## Estimate Std. Error t value Pr(>|t|)
## (Intercept) 37.22727 1.59879 23.285 < 2e-16 ***
## wt -3.87783 0.63273 -6.129 1.12e-06 ***
## hp -0.03177 0.00903 -3.519 0.00145 **
## ---
## Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
##
## Residual standard error: 2.593 on 29 degrees of freedom
## Multiple R-squared: 0.8268, Adjusted R-squared: 0.8148
## F-statistic: 69.21 on 2 and 29 DF, p-value: 9.109e-12
The intercept estimate or beta 0 will be 37.22 which is the average value of mpg when both wt and hp are 0. The estimates for both wt and hp will be -3.88 and -.03 respectively. The previously mentioned estimates are known as beta 1 x1 and beta 2 x2. The numbers in terms of columns can connect to it’s ability to showcase the mean difference. The standard error for the intercept estimate is 1.60 while the wt has .63 coupled with hp being .01. These numbers state how they showcase the standard deviation of the sample. These numbers in terms of standard deviation are also ideal. The intercept estimate contains a t value of 23.29 which allows for a distance of the standard error from the mean. The other numbers for the columns of hp and wt include for t value are -3.52 to -6.13. The p values for all of the terms are all below .05 which defines significance, they are all great predictors to use.
The columns also contain no negative values as well.
par(mfrow = c(2, 2))
plot(LR)
par(mfrow = c(1, 1))
When looking at the standard analysis of the 4 beginning graphs at first for linearity of the data, I can see how most of the fitted values for the residual graph are centered around 15 to 25. The line here is curved and most are spread out, thus confirming that the data is not normal. In the case for Q-Q Residuals or normality of residuals, most of the points do follow the line which is satisfactory due to it’s ability to be normal, note some points though are outliers and are apart from the trend. When looking at Scale-Location data through homogeneity of variance, the line is curved similar to the first graph yet the points will be scattered from the line, proof of high variance. In the graph of Residuals vs Leverage, the Cook’s distance is large through most of the points being focused around the .05 area. There are some points though that are outliers starting at .1 to almost .4.
lm_mse <- mean(residuals(LR)^2)
print(paste("Mean Squared Error for Linear Model:", round(lm_mse, 2)))
## [1] "Mean Squared Error for Linear Model: 6.1"
For the Mean Square Error for this assignment, it will be set at 6.1. This number tells us that it is quite high and as such it defines over-fitting.
summary(LR)$r.squared
## [1] 0.8267855
The information can be found inside of the beginning of the linear regression model such that the R^2 will be .8268, which is a proof of concept of being a larger number and a good sight.
LRint <- lm(mpg ~ wt * hp, data = mtcars)
summary(LRint)
##
## Call:
## lm(formula = mpg ~ wt * hp, data = mtcars)
##
## Residuals:
## Min 1Q Median 3Q Max
## -3.0632 -1.6491 -0.7362 1.4211 4.5513
##
## Coefficients:
## Estimate Std. Error t value Pr(>|t|)
## (Intercept) 49.80842 3.60516 13.816 5.01e-14 ***
## wt -8.21662 1.26971 -6.471 5.20e-07 ***
## hp -0.12010 0.02470 -4.863 4.04e-05 ***
## wt:hp 0.02785 0.00742 3.753 0.000811 ***
## ---
## Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
##
## Residual standard error: 2.153 on 28 degrees of freedom
## Multiple R-squared: 0.8848, Adjusted R-squared: 0.8724
## F-statistic: 71.66 on 3 and 28 DF, p-value: 2.981e-13
The estimate for the interaction is based through the value being .008. That is definitely less than .05 with being a sign to use it as a predictor.
summary(LRint)$r.squared
## [1] 0.8847637
The interaction term increased the traditional R^2 by .20 which is a positive indicator of including an interaction will result in more desirable results.
lower_bound_hp <- quantile(mtcars$hp, 0.05, na.rm = TRUE)
upper_bound_hp <- quantile(mtcars$hp, 0.95, na.rm = TRUE)
mtcars$hp[mtcars$hp < lower_bound_hp] <- lower_bound_hp
mtcars$hp[mtcars$hp > upper_bound_hp] <- upper_bound_hp
summary(mtcars[, "hp"])
## Min. 1st Qu. Median Mean 3rd Qu. Max.
## 63.65 96.50 123.00 144.23 180.00 253.55
LRlhp <- lm(mpg ~ wt + hp, data = mtcars)
summary(LRlhp)
##
## Call:
## lm(formula = mpg ~ wt + hp, data = mtcars)
##
## Residuals:
## Min 1Q Median 3Q Max
## -3.8825 -1.6545 -0.0968 0.8367 5.7259
##
## Coefficients:
## Estimate Std. Error t value Pr(>|t|)
## (Intercept) 37.31722 1.56964 23.774 < 2e-16 ***
## wt -3.58279 0.66427 -5.394 8.5e-06 ***
## hp -0.03952 0.01059 -3.732 0.000824 ***
## ---
## Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
##
## Residual standard error: 2.546 on 29 degrees of freedom
## Multiple R-squared: 0.833, Adjusted R-squared: 0.8215
## F-statistic: 72.34 on 2 and 29 DF, p-value: 5.348e-12
summary(LRlhp)$r.squared
## [1] 0.8330309
You assign a singular winsorized hp variable based upon it’s adaptation to the percents of the above step. When this is done, your beta 0 of the intercept estimate will now be 37.32. While both the beta 1 x1 and beta 2 x2 are -3.58 to -.04. The values stated here were for wt and hp. Intercept increased by .10 while beta 1 x1 decreased compared to beta 2 x2 increasing. In addition to this information, the R^2 will be .833, also by a .01 increase. For the formula in terms of the bounds, each bound will be set to the bound either greater than or less than the bound in question.
car::vif(LR)
## wt hp
## 1.766625 1.766625
For the beta 1 x1 and beta 2 x2, they have the same colinearity. Along with this fact, they also have low values which makes them usable.
An R^2 will likely not turn a decision based on other factors that could impact predictability such as coefficients, p values, and residual graphs.