cars_data <- datasets::cars
str(cars_data)
## 'data.frame': 50 obs. of 2 variables:
## $ speed: num 4 4 7 7 8 9 10 10 10 11 ...
## $ dist : num 2 10 4 22 16 10 18 26 34 17 ...
summary(cars_data)
## speed dist
## Min. : 4.0 Min. : 2.00
## 1st Qu.:12.0 1st Qu.: 26.00
## Median :15.0 Median : 36.00
## Mean :15.4 Mean : 42.98
## 3rd Qu.:19.0 3rd Qu.: 56.00
## Max. :25.0 Max. :120.00
Type put your estimating equation i.e. I am expecting to see subscripts i on your y, x and error term professionally done.
Make sure to describe these two variables.
The independent variable in this dataset is $speed, as in the speed of a car traveling at “x” miles per hour, while $dist, the stopping distance measured in feet, is the dependent variable. Speed is the independent variable because it occurs before the stopping distance is observed. In this relationship, we are looking to determine whether a car’s speed is an accurate predictor of the braking distance required for the car to come to a stop.
\[dist_{i} = \beta_{0} + \beta_{1}speed_{i} + \epsilon_{i}\]
linear_regression <- lm(dist ~ speed, data = cars_data)
summary(linear_regression)
##
## Call:
## lm(formula = dist ~ speed, data = cars_data)
##
## Residuals:
## Min 1Q Median 3Q Max
## -29.069 -9.525 -2.272 9.215 43.201
##
## Coefficients:
## Estimate Std. Error t value Pr(>|t|)
## (Intercept) -17.5791 6.7584 -2.601 0.0123 *
## speed 3.9324 0.4155 9.464 1.49e-12 ***
## ---
## Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
##
## Residual standard error: 15.38 on 48 degrees of freedom
## Multiple R-squared: 0.6511, Adjusted R-squared: 0.6438
## F-statistic: 89.57 on 1 and 48 DF, p-value: 1.49e-12
\[dist_{i} = -17.5791 + 3.9324(speed_{i})\]
The estimated slope parameter (\(\beta_{1}\)) is 3.9324. This means that for every 1 mph increase in speed, the predicted stopping distance increases by 3.9324 feet on average. Since the slope is positive, it is indicated that higher speeds are associated with longer stopping distances.
The estimated intercept parameter (\(\beta_{0}\)) is -17.5791. According to the regression equation, this means that when the car’s speed is 0 mph, the estimated stopping distance is -17.5791 feet. However, since a negative stopping distance is not possible, the intercept is just a mathematical component of the equation and has no meaningful predictive value.
cov_var_slope <- cov(cars$speed, cars$dist) / var(cars$speed)
paste("The slope (\u03b2\u2081) is", signif(cov_var_slope, digits = 5))
## [1] "The slope (β₁) is 3.9324"
cov_ver_intercept <- mean(cars$dist) - (cov(cars$speed, cars$dist) / var(cars$speed)) * mean(cars$speed)
paste("The intercept (\u03b2\u2080) is", signif(cov_ver_intercept, digits = 6))
## [1] "The intercept (β₀) is -17.5791"
We have not gone through these in our lecture yet, but these are the conditions under which OLS is BLUE. You can even refer to any standard Econometrics textbooks or online resources.
Ordinary Least Squares (OLS) is a method used to determine the best-fitting line for a linear regression model by calculating the smallest possible sum of the squared vertical distances (residuals) between every data point and the line. The slope of the OLS line is calculated by the equation \(b_{1} = \frac{s_{y}}{s_{x}}R\), where R is the correlation between the two variables and \(s_{x}\) and \(s_{y}\) are the standard deviations of the explanatory and response variable, respectively. The OLS line passes through the mean point (\(\bar{x}, \bar{y}\)), which is used to calculate the intercept using the point-slope equation (\(b_{0} = \bar{y} - b_{1}\bar{x}\)). The Gauss-Markov Assumptions are conditions under which OLS gives the Best Linear Unbiased Estimator (BLUE). These assumptions include linearity, independent observations, exogeneity, constant variability, and lack of autocorrelation. The assumption of linearity requires the data to obey a linear trend; other, advanced regression models should be applied to non-linear trends. The assumption of independent observations states that the value of one data point should not be influenced by any other data point in the model. The assumption of exogeneity states that the errors within the model have zero conditional mean given the predictors. The assumption of constant variability, or homoscedasticity, requires that the spread of the residuals stays roughly constant across fitted values. Lastly, the assumption of no autocorrelation requires that errors from different observations are uncorrelated. When an OLS line fails to satisfy these assumptions, the resulting linear regression model can yield biased or unreliable estimates. Specifically, violating linearity or exogeneity can bias the slope or intercept coefficients. Violating constant variance or independence leaves OLS unbiased but no longer the best fit, and the standard errors become wrong. In order to test the linear regression model against the Gauss-Markov Assumptions, we use a few diagnostic plots to check the behavior of the model’s residuals. Linearity and exogeneity are tested using a plot with the residuals on the vertical axis and the fitted values on the horizontal axis. In this residuals-versus-fitted plot, we are looking for a random, evenly distributed cloud of points centered horizontally around the zero line. A U-shape, curve, or wave pattern means that these assumptions have been violated. To test the homoscedasticity of the model, the square root of the residuals are plotted against the fitted values. This scale-location plot is looking for a random, uniform spread of points around a horizontal reference line. A widening or narrowing spread indicates heteroscedasticity and a violation of this assumption. The assumption of independent observations is tested using a residuals-versus-leverage plot. This plot looks for points sitting outside of the Cook’s distance lines, which indicate observations with values that greatly influence on the OLS and violate this assumption.