2024-06-09

Introduction to Linear Regression

Linear regression is a statistical method used to model relationships between dependent variables independent variables. Their main goal is to predict outcomes and understand the influence of predictor variables.

Mathematical Foundation

Linear regression is based on the equation:

\[ Y = \beta_0 + \beta_1X + \epsilon \]

  • Y: Dependent variable
  • X: Independent variable
  • β₀: Intercept of the regression line
  • β₁: Slope of the regression line
  • ε: Error term, representing the difference between observed and predicted values

Assumptions of Linear Regression

Linear regression relies on several key assumptions:

  • Linearity: The relationship between the independent and dependent variable has to be linear.
  • Independence: Observations should be independent of one another.
  • Homoscedasticity: The variance of residual is the same for any value of the independent variables.
  • Normality: For any fixed value of an independent variable, Y is normally distributed.

R Code for Linear Regression

The following slide demonstrates how to perform a simple linear regression analysis in R. I’ll use the mtcars dataset, and predict mpg (miles per gallon) as a function of wt (weight of the car in 1000 lbs).

# Load the mtcars dataset
data(mtcars)

# Fit linear regression model where mpg is predicted based on wt
model <- lm(mpg ~ wt, data = mtcars)

# Display a summary of the model to see coefficients and statistics
summary(model)
## 
## Call:
## lm(formula = mpg ~ wt, data = mtcars)
## 
## Residuals:
##     Min      1Q  Median      3Q     Max 
## -4.5432 -2.3647 -0.1252  1.4096  6.8727 
## 
## Coefficients:
##             Estimate Std. Error t value Pr(>|t|)    
## (Intercept)  37.2851     1.8776  19.858  < 2e-16 ***
## wt           -5.3445     0.5591  -9.559 1.29e-10 ***
## ---
## Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
## 
## Residual standard error: 3.046 on 30 degrees of freedom
## Multiple R-squared:  0.7528, Adjusted R-squared:  0.7446 
## F-statistic: 91.38 on 1 and 30 DF,  p-value: 1.294e-10

Additional Analysis with ggplot2

The following slide demonstrates the distribution of fuel efficiency across different cars in the mtcars dataset.

library(ggplot2)
ggplot(mtcars, aes(x = mpg)) +
  geom_histogram(binwidth = 2, fill = "skyblue", color = "black") +
  labs(title = "Histogram of MPG", x = "Miles Per Gallon (MPG)", y = "Frequency")

Interactive Visualization with plotly

You can convert the ggplot2 plot to an interactive plotly plot to enhance engagement. This allows the users to hover over points and see more detailed information.

library(plotly)

# Convert the ggplot2 object to a plotly object
p <- ggplot(mtcars, aes(x = wt, y = mpg)) +
  geom_point() +
  geom_smooth(method = "lm", se = FALSE, color = "lightblue")
  
# Use ggplotly to make the plot interactive
ggplotly(p)

Conclusion and Application

Simple linear regression is a powerful statistical tool used to predict an outcome based on a single predictor variable. It’s widely applicable in many fields including:

  • Economics: Predicting economic indicators based on factors like interest rates or consumer confidence.
  • Health Sciences: Estimating medical outcomes based on treatment variables.
  • Environmental Science: Forecasting pollution levels based on industrial activity metrics.
  • Market Research: Assessing the impact of marketing spend on sales growth.

This method provides valuable insights by quantifying the relationship between variables, making it essential for data-driven decision-making.