2024/1/6

Introduction to simple linear regression

Simple linear regression is a statistical method that allows us to estimate the relationship between two variables. A predictor variable and a response variable. In the context of this project, the predictor is the amount of attended office hours by student, and the response is their respective performance (grade) on a midterm.

The Linear Regression Model

The simple linear regression model is given by: \[ y = \beta_0 + \beta_1 x + \epsilon \]

Where: - \(y\) is the dependent variable (e.g., test scores) - \(x\) is the independent variable (e.g., attendance) - \(\beta_0\) is the intercept - \(\beta_1\) is the slope - \(\epsilon\) is the error term

The slope \(\beta_1\) represents the change in \(y\) for a one-unit change in \(x\).

The formula for the least squares estimates of \(\beta_0\) and \(\beta_1\) are: \[ \hat{\beta}_1 = \frac{\sum (x_i - \bar{x})(y_i - \bar{y})}{\sum (x_i - \bar{x})^2} \] \[ \hat{\beta}_0 = \bar{y} - \hat{\beta}_1 \bar{x} \]

The coefficient of determination (\(R^2\)) is calculated as: \[ R^2 = 1 - \frac{\sum (y_i - \hat{y}_i)^2}{\sum (y_i - \bar{y})^2} \]

Loading the data

Pre-Processing

scores_filtered=scores %>%
  mutate(full_name = paste(`First Name`, `Last Name`, sep = " "))

attendance_filtered=office_hours %>%
  mutate(full_name = paste(`First Name`, `Last Name`, sep = " "))

# Merge the data frames based on the 'full_name' column
combined_data=inner_join(scores_filtered, attendance_filtered, by = "full_name") %>%
  select(full_name, Total, `Total Score`)

# Categorize students based on attendance (0 to 6 hours)
combined_data=combined_data %>%
  mutate(attendance_category = cut(Total, breaks = c(-Inf, 3, 6), labels = c("0-3 hours", "4-6 hours")))

# Calculate average score for each attendance category
score_summary=combined_data %>%
  group_by(attendance_category) %>%
  summarise(average_score = mean(`Total Score`, na.rm = TRUE))

Distribution of Attendance count

This histogram allows me to visualize whether or not there was a certain bias towards high attendance or low attendance. These bin counts are very similar. This will make it easier to effectively dive into the performance of students who fall in one of these two categories.

Average Midterm Scores

Just Based on this histogram, there a clear difference between the high attendance students and the low attendance students. The students that attend more office hours, seem to outscore the students who attended less. In order to determine whether or not this is statistically significant, we can look into creating a linear model that estimates the relationship between these two variables.

Linear Regression

`geom_smooth()` using formula = 'y ~ x'

Is this significant

[1] "The p-value of the slope is:  0.28"

With a traditional significance level of 0.05, the relationship between a students attendance of office hours and their performance is not statistically significant. This does not necessarily mean that students do not benefit from attending these extra study hours. As we saw with the histogram, students who attended more office hours (in the high bin category) did have a higher test average. However, there are many factors that can predict a students performance on a midterm.

Statistical significant of Slope

To determine if the slope estimate (\(\hat{\beta}_1\)) is statistically significant, we perform a hypothesis test:

Hypotheses: - Null hypothesis (\(H_0\)): \(\beta_1 = 0\) (no relationship between \(x\) and \(y\)) - Alternative hypothesis (\(H_A\)): \(\beta_1 \neq 0\) (a relationship exists between \(x\) and \(y\))

Test Statistic: The test statistic for the slope is given by: \[ t = \frac{\hat{\beta}_1}{\text{SE}(\hat{\beta}_1)} \] where \(\text{SE}(\hat{\beta}_1)\) is the standard error of the slope estimate.

P-value: - The p-value indicates the probability of observing a test statistic as extreme as, or more extreme than, the observed value under the null hypothesis. - If the p-value is less than the chosen significance level (e.g., 0.05), we reject the null hypothesis, indicating that the slope is statistically significant.

Confidence Interval: - A 95% confidence interval for \(\beta_1\) can also be used to assess significance. - If the interval does not include 0, it indicates that the slope is significantly different from 0.

Since the p-value > 0.05, we conclude that there is not a statistically significant relationship between office hours attended and midterm test scores.