This analysis examines Dodgers attendance using the DodgersData dataset. The analysis uses descriptive statistics, a scatter plot, and multiple linear regression.
DodgersData[25, c("temp", "attend", "opponent", "bobblehead")]
## temp attend opponent bobblehead
## 25 61 36561 Astros NO
The 25th home game had a temperature of 61 degrees, attendance of 36,561 people, the Astros as the opponent, and no bobblehead promotion.
median(DodgersData$attend)
## [1] 40284
The median attendance was 40,284 people.
sum(DodgersData$day_night == "Night")
## [1] 66
The Dodgers had 66 night games.
plot(DodgersData$temp,
DodgersData$attend / 1000,
main = "Dodgers Attendance by Temperature",
xlab = "Temperature (°F)",
ylab = "Attendance (Thousands)",
pch = 19,
col = "blue")
The scatter plot shows that temperature does not have a strong straight-line relationship with Dodgers attendance. Attendance varies widely at many temperatures, suggesting that other factors, such as promotions, opponents, month, and day of the week, may be more important.
model <- lm(attend ~ month + day_of_week + bobblehead,
data = DodgersData)
summary(model)
##
## Call:
## lm(formula = attend ~ month + day_of_week + bobblehead, data = DodgersData)
##
## Residuals:
## Min 1Q Median 3Q Max
## -10786.5 -3628.1 -516.1 2230.2 14351.0
##
## Coefficients:
## Estimate Std. Error t value Pr(>|t|)
## (Intercept) 38792.98 2364.68 16.405 < 2e-16 ***
## monthAUG 2377.92 2402.91 0.990 0.3259
## monthJUL 2849.83 2578.60 1.105 0.2730
## monthJUN 7163.23 2732.72 2.621 0.0108 *
## monthMAY -2385.62 2291.22 -1.041 0.3015
## monthOCT -662.67 4046.45 -0.164 0.8704
## monthSEP 29.03 2521.25 0.012 0.9908
## day_of_weekMonday -4883.82 2504.65 -1.950 0.0554 .
## day_of_weekSaturday 1488.24 2442.68 0.609 0.5444
## day_of_weekSunday 1840.18 2426.79 0.758 0.4509
## day_of_weekThursday -4108.45 3381.22 -1.215 0.2286
## day_of_weekTuesday 3027.68 2686.43 1.127 0.2638
## day_of_weekWednesday -2423.80 2485.46 -0.975 0.3330
## bobbleheadYES 10714.90 2419.52 4.429 3.59e-05 ***
## ---
## Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
##
## Residual standard error: 6120 on 67 degrees of freedom
## Multiple R-squared: 0.5444, Adjusted R-squared: 0.456
## F-statistic: 6.158 on 13 and 67 DF, p-value: 2.083e-07
The regression model explained approximately 54.44% of the variation in attendance. The bobblehead promotion coefficient was approximately 10,715, meaning attendance was predicted to be about 10,715 people higher during bobblehead games after accounting for month and day of the week. The bobblehead p-value was less than .001, so the relationship was statistically significant.
Customers will be more likely to purchase avocados when the average price is lower because lower prices reduce the perceived financial risk of purchasing the product.
Organic avocados will have a higher average price than conventional avocados because consumers may be willing to pay more for products perceived as healthier or higher quality.
At first I thought regression analysis was mainly used to find simple relationships between two variables. But now I think regression is more useful because it measures how multiple factors influence an outcome while holding other factors constant.
My question is: How can we tell whether a statistically significant relationship also creates a meaningful real-world business impact?