In this project I researched the Poisson distribution and how it applies to generalized linear models. Using this information, I will be looking to see players salary can predict player’s status as free agents and for arbitration. I will see if there is a significant relationship between free agent eligibility and a players salary. I will also see if there is a significant relationship between a players arbitration eligibility and their salary. If the results prove significant, I will predict the probability of free agent eligibility, given a specific salary.
This data set contains offensive statistics from 337 Major League Baseball players from the 1991 season. It also includes indicators of the player’s eligibility for free agency and arbitration and their salary in 1992. The population is players who played at least one game in both the 1991 and 1992 seasons, excluding pitchers.
This data set has 337 observations and 18 variables (13 quantitative, 5 categorical). Each data point has observations for all variables (no NA values). The data is observational and can be represented as count data over the course of a season. Because the data can be classified as count data, this allows the Poisson distribution to be used.
#read the final project data (imported as a dataset from Excel)
library(readxl)
FP_Data <- read_excel("~/Desktop/MATH 248 Final Project/Math 248 Final Project Data (2).xlsx")
head(FP_Data, n=3)
The variables of interest are Salary, FreeAgentElig, and ArbElig. Salary is a player’s salary in 1992, represented in thousands of dollars. FreeAgentElig represents a player’s free agent eligibility in the 1992 season. Free agency eligibility refers to a player not under contract with the ability to sign to other teams. This variable is binary categorical with 1 = eligible and 0 = ineligible. ArbElig represents a player’s arbitration eligibility in the 1992 season. Arbitration eligibility refers to a a player who has reached a specific amount of time in the league and who has the ability to negotiate their own salary. This variable is binary categorical with 1 = eligible and 0 = ineligible.
While we have mainly looked at data that is normally distributed, sometimes data does not behave in that manner. Another distribution that is used commonly is the Poisson distribution.
The Poisson distribution is a discrete probability distribution. It is used to determine the probability of events occurring during a fixed time frame or by using a fixed interval. This makes it useful for looking at count data, like the data set above that is set over a particular amount of time (the 1991 MLB season).
The Poisson distribution formula is represented below.
\[P(X=x) = \frac{-e^\lambda*\mu^x}{x!}\]
Small x is the exact value occurring during an interval of time. P(X=x) is the probability that the random variable, X, will hit the exact value, x, during the given time frame. \(\mu\) is the mean value over a fixed interval of time. The Poisson distribution will be helpful in accessing the probability of free agent eligibility during a season through it’s application in a generalized linear model.
We learned that generalized linear models are used when data does not follow a normal distribution. There are three parts to a generalized linear model; the linear predictor, the link function, and the variance function.
The linear predictor looks very similar to linear regression models discussed earlier and continues to represent the explanatory variables. It takes the form
\[\eta_i = \beta_0 + \beta_1X_{1i} + ... + \beta_pX_{pi}\]
The link function explains how the mean of the random variable, Y, is dependent on the linear predictor function above. It takes the form
\[g(\mu_i) = \eta_i\]
where \(\mu_i\) is the mean or expected value of the random variable, Y or \(E(Y_i)=\mu_i\).
The variance function describes how the variance of the random variable, Y, is dependent on the linear predictor function. It takes the form
\[Var(Y_i) = \phi V(\mu_i)\] where \(Var(Y_i)\) is the variance of the random variable, Y.
A generalized linear model using a Poisson random variable takes the form of
\[log(\mu) = \beta_0 + \beta_1X_1 + \beta_2X_2 + ... + \beta_kX_k + \text{ e}\]
where the linear predictor is
\[\eta_i = \beta_0 + \beta_1X_1 + \beta_2X_2 + ... + \beta_kX_k + \text{ e}\]
the variance function is
\[E(Y) = \mu\]
and the link function is
\[\log(\mu) = \eta_i\]
One unique assumption about the Poisson distribution is that it has equidispersion, which means that the variance and expected value of the random variable, Y, are equal.
In this project, we want to estimate the probability that a person will be free agent eligible given their salary. To get probability from the above model, we need to use the inverse link.
If \(\pi_i\) is the probability that a person will be free agent eligible, then,
\[\eta_i = \beta_0 + \beta_1X_{1i} = log(\frac{\pi_i}{1 - \pi_i})\] where
\[\pi_i = \frac{1}{1 + exp(\eta_i)} = \frac{1}{1+exp(-beta_0 - beta_1X_{1i})}\] This transformation is known as the inverse link function.
Given all of this model information, we can use it in the context of our problem.
To estimate the probability of free agent eligibility, I want to look at its relationship with two different variables, Salary and ArbElig (Arbitration Eligibility).
#dot plot of salary given free agent eligibility
histogram(~Salary|FreeAgentElig, data=FP_Data)
In the graph above, it appears that there are different salary distributions for players who are free agent eligible (left) and those who aren’t (right). It appears that players who are free agent eligible have a right skewed distribution with some extremely high salary values, but most under $1 million. Players who are not free agent eligible appear to have a wider and flatter spread of salaries. Both distributions appear to resemble Poisson disitributions, just with different average values.
#side-by-side histogram of salary given arbitration eligibility
histogram(~Salary|ArbElig, data=FP_Data)
In the histogram above, it is comparing the salary distributions of players who are eligible for arbitration (left) and those who are not (right). For players who are arbitration eligible, the salary distribution appears to be right-skewed and following some-what of a Poisson distribution. For players who are arbitration ineligible, the data appears to be right-skewed and peaking around $1.5 million. This distribution does not appear to follow the Poisson pattern, but it will be interesting to see if this variable has any affect on FreeAgentElig.
I will be comparing two models for FreeAgentElig. I will compare the amount of significant predictors in each, as well as the AIC to choose which model best explains the variation in FreeAgentElig. If there are any significant predictors, I will choose the preferred model and use it to predict FreeAgentElig probabilities given example salary amounts.
The first model is the model of FreeAgentElig given Salary, represented below.
\[log(FreeAgentElig) = \beta_0 + \beta_1(Salary) + e\] The second model is the model of FreeAgentElig given Salary and ArbElig, with the interaction between the two, represented below.
\[log(FreeAgentElig) = \beta_0 + \beta_1(Salary) + \beta_2(ArbElig) + \beta_3(Salary:ArbElig) + e\] We can generate the two models using r.
baseball_model = glm(FreeAgentElig~Salary, family=poisson(link = "log"), data=FP_Data)
summary(baseball_model)
##
## Call:
## glm(formula = FreeAgentElig ~ Salary, family = poisson(link = "log"),
## data = FP_Data)
##
## Coefficients:
## Estimate Std. Error z value Pr(>|z|)
## (Intercept) -1.624e+00 1.445e-01 -11.237 <2e-16 ***
## Salary 4.268e-04 5.495e-05 7.768 8e-15 ***
## ---
## Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
##
## (Dispersion parameter for poisson family taken to be 1)
##
## Null deviance: 247.16 on 336 degrees of freedom
## Residual deviance: 194.06 on 335 degrees of freedom
## AIC: 466.06
##
## Number of Fisher Scoring iterations: 5
For the first model (single predictor with Salary), Salary has a p-value of 8e-15 which is significant at all levels. This means that Salary has a significant relationship with FreeAgentElig. The AIC for this model is 466.06.
baseball_model_2 = glm(FreeAgentElig ~ Salary*ArbElig, family=poisson(link = "log"), data=FP_Data)
summary(baseball_model_2)
##
## Call:
## glm(formula = FreeAgentElig ~ Salary * ArbElig, family = poisson(link = "log"),
## data = FP_Data)
##
## Coefficients:
## Estimate Std. Error z value Pr(>|z|)
## (Intercept) -1.404e+00 1.417e-01 -9.914 < 2e-16 ***
## Salary 4.344e-04 5.323e-05 8.161 3.33e-16 ***
## ArbElig -1.790e+01 2.140e+03 -0.008 0.993
## Salary:ArbElig -4.344e-04 1.143e+00 0.000 1.000
## ---
## Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
##
## (Dispersion parameter for poisson family taken to be 1)
##
## Null deviance: 247.16 on 336 degrees of freedom
## Residual deviance: 130.84 on 333 degrees of freedom
## AIC: 406.84
##
## Number of Fisher Scoring iterations: 17
In the second model (double predictor with Salary and ArbELig, plus their interaction), Salary has a p-value of 3.33e-16 which is significant at all levels. ArbElig has a p-value of 0.993 and Salary:ArbElig has a p-value of 1.000. Both p-values are quite high, meaning that neither of those variables are significant predictors for FreeAgentElig at any level. The AIC for this model is 406.84.
We can see that the AIC for the second model is slightly lower than the first one, with 406.84 to 466.06 respectively. However, because ArbElig and the interaction between ArbElig and Salary are not significant at any level, they are not contributing to predicting the variability in FreeAgentElig. Therefore, the model I am selecting is the first model (single predictor with Salary) because all predictors are significant.
Now that we have chosen our preferred model, we can use it to predict the probability of a player being free agent eligible given their salary. For the first prediction, let’s start with a low amount, $100,000.
#predict the probability of free agent eligibility with a salary of $100,000
newdata = data.frame(Salary = 100)
predict.glm(baseball_model, newdata, type = "response")
## 1
## 0.2056523
The probability that a player with a salary of $100,000 is free agent eligible is 0.206.
Let’s increase to $250,000 and see how the probabilities change.
#predict the probability of free agent eligibility with a salary of $250,000
newdata_2 = data.frame(Salary = 250)
predict.glm(baseball_model, newdata_2, type = "response")
## 1
## 0.2192496
The probability that a player with a salary of $250,000 is free agent eligible is 0.219. The probability increased, but not by much at all.
Let’s increase the salary somewhat drastically to see its effect on the probability of free agent eligibility. Let’s look at a player who has a salary of $2,500,000.
#predict the probability of free agent eligibility with a salary of $2,500,000
newdata_3 = data.frame(Salary = 2500)
predict.glm(baseball_model, newdata_3, type = "response")
## 1
## 0.5728175
The probability that a player with a salary of $2,500,000 is free agent eligible is 0.573. As salary increased drastically, the probability that a player is free agent eligible also increased drastically.
Finally, let’s look at a player whose salary is $3,500,000.
#predict the probability of free agent eligibility with a salary of $3,500,000
newdata_3 = data.frame(Salary = 3500)
predict.glm(baseball_model, newdata_3, type = "response")
## 1
## 0.8777768
The probability that a player with a salary of $3,500,000 is free agent eligible is 0.878. The jump in probability was about 0.3 with the addition of $1 million. From $250,000 to $2,500,000, the jump in probability was about 0.36, but the increase in salary was $2,250,000. We can see that there were drastically different changes in salary from those two jumps, but the change in probability was almost the same. This is showing that as salaries are getting much higher, the increase in probability of free agent eligibility is also growing larger. All this to say, players with very large salaries are highly likely to be free agent eligible.
plot(x=FP_Data$Salary, y=baseball_model$fitted.values, xlab='Salary',
ylab='Probability of Free Agent Eligibility', ylim=c(-0.1, 1.1))
points(x=FP_Data$Salary, y=FP_Data$FreeAgentElig, pch=2, col=2)
In the graph above, we can see the pattern we were seeing in the predicted values. As salary increases, the probability of free agent eligibility is increasing and at an increasing rate.
I set out to investigate if there was a significant relationship between FreeAgentELig (free agent eligibility) and two predictors, Salary and ArbElig (arbitration eligibility). I determined that there was a significant relationship between FreeAgentElig and Salary, but not with ArbElig or the interaction between ArbElig and Salary. I also determined that there is a positive relationship between Salary and probability of FreeAgentElig, meaning that as salary increases, the probability of a player being free agent eligible also increases.
I wanted to determine if there was a significant relationship because it could give valuable insight for teams when forming their rosters. It is important information for organizations to know when they are accessing who they should sign to their team. If a player is free agent eligible, it is likely that their salary will be higher and be more of a cost to the organization. Using this information, teams can then do analysis on other aspects of the player’s game to access if they are worth it to bring on. It is the first step in doing analysis for sports teams and building their rosters and it is something that I would like to explore more.
In this project, I think a weakness was determining if this model was accurately predicted using the Poisson model. The distribution appeared to take the Poisson form and it was count data over a given time period (the 1991 season), however, I would have liked more confirmation that the distribution was indeed the best fit for the question I was attempting to answer. This is something I could investigate more in the future. As for the strengths of the project, I think I did a good job of thoroughly explaining both the Poisson distribution and it’s applications in generalized linear models. I also think I successfully applied it to a real-world example and was able to take away tangible results from it. It was a great introduction to sports analysis and only made me want to pursue it more.
Dataset found in Journal of Statistics Education Data Archive, https://jse.amstat.org/jse_data_archive.htm
Data sourced from CNN/Sports Illustrated at http://www.cnnsi.com/baseball/mlb/historical_profiles/
Sacramento Bee, 10/15/91.
The New York Times, 11/13/91, 2/23/92, 11/19/92.
The Society for American Baseball Research (SABR) at ftp://skypoint.com/pub/members/a/ashbury/sabr/SALARIES /1992_salaries_baseball
“9.2 - Modeling Count Data: Stat 504.” PennState: Statistics Online Courses, online.stat.psu.edu/stat504/lesson/9/9.2-0. Accessed 8 Dec. 2024.
Chugani, Vinod. “Poisson Distribution: A Comprehensive Guide.” DataCamp, DataCamp, 11 Sept. 2024, www.datacamp.com/tutorial/poisson-distribution.
“Free Agency: Glossary.” MLB.Com, www.mlb.com/glossary/transactions/free-agency. Accessed 8 Dec. 2024.
hy9fesh, et al. “Why Is the Coefficient for My Interaction Term in My Glm Equal to Na?” Stack Overflow, 1 Mar. 2023, stackoverflow.com/questions/77869640/why-is-the-coefficient-for-my-interaction-term-in-my-glm-equal-to-na.
Jabeen, Hafsa. “Learn to Use Poisson Regression in R.” Dataquest, 9 Apr. 2023, www.dataquest.io/blog/tutorial-poisson-regression-in-r/.
Lecture 13: GLM for Poisson Data, people.musc.edu/~bandyopd/bmtry711.11/lecture_13.pdf. Accessed 9 Dec. 2024.
“Poisson Regression.” Poisson Regression - an Overview | ScienceDirect Topics, www.sciencedirect.com/topics/psychology/ poisson-regression. Accessed 8 Dec. 2024.
Turney, Shaun. “Poisson Distributions: Definition, Formula & Examples.” Scribbr, 21 June 2023, www.scribbr.com/statistics/poisson-distribution/.