At first, we want to install some needed libraries:
library(car)
library(lmtest)
We got the data from Kaggle, the dataset name is Human Resources Data Set, and the variables are:
Employee Name Employee’s full name
EmpID Employee ID is unique to each employee
MarriedID Is the person married (1 or 0 for yes or no)
MaritalStatusID Marital status code that matches the text field MaritalDesc
EmpStatusID Employment status code that matches text field EmploymentStatus
DeptID Department ID code that matches the department the employee works in
PerfScoreID Performance Score code that matches the employee’s most recent performance score
FromDiversityJobFairID Was the employee sourced from the Diversity job fair? 1 or 0 for yes or no
Salary The person’s yearly salary. $ U.S. Dollars
Termd Has this employee been terminated - 1 or 0
PositionID An integer indicating the person’s position
Position The text name/title of the position the person has
State The state that the person lives in
Zip The zip code for the employee
DOB Date of Birth for the employee
Sex Sex - M or F
MaritalDesc The marital status of the person (divorced, single, widowed, separated, etc)
CitizenDesc Label for whether the person is a Citizen or Eligible NonCitizen
HispanicLatino Yes or No field for whether the employee is Hispanic/Latino
RaceDesc Description/text of the race the person identifies with
DateofHire Date the person was hired
DateofTermination Date the person was terminated, only populated if, in fact, Termd = 1
TermReason A text reason / description for why the person was terminated
EmploymentStatus A description/category of the person’s employment status. Anyone currently working full time = Active
Department Name of the department that the person works in
ManagerName The name of the person’s immediate manager
ManagerID A unique identifier for each manager.
RecruitmentSource The name of the recruitment source where the employee was recruited from
PerformanceScore Performance Score text/category (Fully Meets, Partially Meets, PIP, Exceeds)
EngagementSurvey Results from the last engagement survey, managed by our external partner
EmpSatisfaction A basic satisfaction score between 1 and 5, as reported on a recent employee satisfaction survey
SpecialProjectsCount The number of special projects that the employee worked on during the last 6 months Integer
LastPerformanceReviewDate The most recent date of the person’s last performance review.
DaysLateLast30 The number of times that the employee was late to work during the last 30 days
Absences The number of times the employee was absent from work.
We made some statistical summaries:
adot <- read.csv2("hrdataset.csv", sep = ",", dec = ".", row.names = NULL)
adot
summary(adot)
row.names Employee_Name EmpID MarriedID
Length:311 Length:311 Min. :10001 Min. :0.0000
Class :character Class :character 1st Qu.:10078 1st Qu.:0.0000
Mode :character Mode :character Median :10155 Median :0.0000
Mean :10155 Mean :0.3974
3rd Qu.:10232 3rd Qu.:1.0000
Max. :10311 Max. :1.0000
NA's :4 NA's :4
MaritalStatusID GenderID EmpStatusID DeptID
Min. :0.0000 Min. :0.0000 Min. :1.000 Min. :1.000
1st Qu.:0.0000 1st Qu.:0.0000 1st Qu.:1.000 1st Qu.:5.000
Median :1.0000 Median :0.0000 Median :1.000 Median :5.000
Mean :0.8143 Mean :0.4365 Mean :2.381 Mean :4.632
3rd Qu.:1.0000 3rd Qu.:1.0000 3rd Qu.:5.000 3rd Qu.:5.000
Max. :4.0000 Max. :1.0000 Max. :5.000 Max. :6.000
NA's :4 NA's :4 NA's :4 NA's :4
PerfScoreID FromDiversityJobFairID Salary Termd
Min. :1.00 Min. :0.00000 Min. : 45046 Min. :0.0000
1st Qu.:3.00 1st Qu.:0.00000 1st Qu.: 55502 1st Qu.:0.0000
Median :3.00 Median :0.00000 Median : 62810 Median :0.0000
Mean :2.98 Mean :0.09446 Mean : 68820 Mean :0.3257
3rd Qu.:3.00 3rd Qu.:0.00000 3rd Qu.: 71913 3rd Qu.:1.0000
Max. :4.00 Max. :1.00000 Max. :250000 Max. :1.0000
NA's :4 NA's :4 NA's :4 NA's :4
PositionID Position State Zip
Min. : 1.00 Length:311 Length:311 Min. : 1013
1st Qu.:18.00 Class :character Class :character 1st Qu.: 1896
Median :19.00 Mode :character Mode :character Median : 2132
Mean :16.94 Mean : 6613
3rd Qu.:20.00 3rd Qu.: 2355
Max. :30.00 Max. :98052
NA's :4 NA's :4
DOB Sex MaritalDesc CitizenDesc
Length:311 Length:311 Length:311 Length:311
Class :character Class :character Class :character Class :character
Mode :character Mode :character Mode :character Mode :character
HispanicLatino RaceDesc DateofHire DateofTermination
Length:311 Length:311 Length:311 Length:311
Class :character Class :character Class :character Class :character
Mode :character Mode :character Mode :character Mode :character
TermReason EmploymentStatus Department ManagerName
Length:311 Length:311 Length:311 Length:311
Class :character Class :character Class :character Class :character
Mode :character Mode :character Mode :character Mode :character
ManagerID RecruitmentSource PerformanceScore EngagementSurvey
Min. : 1.00 Length:311 Length:311 Min. :1.120
1st Qu.:10.00 Class :character Class :character 1st Qu.:3.695
Median :15.00 Mode :character Mode :character Median :4.280
Mean :14.68 Mean :4.116
3rd Qu.:19.00 3rd Qu.:4.700
Max. :39.00 Max. :5.000
NA's :12 NA's :4
EmpSatisfaction SpecialProjectsCount LastPerformanceReview_Date
Min. :1.000 Min. :0.000 Length:311
1st Qu.:3.000 1st Qu.:0.000 Class :character
Median :4.000 Median :0.000 Mode :character
Mean :3.899 Mean :1.186
3rd Qu.:5.000 3rd Qu.:0.000
Max. :5.000 Max. :8.000
NA's :4 NA's :4
DaysLateLast30 Absences
Min. :0.0000 Min. : 1.00
1st Qu.:0.0000 1st Qu.: 5.00
Median :0.0000 Median :10.00
Mean :0.4039 Mean :10.23
3rd Qu.:0.0000 3rd Qu.:15.00
Max. :6.0000 Max. :20.00
NA's :4 NA's :4
The linear regression model investigates how employees’ salaries are influenced by their satisfaction level, the number of special projects, absences time, and gender. We use one company’s HR database (adot), where Salary is the dependent variable, and satisfaction level (EmpSatisfaction), number of special projects (SpecialProjectsCount), absences (Absences), and gender (GenderID) are the independent variables.
The linear regression formula is: Salary=β0+β1EmpSatisfaction+β2SpecialProjectsCount+β3Absences+β4GenderID+ε
aaares <- lm(Salary ~ EmpSatisfaction + SpecialProjectsCount + Absences + GenderID, data = adot)
summary(aaares)
Call:
lm(formula = Salary ~ EmpSatisfaction + SpecialProjectsCount +
Absences + GenderID, data = adot)
Residuals:
Min 1Q Median 3Q Max
-48200 -8963 -1884 4538 188809
Coefficients:
Estimate Std. Error t value Pr(>|t|)
(Intercept) 54365.0 5871.3 9.259 <2e-16 ***
EmpSatisfaction 1004.3 1364.8 0.736 0.4624
SpecialProjectsCount 5384.6 533.4 10.094 <2e-16 ***
Absences 381.3 211.5 1.803 0.0725 .
GenderID 583.5 2501.3 0.233 0.8157
---
Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1
Residual standard error: 21630 on 302 degrees of freedom
(4 observations deleted due to missingness)
Multiple R-squared: 0.2626, Adjusted R-squared: 0.2528
F-statistic: 26.88 on 4 and 302 DF, p-value: < 2.2e-16
plot(aaares)
This summary presents the results of a linear regression model fitted to predict salaries based on employee satisfaction (EmpSatisfaction), special projects count (SpecialProjectsCount), and absences (Absences).
The coefficients indicate that for every unit increase in EmpSatisfaction, SpecialProjectsCount, and Absences, there are corresponding increases in estimated salary, although only SpecialProjectsCount is statistically significant (p < 0.001).
The R-squared value suggests that approximately 26.26% of the variance in salaries is explained by these variables.
The F-statistic’s low p-value indicates that the model as a whole is statistically significant in predicting salaries.
We eliminate the GenderID variable, as it is the least significant.
The new linear regression formula is: Salary=β0+β1EmpSatisfaction+β2SpecialProjectsCount+β3Absences+ε
aaares <- lm(Salary ~ EmpSatisfaction + SpecialProjectsCount + Absences, data = adot)
summary(aaares)
Call:
lm(formula = Salary ~ EmpSatisfaction + SpecialProjectsCount +
Absences, data = adot)
Residuals:
Min 1Q Median 3Q Max
-47907 -9086 -2093 4527 188553
Coefficients:
Estimate Std. Error t value Pr(>|t|)
(Intercept) 54668.1 5716.7 9.563 <2e-16 ***
EmpSatisfaction 988.1 1360.9 0.726 0.4684
SpecialProjectsCount 5395.5 530.6 10.169 <2e-16 ***
Absences 381.5 211.2 1.806 0.0719 .
---
Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1
Residual standard error: 21600 on 303 degrees of freedom
(4 observations deleted due to missingness)
Multiple R-squared: 0.2624, Adjusted R-squared: 0.2551
F-statistic: 35.94 on 3 and 303 DF, p-value: < 2.2e-16
plot(aaares)
The model explains approximately 26.24% of the variance in Salary.
The F-statistic is significant (p < 0.05), indicating that the overall model is statistically significant.
Overall, SpecialProjectsCount seems to have a significant positive impact on Salary, while EmpSatisfaction and Absences don’t appear to have a significant effect.
Now EmpSatisfaction is the least statistically significant variable, we exclude it.
The new formula is: Salary=β0+β1SpecialProjectsCount+β2Absences+ε
aaares <- lm(Salary ~ SpecialProjectsCount + Absences, data = adot)
summary(aaares)
Call:
lm(formula = Salary ~ SpecialProjectsCount + Absences, data = adot)
Residuals:
Min 1Q Median 3Q Max
-46973 -8869 -2001 4541 187688
Coefficients:
Estimate Std. Error t value Pr(>|t|)
(Intercept) 58376.9 2564.8 22.76 <2e-16 ***
SpecialProjectsCount 5413.0 529.6 10.22 <2e-16 ***
Absences 393.5 210.4 1.87 0.0624 .
---
Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1
Residual standard error: 21580 on 304 degrees of freedom
(4 observations deleted due to missingness)
Multiple R-squared: 0.2612, Adjusted R-squared: 0.2563
F-statistic: 53.73 on 2 and 304 DF, p-value: < 2.2e-16
plot(aaares)
The model shows that both “SpecialProjectsCount” and “Absences” have statistically significant effects on salary, with “SpecialProjectsCount” having a larger impact.
The R-squared value is 0.2612, suggesting that around 26.12% of the variance in Salary can be explained by the predictors.
The F-statistic and its associated p-value (< 2.2e-16) indicate that the overall model is statistically significant.
The higher the Adjusted R-squared value, the better the model fits the data. The Adjusted R-squared values are as follows:
For the first model (Salary=β0+β1EmpSatisfaction+β2SpecialProjectsCount+β3Absences+β4GenderID+ε): 0.2528
For the second model (Salary=β0+β1EmpSatisfaction+β2SpecialProjectsCount+β3Absences+ε): 0.2551
For the third model (Salary=β0+β1SpecialProjectsCount+β2Absences+ε): 0.2563
Based on this, the third model - Salary=β0+β1SpecialProjectsCount+β2Absences+ε - is the best since it has the highest Adjusted R-squared value, indicating the best fit to the data while utilizing fewer independent variables. While the R-squared values are similar across all three models, the Adjusted R-squared takes into account the complexity of the model, so we choosing the third model.