2024-05-30

Salary Vs. Years of Experience

At most places of employment, more years of experience means more income. However, with most things, there are always exceptions to the rule.

Education, location, and sometimes unfortunately even favoritism plays a role in determining the overall pay for employees.

For this Homework however, we will just be focusing on data obtained from kaggle.com

Data Summary

The summary of the data we used is shown below. We have obtained data from a minimum Years of Experience of 1.1 and a max of 10.5. But will these values correlate to more experience means more money?

 YearsExperience      Salary      
 Min.   : 1.100   Min.   : 37731  
 1st Qu.: 3.200   1st Qu.: 56721  
 Median : 4.700   Median : 65237  
 Mean   : 5.313   Mean   : 76003  
 3rd Qu.: 7.700   3rd Qu.:100545  
 Max.   :10.500   Max.   :122391  

Graphs of Data Set

Here we have the graph of the data obtained plotted as multiple points, along with the fitted line through the middle, which shows the linear regression equation. \[\begin{equation} y=\underbrace{ \overbrace{25792.2}^\text{Intercept} + \overbrace{9450x}^\text{Value} }_\text{Linear Regression Equation} \end{equation}\]

Code of plot_ly

plot_ly is a very useful package where we can visualize data in a graph and customize it, below is the code we used to create the plot_ly graph on the previous slide.

sal = lm(Salary ~ YearsExperience, data=id) y = id\(Salary; x = id\)YearsExperience sal1 = list(title = “Salary”, titlefont = list(family = “Modern Computer Roman”)) sal2 = list(title = “Years of Experience”, titlefont = list(family = “Modern Computer Roman”), range = c(0,11))

fig = plot_ly(x=x, y=y, type=“scatter”, mode=“markers”, name=“data”, width=800, height=430) %>% add_lines(y=fitted(sal), x=x, name=“fitted”) %>% layout(xaxis=sal2, yaxis=sal1) %>% layout(margin=list(l=25, r=25, b=25, t=25))

config(fig, displaylogo=FALSE)

Linear Regression Equation

To find the Linear Regression of any given data set we use the equation below.

\[\begin{equation} \hat{Y}_i = \hat{\beta}_0 + \hat{\beta}_1 X_i + \hat{\epsilon}_i \end{equation}\]

The parts of the equation is also given below.

\[\begin{equation} \hat{Y}_i = \underbrace{ \overbrace{\hat{\beta}_0}^\text{Y-intercept} + \overbrace{\hat{\beta}_1}^\text{Slope/Coefficient} \overbrace{X_i}^\text{Independent Variable}}_\text{Linear Component} + \underbrace{\overbrace{\hat{\epsilon}_i}^\text{Random Error Term}}_\text{Random Error Component} \end{equation}\]

ggplots

Here we have a ggplot of data set.

Linear Regression Date

And here we have the results from the Linear Regression calculations from the obtained data set.

Call:
lm(formula = Salary ~ YearsExperience, data = id)

Residuals:
    Min      1Q  Median      3Q     Max 
-7958.0 -4088.5  -459.9  3372.6 11448.0 

Coefficients:
                Estimate Std. Error t value Pr(>|t|)    
(Intercept)      25792.2     2273.1   11.35 5.51e-12 ***
YearsExperience   9450.0      378.8   24.95  < 2e-16 ***
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

Residual standard error: 5788 on 28 degrees of freedom
Multiple R-squared:  0.957, Adjusted R-squared:  0.9554 
F-statistic: 622.5 on 1 and 28 DF,  p-value: < 2.2e-16

Linear Regression Models

The following graphs show the different parts of the linear regression model itself. From Residual Standard and Residual Standard Error, Residual squared, and the square root of the Residual.

Summary

In conclusion, the overall correlation between years of experience and the amount of money made per year shows a steady relationship with the increase in years of experience.

While there are always different factors that can contribute to this type of data and other data sets, simple linear regression equations and models give a very good sample of data that can be used to help make life decisions.