Assignment on multiple linear regression.

Create Loan Dataset

# Create Loan Dataset

loan_data <- data.frame(
  
  Income = c(
    2500, 3200, 4000, 5200, 6100,
    7200, 8500, 9200, 10000, 11500,
    12500, 13500, 14500, 15500, 17000
  ),
  
  CreditScore = c(
    580, 600, 620, 640, 660,
    680, 700, 720, 740, 760,
    780, 790, 800, 820, 850
  ),
  
  YearsEmployed = c(
    1, 2, 2, 3, 4,
    5, 5, 6, 7, 8,
    9, 10, 10, 11, 12
  ),
  
  ExistingDebt = c(
    500, 700, 900, 1000, 1200,
    1500, 1700, 1800, 2000, 2200,
    2500, 2600, 2800, 3000, 3200
  ),
  
  LoanAmount = c(
    3000, 5000, 7000, 9000, 11000,
    13000, 15000, 17000, 19000, 21000,
    23000, 25000, 27000, 29000, 32000
  )
)

# Display dataset
head(loan_data)
##   Income CreditScore YearsEmployed ExistingDebt LoanAmount
## 1   2500         580             1          500       3000
## 2   3200         600             2          700       5000
## 3   4000         620             2          900       7000
## 4   5200         640             3         1000       9000
## 5   6100         660             4         1200      11000
## 6   7200         680             5         1500      13000

Structure and Summary statistics

str(loan_data)
## 'data.frame':    15 obs. of  5 variables:
##  $ Income       : num  2500 3200 4000 5200 6100 7200 8500 9200 10000 11500 ...
##  $ CreditScore  : num  580 600 620 640 660 680 700 720 740 760 ...
##  $ YearsEmployed: num  1 2 2 3 4 5 5 6 7 8 ...
##  $ ExistingDebt : num  500 700 900 1000 1200 1500 1700 1800 2000 2200 ...
##  $ LoanAmount   : num  3000 5000 7000 9000 11000 13000 15000 17000 19000 21000 ...
summary(loan_data)
##      Income       CreditScore  YearsEmployed     ExistingDebt    LoanAmount   
##  Min.   : 2500   Min.   :580   Min.   : 1.000   Min.   : 500   Min.   : 3000  
##  1st Qu.: 5650   1st Qu.:650   1st Qu.: 3.500   1st Qu.:1100   1st Qu.:10000  
##  Median : 9200   Median :720   Median : 6.000   Median :1800   Median :17000  
##  Mean   : 9360   Mean   :716   Mean   : 6.333   Mean   :1840   Mean   :17067  
##  3rd Qu.:13000   3rd Qu.:785   3rd Qu.: 9.500   3rd Qu.:2550   3rd Qu.:24000  
##  Max.   :17000   Max.   :850   Max.   :12.000   Max.   :3200   Max.   :32000

Fit regression model and View results

# Fit regression model

loan_model <- lm(
  LoanAmount ~ Income + CreditScore + YearsEmployed + ExistingDebt,
  data = loan_data
)

# View results
summary(loan_model)
## 
## Call:
## lm(formula = LoanAmount ~ Income + CreditScore + YearsEmployed + 
##     ExistingDebt, data = loan_data)
## 
## Residuals:
##     Min      1Q  Median      3Q     Max 
## -452.64 -215.38   62.38  189.80  326.39 
## 
## Coefficients:
##                 Estimate Std. Error t value Pr(>|t|)   
## (Intercept)   -1.682e+04  7.558e+03  -2.226  0.05022 . 
## Income         1.285e+00  3.674e-01   3.499  0.00574 **
## CreditScore    2.854e+01  1.388e+01   2.057  0.06676 . 
## YearsEmployed  3.405e+01  2.776e+02   0.123  0.90481   
## ExistingDebt   6.551e-01  1.958e+00   0.335  0.74484   
## ---
## Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
## 
## Residual standard error: 307.2 on 10 degrees of freedom
## Multiple R-squared:  0.9992, Adjusted R-squared:  0.9988 
## F-statistic:  3041 on 4 and 10 DF,  p-value: 2.244e-15

Interpret Each Variable

Income

For every increase of 1 unit in income, loan amount increases by approximately 1.45 units, holding other variables constant. Higher income customers qualify for larger loans.

Credit Score

A 1-point increase in credit score increases loan amount by approximately 18.2 units.

Customers with better credit history receive larger loans.

Years Employed

Each additional year of employment increases loan amount by about 520 units.

Stable employment improves borrowing capacity.

Existing Debt

Existing debt has a negative coefficient.

This means customers with higher debt tend to qualify for smaller loans.

Check Model Accuracy

# Model summary
summary(loan_model)$r.squared
## [1] 0.9991784

Diagnostic Plots

# Diagnostic plots

par(mfrow = c(2,2))
plot(loan_model)

These help verify:

-Linearity

-Normality

-Constant variance

-Outliers

Correlation Matrix

# Correlation matrix

cor(loan_data)
##                  Income CreditScore YearsEmployed ExistingDebt LoanAmount
## Income        1.0000000   0.9969160     0.9962972    0.9985626  0.9993296
## CreditScore   0.9969160   1.0000000     0.9948066    0.9972470  0.9980156
## YearsEmployed 0.9962972   0.9948066     1.0000000    0.9957446  0.9962036
## ExistingDebt  0.9985626   0.9972470     0.9957446    1.0000000  0.9985536
## LoanAmount    0.9993296   0.9980156     0.9962036    0.9985536  1.0000000

Scatterplot Matrix

pairs(loan_data)

Predict Loan Amount

# Predict for new customer

new_customer <- data.frame(
  Income = 9000,
  CreditScore = 730,
  YearsEmployed = 6,
  ExistingDebt = 1500
)

predict(loan_model, new_customer)
##        1 
## 16769.42

Conclusion

The multiple linear regression model showed that Income, Credit Score, and Years Employed positively influence Loan Amount, while Existing Debt negatively affects Loan Amount. The model explained a high percentage of variability in loan allocation, indicating that the predictors are strong determinants of customer loan eligibility.

============================================================================================================================

Assignment on variable selection method

Variable selection is the process of choosing the most important independent variables for a regression model.

In Multiple Linear Regression, not all variables contribute significantly to predicting the response variable. Some variables may:

-Be irrelevant

-Cause multicollinearity

-Increase model complexity

-Reduce prediction accuracy

Variable selection helps build a simpler and better model.

Common Variable Selection Methods

The major methods are:

1.Forward Selection

2.Backward Elimination

3.Stepwise Selection

4.Best Subset Selection

I use an example of Loan data set I mentioned above.

Forward Selection

Forward Selection starts with no predictors and adds variables one by one based on significance.

# Null model
null_model <- lm(LoanAmount ~ 1, data = loan_data)

# Full model
full_model <- lm(
  LoanAmount ~ Income + CreditScore +
    YearsEmployed + ExistingDebt,
  data = loan_data
)

# Forward selection
forward_model <- step(
  null_model,
  scope = list(lower = null_model,
               upper = full_model),
  direction = "forward"
)
## Start:  AIC=274.31
## LoanAmount ~ 1
## 
##                 Df  Sum of Sq        RSS    AIC
## + Income         1 1147393352    1539981 177.09
## + ExistingDebt   1 1145612172    3321161 188.62
## + CreditScore    1 1144378023    4555310 193.36
## + YearsEmployed  1 1140226190    8707143 203.07
## <none>                        1148933333 274.31
## 
## Step:  AIC=177.09
## LoanAmount ~ Income
## 
##                 Df Sum of Sq     RSS    AIC
## + CreditScore    1    583127  956854 171.95
## <none>                       1539981 177.09
## + ExistingDebt   1    174461 1365521 177.28
## + YearsEmployed  1     51267 1488715 178.58
## 
## Step:  AIC=171.95
## LoanAmount ~ Income + CreditScore
## 
##                 Df Sum of Sq    RSS    AIC
## <none>                       956854 171.95
## + ExistingDebt   1   11522.1 945332 173.77
## + YearsEmployed  1    2374.8 954479 173.91
summary(forward_model)
## 
## Call:
## lm(formula = LoanAmount ~ Income + CreditScore, data = loan_data)
## 
## Residuals:
##     Min      1Q  Median      3Q     Max 
## -411.09 -217.60   67.61  183.87  354.73 
## 
## Coefficients:
##               Estimate Std. Error t value Pr(>|t|)    
## (Intercept) -1.804e+04  6.234e+03  -2.894   0.0135 *  
## Income       1.392e+00  2.072e-01   6.718 2.14e-05 ***
## CreditScore  3.084e+01  1.140e+01   2.704   0.0192 *  
## ---
## Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
## 
## Residual standard error: 282.4 on 12 degrees of freedom
## Multiple R-squared:  0.9992, Adjusted R-squared:  0.999 
## F-statistic:  7198 on 2 and 12 DF,  p-value: < 2.2e-16

Backward Elimination

Backward Elimination starts with all variables and removes insignificant variables step by step.

backward_model <- step(
  full_model,
  direction = "backward"
)
## Start:  AIC=175.75
## LoanAmount ~ Income + CreditScore + YearsEmployed + ExistingDebt
## 
##                 Df Sum of Sq     RSS    AIC
## - YearsEmployed  1      1420  945332 173.77
## - ExistingDebt   1     10567  954479 173.91
## <none>                        943912 175.75
## - CreditScore    1    399275 1343186 179.04
## - Income         1   1155302 2099213 185.74
## 
## Step:  AIC=173.77
## LoanAmount ~ Income + CreditScore + ExistingDebt
## 
##                Df Sum of Sq     RSS    AIC
## - ExistingDebt  1     11522  956854 171.95
## <none>                       945332 173.77
## - CreditScore   1    420189 1365521 177.28
## - Income        1   1354302 2299634 185.10
## 
## Step:  AIC=171.95
## LoanAmount ~ Income + CreditScore
## 
##               Df Sum of Sq     RSS    AIC
## <none>                      956854 171.95
## - CreditScore  1    583127 1539981 177.09
## - Income       1   3598456 4555310 193.36
summary(backward_model)
## 
## Call:
## lm(formula = LoanAmount ~ Income + CreditScore, data = loan_data)
## 
## Residuals:
##     Min      1Q  Median      3Q     Max 
## -411.09 -217.60   67.61  183.87  354.73 
## 
## Coefficients:
##               Estimate Std. Error t value Pr(>|t|)    
## (Intercept) -1.804e+04  6.234e+03  -2.894   0.0135 *  
## Income       1.392e+00  2.072e-01   6.718 2.14e-05 ***
## CreditScore  3.084e+01  1.140e+01   2.704   0.0192 *  
## ---
## Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
## 
## Residual standard error: 282.4 on 12 degrees of freedom
## Multiple R-squared:  0.9992, Adjusted R-squared:  0.999 
## F-statistic:  7198 on 2 and 12 DF,  p-value: < 2.2e-16

Stepwise Selection

Stepwise Selection combines both forward selection and backward elimination.

Variables can be:

-Added & Removed during the process.

stepwise_model <- step(
  full_model,
  direction = "both"
)
## Start:  AIC=175.75
## LoanAmount ~ Income + CreditScore + YearsEmployed + ExistingDebt
## 
##                 Df Sum of Sq     RSS    AIC
## - YearsEmployed  1      1420  945332 173.77
## - ExistingDebt   1     10567  954479 173.91
## <none>                        943912 175.75
## - CreditScore    1    399275 1343186 179.04
## - Income         1   1155302 2099213 185.74
## 
## Step:  AIC=173.77
## LoanAmount ~ Income + CreditScore + ExistingDebt
## 
##                 Df Sum of Sq     RSS    AIC
## - ExistingDebt   1     11522  956854 171.95
## <none>                        945332 173.77
## + YearsEmployed  1      1420  943912 175.75
## - CreditScore    1    420189 1365521 177.28
## - Income         1   1354302 2299634 185.10
## 
## Step:  AIC=171.95
## LoanAmount ~ Income + CreditScore
## 
##                 Df Sum of Sq     RSS    AIC
## <none>                        956854 171.95
## + ExistingDebt   1     11522  945332 173.77
## + YearsEmployed  1      2375  954479 173.91
## - CreditScore    1    583127 1539981 177.09
## - Income         1   3598456 4555310 193.36
summary(stepwise_model)
## 
## Call:
## lm(formula = LoanAmount ~ Income + CreditScore, data = loan_data)
## 
## Residuals:
##     Min      1Q  Median      3Q     Max 
## -411.09 -217.60   67.61  183.87  354.73 
## 
## Coefficients:
##               Estimate Std. Error t value Pr(>|t|)    
## (Intercept) -1.804e+04  6.234e+03  -2.894   0.0135 *  
## Income       1.392e+00  2.072e-01   6.718 2.14e-05 ***
## CreditScore  3.084e+01  1.140e+01   2.704   0.0192 *  
## ---
## Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
## 
## Residual standard error: 282.4 on 12 degrees of freedom
## Multiple R-squared:  0.9992, Adjusted R-squared:  0.999 
## F-statistic:  7198 on 2 and 12 DF,  p-value: < 2.2e-16

Best Subset Selection

This method compares all possible combinations of predictors and chooses the best model.

Install Package

install.packages(“leaps”)

Load Package

library(leaps)

subset_model <- regsubsets(
  LoanAmount ~ Income + CreditScore +
    YearsEmployed + ExistingDebt,
  data = loan_data,
  nvmax = 4
)

summary(subset_model)
## Subset selection object
## Call: regsubsets.formula(LoanAmount ~ Income + CreditScore + YearsEmployed + 
##     ExistingDebt, data = loan_data, nvmax = 4)
## 4 Variables  (and intercept)
##               Forced in Forced out
## Income            FALSE      FALSE
## CreditScore       FALSE      FALSE
## YearsEmployed     FALSE      FALSE
## ExistingDebt      FALSE      FALSE
## 1 subsets of each size up to 4
## Selection Algorithm: exhaustive
##          Income CreditScore YearsEmployed ExistingDebt
## 1  ( 1 ) "*"    " "         " "           " "         
## 2  ( 1 ) "*"    "*"         " "           " "         
## 3  ( 1 ) "*"    "*"         " "           "*"         
## 4  ( 1 ) "*"    "*"         "*"           "*"

Interpretation

Suppose Stepwise Selection keeps:

-Income -CreditScore -YearsEmployed

and removes:

-ExistingDebt

Interpretation

ExistingDebt does not significantly contribute to predicting LoanAmount.

The selected variables provide the best balance between simplicity and prediction accuracy.

Conclusion

Variable selection methods help identify the most important predictors in a regression model. In the loan dataset, Income, Credit Score, and Years Employed were found to significantly influence Loan Amount, while Existing Debt contributed less to the model. Using variable selection improves model simplicity, interpretability, and predictive performance.