Introduction

Artificial intelligence (AI) is a field of technology that focuses on simulating human intelligence and automating tasks. AI can have a significant impact on academic productivity in a number of ways, such as automating tasks, helping with data analysis, and providing personalized learning experiences. However, AI can also lead to technostress among academic workers. Technostress is an emotional state that is caused by the use of technology, and is characterized by feelings of stress, anxiety, and frustration. Technostress can be caused by a number of factors, such as overload, technological problems, and uncertainty about how to use technology.

Technostress can have a negative impact on academic productivity. Those who experience technostress may be less productive, more likely to make mistakes, and more likely to feel burned out. Technostress can also have a negative impact on the learning experience. Those who experience technostress are less likely to be interested in learning, less likely to participate actively in class, and more likely to feel overwhelmed.

It is important to research the relationship between AI, technostress, and academic productivity in order to understand how to use AI safely and effectively in the academic environment.

Load Data

# Load the dataset
data <- read_excel("encoded_data.xlsx")

# Define dependent and independent variables for the model
dependent_var <- 'ai_affect_productity'
independent_vars <- c('how_often_experience_technostress', 'recommendation_coded', 'areas_use_ai_coded', 
                    'ai_platforms_coded', 'tech_solution_coded', 'age', 'field_coded', 'education')

# Drop rows with missing values in the selected columns
data <- na.omit(data[, c(dependent_var, independent_vars)])

Fit Regression Model

# Fit a regression model
model <- lm(ai_affect_productity ~ how_often_experience_technostress + recommendation_coded + 
            areas_use_ai_coded + ai_platforms_coded + tech_solution_coded + 
            age + field_coded + education, data=data)
summary(model)
## 
## Call:
## lm(formula = ai_affect_productity ~ how_often_experience_technostress + 
##     recommendation_coded + areas_use_ai_coded + ai_platforms_coded + 
##     tech_solution_coded + age + field_coded + education, data = data)
## 
## Residuals:
##     Min      1Q  Median      3Q     Max 
## -1.1381 -0.5102 -0.2553  0.2043  1.8292 
## 
## Coefficients:
##                                    Estimate Std. Error t value Pr(>|t|)    
## (Intercept)                        1.768547   0.313759   5.637 1.08e-07 ***
## how_often_experience_technostress -0.222235   0.071795  -3.095  0.00242 ** 
## recommendation_coded              -0.007376   0.018546  -0.398  0.69150    
## areas_use_ai_coded                -0.017592   0.036176  -0.486  0.62761    
## ai_platforms_coded                 0.089350   0.055861   1.600  0.11221    
## tech_solution_coded               -0.036018   0.025930  -1.389  0.16726    
## age                                0.006042   0.068789   0.088  0.93014    
## field_coded                        0.032151   0.010588   3.037  0.00291 ** 
## education                          0.002151   0.131726   0.016  0.98700    
## ---
## Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
## 
## Residual standard error: 0.7833 on 126 degrees of freedom
## Multiple R-squared:  0.1535, Adjusted R-squared:  0.0998 
## F-statistic: 2.857 on 8 and 126 DF,  p-value: 0.005945

Diagnostic Plots

Residuals vs Fitted

# Residuals vs Fitted Plot
plot(model, which=1)

This chart shows the relationship between the residuals and the fitted values for a linear regression model. The chart has the following characteristics:

Residuals vs. Fitted Values: The y-axis represents the residuals, and the x-axis represents the fitted values. Residuals show the differences between the observed values and the values predicted by the model.

Model: The regression model uses the following variables: ai_affect_productivity, how_often_experience_technostress, and recommendatio … (the full variable names are not entirely visible at the bottom of the chart).

Trend: There is a red line on the chart showing the trend of the residuals as a function of the fitted values. The curvature of the line suggests that the model is likely not fully linear and might not fit the data appropriately.

Distribution of Residuals: The chart indicates that the residuals are not evenly distributed along the fitted values, which might suggest that there are certain nonlinear relationships in the data that the current model does not capture.

Notable Outliers: There are some outliers on the chart that significantly deviate from the trend, and these points are marked with numbers (e.g., 19).

Based on this chart, it might be useful to conduct further analyses, such as applying nonlinear models or examining the outliers more closely, to improve the model’s fit and accuracy.

Normal Q-Q

# Normal Q-Q Plot
plot(model, which=2)

This chart shows a quantile-quantile (Q-Q) plot for the standardized residuals of the regression model. The chart has the following characteristics:

Q-Q Residuals: The y-axis represents the standardized residuals, while the x-axis shows the theoretical quantiles. The standardized residuals are the normalized values of the residuals compared to the theoretical quantiles of a normal distribution.

Purpose: The purpose of the Q-Q plot is to check whether the residuals follow a normal distribution. If the residuals follow a normal distribution, the data points should lie roughly along a straight line.

Deviations: The plot shows that the data points significantly deviate from the straight line, especially at the extremes (top right and bottom left corners of the graph). This suggests that the residuals do not perfectly follow a normal distribution, which can be problematic for the assumptions of the model.

Trend: The plot indicates that the distribution of the residuals may be asymmetric and potentially follow a skewed distribution, requiring further analysis.

Based on this chart, it might be useful to re-examine the assumptions of the model and possibly apply further transformations or different modeling techniques to achieve a better fit to the data.

Scale Location

# Scale-Location Plot
plot(model, which=3)

This chart shows a Scale-Location plot for the standardized residuals of the regression model. The chart has the following characteristics:

Scale-Location: The y-axis represents the square root of the standardized residuals (sqrt(standardized residuals)), while the x-axis shows the fitted values.

Purpose: The purpose of the Scale-Location plot is to check whether the residuals have constant variance across the fitted values (homoscedasticity). If the residuals have constant variance, the points should be randomly scattered on the graph and not show any specific pattern.

Trend: There is a red trend line on the chart showing the spread of the residuals as a function of the fitted values. The upward slope of the line suggests that the residuals’ variance increases with the fitted values, indicating heteroscedasticity.

Residuals vs Leverage

# Residuals vs Leverage Plot
plot(model, which=5)

This chart shows the relationship between residuals and leverage in a regression model. The chart has the following characteristics:

Residuals vs. Leverage: The y-axis represents the standardized residuals, and the x-axis represents the leverage values. Leverage measures how much influence a particular observation has on the fit of the regression model.

Cook’s Distance: A gray dashed line indicates Cook’s distance on the chart. This line shows which points have a significant influence on the model parameters. Points with a large Cook’s distance require special attention because they can significantly affect the model’s results.

Trend: A red trend line on the chart shows the distribution of the residuals as a function of leverage. This line indicates whether there is any pattern or trend between the residuals and leverage.