Introduction
Artificial intelligence (AI) is a field of technology that focuses on simulating human intelligence and automating tasks. AI can have a significant impact on academic productivity in a number of ways, such as automating tasks, helping with data analysis, and providing personalized learning experiences. However, AI can also lead to technostress among academic workers. Technostress is an emotional state that is caused by the use of technology, and is characterized by feelings of stress, anxiety, and frustration. Technostress can be caused by a number of factors, such as overload, technological problems, and uncertainty about how to use technology.
Technostress can have a negative impact on academic productivity. Those who experience technostress may be less productive, more likely to make mistakes, and more likely to feel burned out. Technostress can also have a negative impact on the learning experience. Those who experience technostress are less likely to be interested in learning, less likely to participate actively in class, and more likely to feel overwhelmed.
It is important to research the relationship between AI, technostress, and academic productivity in order to understand how to use AI safely and effectively in the academic environment.
# Load the dataset
data <- read_excel("encoded_data.xlsx")
# Define dependent and independent variables for the model
dependent_var <- 'ai_affect_productity'
independent_vars <- c('how_often_experience_technostress', 'recommendation_coded', 'areas_use_ai_coded',
'ai_platforms_coded', 'tech_solution_coded', 'age', 'field_coded', 'education')
# Drop rows with missing values in the selected columns
data <- na.omit(data[, c(dependent_var, independent_vars)])
# Fit a regression model
model <- lm(ai_affect_productity ~ how_often_experience_technostress + recommendation_coded +
areas_use_ai_coded + ai_platforms_coded + tech_solution_coded +
age + field_coded + education, data=data)
summary(model)
##
## Call:
## lm(formula = ai_affect_productity ~ how_often_experience_technostress +
## recommendation_coded + areas_use_ai_coded + ai_platforms_coded +
## tech_solution_coded + age + field_coded + education, data = data)
##
## Residuals:
## Min 1Q Median 3Q Max
## -1.1381 -0.5102 -0.2553 0.2043 1.8292
##
## Coefficients:
## Estimate Std. Error t value Pr(>|t|)
## (Intercept) 1.768547 0.313759 5.637 1.08e-07 ***
## how_often_experience_technostress -0.222235 0.071795 -3.095 0.00242 **
## recommendation_coded -0.007376 0.018546 -0.398 0.69150
## areas_use_ai_coded -0.017592 0.036176 -0.486 0.62761
## ai_platforms_coded 0.089350 0.055861 1.600 0.11221
## tech_solution_coded -0.036018 0.025930 -1.389 0.16726
## age 0.006042 0.068789 0.088 0.93014
## field_coded 0.032151 0.010588 3.037 0.00291 **
## education 0.002151 0.131726 0.016 0.98700
## ---
## Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
##
## Residual standard error: 0.7833 on 126 degrees of freedom
## Multiple R-squared: 0.1535, Adjusted R-squared: 0.0998
## F-statistic: 2.857 on 8 and 126 DF, p-value: 0.005945
This regression analysis summary models the dependent variable
ai_affect_productivity using multiple independent
variables. Below are the detailed results:
Min: -1.1381
1Q: -0.5102
Median: -0.2553
3Q: 0.2043
Max: 1.8292
The table shows the regression model coefficients for the following variables:
Intercept: 1.768547 (p < 0.001), which is significant.
how_often_experience_technostress: -0.222235 (p
< 0.01), which is significant. This means that for every unit
increase in this variable, the ai_affect_productivity
decreases by an average of 0.222235 units.
recommendation_coded: -0.007376 (not significant, p = 0.69150), has no significant effect.
areas_use_ai_coded: -0.017592 (not significant, p = 0.62761), has no significant effect.
ai_platforms_coded: 0.089350 (not significant, p = 0.11221), has no significant effect.
tech_solution_coded: -0.036018 (not significant, p = 0.16726), has no significant effect.
age: 0.006042 (not significant, p = 0.93014), has no significant effect.
field_coded: 0.032151 (p < 0.01), which is
significant. This means that for every unit increase in this variable,
the ai_affect_productivity increases by an average of
0.032151 units.
education: 0.002151 (not significant, p = 0.98700), has no significant effect.Diagnostic Plots
# Residuals vs Fitted Plot
plot(model, which=1)
This chart shows the relationship between the residuals and the fitted values for a linear regression model. The chart has the following characteristics:
Residuals vs. Fitted Values: The y-axis represents the residuals, and the x-axis represents the fitted values. Residuals show the differences between the observed values and the values predicted by the model.
Model: The regression model uses the following variables: ai_affect_productivity, how_often_experience_technostress, and recommendatio … (the full variable names are not entirely visible at the bottom of the chart).
Trend: There is a red line on the chart showing the trend of the residuals as a function of the fitted values. The curvature of the line suggests that the model is likely not fully linear and might not fit the data appropriately.
Distribution of Residuals: The chart indicates that the residuals are not evenly distributed along the fitted values, which might suggest that there are certain nonlinear relationships in the data that the current model does not capture.
Notable Outliers: There are some outliers on the chart that significantly deviate from the trend, and these points are marked with numbers (e.g., 19).
Based on this chart, it might be useful to conduct further analyses, such as applying nonlinear models or examining the outliers more closely, to improve the model’s fit and accuracy.
# Normal Q-Q Plot
plot(model, which=2)
This chart shows a quantile-quantile (Q-Q) plot for the standardized residuals of the regression model. The chart has the following characteristics:
Q-Q Residuals: The y-axis represents the standardized residuals, while the x-axis shows the theoretical quantiles. The standardized residuals are the normalized values of the residuals compared to the theoretical quantiles of a normal distribution.
Purpose: The purpose of the Q-Q plot is to check whether the residuals follow a normal distribution. If the residuals follow a normal distribution, the data points should lie roughly along a straight line.
Deviations: The plot shows that the data points significantly deviate from the straight line, especially at the extremes (top right and bottom left corners of the graph). This suggests that the residuals do not perfectly follow a normal distribution, which can be problematic for the assumptions of the model.
Trend: The plot indicates that the distribution of the residuals may be asymmetric and potentially follow a skewed distribution, requiring further analysis.
Based on this chart, it might be useful to re-examine the assumptions of the model and possibly apply further transformations or different modeling techniques to achieve a better fit to the data.
Scale Location
# Scale-Location Plot
plot(model, which=3)
This chart shows a Scale-Location plot for the standardized residuals of the regression model. The chart has the following characteristics:
Scale-Location: The y-axis represents the square root of the standardized residuals (sqrt(standardized residuals)), while the x-axis shows the fitted values.
Purpose: The purpose of the Scale-Location plot is to check whether the residuals have constant variance across the fitted values (homoscedasticity). If the residuals have constant variance, the points should be randomly scattered on the graph and not show any specific pattern.
Trend: There is a red trend line on the chart showing the spread of the residuals as a function of the fitted values. The upward slope of the line suggests that the residuals’ variance increases with the fitted values, indicating heteroscedasticity.
# Residuals vs Leverage Plot
plot(model, which=5)
This chart shows the relationship between residuals and leverage in a regression model. The chart has the following characteristics:
Residuals vs. Leverage: The y-axis represents the standardized residuals, and the x-axis represents the leverage values. Leverage measures how much influence a particular observation has on the fit of the regression model.
Cook’s Distance: A gray dashed line indicates Cook’s distance on the chart. This line shows which points have a significant influence on the model parameters. Points with a large Cook’s distance require special attention because they can significantly affect the model’s results.
Trend: A red trend line on the chart shows the distribution of the residuals as a function of leverage. This line indicates whether there is any pattern or trend between the residuals and leverage.
# Cook's Distance
cooksd <- cooks.distance(model)
plot(cooksd, ylab="Cook's distance", type="h")
abline(h = 4/(nrow(data) - length(model$coefficients) - 1), col="red")
library(ggplot2)
library(GGally)
## Registered S3 method overwritten by 'GGally':
## method from
## +.gg ggplot2
library(reshape2) # Make sure to load the reshape2 library
# Load the data
data <- read.csv("encoded_data.csv", sep=";")
# Convert relevant columns to numeric
cols <- c('ai_affect_productity', 'how_often_experience_technostress', 'recommendation_coded',
'areas_use_ai_coded', 'ai_platforms_coded', 'tech_solution_coded', 'age', 'field_coded', 'education')
data[cols] <- lapply(data[cols], as.numeric)
# Drop rows with missing values
data_clean <- na.omit(data[cols])
# Pair plot
ggpairs(data_clean, lower = list(continuous = wrap("smooth", alpha = 0.5, color = "blue")),
diag = list(continuous = wrap("barDiag", fill = "blue")),
upper = list(continuous = wrap("cor", size = 3)))
## `stat_bin()` using `bins = 30`. Pick better value with `binwidth`.
## `stat_bin()` using `bins = 30`. Pick better value with `binwidth`.
## `stat_bin()` using `bins = 30`. Pick better value with `binwidth`.
## `stat_bin()` using `bins = 30`. Pick better value with `binwidth`.
## `stat_bin()` using `bins = 30`. Pick better value with `binwidth`.
## `stat_bin()` using `bins = 30`. Pick better value with `binwidth`.
## `stat_bin()` using `bins = 30`. Pick better value with `binwidth`.
## `stat_bin()` using `bins = 30`. Pick better value with `binwidth`.
## `stat_bin()` using `bins = 30`. Pick better value with `binwidth`.
library(reshape2) # Ensure reshape2 is loaded here as well
# Correlation matrix
correlation_matrix <- cor(data_clean)
melted_cormat <- melt(correlation_matrix)
ggplot(data = melted_cormat, aes(x=Var1, y=Var2, fill=value)) +
geom_tile() +
scale_fill_gradient2(low = "blue", high = "red", mid = "white",
midpoint = 0, limit = c(-1, 1), space = "Lab",
name="Correlation") +
theme_minimal() +
theme(axis.text.x = element_text(angle = 45, vjust = 1,
size = 12, hjust = 1)) +
coord_fixed()
This image shows a correlation matrix heatmap, which illustrates the correlation coefficients between various variables in the dataset. Here’s a detailed summary in English:
The correlation matrix heatmap visually represents the strength and direction of relationships between pairs of variables. The color intensity indicates the strength of the correlation, with red indicating positive correlations and blue indicating negative correlations. Key Observations:
Strong Positive Correlations: recommendation_coded and ai_affect_productity show a strong positive correlation. education and field_coded also exhibit a strong positive correlation. recommendation_coded has a strong positive correlation with areas_use_ai_coded. Strong Negative Correlations: how_often_experience_technostress and ai_affect_productity show a strong negative correlation. Moderate Correlations: age and field_coded show a moderate positive correlation. tech_solution_coded and ai_platforms_coded exhibit a moderate positive correlation. areas_use_ai_coded and recommendation_coded have a moderate positive correlation. Weak or No Correlation: Several variable pairs show little to no correlation (near white or light-colored squares), indicating weak or no linear relationship between those variables. Specific Variable Relationships:
ai_affect_productity: Positively correlated with recommendation_coded. Negatively correlated with how_often_experience_technostress. how_often_experience_technostress: Negatively correlated with ai_affect_productity. education: Positively correlated with field_coded. recommendation_coded: Positively correlated with ai_affect_productity and areas_use_ai_coded. Interpretation:
Variables with strong positive correlations move together, meaning an increase in one variable is associated with an increase in the other. Variables with strong negative correlations move in opposite directions, meaning an increase in one variable is associated with a decrease in the other. Understanding these relationships helps in identifying key variables that influence each other, which is crucial for building predictive models and making informed decisions. Visual Features: Color Scale: The color gradient ranges from blue (negative correlation) to red (positive correlation). The intensity of the color represents the strength of the correlation, with deeper colors indicating stronger correlations. Axes: The variables are listed along both the x and y axes. Diagonal: The diagonal represents the correlation of each variable with itself, which is always 1 and is typically displayed in dark red. This correlation matrix provides a comprehensive overview of how variables in the dataset relate to each other, offering valuable insights for further statistical analysis and model building.
Conclusion of the Technostress and Artificial Intelligence Analysis The regression analysis aimed to understand the impact of various factors on the perceived effect of artificial intelligence (AI) on productivity, focusing on technostress as a significant variable.
Key Findings: Technostress Impact:
The variable how_often_experience_technostress showed a significant negative effect on ai_affect_productivity (Estimate: -0.222235, p < 0.01). This indicates that increased technostress is associated with a decrease in the positive impact of AI on productivity. Field of Work:
The variable field_coded was also significant (Estimate: 0.032151, p < 0.01). This suggests that individuals from different fields experience varying impacts of AI on productivity, with some fields showing a more positive association. Other Variables:
Variables such as recommendation_coded, areas_use_ai_coded, ai_platforms_coded, tech_solution_coded, age, and education did not show significant effects on ai_affect_productivity. Residual Analysis: The residuals of the model ranged from -1.1381 to 1.8292, with a median value of -0.2553. The distribution of residuals indicates some variance around the predicted values, suggesting areas for model improvement. Model Performance: The model’s R-squared value was 0.1535, indicating that approximately 15.35% of the variance in ai_affect_productivity is explained by the included variables. The adjusted R-squared value was 0.0998, reflecting the model’s performance when considering the number of predictors. Conclusion: The analysis highlights that technostress significantly impacts the perceived effectiveness of AI on productivity, with higher levels of technostress correlating with lower perceived productivity gains from AI. Additionally, the field of work plays a crucial role in how AI’s impact on productivity is perceived.
However, the overall model explains a modest portion of the variance in AI’s impact on productivity, suggesting that other unexamined factors might also be influential. Future research could explore additional variables and non-linear relationships to improve the model’s explanatory power and better understand the multifaceted relationship between technostress and AI.
Contact Information:
Alžbeta Simon, PhDr.
Doctoral Student
Position: Selye János University, Faculty of Economics and Informatics, Department of Management
Address: Hradná ul. 21. P.O. BOX 54. 945 01, Komárno, Slovakia
Email: csuhajova93@gmail.com