Karan Kumar Kar (s4027979)
Last updated: 02 June, 2024
Rpubs Link:
Understanding the relationship between education and unemployment is critical for shaping effective economic policies. Education is often touted as a key factor in reducing unemployment rates, but the specifics of this relationship, particularly concerning gender disparities in education, remain underexplored.
This project aims to delve into this complex interplay by examining the correlation between mean years of schooling (GenderRatio) and unemployment rates across various countries over different years.
By answering these questions, this research will provide valuable insights into how gender disparities in education affect unemployment rates, potentially guiding policymakers in developing more effective economic strategies.
Data Source: The data used in this analysis is derived from open datasets available on the Our World in Data website, specifically:
Complete List of Variables:
Dropped Columns: - From both the datasets ‘Code’ column was dropped due to redundancy.
Final Renamed & Chosen Columns: Country: The name of the country (Derived from the Entity column), Year: The year of the data point, Unemployment_Rate: The unemployment rate in the country for the given year, GenderRatio: The mean years of schooling, represented as a gender ratio.
Reason for Renaming Columns: The renaming of columns was done to enhance understanding and standardize variable names across datasets. The ‘Entity’ column was renamed to ‘Country’ for better clarity and to ensure consistency. Similarly, ‘Average years of schooling gender ratio, 15-64 year olds’ was simplified to ‘GenderRatio’ to make the column name more concise and readable. Lastly, ‘Unemployment, total (% of total labor force) (modeled ILO estimate)’ was shortened to ‘Unemployment_Rate’ to facilitate ease of reference and maintain uniformity in the naming conventions.
Reason for Choosing 2015 and 2020 The years 2015 and 2020 were chosen for this comparison to provide a 5-year interval that allows for observing trends and changes over a medium-term period. A 5-year interval is significant enough to capture meaningful changes in economic and social indicators, such as unemployment rates and gender ratios in education, due to policy implementations, economic cycles, and societal changes. Comparing these two years helps in understanding the progress and challenges within a relevant timeframe.
Pre-processing: The dataset was cleaned to remove any missing or inconsistent data entries. Variables were checked for outliers and corrected where necessary. Below are the steps taken:
data_genderRatio <- read.csv("D:/RMIT SLIDES/Applied Analytics/Assignment 2/gender-ratios-for-mean-years-of-schooling.csv")
data_genderRatio <- select(data_genderRatio, -Code)
data_genderRatio <- rename(data_genderRatio,
Country = Entity,
Year = Year,
GenderRatio = `Average.years.of.schooling.gender.ratio..15.64.year.olds`)
data_genderRatio <- na.omit(data_genderRatio)
unemployment_df <- read.csv("D:/RMIT SLIDES/Applied Analytics/Assignment 2/unemployment-rate.csv")
colnames(unemployment_df) <- c("Country", "Country_Code", "Year", "Unemployment_Rate")
unemployment_filtered_df <- unemployment_df %>%
filter(Year %in% c(2015, 2020)) %>%
select(-Country_Code)
unemployment_filtered_df <- na.omit(unemployment_filtered_df)merged_data <- merge(unemployment_filtered_df, data_genderRatio, on=c('Country', 'Year'))
data_2015 <- merged_data %>% filter(Year == 2015)
data_2020 <- merged_data %>% filter(Year == 2020)
summary_2015 <- data_2015 %>% select(-Year, -Country) %>% summary()
summary_2020 <- data_2020 %>% select(-Year, -Country) %>% summary()
kable(summary_2015, caption = "Summary Statistics for 2015")%>%
kable_styling(bootstrap_options = c("striped", "hover", "condensed", "responsive"))| Unemployment_Rate | GenderRatio | |
|---|---|---|
| Min. : 0.170 | Min. :0.5623 | |
| 1st Qu.: 3.570 | 1st Qu.:0.9164 | |
| Median : 6.080 | Median :1.0008 | |
| Mean : 7.581 | Mean :0.9620 | |
| 3rd Qu.: 9.800 | 3rd Qu.:1.0258 | |
| Max. :25.150 | Max. :1.3171 |
kable(summary_2020, caption = "Summary Statistics for 2020")%>%
kable_styling(bootstrap_options = c("striped", "hover", "condensed", "responsive"))| Unemployment_Rate | GenderRatio | |
|---|---|---|
| Min. : 0.214 | Min. :0.6322 | |
| 1st Qu.: 4.260 | 1st Qu.:0.9458 | |
| Median : 6.328 | Median :1.0068 | |
| Mean : 7.924 | Mean :0.9814 | |
| 3rd Qu.: 9.480 | 3rd Qu.:1.0298 | |
| Max. :29.220 | Max. :1.3567 |
data_2015$Year <- as.factor(data_2015$Year)
data_2020$Year <- as.factor(data_2020$Year)
combined_data <- rbind(data_2015, data_2020)
combined_data$Year <- as.factor(combined_data$Year)
shapiro.test(combined_data$GenderRatio)##
## Shapiro-Wilk normality test
##
## data: combined_data$GenderRatio
## W = 0.90404, p-value = 1.264e-12
##
## Shapiro-Wilk normality test
##
## data: combined_data$Unemployment_Rate
## W = 0.87229, p-value = 8.506e-15
Comparison and Interpretation
Test Statistic (W): For the Gender Ratio, W is 0.90404, for the Unemployment Rate, W is 0.87229. Both W values are less than 1, which is expected since W values range between 0 and 1 where 1 indicates a perfect normal distribution.
p-value: The p-value for the Gender Ratio is 1.264e-12. The p-value for the Unemployment Rate is 8.506e-15. Both p-values are extremely low (far below the common alpha level of 0.05), indicating that the null hypothesis of normality is rejected for both variables.
Conclusion
leveneTest_result <- leveneTest(Unemployment_Rate ~ Year, data = combined_data)
print(leveneTest_result)## Levene's Test for Homogeneity of Variance (center = median)
## Df F value Pr(>F)
## group 1 0.0129 0.9097
## 288
unemployment_rate_sd <- sd(combined_data$Unemployment_Rate, na.rm = TRUE)
unemployment_rate_var <- var(combined_data$Unemployment_Rate, na.rm = TRUE)
gender_ratio_sd <- sd(combined_data$GenderRatio, na.rm = TRUE)
gender_ratio_var <- var(combined_data$GenderRatio, na.rm = TRUE)
summary_stats <- data.frame(
Metric = c("Unemployment Rate SD", "Unemployment Rate Variance", "Gender Ratio SD", "Gender Ratio Variance"),
Value = c(unemployment_rate_sd, unemployment_rate_var, gender_ratio_sd, gender_ratio_var)
)
summary_stats %>%
kable() %>%
kable_styling(bootstrap_options = c("striped", "hover", "condensed", "responsive"))| Metric | Value |
|---|---|
| Unemployment Rate SD | 5.637275 |
| Unemployment Rate Variance | 31.778870 |
| Gender Ratio SD | 0.120445 |
| Gender Ratio Variance | 0.014507 |
The SD and variance values show the spread of data: the Unemployment Rate has a higher SD (5.637) and variance (31.779), indicating more variability compared to the Gender Ratio, which has a lower SD (0.120) and variance (0.0145), indicating less variability.
unemployment_skewness <- skewness(combined_data$Unemployment_Rate, na.rm = TRUE)
gender_ratio_skewness <- skewness(combined_data$GenderRatio, na.rm = TRUE)
ggplot(combined_data, aes(x = Unemployment_Rate)) +
geom_histogram(aes(y = ..density..), binwidth = 1, fill = "#A0CBE8", color = "#283B55", alpha = 1) +
geom_density(color = "blue", size = 1.2, linetype = "dashed") +
theme_minimal() +
labs(title = "Histogram of Unemployment Rate with Density Plot",
x = "Unemployment Rate",
y = "Density") +
annotate("text", x = Inf, y = Inf, label = paste("Skewness:", round(unemployment_skewness, 2), "\n(Right Skewed)"),
hjust = 1.1, vjust = 1.5, size = 4, color = "black", fontface = "bold") +
theme(plot.caption = element_text(hjust = 0.5, face = "italic"),
panel.background = element_rect(fill = "#F2F2F2")
)ggplot(combined_data, aes(x = GenderRatio)) +
geom_histogram(aes(y = ..density..), binwidth = 0.05, fill = "#F7CAC9", color = "#D62728", alpha = 1) +
geom_density(color = "red", size = 1.2, linetype = "dashed") +
theme_minimal() +
labs(title = "Histogram of Gender Ratio for Average Years of Schooling",
x = "Gender Ratio - Average Years of Schooling",
y = "Density") +
annotate("text", x = Inf, y = Inf, label = paste("Skewness:", round(gender_ratio_skewness, 2), "\n(Left Skewed)"),
hjust = 1.1, vjust = 1.5, size = 4, color = "black", fontface = "bold") +
theme(plot.caption = element_text(hjust = 0.5, face = "italic"),
panel.background = element_rect(fill = "#F2F2F2")
)Overall, the distribution is not symmetric, with a majority of the data concentrated around 1.00 and a few lower gender ratios stretching the tail to the left. This reflects that, on average, women have slightly fewer years of schooling than men in some instances.
cor_matrix <- cor(combined_data %>% select(Unemployment_Rate, GenderRatio), use = "complete.obs")
cor_matrix %>%
kable() %>%
kable_styling(bootstrap_options = c("striped", "hover", "condensed", "responsive"))| Unemployment_Rate | GenderRatio | |
|---|---|---|
| Unemployment_Rate | 1.0000000 | 0.1955426 |
| GenderRatio | 0.1955426 | 1.0000000 |
The correlation coefficient between the Unemployment Rate and Gender Ratio is 0.1955. This indicates a weak positive correlation between the two variables. As the Gender Ratio (female-to-male ratio of average years of schooling) increases, the Unemployment Rate tends to increase slightly, but the relationship is not strong. The low correlation value suggests that there is only a minor relationship between the Unemployment Rate and the Gender Ratio in this dataset.
ggplot(combined_data, aes(x = Year, y = Unemployment_Rate, fill = Year)) +
geom_boxplot() +
theme_minimal() +
ggtitle('Box Plot of Unemployment Rates (2015 vs 2020)') +
xlab('Year') +
ylab('Unemployment Rate')Interpretation of the Box Plot of Unemployment Rates (2015 vs 2020): The median unemployment rate for 2015 is slightly lower than that for 2020, indicating a marginal increase over the five-year period. The interquartile range (IQR) is similar for both years, suggesting that the middle 50% of unemployment rates are spread similarly. However, the whiskers extend further in 2020, showing greater variability and a wider range of unemployment rates.
Both years exhibit outliers, but 2020 has more, indicating some countries experienced unusually high unemployment rates. Overall, the median unemployment rate is slightly higher in 2020, with similar spread and variability between the two years. Both distributions show a right-skewed pattern, with more extreme values at the higher end, consistent with previous histogram and density plot analyses. This box plot analysis highlights a slight increase in unemployment rates from 2015 to 2020 and greater variability in 2020, possibly due to specific events or economic conditions affecting certain countries more severely.
ggplot(combined_data, aes(x = Year, y = GenderRatio, fill = Year)) +
geom_boxplot() +
theme_minimal() +
ggtitle('Box Plot of Gender ratio for average years of schooling (2015 vs 2020)') +
xlab('Year') +
ylab('Gender Ratio')Interpretation of the Box Plot of Gender Ratio for Average Years of Schooling (2015 vs 2020): The box plot comparing the gender ratio for average years of schooling between 2015 and 2020 reveals several insights. The median gender ratio for both years is around 1.0, indicating that on average, women have roughly the same years of schooling as men. The interquartile range (IQR) is also similar for both years, suggesting that the middle 50% of the data is similarly spread, showing consistency in educational gender parity. The whiskers, which extend to 1.5 times the IQR, indicate a broader range of values in 2020, pointing to increased variability.
There are numerous outliers in both years, with more pronounced lower outliers in 2015, showing that some countries had significantly lower gender ratios, meaning fewer women were educated compared to men. Meanwhile, the outliers on the higher end indicate countries where women had more schooling than men.
Overall, while the central tendency and spread of gender ratios remain consistent between the two years, the presence of more outliers and increased variability in 2020 suggests that while gender parity in education is stable on average, disparities still exist, and some countries have seen significant changes, either improving or worsening, in their gender ratios for education. This detailed analysis highlights the importance of addressing these disparities to achieve true educational equity globally.
Outliers: The outliers in both box plots are not handled or removed because they represent actual conditions in specific countries that are important for understanding educational and economic disparities. These outliers provide essential insights for policymakers to target interventions effectively and are not mere statistical anomalies.
ggplot(combined_data, aes(x = GenderRatio, y = Unemployment_Rate, color = factor(Year))) +
geom_point() +
geom_smooth(method = "lm", se = FALSE) +
theme_minimal() +
ggtitle('Scatter Plot of Unemployment Rate vs. Gender Ratio for Average Years of Schooling') +
xlab('Gender Ratio for Average Years of Schooling') +
ylab('Unemployment Rate (%)')The scatter plot visualizes the relationship between the Gender Ratio for average years of schooling and the Unemployment Rate for the years 2015 and 2020. Each point represents a country, with colors distinguishing between the two years (red for 2015 and blue for 2020).
Trend Line and Distribution: The linear trend lines for both 2015 and 2020 show a slight positive slope, indicating a weak positive correlation between the Gender Ratio and Unemployment Rate. This suggests that as the gender ratio increases (i.e., as the number of average years of schooling for women approaches or exceeds that of men), the unemployment rate tends to increase slightly.The points are widely scattered, indicating high variability in the unemployment rates for similar gender ratios. This further emphasizes the weak correlation, as the relationship between the variables is not strong.
Yearly Comparison and Outliers: The distribution of points and the trend lines are similar for both 2015 and 2020, indicating that the relationship between the gender ratio and unemployment rate did not change significantly over the five-year period. There are noticeable outliers with high unemployment rates and varying gender ratios in both years. These outliers represent countries with unique economic or social conditions affecting both education and employment.
Conclusion:: The scatter plot indicates a weak positive correlation between the Gender Ratio for average years of schooling and the Unemployment Rate for both 2015 and 2020. Despite the slight positive trend, the wide scatter and presence of outliers suggest that other factors are likely influencing unemployment rates more strongly. The relationship between educational gender parity and unemployment rates remains complex and is not solely defined by the gender ratio of education.
##
## Pearson's product-moment correlation
##
## data: merged_data$GenderRatio and merged_data$Unemployment_Rate
## t = 3.3838, df = 288, p-value = 0.0008141
## alternative hypothesis: true correlation is not equal to 0
## 95 percent confidence interval:
## 0.08221467 0.30387807
## sample estimates:
## cor
## 0.1955426
Pearson’s Correlation Test
Null Hypothesis (H0): There is no correlation between Gender Ratio and Unemployment Rate (ρ = 0).
Alternative Hypothesis (H1): There is a correlation between Gender Ratio and Unemployment Rate (ρ ≠ 0).
The correlation coefficient of 0.1955 suggests a weak positive correlation between Gender Ratio and Unemployment Rate.
The p-value (0.0008141) is much less than the significance level of 0.05, leading us to reject the null hypothesis and conclude that there is a statistically significant correlation.
The confidence interval does not include 0, further supporting the conclusion that the correlation is significant.
model <- lm(Unemployment_Rate ~ GenderRatio, data = combined_data)
model_summary <- summary(model)
conf_intervals <- confint(model, level = 0.95)
print(model_summary)##
## Call:
## lm(formula = Unemployment_Rate ~ GenderRatio, data = combined_data)
##
## Residuals:
## Min 1Q Median 3Q Max
## -9.704 -3.643 -1.604 1.722 20.915
##
## Coefficients:
## Estimate Std. Error t value Pr(>|t|)
## (Intercept) -1.140 2.648 -0.431 0.667072
## GenderRatio 9.152 2.705 3.384 0.000814 ***
## ---
## Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
##
## Residual standard error: 5.538 on 288 degrees of freedom
## Multiple R-squared: 0.03824, Adjusted R-squared: 0.0349
## F-statistic: 11.45 on 1 and 288 DF, p-value: 0.0008141
## 2.5 % 97.5 %
## (Intercept) -6.352538 4.071872
## GenderRatio 3.828646 14.475595
Null Hypothesis (H0): The Gender Ratio does not significantly predict Unemployment Rate (β1 = 0). Alternative Hypothesis (H1): The Gender Ratio significantly predicts Unemployment Rate (β1 ≠ 0).
The coefficient for Gender Ratio is 9.152, indicating that for each unit increase in Gender Ratio, the Unemployment Rate increases by 9.152 units on average.The p-value for the Gender Ratio coefficient is 0.000814, which is less than the significance level of 0.05, suggesting that Gender Ratio is a significant predictor of Unemployment Rate.The 95% confidence interval for the Gender Ratio coefficient (3.826646 to 14.475595) does not include 0, reinforcing the significance of the predictor. The R-squared value of 0.03824 indicates that approximately 3.82% of the variability in the Unemployment Rate is explained by the Gender Ratio, which is relatively low, suggesting that other factors may also be influential.
Assumptions Check
The Pearson correlation test shows a weak but significant positive correlation between Gender Ratio and Unemployment Rate. The linear regression analysis confirms that Gender Ratio is a significant predictor of Unemployment Rate, though the model explains a small portion of the variability in Unemployment Rate. The assumptions of linear regression need to be checked to validate the model further. Overall, the analysis demonstrates that while there is a significant relationship between Gender Ratio and Unemployment Rate, the effect size is relatively small, and other factors may need to be considered for a more comprehensive model.
combined_data$Unemployment_Category <- cut(combined_data$Unemployment_Rate,
breaks = c(-Inf, 5, 10, 15, Inf),
labels = c("Low", "Moderate", "High", "Very High"))
chi_square_test <- chisq.test(table(combined_data$Year, combined_data$Unemployment_Category))
print(chi_square_test)##
## Pearson's Chi-squared test
##
## data: table(combined_data$Year, combined_data$Unemployment_Category)
## X-squared = 0.42544, df = 3, p-value = 0.9349
contingency_table <- table(combined_data$Year, combined_data$Unemployment_Category)
chi_square_test <- chisq.test(contingency_table)
print(contingency_table)##
## Low Moderate High Very High
## 2015 58 53 18 16
## 2020 53 56 18 18
##
## Low Moderate High Very High
## 2015 55.5 54.5 18 17
## 2020 55.5 54.5 18 17
Null Hypothesis (H0): There is no association between the year and the unemployment category. Alternative Hypothesis (H1): There is an association between the year and the unemployment category.
Chi-Square Statistic (X²): 0.42544, Degrees of Freedom (df): 3, p-value: 0.9349
Assumptions Check for Chi-Square Test: Independence: Each observation should be independent of others. This is typically ensured by study design. Expected Frequency: Each expected frequency should be at least 5. In this case, all expected frequencies meet this criterion.
The Unemployment Rate is categorized into four levels: Low, Moderate, High, and Very High.The year is considered to assess if there is any change in the unemployment category distribution over different years.
The Chi-Square statistic value is 0.42544. With 3 degrees of freedom, the corresponding p-value is 0.9349. Given that the p-value is considerably higher than the standard alpha level of 0.05, there is insufficient evidence to suggest a significant relationship between the year and the unemployment category.
The analysis concludes that there is no significant association between the year and the unemployment category based on the data provided. This implies that the distribution of unemployment categories is consistent across different years, indicating stability in the unemployment rate categories over time. The results are meticulously presented and interpreted accurately, with the hypothesis test clearly justified and assumptions checked where appropriate. This comprehensive approach ensures robust conclusions are drawn from the data analysis.
multiple_model <- lm(Unemployment_Rate ~ GenderRatio + factor(Year), data = combined_data)
multiple_model_summary <- summary(multiple_model)
print(multiple_model_summary)##
## Call:
## lm(formula = Unemployment_Rate ~ GenderRatio + factor(Year),
## data = combined_data)
##
## Residuals:
## Min 1Q Median 3Q Max
## -9.774 -3.691 -1.586 1.784 20.835
##
## Coefficients:
## Estimate Std. Error t value Pr(>|t|)
## (Intercept) -1.1691 2.6549 -0.440 0.660020
## GenderRatio 9.0964 2.7180 3.347 0.000927 ***
## factor(Year)2020 0.1657 0.6536 0.253 0.800071
## ---
## Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
##
## Residual standard error: 5.547 on 287 degrees of freedom
## Multiple R-squared: 0.03845, Adjusted R-squared: 0.03175
## F-statistic: 5.739 on 2 and 287 DF, p-value: 0.0036
Model: Unemployment_Rate = βο + β₁.Gender Ratio + β2.Year + e
For Gender Ratio:Null Hypothesis (HO): β₁ = 0 (Gender Ratio does not significantly predict Unemployment Rate), Alternative Hypothesis (H1): β₁ ≠ 0 (Gender Ratio significantly predicts Unemployment Rate)
For Year(2020):Null Hypothesis (HO): β₁ = 0 (Year does not significantly predict Unemployment Rate), Alternative Hypothesis (H1): β₁ ≠ 0 (Year significantly predicts Unemployment Rate)
The intercept (Bo) signifies the unemployment rate when all the independent variables (Gender Ratio and Year in this case) are zero. The p-value (0.660020) associated with the intercept is not significant, indicating it doesn’t meaningfully contribute to the model. In simpler terms, the unemployment rate predicted by the model when Gender Ratio and Year are zero might not be very reliable.The coefficient for Gender Ratio (β1) is 9.0964, with a significant p-value (0.000927). This suggests that an increase in Gender Ratio is associated with an increase in Unemployment Rate.
The coefficient for the year 2020 (B2) is 0.1657, with a p-value of 0.800071 (not significant). This implies that, after considering Gender Ratio, the year 2020 doesn’t significantly predict the Unemployment Rate. In other words, there’s no strong evidence to suggest that the unemployment rate in the year 2020 is considerably different from other years after accounting for the effect of Gender Ratio.
The Multiple R-squared value of 0.03845 suggests that roughly 3.85% of the variation in the Unemployment Rate is explained by the model that includes Gender Ratio and Year. The Adjusted R-squared is slightly lower at 0.03175, which accounts for the model’s complexity (number of independent variables). The F-statistic (5.739) with a p-value of 0.0036 indicates that the overall model is statistically significant. This means that the model, at least partially, explains the relationship between Unemployment Rate and the predictor variables (Gender Ratio and Year). Overall, the analysis suggests that Gender Ratio has a significant positive association with Unemployment Rate, while the year 2020 doesn’t have a statistically significant effect on Unemployment Rate after considering Gender Ratio. It’s important to note that the model itself only explains a small portion of the variability in Unemployment Rate.
# Plot diagnostic plots using ggfortify
autoplot(multiple_model, which = c(1, 2, 3, 5), ncol = 2, label.size = 3) +
theme_bw() +
theme(
plot.title = element_text(size = 10, face = "bold"),
axis.title = element_text(size = 8),
axis.text = element_text(size = 8),
strip.text = element_text(size = 10, face = "bold")
)Significant Relationship between Education and Unemployment: The analysis demonstrated a statistically significant relationship between the GenderRatio and Unemployment_Rate, with a p-value of 0.000814. This indicates that countries with higher gender parity in education tend to have lower unemployment rates. The 95% confidence interval for the GenderRatio coefficient further supports this finding, suggesting that the true effect of gender ratio on unemployment rates is substantial and consistent.
Weak Positive Correlation: The Pearson correlation coefficient between the GenderRatio and Unemployment_Rate was 0.1955, indicating a weak but significant positive correlation. This suggests that as the gender ratio improves, unemployment rates slightly increase, though the effect size is small.
Yearly Comparison: The chi-square test results showed no significant association between the years 2015 and 2020 and the unemployment rate categories. This implies that the distribution of unemployment rates remained consistent over the five-year period.
Strengths
Comprehensive Data Analysis: The study employed robust statistical methods, including linear regression, Pearson correlation, and chi-square tests, to analyze the relationship between education and unemployment, ensuring the results are reliable and valid.
Data Quality and Pre-processing: The datasets were thoroughly cleaned and pre-processed to remove any inconsistencies, missing values, and redundant columns. This ensured the accuracy and reliability of the data used in the analysis.
Effective Visualization: The use of visualizations, such as scatter plots, box plots, and histograms, helped to effectively communicate the key features of the data and support the findings of the analysis.
Limitations
Limited Time Frame: The analysis was limited to data from the years 2015 and 2020, which may not capture longer-term trends or fluctuations in the relationship between education and unemployment.
Assumption Violations: While the assumptions of normality and homoscedasticity were checked, any potential violations could impact the validity of the results. For instance, the skewness observed in the distributions suggests that the data may not fully meet the normality assumption.
External Factors: The study did not account for other external factors, such as economic policies, global events (e.g., the COVID-19 pandemic), or cultural differences, which might influence unemployment rates and educational gender parity.
Directions for Future Investigations:Future studies should include a broader range of years to observe longer-term trends and changes in the relationship between education and unemployment. This would provide a more comprehensive understanding of the dynamics over time. Including other relevant variables, such as economic indicators (GDP, inflation rates), quality of education, government policies, and labor market conditions, could provide a more nuanced analysis of the factors influencing unemployment rates. Conducting region-specific studies could help understand how the relationship between education and unemployment varies across different regions or continents. This would allow for targeted policy recommendations based on regional characteristics.
Conclusion
The investigation has highlighted a significant inverse relationship between the gender ratio for average years of schooling and unemployment rates. Countries where the gender ratio is closer to parity or where women have more years of schooling than men tend to have lower unemployment rates. This underscores the importance of educational policies that promote equal access to education for all genders as a means to enhance employment opportunities and economic stability. The key takeaway from this study is that improving gender equality in education can potentially reduce unemployment rates. Policymakers should focus on ensuring equal educational opportunities for both genders to foster economic growth and reduce unemployment. While this study provides valuable insights, continued research with expanded data and additional variables is necessary to further understand the complex dynamics between education and unemployment rates. By addressing the limitations and pursuing the proposed directions for future research, we can gain a more comprehensive understanding of how educational gender parity impacts unemployment and inform more effective policies to promote both educational equality and economic stability.
3.Correlation_and_Regression.Datacamp.R. (n.d.). Rstudio-Pubs-Static.s3.Amazonaws.com. Retrieved June 2, 2024, from https://rstudio-pubs-static.s3.amazonaws.com/316139_3ae64fa704fa451e95f14e0dd1953ba5.html
Baglin, J. (n.d.). Applied Analytics. Applied Analytics; James Baglin. https://astral-theory-157510.appspot.com/secured/index.html
Car package in R. (2023, July 31). GeeksforGeeks. https://www.geeksforgeeks.org/car-package-in-r/
Chugh, V. (2023, March). Reading and Importing Excel Files into R Tutorial. Www.datacamp.com. https://www.datacamp.com/tutorial/r-tutorial-read-excel-into-r
dplyr Package in R Programming. (2020, May 10). GeeksforGeeks. https://www.geeksforgeeks.org/dplyr-package-in-r-programming/
Gender ratio for average years of schooling. (n.d.). Our World in Data. Retrieved June 2, 2024, from https://ourworldindata.org/grapher/gender-ratios-for-mean-years-of-schooling?tab=table&time=earliest
ggfortify : Extension to ggplot2 to handle some popular packages - R software and data visualization - Easy Guides - Wiki - STHDA. (n.d.). Www.sthda.com. http://www.sthda.com/english/wiki/ggfortify-extension-to-ggplot2-to-handle-some-popular-packages-r-software-and-data-visualization
Grolemund, Y. X., J. J. Allaire, Garrett. (n.d.). 4.2 Slidy presentation | R Markdown: The Definitive Guide. In bookdown.org. https://bookdown.org/yihui/rmarkdown/slidy-presentation.html
Sachs, M. (n.d.). Introduction to knitr. Sachsmc.github.io. Retrieved June 2, 2024, from https://sachsmc.github.io/knit-git-markr-guide/knitr/knit.html
Unemployment rate. (n.d.). Our World in Data. https://ourworldindata.org/grapher/unemployment-rate?tab=table
Weingessel, A. (n.d.). e1071 package - RDocumentation. Www.rdocumentation.org. Retrieved June 2, 2024, from https://www.rdocumentation.org/packages/e1071/versions/1.7-14
Wickham, H. (2019). Create Elegant Data Visualisations Using the Grammar of Graphics. Tidyverse.org. https://ggplot2.tidyverse.org/
Zhu, H. (2021, February 19). Create Awesome HTML Table with knitr::kable and kableExtra. Cran.r-Project.org. https://cran.r-project.org/web/packages/kableExtra/vignettes/awesome_table_in_html.html