Jayamanna Mohottilage Irushi Chanika Wickramarathna (s4024983) & Samaraweera Mudalige Dona Bawanthi Nayanathari (s4040017)
Last updated: 15 October, 2023
This data set contains a diverse variety of features that provide a full picture of students’ accomplishments, behaviors, and characteristics. Data consists of variables of different types characters as well as numeric.
For this analysis, we will use open data obtained from the https://www.https://www.kaggle.com/datasets/spscientist/students-performance-in-exams/data
## [1] "E:/RMIT/APPLIED ANALYTICS/Assignments/final"
# set directory
setwd("E:/RMIT/APPLIED ANALYTICS/Assignments/final")
# load student performance data
st_performance <- read.csv("E:/RMIT/APPLIED ANALYTICS/Assignments/final/StudentsPerformance.csv")
# view the data set
View(st_performance)
# produce the first 1-6 data in the data set
print.data.frame(head(st_performance)) ## gender race.ethnicity parental.level.of.education lunch
## 1 female group B bachelor's degree standard
## 2 female group C some college standard
## 3 female group B master's degree standard
## 4 male group A associate's degree free/reduced
## 5 male group C some college standard
## 6 female group B associate's degree standard
## test.preparation.course math.score reading.score writing.score
## 1 none 72 72 74
## 2 completed 69 90 88
## 3 none 90 95 93
## 4 none 47 57 44
## 5 none 76 78 75
## 6 none 71 83 78
## 'data.frame': 1000 obs. of 8 variables:
## $ gender : chr "female" "female" "female" "male" ...
## $ race.ethnicity : chr "group B" "group C" "group B" "group A" ...
## $ parental.level.of.education: chr "bachelor's degree" "some college" "master's degree" "associate's degree" ...
## $ lunch : chr "standard" "standard" "standard" "free/reduced" ...
## $ test.preparation.course : chr "none" "completed" "none" "none" ...
## $ math.score : int 72 69 90 47 76 71 88 40 64 38 ...
## $ reading.score : int 72 90 95 57 78 83 95 43 64 60 ...
## $ writing.score : int 74 88 93 44 75 78 92 39 67 50 ...
# gender converted from character to factor using the as.factor() function and the levels() function
st_performance$gender <- as.factor(st_performance$gender)
levels(st_performance$gender) ## [1] "female" "male"
# test preparation course converted from character to factor using the as.factor() function and the levels() function
st_performance$test.preparation.course <- as.factor(st_performance$test.preparation.course)
levels(st_performance$test.preparation.course) ## [1] "completed" "none"
## [1] 0
## female male
## 518 482
# print test preparation count on "completed" & "none"
test_preparation_count <- summary(st_performance$test.preparation.course)
test_preparation_count## completed none
## 358 642
# Statistical summary of test preparation course with respect to male and female.
# Convert the 'test.preparation.course' variable to numeric
st_performance$test.preparation.course <- as.numeric(st_performance$test.preparation.course)
# perform the summarization
st_performance %>%
group_by(gender) %>%
summarise(Min = min(test.preparation.course, na.rm = TRUE),
Q1 = quantile(test.preparation.course, probs = 0.25, na.rm = TRUE),
Median = median(test.preparation.course, na.rm = TRUE),
Q3 = quantile(test.preparation.course, probs = 0.75, na.rm = TRUE),
Max = max(test.preparation.course, na.rm = TRUE),
Mean = mean(test.preparation.course, na.rm = TRUE),
SD = sd(test.preparation.course, na.rm = TRUE),
n = n(),
Missing = sum(is.na(test.preparation.course))) -> table1
knitr::kable(table1)| gender | Min | Q1 | Median | Q3 | Max | Mean | SD | n | Missing |
|---|---|---|---|---|---|---|---|---|---|
| female | 1 | 1 | 2 | 2 | 2 | 1.644788 | 0.4790402 | 518 | 0 |
| male | 1 | 1 | 2 | 2 | 2 | 1.639004 | 0.4807883 | 482 | 0 |
# visualize the relationship between gender and test preparation course
table <-
table(st_performance$test.preparation.course, st_performance$gender)%>%
prop.table(margin = 2)
knitr::kable(table)| female | male |
|---|---|
| 0.3552124 | 0.3609959 |
| 0.6447876 | 0.6390041 |
# Descriptive visualization - Bar plot
barplot(table, ylab = "Preparation within group", ylim = c(0, .9),
legend = rownames(table), beside = TRUE,
args.legend = c(x = "top", horiz = TRUE, title = "Test Preparation Course"),
xlab = "Gender", col = c("yellow", "green"), border = "#69b3a2")The Null Hypothesis is as follows:
H0:There is no association in the population between gender and completion of test preparation course.
The Alternative Hypothesis is as follows:
HA:There is an association in the population between gender and completion of test preparation course.
The Chi-square Test of Association was chosen as the hypothesis test for this data set because we are interested in the link between gender and exam preparation course, both of which are categorical variables.
# calculate the chi-squared value using the chisq.test() function.
chi1 <- chisq.test(
table(st_performance$gender, st_performance$test.preparation.course))
chi1##
## Pearson's Chi-squared test with Yates' continuity correction
##
## data: table(st_performance$gender, st_performance$test.preparation.course)
## X-squared = 0.015529, df = 1, p-value = 0.9008
##
## 1 2
## female 184 334
## male 174 308
##
## 1 2
## female 185.444 332.556
## male 172.556 309.444
## [1] 3.841459
## [1] 0.900828
The experiment began with the question of whether females or males were more likely to complete a test preparation course. The Chi-square Test of Association was conducted to investigate if there was a relationship between gender and test preparation course. The findings failed to reject the Null Hypothesis since they demonstrated no correlation between the two variables and were not statistically significant. As a result, they are inconclusive. The investigation’s strength is its relatively high sample size (n = 1000). However, there are numerous limits to this analysis. Some examples include not knowing if the students are from the same school, not knowing the students’ ages, the students being only from the United States, and so on. Future research investigating the relationship between gender and test preparation course should include additional information such as age, school, and socioeconomic level to produce a clearer and more significant result. There was no correlation between gender and completion of a test preparation course in this study. The findings cannot be generalized to the whole population due to a lack of statistical significance.