MATH1324 Applied Analytics

Applied data project 2

Jayamanna Mohottilage Irushi Chanika Wickramarathna (s4024983) & Samaraweera Mudalige Dona Bawanthi Nayanathari (s4040017)

Last updated: 15 October, 2023

Introduction

Introduction Cont.

Problem Statement

Data

This data set contains a diverse variety of features that provide a full picture of students’ accomplishments, behaviors, and characteristics. Data consists of variables of different types characters as well as numeric.

For this analysis, we will use open data obtained from the https://www.https://www.kaggle.com/datasets/spscientist/students-performance-in-exams/data

Descriptive Statistics and Visualisation

#check current directory
getwd() 
## [1] "E:/RMIT/APPLIED ANALYTICS/Assignments/final"
# set directory
setwd("E:/RMIT/APPLIED ANALYTICS/Assignments/final") 

# load student performance data
st_performance <- read.csv("E:/RMIT/APPLIED ANALYTICS/Assignments/final/StudentsPerformance.csv") 

# view the data set
View(st_performance) 

# produce the first 1-6 data in the data set
print.data.frame(head(st_performance)) 
##   gender race.ethnicity parental.level.of.education        lunch
## 1 female        group B           bachelor's degree     standard
## 2 female        group C                some college     standard
## 3 female        group B             master's degree     standard
## 4   male        group A          associate's degree free/reduced
## 5   male        group C                some college     standard
## 6 female        group B          associate's degree     standard
##   test.preparation.course math.score reading.score writing.score
## 1                    none         72            72            74
## 2               completed         69            90            88
## 3                    none         90            95            93
## 4                    none         47            57            44
## 5                    none         76            78            75
## 6                    none         71            83            78
# checking variable types
str(st_performance) 
## 'data.frame':    1000 obs. of  8 variables:
##  $ gender                     : chr  "female" "female" "female" "male" ...
##  $ race.ethnicity             : chr  "group B" "group C" "group B" "group A" ...
##  $ parental.level.of.education: chr  "bachelor's degree" "some college" "master's degree" "associate's degree" ...
##  $ lunch                      : chr  "standard" "standard" "standard" "free/reduced" ...
##  $ test.preparation.course    : chr  "none" "completed" "none" "none" ...
##  $ math.score                 : int  72 69 90 47 76 71 88 40 64 38 ...
##  $ reading.score              : int  72 90 95 57 78 83 95 43 64 60 ...
##  $ writing.score              : int  74 88 93 44 75 78 92 39 67 50 ...
# gender converted from character to factor using the as.factor() function and the levels() function
st_performance$gender <- as.factor(st_performance$gender)
levels(st_performance$gender) 
## [1] "female" "male"
# test preparation course converted from character to factor using the as.factor() function and the levels() function
st_performance$test.preparation.course <- as.factor(st_performance$test.preparation.course)
levels(st_performance$test.preparation.course) 
## [1] "completed" "none"
# check missing values
sum(is.na(st_performance)) 
## [1] 0
# print gender count
gender_count <- summary(st_performance$gender) 
gender_count
## female   male 
##    518    482
# print test preparation count on "completed" & "none"
test_preparation_count <- summary(st_performance$test.preparation.course)  
test_preparation_count
## completed      none 
##       358       642

Decsriptive Statistics Cont.

# Statistical summary of test preparation course with respect to male and female.

# Convert the 'test.preparation.course' variable to numeric 

st_performance$test.preparation.course <- as.numeric(st_performance$test.preparation.course)

#  perform the summarization

st_performance %>%
  group_by(gender) %>%
  summarise(Min = min(test.preparation.course, na.rm = TRUE),
            Q1 = quantile(test.preparation.course, probs = 0.25, na.rm = TRUE),
            Median = median(test.preparation.course, na.rm = TRUE),
            Q3 = quantile(test.preparation.course, probs = 0.75, na.rm = TRUE),
            Max = max(test.preparation.course, na.rm = TRUE),
            Mean = mean(test.preparation.course, na.rm = TRUE),
            SD = sd(test.preparation.course, na.rm = TRUE),
            n = n(),
            Missing = sum(is.na(test.preparation.course))) -> table1
knitr::kable(table1)
gender Min Q1 Median Q3 Max Mean SD n Missing
female 1 1 2 2 2 1.644788 0.4790402 518 0
male 1 1 2 2 2 1.639004 0.4807883 482 0
# visualize the relationship between gender and test preparation course

table <- 
  table(st_performance$test.preparation.course, st_performance$gender)%>% 
  prop.table(margin = 2)
knitr::kable(table)
female male
0.3552124 0.3609959
0.6447876 0.6390041
# Descriptive visualization - Bar plot

barplot(table, ylab = "Preparation within group", ylim = c(0, .9), 
        legend = rownames(table), beside = TRUE, 
        args.legend = c(x = "top", horiz = TRUE, title = "Test Preparation Course"), 
        xlab = "Gender", col = c("yellow", "green"), border = "#69b3a2")

Hypothesis Testing

# calculate the chi-squared value using the chisq.test() function.

chi1 <- chisq.test(
  table(st_performance$gender, st_performance$test.preparation.course))
chi1
## 
##  Pearson's Chi-squared test with Yates' continuity correction
## 
## data:  table(st_performance$gender, st_performance$test.preparation.course)
## X-squared = 0.015529, df = 1, p-value = 0.9008
# check the observed and expected values

chi1$observed
##         
##            1   2
##   female 184 334
##   male   174 308
chi1$expected
##         
##                1       2
##   female 185.444 332.556
##   male   172.556 309.444
# find the critical value

qchisq(p = .95, df = 1)
## [1] 3.841459
# find the p-value

pchisq(q = 0.015529, df = 1, lower.tail = FALSE)
## [1] 0.900828

Discussion

The experiment began with the question of whether females or males were more likely to complete a test preparation course. The Chi-square Test of Association was conducted to investigate if there was a relationship between gender and test preparation course. The findings failed to reject the Null Hypothesis since they demonstrated no correlation between the two variables and were not statistically significant. As a result, they are inconclusive. The investigation’s strength is its relatively high sample size (n = 1000). However, there are numerous limits to this analysis. Some examples include not knowing if the students are from the same school, not knowing the students’ ages, the students being only from the United States, and so on. Future research investigating the relationship between gender and test preparation course should include additional information such as age, school, and socioeconomic level to produce a clearer and more significant result. There was no correlation between gender and completion of a test preparation course in this study. The findings cannot be generalized to the whole population due to a lack of statistical significance.

References