Research question Is there a difference in body weight (kg) before versus after all participants tried the keto diet?

library(ggpubr)
## Loading required package: ggplot2
library(effsize)
library(rstatix)
## 
## Attaching package: 'rstatix'
## The following object is masked from 'package:stats':
## 
##     filter
library(readxl)

Import dataset

A6Q2 <- read_excel("A6Q2.xlsx")

Create before versus after variables

Before <- A6Q2$Before
After <- A6Q2$After

Calculate difference scores

Differences <- After - Before

Calculate descriptive statistics

mean(Before, na.rm = TRUE)
## [1] 76.13299
median(Before, na.rm = TRUE)
## [1] 75.95988
sd(Before, na.rm = TRUE)
## [1] 7.781323
mean(After, na.rm = TRUE)
## [1] 57.17874
median(After, na.rm = TRUE)
## [1] 58.36459
sd(After, na.rm = TRUE)
## [1] 14.39364

Create histogram

hist(Differences, main = "Histogram of Difference Scores", xlab = "Value", ylab = "Frequency", col = "blue", border = "black", breaks = 20)

Histogram of Difference Scores The difference scores look abnormally distributed. The data is negatively skewed. The data does not have a proper bell curve. Boxplot to check normality and outliers

boxplot(Differences, main = "Distribution of Score Differences (After - Before)", ylab = "Difference in Scores", col = "blue", border = "darkblue")

Boxplot There is one dot outside the boxplot. The dot is not close to the whiskers. Based on these findings, the boxplot is not normal. Update made to the anlaysis on where this was normal. Statistically check normality

shapiro.test(Differences)
## 
##  Shapiro-Wilk normality test
## 
## data:  Differences
## W = 0.89142, p-value = 0.02856

Shaprio-Wilk Difference Scores The data is abnormally distributed, (p=0.02856).

wilcox.test(Before, After, paired = TRUE, na.action = na.omit)
## 
##  Wilcoxon signed rank exact test
## 
## data:  Before and After
## V = 210, p-value = 1.907e-06
## alternative hypothesis: true location shift is not equal to 0
df_long <- data.frame(id = rep(1:length(Before) , 2), time = rep(c("Before", "After"), each = length(Before)), score = c(Before, After)) 
wilcox_effsize(df_long, score ~ time, paired = TRUE)
## # A tibble: 1 × 7
##   .y.   group1 group2 effsize    n1    n2 magnitude
## * <chr> <chr>  <chr>    <dbl> <int> <int> <ord>    
## 1 score After  Before   0.877    20    20 large

A Wilcoxon Signed-Rank Test was conducted to determine if there was a difference in Outcome Variable before Independent Variable versus after Independent Variable. Before scores (Mdn = 75.96 were significantly different from after scores (Mdn = 58.36), V = 210, p < .001. Update made to rounding. The effect size was large, r = .877