This report analyzes customer behavior at a fictional online retail store. The two datasets used here are synthetic: they were generated with an AI tool (Claude) and contain no real people or private information. Patterns were built into the data on purpose so that the statistical tests could be demonstrated.
Dataset 1 (between-subjects, N = 150) contains different customers in different groups. It is used for the chi-square test, the Pearson correlation, and the independent-samples t-test.
Dataset 2 (within-subjects, N = 60) contains the same customers measured twice, before and after a checkout redesign. It is used for the dependent (paired) t-test.
head(d1)
## customer_id membership_type repeat_purchase monthly_visits monthly_spending
## 1 C001 Standard No 15 153.80
## 2 C002 Standard Yes 12 100.45
## 3 C003 Premium Yes 7 59.78
## 4 C004 Premium Yes 12 114.99
## 5 C005 Premium Yes 13 87.81
## 6 C006 Standard Yes 16 91.10
## contacted_support satisfaction_score
## 1 No 73
## 2 No 78
## 3 No 85
## 4 No 66
## 5 No 93
## 6 No 79
head(d2)
## customer_id satisfaction_before satisfaction_after
## 1 W001 65 80
## 2 W002 37 38
## 3 W003 51 54
## 4 W004 55 63
## 5 W005 57 72
## 6 W006 64 75
All tests use a significance level of α = .05.
Is membership type associated with whether a customer made a repeat purchase?
A chi-square test of independence is appropriate because both variables are categorical. The assumption that all expected counts are at least 5 is checked below.
chi_table <- table(d1$membership_type, d1$repeat_purchase)
chi_table
##
## No Yes
## Free 39 21
## Premium 4 31
## Standard 28 27
# Row percentages
round(prop.table(chi_table, margin = 1) * 100, 1)
##
## No Yes
## Free 65.0 35.0
## Premium 11.4 88.6
## Standard 50.9 49.1
chi_result <- chisq.test(chi_table)
chi_result
##
## Pearson's Chi-squared test
##
## data: chi_table
## X-squared = 25.894, df = 2, p-value = 2.384e-06
# Assumption check: expected counts
chi_result$expected
##
## No Yes
## Free 28.40000 31.60000
## Premium 16.56667 18.43333
## Standard 26.03333 28.96667
# Effect size: Cramer's V
n_chi <- sum(chi_table)
k_chi <- min(nrow(chi_table) - 1, ncol(chi_table) - 1)
cramers_v <- sqrt(as.numeric(chi_result$statistic) / (n_chi * k_chi))
cramers_v
## [1] 0.4154816
barplot(t(prop.table(chi_table, margin = 1)),
beside = TRUE, legend.text = TRUE,
main = "Repeat Purchase by Membership Type",
xlab = "Membership Type", ylab = "Proportion",
col = c("gray70", "steelblue"))
All expected counts were above 5 (the smallest was 16.57), so the assumption was met.
A chi-square test of independence showed a significant association between membership type and repeat purchase, χ²(2, N = 150) = 25.89, p < .001, Cramér’s V = .42.
Membership type and repeat purchasing are related. We reject the null hypothesis. Premium members were the most likely to buy again (88.6%), followed by Standard members (49.1%), while Free members were the least likely (35.0%). Cramér’s V of .42 indicates a large effect. In plain terms, customers with better memberships are much more likely to come back and buy again.
Is the number of monthly website visits related to monthly spending?
A Pearson correlation is appropriate because both variables are numeric. Assumptions are normality of each variable (Shapiro-Wilk test) and a linear relationship (scatterplot).
# Normality checks (p > .05 = normal)
shapiro.test(d1$monthly_visits)
##
## Shapiro-Wilk normality test
##
## data: d1$monthly_visits
## W = 0.99174, p-value = 0.5363
shapiro.test(d1$monthly_spending)
##
## Shapiro-Wilk normality test
##
## data: d1$monthly_spending
## W = 0.99278, p-value = 0.6539
cor_result <- cor.test(d1$monthly_visits, d1$monthly_spending,
method = "pearson")
cor_result
##
## Pearson's product-moment correlation
##
## data: d1$monthly_visits and d1$monthly_spending
## t = 11.26, df = 148, p-value < 2.2e-16
## alternative hypothesis: true correlation is not equal to 0
## 95 percent confidence interval:
## 0.5823912 0.7570995
## sample estimates:
## cor
## 0.6792546
# Effect size: r squared
cor_result$estimate^2
## cor
## 0.4613868
plot(d1$monthly_visits, d1$monthly_spending,
main = "Monthly Visits vs. Monthly Spending",
xlab = "Monthly Visits", ylab = "Monthly Spending ($)",
pch = 19, col = "steelblue")
abline(lm(monthly_spending ~ monthly_visits, data = d1), col = "red", lwd = 2)
Both variables met the normality assumption (visits: p = .536; spending: p = .654), and the scatterplot shows a linear pattern, so Pearson is appropriate.
A Pearson correlation showed a strong, positive, significant relationship between monthly website visits and monthly spending, r(148) = .68, p < .001, 95% CI [.58, .76]. Visits explained about 46% of the variance in spending (r² = .46).
We reject the null hypothesis. Customers who visit the website more often tend to spend more money. The relationship is strong and positive. Note that correlation does not prove that visiting causes spending; it only shows that the two go up together.
Do customers who contacted customer support differ in satisfaction from customers who did not?
An independent-samples t-test (Welch version) is appropriate because two separate groups of customers are compared on a numeric outcome. Assumptions checked are normality within each group (Shapiro-Wilk) and equal variances (F test).
# Descriptive statistics
tapply(d1$satisfaction_score, d1$contacted_support, mean)
## No Yes
## 74.73 65.22
tapply(d1$satisfaction_score, d1$contacted_support, sd)
## No Yes
## 10.193933 8.748854
table(d1$contacted_support)
##
## No Yes
## 100 50
# Normality within each group (p > .05 = normal)
shapiro.test(d1$satisfaction_score[d1$contacted_support == "Yes"])
##
## Shapiro-Wilk normality test
##
## data: d1$satisfaction_score[d1$contacted_support == "Yes"]
## W = 0.95709, p-value = 0.06712
shapiro.test(d1$satisfaction_score[d1$contacted_support == "No"])
##
## Shapiro-Wilk normality test
##
## data: d1$satisfaction_score[d1$contacted_support == "No"]
## W = 0.98859, p-value = 0.5525
# Equal variances (p > .05 = equal)
var.test(satisfaction_score ~ contacted_support, data = d1)
##
## F test to compare two variances
##
## data: satisfaction_score by contacted_support
## F = 1.3576, num df = 99, denom df = 49, p-value = 0.237
## alternative hypothesis: true ratio of variances is not equal to 1
## 95 percent confidence interval:
## 0.8159788 2.1686570
## sample estimates:
## ratio of variances
## 1.357629
# Welch t-test
ind_t <- t.test(satisfaction_score ~ contacted_support, data = d1,
var.equal = FALSE)
ind_t
##
## Welch Two Sample t-test
##
## data: satisfaction_score by contacted_support
## t = 5.9322, df = 112.46, p-value = 3.366e-08
## alternative hypothesis: true difference in means between group No and group Yes is not equal to 0
## 95 percent confidence interval:
## 6.333753 12.686247
## sample estimates:
## mean in group No mean in group Yes
## 74.73 65.22
# Effect size: Cohen's d
g_yes <- d1$satisfaction_score[d1$contacted_support == "Yes"]
g_no <- d1$satisfaction_score[d1$contacted_support == "No"]
sd_pooled <- sqrt(((length(g_yes) - 1) * var(g_yes) +
(length(g_no) - 1) * var(g_no)) /
(length(g_yes) + length(g_no) - 2))
cohens_d_ind <- abs(mean(g_yes) - mean(g_no)) / sd_pooled
cohens_d_ind
## [1] 0.9764596
boxplot(satisfaction_score ~ contacted_support, data = d1,
main = "Satisfaction by Support Contact",
xlab = "Contacted Support", ylab = "Satisfaction Score",
col = c("lightblue", "lightgreen"))
Normality was met in both groups (No: p = .553; Yes: p = .067), and variances were equal (p = .237), so the assumptions were satisfied.
An independent-samples t-test showed that customers who did not contact support (M = 74.73, SD = 10.19) reported significantly higher satisfaction than customers who did contact support (M = 65.22, SD = 8.75), t(112.46) = 5.93, p < .001, 95% CI [6.33, 12.69], Cohen’s d = 0.98.
We reject the null hypothesis. Customers who contacted support were, on average, about 9.5 points less satisfied than those who did not. Cohen’s d of 0.98 is a large effect. This may suggest that customers who need support are having problems, which lowers their satisfaction.
Did customer satisfaction change from before to after the checkout redesign?
The same 60 customers were measured both times.
A dependent (paired) t-test is appropriate because the same customers were measured twice. The assumption is that the difference scores are normally distributed (Shapiro-Wilk test and histogram).
# Descriptive statistics
mean(d2$satisfaction_before)
## [1] 63.83333
sd(d2$satisfaction_before)
## [1] 11.50117
mean(d2$satisfaction_after)
## [1] 69.76667
sd(d2$satisfaction_after)
## [1] 14.35549
# Difference scores (after minus before)
d2$difference <- d2$satisfaction_after - d2$satisfaction_before
mean(d2$difference)
## [1] 5.933333
sd(d2$difference)
## [1] 5.871121
# Normality of the difference scores (p > .05 = normal)
shapiro.test(d2$difference)
##
## Shapiro-Wilk normality test
##
## data: d2$difference
## W = 0.97678, p-value = 0.3078
# Paired t-test
paired_t <- t.test(d2$satisfaction_after, d2$satisfaction_before,
paired = TRUE)
paired_t
##
## Paired t-test
##
## data: d2$satisfaction_after and d2$satisfaction_before
## t = 7.828, df = 59, p-value = 1.069e-10
## alternative hypothesis: true mean difference is not equal to 0
## 95 percent confidence interval:
## 4.416662 7.450005
## sample estimates:
## mean difference
## 5.933333
# Effect size: Cohen's d for paired data
cohens_d_paired <- mean(d2$difference) / sd(d2$difference)
cohens_d_paired
## [1] 1.010596
hist(d2$difference,
main = "Change in Satisfaction (After - Before)",
xlab = "Difference Score", col = "steelblue", border = "white")
The difference scores were normally distributed (Shapiro-Wilk p = .308), so the assumption was met.
A paired-samples t-test showed that satisfaction was significantly higher after the redesign (M = 69.77, SD = 14.36) than before the redesign (M = 63.83, SD = 11.50), t(59) = 7.83, p < .001, mean difference = 5.93, 95% CI [4.42, 7.45], Cohen’s d = 1.01.
We reject the null hypothesis. On average, customers’ satisfaction increased by about 6 points after the checkout redesign. Cohen’s d of 1.01 is a large effect. Because the same customers were measured both times, this suggests the redesign was followed by higher satisfaction. Without a control group, we cannot be certain the redesign itself caused the increase.
All four analyses produced statistically significant results:
| Analysis | Test | Key Result |
|---|---|---|
| Membership type and repeat purchase | Chi-square | χ²(2, N = 150) = 25.89, p < .001, V = .42 |
| Visits and spending | Pearson correlation | r(148) = .68, p < .001 |
| Support contact and satisfaction | Independent t-test | t(112.46) = 5.93, p < .001, d = 0.98 |
| Satisfaction before and after redesign | Paired t-test | t(59) = 7.83, p < .001, d = 1.01 |
Higher membership levels were linked to more repeat purchases, more visits were linked to more spending, customers who contacted support were less satisfied, and satisfaction rose after the checkout redesign.
Limitations: The data are synthetic, and the effects were built in on purpose, so real-world data would likely show weaker and messier patterns. These results also show associations and differences, not proof of cause and effect.