title: “Sampling Distributions and the CLT: Diamond Prices” author: “[Yasmeen Shamsulddin]” date: “2026-09-30” output: html_document —
In this report I treat the Diamonds dataset from the Stat2Data package as a population, and I study the variable TotalPrice, measured in dollars. I use simulation and the Central Limit Theorem to study the sampling distribution of the sample mean for samples of size n = 30.
library(Stat2Data)
library(dplyr)
library(ggplot2)
data(Diamonds)
ggplot(Diamonds, aes(x = TotalPrice)) +
geom_histogram(bins = 30, fill = "lightblue", color = "black") +
labs(title = "Diamond Prices (Population)", x = "Total price ($)", y = "Count") +
theme_minimal()
mu <- mean(Diamonds$TotalPrice)
sigma <- sd(Diamonds$TotalPrice)
mu
## [1] 7450.012
sigma
## [1] 7780.894
The shape of the data is unimodal and skewed right. Mu is population mean and sigma is standard deviation.
n <- 30
B <- 1000
set.seed(123)
sample_means <- replicate(B, mean(sample(Diamonds$TotalPrice, size = n, replace = TRUE)))
mean(sample_means)
## [1] 7459.631
sd(sample_means)
## [1] 1405.478
ggplot(data.frame(xbar = sample_means), aes(x = xbar)) +
geom_histogram(bins = 30, fill = "lightgreen", color = "black") +
labs(title = "Simulated Sampling Distribution of Mean Diamond Price (n = 30)",
x = "Sample mean price ($)", y = "Count") +
theme_minimal()
The shape is also unimodal but it’s symmetric and more spread out.
SE <- sigma / sqrt(n)
SE
## [1] 1420.59
ggplot(data.frame(xbar = sample_means), aes(x = xbar)) +
geom_histogram(aes(y = after_stat(density)), bins = 30,
fill = "lightgreen", color = "black") +
stat_function(fun = dnorm, args = list(mean = mu, sd = SE),
color = "red", linewidth = 1) +
labs(title = "Simulated Mean Diamond Prices with the CLT's Normal Model",
x = "Sample mean price ($)", y = "Density") +
theme_minimal()
The CLT model matches really well and is pretty accurate. The mean and mu are also very close as well as the sd and sigma.
What is the probability that a random sample of 30 diamonds has a mean price greater than $10,000?
1 - pnorm(10000, mean = mu, sd = SE) # CLT answer
## [1] 0.03632525
mean(sample_means > 10000) # simulation answer
## [1] 0.05
The CLT answer was 0.03632525 and the simulation answer was 0.05. Although they are similar, they’re not identical because the CLT model doesn’t take the original population into account since it’s more theoretical.
90% of samples of 30 diamonds will have a mean price above what value?
qnorm(0.10, mean = mu, sd = SE) # CLT answer
## [1] 5629.452
quantile(sample_means, 0.10) # simulation check
## 10%
## 5816.595
The average price of a sample of 30 diamonds is $5629.45.
The sampling distribution and CLT are very accurate to each other so that’s good to know. I’d want to investigate CLT models more.