title: “Sampling Distributions and the CLT: Diamond Prices” author: “[Yasmeen Shamsulddin]” date: “2026-09-30” output: html_document —

Introduction

In this report I treat the Diamonds dataset from the Stat2Data package as a population, and I study the variable TotalPrice, measured in dollars. I use simulation and the Central Limit Theorem to study the sampling distribution of the sample mean for samples of size n = 30.

The Population

library(Stat2Data)
library(dplyr)
library(ggplot2)

data(Diamonds)

ggplot(Diamonds, aes(x = TotalPrice)) +
  geom_histogram(bins = 30, fill = "lightblue", color = "black") +
  labs(title = "Diamond Prices (Population)", x = "Total price ($)", y = "Count") +
  theme_minimal()

mu <- mean(Diamonds$TotalPrice)
sigma <- sd(Diamonds$TotalPrice)
mu
## [1] 7450.012
sigma
## [1] 7780.894

The shape of the data is unimodal and skewed right. Mu is population mean and sigma is standard deviation.

A Simulated Sampling Distribution

n <- 30
B <- 1000
set.seed(123)

sample_means <- replicate(B, mean(sample(Diamonds$TotalPrice, size = n, replace = TRUE)))

mean(sample_means)
## [1] 7459.631
sd(sample_means)
## [1] 1405.478
ggplot(data.frame(xbar = sample_means), aes(x = xbar)) +
  geom_histogram(bins = 30, fill = "lightgreen", color = "black") +
  labs(title = "Simulated Sampling Distribution of Mean Diamond Price (n = 30)",
       x = "Sample mean price ($)", y = "Count") +
  theme_minimal()

The shape is also unimodal but it’s symmetric and more spread out.

The CLT’s Prediction

SE <- sigma / sqrt(n)
SE
## [1] 1420.59
ggplot(data.frame(xbar = sample_means), aes(x = xbar)) +
  geom_histogram(aes(y = after_stat(density)), bins = 30,
                 fill = "lightgreen", color = "black") +
  stat_function(fun = dnorm, args = list(mean = mu, sd = SE),
                color = "red", linewidth = 1) +
  labs(title = "Simulated Mean Diamond Prices with the CLT's Normal Model",
       x = "Sample mean price ($)", y = "Density") +
  theme_minimal()

The CLT model matches really well and is pretty accurate. The mean and mu are also very close as well as the sd and sigma.

A Probability Using pnorm()

What is the probability that a random sample of 30 diamonds has a mean price greater than $10,000?

1 - pnorm(10000, mean = mu, sd = SE)   # CLT answer
## [1] 0.03632525
mean(sample_means > 10000)             # simulation answer
## [1] 0.05

The CLT answer was 0.03632525 and the simulation answer was 0.05. Although they are similar, they’re not identical because the CLT model doesn’t take the original population into account since it’s more theoretical.

A Cutoff Using qnorm()

90% of samples of 30 diamonds will have a mean price above what value?

qnorm(0.10, mean = mu, sd = SE)   # CLT answer
## [1] 5629.452
quantile(sample_means, 0.10)      # simulation check
##      10% 
## 5816.595

The average price of a sample of 30 diamonds is $5629.45.

Conclusion

The sampling distribution and CLT are very accurate to each other so that’s good to know. I’d want to investigate CLT models more.