Step 1: Population, Histogram, and Parameters

library(Stat2Data)
library(dplyr)
library(ggplot2)

data(Diamonds)
mu <- mean(Diamonds$TotalPrice, na.rm = TRUE)
sigma <- sd(Diamonds$TotalPrice, na.rm = TRUE)

ggplot(Diamonds, aes(x = TotalPrice)) +
  geom_histogram(binwidth = 500, fill = "lightblue", color = "black") +
  labs(title = "Population Distribution of TotalPrice",
       x = "TotalPrice", y = "Count")

mu
## [1] 7450.012
sigma
## [1] 7780.894

Student interpretation:
The population distribution is right‑skewed and bimodal.


Step 2: Sampling Distribution Simulation

set.seed(1)
n <- 30

sample_means <- replicate(1000, {
  mean(sample(Diamonds$TotalPrice, n, replace = TRUE))
})

ggplot(data.frame(sample_means), aes(x = sample_means)) +
  geom_histogram(binwidth = 100, fill = "lightblue", color = "black") +
  labs(title = "Sampling Distribution of Sample Means",
       x = "Sample Mean", y = "Count")

mean(sample_means)
## [1] 7490.552
sd(sample_means)
## [1] 1477.04

Student interpretation:
The distribution of sample means is more symmetric looking compared to the population distribution earlier.
The graph looks slightly symmetric and unimodal, maybe slightly skewed right.


Step 3: CLT Model and Overlay

SE <- sigma / sqrt(n)
SE
## [1] 1420.59
ggplot(data.frame(sample_means), aes(x = sample_means)) +
  geom_histogram(aes(y = ..density..), binwidth = 100,
                 fill = "lightblue", color = "black") +
  stat_function(fun = dnorm,
                args = list(mean = mu, sd = SE),
                color = "red", size = 1) +
  labs(title = "Sampling Distribution with CLT Normal Curve",
       x = "Sample Mean", y = "Density")

Student interpretation:
The mu is about 7450.012, while the mean(sample_means) is 7459.631.
So the CLT curve is slightly moved to the right of the sample mean curve, but they are close.
The SE is 1420.59 and the sd(sample_means) is 1405.478.
So the values are further apart — the CLT curve has a bigger SD, while the sample mean graph has a lower SD.
The CLT curve doesn’t fit the best, with some of the sample mean values being way above the red line.


Step 4: Probability Above 10,000

CLT_prob <- 1 - pnorm(10000, mean = mu, sd = SE)
CLT_prob
## [1] 0.03632525
sim_prob <- mean(sample_means > 10000)
sim_prob
## [1] 0.057

Student interpretation:
They are close with a difference of about .014 because they are similar in shape but still have differences.
The sampling means come from repeated samples of size 30, while the CLT curve is an idealized normal model.


Step 5: 90% Cutoff Value

cutoff_90 <- qnorm(0.10, mean = mu, sd = SE)
cutoff_90
## [1] 5629.452
quantile(sample_means, 0.10)
##      10% 
## 5726.127

Student interpretation:
The cutoff value means that 90% of samples will have a mean above 5816.595.
Cutoff at 90 because you want the above number, not all the values below.