library(Stat2Data)
library(dplyr)
library(ggplot2)
data(Diamonds)
mu <- mean(Diamonds$TotalPrice, na.rm = TRUE)
sigma <- sd(Diamonds$TotalPrice, na.rm = TRUE)
ggplot(Diamonds, aes(x = TotalPrice)) +
geom_histogram(binwidth = 500, fill = "lightblue", color = "black") +
labs(title = "Population Distribution of TotalPrice",
x = "TotalPrice", y = "Count")
mu
## [1] 7450.012
sigma
## [1] 7780.894
Student interpretation:
The population distribution is right‑skewed and bimodal.
set.seed(1)
n <- 30
sample_means <- replicate(1000, {
mean(sample(Diamonds$TotalPrice, n, replace = TRUE))
})
ggplot(data.frame(sample_means), aes(x = sample_means)) +
geom_histogram(binwidth = 100, fill = "lightblue", color = "black") +
labs(title = "Sampling Distribution of Sample Means",
x = "Sample Mean", y = "Count")
mean(sample_means)
## [1] 7490.552
sd(sample_means)
## [1] 1477.04
Student interpretation:
The distribution of sample means is more symmetric looking compared to
the population distribution earlier.
The graph looks slightly symmetric and unimodal, maybe slightly skewed
right.
SE <- sigma / sqrt(n)
SE
## [1] 1420.59
ggplot(data.frame(sample_means), aes(x = sample_means)) +
geom_histogram(aes(y = ..density..), binwidth = 100,
fill = "lightblue", color = "black") +
stat_function(fun = dnorm,
args = list(mean = mu, sd = SE),
color = "red", size = 1) +
labs(title = "Sampling Distribution with CLT Normal Curve",
x = "Sample Mean", y = "Density")
Student interpretation:
The mu is about 7450.012, while the mean(sample_means) is
7459.631.
So the CLT curve is slightly moved to the right of the sample mean
curve, but they are close.
The SE is 1420.59 and the sd(sample_means) is 1405.478.
So the values are further apart — the CLT curve has a bigger SD, while
the sample mean graph has a lower SD.
The CLT curve doesn’t fit the best, with some of the sample mean values
being way above the red line.
CLT_prob <- 1 - pnorm(10000, mean = mu, sd = SE)
CLT_prob
## [1] 0.03632525
sim_prob <- mean(sample_means > 10000)
sim_prob
## [1] 0.057
Student interpretation:
They are close with a difference of about .014 because they are similar
in shape but still have differences.
The sampling means come from repeated samples of size 30, while the CLT
curve is an idealized normal model.
cutoff_90 <- qnorm(0.10, mean = mu, sd = SE)
cutoff_90
## [1] 5629.452
quantile(sample_means, 0.10)
## 10%
## 5726.127
Student interpretation:
The cutoff value means that 90% of samples will have a mean above
5816.595.
Cutoff at 90 because you want the above number, not all the values
below.