==============================================================================
Unit 6 R Publication: Sampling Distributions & Central Limit
Theorem
Dataset: Diamonds (Variable: TotalPrice)
==============================================================================
——————————————————————————
STEP 1: Population Parameters and Distribution
——————————————————————————
Load required libraries
library(Stat2Data) library(dplyr) library(ggplot2)
Load the Diamonds dataset
data(“Diamonds”)
Create a histogram of TotalPrice
ggplot(Diamonds, aes(x = TotalPrice)) + geom_histogram(fill =
“skyblue”, color = “black”, bins = 20) + labs(title = “Population
Distribution of Diamond Prices”, x = “Total Price (in $)”, y = “Count”)
+ theme_minimal()
Calculate and store population parameters
mu <- mean(Diamonds\(TotalPrice) sigma
<- sd(Diamonds\)TotalPrice)
View stored parameters
mu sigma
REFLECTION 1:
- Population Mean (mu): $7,450.01
- Population Standard Deviation (sigma): \(7,780.89 # - Shape & Units: The population
distribution is strongly right-skewed with # units in dollars
(\)).
——————————————————————————
STEP 2: Sampling Distribution via Simulation
——————————————————————————
Set seed for reproducibility
set.seed(123)
Simulate 1,000 sample means, each with n = 30
sample_means <- replicate(1000, mean(sample(Diamonds$TotalPrice,
size = 30, replace = TRUE)))
Plot histogram of the sampling distribution
ggplot(data.frame(sample_means), aes(x = sample_means), aes(x =
sample_means)) + geom_histogram(fill = “lightgreen”, color = “black”,
bins = 30) + labs(title = “Sampling Distribution of Sample Means (n =
30)”, x = “Sample Means (in $)”, y = “Count”) + theme_minimal()
Calculate mean and standard error of the simulated sample means
mean(sample_means) sd(sample_means)
REFLECTION 2:
- Shape Change: The sampling distribution becomes unimodal and much
more
symmetric/bell-shaped compared to the strongly skewed
population.
- Center Comparison: The mean of the simulated sample means
($7,459.63)
closely targets the true population mean ($7,450.01).
——————————————————————————
STEP 3: Theoretical vs. Empirical Standard Error
——————————————————————————
Calculate theoretical standard error for n = 30
n <- 30 SE <- sigma / sqrt(n)
View simulated SE again for comparison
sd(sample_means)
REFLECTION 3:
- Standard Error Comparison: The simulated SE ($1,405.48) is very
close to
the theoretical SE ($1,420.59).
- Interpretation: A typical sample mean of size n = 30 will fall
within roughly
$1,420 of the true population mean mu.
——————————————————————————
STEP 4: Finding Probabilities using pnorm()
——————————————————————————
Find the probability that a sample mean is less than $6,000
pnorm(6000, mean = mu, sd = SE)
REFLECTION 4:
- Output: 0.1536 (or ~15.4%)
- Interpretation: There is approximately a 15.4% chance of drawing a
random
sample of 30 diamonds with a sample mean price of less than
$6,000.
——————————————————————————
STEP 5: Finding Quantiles using qnorm()
——————————————————————————
Find the theoretical 90th percentile of the sampling
distribution
qnorm(0.90, mean = mu, sd = SE)
REFLECTION 5:
- Output: $9,270.57
- Interpretation: 90% of all random samples of size 30 will have an
average
diamond price of $9,270.57 or less.
——————————————————————————
STEP 6: Empirical vs. Theoretical Model Validation
——————————————————————————
Find the empirical 90th percentile from simulated sample means
quantile(sample_means, 0.90)
REFLECTION 6:
- Comparison: The empirical 90th percentile ($9,385.17) aligns very
closely
with the theoretical 90th percentile calculated via qnorm()
($9,270.57),
confirming that the Central Limit Theorem normal approximation is
accurate.