==============================================================================

Unit 6 R Publication: Sampling Distributions & Central Limit Theorem

Dataset: Diamonds (Variable: TotalPrice)

==============================================================================

——————————————————————————

STEP 1: Population Parameters and Distribution

——————————————————————————

Load required libraries

library(Stat2Data) library(dplyr) library(ggplot2)

Load the Diamonds dataset

data(“Diamonds”)

Create a histogram of TotalPrice

ggplot(Diamonds, aes(x = TotalPrice)) + geom_histogram(fill = “skyblue”, color = “black”, bins = 20) + labs(title = “Population Distribution of Diamond Prices”, x = “Total Price (in $)”, y = “Count”) + theme_minimal()

Calculate and store population parameters

mu <- mean(Diamonds\(TotalPrice) sigma <- sd(Diamonds\)TotalPrice)

View stored parameters

mu sigma

REFLECTION 1:

- Population Mean (mu): $7,450.01

- Population Standard Deviation (sigma): \(7,780.89 # - Shape & Units: The population distribution is strongly right-skewed with # units in dollars (\)).

——————————————————————————

STEP 2: Sampling Distribution via Simulation

——————————————————————————

Set seed for reproducibility

set.seed(123)

Simulate 1,000 sample means, each with n = 30

sample_means <- replicate(1000, mean(sample(Diamonds$TotalPrice, size = 30, replace = TRUE)))

Plot histogram of the sampling distribution

ggplot(data.frame(sample_means), aes(x = sample_means), aes(x = sample_means)) + geom_histogram(fill = “lightgreen”, color = “black”, bins = 30) + labs(title = “Sampling Distribution of Sample Means (n = 30)”, x = “Sample Means (in $)”, y = “Count”) + theme_minimal()

Calculate mean and standard error of the simulated sample means

mean(sample_means) sd(sample_means)

REFLECTION 2:

- Shape Change: The sampling distribution becomes unimodal and much more

symmetric/bell-shaped compared to the strongly skewed population.

- Center Comparison: The mean of the simulated sample means ($7,459.63)

closely targets the true population mean ($7,450.01).

——————————————————————————

STEP 3: Theoretical vs. Empirical Standard Error

——————————————————————————

Calculate theoretical standard error for n = 30

n <- 30 SE <- sigma / sqrt(n)

View theoretical SE

SE

View simulated SE again for comparison

sd(sample_means)

REFLECTION 3:

- Standard Error Comparison: The simulated SE ($1,405.48) is very close to

the theoretical SE ($1,420.59).

- Interpretation: A typical sample mean of size n = 30 will fall within roughly

$1,420 of the true population mean mu.

——————————————————————————

STEP 4: Finding Probabilities using pnorm()

——————————————————————————

Find the probability that a sample mean is less than $6,000

pnorm(6000, mean = mu, sd = SE)

REFLECTION 4:

- Output: 0.1536 (or ~15.4%)

- Interpretation: There is approximately a 15.4% chance of drawing a random

sample of 30 diamonds with a sample mean price of less than $6,000.

——————————————————————————

STEP 5: Finding Quantiles using qnorm()

——————————————————————————

Find the theoretical 90th percentile of the sampling distribution

qnorm(0.90, mean = mu, sd = SE)

REFLECTION 5:

- Output: $9,270.57

- Interpretation: 90% of all random samples of size 30 will have an average

diamond price of $9,270.57 or less.

——————————————————————————

STEP 6: Empirical vs. Theoretical Model Validation

——————————————————————————

Find the empirical 90th percentile from simulated sample means

quantile(sample_means, 0.90)

REFLECTION 6:

- Comparison: The empirical 90th percentile ($9,385.17) aligns very closely

with the theoretical 90th percentile calculated via qnorm() ($9,270.57),

confirming that the Central Limit Theorem normal approximation is accurate.