5a.Apply the CLT on the sample mean of this chosen distribution in R
set.seed(64)
##Defining parameters
lambda<-1
n<-40
my_exp<-rexp(n,rate = lambda)
mu<-mean(my_exp)
sd<-sd(my_exp)
#Histogram
hist(x = my_exp, main = "Histogram of Exponential Distribution (lambda=1)")
num_sims<-10000 ##number of simulations
n2<-c(2,6,30,50,200,1000) ##sample sizes
num_cols<-length(n2)
##Generate matrix of exponential data using 0 entries only
z<-matrix(data = 0,nrow = num_sims,ncol = num_cols)
colnames(z)<-c("Sample size = 2","Sample size = 6","Sample size = 30", "Sample size = 50", "Sample size = 200", "Sample size = 1000")
z[1:16]
## [1] 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0
##Take 10,000 random samples with multiple sample sizes, compute means, and place sample mean in matrix
for (j in 1:num_cols) {
for (i in 1:num_sims) {
n2<-c(2,6,30,50,200,1000)
z[i,j]<-mean(sample(x=my_exp,size = n2[j],replace = TRUE))
}
}
colnames(z)<-paste0("n_",n2)
print(head(z))
## n_2 n_6 n_30 n_50 n_200 n_1000
## [1,] 1.80376410 1.5868425 0.7667894 0.8597662 0.8731309 0.8938599
## [2,] 0.53031250 0.6942851 0.7652580 0.8588917 0.9042555 0.9204449
## [3,] 1.60992510 1.1367223 1.2331048 0.6284774 0.8263410 0.8903453
## [4,] 0.08005113 1.2635194 0.6678671 0.7771887 0.8413882 0.9024281
## [5,] 0.47353444 0.9597628 1.0083891 0.8052826 0.9936457 0.8903005
## [6,] 0.52510615 1.6347877 0.6943726 0.8525975 0.8847521 0.9441525
summary(z)
## n_2 n_6 n_30 n_50
## Min. :0.01211 Min. :0.07584 Min. :0.3915 Min. :0.4616
## 1st Qu.:0.38041 1st Qu.:0.60427 1st Qu.:0.7815 1st Qu.:0.8144
## Median :0.71668 Median :0.85463 Median :0.9013 Median :0.9026
## Mean :0.91985 Mean :0.90499 Mean :0.9119 Mean :0.9099
## 3rd Qu.:1.25288 3rd Qu.:1.14934 3rd Qu.:1.0302 3rd Qu.:1.0011
## Max. :4.44794 Max. :3.18559 Max. :1.8590 Max. :1.4773
## n_200 n_1000
## Min. :0.6588 Min. :0.7890
## 1st Qu.:0.8639 1st Qu.:0.8902
## Median :0.9103 Median :0.9113
## Mean :0.9121 Mean :0.9114
## 3rd Qu.:0.9579 3rd Qu.:0.9321
## Max. :1.1810 Max. :1.0401
#Histograms
for (k in 1:num_cols) {hist(x=z[,k],main="Histogram of Mean of Exponential Distribution (lambda = 1)",xlim=c(-2,4),xlab=paste0("Sample size: ",n2[k]))}
##Compare means
round(mu,3)
## [1] 0.911
round(colMeans(z),3)
## n_2 n_6 n_30 n_50 n_200 n_1000
## 0.920 0.905 0.912 0.910 0.912 0.911
The central limit theorem holds as expected. As the sample size increases from n=2 to n=1000, the sample mean begins to more closely approximate the population mean (0.911). The original population follows the exponential distribution, which features the results skewed to the right. Using small sample sizes, the histograms also show right-sided skewness. When the sample sizes begin to increase (n>30), the distributions more closely resemble a normally-distributed bell curve.Finally, the spread narrows as sample size increases.
5b. Apply CLT on median
set.seed(64)
##Defining parameters
lambda<-1
n<-40
my_exp<-rexp(n,rate = lambda)
#Histogram
hist(x = my_exp, main = "Histogram of Exponential Distribution (lambda=1)")
num_sims<-10000 ##number of simulations
n2<-c(2,6,30,50,200,1000) ##sample sizes
num_cols<-length(n2)
#Take 10k,000 random samples and comupte medians
z2<-sapply(n2, function(size){replicate(num_sims,median(sample(x=my_exp,size=size,replace = TRUE)))})
colnames(z2)<-paste0("n_",n2)
print(head(z2))
## n_2 n_6 n_30 n_50 n_200 n_1000
## [1,] 1.80376410 1.9022397 0.4801917 0.5671803 0.5671803 0.5612876
## [2,] 0.53031250 0.6078501 0.5877427 0.6141979 0.5612876 0.5612876
## [3,] 1.60992510 0.7523529 0.6544126 0.4393665 0.5572607 0.5612876
## [4,] 0.08005113 0.9655788 0.5163731 0.5671803 0.5730730 0.5730730
## [5,] 0.47353444 0.6175727 0.6569089 0.5592741 0.5612876 0.5572607
## [6,] 0.52510615 1.2207894 0.3154113 0.5592741 0.5230927 0.5730730
#Histograms
for (k in 1:length(n2)) {hist(x=z2[, k],main = paste0("Histogram of Sample Medians (n = ", n2[k], ")"),xlim=c(0,3),xlab=paste0("Sample Size: ",n2[k]))}
##Compare medians
cat("Original Data Median: ", round(median(my_exp), 3), "\n\n")
## Original Data Median: 0.567
cat("Simulation Medians by Sample Size:\n")
## Simulation Medians by Sample Size:
# Convert z2 to a data frame so sapply calculates the median for each column
medians<-(sapply(as.data.frame(z2), median))
print(round(medians,3))
## n_2 n_6 n_30 n_50 n_200 n_1000
## 0.717 0.569 0.567 0.567 0.567 0.567
The central limit theorem holds as expected. As the sample size increases from n=2 to n=1000, the sample median begins to more closely approximate the population median (0.567). The original population follows the exponential distribution, which features the results skewed to the right. Using small sample sizes, the histograms also show right-sided skewness. When the sample sizes begin to increase (n>30), the distributions more closely resemble a normally-distributed bell curve.Finally, the spread narrows as sample size increases.