5a.Apply the CLT on the sample mean of this chosen distribution in R

set.seed(64)

##Defining parameters
lambda<-1
n<-40
my_exp<-rexp(n,rate = lambda)
mu<-mean(my_exp)
sd<-sd(my_exp)

#Histogram
hist(x = my_exp, main = "Histogram of Exponential Distribution (lambda=1)")

num_sims<-10000 ##number of simulations
n2<-c(2,6,30,50,200,1000) ##sample sizes
num_cols<-length(n2)

##Generate matrix of exponential data using 0 entries only
z<-matrix(data = 0,nrow = num_sims,ncol = num_cols)
colnames(z)<-c("Sample size = 2","Sample size = 6","Sample size = 30", "Sample size = 50", "Sample size = 200", "Sample size = 1000")
z[1:16]
##  [1] 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0
##Take 10,000 random samples with multiple sample sizes, compute means, and place sample mean in matrix
for (j in 1:num_cols) {
  for (i in 1:num_sims) {
    n2<-c(2,6,30,50,200,1000)
    z[i,j]<-mean(sample(x=my_exp,size = n2[j],replace = TRUE))
    
  }
  
}
colnames(z)<-paste0("n_",n2)

print(head(z))
##             n_2       n_6      n_30      n_50     n_200    n_1000
## [1,] 1.80376410 1.5868425 0.7667894 0.8597662 0.8731309 0.8938599
## [2,] 0.53031250 0.6942851 0.7652580 0.8588917 0.9042555 0.9204449
## [3,] 1.60992510 1.1367223 1.2331048 0.6284774 0.8263410 0.8903453
## [4,] 0.08005113 1.2635194 0.6678671 0.7771887 0.8413882 0.9024281
## [5,] 0.47353444 0.9597628 1.0083891 0.8052826 0.9936457 0.8903005
## [6,] 0.52510615 1.6347877 0.6943726 0.8525975 0.8847521 0.9441525
summary(z)
##       n_2               n_6               n_30             n_50       
##  Min.   :0.01211   Min.   :0.07584   Min.   :0.3915   Min.   :0.4616  
##  1st Qu.:0.38041   1st Qu.:0.60427   1st Qu.:0.7815   1st Qu.:0.8144  
##  Median :0.71668   Median :0.85463   Median :0.9013   Median :0.9026  
##  Mean   :0.91985   Mean   :0.90499   Mean   :0.9119   Mean   :0.9099  
##  3rd Qu.:1.25288   3rd Qu.:1.14934   3rd Qu.:1.0302   3rd Qu.:1.0011  
##  Max.   :4.44794   Max.   :3.18559   Max.   :1.8590   Max.   :1.4773  
##      n_200            n_1000      
##  Min.   :0.6588   Min.   :0.7890  
##  1st Qu.:0.8639   1st Qu.:0.8902  
##  Median :0.9103   Median :0.9113  
##  Mean   :0.9121   Mean   :0.9114  
##  3rd Qu.:0.9579   3rd Qu.:0.9321  
##  Max.   :1.1810   Max.   :1.0401
#Histograms
for (k in 1:num_cols) {hist(x=z[,k],main="Histogram of Mean of Exponential Distribution (lambda = 1)",xlim=c(-2,4),xlab=paste0("Sample size: ",n2[k]))}

##Compare means
round(mu,3)
## [1] 0.911
round(colMeans(z),3)
##    n_2    n_6   n_30   n_50  n_200 n_1000 
##  0.920  0.905  0.912  0.910  0.912  0.911

The central limit theorem holds as expected. As the sample size increases from n=2 to n=1000, the sample mean begins to more closely approximate the population mean (0.911). The original population follows the exponential distribution, which features the results skewed to the right. Using small sample sizes, the histograms also show right-sided skewness. When the sample sizes begin to increase (n>30), the distributions more closely resemble a normally-distributed bell curve.Finally, the spread narrows as sample size increases.

5b. Apply CLT on median

set.seed(64)

##Defining parameters
lambda<-1
n<-40
my_exp<-rexp(n,rate = lambda)

#Histogram
hist(x = my_exp, main = "Histogram of Exponential Distribution (lambda=1)")

num_sims<-10000 ##number of simulations
n2<-c(2,6,30,50,200,1000) ##sample sizes
num_cols<-length(n2)

#Take 10k,000 random samples and comupte medians
z2<-sapply(n2, function(size){replicate(num_sims,median(sample(x=my_exp,size=size,replace = TRUE)))})
colnames(z2)<-paste0("n_",n2)
print(head(z2))
##             n_2       n_6      n_30      n_50     n_200    n_1000
## [1,] 1.80376410 1.9022397 0.4801917 0.5671803 0.5671803 0.5612876
## [2,] 0.53031250 0.6078501 0.5877427 0.6141979 0.5612876 0.5612876
## [3,] 1.60992510 0.7523529 0.6544126 0.4393665 0.5572607 0.5612876
## [4,] 0.08005113 0.9655788 0.5163731 0.5671803 0.5730730 0.5730730
## [5,] 0.47353444 0.6175727 0.6569089 0.5592741 0.5612876 0.5572607
## [6,] 0.52510615 1.2207894 0.3154113 0.5592741 0.5230927 0.5730730
#Histograms
for (k in 1:length(n2)) {hist(x=z2[, k],main = paste0("Histogram of Sample Medians (n = ", n2[k], ")"),xlim=c(0,3),xlab=paste0("Sample Size: ",n2[k]))}

##Compare medians
cat("Original Data Median: ", round(median(my_exp), 3), "\n\n")
## Original Data Median:  0.567
cat("Simulation Medians by Sample Size:\n")
## Simulation Medians by Sample Size:
# Convert z2 to a data frame so sapply calculates the median for each column
medians<-(sapply(as.data.frame(z2), median))
print(round(medians,3))
##    n_2    n_6   n_30   n_50  n_200 n_1000 
##  0.717  0.569  0.567  0.567  0.567  0.567

The central limit theorem holds as expected. As the sample size increases from n=2 to n=1000, the sample median begins to more closely approximate the population median (0.567). The original population follows the exponential distribution, which features the results skewed to the right. Using small sample sizes, the histograms also show right-sided skewness. When the sample sizes begin to increase (n>30), the distributions more closely resemble a normally-distributed bell curve.Finally, the spread narrows as sample size increases.