https://rpubs.com/evelynebrie/bootstrapping
This tutorial outlines how to generate data with certain parameters (things you are creating but could very easily determine for yourself from your own sample data).
set.seed(300)
myData <- rnorm(2000,20,4.5)
length(myData)
## [1] 2000
mean(myData)
## [1] 20.25773
sd(myData)
## [1] 4.590852
set.seed(200)
sample.size <- 2000
n.samples <- 1000
bootstrap.results <- c()
for (i in 1:n.samples)
{
obs <- sample(1:sample.size, replace=TRUE)
bootstrap.results[i] <- mean(myData[obs])
}
length(bootstrap.results)
## [1] 1000
summary(bootstrap.results)
## Min. 1st Qu. Median Mean 3rd Qu. Max.
## 19.92 20.19 20.26 20.26 20.33 20.57
sd(bootstrap.results)
## [1] 0.1021229
par(mfrow=c(2,1), pin=c(5.8,0.98))
hist(bootstrap.results,
col="#d83737",
xlab="Mean",
main=paste("Means of 1000 bootstrap samples from myData"))
hist(myData,
col="#37aad8",
xlab="Value",
main=paste("Distribution of myData"))
set.seed(200)
sample.size <- 2000
n.samples <- 1000
bootstrap.results <- c()
for (i in 1:n.samples)
{
bootstrap.results[i] <- mean(rnorm(2000,20,4.5))
}
length(bootstrap.results)
## [1] 1000
summary(bootstrap.results)
## Min. 1st Qu. Median Mean 3rd Qu. Max.
## 19.64 19.93 20.00 20.00 20.07 20.32
sd(bootstrap.results)
## [1] 0.1041927
par(mfrow=c(2,1), pin=c(5.8,0.98))
hist(bootstrap.results,
col="#d83737",
xlab="Mean",
main=paste("Means of 1000 bootstrap samples from the DGP"))
hist(myData,
col="#37aad8",
xlab="Value",
main=paste("Distribution of myData"))
There are 2 EXERCISES listed at the bottom of the linked page, please complete those 2 exercises as part of this assignment and include th code and output for them below.
set.seed(150)
data <- rnorm(1000, mean = 30, sd = 2.5)
results <- numeric(50)
for(i in 1:50) {
sample_data <- sample(data, 1000, replace = TRUE)
results[i] <- mean(sample_data)
}
results
## [1] 29.91217 29.95111 29.88946 29.83869 29.83717 29.96443 29.87997 29.98104
## [9] 29.88365 29.88761 29.92220 29.88761 29.92020 29.88494 29.90298 29.91531
## [17] 29.87105 29.87878 30.02194 29.94590 29.95923 29.74643 29.92373 29.94112
## [25] 29.83525 30.01097 29.99350 30.06145 29.97009 29.92336 29.84687 29.92258
## [33] 29.91140 29.99878 29.96903 29.90596 29.86557 29.97388 29.96784 29.97746
## [41] 29.87797 30.01480 30.10231 29.91795 29.89808 29.99120 29.93832 29.85962
## [49] 29.90181 29.95929
par(mfrow = c(1, 2))
hist(results,
main = "Distribution of Sample Means",
xlab = "Sample Mean",
col = "lightblue")
hist(data,
main = "Distribution of Original Data",
xlab = "Values",
col = "lightgreen")
Download and bring in the “trees.csv” file from Canvas, call it “trees”. This dataset contains tree measurements from 31 individual trees. We are going to apply bootstrapping to repeatedly sample from this sample!
Perform the following tasks: 1. Code: determine the sample size, the SD of height, and the median of height 2. Code: copy/paste all of the code from section 2.1 in th link above, but you will need to replace certain values: a. you do not need the set.seed() code b. sample.size should be the sample size from our trees data c. any mention of “myData” should be replaced with “trees$Height”
trees <- read.csv("~/Desktop/BIN510-assignments/trees.csv")
length(trees$Height)
## [1] 31
sd(trees$Height)
## [1] 6.371813
median(trees$Height)
## [1] 76
sample.size <- 31
n.samples <- 1000
bootstrap.results <- c()
for (i in 1:n.samples)
{
obs <- sample(1:sample.size, replace=TRUE)
bootstrap.results[i] <- mean(trees$Height[obs])
}
length(bootstrap.results)
## [1] 1000
sd(bootstrap.results)
## [1] 1.096471
par(mfrow=c(2,1), pin=c(5.8,0.98))
hist(bootstrap.results,
col="green",
xlab="Mean",
main=paste("Means of 1000 bootstrap samples from trees"))
hist(trees$Height,
col="#37aad8",
xlab="Height",
main=paste("Distribution of trees$Height"))