https://rpubs.com/evelynebrie/bootstrapping
This tutorial outlines how to generate data with certain parameters (things you are creating but could very easily determine for yourself from your own sample data).
Generating Data
set.seed(300)
myData <- rnorm(2000,20,4.5)
length(myData)
## [1] 2000
mean(myData)
## [1] 20.25773
sd(myData)
## [1] 4.590852
Bootstrapping
#Resampling
set.seed(200)
sample.size <- 2000
n.samples <- 1000
bootstrap.results <- c()
for (i in 1:n.samples)
{
obs <- sample(1:sample.size, replace=TRUE)
bootstrap.results[i] <- mean(myData[obs])
}
length(bootstrap.results)
## [1] 1000
summary(bootstrap.results)
## Min. 1st Qu. Median Mean 3rd Qu. Max.
## 19.92 20.19 20.26 20.26 20.33 20.57
sd(bootstrap.results)
## [1] 0.1021229
#Graphing
hist(bootstrap.results,
col="#d83737",
xlab="Mean",
main=paste("Means of 100 bootstrap samples from myData"))
hist(myData,
col="#37aad8",
xlab="Value",
main=paste("Distribution of myData"))
#Resampling again
set.seed(200)
sample.size <- 2000
n.samples <- 1000
bootstrap.results <- c()
for (i in 1:n.samples)
{
bootstrap.results[i] <- mean(rnorm(2000,20,4.5))
}
length(bootstrap.results)
## [1] 1000
summary(bootstrap.results)
## Min. 1st Qu. Median Mean 3rd Qu. Max.
## 19.64 19.93 20.00 20.00 20.07 20.32
sd(bootstrap.results)
## [1] 0.1041927
#I was having trouble loading the historgrams with the par() code so I omitted it
#Graphing
hist(bootstrap.results,
col="#d83737",
xlab="Mean",
main=paste("Means of 1000 bootstrap samples from the DGP"))
hist(myData,
col="#37aad8",
xlab="Value",
main=paste("Distribution of myData"))
There are 2 EXERCISES listed at the bottom of the linked page, please complete those 2 exercises as part of this assignment and include th code and output for them below.
Exercise 1
set.seed(150)
newData <- rnorm(1000,30,2.5)
length(newData)
## [1] 1000
mean(newData)
## [1] 29.92068
sd(newData)
## [1] 2.475175
#Resampling
sample.size <- 1000
n.samples <- 1000
bootstrap.results <- c()
for (i in 1:n.samples)
{obs <- sample(1:sample.size,replace = TRUE)
bootstrap.results[i] <- mean(newData[obs])
}
length(bootstrap.results)
## [1] 1000
summary(bootstrap.results)
## Min. 1st Qu. Median Mean 3rd Qu. Max.
## 29.68 29.87 29.92 29.92 29.98 30.17
mean(bootstrap.results)
## [1] 29.92357
Exercise 2
par(
mfrow = c(2,1),
mar = c(4,4,2,1)
)
hist(bootstrap.results,
col = "tomato",
xlab = "Mean",
main = "Means of 50 Samples from exerciseData"
)
hist(newData,
col = "skyblue",
xlab = "Value",
main = "Distribution of Original Data"
)
Download and bring in the “trees.csv” file from Canvas, call it “trees”. This dataset contains tree measurements from 31 individual trees. We are going to apply bootstrapping to repeatedly sample from this sample!
Perform the following tasks: 1. Code: determine the sample size, the SD of height, and the median of height 2. Code: copy/paste all of the code from section 2.1 in th link above, but you will need to replace certain values: a. you do not need the set.seed() code b. sample.size should be the sample size from our trees data c. any mention of “myData” should be replaced with “trees$Height”
trees <- read.csv("trees.csv", header = TRUE)
str(trees)
## 'data.frame': 31 obs. of 4 variables:
## $ rownames: int 1 2 3 4 5 6 7 8 9 10 ...
## $ Girth : num 8.3 8.6 8.8 10.5 10.7 10.8 11 11 11.1 11.2 ...
## $ Height : int 70 65 63 72 81 83 66 75 80 75 ...
## $ Volume : num 10.3 10.3 10.2 16.4 18.8 19.7 15.6 18.2 22.6 19.9 ...
length(trees$Height)
## [1] 31
sd(trees$Height)
## [1] 6.371813
median(trees$Height)
## [1] 76
sample.size <- length(trees$Height)
n.samples <- 1000
bootstrap.results <- c()
for (i in 1:n.samples)
{obs <- sample(1:sample.size,replace = TRUE)
bootstrap.results[i] <- mean(trees$Height[obs])
}
length(bootstrap.results)
## [1] 1000
summary(bootstrap.results)
## Min. 1st Qu. Median Mean 3rd Qu. Max.
## 72.42 75.22 76.00 75.97 76.77 79.97
mean(bootstrap.results)
## [1] 75.96871
par(
mfrow = c(2,1),
mar = c(4,4,2,1)
)
hist(
bootstrap.results,
col = "orchid",
xlab = "Mean Height",
main = "Means of 1000 Bootstrap Samples of Tree Height"
)
hist(
trees$Height,
col = "olivedrab3",
xlab = "Tree Height",
main = "Distribution of Tree Height"
)