I will begin by obtaining the working directory.

The working directory says “/Users/owner/Downloads” and I would like for this to be in a specific file for this assignment so it is easier to find. I will now set the working directory so that it is easier for me to find on my desktop home screen.

setwd("/Users/owner/Desktop/Lab Assignment 6 PSYC 2020")

1) Population Distribution (2.5 pts)

1a. Save a dataset by using the code below. Let’s assume it is a population distribution.

Saving the dataset into an object:

objectname <- c(rep(2,200), rep(4, 200), rep(6, 200), rep(8,200), rep(10,200))

1b. Obtain a histogram. What kind of distribution does it have? (1.5 pts)

Obtaining a histogram using hist():

hist(objectname)

## After running the above code, it can be seen that it has a uniform distribution.

1c. What is the mean and SD of the population? (1 pt)

Obtaining the mean and standard deviation using mean() and sd():

mean(objectname)
## [1] 6
sd(objectname)
## [1] 2.829842
## The mean of the population is 6. The standard deviation is 2.829842. 

2) Sample Distribution (4 pts)

2a. Obtain five different samples with n = 30 by using sample() code. Make sure you use the set.seed(5302025) and set different names to those five (5) samples to proceed the next step. (2 pt)

Obtaining the samples:

set.seed(5302025)

sample1 <- sample(objectname, 30)
sample2 <- sample(objectname, 30)
sample3 <- sample(objectname, 30)
sample4 <- sample(objectname, 30)
sample5 <- sample(objectname, 30)

2b. What are the five (5) means? (1 pt)

Using mean() to find the means:

mean(sample1)
## [1] 5.8
mean(sample2)
## [1] 5.6
mean(sample3)
## [1] 5.533333
mean(sample4)
## [1] 4.666667
mean(sample5)
## [1] 6.133333
## The means for sample 1, 2, 3, 4, and 5 are 5.8, 5.6, 5.533333, 4.666667, and 6.133333, respectively.

2c. What are the five (5) SDs? (1 pt)

finding the five standard deviations using sd():

set.seed(5302025)
sd(sample1)
## [1] 2.940795
sd(sample2)
## [1] 2.485822
sd(sample3)
## [1] 3.002681
sd(sample4)
## [1] 2.590877
sd(sample5)
## [1] 2.096521
## The standard deviations for sample 1, 2, 3, 4, and 5 are 2.940795, 2.485822, 3.002681, 2.590877, and 2.096521, respectively.

3) Sampling distribution (3 pts)

3a. What is the average of the five (5) sample means above? (1 pt)

finding the average of the sample means:

mean(c(mean(sample1), mean(sample2), mean(sample3), mean(sample4), mean(sample5)))
## [1] 5.546667
## The average of the five sample means is 5.546667.

3b. What is the SD of the five (5) sample means? (1 pt)

finding the standard deviation of the five sample means using sd():

sd(c(mean(sample1), mean(sample2), mean(sample3), mean(sample4), mean(sample5)))
## [1] 0.5444671
## The standard deviation of the five sample means is 0.5444671.

3c. Compare the variance of population distribution, sample distribution, and sampling distribution. What are the differences? Which one has the least variance? (1 pt)

var(objectname)
## [1] 8.008008
var(c(sample1, sample2, sample3, sample4, sample5))
## [1] 7.041432
var(c(mean(sample1), mean(sample2), mean(sample3), mean(sample4), mean(sample5)))
## [1] 0.2964444
## The variance of the population distribution is 8.008008, the variance of the sample distribution is 7.041432, and the variance of the sampling distribution is 0.2964444. The population distribution contains all of the values in the population, so it has the largest spread and variance. The sample distribution contains the 30 observations from one sample, and the sampling distribution has five sample means. Since the sampling distribution has the lowest value in this category, it has the least amount of variance. 

4) Population Distribution (3 pts)

4a. Create a distribution by using the code below and assume it is a population distribution.

creating a distribution

objectname2 <- rnorm(999, 110, 10)

4b. Based on the code, what kind of distribution this population has? (1 pt)

This population has a normal distribution.

4c. Based on the code, what is the mean and SD of this population? (1 pt)

Finding the mean and standard deviation using mean() and sd():

mean(objectname2)
## [1] 109.4424
sd(objectname2)
## [1] 10.20237
## The mean is 109.44240, and the standard deviation is 10.20237, or about 10. 

4d. Obtain a histogram. (1 pt)

Using hist() to obtain a histogram:

hist(objectname2)

5) Sampling Distribution (4 pts)

5a. Obtain 10 sample means with 45 observations per each drawn from the population above.

creating the sampling distribution of sample means:

samples = rep(NA, 10)
for (i in 1:10) {
  samples[i] = mean(sample(objectname2, 45, replace=T))
}

Obtaining the means:

sample1 <- sample(objectname2, 45)
sample2 <- sample(objectname2, 45)
sample3 <- sample(objectname2, 45)
sample4 <- sample(objectname2, 45)
sample5 <- sample(objectname2, 45)
sample6 <- sample(objectname2, 45)
sample7 <- sample(objectname2, 45)
sample8 <- sample(objectname2, 45)
sample9 <- sample(objectname2, 45)
sample10 <- sample(objectname2, 45)

means <- c(
  mean(sample1),
  mean(sample2),
  mean(sample3),
  mean(sample4),
  mean(sample5),
  mean(sample6),
  mean(sample7),
  mean(sample8),
  mean(sample9),
  mean(sample10))

5b. Based on the code, how many samples did you draw? (1 pt)

## I drew 10 samples and each had 45 observations.

5c. What is the mean and SD of this sampling distribution? (1 pt)

Finding mean and standard deviation

mean(samples)
## [1] 109.0191
sd(samples)
## [1] 0.8875673
## The mean is 109.4314 and the standard deviation is 1.773058.

5d. Obtain a histogram. (1 pt)

Obtaining a histogram using hist()

hist(samples)

5e. Compare the variance of population distribution and sampling distribution. What are the differences? Which one has the smaller variance? (1 pt)

Comparing the variances using var():

var(objectname2)
## [1] 104.0883
var(samples)
## [1] 0.7877758
## The variation of population distribution is 104.0883, but the variation of sampling distribution is substantially lower at 3.143735, meaning that the sampling distribution has less variance than the population distribution.

6) Population Distribution (3 pts)

6a. Create a new dataset by using the code below and assume it is a population.

Creaing a new dataset based on code given in class:

objectname3 <- rexp(1200, 0.5)

6b. Obtain a histogram. (1 pt)

Obtaining a histogram using hist():

hist(objectname3)

6c. Based on the code, what kind of distribution this population has? (1 pt)

The population is positively skewed or right-skewed distribution, as confirmed by the histogram above.

6d. Draw a sample with 30 observations from this population by using sample() code. (1 pt)

Drawing the sample:

sample3 <- sample(objectname3, 30)

7) Sampling Distribution (4 pts)

7a. Obtain 50 sample means with 10 observations per each drawn from the population above. You can use the code below.

samples = rep(NA, 50)
for (i in 1:50) {
  samples[i] = mean(sample(objectname3, 10, replace = T))
}

7b. Obtain a histogram for this sampling distribution with 10 observations per sample. (1 pt)

Using hist() to obtain the histogram:

hist(samples)

## This results in a histogram of the 50 sample means and each mean is based on 10 observations. From running the histogram, it can be determined that this is a positively skewed distribution.

7c. Modify the code to get 50 sampling with 40 observations per sample. (1 pt)

Modifying to get the instructed sampling with observations:

samples <- rep(NA, 50)

for (i in 1:50) {
  samples[i] <- mean(sample(objectname3, 40))
}

7d. Modify the code to get 50 sampling with 200 observations per sample. (1 pt)

Modifying to get the instructed sampling with observations:

samples <- rep(NA, 50)

for (i in 1:50) {
  samples[i] <- mean(sample(objectname3, 200))
}

7e. Compare three histograms for sampling distributions with 10, 40, and 200 observations. Which one is peaked the most? (Make sure to use the same range for x-axis of three (3) histograms. Use the code ‘xlim = c(minimum, maximum)’ as subargument when making histograms.) (1 pt)

First, creating three separate objects for each observation amount:

## for 10 observations per sample

samples10 <- rep(NA, 50) 
for (i in 1:50) {
  samples10[i] <- mean(sample(objectname3, 10))
}

## for 40 observations per sample

samples40 <- rep(NA, 50)
for (i in 1:50) {
  samples40[i] <- mean(sample(objectname3, 40))
}

## for 200 observations per sample

samples200 <- rep(NA, 50)
for (i in 1:50) {
  samples200[i] <- mean(sample(objectname3, 200))
}

Comparing the three histograms by making them using hist() for the different sampling distributions:

hist(samples10, xlim = c(0, 10))

hist(samples40, xlim = c(0, 10))

hist(samples200, xlim = c(0,10))

Based on the three histograms,the sampling distribution with 200 observations per sample is the most peaked.