Welcome to the PSYC3361 coding W2 self test. The test assesses your ability to use the coding skills covered in the Week 2 online coding modules.
In particular, it assesses your ability to…
It is IMPORTANT to document the code that you write so that someone who is looking at your code can understand what it is doing. Above each chunk, write a few sentences outlining which packages/functions you have chosen to use and what the function is doing to your data. Where relevant, also write a sentence that interprets the output of your code.
Your notes should also document the troubleshooting process you went through to arrive at the code that worked.
For each of the challenges below, the documentation is JUST AS IMPORTANT as the code.
Good luck!!
Jenny
PS- if you get stuck have a look in the /images folder for inspiration
The packages needed for this self-test are ‘tidyverse’ and ‘here’
library(tidyverse)
library(here)
I have read the dino.csv file from the data folder using the here and read_csv commands. I have given this data to the variable dino.
dino <- read_csv(here("data", "dino.csv"))
The dino dataset comes from a paper illustrating the importance of plotting your data. In each of these datasets, the mean and variance of x and y are identical and the two variables are correlated in the same way (R = -0.06). When plotted, however, each reveals a very different pattern
Using the ggplot2 package I have assigned the dino variable, with mapping x = x and y = y, to the new variable dino_plot. Within this, I have used geom_point, geom_smooth and facet_wrap to plot the actual data and regression line, and seperate them by the dataset variable.
I had issues with implementing the geom_smooth function as the method was originally in ‘loess’, after discussing this issue in my lab I was advised to change this to glm and the problem was fixed!
dino_plot <- ggplot(data = dino, mapping = aes(x = x, y = y)) +
geom_point() +
geom_smooth(method = glm) +
facet_wrap(vars(dataset))
plot(dino_plot)
## `geom_smooth()` using formula = 'y ~ x'
HINT: add some colour, play with palettes, try a different theme, add a title, subtitle, caption
To make the plot prettier, I have coloured the points according to x value, added a title and subtitle, changed the theme of the plot and edited the palette of the colours (using scale_color_viridis_c - as x is a continuous value). Finally I have plotted this.
dino_plot <- dino_plot +
geom_point(mapping = aes(x = x, y = y, colour = x)) +
ggtitle(label = "Dino Dataset Point Plot",
subtitle = "Source: PSYC3361"
) +
theme_grey() +
scale_color_viridis_c(alpha = 0.5, name = "x")
plot(dino_plot)
## `geom_smooth()` using formula = 'y ~ x'
Can you write code to show that the mean, variance, and correlation between x and y is the same for each of the datasets?? HINT: this is a group_by and summarise problem
Here I create a pipe, called dino_pipe, using the data from Dino.csv, I group this by dataset and then summarise the mean and variance for x and y for each dataset, as well as, showing the correlation between x and y for each. Finally, I ungroup and print this info.
As you can see, all of these conditions largely show the same data across each dataset.
dino_pipe <- dino %>%
group_by(dataset) %>%
summarise(mean_x = mean(x),
variance_x = var(x),
mean_y = mean(y),
variance_y = var(y),
correlation = cor(x, y)
) %>%
ungroup()
print(dino_pipe)
## # A tibble: 13 × 6
## dataset mean_x variance_x mean_y variance_y correlation
## <chr> <dbl> <dbl> <dbl> <dbl> <dbl>
## 1 away 54.3 281. 47.8 726. -0.0641
## 2 bullseye 54.3 281. 47.8 726. -0.0686
## 3 circle 54.3 281. 47.8 725. -0.0683
## 4 dino 54.3 281. 47.8 726. -0.0645
## 5 dots 54.3 281. 47.8 725. -0.0603
## 6 h_lines 54.3 281. 47.8 726. -0.0617
## 7 high_lines 54.3 281. 47.8 726. -0.0685
## 8 slant_down 54.3 281. 47.8 726. -0.0690
## 9 slant_up 54.3 281. 47.8 726. -0.0686
## 10 star 54.3 281. 47.8 725. -0.0630
## 11 v_lines 54.3 281. 47.8 726. -0.0694
## 12 wide_lines 54.3 281. 47.8 726. -0.0666
## 13 x_shape 54.3 281. 47.8 725. -0.0656