Welcome to the PSYC3361 coding W2 self test. The test assesses your ability to use the coding skills covered in the Week 2 online coding modules.

In particular, it assesses your ability to…

  • choose packages/functions
  • read in data
  • make a scatter plot
  • use facet_wrap
  • customise your plot with themes, colours and labels
  • use group_by and summarise

It is IMPORTANT to document the code that you write so that someone who is looking at your code can understand what it is doing. Above each chunk, write a few sentences outlining which packages/functions you have chosen to use and what the function is doing to your data. Where relevant, also write a sentence that interprets the output of your code.

Your notes should also document the troubleshooting process you went through to arrive at the code that worked.

For each of the challenges below, the documentation is JUST AS IMPORTANT as the code.

Good luck!!

Jenny

PS- if you get stuck have a look in the /images folder for inspiration

load the packages you need

# Load libraries for plots and for group_by and summarise later on 
library(tidyverse)
library(ggplot2)
library(dplyr)

read in the dino data

# Read dino.csv into variable "dino"

dino <- read.csv("data/dino.csv")

dino %>% 
  ggplot(aes(x,y))+
  geom_point(aes(colour = x))+
  geom_smooth()+
  scale_colour_steps2(low = "royalblue3", mid = "black", high = "red3",  midpoint = 50)

  # scale_colour_gradient2(low = "royalblue3", mid = "black", high = "red3",  midpoint = 60)

reproduce this plot

The dino dataset comes from a paper illustrating the importance of plotting your data. In each of these datasets, the mean and variance of x and y are identical and the two variables are correlated in the same way (R = -0.06). When plotted, however, each reveals a very different pattern

dino %>% 
  ggplot(aes(x,y))+
  geom_point()+
  facet_wrap(vars(dataset))+
  geom_smooth(method = "lm")
## `geom_smooth()` using formula = 'y ~ x'

what can you do to make it prettier

HINT: add some colour, play with palettes, try a different theme, add a title, subtitle, caption

dino %>% 
  ggplot(aes(x,y))+
  geom_point(aes(colour = dataset))+
  facet_wrap(vars(dataset))+
  geom_smooth(aes(colour = dataset),method = "lm")+
  theme_minimal()+
  labs(title = "Plots",
       subtitle = "By Jeremy",
       x = "X",
       y = "Y",
       caption = "dino plots")
## `geom_smooth()` using formula = 'y ~ x'

extra challenge

Can you write code to show that the mean, variance, and correlation between x and y is the same for each of the datasets?? HINT: this is a group_by and summarise problem

dino %>% 
  group_by(dataset) %>% 
  summarise(mean(x),mean(y),var(x),var(y),cor(x,y)) %>% 
  ungroup()
## # A tibble: 13 × 6
##    dataset    `mean(x)` `mean(y)` `var(x)` `var(y)` `cor(x, y)`
##    <chr>          <dbl>     <dbl>    <dbl>    <dbl>       <dbl>
##  1 away            54.3      47.8     281.     726.     -0.0641
##  2 bullseye        54.3      47.8     281.     726.     -0.0686
##  3 circle          54.3      47.8     281.     725.     -0.0683
##  4 dino            54.3      47.8     281.     726.     -0.0645
##  5 dots            54.3      47.8     281.     725.     -0.0603
##  6 h_lines         54.3      47.8     281.     726.     -0.0617
##  7 high_lines      54.3      47.8     281.     726.     -0.0685
##  8 slant_down      54.3      47.8     281.     726.     -0.0690
##  9 slant_up        54.3      47.8     281.     726.     -0.0686
## 10 star            54.3      47.8     281.     725.     -0.0630
## 11 v_lines         54.3      47.8     281.     726.     -0.0694
## 12 wide_lines      54.3      47.8     281.     726.     -0.0666
## 13 x_shape         54.3      47.8     281.     725.     -0.0656

knit your document to pdf