Unit 4 Final Part II: Heatmaps

Another very common visual tool for analysis is a heatmap. We’ve seen through our previous assignments and readings that these are useful for taking frequency or count data and showing their distribution. There are many applications for this type of figure, including tracking abundances of species across functional units, making comparisons between treatments, etc. The data set we will be working with for this assignment comes from drug/substance data. Columns display the measurements as percentage of the age group that used a certain substance in a 12-month period, as well as the median number of times an individual used in the 12-month period.

Complete the prompts below and be sure to answer all questions as completely as possible!

Part 1: loading packages, data, and manipulating data

  1. Code: install and load the following package: “tidyverse”, “patchwork”, “ggplot2”, and “RColorBrewer”
library(tidyverse)
## ── Attaching core tidyverse packages ──────────────────────── tidyverse 2.0.0 ──
## ✔ dplyr     1.2.1     ✔ readr     2.2.0
## ✔ forcats   1.0.1     ✔ stringr   1.6.0
## ✔ ggplot2   4.0.3     ✔ tibble    3.3.1
## ✔ lubridate 1.9.5     ✔ tidyr     1.3.2
## ✔ purrr     1.2.2     
## ── Conflicts ────────────────────────────────────────── tidyverse_conflicts() ──
## ✖ dplyr::filter() masks stats::filter()
## ✖ dplyr::lag()    masks stats::lag()
## ℹ Use the conflicted package (<http://conflicted.r-lib.org/>) to force all conflicts to become errors
library(patchwork)
library(ggplot2)
library(RColorBrewer)
  1. Code: load the drug use data set with the following code, this code will also alter the data types for certain variables
druguse<-read.csv('drug-use-by-age.csv',
                  header=T,
                  na.strings = "-")

druguse[,3:28] <- sapply(druguse[,3:28],as.numeric)
  1. One of the biggest challenges we run into with data analysis is the formatting of our data and making sure it’s set up correctly for use to run analysis and plot figures. This course didn’t spend a lot of time on this, this will be something covered in more details in other courses in this program. But for the druguse data, columns display the same measurements (use as percentage of the age group that used in the 12-month period, and frequency as the median number of times an individual used in the 12-month period), but categories are split among columns.

One tidy format would have use and frequency as their own columns, with drug_type as indicating the type of drug. Converting to this format would allow for easier and more direct analysis. Run the code chunk below to rework the data:

Our steps will be:

  1. Use pivot longer on all columns except age and n.
  2. Use the names_pattern argument to split these column titles by drug type and measurement type (“use” versus “frequency”)
  3. Then we will pivot wider to put use and frequency as their own variables.
druguse %>%
    pivot_longer(cols = -c(age,n),
                 names_pattern = "(.*)_(.*)",
                 names_to=c("drug_type","measurement")) %>%
    pivot_wider(names_from = measurement,
                values_from = value)->drugusetidy

Part II: creating, manipulating, and interpretting the heatmap

  1. Code: run the following code to create a heatmap for all of the substances across age
drugusetidy %>%
    ggplot(data =.,
           aes(x = age,
               y = drug_type,
               fill = use))+
    geom_tile()+
    scale_fill_gradientn("Percentage Of\nAge Group Using\nIn The Past Year",
                         colors = brewer.pal(9,name="YlOrRd"))+
    theme(axis.text.x = element_text(angle = 60, hjust=1))+
    labs(x="Age Group",
         y="Drug Type",
         title="Drug-use Percentages by Age")+
    theme(plot.background = element_rect(fill="gray"),
          legend.background = element_rect(fill="gray"))

  1. Question: How do you interpret the heat map? What substances seem to have the highest percentage of use? At what ages dow we see a lot of use?
  1. Code: the complete heatmap has a lot of lower percentages for certain substances, let’s run the code below to hone in on just a few of the substances (yes, I know reliever is spelled wrong! but that’s how it is in the dataset).
drugusetidy %>%
  filter(drug_type %in% c("marijuana","pain_releiver")) %>%
  ggplot(data =.,
         aes(x = age,
             y = drug_type,
             fill = use))+
  geom_tile()+
  scale_fill_gradientn("Percentage Of\nAge Group Using\nIn The Past Year",
                       colors = brewer.pal(9,name="YlOrRd"))+
  theme(axis.text.x = element_text(angle = 60, hjust=1))+
  labs(x="Age Group",
       y="Drug Type",
       title="Drug-use Percentages by Age")+
  theme(plot.background = element_rect(fill="gray"),
        legend.background = element_rect(fill="gray"))

  1. Code: recreate the code from 3 to hone in on 3 substances of interest to you!
drugusetidy %>%
  filter(drug_type %in% c("hallucinogen","cocaine","heroin")) %>%
  ggplot(data =.,
         aes(x = age,
             y = drug_type,
             fill = use))+
  geom_tile()+
  scale_fill_gradientn("Percentage Of\nAge Group Using\nIn The Past Year",
                       colors = brewer.pal(9,name="YlOrRd"))+
  theme(axis.text.x = element_text(angle = 60, hjust=1))+
  labs(x="Age Group",
       y="Drug Type",
       title="Drug-use Percentages by Age")+
  theme(plot.background = element_rect(fill="gray"),
        legend.background = element_rect(fill="gray"))

  1. Question: how would you interpret the heatmap you just made?
  1. Code: The two heatmaps you just made looked at the use , recreate the figure from 3 but instead of using use, use frequency (NOTE: you may need to look at the columns in drugusetidy to get the variable correct)

BONUS (2 pts): change the legend title to reflect that we are now looking at a median frequency by age!

drugusetidy %>%
  filter(
    drug_type %in% c(
      "marijuana",
      "pain_releiver"
    )
  ) %>%
  ggplot(
    data = .,
    aes(
      x = age,
      y = drug_type,
      fill = frequency
    )
  ) +
  geom_tile() +
  scale_fill_gradientn(
    "Median Frequency\nOf Use\nIn The Past Year",
    colors = brewer.pal(9, name="YlOrRd")
  ) +
  theme(
    axis.text.x = element_text(
      angle = 60,
      hjust = 1
    )
  ) +
  labs(
    x = "Age Group",
    y = "Drug Type",
    title = "Drug-use Frequency by Age"
  ) +
  theme(
    plot.background = element_rect(fill="gray"),
    legend.background = element_rect(fill="gray")
  )

  1. Question: how would you interpret the new heat map you just made?
  1. Code: recreate the code from 6 to hone in on 3 substances of interest to you!
drugusetidy %>%
  filter(
    drug_type %in% c(
      "alcohol",
      "marijuana",
      "cocaine"
    )
  ) %>%
  ggplot(
    data = .,
    aes(
      x = age,
      y = drug_type,
      fill = frequency
    )
  ) +
  geom_tile() +
  scale_fill_gradientn(
    "Median Frequency\nOf Use\nIn The Past Year",
    colors = brewer.pal(9, name="YlOrRd")
  ) +
  theme(
    axis.text.x = element_text(
      angle = 60,
      hjust = 1
    )
  ) +
  labs(
    x = "Age Group",
    y = "Drug Type",
    title = "Drug-use Frequency by Age"
  ) +
  theme(
    plot.background = element_rect(fill="gray"),
    legend.background = element_rect(fill="gray")
  )

  1. Question: how would you interpret the heatmap you just made?