Unit 4 Final Part I: Volcano Plots

As we’ve seen, volcano plots are incredibly useful visuals to create when working with expression data for a large number of proteins. The data set that we will work with for this assignment has data comparing protein expression between two different types of skin cancer. Each row of the data set represents a comparison between expression of a specific protein with respect to the two types of cancers.

Complete the prompts below and be sure to answer all questions as completely as possible!

NOTE: given that I ask you to overwrite the variable name for the dataframe several times, if you run into errors with your code, try putting ALL of your code together in one code chunk!

Part I: Exploring the data

  1. Code: Bring in the “RDEB.csv” file and call it “mySig”
mySig <- read.csv("RDEB.csv",  )
  1. Code: Explore the data set by looking at the variables and variable types
str(mySig)
## 'data.frame':    1764 obs. of  11 variables:
##  $ Protein             : chr  "sp|A0A0C4DH31|HV118_HUMAN" "sp|A0A0C4DH33|HV124_HUMAN" "sp|A0A0C4DH38|HV551_HUMAN" "sp|A0A0C4DH67|KV108_HUMAN" ...
##  $ Label               : chr  "metast-RDEB" "metast-RDEB" "metast-RDEB" "metast-RDEB" ...
##  $ log2FC              : chr  "-0.260261457" "-0.302608749" "-0.063104838" "-2.119160915" ...
##  $ SE                  : num  0.789 0.224 0.973 0.516 0.759 ...
##  $ Tvalue              : num  -0.3297 -1.3519 -0.0649 -4.1087 0.2024 ...
##  $ DF                  : int  11 11 7 6 6 17 4 6 13 6 ...
##  $ pvalue              : num  0.7478 0.2036 0.9501 0.0063 0.8463 ...
##  $ adj.pvalue          : num  0.8526 0.3927 0.9699 0.0703 0.9088 ...
##  $ issue               : chr  NA NA NA NA ...
##  $ MissingPercentage   : num  0.579 0.526 0.719 0.711 0.754 ...
##  $ ImputationPercentage: num  0.263 0.211 0.193 0.132 0.175 ...
  1. Question: What do you notice about the columns in this dataset? What columns/variables sound familiar given our previous content on hypothesis testing?
  1. Question: If each row represents a comparison for expression between two types of cancer, what type of test was likely run to generate the data in a few of the columns? Why?

Part II: Manipulating the data

  1. Code: to generate a volcano plot, we really only need 3 columns of data, write a line of code that overwrites the “mySig” variable but keeps on columns 1, 3, and 7. NOTE: in one of our first R labs, you created code that removed ROWS, the code will be similar and follow the example below. Keep the name of the dataframe the same!

Ex: myData <- myData[, c(4, 8, 9)] ### this code would rewrite myData but keep only columns 4,8 and 9.

mySig <- mySig[, c(1, 3, 7)]
str(mySig)
## 'data.frame':    1764 obs. of  3 variables:
##  $ Protein: chr  "sp|A0A0C4DH31|HV118_HUMAN" "sp|A0A0C4DH33|HV124_HUMAN" "sp|A0A0C4DH38|HV551_HUMAN" "sp|A0A0C4DH67|KV108_HUMAN" ...
##  $ log2FC : chr  "-0.260261457" "-0.302608749" "-0.063104838" "-2.119160915" ...
##  $ pvalue : num  0.7478 0.2036 0.9501 0.0063 0.8463 ...
  1. Code: given that this dataset is so large, there are likely instances where “NA” or “INf” values exist, if these stay in your dataframe, R will trip over itself while trying to make a plot. You need to search for and seek out rows in your data that are complete (no NA values), use the example code below and recreate your own

Ex: myData <- myData[complete.cases(myData),] ### this code will search for and pick out complete cases for each row, meaning all of the data are there.

mySig <- mySig[complete.cases(mySig),]
  1. Code: change the log2FC variable to a numeric
mySig$log2FC <- as.numeric(mySig$log2FC)

Part III: Creating, changing, and interpretting the volcano plot

  1. Code: load the “EnhancedVolcano” package (NOTE: if you get an error saying that the package won’t load because of issues in software version, you can use the following code to make sure it loads)
library(EnhancedVolcano)
## Loading required package: ggplot2
## Loading required package: ggrepel

install.packages(“EnhancedVolcano”, force = T)

  1. Code: run the following code chunk to create the plot
rownames(mySig) <- mySig$Protein
mySig <- mySig[,-1]

EnhancedVolcano(mySig,
                lab = rownames(mySig),
                x = 'log2FC',
                y = 'pvalue',
                pCutoff = 0.05,
                FCcutoff = 2,
                pointSize = 4,
                labSize = 3)
## Warning: Using `size` aesthetic for lines was deprecated in ggplot2 3.4.0.
## ℹ Please use `linewidth` instead.
## ℹ The deprecated feature was likely used in the EnhancedVolcano package.
##   Please report the issue to the authors.
## This warning is displayed once per session.
## Call `lifecycle::last_lifecycle_warnings()` to see where this warning was
## generated.
## Warning: The `size` argument of `element_line()` is deprecated as of ggplot2 3.4.0.
## ℹ Please use the `linewidth` argument instead.
## ℹ The deprecated feature was likely used in the EnhancedVolcano package.
##   Please report the issue to the authors.
## This warning is displayed once per session.
## Call `lifecycle::last_lifecycle_warnings()` to see where this warning was
## generated.

  1. Code: The settings for the code above had created visual cutoffs for the p-value and the log fold change, create a plot that changes the cutoffs for both of these variables, make the pCutoff 0.005 and the FCcutoff 2.5
EnhancedVolcano(mySig,
                lab = rownames(mySig),
                x = 'log2FC',
                y = 'pvalue',
                pCutoff = 0.005,
                FCcutoff = 2.5,
                pointSize = 4,
                labSize = 3)

  1. Question: How do you interpret the plot you made? Think about explaining the colors and locations!
  1. Question: If you were to use this information to hone in one specific proteins that experience drastic differences in fold change and significant p-values, what proteins would they be? How would you use this information to determine future research directions?