Please follow the tutorial linked here: https://bioinfo4all.wordpress.com/2021/01/31/tutorial-6-how-to-do-principal-component-analysis-pca-in-r/
This tutorial does a good job of describing why we use PCA and applies it to common data types for bioinformatics. As you go through the tutorial, copy and paste code below to perform the tasks. Additionally, I have a few questions you need to answer.
NOTE: I have already altered the .csv file you need for this, it is found on Canvas as “cell.csv”, bring in this file using the code below (your working directory will need to be set up correctly):
data <- read.csv("cell.csv", row.names = 1)
The code above calls the dataframe “data”, if for some reason, you end up calling your data someting else, you will need to alter all of the rest of the code accordingly. Add code chunks below this starting the tutorial with the “install.packages” code on the website:
library("factoextra")
## Loading required package: ggplot2
## Welcome to factoextra!
## Want to learn more? See two factoextra-related books at https://www.datanovia.com/library/principal-component-methods
library("FactoMineR")
pca.data <- PCA(data[,-1], scale.unit = TRUE, graph = FALSE)
# Make sure most of the data will be present and plot to show with of the dimensions contributes to how much variation.
fviz_eig(pca.data, addlabels = TRUE, ylim = c(0, 70))
## Warning: This FactoMineR PCA result contains only 5 eigenvalues and does not
## include the complete spectrum. Refit the PCA with a larger `ncp` before drawing
## a complete scree plot.
# Correlation plot
fviz_pca_var(pca.data, col.var = "cos2",
gradient.cols = c("#FFCC00", "#CC9933", "#660033", "#330033"),
repel = TRUE)
# Plot cell types and flip to put cell types as rows
pca.data <- PCA(t(data[,-1]), scale.unit = TRUE, graph = FALSE)
# Visualize
fviz_pca_ind(pca.data, col.ind = "cos2",
gradient.cols = c("#FFCC00", "#CC9933", "#660033", "#330033"),
repel = TRUE)
# Add labels
library(ggpubr)
a <- fviz_pca_ind(pca.data, col.ind = "cos2",
gradient.cols = c("#FFCC00", "#CC9933", "#660033", "#330033"),
repel = TRUE)
ggpar(a,
title = "Principal Component Analysis",
xlab = "PC1", ylab = "PC2",
legend.title = "Cos2", legend.position = "top",
ggtheme = theme_minimal())
pca.data <- PCA(data[,-1], scale.unit = TRUE,ncp = 2, graph = FALSE)
# Convert lineage to a factor
data$lineage <- as.factor(data$lineage)
# Color palette
library(RColorBrewer)
nb.cols <- 3
mycolors <- colorRampPalette(brewer.pal(3, "Set1"))(nb.cols)
# Create PCA plot and add labels
a <- fviz_pca_ind(pca.data, col.ind = data$lineage,
palette = mycolors, addEllipses = TRUE)
ggpar(a,
title = "Principal Component Analysis",
xlab = "PC1", ylab = "PC2",
legend.title = "Cell type", legend.position = "top",
ggtheme = theme_minimal())
QUESTIONS: 1. Why does the author of this tutorial suggest PCA is a powerful analysis tool?