Abstract

The American Gut Project collects microbiome samples and survey information for research. They have over 40,000 samples that have been sequenced from human gut and oral microbiomes. This specific project aims to make a human readable output so that it can be interpreted by clients.

The Workflow

A text file with a sample from The American Gut project was imported into R. The necessary packages were downloaded. The file was then read as invividual lines to later be organized

library(dplyr)
## Warning: package 'dplyr' was built under R version 4.4.3
## 
## Attaching package: 'dplyr'
## The following objects are masked from 'package:stats':
## 
##     filter, lag
## The following objects are masked from 'package:base':
## 
##     intersect, setdiff, setequal, union
biosample <- readLines("biosample_result (1).txt")
head(biosample)
## [1] "1: american gut project; 10317.X00185902\t"                
## [2] "Identifiers: BioSample: SAMEA112990854; SRA: ERS14985877\t"
## [3] "Organism: human gut metagenome\t"                          
## [4] "Attributes:\t"                                             
## [5] "    /ENA-CHECKLIST\tERC000011"                             
## [6] "    /ENA-FIRST-PUBLIC\t4/28/2023"

The raw file has many types of information including survey attributes. These lines were chosen to be showcased for analysis and cleanup later on

survey <- biosample[grepl("/", biosample)]
head(survey)
## [1] "    /ENA-CHECKLIST\tERC000011"                                                   
## [2] "    /ENA-FIRST-PUBLIC\t4/28/2023"                                                
## [3] "    /ENA-LAST-UPDATE\t4/28/2023"                                                 
## [4] "    /External Id\tSAMEA112990854"                                                
## [5] "    /INSDC center alias\tUCSDMI"                                                 
## [6] "    /INSDC center name\tUniversity of California San Diego Microbiome Initiative"

The survey itself with its attributes and responses were then separated and converted into a dataframe with the addition of column names manually provided

survey <- read.delim(
  text = paste(survey, collapse = "\n"),
  header = FALSE,
  sep = "\t")
colnames(survey) <- c("Attribute", "Response")

*The data was then cleaned. To do this, “/” and “_” were replaced within the attribute colum to make them more readable and easier to interpret*

survey$Attribute <- gsub("/", "", survey$Attribute)
survey$Attribute <- gsub("_", " ", survey$Attribute)
head(survey)
##                Attribute
## 1          ENA-CHECKLIST
## 2       ENA-FIRST-PUBLIC
## 3        ENA-LAST-UPDATE
## 4            External Id
## 5     INSDC center alias
## 6      INSDC center name
##                                                   Response
## 1                                                ERC000011
## 2                                                4/28/2023
## 3                                                4/28/2023
## 4                                           SAMEA112990854
## 5                                                   UCSDMI
## 6 University of California San Diego Microbiome Initiative

Result

The text file given was uploaded and converted into a dataset that is human readable. The final CSV file can be used for research purposes and as a starting point for further microbiome analysis