20.1 Primary Tool Papers
- Andrews S. (2010). FastQC. Babraham Bioinformatics. [Link]
- Cock P.J.A. et al. (2010). “The Sanger FASTQ file format.” Nucleic Acids Research 38(6):1767–1771. [Link]
- Ewing B. & Green P. (1998). “Base-calling of automated sequencer traces using phred II.” Genome Research 8:186–194. [PubMed]
- Chen S. et al. (2018). “fastp: an ultra-fast all-in-one FASTQ preprocessor.” Bioinformatics 34(17):i884–i890. [Link]
- Ewels P. et al. (2016). “MultiQC: summarize analysis results for multiple tools and samples.” Bioinformatics 32(19):3047–3048. [Link]
- Li H. (2013). “Aligning sequence reads, clone sequences and assembly contigs with BWA-MEM.” arXiv:1303.3997. [arXiv]
- Vasimuddin Md. et al. (2019). “Efficient Architecture-Aware Acceleration of BWA-MEM for Multicore Systems.” IPDPS. [arXiv]
- Li H. (2018). “Minimap2: pairwise alignment for nucleotide sequences.” Bioinformatics 34(18):3094–3100. [Link]
- Li H. et al. (2009). “The Sequence Alignment/Map format and SAMtools.” Bioinformatics 25(16):2078–2079. [Link]
- DePristo M.A. et al. (2011). “A framework for variation discovery and genotyping using next-generation DNA sequencing data.” Nature Genetics 43:491–498. [Link]
- Van der Auwera G.A. & O’Connor B.D. (2020). Genomics in the Cloud. O’Reilly Media.
- Poplin R. et al. (2018). “Scaling accurate genetic variant discovery to tens of thousands of samples.” bioRxiv 201178. [bioRxiv]
- Poplin R. et al. (2018). “A universal SNP and small-indel variant caller using deep neural networks.” Nature Biotechnology 36:983–987. [Link]
- Benjamin D. et al. (2019). “Calling Somatic SNVs and Indels with Mutect2.” bioRxiv 861054. [bioRxiv]
- Cibulskis K. et al. (2013). “Sensitive detection of somatic point mutations in impure and heterogeneous cancer samples.” Nature Biotechnology 31:213–219.
- Chen X. et al. (2016). “Manta: rapid detection of structural variants and indels.” Bioinformatics 32(8):1220–1222.
- Talevich E. et al. (2016). “CNVkit: Genome-Wide Copy Number Detection.” PLOS Computational Biology 12(4):e1004873. [Link]
- Danecek P. et al. (2011). “The variant call format and VCFtools.” Bioinformatics 27(15):2156–2158.
- McLaren W. et al. (2016). “The Ensembl Variant Effect Predictor.” Genome Biology 17:122. [Link]
- Chen S. et al. (2024). “A genomic mutational constraint map using variation in 76,156 human genomes.” Nature 625:92–100 (gnomAD v4). [Link]
- Landrum M.J. et al. (2018). “ClinVar: improving access to variant interpretations.” Nucleic Acids Research 46(D1):D1062–D1067.
- Richards S. et al. (2015). “Standards and guidelines for the interpretation of sequence variants.” Genetics in Medicine 17:405–424. [Link]
- Dobin A. et al. (2013). “STAR: ultrafast universal RNA-seq aligner.” Bioinformatics 29(1):15–21.
- Patro R. et al. (2017). “Salmon provides fast and bias-aware quantification.” Nature Methods 14:417–419. [Link]
- Love M.I., Huber W., Anders S. (2014). “Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2.” Genome Biology 15:550. [Link]
- Robinson M.D., McCarthy D.J., Smyth G.K. (2010). “edgeR: a Bioconductor package for differential expression analysis.” Bioinformatics 26(1):139–140.
- Subramanian A. et al. (2005). “Gene set enrichment analysis.” PNAS 102(43):15545–15550. [Link]
- Liao Y., Smyth G.K., Shi W. (2014). “featureCounts.” Bioinformatics 30(7):923–930.
- Zheng G.X.Y. et al. (2017). “Massively parallel digital transcriptional profiling of single cells.” Nature Communications 8:14049.
- Kaminow B., Yunusov D., Dobin A. (2021). “STARsolo.” bioRxiv 2021.05.05.442755.
- Wolf F.A., Angerer P., Theis F.J. (2018). “SCANPY: large-scale single-cell gene expression data analysis.” Genome Biology 19:15.
- Luecken M.D. & Theis F.J. (2019). “Current best practices in single-cell RNA-seq analysis.” Molecular Systems Biology 15:e8746.
- Hafemeister C., Satija R. (2019). “Normalization and variance stabilization of single-cell RNA-seq data using SCTransform.” Genome Biology 20:296.
- McInnes L., Healy J., Melville J. (2018). “UMAP: Uniform Manifold Approximation and Projection.” arXiv:1802.03426.
- Traag V.A., Waltman L., van Eck N.J. (2019). “From Louvain to Leiden.” Scientific Reports 9:5233.
- Hao Y. et al. (2021). “Integrated analysis of multimodal single-cell data.” Cell 184(13):3573–3587.
- Korsunsky I. et al. (2019). “Fast, sensitive and accurate integration of single-cell data with Harmony.” Nature Methods 16:1289–1296.
- Bergen V. et al. (2020). “Generalizing RNA velocity to transient cell states through dynamical modeling.” Nature Biotechnology 38:1408–1414.
- Jin S. et al. (2021). “Inference and analysis of cell-cell communication using CellChat.” Nature Communications 12:1088.
- Efremova M. et al. (2020). “CellPhoneDB.” Nature Protocols 15:1484–1506.
- Tirosh I. et al. (2016). “Dissecting the multicellular ecosystem of metastatic melanoma by single-cell RNA-seq.” Science 352:189–196.
- Gao R. et al. (2021). “Delineating copy number and clonal substructure from single-cell transcriptomes.” Nature Biotechnology 39:599–608 (CopyKAT).
- Büttner M. et al. (2021). “scCODA is a Bayesian model for compositional single-cell data analysis.” Nature Communications 12:6876.
- Wherry E.J., Kurachi M. (2015). “Molecular and cellular insights into T cell exhaustion.” Nature Reviews Immunology 15:486–499.
- Miller B.C. et al. (2019). “Subsets of exhausted CD8+ T cells differentially mediate tumor control.” Nature Immunology 20:326–336.
- Di Tommaso P. et al. (2017). “Nextflow enables reproducible computational workflows.” Nature Biotechnology 35:316–319.
- Ewels P.A. et al. (2020). “The nf-core framework for community-curated bioinformatics pipelines.” Nature Biotechnology 38:276–278.
- Nurk S. et al. (2022). “The complete sequence of a human genome.” Science 376:44–53. [Link]
- Wang Y. et al. (2021). “Nanopore sequencing technology, bioinformatics and applications.” Nature Biotechnology 39:1348–1365.
- Wenger A.M. et al. (2019). “Accurate circular consensus long-read sequencing improves variant detection.” Nature Biotechnology 37:1155–1162.
- Benjamini Y., Hochberg Y. (1995). “Controlling the False Discovery Rate.” JRSS-B 57(1):289–300. [Link]
- Li M.M. et al. (2017). “Standards and Guidelines for the Interpretation and Reporting of Sequence Variants in Cancer.” JMD 19(1):4–23.
- Burrows M., Wheeler D.J. (1994). “A block-sorting lossless data compression algorithm.” DEC Technical Report 124. [PDF]
- Sanger F., Nicklen S., Coulson A.R. (1977). “DNA sequencing with chain-terminating inhibitors.” PNAS 74(12):5463–5467. [Link]
- Lander E.S. et al. (2001). “Initial sequencing and analysis of the human genome.” Nature 409:860–921.
- Wolock S.L., Lopez R., Klein A.M. (2019). “Scrublet: Computational Identification of Cell Doublets.” Cell Systems 8(4):281–291.
- Aran D. et al. (2019). “Reference-based analysis of lung single-cell sequencing.” Nature Immunology 20:163–172 (SingleR).
- Stuart T. et al. (2019). “Comprehensive Integration of Single-Cell Data.” Cell 177(7):1888–1902.
- Lun A.T.L. et al. (2019). “EmptyDrops.” Genome Biology 20:63.
- Cheng J. et al. (2023). “Accurate proteome-wide missense variant effect prediction with AlphaMissense.” Science 381:eadg7492.
- Kircher M. et al. (2014). “A general framework for estimating the relative pathogenicity of human genetic variants.” Nature Genetics 46:310–315 (CADD).
- Jaganathan K. et al. (2019). “Predicting Splicing from Primary Sequence with Deep Learning.” Cell 176:535–548 (SpliceAI).
- Li M.M. et al. (2017). “Standards and Guidelines for Somatic Variant Interpretation.” JMD 19(1):4–23.