Literature
Curated papers on sample prep, primer design, sequencing platforms, bioinformatics, and analysis, with a short note on why each matters.
This isn't an exhaustive bibliography, it's the papers we actually reach for when a project runs into one of these questions, or that back up a specific claim made elsewhere on this site. Where a paper connects directly to something we discuss in more depth, we've linked to it.
Sample Preparation (Wet-Lab)
A noninvasive, cost- and time-effective alternative to actively filtering DNA from water samples.
Bessey et al. (2021) Passive eDNA collection enhances aquatic biodiversity analysis. Communications Biology 4:236.
PCR inhibitors can originate from the sample itself, the extraction method, or even the plastics used during sample preparation.
Schrader et al. (2012) PCR inhibitors: occurrence, properties and removal. Journal of Applied Microbiology 113:1014-1026.
Synthetic spike-ins to improve the reliability and accuracy of amplicon sequencing microbiome studies.
Tourlousse et al. (2017) Synthetic spike-in standards for high-throughput 16S rRNA gene amplicon sequencing. Nucleic Acids Research 45(4):e23.
Primer Design
The flagellin gene (fliC) as a single-copy alternative to 16S amplicon sequencing. Directly relevant to the copy-number problem discussed on the Primer Design page: a single-copy marker sidesteps the issue structurally, at the cost of a much thinner reference database.
Hu et al. (2022) Single-gene long-read sequencing illuminates Escherichia coli strain dynamics in the human intestinal microbiome. Cell Reports 38.
The database behind the 16S copy-number numbers we cite on the Primer Design page (1 to 15+ copies across bacteria), and the basis for copy-number correction tools like PICRUSt2.
Stoddard, Smith, Hein, Roller, and Schmidt (2015) rrnDB: improved tools for interpreting rRNA gene abundance in bacteria and archaea and a new foundation for future development. Nucleic Acids Research 43(D1):D593-D598.
The original peptide nucleic acid (PNA) clamp paper behind the plant-DNA and chloroplast/mitochondria blocking approaches discussed on the Primer Design page.
Lundberg, Yourstone, Mieczkowski, Feng, Vogel, and Dangl (2013) Practical innovations for high-throughput amplicon sequencing. Nature Methods 10:999-1002.
Sequencing Platforms & Technology
LoopSeq, a commercially available synthetic long-read (SLR) sequencing technology for Illumina short-read platforms, an alternative worth knowing about alongside the PacBio-vs-Illumina tradeoffs discussed on the Long-Read vs. Short-Read Sequencing page.
Callahan et al. (2021) Ultra-accurate microbial amplicon sequencing with synthetic long reads. Microbiome 9:130.
Bioinformatics & Clustering
The original UPARSE paper: 97% OTU clustering, as implemented in USEARCH and discussed throughout the Data Prep Short-Reads (PE) page.
Edgar (2013) UPARSE: highly accurate OTU sequences from microbial amplicon reads. Nature Methods 10:996-998.
UNOISE2: the denoising (zOTU) approach behind our default clustering recommendation.
Edgar (2016) UNOISE2: improved error-correction for Illumina 16S and ITS amplicon sequencing. bioRxiv 081257.
Chimera detection and why "error-free chimera prediction from sequence is impossible in principle," relevant background for the chimera risk discussion on the Long-Read vs. Short-Read Sequencing page.
Edgar (2016) UCHIME2: improved chimera prediction for amplicon sequencing. bioRxiv 074252.
Taxonomic Annotation & Reference Databases
A conceptual framework for the identification of fungi.
Lücking et al. (2020) Unambiguous identification of fungi: where do we stand and how accurate and precise is fungal DNA barcoding? IMA Fungus 11:14.
The SINTAX classifier we use for taxonomic assignment, described in full on the Data Prep Short-Reads (PE), Step F page.
Edgar (2016) SINTAX, a simple non-Bayesian taxonomy classifier for 16S and ITS sequences. bioRxiv 074161.
Open versus closed taxonomic databases for eDNA community assignment, cited throughout this site wherever reference database choice comes up.
Blackman, Walser, Rüber, Brantschen, Villalba, Brodersen, Seehausen, and Altermatt (2023) General principles for assignments of communities from eDNA: Open versus closed taxonomic databases. Environmental DNA 5(2):326-342.
Data Analysis & Statistics
A critique of the rapid growth of interest in microbiota research, highlighting concerns about oversimplification, misinterpretation, and the need for real expertise in the field.
Boers, Boucher, and Wielinga (2016) Suddenly everyone is a microbiota specialist. Clinical Microbiology and Infection 22(7):581-582.
A discussion of different approaches to defining a core microbiome.
Neu, Allen, and Roy (2021) Defining and quantifying the core microbiome: Challenges and prospects. PNAS 118(51).
Rarefaction as an approach to controlling for uneven sequencing depth. See the OTU/Count Table page for why this is a genuinely contested question, not a settled one, and for the counterpoint immediately below.
Schloss (2024) Rarefaction is currently the best approach to control for uneven sequencing effort in amplicon sequence analyses. mSphere 9(2).
The other side of the rarefaction debate. Read alongside Schloss above, not instead of it.
McMurdie and Holmes (2014) Waste not, want not: why rarefying microbiome data is inadmissible. PLoS Computational Biology 10:e1003531.
PICRUSt2: predicting functional potential from marker-gene data, including the kind of copy-number-aware correction relevant to the Primer Design page's discussion of quantitative bias.
Douglas, Maffei, Zaneveld, Yurgel, Brown, Taylor, Huttenhower, and Langille (2020) PICRUSt2 for prediction of metagenome functions. Nature Biotechnology 38:685-688.