Skip to content

Primer Design for Amplicon Sequencing

Overview

Primer design is one of the most consequential decisions in any amplicon sequencing study. A poorly chosen primer pair can introduce systematic taxonomic bias, coamplify nontarget organisms, produce excessive nonspecific products, or fail entirely on certain template types. None of these problems will be apparent from library preparation metrics alone. This page covers the key principles and common pitfalls for designing or selecting primers for community profiling by amplicon sequencing.


Primer Placement and Hypervariable Region Selection

The 16S rRNA gene contains nine hypervariable regions (V1 to V9) interspersed with conserved flanking regions used as primer binding sites. These regions aren't interchangeable: each has different discriminatory power depending on the taxonomic group, a primer pair that resolves one phylum well can perform poorly on another, simply because the region it targets happens to be more conserved in that lineage. No single primer pair captures all phylogenetically informative variation across every group, and the choice of target region is therefore a real decision, not a formality: it directly determines which taxa in your sample end up well resolved and which don't.

The general rule is: the longer the amplicon, the better the taxonomic resolution, but the greater the technical challenges. Longer amplicons increase chimera formation risk during PCR, accumulate more errors in absolute terms even at a constant per-base error rate, and cost more to sequence to sufficient depth (see Long-Read vs. Short-Read Sequencing for each of these in detail). Primer placement should therefore reflect a deliberate tradeoff between resolution and data quality rather than defaulting to the most commonly published pair.

DNA Quality Can Take This Decision Out of Your Hands

All of the above assumes you're actually free to choose. Degraded template, common with environmental DNA, formalin-fixed tissue, or older museum specimens, may simply not contain intact fragments long enough to support a longer amplicon, no matter how much resolution you'd prefer. In that situation, targeting a short region like V4 alone isn't a tradeoff between resolution and technical challenges, it's the only option that can physically amplify from what's left of the DNA. Check the fragment size distribution of your extracted DNA (a Bioanalyzer or TapeStation trace, or even a simple gel, will show you) before committing to primers spanning a long region. No amount of polymerase or primer optimisation recovers template that isn't there.

In our experience, particularly with the longer reads modern platforms like the MiSeq i100 make practical (see Overview Workflows), amplicons spanning V1 to V5 (~850 to 900bp) offer the best balance for most bacterial community profiling studies: enough hypervariable regions to resolve most taxa well, while staying short enough to sequence and process reliably. This isn't a universal rule, other labs weighing different priorities (literature continuity, cost, established pipelines) reasonably land on V3-V4 or V4-V5 instead, and that's a legitimate choice too. Studies with specific constraints on amplicon length, such as when using paired short reads with limited overlap capacity, often target V3 to V4 (~460bp) or V4 alone (~250bp), accepting reduced resolution in exchange for higher quality data.


Multicopy vs. Single-Copy Genes: A Hidden Source of Quantitative Bias

Primer and marker choice affects more than resolution and specificity, it also determines how directly your read counts relate to the actual number of organisms in your sample, and for the most commonly used markers, that relationship is weaker than it looks.

16S rRNA Gene Copy Number Varies Enormously

The 16S rRNA gene is not single-copy. Bacterial genomes carry anywhere from 1 to more than 15 copies of it, and archaea carry 1 to 4. A species with 10 copies produces roughly ten times the PCR template, and therefore roughly ten times the read count, of an equally abundant species with a single copy, purely as an artefact of genome structure, with no difference in actual cell number. This systematically inflates the apparent abundance of high-copy-number lineages (several Bacillota, for example, routinely carry 7 or more copies) relative to low-copy-number ones, and it affects diversity estimates as well as relative abundance, since it distorts the evenness of the community you appear to see.

The same underlying problem applies to fungal ITS, the standard marker for fungal community profiling: rDNA copy number varies substantially between fungal taxa, and can even vary somewhat within a species, for the same reason as 16S. Markers derived from tandemly repeated ribosomal RNA gene clusters, the rrn operon in bacteria and archaea, the analogous rDNA repeat unit in fungi, are, by their biological nature, prone to this.

Correction is possible, but imperfect. Tools like rrnDB, PICRUSt2, and CopyRighter estimate or predict copy number per taxon and rescale read counts accordingly. These corrections measurably improve agreement between amplicon-based and shotgun metagenomic abundance estimates in validation studies, but they rely on copy number being known or accurately predictable for your specific taxa, which is weakest exactly where it matters most: novel, underrepresented, or poorly characterised organisms, the same ones already hardest to classify taxonomically in the first place (see Data Prep Short-Reads (PE), Step F).

Single-copy marker genes avoid this problem structurally, since by definition there's no copy-number multiplier to correct for. The tradeoff is that single-copy protein-coding genes are typically less conserved than rRNA genes, making truly universal primers harder to design, and reference database coverage is usually far thinner than for 16S or ITS. This is part of why the functional marker genes on the Long-Read vs. Short-Read Sequencing page (ppk1, nxrB, amoA, dsrB, mcrA) are attractive for resolving specific lineages, but not a wholesale replacement for 16S as a general community-profiling marker.

The practical takeaway: treat amplicon-based relative abundance as approximate, particularly when comparing taxa with very different copy numbers, rather than as a direct proxy for relative cell counts. This is a property of the marker gene itself, not something better primers, more sequencing depth, or a different clustering method can fix.


Degeneracy: Coverage vs. Specificity

Because conserved primer binding regions still show sequence variation across the breadth of the bacterial domain, most universal primers are designed with degenerate bases, meaning positions where a mixture of nucleotides is incorporated to allow the primer pool to amplify a broader range of template sequences.

IUPAC degeneracy codes

The most commonly encountered degenerate codes in 16S primers are:

Code Bases Code Bases
M A/C Y C/T
R A/G W A/T
S C/G K G/T
H A/C/T B C/G/T
N A/C/G/T

Degeneracy is essential for broad phylogenetic coverage but comes at a cost. A highly degenerate primer creates a complex mixture of molecules in the PCR reaction, and variants that match abundant or efficiently amplified templates tend to become overrepresented during early PCR cycles, introducing amplification bias against rare or phylogenetically divergent taxa. The degree of degeneracy should therefore be kept as low as possible while still achieving adequate coverage of the target lineages.

The Actual Primer Mix Is Never Quite What You Ordered, and Never Quite the Same Twice

A degenerate position isn't synthesised by picking one base per molecule to match a target ratio. It's made by mixing the relevant phosphoramidite reagents, nominally in equal proportion, at that coupling step, and letting the chemistry sort out which base actually ends up in each individual oligo molecule. In practice this mixing is rarely perfectly equimolar: different bases couple with slightly different efficiency (G, for instance, tends to couple somewhat less efficiently than the others), and synthesiser-specific factors like reagent bottle position and distance from the synthesis column add further variation. The result is that the real, delivered ratio of primer variants in a "degenerate" primer pool deviates from the nominal ratio implied by its IUPAC code, and because each synthesis run is its own random process, that deviation isn't identical from one order to the next, even for the exact same primer sequence from the same supplier.

This matters beyond an abstract QC concern: if the relative proportion of primer variants shifts between batches, the relative amplification efficiency across your target taxa can shift with it, in a way that has nothing to do with your samples. A result that looked reproducible within one sequencing run might not reproduce as cleanly with a new primer order for a follow-up study. This is a good argument for ordering enough primer for an entire project in a single synthesis batch where feasible, and for treating primer batch as a variable worth recording alongside your other metadata, the same way you'd record a reagent lot number.

For particularly divergent lineages, the alternative to high degeneracy is primer cocktails: separate oligos designed for different subgroups, pooled before PCR. This avoids the dilution effect of degenerate pools but requires careful normalisation of component concentrations.


3' End Stability and Protection Against Exonuclease Activity

The proofreading problem

High fidelity polymerases used in amplicon sequencing workflows, such as Phusion, Q5, or KAPA HiFi, possess 3' to 5' exonuclease (proofreading) activity. This activity is the source of their accuracy in replication, but it creates a significant problem when amplifying diverse templates with degenerate primers.

The mechanism depends on the number of consecutive mismatches at the 3' end of the primer. When only a single terminal base is mismatched, the proofreading polymerase removes it, leaving the new 3' end correctly matched. Extension then proceeds from the slightly shortened primer, meaning the template is still amplified. In this case proofreading is not a problem and the successful amplification of those templates can actually be observed in the sequencing data: because the primer sequence is physically incorporated into the amplicon, reads originating from templates with a mismatch at the primer site will carry that variant in the primer binding region at a consistent position and frequency, distinguishable from the low and randomly distributed background of sequencing errors.

One important failure mode is thought to arise when multiple consecutive mismatches are present at the 3' end, as is likely for degenerate primers amplifying highly divergent templates. The polymerase removes one base, encounters another mismatch, removes that too, and continues chewing back progressively. Each removal shortens the primer and reduces its melting temperature, until the primer dissociates from the template entirely before extension can begin. These templates are silently excluded from the library: they fail to amplify, produce no reads, and their absence is indistinguishable from genuine biological absence. This is likely one contributor to taxonomic bias, and it disproportionately affects rare or phylogenetically divergent lineages whose sequences diverge most from the primer consensus.

This chew-back behaviour, more formally called primer editing, has been directly measured rather than just inferred: Gohl, Auch, Certano, LeFrançois, Bouevitch, Doukhanine, Fragel, Macklaim, Hollister, Garbe, and Beckman (2021) Nucleic Acids Research 49(15):e87 built synthetic DNA standards specifically to quantify how much proofreading polymerases edit primers, and confirmed the effect is concentration-dependent.

Perfect match:           |||||||||| (10/10 bases matched)
                          → extension proceeds normally

Single 3' mismatch:      |||||||||X (1 mismatch at 3' end)
                          → proofreading removes the mismatched base
                          → extension proceeds from the shortened primer
                          → template is amplified, mismatch visible in the reads

Three 3' mismatches:      |||||||XXX (3 consecutive mismatches at 3' end)
                          → proofreading removes one base, hits another mismatch
                          → removes that too, hits a third
                          → primer progressively shortens and its Tm drops
                          → primer dissociates before extension
                          → template is silently excluded, no reads, no trace

Proofreading bias is partially detectable but never fully visible

Templates that were amplified despite a single 3' mismatch leave a detectable trace: consistent positional variants in the primer binding region of the reads. Templates that failed to amplify due to progressive 3' chew-back leave no trace at all. Studies comparing high fidelity and nonproofreading polymerases on the same environmental samples have documented substantial differences in recovered community composition, particularly for rare or phylogenetically divergent taxa, reflecting both effects simultaneously.

See Insights & Solutions for how we investigate suspected false priming directly in real sequencing data, by examining mismatch patterns at the primer binding site.

Phosphorothioate modifications

The standard solution is to incorporate phosphorothioate (PS) bonds at the 3' end of the primer. In a phosphorothioate linkage, one of the nonbridging oxygen atoms in the phosphodiester backbone is replaced by a sulphur atom. This modification is invisible to the DNA polymerase during extension but renders the bond resistant to degradation by both 3' to 5' exonucleases (proofreading) and 5' to 3' exonucleases (such as those involved in nick translation).

In practice, two to three phosphorothioate bonds at the 3' terminus are sufficient to protect the primer from proofreading degradation while minimally affecting hybridisation thermodynamics. The modification is incorporated during chemical synthesis and is a standard offering from most oligo suppliers. Gohl et al. (2021), cited above, directly demonstrated this tuning effect using their primer-editing standards: phosphorothioate linkages measurably reduced the extent of primer editing across multiple proofreading polymerases (KAPA HiFi, Q5, and Phusion), which is the direct experimental support behind this recommendation, not just a theoretical expectation from the chemistry.

The tradeoff is a modest reduction in specificity: by preventing the polymerase from removing a mismatched 3' base, you allow extension from imperfectly matched primer to template duplexes that would otherwise be aborted. For environmental amplicon sequencing, where broad taxonomic coverage is the goal, this is an acceptable and generally desirable tradeoff. In applications where specificity is paramount, for example detection of a specific pathogen in a clinical sample, unprotected primers with a proofreading polymerase may be preferable.

Polymerase choice

An alternative to chemical modification is to use a nonproofreading polymerase such as Taq. Taq lacks 3' to 5' exonuclease activity and will extend from mismatched 3' ends, providing broader template coverage. However, Taq also lacks the accuracy of high fidelity enzymes, introducing substitution errors at a rate approximately tenfold higher than proofreading polymerases. For short amplicons (<500bp) this is often acceptable; for longer amplicons (>800bp) the accumulation of errors becomes problematic, particularly for ASV based workflows where exact sequences are used as biological units.

The practical recommendation for long amplicons in diverse communities is therefore to use a high fidelity polymerase with phosphorothioate protected degenerate primers, combining broad template coverage with low per base error rates.


Nontarget Coamplification

One of the most frequently overlooked aspects of primer design is specificity toward the intended organismal group. Universal primers designed to amplify broad prokaryotic or eukaryotic diversity will, by definition, amplify any template that matches their binding site, including templates from organisms that the researcher did not intend to sequence. This is not a sequencing artefact but a genuine PCR outcome, and it can dominate libraries in samples where nontarget organisms are abundant.

ITS primers and plant coamplification

The internal transcribed spacer (ITS) region is the standard marker for fungal community profiling. ITS1 and ITS2 are widely used, and the most common primer pairs (ITS1F/ITS2, ITS3/ITS4, and their variants) are often described as "fungal specific" in the literature. In practice, however, their binding sites are present in the ribosomal RNA gene cluster of all eukaryotes, including plants (a tandem rDNA repeat unit, not a bacterial-style operon), and plant ITS sequences are frequently coamplified when environmental samples contain plant material.

Plant coamplification in ITS studies

In soil, litter, root, phyllosphere, or any plant associated sample, plant ITS sequences can represent the majority of recovered reads when standard fungal primers are used, often outcompeting fungal templates due to the high copy number of plant nuclear rDNA. This can reduce fungal sequencing depth by an order of magnitude and, if not recognised, leads to severe underestimation of fungal diversity.

The problem is particularly acute with:

ITS1 (ITS1F / ITS2 primer pair): Plant ITS1 sequences are coamplified at high efficiency with most standard fungal primers, especially in root and soil samples.

ITS2 (ITS3 / ITS4 primer pair): ITS2 shows somewhat better fungal to plant discrimination but is not immune, particularly in samples with high plant biomass.

Several approaches exist to reduce nontarget coamplification, each with practical limitations discussed below.

PNA blocking oligonucleotides

The most widely cited approach is the use of peptide nucleic acid (PNA) blocking probes: nonextendable synthetic oligomers that competitively occupy the nontarget binding site and prevent primer extension from those templates. Well known examples include the plant ITS blocking PNAs for fungal studies and the chloroplast blocking PNA (PGBC) and mitochondria blocking PNA (PMBC) probes used with the 515F/806R primer pair for bacterial 16S studies.

When they work, PNA blockers can suppress nontarget reads from over 80% to under 5% of the library. In practice, however, they are among the more technically demanding reagents in amplicon workflows and frequently require substantial optimisation before they perform reliably.

Without PNA blocker:              With PNA blocker:

Primer site (plant + fungal)      Primer site (plant + fungal)
     ↓                                 ↓
 both bind, both amplify           PNA occupies plant site first
     ↓                                 ↓
 library dominated by plant        primer can't bind there
 reads (up to ~90%+)                   ↓
                                    primer binds fungal site instead
                                        ↓
                                    fungal reads recovered at
                                    much higher proportion

PNA blockers are difficult to work with in practice

Several factors make PNA blockers unreliable without careful empirical calibration.

Concentration is critical and sample dependent. PNA blockers compete with primers for the same binding site. Too little PNA leaves blocking incomplete; too much suppresses legitimate target amplification. The optimal concentration window is often narrow and shifts with the ratio of target to nontarget DNA, which varies between samples in an environmental study. Reoptimisation for each new sample type is often necessary.

Annealing temperature requires careful calibration. PNA hybridisation kinetics differ from DNA, and most protocols include a dedicated PNA annealing step at a temperature above the primer annealing temperature to ensure the PNA occupies its site before the primer competes. The temperature window where the PNA blocks efficiently without disrupting primer binding can be very tight and must be determined empirically.

Sequence matching to local flora is often imperfect. Most published PNA blockers were designed against Arabidopsis type or common crop plant sequences. In samples dominated by other plant groups, grasses, tropical taxa, or locally abundant species, the PNA may carry mismatches against the actual nontarget sequences present, substantially reducing blocking efficiency. A blocker validated in an agricultural soil study in central Europe may fail in a sample from a different ecosystem.

Stability and cost. PNA synthesis requires specialised chemistry unavailable in most molecular biology labs, making researchers dependent on commercial suppliers at relatively high per-oligo cost compared to standard DNA primers. PNA blockers are stable under PCR conditions but degrade with repeated freeze-thaw cycles, and loss of efficacy is not readily detectable without running parallel controls.

Alternative approaches

Where PNA blockers are impractical or underperform, two alternatives are worth considering.

Restriction enzyme digestion after PCR can selectively eliminate nontarget amplicons when a restriction site is conserved in the nontarget sequence but absent in the target. This approach is less universal but more robust and cheaper when an appropriate restriction site exists. It requires prior knowledge of the nontarget sequences likely to be present and verification that the chosen site is absent across the full diversity of target templates.

Bioinformatic filtering is increasingly feasible as reference databases improve. Reads assigned to nontarget taxa are simply removed after classification, accepting that their presence in the library reduces the sequencing depth available for target organisms. This approach is least satisfying when nontarget reads dominate the library, as useful target depth may be too low to recover, but it is a practical fallback when wet lab solutions have failed or when the degree of contamination was not anticipated at the study design stage. See Data Prep Short-Reads (PE), Step F for how taxonomic assignment, the step this filtering relies on, actually works.

16S primers and mitochondrial/chloroplast coamplification

Mitochondria and chloroplasts retain 16S like rRNA genes with sequences similar enough to bind many universal bacterial primers. In samples containing eukaryotic cells, including many environmental and host associated samples, mitochondrial and chloroplast 16S sequences can constitute a substantial fraction of the library. This is particularly relevant in plant associated microbiome studies (rhizosphere, phyllosphere, endophytes) where chloroplast sequences routinely represent 20 to 60% of reads with standard primers, in animal gut microbiome studies using biopsy or tissue samples where mitochondrial reads can be abundant, and in eukaryotic cell culture experiments where bacterial contamination is being assessed.

The same blocking approaches described above apply here, with the same practical caveats. PNA blockers for chloroplast and mitochondrial sequences require the same careful concentration and temperature optimisation, and their sequence coverage of diverse plant lineages is similarly imperfect. Bioinformatic filtering is straightforward for mitochondrial and chloroplast reads since their taxonomy is well represented in reference databases.

18S primers and metazoan coamplification

Studies targeting protists or microeukaryotes using 18S rRNA primers face analogous problems in samples containing metazoan biomass. Animal 18S sequences are highly conserved and bind most universal eukaryotic primers efficiently, often dominating libraries from sediment, water, or gut samples. As with ITS and 16S, the choice between blocking probes, restriction digestion, and bioinformatic filtering depends on the degree of contamination expected and the sequencing depth available.

12S primers and off target vertebrate amplification

12S rRNA primers used for vertebrate eDNA metabarcoding (fish, mammals) in mixed environmental samples may coamplify nontarget vertebrate groups depending on primer placement. A primer pair optimised for fish detection in aquatic samples may also amplify amphibian, bird, or mammal templates present in the same sample, requiring bioinformatic filtering or blocking probes for unwanted groups.


Primer Evaluation Before Sequencing

Before committing to a primer pair for a large sequencing experiment, several evaluation steps are worth performing.

In silico coverage analysis using tools such as TestPrime (SILVA) or Primer Prospector allows estimation of how well a primer matches reference sequences across different taxonomic groups. Coverage should be evaluated not only at the domain level but for the specific phyla and classes expected in the sample. A primer may show 95% overall bacterial coverage while poorly matching the Planctomycetes or Verrucomicrobia that are the focus of the study.

In silico amplification using tools such as PrimerProspector or ecoPCR (from the OBITools suite) against a local database allows identification of nontarget taxa that would be coamplified, as well as estimation of expected amplicon size distributions across taxa.

Empirical testing on mock communities with known composition, ideally including phylogenetically divergent taxa and nontarget organisms expected in real samples, provides a direct measure of amplification bias before the experiment scales up.

Pilot sequencing on a small number of real samples, even just four or five, is worth doing before committing to a full run. It often reveals problems that neither in silico prediction nor mock communities catch, since mock communities are, by design, easier templates than messy real environmental DNA. See Insights & Solutions for what to look for once you have pilot data, primer mismatch patterns in particular.


Summary of Key Recommendations

Consideration Recommendation
Amplicon length Match to platform and diversity; V1 to V5 (~900bp) for broad bacterial diversity on MiSeq i100, in our experience, though V3-V4/V4-V5 remain reasonable choices depending on your priorities
Marker copy number Treat relative abundance as approximate; correct with rrnDB/PICRUSt2/CopyRighter where feasible, or consider a single-copy marker if quantitative accuracy is the priority
Degeneracy Use the lowest degeneracy sufficient to cover target lineages; consider primer cocktails for highly divergent groups
Primer batch variability Order enough primer for a whole project in one synthesis batch where feasible; record batch/lot alongside other metadata
3' protection Add two to three phosphorothioate bonds at the 3' end when using high fidelity polymerases with degenerate primers targeting diverse templates
Polymerase High fidelity with PS protected primers for long amplicons; Taq acceptable for short amplicons where error rate is less critical
Nontarget coamplification Evaluate all primers in silico against expected nontarget taxa before use; use blocking probes in samples with high nontarget biomass
ITS / plant samples Always test for plant coamplification; use PNA blockers in plant associated samples
16S / eukaryotic samples Use chloroplast and mitochondrial blocking probes in plant associated or tissue samples
Pre-experiment validation Run in silico coverage analysis, mock community tests, and pilot sequencing before large scale sequencing