Skip to content

The GDC has been helping students and researchers with amplicon-based sequencing (metabarcoding) for many years. We take every project seriously, because it matters, to you and to us. We've worked closely with our users to develop lab protocols, solve problems with tricky samples, refine sampling designs, streamline data preparation and assist with data analysis.

Everything starts with your scientific question. It shapes the sampling design, the sample preparation, the sequencing technology you choose, and even how you analyse the data afterward. Get the question clear early, and the rest of the project has something solid to build on.

We like to help as much as we can, but honestly, sometimes the data doesn't meet expectations, or a run simply doesn't deliver enough, or good enough, data. Things don't always go as planned, and no amount of careful planning eliminates that entirely. When it happens, we don't leave you out in the rain: we'll help you understand what went wrong and figure out what to do next.

While every project is unique, some challenges are common. This website is a growing collection of frequently requested help topics. It's not exhaustive, and it's certainly open to improvement, but we hope it proves helpful.

What You Will Find Here

The site follows the natural path of an amplicon sequencing project: from your first scientific question, through generating and processing raw reads, to finished analysis, plus a place to go when something along the way doesn't work out.

Scientific question
        │
        ▼
Study & primer design   ◄──  Planning Your Project
        │
        ▼
Sampling, extraction, PCR, sequencing
        │
        ▼
Raw reads                ◄──  Raw-Read Preparation
        │
        ▼
Count (OTU/zOTU) table
        │
        ▼
R / phyloseq              ◄──  Downstream Analysis
        │
        ▼
Results

     Troubleshooting: relevant at any stage above
  • Planning Your Project


    Primer design, reference database choice, sample ID encoding, and project scope, decisions worth making before you commit to a sequencing strategy.

    Start planning

  • Raw-Read Preparation


    Get from raw sequencing output to a clean count (OTU) table: download, demultiplexing, primer trimming, quality filtering, and clustering, for both short and long reads.

    Start here

  • Downstream Analysis


    Set up your R environment, import your count table, check your controls, and get ready for downstream statistics.

    Get ready

  • Troubleshooting


    Practical notes from real projects: PCR issues, platform-specific quirks, contamination, and how to interpret your results correctly.

    Browse topics

Not sure where to start? Head to Overview Workflows for a bird's-eye view of the whole process. For background reading rather than a specific problem, see the Literature Collection at the bottom of the site map below.

Full Site Map

Planning Your Project

  • Primer Design: primer placement, degeneracy and 3' protection, avoiding nontarget coamplification
  • Choosing a Reference Database: coverage, reliability, and outgroups when choosing a taxonomic reference
  • Sample IDs: building short, systematic codes for your samples, not just descriptive names
  • Project Scope: why combining more samples into one project isn't just a storage problem, it's a complexity and bias one
  • Long versus Short: weighing PacBio full-length 16S against extended Illumina amplicons for diverse communities

Raw-Read Preparation

  • Overview Workflows: how Illumina, PacBio and AVITI data differ, and how to pick a clustering strategy for each
  • Data Download: a checklist and terminal commands to sanity-check your raw files the moment they arrive
  • Data Prep Short-Reads (PE): the full paired-end Illumina workflow, from quality control to taxonomic assignment
  • Data Prep Short-Reads (SE): the fallback workflow for poor read overlap, based on a trimmed R1 read alone
  • Data Prep Long-Reads: the PacBio CCS/HiFi workflow, from deconcatenation and orientation to annotated count tables
  • Data Prep Output: what's in your results package, and where to find the files you actually need
  • Count (OTU) Table: what a count table looks like, why it behaves the way it does, and how to normalise it
  • Understanding Your Summary: what the numbers in your data preparation summary email actually mean

Downstream Analysis

  • Get Ready: installing R packages and structuring your scripts before you start analysing
  • Data Import: getting your count table, map file and tree into a phyloseq object
  • Controls: what your positive and negative controls tell you, and the first things to check once your data is imported

Troubleshooting

  • Insights & Solutions: practical notes from real projects, PCR troubleshooting, platform-specific quirks, and classifiers worth knowing about
  • Identifying Contaminants: statistical contaminant identification with decontam, once you've checked your controls directly
  • Cluster vs. Species: why neither OTUs nor zOTUs are species, and what that means for richness and diversity comparisons

Literature Collection

  • Literature: curated papers on sample prep, primer design, sequencing platforms, bioinformatics, and analysis, with a short note on why each matters

Talk to Us

We have a lot of experience with amplicon sequencing, but not all the answers. Every sample, every study system, and every research question brings something we haven't seen before, and we're happy to learn alongside you rather than pretend otherwise. What we can offer is a second pair of eyes: help thinking through a sampling design, troubleshooting a difficult sample, or working out why a run didn't go the way you expected. Some of the best solutions we've found came out of working through a problem together, not from us already knowing the answer going in. The earlier you talk to us, the more we can actually do, while decisions are still open rather than after they're locked in.