Skip to content

Sequencing Methods

High-throughput sequencing has transformed biology over the past two decades, giving rise to a bewildering variety of methods, acronyms and applications. This page provides a structured overview of the most widely used sequencing approaches, explaining what each method does, how it works in principle, and what kinds of questions it is designed to answer. Methods are listed alphabetically for easy reference.

A note on terminology

Many of these methods are platform-independent in principle but are most commonly implemented on Illumina short-read or Oxford Nanopore long-read platforms. Where platform choice matters, this is noted in the relevant entry.


AmpSeq (Amplicon Sequencing) and Metabarcoding

What it is: Targeted sequencing of specific genomic regions that have been amplified by PCR prior to sequencing. When applied to environmental samples for species identification using a standardised marker gene, this approach is known as metabarcoding.

How it works: Primers are designed to flank one or more regions of interest. PCR amplifies these regions from all samples simultaneously, and the resulting amplicons are sequenced in a highly multiplexed fashion. Because sequencing is focused on a small target region, very high sequencing depth is achievable at low cost.

Typical applications: Metabarcoding uses marker genes such as 16S rRNA (bacteria and archaea), ITS (fungi), or COI (animals) to identify which species are present in a mixed environmental sample such as soil, water, or gut content. More broadly, AmpSeq is also used for genotyping at known SNP panels, pathogen detection, and population genetics using targeted markers.

Key outputs: Species composition tables, allele frequencies, haplotype lists.

Limitations: Restricted to pre-defined target regions. PCR amplification introduces bias, and primer mismatches can cause certain variants or taxa to be under-represented or missed entirely.


ATAC-Seq (Assay for Transposase-Accessible Chromatin)

What it is: A method for mapping open chromatin regions across the genome, revealing where the DNA is accessible to transcription factors and other regulatory proteins.

How it works: A hyperactive transposase enzyme (Tn5) is used to simultaneously cut and tag accessible regions of chromatin with sequencing adapters. These tagged fragments are then sequenced. Regions that are densely packed (heterochromatin) are inaccessible and therefore not tagged, while open regions (euchromatin) are preferentially captured.

Typical applications: Mapping regulatory elements such as enhancers and promoters, studying chromatin remodelling during development or in response to stimuli, and identifying transcription factor binding sites.

Key outputs: Genome-wide maps of chromatin accessibility, peak files indicating open regions.

Limitations: Requires very fresh, intact nuclei, making sample preparation technically demanding. Results are sensitive to the quality of the input material.


Bisulfite Sequencing (BS-Seq / WGBS)

What it is: A method for mapping DNA methylation at single-base resolution across the entire genome.

How it works: DNA is treated with sodium bisulfite, which chemically converts unmethylated cytosines to uracil (read as thymine after sequencing), while methylated cytosines remain unchanged. By comparing bisulfite-treated sequences to a reference genome, the methylation status of each cytosine can be determined.

Typical applications: Epigenetic studies, imprinting, cancer epigenomics, development, and the study of gene regulation through promoter methylation.

Key outputs: Genome-wide methylation maps at single-CpG resolution.

Limitations: Bisulfite treatment degrades DNA significantly, requiring high input quantities. The method also introduces severe sequence complexity reduction, making alignment and analysis more challenging. Targeted versions (RRBS, Reduced Representation Bisulfite Sequencing) exist for lower-cost alternatives.


ChIP-Seq (Chromatin Immunoprecipitation Sequencing)

What it is: A method for identifying the genomic locations where specific proteins, such as transcription factors or histone modifications, bind to DNA.

How it works: Proteins are cross-linked to DNA in living cells, the chromatin is fragmented, and an antibody specific to the protein of interest is used to immunoprecipitate the protein-DNA complexes. The associated DNA is then purified and sequenced.

Typical applications: Mapping transcription factor binding sites, characterising histone modification patterns (e.g., H3K4me3 for active promoters, H3K27me3 for repressed regions), and studying gene regulation.

Key outputs: Genome-wide maps of protein-DNA interactions, peak files.

Limitations: Requires a high-quality, specific antibody for the protein of interest. Results can vary substantially between antibodies and cell types.


Hi-C (Chromosome Conformation Capture)

What it is: A method for mapping the three-dimensional organisation of the genome, revealing which genomic regions are physically close to each other in the nucleus even if they are far apart in linear sequence.

How it works: Chromatin is cross-linked in intact cells, digested with a restriction enzyme, and the cut ends are ligated together. Only ends that are physically proximate in the nucleus get ligated. The resulting DNA junctions are sequenced, and the frequency of contact between any two genomic loci is used to reconstruct the spatial organisation of chromosomes.

Typical applications: Identifying topologically associating domains (TADs), characterising chromatin loops, genome assembly scaffolding using chromosome-scale contact information, and studying the relationship between 3D genome organisation and gene regulation.

Key outputs: Contact matrices, TAD boundaries, loop anchors.

Limitations: Computationally intensive. Requires deep sequencing for high-resolution maps. Very sensitive to cross-linking efficiency and cell quality.


Long-Read Sequencing (PacBio and Oxford Nanopore)

What it is: Sequencing technologies capable of producing reads tens of thousands of base pairs long, in contrast to the 150 to 300 bp reads produced by standard Illumina sequencing.

How it works: PacBio (SMRT Sequencing) observes the incorporation of fluorescently labelled nucleotides by a single DNA polymerase in real time, producing reads averaging 10 to 25 kb. Oxford Nanopore Technology (ONT) threads a DNA or RNA strand through a protein nanopore and measures changes in ionic current as each base passes through, producing reads that can exceed 1 Mb.

Typical applications: De novo genome assembly, resolving repetitive regions, phasing haplotypes, structural variant detection, direct RNA sequencing (Nanopore), and base modification detection without bisulfite treatment (Nanopore).

Key outputs: Long-read assemblies, phased haplotype sequences, structural variant calls.

Limitations: Higher error rates per base compared to Illumina (though PacBio HiFi reads now approach Illumina accuracy). Higher cost per base. Oxford Nanopore requires careful library preparation to maximise read length.


MPS (Massively Parallel Sequencing)

What it is: The umbrella term for all modern high-throughput sequencing approaches, also referred to as Next-Generation Sequencing (NGS). Rather than sequencing one molecule at a time, MPS platforms sequence millions to billions of DNA fragments simultaneously.

How it works: In the most widely used implementation (Illumina), DNA fragments are attached to a flow cell surface and amplified into clusters. Each cluster is sequenced by synthesis, where fluorescently labelled nucleotides are incorporated one at a time and imaged. The resulting colour signals are decoded into a sequence.

Typical applications: Essentially any genomic, transcriptomic, or epigenomic application. MPS is the foundation on which all other methods in this list are built.

Key outputs: FASTQ files containing raw reads with quality scores.

Limitations: Standard Illumina reads are short (150 to 300 bp), which limits the ability to resolve repetitive regions and phase variants over long distances.


Metagenomics (Shotgun Metagenomics)

What it is: Untargeted sequencing of all DNA present in an environmental or clinical sample, without prior amplification of any specific target region.

How it works: Total DNA is extracted directly from the sample, sheared into fragments, and sequenced. The resulting reads can be assembled into contigs or mapped against reference databases to identify the organisms and genes present.

Typical applications: Characterising the taxonomic composition and functional potential of microbial communities in soil, water, gut, and other environments. Unlike 16S amplicon sequencing, metagenomics provides functional as well as taxonomic information.

Key outputs: Taxonomic profiles, assembled metagenome-assembled genomes (MAGs), gene catalogues.

Limitations: Requires substantially more sequencing depth than amplicon approaches to detect rare community members. Host DNA contamination can dominate the signal in host-associated samples. Computationally demanding.


Metatranscriptomics

What it is: Sequencing of all RNA present in an environmental or clinical sample to capture the active gene expression of an entire microbial community at a given moment.

How it works: Total RNA is extracted from the sample, ribosomal RNA is depleted (rRNA makes up the vast majority of RNA in a cell), and the remaining messenger RNA is converted to cDNA and sequenced. The resulting reads reveal which genes are being expressed and at what level across all organisms in the community.

Typical applications: Studying the functional activity of microbial communities in response to environmental conditions, identifying metabolic pathways that are actively expressed, and comparing functional activity across conditions or time points.

Key outputs: Expression profiles of community-wide genes, active metabolic pathway maps.

Limitations: RNA is highly unstable and requires careful and rapid sample preservation. Ribosomal RNA depletion is imperfect, and residual rRNA reads can dominate the dataset. Linking transcripts to specific organisms is computationally challenging.


RADseq (Restriction-site Associated DNA Sequencing)

What it is: A reduced-representation sequencing method that sequences only the genomic regions flanking specific restriction enzyme cut sites, enabling cost-effective genotyping of large numbers of individuals across thousands of loci.

How it works: Genomic DNA is digested with one or two restriction enzymes, adapters are ligated to the cut ends, and only the fragments adjacent to cut sites are selected and sequenced. Because the same restriction sites are cut reproducibly across individuals, the same loci are sequenced in every sample.

Typical applications: Population genetics, phylogeography, linkage mapping, and studies of local adaptation, particularly in non-model organisms where no reference genome exists.

Key outputs: Genotype matrices with thousands of SNPs across many individuals.

Limitations: Loci are lost if a restriction site is disrupted by a mutation (allelic dropout). The number of loci generated depends on the choice of restriction enzyme and genome size. Not well suited to very large or highly repetitive genomes.


RNA-Seq (RNA Sequencing)

What it is: Sequencing of the transcriptome, the complete set of RNA molecules present in a cell or tissue at a given time, to measure gene expression levels and discover novel transcripts.

How it works: RNA is extracted, ribosomal RNA is typically depleted or poly-A selected to enrich for messenger RNA, and the RNA is reverse-transcribed into cDNA. The cDNA library is then sequenced. Read counts mapping to each gene serve as a proxy for expression level.

Typical applications: Differential gene expression analysis between conditions, transcript discovery and annotation, alternative splicing analysis, and allele-specific expression.

Key outputs: Gene expression count matrices, lists of differentially expressed genes, transcript assemblies.

Limitations: Captures a snapshot of expression at the moment of RNA extraction. The poly-A selection step enriches for mature mRNA but misses non-polyadenylated transcripts. Strand information can be lost unless strand-specific library preparation is used.


Sanger Sequencing

What it is: The original DNA sequencing method developed by Frederick Sanger in 1977, and the dominant sequencing technology for nearly three decades before the advent of massively parallel sequencing.

How it works: A DNA template is replicated in the presence of chain-terminating dideoxynucleotides (ddNTPs), each labelled with a different fluorescent dye. Replication stops randomly wherever a ddNTP is incorporated, generating a population of fragments of all possible lengths. These fragments are separated by capillary electrophoresis and the fluorescent signal at each position is read to determine the sequence.

Typical applications: Sequencing of individual genes or short target regions, verification of cloning results, confirmation of variants identified by MPS, and clinical diagnostics where a small number of known mutations need to be checked.

Key outputs: A chromatogram showing fluorescent peaks corresponding to each base, typically covering 600 to 1000 bp per read.

Limitations: Low throughput — only one fragment is sequenced per reaction. Not suitable for whole-genome or transcriptome-scale projects. Cost per base is much higher than MPS at scale, though for small numbers of targets it remains competitive and highly accurate.


scRNA-Seq (Single-Cell RNA Sequencing)

What it is: RNA sequencing performed at the resolution of individual cells, allowing gene expression to be measured cell by cell rather than as an average across a bulk sample.

How it works: Individual cells are isolated, typically using microfluidics or droplet-based systems such as 10x Genomics Chromium. Each cell is tagged with a unique barcode, RNA is captured and reverse-transcribed, and the resulting cDNA library is sequenced. Bioinformatic analysis then separates reads by barcode to reconstruct the expression profile of each cell.

Typical applications: Identifying and characterising cell types within a tissue, studying developmental trajectories, discovering rare cell populations, and analysing cellular heterogeneity in disease.

Key outputs: Cell by gene expression matrices, UMAP or t-SNE visualisations of cell clusters, cell type annotations.

Limitations: Captures only a fraction of each cell's transcriptome (dropout). High cost relative to bulk RNA-Seq. Requires specialised equipment and significant computational resources for analysis.


Spatial Transcriptomics

What it is: A method that combines RNA sequencing with spatial information, allowing gene expression to be measured while preserving the physical location of cells within a tissue section.

How it works: A tissue section is placed on a slide containing spatially barcoded capture probes. RNA from the tissue diffuses into the slide and is captured at defined coordinates. The captured RNA is sequenced, and expression data is mapped back to the spatial coordinates of the tissue.

Typical applications: Studying gene expression in the context of tissue architecture, identifying spatially restricted cell populations, and understanding how cell signalling and communication depend on tissue organisation.

Key outputs: Spatially resolved expression maps overlaid on tissue images.

Limitations: Resolution is improving rapidly but currently does not always reach single-cell resolution. Tissue handling and RNA preservation are critical.


WES (Whole Exome Sequencing)

What it is: Sequencing of only the protein-coding regions of the genome (the exome), which represent approximately 1 to 2% of the total genome but contain the majority of disease-causing variants.

How it works: Genomic DNA is sheared and the exonic regions are captured using hybridisation probes that selectively bind to exon sequences. The captured fragments are then sequenced at high depth.

Typical applications: Clinical genetics, rare disease diagnosis, cancer genomics, and population-scale variant discovery focused on coding regions.

Key outputs: Variant call files (VCF) containing SNPs and small indels in coding regions.

Limitations: Misses regulatory and non-coding variants, which can be biologically important. Capture efficiency is uneven across exons, leading to variable coverage.


WGS (Whole Genome Sequencing)

What it is: Sequencing of the complete genome of an organism, providing a comprehensive view of all genetic variation without any prior selection of target regions.

How it works: Genomic DNA is sheared into fragments of a defined size, adapters are ligated, and the library is sequenced. Reads are aligned to a reference genome and variants are called across the entire sequence.

Typical applications: De novo genome assembly, comprehensive variant detection including SNPs, indels, copy number variants and structural variants, population genomics, and evolutionary studies.

Key outputs: Genome assemblies or aligned BAM files, comprehensive variant call sets.

Limitations: The most expensive per-sample approach in terms of sequencing depth required. Analysis of structural variants and repetitive regions remains challenging with short reads alone.


This list is not exhaustive

New sequencing methods and variations appear regularly. If you encounter a method not listed here, the original publication describing it is usually the best starting point. Review articles in journals such as Nature Methods and Nature Reviews Genetics are also excellent resources for staying current.