Skip to content

UNOISE3: From sequencing errors to ZOTUs

UNOISE3 is a denoising method in USEARCH that tries to distinguish genuine biological sequence variants from sequencing and PCR errors.

The basic idea is:

A rare sequence that is very similar to a much more abundant sequence may be an error rather than a genuine biological variant.

UNOISE3 uses sequence similarity and abundance to make this decision.


From reads to ZOTUs

After quality filtering, identical reads are usually dereplicated. Instead of storing every copy separately, we keep each unique sequence together with its abundance.

For example:

A  ACGTACGTACGT    20,000 reads
B  ACGTACGTTCGT         5 reads
C  TGCATGCTAGCA       800 reads

UNOISE3 then considers whether rare sequences such as B could simply be erroneous versions of abundant sequences such as A.

flowchart TD
    A[Sequencing reads] --> B[Quality filtering]
    B --> C[Dereplication]
    C --> D["Unique sequences + abundance"]
    D --> E["minsize"]
    E --> F[UNOISE3]
    F --> G[Error-like sequences removed]
    F --> H[Retained sequence variants]
    H --> I[Chimera removal]
    I --> J[ZOTUs]

A ZOTU is therefore an inferred sequence variant after denoising. It is not automatically a species.


minsize: which sequences are considered?

minsize is an abundance filter applied before UNOISE3 evaluates the sequences.

For example:

-minsize 8

means that a unique sequence must occur at least 8 times to enter the denoising step.

Consider:

A  ACGTACGTACGT    20,000
B  ACGTACGTTCGT         5

With minsize = 8:

A  → enters UNOISE3
B  → excluded

With minsize = 4:

A  → enters UNOISE3
B  → enters UNOISE3

But letting B enter UNOISE3 does not mean that B becomes a ZOTU. UNOISE3 still has to decide whether B looks like a genuine variant or an error.

Think of minsize as the entrance gate:

"Is this sequence abundant enough to be considered?"


unoise_alpha: when is a rare sequence considered an error?

Now consider two sequences differing by only one nucleotide:

A  ACGTACGTACGT    20,000 reads
B  ACGTACGTTCGT         B reads

UNOISE3 compares their abundances.

For this simplified one-nucleotide example, the error criterion can be expressed as:

\frac{B}{A} \leq \frac{1}{2^{\alpha+1}}

where alpha is the unoise_alpha parameter.

The default is:

unoise_alpha = 2

The larger alpha becomes, the easier it is for a rare sequence to escape this error criterion.


A concrete example

Let's keep A fixed at 20,000 reads and vary the abundance of B.

B reads B/A α = 1 α = 2 α = 3 α = 4
5 0.00025 error error error error
50 0.0025 error error error error
500 0.025 error error error potential ZOTU
5,000 0.25 potential ZOTU potential ZOTU potential ZOTU potential ZOTU

For the default alpha = 2, the threshold is:

\frac{1}{2^{2+1}} = \frac{1}{8} = 0.125

With A = 20,000:

20,000 \times 0.125 = 2,500

So, in this simplified example, B needs to reach roughly 2,500 reads before it escapes the abundance criterion for being treated as an error.

The effect of alpha

alpha Error threshold B/A Equivalent threshold if A = 20,000
1 0.25 5,000
2 0.125 2,500
3 0.0625 1,250
4 0.03125 625

Thus:

lower alpha
    ↓
more stringent error interpretation
    ↓
fewer rare variants retained


higher alpha
    ↓
less stringent error interpretation
    ↓
more rare variants retained

Why this matters

Imagine that B is not an error. It is actually a rare biological variant:

A  common variant      20,000 reads
B  genuine rare variant   500 reads

The two sequences differ by only one nucleotide.

With minsize = 8, B is allowed into UNOISE3. But with the default alpha = 2, B has only:

500 / 20,000 = 0.025

or 2.5% of A's abundance.

That is below the simplified UNOISE3 threshold of 12.5%, so B may be interpreted as an error and not retained as a separate ZOTU.

This illustrates an important principle:

Denoising cannot know whether a rare sequence is biologically real. It infers this from the pattern of abundance and sequence similarity.


minsize and alpha do different things

The two parameters should not be confused:

Parameter Question
minsize Should this sequence enter the denoising analysis at all?
unoise_alpha Given that it enters, how readily can it be distinguished from an abundant, similar sequence?

Conceptually:

Unique sequences
       │
       ▼
   minsize
       │
       │  Too rare?
       ├──────────────► excluded
       │
       ▼
   UNOISE3
       │
       │  Error-like?
       ├──────────────► removed
       │
       ▼
  retained sequence
       │
       ▼
     ZOTU

ZOTUs are not species

UNOISE3 produces sequence variants, not species identifications.

For example:

ZOTU1 ──┐
        ├── Species A
ZOTU2 ──┘

ZOTU3 ───── Species B

Several ZOTUs can therefore belong to the same species, while a ZOTU can also remain taxonomically unidentified.

Taxonomic assignment is a separate step after denoising.


The main take-home message

UNOISE3 uses abundance and sequence similarity to identify sequences that are likely to be errors.

  • minsize determines which low-abundance sequences are considered.
  • unoise_alpha influences how readily rare sequences are interpreted as errors.
  • Higher alpha generally allows more rare sequence variants to survive denoising.
  • A ZOTU is an inferred sequence variant, not necessarily a species.
  • A genuinely rare biological variant can be mistaken for an error if it closely resembles a much more abundant sequence.

A simplified model

The abundance calculations above illustrate the core effect of unoise_alpha for two sequences differing by one nucleotide. The complete UNOISE3 algorithm considers the abundance ordering and relationships among all unique sequences, so the table should be understood as a conceptual example rather than a complete prediction of every ZOTU assignment.