UNOISE3: From sequencing errors to ZOTUs
UNOISE3 is a denoising method in USEARCH that tries to distinguish genuine biological sequence variants from sequencing and PCR errors.
The basic idea is:
A rare sequence that is very similar to a much more abundant sequence may be an error rather than a genuine biological variant.
UNOISE3 uses sequence similarity and abundance to make this decision.
From reads to ZOTUs
After quality filtering, identical reads are usually dereplicated. Instead of storing every copy separately, we keep each unique sequence together with its abundance.
For example:
A ACGTACGTACGT 20,000 reads
B ACGTACGTTCGT 5 reads
C TGCATGCTAGCA 800 reads
UNOISE3 then considers whether rare sequences such as B could simply be erroneous versions of abundant sequences such as A.
flowchart TD
A[Sequencing reads] --> B[Quality filtering]
B --> C[Dereplication]
C --> D["Unique sequences + abundance"]
D --> E["minsize"]
E --> F[UNOISE3]
F --> G[Error-like sequences removed]
F --> H[Retained sequence variants]
H --> I[Chimera removal]
I --> J[ZOTUs]
A ZOTU is therefore an inferred sequence variant after denoising. It is not automatically a species.
minsize: which sequences are considered?
minsize is an abundance filter applied before UNOISE3 evaluates the sequences.
For example:
-minsize 8
means that a unique sequence must occur at least 8 times to enter the denoising step.
Consider:
A ACGTACGTACGT 20,000
B ACGTACGTTCGT 5
With minsize = 8:
A → enters UNOISE3
B → excluded
With minsize = 4:
A → enters UNOISE3
B → enters UNOISE3
But letting B enter UNOISE3 does not mean that B becomes a ZOTU. UNOISE3 still has to decide whether B looks like a genuine variant or an error.
Think of minsize as the entrance gate:
"Is this sequence abundant enough to be considered?"
unoise_alpha: when is a rare sequence considered an error?
Now consider two sequences differing by only one nucleotide:
A ACGTACGTACGT 20,000 reads
B ACGTACGTTCGT B reads
UNOISE3 compares their abundances.
For this simplified one-nucleotide example, the error criterion can be expressed as:
where alpha is the unoise_alpha parameter.
The default is:
unoise_alpha = 2
The larger alpha becomes, the easier it is for a rare sequence to escape this error criterion.
A concrete example
Let's keep A fixed at 20,000 reads and vary the abundance of B.
| B reads | B/A | α = 1 | α = 2 | α = 3 | α = 4 |
|---|---|---|---|---|---|
| 5 | 0.00025 | error | error | error | error |
| 50 | 0.0025 | error | error | error | error |
| 500 | 0.025 | error | error | error | potential ZOTU |
| 5,000 | 0.25 | potential ZOTU | potential ZOTU | potential ZOTU | potential ZOTU |
For the default alpha = 2, the threshold is:
With A = 20,000:
So, in this simplified example, B needs to reach roughly 2,500 reads before it escapes the abundance criterion for being treated as an error.
The effect of alpha
alpha |
Error threshold B/A | Equivalent threshold if A = 20,000 |
|---|---|---|
| 1 | 0.25 | 5,000 |
| 2 | 0.125 | 2,500 |
| 3 | 0.0625 | 1,250 |
| 4 | 0.03125 | 625 |
Thus:
lower alpha
↓
more stringent error interpretation
↓
fewer rare variants retained
higher alpha
↓
less stringent error interpretation
↓
more rare variants retained
Why this matters
Imagine that B is not an error. It is actually a rare biological variant:
A common variant 20,000 reads
B genuine rare variant 500 reads
The two sequences differ by only one nucleotide.
With minsize = 8, B is allowed into UNOISE3. But with the default alpha = 2, B has only:
or 2.5% of A's abundance.
That is below the simplified UNOISE3 threshold of 12.5%, so B may be interpreted as an error and not retained as a separate ZOTU.
This illustrates an important principle:
Denoising cannot know whether a rare sequence is biologically real. It infers this from the pattern of abundance and sequence similarity.
minsize and alpha do different things
The two parameters should not be confused:
| Parameter | Question |
|---|---|
minsize |
Should this sequence enter the denoising analysis at all? |
unoise_alpha |
Given that it enters, how readily can it be distinguished from an abundant, similar sequence? |
Conceptually:
Unique sequences
│
▼
minsize
│
│ Too rare?
├──────────────► excluded
│
▼
UNOISE3
│
│ Error-like?
├──────────────► removed
│
▼
retained sequence
│
▼
ZOTU
ZOTUs are not species
UNOISE3 produces sequence variants, not species identifications.
For example:
ZOTU1 ──┐
├── Species A
ZOTU2 ──┘
ZOTU3 ───── Species B
Several ZOTUs can therefore belong to the same species, while a ZOTU can also remain taxonomically unidentified.
Taxonomic assignment is a separate step after denoising.
The main take-home message
UNOISE3 uses abundance and sequence similarity to identify sequences that are likely to be errors.
minsizedetermines which low-abundance sequences are considered.unoise_alphainfluences how readily rare sequences are interpreted as errors.- Higher
alphagenerally allows more rare sequence variants to survive denoising. - A ZOTU is an inferred sequence variant, not necessarily a species.
- A genuinely rare biological variant can be mistaken for an error if it closely resembles a much more abundant sequence.
A simplified model
The abundance calculations above illustrate the core effect of unoise_alpha for two sequences differing by one nucleotide. The complete UNOISE3 algorithm considers the abundance ordering and relationships among all unique sequences, so the table should be understood as a conceptual example rather than a complete prediction of every ZOTU assignment.