Data Download
A checklist and terminal commands to sanity-check your raw files the moment they arrive.
Once your MiSeq, i100, PacBio, or AVITI run is complete, you need to obtain your sequences. At the GDC, your data set is uploaded to our server, where it is either processed for you or made available for download.
Regardless of where you get your data from, download it as soon as it is available and check the quality of the run before you do anything else. Catching a problem now, while a rerun or a quick fix is still possible, is far cheaper than catching it after weeks of downstream analysis. Here are the checks we go through, and why each one matters.
Check-List
- Download data via terminal (e.g. sftp), not a browser, whenever possible
- Verify file integrity (e.g. md5sum)
- Verify file counts: N_{Files} = 2 \times N_{Samples} for paired-end R1/R2 data
- Check FASTQ headers: how many runs, how many lanes?
- Get a few random reads and BLAST them
- Read the basecall and/or QC report(s)
- Check the read length distribution
- Check for possible contamination (e.g. PhiX)
- Archive a copy of your raw data
- Upload the raw data to a public archive (e.g. ENA)
1. Download via Terminal
We recommend using the terminal (e.g. sftp, scp, or wget) to download your sequences whenever possible, rather than a browser. It is usually faster, more reliable for large files, and much easier to script for batch downloads across many samples.
2. Verify File Integrity with md5sum
Use the md5sum command to check the integrity of files after transfer. Any change to a file, even a single corrupted byte, changes its MD5 hash, so comparing hashes before and after transfer catches a corrupted download that a file size check alone might miss.
# Generate md5sum checksums for all R1/R2 files in the current directory
find . -name "*_R[12]*.fastq.gz" -exec md5sum {} \; > md5sums.txt
Naming Your Checksum File
Give the output file a name that identifies the project and run, for example p1197_run260117_16S_md5sums.txt, rather than something generic. You will likely accumulate many of these over time, and a clear name saves you from opening files just to figure out what they belong to.
If we provide you with a checksum file alongside your data, verify against it directly instead of just generating your own:
# Verify files against a checksum file we provided
md5sum -c p1197_run260117_16S_md5sums.txt
A Quick Structural Check, No Reference Needed
A checksum only tells you a file matches what the sender computed, which is no help if the file was already truncated before the checksum was generated, or if you have no reference checksum to compare against at all. FASTQ has a fixed structure: every read is exactly 4 lines (header, sequence, + separator, quality). A file that isn't a multiple of 4 lines long is truncated or otherwise corrupted, no reference file required to know that:
for f in a_data/gz/*.f*q.gz; do
n=$(zcat "$f" | wc -l)
if (( n % 4 != 0 )); then
echo "$f: $n lines, not a multiple of 4, likely truncated"
fi
done
This only catches missing lines at the end, not corruption that preserves the line count. For a stricter check, confirm the record structure holds throughout the file. Don't anchor this on header lines starting with @, that character is also a valid Phred+33 quality score, so a quality line can legitimately start with @ too, which makes it a weak signal. The separator line is much safer: on line 3 of every record, raw read data has essentially nothing but a bare + and nothing else, so matching the full line, not just its first character, is a near-guaranteed check:
for f in a_data/gz/*.f*q.gz; do
zcat "$f" | awk -v f="$f" \
'NR%4==3 && $0!="+" {print f": bad separator at line "NR; bad=1}
END {if (!bad) print f": OK"}'
done
3. Verify File Counts
Before going further, confirm you actually received the number of files you expect. For standard paired-end Illumina or AVITI data with only R1 and R2 exported, that means:
N_{Files} = 2 \times N_{Samples}
If index reads were also exported (see the Illumina section of the workflow overview for when that happens), expect 4 files per sample instead.
One Small Adjustment for GDC-Delivered Data
If your data comes from us, we also include the Undetermined_R1/Undetermined_R2 files (see Step 8 for why they're worth keeping), so your actual file count will be N_{Files} = (2 \times N_{Samples}) + 2. Not a problem, just don't be thrown off if your count is two files higher than the formula above suggests.
# Count R1 and R2 files separately, they should match
ls a_data/gz/*_R1*.f*q.gz | wc -l
ls a_data/gz/*_R2*.f*q.gz | wc -l
# n(R1) should equal n(R2), and both should equal your number of samples
# (plus one each, for Undetermined_R1/R2, on GDC-delivered data)
If the counts don't match your sample sheet, stop here and sort it out before proceeding. A missing file is much easier to chase down now than after you have already started processing.
4. Check FASTQ Headers
The first line of every FASTQ read is a header that tells you which instrument, run, and lane the read came from. For a standard Illumina header this looks like:
@M01761:234:000000000-B32NW:1:2107:10522:1813 2:Y:0:CCTAAGAC+TAGCCTTA
Here M01761 is the instrument ID and 234 is the run number.
# Print the first header of every file, one per file
for f in a_data/gz/*.f*q.gz; do zcat "$f" | head -n 1; done
A Common Mistake
It's tempting to pipe everything into zcat at once and pull the first line: zcat a_data/gz/*.f*q.gz | head -n 1. Don't. Because head stops reading as soon as it has one line, this command only ever shows you the header of the very first file in the glob, not one header per file. Use the loop above instead if you want to check every file.
No Index Files on the MiSeq i100
If you're expecting I1/I2 index read files alongside R1/R2, check which instrument generated your data first. The older MiSeq can export them; the MiSeq i100 currently can't, see the Illumina section of the Overview Workflows page for why. Don't spend time chasing files that were never going to be there.
Once you have headers from all files, check whether your data actually comes from a single run:
# List the unique instrument:run combinations across all files
zgrep -h " " a_data/gz/*.f*q.gz | cut -d : -f 1,2 | sort -u
Don't Forget -h
zgrep prepends the filename to every match when it searches more than one file. Without the -h flag, that filename becomes an extra field before the header itself, and cut -d: -f1,2 will grab the wrong columns. Always add -h when running this on multiple files.
A Faster Shortcut for Large Files
The command above scans every single header in every file, which gets slow once your files are large or numerous. Since files from different runs are normally mixed by concatenation, one run's reads appended after another's, rather than interleaved read by read, checking just the first and last read of each file is enough to catch the same problem:
for f in a_data/gz/*.f*q.gz; do
zcat "$f" | awk 'NR%4==1 {last=$0; if (NR==1) first=$0} END {print first; print last}'
done | cut -d : -f 1,2 | sort -u
This still decompresses each file once, that part is unavoidable, but it avoids matching and cutting every header line, and feeds only two lines per file into sort. If the first and last read of a file report different instrument:run combinations, you know that file mixes runs, without needing to check what's in between. This assumes concatenation rather than interleaving; if you have reason to suspect reads were shuffled between runs rather than appended, fall back to the full scan above.
If either check returns more than one instrument:run combination, your files span more than one sequencing run, which is worth knowing before you pool them for analysis. It can be entirely expected (e.g. large projects sequenced in batches) or a sign that files from different projects got mixed into the same folder by mistake.
5. Get a Few Random Reads and BLAST Them
This is the fastest sanity check available, and the one most often skipped. Before investing time in a full analysis, pull out a handful of reads and BLAST them against NCBI's nt database.
# Extract 10 random reads from an R1 file as FASTA, no external tools needed
zcat sample_R1.fastq.gz | paste - - - - | shuf -n 10 | \
awk -F'\t' '{print ">"substr($1,2)"\n"$2}' > random10.fasta
Paste the sequences into the NCBI BLAST web interface (blastn, against nt), or run them locally if you have a BLAST database set up.
What You're Looking For
You want to see hits broadly consistent with your expected target group and marker gene, for example bacterial 16S hits for a soil microbiome 16S project. A handful of unexpected top hits isn't necessarily alarming, but if most or all of your random reads come back as something completely different from what you expected (wrong organism group entirely, vector or adapter sequence, or another marker gene altogether), that's worth chasing down immediately. It can point to a mislabeled sample, a mixed-up index, or a primer issue, all things you would much rather catch now than after clustering and taxonomic assignment.
6. Read the Basecall and/or QC Report(s)
QC Report. Your data is usually accompanied by a quality report. FastQC in combination with MultiQC is our preferred choice, and both are convenient and quick to run. Both also come with detailed manuals, worth reading carefully so you understand their limitations.
FastQC Flags Are Not All Relevant to Amplicon Data
FastQC was designed with genome and transcriptome sequencing in mind, and several of its default red flags are misleading for amplicon projects. High duplication levels are the classic example: FastQC will often flag this as a failure, but in an amplicon run you are deliberately sequencing the same short target region across many templates and many samples, so a high duplication rate is expected and not a problem on its own. The same goes for some of its per-base sequence content and GC content warnings, which assume a diverse genomic library rather than a single, narrow amplicon. It takes a bit of practice to read a FastQC report correctly for amplicon data, learning which red and orange flags actually matter for your project and which are just noise from a tool tuned for a different use case. If a report looks alarming and you're not sure whether it's a real problem, send it our way rather than guessing.
Adapter Content Is Worth Checking Directly, Unlike Duplication
One FastQC flag that genuinely matters for amplicon data, unlike duplication levels, is Adapter Content. If your expected amplicon is shorter than your read length (a staggered-read situation, see Step C of the PE workflow for when this comes up), the sequencer reads straight through the end of your insert and into the adapter sequence attached during library prep. FastQC's Adapter Content module picks this up directly, and a rising adapter-content curve towards the end of your reads is a real, actionable signal, not noise from a tool tuned for a different use case.
This is a more common and easily overlooked problem than it sounds. The GDC's MPS course page includes a challenge showing adapter contamination turning up even in reference sequences submitted to NCBI, so it's worth actually checking for it yourself rather than assuming it isn't there, with FastQC, or with a simple direct search for known adapter sequences in your reads.
Beyond cleanup, checking for it also tells you something diagnostic: where the adapter read-through consistently starts is where your actual amplicon ends. If that position doesn't match what you expected, that's a direct, independent signal about your true amplicon size, worth cross-checking against the read length distribution above. It gets trimmed away during primer and adapter trimming later in processing (see the Cutadapt reference on that page), but worth knowing it's there, why, and what it's telling you, rather than being surprised by it.
Basecaller Report. This is particularly relevant for PacBio data. The report is not easy to parse at first glance, but with a bit of experience, or a bit of help from us, it tells you a lot about run quality before you commit to downstream processing: number of passes per CCS read, read quality distributions, and adapter or barcode statistics among them.
7. Check the Read Length Distribution
# Read-length distribution for one file, single pass, no sort needed
zcat a_data/gz/sample_R1.f*q.gz | \
awk 'NR % 4 == 2 {c[length($0)]++} END {for (l in c) print c[l], l}' | sort -k2,2n
Building the histogram directly in awk and only sorting the handful of resulting length categories, rather than piping every single read length through sort -n, matters once you have serious read numbers. Sorting hundreds of millions of length values, which is exactly what you get on a full AVITI run, is slow and mostly pointless.
For Fixed-Length Platforms, Sample Instead
On Illumina and AVITI, every read is sequenced for the same fixed number of cycles, so you already know roughly what length to expect, and you're really just confirming there's no widespread truncation or adapter read-through, not discovering a real distribution. There's little point scanning every read on a 600-plus-million-read AVITI file for that. Check a sample instead:
zcat a_data/gz/sample_R1.f*q.gz | head -n 40000 | \
awk 'NR % 4 == 2 {c[length($0)]++} END {for (l in c) print c[l], l}' | sort -k2,2n
That's 10,000 reads, plenty to catch a systematic problem. Save the full, exhaustive scan for PacBio CCS data, where the read length distribution reflects real biological variation in amplicon length and is worth looking at in full.
For Illumina or AVITI data, you should see a sharp peak at (or very close to) the cycle number you sequenced. A wide spread, an unexpected second peak, or a peak noticeably shorter than expected can indicate adapter read-through, poor cluster quality late in the run, or a library insert size problem. For PacBio CCS data, expect a broader distribution centred on your amplicon length, since read length there reflects the biology of your amplicon rather than a fixed cycle count.
Coding vs. Non-Coding Markers Look Different Here Too
The shape of that biological distribution itself depends on your marker. Protein-coding loci like COI tend to show a much sharper length peak than non-coding markers like 16S or ITS, since insertions or deletions that break the reading frame are strongly selected against in real sequences (see Data Prep Short-Reads (PE), Step G for the ORF-based filtering this makes possible, applied after clustering). A broad, fuzzy peak is normal for 16S or ITS; the same shape on a COI amplicon is worth a second look.
8. Check for Possible Contamination (e.g. PhiX)
PhiX is a control library Illumina spikes into runs to add base diversity, particularly important for low-diversity libraries like amplicons (see the AVITI section of the workflow overview for how newer chemistries reduce this need). PhiX reads don't carry your sample's index, so they normally end up in the Undetermined bin during demultiplexing rather than in your per-sample files. If PhiX-derived sequences turn up in your actual sample files instead, that points to index misassignment or cross-talk on the flow cell.
# Count undetermined (non-demultiplexed) reads
zgrep -c "^+$" a_data/gz/Undetermined*_R[12]*.f*q.gz
Watch the File Pattern
Match the Undetermined files with the same wildcard pattern you use elsewhere, _R[12]*.f*q.gz, not _R[12].f*q.gz. The version without the trailing * misses the _001 suffix that real Illumina filenames like Undetermined_S0_L001_R1_001.fastq.gz have, and will silently match nothing.
A high proportion of undetermined reads relative to your total is worth a closer look. To confirm a suspected PhiX signal directly, align a sample of your reads against the PhiX genome (NCBI accession NC_001422.1) with a fast aligner such as minimap2 or bowtie2. A cluster of clean, high-identity hits confirms PhiX contamination rather than a coincidental partial match.
Ask for the Undetermined Reads Themselves, Not Just the Count
We, and many other sequencing centres, will provide you with the actual Undetermined FASTQ files alongside your demultiplexed data, not just a read count. Keep them. Beyond confirming PhiX, they're a genuinely useful troubleshooting resource if a specific sample comes back with far fewer reads than expected: those reads may not be missing at all, they may be sitting in the Undetermined bin under a barcode combination that doesn't match your sample sheet, from a typo, a swapped index pair, or a mismatch between what was ordered and what was recorded. The same random-read BLAST approach from Step 5 works just as well applied to a handful of Undetermined reads, and can sometimes explain exactly where a sample's missing reads actually went.
FastQ Screen for a Broader Check
FastQ Screen does this kind of check more systematically. Rather than testing against one reference at a time, it screens a subset of your reads against a whole panel of reference genomes in one go, commonly including PhiX, common vectors, and whichever organisms are relevant to your work, and reports what fraction of reads match each. It's a good complement to the targeted PhiX alignment above, especially if you suspect contamination from something other than PhiX, or just want a routine first-pass check without deciding in advance what to look for.
9. Archive a Copy of Your Raw Data
Keep an untouched copy of your raw files somewhere separate from your working analysis directory, and never overwrite or edit that copy directly, always copy before you process. Store it in at least two places (for example, institutional storage plus an external backup), and keep your checksum file alongside it so you can verify integrity again later if you ever need to. Raw sequencing data is expensive and, for many sample types, effectively impossible to regenerate, so treat it accordingly.
10. Upload the Raw Data to a Public Archive (e.g. ENA)
Many journals and funders now require raw sequencing reads to be publicly archived alongside a publication. The European Nucleotide Archive (ENA) is a common choice; NCBI's Sequence Read Archive (SRA) is the equivalent in the US. Submission requires registering a study (BioProject) and sample metadata (BioSample) before you can upload reads, and for larger batches, ENA's Webin-CLI tool is much less painful than uploading through the browser.
Talk to Us Early
Setting up a submission correctly the first time, with consistent sample metadata that matches your manuscript, saves a lot of back-and-forth later. If you're planning to publish, get in touch with us before you submit so we can help make sure the metadata is structured correctly from the start.