selection guide
Quality scores and a FASTQ file
Read the four lines of a FASTQ record, treat Q20 and Q30 as planning thresholds, and expect quality to fall toward the end of a read.
- Author
- EVRINTH Editorial Team
- Published
- 8 October 2026
- Updated
- 8 October 2026
- Reading time
- 10 min

A FASTQ file is the usual way a sequencing run hands you bases and the caller's confidence in each of them. Selecting what to trust in that file is a different skill from admiring a high percentage on a summary plot. This guide is the layout of the record, the meaning of the quality line, and the decisions those characters do and do not support. How the file comes out of a library is next-generation sequencing from library to reads. Why a confident base can still be the wrong observation is coverage depth is not the same as accuracy.
You will meet FASTQ from short-read instruments most often, and also from long-read workflows that choose to emit it. The encoding story below is the one modern short-read files use. Confirm it before you interpret a file from an older run.
Four lines, and only one of them is sequence
A record is four lines. Line one begins with @ and a read identifier. The identifier may contain the instrument, the run, the flow cell, the cluster coordinate, and, depending on the software, a hint about whether the read passed a chastity or purity filter. Do not parse it casually. Colon-separated fields have changed between software versions. Treat the whole string as an identifier unless you have the manual for the version that wrote it.
Line two is the sequence, in the letters the basecaller chose. It may include N where the caller declined to choose. An N is a decision, not a missing byte.
Line three begins with + . Some writers repeat the identifier after the plus. Many write a bare plus. Both are legal. A line that does not start with + means the record is broken, or you have slipped by a line.
Line four is quality, and it must be the same length as line two. Each character encodes a Phred score for the base above it. In the encoding used by current Illumina-style files, and by Sanger FASTQ before that, the score is the character's ASCII code minus 33. This is called Phred+33. The character whose code is 63, a question mark, is Q30. The digit 5 is Q20. A letter I is Q40. People who read the quality line as sequence will find phantom bases and a nonsense alignment. It is a row of scores wearing printable characters.
Older Illumina pipelines used a Phred+64 encoding. If a historical file shows almost no punctuation in the quality line and a wall of letters, check the encoding before you celebrate universal high quality. Subtracting 33 from a Phred+64 file invents spectacular scores. The Sequence Read Archive and the European Nucleotide Archive record platform metadata that helps, and they do not remove the need to look at the characters.
Paired-end data are two files, or a interleaved arrangement you should not assume. The mate is not hiding in the quality line. Records are mates when their identifiers and their order say so, under the convention of the tool that will align them. A sort that reorders one file and not the other creates confident chimeric pairs. The quality scores will look fine.
What Q20 and Q30 are for
A Phred score Q means the caller estimates an error probability of ten raised to the power of minus Q over ten. Q20 is an estimated one error in a hundred. Q30 is an estimated one error in a thousand. Those two thresholds are widely used in planning and in run summaries: people report the fraction of bases at or above Q30, and they sometimes trim the tail of a read where scores fall below Q20. Both habits are reasonable. Neither is a universal pass rule.
A project can need a stricter or a looser tail. A regulatory or clinical validation, which this article is not, would set its own threshold and defend it. A research run can keep bases below Q20 and let the aligner and the variant model down-weight them, which is often kinder than hard-trimming a read into a dangerous stub. Hard-trimming everything below Q30 can make short reads shorter than the mapper can place, and it can bias the ends of amplicons. Decide, and write the decision. Do not inherit a trim from a tutorial that had a different genome.
The percentage of Q30 bases is a property of the run's signal. It is not coverage, and it is not sample identity. A completely wrong library, perfectly sequenced, has a beautiful Q30 fraction. A swapped index has a beautiful Q30 fraction. Use the score to judge the basecaller, then use the other checks to judge the experiment. Run-level summaries are a screen. The per-base plot along the read is the more honest picture.
The shape of a healthy decline, and the shapes that are not
On reversible-terminator sequencing, quality often starts high and falls toward the end of the read. Later cycles are harder: some molecules in a cluster fall behind, signal dims, and the caller spreads its probability. A per-base drop toward the end is expected. You plan for it by not staking a difficult variant on the last few cycles alone, and by pairing reads so the far end of one read is the near end of its mate when the insert is short enough.
A drop to useless scores in the first ten cycles is not that pattern. It suggests a cluster problem, a reagent or focus failure, or a library so full of low-complexity sequence that the instrument could not calibrate. Adapter dimers produce a related ugly plot: the read becomes adapter, and quality behaviour depends on that sequence. Look at over-represented sequences before you repeat the sequencing and leave the library unchanged. The prep branch is in library prep is where most runs are won.
Long-read FASTQ, when a workflow emits one, does not owe you the same end-of-read curve. The error mode belongs to that chemistry, as how short reads and long reads differ insists. Applying a short-read Q30 trim recipe to a long-read file can delete the evidence you bought the chemistry for. Read the exporter's note on what its quality values mean. Some long-read scores are not interchangeable with an Illumina Phred in a variant caller tuned for the other.
Sanger chromatogram quality is a cousin, not a twin. Phred scores were born on Sanger traces. A FASTQ exported from a chromatogram still describes one reaction's peaks, including dye blobs and mixed peaks, and it is still not a multi-molecule variant call. The trace reading is Sanger sequencing for a single amplicon.
| Choice | A sound reason to make it | A poor reason to make it |
|---|---|---|
| Keep the raw FASTQ | You may retrim or realign later | The summary already said Q30, so the file is done |
| Trim a failing tail | Late cycles are dominating downstream noise | A tutorial used Q30 and you have not looked at this plot |
| Allow N bases | The caller was honestly unsure | You hope N will map as a match |
| Reject a whole read on low mean quality | The read is mostly unusable and short | One low base in an otherwise placed read scares you |
| Demand Phred+33 in the specification | Mixed encodings silently rescale every score | You assume every file since the invention of FASTQ used it |
| Archive mates together | Order and identifiers stay paired | Filenames look similar, so they must be mates |
Selecting tools and deliverables without fooling yourself
Viewers that plot per-base quality, per-base sequence content, adapter content, and duplication are how you see the file. They are not a pass certificate. A green summary beside a failed blank is still a failed blank. Keep the plot with the record so a collaborator can see the decline you decided to keep.
Compression and file integrity belong in the selection too. A FASTQ that lost its last lines in a copy will fail a parse because a record is incomplete, or, worse, will shift the four-line frame if a line was truncated in the middle. Check that records parse before a long alignment. A checksum written at the source, checked at the destination, is the selection criterion for the transfer itself.
When you choose a trimming tool, select one that can keep mates in register. Trimming one file more aggressively than its pair, then aligning as pairs, creates a systematic mess at the ends. When you choose an aligner, confirm it expects Phred+33. Most current ones do. A pipeline assembled from an old blog and a new file is how encodings get crossed.
Deposition completes the selection. Public archives such as the Sequence Read Archive and the European Nucleotide Archive expect the reads and enough metadata to know the platform. A quality score without a platform is an orphaned number.
Scores do not decontaminate a sample
A high-quality FASTQ of an infectious organism's genome is still that organism's sequence data, and the sample that produced it may have been infectious on the bench. Institutional biosafety rules cover the work. Data-handling rules and the ethics approval cover the file if the source was a person. This guide does not approve diagnostic reporting. Do not email a FASTQ of a human sample around as if quality scores had anonymised it. Identifiers in the read names can also carry instrument details you did not intend to publish. Look at them before a public deposit.
A cut in power, and a file that is short by a third
Copying a large FASTQ when the power fails leaves a partial file that can look plausible until a parser hits the end. Make the checksum part of the handoff between buildings, not an optional extra. Heat matters less to the bytes than to the disk they live on: a drive left in a hot vehicle is a physical risk to the only copy. Keep two copies, and do not delete the instrument copy until the checksum matches. In the written specification, name Phred+33, whether mates are separate files, and the minimum parse check the receiver will run. A project that accepts any file ending in fq will eventually accept a Phred+64 relic or a truncated copy and call the resulting variants a discovery.
What to put in the enquiry
Ask for FASTQ as a deliverable when you need the calls, and say Phred+33 explicitly. Ask whether pairs will be delivered, whether adapters will already have been trimmed, and which summary plots accompany the file. Ask for the checksum method. If you also need alignments, name the reference build in the same request so the FASTQ and the alignment describe the same plan.
A genome-scale file set can be discussed against the whole-genome sequencing enquiry reference. A single-trace alternative can be discussed against the Sanger DNA sequencing enquiry reference. Consumable classes are in the genomics and sequencing catalogue, with the study context on the genomics research pathway. Put the file format and the encoding in the quote request. The checks before anyone relies on the file are in what to check before a sequencing run, and the specification habit is in commissioning a sequencing or proteomics study.
Questions from the bench
What are the four lines of a FASTQ record?
The first line starts with @ and carries the read identifier. The second line is the base sequence. The third line starts with + and may repeat the identifier or stay otherwise empty. The fourth line is the quality string, one character per base, encoded as a Phred score. The quality line is not DNA, even when some of its characters look like bases.
Does a Q30 base mean the base is correct?
It means the caller assigned an estimated error probability of one in a thousand for that base, because a Phred score is minus ten times the log of the error probability. The base can still be wrong. Across a genome, many Q30 bases will be wrong somewhere. Q30 is a per-base estimate, not a seal on the read and not a seal on a later variant call.
Why do qualities fall toward the end of a short read?
On sequencing-by-synthesis, later cycles accumulate phasing, fading signal, and overlapping clusters. The caller becomes less sure, and the quality characters decline. A gentle decline is ordinary. A collapse in the first few cycles is a different fault, often cluster generation, focus, or a library of dimers, and it should not be shrugged off as the usual end-of-read pattern.
Which file should a project request, FASTQ or something already aligned?
Ask for FASTQ when you need to repeat trimming, alignment, or a check of the raw calls. Ask for an alignment as well when you also need a defined reference and duplicate marking. A BAM without the FASTQ can be enough for some reviews and is useless if the reference build was wrong. Name the build either way. Quality scores inside a FASTQ do not record which genome you intended.
References
Manufacturer names identify published method classes. Trademarks remain with their owners. Catalogue records on this site are independent references for enquiry. They are not a statement of inventory, distribution rights or a supply commitment. This page is educational. It is not medical advice, a diagnostic protocol or a biosafety approval.
Catalogue
Related products and categories
These links follow the subject of the article into published manufacturer references. A listing is a reference for an enquiry, not a statement of stock or distribution rights.
Continue in this cluster
Related reading
Next-generation sequencing from library to readsHow a sequencing library becomes reads: adapters, flow cells, quality scores and the checks that stop a bad library from wasting a run.
16S profiling and its taxonomic limitsJudge a 16S profile at the rank the marker supports, and see where copy number, primer bias and species names stop being honest.
A glossary of sequencing termsWorking definitions of read, coverage, depth, MAPQ, Phred, VCF, BAM and the related words, each tied to the mistake that word prevents.