guide
A glossary of sequencing terms
Working definitions of read, coverage, depth, MAPQ, Phred, VCF, BAM and the related words, each tied to the mistake that word prevents.
- Author
- EVRINTH Editorial Team
- Published
- 8 October 2026
- Updated
- 8 October 2026
- Reading time
- 9 min

People talk past each other in sequencing meetings because coverage and depth are not the same noun, and because a VCF line is not a read. This glossary gives working definitions for the words that show up between a library and a claim. Each definition includes the mistake it is meant to prevent. The physical path from library to read is next-generation sequencing from library to reads. How those files become a result you can argue with is from reads to a result a biologist can question. The definitions below are research vocabulary. They do not certify an assay.
A read is one observation, and a contig is a consensus stretch
A read is the instrument's record of one pass along a fragment: a string of bases plus, for most short-read and many long-read formats, a quality for each base. One read is not one independent molecule of the genome if PCR made many copies and the sequencer saw several of them. Treating every read as an independent biological event is the mistake this word prevents. When you mean the fragment that went into the library, say fragment or molecule, and then show how you collapsed copies.
A contig is a contiguous sequence built by assembly, a consensus of overlapping reads. It is not a chromosome, and it is not a proof that the genome is finished. The mistake it prevents is quoting a contig length as the size of a chromosome arm, or treating a gap between contigs as a biological deletion. Gaps are often places the assembler refused to guess. Sanger sequencing still produces a kind of read, a trace of one reaction. Calling that trace a contig confuses a single-molecule chromatogram with an assembly.
Coverage, depth, duplicates, and insert size
Coverage, used carefully, is breadth: the fraction of the reference that received reads. Depth is the count at a base, and mean depth is that count averaged over a region you must specify. The mistake "30x coverage" prevents, when you unpack it, is believing the whole genome is uniformly 30 reads deep. It may be 30 on average and zero in the repeats. Say mean depth and say breadth when you need both.
A duplicate is a read pair, or a read, judged to be a copy of another from PCR or from an optical accident on the flow cell. The mistake it prevents is counting those copies as extra evidence. Marking duplicates and then still using them to call a rare allele invents support. Some applications, such as certain counting assays, have their own rules. Write the rule down. Do not inherit a variant-calling habit.
Insert size is the length of the genomic fragment between the adapters, not the length of the read. In paired-end sequencing the two reads spy on the ends and the insert size is inferred. The mistake it prevents is designing a "150 base pair insert" when you meant a 150 base pair read, and then wondering why mates overlap or why they span a different distance from the one the structural caller expected.
Phred, MAPQ, adapter, and index
A Phred score is a base quality. In the common definition, a score of Q means the estimated probability of an error at that base is ten to the power of minus Q over ten. A high score is confidence in the base call. It is not confidence in where the read belongs. The mistake it prevents is throwing away a whole read because one base is weak, or, in the other direction, trusting a variant base solely because the quality character looked impressive in a text file.
MAPQ is a mapping quality, a score for how confidently the aligner placed the read. Low MAPQ is what repeats produce when several genomic homes look equally good. The mistake it prevents is summing a multi-mapping read into a depth track as if it had one true address. Filters that ignore MAPQ turn repeats into false peaks and false variants.
An adapter is the synthetic sequence added so the fragment can bind the instrument's chemistry. It is not biological sequence. The mistake it prevents is aligning adapter bases to the genome and calling them a variant at the end of a read. Trim adapters, and record that you did. An index, sometimes called a barcode in the sample-sheet sense, is a short sequence that marks which sample a read came from. It is not an adapter, though it is often read from the same synthetic end, and it is not a unique molecular identifier unless your protocol added one on purpose. The mistake it prevents is demultiplexing on the wrong bases and then analysing a mixture under a single name. Collisions between indexes are treated in index collisions in multiplexed pools.
VCF, BAM, CRAM, and demultiplex
A VCF is a variant call format file: positions where the sample is claimed to differ from a reference, with fields for genotype, quality and filters. The mistake it prevents is treating the file as raw data. It is an argument. A different caller, or a looser filter, produces a different VCF from the same reads. Coordinates assume a reference build.
A BAM is a binary alignment of reads to a reference. A CRAM is a more compressed alignment that typically needs that same reference to be expanded. The mistake both terms prevent is archiving "the alignment" without the build, and the further mistake of assuming a CRAM can be opened years later on a machine that no longer has that genome. A FASTQ is the usual file of unaligned reads and qualities. Demultiplex is the split of a pooled run into per-sample files using the index. The mistake it prevents is believing the split checks biology. It checks the index against the sheet. A wrong sheet demultiplexes cleanly into the wrong names.
You can view a coordinate from these files on the UCSC Genome Browser only after the build matches. Reads you later share may belong in the Sequence Read Archive, under the consent you have, which is a policy decision and not a property of the file format.
| Pair of words | Keep them apart because | The mix-up it causes |
|---|---|---|
| Coverage and depth | Breadth is not pile-up height | A hole hidden inside a flattering mean |
| Phred and MAPQ | A base can be right and the address wrong | Variant calls inside repeats |
| Insert size and read length | The fragment is longer than what was read | Libraries built to the wrong geometry |
| VCF and BAM | Calls are filtered arguments; alignments are the reads placed | A result with nothing left to reanalyse |
| Index and adapter | One sorts samples; the other sticks the fragment down | Trimming the barcode, or binning on adapter sequence |
Where the words sit on a pile of reads
Picture one locus. Several reads overlap a base. Some are marked duplicate. One aligns with a low MAPQ because the next exon looks similar. Adapter sequence, if you forgot to trim, dangles off the end and fails to match. The depth you quote depends on which of those reads you keep. The glossary is a way to say that sentence in labels.
Using the words when a report goes wrong
If a collaborator says the sample has excellent quality, ask for the Phred distribution and the MAPQ distribution separately. If they say the gene is covered, ask for breadth over that gene and for depth at the exon you care about. If they say the variant is in the VCF, ask which filter it passed and whether the BAM still shows it on both strands. If a demultiplex report looks perfect and the biology looks swapped, the glossary will not mention the plate map, but sample identity mix-ups and plate maps will. The words stop you curing an identity problem with a stricter Phred cut.
The short-read method class these formats grew up around is outlined by Illumina. Long reads use the same nouns with different error shapes. Do not copy a short-read MAPQ habit onto a long-read alignment without reading what that aligner means by the score.
Research vocabulary is not a clinical grade
None of these terms turn a file into a diagnosis, a biosafety clearance, or an approved test. A VCF of a human sample in a research project remains a research file. Reporting it to a patient is a different activity, under different rules, and this glossary does not authorise it. Handle the biological sample under your institution's rules before any of the files exist.
Shared drives and mixed dialects
When a core facility and a research group in another city exchange files, they often use coverage for depth and barcode for index. Agree the glossary in the statement of work before the drive is copied. A partial download after a dropped connection should be checked against a checksum, and the file names should already contain the build and the filter. A folder called final_final is how the wrong VCF becomes the archive. The statement of work is the place those nouns get pinned down.
How to use the words in an enquiry
Say the deliverable with the glossary's nouns: FASTQ, BAM or CRAM against a named build, VCF with the filter described, and demultiplex rules for the indexes. Say whether mean depth or breadth is the acceptance metric. Consumables that feed the run are in the genomics and sequencing catalogue. The research context is the genomics research pathway. A genome-scale file set can be discussed through the whole-genome sequencing enquiry reference. A single-locus trace, which is a different kind of read, can be discussed through the Sanger DNA sequencing enquiry reference. Put the vocabulary in the quote request so the files that come back still answer the question you asked.
Read a sequencing report with the glossary open
- 01Separate the molecule from the fileSay whether you are holding a read, a contig, an alignment or a variant call. A FASTQ base and a VCF allele are different claims. Write down which file the sentence in the report actually came from.
- 02Name the reference and the filterRecord the genome build beside any coordinate, and record the filter string beside any variant. A position without a build cannot be checked in a browser. A call without a filter cannot be compared with next month's call.
- 03Match the quality word to the error it describesUse Phred for a base and MAPQ for a placement. If the report says only quality, ask which one. Loading a high-Phred read that mapped to three places will inflate a pileup you should not trust.
- 04Keep the file that can still answer the questionIf you may need to refilter, keep the VCF unfiltered or the BAM. If you may need to realign, keep the FASTQ. A slide of gene names cannot support any of those second looks.
Questions from the bench
Is coverage the same number as depth?
No. Coverage in the careful sense is how much of the reference has at least some reads, a breadth. Depth is how many reads pile onto a base. A sample can have high mean depth and still leave a tenth of the genome uncovered. Ask which quantity the report plotted.
Does a high Phred score mean the read was mapped to the right place?
It means the base caller was confident in that base. Placement is MAPQ, or an equivalent mapping score. A beautifully called base inside a repeat can still be assigned to the wrong copy. Keep the two scores in separate columns in your notes.
Why keep a BAM if a VCF has already been delivered?
The VCF is the set of differences someone chose to report after filtering. The BAM is the alignments those calls were made from, on a named reference. Without it you cannot see the strand balance, the duplicates, or the reads that were almost a variant and were dropped.
References
Manufacturer names identify published method classes. Trademarks remain with their owners. Catalogue records on this site are independent references for enquiry. They are not a statement of inventory, distribution rights or a supply commitment. This page is educational. It is not medical advice, a diagnostic protocol or a biosafety approval.
Catalogue
Related products and categories
These links follow the subject of the article into published manufacturer references. A listing is a reference for an enquiry, not a statement of stock or distribution rights.
Continue in this cluster
Related reading
Next-generation sequencing from library to readsHow a sequencing library becomes reads: adapters, flow cells, quality scores and the checks that stop a bad library from wasting a run.
16S profiling and its taxonomic limitsJudge a 16S profile at the rank the marker supports, and see where copy number, primer bias and species names stop being honest.
Adapters indexes and barcode hoppingSeparate adapter sequence from i5 and i7 indexes, and see why unique dual indexes resist barcode hopping better than one shared index.