comparison
Coverage depth is not the same as accuracy
Separate coverage depth from base accuracy, mapping quality, strand bias and PCR copies so a deep pile of one molecule is not truth.
- Author
- EVRINTH Editorial Team
- Published
- 8 October 2026
- Updated
- 8 October 2026
- Reading time
- 9 min

Coverage depth counts reads. Accuracy asks whether those reads are right, whether they are independent, and whether they landed on the correct place. The two numbers are sold as if they were one purchase. They are not. This comparison is how to keep them apart when you write a plan, a specification, or a sentence in a paper. The path that produces the reads is next-generation sequencing from library to reads. The base qualities inside each read are a further file format, set out in quality scores and a FASTQ file.
A mean of one hundred fold over a target sounds decisive. One hundred copies of a wrong molecule are still that molecule. The rest of this page is what else has to be true before the count deserves a biological sentence.
What a depth number is actually counting
At a single reference position, raw depth is the number of reads whose alignment overlaps that position. Mean depth is the average of that count across a region. Both depend on the reference you mapped to and on whether you counted duplicate reads, supplementary alignments, and reads with poor mapping scores. A methods sentence that says coverage without those choices cannot be compared with another paper's coverage.
Breadth is the other axis. A genome sequenced to a high mean can still leave a medically or biologically interesting exon at zero if the library, the capture, or the GC content refused that sequence. The mean does not confess the hole. Plot the fraction of bases at or above a few thresholds, or look at the per-exon distribution, before you accept the mean as a description of the experiment. Capture designs make this failure ordinary. It is one reason whole genome versus whole exome concepts treats evenness as part of the assay, not as a footnote.
Sampling depth does real work against random absence. If independent reads truly arrive from both alleles of a diploid site, more of them make it less likely that you simply missed the second allele. That is a sampling statement. It assumes the reads are independent draws from the alleles in the sample. PCR duplicates break the assumption. So does a library made from three starting molecules. So does a mapping that piles every copy of a repeat onto one coordinate.
A common planning figure for a short-read diploid genome sits on the order of a few tens of reads of mean depth. That figure is a sampling plan for heterozygous sites under a particular error model. It is not an accuracy guarantee, and it is not a reason to ignore duplicates. An exome plan often aims higher because capture is uneven and some exons would otherwise fall to nothing. Higher, in that sentence, still means more counts. It does not convert the counts into truth.
The properties that are not depth
Base accuracy is the quality score on a base call, the estimated chance that this base is the wrong letter. A deep pile of low-quality bases is a deep pile of unreliable letters. Thresholds such as Q20 and Q30 are planning language for that estimate. They are defined in the FASTQ page. They do not know whether the read is a duplicate.
Mapping quality is a different estimate: the chance the read is aligned in the wrong place. Repeats, short inserts, and reference gaps produce reads with confident bases and uncertain addresses. A variant caller that trusts the address will report a difference at a coordinate the read may not belong to. Record how you filtered mapping quality. Leaving the filter at the tool's default is a choice. Say so.
Strand bias is a clue that the observation is an artefact of the chemistry rather than a molecule that existed on both strands. Some error modes appear at the ends of reads, or on one strand, because of a context the polymerase or the sequencer mishandles. A real heterozygous allele in a ordinary library is usually seen on both strands. A site that is one hundred reads of one strand and one read of the other is not a stronger heterozygote because the one hundred is large. It is a more emphatic bias.
Duplicates are PCR copies or optical copies. Marking them is not a decorative step. Once a molecule has been copied, further reads repeat its errors, including a polymerase mistake in an early cycle. The mistake can reach a high apparent allele fraction inside that family. Duplicate-unaware depth will defend the mistake with a large integer. Unique depth, or a count of unique molecular identifiers when the library has them, is the integer that matches the sampling story.
PCR error and contamination sit in the same family of problems. They are systematic relative to the library even when they are random relative to the genome. Library prep is where most runs are won is where those copies are created. Index misassignment, in adapters indexes and barcode hopping, can add a thin layer of someone else's genome. At high depth that thin layer becomes visible and can be called as a low allele fraction. The blank tells you the floor. Depth without a blank turns the floor into discoveries.
Where a high number should stop being used
Use raw depth to talk about loading and about whether the sequencer did work. Use duplicate-aware depth to talk about how thoroughly you sampled independent molecules. Use breadth to talk about whether the target was actually in the file. Use base quality, mapping quality, and strand balance to talk about whether a particular site is believable. Use an orthogonal method when the site changes a conclusion you cannot afford to retract.
Sanger sequencing of a fresh amplicon is that orthogonal method for a high-fraction allele. It is not deeper coverage of the same library. It is a different chemistry on, ideally, a new PCR. Its limits, including mixtures and the short useful window, are in Sanger sequencing for a single amplicon. A low-fraction allele that Sanger cannot see is not refuted by a clean Sanger peak. Match the check to the fraction you claimed.
Branch when the duplicate rate is high and the unique depth is below the plan. The run may have produced enormous files. The experiment did not reach its sampling goal. Repeating the analysis with a looser filter does not create molecules. A new library from more input does, when more input exists.
| Number people quote | What it measures | What can still be wrong |
|---|---|---|
| Mean raw depth | Average overlapping reads, duplicates included | Copies of one molecule, or of a contaminant |
| Unique depth | Reads judged independent after duplicate marking | Marking can miss or over-collapse families |
| Breadth at a threshold | Fraction of the target with enough reads | The threshold can be too low for the allele you need |
| Percent of bases at Q30 | Base-call confidence across the file | Confident bases can be mapped to the wrong locus |
| Mapping quality filter | How strictly unplaced reads were dropped | A hard filter does not fix a bad reference |
| Strand balance at a site | Whether both strands contributed | Balance does not prove the right sample was sequenced |
Failure modes that wear a large integer
Jackpot amplification of one fragment makes a local depth spike. In an amplicon or a capture, that spike is often a small, happy molecule, not a biological copy-number change. Compare the spike with the insert size and with the duplicate flag before you report an amplification in the genome.
Mismapping of a paralogue creates a confident difference from the reference at high depth. The reads are real. The address is shared with a relative the reference also contains, or fails to contain. Reference genomes and why the build matters is the other half of that trap. Adding depth makes the false difference cleaner.
Soft clipping at a site, or variants that appear only in the last few cycles, follow the per-base decline of short-read chemistry. A call supported only by read ends is weaker than the depth suggests, because those cycles are the least accurate. Look at where in the read the alternate allele sits.
A specification that pays for a round number and never mentions these checks will receive the round number. The scientific failure will be compliant.
Research claims, infectious material, and no diagnostic badge
Depth is not a clinical validation. A research alignment, however deep, stays inside the ethics approval and the question that approval named. Samples may be infectious before they are sequence. The institutional biosafety decision covers the handling. This comparison does not. Do not describe a high-depth research file as a diagnosis, and do not describe a low-depth file as a clean bill of health. Absence of a read is a statement about the library and the threshold, which variant calling is a model not a fact keeps in the realm of models.
Put independence in the written brief
Across collaborating benches, the sentence "we need 100x" is how duplicated libraries pass a handoff. Write unique depth or a stated duplicate policy, the minimum breadth, the reference build, and the mapping-quality floor into the specification. Name who calculates those figures, because two tools mark duplicates differently and both will claim to have met the integer. If a power cut or a heat excursion forced a repeat library, do not average its depth with the failed partial run to reach the integer. The molecules were not one experiment.
What to ask for when depth is the commercial question
Ask what will be delivered besides a mean: a breadth histogram, a duplicate rate, a mapping-rate, and the reference FASTA. Ask how low-depth regions of the target will be reported rather than averaged away. Ask which orthogonal check is included when a site becomes a claim.
Those design choices can be discussed against the whole-genome sequencing enquiry reference. A claim that is really one amplicon can be discussed against the Sanger DNA sequencing enquiry reference, where depth in the short-read sense is the wrong purchase. Reagent classes sit in the genomics and sequencing catalogue, and the study shape sits on the genomics research pathway. Send the acceptance numbers, not a single fold, with the quote request. Commissioning a sequencing or proteomics study is the longer habit of writing those numbers before the library exists.
Questions from the bench
If every read agrees, does the depth make the base true?
Agreement among reads is only as good as the independence of those reads. PCR copies of one molecule agree with each other even when the molecule already carried an error. A contaminant that dominated the library will also agree with itself. Depth multiplies whatever was in the tube.
What is the difference between depth and breadth?
Depth is how many reads cover a position. Breadth is how much of the target is covered at all. A mean depth of a hundred can hide an exon with none. Report both, and report the fraction of the target above the depth your question needs, or the mean will flatter a uneven library.
Where does mapping quality fit?
Mapping quality is the aligner's statement that this read might belong somewhere else. A base can have an excellent call quality and a poor mapping quality when the sequence is a repeat. Counting that read toward a variant treats a placement guess as an observation. Filters that ignore mapping quality will call variants in the repeats you cannot resolve.
Should a specification name a single fold of coverage?
A single fold number is a loading target, not an acceptance test for truth. Specify unique molecules or duplicate-marked depth, the fraction of the target covered, and what you will do about strand bias and low mapping quality. A file that hits a round number and fails those checks has not met a scientific specification.
References
Manufacturer names identify published method classes. Trademarks remain with their owners. Catalogue records on this site are independent references for enquiry. They are not a statement of inventory, distribution rights or a supply commitment. This page is educational. It is not medical advice, a diagnostic protocol or a biosafety approval.
Catalogue
Related products and categories
These links follow the subject of the article into published manufacturer references. A listing is a reference for an enquiry, not a statement of stock or distribution rights.
Continue in this cluster
Related reading
Next-generation sequencing from library to readsHow a sequencing library becomes reads: adapters, flow cells, quality scores and the checks that stop a bad library from wasting a run.
16S profiling and its taxonomic limitsJudge a 16S profile at the rank the marker supports, and see where copy number, primer bias and species names stop being honest.
A glossary of sequencing termsWorking definitions of read, coverage, depth, MAPQ, Phred, VCF, BAM and the related words, each tied to the mistake that word prevents.