Skip to content
EVRINTH

comparison

Coverage depth is not the same as accuracy

Separate coverage depth from base accuracy, mapping quality, strand bias and PCR copies so a deep pile of one molecule is not truth.

Author
EVRINTH Editorial Team
Published
8 October 2026
Updated
8 October 2026
Reading time
9 min
Benchtop sequencing instrument with a teal status light and a flow cell cartridge in a genomics lab
Benchtop sequencing instrument with a teal status light and a flow cell cartridge in a genomics lab

Coverage depth counts reads. Accuracy asks whether those reads are right, whether they are independent, and whether they landed on the correct place. The two numbers are sold as if they were one purchase. They are not. This comparison is how to keep them apart when you write a plan, a specification, or a sentence in a paper. The path that produces the reads is next-generation sequencing from library to reads. The base qualities inside each read are a further file format, set out in quality scores and a FASTQ file.

A mean of one hundred fold over a target sounds decisive. One hundred copies of a wrong molecule are still that molecule. The rest of this page is what else has to be true before the count deserves a biological sentence.

What a depth number is actually counting

At a single reference position, raw depth is the number of reads whose alignment overlaps that position. Mean depth is the average of that count across a region. Both depend on the reference you mapped to and on whether you counted duplicate reads, supplementary alignments, and reads with poor mapping scores. A methods sentence that says coverage without those choices cannot be compared with another paper's coverage.

Breadth is the other axis. A genome sequenced to a high mean can still leave a medically or biologically interesting exon at zero if the library, the capture, or the GC content refused that sequence. The mean does not confess the hole. Plot the fraction of bases at or above a few thresholds, or look at the per-exon distribution, before you accept the mean as a description of the experiment. Capture designs make this failure ordinary. It is one reason whole genome versus whole exome concepts treats evenness as part of the assay, not as a footnote.

Sampling depth does real work against random absence. If independent reads truly arrive from both alleles of a diploid site, more of them make it less likely that you simply missed the second allele. That is a sampling statement. It assumes the reads are independent draws from the alleles in the sample. PCR duplicates break the assumption. So does a library made from three starting molecules. So does a mapping that piles every copy of a repeat onto one coordinate.

A common planning figure for a short-read diploid genome sits on the order of a few tens of reads of mean depth. That figure is a sampling plan for heterozygous sites under a particular error model. It is not an accuracy guarantee, and it is not a reason to ignore duplicates. An exome plan often aims higher because capture is uneven and some exons would otherwise fall to nothing. Higher, in that sentence, still means more counts. It does not convert the counts into truth.

The properties that are not depth

Base accuracy is the quality score on a base call, the estimated chance that this base is the wrong letter. A deep pile of low-quality bases is a deep pile of unreliable letters. Thresholds such as Q20 and Q30 are planning language for that estimate. They are defined in the FASTQ page. They do not know whether the read is a duplicate.

Mapping quality is a different estimate: the chance the read is aligned in the wrong place. Repeats, short inserts, and reference gaps produce reads with confident bases and uncertain addresses. A variant caller that trusts the address will report a difference at a coordinate the read may not belong to. Record how you filtered mapping quality. Leaving the filter at the tool's default is a choice. Say so.

Strand bias is a clue that the observation is an artefact of the chemistry rather than a molecule that existed on both strands. Some error modes appear at the ends of reads, or on one strand, because of a context the polymerase or the sequencer mishandles. A real heterozygous allele in a ordinary library is usually seen on both strands. A site that is one hundred reads of one strand and one read of the other is not a stronger heterozygote because the one hundred is large. It is a more emphatic bias.

Duplicates are PCR copies or optical copies. Marking them is not a decorative step. Once a molecule has been copied, further reads repeat its errors, including a polymerase mistake in an early cycle. The mistake can reach a high apparent allele fraction inside that family. Duplicate-unaware depth will defend the mistake with a large integer. Unique depth, or a count of unique molecular identifiers when the library has them, is the integer that matches the sampling story.

PCR error and contamination sit in the same family of problems. They are systematic relative to the library even when they are random relative to the genome. Library prep is where most runs are won is where those copies are created. Index misassignment, in adapters indexes and barcode hopping, can add a thin layer of someone else's genome. At high depth that thin layer becomes visible and can be called as a low allele fraction. The blank tells you the floor. Depth without a blank turns the floor into discoveries.

Where a high number should stop being used

Use raw depth to talk about loading and about whether the sequencer did work. Use duplicate-aware depth to talk about how thoroughly you sampled independent molecules. Use breadth to talk about whether the target was actually in the file. Use base quality, mapping quality, and strand balance to talk about whether a particular site is believable. Use an orthogonal method when the site changes a conclusion you cannot afford to retract.

Sanger sequencing of a fresh amplicon is that orthogonal method for a high-fraction allele. It is not deeper coverage of the same library. It is a different chemistry on, ideally, a new PCR. Its limits, including mixtures and the short useful window, are in Sanger sequencing for a single amplicon. A low-fraction allele that Sanger cannot see is not refuted by a clean Sanger peak. Match the check to the fraction you claimed.

Branch when the duplicate rate is high and the unique depth is below the plan. The run may have produced enormous files. The experiment did not reach its sampling goal. Repeating the analysis with a looser filter does not create molecules. A new library from more input does, when more input exists.

Number people quoteWhat it measuresWhat can still be wrong
Mean raw depthAverage overlapping reads, duplicates includedCopies of one molecule, or of a contaminant
Unique depthReads judged independent after duplicate markingMarking can miss or over-collapse families
Breadth at a thresholdFraction of the target with enough readsThe threshold can be too low for the allele you need
Percent of bases at Q30Base-call confidence across the fileConfident bases can be mapped to the wrong locus
Mapping quality filterHow strictly unplaced reads were droppedA hard filter does not fix a bad reference
Strand balance at a siteWhether both strands contributedBalance does not prove the right sample was sequenced
Depth of copies versus independent molecules The same apparent depth, two histories 100 reads, one parent molecule An early PCR error is on every copy. Independent molecules Disagreement can be real alleles.
One hundred reads that are copies of a single wrong molecule agree with each other, while a handful of independent molecules can disagree and still be the more honest sample.

Failure modes that wear a large integer

Jackpot amplification of one fragment makes a local depth spike. In an amplicon or a capture, that spike is often a small, happy molecule, not a biological copy-number change. Compare the spike with the insert size and with the duplicate flag before you report an amplification in the genome.

Mismapping of a paralogue creates a confident difference from the reference at high depth. The reads are real. The address is shared with a relative the reference also contains, or fails to contain. Reference genomes and why the build matters is the other half of that trap. Adding depth makes the false difference cleaner.

Soft clipping at a site, or variants that appear only in the last few cycles, follow the per-base decline of short-read chemistry. A call supported only by read ends is weaker than the depth suggests, because those cycles are the least accurate. Look at where in the read the alternate allele sits.

A specification that pays for a round number and never mentions these checks will receive the round number. The scientific failure will be compliant.

Research claims, infectious material, and no diagnostic badge

Depth is not a clinical validation. A research alignment, however deep, stays inside the ethics approval and the question that approval named. Samples may be infectious before they are sequence. The institutional biosafety decision covers the handling. This comparison does not. Do not describe a high-depth research file as a diagnosis, and do not describe a low-depth file as a clean bill of health. Absence of a read is a statement about the library and the threshold, which variant calling is a model not a fact keeps in the realm of models.

Put independence in the written brief

Across collaborating benches, the sentence "we need 100x" is how duplicated libraries pass a handoff. Write unique depth or a stated duplicate policy, the minimum breadth, the reference build, and the mapping-quality floor into the specification. Name who calculates those figures, because two tools mark duplicates differently and both will claim to have met the integer. If a power cut or a heat excursion forced a repeat library, do not average its depth with the failed partial run to reach the integer. The molecules were not one experiment.

What to ask for when depth is the commercial question

Ask what will be delivered besides a mean: a breadth histogram, a duplicate rate, a mapping-rate, and the reference FASTA. Ask how low-depth regions of the target will be reported rather than averaged away. Ask which orthogonal check is included when a site becomes a claim.

Those design choices can be discussed against the whole-genome sequencing enquiry reference. A claim that is really one amplicon can be discussed against the Sanger DNA sequencing enquiry reference, where depth in the short-read sense is the wrong purchase. Reagent classes sit in the genomics and sequencing catalogue, and the study shape sits on the genomics research pathway. Send the acceptance numbers, not a single fold, with the quote request. Commissioning a sequencing or proteomics study is the longer habit of writing those numbers before the library exists.

Questions from the bench

If every read agrees, does the depth make the base true?

Agreement among reads is only as good as the independence of those reads. PCR copies of one molecule agree with each other even when the molecule already carried an error. A contaminant that dominated the library will also agree with itself. Depth multiplies whatever was in the tube.

What is the difference between depth and breadth?

Depth is how many reads cover a position. Breadth is how much of the target is covered at all. A mean depth of a hundred can hide an exon with none. Report both, and report the fraction of the target above the depth your question needs, or the mean will flatter a uneven library.

Where does mapping quality fit?

Mapping quality is the aligner's statement that this read might belong somewhere else. A base can have an excellent call quality and a poor mapping quality when the sequence is a repeat. Counting that read toward a variant treats a placement guess as an observation. Filters that ignore mapping quality will call variants in the repeats you cannot resolve.

Should a specification name a single fold of coverage?

A single fold number is a loading target, not an acceptance test for truth. Specify unique molecules or duplicate-marked depth, the fraction of the target covered, and what you will do about strand bias and low mapping quality. A file that hits a round number and fails those checks has not met a scientific specification.

References

  1. Ensembl genome browser
  2. UCSC Genome Browser
  3. European Nucleotide Archive

Manufacturer names identify published method classes. Trademarks remain with their owners. Catalogue records on this site are independent references for enquiry. They are not a statement of inventory, distribution rights or a supply commitment. This page is educational. It is not medical advice, a diagnostic protocol or a biosafety approval.

Catalogue

Related products and categories

These links follow the subject of the article into published manufacturer references. A listing is a reference for an enquiry, not a statement of stock or distribution rights.