Skip to content
EVRINTH

explainer

Whole genome versus whole exome concepts

Compare whole-genome and whole-exome designs on non-coding sequence, exon evenness, structural variants and which capture kit was used.

Author
EVRINTH Editorial Team
Published
8 October 2026
Updated
8 October 2026
Reading time
9 min
Benchtop sequencing instrument with a teal status light and a flow cell cartridge in a genomics lab
Benchtop sequencing instrument with a teal status light and a flow cell cartridge in a genomics lab

Whole-genome sequencing and whole-exome sequencing are different assays that share a sequencer and a reputation. One fragments the genome and reads across it. The other pulls down a designed subset, mostly annotated exons, and reads that subset very thoroughly in places and poorly in others. This explainer is how to choose, and how to stop calling both of them "the sequence" of a person, an animal, or a cell line. The shared mechanics of libraries and reads are in next-generation sequencing from library to reads.

The choice is a scientific one before it is a file-size one. A question about a deep intronic change, a regulatory region, or a structural variant with breakpoints outside exons is a genome question. A question that is honestly limited to coding sequence, in a design where many samples matter more than evenness, can be an exome question. A question about one known exon is often a PCR question, and then Sanger sequencing for a single amplicon or a small amplicon panel is the proportionate tool.

What each method actually isolates

A whole-genome library is genomic DNA fragmented to an insert length, ligated to adapters, and sequenced. PCR may be omitted when input allows. The read budget is spread across the callable genome, including non-coding sequence. Evenness is relatively high in ordinary unique sequence because no bait had to like the fragment. Repeats, centromeres, and segmental duplications remain hard, especially for short reads, for the reasons in how short reads and long reads differ. Mitochondrial genomes often arrive at high depth as a side effect of their copy number. That depth is useful and is not a mitochondrial assay you designed.

A whole-exome library, in the method class this page means, hybridises fragmented DNA to a probe or bait set aimed at annotated exons, captures the hybrids, and sequences what came down. Older designs used PCR amplicons across exons. Those are amplicon panels with a large primer list, and they inherit primer dropout as described in amplicon sequencing for a defined region. If your "exome" is a PCR panel, say so. Capture and PCR do not fail the same way.

Capture baits are designed against a genome build and an annotation. They typically include the exon and a limited flank, so a splice-site base just inside the intron may be present and a deep intronic base will not. Some kits add selected non-coding loci. The added loci are a list, not a principle. You cannot infer them from the word exome. Record the kit name and version. A kit designed on an older build can under-represent genes the annotation later gained, and it can target coordinates that reference genomes and why the build matters would now place elsewhere. Two laboratories that both report an exome and used different kits have not run the same experiment.

Evenness, data volume, and the holes

Exome depth is bumpy. Baits differ in efficiency, GC changes hybridisation, and small exons offer a poor target. Laboratories often plan a higher mean depth on an exome than on a genome so that the weaker exons still clear a floor. The mean is not the floor. Report the fraction of target bases above the depth you need, or the holes will not appear in the paper. Coverage depth is not the same as accuracy is the general case. On an exome the target definition is itself a kit decision: off-target reads exist, and counting them into a genome-wide mean flatters the assay.

Genome data are larger. More bases are sequenced, files are heavier, and alignment costs more computation and more disk. That is an operational fact, not a scientific virtue. A genome that was sequenced too thinly to call heterozygous sites in unique sequence is not superior to a careful exome. It is a thin genome. Match the depth plan to the allele fraction you need to see, after duplicates, and do not spend the entire advantage of evenness by under-loading.

Off-target capture reads are sometimes mined for extra information, including a hint of copy number or mitochondrial sequence. Treat that mining as exploratory. The assay was not designed for uniform coverage of those regions, and the statistics of a bait-enriched library do not transfer.

Structural variants and non-coding sequence

A deletion that removes three exons can change read depth in an exome, and capture noise changes read depth too. Calling copy number from exome depth is possible in careful hands and is full of false steps in careless ones. A breakpoint in an intron is simply not in the exome file. Split reads and discordant pairs cannot land on sequence that was never captured.

A whole-genome short-read library can detect many of those events from discordant pairs, split reads, and more even depth, and it will still miss events buried in repeats longer than the insert. A long-read genome is a further step up in span, and it is a different specification from "WGS" typed as a single acronym. Write the read class. Non-coding single-nucleotide changes are in the genome file to the extent the depth and the mapping allow, and absent from the exome unless a bait happened to cover them. If the hypothesis is regulatory, an exome is the wrong search.

Neither design replaces a model of the variant. Calls remain hypotheses, as variant calling is a model not a fact insists. Germline versus somatic settings matter more, not less, when the file is large, because a wrong prior scaled to a whole genome produces a long list of confident mistakes.

Design choiceWhat you gainWhat you give up
Short-read whole genomeMore even unique sequence, non-coding bases, better SV cluesA large file, and repeats you still cannot place
Long-read whole genomeSpan across structures short inserts missA different error mode and a different prep
Capture exomeRead budget spent on a coding targetMost non-coding sequence, uneven exons, kit dependence
PCR exome panelA defined primer list you can auditPrimer dropout and chimera behaviour
One Sanger ampliconA direct trace of one windowAny claim about the rest of the gene
Genome-wide ticks versus exon baits Same locus, two designs Whole genome Whole exome Baits sit on exons. Introns stay thin. Exon height would vary with bait efficiency. A kit version decides which exons exist in the lower track.
Whole-genome coverage ticks across coding and non-coding sequence, while exome baits leave the introns unread and the exons uneven.

Failure modes that follow the design

An exome that used a kit built on a retired gene model will miss a gene you care about and will not warn you with a low-quality flag. The target bed file simply lacks the interval. Ask for the bed, and compare it with the annotation release you intend to cite. Ensembl and the UCSC Genome Browser are where you check that the interval still means what you think.

A genome aligned to a reference without decoys or with an unexpected alt policy will move depth around in ways the exome, which never targeted those regions, will not reproduce. Do not compare a genome depth track with an exome depth track and call the difference biology.

Sample swaps and index hops hurt both designs and are easier to notice in a genome because off-target sex chromosomes, mitochondrial haplogroups, or a contamination estimate have more sequence to work with. An exome can still run those checks inside its target. Run them. A beautiful mean depth of the wrong person is a successful capture.

PCR duplicates in a low-input genome or a heavily amplified exome create the false depth discussed elsewhere. Capture protocols include PCR. The cycle count is part of the kit version's behaviour. Record it.

A human genome or exome is identifiable research data under the approval that permitted the sample, not a clinical report this explainer can issue. Do not return a research call as a medical result. Animal and microbial genomes extracted from material that may be infectious stay under the institutional biosafety decision through extraction and library prep. The assay type does not change the risk group. A larger file is not a more approved file.

Write the kit, the build, and the bytes

The specification between sites should name genome or exome, the capture kit and version if it is an exome, the bait bed or its identifier, the reference build and FASTA, the read length and pairing, and the deliverable. FASTQ alone, or FASTQ plus alignments, is a real fork: an alignment without the build named in the header will be orphaned. Genome files are large. A handoff on a drive that crossed a city, or a transfer interrupted by a power cut, needs a checksum before anyone deletes the source. Do not accept a partial FASTQ because the mean depth of the completed portion already looks like the planned figure. Heat during transport belongs in the same note as any other library: the drive and the tube both have a temperature they were not meant to sit at in a closed vehicle.

What to ask in the enquiry

Ask which design is proposed and why the other would fail the question. Ask for the kit version and the target file for any exome. Ask how structural variants will be reported, or say that they will not. Ask for the duplicate policy and the breadth metric, not only a mean fold.

A whole-genome design can be discussed against the whole-genome sequencing enquiry reference. An orthogonal single-locus check can be discussed against the Sanger DNA sequencing enquiry reference. Library reagents are in the genomics and sequencing catalogue, and the study frame is the genomics research pathway. Put the kit, the build, and the file type in the quote request. Commissioning a sequencing or proteomics study keeps that list in the statement of work, and what to check before a sequencing run is the gate on the library that claims to be one of these designs.

Questions from the bench

Does a whole exome include the whole gene?

It includes the exons the capture kit targeted, often with a short flank into the intron, and any extra loci the vendor added. Introns, most regulatory sequence, and intergenic sequence are not in the design. Two kits sold as exomes target different exon sets. The gene is not the unit that was sequenced. The bait set is.

Why is exome coverage uneven when the mean looks high?

Hybridisation capture prefers some sequences. GC-rich exons, very short exons, and regions with weak baits come down at lower depth than the mean. A high mean can sit on top of holes. Genome sequencing of a sheared library is more even, and it still has holes in repeats and in sequence the reference represents badly.

Which design sees a structural variant?

A rearrangement whose breakpoints sit in introns is invisible to an exome that never captured those introns, except as a change in exon depth that is easy to confuse with capture noise. A whole-genome library can span or flank the breakpoint if the reads and the insert are long enough. Sensitivity still depends on repeats at the junction and on whether the reads are short or long.

Do I need Sanger if I already chose a genome or an exome?

You need an orthogonal check when a single site carries the conclusion. Sanger of a fresh amplicon does that for a high-fraction allele and does it more cheaply than a genome. It does not replace either design for discovery. Choose the genome or the exome for the search, and keep Sanger for the site you will stand on.

References

  1. Ensembl genome browser
  2. UCSC Genome Browser
  3. Illumina overview of next-generation sequencing

Manufacturer names identify published method classes. Trademarks remain with their owners. Catalogue records on this site are independent references for enquiry. They are not a statement of inventory, distribution rights or a supply commitment. This page is educational. It is not medical advice, a diagnostic protocol or a biosafety approval.

Catalogue

Related products and categories

These links follow the subject of the article into published manufacturer references. A listing is a reference for an enquiry, not a statement of stock or distribution rights.