Skip to content
EVRINTH

application

Variant calling is a model not a fact

Treat each VCF row as a model after alignment, and decide when a germline or somatic call still needs an orthogonal Sanger check.

Author
EVRINTH Editorial Team
Published
8 October 2026
Updated
8 October 2026
Reading time
9 min
Benchtop sequencing instrument with a teal status light and a flow cell cartridge in a genomics lab
Benchtop sequencing instrument with a teal status light and a flow cell cartridge in a genomics lab

Variant calling is a model not a fact. Reads are observations. An aligner proposes where they sit. A caller proposes which differences from the reference are likelier than noise, under a prior you chose. The VCF file is the list of those proposals. This page is how to use that list without promoting a row into a fact the experiment did not earn. The reads themselves start in next-generation sequencing from library to reads. The reference they are compared with has to be named, which is reference genomes and why the build matters.

Nothing in the caller creates information the library never contained. A systematic PCR error, a hop into the wrong sample, or a duplicate family will be modelled faithfully. Faithful modelling of an artefact is still an artefact. Library preparation and the independence of reads, in coverage depth is not the same as accuracy, are upstream of any genotype.

Alignment comes first, and it is already a model

A short-read aligner places each read on the reference, or refuses to. Placement in a repeat is uncertain even when the bases are sharp. Mapping quality records that uncertainty. Soft clips record ends that did not match. Those fields enter the caller if you let them. If you strip them, you have thrown away the aligner's doubt and kept its guess.

The reference allele at a site is whatever the FASTA says. A difference is only a difference from that FASTA. Callers do not know which allele is ancestral, common, or functional. Annotation, added later, borrows a gene model and a transcript. Change the transcript and a coding label can change while the genomic alleles stay put. Keep the genomic description as the durable one, and treat the amino-acid label as a view.

Long reads use different aligners and different error models. A caller tuned for short-read substitutions will scold a long-read alignment for indels the chemistry tends to make, or it will miss the structure the long read was for. Match the model class to the read class in how short reads and long reads differ.

Germline settings and somatic settings

A germline model for a diploid organism scores genotypes such as homozygous reference, heterozygous, and homozygous alternate. It expects roughly balanced alleles at a heterozygote, with scatter, and it may use a population prior so that a known common allele needs less new evidence than a never-seen one. That prior is a scientific choice. It helps in humans when the prior matches the ancestry the panel actually represents, and it harms a call when the sample comes from a population the panel barely contains. Say which prior you used, or say that you used none.

A somatic model answers a different question: which alleles are present at fractions a diploid germline model would call noise, and which of those are absent from a matched normal. Tumour purity and subclones mean a real acquired allele can be a small minority of reads. The matched normal is what lets the model avoid reporting the person's germline as if the tumour had acquired it. Tumour-only calling is possible and has a higher burden of germline leakage and of artefacts that a normal would have vetoed. If you call tumour-only, the VCF is a hypothesis with that limitation written on it.

Do not run one setting and describe the other. A germline call set on a mixed infection, a pooled sample, or a tumour is the wrong model even when the VCF is full of neatly scored rows. Ploidy that is not diploid, including cell lines with unstable karyotypes, breaks the heterozygote expectation. The model will still emit genotypes. They will be the nearest diploid story, which may be fiction.

A VCF row is a hypothesis with columns

The durable columns are the chromosome, the position, the reference allele, and the alternate allele, plus the build you must store beside the file because the VCF header is only as good as the person who wrote it. QUAL is the caller's confidence, defined by that caller. FILTER is the decision of a filter set: PASS or a list of reasons. INFO and FORMAT carry depth, allele counts, genotype likelihoods, and whatever else the tool emits. A genotype of 0/1 is the model's best diploid label. It is not a photograph of two chromosomes.

Hard filters cut on thresholds: minimum depth, minimum allele count, strand bias, low complexity. They are understandable and crude. Learning filters trained on a set of known variants are a different class. They require a training set that resembles your data. A filter trained on one platform and applied to another punishes the second platform's normal error mode. Either way, a filtered-out row can be true, and a PASS row can be false. Filters change the false-positive and false-negative balance. They do not step outside probability.

Multiple alternate alleles, indels near repeats, and clusters of calls in a few bases are where models strain. A complex haplotype can be represented as several nearby rows or as one haplotype, depending on the caller. Comparing two VCFs as if each row were an independent fact will double-count one event or call a discordance that is only a representation difference. Normalise representation before you compare, and still look at the reads when a discordance matters.

Column or flagA disciplined readingAn over-reading
CHROM and POSAn address on the build in the headerAn address on whichever build the reader assumes
REF and ALTThe alleles the model comparedA functional or clinical label
QUALThis caller's confidenceA universal probability you can threshold forever
FILTER PASSSurvived this filter setConfirmed by an orthogonal method
Genotype 0/1Best diploid label under the priorProof the cells are heterozygous and clonal
Allele depthReads assigned to each allele after the alignerIndependent molecules, unless duplicates were handled
Somatic flagThe model used a somatic modeThe allele is acquired, causal, or diagnostic
From alignments to a hypothesis and an orthogonal check Align reads Statistical model VCF row Filters Sanger if it can see The row remains a hypothesis. Orthogonal peaks are a second experiment, not a new column.
Reads are aligned, a statistical model writes a VCF hypothesis, filters score it, and a claim that matters can still go to an orthogonal Sanger trace.

Where the model should be challenged

Challenge a call that decides a conclusion: a genotype you will build a figure on, a site you will edit, a difference you will treat as the result of a treatment. For a high-fraction germline allele, the orthogonal check is Sanger sequencing of a new amplicon. Design primers on the same build, outside the bases you are judging, clean the product, and read peaks as Sanger sequencing for a single amplicon describes. Agreement supports the hypothesis. Mixed peaks mean the sample or the primer is not the single template you assumed. They do not mean you should pick the peak that matches the VCF.

For a low-fraction somatic hypothesis, that same Sanger reaction is the wrong court. It reports the majority molecules. Absence of a minor peak is expected even when the minor allele is real. A second independent library, or a method class built for small fractions, is the check that can actually fail. Do not cite a negative Sanger as confirmation that the somatic model was wrong, and do not cite it as confirmation that the allele is real.

Also challenge calls that live only at read ends, only on one strand, only inside a duplicate family, or only in a region of low mapping quality. The model may have down-weighted them already. If it did not, your filter is softer than you think. Homopolymers and short tandem repeats produce indel chatter. A one-base indel in a long homopolymer needs the chemistry's error mode in the interpretation, not only the genotype field.

Reference bias belongs on the list. Reads carrying the alternate allele can map less readily than reference reads, so the allele fraction shrinks. A heterozygous site can look like a low-fraction artefact, and a real low fraction can sink below the filter. More coverage of the same biased library repeats the bias.

Not a diagnosis, and not a ruling on the organism

A research VCF is not a medical interpretation. It does not assign a condition, a carrier status, or a treatment. Human reads stay inside the consent and the review that produced them. If the organism in the tube may be infectious, the institutional biosafety decision governs the work up to and including the extract. The VCF does not change that decision after the fact, and this page does not issue one. Resist the slide from "the model prefers an alternate allele" to "the patient has" or "the isolate is resistant". Those sentences need assays and authorities this file does not contain.

Tie the file to the tube in the specification

Write, in the study specification, the caller name and version, germline or somatic settings, the matched-normal policy, the filter set, the build, and which sites will receive an orthogonal check. A collaboration that receives only a spreadsheet of gene names has lost the model. Sample identity travels with the file: the tube label, the index, and the VCF sample column are one chain. A power cut that forces a re-run produces a second library, not a second half of the same likelihood. Do not merge the VCFs to increase depth without saying they are replicates. Heat and delayed handoff matter because a drive of BAMs without the header that names the FASTA will be reanalysed on someone else's default genome, and the coordinates will drift exactly as the reference page warns.

What to ask before anyone calls

Ask which model will be run, what the deliverable VCF is filtered by, and whether unfiltered records are retained. Ask how duplicates are marked before calling. Ask which claims include a Sanger check or a second library, and which claims will be left as model output. The genomics research pathway frames that design. Reagent classes for any new amplicon are in the genomics and sequencing catalogue.

A call set on a whole genome can be discussed against the whole-genome sequencing enquiry reference. The orthogonal chromatogram can be discussed against the Sanger DNA sequencing enquiry reference. Put the model, the build, and the check in the quote request. What to check before a sequencing run still applies: a caller cannot repair a library that should not have been loaded. Commissioning a sequencing or proteomics study is the place the model name enters the statement of work.

Questions from the bench

Is a PASS value in a VCF a statement of biological truth?

PASS means the record survived the filters this pipeline applied. Those filters encode assumptions about depth, strand, and allele fraction. A second pipeline can fail the same site, or pass a site this one rejected. Report the caller, the version, and the filter, and keep PASS in that scope.

Why can the same alignments yield different calls under germline and somatic settings?

Germline settings expect a ploidy, often diploid, and allele fractions near zero, one half, or one. Somatic settings allow lower fractions and, when a matched normal exists, subtract the germline. Running a germline model on a tumour mixture will miss low-fraction alleles and will happily call germline sites as if they were the tumour's news. The setting is part of the hypothesis.

When is Sanger the right orthogonal check?

When the claim is a high-fraction allele that a fresh PCR can show as peaks, Sanger is a direct second method. Sequence a new amplicon, ideally with primers outside any earlier amplicon, and read the chromatogram. When the claim is a low-fraction somatic allele, Sanger often cannot see it. A clean single peak does not refute that claim. Use a second library or a method that can see small fractions.

Should I call variants from a pileup by eye?

A pileup is a useful way to notice a site the model might be mishandling. It is not a caller. Your eye does not apply the mapping-quality filter, the duplicate policy, or the allele-fraction prior in a way a second person can rerun. If the eye and the VCF disagree, that disagreement is the start of a check, not a licence to edit the VCF by hand and keep the model's score.

References

  1. Ensembl genome browser
  2. UCSC Genome Browser
  3. NCBI Sequence Read Archive

Manufacturer names identify published method classes. Trademarks remain with their owners. Catalogue records on this site are independent references for enquiry. They are not a statement of inventory, distribution rights or a supply commitment. This page is educational. It is not medical advice, a diagnostic protocol or a biosafety approval.

Catalogue

Related products and categories

These links follow the subject of the article into published manufacturer references. A listing is a reference for an enquiry, not a statement of stock or distribution rights.