glossary
Reference genomes and why the build matters
Keep the genome build with every coordinate, because GRCh38, hg19, patches, alt contigs and decoys do not share one numbering.
- Author
- EVRINTH Editorial Team
- Published
- 8 October 2026
- Updated
- 8 October 2026
- Reading time
- 9 min

A reference genome is a shared coordinate system, not the genome of the person or the strain in your tube. Builds are successive editions of that system. The same physical base can carry a different number in each edition, and a variant written as a number without the edition is not a variant anyone else can find. This glossary is the vocabulary that keeps those editions apart. Aligning reads to whatever FASTA was on the server is the step that comes after next-generation sequencing from library to reads. Calling differences against it is variant calling is a model not a fact.
Human examples dominate the argument because the public builds are named and easy to mix up. The rule is the same for a bacterial strain, a crop, or a plasmid: name the accession and the version. GenBank and Ensembl are the public places those versions live. The UCSC Genome Browser is where many people look at the human ones and where the hg nicknames come from.
Builds, and why a number moves
GRCh38 is the Genome Reference Consortium human build 38. UCSC calls that generation hg38. GRCh37 is the previous consortium build. UCSC's hg19 is the browser view most people mean by the older human coordinates, with the caveat in the questions above that hg19 and GRCh37 are close and not identical. A coordinate copied from a 2013 supplementary table is usually an hg19-era number. The same base on GRCh38 is often a different integer, because gaps were filled, sequence was corrected, and centromeric and segmental regions were rebuilt. Downstream of any insertion in the reference, every coordinate shifts.
That is the whole content of the warning. Chromosome, position, reference allele, and alternate allele are an address on a named map. Change the map and the address points at a different house, or at a house that was not built yet. A liftover tool applies a chain file to translate intervals. The translation can succeed, fail, or hit more than one place. A successful liftover is a new record that must cite the chain. It is not evidence that the original authors aligned to the new build. If the chain cannot place an interval, the honest result is unmapped, not a nearest guess you type by hand.
Chromosome names are part of the build's dialect. Some FASTA files call the first chromosome 1, others chr1. Mitochondrial sequence may be MT, chrM, or a separate accession. A pipeline that maps with one dialect and annotates with the other will report no overlap and a calm empty table. That empty table is a naming bug. It will be described as a negative scientific result if nobody checks the contig names.
Patches, alt contigs, and decoys
Patch releases update GRCh38 in place. A paper that says GRCh38 and a FASTA that is GRCh38.p14 are only the same reference if you say so. Novel patch sequence is invisible to an aligner whose FASTA lacks it. Reads from that sequence become unmapped, or they stick to a similar primary region and invent differences. When you record the build, record the patch and the file name, not only the two letters and two digits everyone remembers.
Alt contigs represent haplotypes that differ enough from the primary path that a single linear sequence would hide them. The major histocompatibility complex is the example people repeat because the differences are large. Including alts in the FASTA without a matching analysis policy causes two symmetric failures. Reads from a person who matches the alt may align there and vanish from the primary chromosome, so a primary-only variant call looks like a deletion or a no-call. Reads that align to both, if the caller does not handle the pair, can be counted twice. The remedy is not a slogan for or against alts. It is to use a documented analysis set and a caller that knows that set. Write down which choice you made.
Decoy sequences are sink contigs. The decoy set commonly paired with GRCh38 gives reads from unplaced human sequence and from some non-human material, including a viral decoy, a place to align that is not a protein-coding gene they faintly resemble. Without decoys, those reads are forced onto the primary assembly or left unmapped in a way that still confuses depth. With decoys, you must remember that a read on a decoy is not a variant in the gene it almost matched. Depth plots that ignore decoy placements will under-call contamination and over-call similarity.
Analysis sets that bundle primary chromosomes, alts, decoys, and a viral contig are a defined object. "The human genome" is not. If two collaborators both used GRCh38 and one included the decoy set, their unmapped fractions are not comparable. Coverage depth is not the same as accuracy starts from that non-comparability: a mean depth on different FASTA files is two means.
How to write the map down
Record, in the notebook and in the methods, all of the following when the organism is human: the build nickname and the consortium name, the patch, whether alts were included, whether decoys were included, the FASTA file name or a checksum, the chromosome-naming dialect, and the gene annotation release used to name transcripts. Ensembl and UCSC annotation releases are not interchangeable just because the underlying build matches. A transcript identifier with a version is part of the address of a coding change.
For other organisms, record the accession and the assembly version. A genus name is not a build. Two strains of the same species can differ by rearrangements that make a coordinate from one FASTA actively misleading on the other. Plasmid maps have versions too. A Sanger trace compared with the wrong map will "find" the engineering you already did, or miss it. That comparison is only as good as the map, which Sanger sequencing for a single amplicon assumes you named.
A database identifier such as an rs number is not a substitute for the build. Identifiers are merged, split, and retired. The alleles at that identifier still need a genomic address if a second person is to see them in an alignment.
| Phrase in a report | What it fixes | What it leaves open |
|---|---|---|
| hg19 or GRCh37, said carefully | The older human coordinate generation | Which exact FASTA, and any extra contigs |
| GRCh38 | The current primary generation people mean | Patch level, alts, decoys, chr prefix |
| A lifted coordinate | An interval translated by a named chain | Whether the chain failed nearby or split the locus |
| Gene symbol plus amino-acid change | A human-readable label | Transcript version and genomic build |
| Strain accession | The organism's own map | Whether the reads were actually aligned to it |
| rs identifier | A database cross-reference | Retirement, merges, and the build of the alleles |
When the translation itself is the error
Liftover across a region that was rebuilt, not merely shifted, can map a variant onto a related but different sequence. The alleles may no longer match the new reference base. A caller, or a person, who then "corrects" the allele to fit the new reference has created a third variant that nobody sequenced. If the reference allele in the old call does not match the old FASTA, the old call was already broken. Check alleles against the FASTA of the stated build before you lift anything.
Mixing annotation from one build with alignments from another produces confident-looking genes on the wrong exons. The track and the reads must share a coordinate system. Browser sessions that overlay hg19 variants on an hg38 view without a lift are a picture of that mistake.
A Sanger result compared with a primer designed on a different build can fail to sit down, or sit down in the wrong place, because the primer was never a sequence in the genome you amplified. Primer design tools need the same build you will report. NCBI Primer-BLAST will follow the reference you select. Selecting is the step.
Sequences from people, and sequences from pathogens
A human reference build is not a consent document. Aligning a person's reads to GRCh38 is still human data under the approval that allowed the sample. A pathogen reference is not a biosafety approval. Culture, extraction, and sequencing of an organism that may be infectious follow the institutional decision, with biosafety basics for research benches as the bench-level frame. This glossary does not identify an isolate for clinical action. Naming the build makes a research coordinate reproducible. It does not make it a diagnosis.
Name the FASTA where the collaboration is written
When two laboratories share the analysis, a loose specification is what lets them analyse the human genome on different files. Put the FASTA name, the patch, alt and decoy choices, and the chromosome dialect in the study specification before either site aligns. A coordinate list emailed later, without that header, cannot be merged safely with the other site's table. If files move on a drive because the network is the unreliable part, the checksum of the FASTA travels with the drive. Two files that share a nickname and differ by a patch will both open in a browser and will not agree on a variant.
What the enquiry has to pin down
State the build, the patch or accession, and whether the deliverable is coordinates on that build only. Say if liftover to an older build is required for a historical comparison, and who will run it. Ask which annotation release will name the genes. A request that says human reference will be answered with someone's default.
Whole-genome alignments depend on this choice and can be discussed against the whole-genome sequencing enquiry reference. A one-locus Sanger comparison can be discussed against the Sanger DNA sequencing enquiry reference. The surrounding design is the genomics research pathway, and reagent classes are in the genomics and sequencing catalogue. Put the FASTA identity in the quote request. Commissioning a sequencing or proteomics study is where that identity belongs in the statement of work, next to the biological question rather than after the VCF has already been emailed.
Questions from the bench
Is hg19 the same file as GRCh37?
They are the same generation of the human reference and they are not byte-identical. UCSC's hg19 and the Genome Reference Consortium's GRCh37 differ in a handful of sequences, including how some extra contigs and the mitochondrial sequence have been handled. Write the exact FASTA name you aligned to. Saying hg19 in a paper and GRCh37 in a table is how a later liftover uses the wrong chain.
What is a patch on GRCh38?
A patch release, such as a numbered GRCh38 patch, adds or corrects sequence without pretending the whole primary assembly was rebuilt from nothing. Two pipelines that both say GRCh38 can still be on different patches. Reads that belong on patch sequence will look unmapped or mismapped if the FASTA omitted the patch. Record the patch level with the build.
Why do alt contigs and decoys change variant calls?
Alt contigs carry alternate haplotypes, so a read can place on the primary chromosome or on the alt. If the aligner is allowed to use alts and the variant caller is not prepared for that, alleles can disappear or be counted twice. Decoy sequences give reads from unplaced or non-human material, such as a viral decoy used with human builds, somewhere to land that is not a human gene. Leaving them out pushes those reads onto the nearest human lookalike.
Can I report a gene variant as the gene name alone?
The gene name does not fix the coordinate system. Transcripts have versions, exons move between annotation releases, and a coding change needs the transcript identifier as well as the genomic build. A sentence with a gene symbol and a bare amino-acid change is not yet a coordinate a second laboratory can find.
References
Manufacturer names identify published method classes. Trademarks remain with their owners. Catalogue records on this site are independent references for enquiry. They are not a statement of inventory, distribution rights or a supply commitment. This page is educational. It is not medical advice, a diagnostic protocol or a biosafety approval.
Catalogue
Related products and categories
These links follow the subject of the article into published manufacturer references. A listing is a reference for an enquiry, not a statement of stock or distribution rights.
Continue in this cluster
Related reading
Next-generation sequencing from library to readsHow a sequencing library becomes reads: adapters, flow cells, quality scores and the checks that stop a bad library from wasting a run.
16S profiling and its taxonomic limitsJudge a 16S profile at the rank the marker supports, and see where copy number, primer bias and species names stop being honest.
A glossary of sequencing termsWorking definitions of read, coverage, depth, MAPQ, Phred, VCF, BAM and the related words, each tied to the mistake that word prevents.