Skip to content
EVRINTH

comparison

Long-read sequencing for structural variants

How to judge whether a long read can span an insertion, deletion, inversion or repeat that short reads only split or fail to cross.

Author
EVRINTH Editorial Team
Published
8 October 2026
Updated
8 October 2026
Reading time
9 min
Benchtop sequencing instrument with a teal status light and a flow cell cartridge in a genomics lab
Benchtop sequencing instrument with a teal status light and a flow cell cartridge in a genomics lab

An insertion of two thousand bases does not fit inside a read of one hundred and fifty. That mismatch is why a laboratory reaches for a long read when the question is a structural variant: an insertion, a deletion, an inversion, or a repeat that short reads never walk through. The chemistry of short-read libraries, adapters and flow cells is the subject of next-generation sequencing from library to reads. This page is the comparison that comes after those reads exist. It is a research explainer. It does not authorise a clinical karyotype or a diagnostic report.

A rearrangement the short read never enters

Structural variants are differences in how sequence is arranged, large enough that a single short read cannot contain them. A deletion removes a stretch. An insertion adds sequence that may be new to the reference. An inversion flips a stretch so that the flanks point the wrong way. A tandem repeat copies a motif until the array is longer than the read, so every short read that starts inside the array looks the same.

Short reads still notice some of these events, indirectly. A read that crosses a breakpoint can align in two pieces. That split read is evidence of a junction, and the bases right at the break can be real. The rest of a long insertion is simply absent from that read. A pair of reads from the two ends of a fragment can also land further apart than the library intended, or point in an unexpected orientation. Those discordant pairs are a hypothesis about the fragment, not a copy of the rearranged sequence.

The decision this supports is practical. If you need the inserted bases, the inverted order, or a repeat array that is longer than your read and your insert, a short-read split is the wrong object to stop on. If you only need to know that a breakpoint exists in a region you will then amplify, the split may be enough to design the next assay.

Alignment sees a break; a span sees the event

Alignment places a read on a reference genome. You should name the build, the way you would name it in Ensembl or the UCSC Genome Browser, because a breakpoint coordinate is meaningless on an unnamed assembly. A short read that matches on both sides of a deletion aligns as two segments with a gap between them. A short read that falls inside a novel insertion often fails to align, or aligns only at one edge.

A long read changes the geometry. Platforms in this class produce reads of kilobases, and often much longer, from a single DNA molecule. Two common families are single-pass nanopore reads and consensus reads built from a circular template. They differ in length distribution and in the kind of errors they make. Both can contain an event that a short read can only punctuate. When the read is longer than the variant, the alignment shows reference, then the variant sequence, then reference again, in one record. You can read the junction instead of inferring it from two anchors.

That span is still an alignment against a reference. Sequence that is missing from the reference appears as an insertion in the read. Sequence that is missing from the sample appears as a deletion relative to the reference. If your reference itself has a collapsed repeat, a perfectly good long read will disagree with it for a boring reason. Look at the reference before you invent a mutation.

Assembly is a different claim from a split alignment

Assembly builds contigs from overlaps between reads, then compares those contigs with the reference. A structural variant can show up as a contig that takes a different path. Assembly is how you recover an insertion that has no home in the reference, and how you try to write a repeat in the order the molecules actually had.

Assembly can also hide the event you wanted. If every read is still shorter than a repeat, the assembler has many equally good overlaps and may collapse them. If coverage is thin, a chimeric molecule can become a false join. A split-read call and an assembly call are allowed to disagree. The honest report says which representation you used, which reference you compared, and which reads support the junction.

Short-read methods remain the right comparison class for counting many samples at a breakpoint you already know. The public overview of that short-read class from Illumina is background for the method, not a pipeline for your structural calls. Choosing length and pairing for those short reads is a separate decision, covered in choosing read length and paired ends.

What you prepare when the question is structure

Long reads are only as long as the DNA you give them. The reagent class that matters is high-molecular-weight DNA, extracted so that the molecules stay long, then a library method matched to the instrument. Shearing is sometimes deliberate, to a size window the instrument prefers, and sometimes the enemy, if the question was a repeat longer than the fragments you just made. Follow the extraction and library instructions for that instrument class. Do not invent a shearing time.

The equipment class is the long-read instrument plus a way to check molecule length before you commit the sample. A short fragment analyser trace that tops out far below the variant you hope to span is a warning. So is a smear of degraded DNA. Coverage still matters. One spanning read is an anecdote. Several independent molecules that agree are the start of a call. Depth and the holes that chemistry leaves in extreme sequence are discussed in GC bias and coverage holes.

Variant classWhat short reads can showWhat a spanning long read adds
DeletionA split read or a pair whose insert looks too largeThe flanking sequence in order, on one molecule
Insertion longer than the readAn alignment that breaks, often without the new basesThe inserted sequence, when the read is longer than the insert
InversionMates in an unexpected orientationThe flipped segment between the junctions
Repeat longer than the readReads that stop at the edge of the arrayA walk through the array, if the molecule is longer than the array

When the control for a span fails

Treat a known structural allele, or a previously characterised cell line, as the positive control when you have one. If those molecules do not span the event you already trust, the new sample's negative result is about the library or the instrument, not about biology. If the control spans and the sample does not, you still ask whether the sample DNA was shorter. A failed length check stops the structural claim. You return to extraction or you change the question to a breakpoint assay that short reads or Sanger sequencing of a junction amplicon can actually finish.

A run that produces long reads from the control and short, fragmented reads from the sample is a sample-handling result. Do not average them into one quality number and move on.

Short read split beside a long read spanning an inversion Reference with an inverted segment inverted Short read: two anchors, no interior split alignment Long read: flanks plus the flipped segment span
A short read splits at an inversion breakpoint, while a long read can carry the flipped segment between the two flanks.

How a false structural call is produced

Repeats create alignments with many legal positions. A short read that is forced into one of them can look like a breakpoint. A long read reduces that ambiguity only when it actually leaves the repeat. If it ends inside the array, you have a longer anecdote of the same problem.

Chimeric molecules, from ligation accidents or from pore-level artefacts, join sequences that were never neighbours in the genome. They look like elegant inversions. Ask whether independent molecules, from an independent library, share the junction. A single spectacular read is a lead.

Reference error is the quiet false call. A missing segment in the assembly becomes an "insertion" in every sample you align, including the ones you thought were wild type. Align a control you trust. If the control carries the same event, you are looking at the reference.

Research use, and the biosafety decision you do not outsource

Human, animal, plant and microbial genomes are all sequenced in research. The structural call describes molecules and a reference. It does not diagnose a person, clear a pathogen, or approve a release. If the DNA comes from a human participant or from an organism your institution treats as a biosafety concern, the approval sits with that institution before the library is made. This article is not that approval.

Deposited reads, when you share them, belong in a public archive such as the European Nucleotide Archive only under the consent and policy you actually have. A structural figure in a slide is not a substitute for the reads.

High-molecular-weight DNA does not enjoy a hot courier

Long-read questions die when the DNA breaks. In a hot season, a shipment that sits on a dock can shear the molecules you needed. Ask the receiving laboratory how they want high-molecular-weight DNA packed, and read dry-ice packing for sample shipments as a shipping practice, not as a promise about length. A power cut that warms a freezer does the same damage more quietly. Record the extraction and the length check so a short library can be traced to handling rather than to a genome that lacked the variant.

What to send when you ask about a structural study

State the organism, the reference build, the variant classes you care about, and the size of the largest event you need to contain in one read. Say whether you need the inserted bases or only a breakpoint. Say how intact the DNA is, if you know it. Consumables for library preparation sit in the genomics and sequencing catalogue. The scientific context for the work is the genomics research pathway.

A genome-wide design can be discussed through the whole-genome sequencing enquiry reference. A long-read design can be discussed through the long-read DNA sequencing enquiry reference. A single junction that you can amplify is often a better fit for the Sanger DNA sequencing enquiry reference. Send the requirement with the quote request. These pages are how a method is specified. They are not a claim that a particular instrument will be run for you on a particular day.

Questions from the bench

Can a long read replace a short-read structural variant call?

It answers a different geometric question. A short read can propose a breakpoint from a split alignment or a discordant pair, while a long read can carry the event in one molecule when the read is longer than the event. Keeping both is reasonable when the short-read call is the screen and the long read is the span you want to inspect.

Does a long read that crosses a repeat prove the copy number?

It proves that this molecule contained a run of that repeat long enough to be read. Copy number in the genome still depends on how many distinct molecules support each length, and on whether the assembler collapsed siblings. Report the span you observed and the coverage behind it.

Is a long-read structural variant a clinical diagnosis?

No. This page is a research comparison of what read length can show. A diagnostic claim needs a validated assay, a reference range, and the legal framework that applies where the sample was taken. A research call is a statement about reads and a named reference.

References

  1. Ensembl genome browser
  2. UCSC Genome Browser
  3. Illumina overview of next-generation sequencing
  4. European Nucleotide Archive

Manufacturer names identify published method classes. Trademarks remain with their owners. Catalogue records on this site are independent references for enquiry. They are not a statement of inventory, distribution rights or a supply commitment. This page is educational. It is not medical advice, a diagnostic protocol or a biosafety approval.

Catalogue

Related products and categories

These links follow the subject of the article into published manufacturer references. A listing is a reference for an enquiry, not a statement of stock or distribution rights.