Skip to content
EVRINTH

explainer

How short reads and long reads differ

Choose short paired reads or long reads by the repeat, structural variant or homopolymer the experiment actually has to resolve.

Author
EVRINTH Editorial Team
Published
8 October 2026
Updated
8 October 2026
Reading time
9 min
Benchtop sequencing instrument with a teal status light and a flow cell cartridge in a genomics lab
Benchtop sequencing instrument with a teal status light and a flow cell cartridge in a genomics lab

Read length is a choice about which structures you are willing to see in one piece. Short paired reads and long reads are method classes, not two settings of one instrument. This page is how to pick between them before a library is made. The shared path into a read file is next-generation sequencing from library to reads. A single plasmid or a single PCR product still belongs on a capillary, which is a third class and is set out in Sanger sequencing for a single amplicon.

The practical fork is simple to say and easy to fake. If the feature you care about is longer than the read, and longer than the fragment the two ends came from, a short-read file will describe that feature by absence, by broken pairs, or by a drop in uniquely mapped coverage. A long read can include the feature and the unique sequence on both sides. If the feature is a single base in ordinary unique sequence, the extra length buys less than people hope, and the error mode of the chemistry matters more than the brochure's maximum read.

Two chemistries, two ways of noticing a base

Illumina-style short reads are sequencing by synthesis on a flow cell. Many copies of each fragment form a cluster. Reversible terminators add one base per cycle, a dye is imaged, and the block is removed so the next cycle can proceed. The public account of that method class is the Illumina technology overview. Read lengths in common use sit in the tens to a few hundred bases. A paired-end run sequences both ends of the same fragment. The unsequenced middle is an unknown stretch whose length you chose at library size selection. The pair tells you the two ends belonged together. It does not tell you the sequence between them.

Oxford Nanopore-style long reads watch a single strand pass through a pore and infer bases from the electrical signal. The molecule can be as long as the DNA you managed to keep intact. PacBio-style long reads watch a polymerase in a small well. One version of that chemistry circles a template and builds a consensus from many passes over an insert of modest length, often discussed as highly accurate long reads. Another version keeps a continuous observation of a longer molecule. Those are different products. Calling all of them long reads is fair at the level of length and unfair at the level of accuracy profile.

Do not invent an error rate for any of them in a protocol or a paper. Platform chemistry dominates the error mode, and the current chemistry note is the manufacturer's. What you can say without a borrowed percentage is the shape of the mistakes each class has been known for. Short reversible-terminator reads have tended toward substitution errors, with quality falling as cycles accumulate and clusters drift out of phase. Nanopore signal has been particularly awkward in homopolymers, where a run of the same base changes the current by a small step, and the characteristic mistakes have included insertions and deletions. PacBio's raw continuous reads and its consensus reads do not share one accuracy profile. Consensus of many subreads answers a different question from a single pass. Quote the product you loaded.

Libraries and rooms are not interchangeable

A short-read library is fragmented, sized, and fitted with adapters the flow cell can hold. The fragment length is a design variable: short inserts can make the two reads overlap, which helps a small amplicon and hurts a structural question; longer inserts give the pair more span and can cluster less predictably if you exceed what that flow cell likes. The prep decisions are in library prep is where most runs are won.

A long-read library is a different discipline. The enemy is a break. Pipette shear, a harsh extraction, or a cleanup that selects short molecules will quietly turn a long-read plan into an expensive short-read plan. High-molecular-weight DNA extraction, gentle mixing, and a size selection that discards the small fragments are the reagent and handling classes that matter. The instrument will sequence whatever length you actually gave it.

Neither library is a Sanger reaction. You do not add one primer and read one chromatogram. You add adapters and accept a population of molecules. Catalogue classes for extraction, cleanup, and the plastics those preps consume are under genomics and sequencing.

Where the scientific fork actually is

Write the feature down before you pick the instrument.

A single-nucleotide difference in unique sequence, counted across many samples, is a short-read strength. The reads are numerous, the error mode is well modelled, and the mapping is stable when the genome has a place for them. Coverage here means how many independent reads land, which is a separate argument from accuracy and is taken up in coverage depth is not the same as accuracy.

A tandem repeat longer than the read, and longer than the paired fragment, cannot be placed. Every read that sits inside the repeat maps to many coordinates or is given a poor mapping score. You can notice that the region is difficult. You cannot count the repeat units from those reads. A molecule that starts in unique sequence, crosses the repeat, and ends in unique sequence carries the count in one string.

A structural variant is the same geometry at a larger scale. A deletion of a few kilobases may show up in short reads as pairs that map further apart than the library expected, or as reads that split across the junction. That inference can be right, and it can be confused by repeats at the breakpoint. A long read that contains the junction and both flanks is a direct observation of the rearrangement. Inversions and insertions of sequence that is not in the reference follow the same logic.

Homopolymers trouble more than one class. Short-read indel calling slips on long runs of A or T. Nanopore-style signal has its own homopolymer difficulty. If the claim is the exact length of a homopolymer, say so in the design and pick the chemistry whose current note addresses that mode. Do not assume length alone solves it.

Branch when the sample cannot yield long DNA. Formalin-damaged tissue, ancient fragments, and heavily sheared extracts will not become long reads because the instrument is capable of them. Sequence what the molecules are. Pretending otherwise produces a short library with long-read adapters and a confused specification.

Question you will write in the paperShort paired readsLong reads
A base change in unique sequence, many samplesDirect, and usually the economical designPossible, and often unnecessary
Number of copies inside a long tandem repeatUnplaced reads and an unresolved countA spanning molecule can carry the count
A rearrangement of several kilobasesIndirect signals: discordant pairs, split reads, depthA read can contain the junction
Exact length of a homopolymerIndel calling is fragileDepends on the chemistry; read the current note
One cleaned amplicon under one kilobaseA heavy instrument for a small questionStill the wrong witness if a chromatogram would do
Short pairs versus a spanning long read Same locus, two method classes Unique Repeat longer than a short read Unique Paired short reads sit inside the repeat One long read spans unique, repeat, unique Length helps only when the molecule crosses the feature. Chemistry still sets the error mode.
Short paired reads fall inside a long repeat and cannot place it, while one long read that reaches unique sequence on both sides carries the span.

What people misread in the file

A high mapping score on a short read proves that this read preferred one place in the reference you supplied. It does not prove the read would be unique in a more complete reference, and it does not repair a homopolymer slip. Long reads with a characteristic indel mode can look messy in a short-read viewer that was tuned to substitutions. Switching viewer settings is not a new chemistry.

Assembly contigs that break at the same repeat in every short-read library are describing the library, not a biological truncation. Feeding those contigs to a gene caller and reporting a truncated gene is how a methods limit becomes a false discovery. Deposit the reads where the record can show the platform, for example the Sequence Read Archive or the European Nucleotide Archive, and name the chemistry in the methods so a later reader does not treat a 2016 error mode as yours.

Coverage comparisons across platforms are easy to fake. A long read covers more reference bases per read and fewer independent molecules per gigabase of sequence when the molecules are very long. Depth for a small variant and span for a rearrangement are different currencies. Write the currency you spent.

Research use, and material that may be infectious

Choosing a platform does not reclassify the organism. High-molecular-weight extraction of a pathogen is still work with that pathogen until the institutional biosafety decision says the lysate is safe to move. This explainer is not that decision and not a diagnostic claim about anyone's genome. Human samples stay inside the ethics approval that allowed the tube to exist.

A long run and the mains

Nanopore-style runs are often left for many hours. A power cut in that window is part of the experiment's history, not a footnote the basecaller will ignore. Ask, before the run, whether the sequencer and the disk that receives the signal are on a supply that survives an outage, and whether a partial signal file is analysable under the manufacturer's resume rules. A laptop battery is not a data plan. Short-read flow cells also stop when the room loses power, and a partial cycle set is a different library of signal from a finished one. Record the outage. Do not pool the fragment with a later successful run and call the quality scores comparable. Heat belongs in the same note: an instrument room that climbs through the afternoon can leave the manufacturer's temperature range, and the chemistry note, not a local habit, sets that range.

What the enquiry has to name

Name the structure you need to resolve, the sample's likely fragment length, whether a short paired design can span it, and which long-read product you mean if you mean one. Say whether the deliverable is reads, an assembly, or both. A sentence that says only long-read sequencing has not chosen between a consensus chemistry and a single-pass chemistry.

A whole-genome design that depends on this choice can be discussed against the whole-genome sequencing enquiry reference. A one-amplicon confirmation can be discussed against the Sanger DNA sequencing enquiry reference. Put the platform, the insert size, and the feature you must span in the quote request. The surrounding study design sits on the genomics research pathway. How to write that brief so two sites do not assume different products is part of commissioning a sequencing or proteomics study.

Questions from the bench

Are long reads simply short reads joined together?

No. A long read is one molecule observed for a long stretch by a different chemistry. Paired short reads are two observations from the ends of a fragment, with a gap in the middle whenever the fragment is longer than the two reads. Joining short reads in software is assembly, and it fails where the repeats are longer than the fragments.

Which error rate should a methods section quote?

Quote the chemistry and the basecaller version the manufacturer documents for the run you did. Platform chemistry dominates the error mode, and that note moves when the chemistry moves. A number copied from an older paper describes that paper's reagents. It is not a property of the read length itself.

When is a short read the better witness?

When the question is a small variant in unique sequence, or a count of many molecules, short paired reads are a strong and familiar design. They place badly inside long repeats and they infer large rearrangements from indirect signals. Use the read that spans the structure you intend to claim.

Does a long read replace a Sanger check of one exon?

It replaces the need to assemble that exon from tiny pieces. It does not replace a dedicated chromatogram when the claim is the exact base sequence of one cleaned amplicon. Sanger remains a single-primer reading, with its own length limit, described in the single-amplicon guide.

References

  1. Illumina overview of next-generation sequencing
  2. NCBI Sequence Read Archive
  3. European Nucleotide Archive

Manufacturer names identify published method classes. Trademarks remain with their owners. Catalogue records on this site are independent references for enquiry. They are not a statement of inventory, distribution rights or a supply commitment. This page is educational. It is not medical advice, a diagnostic protocol or a biosafety approval.

Catalogue

Related products and categories

These links follow the subject of the article into published manufacturer references. A listing is a reference for an enquiry, not a statement of stock or distribution rights.