Skip to content
EVRINTH

selection guide

Choosing read length and paired ends

Choose read length and paired or single ends from the sequence you must place. Longer short reads help repeats only a little; mates rescue mapping.

Author
EVRINTH Editorial Team
Published
8 October 2026
Updated
8 October 2026
Reading time
7 min
Benchtop sequencing instrument with a teal status light and a flow cell cartridge in a genomics lab
Benchtop sequencing instrument with a teal status light and a flow cell cartridge in a genomics lab

A default of one hundred and fifty bases, paired end, is a habit of a kit menu. Read length and pairing are answers to a placement problem, and the problem comes first. What sequence must be uniquely found? Do you need the distance between two ends of a fragment? Is the molecule you care about shorter than a read? The chemistry that turns that choice into a library is next-generation sequencing from library to reads. This page is the selection guide for the geometry. It does not rank instruments, and it does not talk about money.

What extra bases on a short read actually buy

A short read is placed by finding where its sequence matches a reference, or by overlapping it with other reads in an assembly. Extra bases add information only while they keep being unique. In unique sequence, a modest read is often enough, and extra length mostly spends cycles you did not need. In mildly repetitive sequence, a longer read can reach a unique k-mer that a shorter read never includes, so the aligner stops guessing.

The gain saturates. A read of a few hundred bases still cannot span a transposon, a centromeric array, or an insertion of a kilobase. Buying the longest short-read setting on the menu and calling the repeat "solved" is how projects discover the saturation after the run. When the event itself must sit inside one molecule, the comparison shifts to long reads, set out in long-read sequencing for structural variants.

Length also interacts with quality. On many short-read platforms the later cycles are the weaker ones. A nominal long read whose tail is unusable is a shorter read with extra effort in the trimmer. Look at quality by cycle for the instrument class you will use. The public introduction from Illumina describes that class. Other classes publish their own cycle behaviour. Follow those documents rather than a remembered default.

What the mate buys

Paired-end sequencing reads both ends of a fragment. The middle is not read. It is implied by the fact that the two ends belonged to one molecule. That implication is useful in three ways.

First, rescue. One end falls in a repeat and could belong in twenty places. The other end falls in unique sequence. The pair has one plausible address, provided the insert size is about what the library claimed. Second, structure. If the two ends land much further apart than the library's insert, or point the wrong way, the fragment disagrees with the reference. That disagreement is a hypothesis about a deletion, an insertion or an inversion. It is not the sequence of the event, unless the event is small enough to sit in the gap you can still infer. Third, duplicates and PCR copies. Pairs give a clearer signature of the fragment they came from than a single read does, which matters when you later decide what "one molecule" means.

The insert size is a library choice, distinct from read length. A long insert with short reads is a different experiment from a short insert with long reads that overlap in the middle. Overlapping pairs are wonderful for a small amplicon and wasteful if you thought you were spanning a structural variant. State both numbers.

When a single end is the honest design

Small RNA is the clearest case. The fragment is short, often around the length of the read you would have chosen anyway, and a second end may be adapter rather than biology. Counting applications can be similar when the feature you count is unique and you have no use for insert size. Some older chromatin assays were designed as single-end because the question was occupancy at a site, not the fragment's other end. Modern designs often pair them. The point is not nostalgia. The point is to notice when the mate would be empty information.

Single-end is a poor default for a new genome assembly, for structural variants, or for any placement inside recent duplications. If you are unsure, write the question in one line. If the line does not mention distance, orientation, or a repeat, single-end may be enough. If it mentions any of those, pair the reads.

Amplicon sequencing has its own geometry. If the amplicon is shorter than a read, a long read setting sequences into the adapter. Trim will remove it, and you will have paid attention for bases you discard. Match read length to amplicon length. That design conversation can be held through the amplicon sequencing enquiry reference.

Sanger length is a different geometry

A Sanger read is one reaction on a population of copies of one template, often several hundred bases of useful trace, sometimes more when the template is kind. It is not a paired-end short read, and it is not a long read from a single native molecule in the nanopore or consensus-long-read sense. It is the right geometry for confirming a plasmid junction, a clone, or a single amplicon where you want to see the trace. Choosing it because "longer is better" muddles the methods. Choosing it because you have one locus and you want a chromatogram is coherent. The questions that separate the methods are in questions for a Sanger versus NGS decision.

Question you actually havePairing that serves itWhat length is doing
Count features that are already uniqueSingle-end is often enoughOnly enough to be unique
Place a read that starts in a short repeatPaired ends, so the mate can anchorLonger reads help until the repeat exceeds them
Infer a deletion or an inversion from distancePaired ends with a known insertRead length does not contain the event
Confirm one plasmid junctionUsually one Sanger reactionA trace across the junction, not a short-read default
Contain a repeat or a long insertionA long-read classShort-read length will not arrive there

Branch before you accept the menu default

Sketch the locus, or name the accession in GenBank if the sequence is public. Mark the repeats that are longer than the read you are about to order. If your question sits inside one of those repeats, stop and change class. If your amplicon is shorter than the read, shorten the read or accept that trimming will eat the tail. If you need insert-size evidence, specify the insert window, not only the cycle kit. If a collaborator says "use what we always run", ask them to point at the question that made "always" correct. Methods people have written down on protocols.io sometimes record that question. A protocol that omits it is a recipe, not a reason.

Failed geometry looks like successful chemistry. The lane clusters. The qualities pass. The alignments are simply ambiguous, or the mates overlap when you needed them far apart. QC will not flag "wrong design".

Paired reads on one insert versus a single end One fragment read 1 read 2 insert, mostly unread Single end in a repeat no mate, many legal addresses
Paired ends read two flanks of one insert and leave the middle implied; a single end has no mate to rescue a repeat.

Research use of a geometry

The choice of length and pairing does not make an assay clinical, and it does not approve sequencing of a regulated organism. It is a research design choice. A human sample still needs the approval you already hold. A pathogen genome still needs the containment your institution assigned. State those boundaries in the same document as the read length so a receiving laboratory sees one specification.

Write the geometry down so another site cannot invent a usual

Facilities in different cities accumulate different habits. "Paired end" without a length, or a length without an insert, will be filled in locally. Put both in the written specification, next to the biological question. Heat and shipping do not change the correct geometry, but a degraded sample can make a long-insert library impossible. If the DNA is already fragmented, a design that needed long inserts has to be redrawn. Say so before the kit is opened. The statement of work is the document that should carry the numbers.

What to send with the request

Send the question, the reference or accession, the read length, paired or single, the insert window, and the reason a shorter or unpaired design would fail. Note if the template is one plasmid that should be a Sanger trace instead. Reagent classes live in the genomics and sequencing catalogue. The setting is the genomics research pathway.

Discuss a genome design through the whole-genome sequencing enquiry reference, an amplicon design through the amplicon sequencing enquiry reference, and a single-template trace through the Sanger DNA sequencing enquiry reference. Attach the geometry to the quote request. These pages are how the method is discussed. A menu default is not a design.

Questions from the bench

When is a single end enough?

When the biological sequence is already unique enough to place with one read, and you do not need insert size. Small-RNA sequencing is the classic case, because the fragment itself is short. Some counting assays are in the same position if the transcript or the gene you count is unambiguous without a mate.

Will a longer short read solve a repetitive region?

It helps until the repeat is longer than the read. Moving from a short single end to a longer single end recovers placements that were ambiguous at the shorter length. A repeat of kilobases remains ambiguous. Paired ends add the mate's address, and a long-read method is the class that can contain the repeat itself.

How do paired ends help if each read is still short?

The two reads come from the two ends of one fragment, so the distance and the orientation between them are data. A read that lands in a repeat can still be placed if its mate lands in unique sequence. The pair also reveals insertions, deletions and inversions as unexpected distances or orientations, without containing the whole event.

References

  1. Illumina overview of next-generation sequencing
  2. NCBI GenBank
  3. protocols.io method repository

Manufacturer names identify published method classes. Trademarks remain with their owners. Catalogue records on this site are independent references for enquiry. They are not a statement of inventory, distribution rights or a supply commitment. This page is educational. It is not medical advice, a diagnostic protocol or a biosafety approval.

Catalogue

Related products and categories

These links follow the subject of the article into published manufacturer references. A listing is a reference for an enquiry, not a statement of stock or distribution rights.