Skip to content
EVRINTH

Pillar guide

Next-generation sequencing from library to reads

How a sequencing library becomes reads: adapters, flow cells, quality scores and the checks that stop a bad library from wasting a run.

Author
EVRINTH Editorial Team
Published
8 October 2026
Updated
8 October 2026
Reading time
8 min
Benchtop sequencing instrument with a teal status light and a flow cell cartridge in a genomics lab
Benchtop sequencing instrument with a teal status light and a flow cell cartridge in a genomics lab

Next-generation sequencing reads many DNA fragments at once and writes a base call for each cycle it can see. The library is fragments of a chosen size, fitted with adapters the instrument can hold and prime. A run is only as relevant as that library. This page follows the path from a nucleic-acid tube to a read file. It is a research explainer, not a manufacturer's run manual and not a clinical sequencing protocol.

The instrument class many benches picture is short-read sequencing by synthesis, with fluorescent reversible terminators and a flow cell, as described in the public Illumina technology overview. Long-read platforms ask different questions and use different libraries. Do not mix their quality rules into a short-read figure. Catalogue classes for the work sit under genomics and sequencing. The requirement travels with the quote request.

What a library is

Genomic DNA, or amplicons, or cDNA, is not yet a library. Library preparation breaks the material into fragments the chemistry can handle, repairs the ends so adapters can ligate, and attaches those adapters. The adapters carry the sequences the flow cell binds and the primer sites the sequencing reaction uses. In a multiplexed run they also carry indexes, so reads from different samples can be sorted after the fact.

Fragmentation is mechanical or enzymatic. The size of the insert decides whether paired reads will overlap, whether a repetitive region is spanned, and whether clusters amplify well. A library of very short inserts answers a different question from a library of long ones.

PCR during library preparation increases material and also copies whatever amplified most easily. Those copies become duplicate reads. PCR-free libraries avoid that step when the input is large enough. Unique molecular identifiers help only when the analysis actually uses them to tell copies from starting molecules.

Adapter dimers are short adapter-adapter molecules with almost no insert. They cluster very efficiently and produce reads that are mostly adapter. A size trace that shows a sharp peak near the length of two adapters is a warning to clean the library again, not a minor cosmetic. Primer or adapter leftovers are the same family of problem as primer-dimer on a PCR gel.

From molarity to a flow cell

Instruments consume a number of molecules, not a mass in a notebook. A mass measurement that ignores the average fragment length will overload a library of short pieces and underload a library of long ones. Fluorometric assays that see double-stranded DNA, plus a size distribution, are the usual pair. Absorbance alone cannot see adapter dimer and cannot tell DNA from free nucleotides. The habit of trusting one number is how flow cells are wasted.

Too many molecules crowd the clusters so their signals overlap and base calling fails. Too few waste the flow cell. The acceptable window belongs to the instrument and the chemistry version you are running. A number copied from a different model is a common route to a dim run.

Indexes need a plan. If two samples share an index, they cannot be separated. Index hopping, where a barcode is misassigned to another sample's fragment, is a known multiplex failure. Unique dual indexes, when the chemistry supports them, make that hop visible and filterable. A water-only library, or a blank carried from extraction, shows you the contamination and hopping floor of the run. Extraction blanks are discussed with the logic in how DNA extraction methods differ.

A control library or a known spike-in separates an instrument failure from an empty sample library. A failed control means the sample reads should not be analysed as if the cycle images were fine.

From nucleic acid to checked reads Fragments of known size Adapters and indexes Size and amount reject dimers Flow cell Reads then QC Demultiplex only after you know which index belonged to which tube.
A library is fragmented DNA with adapters. Only after a size and amount check should it be loaded. Reads still need a quality inspection before any biological claim.

What the read file does and does not contain

Each cluster produces a read, or a pair of reads, with a quality score per base. Those scores are the caller's confidence. They drop for the usual chemical reasons: a difficult sequence context, a fading cluster, a cycle that was disturbed. FastQC is a public tool that plots per-base quality, adapter content, duplication and a few other signs that the FASTQ is not what you hoped. It does not trim, map, or decide whether a variant is real. A passing plot is permission to start analysis, not the end of it.

Mapping places reads on a reference. The reference has a version. Ensembl and GenBank are public homes for those sequences. A variant called against the wrong assembly is a bookkeeping error that looks like biology. Unmapped reads can be contamination, circular DNA the reference omitted, or a library so full of adapter that there was nothing to map. Look at them before you discard the fraction as "low quality DNA" in the sample.

Coverage depth and coverage breadth are different. Reads piled on one exon can be an amplicon or a capture failure. A genome-wide claim needs breadth, and a variant claim needs depth at that locus plus an error model. PCR duplicates make depth look higher than the number of independent molecules.

Check before you trust the fileFavours this conclusionDoes not mean
Size trace without a dimer peakInserts, not adapter pairs, dominate the libraryThe insert is the genome region you wanted
Loading control or spike-in behavesThe instrument cycles produced usable signalEvery sample on the run was equally good
Per-base quality mostly highBase calls are confident on those cyclesThe read mapped to the right locus
Low adapter content after trimming checkAdapter sequence is not the whole readMapping and duplicates are fine
Blank library is nearly emptyGross contamination or hopping is limitedThe biological sample is pure

Choosing the shape of the experiment

Whole-genome libraries, targeted capture and amplicon sequencing spend reads on different fractions of the sample and fail for different reasons. A single sequencing label does not say which claim you ran. Amplicon runs inherit PCR bias and any contaminated no-template well.

Sanger sequencing remains the straightforward check for one plasmid or one edited amplicon. The Sanger DNA sequencing enquiry reference is a way to specify that job. A whole-genome question is a different specification, and the whole-genome sequencing enquiry reference is a prompt for it. Both are independent method references. Ask whether a quotation is possible. Neither page means a sequencer is already running for you.

A swapped index spreadsheet is indistinguishable from a surprising biological result. Demultiplex from the sheet that was written when the tubes were labelled.

Rooms, power and the limit of the claim

A library left on a warm bench is being incubated. Keep the instrument room inside the range stated for that instrument, and do not invent a new loading concentration because the room is hot. Let a cold flow cell meet room temperature with the lid closed before you open it. Condensation is water you did not mean to add.

A power cut during a run is an instrument event, not a minor pause. Some platforms can resume in ways their own manual describes. A resumed run is not automatically comparable to an uninterrupted one. Record it, and do not pool those reads with a later batch as if the cycle history were the same. Data files need a copy that survives the same power cut. A FASTQ that exists only on the instrument workstation is one disk away from disappearing.

Research sequencing of human or animal samples stays inside the ethics approval and the biosafety level of those samples. Reads are not a diagnosis. They are not a licence to identify a person outside the consent that covered the tube. This article does not provide that consent framework.

What the enquiry has to specify

EVRINTH can take a sourcing question. State the organism, the library type, the read length and pairing you need, the number of samples, the input integrity, and whether the nucleic acid is already a library or still a lysate. The genomics research pathway connects that specification to the surrounding design. Ask whether a quotation is possible. Include the reference genome you intend to map against. A request that says only "sequence this" has not yet chosen a method.

Take a nucleic-acid sample to a read file you can interpret

  1. 01Define the biological unit the reads must representDecide whether you need a whole genome, an enriched set of regions, or an amplicon. Coverage, fragment length and how much PCR you can tolerate follow from that choice.
  2. 02Build a library and measure the molecules, not only the massFragment, repair ends and add adapters by the chemistry class you are using. Then check amount and size distribution. An adapter-dimer peak is a small, real library that will steal clusters from the fragments you wanted.
  3. 03Match the loading and the indexes to the instrumentLoad in the molar range that instrument and that flow cell expect. Use indexes that let you pull mixed samples apart afterwards. A control library, when you have one, shows whether the run itself worked.
  4. 04Separate read quality from a biological conclusionInspect per-base quality, adapter content and duplication before you trust variant calls or counts. A FASTQ file is a measurement of the library that was loaded. It is not yet an answer about the organism.

Questions from the bench

Is a large number of reads the same as deep coverage of my genome?

No. Reads can be adapter dimers, duplicates of one PCR fragment, or DNA from the wrong organism. Coverage is how those reads land on the reference after you have filtered and mapped them. A run report with many reads and a tiny mapped fraction is a failed library, not a deeply sequenced genome.

What does a Q30 score actually mean?

On the usual Phred scale, a quality of 30 assigned to a base means the caller estimates a 1 in 1,000 chance that the base is wrong. It is a property of that base call, not of the whole read, and not of the variant you might later report. A read can be mostly high quality and still be in the wrong place because the insert was a repeat.

When is Sanger the better tool?

When you need a careful read of one known amplicon or one plasmid, Sanger sequencing is direct and the failure modes are easier to see on a trace. Next-generation sequencing earns its place when you need many molecules, many loci, or a genome. Using a flow cell to check a single clone is rarely the honest design.

Can these reads be treated as a medical diagnosis?

No. Research reads support a research claim after analysis, controls and whatever ethics review covered the samples. Diagnostic sequencing is a different quality system and a different legal setting. This page does not interpret anyone's clinical genome.

References

  1. Illumina overview of next-generation sequencing
  2. FastQC: a quality control tool for high throughput sequence data
  3. Ensembl genome browser
  4. protocols.io method repository
  5. NCBI GenBank

Manufacturer names identify published method classes. Trademarks remain with their owners. Catalogue records on this site are independent references for enquiry. They are not a statement of inventory, distribution rights or a supply commitment. This page is educational. It is not medical advice, a diagnostic protocol or a biosafety approval.

Catalogue

Related products and categories

These links follow the subject of the article into published manufacturer references. A listing is a reference for an enquiry, not a statement of stock or distribution rights.