Pillar guide
Next-generation sequencing from library to reads
How a sequencing library becomes reads: adapters, flow cells, quality scores and the checks that stop a bad library from wasting a run.
- Author
- EVRINTH Editorial Team
- Published
- 8 October 2026
- Updated
- 8 October 2026
- Reading time
- 8 min

Next-generation sequencing reads many DNA fragments at once and writes a base call for each cycle it can see. The library is fragments of a chosen size, fitted with adapters the instrument can hold and prime. A run is only as relevant as that library. This page follows the path from a nucleic-acid tube to a read file. It is a research explainer, not a manufacturer's run manual and not a clinical sequencing protocol.
The instrument class many benches picture is short-read sequencing by synthesis, with fluorescent reversible terminators and a flow cell, as described in the public Illumina technology overview. Long-read platforms ask different questions and use different libraries. Do not mix their quality rules into a short-read figure. Catalogue classes for the work sit under genomics and sequencing. The requirement travels with the quote request.
What a library is
Genomic DNA, or amplicons, or cDNA, is not yet a library. Library preparation breaks the material into fragments the chemistry can handle, repairs the ends so adapters can ligate, and attaches those adapters. The adapters carry the sequences the flow cell binds and the primer sites the sequencing reaction uses. In a multiplexed run they also carry indexes, so reads from different samples can be sorted after the fact.
Fragmentation is mechanical or enzymatic. The size of the insert decides whether paired reads will overlap, whether a repetitive region is spanned, and whether clusters amplify well. A library of very short inserts answers a different question from a library of long ones.
PCR during library preparation increases material and also copies whatever amplified most easily. Those copies become duplicate reads. PCR-free libraries avoid that step when the input is large enough. Unique molecular identifiers help only when the analysis actually uses them to tell copies from starting molecules.
Adapter dimers are short adapter-adapter molecules with almost no insert. They cluster very efficiently and produce reads that are mostly adapter. A size trace that shows a sharp peak near the length of two adapters is a warning to clean the library again, not a minor cosmetic. Primer or adapter leftovers are the same family of problem as primer-dimer on a PCR gel.
From molarity to a flow cell
Instruments consume a number of molecules, not a mass in a notebook. A mass measurement that ignores the average fragment length will overload a library of short pieces and underload a library of long ones. Fluorometric assays that see double-stranded DNA, plus a size distribution, are the usual pair. Absorbance alone cannot see adapter dimer and cannot tell DNA from free nucleotides. The habit of trusting one number is how flow cells are wasted.
Too many molecules crowd the clusters so their signals overlap and base calling fails. Too few waste the flow cell. The acceptable window belongs to the instrument and the chemistry version you are running. A number copied from a different model is a common route to a dim run.
Indexes need a plan. If two samples share an index, they cannot be separated. Index hopping, where a barcode is misassigned to another sample's fragment, is a known multiplex failure. Unique dual indexes, when the chemistry supports them, make that hop visible and filterable. A water-only library, or a blank carried from extraction, shows you the contamination and hopping floor of the run. Extraction blanks are discussed with the logic in how DNA extraction methods differ.
A control library or a known spike-in separates an instrument failure from an empty sample library. A failed control means the sample reads should not be analysed as if the cycle images were fine.
What the read file does and does not contain
Each cluster produces a read, or a pair of reads, with a quality score per base. Those scores are the caller's confidence. They drop for the usual chemical reasons: a difficult sequence context, a fading cluster, a cycle that was disturbed. FastQC is a public tool that plots per-base quality, adapter content, duplication and a few other signs that the FASTQ is not what you hoped. It does not trim, map, or decide whether a variant is real. A passing plot is permission to start analysis, not the end of it.
Mapping places reads on a reference. The reference has a version. Ensembl and GenBank are public homes for those sequences. A variant called against the wrong assembly is a bookkeeping error that looks like biology. Unmapped reads can be contamination, circular DNA the reference omitted, or a library so full of adapter that there was nothing to map. Look at them before you discard the fraction as "low quality DNA" in the sample.
Coverage depth and coverage breadth are different. Reads piled on one exon can be an amplicon or a capture failure. A genome-wide claim needs breadth, and a variant claim needs depth at that locus plus an error model. PCR duplicates make depth look higher than the number of independent molecules.
| Check before you trust the file | Favours this conclusion | Does not mean |
|---|---|---|
| Size trace without a dimer peak | Inserts, not adapter pairs, dominate the library | The insert is the genome region you wanted |
| Loading control or spike-in behaves | The instrument cycles produced usable signal | Every sample on the run was equally good |
| Per-base quality mostly high | Base calls are confident on those cycles | The read mapped to the right locus |
| Low adapter content after trimming check | Adapter sequence is not the whole read | Mapping and duplicates are fine |
| Blank library is nearly empty | Gross contamination or hopping is limited | The biological sample is pure |
Choosing the shape of the experiment
Whole-genome libraries, targeted capture and amplicon sequencing spend reads on different fractions of the sample and fail for different reasons. A single sequencing label does not say which claim you ran. Amplicon runs inherit PCR bias and any contaminated no-template well.
Sanger sequencing remains the straightforward check for one plasmid or one edited amplicon. The Sanger DNA sequencing enquiry reference is a way to specify that job. A whole-genome question is a different specification, and the whole-genome sequencing enquiry reference is a prompt for it. Both are independent method references. Ask whether a quotation is possible. Neither page means a sequencer is already running for you.
A swapped index spreadsheet is indistinguishable from a surprising biological result. Demultiplex from the sheet that was written when the tubes were labelled.
Rooms, power and the limit of the claim
A library left on a warm bench is being incubated. Keep the instrument room inside the range stated for that instrument, and do not invent a new loading concentration because the room is hot. Let a cold flow cell meet room temperature with the lid closed before you open it. Condensation is water you did not mean to add.
A power cut during a run is an instrument event, not a minor pause. Some platforms can resume in ways their own manual describes. A resumed run is not automatically comparable to an uninterrupted one. Record it, and do not pool those reads with a later batch as if the cycle history were the same. Data files need a copy that survives the same power cut. A FASTQ that exists only on the instrument workstation is one disk away from disappearing.
Research sequencing of human or animal samples stays inside the ethics approval and the biosafety level of those samples. Reads are not a diagnosis. They are not a licence to identify a person outside the consent that covered the tube. This article does not provide that consent framework.
What the enquiry has to specify
EVRINTH can take a sourcing question. State the organism, the library type, the read length and pairing you need, the number of samples, the input integrity, and whether the nucleic acid is already a library or still a lysate. The genomics research pathway connects that specification to the surrounding design. Ask whether a quotation is possible. Include the reference genome you intend to map against. A request that says only "sequence this" has not yet chosen a method.
Take a nucleic-acid sample to a read file you can interpret
- 01Define the biological unit the reads must representDecide whether you need a whole genome, an enriched set of regions, or an amplicon. Coverage, fragment length and how much PCR you can tolerate follow from that choice.
- 02Build a library and measure the molecules, not only the massFragment, repair ends and add adapters by the chemistry class you are using. Then check amount and size distribution. An adapter-dimer peak is a small, real library that will steal clusters from the fragments you wanted.
- 03Match the loading and the indexes to the instrumentLoad in the molar range that instrument and that flow cell expect. Use indexes that let you pull mixed samples apart afterwards. A control library, when you have one, shows whether the run itself worked.
- 04Separate read quality from a biological conclusionInspect per-base quality, adapter content and duplication before you trust variant calls or counts. A FASTQ file is a measurement of the library that was loaded. It is not yet an answer about the organism.
Questions from the bench
Is a large number of reads the same as deep coverage of my genome?
No. Reads can be adapter dimers, duplicates of one PCR fragment, or DNA from the wrong organism. Coverage is how those reads land on the reference after you have filtered and mapped them. A run report with many reads and a tiny mapped fraction is a failed library, not a deeply sequenced genome.
What does a Q30 score actually mean?
On the usual Phred scale, a quality of 30 assigned to a base means the caller estimates a 1 in 1,000 chance that the base is wrong. It is a property of that base call, not of the whole read, and not of the variant you might later report. A read can be mostly high quality and still be in the wrong place because the insert was a repeat.
When is Sanger the better tool?
When you need a careful read of one known amplicon or one plasmid, Sanger sequencing is direct and the failure modes are easier to see on a trace. Next-generation sequencing earns its place when you need many molecules, many loci, or a genome. Using a flow cell to check a single clone is rarely the honest design.
Can these reads be treated as a medical diagnosis?
No. Research reads support a research claim after analysis, controls and whatever ethics review covered the samples. Diagnostic sequencing is a different quality system and a different legal setting. This page does not interpret anyone's clinical genome.
References
Manufacturer names identify published method classes. Trademarks remain with their owners. Catalogue records on this site are independent references for enquiry. They are not a statement of inventory, distribution rights or a supply commitment. This page is educational. It is not medical advice, a diagnostic protocol or a biosafety approval.
Catalogue
Related products and categories
These links follow the subject of the article into published manufacturer references. A listing is a reference for an enquiry, not a statement of stock or distribution rights.
Continue in this cluster
Related reading
What to check before a sequencing runWhat to confirm before a sequencing run: sample identity, amount, fragment size, indexes, and which failed check should stop the instrument.
16S profiling and its taxonomic limitsJudge a 16S profile at the rank the marker supports, and see where copy number, primer bias and species names stop being honest.
A glossary of sequencing termsWorking definitions of read, coverage, depth, MAPQ, Phred, VCF, BAM and the related words, each tied to the mistake that word prevents.