protocol overview
Metagenomic sequencing overview
Plan shotgun sequencing of community DNA, including host contamination, rare taxa, and assembly versus read classification.
- Author
- EVRINTH Editorial Team
- Published
- 8 October 2026
- Updated
- 8 October 2026
- Reading time
- 11 min

A metagenome, in the sense this overview uses, is shotgun sequencing of DNA from a community. Many genomes enter the library because many organisms were in the tube, plus whatever host and whatever reagent background extraction brought along. The decision it supports is which organisms and which genes you can defend at a stated depth, and which absences are only the limit of that depth. It is not a 16S experiment. Marker-gene profiling has a different bias and a lower ceiling, set out in 16S profiling and its taxonomic limits. The instrument side is still a library and a run, as in next-generation sequencing from library to reads.
People commission this method for soil, water, stool, biofilms, and swabs. The science is the same shape in each case and the host fraction is not. Plan the host first or the read budget will surprise you after the invoice, not before.
Shotgun fragments have no primer to define them
There is no pair of primers choosing a gene. Fragmentation and adapter ligation, the path in library prep is where most runs are won, turn total extracted DNA into a library. Every genome that lysed can contribute, in proportion to its abundance, its genome size, and how willingly it broke open. A large genome at moderate abundance can out-read a small genome that was more numerous. Counts of reads are not counts of cells until you model genome size, and even then extraction bias remains.
Short or long reads are available, with the trade-offs in how short reads and long reads differ. Short reads classify well against databases and struggle to assemble repeats and strain mixtures. Long reads can span those repeats when the extraction kept the DNA long, which environmental samples often do not. Choose the read class for the assembly you will actually attempt. A long-read specification on a harsh bead-beating lysate is a short-read experiment wearing the wrong name.
Host DNA is community DNA from the point of view of the sequencer. A tissue swab can be nearly all host. Those reads are high quality and biologically real and they are not the microbiome. Map or classify them against a host reference and set them aside before you interpret the remainder. The reference has a build, even for the host, which reference genomes and why the build matters already insisted on for simpler experiments. A poor host reference leaves host reads in the microbial pile, where they become fake microbes that match the nearest database neighbour.
Extraction, blanks, and the community you did not lyse
Extraction is the first analysis. A protocol harsh enough to break spores and Gram-positive walls will shear softer genomes. A gentle protocol does the reverse. The community in the FASTQ is the community you lysed, not the community a microscope would have counted. How DNA extraction methods differ is the general comparison. For a metagenome, write the method down to the bead and the lysis class, because a later paper cannot reinterpret your species list without it.
Sequence an extraction blank and a library blank. Kit reagents carry their own low-level DNA. In a high-biomass stool sample the reagent signal may be a small fraction. In a low-biomass swab it can be the sample. Any taxon you report that is also in the blank, at a similar level, is not a discovery. This is not optional hygiene. It is the control that makes rarity interpretable.
A mock community, when you have one of known composition, tells you what this extraction and this classifier do to a truth you planted. It does not validate every taxon in a real sample. It shows the direction of the bias. Use it that way.
Inhibitors from soil and stool, humic substances and polysaccharides among them, can stall a library even when a dye says DNA is present. A failed library of a real sample next to a successful blank is inhibition or loss, not a sterile ecosystem. Branch back to cleanup or dilution of the extract. Do not interpret an empty FASTQ as ecology.
Classification and assembly answer different questions
Read classification compares reads, or short strings drawn from them, to a database and assigns a taxon at whatever rank the match supports. It can work from a handful of reads, which is why it sees rarer organisms than assembly does. It cannot see organisms that are not in the database, except as an unassigned fraction or a nearest neighbour that may be wrong. The database version is part of the result. A species called in 2018 and a species called from the same reads against a later database are two results. Deposit the reads, for example at the European Nucleotide Archive or the Sequence Read Archive, so the classification can be repeated.
Assembly overlaps reads into contigs and may bin contigs into metagenome-assembled genomes. Bins are hypotheses about which contigs share a genome. Completeness and contamination estimates are models of single-copy marker genes, not a guarantee that the bin is one organism you could culture. Strain mixtures with similar sequences fracture or collapse. A bin at low coverage is a sketch. Depth for assembly has to cover that genome many times. A community plan that bought enough reads to classify the top taxa has not bought enough reads to assemble the rare ones. Say which of the two you paid for.
Gene catalogues and mapping reads to known proteins are a third path: functional potential without claiming a bin. The gene was present as DNA. It may not be expressed. Expression is metatranscriptomics, a different library with different degradation rules, closer in handling to protecting RNA during extraction. Do not report RNA claims from a DNA metagenome.
Branch when the host fraction dominates. Either the biological story is about a small microbial remainder, in which case you need more sequencing or a justified depletion step, or the sample type cannot answer the question. Depletion kits remove some microbial DNA too. If you deplete, the abundances are no longer comparable to an undepleted study. Write the step into the comparison or do not make the comparison.
Branch when classification and assembly disagree. A species with many assigned reads and no decent bin may be a database relative of several real strains, or a depth problem. Look at the alignments before you pick the list that fits the abstract.
| Choice | What a positive result can mean | What it cannot mean |
|---|---|---|
| Extraction blank is nearly empty | Reagent background is low relative to samples | Every rare taxon is biological |
| Host fraction reported first | Microbial depth is the remainder | Raw read count equals microbial coverage |
| Classifier hit at genus | The read resembles that clade in this database | A cultured species identification |
| Classifier hit at species | The database and threshold allowed it | The same species another database would name |
| Metagenome-assembled genome | Contigs were grouped under a stated model | A complete, pure, culturable isolate |
| Gene present in DNA | Functional potential in the extract | Expression, activity, or a phenotype |
| Taxon absent | Not seen at this depth after filters | Absent from the ecosystem |
Rarity, contamination, and a function you did not sequence
A rare taxon requires many more reads before one of its fragments is likely to appear. Failure to see it is a statement about depth and about lysis. Stacking samples in a figure and treating every zero as biological absence is how under-sequenced libraries become ecology. Publish the microbial read count after host removal, not only the raw total.
Contamination from the bench, from the index hop described in adapters indexes and barcode hopping, and from the kit all land in the same rare fraction where the interesting biology is supposed to sit. Unique dual indexes and a sequenced blank are the practical defence. A low-level pathogen hit in a single sample, absent from the blank and absent from replicates, is a lead. It is not a finding you attach a clinical verb to.
Functional over-claim is the other common failure. A genus associated in the literature with a pathway does not put that pathway in your sample. The gene sequence has to be in the reads or the contigs. Even then you have potential, not activity, and not a metabolite. Pair the claim to the analysis you ran.
Quality scores still matter and still do not rescue composition. A high-Q30 file of mostly host DNA is a good run of the wrong fraction. Judge quality scores and a FASTQ file first, then the host fraction, then the taxa.
Samples may be infectious
Stool, sewage, clinical swabs, and animal samples can contain infectious organisms. Shotgun extraction may or may not inactivate them. The institutional biosafety decision, informed by documents such as the Laboratory biosafety manual, chooses the room and the inactivation. This overview does not. Do not culture an organism out of a metagenomic "hit" without that decision. Research sequence is not a diagnostic identification and not permission to handle a higher risk group on an open bench. Biosafety basics for research benches is the local frame for the work that precedes the sequencer.
Heat on the way to the freezer
Community composition is not stable in a warm tube. Some cells grow, some lyse, and the DNA you extract hours later can be a different mixture from the sample that was collected. The holding rule belongs in the protocol: freeze, or the preservative the project chose, and the time allowed in between. A courier handoff across a hot afternoon without that preservative is a new experiment. Record the time out of the cold. A power cut that warms a freezer full of stool or soil samples is an excursion, in the sense of storing biological samples from fridge to freezer. Do not pool those extracts with an uninterrupted batch and call the batch one community. Write the holding condition into the specification so the receiving laboratory can reject a warm shipment instead of sequencing it politely.
What the enquiry has to decide
State the matrix, the host, the extraction class, whether a blank and a mock will be sequenced, the read class, and whether the deliverable is classifications, assemblies, or both. Name the database if you already require one. Say that 16S is a different request if someone has used the words interchangeably.
A shotgun community design can be discussed against the whole-genome sequencing enquiry reference if the brief makes clear the genome is a community, not one organism. A single cultured isolate confirmed at one locus can be discussed against the Sanger DNA sequencing enquiry reference, which is the wrong tool for a community and the right tool for that isolate. Reagent classes are in the genomics and sequencing catalogue. The study shape sits on the genomics research pathway. Put matrix, host fraction, blanks, and analysis type in the quote request. Commissioning a sequencing or proteomics study is how those stay fixed, and what to check before a sequencing run is the gate that stops a host-dominated library from being described as deep microbial coverage.
Plan shotgun sequencing of a mixed community
- 01Name the community and the host it rides in onWrite whether the tube is soil, water, stool, a swab, or a tissue, and which host genome will consume reads. A skin swab and a stool sample are different host fractions. The plan for depth starts from that fraction, not from a generic microbiome number.
- 02Choose an extraction and sequence a blank beside itHard-to-lyse cells and easy-to-lyse cells do not enter the tube in the proportions they had in the sample. Record the extraction class and run an extraction blank through library prep. Reads in the blank are the reagent background your samples must exceed.
- 03Decide classification, assembly, or both before loadingRead classification answers which known sequences the reads resemble. Assembly asks for enough depth of a genome to overlap it into contigs. A rare organism may classify from a few reads and still be impossible to assemble. Do not promise bins from a depth chosen for classification.
- 04Bound every absence and every functionAbsence means not detected at this depth after host reads were removed. Function means a gene sequence was present, and only if you actually analysed genes. A taxon name is not a metabolic pathway. Keep 16S results in a separate column if you also ran a marker.
Questions from the bench
How is shotgun metagenomics different from 16S profiling?
Shotgun sequencing fragments the community DNA that extraction released, without primers that define a single gene. 16S profiling amplifies one marker and classifies that marker. A shotgun file can contain functional genes and strain-level differences when depth allows. A 16S file cannot, which is why the marker method has its own limits and its own page.
What should we do when most reads are host?
Report the host fraction before you report the community. The microbial depth is what remains. You can sequence more deeply, deplete host material before the library, or change the sample type. Depletion is a biased extra step and belongs in the methods. Ignoring the host fraction and quoting the raw read count as microbial coverage overstates the experiment.
Why can two reasonable analyses list different species?
Classifiers use different databases and different thresholds. A read may be placed at genus in one database and forced to a species in another. Assembly-based bins may split or join those same reads. Publish the database version and the rank you trust. A species list without that version is not reproducible, even when both lists came from the same FASTQ.
Can a metagenome identify a pathogen for clinical action?
Not on the authority of this overview. A research classification is a database match. Clinical identification needs a validated assay, a quality system, and the legal framework that applies where you work. Treat an unexpected pathogen-like hit as a reason to check blanks and to talk to the biosafety process, not as a diagnosis.
References
Manufacturer names identify published method classes. Trademarks remain with their owners. Catalogue records on this site are independent references for enquiry. They are not a statement of inventory, distribution rights or a supply commitment. This page is educational. It is not medical advice, a diagnostic protocol or a biosafety approval.
Catalogue
Related products and categories
These links follow the subject of the article into published manufacturer references. A listing is a reference for an enquiry, not a statement of stock or distribution rights.
Continue in this cluster
Related reading
Next-generation sequencing from library to readsHow a sequencing library becomes reads: adapters, flow cells, quality scores and the checks that stop a bad library from wasting a run.
16S profiling and its taxonomic limitsJudge a 16S profile at the rank the marker supports, and see where copy number, primer bias and species names stop being honest.
A glossary of sequencing termsWorking definitions of read, coverage, depth, MAPQ, Phred, VCF, BAM and the related words, each tied to the mistake that word prevents.