explainer
Single-cell RNA-seq in outline
What single-cell RNA-seq adds beyond a bulk average: droplets or wells, empty droplets, doublets, UMIs, sparse counts, and clustering as a parameter.
- Author
- EVRINTH Editorial Team
- Published
- 8 October 2026
- Updated
- 8 October 2026
- Reading time
- 7 min

A bulk RNA-seq tube reports the average of every cell you lysed together. If half the cells induce a transcript and half silence it, the average can look unchanged. Single-cell RNA-seq tries to give each cell, or each nucleus, its own barcode so that composition and per-cell state can be told apart. This explainer is the outline of that idea: what the wet lab is doing, what the matrix contains, and which claims the outline cannot carry. Bulk design still starts from from cells to a gene expression result. A short-read technology class is described in the Illumina sequencing overview.
The decision the outline supports
Choose single-cell methods when the question is about which cells are present, or whether a change is a shift inside cells rather than a shift in how many of those cells you captured. Choose bulk methods when the question is the average of a sample you can define, and you would rather spend effort on biological replicates than on dissociation. The single-cell RNA sequencing enquiry reference is a page where that choice can be discussed. It is an enquiry reference. The biology of the question still has to be written by you.
Droplets, wells, and a barcode per compartment
Two families dominate the practical outline. In droplet methods, cells are partitioned into tiny aqueous compartments in oil, together with beads that carry barcoded primers. In plate or well methods, cells are sorted or pipetted into wells and a barcode is added there. Droplets favour large cell numbers and a shallower look at each cell. Wells can go deeper on fewer cells and can pair more easily with a known surface marker from the sort. Both end as a table of counts per barcode per gene. Neither is a microscope.
The chemistry after partitioning is still a tiny RNA-seq library: reverse transcription, amplification, and sequencing. Many workflows add a unique molecular identifier, a UMI, as a random tag on each captured molecule. PCR then copies that tagged molecule many times. Collapsing reads that share a UMI, a cell barcode and a gene is how you estimate molecules rather than amplification copies. UMIs do not correct a cell that was stressed before it was captured. They correct a counting bias that the PCR would otherwise invent.
Empty droplets are compartments that received a barcode and ambient RNA, and no cell. They produce low, messy counts that reflect whatever RNA leaked into the suspension. Calling them is a threshold or a model, often visible as a knee in a barcode-rank plot. Set the threshold with the plot in hand, and record it. Doublets are compartments that received two cells. They can look like an exciting intermediate state because they express markers of both parents. Doublet-finding tools and a sense of unexpected marker combinations are the checks. A doublet rate rises when you overload the partition.
Sparse counts are the normal matrix
A mammalian cell contains a wide transcriptome, and a single-cell library samples a thin slice of it. Most genes in most barcodes are zero. That sparsity is expected. Downstream steps often normalise, select variable genes, and reduce dimensions before anyone draws a plot. Each of those steps has parameters. The plot is the parameters plus the data.
Clustering, commonly a graph-based method with a resolution setting, groups barcodes in that reduced space. The resolution is a knob. Turn it up and you get more clusters. Turn it down and populations merge. There is no factory setting that reveals the true number of cell types. Marker genes, known from prior work or checked on a genome browser, are how you argue that a cluster matches a named population. A cluster without a marker story is a group of similar barcodes. Ambient RNA can also paint the same contaminant gene onto many clusters, which looks like shared biology.
The biological replicate is easy to forget because the cell count is large. One well of cells from one donor is one donor. Treatment claims need donors or independent cultures, with the same logic as biological versus technical replicates. Comparing two captures done on different days adds a batch. Include the capture identity in the analysis when you have more than one.
| Feature | What bulk RNA-seq gives you | What the single-cell outline adds |
|---|---|---|
| Object in the tube | A pooled lysate | A barcode meant to represent one cell or nucleus |
| Typical matrix | Dense enough for gene-level contrasts | Sparse, with many zeros from limited capture |
| Molecule counting | Deduplication depends on the protocol | UMIs are there to collapse PCR copies |
| Main artefacts | Batch, DNA, the wrong library selection | Empty droplets, doublets, dissociation stress |
| Clustering | Sometimes used on samples | A parameter on cells, not a cell type by default |
| Sample size for a donor-level claim | Number of donors or cultures | Still the number of donors, not the number of cells |
Where the outline sends you back to bulk, or forward with care
If dissociation takes so long that stress genes dominate every cluster, you have measured the protocol. Nuclei instead of whole cells are one branch when intact cells will not survive, and nuclei measure a different RNA population. If the suspension is full of debris, empty-droplet and doublet calls become guesses. Clean the suspension or repeat, rather than naming every cloud on the plot.
If the question collapses to "did this gene change on average", a bulk library or a targeted RT-qPCR may answer it with less theatre. qPCR on sorted populations is a strong follow-up because it ties a marker you can defend to a transcript measurement with its own controls. It is not automatic validation of every cluster.
Read other people's matrices in the Sequence Read Archive with their chemistry and their donor count in view. A pretty embedding from one donor is a case study.
Research use, and biosafety for live cells
Single-cell work often starts with live human or animal cells. Containment, infection risk and ethics approval are institutional decisions. The WHO laboratory biosafety manual is a public orientation to laboratory biosafety thinking, not a permit for your project. This page does not diagnose a patient from a cluster, and it does not authorise clinical action from a research embedding. Dissociation enzymes and fixation reagents have their own hazard notes. Follow those notes for the products you handle.
Dissociation in a warm room
Cells keep changing until they are lysed or fixed. In a warm laboratory, a dissociation that overruns because a centrifuge queue is long will induce stress transcripts and kill fragile populations. The cluster you then call a cell type may be the cells that survived the wait. Keep the suspension cold in the way the protocol specifies, write the time from tissue to capture, and treat a power cut that warms the room or the reagents as part of that time. Humidity and open tubes are a smaller issue than temperature here, but a dried pellet is a lost capture. None of this is repaired by adding cells from the same stressed suspension and calling them replicates.
What an enquiry should already say
State the organism, the tissue, whether you can bring a single-cell suspension or need nuclei, the number of donors or cultures, and the question in one sentence: composition, a within-cell response, or an average you might be better off measuring in bulk. Instrument classes sit in the genomics and sequencing catalogue. The wider nucleic acid path is the nucleic acid analysis pathway.
Use the single-cell RNA sequencing enquiry reference when that method is the one you want discussed, and the mRNA sequencing enquiry reference or the differential expression analysis enquiry reference when the question is still bulk. Put the specimen, the donor count and the decision on the quote request, and ask whether a quotation is possible. A single-cell design can be discussed from that note. The cluster names remain a scientific claim you support with markers and with replicates at the right unit.
Questions from the bench
Is a cluster in a single-cell plot a cell type?
A cluster is a group of barcodes that sit together under the parameters you chose: which genes, how many components, how many neighbours, and what resolution. Those settings can split one biological population or merge two. A cell-type name belongs on a cluster only after markers, and preferably a second line of evidence, support it. Changing the resolution is allowed to change the picture.
Do thousands of cells mean thousands of biological replicates?
They mean thousands of cells, usually from a much smaller number of captures and donors. The unit for a claim about a treatment in a population is still the donor, the animal, or the independent culture. Cells from one donor are not independent samples of the species. Report both numbers, and do not let the cell count become n.
Why are most of the counts zero?
Each cell yields a small sample of its RNA, so many transcripts are simply not captured. The resulting matrix is sparse. A zero can be a gene the cell does not express, or a gene it expresses below the capture rate. Treating every zero as a confident biological off-switch overstates the chemistry. UMIs help you count molecules rather than PCR copies, and they do not fill in the missing genes.
When is bulk RNA-seq the better outline?
When you need a robust average of a defined sample, when the tissue cannot be dissociated without destroying the biology, or when the question is a straightforward group contrast with limited material and limited budget. Single-cell methods answer composition and cell-to-cell variation. They are a poor replacement for a bulk design you have not yet replicated.
References
Manufacturer names identify published method classes. Trademarks remain with their owners. Catalogue records on this site are independent references for enquiry. They are not a statement of inventory, distribution rights or a supply commitment. This page is educational. It is not medical advice, a diagnostic protocol or a biosafety approval.
Catalogue
Related products and categories
These links follow the subject of the article into published manufacturer references. A listing is a reference for an enquiry, not a statement of stock or distribution rights.
Continue in this cluster
Related reading
From cells to a gene expression resultHow a laboratory goes from cells to a gene expression result, and how RNA quality decides between a focused RT-qPCR assay and RNA-seq.
Spatial transcriptomics as a conceptHow spatial transcriptomics attaches expression to a tissue coordinate, why spot and cell resolution differ, and what section quality and permeabilization
Biological versus technical replicatesHow to tell a biological replicate from a technical one, and which unit of inference actually supports a treatment claim in an expression study.