Skip to content
EVRINTH

glossary

Normalisation is a choice

Why library size, median-ratio, TMM and TPM answer different questions, and why the normalisation method belongs in the result.

Author
EVRINTH Editorial Team
Published
8 October 2026
Updated
8 October 2026
Reading time
9 min
Researcher viewing gene expression heatmaps and genomic tracks on two monitors at night
Researcher viewing gene expression heatmaps and genomic tracks on two monitors at night

Change the normalisation and a gene can climb or fall a ranked list without a single read being added or removed. The method is not a polite prelude to the biology. It is part of the result. This glossary separates four scalings people collapse into the word normalised, and it says when the usual assumption fails. The table those scalings feed is discussed among the columns in a differential expression result, on the path in from cells to a gene expression result.

Rank is a consequence of scaling

Libraries do not arrive with equal depth. A deeper sample would look "up" for every gene if you compared raw counts. Scaling removes that technical tilt under an assumption. Change the assumption and you remove a different tilt. Genes move. None of this requires a laboratory error. It requires two honest methods pointed at one matrix.

The words below are method classes. Software implementations differ in defaults and in names. This page does not paste parameters. It tells you what question each class is allowed to answer.

Four scalings, four assumptions

Library size. Divide by the total assigned counts, or report counts per million. The assumption is that depth is the main difference between samples and that overall composition is similar. It fails when one condition is dominated by a few enormous transcripts, a viral RNA, leftover ribosomal RNA, or a strongly induced gene family. Those molecules inflate the denominator and push every other gene down. The table calls them repressed. They may simply be a smaller share of a library that one transcript filled.

Median ratio. This is the size-factor class associated with DESeq2-style workflows. For each gene, form the ratio of this sample's count to a reference level across samples, often a geometric mean, and take the median of those ratios as the sample's size factor. Genes with zero counts drop out of the median. The assumption is that most genes are not differentially expressed, so the middle of the ratio distribution reflects depth rather than biology. A minority of truly changing genes do not drag the median the way they drag a sum. A global shift, where most genes move together, does drag it, and the method will partly erase the shift you wanted to see.

TMM. The trimmed mean of M-values, the class associated with edgeR-style workflows, looks at log fold changes between samples and at average abundance. It trims genes with extreme fold changes and extreme abundance, then estimates a scaling from what remains. The assumption is again that most genes do not change, with an explicit decision to ignore the extremes when setting the factor. TMM and a median ratio often agree. They can disagree when the distribution of fold changes is awkward, when composition is wildly unbalanced, or when many zeros make the median noisy. "We used a standard method" is not a record of which one happened.

TPM. Transcripts per million adjust each gene for length and then scale the sample so those length-adjusted values sum to one million. TPM answers a within-sample composition question: the share of transcript mass, corrected so a long gene is not louder only because it is long. It is not a between-sample size factor for a count model. Feeding TPM into a test built for raw counts double-counts the scaling and throws away the mean-variance relationship the test expected. Comparing TPM of one gene across samples is common and still compositional. If one sample grows a huge unrelated transcript, other genes' TPM fall even when their molecule counts did not.

FPKM is a related length-and-depth scaling, not a fifth miracle. Treat it as a cousin of TPM when you are reading an older table. Do not mix FPKM and TPM in one ranked list and call the list a replicate.

Relative RT-qPCR normalises with reference genes, which is a different operation aimed at a few assays. It is not a substitute sentence when the data are counts. The assumptions are in RT-qPCR for relative expression.

When most genes do move

Median ratio and TMM are the wrong yardstick when the biology really is a global change: a treatment that shuts transcription down, a comparison of cell types with different RNA content per cell, or a host sample overwhelmed by pathogen RNA. The methods will try to hold the bulk of genes steady. They will succeed at the arithmetic and fail at the question.

Spike-ins are the class of answer. Exogenous RNA, added in a known amount per cell or per sample before the step you care about, gives a ladder that is not part of the endogenous shift. Use them when you have already decided the stable-majority assumption is false. Do not add them as a footnote after a confusing table. Spike-ins fail when the addition is uneven, when they extract differently from the sample RNA, or when the library chemistry treats the spike sequence unlike endogenous fragments. A bad spike is not a quiet correction. It is a new batch effect. State how the spike was added if the result depends on it.

There is no shame in library-size scaling when the question is explicitly about composition, including "what fraction of this library is ribosomal". There is a problem when that scaling is used to claim repression of every other gene.

Say the method in the same sentence as the list

A ranked list without a method is not reusable. Write the scaling, whether fold changes were shrunk, and the assumption. If you show TPM in a supplement and a median-ratio test in the main figure, label both axes with those words. Readers will subtract one column from the other if you do not.

Before you defend a gene, recompute the size factors under a second class if the biology is anywhere near the failure modes above. If the gene only exists under one scaling, the result is the scaling. If it survives both, you have a sturdier sentence. This is a branch in interpretation, not a demand to run every package.

Annotation still matters. Length for TPM comes from a transcript model. The wrong model, or a gene id joined to a length from a different species, scales the wrong feature. Check the release on Ensembl. A technology overview of the reads themselves is the Illumina sequencing introduction. It does not choose your size factor.

Gene rank under three scalings Library size 1. Gene A 2. Gene C 3. Gene B Median ratio 1. Gene A 2. Gene B 3. Gene C TPM 1. Gene C 2. Gene A 3. Gene B The reads did not change. The question each scaling answers did.
The same three genes can change order when library size, a median-ratio factor and TPM are used as the scaling.

Scaling against the question it can answer

ScalingQuestion it can carryAssumption that breaks it
Library size or counts per millionDepth differs, composition is broadly similarOne transcript dominates the library
Median ratioDifferential expression when most genes are stableA global shift of most genes
TMMThe same family of question, with extremes trimmedThe same global shift, or a very strange fold distribution
TPMShare of length-adjusted mass inside one sampleUsing it as a count, or as a between-sample proof
Spike-in ladderThe stable-majority assumption is false on purposeUneven addition of the spike

Rankings that moved for a clerical reason

The failure mode is a methods sentence that says normalised and a figure that mixes two of the rows above. A reviewer, or a future you, will not be able to tell whether a gene fell because the treatment repressed it or because TPM punished it for a long isoform. Another failure is to normalise away a ribosomal failure. If half the library is ribosomal, a size factor will shrink the informative counts and the QC report was the right place to stop. Scaling is not a detergent for a bad library.

Zeros interact with the median. A gene that is zero in many samples contributes nothing to those size factors, which is intended, and then its own ratio is estimated from a thin remainder. Do not narrate that ratio as precise just because the column has many digits.

Deposited counts in the European Nucleotide Archive are far more reusable as raw assignments plus a sample sheet than as a pre-scaled matrix with the method lost. Keep the integers. The scaling can be repeated. A TPM-only deposit cannot be put back into a count model honestly.

Not a clinical normal range

None of these methods defines a normal clinical range, a diagnostic threshold, or a biosafety level. They are research scalings for research counts. Human and infectious samples follow institutional containment regardless of which size factor you prefer. Do not describe a TPM share as a medical percentage.

Two sites, two size factors

Collaborators often exchange a spreadsheet that has already been scaled, then each scales it again. The second scaling does not cancel the first. Agree the integer matrix as the object that moves, and agree the sentence, median ratio or TMM or TPM, as a choice each analysis must restate. If one site's samples contain a pathogen RNA and the other's do not, say so before either site computes a factor. The factor will treat that RNA as depth.

A warm shipment that degrades one site's RNA changes composition toward fragments that survived. Normalisation will not recognise the journey. Integrity belongs in the sample sheet beside the size factor, or the factor will explain a courier as if it were a pathway.

What to specify when counts will be compared

Say whether you need raw counts, a named size-factor class, or a within-sample composition such as TPM, and say if a global shift is plausible enough that spike-ins should be discussed before the library is made. Sequencing chemistry can be raised through the mRNA sequencing enquiry reference. The scaling and the test can be raised through the differential expression analysis enquiry reference. Those pages are a prompt to name the method. They are not a default normalisation applied in silence.

Related consumables sit in the genomics and sequencing catalogue. The surrounding path is nucleic acid analysis. Put the scaling sentence in the quote request.

Questions from the bench

Is TPM a normalisation for differential expression?

It is a within-sample scaling to composition, after a length adjustment, set to a million. That answers what share of this sample a gene represents. It is the wrong input for a count-based differential test, which needs the integer assignments and a variance model that matches counts. Using TPM as if it were a size factor will change the ranking for a clerical reason.

When do spike-ins belong in the design?

When you cannot assume that most genes stay put. A global transcriptional shutdown, or a comparison of cells with very different RNA content, breaks median-ratio and TMM assumptions, because those methods treat the bulk of genes as a stable yardstick. Exogenous RNA added in a known amount is one way to bring an external yardstick. The addition has to be even, or the yardstick is worse than the assumption it replaced.

Will a median-ratio method and TMM rank genes the same way?

Often they are close, because both assume that most genes are not differentially expressed and both try to ignore a minority of large changes. They are not the same calculation. On a dataset dominated by a few huge transcripts, or by a global shift, they can part company. State which one you used rather than saying the counts were normalised.

What sentence should sit next to a ranked list?

Name the scaling, the test, and the assumption. For example, median-ratio size factors with a count model, assuming most genes do not change. A reader can then judge whether that assumption fits the biology. A list with no sentence cannot be compared with next month's list.

References

  1. Ensembl genome browser
  2. Illumina overview of next-generation sequencing
  3. European Nucleotide Archive

Manufacturer names identify published method classes. Trademarks remain with their owners. Catalogue records on this site are independent references for enquiry. They are not a statement of inventory, distribution rights or a supply commitment. This page is educational. It is not medical advice, a diagnostic protocol or a biosafety approval.

Catalogue

Related products and categories

These links follow the subject of the article into published manufacturer references. A listing is a reference for an enquiry, not a statement of stock or distribution rights.