Skip to content
EVRINTH

guide

A glossary of expression terms

Plain definitions of Cq, TPM, counts, FDR, UMIs and the other expression words that cause arguments when they are used loosely.

Author
EVRINTH Editorial Team
Published
8 October 2026
Updated
8 October 2026
Reading time
11 min
Researcher viewing gene expression heatmaps and genomic tracks on two monitors at night
Researcher viewing gene expression heatmaps and genomic tracks on two monitors at night

Cq, TPM and FDR are not synonyms for confidence. Each names a different operation, and mixing them is how a careful experiment becomes a loose sentence. This glossary pins twelve words to the mistake each one usually hides. The assay path they belong to is from cells to a gene expression result. Definitions below are research usage. They are not a diagnostic dictionary.

Terms from the qPCR bench

Cq and Ct. The quantification cycle is the cycle number at which fluorescence crosses a threshold you, or the instrument, chose. Ct means the same crossing in older papers. A lower Cq means more starting target in that assay, if efficiency is comparable and the threshold is honest. It does not mean 18 percent when the Cq is 18. It does not travel across unrelated genes without a standard. Reporting expectations for this number are part of the MIQE guidelines.

RT. In this cluster, RT means reverse transcription: an enzyme copying RNA into complementary DNA. The mistake is to call the whole RT-qPCR experiment "PCR" and then forget that the reverse transcription can fail while the later polymerase is fine. A second mistake is the abbreviation itself. In a written protocol, RT can also mean room temperature. Say reverse transcription in full wherever a freezer note could be misread.

The priming of that step matters. Oligo-dT, random primers and gene-specific primers do not copy the same population of RNA. A Cq downstream of one priming method is not the same measurement as a Cq downstream of another.

Primer efficiency. Efficiency describes how completely the product doubles each cycle, usually inferred from the slope of a dilution series of a clean template. One hundred percent means a doubling. The common mistake is to assume that number and apply a comparative Cq formula when the target slope and the reference slope disagree, or when the sample is inhibited and the clean standard was not. Efficiency belongs to a stated mix and a stated primer pair. It does not transfer silently to a new master mix.

Reference gene. A reference gene is a transcript used to scale a target in relative RT-qPCR, on the claim that the reference did not move under the treatment. "Housekeeping" is the older hope that some transcripts never move. The mistake is to treat GAPDH or a similar symbol as a universal normaliser because a previous tissue used it. Stability is a result you measure in the samples you have. How to test that claim is in RT-qPCR for relative expression. A reference gene is not a column that RNA-seq is required to contain.

Terms from a count matrix

Counts. A count is an integer: reads or fragments assigned to a feature by a stated pipeline. The mistake is to average counts from libraries of different depth and call the average an expression. Another mistake is to treat a count as a molecule count when no unique tag was collapsed and no calibration was performed. Zero means no read was assigned. It does not, on its own, mean the cell contained no transcript.

TPM. Transcripts per million scale each gene by length and then scale the sample so the length-adjusted values sum to one million. TPM answers a composition question inside one sample: what share of the transcript mass this gene represents, with a length correction so long genes are not automatically louder. The mistake is to use TPM as if it were a normalised count for a differential test, or to compare TPM across samples and ignore that a change in one abundant gene moves everyone else's share.

FPKM. Fragments per kilobase per million mapped fragments are the older paired-end cousin of that idea. The mistake is the same family: treating FPKM as an absolute concentration, or summing it across genes and expecting a meaningful total. Prefer to say which scaling you used rather than to treat FPKM and TPM as interchangeable labels for "normalised".

Normalisation. Normalisation is whatever scaling you chose so that a comparison is fair under a stated assumption. Library size, a median ratio, a trimmed mean of fold changes, and a reference-gene subtraction are all normalisations, and they are not the same operation. The mistake is the sentence "we normalised the data" with no method. If the assumption is wrong, the scaling invents differences. Name the method in the same sentence as the result.

Terms about belief and design

FDR. The false discovery rate is the expected share of rejected null hypotheses that are false, as estimated by a procedure such as Benjamini-Hochberg. It is a property of the list you rejected, at the threshold you set. The mistake is to read an adjusted p of 0.04 as "a 4 percent chance this gene is wrong" for that row alone, or to say FDR when only raw p values were computed. Effect size lives in the fold change. The FDR does not replace it.

Biological replicate. A biological replicate is an independent biological unit: a separate animal, donor, plant, or independently established culture. The mistake is to label three wells from one flask as n equals 3. Those wells estimate pipetting and instrument noise. They do not estimate how the next flask would behave. A technical replicate is still worth running. It is not the unit of a treatment claim.

Strandedness. Strandedness is whether the library chemistry preserved the direction of the original RNA. The mistake is to analyse a stranded library as unstranded and throw away overlaps you paid to separate, or to set the orientation flag backwards so antisense is counted as sense. Unstranded is a real class, not a synonym for broken. The chemistry has to be in the sample note.

UMI. A unique molecular identifier is a random tag attached to a nucleic acid molecule before amplification, so later copies can be collapsed to one. The mistake is to believe the tag erases every bias. It stops PCR duplicates being counted as depth. It does not recover molecules that were never captured, and a short tag will be shared by chance among abundant molecules. Tags added after PCR do not identify original molecules.

Gene models that these counts refer to should be a named release. Ensembl is one place to check that the identifier in the table still means the locus you think it means.

Classes of reagent and file these words point at

Cq, RT, efficiency and reference gene point at a reverse transcriptase, a qPCR mix, oligonucleotides and a real-time instrument. Counts, TPM, FPKM, normalisation, FDR, strandedness and UMI point at a library kit, a sequencer and a quantification file. The reagent classes live in the molecular biology catalogue. Neither class is made more exact by using a word from the other without a conversion you can defend.

A short-read technology overview from Illumina is useful background for the sequencing words. It does not define Cq.

Three expression terms that are not interchangeable Cq one assay, one threshold is not TPM share inside one sample is not raw p not an FDR Write the unit on the figure. A reader cannot recover it from the word expression.
Cq, TPM and a raw p value answer different questions, so none of them is a synonym for the others.

A disagreement has a branch

When two people cite different numbers for the same sample, do not average them. Ask which word each person used.

If one has a Cq and the other has a fold change, the fold change already subtracted a reference and compared a calibrator. The Cq did not. Go back to the wells.

If one has TPM and the other has a count, they are answering share-within-sample versus assigned reads. Choose the question, then the column. Do not plot them as replicates of one measurement.

If one quotes a raw p and the other quotes an adjusted p, only one of them faced the number of tests. Use the procedure that matches the list you intend to claim. If the list is long and only raw p values exist, the branch is to run an adjustment or to narrow the claim to a predeclared set of genes. It is not to relabel the raw column.

If one count collapsed UMIs and the other did not, the depths are not comparable. Say so in the legend.

Word, use, and the mistake beside it

WordHonest useUsual mistake
Cq or CtCrossing cycle for this assay and thresholdRead as a concentration or compared across genes raw
RTReverse transcription that made the cDNAConfused with room temperature, or skipped as a failure point
Primer efficiencyDoubling behaviour of this pair in this mixAssumed perfect in an inhibited sample
Reference geneA transcript shown to be stable hereA famous symbol copied from another tissue
CountsAssigned reads under a named pipelineAveraged across unequal depths
TPM or FPKMLength-scaled compositionUsed as counts in a differential test
NormalisationA named scaling with a stated assumption"Normalised" with the method omitted
FDRError rate of the rejected listRead as the odds that one gene is false
Biological replicateAn independent unitWells from one flask called n equals 3
StrandednessDirection kept, and which orientationFlag omitted or copied from another kit
UMITag used to collapse pre-tagged copiesTreated as proof the library is unbiased

Loose language, wrong sentence

The failures are clerical and then biological. A figure axis labelled "expression" will be quoted as a concentration. A methods paragraph that says FDR when the software column was a raw p value will be over-trusted. A stranded library described only as RNA-seq will be reanalysed without direction. None of these require a bad reagent. They require a word used one size larger than the measurement.

Write the word at the size it earns. "Relative Cq difference against a reference we checked" is longer than "expression" and much harder to misuse.

Research vocabulary is not a diagnostic claim

These terms do not approve a medical test, a reference interval, or a biosafety level. Samples from people or from infected sources follow the containment your institution sets. A small adjusted p value is not clinical significance. Enzymes and dyes follow the safety notes on the products you actually open.

Putting the words in a shared brief

Laboratories that share a study need a one-page specification that fixes the words before the figures exist. Which unit will be exchanged, which identifier release, whether FDR will be computed, what a biological replicate is in this design, and whether strand and tags are in scope. A warm shipping leg or a power cut belongs in the sample history, not inside the definition of Cq. The glossary stays stable. The sample note records the exception.

That specification is also how a later reader avoids redefining TPM in a footnote. Agree the list when the design is written. Changing a definition after the plot is made is how two true statements become an argument.

What to name in an enquiry

Say which of these words you need in the deliverable. A Cq table, a count matrix, a differential table with a named adjustment, or a length-scaled abundance are different requests. Sequencing scope can be discussed from the mRNA sequencing enquiry reference. The statistical words in the table can be discussed from the differential expression analysis enquiry reference. Use the pages to frame the method. Name the unit you want back.

The sample-to-result context is the nucleic acid analysis pathway. Send the word list with the quote request so the reply can say which method class produces it.

Pin an expression result to the words it actually used

  1. 01Name the unit before any two numbers are comparedWrite whether the figure shows Cq, a fold change from a reference gene, raw counts, or a length-scaled abundance such as TPM. If the unit is missing, stop the comparison. Two columns with different units can be plotted on one axis and still be meaningless.
  2. 02Separate the biological unit from a technical splitState what was independently cultured, collected or treated. A second well from the same reverse transcription is a technical replicate. Call it that in the legend so a reader does not treat pipette repeats as the sample size.
  3. 03Match the uncertainty word to the procedureIf many features were tested, the word that belongs beside a threshold is the false discovery procedure you ran, not a raw p value copied from a single-test habit. If no adjustment was run, say so rather than borrowing the abbreviation FDR.
  4. 04Record strand and tags if the data are sequencingWrite whether the library was stranded and whether unique molecular identifiers were collapsed. A count that still includes PCR copies is a different object from a molecule count. The glossary words only help if the figure uses them in that strict sense.

Questions from the bench

Is Cq the same thing as Ct?

They are two names for the cycle at which a real-time signal crosses a chosen threshold. Cq is the more general term. Ct is the older one. The mistake is not the choice of letters. The mistake is treating either number as a concentration, or comparing it across genes without an efficiency and a reference.

Can TPM values be used as the input to a count-based differential test?

No. TPM is a within-sample composition scaled to a million after a length adjustment. Count-based tests expect integer assignments and a variance model that matches counts. Feeding them TPM double-transforms the data and hides the sampling noise. Keep TPM for composition questions and counts for the test.

Is an adjusted p of 0.05 the probability that this one gene is a false positive?

No. A false discovery rate describes the list of rejections, not a personal probability for a single row. An adjusted value near a threshold means the gene sits at the edge of the procedure you chose. It is not a 5 percent chance that the biology is wrong, and it is not a measure of effect size.

Does a unique molecular identifier make a low-input library unbiased?

It lets you collapse PCR copies of a molecule that was tagged before amplification. Bias from capture or from reverse transcription before the tag remains. A tag that is too short also collides, so two molecules can be counted as one. The identifier fixes double-counting of copies. It does not fix the library.

References

  1. MIQE guidelines for quantitative real-time PCR
  2. Ensembl genome browser
  3. Illumina overview of next-generation sequencing

Manufacturer names identify published method classes. Trademarks remain with their owners. Catalogue records on this site are independent references for enquiry. They are not a statement of inventory, distribution rights or a supply commitment. This page is educational. It is not medical advice, a diagnostic protocol or a biosafety approval.