guide
Pathway lists and the multiple-testing problem
How testing thousands of genes creates a multiple-testing problem, and why a pathway table built on an already chosen list is not yet a mechanism.
- Author
- EVRINTH Editorial Team
- Published
- 8 October 2026
- Updated
- 8 October 2026
- Reading time
- 10 min

Testing one gene at a threshold of 0.05 is a statement about that gene. Testing a transcriptome at the same threshold is a statement about a crowd. Pathway tools then take a list from that crowd and test gene sets, sometimes using the very selection that created the list. The long table that comes back is easy to read as a mechanism. It is usually a summary of overlap. This guide shows how to keep those layers apart. The experiment that produced the counts is from cells to a gene expression result.
Who has to make the call
The person who will write "we observed activation of" a pathway name needs this distinction, and so does anyone asked to approve that sentence. The decision is what the differential table supports, what the enrichment table adds, and what still requires a new experiment. Analysis scope of this kind can be discussed through the differential expression analysis enquiry reference. Bring a contrast, not a hope that the software will choose the biology.
What a single p-value was built to do
A p-value, in the usual sense, is the probability of a result at least as extreme as the one you saw, if that gene's null hypothesis were true and the model held. It is not the probability that the null is true, and it is not an effect size. A gene can have a tiny p-value and a fold change too small to matter in your system, or a large fold change and a p-value that the replicates cannot support. Report the estimate and the uncertainty beside any threshold.
Thresholds become a different instrument when you apply them many times. A planning picture used throughout transcriptome work is a mammalian gene count on the order of twenty thousand features. If every null were true, and each test were independent and well calibrated, a per-test cutoff of 0.05 would be expected to mark about one in twenty features, on the order of a thousand calls, even though nothing changed. Real studies are messier than that sketch: tests are correlated, some nulls are false, and effect sizes vary. The sketch still explains why an unadjusted p-value column is a poor place to stop.
False-discovery-rate methods, including the Benjamini-Hochberg procedure, aim at a different error. They try to control the expected proportion of false calls among the calls you reject. A q-value or an adjusted p-value is the number you read when that is the error you care about. Saying "significant" without saying which error rate you controlled leaves the reader guessing. The biological n behind the test is still the biological replicate, as in biological versus technical replicates. An adjusted p-value computed on pseudoreplicated wells does not become trustworthy because it was adjusted.
Two kinds of pathway question
Gene-set collections group symbols under names: signalling maps, disease labels, transcription-factor targets, cellular components. The names are curated, incomplete and overlapping. A gene can sit in dozens of sets. Enrichment software will report many of those sets together because they contain the same genes. That is redundancy, not twenty independent mechanisms.
Over-representation analysis starts from a list you already cut, often genes that passed an FDR bar, and asks whether a set is more common in that list than in a background. The background matters. Using the whole genome when you only measured expressed genes inflates some results. Using the list itself as if it were a new sample of genes double-counts the first test. This is the double-dip: the pathway p-value is not an independent replication of the differential expression. It is a rearrangement of a list the differential expression already selected. It can still be a useful description of that list. It cannot be written as confirmation.
Rank-based methods, in the family of gene-set enrichment that uses every gene's statistic from strongest to weakest, do not require a hard cutoff. They ask whether the members of a set tend to sit toward one end of the ranking. That is a better match to some questions, and it is still not a mechanism. It inherits the ranking's batch effects, annotation errors and model mistakes. Competitive nulls and self-contained nulls answer different statistical sentences. If you do not know which one the tool used, you do not yet know what "enriched" means in your file.
A pathway diagram in a slide is a further view. Like a heatmap, it can be arranged to look inevitable. The discipline in a heatmap is a picture, not a conclusion applies here. The table of gene-level results remains the result. The set table is a summary you interpret with the overlap in mind.
A workflow that keeps the layers labelled
Define the contrast before you see the pathways: which groups, which covariates, which features are in the universe. Run the differential model and save the full table, including genes that do not pass. Choose an error rate and an effect-size filter you can defend, and produce the primary gene list from those rules.
Only then run enrichment, and write down the tool, the gene-set collection, the background and whether the method saw a cutoff list or a ranking. When the table is long, cluster the redundant set names or note the genes that drive many of them. A set that lights up only because one famous gene is in it is a different finding from a set whose members move together.
Branch when the story is too smooth. If every significant pathway is a rewording of "immune response" carried by the same dozen genes, you have one signal and many labels. If the top pathway disappears when you change the background or the FDR cutoff within a range you consider reasonable, it was fragile. If the leading genes are poorly annotated in your species, check them on Ensembl or GenBank before you let a human gene-set name speak for them. Orthologue mapping errors create confident pathways about the wrong genes.
A follow-up that can carry weight is a short RT-qPCR panel, or a perturbation, chosen for a reason you can state. The qPCR needs its own reference genes and efficiencies, as in RT-qPCR for relative expression. Re-running a second enrichment website on the same genes does not add a biological replicate. Public count archives such as the Sequence Read Archive are places you might look for a similar contrast. A match there is still someone else's design, batches included.
| Object | Question it answers | Easy way to over-read it |
|---|---|---|
| Raw p-value | How surprising is this gene under its null? | Treating 0.05 as a transcriptome-wide pass |
| FDR or q-value | What false-call share do you expect in the kept list? | Calling it the probability a given gene is true |
| Log fold change | How large is the estimate? | Ignoring it because the adjusted p-value is small |
| Over-representation | Is this set common in an already selected list? | Calling the set p-value an independent confirmation |
| Rank-based enrichment | Do set members gather at one end of the full ranking? | Treating the enrichment as proof of a mechanism |
| Pathway cartoon | How you chose to draw the set | Hiding genes in the set that did not change |
A long table is still a list of names
Failure looks like a discussion section that cites twenty pathway names and no gene. Readers cannot tell whether the signal is broad or whether five interferon-stimulated genes are dragging every set that contains them. Count how many unique genes do the work. Check the direction: a set can be "enriched" when half its members go up and half go down, which is a poor match for a sentence about activation. Check the contrast: enrichment on a batch-confounded comparison enriches the batch.
Another failure is shopping. If you lower the FDR cutoff, switch collections, and change the background until a favoured pathway appears, the procedure is no longer the one you would have reported at the start. Pre-state the collection and the error rate, or label the search as exploratory and keep it out of the abstract.
Research limits, not a licence to name a disease mechanism
Pathway results from a research model support hypotheses about that model. They do not diagnose a person, and they do not show that a drug hit the pathway you named. Human expression data remain under the ethics approval and the biosafety rules of the institution that holds the samples. This guide does not supply either one. A gene-set name borrowed from a human disease collection is particularly easy to over-read when the samples were a cell line.
Writing the statistical claim into the specification
When you ask another group to analyse a matrix, the multiple-testing rule and the gene-set collection are part of the specification, in the same way a buffer is part of a bench protocol. Say whether you want FDR control, what contrast defines the ranking, and whether enrichment is in scope. In a laboratory that queues samples across a hot season, also say which batch column must stay in the model, so a pathway list does not describe a shipment. A clear specification is shorter than a retraction of a mechanism sentence.
What to put on the enquiry
Name the species, the contrast, the biological n, the error rate you care about, and whether a pathway summary is wanted as a description or is out of scope. Sequencing reagent classes sit in the genomics and sequencing catalogue. The sample path is the nucleic acid analysis pathway.
The mRNA sequencing enquiry reference and the differential expression analysis enquiry reference are enquiry references for the measurement and the model. Send the contrast and the error-rate preference through the quote request and ask whether a quotation is possible. A pathway summary can be discussed as a description of a tested list. The mechanism, if you need one, is a further experiment you specify yourself.
Read a pathway table without promoting it to a mechanism
- 01Name the contrast and the universe of testsWrite which groups were compared and how many features were tested. A p-value threshold applied to one gene does not mean the same thing when it is applied to a whole transcriptome. Record the multiple-testing procedure you will trust.
- 02Separate the differential table from the enrichment tableKeep the gene-level estimates, with their fold changes and adjusted values, as the primary result. Treat a pathway table as a summary of gene sets inside that result, and note whether the enrichment reused a list you had already selected.
- 03Ask what the enrichment statistic actually testedOver-representation asks whether a pre-cut gene list contains more members of a set than a background would suggest. A rank-based method uses the full ordering of genes and asks a different question. Neither test demonstrates that the pathway operated as a mechanism in your cells.
- 04Decide what would count as a follow-upA follow-up is a new measurement aimed at a pre-stated gene or perturbation, such as an RT-qPCR panel designed before you re-read the list. Another enrichment tool on the same list is a second view, not an independent confirmation.
Questions from the bench
Why is a p-value of 0.05 a weak bar for twenty thousand genes?
A p-value is calibrated to one test. If every null were true and the tests behaved as the threshold assumes, a cutoff of 0.05 would flag about one test in twenty. Across a mammalian-scale transcriptome that is a large list of expected false calls. An FDR procedure asks a different question: among the calls you keep, what share do you expect to be false.
Is a q-value the same object as a p-value?
Both are probabilities in a broad sense, and they answer different errors. The p-value refers to a single test under its null. An FDR-adjusted value, often shown as a q-value or an adjusted p-value, refers to the list of tests you are willing to call significant together. Reporting one in the column meant for the other changes the claim.
Our pathway tool returned fifty significant pathways. Which mechanism is real?
The table says those gene sets overlap the statistical result more than the tool's null expected. Sets overlap each other, names are curated, and many pathways share the same famous genes. The list is a hypothesis menu. A mechanism needs an experiment that perturbs or measures the biology, not a second sort of the same genes.
Does confirming the top genes by RT-qPCR validate the whole pathway?
It checks those transcripts with a second method, if the primers, reference genes and efficiencies are sound. It does not show that every member of the pathway changed, and it does not show that the pathway caused the phenotype. Phrase the confirmation as the genes you actually remeasured.
References
Manufacturer names identify published method classes. Trademarks remain with their owners. Catalogue records on this site are independent references for enquiry. They are not a statement of inventory, distribution rights or a supply commitment. This page is educational. It is not medical advice, a diagnostic protocol or a biosafety approval.
Catalogue
Related products and categories
These links follow the subject of the article into published manufacturer references. A listing is a reference for an enquiry, not a statement of stock or distribution rights.
Continue in this cluster
Related reading
From cells to a gene expression resultHow a laboratory goes from cells to a gene expression result, and how RNA quality decides between a focused RT-qPCR assay and RNA-seq.
A heatmap is a picture not a conclusionHow clustering, colour scale and row scaling change a gene-expression heatmap, and why the differential table and the contrast remain the result.
Biological versus technical replicatesHow to tell a biological replicate from a technical one, and which unit of inference actually supports a treatment claim in an expression study.