Skip to content
EVRINTH

application

What a differential expression table contains

Which columns in a differential expression table a biologist should question, including fold-change sign, adjusted p and identifier collisions.

Author
EVRINTH Editorial Team
Published
8 October 2026
Updated
8 October 2026
Reading time
9 min
Researcher viewing gene expression heatmaps and genomic tracks on two monitors at night
Researcher viewing gene expression heatmaps and genomic tracks on two monitors at night

A differential expression table is a list of arguments, one row per feature, and most of the argument is in the column headers. A biologist who can question those columns will notice a missing adjustment, a silent contrast, or a gene symbol that joined the wrong locus. A biologist who sorts on the first interesting-looking number will defend a ranking the table did not actually make. This page is the application of that habit. Where the table sits in the path from cells to a claim is covered in from cells to a gene expression result.

A row is a claim

Each row says, in effect: under this annotation, this contrast, this scaling and this test, here is how we summarise the evidence. The row does not say the protein changed, and it does not say the transcript is a validated marker. It says the counts, as processed, produced these statistics. Questioning a table means asking whether that sentence is the sentence you meant to write.

The same spreadsheet can be honest or misleading without a single corrupted cell. The usual trouble is a header that was abbreviated until the contrast, the logarithm base, and the adjustment disappeared.

The numbers and what they are allowed to mean

The gene identifier is the join key. For many animal studies it is an Ensembl gene id, sometimes with a version suffix after a dot. A transcript identifier is a different object. A gene symbol is a label for humans. Stable identifiers from a named release, which you can check on Ensembl or in the UCSC Genome Browser, survive a synonym change. Symbols do not.

The log fold change is a ratio on a log scale, almost always base 2 in RNA-seq tools, but the table should say so. The sign depends entirely on which group was coded as the numerator. A value near zero means the model did not find a shift in the mean. A very large value from a gene with tiny counts is often an unstable ratio. Some tables report a shrunk log fold change that pulls those unstable ratios toward zero. An unshrunken column and a shrunk column will not rank genes in the same order. Know which one you sorted.

Average expression is the abundance the test used as context. In one widely used count model it is a mean of normalised counts across the samples in the contrast. In other software it has another name and another scaling. Low average expression means the gene was barely measured. A spectacular ratio at a tiny average is a weak place to build a story. The column does not, by itself, say whether a reference gene was stable. Reference genes are a qPCR concept that people sometimes expect to see here. They are not a required column. If someone ranked RNA-seq results by a housekeeping symbol, they imported a habit from a different assay.

The p value is the result of the test for that row under the model's assumptions. It is not the probability that the gene is biologically important. With thousands of rows, many small p values appear even when every null hypothesis is true.

The adjusted p value is the column that faces that fact. A Benjamini-Hochberg procedure, or a related false discovery rate method, is the usual class. A table of fifteen thousand genes and only raw p values will be full of false hits if you threshold it as though one test had been run. Missing adjusted p values need a reason, as the questions above spell out. A blank is sometimes "not tested after filtering", which is different from "tested and not small".

Counts, when they are present, should say whether they are raw assigned reads or already scaled. A zero raw count means this pipeline assigned no read to that feature in that sample. Downstream, some tools refuse a test on all-zero rows, and some imputation or shrinkage steps move the reported ratio away from infinity. Do not read a zero as a calibrated absence inside the cell. Do not take a logarithm of zero without the offset the methods actually used.

Tables come from method classes

The wet-lab side is a library and a sequencer. The numerical side is a quantification class and a statistical class. Alignment against a genome, alignment against a transcriptome, and pseudo-alignment are different ways to produce a count. A negative-binomial count model and a linear model on transformed values are different tests. This page does not pick a software package for you or paste its parameters. It asks you to keep the class name beside the file.

Filtering is part of the method. Independent filtering may remove low-count genes before adjustment, which is one reason adjusted p values go missing. If you restore those rows by hand and reuse the old adjusted values, the adjustment no longer matches the tests that remain.

Normalisation decides the size factors that enter the model. Change it, and both the average expression and the ranking can move. Treat "the data were normalised" as an incomplete sentence.

Columns a biologist should be able to question gene id log fold change average expression p adjusted p present? counts, if kept Sign of the log depends on which group is the numerator. A blank adjusted p is untested or unadjusted. It is not a result. Join on the stable gene id. A symbol is a label, not a key.
A differential expression row is unreadable until the contrast sign and the adjusted p column are both present and defined.

Walk one gene before you believe the list

Pick a gene you understand biologically and find its row before you sort the whole file. Confirm the identifier is from the species you sequenced. Read the contrast aloud. If the log fold change is positive, say which group that calls higher, using the design note rather than the colour on a heatmap. Check the average expression. If the gene was barely detected, do not promote it because the ratio looks dramatic.

Check the adjusted p value. If it is missing, find out whether the gene was filtered. If the column does not exist at all, the table is not ready for a threshold. Then look at the counts across samples, not only the summary. A single outlier sample can create a p value. The summary columns will not show you that until you look.

If that one gene makes sense, look at the distribution of the rest. If it does not, the pipeline or the sample sheet is the problem, and sorting will only produce a longer version of the same mistake.

Columns worth questioning

ColumnQuestion to ask before you sort on itA bad reading
Gene idWhich release, and is this a gene or a transcript?Joining on symbol and merging two loci
Log fold changeWhich group is the numerator, which base, shrunk or not?Calling every positive value "up in treatment" by habit
Average expressionIs this gene actually measured?Trusting a huge ratio from almost no counts
p valueWas this one test or one of many thousands?A raw threshold after a genome-wide search
Adjusted pWas adjustment run, and were filtered genes left blank?Treating a blank as not significant, or dropping the column
CountsRaw or scaled, and what does zero mean here?A zero reported as a concentration of none

Four ways a tidy table still lies

The adjusted column is missing and a raw threshold is applied anyway. With a transcriptome-sized list, that threshold does not mean what it means in a paper that tested four genes.

The sign is read from memory. Someone coded control as the numerator months ago. Every "up" gene in the slide is down, and the pathway story is the reverse of the samples.

Zeros are narrated as biological off-switches. A gene with zero counts in one group and a handful in the other can be a real difference or a sampling miss, especially in a thin library. The table's job is to show the uncertainty. Replacing zero with a confident sentence removes it.

Symbols collide. One symbol maps to more than one Ensembl gene, version suffixes fail an exact join, and a mouse identifier is pasted into a human table because the letters looked familiar. Spreadsheets also rewrite some gene symbols as dates. The identifier that went in is not the identifier that comes back. Keep the stable id in a text column and check a handful of rows against the genome browser after any round trip through a spreadsheet.

A deposited experiment in the European Nucleotide Archive without the contrast and the annotation release cannot be re-read from the count file alone. Your table should carry those facts in a header or a sibling metadata file so it does not become that kind of orphan.

The table is not a medical report

A research differential expression table is not a diagnostic panel and not a biosafety determination. Human samples and infectious material follow institutional containment rules whether or not a p value is small. Do not present adjusted p values to a collaborator as clinical significance. The word significant, if you use it, belongs to the statistical procedure you named, at the threshold you set, for the contrast you defined.

Files moving between laboratories

When a table leaves one site for another, the failure mode is clerical. Locales change decimal marks. Spreadsheet software guesses types. A collaborator filters to "the significant tab" and the blank adjusted p values vanish without a note. Send a delimited text file plus a short note that states the contrast, the logarithm, the adjustment, the annotation release and the meaning of a blank. Ask the receiver to join on the gene id and to count how many rows fail the join. A failed join is a finding about the file, not a nuisance to hide.

That note is also the specification you wanted at the start. Analysis scope belongs in the same place as the library chemistry. A heatmap picture without the columns above is not the deliverable.

What to send when you want a table discussed

Name the organism, the annotation release, the contrast, whether you need counts only or a full differential table, and how missing adjusted values should be handled. Sequencing can be discussed from the mRNA sequencing enquiry reference. The shape of the table and the statistical class can be discussed from the differential expression analysis enquiry reference. Those pages are a prompt for the method conversation. They do not mean a pipeline has already been chosen for you.

Library reagents are a catalogue class under genomics and sequencing. The path around the table is nucleic acid analysis. Put the contrast in writing and send it with the quote request.

Questions from the bench

What does a missing adjusted p value mean?

It means this row was not given a multiplicity-adjusted result. In some pipelines that happens because the gene was filtered for low count and was never tested, so the blank is not a quiet non-significant call. It can also mean the adjustment was never run. Those two causes are different, and the methods sentence has to say which one you are looking at.

Which group is higher when the log fold change is positive?

The numerator of the contrast is higher. Positive and negative have no biological meaning until someone writes treated-versus-control or the reverse. A methods note that says only log fold change, without the contrast and without the base of the logarithm, cannot be read safely.

Can I join two result tables on gene symbol?

Not if you need a one-to-one match. Symbols collide, change between releases, and can point at several stable identifiers. Join on the versioned gene identifier from a named annotation, then display the symbol for reading. A symbol-only join will merge distinct loci and drop genes whose names differ by a synonym.

Does a count of zero mean the transcript is absent from the cell?

It means no read was assigned to that feature by this pipeline. The gene may be absent, the RNA may have been too scarce to sample, or the annotation may not match the transcript you care about. Zero is a statement about the file. It is not a measured concentration of nothing.

References

  1. Ensembl genome browser
  2. UCSC Genome Browser
  3. European Nucleotide Archive

Manufacturer names identify published method classes. Trademarks remain with their owners. Catalogue records on this site are independent references for enquiry. They are not a statement of inventory, distribution rights or a supply commitment. This page is educational. It is not medical advice, a diagnostic protocol or a biosafety approval.

Catalogue

Related products and categories

These links follow the subject of the article into published manufacturer references. A listing is a reference for an enquiry, not a statement of stock or distribution rights.