glossary
False discovery rates and decoy searches
A decoy search estimates the false discovery rate of a whole accepted list. It is not the probability that one protein is true, and it does not repair a wrong
- Author
- EVRINTH Editorial Team
- Published
- 8 October 2026
- Updated
- 8 October 2026
- Reading time
- 8 min

A false discovery rate is an estimate about a list. A decoy search is the usual way proteomics builds that estimate. The sentence people hear, that a reported protein is ninety-nine percent true, is a different sentence and usually a wrong one. Bottom-up proteomics in plain language introduces the idea. This glossary page says what the number estimates, what a one-percent habit is for, and what decoys leave untouched.
How a decoy enters the competition
The search engine compares fragment spectra with candidate peptides from a target database, often a UniProt proteome plus contaminants. Beside those real sequences it scores decoys: the same sequences reversed, or shuffled so they keep a similar composition without remaining real proteins. A decoy should not be the biological source of the spectrum. When a decoy nevertheless outscores the targets, the match is a picture of chance.
You then accept every match above some score. Among the accepted matches, some are targets and some are decoys. The false discovery rate estimates the fraction of that accepted list which is decoy-like, and therefore wrong. Laboratories often compute it as a function of the decoy count and the target count above the line. The exact formula varies by tool. The meaning does not: it is a property of the list you kept, at the threshold you chose.
Raise the score line and the list shrinks. Fewer decoys remain, and the estimated rate falls. You also throw away true peptides that happened to score modestly. That trade is the whole point of choosing a threshold. There is no line that keeps every true peptide and no false one.
Peptide lists and protein lists are different counts
A peptide false discovery rate controls the peptide matches. A protein false discovery rate controls the proteins you infer from those peptides. The second problem is harder because of the way peptides are grouped. Several weak peptides can nominate one protein. Shared peptides can nominate a whole family. A parsimonious grouping tries to explain the observed peptides with the smallest honest set of proteins, and unique peptides are the ones that pin a sequence down. Change the grouping rules and the protein list changes, even if every peptide score stays still. The protein-level error rate moves with that choice.
Report both levels when you have them, and do not let a reader apply the peptide threshold to a protein table. A protein supported by one modest peptide is not in the same position as a protein supported by many peptides that each passed a strict peptide threshold. Say which evidence you required. One unique peptide, two peptides, or a protein-level error rate: pick a rule and write it down. Switching the rule after you have seen which proteins appear is how a false discovery rate becomes decorative.
A one-percent habit is not a statute
In discovery proteomics, a one-percent protein false discovery rate is a common reporting threshold. It is a convention that makes lists comparable in conversation. It is not a law, not a journal's universal rule, and not the right bar for every decision. A targeted follow-up on a short candidate list can justify a different standard, declared in advance. A claim that will be expensive to chase may deserve a stricter list. A search of a tiny database, such as a single recombinant protein plus contaminants, behaves differently from a search of a whole proteome, and a one-percent figure there can mean something quite fragile. State the database size class and the threshold together.
Some software also reports a posterior error probability for an individual match. That number tries to say how likely this one match is wrong. It is useful and it is not the false discovery rate. Quoting one in the sentence where the other belongs is how talks become unreproducible. ProteomeXchange deposits are easier to reuse when the threshold and the level are written in the method rather than implied by a software default.
The drawing is a sketch of the logic, not a scale of any real search. Decoy bars that cross the line are the ones that warn you. Target bars below the line are the true peptides you chose to leave behind.
Words, and what people wrongly hear
| Term | What it estimates | What people wrongly hear |
|---|---|---|
| Target-decoy search | How often chance matches, modelled by reversed or shuffled sequences, pass the same scoring | That every decoy hit was a real protein from the sample |
| False discovery rate | The estimated fraction of the accepted list that is wrong, at this threshold | The probability that one named protein is false, or that it is ninety-nine percent true |
| Peptide-level rate | Error among accepted peptide-spectrum matches | Error among proteins |
| Protein-level rate | Error among inferred proteins, after grouping | A stricter peptide filter, automatically |
| Posterior error on one match | A tool's estimate for that match alone | The false discovery rate of the whole table |
What decoys cannot repair
A wrong database is invisible to the decoys built from it. Searching a human proteome against a bacterial lysate returns human-looking matches and a false discovery rate computed entirely inside that mistake. Searching without keratin and without the digestion enzyme hides contaminants inside sample proteins that happen to share short stretches, or it forces them onto the nearest sample sequence. Add the contaminant sequences, then decoy the enlarged database.
A wrong modification list is the same kind of miss. If the spectra are full of a chemical side reaction you did not allow, the engine explains them with ordinary peptides that are slightly wrong, and the decoys judge that wrong explanation. Allowing every modification you can think of is the opposite mistake: the search space grows, decoys pass more easily, and the rate you report may be honest about a question you made too loose. Sample preparation belongs in this paragraph because the modifications and the contaminants are created at the bench. The statistics do not wash the tube.
Grouping rules can also manufacture certainty. If you count a protein as identified from any shared peptide, you will report a family as if it were a gene. Parsimony and a demand for unique peptides shrink that list. They can also hide a real protein that only produced shared peptides. Say which choice you made. The false discovery rate does not choose it for you.
Standards homes
The Human Proteome Organization and the Proteomics Standards Initiative are where identification guidelines and file standards are developed in public. Read them when you need the community's language. Quoting the organisations in a report does not certify the report. The certifying act, such as it is, is local: database version, decoy method, threshold, peptide or protein level, and grouping rule, written where a reader can find them.
Research claims, stated at list strength
A one-percent list is still a research list. It will contain mistakes at about the rate you accepted, and you do not know which rows they are. It does not diagnose disease, assign function, or authorise a clinical decision. Biosafety of how the sample was generated is an institutional matter and is untouched by the score line. Treat the threshold as a claim about error in bulk. Treat any single row you intend to spend a year on as a row that needs more evidence than the list threshold.
Write the threshold into the specification
Two groups sharing an instrument can both say they "used one percent" and still be incomparable, because one controlled peptides and the other controlled proteins, or one grouped isoforms and the other did not. In the analysis specification, name the level, the numerical threshold, the decoy style, and the database build. How to write a laboratory sourcing enquiry is a useful pattern for putting those sentences where a collaborator will see them before the search is run. Changing the threshold after the proteins of interest are known is a different study. Call it that.
What to ask for in an identification enquiry
Database, species, contaminant set, modification classes, missed-cleavage allowance, decoy method, and the peptide and protein thresholds you want reported. The protein identification by LC-MS/MS reference is the method page closest to this glossary. The shotgun discovery proteomics and differential abundance references cover the study around the list. Send the threshold language with the quote request. A method can be discussed at the strength of the list you are willing to defend. A request for "significant proteins" is not yet a false discovery rate.
Questions from the bench
Does a one-percent false discovery rate mean each protein is ninety-nine percent certain?
It means that, among the accepted matches at that threshold, about one percent are estimated to be decoys and therefore wrong. The estimate belongs to the list. A single protein at the top of the list is not assigned a personal probability of ninety-nine percent by that sentence. Some tools also report a posterior error for one match. That is a different number. Do not translate one into the other in a talk.
Why can the peptide false discovery rate and the protein false discovery rate disagree?
Peptides are scored one by one. Proteins are groups of peptides, and a group can gather several mediocre matches into something that looks identified. A threshold that is honest for peptides can be too loose for proteins, or the other way around, depending on how groups are built. Report the level you actually controlled. A slide that says one percent without saying peptide or protein has not reported the threshold.
Will a stricter decoy threshold fix a search against the wrong species?
No. Decoys estimate how often chance matches pass your score cutoff inside the database you provided. If that database is the wrong organism, or it lacks the contaminant sequences you injected, the true sequences are simply absent. The search will still return the best wrong answers and a tidy false discovery rate against its own decoys. Change the database, then estimate the error again.
Where do HUPO and the Proteomics Standards Initiative sit in this?
They are public homes for how the field argues about identification standards and data formats. Citing them shows you know where the conversation lives. It does not certify a particular search, a particular software default, or a particular laboratory. Your report still has to say which decoy method and which threshold you used.
References
Manufacturer names identify published method classes. Trademarks remain with their owners. Catalogue records on this site are independent references for enquiry. They are not a statement of inventory, distribution rights or a supply commitment. This page is educational. It is not medical advice, a diagnostic protocol or a biosafety approval.
Catalogue
Related products and categories
These links follow the subject of the article into published manufacturer references. A listing is a reference for an enquiry, not a statement of stock or distribution rights.
Continue in this cluster
Related reading
Bottom-up proteomics in plain languageBottom-up proteomics in plain language: proteins are digested to peptides, a mass spectrometer fragments them, and a database search names candidates.
How to write a laboratory sourcing enquiryWrite a laboratory enquiry as measurable performance, documents and an acceptance check, so a quotation can be compared with the scientific need.
A protein ID is not a proof of functionA proteomics identification is a sequence match under a false discovery rate. Function still needs activity, phenotype, localisation or a perturbation, plus