application
A protein ID is not a proof of function
A proteomics identification is a sequence match under a false discovery rate. Function still needs activity, phenotype, localisation or a perturbation, plus
- Author
- EVRINTH Editorial Team
- Published
- 8 October 2026
- Updated
- 8 October 2026
- Reading time
- 8 min

An identification is a sequence match that survived a false discovery threshold. Function is a claim that the protein does something: an activity you measured, a place you localised it, a phenotype that moves when you perturb it. The two claims get stapled together in talks because the accession number comes with a name that already sounds functional. Bottom-up proteomics in plain language is the identification half. This page is an application walk-through of a pull-down list, and of how far that list is allowed to reach.
What the accession actually certifies
The search says: these fragment spectra fit peptides from this database entry better than they fit the decoys, at the threshold you set. UniProt and NCBI Protein will then tell you the name, the gene, and a summary of what other people have concluded about function. That summary is literature curated into a database. It is not a result from your tube. Quoting it is fair only if you quote it as background. Your result remains "peptides from this sequence were present in this sample".
Presence itself is qualified. Bottom-up data rarely cover the whole chain. You saw peptides. You inferred a protein. A contaminant keratin identification is a sequence match of the same logical kind, and it is not a biological claim about your cells. Background proteins that bind beads are sequence matches too. The false discovery rate does not label them as boring. You do that, with controls.
Groups, isoforms, and the gene name you wanted
Shared peptides make protein groups. If two isoforms contain the same peptide, the peptide supports the group. It does not elect a single splice form. Isoform-specific peptides are often short, modified, or simply never picked. A report that prints one gene symbol has usually chosen a representative accession by a rule: the one with most peptides, the reviewed entry, the first row. Ask for that rule. Then ask whether any peptide was unique to the isoform your story needs.
Parsimony tries not to list two proteins when one would explain the peptides. That keeps the table smaller and can hide a second protein that was truly present but only shared its peptides. Razor peptides, assigned to the group that already has more evidence, push intensity toward the winner. For identification, say "group". For a functional story about one isoform, show the distinguishing peptide or stop claiming the isoform.
NCBI Bookshelf chapters on protein function are a reminder that function, in cell biology, is argued from experiments that change a protein and watch a process. A table of accessions is not one of those experiments. HUPO guidance on human protein evidence is similarly cautious about what an identification contributes to a protein's known existence. Existence is still not mechanism.
A pull-down, walked through at the right strength
Imagine a bait protein on a bead, incubated with a lysate, washed, and digested. The list that comes back has three populations, and the analysis has to separate them before anyone draws arrows.
The bait itself should be abundant. If the bait is missing, the pull-down did not capture the bait, and every prey is a story about the beads. Sample preparation is the first control: was the fusion soluble, was the tag intact, did the wash leave any bait to digest?
Prey candidates are proteins enriched in the bait sample relative to a control that has beads without bait, or an unrelated bait of the same type. Enrichment can be qualitative in a first experiment, seen versus not seen, and it becomes a quantitative claim only with replicates and a fair comparison. Proteins that appear in both samples at similar evidence are background until a better wash or a better control says otherwise. Ribosomal proteins, chaperones and cytoskeletal proteins are classic background because they are abundant and sticky. Their appearance is expected. Their appearance as "the pathway" is a choice you must justify with the control.
The third population is junk from the bench: keratin, the digestion enzyme, antibody fragments if you used one. They belong on a contaminant list. They can be correctly identified and still be irrelevant to the bait. A blank digestion of the beads alone is the cleanest way to see them.
Only after that sort do you earn a short candidate list. That list is input to a hypothesis. It is not the hypothesis confirmed.
A pathway cartoon is a hypothesis
Taking the gene symbols that remain, dropping them into a pathway drawing, and colouring the ones you found is a picture of overlap with someone else's model. Overlap can be useful. It can also be forced: large pathways and sticky proteins overlap by chance, and the eye is kind to arrows that match a hope. The cartoon becomes evidence when a perturbation moves the process, when an activity assay shows the reaction, or when localisation shows the proteins meet. Until then, label the figure as a hypothesis generated by a pull-down. Readers can respect a hypothesis. They cannot check a claim that was drawn as if the mass spectrometer had measured causality.
Claim type, and the extra evidence it needs
| Claim you might want to make | What the identification already gave you | Extra evidence the claim needs |
|---|---|---|
| This sequence was in the sample | Peptides matched under a stated false discovery rate | A database that actually contains it, and a contaminant check |
| This isoform, not its siblings | Only if a unique peptide was observed | The distinguishing peptide, or an orthogonal isoform assay |
| The bait binds this prey | Co-presence in a digested eluate | A control matrix that shows enrichment over beads and background |
| The protein carries out a reaction | A name associated with that reaction in a database | An activity assay in your material |
| The protein causes the phenotype | Nothing causal | A perturbation and a phenotype measured with controls |
| The band and the list agree | A candidate mass | A gel read as a gel, which still is not function |
Reading a protein gel belongs in the last row. A band at the bait's molecular weight supports the idea that the bait was present and reasonably intact. It does not prove the prey list, and it does not prove function. Use it as a check that the pull-down happened, then return to the controls.
Failure modes of interpretation
The common failure is a slide that promotes every accession to a verb. "We found that X regulates Y" when the data say "peptides from X were enriched over the bead control". Rewrite the verb to match the evidence. The experiment did not become weaker. The sentence became checkable.
A second failure is ignoring protein groups until a reviewer notices that the functional isoform was never uniquely seen. Do that check before the cartoon. If the unique peptide is absent, the honest claim is the group.
A third failure is treating keratin, serum albumin from medium, and the antibody used for capture as biological hits because the search named them confidently. Confidence is about the sequence match. Relevance is about the control and the sample-preparation blanks. Keep a contaminant column in the table so the audience can see you separated the two.
Research limits
Nothing in a discovery list is a diagnosis or a licence to change a treatment. Functional claims that touch health need evidence and a regulatory context this page does not provide. Biosafety of the lysate, the expression system and the affinity resin is an institutional decision. HUPO language about protein evidence can help you describe existence. It does not supply the missing activity assay.
Shared reports, and how strong a sentence may be
When several groups share one facility report, the dangerous habit is a results table that travels by email and grows verbs on the way. Specify, in the analysis request, which sentence the table is allowed to support: presence, enrichment over a named control, or a quantitative change with replicates. Ask for protein groups to stay visible rather than collapsed to a single favourite gene. A specification of that kind is ordinary project hygiene, and it matters more when the person who will present the cartoon did not see the controls. Put the allowed sentence in the same document as the sample list.
What to put in an enquiry about a hit list
State the experiment class, such as a pull-down or a gel band, the control you will run beside the bait, the species database, and whether isoforms must be distinguished. Say if the decision you need is identification only, or identification plus a quantitative comparison. Ask for a report that keeps protein groups, contaminant entries and the false discovery rule visible.
The protein identification by LC-MS/MS reference, the shotgun discovery proteomics reference and the differential abundance reference are method pages around that report. A method can be discussed from the quote request. Bring the claim you hope to make, and expect the useful reply to say what extra evidence that claim still needs.
Questions from the bench
If a protein is identified with many peptides, has its function been shown?
Many peptides make the sequence identification stronger. They say the protein, or a close sequence, was in the sample under your false discovery rule. They do not say the protein was active, that it caused a phenotype, or that it sits in the pathway drawn on the last slide. Function is an extra experiment: an assay, a localisation, or a perturbation with controls.
What is a protein group, in practical terms?
It is the set of sequences that share the peptides you observed, collapsed because the data cannot decide among them. A gene name picked from the group is a convenience. Isoforms often remain unresolved because the peptides that would distinguish them were never seen. Report the group, and say when you are guessing a single gene.
A pull-down found a famous protein. Is that a mechanism?
It is a candidate enriched under the conditions you ran, if the controls say so. Famous proteins also stick to beads. Compare the bait sample with a bead-only or unrelated-bait control, and separate the bait itself from prey and from background. A pathway cartoon built from the unfiltered list is a hypothesis you have not tested yet.
Does a band on a gel prove the function the mass list suggests?
A gel can corroborate enrichment or a molecular weight. [Reading a protein gel](/blog/reading-a-protein-gel) explains what a band can honestly show. A band at the expected size is still not an activity assay. Use it as an orthogonal check of abundance or purity, and keep the functional claim for an experiment that measures function.
References
Manufacturer names identify published method classes. Trademarks remain with their owners. Catalogue records on this site are independent references for enquiry. They are not a statement of inventory, distribution rights or a supply commitment. This page is educational. It is not medical advice, a diagnostic protocol or a biosafety approval.
Catalogue
Related products and categories
These links follow the subject of the article into published manufacturer references. A listing is a reference for an enquiry, not a statement of stock or distribution rights.
Continue in this cluster
Related reading
Bottom-up proteomics in plain languageBottom-up proteomics in plain language: proteins are digested to peptides, a mass spectrometer fragments them, and a database search names candidates.
Reading a protein gelHow to read a protein gel: what SDS-PAGE bands, smears, ladders and loading differences can support, and what they cannot identify.
False discovery rates and decoy searchesA decoy search estimates the false discovery rate of a whole accepted list. It is not the probability that one protein is true, and it does not repair a wrong