Skip to content
EVRINTH

Pillar guide

Bottom-up proteomics in plain language

Bottom-up proteomics in plain language: proteins are digested to peptides, a mass spectrometer fragments them, and a database search names candidates.

Author
EVRINTH Editorial Team
Published
8 October 2026
Updated
8 October 2026
Reading time
7 min
Mass spectrometer coupled to a liquid chromatography system with sample vials in the foreground
Mass spectrometer coupled to a liquid chromatography system with sample vials in the foreground

Bottom-up proteomics in plain language starts with a blunt fact. A mass spectrometer is very good at weighing and breaking peptides, and a protein mixture is easier to identify after a protease has cut it into those peptides. The laboratory then asks a database which proteins could have produced the pieces. That is the whole idea. The care lives in the cutting, the searching, and the restraint about what a match means. How the peptides are made is the companion guide preparing peptides for mass spectrometry. This page is a research explainer, not a clinical assay.

What question the method can answer

A discovery experiment asks which proteins are present above the method's ability to see them, in this sample, given this database. A quantitative experiment asks whether those proteins differ between conditions, and only if the design included replicates and a fair way to compare signals. Identifying a gel band is a narrower question: what is enriched in this slice? Each question needs a different amount of caution in the report.

Bottom-up work is the peptide route. Top-down work measures intact proteins and answers a different set of questions about proteoforms. If someone promises "the proteome" from a single injection of a complex lysate, they are describing an aspiration. Dynamic range is wide. Abundant proteins occupy the instrument. Low-abundance proteins can be real and still absent from the table.

The community that argues about these standards includes the Human Proteome Organization. Their public pages are a place to see how the field talks about identification and about human protein evidence. They are not a vendor protocol.

From a protein to a scored peptide

A typical enzyme is trypsin, which cuts on the carboxyl side of lysine and arginine when the next residue does not block it. Each protein becomes a set of peptides with predictable masses if you know the sequence and the modifications you allow. Liquid chromatography separates those peptides over time so the instrument does not see the entire mixture in one instant. The mass spectrometer measures peptide masses and then fragments selected peptides. The fragment pattern is the evidence.

A search engine compares that pattern with theoretical patterns from a sequence database. UniProt is a common source of protein sequences. Genome browsers such as Ensembl matter when you need the gene model behind a protein accession. The search score is a ranking. A decoy database of reversed or shuffled sequences estimates how often chance alone would have produced a score that high. The false discovery rate is that estimate applied to the list you keep. Report the threshold. A protein that appears only because the threshold was relaxed is not a quiet extra. It is a different claim.

Protein inference is the step that groups peptides back into proteins. A peptide that sits in one sequence is strong evidence for that sequence. A peptide shared by several isoforms supports the family, not a favourite isoform, unless a unique peptide is also present. Contaminant databases belong in the search so that keratin and common laboratory proteins are named as themselves.

Classes of instrument and experiment

The hardware class most people mean is liquid chromatography coupled to tandem mass spectrometry. Quadrupole time-of-flight and ion-trap Orbitrap-style instruments are families, not synonyms. Data-dependent acquisition picks abundant peptides for fragmentation. Data-independent acquisition fragments broader windows and asks the software to disentangle them. Label-free quantification compares signals across separate runs and is sensitive to how well those runs match. Isobaric labelling compares samples mixed inside one run and has its own ratio distortions. None of these names is a result. They are the method section.

Shotgun discovery is the broad survey. A targeted method watches peptides you chose in advance and is the better route when the question is already narrow. Public workflow notes on protocols.io show how laboratories write these choices down. Use them as examples of structure, not as a method you can skip recording.

Bottom-up proteomics from protein to list Protein Peptides LC MS/MS Database match with a stated FDR A match is a scored hypothesis. Abundance is a further claim.
Bottom-up proteomics cuts proteins into peptides, separates and fragments them, then accepts matches under a false-discovery threshold.

How to read the table

Start with the method block, not the first protein name. Which enzyme, which database, which modifications, which false-discovery threshold, which quantification if any? Then look at peptide counts and whether peptides are unique. A protein with one shared peptide is a hint. A protein with several unique peptides is a much stronger identification. Modified peptides support a modification only when that modification was searched and the spectrum really requires it. A variable modification that is allowed everywhere will be assigned sometimes by chance.

Comparisons between samples need the same search settings and a design that can see batch. A pathway picture drawn from a list of names is a hypothesis generator. It inherits every weak identification you left in the list.

Report itemWhat it can supportWhat it leaves open
Peptide-spectrum match at a stated FDRThe spectrum is consistent with that peptide sequence at the list-level error rateThat every residue was observed, or that the protein is abundant
Unique peptidesThe sequence, not only a shared family, is presentIsoforms that differ elsewhere
Intensity or reporter ratioA relative comparison if the design and the normalisation allow itAn absolute concentration in the cell
Missing valueThe peptide was not confidently seen in that runThat the protein is biologically absent
Contaminant hitThe search recognised a common laboratory proteinThat the rest of the list is contaminated in the same way

Failure modes

The wrong database is a silent failure. A human search of a mouse sample will still return matches, because many peptides are identical, and the protein names will be the human ones. Digestion problems, discussed in the peptide guide, show up as extreme missed cleavage or as almost no peptides at all. Detergent and polymer contamination suppress ionisation, so the chromatogram looks empty or full of repeating peaks. Over-loading the column broadens peaks and makes the search look noisy.

A relaxed false discovery rate to "keep interesting proteins" trains the list to confirm a hope. If a protein matters, confirm it with a unique peptide you can inspect, or with an orthogonal method. Keratin and albumin are handling and culture-medium stories. Reduce them with cleaner preparation. Do not delete them from the search so the table looks pure.

Safety and operational limits

Solvents, acids, and proteases are chemical hazards. Source material from people or from pathogens is a biosafety decision for your institution. This article assigns neither a containment level nor a diagnostic meaning. Mass spectrometers are operator-trained instruments. A research list is not a medical panel.

Long chromatographic runs hate power cuts. If a run stops mid-gradient, that file is incomplete. In a humid laboratory, vials and plate seals matter because evaporation changes the concentration the autosampler thinks it injected. Samples that travelled warm should be treated as a stability question before they are injected. Record the handoff.

What to put in an enquiry

Name the organism, the material (lysate, pull-down, or gel slice), whether you need identification only or a comparison between groups, the number of biological replicates, and any detergent or denaturant still in the sample. The shotgun discovery proteomics reference, the protein identification reference, and the differential abundance reference are independent method pages. Use them to phrase the scientific question, and use the quote request to ask whether a quotation is possible. Do not read them as a statement that EVRINTH operates the mass spectrometer or holds a particular column chemistry. The useful reply is the method that would actually be run, including the database and the false-discovery rule.

Read a bottom-up proteomics result at the right strength

  1. 01Separate identification from abundanceA named protein means peptides matched a sequence under a stated false-discovery threshold. Abundance is a second experiment with its own design and its own error model.
  2. 02Ask which peptides did the namingPrefer proteins supported by peptides that are unique to that sequence. A shared peptide can belong to a family, an isoform, or a contaminant with a similar stretch.
  3. 03Check the database and the modificationsThe search can only name sequences that were in the database, with the modifications the search allowed. The wrong species database produces a tidy and irrelevant list.
  4. 04Treat the preparation as part of the resultMissed cleavages, detergents and keratin change what the instrument can see. The peptide-preparation guide is part of reading the table, not a footnote.

Questions from the bench

Is bottom-up proteomics the same as sequencing the protein like a genome?

No. The instrument usually fragments peptides and matches the spectra to a sequence database. You infer the protein from pieces. You do not read every residue of every molecule in the tube.

What does a one-percent false discovery rate mean?

It is a list-level estimate of how many matches are expected to be wrong at the threshold you chose, judged with decoy sequences. It does not mean each reported protein is ninety-nine percent certain in ordinary language.

Why is albumin or keratin always on the list?

Serum albumin from culture medium and keratin from skin and dust are abundant and ionise well. Their presence is a prompt to look at sample handling. It does not, by itself, invalidate every other identification.

Can a research proteomics table be used as a diagnosis?

Not on the strength of this explainer. Clinical protein tests need a validated method and the legal framework that applies to reporting. A discovery list is a research result.

References

  1. Human Proteome Organization (HUPO)
  2. UniProt
  3. protocols.io
  4. Ensembl genome browser

Manufacturer names identify published method classes. Trademarks remain with their owners. Catalogue records on this site are independent references for enquiry. They are not a statement of inventory, distribution rights or a supply commitment. This page is educational. It is not medical advice, a diagnostic protocol or a biosafety approval.

Catalogue

Related products and categories

These links follow the subject of the article into published manufacturer references. A listing is a reference for an enquiry, not a statement of stock or distribution rights.