selection guide
What a proteomics report should state
A proteomics report a second lab can audit names groups, digestion, the database version, FDR cut-offs, normalisation and the missing-value rule.
- Author
- EVRINTH Editorial Team
- Published
- 8 October 2026
- Updated
- 8 October 2026
- Reading time
- 8 min

Accepting a proteomics file is a selection decision you make before the spreadsheet arrives. A long table can still be unauditable if the grouping, the search settings and the missing-value rule are absent. The fields below are the ones a second laboratory can check without sitting beside the instrument. A decorative pathway map is optional. The identification logic those fields sit on is explained in bottom-up proteomics in plain language. If you are still locking the question, the sample plan and the deliverable, read commissioning a sequencing or proteomics study before you treat any vendor template as complete.
What you are selecting when you accept a file
The decision is which blocks you will refuse to do without. A gel-band identification needs the chemistry, the database and the false-discovery thresholds. A comparison between treatments also needs the sample groups, the quantification method, the normalisation and the rule for values the instrument never recorded. If you accept a file that has the first set and lacks the second, you have a list of names. You do not have a contrast.
Public repositories show what a reusable record looks like when it is deposited with its metadata. The Proteomics Standards Initiative describes the community expectations for that metadata. PRIDE and the ProteomeXchange consortium are where many of those records are actually filed. You are not obliged to deposit a private study. You are obliged, if you want the file to survive a staff change, to keep the same class of facts.
The record an auditor can re-read
Start with the sample list. Every injection needs an identifier, a biological group, and a flag if it is a technical reinjection, a pool, or a quality-control insert. Anonymous filenames cannot be matched back to a freezer box. The grouping is the contrast. Write it in words as well as in a column: which samples are the treatment, which are the control, and which replicates are biological rather than repeated injections of one digest.
Digestion chemistry comes next. Name the protease, the reductant class, the alkylator class, and whether the digest was in solution, on a bead or from a gel piece. The cysteine mass in the search has to be the mass that chemistry created. Iodoacetamide is commonly searched as carbamidomethyl-cysteine, a fixed modification. A different alkylator is a different mass. "Standard digest" does not say which one was used.
Cleanup is its own line. Reversed-phase desalting, precipitation and bead workflows leave different residues. A report that skips this line leaves the next person unable to explain a polymer series or a salt plug. Liquid chromatography needs a class as well: reversed-phase separation of peptides is the usual analytical step, with the gradient length and whether fractions were injected separately. Quote the method name. Do not paste a microlitre table from a different instrument.
Instrument class and acquisition mode are evidence types, not branding. Orbitrap, time-of-flight and triple-quadrupole instruments answer different questions, and data-dependent, data-independent and targeted acquisitions do not share one error model. Name the class and the mode. The exact model can be named when you know it.
The database block is where repeatability lives. Record the database name, the version or download date, the taxon, whether isoforms were included, and whether a contaminant collection was appended. Record the enzyme and the missed-cleavage allowance. Record every fixed modification and every variable modification. A later search with a quieter modification list is a different experiment.
False discovery rates belong at two levels when both were applied: peptide or peptide-spectrum match, and protein. Name the software and the version string. Then say how protein groups were built. Shared peptides can be collapsed by parsimony, assigned by a razor rule, or kept only when a unique peptide exists. Those choices change which accession is allowed to carry a name.
Quantification is the last analytical block, and only when the study claimed amounts. State whether the comparison was label-free or used a labelling chemistry, how intensities were normalised, and what was done with missing values. Left as missing, imputed, or dropped unless seen in a stated fraction of replicates are three different analyses. The pathway picture does not choose among them.
Chemistry, chromatography and the instrument method
You do not need a kit insert in the PDF. You do need classes a second person can look up. Protease, denaturant, alkylator, cleanup, column chemistry, acquisition. If a peptide standard or an intrastudy pool was spiked, it belongs in the sample list with its role, not in a footnote that only the operator remembers.
Many laboratories report trypsin, one or two missed cleavages, carbamidomethyl-cysteine as fixed when iodoacetamide was used, oxidation of methionine as variable, and list-level false-discovery thresholds near one percent at peptide and protein level. Treat that set as a familiar window. The report should cite the settings that were actually run, taken from the software session and from the protocol the enzyme supplier describes. Copying this paragraph into a report as if it were your method would make the file look complete and leave it false.
Branches when a field is blank
If the database version is missing, stop before you compare the list with a paper that used another release. Accessions and isoforms move between releases. If the false-discovery threshold is missing, treat the names as unthresholded candidates. If the grouping is missing, keep the identifications only as identifications. If normalisation is missing, do not rank fold changes. If the missing-value rule is missing, a protein seen only in the treated samples may be a protein the sampler missed in the controls.
A second branch is the file that contains a pathway drawing and no protein-group table. Ask for the table and for the list that was fed to the drawing. The drawing cannot audit itself. A third branch is a renamed raw file with no link to the row in the sample sheet. Restore the link or drop the row. An intensity without a sample identity is not a result.
| Field | Why an auditor asks | What a vague answer looks like |
|---|---|---|
| Sample grouping | The contrast is the biology you will claim | Treated and control, with no identifiers |
| Digestion chemistry | Cysteine mass and cleavage must match the search | Standard trypsin digest |
| Cleanup | Polymers and salts explain empty or noisy runs | Cleaned up |
| LC and instrument class | Acquisition limits the claim | LC-MS |
| Database version and taxon | Only supplied sequences can be returned | Human database |
| Enzyme and missed cleavages | A missed site is evidence only if it was allowed | Trypsin |
| Modifications | An unlisted mass shift is invisible | Default modifications |
| FDR at peptide and protein level | List length is a threshold decision | High confidence |
| Software and versions | Scores are not comparable across unnamed builds | The usual pipeline |
| Protein grouping rule | Isoforms and families collapse differently | Protein IDs |
| Normalisation | Injection scale can look like biology | Normalised |
| Missing values | Absence and non-detection are different events | Complete table |
A tidy table that still cannot be audited
The usual failures are clerical, and they change the science. Gene symbols and accessions mixed in one column, with no statement of which column is the key, make a later join ambiguous. Contaminants deleted by hand, with no rule written down, produce a different dataset from contaminants retained and flagged. A fold-change column with no replicate count behind it invites a ranking that the design cannot support.
When two reports of the same samples disagree, compare database version, missed-cleavage allowance and false-discovery thresholds before you compare biology. Those three fields move the list on their own. Keratin that was filtered in one report and kept in the other will also move every abundance rank that was scaled to total intensity.
Research records, identifiers and biosafety
A proteomics report is a research record. It does not diagnose a person, and it does not assign a function to a protein on the strength of a name. Human material may carry direct identifiers. The report should say whether those identifiers were removed, and your institution decides whether the work required ethics review. Infectious or otherwise hazardous samples are a biosafety decision made under institutional rules. The Human Proteome Organization is a place to see how the field talks about protein evidence. It is not a permit to report a clinical result.
Specification writing when samples travel
Write the required fields into the specification before tubes leave the building. A courier handoff in hot weather belongs in the sample record: temperature class on dispatch, whether the box used dry ice or cold packs, and the condition on receipt. If a box warmed because power failed at a receiving bench, that fact is a sample-quality field. Months later the file can still be audited only if the database filename, the software version and the grouping were written while the people who knew them were available. A verbal note that the database was "a recent human set" will not survive a staff change.
Fields to lock before a quotation
Say which of the blocks you will insist on, and which question the file must answer. A band identification and a multi-group comparison do not share a deliverable. Ask for the method class that fits the matrix, and ask for the report fields by name: grouping, digestion, cleanup, acquisition, database version, modifications, false-discovery thresholds, grouping rule, normalisation and missing values.
The shotgun discovery proteomics page, the protein identification by LC-MS/MS page and the differential abundance page are references you can point at while that scope is discussed. They are not a statement that a study is already under way. Send the sample matrix, the contrast and the field list with the quote request.
Questions from the bench
Does a pathway diagram replace the protein table?
No. A diagram is a picture of whatever list was handed to the drawing tool. An auditor still needs the protein groups, the threshold and the contrast. Ask for the table if the file contains only the picture.
Which false-discovery fields belong in the report?
State the peptide or spectrum threshold and the protein threshold separately, and name the software that computed them. A phrase such as high confidence does not tell a second laboratory where the list was cut.
Should the report name every reagent lot?
Name the enzyme class, the cysteine chemistry and the cleanup class, and record lots you actually used when your quality system asks for them. A pasted kit insert is not a substitute for those fields.
Can the report be treated as a clinical result?
Not from this guide. A research identification or abundance table is not a diagnosis. Clinical reporting needs a validated method and the legal framework that applies where you work.
References
Manufacturer names identify published method classes. Trademarks remain with their owners. Catalogue records on this site are independent references for enquiry. They are not a statement of inventory, distribution rights or a supply commitment. This page is educational. It is not medical advice, a diagnostic protocol or a biosafety approval.
Catalogue
Related products and categories
These links follow the subject of the article into published manufacturer references. A listing is a reference for an enquiry, not a statement of stock or distribution rights.
Continue in this cluster
Related reading
Bottom-up proteomics in plain languageBottom-up proteomics in plain language: proteins are digested to peptides, a mass spectrometer fragments them, and a database search names candidates.
Commissioning a sequencing or proteomics studyCommission sequencing or proteomics by locking the question, the reference database, sample QC and the deliverable before a library or a digest is made.
A glossary of proteomics termsA working glossary of proteomics terms, from precursor and PSM to razor peptide, FDR and batch effect, and how each term changes a pull-down or abundance claim.