selection guide
Database choice changes the identifications
How taxon, reviewed versus unreviewed sequences, contaminants and a custom fusion entry change proteomics identifications, and why the decoy must match.
- Author
- EVRINTH Editorial Team
- Published
- 8 October 2026
- Updated
- 8 October 2026
- Reading time
- 8 min

A search engine returns sequences you placed in front of it. A peptide from the wrong species, or from a fusion you forgot to add, does not appear as a near miss with a helpful flag. It is unidentified, or it is forced onto a similar sequence that happens to be in the file. Choosing the database is therefore part of selecting the study, not a software default you leave untouched. The way matches become protein names is bottom-up proteomics in plain language. The digest that produces the masses is preparing peptides for mass spectrometry.
Reviewed, unreviewed, and a sequence you built
UniProt distinguishes reviewed records, the Swiss-Prot section, from unreviewed records, the TrEMBL section. Reviewed entries are fewer, better annotated, and less redundant. They are a strong default when the organism is well studied and you want protein groups that a biologist can read. They are incomplete for a rare isoform, a newly predicted gene, or an organism whose reviewed set is thin. Unreviewed entries add breadth and add noise: fragments, duplicates and hypothetical sequences that share peptides with the proteins you care about. The false-discovery procedure then has a larger space in which to be wrong, and protein grouping becomes harder.
The NCBI Protein database is another legitimate source, with its own accession style and its own redundancy. Mixing NCBI and UniProt accessions in one search without a rule produces a report nobody can join to a later paper. Pick one source, or map them with a stated table. Do not paste two proteomes together and call the result comprehensive.
A custom database is required when the biology is not in the public proteome. A heterologous enzyme, a variant with a single amino-acid change, a tag, a fusion, or a codon-optimised sequence that does not match the natural entry all need the exact chain you expressed. Include the linker. The junction peptide is often the evidence that the construct, rather than a host lookalike, was present. Add the host proteome as well if host cells were the background, or you will assign host peptides to nothing useful. Keep the custom file small and documented. A huge private file of every sequence the group has ever cloned, appended "just in case", recreates the redundancy problem.
Contaminants, isoforms, and decoys cut from the same file
Contaminants are sequences you expect from handling: human keratin, the protease you added, serum albumin from medium, the affinity protein on a bead. Search them inside the same database so their peptide-spectrum matches compete fairly and can be flagged. A contaminant collection maintained for this purpose is the usual tool. If it is missing, keratin is reported as a surprising human protein in a yeast study, or it steals shared peptides from a real hit. If you delete contaminant rows after the search and then renormalise intensities, say so. Silent deletion changes the quantitative table.
Isoform handling is the other lever. Searching only canonical sequences makes unique peptides look common and hides real splice forms. Searching every isoform makes almost every peptide shared and leaves the protein group as the honest unit. Choose from the claim you intend to write. A specification that says "we need isoform resolution" is also a specification that says "we will only claim an isoform when a unique peptide survives this larger database".
Decoys estimate the false-discovery rate. They have to be generated from the exact target file you searched, including contaminants and custom entries, with the same enzyme rules. Reversed or shuffled sequences are the usual constructions. If you add a contaminant FASTA after the decoys were built, those new sequences have no matching decoy space. If you swap in next month's UniProt release and reuse last month's decoys, the estimate is no longer paired. Build targets and decoys as one step, and store both filenames.
Redundancy has a quantitative face. Two identical sequences split or merge depending on the software's grouping. Intensities and spectral counts then move. A "clean" non-redundant file and a "complete" file of the same taxon are different experiments. Record which one you chose.
A selection you can repeat
Select in this order. Taxon first, including every organism that contributed protein, such as a human cell and a viral infection you intend to see. Reviewed or unreviewed next, with a reason. Isoforms explicit. Contaminants appended. Custom sequences added as their own entries with the amino-acid string in the report or in a supplementary file. Decoys rebuilt from that combined target. Version, download date and filename written into the report.
A too-small database is a selection error in the other direction. A single recombinant sequence with no host proteome assigns nothing to the host and can look wonderfully clean. A mammal file used for a bacterial sample does the reverse. Match the file to the tube.
When you must agree with an older study, use that study's database version if you can still obtain it, and say you did. A new search of old raw files against a new release is a new analysis. Deposits linked from the ProteomeXchange consortium sometimes include the FASTA and sometimes only name it. Prefer the ones that include it when you are learning what a repeatable record looks like. The Human Proteome Organization is a route into debates about protein evidence once the accession is chosen. It does not pick your FASTA.
| Database mistake | False story it tells |
|---|---|
| Wrong taxon | Homologs of another species look like a tidy proteome |
| Reviewed only, isoform claimed | The splice form was not in the file, so a canonical name stands in |
| Unreviewed dump, no grouping rule | Fragments and duplicates look like many proteins |
| No contaminant sequences | Keratin or the bead protein is reported as a biological hit |
| Contaminants added after decoys | The error rate no longer matches the search space |
| Fusion or variant omitted | The junction is missing, and a partial host match is over-read |
| Two sources mixed | Accessions cannot be joined, and shared peptides double-count |
| Next month's release, last month's report | The list moved and the paper can no longer be repeated |
A tidy list from the wrong taxon
The failure that fools a busy reader is a clean table. Mouse tissue searched as human returns human accessions with convincing scores wherever the sequences match. The false-discovery rate can look well behaved because the decoys were built from that same wrong file. The rate does not know you chose the wrong organism. Taxonomy is a sample-sheet fact, checked before the search, not a parameter to tune until the list looks familiar.
A second failure is the redundant file that turns one protein into a crowd of groups. The meeting then counts groups as if they were independent discoveries. Ask for the grouping rule and for a non-redundant comparison before you accept the count.
A third failure is the variant searched as wild type. The changed peptide is unidentified. The unchanged peptides support the wild-type entry. The report says the protein was present, which is true and incomplete. The custom sequence is how you see the variant peptide.
A fourth failure is quieter. The correct taxon is searched twice, once with a contaminant collection and once without, and the two protein-group counts are compared as if they were biological replicates. They are not. The second search had a different space of sequences and a different decoy set. Compare counts only across searches that used the same combined file.
Research limits
Database choice improves a research identification. It does not create a diagnostic claim, and it does not prove that a matched protein caused a phenotype. Biosafety of the organism you grew sits with your institution. Adding a viral sequence to the FASTA so you can see viral proteins is an analytical act. Clearance to work with that virus is a separate act.
Write the filename into the specification
Public releases move. A study that says "human Swiss-Prot" without a date cannot be repeated after the next update, including by you. Put the filename, the download date and a checksum or an equivalent note into the specification and into the report. That habit matters when the analysis is repeated months later by someone who was not in the room. It also matters when a power cut or a staff change interrupts the search and the job is rerun on another machine. The second machine should be given the stored FASTA, not whatever a script downloads that morning.
Species, fusions and isoforms in the enquiry
State the organism, every other organism whose proteins you must see, whether you need reviewed sequences or a broader unreviewed set, whether isoforms will be claimed, and the exact custom sequences. Ask for contaminants and for decoys built from the final file. Ask for the filename in the deliverable.
The protein identification by LC-MS/MS reference is the naming task this choice dominates. The shotgun discovery proteomics reference is the wider screen that still depends on the same file. The differential abundance reference inherits every identification decision when it ranks groups. Send the taxon and the engineered sequences with the quote request. The pages are references for that discussion. They are not a statement that a search has already been run.
Questions from the bench
Why can a smaller database produce more unique peptides?
Uniqueness is counted inside the file you searched. Drop the isoforms and a shared peptide suddenly looks specific. That is a database artefact. State the file before you call a peptide unique.
Should contaminants be searched in the same file as the sample?
Yes. Append a contaminant collection and generate decoys from the combined file. Searching contaminants afterwards, or deleting them with no record, changes both the names and the false-discovery accounting.
What if the protein is a fusion we designed?
Add the exact amino-acid sequence, including the tag and the linker, as a custom entry. A host-only database can only assign the pieces that happen to match a host protein, and it will miss the junction peptide.
Does the database choice change a clinical interpretation?
This page is about research identification. A database does not make a discovery list into a diagnosis. Clinical tests use a validated method and the reporting rules of the place you work.
References
Manufacturer names identify published method classes. Trademarks remain with their owners. Catalogue records on this site are independent references for enquiry. They are not a statement of inventory, distribution rights or a supply commitment. This page is educational. It is not medical advice, a diagnostic protocol or a biosafety approval.
Catalogue
Related products and categories
These links follow the subject of the article into published manufacturer references. A listing is a reference for an enquiry, not a statement of stock or distribution rights.
Continue in this cluster
Related reading
Bottom-up proteomics in plain languageBottom-up proteomics in plain language: proteins are digested to peptides, a mass spectrometer fragments them, and a database search names candidates.
Preparing peptides for mass spectrometryHow laboratories prepare peptides for mass spectrometry: reduction, alkylation, digestion, cleanup, and the clues that a digest failed.
A glossary of proteomics termsA working glossary of proteomics terms, from precursor and PSM to razor peptide, FDR and batch effect, and how each term changes a pull-down or abundance claim.