Skip to content
EVRINTH

explainer

Comparing DIA and DDA for a first study

Choosing DIA or DDA for a first study of twelve samples and two conditions, including missing values, a predicted library and what a power cut does to the

Author
EVRINTH Editorial Team
Published
8 October 2026
Updated
8 October 2026
Reading time
7 min
Mass spectrometer coupled to a liquid chromatography system with sample vials in the foreground
Mass spectrometer coupled to a liquid chromatography system with sample vials in the foreground

A first proteomics study often arrives as twelve samples, two conditions, no spectral library and no pilot. Six treated cultures and six controls are enough to ask what changed, and they are not enough to discover the method by accident. Data-dependent acquisition and data-independent acquisition answer that question with different holes. The identification background is bottom-up proteomics in plain language. Locking the deliverable before the queue starts is commissioning a sequencing or proteomics study.

Stochastic selection and windowed fragmentation

In data-dependent acquisition, DDA, the instrument measures precursor masses and then fragments a limited set of the most intense ones before it looks again. The set is not the same in every run. A peptide near the intensity boundary is chosen in one injection and ignored in the next. The spectra are relatively simple, and a person can open one of them and see whether the fragments match the sequence. The cost is the empty cell.

In data-independent acquisition, DIA, the instrument steps through mass windows and fragments what is inside each window, intense or not. Fewer precursors are skipped because of rank. The spectra are mixed, because several peptides fragment together. Identification depends on untangling those mixtures with a spectral library or with a search that builds the equivalent evidence. You gain a fuller grid of values. You take on a harder false-discovery problem, and you should be able to name the tool that handles it.

Neither mode invents proteins the digest did not make. Both still need the same sample preparation. A detergent that survives cleanup will look bad in both, and a wrong database will name the wrong taxon in both.

Fractions, a predicted library, and an FDR you can explain

With no pilot, choose the path you can defend in a lab meeting.

DDA with offline fractions is the understandable path when the team wants inspectable spectra and accepts more injections. Each sample becomes several runs. Twelve samples at six fractions are already a long queue. Depth improves because the mixture in each injection is simpler and the sampler sees further down the abundance scale. You still have stochastic gaps inside each fraction. You also consume more peptide. If the cultures were small, you may not have the material. Say so before you write "deep profiling" into the plan.

DIA is a reasonable first path when you want one injection per sample, a more complete quantitative grid, and a library strategy you can explain. An experimental library built from your own fractions is ideal and spends the instrument time you may not have. A predicted library, generated from the sequence database, lets you start without those fractions. It also offers the search a very large space of theoretical peptides. The false-discovery procedure has to be one that remains calibrated in that space. If nobody on the team can say how the tool estimates error for a predicted library, do not make that the first study. Stay with DDA, or budget time to learn the DIA control on a pooled sample before the twelve are committed.

Public datasets in PRIDE and ProteomeXchange show both styles of deposition. Reading one well-documented example of each is more useful than a slogan about which mode is modern. The Proteomics Standards Initiative is where the metadata for either acquisition ought to be described. The Human Proteome Organization is a route into how protein evidence is discussed once the list exists.

The first injection looks empty

Run a pool or a single representative digest before the twelve. If the total-ion trace is flat, stop. The acquisition mode is not the fault. Fix digestion, cleanup or injection, then return. If DDA returns only a few dozen protein groups from a whole-cell lysate, look at load and chromatography before you add fractions. If DIA returns a huge list and the decoy or entrapment behaviour of the tool looks wrong, stop and fix the error control. A long list is not a successful first study.

A second branch: the pool looks healthy and one biological sample looks empty. Treat that sample as a preparation failure. Do not impute it into the contrast. Replace it only if the design still has a clear rule for replacement. Quietly dropping the awkward control is a different experiment.

Question on a first passWhat DDA can answerWhat DIA can answer
What is present in this lysate above the sampling limit?Yes, with inspectable spectra, and with gapsYes, if the library strategy and its FDR are explicit
Which protein groups differ between these two conditions?Yes, where the peptide was sampled often enoughYes, on a fuller grid, if quantification is in the DIA tool you trust
Which isoform is the one that moved?Only with a unique peptideOnly with a unique peptide
Can we name a phosphosite?Not without enrichment and localisationNot without enrichment and localisation
Can we claim a pathway was activated?No. You can list candidatesNo. You can list candidates
What if a peptide is missing in half the runs?Expected. Show the patternLess common. Still check the chromatogram
DDA gaps versus DIA windows DDA, six runs DIA windows m/z m/z m/z m/z Filled squares are precursors DDA chose. Empty squares are the gaps. DIA windows fragment the range.
DDA leaves stochastic gaps where precursors were not chosen, while DIA fragments successive mass windows and fills more of the grid.

The diagram uses two pale fills so the windows are visible. The scientific point is the pattern: DDA samples a subset, DIA commits to windows.

Missing values are the result, not a nuisance

In this twelve-sample DDA design, expect peptides that appear in four of six treated samples and two of six controls. That pattern can be biology, stochastic sampling, or both. Show it. A protein group kept only when it is seen in every sample is a conservative list and a biased one, because scarce proteins fail that test first. Imputing every hole with a low number can manufacture differences. Imputing with the row mean can hide them. Neither choice is mandatory. The first-study report should say which rule was used and should show the missingness before anything is filled.

DIA reduces that particular hole and introduces others: a fragment group that fails a quality score, a window boundary that splits a precursor, a library peptide that is theoretical and poorly measured. Treat a failed quantitative feature as missing. Do not silently replace it because the grid looked untidy in the meeting.

With six and six and no pilot, you also have no estimate of how large a difference you can see. The honest product is a ranked, thresholded comparison with the acquisition and the missing-value rule written on it. It is not a complete map of the response.

A research comparison, not a diagnosis

Twelve culture extracts are a research design. The result does not diagnose a donor and does not prove that a pathway caused the phenotype. Biosafety of the cultures, including any viral vector used to create them, stays with the institution. Choose acquisition because you can explain its errors, not because one mode sounds more complete.

A power cut in the middle of a randomised queue

Instrument time is part of the design. A randomised order of the twelve is worth keeping so that condition is not the same thing as "morning versus afternoon". If the power fails at injection seven, record the clock time, the last completed raw file, and whether the column was re-equilibrated before the restart. The six files after the restart are a second batch even if the randomisation list is unchanged. Note the break in the sample sheet. Do not rename the restarted files as if the queue had been continuous. A blocked design, all treated samples first, is hurt more by the same cut, because the restart then coincides with the control group. Record the split either way. You can still describe the contrast inside each half. You should not pretend the halves are one uninterrupted measurement.

Asking which acquisition can be discussed

State the twelve-sample contrast, the protein amount you actually have, and whether anyone can explain the DIA error control you would rely on. Ask whether a single-shot method or a fractionated DDA queue fits that amount. Ask for the missing-value rule in the deliverable.

The shotgun discovery proteomics reference is the wide screen. The protein identification by LC-MS/MS reference is the naming question. The differential abundance reference is the comparison you want the twelve samples to support. Raise the acquisition choice through the quote request. Those pages frame the discussion. They are not a claim that the queue has been reserved.

Questions from the bench

Do we need a spectral library before a first DIA study?

You need a library strategy. That can be an experimental library or a predicted one. The people who will sign the result need to explain how the false-discovery tool treats that choice. A predicted library without that explanation is a weak plan.

Why do DDA tables contain so many empty cells?

Data-dependent acquisition fragments a limited number of precursors each cycle, and the choice changes between runs. A peptide can be real and still be missing from a given injection. Those gaps are part of the method.

Will fractions fix a first study that has no pilot?

Fractions buy depth and spend time and sample. They do not replace biological replicates. With twelve samples already split across two conditions, measure whether you can afford several injections of each before you promise a deep proteome.

Is one acquisition mode the clinical one?

Neither mode is a diagnostic test because this page describes research design. A first comparison of two culture conditions stays a research comparison under whichever acquisition you can explain.

References

  1. PRIDE proteomics data repository
  2. ProteomeXchange consortium
  3. Human Proteome Organization (HUPO)
  4. Proteomics Standards Initiative

Manufacturer names identify published method classes. Trademarks remain with their owners. Catalogue records on this site are independent references for enquiry. They are not a statement of inventory, distribution rights or a supply commitment. This page is educational. It is not medical advice, a diagnostic protocol or a biosafety approval.

Catalogue

Related products and categories

These links follow the subject of the article into published manufacturer references. A listing is a reference for an enquiry, not a statement of stock or distribution rights.