troubleshooting
Sample randomisation and batch design
How sample randomisation and blocking stop a processing day from masquerading as biology, and why a confounded proteomics batch often cannot be rescued later.
- Author
- EVRINTH Editorial Team
- Published
- 8 October 2026
- Updated
- 8 October 2026
- Reading time
- 7 min

If every treated lysate is digested on Monday and every control on Tuesday, the contrast you will test is the day. Randomisation and blocking exist so that a processing signature does not wear the name of the biology. This is a troubleshooting page for designs you can still fix, and for datasets you should stop trying to rescue. How peptide intensities become protein claims is bottom-up proteomics in plain language. How long a tube can wait between steps is part of storing biological samples from fridge to freezer.
What a batch writes into the data
A batch is any set of samples that share a technical history the others do not. The same digest day, the same trypsin vial, the same column after a regeneration, the same operator, the same labelling reagent mix. The mass spectrometer records that history as a shift in intensity, retention and missingness. The shift is often small on any one peptide and coherent across hundreds of peptides. A principal-component picture coloured by batch then separates more cleanly than the same picture coloured by treatment. That picture is a check. It is not a mandatory statistical method, and it is not a correction.
Labelled experiments hide the problem only when the labels mix conditions inside one multiplex. If each multiplex contains only treated samples or only controls, the multiplex is the batch and the treatment at the same time. Label-free queues show the problem as a run-order trend. Both designs can be honest. Both can be confounded.
Randomise, or block on purpose
Randomise when samples can be handled as one set. Assign digestion order and injection order by a random draw that is written down before anyone looks at a result. Keep the key. A randomisation that you "adjust" because the treated samples were more convenient in the morning is the Monday design again.
Block when a real constraint exists. A rotor holds sixteen tubes, or a labelling kit has a fixed number of channels, or the tissue collection spans two weeks. Put every condition inside every block as evenly as the counts allow. The block is then a variable you can see, and the treatment still occurs on both days. Record the block the way you record the treatment. Balancing six and six across two digest days means three of each condition each day, not six and then six.
Do the same for injection order. A balanced digest that is then injected as all of condition A followed by all of condition B hands the contrast to the column's ageing. Interleave. Place a quality-control injection, usually a small amount of a pooled digest, at planned intervals. The pool is not a biological sample. Its drift measures the queue. NIST describes reference-material practice for measurement science. A homemade pool is not a certified reference material. It is still the right local check if you say what it is.
Preparation has to allow the randomisation. If controls are still growing while treatments are harvested, you do not yet have a set to randomise. Freeze both as they are collected, under the storage rule you trust, and digest them together later. Harvest day can still be a batch. Nest it or balance it. Do not pretend it was simultaneous.
A plot coloured by day, and a pool that drifts
After the run, before the pathway meeting, colour the samples by digest day, by column, and by treatment. If day separates and treatment does not, believe the day. If both separate and they were balanced, you can discuss a batch term with someone who knows the design. If they separate and they were the same variable, stop.
Watch the interleaved pool. A pool that jumps after a column change marks a batch boundary even if the calendar date did not change. A pool that drifts smoothly warns you that run order matters. Neither plot replaces a designed contrast. They tell you whether the contrast is still visible.
False-discovery thresholds do not solve this. They decide which peptide-spectrum matches you keep. A confounded batch changes the intensities of peptides that are correctly identified. Filtering harder removes names. It does not unmix Monday and the treatment.
| Design | Confound | What you can still conclude |
|---|---|---|
| Conditions mixed within each digest day and interleaved on the column | Residual day effects can be seen and discussed | A treatment contrast that is not the same as the day |
| All treated samples digested and injected first | Day, order and treatment are one thing | Almost nothing about the treatment. The day is a possible full explanation |
| Balanced digest, then injection blocked by condition | Column ageing matches the biology | Digestion is fair. Acquisition order is not. Re-inject in a mixed order if vials remain |
| Each multiplex label set contains only one condition | The label batch is the treatment | You cannot separate kit variation from biology |
| Two balanced blocks, one lost to a failed column | One block remains | The remaining block, if it still contains both conditions, can support a narrower claim |
| Pool never run, order not recorded | You cannot see the batch | Do not invent a correction. Describe the study as confounded if the sheet shows it |
A confound you should not try to algebra away
People reach for a batch-correction tool because the figure is already in a draft. If treatment and batch are identical, the tool has no untreated sample on Monday and no treated sample on Tuesday. It cannot see the treatment except through the day. Removing the day removes the contrast. Leaving the day in and calling it biology publishes the processing. Say that in the notebook. The troubleshooting action is to repeat the digest in a mixed order if sample remains, or to narrow the claim to observations that do not depend on the contrast. A protein identified in both groups can still be an identification. It cannot be a fold change.
Partial confounding is more tempting and still sharp. Five of six treated samples on Monday and one on Tuesday is not balance. A single swapped sample does not create a design. Either reprocess or report the limitation in plain language.
Deposited studies in PRIDE sometimes document run order well enough to see this, and sometimes do not. The Proteomics Standards Initiative expectation is that the design metadata exists. Your local sheet is the version you can still repair before you deposit anything.
Research limits
A balanced design does not make the result clinical, and it does not prove a mechanism. It only stops one class of false treatment effect. Biosafety of the material, and the ethics of human samples, remain institutional. Randomising tubes does not de-identify donors. Keep identifiers off the vial labels you were not supposed to ship.
A power cut that splits one queue into two days
A cut at injection nine creates a batch even when the randomisation was excellent. Record the last good file, the clock time, the column wash or re-equilibration, and the first file after power returned. The two halves can still contain both conditions if you interleaved. You may describe each half, or you may include the break as a visible factor if both halves are large enough to show it. Do not splice the chromatograms. Do not restart the sequence and overwrite the fact of the stop. In a building where cuts are routine, plan the queue in blocks that already match the length of a stable power window, with both conditions inside each block. The cut then lands on a boundary you already named.
What to declare when you commission the run
Declare the condition counts, the blocking constraint, and whether a pool exists. Ask for injection order to be recorded in the report. Ask what happens to the queue record if a run stops overnight. Those facts decide whether a later abundance table can be read as biology.
The differential abundance reference is the comparison this design serves. The shotgun discovery proteomics and protein identification by LC-MS/MS references cover the acquisition underneath it. Send the grouping sheet, not only the sample count, with the quote request. The pages are how you name the method in that request. They do not mean the samples have been scheduled.
Questions from the bench
Can a statistical formula remove a batch that matches the treatment?
Usually it cannot. If every treated sample was processed on one day and every control on another, day and treatment are the same variable. Adjusting for day removes the contrast you wanted, or it pretends to separate things that were never observed separately.
Is a quality-control pool required in every queue?
It is a strong practical check, not a law of statistics. Interleaved injections of one digested pool show whether the column drifted. If you did not make a pool, you can still plot by batch, and you have a weaker check.
Do technical replicates replace randomisation?
No. Repeating one digest tells you about injection noise. It does not put treated and control samples into the same processing day. Randomise the biological samples, or block them on purpose and say so.
Should identifications be filtered before we look at batch?
Look at the pattern in the quantitative table and in a simple score plot before you interpret biology. A false-discovery filter chooses which names exist. It does not remove a technical signature shared by a whole day.
References
Manufacturer names identify published method classes. Trademarks remain with their owners. Catalogue records on this site are independent references for enquiry. They are not a statement of inventory, distribution rights or a supply commitment. This page is educational. It is not medical advice, a diagnostic protocol or a biosafety approval.
Catalogue
Related products and categories
These links follow the subject of the article into published manufacturer references. A listing is a reference for an enquiry, not a statement of stock or distribution rights.
Continue in this cluster
Related reading
Bottom-up proteomics in plain languageBottom-up proteomics in plain language: proteins are digested to peptides, a mass spectrometer fragments them, and a database search names candidates.
Storing biological samples from fridge to freezerFridge, minus-20, minus-80 and nitrogen storage slow different kinds of damage. A warm box is a quarantine decision, not a hope.
A glossary of proteomics termsA working glossary of proteomics terms, from precursor and PSM to razor peptide, FDR and batch effect, and how each term changes a pull-down or abundance claim.