Skip to content
EVRINTH

glossary

Index collisions in multiplexed pools

Same index twice, indexes one edit apart, and sheets written as the reverse complement. How those collisions misassign reads in a multiplexed pool.

Author
EVRINTH Editorial Team
Published
8 October 2026
Updated
8 October 2026
Reading time
8 min
Benchtop sequencing instrument with a teal status light and a flow cell cartridge in a genomics lab
Benchtop sequencing instrument with a teal status light and a flow cell cartridge in a genomics lab

Two samples that share an index do not become two FASTQ files with a warning in the header. They become one mixture, and the demultiplexer can look successful. An index collision is that failure, plus its cousins: indexes a single edit apart, indexes too short to tolerate an error, and a sample sheet written in the wrong orientation. The library that carries the index is described in next-generation sequencing from library to reads. The plate that was supposed to make the sheet true is sample identity mix-ups and plate maps. Checking the sheet is part of what to check before a sequencing run.

The index is a sequence, not a friendly name

An index, or sample barcode, is a short synthetic sequence read alongside the insert. Demultiplexing bins the insert read according to which index sequence was seen. Kit worksheets give those sequences friendly names. The names are not unique across chemistry versions. The bases are the identity. A collision in the strict sense is two wells whose bases are the same. The software then has one bin and two histories. If it refuses the duplicate and dumps both wells into an undetermined file, you lose the samples cleanly. If a handmade sheet hides the duplicate under two names that map to one sequence, you may get one FASTQ that is a blend, labelled as whoever appeared last.

Edit distance is the number of bases at which two indexes differ. Distance one means a single substitution, a common sequencing error, converts one index into the other. The demultiplexer may then misassign the read, or it may mark the index as ambiguous if it was told to require a minimum distance. Too-short indexes have fewer bases, so the same absolute number of errors consumes the distance faster. A validated kit set is a set of sequences chosen to keep distance. An improvised set, built by picking "index 1" from two different product generations, is not.

Dual indexes and the ghost of a hopped read

Many current kits read an index at each end. Unique dual indexing gives every sample a pair that appears once. Combinatorial dual indexing reuses the same i5 sequences and i7 sequences in a grid, so the pair is unique only if both ends are read correctly and no base hops from one cluster to another.

Index hopping is a real failure mode on some patterned flow-cell instruments: a fragment acquires the index of a neighbour and appears as a sample that was never mixed in a well. With combinatorial indexes, a hopped end can create a pair that matches a different real sample, and the read is confidently misassigned. With unique dual indexes, a hopped end usually creates a pair that is not in the sheet, and the read falls into undetermined. The glossary point is simple. "Dual index" is not one design. Say which design you loaded.

The short-read instrument behaviour this refers to is part of the method class introduced by Illumina. If you use another instrument class, use its barcode rules. The word index does not travel unchanged.

What demultiplex will and will not notice

Demultiplex compares observed index reads with the sheet. It can count how many reads fell into undetermined. A large undetermined fraction is a result. It may mean the sheet's sequences do not match the kit, the orientation is wrong, or the index read failed. A small undetermined fraction is not proof of correct assignment. Hopping into a legal combinatorial pair looks assigned. A swapped column of DNA, with the indexes correctly written, looks assigned. The software did its job.

Reverse complement is the classic sheet error. One end of the index is read in an orientation that is not the orientation printed on the tube label. If you type the label and the instrument reads the complement, your samples vanish into undetermined, or they collide with another row that happens to be that complement. The check is mechanical. Generate the sheet from the vendor's tool for that kit and that instrument, or compare every base with the kit's sequence table for the chemistry you opened. Do not retype from a photograph.

Check the sheet before the pool is loaded

Lay the index sequences out as bases. Sort them. Confirm that each sample's pair occurs once. Confirm the edit distance you rely on, using the kit documentation rather than a glance. Confirm orientation against the instrument, not against a neighbouring laboratory's habit. Confirm that the well in the sheet is the well that received that index, which is the plate-map problem again. If any check fails, do not load. A pool that has already been mixed with a duplicated index cannot be unmixed by software.

If you discover the collision after the run, separate what you can. Undetermined reads might be recovered if the orientation was uniformly wrong and you rebuild the sheet. A true duplicate index cannot be split back into donors. Those samples are a mixture. Treat any rare allele in them as untrustworthy. Say so in the record.

Sheet or design errorWhat demultiplex doesWhat the FASTQ then is
Same index on two wellsOne bin, or both undeterminedA mixture, or a loss
Indexes one base apartMisassignment when that base errsLow-level foreign alleles
Index written as its reverse complementUndetermined, or a hit on the wrong rowA sample that looks failed, and a neighbour that looks contaminated
Combinatorial pairs plus hoppingA legal pair that was never preparedA ghost sample with a confident name
Index edit distance and an identical collision Distance one ATGCATGA ATGCATCA one base converts them Identical in one pool GCTAGCTA GCTAGCTA one bin, two samples Check the bases before loading. Names are not sequences.
Two index sequences one base apart can be swapped by a single error; an identical pair in the same pool cannot be split after demultiplex.

Failure modes that survive a tidy undetermined percentage

Hopping into a combinatorial index was one. Another is a short index on a high-error index read, so distance collapses only for the samples with the weakest index qualities. Another is reuse of an index across two pools that are later analysed as one cohort: the sequences do not collide inside a lane, and they collide in the scientist's head, producing two donors with the same barcode in one table. Another is a unique dual index used correctly, then a bioinformatics script that looks only at i7 and ignores i5. The script reintroduces the collision the kit had removed.

Amplicon pools are where this hurts twice. A misassigned read is a wrong organism or a wrong replicate in a counting table, and the counts still sum neatly. Genome pools hide the same error inside a convincing BAM. The design of the amplicon itself can be discussed through the amplicon sequencing enquiry reference. The index check does not become optional because the insert is short.

Laboratories often record the index layout on protocols.io. A protocol that lists index names without bases is incomplete. When you deposit reads, archives such as the Sequence Read Archive ask for sample metadata. The index sequence belongs in your own records even when the archive does not display it.

Misassignment is a governance problem when the samples are people

A hopped or collided index can place one person's variants in another person's file at a low level. In research that is a contaminated genotype. It is also a reason to have taken ethics and data-handling rules seriously before the pool existed. This glossary does not set those rules. It tells you not to describe a misassigned rare allele as a somatic mutation until the index sheet and the index design have been excluded. Biosafety of the organism is likewise unchanged by a clean demultiplex plot.

Sheets that travel between campuses

A core that receives plates from several cities will be handed photographs, half-filled spreadsheets, and kit names from a generation it does not run. Require the bases, the kit identity, the instrument orientation, and a statement that each pair is unique in that pool. Humidity-smeared labels on the plate are not a backup. A second person should diff the sheet against the vendor table before loading. If the diff is inconvenient, it still prevents a mixture. Sanger confirmation of a surprising genotype, discussed through the Sanger DNA sequencing enquiry reference, can rescue a biological claim. It cannot split a FASTQ that was mixed at the index.

What to include when you ask about a multiplex

Send the index bases, the kit version, whether the design is unique dual or combinatorial, the instrument class, and the duplicate check you already ran. Say how many samples share the lane. Reagents and index kits are classes in the genomics and sequencing catalogue. The research context is the genomics research pathway. A genome pool can be discussed through the whole-genome sequencing enquiry reference. An amplicon pool can be discussed through the amplicon sequencing enquiry reference. A single template that should not have been pooled can be discussed through the Sanger DNA sequencing enquiry reference. Put the sheet rules in the quote request so the pool is specified before it is mixed.

Questions from the bench

What is an index collision in a sequencing pool?

It is any situation in which the barcode sequence no longer points at one sample. The strict case is the same index on two wells. The practical cases include indexes so similar that a single base error turns one into the other, and a sheet that records the reverse complement of the bases the instrument reads.

How different should two indexes in one pool be?

Different enough that the errors you expect in index reads do not convert one index into another. Dual indexes raise the number of bases that must fail together. A vendor's validated set is designed with that distance in mind. Mixing leftover indexes from two kits because the names look different is how the distance disappears.

Will the demultiplexer warn me that the sample sheet lists the wrong DNA?

It will warn about some sheet problems, such as a duplicated index sequence, if you look at the log. It will not warn that the DNA in the well is not the DNA the name implies. A reverse complement that still matches some other sample's index can look like a successful assignment of the wrong well.

References

  1. Illumina overview of next-generation sequencing
  2. protocols.io method repository
  3. NCBI Sequence Read Archive

Manufacturer names identify published method classes. Trademarks remain with their owners. Catalogue records on this site are independent references for enquiry. They are not a statement of inventory, distribution rights or a supply commitment. This page is educational. It is not medical advice, a diagnostic protocol or a biosafety approval.

Catalogue

Related products and categories

These links follow the subject of the article into published manufacturer references. A listing is a reference for an enquiry, not a statement of stock or distribution rights.