Skip to content
EVRINTH

application

What sequence data to request after an edit

After an edit, request coordinates, the genome build, FASTQ files, an allele table, coverage and controls. A PDF of the on-target site is not the dataset.

Author
EVRINTH Editorial Team
Published
8 October 2026
Updated
8 October 2026
Reading time
8 min
Benchtop sequencing instrument with a teal status light and a flow cell cartridge in a genomics lab
Benchtop sequencing instrument with a teal status light and a flow cell cartridge in a genomics lab

Someone emails a one-page PDF. A coloured trace stops being readable three bases after the cut, and a caption says edited. That page cannot be reanalysed, cannot be compared with the parent, and cannot be deposited. What sequence data to request after an edit is the list that prevents that file from being the only record. The biology of the cut is in how CRISPR-Cas9 editing works in research. How to interpret alleles once you have them is in checking whether a genome edit worked. Library and read concepts that sit under a FASTQ file are in next-generation sequencing from library to reads.

Coordinates that point at one locus

Ask for a species, a genome assembly, a chromosome or accession, and a start and end that enclose the site. Add the strand if the laboratory uses stranded names. A gene symbol is not a coordinate. Symbols change, and paralogues share them. If the edit is in a custom plasmid or a bacterial genome that is not in the usual browser, send the accession or the map identifier and mark the window in bases from a named feature.

Ensembl is one public place to copy an assembly name from. Write the name as displayed, not "the human genome". Two builds can both be current in different laboratories. A variant that exists in the parental line has to be distinguishable from a scar. That means the parent is sequenced against the same build, or the parent's own sequence is the reference for the allele table.

Primers belong in the request as well. An amplicon that starts inside a possible deletion will hide the deleted allele. State the primer sequences and the expected unedited length so the person analysing reads can see whether both alleles could have amplified.

FASTQ, not only a picture

For a deep amplicon or any short-read count, request the FASTQ files, the sample sheet that maps read names to tubes, and a short note of the library method class. A PDF of a pile-up is a view. The FASTQ is the data. Public archives such as the Sequence Read Archive and the European Nucleotide Archive take read files for a reason: later readers need the bases and the quality scores, not a screenshot.

For Sanger sequencing of a clone, request the chromatogram file, not only a printed trace. Mixed peaks after the cut are the result in a diploid or a pool. A PDF that someone has already base-called into a single letter hides that split. If a software tool proposed two alleles from a mixed trace, ask for the proposal and the raw file. Treat the proposal as an estimate until alleles are separated or counted.

A consensus FASTA without frequencies is the wrong object for a bulk well. Consensus collapses the mixture into one string and then looks certain. If the culture was not cloned, say so, and ask for an allele table rather than a single sequence.

The allele table

An allele table lists each distinct sequence in the window, the change relative to the stated reference, the read count or the fraction, and whether that sequence was also seen in the parent or the no-template control. Indels should be written as bases, not only as "minus one". A frameshift is your interpretation of a coding indel and can sit in a separate column, labelled as an interpretation.

Include a row for the unmodified reference if it is still present. Hiding the wild-type fraction is how a partial edit becomes a knockout in a slide. If a donor was used, the table should show whether the intended substitutions are present and whether the junctions match the donor rather than a random scar.

State the filters. Reads below a quality threshold, primer-dimers, and alignments that start in the wrong place should be counted in a rejected bin or described. A fraction computed after silent filtering is not comparable to another laboratory's fraction.

Coverage and the controls around it

Coverage at an amplicon is how many times the window was read after the filters you stated. A rare indel claimed from a handful of reads sits near the error rate of the polymerase and the sequencer. Ask for the depth on the window and for the same window in an unedited sister sample prepared in parallel. Polymerase mistakes that look like indels appear in that sister sample too. A frequency that shows up only in the edited sample, well above the sister sample, is the observation you can discuss. A frequency near the floor in both is noise.

The no-template control matters if the library could have picked up a previous amplicon. If the control produces reads of the edited allele, the table is contaminated and the fractions are not biological. Say that in the report rather than subtracting a guessed background.

On-target depth does not spill over into off-target sites. Primers define the molecule. The public overview of short-read sequencing from Illumina is background for the method class. It is not a pipeline and it does not add loci you did not amplify.

RequestWhy it has to be in the packetWhat a weaker substitute hides
Assembly and coordinatesSo the window can be found againA gene name that matches several loci
FASTQ or chromatogramSo the call can be repeatedA PDF that already chose one base
Allele table with fractionsSo mixtures stay visibleA single consensus string
Depth on the windowSo rare calls can be judgedA percentage with no denominator
Unedited sister sampleSo polymerase scars are visibleEvery difference called an edit
Statement of which loci were amplifiedSo off-target silence is not impliedAn on-target success treated as genome-wide
Sequence data to request Build and coordinates FASTQ or trace Allele table Depth Parent and no-template On-target reads stay inside the primers. They do not report other loci. A PDF can illustrate the table. It does not replace the files.
A usable return packet holds coordinates, FASTQ or traces, an allele table, depth, and a parental control, not a summary slide alone.

What on-target reads leave unsaid

Write the negative sentence in the request so it comes back in the report. On-target amplicon sequencing does not measure off-target cutting. It does not prove a protein is absent. It does not show that a second, unamplified copy of the gene is intact or gone. It does not describe alleles in cells that were not in the tube. If you want a nominated list of other sites sequenced, name those coordinates in the same build and accept that the list is the whole claim.

Large deletions that remove a primer site drop out of the amplicon. The allele table then over-represents whatever still amplified. If a knockout might be a big deletion, say so, and ask which second assay would see it. A perfect small-indel table can coexist with an unseen deletion allele.

Sample identity is part of the data. A plate map, the clone name, and a note of passage or colony number should travel with the FASTQ. Runs that lose that map are described in the spirit of what to check before a sequencing run. Recovering a genotype from a file name like "sample 4" is how parental variants become edits.

When the return is the wrong shape

If you receive only a PDF, ask for the files before you build a figure. If you receive FASTQ and no reference, you cannot check the alignment. If the allele table lacks the parent, you cannot separate a SNP from a scar. If depth is a single number for the whole run rather than for your window, the number is not coverage of the edit. Send the packet back for the missing piece rather than inventing it in the legend.

A report that calls every non-reference read an edit, including reads in the unedited control, has not finished the analysis. Ask for the control comparison. A report that says "no off-targets" after a single amplicon has overclaimed. Ask for the sentence to be narrowed to the locus that was sequenced.

Shipping notes and the enquiry

DNA amplicons and extracted genomic DNA still need labels that survive a warm courier van. Write the clone name on the tube and in the file map. A cold-chain note belongs in the shipment only when the analyte needs it. Follow the receiver's written conditions rather than a habit from another sample type. Power cuts at either end are a reason to confirm that files transferred completely, checksums included if you use them, before the sequencer is booked again.

Put the assembly, coordinates, primer sequences, clone versus pool, sample count, control lanes, and the file types you want into the quote request. The CRISPR validation sequencing enquiry reference is a discussion prompt for that specification. It is an independent method reference. Ask whether a quotation is possible. Enzymes and purification classes that feed the amplicon sit in the molecular biology catalogue, and the wider experimental setting is the molecular biology pathway.

Sequencing a research edit is not a clinical report and not a biosafety approval. The approval that covered making the edit still covers the DNA when it travels, if your institution says it does. Confirm that locally before you ship.

Questions from the bench

Is a PDF chromatogram enough to archive an edited clone?

It is enough to illustrate a clean Sanger trace in a talk, and it is a poor archive. A PDF cannot be re-base-called, re-aligned, or joined to a later run. Ask for the trace file or, for a counted read, the FASTQ, plus the reference coordinates the analyst used.

What does on-target amplicon sequencing fail to say about off-targets?

It says nothing about them. Reads that start from primers at the intended locus never visit other chromosomes. A nomination list of extra amplicons covers only that list. Genome-wide methods are a separate request. Do not let an on-target allele table stand in for either.

Why does the genome build have to travel with the coordinates?

A coordinate is a position on a named assembly. The same numbers on an older build can point at a different sequence, or at a gap. Ensembl and the UCSC browser both show which assembly is on screen. Write that name in the request so the allele table can be checked against the file you think you ordered.

References

  1. Illumina overview of next-generation sequencing
  2. NCBI Sequence Read Archive
  3. European Nucleotide Archive
  4. Ensembl genome browser

Manufacturer names identify published method classes. Trademarks remain with their owners. Catalogue records on this site are independent references for enquiry. They are not a statement of inventory, distribution rights or a supply commitment. This page is educational. It is not medical advice, a diagnostic protocol or a biosafety approval.

Catalogue

Related products and categories

These links follow the subject of the article into published manufacturer references. A listing is a reference for an enquiry, not a statement of stock or distribution rights.