application
What to put in a sequencing statement of work
The lines a sequencing statement of work needs: organism, geometry, depth, reference build, files, retention and the rule for a failed sample.
- Author
- EVRINTH Editorial Team
- Published
- 8 October 2026
- Updated
- 8 October 2026
- Reading time
- 7 min

A statement of work is the list of decisions that still have to be true after the people who made them have left the call. If the reference build is missing, the variant file that comes back cannot be reopened with confidence. If the fail rule is missing, a weak library becomes a completed sample. Commissioning as a wider habit, including proteomics, is set out in commissioning a sequencing or proteomics study. The route from library to reads, which the statement should name rather than assume, is next-generation sequencing from library to reads.
Organism, application, and the genome you will speak about
Start with the organism and the sample type: cultured cells, tissue, a microbial isolate, a plasmid, an amplicon from a named primer pair. Then name the application in words a second site can execute. Whole-genome sequencing, an amplicon panel, and Sanger sequencing of a single template are different applications. They do not share a default read length, a default file, or a default fail rule. Writing "sequencing" is how those defaults creep in.
If reads will be compared with a reference, name the build and the source. A human analysis might specify a particular GRCh release. Another organism might specify an Ensembl release or a GenBank accession for the contig you care about. "The latest genome" is a moving target. Variants called on one build do not automatically sit on another.
Also write the sample number you will actually send, including controls, and the identifier convention. A statement that says "about a plate" is not a statement. Plate identity itself is easy to get wrong; the checks are in sample identity mix-ups and plate maps.
Read geometry, depth, and what "done" means
For a short-read application, state read length and whether you want paired ends or a single end. Those choices follow the biological question, which is argued in choosing read length and paired ends. For Sanger, state the primer and the region that must be readable, not a cluster density.
State a target depth in the units that match the application. Genome projects often speak of mean coverage across the reference. Amplicon projects often speak of reads per sample or per target. A single number without the denominator invites someone to count reads that were duplicates, or reads that missed the amplicon. Say whether duplicates count, and say the minimum you will accept before a sample is failed.
The fail rule is the most useful sentence in the document. Examples of the form, without pretending they are universal: a sample below the agreed depth is repeated once or it is reported as failed; a library whose trace is mostly adapter-dimer is not loaded; a Sanger trace that is mixed when you asked for a clone is returned as mixed rather than edited into a single base. Pick rules you can observe. A rule you cannot measure will be waived on the day.
Deliverables, the analyst, and the archive
Name the files. FASTQ is the usual raw read deliverable for short-read work. BAM or CRAM is an alignment, and it is only interpretable with the reference you named. VCF is a variant call, which already contains someone else's filters. A PDF of gene names is not a deliverable that can be reanalysed. Say which of these you will receive, and say who creates them.
"Who analyses" prevents a silent handoff. If the sequencing group delivers FASTQ and your group calls variants, write that. If they deliver a filtered VCF, write the filter and write that you may still ask for the reads. Retention belongs in the same paragraph: how long the provider keeps each file class, how you will copy them, and what happens when a download window closes. The custody argument is developed in data retention and who holds the raw files.
Quality thresholds on the reads, such as a fraction of bases above a stated Phred value, are reasonable lines when you know which instrument class you are buying. Copy them from the instrument's own specification or from a threshold you have justified. Do not invent a universal pass mark in the margin.
| Line in the statement | If you leave it blank | What you tend to receive |
|---|---|---|
| Organism and reference build | An analyst picks a genome | Coordinates you cannot match to your tracks |
| Application, length, pairing | A menu default | Reads that do not serve the question |
| Target depth and fail rule | Any completed run counts | Thin samples reported as finished |
| FASTQ, BAM, VCF, and the analyst | A summary slides in | No file you can refilter |
| Sample count and retention | A short link, a short plate | An argument after the files are gone |
A workflow for writing the statement, with a branch
Draft the scientific question in one sentence. Under it, list the lines in the table. If a line has two owners, the sequencing group and your group, mark the owner. If you cannot state a fail rule, you are not ready to send samples. Stop and decide what you would do with a weak library. If the application is really one plasmid confirmation, stop and write a Sanger statement instead of stretching a genome template to fit. If the application is many indexed amplicons, the index kit and the collision rules have to be in the statement, not in a side email.
When the draft is done, a second person should be able to execute it without calling you. That is the test.
How statements fail on ordinary projects
The application line says whole genome and the depth line was copied from an amplicon quote, or the reverse. The reference is a species name with no accession. The deliverable is "results", and the results arrive as a spreadsheet of gene symbols with no filter string. The sample count excludes the control, so the control is sequenced as a gift or not at all. The retention line says the provider will "keep the data", with no file class and no duration. Each of these feels small until you try to reanalyse.
Read quality language belongs in the statement as an observable. "Good data" does not. A fraction of bases above a Phred threshold, a minimum mean depth, a maximum dimer fraction: those can be checked. Tie them to the fail rule.
Ethics and biosafety are not optional extra lines
A statement of work does not replace an ethics review, a material transfer arrangement, or a biosafety decision. If the samples are human, the statement should say that the sequencing is research and should point at the approval you already hold. It should not promise a diagnostic interpretation. If the organism is regulated, say so, so that a receiving laboratory can refuse the package. This article does not grant that approval.
Archives such as the Sequence Read Archive are relevant only if you also decide, in the statement, whether reads may be deposited and under what conditions. Silence is not consent to deposit.
Writing so another city can run the work
Indian collaborations often split the samples, the instrument and the analysis across cities. The statement is what travels when the call does not. Use identifiers that survive a spreadsheet export. State the plate orientation. State how files will move, because a link that fails halfway through a humid week is a retention problem, not a surprise. Avoid verbal amendments. If the depth changes, change the document. Specification writing is the whole task: a second laboratory should be able to order reagents from the genomics and sequencing catalogue and still know which application those reagents serve.
What to attach to a quote request
Attach the statement, not a paragraph that says "please sequence these". Include organism, application, read geometry, target depth, reference build, sample number, deliverables, analyst, retention, and fail rules. The genomics research pathway is the context page for this kind of work. Discuss a genome design through the whole-genome sequencing enquiry reference, an amplicon design through the amplicon sequencing enquiry reference, and a single-template design through the Sanger DNA sequencing enquiry reference. Send it with the quote request. The pages are enquiry references for the method. They do not set a turnaround, and this article does not either.
Questions from the bench
Should the statement of work name the reference genome build?
Yes, whenever reads will be aligned or variants will be called. A coordinate is only defined on a named build, including the release you actually mean. If you leave the line blank, someone will choose a build for you, and later files may not match the browser track your group already uses.
Who should be written down as the group that analyses the reads?
Name the group in the statement, even if that group is your own laboratory. Delivering FASTQ to the people who will analyse it is a different piece of work from delivering a filtered variant file. If both happen, say who may change the filters, and who keeps the unfiltered calls.
Does a target depth mean the biology will be visible?
No. Depth is how many reads you asked to land on the genome or the amplicon. A variant can still hide in a repeat, in a GC-poor stretch, or in a sample that failed library preparation. Pair the depth with a fail rule so a thin sample is repeated or rejected instead of being reported as finished.
References
Manufacturer names identify published method classes. Trademarks remain with their owners. Catalogue records on this site are independent references for enquiry. They are not a statement of inventory, distribution rights or a supply commitment. This page is educational. It is not medical advice, a diagnostic protocol or a biosafety approval.
Catalogue
Related products and categories
These links follow the subject of the article into published manufacturer references. A listing is a reference for an enquiry, not a statement of stock or distribution rights.
Continue in this cluster
Related reading
Next-generation sequencing from library to readsHow a sequencing library becomes reads: adapters, flow cells, quality scores and the checks that stop a bad library from wasting a run.
16S profiling and its taxonomic limitsJudge a 16S profile at the rank the marker supports, and see where copy number, primer bias and species names stop being honest.
A glossary of sequencing termsWorking definitions of read, coverage, depth, MAPQ, Phred, VCF, BAM and the related words, each tied to the mistake that word prevents.