Skip to content
EVRINTH

comparison

Stranded libraries and transcript orientation

How stranded RNA-seq libraries keep transcript direction, and why antisense or overlapping genes are unreadable when the chemistry is unstranded.

Author
EVRINTH Editorial Team
Published
8 October 2026
Updated
8 October 2026
Reading time
9 min
Researcher viewing gene expression heatmaps and genomic tracks on two monitors at night
Researcher viewing gene expression heatmaps and genomic tracks on two monitors at night

Antisense transcription is invisible in an unstranded RNA-seq count. So is the boundary between two genes that sit on opposite strands and overlap. The decision to keep or discard strand is made when the library is built, not when a heatmap is coloured. This comparison is for anyone choosing a library class, or reading a table and needing to know whether direction was even available. The path from cells to a result is set out in from cells to a gene expression result. What follows is only the strand decision.

Direction is a decision, not a default

A transcript has a direction because RNA polymerase moved in one direction along DNA. Many loci also produce RNA from the other strand: antisense transcripts, overlapping reading frames, and genes whose untranslated regions cross. If the assay throws direction away, those molecules land in one pile. Differential expression can still be computed on that pile. The number will not mean "this transcript".

Strand is therefore part of the scientific question. A study of a single non-overlapping protein-coding gene may tolerate an unstranded library. A study of antisense regulation, of a sense-antisense pair, or of a locus you have not yet checked for overlap, will not. Look at the locus on a current gene model, for example in Ensembl, before you treat overlap as a rare special case. It is common enough that the chemistry should be chosen on purpose.

What unstranded chemistry throws away

In a typical short-read RNA-seq library, RNA is copied into complementary DNA and adapters are added so a sequencer can read the fragments. An unstranded protocol keeps both strands of that cDNA through amplification. A read can then align to either genomic strand. Counting software that is told the library is unstranded will add both orientations into the gene total.

That addition is harmless when only one strand of the locus is transcribed. It is not harmless in an overlap. A read that falls in the shared interval is compatible with both genes. The counter has no direction with which to choose. Antisense signal is not reported as its own feature, because the feature table was never given a strand to put it on. You cannot repair that by re-annotating the same reads later. The molecules that would have distinguished the two strands were discarded in the library, or were never marked.

Stranded chemistry marks one strand so that it is removed or so that it is the only strand sequenced. The remaining reads have a known relationship to the original RNA. Overlaps can be split. Antisense can be a separate row. The mark has to be made in the tube. A strand flag typed into an analysis form does not create information the library never stored.

Two stranded classes, and the dUTP method

Stranded kits are not one orientation. Two classes cover most of the short-read methods a laboratory will meet.

The dUTP method is the common second-strand class. First-strand cDNA is made from the RNA with ordinary nucleotides, so that strand is the complement of the RNA. Second-strand synthesis then uses deoxyuridine in place of thymidine, so the new strand contains uracil. A uracil-cutting enzyme destroys that second strand before amplification. What remains, and what is sequenced as read 1, is the first-strand cDNA. Read 1 is therefore antisense to the original RNA. Read 2, in a paired run, falls the other way. People call this pattern by different letters in different tools. Some aligners expect a code that means "read 1 reverse to the RNA". Some quantification tools use another short code for the same chemistry. The letters are properties of the software, not of the biology. Copying a flag from a neighbouring pipeline is a known way to invert an entire table.

A first-strand directional class, used by some ligation or older directional kits, keeps the opposite relationship. Read 1 follows the RNA rather than its complement. The wet-lab steps are different. The analysis flag is different. Naming both classes "stranded" without saying which one is how sense and antisense get swapped and then published.

Unstranded chemistry is the third class, not a failed stranded prep. Both strands survive. There is no honest flag that will recover direction from it. If a stranded prep loses its strand selection, for example because the uracil-cutting step did not work, the library can behave as if it were unstranded. That is a failed stranded library, and it should be described as such, not quietly analysed with a confident flag.

The public orientation to short-read sequencing as a technology class is the Illumina sequencing overview. It is background for the instrument, not a substitute for the kit insert that states which strand survives.

What you hand to the person counting reads

The reagent classes are ordinary: a reverse transcriptase, a second-strand polymerase system, adapters, and, in the dUTP class, a uracil-containing nucleotide mix plus an enzyme that cuts uracil-containing DNA. Exact volumes and temperatures belong to the kit you are holding. Follow that insert. Do not paste a cycle number from a different stranded family into this one.

The record that matters to the analyst is short. State stranded or unstranded. If stranded, state the kit name and the class, dUTP second-strand or first-strand directional. State whether the run is paired and which read is expected to be antisense to the RNA. Attach the annotation release the biological question used, because overlap is a property of a gene model, not of a permanent fact. A later Ensembl release can add a transcript that creates an overlap the old table ignored.

Stranded and unstranded reads on an overlapping pair Locus with two overlapping transcripts Transcript A, sense Transcript B, antisense overlap Stranded library reads assigned by direction A and B stay separate Unstranded library overlap reads have no strand A and B are mixed
In an overlap, stranded reads stay with one transcript, while unstranded reads in the shared interval cannot choose a gene.

Branches when the strand fraction looks wrong

After alignment, a strand check asks what fraction of reads agree with the annotated direction under the flag you chose. Three branches follow.

If nearly all reads agree, the flag matches the chemistry. Proceed, and keep the flag in the methods sentence.

If the fraction agrees only after you reverse the flag, the library is stranded and the note was backwards. Do not average the two settings. Pick the orientation the kit class predicts, confirm it against the fraction, and re-count. A backwards flag does not merely add noise. It can move antisense reads onto the sense gene.

If reads agree about half the time under either flag, the library is behaving as unstranded. Either it was built that way, or strand selection failed. You may still count non-overlapping genes, with the limitation written down. You should not report antisense differentials from that file. Re-making the library is the wet-lab branch when strand was the point of the study.

A second branch sits earlier. If the RNA was already degraded before library prep, strand chemistry cannot invent intact transcripts. Integrity belongs to the extraction, as in protecting RNA during extraction. A stranded protocol on poor RNA still preserves direction of the fragments that remain. It does not restore the parts that were lost.

Library classes side by side

ClassWhat read 1 representsOverlap and antisense
UnstrandedEither strand of the cDNAShared intervals are ambiguous. Antisense is not its own count
dUTP, second-strandAntisense to the original RNA, when the uracil strand was removedDirection can split an overlap, if the flag matches this class
First-strand directionalSense relative to the RNA, in the kits that work this waySame benefit, opposite flag. Do not reuse the dUTP setting

The table is a class comparison, not a command. Confirm the row against the insert for the kit in your hand. A kit from the same manufacturer can move between classes when the product name changes.

Failures that look like antisense biology

A sudden antisense signal in one batch and not another is sometimes biology and sometimes a flag that was set for only one batch. Compare the strand-agreement fraction before you compare the genes. A sample sheet that says "stranded" for every row, while two kits were used, will produce a clean statistical contrast that is really a chemistry contrast.

Overlapping genes can also create a false treatment effect inside an unstranded table. If treatment induces transcript B, unstranded counts of overlapping transcript A rise with it. The table will call A differentially expressed. A stranded recount, or a qPCR primer that distinguishes the two strands, is the check. Relative RT-qPCR does not automatically know the strand either. The primer pair defines the molecule, which is why primer placement belongs in RT-qPCR for relative expression.

Public read archives such as the European Nucleotide Archive show how often strand information is missing from a deposited experiment. If you cannot tell the class from the record, you cannot safely reanalyse the overlaps. Write the class down so the next person is not in that position.

Research use, not a clinical expression test

Library strand is a research design choice. It does not approve a diagnostic expression assay, and it does not set a biosafety level. The organism and the sample type decide containment, and that decision belongs to the institution. Sequencing reagents and cleanup solvents need the chemical assessment that applies in your laboratory. Nothing on this page authorises clinical reporting of an antisense transcript.

The chemistry note has to travel with the tube

Strand is a fact about a tube, and it is easy to separate from the tube. A box that leaves the bench for sequencing should carry the kit class in the same metadata sheet as the sample identifiers, not in a chat message that will not be archived. Heat during a delay damages RNA before any strand mark exists. If a shipment warmed, the integrity question comes before the strand question. A power cut that warms a freezer of finished libraries is a different problem, but the metadata still has to say which class each library was. A rebuilt library after a failed uracil-cutting step needs a new identifier. Reusing the old name hides the branch you took.

Specification writing is the practical habit. The sentence "stranded RNA-seq" is incomplete in a methods note or a sourcing brief. "dUTP second-strand, paired reads, read 1 expected antisense to the RNA" is a sentence an analyst can test.

What the discussion needs

State the organism, whether antisense or overlapping loci are part of the question, the kit class if you have already chosen one, and whether the analyst is expected to infer orientation or to trust your note. Library construction can be raised against the mRNA sequencing enquiry reference. How strand-aware counts would be compared can be raised against the differential expression analysis enquiry reference. Both pages are prompts for a discussion of method. They are not a statement that a particular run is already booked.

Consumables for library work sit with the genomics and sequencing catalogue. The surrounding sample-to-result path is the nucleic acid analysis pathway. Send the scientific requirement through the quote request, including the strand sentence above, and ask which method class would actually be used.

Questions from the bench

Does unstranded data always ruin a differential expression study?

No. Genes that do not overlap another transcription unit on the opposite strand can still be counted usefully from an unstranded library. The loss is specific: antisense signal and the boundary inside an overlap become unreadable. If those features are part of the question, unstranded counts cannot answer it.

What should I tell the analyst about strandedness?

Name whether the library is stranded or unstranded, the kit class, and which read follows the RNA if you know it. A file of reads without that sentence forces the analyst to guess a flag. The guess can be checked against the data later, but the chemistry note is the primary record.

If a gene has no antisense neighbour, does strand still matter?

It matters less for that gene's total, and it still matters for the experiment as a whole. One library chemistry applies to every gene. A backwards strand flag can swap sense and antisense wherever an overlap does exist, even if the gene you plotted looks peaceful.

Can software infer strandedness after the run?

Yes. Tools can compare how reads fall on annotated transcripts and estimate whether the library behaves as stranded, and in which orientation. Treat a mismatch between that estimate and the kit note as a branch: wrong flag, wrong kit recorded, or a strand-selection step that did not hold. Inference does not restore strand to a library that was built unstranded.

References

  1. Illumina overview of next-generation sequencing
  2. Ensembl genome browser
  3. European Nucleotide Archive

Manufacturer names identify published method classes. Trademarks remain with their owners. Catalogue records on this site are independent references for enquiry. They are not a statement of inventory, distribution rights or a supply commitment. This page is educational. It is not medical advice, a diagnostic protocol or a biosafety approval.

Catalogue

Related products and categories

These links follow the subject of the article into published manufacturer references. A listing is a reference for an enquiry, not a statement of stock or distribution rights.