# Direct RNA Sequencing with Oxford Nanopore: Principles and Applications Beyond Genome Assembly


## Key Takeaways

- Direct RNA sequencing with Oxford Nanopore technology enables full-length, single-molecule analysis of native RNA, simultaneously detecting transcript isoforms, poly(A) tail lengths, and base modifications without cDNA conversion. This directly addresses the "evidence gap" in genome assembly by providing transcript-level validation for gene models, unlike DNA sequencing alone.
- The technology leverages the characteristic electrical current disruption caused by each nucleotide base as it passes through a protein pore, with modified bases altering this current signal differently than unmodified ones, allowing for epitranscriptomic profiling. Poly(A) tail length is inferred from signal patterns at the read end.
- Compared to cDNA-based RNA sequencing, direct RNA sequencing preserves base modification information lost during reverse transcription and avoids biases introduced by premature termination at modified bases or secondary structures. However, it currently has lower throughput and higher input requirements.
- Practical workflow considerations include stringent RNA integrity assessment, higher input RNA requirements (necessitating multiplexing for low-input samples), specialized library preparation for non-polyadenylated RNA, and careful selection of flowcell chemistry and basecalling models for optimal performance and analysis compatibility.
- Analysis pipelines involve splice-aware read alignment, direct isoform detection without assembly, comparative base modification calling (e.g., Nanocompore), and poly(A) tail length estimation, all of which are crucial for refining genome annotations by confirming exon-intron boundaries, identifying novel transcripts, and improving gene model accuracy.
- Key limitations include lower sequencing accuracy compared to DNA nanopore or short-read sequencing, higher input requirements, and greater cost per read, necessitating careful experimental design and validation with orthogonal methods (e.g., RT-PCR, targeted sequencing) for critical findings.

---

Direct RNA sequencing with Oxford Nanopore technology reads native RNA molecules in full length without converting them to cDNA, enabling simultaneous detection of transcript isoforms, poly(A) tail lengths, and base modifications. For researchers who have assembled genomes and now need transcript-level evidence to refine gene models, this approach provides a single-molecule view of what is actually expressed and how transcripts are processed. This article explains the technology's operating principles, its practical workflow, and how it complements genome assembly by delivering evidence that DNA sequence alone cannot provide.

## The Evidence Gap in Genome Assembly Projects

Genome assembly produces a reference sequence, but a reference sequence does not reveal which genes are active, which splice variants are used, or which RNA modifications regulate transcript fate. Researchers who complete a genome assembly often face a second challenge: annotating genes with confidence. Gene prediction algorithms can propose models, but those models require experimental evidence to confirm exon boundaries, alternative splice sites, and untranslated regions.

Direct RNA sequencing addresses this evidence gap by reading RNA molecules as they exist in the cell. The technology captures full-length transcripts, preserves chemical modifications, and measures poly(A) tail lengths in the same read. This combination of features makes it distinct from cDNA-based RNA sequencing, which loses modification information and often fragments long transcripts during library preparation.

The technology has matured considerably since its introduction. Nanopore sequencing now supports both DNA and RNA sequencing, and its openness and versatility have made it a preferred option for an increasing number of research teams working on telomere-to-telomere genome assembly, direct RNA sequencing, and metagenomics. For genome annotation projects, direct RNA sequencing provides the transcript-level evidence needed to validate and correct gene models.

## Core Principles of Nanopore Direct RNA Sequencing

### How the Nanopore Reads RNA

Oxford Nanopore sequencing passes a single RNA molecule through a protein pore embedded in a membrane. An electrical current flows across the pore, and each nucleotide base disrupts that current in a characteristic way as it passes through. The resulting current shifts are decoded by basecalling algorithms into a nucleotide sequence.

For direct RNA sequencing, the RNA molecule is sequenced in its native form. This means the molecule retains its chemical modifications, and the read represents the actual RNA transcript instead of a copy. The technology enables full-length, single-molecule sequencing of native RNA, capturing transcript isoforms and preserving epitranscriptomic modifications without cDNA conversion.

The key distinction from DNA nanopore sequencing is that RNA is less stable than DNA and requires careful handling. The RNA molecule must remain intact through library preparation and sequencing, which places demands on sample quality and storage.

### What Direct RNA Sequencing Detects

A single direct RNA sequencing read provides multiple layers of information. The primary sequence reveals exon junctions and splice variants. The electrical current signal contains information about base modifications, because modified bases alter current flow differently than unmodified bases. The time between signals and the signal patterns at the read end provide information about poly(A) tail length.

This multiplexed information is the technology's main advantage. Researchers can ask questions about isoform usage, modification status, and poly(A) tail dynamics in a single experiment. Direct RNA sequencing has redefined transcriptomic studies across diverse systems, from uncovering novel transcripts and alternative splicing events in cancer, plants, and parasites to enabling direct detection of m6A, m5C, pseudouridine, and RNA editing events.

### Comparison With cDNA-Based Approaches

Standard RNA sequencing converts RNA to cDNA before sequencing. This conversion step has two consequences. First, it removes base modifications, because the reverse transcription and amplification process does not preserve them. Second, it introduces potential biases in isoform representation, because reverse transcription can prematurely terminate at modified bases or secondary structures.

Direct RNA sequencing avoids both problems by reading the native molecule. The tradeoff is that direct RNA sequencing currently has lower throughput and higher input requirements than cDNA-based approaches. The technology also has limitations in sequencing accuracy, input requirements, and cost, although ongoing improvements in nanopore chemistry, basecalling algorithms, and machine learning integration are rapidly expanding its utility.

## At a Glance

| Feature | Direct RNA Sequencing | cDNA-Based RNA Sequencing | Genome Assembly Support |
| --- | --- | --- | --- |
| Molecule sequenced | Native RNA | cDNA copy | DNA |
| Base modifications | Detected in current signal | Lost during conversion | Not applicable |
| Read length | Full-length transcripts | Dependent on fragmentation | Long reads for contiguity |
| Isoform resolution | Single-molecule | Often requires assembly | Not applicable |
| Poly(A) tail measurement | Yes | No | Not applicable |
| Primary genome annotation use | Confirm exon models, splice variants | Expression quantification | Contig construction |

## Practical Workflow for Direct RNA Sequencing

### Sample Preparation and RNA Quality

The success of a direct RNA sequencing experiment depends on RNA integrity. Degraded RNA produces truncated reads that cannot support isoform analysis or modification detection. Researchers should assess RNA quality before library preparation using an electrophoretic or fluorometric method that reports RNA integrity.

The input requirement for direct RNA sequencing is higher than for cDNA-based methods. Laboratories working with low-input biological samples face particular challenges, and multiplexing strategies can help address cost-effectiveness and sample throughput concerns. Commercial methods for molecular barcoding of multiple direct RNA sequencing samples have been lacking, but community-driven approaches have emerged to fill this gap.

### Library Preparation Steps

Direct RNA sequencing library preparation involves ligating sequencing adapters to the RNA molecules. The adapter contains the motor protein that drives the RNA through the pore. The poly(A) tail of messenger RNA provides a natural handle for adapter ligation, which is why standard protocols work with polyadenylated RNA.

For RNA species without poly(A) tails, such as transfer RNA, specialized library preparation approaches are required. Custom libraries that do not contain standard poly(A) tails have been demonstrated with adapted protocols, including Nano-tRNAseq libraries. Researchers working with non-polyadenylated RNA should verify that their chosen protocol supports their RNA species of interest.

### Sequencing Run Setup

The sequencing run requires a flowcell with functional nanopores and a basecalling configuration appropriate for RNA. Newer RNA chemistry flowcells have improved performance, but they also require updated analysis tools. Some community-developed analysis methods are not compatible with newer RNA chemistry flowcells and the latest generation of graphics processing units, so researchers should verify tool compatibility before starting a project.

Basecalling can be performed in real time during the run or after completion. Real-time basecalling allows researchers to monitor read quality and abort poor runs early. The basecalling model should be selected based on the RNA chemistry and the downstream analysis requirements.

### Demultiplexing Multiplexed Samples

Multiplexing multiple samples in a single sequencing run reduces cost per sample. The demultiplexing step assigns each read to its sample of origin based on barcode sequences. For direct RNA sequencing, this step requires tools that can handle the specific error profiles and chemistry versions of RNA data.

SeqTagger is a rapid and robust method that can demultiplex direct RNA sequencing data sets with high precision and recall. It works with both RNA002/R9.4 and RNA004/RNA chemistries and performs well for long and short RNA libraries. Increasing multiplexing up to 96 barcodes yields highly accurate demultiplexing models. The availability of an efficient and simple multiplexing strategy improves cost-effectiveness and facilitates analysis of low-input biological samples.

## Analysis Pipeline for Direct RNA Sequencing Data

### Basecalling and Quality Control

The first analysis step after sequencing is basecalling, which converts raw electrical current signals into nucleotide sequences. Basecalling accuracy directly affects all downstream analyses, including isoform detection and modification calling. Researchers should use the basecalling model recommended for their flowcell and chemistry version.

Quality control after basecalling should assess read length distribution, read count, and per-read quality scores. Direct RNA sequencing reads are typically shorter than DNA nanopore reads because RNA molecules are shorter and more fragile. A successful run should produce a read length distribution that matches the expected transcript sizes in the sample.

### Read Alignment to the Genome

The next step is aligning direct RNA sequencing reads to the reference genome. Splice-aware aligners are required because reads span exon junctions. The alignment step produces a BAM file that serves as input for most downstream analyses.

Alignment quality should be assessed by mapping rate, coverage uniformity, and junction support. Reads that fail to map may indicate contamination, misassembly, or novel transcripts that do not match the reference. Researchers working with newly assembled genomes should expect some reads to map poorly if the assembly contains errors or missing regions.

### Isoform Detection and Quantification

Direct RNA sequencing reads that span full-length transcripts provide direct evidence for isoform structure. Each read represents a complete transcript molecule, so isoform detection does not require transcript assembly from short fragments. This simplifies the analysis and reduces ambiguity in isoform assignment.

Isoform quantification from direct RNA sequencing data requires accounting for read depth and potential biases. The technology supports isoform quantification and poly(A) tail measurement in the same experiment. Researchers should compare isoform calls against existing annotations and validate novel isoforms with orthogonal methods when possible.

### Base Modification Detection

Base modification detection from direct RNA sequencing data relies on comparing current signal patterns between modified and unmodified samples. The Nanocompore approach compares an RNA sample of interest against a non-modified control sample, not requiring a training set and allowing the use of replicates. This comparative strategy can detect different RNA modifications with position accuracy in vitro and has been applied to profile m6A in vivo in yeast and human RNAs.

The comparative approach requires careful experimental design. The control sample must be identical to the treatment sample except for the modification of interest. For m6A detection, a common control is a sample treated with an enzyme that removes the modification or a sample from a cell line lacking the modifying enzyme.

### Poly(A) Tail Length Analysis

Poly(A) tail length can be estimated from direct RNA sequencing reads because the tail produces a characteristic signal pattern at the read end. The analysis measures the number of adenine bases in the tail, providing a distribution of tail lengths across transcripts.

Poly(A) tail length correlates with mRNA abundance and stability. Global assessments have indicated a negative association between poly(A) tail length and mRNA abundance, while pathway-specific responses can be observed upon depletion of m6A-forming enzymes. These correlations provide functional context for understanding how transcript features interact.

## Using Direct RNA Sequencing to Improve Genome Annotations

### Confirming Exon-Intron Boundaries

Direct RNA sequencing reads that span exon junctions provide direct evidence for splice site usage. Each junction-spanning read confirms that two exons are joined in a mature transcript. The density of junction-spanning reads indicates the relative usage of alternative splice sites.

For genome annotation, this evidence is valuable because gene prediction algorithms often propose multiple possible splice variants. Direct RNA sequencing data can confirm which variants are actually expressed and at what relative abundance. This evidence is particularly useful for genes with complex alternative splicing patterns.

### Identifying Novel Transcripts

Direct RNA sequencing can uncover novel transcripts that are absent from existing annotations. These may include previously unannotated genes, alternative isoforms of known genes, or non-coding RNAs. The full-length nature of direct RNA sequencing reads makes novel transcript identification straightforward because each read provides a complete transcript model.

Novel transcript discovery has been demonstrated across diverse systems, including cancer, plants, and parasites. Researchers working with newly assembled genomes should expect to find transcripts that do not match predicted gene models. These novel transcripts may represent genes that were missed by prediction algorithms or isoforms that are specific to the sampled condition.

### Detecting Fusion Transcripts

Fusion transcripts result from genomic rearrangements or trans-splicing events that join sequences from different genes. Direct RNA sequencing can identify fusion transcripts because a single read spans the fusion junction. The technology supports fusion transcript identification as one of its analytical capabilities.

Fusion transcript detection from direct RNA sequencing data requires careful filtering to distinguish true fusions from alignment artifacts. Researchers should validate candidate fusions with orthogonal methods, such as PCR or DNA sequencing, before reporting them as novel findings.

### Improving Gene Model Accuracy

The combination of isoform evidence, modification data, and poly(A) tail measurements from direct RNA sequencing can improve gene model accuracy in assembled genomes. Exon boundaries confirmed by full-length reads are more reliable than those predicted by algorithms alone. Alternative isoforms identified by direct RNA sequencing expand the annotation beyond single representative transcripts.

For genome assembly projects, direct RNA sequencing data can also identify assembly errors. Reads that span regions where the assembly is incorrect will show alignment artifacts, such as soft clipping or split alignments. Investigating these patterns can reveal misassembled regions that require correction.

## Options and Tradeoffs in Direct RNA Sequencing

### Chemistry and Flowcell Selection

The choice of RNA chemistry and flowcell affects read length, throughput, and accuracy. Newer chemistries offer improved performance but may require updated analysis tools. Researchers should verify that their planned analysis pipeline supports their chosen chemistry version.

The RNA002/R9.4 chemistry has been widely used and has extensive community tool support. The RNA004/RNA chemistry is newer and offers improved performance but has less mature tool support. Researchers starting new projects should consider whether their analysis tools support the newer chemistry.

### Multiplexing Strategy

Multiplexing multiple samples in a single run reduces cost per sample but introduces complexity in library preparation and demultiplexing. The number of samples that can be multiplexed depends on the barcoding strategy and the sequencing depth required per sample.

Increasing multiplexing up to 96 barcodes yields highly accurate demultiplexing models with current tools. However, higher multiplexing reduces reads per sample, which may be limiting for low-abundance transcripts or modification detection. Researchers should balance multiplexing level against the depth required for their biological questions.

### Input RNA Amount

Direct RNA sequencing requires more input RNA than cDNA-based methods. The technology has limitations in input requirements, which constrains its use for samples with limited material. Researchers working with low-input samples should consider whether their RNA yield is sufficient for direct RNA sequencing or whether alternative approaches are more appropriate.

For low-input biological samples, multiplexing strategies can improve cost-effectiveness and facilitate analysis. However, the input requirement remains a practical constraint that should be considered during experimental design.

### Sequencing Depth Considerations

The sequencing depth required depends on the biological questions. Isoform detection requires sufficient reads per transcript to distinguish true isoforms from noise. Modification detection requires deeper coverage because modification calling relies on comparing signal patterns across many reads.

Researchers should consider their depth requirements when planning multiplexing levels and run duration. Pilot experiments can help determine the depth needed for specific applications.

## Records and Measurements for Direct RNA Sequencing Projects

### Essential Run Metrics

Every direct RNA sequencing run should be documented with standard metrics. These include total read count, read length N50, mean read quality, and mapping rate. These metrics provide a baseline for comparing runs and troubleshooting problems.

The read length N50 is particularly important for direct RNA sequencing because it reflects RNA integrity and library quality. A low N50 suggests RNA degradation or library preparation issues. The mapping rate indicates whether the sequencing produced usable data and whether the reference genome is appropriate.

### Quality Control Thresholds

Quality control thresholds should be established before starting a project. These thresholds define acceptable ranges for read count, read length, and mapping rate. Runs that fall below thresholds should be investigated and potentially repeated.

The specific thresholds depend on the biological system and the analysis goals. A project focused on isoform detection requires longer reads than a project focused on expression quantification. Researchers should document their thresholds and the rationale for choosing them.

### Sample Tracking and Metadata

Each sample should be tracked with complete metadata, including tissue source, treatment condition, RNA extraction method, RNA quality metrics, and library preparation details. This metadata is essential for interpreting results and for reproducing experiments.

The metadata should also include the barcode sequences used for multiplexing and the demultiplexing tool version. This information allows researchers to trace reads back to their sample of origin and to verify that demultiplexing was performed correctly.

### Analysis Pipeline Documentation

The analysis pipeline should be documented with software versions, parameters, and reference genome versions. This documentation enables reproducibility and allows other researchers to understand how results were generated.

Reproducible analysis workflows are available through community platforms. The [Galaxy Training Network](https://training.galaxyproject.org/) provides accessible workflow training and analysis tutorials for bioinformatics methods. The [nf-core documentation](https://nf-co.re/docs) describes community pipeline standards for reproducible workflow configuration. [Bioconductor](https://bioconductor.org/) provides official package and workflow documentation for reproducible genomic analysis.

## Common Failure Patterns and Troubleshooting

### Low Read Yield

Low read yield can result from insufficient RNA input, degraded RNA, or flowcell problems. The first troubleshooting step is to check RNA quality and quantity. If RNA quality is acceptable, the flowcell should be checked for pore count and performance.

Low read yield can also result from library preparation losses. The adapter ligation step is a common source of loss. Researchers should verify that the adapter ligation was successful by checking the library concentration before loading.

### Short Read Lengths

Short read lengths indicate RNA degradation or premature termination during sequencing. RNA degradation can occur during sample storage, library preparation, or sequencing. The RNA should be handled carefully and stored at appropriate temperatures.

Premature termination can result from RNA secondary structures or modified bases that stall the motor protein. This is more common in RNA with extensive secondary structure or high modification density. Researchers should consider whether their RNA species of interest is prone to this problem.

### Poor Mapping Rate

Poor mapping rates can result from contamination, reference genome errors, or novel transcripts. The first step is to check the taxonomic composition of the reads. Contamination from other species will produce reads that do not map to the reference.

Reference genome errors can cause reads to map poorly or to show alignment artifacts. Direct RNA sequencing reads that span misassembled regions will show soft clipping or split alignments. Investigating these patterns can reveal assembly errors that require correction.

### Demultiplexing Failures

Demultiplexing failures can result from barcode errors, adapter contamination, or tool incompatibility with the chemistry version. The demultiplexing tool should be selected based on the chemistry version and the barcoding strategy.

Some community-driven demultiplexing methods are not compatible with newer RNA chemistry flowcells and the latest generation of graphics processing units. Researchers should verify tool compatibility before starting a project and should test demultiplexing performance on a small subset of reads before processing the full data set.

## Limitations and Interpretation Boundaries

### Sequencing Accuracy

Direct RNA sequencing has lower accuracy than DNA nanopore sequencing and lower accuracy than short-read sequencing. This accuracy limitation affects base-level analyses, including variant calling and modification detection. Researchers should interpret base-level findings with caution and validate important results with orthogonal methods.

The accuracy limitation is particularly relevant for modification detection, which relies on subtle differences in current signals. Comparative approaches that compare modified and unmodified samples can improve confidence but require careful experimental design.

### Input Requirements

Direct RNA sequencing requires more input RNA than cDNA-based methods. This limitation constrains its use for samples with limited material, such as clinical biopsies or rare cell populations. Researchers working with limited material should consider whether direct RNA sequencing is feasible or whether alternative approaches are more appropriate.

The input requirement also affects the types of RNA that can be studied. Low-abundance transcripts may not be detected if the input is insufficient. Researchers should consider their target transcripts and their abundance when planning experiments.

### Cost Considerations

Direct RNA sequencing is more expensive per read than cDNA-based sequencing. The cost difference reflects the lower throughput of direct RNA sequencing and the specialized reagents required. Researchers should consider whether the additional information from direct RNA sequencing justifies the additional cost.

Multiplexing can reduce cost per sample, but the cost per read remains higher than for cDNA-based approaches. The cost-effectiveness of direct RNA sequencing depends on the biological questions and the value of the additional information it provides.

### Interpretation Limits for Modification Detection

Modification detection from direct RNA sequencing data has specific interpretation limits. The comparative approach detects differences between samples, but it does not directly identify the type of modification. A difference in current signal could result from multiple modification types.

The position accuracy of modification detection depends on the analysis method and the sequencing depth. Researchers should validate modification calls with orthogonal methods when possible and should report the confidence of their calls.

## Safety and Regulatory Context

### Data Management and Sharing

Direct RNA sequencing data should be managed according to institutional and funder requirements. Raw sequencing data and processed results should be stored securely and backed up regularly. Data sharing should follow community standards and any applicable regulations.

The [NCBI](https://www.ncbi.nlm.nih.gov/) provides official descriptions of databases, search systems, sequence resources, and analysis services for data management and sharing. Researchers should deposit their data in appropriate public repositories to enable reproducibility and reuse.

### Ethical Considerations for Human Samples

Direct RNA sequencing of human samples raises ethical considerations related to consent, privacy, and data sharing. Researchers should ensure that their studies have appropriate ethical approval and that sample collection follows applicable regulations.

The technology has been applied to human transcriptomes, including studies of mRNA modifications and regulatory features. These studies require careful attention to ethical and regulatory requirements for human subjects research.

### Quality Assurance for Clinical Applications

Direct RNA sequencing has potential applications in mRNA vaccine quality control and RNA-based data storage. These applications require quality assurance standards appropriate for their intended use.

For clinical applications, the technology's limitations in sequencing accuracy and input requirements must be considered. Researchers developing clinical applications should validate their methods thoroughly and document their quality assurance procedures.

## Professional Escalation Criteria

### When to Seek Technical Support

Researchers should seek technical support from the sequencing platform provider when they encounter persistent problems with run performance, basecalling, or library preparation. Technical support can help diagnose flowcell issues, reagent problems, or instrument malfunctions.

The platform provider's documentation and support channels should be consulted before troubleshooting independently. Many common problems have documented solutions that can be applied quickly.

### When to Consult Bioinformatics Experts

Researchers should consult bioinformatics experts when they encounter analysis problems that exceed their expertise. This includes problems with basecalling model selection, alignment parameter optimization, or modification detection interpretation.

Bioinformatics training is available through multiple channels. The [EMBL-EBI Training](https://www.ebi.ac.uk/training) program provides bioinformatics learning pathways and practical analysis education. The [Carpentries Lessons](https://carpentries.org/lessons) provide foundational computing, data, shell, Git, and programming training. These resources can help researchers build the skills needed for direct RNA sequencing analysis.

### When to Validate With Orthogonal Methods

Researchers should validate direct RNA sequencing findings with orthogonal methods when the findings are important for their conclusions. This includes novel isoform calls, modification calls, and fusion transcript identifications.

Orthogonal validation methods include PCR, short-read RNA sequencing, mass spectrometry for modifications, and targeted sequencing approaches. The choice of validation method depends on the finding being validated and the available resources.

### When to Repeat the Experiment

Researchers should repeat the experiment when quality control metrics fall below acceptable thresholds or when results are inconsistent with biological expectations. Repeating the experiment can confirm findings and rule out technical artifacts.

The decision to repeat should be documented with the rationale. Repeating experiments is particularly important for modification detection, which is sensitive to experimental conditions and analysis parameters.

## A Practical Decision Framework for Choosing Direct RNA Sequencing in Genome Annotation Projects

Researchers completing a genome assembly face a practical decision: whether direct RNA sequencing is the right investment for their annotation goals. The technology offers unique transcript-level evidence, but it also carries higher input requirements, cost per read, and analysis complexity compared with cDNA-based approaches. A structured decision framework helps researchers determine when direct RNA sequencing is justified, what depth they need, and how to allocate resources across validation methods.

### Step 1: Define the Annotation Questions That Require Transcript Evidence

The first decision point is identifying which annotation problems cannot be solved with existing data. Genome assembly projects typically have three categories of annotation needs, and each has different implications for whether direct RNA sequencing is necessary.

The first category is confirming predicted gene models. If the assembly was annotated with ab initio gene predictors, those models require experimental validation. Short-read RNA sequencing can confirm many exon junctions, but it struggles with genes that have complex alternative splicing patterns or low expression levels. Direct RNA sequencing provides full-length transcript evidence that resolves isoform structure unambiguously.

The second category is discovering novel transcripts. Genes that were missed by prediction algorithms, non-coding RNAs, and condition-specific isoforms require transcriptome sequencing to be identified. Direct RNA sequencing excels at this because each read represents a complete transcript molecule, eliminating the need for transcript assembly from short fragments.

The third category is characterizing transcript features beyond sequence. If the annotation project needs information about poly(A) tail lengths, base modifications, or the relationship between these features and isoform usage, direct RNA sequencing is the only approach that provides this information in a single experiment. cDNA-based methods cannot detect modifications, and they do not preserve poly(A) tail information.

Researchers should write down their specific annotation questions and rank them by importance. If the top questions require full-length isoform evidence or modification information, direct RNA sequencing is justified. If the questions are limited to confirming exon junctions and quantifying expression, short-read RNA sequencing may be sufficient at lower cost.

### Step 2: Assess Sample Availability and RNA Quality Constraints

The second decision point is whether the biological samples can meet the input requirements of direct RNA sequencing. The technology requires more input RNA than cDNA-based methods, and this constraint shapes experimental feasibility.

Researchers should measure their available RNA yield and assess RNA integrity before committing to direct RNA sequencing. The RNA quality assessment should use an electrophoretic or fluorometric method that reports RNA integrity. Degraded RNA produces truncated reads that cannot support isoform analysis or modification detection, so samples with poor integrity should not be used for direct RNA sequencing regardless of yield.

For samples with limited material, such as clinical biopsies or rare cell populations, researchers should consider whether multiplexing strategies can make the experiment feasible. Multiplexing multiple samples in a single sequencing run reduces cost per sample, and current tools support up to 96 barcodes with highly accurate demultiplexing models. However, higher multiplexing reduces reads per sample, which may be limiting for low-abundance transcripts or modification detection.

The RNA species of interest also matters. Standard direct RNA sequencing protocols work with polyadenylated RNA because the poly(A) tail provides a natural handle for adapter ligation. For non-polyadenylated RNA species, such as transfer RNA, specialized library preparation approaches are required. Custom libraries that do not contain standard poly(A) tails have been demonstrated with adapted protocols, but researchers should verify that their chosen protocol supports their RNA species of interest.

### Step 3: Estimate Sequencing Depth Requirements for Each Analysis Goal

The third decision point is determining how much sequencing depth is needed for the annotation questions. Depth requirements vary substantially by analysis goal, and underestimating depth is a common cause of failed experiments.

For isoform detection, the depth requirement depends on the number of transcripts that need to be resolved and their abundance distribution. Full-length reads provide direct evidence for isoform structure, but low-abundance isoforms require deeper sequencing to be captured. Researchers should estimate the dynamic range of expression in their sample and calculate the depth needed to detect isoforms at their abundance threshold.

For base modification detection, depth requirements are higher because modification calling relies on comparing signal patterns across many reads. The comparative approach, such as Nanocompore, compares an RNA sample of interest against a non-modified control sample, and this comparison requires sufficient coverage at each position to detect statistically significant signal differences. Researchers should expect to need several-fold more coverage for modification detection than for isoform detection alone.

For poly(A) tail length analysis, depth requirements depend on whether the analysis is per-transcript or global. Global assessments of poly(A) tail length distributions require less depth than per-transcript measurements. Researchers should decide which level of analysis their biological questions require and plan depth accordingly.

A practical approach is to run a pilot experiment with a small amount of sequencing, assess the read depth achieved, and then calculate whether the depth is sufficient for the planned analyses. Pilot experiments are particularly valuable for modification detection, which is sensitive to experimental conditions and analysis parameters.

### Step 4: Compare Direct RNA Sequencing With Alternative Evidence Sources

The fourth decision point is comparing direct RNA sequencing with alternative approaches for generating transcript evidence. This comparison should consider information content, cost, and time.

Short-read RNA sequencing is the most common alternative. It provides expression quantification and can confirm many exon junctions, but it loses modification information during cDNA conversion and often fragments long transcripts. For annotation projects that primarily need expression data and basic junction confirmation, short-read sequencing may be sufficient at lower cost.

Isoform sequencing, which sequences full-length cDNA molecules, provides isoform structure information but does not preserve modifications. This approach can be useful when isoform discovery is the primary goal and modification information is not needed. However, the cDNA conversion step can introduce biases in isoform representation because reverse transcription can prematurely terminate at modified bases or secondary structures.

Direct RNA sequencing provides the most complete information per read, but it has the highest cost per read and the most demanding input requirements. The decision should be based on whether the additional information justifies the additional cost. For annotation projects where modification information is not needed and isoform structure can be resolved from cDNA-based approaches, direct RNA sequencing may not be the most cost-effective choice.

Researchers should also consider whether they can combine approaches. A common strategy is to use short-read RNA sequencing for expression quantification and direct RNA sequencing for isoform resolution and modification detection on a subset of samples. This hybrid approach balances cost and information content.

### Step 5: Evaluate Analysis Capacity and Tool Compatibility

The fifth decision point is whether the research team has the analysis capacity to process direct RNA sequencing data. The analysis pipeline requires basecalling, alignment, isoform detection, and potentially modification detection, and each step has specific tool requirements.

Tool compatibility is a particular concern because some community-developed analysis methods are not compatible with newer RNA chemistry flowcells and the latest generation of graphics processing units. Researchers should verify that their planned analysis pipeline supports their chosen chemistry version before starting a project. The RNA002/R9.4 chemistry has been widely used and has extensive community tool support, while the RNA004/RNA chemistry is newer and offers improved performance but has less mature tool support.

The analysis pipeline should be documented with software versions, parameters, and reference genome versions. This documentation enables reproducibility and allows other researchers to understand how results were generated. Reproducible analysis workflows are available through community platforms, including the [Galaxy Training Network](https://training.galaxyproject.org/) for accessible workflow training and the [nf-core documentation](https://nf-co.re/docs) for community pipeline standards.

Researchers who lack bioinformatics expertise should plan for training or collaboration. The [EMBL-EBI Training](https://www.ebi.ac.uk/training) program provides bioinformatics learning pathways and practical analysis education. The [Carpentries Lessons](https://carpentries.org/lessons) provide foundational computing, data, shell, Git, and programming training. [Bioconductor](https://bioconductor.org/) provides official package and workflow documentation for reproducible genomic analysis.

### Step 6: Document the Decision and Set Success Criteria

The final decision point is documenting the rationale for choosing direct RNA sequencing and setting success criteria before starting the experiment. This documentation serves multiple purposes: it clarifies the experimental goals, provides a basis for evaluating success, and supports reproducibility.

The decision document should include the specific annotation questions being addressed, the expected depth requirements, the sample availability and quality constraints, the comparison with alternative approaches, and the analysis capacity assessment. This document should be shared with collaborators and referenced when interpreting results.

Success criteria should be defined for each analysis goal. For isoform detection, success might be defined as confirming a certain percentage of predicted gene models or identifying a minimum number of novel isoforms. For modification detection, success might be defined as detecting modifications at known sites or identifying a minimum number of candidate modification sites for validation.

The success criteria should include quality control thresholds for the sequencing run itself. These thresholds define acceptable ranges for read count, read length, and mapping rate. Runs that fall below thresholds should be investigated and potentially repeated. The specific thresholds depend on the biological system and the analysis goals, and they should be documented with the rationale for choosing them.

### Common Decision Errors and How to Avoid Them

Several recurring errors undermine direct RNA sequencing decisions in genome annotation projects. Recognizing these patterns helps researchers avoid costly mistakes.

The first error is choosing direct RNA sequencing when the annotation questions only require expression quantification. Researchers sometimes select the technology for its novelty or comprehensiveness without assessing whether the additional information will change their conclusions. The decision framework should be applied honestly, and researchers should be willing to conclude that short-read sequencing is sufficient for their questions.

The second error is underestimating depth requirements for modification detection. Researchers who plan for isoform-level depth and then attempt modification calling often find that their coverage is insufficient. The depth requirements for modification detection should be estimated separately and included in the sequencing plan.

The third error is ignoring tool compatibility issues until after sequencing is complete. Researchers who discover that their analysis tools do not support their chosen chemistry version face delays and additional costs. Tool compatibility should be verified before committing to a chemistry version.

The fourth error is failing to document the decision and success criteria. Without documentation, it is difficult to evaluate whether the experiment achieved its goals or to troubleshoot problems when they arise. The decision document and success criteria should be treated as essential components of the experimental design.

### Records and Measurements for the Decision Process

The decision process itself should generate records that support the eventual interpretation of results. These records include the annotation questions, the depth estimates, the sample quality assessments, the tool compatibility checks, and the success criteria.

Each sample should be tracked with complete metadata, including tissue source, treatment condition, RNA extraction method, RNA quality metrics, and library preparation details. This metadata is essential for interpreting results and for reproducing experiments. The metadata should also include the barcode sequences used for multiplexing and the demultiplexing tool version.

The analysis pipeline should be documented with software versions, parameters, and reference genome versions. This documentation enables reproducibility and allows other researchers to understand how results were generated. The [NCBI](https://www.ncbi.nlm.nih.gov/) provides official descriptions of databases, search systems, sequence resources, and analysis services for data management and sharing, and researchers should deposit their data in appropriate public repositories to enable reproducibility and reuse.

### Professional Escalation Criteria for Decision Uncertainty

When the decision framework does not yield a clear answer, researchers should escalate to appropriate expertise. Uncertainty about depth requirements, tool compatibility, or the value of modification information for annotation goals warrants consultation with bioinformatics experts or experienced direct RNA sequencing users.

Researchers should seek technical support from the sequencing platform provider when they encounter persistent problems with run performance, basecalling, or library preparation. Technical support can help diagnose flowcell issues, reagent problems, or instrument malfunctions. The platform provider's documentation and support channels should be consulted before troubleshooting independently.

Researchers should consult bioinformatics experts when they encounter analysis problems that exceed their expertise. This includes problems with basecalling model selection, alignment parameter optimization, or modification detection interpretation. The [EMBL-EBI Training](https://www.ebi.ac.uk/training) program and the [Carpentries Lessons](https://carpentries.org/lessons) provide training pathways that can help researchers build the skills needed for direct RNA sequencing analysis.

When the decision is uncertain because the biological questions are not well defined, researchers should revisit the annotation goals with their collaborators. The decision framework is only as useful as the clarity of the underlying questions, and refining the questions often clarifies the technology choice.

## Frequently Asked Questions

### What is the main advantage of direct RNA sequencing over cDNA-based RNA sequencing?

Direct RNA sequencing reads native RNA molecules without converting them to cDNA, which preserves base modifications and provides full-length transcript information. A single read can reveal the transcript sequence, splice variants, poly(A) tail length, and modification status. cDNA-based methods lose modification information during reverse transcription and often fragment long transcripts, making isoform analysis more complex.

### How does direct RNA sequencing help improve genome assembly annotation?

Direct RNA sequencing provides transcript-level evidence that confirms exon boundaries, identifies alternative splice variants, and detects novel transcripts. Each full-length read represents a complete transcript molecule, so isoform detection does not require assembly from short fragments. This evidence is used to validate and correct gene models predicted by algorithms, improving the accuracy of genome annotations.

### What are the input requirements for direct RNA sequencing?

Direct RNA sequencing requires more input RNA than cDNA-based methods. The exact amount depends on the library preparation protocol and the RNA species being studied. Researchers working with low-input samples should consider whether their RNA yield is sufficient and whether multiplexing strategies can improve cost-effectiveness.

### How are base modifications detected in direct RNA sequencing data?

Base modifications are detected by comparing current signal patterns between modified and unmodified samples. The comparative approach, such as Nanocompore, compares an RNA sample of interest against a non-modified control sample without requiring a training set. Modified bases alter current flow differently than unmodified bases, and these differences are identified by the analysis software.

### Can direct RNA sequencing detect poly(A) tail lengths?

Yes, direct RNA sequencing can measure poly(A) tail lengths because the tail produces a characteristic signal pattern at the read end. The analysis estimates the number of adenine bases in the tail, providing a distribution of tail lengths across transcripts. Poly(A) tail length correlates with mRNA abundance and stability.

### What are the main limitations of direct RNA sequencing?

The main limitations are sequencing accuracy, input requirements, and cost. Direct RNA sequencing has lower accuracy than DNA nanopore sequencing and short-read sequencing. It requires more input RNA than cDNA-based methods and is more expensive per read. These limitations are being addressed by ongoing improvements in nanopore chemistry, basecalling algorithms, and machine learning integration.

### How many samples can be multiplexed in a single direct RNA sequencing run?

Multiplexing up to 96 barcodes yields highly accurate demultiplexing models with current tools. The number of samples that can be multiplexed depends on the barcoding strategy and the sequencing depth required per sample. Higher multiplexing reduces reads per sample, which may be limiting for low-abundance transcripts or modification detection.

### What training resources are available for direct RNA sequencing analysis?

Training resources are available through multiple platforms. The [EMBL-EBI Training](https://www.ebi.ac.uk/training) program provides bioinformatics learning pathways and practical analysis education. The [Galaxy Training Network](https://training.galaxyproject.org/) provides accessible workflow training and analysis tutorials. The [Carpentries Lessons](https://carpentries.org/lessons) provide foundational computing and programming training. [Bioconductor](https://bioconductor.org/) provides official package and workflow documentation for reproducible genomic analysis.

## Related Bioinformatics Guides

- [Oxford Nanopore Sequencing: From Sample to Base Calls](/knowledge/bioinformatics/oxford-nanopore-sequencing-from-sample-to-base-calls)
- [RNA-Seq Normalization Methods: TPM, RPKM, and Beyond](/knowledge/bioinformatics/rna-seq-normalization-methods-tpm-rpkm-and-beyond)
- [How to Choose a Long-Read Sequencing Platform: PacBio vs Oxford Nanopore](/knowledge/bioinformatics/how-to-choose-a-long-read-sequencing-platform-pacbio-vs-oxford-nanopore)
- [Full-Length Transcript Sequencing: Unraveling Isoform Diversity with Long Reads](/knowledge/bioinformatics/full-length-transcript-sequencing-unraveling-isoform-diversity-with-long-reads)
- [Metagenomic Assembly Overview: Challenges and Applications](/knowledge/bioinformatics/metagenomic-assembly-overview-challenges-and-applications)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)


## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [Advances in nanopore direct RNA sequencing and its impact on biological research.](https://pubmed.ncbi.nlm.nih.gov/40930310). Biotechnology advances, 2025.
- [Nanopore direct RNA sequencing of human transcriptomes reveals the complexity of mRNA modifications and crosstalk between regulatory features.](https://pubmed.ncbi.nlm.nih.gov/40359935). Cell genomics, 2025.
- [Rapid and accurate demultiplexing of direct RNA nanopore sequencing data with SeqTagger.](https://pubmed.ncbi.nlm.nih.gov/39880590). Genome research, 2025.
- [Nanopore sequencing: flourishing in its teenage years.](https://pubmed.ncbi.nlm.nih.gov/39293510). Journal of genetics and genomics = Yi chuan xue bao, 2024.
- [RNA modifications detection by comparative Nanopore direct RNA sequencing.](https://pubmed.ncbi.nlm.nih.gov/34893601). Nature communications, 2021.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.