Oxford Nanopore Sequencing: From Sample to Base Calls
Oxford Nanopore sequencing is a long-read, real-time DNA and RNA sequencing technology that measures changes in electrical current as nucleic acid molecules pass through a protein nanopore embedded in a synthetic membrane. The workflow proceeds through four main stages: nucleic acid extraction and quality assessment, library preparation with adapter ligation, flow cell loading and sequencing on an instrument such as MinION, GridION, or PromethION, and computational base calling that converts raw electrical signals into nucleotide sequences. This article explains each stage in operational detail for students, researchers, analysts, and life-science professionals who need to plan, execute, and interpret Nanopore sequencing runs. The practical outcome is a protocol checklist that supports reproducible runs and informed troubleshooting.
At a Glance
The table below summarizes the core stages of the Oxford Nanopore sequencing workflow, the primary decisions at each stage, and the key quality indicators that determine whether to proceed to the next step.
| Workflow Stage | Primary Decisions | Key Quality Indicators |
|---|---|---|
| Nucleic acid extraction | Choose extraction method suited to sample type and read length goals | DNA concentration, purity ratios, high-molecular-weight integrity |
| Library preparation | Select ligation, rapid, amplicon, or Cas9-targeted approach | Adapter ligation efficiency, input quantity, fragment size distribution |
| Flow cell loading | Determine loading concentration and sequencing duration | Active pore count, read throughput, read length N50 |
| Base calling and analysis | Choose base caller, demultiplexing, and downstream pipeline | Q scores, alignment rate, genome coverage, variant concordance |
Understanding the Oxford Nanopore Sequencing Platform
Oxford Nanopore sequencing belongs to a class of technologies often described as third-generation or long-read sequencing. These platforms were developed to address constraints of earlier short-read methods, particularly the limited read length and the requirement for PCR amplification that can bias representation of complex genomic regions [11]. Nanopore sequencing measures ionic current changes as a single-stranded DNA or RNA molecule translocates through a nanoscale protein pore. The characteristic current disruption pattern produced by each nucleotide or nucleotide combination is recorded in real time and subsequently translated into a base sequence.
The technology has advanced rapidly since its first commercial release. Improvements in chemistry, pore design, and base-calling algorithms have produced substantial gains in accuracy, read length, and throughput [5]. These advances have enabled applications ranging from whole-genome assembly and full-length transcript detection to base modification identification and rapid clinical diagnosis [5]. The platform is also used for outbreak surveillance, where the portability and real-time data streaming of instruments such as the MinION support field-deployable sequencing [5][16].
For researchers deciding whether to adopt Nanopore sequencing, the relevant comparisons are against short-read platforms and other long-read technologies. Short-read sequencing offers high per-base accuracy but produces reads typically too short to resolve repetitive regions, structural variants, or full-length transcripts [11]. Nanopore sequencing generates reads that can span entire viral genomes, complex repeat arrays, and even megabase-scale genomic segments when ultra-long read protocols are used [7]. The tradeoff is that raw read accuracy is generally lower than short-read platforms, although recent chemistry and base-calling improvements have narrowed this gap [5][9].
Core Principles of Nanopore Sequencing Chemistry
The sequencing reaction occurs on a flow cell containing an array of microscaffolds, each with a protein nanopore embedded in an electrically resistant polymer membrane. An ionic current is applied across the membrane, and single-stranded nucleic acid molecules are processively threaded through the pore by a motor protein. As each nucleotide or short nucleotide motif passes through the pore's constriction zone, it modulates the ionic current in a characteristic way. The resulting current trace, often called a squiggle, is recorded thousands of times per second and provides the raw signal from which base calls are derived.
Base calling is the computational process of converting these electrical current measurements into nucleotide sequences. Modern base callers use neural network models trained on large datasets of known sequences to interpret the current traces. The accuracy of base calling depends on several factors, including pore version, sequencing chemistry, library quality, and the specific base-calling model applied. Base modification detection is possible because modified nucleotides, such as methylated cytosine, produce current signatures distinct from unmodified bases [5][7]. This capability allows direct detection of epigenetic marks without bisulfite conversion or immunoprecipitation, which are required by many short-read protocols.
The real-time nature of data acquisition is a defining operational feature. Sequence data become available as the run progresses, allowing researchers to monitor yield and quality during the run and to stop sequencing once sufficient data have been generated [16]. This capability is particularly valuable in diagnostic and outbreak contexts where time to answer is critical [14][16][19].
Nucleic Acid Extraction and Quality Assessment
The first decision in any Nanopore sequencing project is the method used to obtain nucleic acid of sufficient quantity, purity, and length. The optimal approach depends on the sample type, the target organism, and the read length required for the application.
High-Molecular-Weight DNA Extraction
For whole-genome sequencing and de novo assembly, high-molecular-weight (HMW) DNA is essential because read length is directly limited by fragment length. Standard column-based extraction kits often shear DNA into fragments too short for long-read applications. Protocols that use gentle lysis, minimal pipetting, and wide-bore tips help preserve long fragments. For challenging organisms such as filamentous Actinobacteria, which are difficult to lyse and contain high levels of polysaccharides and other contaminants, modified commercial kit protocols have been developed to consistently extract pure HMW DNA suitable for Nanopore sequencing and complete genome assembly [20].
The quality of extracted DNA should be assessed using three complementary measurements: concentration, purity ratios, and fragment size distribution. Concentration is typically measured by fluorometric methods that specifically detect double-stranded DNA. Purity is assessed by spectrophotometric absorbance ratios, with values around 1.8 for A260/A280 and around 2.0 for A260/A230 indicating low protein and organic contaminant carryover. Fragment size distribution is evaluated by gel electrophoresis or microfluidic analysis. For long-read sequencing, the presence of a high proportion of fragments above 20 kb is generally desirable, although the specific target depends on the application.
RNA Extraction and cDNA Synthesis
RNA sequencing with Nanopore platforms requires conversion of RNA to cDNA for most library preparation protocols. The quality of the cDNA synthesis step directly affects the yield and length of sequencing reads. For viral RNA samples with low titers or degraded RNA, the choice of reverse transcriptase and primer strategy is critical. Protocols using reverse transcriptases capable of long-length cDNA synthesis from a wide range of RNA amounts and quality have been shown to support whole-genome sequencing of viruses such as SARS-CoV-2 from low-titer samples [10].
Sample-Specific Considerations
Different sample types present distinct extraction challenges. Clinical specimens such as serum, cerebrospinal fluid, stool, and respiratory swabs contain variable amounts of host nucleic acid and potential PCR inhibitors [12][13][14]. For metagenomic applications, enrichment strategies such as virus-like particle purification and host nuclease digestion can enhance detection of viral nucleic acids [12]. For pathogen sequencing from blood or dried blood spots, host DNA depletion methods including methylation-sensitive restriction digestion and selective whole-genome amplification can increase the proportion of pathogen reads [17].
For Plasmodium falciparum sequencing from dried blood spots and whole blood, extraction method choice affects both yield and read length. One optimized protocol found that a Tween-Chelex method produced higher DNA yields than a commercial kit, while the commercial kit produced longer reads [17]. This tradeoff between yield and fragment length must be weighed according to the sequencing goals. The same study demonstrated that host DNA depletion methods increased the proportion of parasite DNA and enabled whole-genome sequencing with median read lengths above 2 kb and accuracy of 99.8% [17].
Library Preparation Strategies
Library preparation converts extracted nucleic acid into a sequencing-ready format by attaching adapters that enable the motor protein to engage the DNA and thread it through the pore. Several library preparation strategies are available, each with distinct tradeoffs in input requirements, protocol time, and compatibility with downstream applications.
Ligation-Based Library Preparation
The standard ligation approach requires the highest input quantity but provides the greatest flexibility for read length and throughput. DNA ends are repaired and dA-tailed, then sequencing adapters are ligated onto the prepared ends. This approach is compatible with native DNA, meaning that base modifications are preserved and can be detected during sequencing [5]. Ligation-based protocols are commonly used for whole-genome sequencing, including ultra-long read generation [7].
The input requirement for ligation-based protocols is typically higher than for rapid protocols. For targeted sequencing approaches such as nanopore Cas9-targeted sequencing (nCATS), approximately 3 micrograms of genomic DNA is required [6]. This method uses Cas9 to cleave DNA at specific target sites, enabling adapter ligation at those locations and enriching for regions of interest. The nCATS approach can simultaneously assess single-nucleotide variants, structural variations, and CpG methylation at targeted loci, achieving median coverage of 675-fold on a MinION flow cell and 34-fold on a smaller Flongle flow cell [6].
Rapid and Barcoding Kits
Rapid library preparation kits use a transposase-based approach that simultaneously fragments DNA and attaches adapters in a single step. This reduces protocol time and input requirements but introduces fragmentation that limits read length. Rapid kits are well suited to applications where speed is prioritized over maximum read length, such as diagnostic sequencing and amplicon-based workflows [10][18][19].
Barcoding kits enable multiplexing of multiple samples on a single flow cell. Adapters containing unique barcode sequences are ligated to each sample, allowing pooled sequencing and subsequent computational demultiplexing. The choice between native barcoding and rapid barcoding affects both input requirements and read length. Native barcoding preserves long fragments and native modifications, while rapid barcoding trades these features for shorter protocol time [12][18].
Amplicon-Based Approaches
Amplicon-based library preparation uses PCR to amplify specific genomic regions before adapter ligation. This approach is widely used for viral whole-genome sequencing, where tiled primer sets amplify overlapping fragments that cover the entire viral genome [10][16][18]. Amplicon strategies reduce the input requirement substantially and can generate complete genomes from clinical samples with high cycle threshold values, indicating low viral loads [14][19].
The design of primer sets is a critical decision in amplicon workflows. Tiled amplicon strategies use multiplexed primer pools to amplify overlapping fragments of defined length. For SARS-CoV-2, primer sets have been developed that support single- or double-tiled amplicons from 1.2 to 4.8 kb, with the flexibility to adapt to new variants [10]. For influenza viruses, primer sets targeting the hemagglutinin gene or the entire genome have enabled rapid genetic typing of field samples during outbreaks [16]. For hepatitis B virus, tiling-based PCR amplification followed by Nanopore sequencing has generated near-full-length genomes from residual diagnostic samples, providing both genotype information and intra-patient diversity data [18].
Amplicon-based workflows have also been developed for pathogens with segmented genomes. For severe fever with thrombocytopenia syndrome virus, three pairs of primers targeting terminal conserved regions of the three genome segments enabled enrichment of nearly whole viral genomes directly from clinical serum specimens [14]. This workflow achieved genome coverage of 98.69% and sequence identity above 99.91% compared with Sanger sequencing for a simulated sample with a cycle threshold of 35 sequenced for 10 minutes [14].
Cas9-Targeted Enrichment
Cas9-targeted sequencing offers an alternative to PCR-based enrichment that preserves native DNA and its modifications. The nCATS method uses Cas9 to cleave chromosomal DNA at user-defined target sites, creating defined ends for adapter ligation [6]. This approach avoids the amplification bias and loss of native modifications associated with PCR-based methods. The method has been applied to cell lines, xenografts, and primary human breast tissue, demonstrating simultaneous detection of haplotype-resolved single-nucleotide variants, structural variations, and CpG methylation [6].
The choice of enrichment strategy depends on the application. PCR-based amplicon approaches are faster and require less input but introduce amplification artifacts and lose native modifications. Cas9-targeted approaches preserve native information but require higher input and more complex protocol steps. For applications where base modification detection is essential, native approaches are preferred.
Flow Cell Loading and Sequencing
Flow cell loading is a critical step that directly determines sequencing yield and quality. The goal is to achieve a high density of active pores with efficient DNA capture while avoiding pore blockage and adapter dimer contamination.
Flow Cell Types and Capacity
Oxford Nanopore offers several flow cell formats with different pore counts and throughput characteristics. The MinION flow cell is the entry-level format, suitable for small genomes, targeted sequencing, and field applications. The GridION supports up to five flow cells run in parallel, while the PromethION supports larger numbers of high-capacity flow cells for population-scale projects [9][12]. The Flongle is a smaller, lower-cost flow cell designed for applications requiring modest throughput, such as targeted genotyping [6][13].
The choice of flow cell format affects both cost per run and throughput. For diagnostic genotyping of enteroviruses, Flongle sequencing of a region within the VP1 gene provided rapid genotyping of clinical samples [13]. For whole-genome sequencing of human samples at population scale, a single PromethION flow cell was sufficient to detect single-nucleotide polymorphisms with accuracy comparable to short-read sequencing [9].
Loading Concentration and Pore Occupancy
The concentration of the sequencing library loaded onto the flow cell is a key determinant of pore occupancy and throughput. Too low a concentration results in underutilized pores and reduced yield. Too high a concentration can cause multiple DNA molecules to enter a pore simultaneously, leading to pore blockage and reduced data quality. Optimal loading concentrations are typically determined empirically for each library type and flow cell format.
During the run, the number of active pores should be monitored. A healthy run typically shows a high proportion of pores actively sequencing, with a gradual decline over time as pores become blocked or inactive. The real-time data streaming capability of the platform allows operators to assess pore activity and read yield during the run and to decide when sufficient data have been collected [16].
Sequencing Duration and Stopping Criteria
The optimal sequencing duration depends on the application and the rate of data generation. For diagnostic applications where time to answer is critical, runs can be stopped as soon as sufficient coverage is achieved. Real-time alignment of sequences during the run allows operators to monitor genome coverage and stop sequencing once the target is met [16]. For example, a long amplicon-based workflow for severe fever with thrombocytopenia syndrome virus could finish in 10 hours from serum specimen to genome sequence [14].
For whole-genome assembly projects, longer runs generate more data and improve assembly contiguity. Ultra-long read protocols have produced reads with N50 values above 100 kb and maximum read lengths up to 882 kb [7]. Incorporating additional coverage of ultra-long reads more than doubled assembly contiguity in a human genome assembly project [7].
Base Calling and Demultiplexing
Base calling converts raw electrical current signals into nucleotide sequences. This computational step is performed either in real time during the run or post-run using stored raw signal data. The choice of base-calling model and parameters affects read accuracy and downstream analysis outcomes.
Base-Calling Models and Accuracy
Modern base callers use neural network architectures trained on large datasets of known sequences. The accuracy of base calling has improved substantially with successive chemistry and software releases [5][11]. For high-accuracy applications, the choice of base-calling model should be matched to the pore version and chemistry used in the run. Using an outdated model with new chemistry, or vice versa, can reduce accuracy.
Base modification detection is integrated into the base-calling process for some workflows. Modified bases produce characteristic current signatures that can be distinguished from unmodified bases [5]. The accuracy of modification detection depends on coverage depth and the specific modification being detected. For CpG methylation, haplotype-specific methylation calls have been generated at megabase scales using optimized protocols [9].
Demultiplexing Barcoded Samples
When multiple samples are pooled on a single flow cell using barcoded adapters, the resulting reads must be assigned to their source samples computationally. Demultiplexing uses the barcode sequences to classify reads, and the accuracy of this process depends on barcode quality and the stringency of the matching algorithm. Reads with low-quality barcodes may be unassigned or misassigned, which can affect downstream analysis if not properly handled.
The choice of barcoding kit affects demultiplexing outcomes. Native barcoding kits and rapid barcoding kits use different barcode sequences and adapter structures, and the demultiplexing software must be configured accordingly [12][18]. For diagnostic workflows, accurate demultiplexing is essential to ensure that clinical results are correctly attributed to the right patient sample.
Downstream Analysis and Interpretation
The output of base calling is a set of sequence reads in FASTQ format, which serve as input to downstream analysis pipelines. The choice of analysis approach depends on the application and the questions being addressed.
Read Alignment and Variant Detection
For applications that require comparison to a reference genome, reads are aligned to the reference and variants are identified. The long read lengths of Nanopore sequencing enable detection of structural variants that are difficult or impossible to resolve with short reads [5][9]. Small variant detection, including single-nucleotide polymorphisms and small insertions or deletions, is also possible, although accuracy within homopolymers and tandem repeats remains challenging [9].
The accuracy of variant detection depends on coverage depth, read quality, and the variant-calling algorithm used. For population-scale projects, optimized protocols have achieved single-nucleotide polymorphism detection with F1-scores comparable to short-read sequencing [9]. Structural variant detection with Nanopore data has achieved performance on par with state-of-the-art de novo assembly methods [9].
De Novo Assembly
For organisms without a reference genome, or for applications requiring reference-free analysis, de novo assembly reconstructs genome sequences from reads. The long reads generated by Nanopore sequencing are particularly valuable for assembly because they can span repetitive regions and resolve complex genomic structures [7][8].
Assembly software has been developed or adapted to handle Nanopore long reads. The SPAdes assembler, originally developed for short-read microbial genome assembly, has been extended to support hybrid assembly from short and long reads, including Oxford Nanopore data [8]. For human genome assembly, a combination of Nanopore reads and complementary short-read data produced assemblies with accuracy exceeding 99.8% [7].
The contiguity of assemblies is measured by metrics such as N50, which indicates the length of the shortest contig at which 50% of the assembly is contained. Ultra-long reads have been shown to substantially improve assembly contiguity. In one human genome assembly project, incorporating additional coverage of ultra-long reads more than doubled the assembly N50 from approximately 3 Mb to approximately 6.4 Mb [7].
Metagenomic and Targeted Analysis
For metagenomic applications, reads are classified taxonomically by comparison to reference databases. The long reads of Nanopore sequencing can improve taxonomic classification by spanning multiple conserved regions within a single read. Viral metagenomic protocols have generated near-complete genomes from clinical samples, with quality comparable to Sanger sequencing [12].
For targeted applications such as genotyping and antimicrobial resistance profiling, reads are compared to specific reference sequences or databases. Amplicon-based workflows have been developed for rapid genotyping of enteroviruses [13], foot-and-mouth disease virus [19], and hepatitis B virus [18]. These workflows provide genotype information that supports outbreak response and clinical management decisions.
Practical Implementation Steps
The following steps outline a practical approach to planning and executing an Oxford Nanopore sequencing run. These steps are intended as a protocol checklist that can be adapted to specific applications.
Step 1: Define the Sequencing Objective
Clearly define the biological question and the data requirements needed to answer it. Consider the target genome size, the required coverage depth, the read length needed to resolve the features of interest, and whether base modification detection is required. These decisions determine the choice of flow cell format, library preparation method, and sequencing duration.
Step 2: Select the Library Preparation Method
Choose the library preparation approach based on input quantity, read length requirements, and whether native modifications must be preserved. Ligation-based protocols are appropriate for whole-genome sequencing and applications requiring maximum read length. Rapid protocols are appropriate when speed is prioritized. Amplicon-based protocols are appropriate for targeted sequencing of known regions, particularly for viral genomes. Cas9-targeted enrichment is appropriate when native modifications must be preserved at targeted loci.
Step 3: Extract and Assess Nucleic Acid Quality
Extract nucleic acid using a method appropriate for the sample type and the read length goals. Assess concentration, purity, and fragment size distribution before proceeding to library preparation. For long-read applications, confirm that a sufficient proportion of fragments exceed the minimum length required for the application.
Step 4: Prepare the Sequencing Library
Follow the manufacturer protocol for the selected library preparation method. Record the input quantity, the yield after each purification step, and the final library concentration. Assess library quality by measuring concentration and, if possible, fragment size distribution. Confirm that adapter ligation was successful before loading the flow cell.
Step 5: Prime and Load the Flow Cell
Prime the flow cell according to the manufacturer instructions to establish the ionic current and prepare the pores for DNA capture. Load the sequencing library at the recommended concentration. Monitor the active pore count after loading to confirm that the flow cell is healthy and that DNA is being captured.
Step 6: Start Sequencing and Monitor the Run
Start the sequencing run and monitor key metrics including active pore count, read throughput, read length distribution, and base-calling quality. Use real-time data streaming to assess whether sufficient data have been generated for the application. Stop the run when the target coverage or data yield is achieved.
Step 7: Base Call and Demultiplex
Base call the raw signal data using the appropriate model for the pore version and chemistry. If barcoded samples were pooled, demultiplex the reads and assign them to source samples. Assess read quality metrics and filter reads according to the requirements of the downstream analysis.
Step 8: Perform Downstream Analysis
Align reads to a reference genome, assemble reads de novo, or classify reads taxonomically according to the application. Assess coverage, variant concordance, and assembly metrics to evaluate data quality. Document the analysis parameters and results for reproducibility.
Records and Measurements
Maintaining detailed records of each sequencing run is essential for troubleshooting, quality assurance, and reproducibility. The following measurements should be recorded for every run.
Pre-Run Records
Record the sample identifier, sample type, extraction method, and extraction quality metrics including concentration, purity ratios, and fragment size distribution. Record the library preparation method, kit lot numbers, input quantity, and final library concentration. Record the flow cell type, lot number, and the number of active pores before loading.
Run-Time Records
Record the loading concentration, the number of active pores at the start of the run, and the sequencing duration. Monitor and record read throughput over time, read length N50, and base-calling quality scores. Record any anomalies such as sudden drops in pore activity or unusual current traces.
Post-Run Records
Record the total data yield, the final read length distribution, and the base-calling accuracy metrics. For barcoded runs, record the number of reads assigned to each barcode and the proportion of unassigned reads. Record the downstream analysis results including coverage, variant counts, and assembly metrics.
These records support comparison across runs and enable identification of systematic issues. For diagnostic applications, records must also meet regulatory and quality management requirements, including traceability of samples and results [18][21].
Common Failure Patterns and Troubleshooting
Several failure patterns recur across Nanopore sequencing projects. Recognizing these patterns and understanding their causes enables efficient troubleshooting.
Low Pore Occupancy
Low pore occupancy results in reduced throughput and may indicate that the library concentration was too low, that adapters were not ligated efficiently, or that the flow cell was not primed correctly. Check the library concentration and adapter ligation efficiency. Confirm that the priming step was performed correctly and that no air bubbles were introduced into the flow cell.
Short Read Lengths
Read lengths shorter than expected may indicate DNA fragmentation during extraction or library preparation, or may result from using a rapid kit that fragments DNA by design. Assess the fragment size distribution of the input DNA and the final library. For applications requiring long reads, use extraction and library preparation methods that preserve high-molecular-weight DNA.
High Adapter Dimer Content
Adapter dimers are short reads consisting primarily of adapter sequences with little or no insert. High adapter dimer content reduces throughput and may indicate that the adapter concentration was too high relative to the input DNA, or that the purification steps did not effectively remove excess adapters. Optimize the adapter-to-input ratio and confirm that purification steps are performed correctly.
Base-Calling Accuracy Issues
Lower than expected base-calling accuracy may result from using an outdated base-calling model, from poor library quality, or from pore degradation during the run. Confirm that the base-calling model matches the pore version and chemistry. Assess library quality and consider whether the run duration exceeded the useful lifetime of the flow cell.
Barcode Misassignment
Misassignment of reads to the wrong barcode can occur when barcode sequences are similar, when barcode quality is low, or when the demultiplexing parameters are not optimal. Use the recommended demultiplexing parameters for the specific barcoding kit. Assess barcode quality and consider whether the number of samples multiplexed exceeds the recommended limit.
Low Genome Coverage
Insufficient genome coverage may result from low throughput, from uneven amplification in amplicon-based workflows, or from host DNA contamination in metagenomic applications. For amplicon workflows, assess the balance of coverage across amplicons and optimize primer pools if needed. For metagenomic applications, consider host DNA depletion strategies to increase the proportion of target reads [17].
Limitations and Interpretation Boundaries
Oxford Nanopore sequencing has limitations that affect data interpretation and application suitability. Understanding these boundaries is essential for designing experiments and interpreting results.
Error Profiles
Nanopore sequencing has a distinct error profile compared with short-read platforms. Errors are not uniformly distributed across the genome. Homopolymers and tandem repeats are particularly challenging for base calling, and small insertion and deletion calling within these regions remains difficult [9]. Single-nucleotide polymorphism detection accuracy is comparable to short-read sequencing when sufficient coverage is achieved, but small indel calling in repetitive regions requires caution [9].
Input Quantity Requirements
Different library preparation methods have different input requirements. Ligation-based protocols and Cas9-targeted enrichment require higher inputs than rapid or amplicon-based approaches. The nCATS method requires approximately 3 micrograms of genomic DNA [6]. For samples with limited nucleic acid, amplicon-based approaches may be the only viable option, but they introduce amplification bias and lose native modifications.
Throughput and Cost Considerations
The throughput of Nanopore sequencing varies by flow cell format and run duration. While a single PromethION flow cell can support population-scale projects [9], the cost per base may be higher than short-read platforms for some applications. The choice of platform should consider the total cost of the workflow, including library preparation, sequencing, and analysis.
Bioinformatics Complexity
The analysis of Nanopore data requires specialized bioinformatics skills and computational resources. Protocol standardization and the development of easy-to-use pipelines are ongoing needs for routine applications [11]. For diagnostic contexts, the bioinformatics complexity of Nanopore workflows is a recognized challenge that must be addressed through training and standardized analysis pipelines [21].
Safety and Regulatory Context
Nanopore sequencing workflows involve several safety and regulatory considerations that vary by application and jurisdiction.
Laboratory Safety
Standard molecular biology safety practices apply to Nanopore sequencing workflows. This includes proper handling of biological samples, use of personal protective equipment, and adherence to biosafety levels appropriate for the sample types being processed. For clinical samples, additional precautions may be required to prevent exposure to infectious agents.
Data Sharing and Privacy
Genomic data generated by Nanopore sequencing may be subject to data sharing policies and privacy regulations. The National Institutes of Health Genomic Data Sharing Policy outlines expectations for the sharing of genomic data generated with NIH funding [3]. Researchers should be aware of the data sharing requirements applicable to their funding sources and should plan for data deposition in appropriate repositories such as those maintained by the National Center for Biotechnology Information [2].
Data Management and Reproducibility
The FAIR Guiding Principles provide a framework for making data findable, accessible, interoperable, and reusable [4]. Applying these principles to Nanopore sequencing data involves documenting the experimental and computational protocols, depositing raw and processed data in public repositories, and using standard file formats and metadata schemas. Training resources for data management and bioinformatics are available through organizations such as the European Bioinformatics Institute [1].
Diagnostic Applications
For diagnostic applications, Nanopore sequencing workflows must meet regulatory requirements for clinical testing. This includes validation of the analytical and clinical performance of the assay, implementation of quality management systems, and adherence to reporting requirements. The use of Nanopore sequencing in clinical diagnostics is an active area of development, with protocols being optimized for challenging sample types and rapid turnaround times [18][21].
Professional Escalation Criteria
Researchers and analysts should escalate issues to supervisors, collaborators, or technical support when certain conditions are met. The following criteria indicate when professional escalation is appropriate.
Escalate When Data Quality Is Insufficient
If base-calling accuracy, read length, or throughput falls below the thresholds required for the application, escalate to technical support or experienced colleagues. Do not proceed with downstream analysis if the data quality is insufficient to support the biological conclusions.
Escalate When Results Are Inconsistent
If replicate runs produce inconsistent results, or if variant calls conflict with orthogonal validation data, escalate to investigate potential sources of systematic error. Inconsistency may indicate issues with sample handling, library preparation, or analysis parameters.
Escalate When Diagnostic Results Are Ambiguous
For diagnostic applications, ambiguous or unexpected results should be escalated to the responsible clinician or laboratory director. Confirmatory testing with an orthogonal method may be required before reporting results. The concordance of Nanopore sequencing with Sanger sequencing has been demonstrated for several applications [12][14], but confirmatory testing remains appropriate for critical clinical decisions.
Escalate When Regulatory Requirements Are Unclear
If the applicable data sharing, privacy, or diagnostic regulatory requirements are unclear, escalate to institutional compliance officers or regulatory affairs specialists. Do not assume that standard practices are sufficient for all applications and jurisdictions.
Frequently Asked Questions
What is the difference between Oxford Nanopore sequencing and short-read sequencing?
Oxford Nanopore sequencing produces long reads, often exceeding 10 kb and reaching hundreds of kilobases with optimized protocols, while short-read platforms typically produce reads of 150 to 300 bases [7][11]. Nanopore sequencing also enables real-time data streaming and direct detection of base modifications without chemical conversion [5]. Short-read platforms generally offer higher per-base accuracy but cannot resolve repetitive regions, structural variants, or full-length transcripts as effectively as long-read platforms [11].
How much DNA is required for Oxford Nanopore sequencing?
The input requirement depends on the library preparation method. Ligation-based protocols and Cas9-targeted enrichment require higher inputs, with the nCATS method requiring approximately 3 micrograms of genomic DNA [6]. Rapid and amplicon-based protocols require substantially less input, and amplicon approaches can generate complete viral genomes from clinical samples with low viral loads [10][14][19].
How long does an Oxford Nanopore sequencing run take?
Run duration is flexible and depends on the application and the rate of data generation. Diagnostic workflows can be completed in as little as 10 hours from sample to genome sequence [14]. For amplicon-based outbreak tracking, real-time alignment allows the run to be stopped as soon as sufficient data are generated [16]. Whole-genome assembly projects may require longer runs to generate sufficient coverage, particularly when ultra-long reads are needed [7].
Can Oxford Nanopore sequencing detect DNA methylation and other base modifications?
Yes. Nanopore sequencing detects base modifications directly because modified nucleotides produce characteristic current signatures distinct from unmodified bases [5]. This capability has been used to detect CpG methylation at megabase scales with haplotype-specific resolution [9] and to simultaneously assess methylation alongside single-nucleotide variants and structural variations in targeted sequencing approaches [6].
What is the accuracy of Oxford Nanopore sequencing?
Accuracy has improved substantially with advances in chemistry, pore design, and base-calling algorithms [5][11]. Optimized protocols have achieved 99.8% accuracy for Plasmodium falciparum whole-genome sequencing [17] and assembly accuracy exceeding 99.8% for human genomes when combined with complementary short-read data [7]. Single-nucleotide polymorphism detection with Nanopore data has achieved F1-scores comparable to short-read sequencing, although small indel calling in homopolymers and tandem repeats remains challenging [9].
What is the difference between MinION, GridION, and PromethION?
These are different instrument formats with different throughput capacities. The MinION is a portable, entry-level device suitable for small genomes and field applications. The GridION supports up to five flow cells run in parallel. The PromethION supports larger numbers of high-capacity flow cells for population-scale projects [9][12]. The Flongle is a smaller, lower-cost flow cell format suitable for targeted applications requiring modest throughput [6][13].
How do I choose between amplicon-based and native library preparation?
The choice depends on the application. Amplicon-based approaches use PCR to amplify target regions, reducing input requirements and enabling sequencing of low-titer samples [10][14][19]. However, PCR introduces amplification bias and loses native base modifications. Native approaches, including ligation-based and Cas9-targeted methods, preserve native modifications and avoid amplification artifacts but require higher input quantities [6][7]. If base modification detection is required, native approaches are necessary.
What bioinformatics skills are needed for Nanopore data analysis?
Nanopore data analysis requires familiarity with command-line tools for base calling, demultiplexing, read alignment, variant calling, and assembly. The bioinformatics complexity of Nanopore workflows is a recognized challenge, particularly for diagnostic applications [21]. Protocol standardization and the development of easy-to-use pipelines are ongoing needs [11]. Training resources are available through organizations such as the European Bioinformatics Institute [1].
Related Bioinformatics Guides
- Long-Read Sequencing Technologies: PacBio and Oxford Nanopore
- Basecalling Algorithms for Nanopore Sequencing
- Single-Cell RNA Sequencing: From Bulk to Resolution
- Nanopore Adaptive Sampling for Targeted Pathogen Sequencing
- Master Guide: Single-Cell RNA Sequencing Bioinformatics Workflows
References and Further Reading
- EMBL-EBI Training. European Bioinformatics Institute.
- NCBI Data Resources. National Center for Biotechnology Information.
- Genomic Data Sharing Policy. National Institutes of Health.
- The FAIR Guiding Principles. Scientific Data.
- Nanopore sequencing technology, bioinformatics and applications.. Nature biotechnology, 2021.
- Targeted nanopore sequencing with Cas9-guided adapter ligation.. Nature biotechnology, 2020.
- Nanopore sequencing and assembly of a human genome with ultra-long reads.. Nature biotechnology, 2018.
- Using SPAdes De Novo Assembler.. Current protocols in bioinformatics, 2020.
- Scalable Nanopore sequencing of human genomes provides a comprehensive view of haplotype-resolved variation and methylation.. Nature methods, 2023.
- Universal whole-genome Oxford nanopore sequencing of SARS-CoV-2 using tiled amplicons.. Scientific reports, 2023.
- The Third-Generation Sequencing Challenge: Novel Insights for the Omic Sciences.. Biomolecules, 2024.
- A viral metagenomic protocol for nanopore sequencing of group A rotavirus.. Journal of virological methods, 2023.
- Retrospective Genotyping of Enteroviruses Using a Diagnostic Nanopore Sequencing Workflow.. 2024.
- Rapidly obtaining genome sequence of Severe Fever with Thrombocytopenia Syndrome virus directly from clinical serum specimen using long amplicon based nanopore sequencing workflow.. 2025.
- Developing a nanopore sequencing workflow for protein engineering applications. 2023.
- An amplicon-based nanopore sequencing workflow for rapid tracking of avian influenza outbreaks, France, 2020-2022.. 2024.
- Development of an Oxford nanopore sequencing technology-based whole genome sequencing method for Plasmodium falciparum to support malaria molecular surveillance.. Scientific Reports, 2026.
- An Oxford Nanopore Technology-Based Hepatitis B Virus Sequencing Protocol Suitable for Genomic Surveillance Within Clinical Diagnostic Settings. International Journal of Molecular Sciences, 2024.
- Rapid Sequence Identification of Foot-and-Mouth Disease Virus Utilizing FMDV-ONTAPS: The Oxford Nanopore Technologies Amplicon P1 Sequencing Protocol. Viruses, 2026.
- Extraction and Oxford Nanopore sequencing of genomic DNA from filamentous Actinobacteria. STAR Protocols, 2022.
- Oxford Nanopore Sequencing in pediatric emergency infectious diseases: from rapid diagnosis to precision medicine. Frontiers in Cellular and Infection Microbiology, 2026.
- High-performance protocol for ultra-short DNA sequencing using Oxford Nanopore Technology (ONT). Plos One, 2025.
- Protocol for high-throughput processing of fecal samples for long-read metagenomic sequencing using PacBio HiFi or Oxford Nanopore Technologies. STAR Protocols, 2026.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.