# How to Assemble a Genome with Oxford Nanopore Ultra-Long Reads: A Practical Workflow

## Scope and Reader Context

This article addresses researchers and laboratory professionals who need a practical genome assembly workflow for Oxford Nanopore ultra-long reads. The focus is the complete path from high molecular weight DNA extraction through library preparation, sequencing, basecalling, assembly, polishing, and quality assessment. The workflow choices described here apply to bacterial isolate sequencing, clinical pathogen characterization, veterinary diagnostics, and larger eukaryotic genome projects where read length directly affects assembly contiguity. The guidance draws on published interlaboratory evaluations and comparative studies of extraction kits and assemblers, with emphasis on decisions that laboratory staff control directly.

Ultra-long reads, commonly defined as reads exceeding 100 kilobases, provide the mapping span needed to resolve repetitive regions, structural variants, and complex genomic arrangements that short-read platforms cannot bridge. The practical challenge is that achieving and maintaining these read lengths requires coordinated decisions at every step, from cell lysis through sequencing chemistry selection. This workflow treats read length optimization, assembler choice, and polishing as connected decisions instead of isolated steps.

## At a Glance

| Workflow Stage | Primary Decision | Key Observation | Practical Consequence |
| --- | --- | --- | --- |
| DNA extraction | Extraction kit and protocol selection | Extraction method affects yield, integrity, and ultra-long read proportion | Fire Monkey extracts achieved highest N50 values, Genomic Tip gave highest sequencing yields, Nanobind produced the highest proportion of reads over 100 kb |
| DNA quality assessment | Purity and integrity measurement method | Digital PCR linkage assays at 100 kb and 150 kb distances predicted ultra-long read output better than pulse-field gel electrophoresis | Use linkage-based quantification when ultra-long read proportion is the priority |
| Library preparation | Input DNA quantity and fragmentation control | High molecular weight DNA requires gentle handling to preserve read length | Coverage depth, beyond read length, determines structural variant calling success |
| Sequencing | Flow cell and chemistry selection | Real-time data streaming enables adaptive decisions during runs | Monitor yield and read length distribution early to decide whether to extend runtime |
| Assembly | Assembler selection | Assembler choice affects genome completeness and downstream analysis | Flye outperformed Unicycler for antimicrobial resistance determinant detection in comparative testing |
| Polishing | Polishing tool and iteration count | Polishing corrects base-level errors after initial assembly | Genome completeness assessment with CheckM provides an independent quality check |

## Core Principles of Ultra-Long Read Assembly

### Why Read Length Matters for Genome Assembly

Genome assembly reconstructs a complete sequence from overlapping reads. The fundamental constraint is that repetitive regions shorter than the read length cannot be unambiguously resolved. When reads span an entire repeat plus flanking unique sequence, the assembler can place the repeat correctly. Ultra-long reads therefore directly improve assembly contiguity by bridging repeats that shorter reads cannot cross.

The practical consequence is measurable in assembly statistics. Projects using three complementary long-read and long-range datasets achieved N50 values exceeding 100 megabases for gapless contigs in human haplotypes, with substantial improvement in reconstruction of segmentally duplicated complex regions compared to earlier pangenome graphs. This outcome required coordinated use of multiple data types, but the read length contribution was central to resolving regions that remained fragmented in previous assemblies.

For bacterial genomes, the stakes are different but equally concrete. Complete circular chromosomes and plasmids can be assembled when reads span the entire replicon or when the assembly graph resolves the circular structure. The choice of assembler affects whether the full complement of resistance determinants is recovered from clinical isolates.

### The Relationship Between Input DNA and Read Length

The maximum possible read length is set by the DNA fragment size in the library. Fragmented DNA produces short reads regardless of sequencing chemistry improvements. This is why extraction method evaluation is the first decision point in the workflow.

An interlaboratory study tested four extraction methods with a reference cell line containing known chromosomal alterations. All methods produced samples of acceptable purity, but yield varied considerably between laboratories. This finding has a direct operational implication: the same kit can perform differently in different hands, so each laboratory should establish its own baseline performance instead of assuming published yields will transfer directly.

The same study found that a digital PCR assay with duplexes at 100 kb and 150 kb distances was predictive of ultra-long read output and provided a more quantitative readout than pulse-field gel electrophoresis, which varied in performance between instruments and gel dyes. Neither method predicted the proportion of short reads under 10 kb. This distinction matters because the factors that produce ultra-long reads are not the same factors that prevent DNA fragmentation.

### Coverage Requirements for Structural Variant Resolution

Read length alone does not guarantee successful variant detection. The interlaboratory study found that coverage was a key factor in structural variant calling success, but the effect depended on the specific variant caller used. Megabase-scale structural variants were challenging to analyze with standard callers and required confirmation based on coverage plots and mapping of junction sequences.

This finding shapes the practical workflow in two ways. First, target coverage should be set with the downstream analysis in mind, beyond assembly contiguity. Second, structural variant calls from automated tools should be treated as candidate findings that require visual confirmation when the variant is large or clinically significant.

## DNA Extraction for Ultra-Long Reads

### Extraction Method Selection

The extraction method determines the starting DNA quality for the entire workflow. Comparative testing of three commercial kits for clinical isolates of Pseudomonas aeruginosa and Enterobacter cloacae showed that the DNeasy kit yielded up to 4.7 times more DNA and approximately 50 percent higher sequencing output than MagAttract, while MagAttract produced higher DNA integrity and more contiguous assemblies.

This tradeoff between yield and integrity is the central decision in extraction method selection. High yield supports deeper sequencing coverage, but fragmented DNA limits read length and therefore assembly contiguity. The optimal choice depends on whether the project prioritizes total output or maximum read length.

The high GC content of Pseudomonas aeruginosa, approximately 67 percent, complicates DNA extraction and long-read sequencing with downstream effects on assembly and analysis. Laboratories working with high GC organisms should expect extraction to require more optimization and should verify DNA quality before proceeding to library preparation.

### Extraction Method Comparison Evidence

The interlaboratory evaluation of four extraction methods provides the most direct comparison available for ultra-long read workflows. Fire Monkey extracts achieved the highest N50 values, Genomic Tip gave the highest sequencing yields, and Nanobind produced the highest proportion of ultra-long reads over 100 kb. All four methods produced samples of acceptable purity, and library preparation and sequencing were successful for all four.

The practical interpretation is that no single method dominates across all metrics. A project prioritizing maximum read length for resolving complex repeats might choose Nanobind. A project needing maximum total yield for deep coverage might choose Genomic Tip. A project seeking balanced performance might choose Fire Monkey based on its N50 advantage.

The study also demonstrated that extraction method affects structural variant calling outcomes. This means the extraction decision propagates through the entire workflow and cannot be reversed at the assembly stage.

### Quality Assessment Before Library Preparation

DNA quality assessment serves two purposes: confirming that extraction produced usable DNA and predicting whether the library will generate ultra-long reads. The interlaboratory study found that digital PCR linkage assays provided a quantitative readout that predicted ultra-long read proportion, while pulse-field gel electrophoresis performance varied between instruments and gel dyes.

Spectrophotometry, fluorometry, and capillary electrophoresis are commonly applied approaches for evaluating DNA purity and integrity. Each measures a different property. Spectrophotometry detects protein and chemical contamination through absorbance ratios. Fluorometry measures double-stranded DNA concentration specifically. Capillary electrophoresis provides fragment size distribution information.

The limitation of these methods is that they do not directly measure the DNA properties that determine ultra-long read success. The digital PCR linkage approach addresses this gap by measuring whether DNA fragments retain physical linkage across defined distances. Laboratories that routinely perform ultra-long read projects should consider implementing this assay as a pre-sequencing quality gate.

## Library Preparation and Sequencing

### Input DNA Quantity and Handling

Library preparation protocols specify input DNA quantities, and these specifications should be followed precisely. The interlaboratory study found that library preparation and sequencing were successful for all four extraction methods tested, which suggests that the library preparation step is robust across reasonable input quality variation.

Gentle handling during library preparation preserves high molecular weight DNA. Pipetting with wide-bore tips, avoiding vortex mixing, and minimizing freeze-thaw cycles all reduce mechanical shearing. The effort invested in extraction quality is lost if the library preparation step fragments the DNA.

### Sequencing Platform and Chemistry Selection

Oxford Nanopore sequencing operates as a fourth-generation real-time genomic platform with ultra-long read lengths, real-time data streaming, and portability. These features enable comprehensive microbial characterization directly at the point of care, even from low-biomass clinical specimens.

The real-time data streaming capability has a practical workflow consequence: sequencing runs can be monitored and decisions made during the run. If yield is lower than expected early in the run, the runtime can be extended. If the read length distribution shifts toward shorter reads, the library quality issue can be identified before investing additional flow cell time.

For veterinary and field applications, the portability and minimal infrastructure requirements enable rapid on-site sequencing for diagnostics and surveillance. This capability is relevant for outbreak response where samples cannot be shipped to a central laboratory without delay.

### Real-Time Data Streaming and Adaptive Decisions

The real-time nature of nanopore sequencing changes how runs are managed compared to batch sequencing platforms. Data become available as reads are generated, which enables early quality assessment and run adjustment.

The practical decision points during a run include whether to extend runtime to achieve target coverage, whether to reload the flow cell, and whether to stop a run that is producing poor quality data. These decisions require defined thresholds before the run starts so that choices are made consistently instead of reactively.

## Basecalling and Read Processing

### Basecalling Options and Accuracy Considerations

Basecalling converts raw electrical signal data into nucleotide sequences. The choice of basecaller and model affects read accuracy, which in turn affects assembly quality and variant calling confidence.

The published literature on nanopore workflows consistently describes basecalling as a required step before assembly, but specific basecaller recommendations vary by project type and update frequency. Laboratories should establish a basecalling protocol that matches their sequencing chemistry version and verify that the basecaller output format is compatible with their downstream assembly tools.

### Read Length Filtering and Quality Filtering

Read processing typically involves filtering for read length and quality. The filtering thresholds should be set based on the assembly goals. A project targeting complete bacterial genomes might filter out reads below a minimum length threshold to reduce computational load. A project targeting structural variants might retain shorter reads to maximize coverage.

The interlaboratory study found that neither pulse-field gel electrophoresis nor digital PCR predicted the proportion of short reads under 10 kb. This means short read proportion is not predictable from pre-sequencing DNA quality measurements and must be assessed from the sequencing output directly.

## Assembler Selection

### Assembler Comparison Evidence

The choice of assembler has a measurable impact on assembly quality and downstream analysis success. Comparative testing of Unicycler and Flye for antimicrobial resistance determinant detection in Pseudomonas aeruginosa and Enterobacter cloacae found that Flye outperformed Unicycler, increasing detection by 2 to 14 percentage points across workflows.

The best-performing combination in that study was DNeasy extraction with Flye assembly, which achieved 95.2 percent antimicrobial resistance determinant detection compared to 67.8 percent for the MagMAX extraction with Unicycler assembly. This finding demonstrates that the extraction and assembly decisions interact, and the optimal combination is not necessarily the best individual components.

The study also found that the choice of assembly had a greater impact on the detection of antimicrobial resistance determinants than the extraction method alone. This means assembler selection deserves at least as much attention as extraction optimization when the downstream goal is resistance gene detection.

### Assembler Selection Criteria

Assembler selection should consider the genome size, expected complexity, computational resources, and downstream analysis requirements. For bacterial genomes, the choice between Unicycler and Flye can be made based on the comparative evidence available. For larger eukaryotic genomes, the assembler choice may be constrained by computational requirements and the need for specific features such as haplotype phasing.

The pangenome project that achieved N50 values exceeding 100 megabases used three complementary long-read and long-range datasets. This multi-data approach is not necessary for bacterial genomes but becomes relevant for complex eukaryotic projects where a single data type cannot resolve all regions.

### Computational Resource Requirements

Assembly of ultra-long reads requires substantial computational resources. The memory requirements scale with genome size and read depth. Bacterial genome assembly can typically be completed on a high-end workstation, while mammalian genome assembly requires server-class hardware or cloud computing.

The computational cost of assembly should be considered in project planning. Multiple assembler runs for comparison purposes multiply the computational requirement. The comparative evidence for assembler choice should be used to select one primary assembler instead of running multiple assemblers routinely.

## Polishing and Quality Assessment

### Polishing with Medaka

Polishing corrects base-level errors in the initial assembly using the sequencing reads. Medaka is a polishing tool designed for Oxford Nanopore data that uses a neural network model to predict corrected consensus sequences.

The polishing step is applied after the initial assembly and before final quality assessment. The number of polishing iterations should be determined based on the assembly quality metrics. Additional polishing rounds provide diminishing returns once the assembly reaches a plateau in quality scores.

### Assembly Quality Metrics

Assembly quality is assessed using multiple metrics. QUAST provides comprehensive assembly statistics including contiguity measures such as N50 and completeness measures such as the presence of expected genes. CheckM assesses genome completeness and contamination based on marker gene sets.

The comparative study of extraction kits and assemblers used QUAST for assembly quality assessment and CheckM for genome completeness evaluation. These tools provide independent quality checks that complement the assembly statistics reported by the assembler itself.

### Genome Completeness Assessment

Genome completeness assessment is particularly important for clinical and diagnostic applications where missing genetic content could lead to incorrect conclusions. CheckM estimates completeness based on the presence of lineage-specific marker genes.

For antimicrobial resistance determinant detection, genome completeness directly affects the sensitivity of resistance gene identification. An incomplete assembly may lack the plasmid or chromosomal region carrying resistance genes, leading to false negative results.

## Structural Variant Analysis

### Structural Variant Calling Approaches

Structural variant analysis from long-read assemblies uses the read mapping information to identify deletions, insertions, inversions, duplications, and translocations. The interlaboratory study found that structural variant calling success depended on coverage and the specific variant caller used.

The study also found that megabase-scale structural variants were challenging to analyze with standard callers and required confirmation based on coverage plots and mapping of junction sequences. This finding establishes a practical workflow requirement: large structural variant calls should be visually confirmed instead of accepted from automated output.

### Confirmation Strategies for Large Variants

Confirmation of large structural variants requires examining the underlying evidence. Coverage plots show whether read depth changes across the variant boundary. Junction sequence mapping shows whether reads span the predicted breakpoint.

The interlaboratory study noted that the findings of earlier studies were only partially reproduced, which underscores the importance of confirming structural variant calls instead of relying on published expectations. Each project should establish its own confirmation workflow for large variant calls.

## Common Failure Patterns and Troubleshooting

### Low Ultra-Long Read Proportion

The most common failure in ultra-long read projects is a lower than expected proportion of reads over 100 kb. The interlaboratory study found that the proportion of ultra-long reads varied by extraction method, with Nanobind producing the highest proportion. If ultra-long read proportion is low, the extraction method should be the first suspect.

The digital PCR linkage assay provides a pre-sequencing check for this failure mode. If the linkage percentage is low, the library will likely produce few ultra-long reads regardless of sequencing conditions.

### High Short Read Proportion

The proportion of short reads under 10 kb was not predictable from either pulse-field gel electrophoresis or digital PCR measurements in the interlaboratory study. This means short read contamination can appear even when pre-sequencing quality checks pass.

Short reads in the library can result from mechanical shearing during library preparation, nuclease contamination, or degradation during storage. If short read proportion is unexpectedly high, the library preparation handling should be reviewed.

### Low Sequencing Yield

Sequencing yield variation between laboratories was considerable in the interlaboratory study even when the same extraction method was used. This finding indicates that operator technique and laboratory conditions affect yield independently of the kit choice.

Low yield troubleshooting should examine the extraction protocol execution, quantification accuracy, and library preparation efficiency. The DNeasy kit produced up to 4.7 times higher DNA yield than MagAttract in the comparative study, so extraction kit choice is a relevant factor when yield is consistently low.

### Assembly Fragmentation

Assembly fragmentation appears as a lower N50 than expected for the read length distribution. The assembler choice affects contiguity, with Flye producing more contiguous assemblies than Unicycler in the comparative study.

If assembly fragmentation occurs despite good read length distribution, the assembler parameters or version may need adjustment. The comparative evidence supports Flye as the default choice for bacterial genomes when antimicrobial resistance detection is the downstream goal.

## Limitations and Interpretation Boundaries

### Accuracy Limitations of Nanopore Sequencing

Nanopore sequencing has inherent base-level accuracy limitations compared to short-read sequencing. The polishing step corrects many of these errors, but some regions may remain difficult to resolve accurately.

For clinical applications, the accuracy limitations mean that variant calls should be confirmed with an orthogonal method when the result affects patient management. The pediatric emergency infectious disease review noted that bioinformatics complexity and analytical standardization remain challenges for clinical integration.

### Host DNA Interference in Clinical Samples

Clinical and veterinary samples often contain host DNA that competes with pathogen DNA during sequencing. The veterinary pathogen detection review noted limitations in accuracy and host-DNA interference as current challenges.

For low-biomass clinical specimens, the proportion of pathogen reads may be very low, requiring deep sequencing to achieve sufficient pathogen coverage. Adaptive sampling approaches are emerging as a tool to enrich for pathogen reads, but these methods add workflow complexity.

### Reproducibility Considerations

The interlaboratory study found considerable yield variation between laboratories using the same extraction method. This finding has implications for reproducibility across sites and over time within a single laboratory.

Standardized protocols and defined quality thresholds improve reproducibility. The [Galaxy Training Network](https://training.galaxyproject.org/) provides accessible workflow training and analysis tutorials that support reproducible analysis practices. The [nf-core documentation](https://nf-co.re/docs) describes community pipeline standards for reproducible workflow configuration.

## Professional Escalation Criteria

### When to Seek Specialized Support

Laboratories should escalate to specialized support when assembly results consistently fail quality thresholds despite protocol optimization. The specific escalation triggers include persistent low ultra-long read proportion, assembly fragmentation that does not improve with assembler changes, and structural variant calls that cannot be confirmed with available tools.

The [NCBI](https://www.ncbi.nlm.nih.gov/) provides sequence resources and analysis services that can support assembly validation and comparison against reference databases. The [EMBL-EBI training](https://www.ebi.ac.uk/training) program offers bioinformatics learning pathways and data-resource training for laboratories building internal expertise.

### When to Use Managed Workflow Platforms

Managed workflow platforms reduce the bioinformatics burden for laboratories without dedicated computational staff. The [nf-core documentation](https://nf-co.re/docs) describes community pipelines that follow reproducibility standards and can be configured for specific project needs.

The [Galaxy Training Network](https://training.galaxyproject.org/) provides accessible workflow training that supports laboratories transitioning from manual analysis to managed platforms. The [Carpentries lessons](https://carpentries.org/lessons) provide foundational computing, data, shell, Git, and programming training that builds the skills needed to use these platforms effectively.

### When to Seek Collaboration for Complex Genomes

Complex eukaryotic genome projects may require collaboration with specialized genome centers. The pangenome project that achieved N50 values exceeding 100 megabases used multiple complementary datasets and substantial computational resources that would be challenging for a single laboratory to replicate.

The [Bioconductor](https://bioconductor.org/) project provides official package, workflow, installation, and reproducible genomic-analysis documentation that supports advanced analysis needs. Laboratories undertaking complex genome projects should assess whether their computational infrastructure and bioinformatics expertise match the project requirements before starting.

## Records and Documentation

### Required Records for Reproducible Workflows

Reproducible nanopore assembly workflows require documentation of the extraction method, DNA quality measurements, library preparation conditions, sequencing chemistry version, basecaller version and model, assembler version and parameters, and polishing settings.

The interlaboratory study demonstrated that extraction method affects sequencing outcomes, so the extraction method must be recorded with the sample metadata. The comparative study showed that assembler choice affects downstream analysis, so the assembler version and parameters must be documented.

### Quality Control Records

Quality control records should include the DNA concentration measurements, purity ratios, integrity assessments, and linkage assay results where available. These records support troubleshooting when sequencing outcomes deviate from expectations.

The digital PCR linkage assay provides a quantitative record that predicts ultra-long read proportion. Recording this measurement for each sample creates a dataset that can be used to evaluate extraction method performance over time.

### Assembly Quality Records

Assembly quality records should include the N50 value, total assembly length, number of contigs, genome completeness estimate, and contamination estimate. These metrics provide the basis for comparing assemblies across samples and over time.

The QUAST and CheckM outputs should be archived with the assembly files. The comparative study used these tools for quality assessment, establishing them as standard practice for bacterial genome assembly workflows.

## Safety and Regulatory Context

### Biosafety Considerations for Clinical Samples

Clinical samples require biosafety precautions appropriate for the suspected pathogen. The extraction and library preparation steps should be performed in facilities that match the biosafety level required for the organism.

The pediatric emergency infectious disease review described nanopore sequencing applications for acute respiratory, bloodstream, and central nervous system infections. These sample types may contain high-risk pathogens that require specific containment measures during processing.

### Data Handling and Privacy

Genomic sequence data from human samples carries privacy considerations. The pangenome project generated haplotypes from human individuals, and such data require appropriate consent and data access controls.

Laboratories working with human samples should follow applicable data protection requirements for genomic data. The [NCBI](https://www.ncbi.nlm.nih.gov/) provides guidance on data submission and access controls for human sequence data.

### Regulatory Considerations for Diagnostic Use

Nanopore sequencing for clinical diagnostic use may be subject to regulatory requirements that vary by jurisdiction. The veterinary pathogen detection review noted that routine veterinary deployment still faces uncertainty in study design, sample preparation, and interpretation thresholds across diverse hosts and sample matrices.

Laboratories implementing nanopore sequencing for diagnostic purposes should verify the regulatory status of their workflow in their jurisdiction. The bioinformatics complexity and analytical standardization challenges noted in the pediatric review are relevant to regulatory approval considerations.

## Decision Framework for Assembler and Polishing Selection

### Establishing a Structured Selection Process

The comparative evidence from published studies shows that assembler choice and polishing strategy directly affect assembly quality and downstream analysis outcomes. However, selecting the right combination for a specific project requires a structured decision process instead of relying on default choices or habits. This section provides a practical framework for evaluating assembler and polishing options against project-specific requirements.

The framework operates on three tiers. The first tier defines the assembly objectives in measurable terms. The second tier maps those objectives to assembler capabilities using published comparative evidence. The third tier establishes polishing and validation criteria that determine when an assembly is complete. Each tier produces records that support troubleshooting and reproducibility across projects.

### Tier One: Defining Assembly Objectives

Before selecting an assembler, define what the assembly must achieve. The objectives should be specific enough to guide tool selection and quality assessment. For bacterial isolate sequencing, the objectives typically include complete circular chromosome reconstruction, plasmid resolution, and accurate antimicrobial resistance determinant detection. For eukaryotic projects, the objectives may include contiguity thresholds expressed as N50 targets, complete reconstruction of specific complex regions, or haplotype resolution.

The comparative study of extraction kits and assemblers for Pseudomonas aeruginosa and Enterobacter cloacae demonstrated that assembler choice had a greater impact on antimicrobial resistance determinant detection than the extraction method alone. This finding establishes that assembly objectives must be defined before assembler selection because the assembler choice directly affects whether the downstream analysis goals are met.

Write the objectives as measurable criteria. Examples include a minimum N50 value, a maximum number of contigs, complete reconstruction of a specific plasmid, or detection of at least 95 percent of expected antimicrobial resistance determinants. These criteria become the acceptance thresholds for the assembly and the basis for comparing assembler performance.

### Tier Two: Mapping Objectives to Assembler Capabilities

The published comparative evidence provides a starting point for assembler selection. The study comparing Unicycler and Flye for clinical isolates found that Flye outperformed Unicycler, increasing antimicrobial resistance determinant detection by 2 to 14 percentage points across workflows. The best-performing combination was DNeasy extraction with Flye assembly, achieving 95.2 percent antimicrobial resistance determinant detection compared to 67.8 percent for MagMAX extraction with Unicycler assembly.

This evidence supports Flye as the default assembler for bacterial genomes when antimicrobial resistance detection is the downstream goal. However, the framework requires evaluating whether the published evidence applies to the specific project context. Consider the organism characteristics, expected genome complexity, and available computational resources.

For high GC content organisms such as Pseudomonas aeruginosa with approximately 67 percent GC content, the extraction and sequencing steps require additional optimization. The assembler choice interacts with these upstream decisions. The framework should therefore include a step for verifying that the selected assembler handles the expected genome characteristics before committing to a full assembly run.

For larger eukaryotic genomes, the assembler selection criteria differ. The pangenome project that achieved N50 values exceeding 100 megabases used three complementary long-read and long-range datasets. This multi-data approach required assemblers capable of integrating multiple data types and handling the computational load of large genome assembly. The framework should flag eukaryotic projects as requiring additional evaluation beyond the bacterial-focused comparative evidence.

### Tier Three: Polishing and Validation Criteria

Polishing corrects base-level errors in the initial assembly using the sequencing reads. The polishing strategy should be defined before the initial assembly completes so that the workflow proceeds without interruption. The number of polishing rounds and the specific polishing tool should be recorded for each assembly.

The validation criteria determine when the assembly is acceptable for downstream analysis. Assembly quality metrics include N50 value, total assembly length, number of contigs, genome completeness estimate, and contamination estimate. The comparative study used QUAST for assembly quality assessment and CheckM for genome completeness evaluation. These tools provide independent quality checks that complement the assembly statistics reported by the assembler itself.

The validation criteria should be set at the same time as the assembly objectives. If the assembly objectives specify a minimum N50 value, the validation step checks whether the polished assembly meets that threshold. If the objectives specify complete antimicrobial resistance determinant detection, the validation step includes running the resistance determinant identification tool and comparing the results against the expected set.

### Implementing the Decision Framework

The implementation steps for this framework are as follows. First, document the assembly objectives in a project record that includes the organism, expected genome size, downstream analysis requirements, and measurable acceptance criteria. Second, select the assembler based on published comparative evidence and the project context. For bacterial genomes with antimicrobial resistance detection as the goal, Flye is the evidence-supported choice. Third, define the polishing protocol including the tool, model, and number of rounds. Fourth, run the initial assembly and record the assembly statistics. Fifth, apply the polishing protocol and record the post-polishing statistics. Sixth, run the validation tools and compare the results against the acceptance criteria. Seventh, document the final assembly records and any deviations from the planned workflow.

The [Galaxy Training Network](https://training.galaxyproject.org/) provides accessible workflow training and analysis tutorials that support implementing structured assembly workflows. The [nf-core documentation](https://nf-co.re/docs) describes community pipeline standards that can be configured to implement this decision framework in a reproducible manner.

### Record System for Assembly Decisions

A structured record system supports consistent decision-making across projects and enables troubleshooting when assemblies fail to meet acceptance criteria. The record for each assembly should include the assembler name and version, the assembler parameters and any deviations from defaults, the polishing tool and model, the number of polishing rounds, the assembly statistics before and after polishing, the validation tool outputs, and the final acceptance decision.

The record should also include the rationale for the assembler selection. If the selection was based on published comparative evidence, cite the specific study and the relevant findings. If the selection was based on prior experience with similar organisms, document that experience. This rationale supports future decisions and helps other laboratory members understand why a particular assembler was chosen.

The interlaboratory study found considerable yield variation between laboratories using the same extraction method. This finding supports maintaining detailed records of the entire workflow, beyond the assembly step. The extraction method, DNA quality measurements, library preparation conditions, and sequencing parameters all affect assembly outcomes and should be recorded with the assembly records.

### Troubleshooting Within the Decision Framework

When an assembly fails to meet the acceptance criteria, the decision framework provides a structured troubleshooting path. The first step is to determine whether the failure is in the assembly step or an upstream step. Check the read length distribution and coverage against the project requirements. If the read length distribution is poor, the issue is likely in extraction or library preparation instead of assembly.

The interlaboratory study found that the proportion of ultra-long reads over 100 kb varied by extraction method, with Nanobind producing the highest proportion. If the ultra-long read proportion is below expectations, the extraction method should be reviewed before changing assembler settings.

If the read length distribution and coverage are adequate but the assembly is fragmented, the assembler choice or parameters should be reviewed. The comparative evidence supports Flye over Unicycler for bacterial genomes. If Flye produces fragmented assemblies despite adequate input data, check the assembler version and parameters against the documented configuration.

If the assembly is contiguous but the downstream analysis fails to detect expected features such as antimicrobial resistance determinants, the issue may be in the assembly accuracy instead of contiguity. The polishing protocol should be reviewed, and additional polishing rounds may be needed. The comparative study found that the choice of assembly had a greater impact on antimicrobial resistance determinant detection than the extraction method alone, which supports investing effort in assembly optimization when downstream detection is the goal.

### Comparison of Assembler Performance Across Project Types

The published evidence supports specific assembler choices for specific project types. For bacterial genomes with antimicrobial resistance detection as the goal, Flye outperformed Unicycler in the comparative study. For projects requiring complete reconstruction of complex eukaryotic regions, the pangenome project used multiple complementary datasets and achieved N50 values exceeding 100 megabases, but the specific assembler configuration was not the focus of the published description.

The framework should therefore distinguish between evidence-supported assembler choices and choices that require local evaluation. For bacterial genomes, the comparative evidence provides a clear default. For eukaryotic genomes, the assembler choice should be based on the specific project requirements and may require running multiple assemblers on a test dataset to compare performance.

The [Bioconductor](https://bioconductor.org/) project provides official package and workflow documentation that supports advanced genomic analysis needs. The [EMBL-EBI training](https://www.ebi.ac.uk/training) program offers bioinformatics learning pathways that can help laboratories build the expertise needed to evaluate assembler performance for non-standard project types.

### Common Failure Patterns in Assembler Selection

The most common failure pattern is selecting an assembler without defining the assembly objectives first. This leads to assemblies that are contiguous but fail to support the downstream analysis, or assemblies that support the downstream analysis but are unnecessarily fragmented. The decision framework prevents this failure by requiring objectives to be defined before assembler selection.

The second common failure pattern is applying published comparative evidence without checking whether the project context matches the study context. The comparative study of Unicycler and Flye used clinical isolates of Pseudomonas aeruginosa and Enterobacter cloacae. The findings may not transfer directly to other organisms with different genome characteristics. The framework addresses this by requiring a check of whether the published evidence applies to the specific project context.

The third common failure pattern is skipping the validation step. An assembly that meets contiguity thresholds may still miss important genetic content. The comparative study found that genome completeness directly affects the sensitivity of antimicrobial resistance determinant detection. The framework requires validation against the acceptance criteria before the assembly is considered complete.

### Professional Escalation Within the Decision Framework

When assemblies consistently fail to meet acceptance criteria despite following the decision framework, escalation to specialized support is appropriate. The escalation triggers include persistent assembly fragmentation with adequate input data, failure to detect expected genetic content despite contiguous assemblies, and computational resource limitations that prevent running the selected assembler.

The [NCBI](https://www.ncbi.nlm.nih.gov/) provides sequence resources and analysis services that can support assembly validation and comparison against reference databases. The [Carpentries lessons](https://carpentries.org/lessons) provide foundational computing and data training that can help laboratories build the skills needed to troubleshoot assembly issues independently.

For complex eukaryotic genome projects that exceed local computational capacity, collaboration with specialized genome centers may be necessary. The pangenome project that achieved N50 values exceeding 100 megabases used substantial computational resources that would be challenging for a single laboratory to replicate. The decision framework should include an assessment of whether the computational requirements for the selected assembler match the available infrastructure before starting the assembly run.

## Frequently Asked Questions

### What is the minimum DNA quality needed for ultra-long read sequencing?

The interlaboratory study found that all four extraction methods tested produced samples of acceptable purity, but yield varied considerably between laboratories. The key quality factors are DNA integrity, measured by fragment size distribution or linkage assays, and the absence of contaminants that inhibit library preparation. The digital PCR linkage assay at 100 kb and 150 kb distances predicted ultra-long read output and provides a quantitative threshold for accepting or rejecting a DNA sample before sequencing.

### How much coverage is needed for structural variant detection with nanopore reads?

The interlaboratory study found that coverage was a key factor in structural variant calling success, but the effect depended on the specific variant caller used. There is no single coverage threshold that applies across all callers and variant types. Megabase-scale structural variants required confirmation based on coverage plots and mapping of junction sequences, which means sufficient coverage must be available for visual inspection of the variant region.

### Which assembler should I use for bacterial genomes?

The comparative study of Unicycler and Flye for Pseudomonas aeruginosa and Enterobacter cloacae found that Flye outperformed Unicycler, increasing antimicrobial resistance determinant detection by 2 to 14 percentage points. The best-performing combination was DNeasy extraction with Flye assembly. The choice of assembly had a greater impact on resistance determinant detection than the extraction method alone, so assembler selection deserves careful consideration.

### How do I choose between DNA extraction kits for nanopore sequencing?

The interlaboratory study tested Fire Monkey, Nanobind, Puregene, and Genomic-tip methods. Fire Monkey extracts achieved the highest N50 values, Genomic Tip gave the highest sequencing yields, and Nanobind produced the highest proportion of ultra-long reads over 100 kb. The comparative study of three commercial kits found that DNeasy yielded up to 4.7 times more DNA and approximately 50 percent higher sequencing output than MagAttract, while MagAttract produced higher DNA integrity and more contiguous assemblies. The choice depends on whether the project prioritizes yield, read length, or assembly contiguity.

### How many polishing rounds are needed after assembly?

The polishing step corrects base-level errors using the sequencing reads. The number of rounds should be determined based on assembly quality metrics, with additional rounds providing diminishing returns once quality plateaus. The specific number depends on the assembler, basecaller, and sequencing chemistry used, so each laboratory should establish its own polishing protocol based on observed quality improvements.

### Can nanopore sequencing be used for clinical diagnostics?

Nanopore sequencing has been applied to pediatric emergency infectious diseases for rapid etiologic diagnosis, with features including ultra-long read lengths, real-time data streaming, and portability that enable comprehensive microbial characterization directly at the point of care. The veterinary pathogen detection review described applications for viral, bacterial, and parasitic pathogen detection in animals. However, both reviews noted current limitations including bioinformatics complexity, analytical standardization challenges, and host-DNA interference that must be addressed for routine clinical use.

### How do I confirm large structural variant calls?

The interlaboratory study found that megabase-scale structural variants were challenging to analyze with standard callers and required confirmation based on coverage plots and mapping of junction sequences. This means large variant calls should be visually inspected using the read mapping data instead of accepted from automated caller output. The study also noted that findings of earlier studies were only partially reproduced, which supports a confirmation workflow for every large variant call.

### What computational resources are needed for nanopore genome assembly?

Computational requirements scale with genome size and read depth. Bacterial genome assembly can typically be completed on a high-end workstation, while larger eukaryotic genomes require server-class hardware or cloud computing. The pangenome project that achieved N50 values exceeding 100 megabases used substantial computational resources and multiple complementary datasets. Laboratories should assess their computational capacity before starting assembly projects and consider managed workflow platforms such as those described in the [nf-core documentation](https://nf-co.re/docs) when local resources are limited.

## Related Bioinformatics Guides

- [De Novo Genome Assembly with Long Reads: A Practical Workflow](/knowledge/bioinformatics/de-novo-genome-assembly-with-long-reads-a-practical-workflow)
- [How to Choose a Long-Read Sequencing Platform: PacBio vs Oxford Nanopore](/knowledge/bioinformatics/how-to-choose-a-long-read-sequencing-platform-pacbio-vs-oxford-nanopore)
- [Long-Read Metagenome Assembly: Overcoming Challenges with Nanopore and PacBio Data](/knowledge/bioinformatics/long-read-metagenome-assembly-overcoming-challenges-with-nanopore-and-pacbio-data)
- [Long-Read Genome Assembly and Polishing Strategies](/knowledge/bioinformatics/long-read-genome-assembly-and-polishing-strategies)
- [Hybrid Genome Assembly: Combining Short and Long Reads for Better Results](/knowledge/bioinformatics/hybrid-genome-assembly-combining-short-and-long-reads-for-better-results)

## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [Interlaboratory evaluation of high molecular weight DNA extraction methods for long-read sequencing and structural variant analysis.](https://pubmed.ncbi.nlm.nih.gov/40722061). BMC genomics, 2025.
- [Comparison of three commercial DNA extraction kits and assemblers for AMR determinant detection in Pseudomonas aeruginosa and Enterobacter cloacae using long-read sequencing.](https://pubmed.ncbi.nlm.nih.gov/41177336). Journal of microbiological methods, 2025.
- [Oxford Nanopore Sequencing in pediatric emergency infectious diseases: from rapid diagnosis to precision medicine.](https://doi.org/10.3389/fcimb.2026.1793808). 2026.
- [Nanopore Sequencing in Veterinary Pathogen Detection: A Review of Technologies and Applications.](https://doi.org/10.3390/vetsci13030216). 2026.
- [Accessing medically relevant complex regions with a pangenome graph of 20 near-complete Japanese haplotypes.](https://doi.org/10.1038/s41467-026-73461-x). 2026.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.