# Long-Read Assembly Algorithms: How Graph-Based Methods Exploit Nanopore and PacBio Data


## Key Takeaways

- Long-read assembly algorithms, such as Canu (Overlap-Layout-Consensus) and Flye (Repeat Graph), are essential for resolving complex genomes by spanning repetitive regions that short-read assemblers (de Bruijn graph) fail to resolve due to their reliance on k-mer matching and sensitivity to errors.
- Graph-based methods exploit the length of Nanopore and PacBio reads to traverse and resolve repeats, with Canu employing adaptive k-mer weighting and sparse graph construction to avoid collapsing diverged repeats, while Flye builds an explicit repeat graph from read paths.
- Error correction and post-assembly polishing are critical for long-read data, which inherently has higher per-base error rates than short reads; strategies include self-correction using reads themselves (e.g., Canu's approach) or alignment-based correction with short reads or polished contigs (e.g., GoldPolish-Target).
- Repeat resolution is fundamentally enabled by read length, allowing single reads to span repetitive elements and their flanking unique sequences, thereby providing direct evidence for correct genomic arrangement, though highly similar tandem repeats or extremely large repeats remain challenging.
- Practical assembly workflows involve rigorous data quality assessment (read length N50, error rate), read trimming, pre-assembly error correction, assembler selection (Canu for complex repeats, Flye for large genomes/limited memory), post-assembly polishing, and comprehensive quality assessment using metrics like contig N50 and completeness (e.g., BUSCO).
- Parameter tuning for Canu requires accurate genome size estimation, appropriate read length cutoffs, and careful consideration of the corrected error rate, while Flye's performance is influenced by genome size, coverage, and the number of polishing iterations.

---

Researchers and laboratory professionals working with Oxford Nanopore Technologies (ONT) and Pacific Biosciences (PacBio) sequencing platforms face a distinct computational problem: short-read assemblers collapse repeats, misplace contigs, and fail to resolve structural variation because their underlying algorithms assume uniform coverage and short-range connectivity. Long-read assembly algorithms address this by using graph-based methods that exploit the length of individual reads to span repetitive regions and by implementing error correction strategies that accommodate the higher per-base error rates of nanopore and PacBio data. This article explains how assemblers such as Canu and Flye differ from short-read approaches, describes the algorithmic adaptations that make long-read assembly feasible, and provides practical parameter tuning recommendations grounded in the current literature.

The practical outcome for readers is the ability to make informed decisions about assembly strategy, parameter selection, and quality assessment when working with long-read sequencing data. The content covers the algorithmic foundations of overlap-layout-consensus and de Bruijn graph approaches, the specific adaptations for noisy long reads, repeat resolution mechanisms, error correction and polishing strategies, and concrete workflow recommendations. The discussion draws on peer-reviewed literature describing assembly algorithms and on official documentation from bioinformatics training and workflow resources.

## At a Glance

The table below summarizes the key differences between short-read and long-read assembly approaches, the primary algorithmic strategies, and the practical considerations for parameter tuning.

| Assembly Approach | Core Algorithm | Error Handling | Repeat Resolution | Practical Parameter Focus |
| --- | --- | --- | --- | --- |
| Short-read (Illumina) | de Bruijn graph from k-mers | High base accuracy, low error correction need | Relies on coverage depth and paired-end information | K-mer size selection, coverage cutoff, insert size |
| Long-read overlap (Canu) | Overlap-layout-consensus with adaptive k-mer weighting | Explicit error correction before assembly | Read length spans repeats, graph avoids collapsing diverged repeats | Read length cutoff, corrected error rate, genome size estimate |
| Long-read graph (Flye) | Repeat graph constructed from error-prone reads | Iterative assembly graph construction and polishing | Repeat graph resolves repeats using read paths | Genome size, coverage, polishing iterations |

The choice between assembly strategies depends on sequencing platform, coverage depth, genome complexity, and available computational resources. Long-read assemblers generally require higher coverage to achieve comparable accuracy to short-read assemblies, but they produce more contiguous assemblies with better resolution of repetitive regions.

## The Long-Read Assembly Problem

### Why Short-Read Assemblers Fail on Long-Read Data

Short-read assembly algorithms were designed for sequencing data with high base accuracy and short read lengths, typically 150 to 300 base pairs. The de Bruijn graph approach fragments reads into k-mers and constructs a graph where nodes represent k-mers and edges represent overlaps of k-1 bases. This approach works well when reads are accurate and coverage is uniform, but it struggles with long-read data for several reasons.

Long reads from ONT and PacBio platforms have higher error rates than short reads, with error profiles dominated by insertions and deletions instead of substitutions. The de Bruijn graph approach requires exact k-mer matches to construct the graph, and sequencing errors create spurious k-mers that fragment the graph and produce disconnected components. Short-read assemblers typically handle this by filtering low-abundance k-mers, but this strategy becomes less effective when error rates exceed a few percent.

The recent advent of long-read sequencing technologies from PacBio and ONT has led to substantial improvements in accuracy and computational cost for genome sequencing, yet de novo whole-genome assembly still presents significant challenges related to result quality and computational demands. As sequencing accuracy and throughput advance, a continuous stream of new assembly tools has emerged, making platform and tool selection a critical decision for achieving high-quality genome reconstructions.

### The Promise of Long Reads

Long reads provide information that short reads cannot capture. A single read of 10 to 100 kilobases can span repetitive elements that are several kilobases in length, providing direct evidence for the arrangement of unique sequences flanking the repeat. This property is fundamental to repeat resolution in long-read assembly.

The development of long-read sequencing has enabled the near-complete assembly of individual chromosomes, known as telomere-to-telomere assembly, for many organisms. Until recently, genomes were typically assembled into fragments of a few megabases at best, but technological advances in long-read sequencing now permit the reconstruction of complete chromosome sequences. This progress has been driven by both improvements in sequencing technology and the development of assembly algorithms specifically designed for long-read data.

## Core Principles of Graph-Based Assembly

### Overlap-Layout-Consensus

The overlap-layout-consensus (OLC) approach is one of the oldest assembly strategies and has been adapted for long-read data. The algorithm proceeds in three stages. First, all pairs of reads are compared to identify overlaps, which are regions where the suffix of one read matches the prefix of another. Second, a layout is constructed by arranging reads according to their overlaps, creating a graph where nodes represent reads and edges represent overlaps. Third, a consensus sequence is derived from the multiple sequence alignment of reads in each layout path.

Canu, a successor of Celera Assembler, was specifically designed for noisy single-molecule sequences and introduces support for nanopore sequencing. Canu halves depth-of-coverage requirements and improves assembly continuity while reducing runtime by an order of magnitude on large genomes compared to Celera Assembler 8.2. These advances result from new overlapping and assembly algorithms, including an adaptive overlapping strategy based on tf-idf weighted MinHash and a sparse assembly graph construction that avoids collapsing diverged repeats and haplotypes.

The OLC approach is computationally intensive because it requires all-versus-all read comparison, which scales quadratically with the number of reads. For large genomes with deep coverage, this becomes a significant computational burden. Canu addresses this with adaptive k-mer weighting that reduces the number of comparisons needed while maintaining sensitivity for true overlaps.

### De Bruijn Graphs and Their Limitations

De Bruijn graph assemblers fragment reads into k-mers and construct a graph where each node represents a k-mer and each edge represents a k-1 base overlap between two k-mers. The graph is traversed to produce contigs, with branching points indicating repeats or sequencing errors.

For short-read data, de Bruijn graph assemblers are efficient and accurate because reads are highly accurate and k-mer frequencies provide information about coverage. However, for long-read data with higher error rates, the de Bruijn graph approach faces challenges. Sequencing errors create k-mers that do not match the true genome sequence, producing spurious branches and dead ends in the graph. Error correction can mitigate this problem, but it adds computational cost and can introduce its own errors.

### Repeat Graphs for Long Reads

Flye and similar assemblers use a different graph representation called a repeat graph. Instead of building the graph from k-mers of individual reads, the repeat graph is constructed by analyzing the paths that reads take through repetitive regions. This approach exploits the fact that long reads can span entire repeats, providing direct evidence for how repeats are connected to flanking unique sequences.

The repeat graph construction begins with error correction of the reads, followed by the identification of read overlaps and the construction of a graph where nodes represent unique sequences and edges represent connections through repeats. The graph is then simplified by resolving bubbles and other structures that arise from sequencing errors or genuine genomic variation.

This approach is particularly effective for genomes with large repeats, such as those found in plant and animal genomes. The repeat graph can represent complex repeat structures that would be collapsed or fragmented by de Bruijn graph assemblers.

## Error Correction Strategies

### Why Long Reads Need Error Correction

Long reads from ONT and PacBio platforms have higher error rates relative to short reads. If left unaddressed, subsequent genome assemblies may exhibit high base error rates that compromise the reliability of downstream analysis. The error profiles differ between platforms, with ONT reads historically having higher error rates than PacBio reads, although both have improved substantially in recent years.

Error correction in long-read assembly serves two purposes. First, it improves the accuracy of the assembled sequence by correcting errors in individual reads before assembly. Second, it improves the efficiency of the assembly process by reducing the complexity of the overlap graph, since corrected reads produce cleaner overlaps and fewer spurious connections.

### Pre-Assembly Error Correction

Pre-assembly error correction can be performed using either short reads or long reads. When short reads are available for the same genome, they can be used to correct errors in long reads by aligning short reads to long reads and using the consensus of the short-read alignments to identify and correct errors. This approach is effective but requires additional sequencing data and computational resources.

Self-correction uses the long reads themselves to correct errors. The reads are compared to each other, and regions where multiple reads agree are considered reliable, while regions where a single read disagrees with the consensus are flagged as potential errors. Canu uses this approach with its adaptive overlapping strategy, which identifies true overlaps between reads while filtering out spurious overlaps caused by sequencing errors.

### Post-Assembly Polishing

Post-assembly polishing corrects errors in the assembled contigs after the initial assembly is complete. Polishing tools align reads back to the assembled contigs and use the alignment information to identify and correct errors. This step is important because even with pre-assembly error correction, the assembled sequence may contain residual errors.

Several specialized error correction tools for genome assemblies have emerged, employing a range of algorithms and strategies to improve base quality. Despite these efforts, many genome assembly workflows still produce regions with elevated error rates, such as gaps filled with unpolished or ambiguous bases. GoldPolish-Target is a modular targeted sequence polishing pipeline that isolates and polishes user-specified assembly loci, offering a resource-efficient means for polishing targeted regions of draft genomes.

Experiments using Drosophila melanogaster and Homo sapiens datasets demonstrate that GoldPolish-Target can reduce insertion/deletion and mismatch errors by up to 49.2% and 55.4% respectively, achieving base accuracy values upwards of 99.9% with Phred score Q greater than 30. This polishing accuracy is comparable to the current state-of-the-art tool Medaka, while exhibiting up to 27-fold shorter run times and consuming 95% less memory on average. GoldPolish-Target offers the ability to target specific regions of an assembly for polishing, providing a computationally lightweight and highly scalable solution for base error correction.

## Repeat Resolution Mechanisms

### How Read Length Enables Repeat Resolution

The fundamental advantage of long reads for repeat resolution is that a single read can span an entire repetitive element, providing direct evidence for the unique sequences that flank the repeat. When a read spans a repeat, the assembly algorithm can determine the correct path through the repeat without ambiguity.

For example, consider a genome with a repeat element of 5 kilobases flanked by unique sequences A and B on one side and C and D on the other. A short read of 150 base pairs cannot span the repeat and provides no information about which unique sequences are adjacent to which. A long read of 10 kilobases can span the entire repeat and the flanking unique sequences, providing direct evidence that A is adjacent to C and B is adjacent to D.

### Graph-Based Repeat Resolution

Graph-based assemblers use the paths of reads through the graph to resolve repeats. In the overlap graph, a repeat appears as a region where multiple reads overlap in a complex pattern. The assembler must determine which reads represent the same genomic location and which represent different copies of the repeat.

Canu uses a sparse assembly graph construction that avoids collapsing diverged repeats and haplotypes. This is important because repeats that have diverged slightly between copies can be incorrectly merged if the assembler is not careful. The sparse graph construction preserves the distinction between similar but non-identical repeats, allowing the assembler to produce separate contigs for each copy.

The repeat graph approach used by Flye takes a different strategy. Instead of trying to resolve repeats during assembly, the repeat graph explicitly represents the repeat structure of the genome. The graph contains nodes for unique sequences and edges for connections through repeats, and the assembler uses the paths of reads through the graph to determine the correct arrangement of repeats.

### Limitations of Repeat Resolution

Despite the advantages of long reads, repeat resolution remains challenging for certain types of repeats. Tandem repeats, where the same sequence is repeated consecutively, can be difficult to assemble because the number of repeat copies cannot always be determined from read coverage alone. Very large repeats that exceed the length of individual reads cannot be spanned by a single read and require additional information from linked reads or optical mapping.

The telomere-to-telomere assembly era has highlighted both the progress and the remaining challenges in repeat resolution. While near-complete assembly of each chromosome is now possible for many organisms, additional developments are required to resolve remaining assembly gaps and to assemble non-diploid genomes.

## Practical Assembly Workflow

### Step 1: Data Quality Assessment

Before beginning assembly, assess the quality of the sequencing data. This includes checking read length distributions, coverage depth, and error rates. The [NCBI](https://www.ncbi.nlm.nih.gov/) provides resources for sequence data management and quality assessment that can be used to evaluate raw sequencing data.

For long-read data, the key quality metrics are read length N50, which is the read length at which 50% of the sequenced bases are in reads of that length or longer, and the error rate, which can be estimated by aligning reads to a closely related reference genome if one is available. Coverage depth should be calculated based on the estimated genome size, with higher coverage generally producing better assemblies.

### Step 2: Read Trimming and Filtering

Remove low-quality reads and trim adapters before assembly. Long-read sequencing platforms produce reads of variable quality, and filtering out very short reads or reads with high error rates can improve assembly quality. The specific filtering parameters depend on the assembler being used and the characteristics of the data.

### Step 3: Error Correction

Perform error correction on the long reads before assembly. This can be done using the self-correction approach implemented in Canu or using external error correction tools. The choice of error correction strategy depends on whether short-read data is available for the same genome and on the computational resources available.

### Step 4: Assembly

Run the assembly algorithm with parameters appropriate for the data. The key parameters for Canu include the estimated genome size, the minimum read length, and the corrected error rate. For Flye, the key parameters include the genome size and the coverage depth.

The choice of assembler depends on the characteristics of the data and the goals of the assembly project. Canu is well-suited for genomes with complex repeat structures and for projects that require graph-based assembly outputs. Flye is often faster and more memory-efficient, making it suitable for large genomes or projects with limited computational resources.

### Step 5: Polishing

Polish the assembled contigs to correct residual errors. This can be done using the reads that were used for assembly or using additional sequencing data. Polishing tools such as Medaka and GoldPolish-Target align reads to the assembled contigs and use the alignment information to identify and correct errors.

For targeted polishing of specific regions, GoldPolish-Target offers the ability to isolate and polish user-specified assembly loci, providing a resource-efficient means for improving base quality in regions that are known to have elevated error rates.

### Step 6: Quality Assessment

Assess the quality of the final assembly using metrics such as contig N50, which is the contig length at which 50% of the assembled bases are in contigs of that length or longer, and the number of contigs. Additional quality metrics include the completeness of the assembly, which can be assessed by checking for the presence of conserved single-copy genes, and the accuracy of the assembly, which can be assessed by aligning reads back to the assembly and checking for mismatches.

The [EMBL-EBI Training](https://www.ebi.ac.uk/training) provides learning pathways for bioinformatics analysis that cover genome assembly quality assessment and other practical analysis skills. These resources can be useful for researchers who are new to long-read assembly or who want to improve their analysis workflows.

## Parameter Tuning for Canu

### Genome Size Estimation

The genome size estimate is a critical parameter for Canu because it is used to calculate expected coverage and to set thresholds for read filtering and overlap detection. An accurate genome size estimate is important for optimal assembly performance.

Genome size can be estimated using flow cytometry, k-mer analysis of short-read data, or published values for closely related species. If the genome size estimate is too small, Canu may filter out legitimate reads or fail to detect true overlaps. If the estimate is too large, Canu may retain spurious reads and produce a fragmented assembly.

### Read Length Cutoff

The minimum read length parameter determines which reads are used for assembly. Reads shorter than the cutoff are discarded. A higher read length cutoff reduces the number of reads used for assembly, which can improve computational efficiency but may also discard useful information from shorter reads.

For genomes with large repeats, a higher read length cutoff may be beneficial because longer reads are more likely to span repeats. However, if the read length distribution is such that a high cutoff discards too many reads, the coverage may become insufficient for a good assembly.

### Corrected Error Rate

The corrected error rate parameter tells Canu what error rate to expect after error correction. This parameter affects the overlap detection and the assembly graph construction. If the corrected error rate is set too low, Canu may fail to detect true overlaps between reads that have higher error rates. If it is set too high, Canu may accept spurious overlaps and produce an incorrect assembly.

The default corrected error rate in Canu is appropriate for most PacBio and ONT data, but it may need to be adjusted for data with unusual error profiles. The error rate can be estimated by aligning a subset of reads to a reference genome or by examining the overlap statistics produced during the assembly.

### Coverage and Computational Resources

Canu requires substantial computational resources for large genomes. The runtime and memory usage depend on the genome size, the coverage depth, and the number of reads. Canu reduces runtime by an order of magnitude on large genomes compared to Celera Assembler 8.2, but it still requires significant resources for mammalian-sized genomes.

For large genomes, it may be necessary to use a computing cluster or a high-memory server. The [nf-core documentation](https://nf-co.re/docs) provides guidance on configuring and running bioinformatics workflows in cluster environments, which can be useful for researchers who need to run large assembly jobs.

## Parameter Tuning for Flye

### Genome Size and Coverage

Flye requires an estimate of the genome size and the coverage depth. The genome size estimate is used to determine the expected coverage and to set thresholds for read filtering. The coverage estimate is used to distinguish between sequencing errors and genuine genomic variation.

If the genome size estimate is inaccurate, Flye may produce an assembly with incorrect coverage estimates, leading to errors in the repeat graph construction. The coverage estimate can be refined during the assembly process, but an accurate initial estimate improves the quality of the final assembly.

### Read Filtering

Flye filters reads based on length and quality before assembly. The minimum read length parameter determines which reads are used, and the minimum coverage parameter determines the minimum number of reads required to support a region of the assembly.

For genomes with high coverage, a higher minimum coverage threshold can reduce the number of spurious connections in the repeat graph. However, if the coverage is variable across the genome, a high threshold may discard legitimate reads from low-coverage regions.

### Polishing Iterations

Flye includes a polishing step that corrects errors in the assembled contigs. The number of polishing iterations can be specified, with more iterations generally producing more accurate assemblies but requiring more computational time.

The optimal number of polishing iterations depends on the error rate of the sequencing data and the desired accuracy of the final assembly. For data with higher error rates, more polishing iterations may be needed to achieve the desired accuracy.

## Assembly Graph Outputs and Analysis

### GFA Format

Canu provides graph-based assembly outputs in graphical fragment assembly (GFA) format for analysis or integration with complementary phasing and scaffolding techniques. The GFA format represents the assembly graph, including the connections between contigs and the paths through repeats.

The combination of highly resolved assembly graphs with long-range scaffolding information promises the complete and automated assembly of complex genomes. Researchers can use the GFA output to visualize the assembly graph, identify unresolved regions, and integrate additional data such as Hi-C or optical mapping data for scaffolding.

### Analyzing Assembly Graphs

Assembly graphs can be analyzed using specialized tools that visualize the graph structure and identify regions that require additional attention. The graph structure can reveal unresolved repeats, misassemblies, and regions of low coverage that may need additional sequencing.

The [Bioconductor](https://bioconductor.org/) project provides packages for genomic analysis that can be used to analyze assembly graphs and assess assembly quality. These packages offer reproducible workflows for genomic analysis and can be integrated into larger analysis pipelines.

## Common Failure Patterns

### Fragmented Assemblies

A fragmented assembly with many small contigs can result from insufficient coverage, high error rates, or complex repeat structures. If the assembly is fragmented, consider increasing the coverage by sequencing more, adjusting the error correction parameters, or using a different assembler.

### Misassemblies

Misassemblies occur when the assembler incorrectly joins sequences from different genomic locations. This can result from unresolved repeats, chimeric reads, or errors in the overlap detection. Misassemblies can be detected by aligning reads back to the assembly and looking for regions with inconsistent coverage or by comparing the assembly to a reference genome if one is available.

### Collapsed Repeats

Collapsed repeats occur when the assembler merges multiple copies of a repeat into a single sequence. This can result in an assembly that is shorter than the true genome and that has incorrect sequence in the repeat regions. Canu's sparse assembly graph construction is designed to avoid collapsing diverged repeats and haplotypes, but collapsed repeats can still occur with highly similar repeat copies.

### High Error Rates in Specific Regions

Some regions of the assembly may have elevated error rates even after polishing. These regions may correspond to gaps filled with unpolished or ambiguous bases. Targeted polishing with tools such as GoldPolish-Target can be used to improve the accuracy of these regions without the computational cost of polishing the entire assembly.

## Records and Measurements

### Assembly Quality Metrics

Record the following metrics for each assembly project to enable comparison and troubleshooting:

| Metric | Definition | Interpretation |
| --- | --- | --- |
| Contig N50 | Length at which 50% of assembled bases are in contigs of that length or longer | Higher values indicate more contiguous assemblies |
| Number of contigs | Total count of contigs in the assembly | Lower values indicate more complete assemblies |
| Assembly size | Total number of bases in the assembly | Compare to expected genome size to assess completeness |
| Base accuracy | Percentage of bases that match the true sequence | Assessed by read alignment or comparison to reference |
| Completeness | Percentage of conserved single-copy genes present | Assessed with BUSCO or similar tools |

### Computational Resource Records

Record the runtime, memory usage, and disk space required for each assembly. These records are useful for planning future assembly projects and for identifying parameter settings that are computationally efficient.

### Reproducibility Records

Document the exact parameters used for each assembly, including the software version, the parameter settings, and the input data. This documentation is essential for reproducibility and for troubleshooting assembly problems.

The [Carpentries lessons](https://carpentries.org/lessons) provide foundational training in reproducible research practices, including version control with Git and data management, which can be applied to genome assembly projects. The [Galaxy Training Network](https://training.galaxyproject.org/) also offers accessible workflow training that covers reproducible analysis practices.

## Limitations and Interpretation

### Coverage Requirements

Long-read assembly requires higher coverage than short-read assembly to achieve comparable accuracy. Canu halves depth-of-coverage requirements compared to Celera Assembler 8.2, but it still requires substantial coverage for large genomes. The optimal coverage depends on the error rate of the sequencing data and the complexity of the genome.

### Computational Demands

Long-read assembly is computationally intensive, particularly for large genomes. The all-versus-all read comparison in OLC assemblers scales quadratically with the number of reads, and the error correction and polishing steps add additional computational cost. Researchers should plan for substantial computational resources when assembling large genomes.

### Platform-Specific Considerations

The choice between PacBio and ONT sequencing platforms affects the assembly strategy. The two platforms have different error profiles, read length distributions, and throughput characteristics. The recent advances in both platforms have led to substantial improvements in accuracy and computational cost, but the optimal assembly strategy may differ between platforms.

### Interpretation Limits

An assembly is a reconstruction of the genome sequence, not a definitive representation. Regions with complex repeat structures, high error rates, or insufficient coverage may be incorrectly assembled. Researchers should interpret assembly results with caution and validate important findings with additional experiments.

## Professional Escalation Criteria

### When to Seek Additional Expertise

Consider seeking additional expertise or consulting with bioinformatics specialists in the following situations:

- The assembly produces highly fragmented contigs despite adequate coverage and reasonable parameter settings
- The assembly contains many misassemblies that cannot be resolved by parameter adjustment
- The genome contains complex repeat structures that cannot be resolved with the available data
- The computational resources required for assembly exceed the available capacity
- The assembly quality metrics do not meet the requirements for the intended downstream analysis

### When to Consider Additional Data

Additional sequencing data may be needed in the following situations:

- Coverage is insufficient for a complete assembly
- The genome contains very large repeats that exceed the read length
- The assembly contains gaps that cannot be resolved with the available data
- The genome is highly heterozygous or polyploid, requiring specialized assembly approaches

### When to Consider Alternative Tools

Alternative assembly tools may be appropriate in the following situations:

- The current assembler produces poor results with the available data
- The genome has specific characteristics that are better handled by a different assembler
- The computational resources required by the current assembler are prohibitive
- The assembly graph output is needed for integration with other data types

## A Decision Framework for Choosing Between Canu and Flye

### The Selection Problem

Researchers often struggle to choose between Canu and Flye because both assemblers handle long-read data but use fundamentally different algorithmic strategies. The decision is not about which assembler is better overall but about which one matches the specific characteristics of your data, genome, and available computational resources. A structured decision framework helps avoid wasted compute time and produces assemblies that match the intended downstream analysis.

### Decision Criteria and Thresholds

The primary criteria for assembler selection are genome size, repeat complexity, coverage depth, read length distribution, and computational budget. The table below summarizes the decision points.

| Criterion | Favor Canu | Favor Flye |
| --- | --- | --- |
| Genome size | Under 500 Mb | Over 500 Mb |
| Repeat complexity | High, diverged repeats | Moderate, uniform repeats |
| Coverage depth | 30x to 60x | 50x to 100x |
| Read length N50 | Over 15 kb | Over 10 kb |
| Computational budget | High memory available | Limited memory |
| Output requirement | GFA graph for scaffolding | Fast draft assembly |

These thresholds are starting points, not fixed rules. A genome with moderate complexity and ample compute resources may assemble well with either tool. The framework directs you to test the assembler that matches your dominant constraint first.

### Step 1: Characterize Your Data

Before selecting an assembler, quantify three data properties. First, estimate the genome size using k-mer analysis of short reads if available, flow cytometry, or published values for closely related species. Second, calculate the read length N50 from the sequencing summary file. Third, estimate the coverage depth by dividing the total number of sequenced bases by the estimated genome size.

Record these values in a simple table with the date, the sequencing platform, the flow cell or chip identifier, and the software version used for base calling. This record becomes the basis for the assembler decision and for troubleshooting if the assembly fails.

### Step 2: Apply the Primary Decision Rule

The primary decision rule prioritizes the constraint that is hardest to change. Computational resources are often the limiting factor for large genomes. If your compute node has less than 256 GB of RAM and the genome is over 500 Mb, start with Flye because it is generally more memory-efficient. If the genome is under 500 Mb and you have access to a high-memory node, start with Canu because its overlap-based approach tends to produce more contiguous assemblies for complex repeat structures.

Repeat complexity is the second decision point. If the genome is known to contain large segmental duplications, recent transposable element insertions, or highly heterozygous regions, Canu's sparse assembly graph construction is designed to avoid collapsing diverged repeats and haplotypes. Flye's repeat graph approach handles uniform repeats well but may struggle with highly diverged repeat copies.

### Step 3: Run a Pilot Assembly

Do not commit to a full assembly before testing. Extract a subset of reads representing 20x to 30x coverage and run both assemblers with default parameters on this subset. Compare the contig N50, the number of contigs, and the runtime. This pilot run typically takes a few hours and provides direct evidence for which assembler performs better on your specific data.

The pilot run also reveals parameter sensitivity. If Canu produces a fragmented assembly with the default corrected error rate, adjust the value and rerun. If Flye produces many small contigs, check whether the genome size estimate is accurate and whether the coverage threshold is appropriate.

### Step 4: Evaluate Graph Outputs

If the pilot assemblies are similar in contiguity, examine the assembly graph outputs. Canu provides graph-based assembly outputs in GFA format for analysis or integration with complementary phasing and scaffolding techniques. Flye also produces graph information that can be visualized. The graph structure reveals unresolved repeats and regions that may require additional sequencing.

The combination of highly resolved assembly graphs with long-range scaffolding information promises the complete and automated assembly of complex genomes. If you plan to integrate Hi-C or optical mapping data for scaffolding, the GFA output from Canu provides a direct interface for these downstream steps.

### Step 5: Document the Decision

Record the assembler choice, the parameter settings, the pilot assembly metrics, and the rationale for the decision. This documentation supports reproducibility and provides a reference for future assembly projects with similar data characteristics. The [Carpentries lessons](https://carpentries.org/lessons) provide foundational training in reproducible research practices, including version control with Git and data management, which can be applied to genome assembly projects.

### Troubleshooting the Decision Framework

If the chosen assembler produces poor results, work through the following checks before switching tools. First, verify that the genome size estimate is accurate. An inaccurate estimate affects coverage calculations and read filtering thresholds in both assemblers. Second, confirm that the read length distribution supports the assembler's requirements. Canu benefits from longer reads because its overlap detection improves with read length. Third, check whether the error rate of the sequencing data matches the assembler's assumptions. Data with unusually high error rates may require additional pre-assembly error correction.

If all checks pass and the assembly is still poor, consider whether the genome has characteristics that neither assembler handles well. Highly polyploid genomes, genomes with extreme heterozygosity, and genomes with very large tandem repeat arrays may require specialized assembly approaches beyond the scope of Canu and Flye.

### When to Escalate

Seek additional expertise or consider alternative tools when the pilot assemblies from both Canu and Flye produce contig N50 values below 100 kb for a microbial genome or below 1 Mb for a eukaryotic genome with adequate coverage. These thresholds indicate that the data or the genome complexity exceeds the capabilities of standard long-read assembly approaches. Additional sequencing data, such as ultra-long reads or linked reads, may be necessary to resolve the assembly.

The [Galaxy Training Network](https://training.galaxyproject.org/) offers accessible workflow training that covers reproducible analysis practices and can help researchers develop the skills needed to troubleshoot assembly problems. The [nf-core documentation](https://nf-co.re/docs) provides guidance on configuring and running bioinformatics workflows in cluster environments, which is useful when scaling assembly jobs to larger compute resources.

### Integrating the Framework into Your Workflow

The decision framework replaces ad hoc assembler selection with a structured process that produces defensible choices. The pilot assembly step is the most valuable component because it provides direct evidence from your data instead of relying on general recommendations. The documentation step ensures that the decision can be reviewed and improved for future projects.

The [Bioconductor](https://bioconductor.org/) project provides packages for genomic analysis that can be used to analyze assembly graphs and assess assembly quality. These packages offer reproducible workflows for genomic analysis and can be integrated into larger analysis pipelines. The [EMBL-EBI Training](https://www.ebi.ac.uk/training) provides learning pathways for bioinformatics analysis that cover genome assembly quality assessment and other practical analysis skills.

## Frequently Asked Questions

### What is the main difference between short-read and long-read assembly algorithms?

Short-read assembly algorithms use de Bruijn graphs constructed from k-mers, which require exact matches and are sensitive to sequencing errors. Long-read assembly algorithms use overlap-based or repeat graph approaches that can accommodate higher error rates and exploit read length to span repetitive regions. The choice of algorithm depends on the sequencing platform and the characteristics of the genome being assembled.

### Why do long reads require error correction before assembly?

Long reads from ONT and PacBio platforms have higher error rates relative to short reads, with error profiles dominated by insertions and deletions. If left unaddressed, these errors can produce spurious connections in the assembly graph and result in assemblies with high base error rates that compromise downstream analysis. Error correction improves the accuracy of individual reads and reduces the complexity of the assembly graph.

### How does read length help resolve repetitive regions in genomes?

A single long read can span an entire repetitive element and the unique sequences that flank it, providing direct evidence for the correct arrangement of sequences around the repeat. Short reads cannot span most repeats and provide no information about which unique sequences are adjacent to which. This property of long reads is fundamental to repeat resolution in genome assembly.

### What is the difference between pre-assembly error correction and post-assembly polishing?

Pre-assembly error correction corrects errors in individual reads before the assembly is constructed, improving the quality of the overlaps and the assembly graph. Post-assembly polishing corrects errors in the assembled contigs after the initial assembly is complete, by aligning reads back to the contigs and using the alignment information to identify and correct errors. Both steps are important for producing accurate assemblies.

### How do I choose between Canu and Flye for my assembly project?

The choice depends on the characteristics of your data and your assembly goals. Canu is well-suited for genomes with complex repeat structures and provides graph-based assembly outputs in GFA format. Flye is often faster and more memory-efficient, making it suitable for large genomes or projects with limited computational resources. Consider testing both assemblers with a subset of your data to determine which produces better results.

### What is targeted polishing and when should I use it?

Targeted polishing is a strategy for correcting errors in specific regions of an assembly instead of polishing the entire assembly. Tools such as GoldPolish-Target isolate and polish user-specified assembly loci, offering a resource-efficient means for improving base quality in regions with elevated error rates. This approach is useful when only certain regions of the assembly have quality problems, such as gaps filled with unpolished or ambiguous bases.

### What coverage depth do I need for long-read assembly?

The optimal coverage depth depends on the error rate of the sequencing data and the complexity of the genome. Canu halves depth-of-coverage requirements compared to Celera Assembler 8.2, but substantial coverage is still needed for large genomes. Higher coverage generally produces better assemblies, but the computational cost increases with coverage.

### How do I assess the quality of my long-read assembly?

Assembly quality is assessed using metrics such as contig N50, the number of contigs, assembly size, base accuracy, and completeness. Base accuracy can be assessed by aligning reads back to the assembly and checking for mismatches. Completeness can be assessed by checking for the presence of conserved single-copy genes. The [NCBI](https://www.ncbi.nlm.nih.gov/) and [EMBL-EBI Training](https://www.ebi.ac.uk/training) provide resources and training for genome assembly quality assessment.

## Related Bioinformatics Guides

- [Long-Read Metagenome Assembly: Overcoming Challenges with Nanopore and PacBio Data](/knowledge/bioinformatics/long-read-metagenome-assembly-overcoming-challenges-with-nanopore-and-pacbio-data)
- [Evaluating Metagenomic Assembly Tools: A Benchmarking Framework for Short-Read and Long-Read Data](/knowledge/bioinformatics/evaluating-metagenomic-assembly-tools-a-benchmarking-framework-for-short-read-and-long-read-data)
- [Hybrid Genome Assembly: Combining Short and Long Reads for Better Results](/knowledge/bioinformatics/hybrid-genome-assembly-combining-short-and-long-reads-for-better-results)
- [De Novo Genome Assembly with Long Reads: A Practical Workflow](/knowledge/bioinformatics/de-novo-genome-assembly-with-long-reads-a-practical-workflow)
- [How to Choose a Long-Read Sequencing Platform: PacBio vs Oxford Nanopore](/knowledge/bioinformatics/how-to-choose-a-long-read-sequencing-platform-pacbio-vs-oxford-nanopore)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)


## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [Advancements in long-read genome sequencing technologies and algorithms.](https://pubmed.ncbi.nlm.nih.gov/38608738). Genomics, 2024.
- [Genome assembly in the telomere-to-telomere era.](https://pubmed.ncbi.nlm.nih.gov/38649458). Nature reviews. Genetics, 2024.
- [GoldPolish-target: targeted long-read genome assembly polishing.](https://pubmed.ncbi.nlm.nih.gov/40055584). BMC bioinformatics, 2025.
- [Canu: scalable and accurate long-read assembly via adaptive k-mer weighting and repeat separation.](https://pubmed.ncbi.nlm.nih.gov/28298431). Genome research, 2017.
- [Minimap2: pairwise alignment for nucleotide sequences.](https://pubmed.ncbi.nlm.nih.gov/29750242). Bioinformatics (Oxford, England), 2018.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.