Choosing the Right Assembler for Complex Genomes: A Comparison of hifiasm, Canu, and Flye

By Dr. Zubair Khalid, DVM, MS, PhD ·

Choosing the Right Assembler for Complex Genomes: A Comparison of hifiasm, Canu, and Flye

Key Takeaways

  • Read accuracy dictates assembler choice: PacBio HiFi reads (>99% accuracy) enable direct haplotype-aware assembly with hifiasm, bypassing computationally intensive error correction. Oxford Nanopore or older PacBio continuous long reads (85-95% accuracy) necessitate assemblers like Canu or Flye that incorporate error correction steps.
  • Heterozygosity necessitates phased assembly: For diploid genomes with heterozygosity >0.5%, collapsed assemblies create mosaic sequences that break gene models. hifiasm natively produces phased haplotypes (hap1, hap2), while Canu offers a haplotype-aware mode at higher computational cost; Flye produces collapsed assemblies.
  • Repeat content dictates read length requirements: Long reads are crucial to span repetitive regions; HiFi reads (10-25 kb) suffice for many complex genomes, while ultra-long Nanopore reads (>100 kb) are beneficial for genomes with exceptionally large repeats or segmental duplications.
  • Pre-assembly genome characterization is critical: K-mer analysis to estimate genome size, heterozygosity, and repeat content informs assembler selection and coverage targets. For HiFi data, a closed-loop estimation framework integrating FastK and GenomeScope 2.0 offers more robust estimates.
  • Coverage targets vary by assembler and data type: hifiasm typically requires 30-50x HiFi coverage, while Canu and Flye with continuous long reads benefit from 50-100x coverage due to the demands of error correction.
  • Assembly quality is multi-faceted: Assess contiguity (N50), completeness (BUSCO), and accuracy (QV), and validate by aligning reads back to the assembly. Failures in these metrics necessitate parameter adjustment, additional data, or assembler reassessment.

Genome assembly for complex genomes requires a deliberate choice among assemblers because repeat content, heterozygosity, and sequencing platform determine which tool will produce a usable result. This article compares hifiasm, Canu, and Flye for researchers working with plant, animal, and fungal genomes that contain high repeat fractions or elevated heterozygosity. The practical outcome is a decision framework based on input data type, genome characteristics, and project goals, with concrete quality checks and escalation criteria for when an assembly fails validation.

Context for Complex Genome Assembly

Complex genomes present assembly challenges that simple microbial genomes do not. The primary sources of difficulty are unresolved complex repeats and haplotype heterozygosity, both of which generate assembly errors when handled incorrectly. Long-read sequencing technologies are essential for maximizing genome quality in vertebrates and other complex taxa, as demonstrated by the Vertebrate Genomes Project's work across 16 species representing six major vertebrate lineages. That project confirmed that assemblies produced without adequate handling of repeats and heterozygosity contain substantial errors, missing sequence, and false gene duplications that are only corrected when long reads and appropriate assembly strategies are applied.

The practical implication for a researcher starting a genome project is that assembler selection is not a trivial default choice. The assembler determines how repeats are resolved, how haplotypes are separated or collapsed, and how much manual curation will be required downstream. For a genome with 75% repeat content, such as the cephalopod Sepiella japonica with an estimated genome size of approximately 4.3 Gb and heterozygosity near 0.8%, the assembler must handle both massive repeat fractions and moderate heterozygosity simultaneously. For a genome with lower repeat content but high heterozygosity, such as an outcrossing plant species, the assembler must separate haplotypes cleanly or the resulting assembly will contain haplotype switches that break gene models.

The comparison that follows is grounded in the actual behavior of hifiasm, Canu, and Flye on genomes with varying repeat content and heterozygosity. Each tool has a distinct algorithmic approach, input requirement, and output characteristic that makes it suitable for specific scenarios. Understanding these differences before starting an assembly project saves weeks of computational time and manual curation effort.

At a Glance: Assembler Comparison Table

FeaturehifiasmCanuFlye
Primary inputPacBio HiFi (CCS) readsOxford Nanopore or PacBio continuous long readsOxford Nanopore or PacBio continuous long reads
Haplotype handlingPhased assembly with hap1 and hap2 outputsCollapsed assembly with haplotype-aware mode availableCollapsed assembly, no native phasing
Repeat resolutionExcellent with HiFi accuracyGood with sufficient coverageGood with sufficient coverage
Computational demandModerate, optimized for HiFiHigh, requires significant CPU timeModerate, faster than Canu
Best use caseDiploid genomes with HiFi data, heterozygous samplesGenomes sequenced with older long-read platformsQuick draft assemblies, bacterial and fungal genomes, ONT data
Output formatGFA, FASTA for haplotypesFASTA, GFAFASTA, GFA
Polishing requirementMinimal with HiFi inputUsually requiredUsually required

This table reflects the general operational characteristics of each assembler. The sections below explain the algorithmic basis for these differences and provide concrete guidance for matching assembler to genome project.

Core Principles of Assembler Selection

Read Accuracy Determines Assembly Strategy

The accuracy of the input reads is the single most important factor in assembler selection. HiFi reads from Pacific Biosciences provide accuracy above 99% and are produced by circular consensus sequencing. These reads are long enough to span most repeats and accurate enough that the assembler does not need to spend computational effort correcting errors before assembly. hifiasm is designed specifically for this data type and exploits the high accuracy to perform haplotype-aware assembly directly from the read set.

Continuous long reads from Oxford Nanopore or older PacBio platforms have lower per-base accuracy, often in the range of 85% to 95% depending on the chemistry and basecalling model. Canu and Flye both accept these reads and include error correction steps in their pipelines. Canu performs its own error correction using read-to-read overlaps, which is computationally expensive but produces corrected reads that can then be assembled. Flye uses a different approach based on repeat graph construction from raw reads, which is faster but may leave more residual errors that require polishing.

The decision between hifiasm and the other two tools is therefore often decided by the sequencing platform available. If HiFi data are available, hifiasm is the appropriate choice for most complex genomes. If only continuous long reads are available, Canu or Flye are the options, and the choice between them depends on genome size, computational resources, and the tolerance for manual curation.

Heterozygosity Drives Haplotype Handling

Diploid genomes contain two copies of each chromosome, and these copies differ at polymorphic sites. The level of heterozygosity determines whether an assembler should collapse the two haplotypes into a single consensus sequence or separate them into two phased assemblies.

For a highly inbred or homozygous sample, collapsing haplotypes is acceptable and produces a single representative sequence. For an outcrossing or wild sample with heterozygosity above approximately 0.5%, collapsing haplotypes creates a mosaic sequence that switches between the two parental copies at heterozygous sites. This mosaic structure breaks gene models and creates false structural variants when the assembly is compared to other samples.

hifiasm addresses this problem by constructing a phased assembly that separates the two haplotypes into distinct output files. The assembler uses the HiFi read accuracy to identify heterozygous sites and phase reads into two groups before constructing contigs. The result is two haplotype-resolved assemblies, labeled hap1 and hap2, that can be used independently or combined with Hi-C data for chromosome-scale scaffolding.

Canu includes a haplotype-aware mode that can separate haplotypes, but this mode requires higher coverage and more computational time. Flye does not provide native haplotype separation and will collapse heterozygous regions into a single sequence. For a researcher working with a heterozygous sample and continuous long reads, Canu with haplotype-aware settings is the better choice, though the computational cost is substantial.

Repeat Content Determines Read Length Requirements

The repeat content of a genome sets the minimum read length needed to span repetitive regions. A repeat that is longer than the read length cannot be resolved by the assembler, and the resulting assembly will contain gaps or collapsed copies. For genomes with high repeat content, such as the 76% repeat fraction reported for Sepiella japonica, reads must be long enough to span the largest common repeat families.

HiFi reads are typically 10 to 25 kb in length, which is sufficient for many plant and animal genomes but may not span the largest repeats in some species. Continuous long reads from Oxford Nanopore can reach 100 kb or more, which provides better repeat spanning but at lower per-base accuracy. The tradeoff between read length and read accuracy is central to assembler selection.

hifiasm with HiFi reads provides the best balance for most complex genomes because the read accuracy enables correct assembly of the sequence between repeats, and the read length is sufficient for most repeat families. For genomes with very large repeats or segmental duplications, a hybrid approach using both HiFi and ultra-long Oxford Nanopore reads may be necessary, with the ultra-long reads providing the repeat-spanning information and the HiFi reads providing the accuracy.

Practical Workflow for Assembler Selection

Step 1: Characterize the Genome Before Assembly

Before selecting an assembler, estimate the genome size, heterozygosity, and repeat content using k-mer analysis of the sequencing data. This step is essential because it determines which assembler and parameters are appropriate. K-mer-based estimation tools applied to quality-filtered short reads can provide genome size, heterozygosity, repeat content, and GC content estimates, as demonstrated in the Sepiella japonica genome survey.

For HiFi reads, a closed-loop genome size estimation framework that integrates FastK and GenomeScope 2.0 can provide consistent estimates across different k-mer lengths. This approach is more reliable than single k-mer predictions because it identifies the convergence region where genome size estimates stabilize as k-mer length varies. The robustness of this method has been demonstrated across diploid and polyploid species, making it suitable for complex genomes where standard k-mer analysis may be misleading.

The genome survey should produce the following records for your project notebook:

  • Estimated genome size in megabases or gigabases
  • Heterozygosity rate as a percentage
  • Repeat content as a percentage
  • GC content as a percentage
  • Sequencing platform and read length distribution
  • Estimated coverage based on total sequenced bases divided by estimated genome size

These values directly inform assembler choice. A genome with heterozygosity above 0.5% and HiFi data available should use hifiasm. A genome with heterozygosity below 0.5% and continuous long reads should use Canu or Flye. A genome with very high repeat content may require a hybrid approach regardless of assembler.

Step 2: Match Assembler to Data Type

The sequencing platform determines the realistic assembler options. If the project has PacBio HiFi data, hifiasm is the primary choice for complex genomes. The assembler accepts the HiFi reads directly and produces phased assemblies without a separate error correction step. This saves substantial computational time compared to Canu, which must correct errors in continuous long reads before assembly.

If the project has Oxford Nanopore data, Flye is often the fastest path to a draft assembly, but the result will require polishing with short reads or HiFi reads to reach acceptable accuracy. Canu is a more conservative choice for Nanopore data because its error correction step produces higher-quality corrected reads, but the computational cost is significantly higher.

If the project has older PacBio continuous long reads, Canu is the appropriate choice because it was designed for this data type and includes the error correction needed to produce a usable assembly.

Step 3: Set Coverage Targets

Coverage is the ratio of total sequenced bases to the estimated genome size. For HiFi assembly with hifiasm, a coverage of 30x to 50x is typically sufficient for diploid genomes. Higher coverage may be needed for genomes with high heterozygosity because the assembler must have enough reads from each haplotype to phase them correctly.

For continuous long reads with Canu or Flye, a coverage of 50x to 100x is recommended because the error correction step consumes coverage. Canu's error correction requires sufficient read overlap to identify and correct errors, and low coverage results in uncorrected reads that produce fragmented assemblies.

The coverage calculation should be recorded before assembly begins. If the actual coverage is below the target for the chosen assembler, additional sequencing is needed before assembly, or the assembler choice should be reconsidered.

Step 4: Run the Assembly and Record Parameters

Each assembler has specific parameters that affect the output. Record the exact command, version, and parameters used for every assembly run. This record is essential for reproducibility and for troubleshooting when the assembly fails quality checks.

For hifiasm, the key parameters include the read input file, the number of threads, and the output prefix. The assembler produces multiple output files, including the primary assembly and the alternate haplotigs. The primary assembly is the collapsed representation, and the hap1 and hap2 files are the phased haplotypes.

For Canu, the key parameters include the genome size estimate, the read type, and the error correction settings. Canu uses the genome size estimate to calibrate its overlap and correction parameters, so an accurate estimate from the k-mer analysis is important.

For Flye, the key parameters include the read type, the genome size estimate, and the number of iterations for the repeat graph construction. Flye also produces an assembly graph that can be visualized to identify unresolved repeats.

Step 5: Assess Assembly Quality with Standard Metrics

After assembly, evaluate the output using standard quality metrics before proceeding to downstream analysis. The two most important metrics are contiguity and completeness.

Contiguity is measured by N50, which is the contig length at which half of the assembled bases are in contigs of that length or longer. A higher N50 indicates a more contiguous assembly. For example, the Digitalis purpurea genome assembly achieved an N50 of 4.3 Mb, which is considered good for a plant genome with substantial repeat content. The bighead catfish assembly achieved a scaffold N50 of 33.48 Mb after Hi-C scaffolding, demonstrating the value of adding chromatin conformation data for chromosome-scale assembly.

Completeness is measured by BUSCO, which assesses the presence of single-copy orthologs expected to be present in the genome. A BUSCO completeness of 95% or higher is generally considered good for a complex genome. The bighead catfish assembly achieved 95.5% BUSCO completeness, and the Digitalis purpurea assembly achieved approximately 96% complete BUSCO genes.

Additional quality metrics include the quality value (QV), which estimates the base-level accuracy of the assembly. The bighead catfish assembly achieved a QV of 50, which corresponds to an estimated error rate of 1 in 100,000 bases. This level of accuracy is achievable with HiFi data and appropriate polishing.

Step 6: Polish the Assembly if Needed

Polishing is the process of correcting residual errors in the assembly using additional sequencing data. For HiFi assemblies with hifiasm, polishing is often unnecessary because the input reads are already highly accurate. For assemblies from continuous long reads with Canu or Flye, polishing with short reads or HiFi reads is usually required to reach acceptable accuracy.

The bighead catfish assembly used Illumina short-read polishing after HiFi and Nanopore assembly, demonstrating that even with high-quality long reads, short-read polishing can improve accuracy. The polishing step should be recorded with the same rigor as the assembly step, including the tool, version, parameters, and input data.

Options and Tradeoffs: Detailed Assembler Comparison

hifiasm: The HiFi-Optimized Phasing Assembler

hifiasm is designed for PacBio HiFi reads and produces haplotype-resolved assemblies for diploid genomes. The assembler uses the high accuracy of HiFi reads to construct a phased assembly graph that separates the two haplotypes without requiring parental data or Hi-C data for phasing.

The primary advantage of hifiasm is its ability to produce two phased haplotypes from a single HiFi dataset. This is valuable for heterozygous samples where a collapsed assembly would create a mosaic sequence. The phased haplotypes can be used for downstream analysis of structural variation, allele-specific expression, and evolutionary studies.

The primary limitation of hifiasm is its dependence on HiFi data. If the project has only continuous long reads, hifiasm is not the appropriate choice. Additionally, hifiasm may produce more fragmented assemblies than Canu for genomes with very high repeat content, because the phasing process can break contigs at heterozygous sites within repeats.

For a researcher with HiFi data and a heterozygous diploid genome, hifiasm is the recommended choice. The assembler produces high-quality phased assemblies with minimal manual intervention, and the output is suitable for chromosome-scale scaffolding with Hi-C data.

Canu: The Conservative Error-Correcting Assembler

Canu is a long-read assembler that performs its own error correction before assembly. The assembler is designed for continuous long reads from Oxford Nanopore and PacBio platforms and is particularly suitable for genomes where read accuracy is lower and error correction is essential.

The primary advantage of Canu is its robustness. The error correction step produces high-quality corrected reads that can be assembled into contiguous contigs even for genomes with high repeat content. Canu also includes a haplotype-aware mode that can separate haplotypes, though this mode requires higher coverage and more computational time.

The primary limitation of Canu is its computational cost. The error correction step is computationally intensive and can take days or weeks for large genomes. For a genome of 4 Gb or larger, Canu may require substantial computational resources and time.

For a researcher with continuous long reads and a genome with high repeat content, Canu is a reliable choice if computational resources are available. The assembler produces high-quality assemblies but requires patience and careful parameter tuning.

Flye: The Fast Repeat-Graph Assembler

Flye is a long-read assembler that constructs a repeat graph from raw reads and resolves repeats through an iterative process. The assembler is designed for speed and works with both Oxford Nanopore and PacBio continuous long reads.

The primary advantage of Flye is its speed. The assembler is significantly faster than Canu for the same input data, making it suitable for quick draft assemblies and for genomes where computational resources are limited. Flye also handles high repeat content well because the repeat graph approach explicitly models repetitive regions.

The primary limitation of Flye is the accuracy of the output. Assemblies produced by Flye typically contain more residual errors than those produced by Canu or hifiasm, and polishing is almost always required. Flye also does not provide native haplotype separation, so heterozygous genomes will produce collapsed assemblies with mosaic sequences.

For a researcher who needs a quick draft assembly for a bacterial or fungal genome, or for a first-pass assembly of a larger genome before more detailed analysis, Flye is a practical choice. For a final reference-quality assembly of a complex genome, Flye is less suitable unless combined with extensive polishing and manual curation.

Observations and Measurements for Assembler Comparison

Case Example: Digitalis purpurea Genome Assembly

The Digitalis purpurea genome assembly provides a concrete example of what is achievable with long-read sequencing and appropriate assembly methods. The assembly achieved an N50 of 4.3 Mb and approximately 96% complete BUSCO genes, indicating high contiguity and completeness for a plant genome.

The assembly enabled the identification of structural genes in the anthocyanin biosynthesis pathway and the corresponding transcriptional regulators. Comparison of magenta and white flowering plants revealed a large insertion in the anthocyanidin synthase gene in white flowering plants that likely renders the gene non-functional, explaining the loss of anthocyanin pigmentation. A large insertion in the DpTFL1/CEN gene was also identified as likely responsible for the development of large terminal flowers.

This example demonstrates that a high-quality assembly is not an end in itself but a foundation for biological discovery. The choice of assembler and the quality of the assembly directly affect the ability to identify structural variants, gene content, and regulatory elements.

Case Example: Bighead Catfish Genome Assembly

The bighead catfish genome assembly demonstrates the value of a multi-platform approach. The 880 Mb genome was assembled using HiFi long-read sequencing from Pacific Biosciences and Oxford Nanopore Technologies, scaffolded with Hi-C data, and polished with Illumina short reads. The assembly spans 27 pseudo-chromosomes with a scaffold N50 of 33.48 Mb, 95.5% BUSCO completeness, and a QV of 50.

This assembly provides a foundation for studying aquaculture traits, genetic diversity, and structural variation in the species. The multi-platform approach, combining the accuracy of HiFi reads, the length of Nanopore reads, the scaffolding power of Hi-C, and the polishing accuracy of short reads, produced a chromosome-scale assembly that would not have been achievable with any single platform alone.

For a researcher planning a chromosome-scale assembly, this example illustrates the value of combining multiple data types. The assembler choice is only one part of the workflow, and the integration of scaffolding and polishing data is equally important.

Case Example: Sepiella japonica Genome Survey

The Sepiella japonica genome survey provides an example of the characterization step that should precede assembler selection. The study used quality-filtered short reads for k-mer-based estimation of genome size, heterozygosity, repeat content, and GC content. The estimated genome sizes were 4317 Mb for females and 4222 Mb for males, with heterozygosity rates of 0.85% and 0.77% and repeat content of 76.05% and 75.91%.

The draft assemblies produced from short reads alone had N50 values of approximately 500 bp, demonstrating that short-read assembly is inadequate for complex genomes. The study concluded that the data provide fundamental information for subsequent high-quality whole-genome assembly, which would require long-read sequencing and an appropriate assembler.

This example illustrates the importance of the genome survey step. Without accurate estimates of genome size, heterozygosity, and repeat content, the assembler choice and coverage targets cannot be determined rationally.

Records and Measurements for Assembly Projects

Project Notebook Requirements

Maintain a detailed project notebook for every genome assembly project. The notebook should include the following records:

  • Sample identification and source
  • DNA extraction method and quality metrics
  • Sequencing platform, chemistry, and basecalling model
  • Read length distribution and total sequenced bases
  • K-mer analysis results including genome size, heterozygosity, repeat content, and GC content
  • Assembler name, version, and exact command used
  • Assembly parameters and rationale for parameter choices
  • Assembly quality metrics including N50, BUSCO completeness, and QV
  • Polishing steps and parameters
  • Manual curation steps and decisions

These records are essential for reproducibility and for troubleshooting when assembly problems arise. They also provide the documentation needed for publication and for depositing the assembly in public databases such as NCBI.

Quality Control Checkpoints

Establish quality control checkpoints at each stage of the assembly workflow. The checkpoints should include:

  • Read quality assessment before assembly, including read length distribution and estimated accuracy
  • K-mer analysis results before assembler selection
  • Assembly contiguity metrics immediately after assembly
  • BUSCO completeness assessment after assembly
  • QV estimation after polishing
  • Alignment of reads back to the assembly to check for misassemblies
  • Comparison of assembly size to the k-mer-based genome size estimate

Each checkpoint should have a pass or fail criterion. If the assembly fails a checkpoint, the appropriate response is to adjust the assembler parameters, add more data, or switch to a different assembler.

Common Failure Patterns and Responses

Several failure patterns recur in complex genome assembly projects. Recognizing these patterns early saves time and computational resources.

The first failure pattern is an assembly that is significantly smaller than the estimated genome size. This indicates that the assembler collapsed repetitive regions or failed to assemble large portions of the genome. The response is to check coverage, increase read length if possible, or switch to an assembler with better repeat handling.

The second failure pattern is an assembly with high BUSCO completeness but low N50. This indicates that the assembly contains most of the genes but is fragmented into many small contigs. The response is to add scaffolding data such as Hi-C or optical mapping, or to adjust the assembler parameters to produce longer contigs.

The third failure pattern is an assembly with high N50 but low BUSCO completeness. This indicates that the assembly is contiguous but missing substantial portions of the genome. The response is to check for collapsed haplotypes or unassembled repeats, and to consider whether the assembler is appropriate for the genome's characteristics.

The fourth failure pattern is an assembly with a mosaic haplotype structure, where contigs switch between the two parental haplotypes at heterozygous sites. This is detected by aligning reads back to the assembly and looking for regions with high heterozygosity. The response is to use a haplotype-aware assembler such as hifiasm or Canu with haplotype-aware settings.

Limitations and Interpretation Boundaries

Assembler Performance Depends on Input Quality

The performance of any assembler depends on the quality of the input data. Low-quality reads with high error rates or short read lengths will produce poor assemblies regardless of the assembler chosen. The genome survey step is essential for identifying data quality issues before assembly begins.

The k-mer-based genome size estimation methods have known limitations. The results vary substantially with the tools and parameters used, and the trade-off in k-mer length amplifies the signal of genomic characteristics related to repeat content or heterozygosity. Genome size predictions are influenced by genomic heterozygosity and sequencing accuracy when different k-mer lengths are employed. The closed-loop estimation framework using HiFi reads provides more consistent results because it leverages the continuity and accuracy of HiFi reads to identify convergence regions in the genome size estimates.

Assembly Quality Metrics Have Boundaries

The standard assembly quality metrics have interpretation boundaries that should be understood before drawing conclusions. N50 measures contiguity but does not measure correctness. A contig can be long but contain misassemblies that are not detected by N50 alone. BUSCO completeness measures the presence of expected single-copy orthologs but does not measure the correctness of their sequence or their genomic context. QV estimates base-level accuracy but does not detect structural errors such as misjoins or haplotype switches.

For a complex genome, no single metric is sufficient to assess assembly quality. The combination of contiguity, completeness, and accuracy metrics, along with read alignment and comparison to the genome size estimate, provides a more complete picture. Manual curation may be necessary to resolve regions where automated metrics are ambiguous.

The Vertebrate Genomes Project Lessons

The Vertebrate Genomes Project has documented lessons that apply broadly to complex genome assembly. The project confirmed that long-read sequencing technologies are essential for maximizing genome quality and that unresolved complex repeats and haplotype heterozygosity are major sources of assembly error when not handled correctly. The project's assemblies corrected substantial errors and added missing sequence in some of the best historical reference genomes.

These lessons indicate that the choice of assembler is also a technical detail but a decision that affects the biological conclusions that can be drawn from the assembly. False gene duplications, increases in gene sizes, and chromosome rearrangements identified from assemblies can be artifacts of assembly errors instead of real biological features. The assembler choice and the quality assessment process determine whether these artifacts are identified and corrected.

Safety and Regulatory Context for Genome Assembly

Data Management and Reproducibility

Genome assembly projects generate large amounts of data that must be managed carefully. Raw sequencing data, intermediate files, and final assemblies should be stored with clear naming conventions and version control. The computational workflows should be documented and made reproducible using standard tools.

The nf-core documentation provides standards for community pipelines that emphasize reproducibility, configuration, and usage. These standards are applicable to genome assembly workflows, even when the assembly itself is run with individual tools instead of a pipeline. The Galaxy Training Network provides accessible workflow training and analysis tutorials that can help researchers develop reproducible assembly workflows.

The Carpentries lessons provide foundational computing, data, shell, Git, and programming training that is useful for researchers who need to develop the computational skills required for genome assembly. The EMBL-EBI training resources provide bioinformatics learning pathways and practical analysis education that can supplement hands-on assembly experience.

Data Deposition and Public Databases

Assemblies that are intended for publication or public use should be deposited in public databases such as NCBI. The NCBI provides official descriptions of databases, search systems, sequence resources, and analysis services that support genome assembly deposition and retrieval. Depositing the assembly and the associated raw data ensures that the work is accessible to the research community and that the assembly can be verified and improved by others.

The Bioconductor project provides official package, workflow, installation, and reproducible genomic-analysis documentation that can support downstream analysis of assembled genomes. These resources are particularly useful for researchers who need to perform complex analyses on their assemblies and want to ensure that the analyses are reproducible.

Professional Escalation Criteria

Some assembly problems require professional escalation beyond the standard troubleshooting steps. Escalate to a bioinformatics core facility, a computational biology consultant, or a collaboration with an assembly specialist when:

  • The assembly fails multiple quality checkpoints despite parameter adjustment and additional data
  • The genome has unusual characteristics such as polyploidy, extreme heterozygosity, or very high repeat content that exceed the capabilities of standard assemblers
  • The assembly is intended for clinical, regulatory, or commercial use where accuracy requirements exceed what can be achieved with standard methods
  • The project requires chromosome-scale assembly and the standard scaffolding approaches have failed
  • The computational resources required for the assembly exceed the capacity of the available infrastructure

Professional escalation should occur early instead of late in the project. The cost of a failed assembly attempt is substantial, and the expertise of an assembly specialist can often identify problems that are not apparent from standard quality metrics.

Frequently Asked Questions

What is the difference between HiFi reads and continuous long reads for genome assembly?

HiFi reads are produced by Pacific Biosciences circular consensus sequencing and have accuracy above 99% with read lengths typically between 10 and 25 kb. Continuous long reads from Oxford Nanopore or older PacBio platforms have lower per-base accuracy, often between 85% and 95%, but can reach much longer lengths, sometimes exceeding 100 kb. The accuracy of HiFi reads enables assemblers like hifiasm to skip error correction and perform haplotype-aware assembly directly. Continuous long reads require error correction before or during assembly, which is why Canu and Flye include error correction steps in their pipelines.

When should I use hifiasm instead of Canu or Flye?

Use hifiasm when you have PacBio HiFi data and the genome is diploid with meaningful heterozygosity. hifiasm produces phased assemblies that separate the two haplotypes, which is valuable for heterozygous samples where a collapsed assembly would create a mosaic sequence. hifiasm is also faster than Canu because it does not perform a separate error correction step. If you have only continuous long reads, hifiasm is not the appropriate choice, and Canu or Flye should be used instead.

How do I estimate genome size and heterozygosity before choosing an assembler?

Use k-mer analysis on quality-filtered sequencing reads to estimate genome size, heterozygosity, repeat content, and GC content. For short reads, standard k-mer tools can provide these estimates, though the results vary with the tools and parameters used. For HiFi reads, a closed-loop estimation framework that integrates FastK and GenomeScope 2.0 provides more consistent results by identifying the convergence region where genome size estimates stabilize as k-mer length varies. The genome survey should be completed before assembler selection because the estimates determine which assembler and coverage targets are appropriate.

What coverage do I need for each assembler?

For hifiasm with HiFi data, a coverage of 30x to 50x is typically sufficient for diploid genomes, with higher coverage needed for genomes with high heterozygosity. For Canu or Flye with continuous long reads, a coverage of 50x to 100x is recommended because the error correction step consumes coverage. The coverage calculation should be based on the estimated genome size from the k-mer analysis, not on the raw sequencing output.

How do I assess whether my assembly is good enough for downstream analysis?

Assess the assembly using multiple quality metrics including N50 for contiguity, BUSCO completeness for gene content, and QV for base-level accuracy. Compare the assembly size to the k-mer-based genome size estimate to check for collapsed repeats or missing sequence. Align reads back to the assembly to check for misassemblies and haplotype switches. For a complex genome, no single metric is sufficient, and the combination of metrics along with manual inspection of problematic regions provides the most reliable assessment.

What should I do if my assembly fails quality checks?

If the assembly fails quality checks, first verify that the input data quality is adequate and that the coverage is sufficient for the chosen assembler. Adjust the assembler parameters, particularly those related to repeat resolution and haplotype handling. If parameter adjustment does not resolve the problem, consider adding more sequencing data, switching to a different assembler, or combining multiple data types such as HiFi reads with ultra-long Nanopore reads. If the assembly continues to fail quality checks, escalate to a bioinformatics specialist or core facility.

Do I need to polish my assembly after using hifiasm?

Polishing is often unnecessary for hifiasm assemblies because the HiFi input reads are already highly accurate. However, if the assembly will be used for applications that require very high accuracy, such as clinical or regulatory use, polishing with additional data may be beneficial. For assemblies from Canu or Flye with continuous long reads, polishing with short reads or HiFi reads is usually required to reach acceptable accuracy. The bighead catfish assembly used Illumina short-read polishing after HiFi and Nanopore assembly, demonstrating that polishing can improve accuracy even for high-quality long-read assemblies.

How do I choose between Canu and Flye for Oxford Nanopore data?

Choose Canu when you need a high-quality assembly and have sufficient computational resources and time. Canu performs its own error correction and produces higher-quality corrected reads, but the computational cost is substantial. Choose Flye when you need a quick draft assembly and computational resources are limited. Flye is significantly faster than Canu but produces assemblies with more residual errors that require polishing. For a final reference-quality assembly of a complex genome, Canu is the more reliable choice, while Flye is suitable for first-pass assemblies and for genomes where speed is more important than completeness.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.