Hybrid Assembly with Unicycler: Combining Short and Long Reads for Bacterial Genomes

By Dr. Zubair Khalid, DVM, MS, PhD ·

Hybrid Assembly with Unicycler: Combining Short and Long Reads for Bacterial Genomes

Key Takeaways

  • Unicycler hybrid assembly synergistically combines the high base-level accuracy of Illumina short reads with the structural resolving power of long reads (PacBio or Oxford Nanopore) to overcome limitations of each platform, enabling complete bacterial genome reconstruction.
  • The Unicycler workflow involves an initial assembly graph construction from short reads (using SPAdes), followed by alignment of long reads to this graph to simplify it, resolve repeats, and bridge gaps, leveraging short reads for accuracy and long reads for structural integrity.
  • Input quality control for short reads involves adapter trimming and quality filtering to prevent spurious connections, while long reads, despite higher error rates, are primarily used for structural information, making Unicycler robust to variable long-read quality and depth.
  • Assembly modes (Normal, Bold, Conservative) offer flexibility: Normal is suitable for most bacterial genomes, Bold for high-quality data aiming for maximal contiguity, and Conservative for problematic data requiring more fragmented but error-reduced assemblies.
  • Quality assessment of Unicycler output is critical, focusing on assembly statistics (contig count, N50), genome completeness (e.g., using BUSCO), and circularity of chromosomes and plasmids, which is essential for accurate downstream analyses like AMR surveillance and phylogenetics.

Microbiologists who need complete bacterial genome assemblies face a persistent problem: short-read sequencing alone produces accurate but fragmented assemblies, while long-read sequencing alone produces complete but error-prone assemblies. Hybrid assembly with Unicycler resolves this by combining the base-level accuracy of Illumina short reads with the structural resolving power of Pacific Biosciences or Oxford Nanopore Technologies long reads. This article provides a practical workflow for Unicycler hybrid assembly, covering input preparation, assembly execution, quality assessment, and interpretation of results for bacterial genome projects.

The Hybrid Assembly Problem in Bacterial Genomics

Bacterial genome assembly requires balancing two competing demands. Short-read platforms such as Illumina generate highly accurate reads of 150 to 300 base pairs, but these reads cannot span repetitive elements, insertion sequences, or multicopy genes that are common in bacterial chromosomes and plasmids. The result is a draft assembly composed of many contigs with uncertain ordering and orientation. Long-read platforms such as PacBio and Oxford Nanopore generate reads of thousands to tens of thousands of base pairs, which can span repetitive regions and resolve structural arrangements. However, long-read sequencing has historically been more expensive per base and carries a higher error rate than short-read sequencing.

The original Unicycler publication describes this tradeoff directly. Illumina sequencing produces accurate but short reads that yield accurate but fragmented assemblies. PacBio and Oxford Nanopore produce long reads that can generate complete assemblies, but the sequencing is more expensive and error-prone. The authors note significant interest in combining these complementary technologies to generate more accurate hybrid assemblies, and they identify a gap in available tools that truly leverage both data types. Unicycler was designed to fill that gap by using the accuracy of short reads and the structural resolving power of long reads in a single assembly pipeline.

The practical consequence for a microbiology laboratory is that hybrid assembly enables complete genome reconstruction for epidemiological investigations, antimicrobial resistance surveillance, and phylogenetic analyses. A complete circular chromosome with associated plasmids provides substantially more biological information than a fragmented draft assembly. For example, a hybrid assembly of a colistin-resistant Escherichia coli strain from Brazil produced a genome of 5,333,039 base pairs and revealed that the mcr-1.5 resistance gene was carried on an IncI2 plasmid of approximately 65,458 base pairs. This level of structural resolution is difficult to achieve with short-read data alone.

How Unicycler Works

Unicycler operates in three stages that build on the strengths of each sequencing platform. The first stage constructs an initial assembly graph from short reads using the SPAdes de novo assembler. This graph represents the relationships between contigs, including connections that may be ambiguous due to repeats. The second stage aligns long reads to this assembly graph using a novel semi-global aligner. The third stage uses the long-read alignments to simplify the graph, resolve repeats, and bridge gaps between contigs.

The key design principle is that short reads provide accurate sequence information while long reads provide structural information. Unicycler does not simply concatenate the two data types. Instead, it uses the short-read assembly as a scaffold and the long reads as a guide for resolving ambiguous regions. This approach allows Unicycler to produce assemblies with larger contigs and fewer misassemblies than other hybrid assemblers, even when long-read depth and accuracy are low.

The original publication reports that Unicycler was tested on both synthetic and real reads. The results showed that Unicycler could assemble larger contigs with fewer misassemblies than other hybrid assemblers under conditions of low long-read depth and accuracy. This robustness is practically important because long-read sequencing runs can vary in yield and quality, and a laboratory may not always achieve ideal coverage.

At a Glance

The following table summarizes the key decisions in a Unicycler hybrid assembly workflow.

Workflow ComponentPrimary OptionsPractical Consideration
Short-read inputIllumina FASTQ files, paired-end or single-endHigher coverage improves base accuracy, adapter trimming is recommended before assembly
Long-read inputOxford Nanopore FASTQ or PacBio FASTQLow-depth long reads can still resolve structure, quality filtering improves assembly
Assembly modeNormal, bold, or conservativeNormal mode suits most bacterial genomes, bold mode for high-quality data, conservative mode for problematic data
Output assessmentQUAST, BUSCO, assembly graph inspectionCheck contig count, N50, genome completeness, and circularity
Downstream analysisProkka, MLST, AMR prediction, phylogenetic toolsAssembly quality directly affects gene detection and variant calling

Input Preparation for Unicycler

Short-Read Quality Control

Illumina short reads are the foundation of the Unicycler assembly graph. The quality of these reads directly affects base-level accuracy in the final assembly. Before running Unicycler, laboratories should assess read quality using standard metrics such as per-base quality scores, GC content distribution, and adapter contamination. Low-quality bases at read ends and adapter sequences should be removed because they can create spurious connections in the assembly graph.

The NCBI provides official documentation for sequence data resources and analysis services that can help laboratories understand read formats and quality metrics. The European Bioinformatics Institute offers training materials for bioinformatics data analysis that cover quality assessment workflows. These resources are useful for laboratories establishing standard operating procedures for sequencing data handling.

Long-Read Quality Control

Long-read data from Oxford Nanopore and PacBio platforms require different quality considerations than short reads. Nanopore reads have historically had higher error rates than Illumina reads, and the error profile includes insertions and deletions in addition to substitutions. PacBio reads have a different error profile that is largely random. Unicycler is designed to tolerate these errors because it uses long reads primarily for structural information instead of base-level accuracy.

A benchmarking study of hybrid assembly approaches for bacterial pathogens tested Unicycler with simulated reads of mediocre and low quality, as well as real reads from multiple bacterial species. The study found that Unicycler performed best for achieving contiguous genomes among the tools tested. Importantly, the study also found that MaSuRCA was less tolerant of low-quality long reads than SPAdes and Unicycler. This finding suggests that Unicycler is a reasonable choice when long-read quality is uncertain or variable.

Coverage Considerations

Coverage depth for both short and long reads affects assembly quality. The original Unicycler publication notes that the tool can produce good assemblies even when long-read depth and accuracy are low. This is a practical advantage because long-read sequencing can be more expensive than short-read sequencing, and laboratories may seek to minimize long-read coverage to control costs.

For short reads, higher coverage generally improves the accuracy of the initial assembly graph. Most bacterial genome projects aim for at least 50-fold coverage with Illumina data, though the optimal depth depends on the genome size and complexity. For long reads, even modest coverage can resolve structural features if the reads span repetitive regions. The benchmarking study of hybrid assembly approaches found that Unicycler assemblies had high genome completeness, approximately 98.7 percent, when compared to other assembly tools in a study of Mycobacterium tuberculosis genomes.

Running Unicycler

Installation and Dependencies

Unicycler is open source under the GPLv3 license and is available from the GitHub repository maintained by the developer. The tool requires several dependencies, including SPAdes for the initial short-read assembly, and it can use additional tools for read alignment and assembly polishing. Laboratories should install Unicycler in a controlled computing environment and verify that all dependencies are available before starting an assembly run.

The Bioconductor project provides official documentation for reproducible genomic analysis workflows, and the Galaxy Training Network offers accessible tutorials for assembly and analysis. These resources can help laboratories implement Unicycler in a reproducible manner. The nf-core documentation describes community standards for pipeline usage and configuration, which is relevant for laboratories that want to integrate Unicycler into larger automated workflows.

Command Structure

A typical Unicycler command specifies the short-read FASTQ files, the long-read FASTQ file, and an output directory. The tool automatically detects whether reads are from Oxford Nanopore or PacBio platforms based on the data characteristics. The command also accepts options for adjusting the assembly mode and for specifying the number of threads to use.

The assembly mode is an important decision. Normal mode is appropriate for most bacterial genomes and balances speed with accuracy. Bold mode assumes high-quality data and may produce more complete assemblies but can be less robust to data problems. Conservative mode is designed for problematic data and may produce more fragmented assemblies but with fewer errors. Laboratories should start with normal mode and adjust based on the characteristics of their data.

Runtime and Resource Requirements

Unicycler runtime depends on genome size, read depth, and available computing resources. Bacterial genomes of approximately 4 to 6 million base pairs typically assemble in a reasonable time on a standard server with multiple CPU cores. The SPAdes step is often the most computationally intensive part of the pipeline. Laboratories should monitor memory usage during assembly runs and ensure that sufficient resources are available.

A comparison of long-read sequencing technologies in hybrid assembly of complex bacterial genomes used Unicycler to assemble 20 bacterial isolates from the Enterobacteriaceae family. These genomes frequently have highly plastic, repetitive genetic structures, and complete reconstruction is relevant for understanding antimicrobial resistance epidemiology. The study found that hybrid assembly with either PacBio or ONT reads facilitated high-quality genome reconstruction and was superior to long-read-only assembly followed by short-read polishing.

Assembly Modes and Their Tradeoffs

Normal Mode

Normal mode is the default and recommended starting point for most bacterial genome assemblies. It uses a balanced approach to graph simplification that works well for typical bacterial genomes with moderate repeat content. The original Unicycler publication describes the algorithm as building an initial assembly graph from short reads and then simplifying the graph using information from both short and long reads.

Bold Mode

Bold mode is designed for high-quality data where the assembler can be more aggressive in resolving ambiguous regions. This mode may produce more complete assemblies with fewer contigs, but it carries a higher risk of misassembly if the data contain errors or if the genome has unusual structural features. Laboratories should use bold mode only when they have confidence in the quality of both short and long reads.

Conservative Mode

Conservative mode is designed for problematic data, such as low-coverage long reads or genomes with complex repeat structures. This mode produces more fragmented assemblies but with fewer misassemblies. The tradeoff is that the assembly may require additional manual finishing steps to achieve complete genome reconstruction.

The benchmarking study of hybrid assembly approaches for bacterial pathogens found that all SPAdes assemblies were incomplete when compared to Unicycler and MaSuRCA. This finding underscores the importance of choosing an assembler that can effectively use both short and long read data. Unicycler's performance was strongest for achieving contiguous genomes, closely followed by MaSuRCA.

Quality Assessment of Hybrid Assemblies

Assembly Statistics

After Unicycler completes an assembly run, laboratories should assess the output using standard assembly statistics. The number of contigs, the N50 value, and the total assembly size provide a first indication of assembly quality. A complete bacterial genome assembly should ideally consist of one contig per replicon, with the chromosome and each plasmid represented as a single circular sequence.

The genome assembly size should be compared to the expected genome size for the species. For example, the Mycobacterium tuberculosis reference genome H37Rv is approximately 4,411,532 base pairs. A study comparing assembly tools for M. tuberculosis genomes found that Unicycler assemblies had a mean size of 4,377,642 base pairs, which is close to the expected genome size. RagOut assemblies were significantly longer at 4,418,574 base pairs, which may indicate over-assembly or inclusion of spurious sequence.

Genome Completeness

Genome completeness can be assessed using tools that search for conserved single-copy genes. The benchmarking study of hybrid assembly approaches for bacterial pathogens determined genome completeness and accuracy for assemblies of ten bacterial species. Unicycler assemblies had high completeness, and the study found that hybrid assemblies with ONT and PacBio long reads detected more genes than short-read assembly alone.

The comparison of long-read sequencing technologies in hybrid assembly of complex bacterial genomes found that hybrid assembly with either PacBio or ONT reads facilitated high-quality genome reconstruction. This study also found that combining ONT and Illumina reads fully resolved most genomes without additional manual steps, and at a lower consumables cost per isolate in the study setting.

Circularity Assessment

A key advantage of hybrid assembly is the ability to produce circular chromosomes and plasmids. Unicycler can produce circular contigs when the data support circularization. Laboratories should check the assembly output for circular contigs and verify that the circularization is supported by read evidence. A circular chromosome that is actually linear or that has an incorrect join can introduce errors in downstream analysis.

The hybrid assembly of the colistin-resistant E. coli strain from Brazil produced a complete genome that enabled detailed analysis of the resistance plasmid. The mcr-1.5 gene was carried by an IncI2 plasmid, and the full genome SNP-based phylogenetic analysis revealed that the strain was highly related to colistin-resistant ST354 lineages associated with urinary tract infections in Brazil since 2015. This level of analysis depends on a complete and accurate assembly.

Records and Measurements for Assembly Projects

Documentation Standards

Laboratories should maintain detailed records for each assembly project. The records should include the sequencing platform and version, read quality metrics, coverage depth for short and long reads, Unicycler version and parameters, assembly statistics, and the date and personnel responsible for the assembly. This documentation supports reproducibility and troubleshooting.

The Carpentries offers lessons on foundational computing, data, shell, Git, and programming that are useful for laboratories implementing reproducible bioinformatics workflows. The Galaxy Training Network provides accessible workflow training and analysis tutorials that emphasize reproducibility. The nf-core documentation describes community pipeline standards for usage and configuration that support reproducible analysis.

Quality Metrics to Record

The following metrics should be recorded for each assembly project:

MetricPurposeInterpretation
Read count and total basesAssess coverage depthHigher coverage generally improves assembly quality
Read N50Assess read length distributionLonger reads improve structural resolution
Assembly contig countAssess fragmentationFewer contigs indicate more complete assembly
Assembly N50Assess contig length distributionHigher N50 indicates better assembly
Total assembly sizeCompare to expected genome sizeLarge deviations may indicate contamination or misassembly
Genome completenessAssess gene contentValues near 100 percent indicate complete gene representation
Circular contigsAssess structural resolutionCircular chromosomes and plasmids indicate complete assembly

Common Failure Patterns in Unicycler Assemblies

Low Long-Read Coverage

When long-read coverage is too low, Unicycler may fail to resolve repetitive regions, resulting in fragmented assemblies. The original publication notes that Unicycler can assemble larger contigs with fewer misassemblies than other hybrid assemblers even when long-read depth is low, but there is a practical minimum below which structural resolution fails. Laboratories should assess long-read coverage before assembly and consider additional sequencing if coverage is inadequate.

Contaminated Input Data

Contamination from other organisms can produce assembly graphs with spurious connections and inflated genome sizes. The assembly size should be compared to the expected genome size for the target species. A benchmarking study of hybrid assembly approaches for bacterial pathogens found that the MaSuRCA assembly of Staphylococcus aureus with real reads contained antimicrobial resistance genes that were not present in the reference genome or in the Unicycler assembly. This discrepancy may indicate contamination or misassembly in the MaSuRCA output.

Poor Short-Read Quality

Low-quality short reads can introduce errors in the initial assembly graph that propagate through the hybrid assembly. Adapter contamination and low-quality base calls should be removed before assembly. The NCBI provides official descriptions of sequence data resources that can help laboratories understand read quality metrics and data formats.

Complex Repeat Structures

Some bacterial genomes have complex repeat structures that are difficult to resolve even with long reads. The comparison of long-read sequencing technologies in hybrid assembly of complex bacterial genomes selected isolates from the Enterobacteriaceae family because these frequently have highly plastic, repetitive genetic structures. The study found that hybrid assembly with either PacBio or ONT reads facilitated high-quality genome reconstruction, but some genomes may require additional manual finishing steps.

Interpretation Limits and Reporting

What Hybrid Assembly Cannot Resolve

Hybrid assembly with Unicycler resolves many structural features, but it has limitations. Very long repeats that exceed the length of the longest long reads may remain ambiguous. Highly repetitive regions such as ribosomal RNA operons may be collapsed or misassembled. Laboratories should be cautious when interpreting assemblies that contain unresolved regions.

The original Unicycler publication describes the tool as producing assemblies that are accurate, complete, and cost-effective. However, the authors also note that few tools exist that truly leverage the benefits of both types of data. This statement acknowledges that hybrid assembly is an active area of development and that current tools have limitations.

Reporting Standards

When reporting hybrid assembly results, laboratories should describe the sequencing platforms, coverage depths, assembly parameters, and quality metrics. This information allows other researchers to assess the reliability of the assembly and to compare results across studies. The European Bioinformatics Institute offers training materials for bioinformatics data analysis that cover reporting standards and data sharing practices.

The benchmarking study of hybrid assembly approaches for bacterial pathogens reported genome completeness and accuracy, antimicrobial resistance, virulence potential, multilocus sequence typing, phylogeny, and pan genome for assemblies of multiple bacterial species. This comprehensive reporting approach provides a model for laboratories that want to characterize their assemblies thoroughly.

Safety and Regulatory Context

Data Management

Bacterial genome sequence data may be subject to institutional, national, or international data management requirements. Laboratories should be aware of applicable regulations for data storage, sharing, and publication. The NCBI provides official descriptions of sequence databases and data submission procedures that laboratories can use to comply with data sharing requirements.

Antimicrobial Resistance Data

Assemblies of antimicrobial-resistant bacteria may have implications for public health surveillance and clinical decision-making. The hybrid assembly of the colistin-resistant E. coli strain from Brazil revealed a broad resistome that included antibiotics, heavy metals, disinfectants, and glyphosate. Laboratories that assemble genomes of resistant pathogens should consider the public health context and follow applicable reporting requirements.

Professional Escalation Criteria

Laboratories should escalate assembly problems to a bioinformatics specialist or supervisor when they encounter any of the following situations:

SituationAction
Assembly fails to completeReview error logs and check input data quality
Assembly size deviates substantially from expectedCheck for contamination and reassess input data
Genome completeness is below 95 percentConsider additional sequencing or alternative assembly parameters
Circular contigs are not produced for expected repliconsInvestigate repeat structures and long-read coverage
Antimicrobial resistance genes are detected in unexpected contextsVerify assembly accuracy before reporting

A Practical Decision Framework for Choosing Between Hybrid Assembly Strategies

Microbiologists often assume that running Unicycler with any combination of short and long reads will produce a complete genome. In practice, the choice of long-read platform, the depth of sequencing, and the expected genome architecture all influence whether Unicycler is the right tool for a given project. This section provides a structured decision framework that laboratories can use before committing computational resources and sequencing budgets to a hybrid assembly project.

Platform Selection Based on Genome Architecture

The first decision in any hybrid assembly project is which long-read platform to use. Oxford Nanopore Technologies and Pacific Biosciences both generate reads that can resolve repetitive structures, but they differ in throughput, cost, error profile, and practical workflow requirements. A comparison of long-read sequencing technologies in hybrid assembly of complex bacterial genomes examined 20 bacterial isolates from the Enterobacteriaceae family, a group chosen because these genomes frequently have highly plastic, repetitive genetic structures. The study found that hybrid assembly with either PacBio or ONT reads facilitated high-quality genome reconstruction and was superior to long-read-only assembly followed by short-read polishing with respect to accuracy and completeness.

The practical implication is that platform choice should be driven by laboratory infrastructure and project timelines instead of by a belief that one platform is inherently superior for Unicycler. ONT sequencing offers the advantage of real-time data generation and lower upfront instrument costs, which makes it attractive for laboratories that do not have access to a PacBio instrument. The comparison study noted that combining ONT and Illumina reads fully resolved most genomes without additional manual steps and at a lower consumables cost per isolate in the study setting. PacBio sequencing, by contrast, offers higher per-read accuracy and a more uniform error profile, which can simplify downstream analysis even though the consumables cost may be higher.

For laboratories that are establishing a hybrid assembly workflow for the first time, the decision framework should include an assessment of existing sequencing infrastructure. If the laboratory already has access to an Illumina instrument and a Nanopore device, the marginal cost of adding ONT long reads to an existing short-read project is lower than the cost of outsourcing PacBio sequencing. If the laboratory is planning to sequence many isolates over time, the per-sample consumables cost difference between ONT and PacBio becomes a significant factor in the decision.

Coverage Depth Decisions for Long Reads

The original Unicycler publication emphasizes that the tool can assemble larger contigs with fewer misassemblies than other hybrid assemblers even when long-read depth and accuracy are low. This robustness is a practical advantage, but it does not mean that coverage depth is irrelevant. The decision framework should include a minimum coverage threshold based on the expected repeat content of the target genome.

For genomes with modest repeat content, such as many Escherichia coli isolates, low long-read coverage may be sufficient to resolve the chromosome and common plasmids. The hybrid assembly of a colistin-resistant E. coli strain from Brazil used Illumina and Nanopore sequence data and produced a genome of 5,333,039 base pairs with the mcr-1.5 gene carried on an IncI2 plasmid. This result demonstrates that Unicycler can produce complete assemblies with practical levels of long-read coverage.

For genomes with complex repeat structures, such as those found in the Enterobacteriaceae family, higher long-read coverage provides more evidence for resolving ambiguous graph connections. The comparison study of long-read sequencing technologies found that hybrid assembly with either PacBio or ONT reads facilitated high-quality genome reconstruction for these challenging genomes. Laboratories should consider the repeat content of the target species when deciding how much long-read coverage to generate.

A practical approach is to start with a modest long-read coverage target and assess the assembly output. If the assembly is fragmented or contains unresolved regions, additional long-read sequencing can be generated and the assembly rerun. This iterative approach avoids the cost of generating excessive long-read coverage upfront while still allowing for complete assembly when needed.

When to Use Unicycler Versus Alternative Assembly Strategies

The decision framework should also address when Unicycler is the appropriate tool and when alternative strategies may be more suitable. A benchmarking study of hybrid assembly approaches for bacterial pathogens compared Unicycler, MaSuRCA, and SPAdes using simulated reads of mediocre and low quality, as well as real reads from multiple bacterial species. The study found that Unicycler performed the best for achieving contiguous genomes, closely followed by MaSuRCA, while all SPAdes assemblies were incomplete. The study also found that MaSuRCA was less tolerant of low-quality long reads than SPAdes and Unicycler.

These findings support the use of Unicycler as the default hybrid assembler for bacterial genomes, particularly when long-read quality is uncertain. However, the decision framework should include criteria for when to consider alternative approaches. If the laboratory has very high-quality long reads and a simple genome architecture, a long-read-only assembly followed by short-read polishing may be sufficient. The comparison study of long-read sequencing technologies found that this approach was inferior to hybrid assembly with respect to accuracy and completeness, but it may be acceptable for projects that do not require complete genome reconstruction.

If the laboratory has access to a well-validated MaSuRCA pipeline and the long-read data are of high quality, MaSuRCA may produce comparable results. The benchmarking study found that MaSuRCA was less tolerant of low-quality long reads, so laboratories with variable long-read quality should prefer Unicycler. The decision framework should include a quality assessment step for long reads before choosing the assembler.

Decision Matrix for Assembly Strategy Selection

The following table provides a structured decision matrix that laboratories can use to select an assembly strategy based on their specific project characteristics.

Project CharacteristicRecommended StrategyRationale
Complete genome needed, long-read quality variableUnicycler hybrid assemblyRobust to low-quality long reads, produces contiguous assemblies
Complete genome needed, high-quality long reads availableUnicycler or MaSuRCA hybrid assemblyBoth tools perform well with high-quality data
Draft assembly acceptable, short reads onlySPAdes or similar short-read assemblerLower cost, sufficient for many analyses
Complete genome needed, no long-read platform availableConsider outsourcing long-read sequencingHybrid assembly requires long reads for structural resolution
Complex repeat structure expectedUnicycler with higher long-read coverageAdditional coverage improves resolution of ambiguous regions
Many isolates to assembleUnicycler with ONT long readsLower consumables cost per isolate in many settings

Cost-Benefit Analysis for Hybrid Assembly

The decision framework should include a cost-benefit analysis that considers the value of complete genome reconstruction relative to the cost of long-read sequencing. The comparison study of long-read sequencing technologies noted that combining ONT and Illumina reads fully resolved most genomes at a lower consumables cost per isolate in the study setting. This finding is relevant for laboratories that are planning large-scale sequencing projects.

For epidemiological investigations, complete genome reconstruction provides information that is not available from draft assemblies. The hybrid assembly of the colistin-resistant E. coli strain from Brazil enabled full genome SNP-based phylogenetic analysis, which revealed that the strain was highly related to colistin-resistant ST354 lineages associated with urinary tract infections in Brazil since 2015. This level of analysis requires a complete assembly that includes plasmids and other mobile genetic elements.

For antimicrobial resistance surveillance, complete assemblies allow precise localization of resistance genes to chromosomes or plasmids. The benchmarking study of hybrid assembly approaches for bacterial pathogens found that hybrid assemblies of five antimicrobial-resistant strains with simulated reads provided consistent antimicrobial resistance genotypes with the reference genomes. The study also found that the MaSuRCA assembly of Staphylococcus aureus with real reads contained antimicrobial resistance genes that were not present in the reference genome or in the Unicycler assembly, highlighting the importance of assembly accuracy for resistance gene reporting.

Laboratories should weigh the cost of long-read sequencing against the value of complete genome reconstruction for their specific research questions. For projects that require precise localization of mobile genetic elements, complete assemblies justify the additional cost. For projects that only require gene presence or absence calls, draft assemblies may be sufficient.

Implementation Steps for the Decision Framework

The following steps provide a structured approach to implementing the decision framework in a laboratory setting.

Step 1: Characterize the target genome. Determine the expected genome size, repeat content, and plasmid content for the target species. This information can be obtained from published reference genomes or from databases such as those maintained by the NCBI. The NCBI provides official descriptions of sequence databases and analysis services that can help laboratories identify appropriate reference genomes.

Step 2: Assess available sequencing infrastructure. Determine which sequencing platforms are available in the laboratory or through collaborators. Consider the cost, turnaround time, and throughput of each platform. The European Bioinformatics Institute offers training materials for bioinformatics data analysis that can help laboratories understand the strengths and limitations of different sequencing platforms.

Step 3: Estimate coverage requirements. Based on the genome architecture and the expected repeat content, estimate the minimum long-read coverage needed for complete assembly. For genomes with modest repeat content, lower coverage may be sufficient. For genomes with complex repeat structures, higher coverage is recommended.

Step 4: Select the assembly strategy. Use the decision matrix to select the appropriate assembly strategy. For most bacterial genome projects, Unicycler hybrid assembly is the recommended default. Consider alternative strategies only when specific project characteristics support their use.

Step 5: Document the decision. Record the rationale for the assembly strategy selection, including the genome characteristics, sequencing platform, coverage estimates, and expected outcomes. This documentation supports reproducibility and provides a basis for troubleshooting if the assembly does not meet expectations.

Step 6: Execute the assembly and assess the output. Run Unicycler with the selected parameters and assess the assembly using the quality metrics described in the records and measurements section. If the assembly does not meet expectations, revisit the decision framework and consider whether additional sequencing or alternative parameters are needed.

Common Failure Patterns in Strategy Selection

The decision framework helps laboratories avoid common failure patterns that arise from poor strategy selection. One common failure is generating insufficient long-read coverage for a genome with complex repeat structures. The comparison study of long-read sequencing technologies selected isolates from the Enterobacteriaceae family because these frequently have highly plastic, repetitive genetic structures. Laboratories that underestimate the repeat content of their target genome may generate insufficient long-read coverage and produce fragmented assemblies.

Another common failure is choosing a long-read platform based on cost alone without considering the error profile and its interaction with the assembler. The benchmarking study of hybrid assembly approaches for bacterial pathogens found that MaSuRCA was less tolerant of low-quality long reads than SPAdes and Unicycler. Laboratories that use MaSuRCA with low-quality long reads may produce assemblies with more errors than they would obtain with Unicycler.

A third common failure is assuming that all hybrid assemblers produce equivalent results. The benchmarking study found that all SPAdes assemblies were incomplete when compared to Unicycler and MaSuRCA. Laboratories that use SPAdes for hybrid assembly may produce fragmented assemblies that require additional finishing steps.

Professional Escalation Criteria for Strategy Decisions

Laboratories should escalate strategy decisions to a bioinformatics specialist or supervisor when they encounter any of the following situations:

SituationAction
Uncertainty about genome repeat contentConsult published reference genomes or seek specialist advice
Limited access to long-read sequencing platformsDiscuss outsourcing options with supervisor or collaborators
Repeated assembly failures with the selected strategyReassess the decision framework and consider alternative approaches
Budget constraints that limit long-read coverageDiscuss the tradeoff between coverage and assembly completeness
Unusual genome architecture or suspected contaminationSeek specialist advice before proceeding with assembly

Integration with Reproducible Workflow Standards

The decision framework should be integrated with reproducible workflow standards to ensure that assembly strategy decisions are documented and traceable. The nf-core documentation describes community pipeline standards for usage and configuration that support reproducible analysis. The Galaxy Training Network offers accessible workflow training and analysis tutorials that emphasize reproducibility. The Carpentries offers lessons on foundational computing, data, shell, Git, and programming that are useful for laboratories implementing reproducible bioinformatics workflows.

Laboratories should record the decision framework inputs and outputs for each assembly project. This documentation should include the genome characteristics, sequencing platform selection, coverage estimates, assembly strategy selection, and the rationale for each decision. This information supports troubleshooting and allows other researchers to understand the basis for assembly strategy decisions.

The Bioconductor project provides official documentation for reproducible genomic analysis workflows that can help laboratories implement the decision framework in a reproducible manner. The European Bioinformatics Institute offers training materials for bioinformatics data analysis that cover practical analysis education and data-resource training. These resources can help laboratories establish standard operating procedures for assembly strategy selection.

Validation of Strategy Decisions

The decision framework should include a validation step that confirms the selected strategy produces the expected results. The validation step should compare the assembly output to the expected genome characteristics for the target species. The genome assembly size should be compared to the expected genome size. For example, the Mycobacterium tuberculosis reference genome H37Rv is approximately 4,411,532 base pairs. A study comparing assembly tools for M. tuberculosis genomes found that Unicycler assemblies had a mean size of 4,377,642 base pairs, which is close to the expected genome size.

The validation step should also assess genome completeness using tools that search for conserved single-copy genes. The benchmarking study of hybrid assembly approaches for bacterial pathogens found that Unicycler assemblies had high genome completeness, approximately 98.7 percent in a study of M. tuberculosis genomes. The study also found that hybrid assemblies with ONT and PacBio long reads detected more genes than short-read assembly alone.

If the assembly output does not match the expected genome characteristics, the laboratory should revisit the decision framework and consider whether the strategy selection was appropriate. This iterative approach ensures that the decision framework produces reliable results across a range of bacterial genome projects.

Frequently Asked Questions

What are the minimum input requirements for Unicycler?

Unicycler requires short reads, typically from Illumina sequencing, and long reads from either Oxford Nanopore or PacBio platforms. The short reads should be in FASTQ format with quality scores. The long reads should also be in FASTQ format. Unicycler automatically detects the long-read platform based on data characteristics. The original publication describes Unicycler as building an initial assembly graph from short reads using SPAdes and then simplifying the graph using information from short and long reads.

How much long-read coverage is needed for a good hybrid assembly?

The original Unicycler publication reports that the tool can assemble larger contigs with fewer misassemblies than other hybrid assemblers even when long-read depth and accuracy are low. However, the practical minimum coverage depends on the genome complexity and the length of repetitive regions. Laboratories should assess long-read coverage before assembly and consider additional sequencing if the assembly is fragmented.

Can Unicycler use PacBio and Oxford Nanopore reads interchangeably?

Yes, Unicycler can use long reads from either PacBio or Oxford Nanopore platforms. A comparison of long-read sequencing technologies in hybrid assembly of complex bacterial genomes found that hybrid assembly with either PacBio or ONT reads facilitated high-quality genome reconstruction. The study also found that combining ONT and Illumina reads fully resolved most genomes without additional manual steps.

How does Unicycler compare to other hybrid assemblers?

A benchmarking study of hybrid assembly approaches for bacterial pathogens found that Unicycler performed the best for achieving contiguous genomes, closely followed by MaSuRCA, while all SPAdes assemblies were incomplete. The study also found that MaSuRCA was less tolerant of low-quality long reads than SPAdes and Unicycler. The original Unicycler publication reports that Unicycler can assemble larger contigs with fewer misassemblies than other hybrid assemblers.

What should I do if my Unicycler assembly is fragmented?

If the assembly is fragmented, first check the quality of the input reads. Low-quality short reads or insufficient long-read coverage can cause fragmentation. Consider running Unicycler in conservative mode, which is designed for problematic data. If the assembly remains fragmented, additional long-read sequencing may be needed to resolve repetitive regions.

How do I know if my assembly is complete?

A complete bacterial genome assembly should have one contig per replicon, with the chromosome and each plasmid represented as a single circular sequence. Genome completeness can be assessed using tools that search for conserved single-copy genes. The benchmarking study of hybrid assembly approaches for bacterial pathogens found that Unicycler assemblies had high genome completeness, approximately 98.7 percent in a study of Mycobacterium tuberculosis genomes.

Can Unicycler assemble plasmids?

Yes, Unicycler can assemble plasmids when the read data support their reconstruction. The hybrid assembly of the colistin-resistant E. coli strain from Brazil produced a complete genome that included the mcr-1.5 gene carried by an IncI2 plasmid of approximately 65,458 base pairs. Plasmid assembly requires that the plasmids are present in the sequenced sample and that the read data provide sufficient coverage.

What downstream analyses can I perform with a Unicycler hybrid assembly?

A complete hybrid assembly supports a wide range of downstream analyses, including antimicrobial resistance prediction, virulence gene detection, multilocus sequence typing, phylogenomic analysis, and pan-genome analysis. The benchmarking study of hybrid assembly approaches for bacterial pathogens used hybrid assemblies to determine antimicrobial resistance, virulence potential, multilocus sequence typing, phylogeny, and pan genome for multiple bacterial species.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.