# Hybrid Assembly with Unicycler: Combining Short and Long Reads for Bacterial Genomes


## Key Takeaways

- Unicycler hybrid assembly synergistically combines the high base-level accuracy of Illumina short reads with the structural resolving power of long reads (PacBio or Oxford Nanopore) to overcome limitations of each platform, enabling complete bacterial genome reconstruction.
- The Unicycler workflow involves an initial assembly graph construction from short reads (using SPAdes), followed by alignment of long reads to this graph to simplify it, resolve repeats, and bridge gaps, leveraging short reads for accuracy and long reads for structural integrity.
- Input quality control for short reads involves adapter trimming and quality filtering to prevent spurious connections, while long reads, despite higher error rates, are primarily used for structural information, making Unicycler robust to variable long-read quality and depth.
- Assembly modes (Normal, Bold, Conservative) offer flexibility: Normal is suitable for most bacterial genomes, Bold for high-quality data aiming for maximal contiguity, and Conservative for problematic data requiring more fragmented but error-reduced assemblies.
- Quality assessment of Unicycler output is critical, focusing on assembly statistics (contig count, N50), genome completeness (e.g., using BUSCO), and circularity of chromosomes and plasmids, which is essential for accurate downstream analyses like AMR surveillance and phylogenetics.

---

Microbiologists who need complete bacterial genome assemblies face a persistent problem: short-read sequencing alone produces accurate but fragmented assemblies, while long-read sequencing alone produces complete but error-prone assemblies. Hybrid assembly with Unicycler resolves this by combining the base-level accuracy of Illumina short reads with the structural resolving power of Pacific Biosciences or Oxford Nanopore Technologies long reads. This article provides a practical workflow for Unicycler hybrid assembly, covering input preparation, assembly execution, quality assessment, and interpretation of results for bacterial genome projects.

## The Hybrid Assembly Problem in Bacterial Genomics

Bacterial genome assembly requires balancing two competing demands. Short-read platforms such as Illumina generate highly accurate reads of 150 to 300 base pairs, but these reads cannot span repetitive elements, insertion sequences, or multicopy genes that are common in bacterial chromosomes and plasmids. The result is a draft assembly composed of many contigs with uncertain ordering and orientation. Long-read platforms such as PacBio and Oxford Nanopore generate reads of thousands to tens of thousands of base pairs, which can span repetitive regions and resolve structural arrangements. However, long-read sequencing has historically been more expensive per base and carries a higher error rate than short-read sequencing.

The original Unicycler publication describes this tradeoff directly. Illumina sequencing produces accurate but short reads that yield accurate but fragmented assemblies. PacBio and Oxford Nanopore produce long reads that can generate complete assemblies, but the sequencing is more expensive and error-prone. The authors note significant interest in combining these complementary technologies to generate more accurate hybrid assemblies, and they identify a gap in available tools that truly leverage both data types. Unicycler was designed to fill that gap by using the accuracy of short reads and the structural resolving power of long reads in a single assembly pipeline.

The practical consequence for a microbiology laboratory is that hybrid assembly enables complete genome reconstruction for epidemiological investigations, antimicrobial resistance surveillance, and phylogenetic analyses. A complete circular chromosome with associated plasmids provides substantially more biological information than a fragmented draft assembly. For example, a hybrid assembly of a colistin-resistant Escherichia coli strain from Brazil produced a genome of 5,333,039 base pairs and revealed that the mcr-1.5 resistance gene was carried on an IncI2 plasmid of approximately 65,458 base pairs. This level of structural resolution is difficult to achieve with short-read data alone.

## How Unicycler Works

Unicycler operates in three stages that build on the strengths of each sequencing platform. The first stage constructs an initial assembly graph from short reads using the SPAdes de novo assembler. This graph represents the relationships between contigs, including connections that may be ambiguous due to repeats. The second stage aligns long reads to this assembly graph using a novel semi-global aligner. The third stage uses the long-read alignments to simplify the graph, resolve repeats, and bridge gaps between contigs.

The key design principle is that short reads provide accurate sequence information while long reads provide structural information. Unicycler does not simply concatenate the two data types. Instead, it uses the short-read assembly as a scaffold and the long reads as a guide for resolving ambiguous regions. This approach allows Unicycler to produce assemblies with larger contigs and fewer misassemblies than other hybrid assemblers, even when long-read depth and accuracy are low.

The original publication reports that Unicycler was tested on both synthetic and real reads. The results showed that Unicycler could assemble larger contigs with fewer misassemblies than other hybrid assemblers under conditions of low long-read depth and accuracy. This robustness is practically important because long-read sequencing runs can vary in yield and quality, and a laboratory may not always achieve ideal coverage.

## At a Glance

The following table summarizes the key decisions in a Unicycler hybrid assembly workflow.

| Workflow Component | Primary Options | Practical Consideration |
| --- | --- | --- |
| Short-read input | Illumina FASTQ files, paired-end or single-end | Higher coverage improves base accuracy, adapter trimming is recommended before assembly |
| Long-read input | Oxford Nanopore FASTQ or PacBio FASTQ | Low-depth long reads can still resolve structure, quality filtering improves assembly |
| Assembly mode | Normal, bold, or conservative | Normal mode suits most bacterial genomes, bold mode for high-quality data, conservative mode for problematic data |
| Output assessment | QUAST, BUSCO, assembly graph inspection | Check contig count, N50, genome completeness, and circularity |
| Downstream analysis | Prokka, MLST, AMR prediction, phylogenetic tools | Assembly quality directly affects gene detection and variant calling |

## Input Preparation for Unicycler

### Short-Read Quality Control

Illumina short reads are the foundation of the Unicycler assembly graph. The quality of these reads directly affects base-level accuracy in the final assembly. Before running Unicycler, laboratories should assess read quality using standard metrics such as per-base quality scores, GC content distribution, and adapter contamination. Low-quality bases at read ends and adapter sequences should be removed because they can create spurious connections in the assembly graph.

The NCBI provides official documentation for sequence data resources and analysis services that can help laboratories understand read formats and quality metrics. The European Bioinformatics Institute offers training materials for bioinformatics data analysis that cover quality assessment workflows. These resources are useful for laboratories establishing standard operating procedures for sequencing data handling.

### Long-Read Quality Control

Long-read data from Oxford Nanopore and PacBio platforms require different quality considerations than short reads. Nanopore reads have historically had higher error rates than Illumina reads, and the error profile includes insertions and deletions in addition to substitutions. PacBio reads have a different error profile that is largely random. Unicycler is designed to tolerate these errors because it uses long reads primarily for structural information instead of base-level accuracy.

A benchmarking study of hybrid assembly approaches for bacterial pathogens tested Unicycler with simulated reads of mediocre and low quality, as well as real reads from multiple bacterial species. The study found that Unicycler performed best for achieving contiguous genomes among the tools tested. Importantly, the study also found that MaSuRCA was less tolerant of low-quality long reads than SPAdes and Unicycler. This finding suggests that Unicycler is a reasonable choice when long-read quality is uncertain or variable.

### Coverage Considerations

Coverage depth for both short and long reads affects assembly quality. The original Unicycler publication notes that the tool can produce good assemblies even when long-read depth and accuracy are low. This is a practical advantage because long-read sequencing can be more expensive than short-read sequencing, and laboratories may seek to minimize long-read coverage to control costs.

For short reads, higher coverage generally improves the accuracy of the initial assembly graph. Most bacterial genome projects aim for at least 50-fold coverage with Illumina data, though the optimal depth depends on the genome size and complexity. For long reads, even modest coverage can resolve structural features if the reads span repetitive regions. The benchmarking study of hybrid assembly approaches found that Unicycler assemblies had high genome completeness, approximately 98.7 percent, when compared to other assembly tools in a study of Mycobacterium tuberculosis genomes.

## Running Unicycler

### Installation and Dependencies

Unicycler is open source under the GPLv3 license and is available from the GitHub repository maintained by the developer. The tool requires several dependencies, including SPAdes for the initial short-read assembly, and it can use additional tools for read alignment and assembly polishing. Laboratories should install Unicycler in a controlled computing environment and verify that all dependencies are available before starting an assembly run.

The Bioconductor project provides official documentation for reproducible genomic analysis workflows, and the Galaxy Training Network offers accessible tutorials for assembly and analysis. These resources can help laboratories implement Unicycler in a reproducible manner. The nf-core documentation describes community standards for pipeline usage and configuration, which is relevant for laboratories that want to integrate Unicycler into larger automated workflows.

### Command Structure

A typical Unicycler command specifies the short-read FASTQ files, the long-read FASTQ file, and an output directory. The tool automatically detects whether reads are from Oxford Nanopore or PacBio platforms based on the data characteristics. The command also accepts options for adjusting the assembly mode and for specifying the number of threads to use.

The assembly mode is an important decision. Normal mode is appropriate for most bacterial genomes and balances speed with accuracy. Bold mode assumes high-quality data and may produce more complete assemblies but can be less robust to data problems. Conservative mode is designed for problematic data and may produce more fragmented assemblies but with fewer errors. Laboratories should start with normal mode and adjust based on the characteristics of their data.

### Runtime and Resource Requirements

Unicycler runtime depends on genome size, read depth, and available computing resources. Bacterial genomes of approximately 4 to 6 million base pairs typically assemble in a reasonable time on a standard server with multiple CPU cores. The SPAdes step is often the most computationally intensive part of the pipeline. Laboratories should monitor memory usage during assembly runs and ensure that sufficient resources are available.

A comparison of long-read sequencing technologies in hybrid assembly of complex bacterial genomes used Unicycler to assemble 20 bacterial isolates from the Enterobacteriaceae family. These genomes frequently have highly plastic, repetitive genetic structures, and complete reconstruction is relevant for understanding antimicrobial resistance epidemiology. The study found that hybrid assembly with either PacBio or ONT reads facilitated high-quality genome reconstruction and was superior to long-read-only assembly followed by short-read polishing.

## Assembly Modes and Their Tradeoffs

### Normal Mode

Normal mode is the default and recommended starting point for most bacterial genome assemblies. It uses a balanced approach to graph simplification that works well for typical bacterial genomes with moderate repeat content. The original Unicycler publication describes the algorithm as building an initial assembly graph from short reads and then simplifying the graph using information from both short and long reads.

### Bold Mode

Bold mode is designed for high-quality data where the assembler can be more aggressive in resolving ambiguous regions. This mode may produce more complete assemblies with fewer contigs, but it carries a higher risk of misassembly if the data contain errors or if the genome has unusual structural features. Laboratories should use bold mode only when they have confidence in the quality of both short and long reads.

### Conservative Mode

Conservative mode is designed for problematic data, such as low-coverage long reads or genomes with complex repeat structures. This mode produces more fragmented assemblies but with fewer misassemblies. The tradeoff is that the assembly may require additional manual finishing steps to achieve complete genome reconstruction.

The benchmarking study of hybrid assembly approaches for bacterial pathogens found that all SPAdes assemblies were incomplete when compared to Unicycler and MaSuRCA. This finding underscores the importance of choosing an assembler that can effectively use both short and long read data. Unicycler's performance was strongest for achieving contiguous genomes, closely followed by MaSuRCA.

## Quality Assessment of Hybrid Assemblies

### Assembly Statistics

After Unicycler completes an assembly run, laboratories should assess the output using standard assembly statistics. The number of contigs, the N50 value, and the total assembly size provide a first indication of assembly quality. A complete bacterial genome assembly should ideally consist of one contig per replicon, with the chromosome and each plasmid represented as a single circular sequence.

The genome assembly size should be compared to the expected genome size for the species. For example, the Mycobacterium tuberculosis reference genome H37Rv is approximately 4,411,532 base pairs. A study comparing assembly tools for M. tuberculosis genomes found that Unicycler assemblies had a mean size of 4,377,642 base pairs, which is close to the expected genome size. RagOut assemblies were significantly longer at 4,418,574 base pairs, which may indicate over-assembly or inclusion of spurious sequence.

### Genome Completeness

Genome completeness can be assessed using tools that search for conserved single-copy genes. The benchmarking study of hybrid assembly approaches for bacterial pathogens determined genome completeness and accuracy for assemblies of ten bacterial species. Unicycler assemblies had high completeness, and the study found that hybrid assemblies with ONT and PacBio long reads detected more genes than short-read assembly alone.

The comparison of long-read sequencing technologies in hybrid assembly of complex bacterial genomes found that hybrid assembly with either PacBio or ONT reads facilitated high-quality genome reconstruction. This study also found that combining ONT and Illumina reads fully resolved most genomes without additional manual steps, and at a lower consumables cost per isolate in the study setting.

### Circularity Assessment

A key advantage of hybrid assembly is the ability to produce circular chromosomes and plasmids. Unicycler can produce circular contigs when the data support circularization. Laboratories should check the assembly output for circular contigs and verify that the circularization is supported by read evidence. A circular chromosome that is actually linear or that has an incorrect join can introduce errors in downstream analysis.

The hybrid assembly of the colistin-resistant E. coli strain from Brazil produced a complete genome that enabled detailed analysis of the resistance plasmid. The mcr-1.5 gene was carried by an IncI2 plasmid, and the full genome SNP-based phylogenetic analysis revealed that the strain was highly related to colistin-resistant ST354 lineages associated with urinary tract infections in Brazil since 2015. This level of analysis depends on a complete and accurate assembly.

## Records and Measurements for Assembly Projects

### Documentation Standards

Laboratories should maintain detailed records for each assembly project. The records should include the sequencing platform and version, read quality metrics, coverage depth for short and long reads, Unicycler version and parameters, assembly statistics, and the date and personnel responsible for the assembly. This documentation supports reproducibility and troubleshooting.

The Carpentries offers lessons on foundational computing, data, shell, Git, and programming that are useful for laboratories implementing reproducible bioinformatics workflows. The Galaxy Training Network provides accessible workflow training and analysis tutorials that emphasize reproducibility. The nf-core documentation describes community pipeline standards for usage and configuration that support reproducible analysis.

### Quality Metrics to Record

The following metrics should be recorded for each assembly project:

| Metric | Purpose | Interpretation |
| --- | --- | --- |
| Read count and total bases | Assess coverage depth | Higher coverage generally improves assembly quality |
| Read N50 | Assess read length distribution | Longer reads improve structural resolution |
| Assembly contig count | Assess fragmentation | Fewer contigs indicate more complete assembly |
| Assembly N50 | Assess contig length distribution | Higher N50 indicates better assembly |
| Total assembly size | Compare to expected genome size | Large deviations may indicate contamination or misassembly |
| Genome completeness | Assess gene content | Values near 100 percent indicate complete gene representation |
| Circular contigs | Assess structural resolution | Circular chromosomes and plasmids indicate complete assembly |

## Common Failure Patterns in Unicycler Assemblies

### Low Long-Read Coverage

When long-read coverage is too low, Unicycler may fail to resolve repetitive regions, resulting in fragmented assemblies. The original publication notes that Unicycler can assemble larger contigs with fewer misassemblies than other hybrid assemblers even when long-read depth is low, but there is a practical minimum below which structural resolution fails. Laboratories should assess long-read coverage before assembly and consider additional sequencing if coverage is inadequate.

### Contaminated Input Data

Contamination from other organisms can produce assembly graphs with spurious connections and inflated genome sizes. The assembly size should be compared to the expected genome size for the target species. A benchmarking study of hybrid assembly approaches for bacterial pathogens found that the MaSuRCA assembly of Staphylococcus aureus with real reads contained antimicrobial resistance genes that were not present in the reference genome or in the Unicycler assembly. This discrepancy may indicate contamination or misassembly in the MaSuRCA output.

### Poor Short-Read Quality

Low-quality short reads can introduce errors in the initial assembly graph that propagate through the hybrid assembly. Adapter contamination and low-quality base calls should be removed before assembly. The NCBI provides official descriptions of sequence data resources that can help laboratories understand read quality metrics and data formats.

### Complex Repeat Structures

Some bacterial genomes have complex repeat structures that are difficult to resolve even with long reads. The comparison of long-read sequencing technologies in hybrid assembly of complex bacterial genomes selected isolates from the Enterobacteriaceae family because these frequently have highly plastic, repetitive genetic structures. The study found that hybrid assembly with either PacBio or ONT reads facilitated high-quality genome reconstruction, but some genomes may require additional manual finishing steps.

## Interpretation Limits and Reporting

### What Hybrid Assembly Cannot Resolve

Hybrid assembly with Unicycler resolves many structural features, but it has limitations. Very long repeats that exceed the length of the longest long reads may remain ambiguous. Highly repetitive regions such as ribosomal RNA operons may be collapsed or misassembled. Laboratories should be cautious when interpreting assemblies that contain unresolved regions.

The original Unicycler publication describes the tool as producing assemblies that are accurate, complete, and cost-effective. However, the authors also note that few tools exist that truly leverage the benefits of both types of data. This statement acknowledges that hybrid assembly is an active area of development and that current tools have limitations.

### Reporting Standards

When reporting hybrid assembly results, laboratories should describe the sequencing platforms, coverage depths, assembly parameters, and quality metrics. This information allows other researchers to assess the reliability of the assembly and to compare results across studies. The European Bioinformatics Institute offers training materials for bioinformatics data analysis that cover reporting standards and data sharing practices.

The benchmarking study of hybrid assembly approaches for bacterial pathogens reported genome completeness and accuracy, antimicrobial resistance, virulence potential, multilocus sequence typing, phylogeny, and pan genome for assemblies of multiple bacterial species. This comprehensive reporting approach provides a model for laboratories that want to characterize their assemblies thoroughly.

## Safety and Regulatory Context

### Data Management

Bacterial genome sequence data may be subject to institutional, national, or international data management requirements. Laboratories should be aware of applicable regulations for data storage, sharing, and publication. The NCBI provides official descriptions of sequence databases and data submission procedures that laboratories can use to comply with data sharing requirements.

### Antimicrobial Resistance Data

Assemblies of antimicrobial-resistant bacteria may have implications for public health surveillance and clinical decision-making. The hybrid assembly of the colistin-resistant E. coli strain from Brazil revealed a broad resistome that included antibiotics, heavy metals, disinfectants, and glyphosate. Laboratories that assemble genomes of resistant pathogens should consider the public health context and follow applicable reporting requirements.

### Professional Escalation Criteria

Laboratories should escalate assembly problems to a bioinformatics specialist or supervisor when they encounter any of the following situations:

| Situation | Action |
| --- | --- |
| Assembly fails to complete | Review error logs and check input data quality |
| Assembly size deviates substantially from expected | Check for contamination and reassess input data |
| Genome completeness is below 95 percent | Consider additional sequencing or alternative assembly parameters |
| Circular contigs are not produced for expected replicons | Investigate repeat structures and long-read coverage |
| Antimicrobial resistance genes are detected in unexpected contexts | Verify assembly accuracy before reporting |

## A Practical Decision Framework for Choosing Between Hybrid Assembly Strategies

Microbiologists often assume that running Unicycler with any combination of short and long reads will produce a complete genome. In practice, the choice of long-read platform, the depth of sequencing, and the expected genome architecture all influence whether Unicycler is the right tool for a given project. This section provides a structured decision framework that laboratories can use before committing computational resources and sequencing budgets to a hybrid assembly project.

### Platform Selection Based on Genome Architecture

The first decision in any hybrid assembly project is which long-read platform to use. Oxford Nanopore Technologies and Pacific Biosciences both generate reads that can resolve repetitive structures, but they differ in throughput, cost, error profile, and practical workflow requirements. A comparison of long-read sequencing technologies in hybrid assembly of complex bacterial genomes examined 20 bacterial isolates from the Enterobacteriaceae family, a group chosen because these genomes frequently have highly plastic, repetitive genetic structures. The study found that hybrid assembly with either PacBio or ONT reads facilitated high-quality genome reconstruction and was superior to long-read-only assembly followed by short-read polishing with respect to accuracy and completeness.

The practical implication is that platform choice should be driven by laboratory infrastructure and project timelines instead of by a belief that one platform is inherently superior for Unicycler. ONT sequencing offers the advantage of real-time data generation and lower upfront instrument costs, which makes it attractive for laboratories that do not have access to a PacBio instrument. The comparison study noted that combining ONT and Illumina reads fully resolved most genomes without additional manual steps and at a lower consumables cost per isolate in the study setting. PacBio sequencing, by contrast, offers higher per-read accuracy and a more uniform error profile, which can simplify downstream analysis even though the consumables cost may be higher.

For laboratories that are establishing a hybrid assembly workflow for the first time, the decision framework should include an assessment of existing sequencing infrastructure. If the laboratory already has access to an Illumina instrument and a Nanopore device, the marginal cost of adding ONT long reads to an existing short-read project is lower than the cost of outsourcing PacBio sequencing. If the laboratory is planning to sequence many isolates over time, the per-sample consumables cost difference between ONT and PacBio becomes a significant factor in the decision.

### Coverage Depth Decisions for Long Reads

The original Unicycler publication emphasizes that the tool can assemble larger contigs with fewer misassemblies than other hybrid assemblers even when long-read depth and accuracy are low. This robustness is a practical advantage, but it does not mean that coverage depth is irrelevant. The decision framework should include a minimum coverage threshold based on the expected repeat content of the target genome.

For genomes with modest repeat content, such as many Escherichia coli isolates, low long-read coverage may be sufficient to resolve the chromosome and common plasmids. The hybrid assembly of a colistin-resistant E. coli strain from Brazil used Illumina and Nanopore sequence data and produced a genome of 5,333,039 base pairs with the mcr-1.5 gene carried on an IncI2 plasmid. This result demonstrates that Unicycler can produce complete assemblies with practical levels of long-read coverage.

For genomes with complex repeat structures, such as those found in the Enterobacteriaceae family, higher long-read coverage provides more evidence for resolving ambiguous graph connections. The comparison study of long-read sequencing technologies found that hybrid assembly with either PacBio or ONT reads facilitated high-quality genome reconstruction for these challenging genomes. Laboratories should consider the repeat content of the target species when deciding how much long-read coverage to generate.

A practical approach is to start with a modest long-read coverage target and assess the assembly output. If the assembly is fragmented or contains unresolved regions, additional long-read sequencing can be generated and the assembly rerun. This iterative approach avoids the cost of generating excessive long-read coverage upfront while still allowing for complete assembly when needed.

### When to Use Unicycler Versus Alternative Assembly Strategies

The decision framework should also address when Unicycler is the appropriate tool and when alternative strategies may be more suitable. A benchmarking study of hybrid assembly approaches for bacterial pathogens compared Unicycler, MaSuRCA, and SPAdes using simulated reads of mediocre and low quality, as well as real reads from multiple bacterial species. The study found that Unicycler performed the best for achieving contiguous genomes, closely followed by MaSuRCA, while all SPAdes assemblies were incomplete. The study also found that MaSuRCA was less tolerant of low-quality long reads than SPAdes and Unicycler.

These findings support the use of Unicycler as the default hybrid assembler for bacterial genomes, particularly when long-read quality is uncertain. However, the decision framework should include criteria for when to consider alternative approaches. If the laboratory has very high-quality long reads and a simple genome architecture, a long-read-only assembly followed by short-read polishing may be sufficient. The comparison study of long-read sequencing technologies found that this approach was inferior to hybrid assembly with respect to accuracy and completeness, but it may be acceptable for projects that do not require complete genome reconstruction.

If the laboratory has access to a well-validated MaSuRCA pipeline and the long-read data are of high quality, MaSuRCA may produce comparable results. The benchmarking study found that MaSuRCA was less tolerant of low-quality long reads, so laboratories with variable long-read quality should prefer Unicycler. The decision framework should include a quality assessment step for long reads before choosing the assembler.

### Decision Matrix for Assembly Strategy Selection

The following table provides a structured decision matrix that laboratories can use to select an assembly strategy based on their specific project characteristics.

| Project Characteristic | Recommended Strategy | Rationale |
| --- | --- | --- |
| Complete genome needed, long-read quality variable | Unicycler hybrid assembly | Robust to low-quality long reads, produces contiguous assemblies |
| Complete genome needed, high-quality long reads available | Unicycler or MaSuRCA hybrid assembly | Both tools perform well with high-quality data |
| Draft assembly acceptable, short reads only | SPAdes or similar short-read assembler | Lower cost, sufficient for many analyses |
| Complete genome needed, no long-read platform available | Consider outsourcing long-read sequencing | Hybrid assembly requires long reads for structural resolution |
| Complex repeat structure expected | Unicycler with higher long-read coverage | Additional coverage improves resolution of ambiguous regions |
| Many isolates to assemble | Unicycler with ONT long reads | Lower consumables cost per isolate in many settings |

### Cost-Benefit Analysis for Hybrid Assembly

The decision framework should include a cost-benefit analysis that considers the value of complete genome reconstruction relative to the cost of long-read sequencing. The comparison study of long-read sequencing technologies noted that combining ONT and Illumina reads fully resolved most genomes at a lower consumables cost per isolate in the study setting. This finding is relevant for laboratories that are planning large-scale sequencing projects.

For epidemiological investigations, complete genome reconstruction provides information that is not available from draft assemblies. The hybrid assembly of the colistin-resistant E. coli strain from Brazil enabled full genome SNP-based phylogenetic analysis, which revealed that the strain was highly related to colistin-resistant ST354 lineages associated with urinary tract infections in Brazil since 2015. This level of analysis requires a complete assembly that includes plasmids and other mobile genetic elements.

For antimicrobial resistance surveillance, complete assemblies allow precise localization of resistance genes to chromosomes or plasmids. The benchmarking study of hybrid assembly approaches for bacterial pathogens found that hybrid assemblies of five antimicrobial-resistant strains with simulated reads provided consistent antimicrobial resistance genotypes with the reference genomes. The study also found that the MaSuRCA assembly of Staphylococcus aureus with real reads contained antimicrobial resistance genes that were not present in the reference genome or in the Unicycler assembly, highlighting the importance of assembly accuracy for resistance gene reporting.

Laboratories should weigh the cost of long-read sequencing against the value of complete genome reconstruction for their specific research questions. For projects that require precise localization of mobile genetic elements, complete assemblies justify the additional cost. For projects that only require gene presence or absence calls, draft assemblies may be sufficient.

### Implementation Steps for the Decision Framework

The following steps provide a structured approach to implementing the decision framework in a laboratory setting.

**Step 1: Characterize the target genome.** Determine the expected genome size, repeat content, and plasmid content for the target species. This information can be obtained from published reference genomes or from databases such as those maintained by the NCBI. The NCBI provides official descriptions of sequence databases and analysis services that can help laboratories identify appropriate reference genomes.

**Step 2: Assess available sequencing infrastructure.** Determine which sequencing platforms are available in the laboratory or through collaborators. Consider the cost, turnaround time, and throughput of each platform. The European Bioinformatics Institute offers training materials for bioinformatics data analysis that can help laboratories understand the strengths and limitations of different sequencing platforms.

**Step 3: Estimate coverage requirements.** Based on the genome architecture and the expected repeat content, estimate the minimum long-read coverage needed for complete assembly. For genomes with modest repeat content, lower coverage may be sufficient. For genomes with complex repeat structures, higher coverage is recommended.

**Step 4: Select the assembly strategy.** Use the decision matrix to select the appropriate assembly strategy. For most bacterial genome projects, Unicycler hybrid assembly is the recommended default. Consider alternative strategies only when specific project characteristics support their use.

**Step 5: Document the decision.** Record the rationale for the assembly strategy selection, including the genome characteristics, sequencing platform, coverage estimates, and expected outcomes. This documentation supports reproducibility and provides a basis for troubleshooting if the assembly does not meet expectations.

**Step 6: Execute the assembly and assess the output.** Run Unicycler with the selected parameters and assess the assembly using the quality metrics described in the records and measurements section. If the assembly does not meet expectations, revisit the decision framework and consider whether additional sequencing or alternative parameters are needed.

### Common Failure Patterns in Strategy Selection

The decision framework helps laboratories avoid common failure patterns that arise from poor strategy selection. One common failure is generating insufficient long-read coverage for a genome with complex repeat structures. The comparison study of long-read sequencing technologies selected isolates from the Enterobacteriaceae family because these frequently have highly plastic, repetitive genetic structures. Laboratories that underestimate the repeat content of their target genome may generate insufficient long-read coverage and produce fragmented assemblies.

Another common failure is choosing a long-read platform based on cost alone without considering the error profile and its interaction with the assembler. The benchmarking study of hybrid assembly approaches for bacterial pathogens found that MaSuRCA was less tolerant of low-quality long reads than SPAdes and Unicycler. Laboratories that use MaSuRCA with low-quality long reads may produce assemblies with more errors than they would obtain with Unicycler.

A third common failure is assuming that all hybrid assemblers produce equivalent results. The benchmarking study found that all SPAdes assemblies were incomplete when compared to Unicycler and MaSuRCA. Laboratories that use SPAdes for hybrid assembly may produce fragmented assemblies that require additional finishing steps.

### Professional Escalation Criteria for Strategy Decisions

Laboratories should escalate strategy decisions to a bioinformatics specialist or supervisor when they encounter any of the following situations:

| Situation | Action |
| --- | --- |
| Uncertainty about genome repeat content | Consult published reference genomes or seek specialist advice |
| Limited access to long-read sequencing platforms | Discuss outsourcing options with supervisor or collaborators |
| Repeated assembly failures with the selected strategy | Reassess the decision framework and consider alternative approaches |
| Budget constraints that limit long-read coverage | Discuss the tradeoff between coverage and assembly completeness |
| Unusual genome architecture or suspected contamination | Seek specialist advice before proceeding with assembly |

### Integration with Reproducible Workflow Standards

The decision framework should be integrated with reproducible workflow standards to ensure that assembly strategy decisions are documented and traceable. The nf-core documentation describes community pipeline standards for usage and configuration that support reproducible analysis. The Galaxy Training Network offers accessible workflow training and analysis tutorials that emphasize reproducibility. The Carpentries offers lessons on foundational computing, data, shell, Git, and programming that are useful for laboratories implementing reproducible bioinformatics workflows.

Laboratories should record the decision framework inputs and outputs for each assembly project. This documentation should include the genome characteristics, sequencing platform selection, coverage estimates, assembly strategy selection, and the rationale for each decision. This information supports troubleshooting and allows other researchers to understand the basis for assembly strategy decisions.

The Bioconductor project provides official documentation for reproducible genomic analysis workflows that can help laboratories implement the decision framework in a reproducible manner. The European Bioinformatics Institute offers training materials for bioinformatics data analysis that cover practical analysis education and data-resource training. These resources can help laboratories establish standard operating procedures for assembly strategy selection.

### Validation of Strategy Decisions

The decision framework should include a validation step that confirms the selected strategy produces the expected results. The validation step should compare the assembly output to the expected genome characteristics for the target species. The genome assembly size should be compared to the expected genome size. For example, the Mycobacterium tuberculosis reference genome H37Rv is approximately 4,411,532 base pairs. A study comparing assembly tools for M. tuberculosis genomes found that Unicycler assemblies had a mean size of 4,377,642 base pairs, which is close to the expected genome size.

The validation step should also assess genome completeness using tools that search for conserved single-copy genes. The benchmarking study of hybrid assembly approaches for bacterial pathogens found that Unicycler assemblies had high genome completeness, approximately 98.7 percent in a study of M. tuberculosis genomes. The study also found that hybrid assemblies with ONT and PacBio long reads detected more genes than short-read assembly alone.

If the assembly output does not match the expected genome characteristics, the laboratory should revisit the decision framework and consider whether the strategy selection was appropriate. This iterative approach ensures that the decision framework produces reliable results across a range of bacterial genome projects.

## Frequently Asked Questions

### What are the minimum input requirements for Unicycler?

Unicycler requires short reads, typically from Illumina sequencing, and long reads from either Oxford Nanopore or PacBio platforms. The short reads should be in FASTQ format with quality scores. The long reads should also be in FASTQ format. Unicycler automatically detects the long-read platform based on data characteristics. The original publication describes Unicycler as building an initial assembly graph from short reads using SPAdes and then simplifying the graph using information from short and long reads.

### How much long-read coverage is needed for a good hybrid assembly?

The original Unicycler publication reports that the tool can assemble larger contigs with fewer misassemblies than other hybrid assemblers even when long-read depth and accuracy are low. However, the practical minimum coverage depends on the genome complexity and the length of repetitive regions. Laboratories should assess long-read coverage before assembly and consider additional sequencing if the assembly is fragmented.

### Can Unicycler use PacBio and Oxford Nanopore reads interchangeably?

Yes, Unicycler can use long reads from either PacBio or Oxford Nanopore platforms. A comparison of long-read sequencing technologies in hybrid assembly of complex bacterial genomes found that hybrid assembly with either PacBio or ONT reads facilitated high-quality genome reconstruction. The study also found that combining ONT and Illumina reads fully resolved most genomes without additional manual steps.

### How does Unicycler compare to other hybrid assemblers?

A benchmarking study of hybrid assembly approaches for bacterial pathogens found that Unicycler performed the best for achieving contiguous genomes, closely followed by MaSuRCA, while all SPAdes assemblies were incomplete. The study also found that MaSuRCA was less tolerant of low-quality long reads than SPAdes and Unicycler. The original Unicycler publication reports that Unicycler can assemble larger contigs with fewer misassemblies than other hybrid assemblers.

### What should I do if my Unicycler assembly is fragmented?

If the assembly is fragmented, first check the quality of the input reads. Low-quality short reads or insufficient long-read coverage can cause fragmentation. Consider running Unicycler in conservative mode, which is designed for problematic data. If the assembly remains fragmented, additional long-read sequencing may be needed to resolve repetitive regions.

### How do I know if my assembly is complete?

A complete bacterial genome assembly should have one contig per replicon, with the chromosome and each plasmid represented as a single circular sequence. Genome completeness can be assessed using tools that search for conserved single-copy genes. The benchmarking study of hybrid assembly approaches for bacterial pathogens found that Unicycler assemblies had high genome completeness, approximately 98.7 percent in a study of Mycobacterium tuberculosis genomes.

### Can Unicycler assemble plasmids?

Yes, Unicycler can assemble plasmids when the read data support their reconstruction. The hybrid assembly of the colistin-resistant E. coli strain from Brazil produced a complete genome that included the mcr-1.5 gene carried by an IncI2 plasmid of approximately 65,458 base pairs. Plasmid assembly requires that the plasmids are present in the sequenced sample and that the read data provide sufficient coverage.

### What downstream analyses can I perform with a Unicycler hybrid assembly?

A complete hybrid assembly supports a wide range of downstream analyses, including antimicrobial resistance prediction, virulence gene detection, multilocus sequence typing, phylogenomic analysis, and pan-genome analysis. The benchmarking study of hybrid assembly approaches for bacterial pathogens used hybrid assemblies to determine antimicrobial resistance, virulence potential, multilocus sequence typing, phylogeny, and pan genome for multiple bacterial species.

## Related Bioinformatics Guides

- [Hybrid Genome Assembly: Combining Short and Long Reads for Better Results](/knowledge/bioinformatics/hybrid-genome-assembly-combining-short-and-long-reads-for-better-results)
- [De Novo Genome Assembly with Long Reads: A Practical Workflow](/knowledge/bioinformatics/de-novo-genome-assembly-with-long-reads-a-practical-workflow)
- [Metagenomic Assembly and Binning: A Practical Workflow for Recovering Genomes from Complex Microbial Communities](/knowledge/bioinformatics/metagenomic-assembly-and-binning-a-practical-workflow-for-recovering-genomes-from-complex-microb)
- [Evaluating Metagenomic Assembly Tools: A Benchmarking Framework for Short-Read and Long-Read Data](/knowledge/bioinformatics/evaluating-metagenomic-assembly-tools-a-benchmarking-framework-for-short-read-and-long-read-data)
- [Long-Read Sequencing for De Novo Assembly of Complex Genomes: Case Studies and Best Practices](/knowledge/bioinformatics/long-read-sequencing-for-de-novo-assembly-of-complex-genomes-case-studies-and-best-practices)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)


## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [Unicycler: Resolving bacterial genome assemblies from short and long sequencing reads.](https://pubmed.ncbi.nlm.nih.gov/28594827). PLoS computational biology, 2017.
- [Comparisons of genome assembly tools for characterization of Mycobacterium tuberculosis genomes using hybrid sequencing technologies.](https://pubmed.ncbi.nlm.nih.gov/39221271). PeerJ, 2024.
- [Benchmarking hybrid assembly approaches for genomic analyses of bacterial pathogens using Illumina and Oxford Nanopore sequencing.](https://pubmed.ncbi.nlm.nih.gov/32928108). BMC genomics, 2020.
- [Hybrid genome assembly of colistin-resistant mcr-1.5-producing Escherichia coli ST354 reveals phylogenomic pattern associated with urinary tract infections in Brazil.](https://pubmed.ncbi.nlm.nih.gov/38408561). Journal of global antimicrobial resistance, 2024.
- [Comparison of long-read sequencing technologies in the hybrid assembly of complex bacterial genomes.](https://pubmed.ncbi.nlm.nih.gov/31483244). Microbial genomics, 2019.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.