Long-Read Genome Assembly and Polishing Strategies
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Long-read sequencing technologies (e.g., PacBio, Oxford Nanopore) generate reads exceeding tens of kilobases, enabling the resolution of complex genomic structures, repetitive regions, and structural variants that are intractable with short-read sequencing. These long reads, despite higher error rates (5-15%), are crucial for accurate annotation of virulence factors and antimicrobial resistance genes in veterinary pathogens.
- Overlap-Layout-Consensus (OLC) is the predominant paradigm for long-read assembly, with algorithms like Miniasm, FALCON, and NextDenovo employing variations such as minimal OLC, hierarchical correction, and pre-assembly error correction to improve contiguity and accuracy. Verkko2 integrates proximity-ligation data with long-read de Bruijn graphs for telomere-to-telomere assembly.
- Polishing strategies are essential for correcting base-level errors, primarily indels in homopolymer regions and GC-rich sequences. Self-polishing tools like Racon, Medaka, and Homopolish utilize long reads for correction, while hybrid polishing employs highly accurate short reads with tools such as Pilon, Polypolish, and Pypolca to achieve superior consensus accuracy.
- Effective polishing requires careful consideration of short-read coverage; approximately 25x coverage is generally optimal for hybrid polishing, with conservative tools like Polypolish-careful recommended for lower depths (<25x) to mitigate false-positive error introduction. Repeat-aware and iterative polishing strategies are critical for achieving telomere-to-telomere quality and further reducing residual error rates.
- Assembly quality is rigorously evaluated using metrics such as BUSCO for gene completeness and Assembly Quality Value (QV) for per-base accuracy. Tools like CloseRead aid in diagnosing errors in complex genomic regions, and validation often involves aligning short reads or comparing against available reference genomes.
Introduction
The advent of third-generation sequencing technologies has enabled the generation of reads exceeding tens of kilobases in length, fundamentally transforming genome assembly [<a href="#ref-1">1</a>, <a href="#ref-2">2</a>]. Unlike short-read platforms that produce highly accurate but short sequences, long-read platforms yield reads with substantially higher error rates (typically 5-15% per base) but provide the contiguity necessary to resolve repetitive regions, structural variants, and complex genomic architectures [<a href="#ref-3">3</a>, <a href="#ref-4">4</a>]. In veterinary medicine, long-read assembly is critical for constructing complete genomes of livestock pathogens, zoonotic agents, and host species, facilitating accurate annotation of virulence factors, antimicrobial resistance genes, and mobile genetic elements [<a href="#ref-1">1</a>, <a href="#ref-5">5</a>]. The assembly process typically involves two major phases: initial contig construction from raw long reads and subsequent polishing to correct residual errors [<a href="#ref-6">6</a>, <a href="#ref-7">7</a>]. This article provides an exhaustive review of long-read genome assembly algorithms and polishing strategies, with emphasis on methodologies applicable to microbial and veterinary genomes.
Long-Read Assembly Algorithms
Overlap-Layout-Consensus (OLC) Approaches
Long-read assemblers predominantly employ the overlap-layout-consensus (OLC) paradigm, which is well suited for reads with high error rates and long lengths [<a href="#ref-8">8</a>]. In OLC, all pairwise read overlaps are computed, a string graph is constructed, and consensus sequences are derived from the graph. The Miniasm assembler implements a minimal OLC approach that bypasses error correction, relying instead on raw read overlaps and subsequent consensus generation with tools such as Racon [<a href="#ref-8">8</a>]. The FALCON assembler extends OLC with a hierarchical correction step and produces phased diploid assemblies through the FALCON-Unzip algorithm [<a href="#ref-9">9</a>]. NextDenovo employs an efficient error correction module prior to OLC assembly, achieving high accuracy on noisy nanopore long reads [<a href="#ref-10">10</a>]. Verkko2 integrates proximity-ligation data with long-read de Bruijn graphs to achieve telomere-to-telomere assembly, demonstrating substantial improvements in contiguity and phasing [<a href="#ref-11">11</a>].
De Bruijn Graph and Hybrid Approaches
While de Bruijn graphs are standard for short-read assembly, their application to long reads is complicated by high error rates. However, hybrid assemblers combine short-read accuracy with long-read contiguity. The SPAdes assembler incorporates long reads via a hybrid mode that uses short-read de Bruijn graphs and long-read repeat resolution [<a href="#ref-3">3</a>]. The Unicycler pipeline automates hybrid assembly for bacterial genomes, using short reads to correct long-read assemblies [<a href="#ref-1">1</a>]. LRScaf and SSPACE-LongRead are scaffolding tools that leverage long reads to order and orient contigs produced by short-read assemblers [<a href="#ref-12">12</a>, <a href="#ref-13">13</a>]. Gap filling is addressed by LR_Gapcloser, which uses raw long reads to close assembly gaps with high efficiency and low memory usage [<a href="#ref-14">14</a>], and by gapFinisher, which processes SSPACE-LongRead output to fill gaps [<a href="#ref-15">15</a>].
Multi-Sample and Haplotype-Aware Assembly
For population-scale studies, multi-sample assembly pipelines such as LORA enable simultaneous assembly of multiple genomes, leveraging shared haplotypes to improve contiguity [<a href="#ref-16">16</a>]. StrainCascade provides an automated modular workflow for high-throughput bacterial genome reconstruction, integrating long-read assembly with characterization of plasmids and mobile elements [<a href="#ref-5">5</a>]. Haplotype-aware error correction methods, such as those described by Barak et al., improve consensus accuracy in diploid or polyploid genomes by distinguishing allelic variants during correction [<a href="#ref-17">17</a>].
Polishing Strategies
Polishing is the process of correcting base-level errors in a draft assembly. Errors in long-read assemblies are predominantly insertions and deletions (indels) in homopolymer regions and GC-rich sequences [<a href="#ref-2">2</a>, <a href="#ref-7">7</a>]. Polishing strategies fall into two categories: self-polishing using only long reads, and hybrid polishing using short reads.
Self-Polishing with Long Reads
Self-polishing tools use the same long-read data to correct the assembly. Racon performs partial order alignment of reads against the assembly and generates a consensus [<a href="#ref-8">8</a>]. Medaka uses a neural network model trained on specific sequencing platform error profiles to predict corrections [<a href="#ref-7">7</a>]. NextPolish is a fast k-mer-based polisher that scores and counts k-mers from short reads but can also be used in a self-polishing mode with long reads [<a href="#ref-18">18</a>]. Homopolish specifically targets homopolymer errors and has shown superior performance on nanopore assemblies [<a href="#ref-2">2</a>, <a href="#ref-7">7</a>]. The PEPPER tool uses a deep learning model for margin polishing and is effective when combined with Medaka [<a href="#ref-7">7</a>]. Evaluation of individual tools on microbial genomes indicates that Homopolish, PEPPER, and Medaka yield the best results, but no single tool is universally optimal [<a href="#ref-7">7</a>].
Hybrid Polishing with Short Reads
Hybrid polishing leverages the high accuracy of short reads to correct long-read assemblies. Pilon is a widely used tool that aligns short reads to the assembly and identifies discrepancies [<a href="#ref-19">19</a>]. Polypolish uses a conservative approach that only makes corrections supported by multiple short-read alignments, minimizing false positives [<a href="#ref-19">19</a>]. Pypolca offers default and careful modes, with the careful mode avoiding false-positive errors at low coverage [<a href="#ref-19">19</a>]. NextPolish outperforms Pilon in speed and correction accuracy when using short-read data [<a href="#ref-18">18</a>]. The depth of short-read coverage significantly affects polishing performance: most benefits are achieved by 25x depth, and Polypolish-careful introduces no false-positive errors at any depth [<a href="#ref-19">19</a>]. For very low depth (<5x), Polypolish-careful alone is recommended; for low depth (5-25x), a combination of Polypolish-careful and Pypolca-careful is optimal; for sufficient depth (>25x), Polypolish-default and Pypolca-careful are preferred [<a href="#ref-19">19</a>].
Repeat-Aware and Iterative Polishing
Standard polishing tools can overcorrect in repetitive regions, introducing new errors. The telomere-to-telomere consortium developed a repeat-aware polishing strategy that uses a diverse panel of sequencing technologies to correct errors without overcorrection [<a href="#ref-6">6</a>]. This approach improved the assembly quality value from 70.2 to 73.9 in the CHM13 human genome [<a href="#ref-6">6</a>]. Iterative polishing, where multiple rounds of correction are applied, can further reduce error rates. For example, a second round of Homopolish after initial polishing improves accuracy [<a href="#ref-7">7</a>]. However, iterative polishing must be carefully monitored to avoid introducing systematic biases [<a href="#ref-6">6</a>].
Evaluation and Validation
Assessing assembly quality requires multiple metrics. BUSCO (Benchmarking Universal Single-Copy Orthologs) evaluates gene completeness [<a href="#ref-7">7</a>]. Assembly quality value (QV) estimates per-base accuracy based on k-mer counts [<a href="#ref-6">6</a>]. CloseRead is a tool that visualizes local assembly quality and diagnoses errors in structurally complex regions such as immunoglobulin loci, enabling targeted re-assembly [<a href="#ref-20">20</a>]. Validation strategies include aligning short reads to the assembly and counting discordant k-mers, as well as comparing to reference genomes when available [<a href="#ref-1">1</a>, <a href="#ref-6">6</a>]. For microbial genomes, gene prediction with Prokka and comparison to known sequences provides additional validation [<a href="#ref-7">7</a>].
Workflow Diagram
The following Mermaid diagram illustrates a typical long-read genome assembly and polishing workflow.
flowchart TD
A["Raw Long Reads"] --> B["Error Correction / Self-Correction"]
B --> C["Overlap-Layout-Consensus Assembly"]
C --> D["Draft Assembly"]
D --> E{"Polishing Strategy"}
E --> F["Self-Polishing: Racon, Medaka, NextPolish"]
E --> G["Hybrid Polishing: Short Reads + Pilon, Polypolish, Pypolca"]
F --> H["Iterative Polishing"]
G --> H
H --> I["Repeat-Aware Polishing"]
I --> J["Validation: BUSCO, QV, CloseRead"]
J --> K{"Errors Detected?"}
K -->|"Yes"| L["Targeted Re-assembly / Gap Filling"]
L --> D
K -->|"No"| M["Final Polished Assembly"]
Frequently Asked Questions
What is the difference between self-polishing and hybrid polishing?
Self-polishing uses only long reads to correct the assembly, while hybrid polishing incorporates short reads to achieve higher accuracy, particularly in homopolymer regions [<a href="#ref-7">7</a>, <a href="#ref-19">19</a>].
Which polishing tool is best for bacterial genomes?
The optimal tool depends on coverage and error profile. For nanopore-only assemblies, Homopolish, PEPPER, and Medaka perform well; for hybrid polishing, Pypolca-careful and Polypolish-careful are recommended [<a href="#ref-7">7</a>, <a href="#ref-19">19</a>].
How much short-read coverage is needed for effective polishing?
Most polishing benefits are achieved at 25x short-read coverage; lower depths require conservative tools like Polypolish-careful to avoid false positives [<a href="#ref-19">19</a>].
Can long-read assembly achieve telomere-to-telomere quality?
Yes, with advanced assemblers like Verkko2 and repeat-aware polishing, telomere-to-telomere assemblies are achievable for complex genomes [<a href="#ref-6">6</a>, <a href="#ref-11">11</a>].
What are common errors in long-read assemblies?
The most common errors are indels in homopolymer runs and GC-rich regions, which can be addressed by specialized polishers like Homopolish [<a href="#ref-2">2</a>, <a href="#ref-7">7</a>].