Assembly Polishing with Long Reads: How to Use PacBio and Nanopore Data to Correct Errors
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Assembly polishing corrects residual base errors in draft genomes, primarily indels in homopolymer regions and low-complexity sequences, which persist due to the inherent error rates of long-read sequencing platforms (PacBio and Nanopore).
- The choice of polishing tool is dictated by the sequencing platform and available data; Racon is a versatile first step for both PacBio and ONT, Medaka offers high accuracy for ONT data using neural networks, and Pilon leverages short reads for residual error correction.
- Iterative polishing is crucial, with multiple rounds of tools like Racon followed by platform-specific refinement (e.g., Medaka for ONT) and potentially short-read correction (e.g., Pilon), but evaluation with metrics like Merqury's QV score is essential to prevent over-polishing and error introduction.
- Evaluating assembly quality post-polishing is paramount; Merqury's QV score quantifies base accuracy improvement, while completeness metrics must be monitored to ensure polishing does not introduce structural errors or remove contigs.
- Common failure patterns include lack of quality improvement (due to uninformative reads or incorrect parameters), introduction of new errors (often from model mismatches or excessive iterations), and computational limitations, necessitating careful parameter tuning and tool selection.
- Polishing addresses base-level errors but cannot rectify structural misassemblies or recover missing sequence; these require reassembly or complementary data like optical maps, and the process is fundamentally limited by the quality of the input sequencing reads.
Direct Answer and Scope
Assembly polishing is the process of correcting residual base errors in a draft genome assembly by realigning sequencing reads to the assembled contigs and generating a corrected consensus sequence. After initial assembly with long reads from Pacific Biosciences (PacBio) or Oxford Nanopore Technologies (ONT), most draft genomes still contain systematic errors, particularly insertion and deletion (indel) errors in homopolymer regions and low-complexity sequences. These errors arise because long-read platforms historically produced raw reads with error rates around 10 percent, and assemblers cannot fully resolve every base position during contig construction. Polishing tools such as Racon, Medaka, and Pilon address this problem by using the original read data to refine the assembly, and the choice of tool depends on the sequencing platform, the availability of short-read data, and the computational resources at hand. This article provides a practical workflow for selecting and running polishing tools, evaluating the resulting improvements with Merqury, and recognizing when polishing has reached its practical limits.
The intended reader is a biology student, researcher, or laboratory professional who has produced a draft assembly and needs to reduce residual errors before downstream analysis. The guidance assumes familiarity with basic command-line operations and the ability to install bioinformatics software. The focus is on concrete decisions: which reads to use for polishing, which tool to run first, how to iterate, and how to verify that the polishing step actually improved the assembly instead of introducing new errors.
Why Long-Read Assemblies Need Polishing
Long-read sequencing technologies enable de novo genome assembly with high contiguity, but the raw reads carry a substantially higher error rate than short reads. Third-generation sequencing platforms produce reads of tens of kilobases, which solve many assembly problems related to repetitive regions and structural variation, yet the error rate of these reads is currently capped around 10 percent according to published evaluations of self-correction methods. If these errors are left unaddressed, the resulting genome assemblies may exhibit high base error rates that compromise the reliability of downstream analysis, including gene prediction, variant calling, and comparative genomics.
The error profiles differ between platforms. PacBio circular consensus sequencing (CCS) reads achieve higher per-base accuracy through multiple passes over the same molecule, while ONT reads have historically shown higher error rates with a particular tendency toward homopolymer-length errors. Both platforms produce errors that are not randomly distributed across the genome. Published comparisons of long-read assembly methods note that errors are predominantly indels and that the error distribution varies across different regions of an assembly. Low-complexity regions, homopolymer runs, and tandem repeats are especially prone to residual errors after the initial assembly step.
The practical consequence is that a draft assembly produced directly from long reads will contain hundreds or thousands of errors even in a small bacterial genome, and millions in a mammalian genome. Polishing is the procedure that fixes these errors in the draft assembly and improves the reliability of genomic analysis. The process works by aligning the original sequencing reads back to the assembled contigs and using the alignment information to compute a corrected consensus sequence at each position.
Core Principles of Polishing
Read Selection and Platform Considerations
The choice of reads for polishing is the first major decision. The most straightforward approach is to use the same long reads that were used for the initial assembly. This is called self-polishing, and it works because the assembler may have made local consensus errors even when the underlying read data contained the correct base at a given position. Aligning the reads back to the assembly and computing a new consensus can correct these errors.
The alternative is to use short reads for polishing. Short-read polishing tools can use short reads to fix errors that remain after long-read polishing, most commonly homopolymer-length errors. However, most short-read polishing tools rely on short-read alignment, which is unreliable in repeat regions. Errors in such regions are therefore challenging to fix and often remain after short-read polishing. Some tools address this limitation by using all-per-read alignments instead of only the best alignment for each read, which allows them to repair errors in repeat sequences that other polishers cannot.
The decision between long-read polishing and short-read polishing is not either-or. Many production workflows use both in sequence: first polish with long reads to correct the bulk of indel errors, then polish with short reads to fix residual homopolymer errors and other platform-specific artifacts. The availability of short-read data for the same sample is a practical constraint. If short reads were generated as part of the sequencing project, they should be retained for this purpose.
Tool Selection by Platform and Data Type
Racon is a widely used polishing tool that works with both PacBio and ONT data. It uses alignments of reads to the assembly and computes a consensus sequence from the aligned reads. Racon is fast and memory-efficient, making it a reasonable first polishing step for most assemblies. It can be run iteratively, with each round using the output of the previous round as the input assembly.
Medaka is a polishing tool developed by the Oxford Nanopore community. It uses a neural network model trained on ONT data to predict the correct consensus sequence from read alignments. Medaka generally produces higher accuracy than Racon on ONT data, but it requires more computational resources and is specific to the ONT platform. Published comparisons show that Medaka achieves base accuracy values upwards of 99.9 percent, corresponding to a Phred score above Q30, when used appropriately.
Pilon is a polishing tool that uses short-read alignments to correct assembly errors. It was originally developed for Illumina data and is commonly used to polish assemblies that were initially constructed from long reads. Pilon can correct both single-nucleotide errors and indel errors, and it also reports regions of the assembly that lack sufficient read coverage for confident correction.
The choice among these tools depends on the data available. For an ONT-only assembly, the standard workflow is Racon followed by Medaka. For a PacBio assembly, Racon is often sufficient, and some workflows add a short-read polishing step with Pilon if Illumina data are available. For bacterial genomes, a short-read polishing step with a tool designed for repeat-aware polishing can catch errors that long-read polishing misses.
Iterative Polishing and Convergence
Polishing is not a single pass operation. Running a polishing tool once will correct many errors, but a second or third round can correct additional errors that were masked by the first round. The practical question is how many rounds to run before the improvements plateau.
The standard approach is to run two to three rounds of Racon, then one round of Medaka for ONT data, then one round of Pilon if short reads are available. After each round, the assembly should be evaluated with a quality metric such as Merqury to determine whether the polishing step actually improved accuracy. If the quality score stops improving or begins to decline, further polishing is unlikely to help and may introduce new errors.
The risk of over-polishing is real. Each polishing round uses the output of the previous round as input, and errors in the input can be propagated or amplified. Tools that use neural network models, such as Medaka, are trained on specific error profiles, and applying them to data that do not match the training distribution can produce unexpected results. The safest practice is to evaluate the assembly after each round and stop when the quality metric plateaus.
At a Glance: Polishing Tool Selection
| Tool | Input Reads | Platform Focus | Strengths | Limitations | Typical Use |
|---|---|---|---|---|---|
| Racon | Long reads (PacBio or ONT) | Both | Fast, memory-efficient, works with any read type | Lower accuracy than model-based tools | First polishing pass, iterative correction |
| Medaka | ONT reads | ONT | High accuracy, neural network model, corrects homopolymer errors | ONT-specific, higher computational cost | Final long-read polishing pass for ONT assemblies |
| Pilon | Short reads (Illumina) | Short-read correction | Corrects residual errors after long-read polishing, reports low-coverage regions | Requires short-read data, unreliable in repeats with standard alignment | Short-read polishing step after long-read polishing |
| Polypolish | Short reads | Repeat regions | Uses all-per-read alignments, repairs errors in repeats | Best used in combination with other short-read polishers | Repeat-aware short-read polishing for bacterial genomes |
| GoldPolish-Target | Long reads | Targeted regions | Polishes user-specified loci, up to 27-fold faster than Medaka, 95 percent less memory | Targets only specified regions, not whole-genome polishing | Resource-efficient polishing of problem regions |
| BlockPolish | Long reads | Both PacBio and ONT | Divides contigs into blocks by complexity, corrects indels well | Requires multiple sequence alignment in complex blocks | High-accuracy polishing with region-aware strategy |
| CONSENT | Long reads | Both, scales to ultra-long reads | Self-correction and polishing, scales to human datasets | Multiple sequence alignment approach may be slower on small datasets | Polishing assemblies from ultra-long reads |
Practical Workflow for Polishing
Step 1: Prepare the Input Data
The polishing workflow requires two inputs: the draft assembly in FASTA format and the sequencing reads in FASTA or FASTQ format. The reads should be the same reads that were used for the initial assembly, or a subset of those reads if the full dataset is too large for practical processing.
Before starting the polishing workflow, verify that the read files are not corrupted and that the assembly file contains only the contigs that should be polished. Some assemblers produce contigs that are shorter than a minimum length threshold, and these may be excluded from polishing to save computational time. The decision to exclude short contigs should be made before polishing begins, because the polishing tools will process every sequence in the input file.
For ONT data, the read file should contain the basecalled reads in FASTQ format. For PacBio data, the read file should contain the subreads or the circular consensus sequences, depending on the sequencing protocol. The choice between subreads and CCS reads affects the polishing result, because CCS reads have higher per-base accuracy and may produce a cleaner consensus.
Step 2: Align Reads to the Assembly
All polishing tools require an alignment of the reads to the assembly. The alignment is typically generated with a long-read aligner such as minimap2, which is designed to handle the high error rates and long read lengths of PacBio and ONT data. The alignment output should be in SAM or BAM format, sorted and indexed for efficient access.
The alignment parameters matter. For Racon, the recommended practice is to use minimap2 with the preset appropriate for the read type: map-ont for ONT reads and map-pb for PacBio reads. These presets set the alignment parameters to account for the expected error profiles of each platform. Using the wrong preset can produce poor alignments and reduce the effectiveness of polishing.
For short-read polishing with Pilon, the short reads should be aligned to the assembly with a short-read aligner such as BWA-MEM. The alignment should be sorted and indexed, and duplicate reads should be marked or removed according to the tool documentation.
Step 3: Run Racon for Initial Polishing
Racon is typically the first polishing tool in the workflow because it is fast and works with any read type. The command takes three inputs: the assembly FASTA file, the read file, and the alignment file in SAM format. Racon produces a polished assembly as output.
The recommended practice is to run Racon for two or three rounds. Each round uses the output of the previous round as the input assembly, and the reads are realigned to the new assembly before each round. The realignment step is important because the improved assembly may produce better alignments, which in turn produce a better consensus.
The number of rounds should be guided by the quality evaluation. If the quality metric improves substantially between round one and round two but only marginally between round two and round three, the polishing has likely converged. Running additional rounds beyond convergence wastes computational time and risks introducing errors.
Step 4: Run Medaka for ONT Assemblies
For ONT assemblies, Medaka is the recommended final long-read polishing step. Medaka uses a neural network model to predict the consensus sequence from the read alignments, and it is trained on ONT-specific error profiles. The tool requires the reads to be aligned to the assembly with minimap2, and it produces a polished assembly as output.
Medaka is more computationally intensive than Racon, so it should be run after Racon has already corrected the bulk of the errors. Running Medaka on an unpolished assembly will still produce improvements, but the tool performs best when the input assembly is already close to the final sequence.
The Medaka model must match the basecalling model used to produce the reads. If the reads were basecalled with a specific model, the corresponding Medaka model should be selected. Using the wrong model can reduce the accuracy of the polishing result.
Step 5: Run Pilon for Short-Read Correction
If short reads are available for the same sample, a final polishing step with Pilon can correct residual errors that long-read polishing missed. Pilon aligns the short reads to the assembly and uses the alignments to identify and correct errors.
Pilon is particularly useful for correcting homopolymer-length errors, which are common in ONT assemblies and are not always fully corrected by Racon or Medaka. However, Pilon relies on short-read alignment, which is unreliable in repeat regions. Errors in repeat regions may remain after Pilon polishing, and some tools are specifically designed to address this limitation.
The Pilon output includes a FASTA file with the polished assembly and a report file that lists the changes made. The report should be reviewed to understand what errors were corrected and whether any regions were flagged for low coverage.
Step 6: Evaluate the Polished Assembly with Merqury
Merqury is a tool that evaluates assembly quality by comparing the assembly to a k-mer spectrum derived from the sequencing reads. It produces a quality value (QV) score that estimates the base accuracy of the assembly, along with other metrics such as completeness and consistency.
The Merqury workflow requires a k-mer database built from the sequencing reads. For a haploid assembly, the k-mer database is built from the reads used for polishing. For a diploid assembly, the k-mer database should include reads from both haplotypes, and the evaluation accounts for the expected heterozygosity.
The QV score from Merqury provides a quantitative measure of polishing improvement. A typical unpolished long-read assembly may have a QV score in the range of 30 to 40, corresponding to an error rate of 0.1 to 0.01 percent. After polishing, the QV score should increase, with well-polished assemblies reaching QV scores above 50 for bacterial genomes and above 40 for larger genomes.
The Merqury evaluation should be run before and after each polishing round to track the improvement. If the QV score does not improve after a polishing round, the polishing step did not help and may have introduced errors. In this case, the previous assembly should be retained and the workflow should be reviewed for potential issues.
Options and Tradeoffs in Polishing Strategies
Whole-Genome versus Targeted Polishing
The standard polishing workflow processes the entire assembly, but this is not always the most efficient approach. Published work on targeted polishing notes that many genome assembly workflows still produce regions with elevated error rates, such as gaps filled with unpolished or ambiguous bases. These problem regions may be a small fraction of the total assembly, yet they can dominate the error count.
Targeted polishing tools such as GoldPolish-Target isolate and polish user-specified assembly loci, offering a resource-efficient means for polishing targeted regions of draft genomes. This approach is useful when the problematic regions are known, such as regions that failed to polish in a previous round or regions that are flagged by quality evaluation tools.
The tradeoff is that targeted polishing requires prior knowledge of which regions need polishing. If the entire assembly has uniformly distributed errors, whole-genome polishing is the appropriate choice. If the errors are concentrated in specific regions, targeted polishing can save substantial computational time and memory.
Self-Correction versus Assembly Polishing
An alternative to polishing the assembly is to correct the reads before assembly. Self-correction methods such as CONSENT use multiple sequence alignment and local de Bruijn graphs to correct errors in the raw reads, and the corrected reads are then used for assembly.
Published comparisons show that self-correction can improve the quality of assemblies produced by tools such as Flye. However, the computational cost of self-correction is substantial, particularly for large genomes. On a human dataset, assembling the raw data and polishing the assembly is less resource consuming than correcting and then assembling the reads, while providing better results.
The practical implication is that assembly polishing is generally preferred over read self-correction for most projects. Self-correction may be useful when the read data are particularly noisy or when the assembly quality is poor even after polishing, but it should not be the default approach.
Combining Multiple Polishing Tools
The best polishing results are often achieved by combining multiple tools. Published benchmarking of short-read polishing tools found that the best results were achieved by using Polypolish in combination with other short-read polishers. Similarly, long-read polishing workflows often combine Racon with Medaka, and some workflows add a short-read polishing step with Pilon.
The order of tools matters. Long-read polishing should be performed first, because it corrects the bulk of the errors and produces an assembly that is close to the final sequence. Short-read polishing should be performed last, because it corrects the residual errors that long-read polishing missed.
The combination of tools should be guided by the quality evaluation. If the QV score plateaus after long-read polishing, adding a short-read polishing step may still improve the assembly by correcting homopolymer errors. If the QV score does not improve after short-read polishing, the short-read data may not be informative for the remaining errors.
Observations and Measurements
Tracking Polishing Progress
The most important measurement in a polishing workflow is the change in assembly quality before and after each polishing round. Merqury provides a QV score that estimates the base accuracy of the assembly, and this score should be recorded for each polishing round.
A typical polishing workflow for an ONT bacterial assembly might produce the following progression:
| Polishing Stage | QV Score | Notes |
|---|---|---|
| Unpolished assembly | 35 | Baseline from initial assembly |
| After Racon round 1 | 42 | Major improvement from indel correction |
| After Racon round 2 | 45 | Smaller improvement from additional correction |
| After Racon round 3 | 45 | No improvement, polishing converged |
| After Medaka | 52 | Neural network model corrects residual errors |
| After Pilon | 55 | Short reads correct homopolymer errors |
The exact values will vary depending on the genome, the sequencing platform, and the assembly tool, but the pattern of diminishing returns is typical. The largest improvement comes from the first polishing round, and subsequent rounds produce smaller gains until the quality plateaus.
Recording Polishing Parameters
For reproducibility, the polishing workflow should be documented with the exact parameters used at each step. The documentation should include:
- The version of each polishing tool
- The version of the aligner and the preset used
- The read files used for each polishing round
- The number of polishing rounds for each tool
- The QV score before and after each round
- The computational time and memory usage for each round
This documentation is essential for reproducing the polishing results and for troubleshooting if the polishing does not produce the expected improvement.
Evaluating Completeness
Base accuracy is not the only quality metric that matters. The polishing step should not remove or truncate contigs, and it should not introduce structural errors. Merqury provides a completeness metric that estimates the fraction of the genome that is represented in the assembly, and this metric should be monitored alongside the QV score.
If the completeness metric decreases after polishing, the polishing step may have introduced errors that caused contigs to be broken or removed. In this case, the polishing parameters should be reviewed and the previous assembly should be retained.
Common Failure Patterns in Polishing
Failure Pattern 1: Polishing Does Not Improve Quality
The most common failure pattern is that the QV score does not improve after a polishing round. This can happen for several reasons:
- The reads used for polishing are not informative for the errors in the assembly. This can occur if the reads are too short, too noisy, or if the error profile of the reads does not match the error profile of the assembly.
- The alignment parameters are incorrect. Using the wrong preset in minimap2 can produce poor alignments that do not provide useful information for the consensus calculation.
- The assembly has already been polished to convergence. If the QV score plateaus, additional polishing rounds will not help.
The response to this failure pattern is to review the alignment parameters, verify that the correct reads are being used, and consider whether a different polishing tool might be more effective.
Failure Pattern 2: Polishing Introduces New Errors
A more serious failure pattern is that the QV score decreases after polishing, indicating that the polishing step introduced new errors. This can happen when:
- The polishing tool is applied to data that do not match its training distribution. For example, applying a Medaka model trained on a specific basecalling version to reads basecalled with a different version.
- The polishing tool is run for too many rounds, and errors in the input assembly are amplified in subsequent rounds.
- The reads used for polishing contain systematic errors that are propagated to the consensus.
The response to this failure pattern is to stop polishing immediately and revert to the previous assembly. The polishing parameters should be reviewed, and the workflow should be adjusted to prevent the error from recurring.
Failure Pattern 3: Polishing Is Too Slow or Uses Too Much Memory
Polishing can be computationally expensive, particularly for large genomes and for tools that use neural network models. If the polishing step takes too long or uses too much memory, the workflow may be impractical.
The response to this failure pattern is to consider targeted polishing instead of whole-genome polishing. Tools such as GoldPolish-Target can polish user-specified assembly loci with substantially lower computational cost. Alternatively, the assembly can be split into smaller chunks for polishing, with the results merged afterward.
Failure Pattern 4: Short-Read Polishing Misses Errors in Repeats
Short-read polishing tools that rely on standard alignment methods are unreliable in repeat regions. Errors in these regions often remain after short-read polishing, and they may be the dominant source of residual errors in the final assembly.
The response to this failure pattern is to use a repeat-aware short-read polisher such as Polypolish, which uses all-per-read alignments to repair errors in repeat sequences. Published benchmarking shows that Polypolish performed well in tests using both simulated and real reads, and it almost never introduced errors during polishing.
Limitations of Polishing
Polishing Cannot Fix Structural Errors
Polishing corrects base errors, but it cannot fix structural errors in the assembly. If the assembler produced an incorrect contig order, a misassembly at a repeat boundary, or a missing region, polishing will not correct these problems. Structural errors must be addressed by reassembly or by using additional data such as optical maps or Hi-C contact maps.
The distinction between base errors and structural errors is important for interpreting polishing results. A high QV score indicates that the bases in the assembly are accurate, but it does not indicate that the assembly is structurally correct. Both types of errors must be evaluated for a complete assessment of assembly quality.
Polishing Cannot Recover Missing Sequence
If a region of the genome is not represented in the assembly, polishing cannot recover it. The polishing tools work by aligning reads to the assembly, and reads that do not align to any contig are not used in the consensus calculation. Missing sequence must be addressed by gap closing or by incorporating additional sequencing data.
Polishing Is Limited by Read Quality
The accuracy of the polished assembly is limited by the accuracy of the reads used for polishing. If the reads contain systematic errors, such as base modifications that are miscalled by the basecaller, the polishing tool may propagate these errors to the consensus. The published error rate of long reads is currently capped around 10 percent, and while polishing reduces this error rate substantially, it cannot eliminate all errors.
Polishing Tools Have Platform-Specific Assumptions
Polishing tools are often trained or optimized for specific sequencing platforms. Medaka is trained on ONT data, and its accuracy on PacBio data is not established. Similarly, tools that are optimized for PacBio data may not perform optimally on ONT data. The tool selection should match the sequencing platform used for the project.
Safety and Reproducibility Context
Reproducibility Standards
Reproducibility is a core requirement for bioinformatics workflows. The polishing step should be documented with sufficient detail that another researcher can reproduce the results. This documentation should include the software versions, the parameters used, and the input files.
Community resources provide guidance on reproducible analysis practices. The Galaxy Training Network offers accessible workflow training and analysis tutorials that emphasize reproducibility. The nf-core documentation describes community pipeline standards for usage, configuration, and reproducible workflow context. Bioconductor provides official package, workflow, installation, and reproducible genomic-analysis documentation. These resources can help researchers structure their polishing workflows for reproducibility.
Data Management
The reads used for polishing should be archived and associated with the assembly in a way that allows the polishing step to be reproduced. The NCBI provides official descriptions of databases, search systems, sequence resources, and analysis services that can be used for data archiving and retrieval. The EMBL-EBI Training resources provide bioinformatics learning pathways and data-resource training that cover data management best practices.
Professional Escalation Criteria
The polishing workflow has limits, and there are situations where the appropriate response is to escalate the problem instead of to continue polishing. These situations include:
- The QV score does not improve after multiple polishing rounds with different tools. This suggests that the assembly has fundamental problems that polishing cannot address.
- The completeness metric decreases after polishing. This suggests that the polishing step is damaging the assembly.
- The polishing step introduces errors that cannot be corrected by subsequent rounds. This suggests that the reads or the alignment parameters are problematic.
- The assembly contains structural errors that are detected by other quality evaluation methods. Polishing cannot fix these errors, and reassembly may be necessary.
In these situations, the appropriate response is to consult with colleagues who have experience with the specific sequencing platform and assembly tools, or to seek guidance from community resources such as the Galaxy Training Network or the Carpentries lessons on foundational computing and data skills.
Decision Framework for Polishing Tool Selection
Define the Error Profile Before Choosing a Tool
The first decision in any polishing workflow is not which tool to run, but which error type dominates the assembly. Published analyses of long-read assemblies consistently show that errors are predominantly indels, with homopolymer-length errors being the most common residual error type in ONT assemblies. PacBio assemblies tend to have a different error distribution, with fewer homopolymer errors but potentially more mismatches in specific sequence contexts. Before selecting a polishing tool, quantify the error profile of the draft assembly by aligning a sample of reads back to the contigs and categorizing the mismatches, insertions, and deletions. This diagnostic step takes less than an hour for most bacterial genomes and provides the evidence needed to choose between a general-purpose polisher and a specialized tool.
The error profile also determines whether short-read polishing will be effective. If the dominant error type is homopolymer-length variation, short-read polishing with a repeat-aware tool is likely to help. If the dominant error type is mismatches in complex or repetitive regions, short-read alignment may be unreliable and the polishing effort should focus on long-read tools with region-aware strategies. The error profile should be recorded in the project documentation alongside the assembly statistics, because it explains why specific polishing decisions were made and provides a baseline for comparing future polishing runs.
Match the Tool to the Region Complexity
The error distribution across an assembly is not uniform. Published work on BlockPolish demonstrates that there are fundamental differences between the error distributions of different regions, and that treating all regions equally during polishing leaves systematic errors uncorrected. Low-complexity regions, homopolymer runs, and tandem repeats accumulate errors differently than unique sequence, and the optimal polishing strategy differs by region type.
For assemblies where the problematic regions are known, targeted polishing offers a resource-efficient alternative to whole-genome polishing. GoldPolish-Target isolates and polishes user-specified assembly loci, and published benchmarks using Drosophila melanogaster and Homo sapiens datasets show that it can reduce indel and mismatch errors by up to 49.2 percent and 55.4 percent respectively, achieving base accuracy values upwards of 99.9 percent. The computational savings are substantial, with up to 27-fold shorter run times and 95 percent less memory consumption on average compared to Medaka. This approach is particularly valuable when a previous polishing round has identified specific regions that remain error-prone, or when the assembly contains gaps filled with unpolished or ambiguous bases.
The decision between whole-genome and targeted polishing should be based on the distribution of errors in the specific assembly. If errors are spread uniformly across the contigs, whole-genome polishing is appropriate. If errors are concentrated in known regions, targeted polishing saves time and memory without sacrificing accuracy. The error profile diagnostic from the previous step provides the evidence for this decision.
Sequence the Polishing Order by Platform and Data Availability
The order of polishing tools should follow a logical progression from general to specific. For ONT assemblies, the standard sequence is Racon for the initial correction pass, followed by Medaka for the neural network-based refinement. For PacBio assemblies, Racon is often sufficient, and a short-read polishing step can be added if Illumina data are available. The rationale for this order is that Racon corrects the bulk of errors quickly and efficiently, while Medaka applies a platform-specific model that performs best when the input assembly is already close to the final sequence.
The availability of short-read data changes the workflow. If short reads exist for the same sample, they should be retained for the final polishing step. Published benchmarking of short-read polishing tools found that the best results were achieved by using Polypolish in combination with other short-read polishers, and that Polypolish almost never introduced errors during polishing. The combination of long-read polishing followed by repeat-aware short-read polishing addresses both the bulk errors and the residual homopolymer errors that long-read tools miss.
Apply the Decision Framework in Practice
The following decision sequence applies the framework to a concrete polishing project:
- Run the error profile diagnostic on the draft assembly to categorize residual errors by type and region.
- If errors are predominantly indels in homopolymer regions and the assembly is from ONT data, start with two rounds of Racon followed by one round of Medaka.
- If errors are predominantly mismatches in complex regions, consider a region-aware tool such as BlockPolish, which divides contigs into blocks by complexity and applies different consensus strategies to each block type.
- If short-read data are available, add a repeat-aware short-read polishing step with Polypolish after the long-read polishing is complete.
- If the error profile shows that problems are concentrated in specific loci, use GoldPolish-Target to polish only those regions instead of running another whole-genome polishing round.
- Evaluate the assembly with Merqury after each step and record the QV score to track whether the polishing order is producing the expected improvements.
Record the Decision Rationale
The polishing decision framework should be documented in the project records alongside the assembly and polishing parameters. The documentation should include the error profile diagnostic results, the rationale for each tool selection, the order of polishing steps, and the QV score after each step. This record serves two purposes. First, it allows another researcher to understand why specific polishing decisions were made and to reproduce the workflow. Second, it provides a basis for troubleshooting if the polishing results are unexpected. If a later analysis reveals errors in the polished assembly, the decision record shows which tools were applied and in what order, allowing the workflow to be revised based on evidence instead of guesswork.
The decision framework also supports professional escalation. If the error profile diagnostic reveals that errors are concentrated in regions that no polishing tool can address, or if the QV score plateaus below the threshold required for the intended downstream analysis, the appropriate response is to escalate the problem to colleagues with experience in the specific sequencing platform or to consult community resources such as the Galaxy Training Network for alternative approaches. The decision record provides the evidence needed for this consultation.
Compare the Framework to Default Workflows
The decision framework differs from a default workflow in one important respect: it requires an initial diagnostic step before any polishing tool is run. A default workflow applies Racon, then Medaka, then Pilon without first characterizing the errors. This approach works for many assemblies, but it can waste computational time on tools that do not address the dominant error type, and it can miss errors that require a specialized tool. The diagnostic step adds a small amount of time to the workflow but provides the evidence needed to make informed decisions about tool selection and order.
The framework also emphasizes the distinction between whole-genome and targeted polishing. Default workflows polish the entire assembly, even when the errors are concentrated in a small fraction of the sequence. Targeted polishing with GoldPolish-Target offers a computationally light-weight and highly scalable solution for base error correction, and it should be considered whenever the error profile shows regional concentration. The published benchmarks demonstrate that targeted polishing achieves accuracy comparable to Medaka while using substantially less computational resources, making it a practical option for large genomes where whole-genome polishing with neural network tools is prohibitively expensive.
Frequently Asked Questions
What is the difference between Racon and Medaka?
Racon is a fast polishing tool that computes a consensus sequence from read alignments using a partial order alignment algorithm. It works with both PacBio and ONT data and is typically used as the first polishing step. Medaka is a neural network-based polishing tool developed for ONT data. It uses a model trained on ONT-specific error profiles to predict the correct consensus sequence. Medaka generally produces higher accuracy than Racon on ONT data but requires more computational resources. The standard workflow for ONT assemblies is to run Racon first, then Medaka.
When should I use Pilon for polishing?
Pilon should be used when short-read data are available for the same sample. It corrects residual errors that remain after long-read polishing, particularly homopolymer-length errors. Pilon aligns short reads to the assembly and uses the alignments to identify and correct errors. It is most effective when used after long-read polishing has already corrected the bulk of the errors. Pilon is less reliable in repeat regions, where short-read alignment is ambiguous.
How many rounds of polishing should I run?
The number of polishing rounds should be guided by the quality evaluation. Run two to three rounds of Racon, then one round of Medaka for ONT assemblies, then one round of Pilon if short reads are available. After each round, evaluate the assembly with Merqury. If the QV score stops improving, further polishing is unlikely to help. If the QV score decreases, stop polishing and revert to the previous assembly.
What is Merqury and how do I use it?
Merqury is a tool that evaluates assembly quality by comparing the assembly to a k-mer spectrum derived from the sequencing reads. It produces a QV score that estimates the base accuracy of the assembly, along with other metrics such as completeness and consistency. To use Merqury, build a k-mer database from the sequencing reads, then run Merqury with the assembly and the k-mer database as inputs. The QV score should be recorded before and after each polishing round to track improvement.
Can I polish an assembly with reads from a different platform?
Yes, but the results may not be optimal. Polishing tools are often trained or optimized for specific sequencing platforms. Medaka is trained on ONT data, and its accuracy on PacBio data is not established. Racon works with any read type, but the alignment parameters must match the read type. If you have reads from multiple platforms, the best approach is to use each platform's reads with the appropriate polishing tool.
What should I do if polishing does not improve my assembly?
If the QV score does not improve after a polishing round, review the alignment parameters and verify that the correct reads are being used. Check that the minimap2 preset matches the read type. If the parameters are correct and the QV score still does not improve, the assembly may have already been polished to convergence, or the reads may not be informative for the remaining errors. Consider trying a different polishing tool or a targeted polishing approach.
How do I know if my assembly is ready for downstream analysis?
An assembly is ready for downstream analysis when the QV score is acceptable for the intended use and the completeness metric indicates that the expected fraction of the genome is represented. For bacterial genomes, a QV score above 50 is generally acceptable. For larger genomes, a QV score above 40 may be acceptable depending on the analysis. The assembly should also be evaluated for structural errors using other methods, because polishing does not fix structural problems.
What are the computational requirements for polishing?
The computational requirements vary by tool and genome size. Racon is fast and memory-efficient, making it suitable for large genomes. Medaka is more computationally intensive and may require a GPU for large genomes. Pilon requires short-read alignment, which is also computationally intensive. For large genomes, targeted polishing tools such as GoldPolish-Target can reduce computational requirements by polishing only the regions that need correction.
Related Bioinformatics Guides
- Long-Read Metagenome Assembly: Overcoming Challenges with Nanopore and PacBio Data
- How to Choose a Long-Read Sequencing Platform: PacBio vs Oxford Nanopore
- Evaluating Metagenomic Assembly Tools: A Benchmarking Framework for Short-Read and Long-Read Data
- Genomic Data Analysis Tools: A Comparative Guide for Researchers
- De Novo Genome Assembly with Long Reads: A Practical Workflow
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- GoldPolish-target: targeted long-read genome assembly polishing.. BMC bioinformatics, 2025.
- BlockPolish: accurate polishing of long-read assembly via block divide-and-conquer.. Briefings in bioinformatics, 2022.
- Polypolish: Short-read polishing of long-read bacterial genome assemblies.. PLoS computational biology, 2022.
- Scalable long read self-correction and assembly polishing with multiple sequence alignment.. Scientific reports, 2021.
- Comparison of long-read methods for sequencing and assembly of a plant genome.. GigaScience, 2020.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.