Assembly Polishing: When and How to Use Racon, Medaka, and Pilon for Accurate Genomes

By Dr. Zubair Khalid, DVM, MS, PhD ·

Assembly Polishing: When and How to Use Racon, Medaka, and Pilon for Accurate Genomes

Key Takeaways

  • Assembly polishing corrects base-level errors (SNPs, indels) in draft genomes by aligning sequencing reads to contigs, crucial for accurate downstream analyses like gene prediction and variant calling.
  • Racon offers platform-agnostic polishing using partial order alignment, excelling in speed and low memory usage but dependent on read mapping quality.
  • Medaka is specifically designed for Oxford Nanopore Technologies (ONT) data, employing a neural network to effectively correct homopolymer errors characteristic of ONT sequencing.
  • Pilon utilizes short-read alignments (e.g., Illumina) to correct errors, fill gaps, and identify local misassemblies, providing broader assembly improvement beyond base correction.
  • The choice of polishing tool hinges on sequencing platform (ONT, PacBio, Illumina), read depth (e.g., >35x for potential polishing necessity), computational resources (Medaka may require GPU), and the error profile of the initial assembly.
  • Polishing is an iterative process, but overpolishing can degrade quality; assess assembly metrics (e.g., mismatches/indel rate, gene completeness via BUSCO) after each round to establish stopping criteria.

Genome assembly polishing is the process of correcting base-level errors in a draft assembly by aligning sequencing reads back to the contigs and using the alignment information to identify and fix mistakes. The three most widely used polishing tools are Racon, Medaka, and Pilon, each with distinct algorithms, input requirements, and performance characteristics. Racon uses partial order alignment to polish assemblies with either long or short reads, Medaka uses a neural network model trained specifically for Oxford Nanopore Technologies (ONT) data, and Pilon uses short-read alignments to correct errors, fill gaps, and identify local assembly problems. The choice among these tools depends on your sequencing platform, read depth, computational resources, and the error profile of your initial assembly. This article provides a practical comparison of these tools, a decision framework for selecting among them, and concrete guidance for implementing polishing workflows that produce accurate genomes for downstream analysis.

Understanding the Polishing Problem in Genome Assembly

Draft genome assemblies contain systematic errors that arise from the limitations of sequencing technologies and the algorithms used to construct contigs. These errors include single nucleotide polymorphisms (SNPs), insertions and deletions (indels), and larger structural artifacts. The error profile varies substantially between sequencing platforms. Short-read technologies such as Illumina produce highly accurate individual reads but struggle with repetitive regions and produce fragmented assemblies. Long-read technologies such as Oxford Nanopore and Pacific Biosciences produce longer contiguous sequences but have higher per-base error rates, particularly in homopolymeric regions where the same nucleotide repeats multiple times.

The practical consequence of assembly errors is that downstream analyses such as gene prediction, variant calling, phylogenetic inference, and comparative genomics can produce misleading results. A single erroneous base within a coding sequence can introduce a premature stop codon, shift the reading frame, or alter an amino acid, leading to incorrect functional annotations. For bacterial outbreak investigations, errors in assemblies can affect genotyping results and the detection of genetic markers, which demand high sequencing accuracy and precise genome assembly for reliable results. The NCBI Data Resources provide reference sequences and analysis tools that researchers use to evaluate assembly quality, but the accuracy of any comparison depends on the quality of the assembly being evaluated.

Polishing addresses these errors by leveraging the redundancy of sequencing data. When multiple reads cover the same genomic position, the consensus of those reads provides a more accurate estimate of the true base than any single read. The challenge is that different polishing tools implement this consensus calculation differently, and the optimal approach depends on the characteristics of your data.

Core Principles of Assembly Polishing

Error Types and Their Origins

Assembly errors fall into several categories that matter for polishing decisions. Substitution errors occur when one base is incorrectly replaced by another. Insertion and deletion errors occur when extra bases are added or existing bases are removed. In long-read data, indels are particularly common in homopolymeric regions because the sequencing signal becomes ambiguous when the same nucleotide is repeated. Methylation can also contribute to errors in ONT assemblies, as modified bases produce different signal patterns that basecallers may interpret incorrectly. One study of ONT assemblies found that methylation caused 6.5% of observed errors, and 81% of errors were located within coding sequences, which underscores the biological significance of polishing for accurate gene prediction.

The error rate of the initial assembly depends on the assembler used and the sequencing chemistry. Modern ONT R10.4.1 flow cells with enhanced basecalling models have generally improved assembly accuracy compared to earlier chemistries. However, the improvement is not uniform across all species. For example, one study of highly pathogenic bacteria found that while perfect genomes were obtained for several species including Klebsiella variicola, Listeria spp., Mycobacterium tuberculosis, Staphylococcus aureus, and Streptococcus pyogenes, other species such as Brucella abortus showed higher accuracy with older basecalling models. This variability means that polishing decisions should be informed by the specific characteristics of your organism and data instead of assumed from platform specifications.

The Role of Read Depth

Read depth, also called coverage, is the average number of reads that align to each position in the genome. Higher coverage provides more evidence for consensus calling and generally improves polishing accuracy, but the relationship is not linear. Beyond a certain coverage threshold, additional reads provide diminishing returns and may even introduce errors if the read set contains systematic biases.

A study of Colletotrichum lini fungal genomes found that at ONT coverage of 35x or higher with corrected reads, additional polishing did not improve assembly accuracy, even when Illumina data was used for polishing. This finding suggests that for high-quality ONT data, there is a coverage threshold beyond which polishing becomes unnecessary. For lower coverage data, polishing can provide substantial improvements, but the results depend on the quality of the reads available for polishing.

For hybrid approaches that combine long-read assemblies with short-read polishing, the depth of the short-read data matters as well. One study of Streptococcus pneumoniae found that long-read assembly followed by short-read polishing was a fast and reliable approach at ONT sequencing depth greater than 100x. At ONT depths below 50x, tools that perform short-read-first assembly such as Unicycler were recommended instead. These thresholds provide practical guidance for deciding whether to polish with short reads or to use a different assembly strategy altogether.

Polishing as an Iterative Process

Polishing is often performed in multiple rounds, with each round using the output of the previous round as input. The rationale is that correcting some errors may reveal additional errors that were previously masked. However, the evidence for iterative polishing is mixed. The bacterial pathogen study found that long-read polishing mainly improves assembly quality with only one round needed, and that additional polishing rounds may actually degrade assembly quality. This degradation can occur when the polishing tool introduces new errors while attempting to correct existing ones, particularly in regions with low coverage or complex repeat structures.

The practical implication is that polishing should be treated as a bounded process with clear stopping criteria instead of an open-ended iterative loop. After each polishing round, assembly quality should be assessed using metrics such as the number of mismatches and indels per 100 kilobases, the completeness of expected genes, and the contiguity of the assembly. If a polishing round does not improve these metrics, additional rounds are unlikely to help and may cause harm.

At a Glance: Polishing Tool Comparison

ToolInput ReadsPrimary PlatformAlgorithmStrengthsLimitations
RaconLong or short readsONT, PacBio, IlluminaPartial order alignmentPlatform agnostic, fast, low memory usageRequires accurate read mapping, may not correct all error types
MedakaONT readsOxford NanoporeNeural network consensusTrained specifically for ONT error profiles, handles homopolymers wellRequires ONT data, needs GPU for reasonable speed on large genomes
PilonShort readsIlluminaVariant detection and consensusCorrects SNPs and indels, identifies local assembly issues, fills gapsRequires high-quality short reads, not designed for long-read error correction

The table above summarizes the key differences among the three tools. Racon is the most flexible option because it can work with any read type, but its accuracy depends on the quality of the read mapping. Medaka is specifically designed for ONT data and uses a neural network that has been trained on the error patterns of ONT sequencing, making it particularly effective for homopolymeric regions. Pilon is designed for short-read data and provides additional functionality beyond error correction, including gap filling and the identification of misassemblies.

Practical Workflow for Assembly Polishing

Step 1: Assess Your Input Data

Before selecting a polishing tool, you need to characterize your sequencing data and initial assembly. Start by determining the sequencing platform and chemistry used. For ONT data, note whether you used R9.4.1 or R10.4.1 flow cells and which basecalling model was applied. For Illumina data, note the read length and whether paired-end sequencing was used. These details affect the expected error profile and the choice of polishing tool.

Next, calculate the read depth for both the long-read and short-read data if you have both. This calculation can be done by aligning reads to the assembly and counting the average number of reads covering each position. Tools such as SAMtools provide depth calculations from alignment files. The coverage values will inform whether polishing is likely to be beneficial and which tool is appropriate.

Finally, evaluate the quality of your initial assembly using metrics such as N50, the number of contigs, and the total assembly size. Compare these metrics to the expected genome size for your organism. Large discrepancies may indicate contamination, misassembly, or incomplete assembly, which polishing cannot fix. The Galaxy Training Network provides accessible tutorials on assembly quality assessment that can guide this evaluation.

Step 2: Select the Polishing Strategy

The decision tree for polishing tool selection depends on your data types and coverage. If you have ONT data with coverage of 35x or higher and used R10.4.1 chemistry with current basecalling models, your assembly may already be accurate enough that polishing provides no benefit. You can verify this by running one round of polishing and comparing the quality metrics before and after.

If you have ONT data with lower coverage or older chemistry, Medaka is the recommended first choice because it is specifically trained for ONT error profiles. Racon can be used as an alternative or in combination with Medaka, particularly if you want to avoid the computational requirements of neural network inference. One study of monkeypox virus sequencing found that the combined use of Canu with Medaka and Homopolish was the best combination for reducing errors in homopolymeric regions without adding mismatches.

If you have both long-read and short-read data, you have two options. The first is to polish the long-read assembly with short reads using Pilon. This approach is effective when the long-read coverage is high enough to produce a contiguous assembly and the short-read coverage is sufficient to correct residual errors. The second option is to use a hybrid assembler such as Unicycler that incorporates both data types during assembly instead of as a post-assembly polishing step. The Streptococcus pneumoniae study found that this approach is recommended when ONT coverage is below 50x.

Step 3: Prepare Read Alignments

All three polishing tools require reads to be aligned to the assembly before polishing can proceed. The alignment step is critical because errors in the alignment will propagate to the polishing step. For Racon and Medaka, long reads are typically aligned using minimap2, which is designed for the high error rates of long-read data. For Pilon, short reads are typically aligned using BWA-MEM or Bowtie2.

The alignment parameters should be appropriate for the read type and the expected error rate. For ONT reads, minimap2 with the map-ont preset is recommended. For PacBio reads, the map-pb preset is appropriate. For Illumina reads, BWA-MEM with default parameters is usually sufficient. After alignment, the resulting SAM or BAM file should be sorted and indexed. For Pilon, the BAM file must be sorted by coordinate and indexed.

It is important to check the alignment statistics before proceeding. The percentage of reads that align, the average alignment identity, and the coverage distribution across the assembly can reveal problems such as contamination, misassembly, or poor read quality. If a substantial fraction of reads fail to align or align with very low identity, the assembly may contain sequences from other organisms or the reads may be of poor quality.

Step 4: Run the Polishing Tool

The specific commands for each tool depend on the version and the input data format. For Racon, the basic workflow is to provide the reads, the overlap file from minimap2, and the assembly sequence. Racon produces a polished assembly as output. Multiple rounds of Racon can be run by using the output of one round as the input to the next, but as noted earlier, additional rounds may not improve quality and can degrade it.

For Medaka, the workflow requires the assembly and the ONT reads. Medaka uses a model that must be selected based on the sequencing chemistry and basecalling model used. The tool aligns the reads to the assembly, performs consensus calling using the neural network, and produces a polished assembly. Medaka can be computationally intensive, particularly for large genomes, and may require a GPU for reasonable runtime.

For Pilon, the workflow requires the assembly FASTA file and a sorted, indexed BAM file of short-read alignments. Pilon analyzes the alignment data to identify discrepancies between the reads and the assembly, then produces a corrected assembly. Pilon also reports changes it made, including base substitutions, indels, and local improvements, which can be reviewed to understand the types of errors that were corrected.

Step 5: Evaluate Polishing Results

After polishing, you must assess whether the polishing improved assembly quality and did not introduce new errors. The primary metrics for this assessment are the number of mismatches and indels per 100 kilobases compared to a reference genome if one is available, the completeness of expected genes, and the overall assembly size and contiguity.

If a reference genome is available for your organism or a close relative, you can align the polished assembly to the reference and calculate the number of differences. This comparison provides a direct measure of accuracy. If no reference is available, you can assess completeness using tools that check for the presence of expected single-copy genes, such as BUSCO or CheckM. These tools compare the assembly against a database of genes that are expected to be present in the organism's taxonomic group.

The EMBL-EBI Training resources provide guidance on genome assembly evaluation and the interpretation of quality metrics. The Bioconductor project offers R packages for genomic analysis that can be used to visualize and compare assemblies before and after polishing.

Tool-Specific Considerations

Racon: Flexible Polishing for Multiple Platforms

Racon is a polishing tool that uses partial order alignment to correct errors in draft assemblies. It accepts reads from any sequencing platform, making it a versatile choice for laboratories that work with multiple technologies. The tool requires an alignment of reads to the assembly, which is typically generated with minimap2 for long reads or with a short-read aligner for Illumina data.

The main advantage of Racon is its speed and low memory usage. It can polish large genomes on a standard workstation without specialized hardware. The main limitation is that its accuracy depends on the quality of the read alignments. If the reads contain systematic errors that the aligner cannot handle, Racon may not correct those errors effectively.

Racon is often used as a first polishing step before applying a more specialized tool such as Medaka. This combination can be effective because Racon corrects the most obvious errors, making the assembly cleaner for Medaka's neural network to process. However, the evidence for the benefit of this combined approach is mixed, and some studies have found that Medaka alone produces equivalent results.

Medaka: ONT-Specific Neural Network Polishing

Medaka is a polishing tool developed by Oxford Nanopore Technologies that uses a neural network to predict the correct consensus sequence from ONT reads aligned to a draft assembly. The neural network is trained on the specific error patterns of ONT sequencing, including the characteristic errors in homopolymeric regions. This training makes Medaka particularly effective for correcting the types of errors that are most common in ONT assemblies.

The main requirement for Medaka is that the input reads must be ONT data. The tool cannot be used with PacBio or Illumina reads. Additionally, Medaka requires the selection of an appropriate model based on the sequencing chemistry and basecalling model used. Using the wrong model can reduce accuracy, so it is important to match the model to your data.

Medaka is computationally intensive compared to Racon. For large genomes, the neural network inference can take substantial time and memory. A GPU can significantly speed up the process, but Medaka can also run on CPU-only systems with longer runtime. The nf-core Documentation describes community pipelines that incorporate Medaka and other polishing tools, providing tested configurations that can simplify the setup process.

Pilon: Short-Read Polishing with Assembly Improvement Features

Pilon is a polishing tool designed for short-read data, typically Illumina. It aligns short reads to the assembly and uses the alignment information to correct base errors, identify and fix local misassemblies, and fill gaps in the assembly. Pilon also reports the changes it makes, which can be useful for understanding the error profile of the original assembly.

The main advantage of Pilon is its ability to do more than simple base correction. It can identify regions where the assembly structure is wrong, such as collapsed repeats or mis-joined contigs, and it can fill gaps using paired-end information. These features make Pilon valuable for improving assembly quality beyond what base-level polishing can achieve.

The main limitation of Pilon is that it requires high-quality short reads. If the short-read data has systematic biases or low coverage, Pilon may introduce errors or fail to correct existing ones. Pilon is not designed to work with long-read data, so it cannot be used to polish assemblies when only ONT or PacBio reads are available.

Hybrid Assembly and Polishing Strategies

Long-Read Assembly with Short-Read Polishing

The most common hybrid strategy is to assemble with long reads to obtain a contiguous draft assembly, then polish with short reads to correct residual errors. This approach leverages the strengths of both technologies: long reads provide the structural framework, and short reads provide accurate base-level information.

The Streptococcus pneumoniae study provides practical guidance for this strategy. At ONT sequencing depth greater than 100x, long-read assembly followed by short-read polishing produced circular and contiguous genomes with high N50 parameters. This approach was fast and reliable. At ONT depth below 50x, the study recommended using short-read-first assembly tools such as Unicycler instead, because the long-read data alone was insufficient to produce a reliable assembly framework.

The coverage of the short-read data also matters. For Pilon to work effectively, the short-read coverage should be sufficient to provide confident consensus calls across the genome. Coverage of 50x or higher is commonly recommended, but the optimal depth depends on the complexity of the genome and the quality of the reads.

Short-Read-First Assembly with Long-Read Polishing

An alternative hybrid strategy is to assemble with short reads first, then use long reads to improve the assembly. This approach is less common because short-read assemblies are typically fragmented, and long reads are better suited to resolving the repeat structures that cause fragmentation. However, for organisms with small genomes or when long-read coverage is limited, this strategy can be effective.

In this approach, the short-read assembly provides the initial contigs, and long reads are aligned to the contigs to identify and correct errors. Racon can be used for this purpose because it accepts any read type. The long reads can also be used to scaffold the short-read contigs, joining them into larger structures.

When Hybrid Polishing Is Not Needed

The Colletotrichum lini study demonstrated that for high-quality ONT data with sufficient coverage, polishing may not be needed at all. The study used ONT R10.4.1 flow cells and corrected the reads with the HERRO algorithm before assembly. At coverage of 35x or higher with the corrected reads, additional polishing did not improve assembly accuracy, even with Illumina data.

This finding has practical implications for cost and time savings. Polishing adds computational time and complexity to the assembly workflow. If your data quality is high enough that polishing provides no benefit, you can skip this step and proceed directly to downstream analysis. The challenge is knowing when your data meets this threshold. The coverage and chemistry of your ONT data, the basecalling model used, and the complexity of your genome all influence whether polishing will help.

Records and Measurements for Polishing Workflows

Documentation Requirements

Reproducible polishing workflows require careful documentation of the tools, versions, parameters, and input data used. This documentation is essential for interpreting results, troubleshooting problems, and publishing methods. At minimum, the following information should be recorded for each polishing run:

  • The version of the polishing tool and the assembler used to create the input assembly
  • The sequencing platform, chemistry, and basecalling model for the reads used in polishing
  • The read alignment tool and parameters used to generate the input alignments
  • The coverage of the reads used for polishing
  • The number of polishing rounds performed
  • The quality metrics of the assembly before and after each polishing round

The The Carpentries Lessons provide training on reproducible research practices, including version control and documentation, that are directly applicable to bioinformatics workflows. The nf-core Documentation describes community standards for pipeline documentation and configuration that can serve as a model for individual workflows.

Quality Metrics to Track

The following metrics should be tracked before and after polishing to evaluate the effectiveness of the process:

  • Number of mismatches per 100 kilobases compared to a reference genome, if available
  • Number of indels per 100 kilobases compared to a reference genome, if available
  • Completeness of expected single-copy genes, as measured by BUSCO or similar tools
  • Total assembly size and number of contigs
  • N50 and L50 values
  • Number of ambiguous bases (Ns) in the assembly

For bacterial genomes, the number of mismatches and indels per 100 kilobases is a standard quality metric. The NCBI Data Resources provide reference genomes and comparison tools that can be used to calculate these values. For eukaryotic genomes, gene completeness is often a more informative metric because reference genomes may not be available for close relatives.

Comparing Polishing Outcomes

To determine whether polishing improved the assembly, compare the quality metrics before and after each polishing round. A polishing round is considered beneficial if it reduces the number of errors without introducing new ones. The bacterial pathogen study found that long-read polishing mainly improves assembly quality with only one round needed, and that additional rounds may degrade quality. This finding suggests that you should stop polishing after the first round unless the quality metrics clearly indicate that additional rounds would help.

When comparing assemblies, it is important to use the same metrics and methods for both the pre-polish and post-polish assemblies. Changes in the evaluation method can confound the comparison and lead to incorrect conclusions about the effectiveness of polishing.

Common Failure Patterns in Polishing Workflows

Overpolishing and Quality Degradation

The most common failure pattern is overpolishing, where multiple rounds of polishing introduce new errors instead of correcting existing ones. This problem occurs because polishing tools can make incorrect corrections when the read evidence is ambiguous or when the tool's model does not match the data. The bacterial pathogen study explicitly found that polishing may degrade assembly quality, and the Colletotrichum study found that polishing provided no benefit beyond a coverage threshold.

To avoid overpolishing, establish clear stopping criteria before you begin. Run one round of polishing, evaluate the quality metrics, and only run additional rounds if the metrics improve. If the metrics do not improve or worsen, stop polishing and use the current assembly.

Incorrect Tool Selection for Data Type

Another common failure is using a polishing tool that is not appropriate for the data type. Pilon requires short-read data and cannot correct long-read errors effectively. Medaka requires ONT data and cannot be used with PacBio or Illumina reads. Racon can work with any read type but may not correct ONT-specific errors as effectively as Medaka.

Before selecting a tool, verify that your data matches the tool's requirements. Check the tool documentation and the version of the tool you are using, as requirements may change between versions. The Bioconductor and Galaxy Training Network resources provide updated documentation and tutorials that can help you select the appropriate tool for your data.

Poor Read Alignment Quality

Polishing tools depend on accurate read alignments. If the reads are not aligned correctly to the assembly, the polishing tool will make incorrect corrections or miss errors entirely. Poor alignment can result from using the wrong aligner for the read type, using incorrect alignment parameters, or having a contaminated or misassembled draft genome.

To diagnose alignment problems, check the alignment statistics before running the polishing tool. The percentage of reads that align, the distribution of alignment identities, and the coverage across the assembly can reveal issues. If a large fraction of reads do not align or align with very low identity, investigate the cause before proceeding with polishing.

Ignoring Coverage Thresholds

The effectiveness of polishing depends on read coverage. At very low coverage, polishing tools do not have enough evidence to make confident corrections. At very high coverage, additional reads provide no benefit and may introduce biases. The Colletotrichum study found that polishing provided no benefit at ONT coverage of 35x or higher with corrected reads, and the Streptococcus study found that different assembly strategies were needed at ONT coverage below 50x compared to above 100x.

Before polishing, calculate the coverage of your reads. If the coverage is very low, consider whether polishing is worthwhile or whether you need additional sequencing. If the coverage is very high, consider whether polishing is necessary at all.

Limitations and Interpretation Constraints

Polishing Cannot Fix Structural Errors

Polishing corrects base-level errors but cannot fix structural errors in the assembly. If the assembler incorrectly joined two regions, collapsed a repeat, or missed a segment, polishing will not correct these problems. Structural errors require different tools and approaches, such as scaffolding, gap closing, or reassembly with different parameters.

The practical implication is that polishing should be performed after you have assessed the structural quality of the assembly. If the assembly has obvious structural problems, such as a much smaller size than expected or a highly fragmented contig set, polishing will not solve these issues. You may need to revisit the assembly step before polishing.

Polishing Results Are Data-Dependent

The effectiveness of polishing depends on the specific characteristics of your data, including the sequencing platform, chemistry, basecalling model, coverage, and the organism being sequenced. Results from one study or one organism may not transfer directly to another. The bacterial pathogen study found that older basecalling models produced higher accuracy for Brucella abortus, which is counterintuitive given that newer models generally improve accuracy. This variability means that you should validate polishing decisions on your own data instead of relying solely on published recommendations.

Reference-Based Evaluation Has Limits

When a reference genome is available, comparing the polished assembly to the reference provides a direct measure of accuracy. However, this comparison has limitations. The reference genome may contain errors, particularly if it was assembled with older technologies. The reference may also differ from your strain due to genuine biological variation, which will appear as mismatches in the comparison even if your assembly is correct.

To interpret reference-based comparisons correctly, you need to distinguish between assembly errors and biological variation. This distinction can be challenging, particularly for organisms with high genetic diversity. The NCBI Data Resources provide tools for comparing genomes and identifying variants, but the interpretation of these comparisons requires biological knowledge of your organism.

Safety and Reproducibility Context

Computational Resource Management

Polishing tools have different computational requirements that affect their practical use. Racon is the least demanding, requiring modest memory and CPU time. Medaka is the most demanding, particularly for large genomes, and may require a GPU for reasonable runtime. Pilon is intermediate, requiring substantial memory for large genomes but no specialized hardware.

Before running a polishing workflow, estimate the computational resources required based on your genome size and read depth. The nf-core Documentation provides guidance on resource estimation and pipeline configuration that can help you plan your computational needs. The Galaxy Training Network offers tutorials on running bioinformatics workflows on shared infrastructure, which can be useful if you do not have access to a dedicated compute cluster.

Reproducibility Standards

Reproducible polishing workflows require fixed tool versions, documented parameters, and recorded input data. Changes in tool versions can alter results, so the exact versions used should be recorded. Containerization tools such as Docker and Singularity can help ensure that the same software environment is used across runs.

The nf-core Documentation describes community standards for reproducible pipelines, including version pinning and containerization. The Bioconductor project provides tools for managing R package versions, which is relevant if you use R for downstream analysis of polished assemblies. The The Carpentries Lessons provide training on version control with Git, which is essential for tracking changes to analysis scripts and workflows.

Data Management

Polishing workflows generate multiple intermediate files, including read alignments, polished assemblies, and quality reports. These files can be large, particularly for eukaryotic genomes with high coverage. A data management plan should specify where these files are stored, how they are named, and how long they are retained.

The EMBL-EBI Training resources provide guidance on data management for bioinformatics projects, including file organization and metadata standards. The NCBI Data Resources provide repositories for depositing final assemblies and raw sequencing data, which is often required by journals and funding agencies.

Professional Escalation Criteria

When to Seek Additional Expertise

Polishing workflows can encounter problems that require specialized expertise to resolve. You should consider escalating to a bioinformatics specialist or a core facility when:

  • The assembly quality metrics do not improve after polishing, or worsen, and you cannot identify the cause
  • The read alignment statistics indicate systematic problems, such as a large fraction of reads failing to align
  • The assembly size is substantially different from the expected genome size, suggesting contamination or misassembly
  • You are working with a novel organism or a complex genome with unusual features such as extreme GC content or extensive repeat structure
  • You need to produce a reference-quality assembly for publication or for a regulated application such as clinical diagnostics

The Galaxy Training Network and EMBL-EBI Training provide learning pathways that can help you build the skills to troubleshoot polishing problems independently. The Bioconductor community provides support forums where you can ask questions about specific analysis problems.

When to Consider Alternative Approaches

Polishing is one step in the assembly workflow, and it may not be the best solution for every problem. You should consider alternative approaches when:

  • The assembly has structural errors that polishing cannot fix, such as collapsed repeats or mis-joined contigs
  • The read coverage is too low for polishing to be effective, and additional sequencing is not feasible
  • The polishing tool consistently introduces new errors, suggesting a mismatch between the tool and the data
  • A different assembly strategy, such as hybrid assembly or a different assembler, might produce a better initial assembly that requires less polishing

The Streptococcus pneumoniae study provides an example of this decision-making. At ONT coverage below 50x, the study recommended using short-read-first assembly tools such as Unicycler instead of polishing a long-read assembly. This recommendation reflects the principle that polishing cannot compensate for insufficient data quality or quantity.

Frequently Asked Questions

What is the difference between assembly polishing and assembly correction?

Assembly polishing and assembly correction are terms that are often used interchangeably, but they can refer to different processes. Polishing typically refers to the post-assembly step of aligning reads to the draft assembly and using the alignments to correct base-level errors. Correction can refer to the pre-assembly step of correcting errors in the reads themselves before assembly, such as the HERRO algorithm used in the Colletotrichum lini study. Both processes aim to improve accuracy, but they operate at different stages of the workflow.

Can I use Racon with Illumina short reads?

Yes, Racon can accept short reads as input. The tool is platform agnostic and can polish assemblies with reads from any sequencing technology. However, for short-read polishing, Pilon is often preferred because it is specifically designed for short-read data and provides additional features such as gap filling and misassembly detection. The choice between Racon and Pilon for short-read polishing depends on your specific needs and the characteristics of your data.

How many rounds of polishing should I perform?

The evidence from published studies suggests that one round of polishing is usually sufficient, and additional rounds may degrade assembly quality. The bacterial pathogen study found that long-read polishing mainly improves assembly quality with only one round needed. The Colletotrichum study found that polishing provided no benefit at high coverage. You should evaluate the quality metrics after each round and stop polishing when the metrics do not improve.

Does polishing work for all genome sizes?

Polishing can be applied to genomes of any size, from small viral genomes to large plant and animal genomes. However, the computational requirements scale with genome size and read depth. For large genomes, Medaka may require substantial computational resources, and Pilon may require large amounts of memory. The practical considerations of polishing large genomes include runtime, memory usage, and storage for intermediate files.

What is the best polishing strategy for ONT-only data?

For ONT-only data, Medaka is generally the recommended polishing tool because it is specifically trained on ONT error profiles. Racon can also be used and may be preferred when computational resources are limited. The monkeypox virus study found that the combined use of Medaka and Homopolish was effective for reducing errors in homopolymeric regions. The optimal strategy depends on your coverage, chemistry, and basecalling model.

How do I know if my assembly needs polishing?

You can assess whether your assembly needs polishing by evaluating its quality metrics. If you have a reference genome available, compare your assembly to the reference and count the number of mismatches and indels. If you do not have a reference, assess gene completeness with tools such as BUSCO. The Colletotrichum study found that at high ONT coverage with corrected reads, polishing provided no benefit, suggesting that some assemblies do not need polishing.

What is the role of read correction before polishing?

Read correction is the process of correcting errors in the reads themselves before assembly or polishing. The Colletotrichum study used the HERRO algorithm to correct ONT reads before assembly, which improved the quality of the initial assembly. Read correction can reduce the burden on polishing by providing cleaner input data. However, read correction adds computational time and may not be necessary for all data types.

Can polishing introduce new errors?

Yes, polishing can introduce new errors, particularly when the read evidence is ambiguous or when the polishing tool's model does not match the data. The bacterial pathogen study explicitly found that polishing may degrade assembly quality. To detect polishing-induced errors, compare the quality metrics before and after polishing and review the changes made by the polishing tool. If polishing introduces errors, consider using a different tool or adjusting the parameters.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.