Polishing with Short Reads: How to Use Pilon to Correct Errors in Long-Read Assemblies
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Pilon is a short-read polishing tool designed to correct base substitutions, small insertions/deletions (indels), and local misassemblies in draft long-read assemblies. It does not extend contigs, close gaps, or resolve structural rearrangements, necessitating prior use of assemblers or scaffolding tools for such tasks.
- Effective Pilon polishing requires high-quality, adapter-trimmed, and low-quality base-trimmed short reads with sufficient depth (e.g., 50-100x for bacterial genomes) to distinguish true errors from sequencing noise. Low coverage regions may remain uncorrected or introduce new errors.
- Alignment of short reads to the draft assembly is a critical prerequisite, typically performed using BWA-MEM or Bowtie2 with parameters that accommodate the expected error rate of the long-read assembly. The resulting BAM file must be sorted and indexed.
- Multiple rounds of Pilon may be necessary to achieve high accuracy, with two to three rounds often required for bacterial genomes to reach reference-like quality. However, excessive polishing (over-polishing) can introduce new errors, particularly in repetitive regions or by correcting true biological variation.
- Persistent errors after Pilon polishing are often indels in homopolymers and repetitive regions where short reads exhibit multi-mapping, leading to ambiguous alignment evidence. Structural errors in the draft assembly, such as misjoins or collapsed repeats, are not corrected by Pilon.
- Pilon is a component of a larger assembly improvement pipeline; for projects requiring high contiguity and accuracy from the outset, considering hybrid assemblers (e.g., Unicycler) that integrate short and long reads during assembly may be more efficient than long-read-only assembly followed by polishing.
Long-read sequencing platforms such as Oxford Nanopore Technologies and Pacific Biosciences produce reads that can span repetitive regions and resolve complex genomic structures, but these reads carry substantially higher error rates than short-read Illumina data. Pilon is a polishing tool that aligns high-quality short reads to a draft long-read assembly and corrects base errors, small indels, and local misassemblies. This article provides a practical protocol for researchers who have a long-read assembly and matching short-read data and need to decide when, how, and how many times to apply Pilon. The guidance covers input preparation, alignment choices, parameter selection, quality assessment, common failure patterns, and the limits of what polishing can fix.
The Role of Pilon in Hybrid Assembly Workflows
Hybrid assembly combines the long contiguity of long reads with the high per-base accuracy of short reads. Long-read-only assemblies often contain errors that matter for downstream biological interpretation. A 2024 benchmarking study of Salmonella enterica serovar Newport outbreak isolates found that Oxford Nanopore assemblies reached approximately 99.95% accuracy, yet even this small error level obscured phylogenetic relationships between closely related isolates. Near-perfect accuracy, defined as roughly five nucleotide errors across a 4.8 Mbp genome, required pipelines that combined both long-read and short-read polishing tools. Pilon was among the short-read polishers that performed similarly to NextPolish, Polypolish, and POLCA in that study. The authors emphasized that polishing order mattered and that using less accurate tools after more accurate ones introduced errors. Indels in homopolymers and repetitive regions where short reads could not be uniquely mapped remained the most challenging errors to correct.
Pilon works by taking a draft assembly in FASTA format, aligning short reads to that assembly, and then using the alignment information to identify discrepancies between the reads and the assembly. It outputs a corrected FASTA file and a changes file that records every edit made. The tool is designed to fix base substitutions, small insertions and deletions, and some local assembly errors. It does not extend contigs, close gaps, or resolve structural rearrangements. Researchers who need those operations must use an assembler or a dedicated scaffolding tool before polishing.
The practical implication is that Pilon is one component of a larger assembly improvement pipeline. It is most effective when the draft assembly is already structurally correct and the remaining errors are small-scale sequence differences. A 2021 study evaluating Pilon and NextPolish on Oxford Nanopore assemblies of bacterial pathogens found that one round of NextPolish produced genome completeness and accuracy parameters similar to reference genomes, whereas two or three rounds of Pilon were needed to reach comparable accuracy. Contiguity did not change after polishing, which confirms that Pilon corrects bases instead of rearranging contigs. The same study found that polishing did not always produce accurate plasmid identification or antimicrobial resistance genotyping, indicating that some biological conclusions remain sensitive to assembly errors even after correction.
Preparing Input Data for Pilon
Short-Read Quality Requirements
Pilon relies on short-read alignments to identify errors. The quality of those alignments depends on the quality of the short reads themselves. Illumina reads with adapter contamination, low-quality tails, or PCR duplicates can produce spurious mismatches that Pilon may interpret as assembly errors. Before running Pilon, trim adapters and low-quality bases from the short reads. Quality trimming tools such as Trimmomatic or fastp are commonly used, and the Galaxy Training Network provides accessible tutorials for these preprocessing steps. The NCBI Data Resources documentation describes the sequence read archive and quality metrics that can help you assess whether your short-read data meet minimum standards.
The depth of short-read coverage matters. Pilon needs enough reads at each position to distinguish true assembly errors from sequencing noise. Low coverage regions will have few supporting reads, and Pilon may leave errors uncorrected or introduce changes based on insufficient evidence. High coverage regions with excessive duplicates can bias the correction toward the dominant read sequence. A typical bacterial genome project might generate 50 to 100-fold coverage of Illumina data, but the optimal depth depends on the genome size, the error rate of the long-read assembly, and the downstream application. The 2024 benchmarking study used a defined set of polishing combinations and did not report a universal coverage threshold, so you should evaluate your own data instead of assume a fixed number.
Long-Read Assembly Quality
Pilon corrects errors in an existing assembly. It does not improve the assembly graph or resolve repeats that were collapsed during assembly. Before polishing, assess the draft assembly with tools such as QUAST, BUSCO, and Merqury. These tools report contiguity statistics, completeness against conserved gene sets, and k-mer-based accuracy estimates. A 2025 benchmarking study of human genome assembly pipelines used these three metrics to evaluate 11 pipelines and found that polishing improved assembly accuracy and continuity, with two rounds of Racon followed by Pilon yielding the best results. The same study noted that the optimal pipeline was implemented on Nextflow to enable parallelization and dependency management, which is relevant if you plan to run polishing at scale.
If the draft assembly has large misjoins, collapsed repeats, or chimeric contigs, Pilon will not fix them. In some cases, polishing can make a structurally incorrect assembly appear more accurate at the base level while leaving the underlying structural errors intact. The 2019 comparison of long-read sequencing technologies in hybrid assembly of complex bacterial genomes found that hybrid assembly with Unicycler was superior to long-read-only assembly with Flye followed by short-read polishing with Pilon, with respect to accuracy and completeness. This result suggests that choosing a hybrid assembler from the start may reduce the need for extensive polishing later. If you already have a long-read-only assembly, Pilon can still improve base accuracy, but you should verify that the assembly structure is correct before investing time in polishing.
Aligning Short Reads to the Long-Read Assembly
Read Alignment Software
Pilon does not align reads itself. You must produce a BAM file of short reads aligned to the draft assembly before running Pilon. The choice of aligner affects the quality of the alignments and therefore the quality of the polishing. BWA-MEM is a common choice for Illumina reads aligned to a reference assembly. Bowtie2 is another option that handles reads with gaps and is often used in variant calling workflows. Both tools are documented in EMBL-EBI Training materials on sequence analysis and in The Carpentries Lessons on bioinformatics workflows.
The aligner must be told that the input is a draft assembly, not a finished reference genome. Draft assemblies contain errors, and reads that span an error site may not align perfectly. Most aligners handle this by allowing mismatches and small gaps, but you should set parameters that permit the expected error rate. For example, if the long-read assembly has an error rate of 1%, the aligner should allow at least that many mismatches per read. If the aligner is too strict, reads will fail to map in error regions, and Pilon will have no evidence to correct those errors.
Alignment Post-Processing
After alignment, the BAM file must be sorted and indexed. Pilon requires a coordinate-sorted BAM file with an associated index. Use samtools to sort the alignments by reference position and to create the index. You should also consider marking or removing duplicate reads. PCR duplicates arise during library preparation and represent the same DNA fragment sequenced multiple times. If duplicates are not removed, they can overrepresent a single fragment and bias the correction. However, duplicate removal is not always necessary for polishing, and some workflows skip it to retain coverage in low-complexity regions. The decision depends on whether your library preparation produced high duplication rates and whether the downstream application requires unbiased allele frequencies.
The nf-core Documentation describes community standards for reproducible bioinformatics pipelines, including alignment and polishing steps. If you are running Pilon as part of a larger pipeline, following these standards can help ensure that your results are reproducible and that others can run the same analysis with the same inputs.
Running Pilon
Basic Command Structure
Pilon is a Java-based tool that takes a FASTA assembly, a BAM file of aligned short reads, and several optional parameters. The basic command specifies the genome file, the BAM file, the output prefix, and the output directory. Pilon writes a corrected FASTA file, a changes file, and a VCF file of variants. The changes file is useful for auditing what Pilon modified and for understanding the error profile of the original assembly.
A typical command might look like this:
java -Xmx16G -jar pilon.jar --genome draft.fasta --frags reads.bam --output polished --outdir pilon_out --changes
The --frags option tells Pilon that the BAM file contains fragment reads, which are short paired-end reads. Pilon also supports --unpaired for unpaired reads and --bam for a general BAM file. The --changes flag writes a detailed report of every change made. The --vcf flag writes a VCF file of the variants Pilon identified. The memory setting -Xmx16G should be adjusted based on the genome size and the depth of coverage. Large genomes and high coverage require more memory.
Parameter Recommendations
Pilon has several parameters that control its behavior. The --fix parameter specifies which types of errors to correct. The default is all, which includes substitutions, insertions, deletions, and local misassemblies. You can restrict Pilon to specific error types with options such as --fix bases or --fix indels. For most applications, the default is appropriate, but if you suspect that the assembly has structural errors, you may want to run Pilon with --fix bases first and then examine the changes file before allowing it to make larger corrections.
The --minmq parameter sets the minimum mapping quality for reads to be used in polishing. Reads with low mapping quality are often multi-mapping reads that align to multiple locations in the assembly. These reads provide ambiguous evidence and can introduce errors if used. The 2024 benchmarking study found that indels in repetitive regions where short reads could not be uniquely mapped remained the most challenging errors to correct. Setting a minimum mapping quality threshold can reduce the influence of multi-mapping reads, but it also reduces coverage in repetitive regions. A threshold of 20 is a common starting point, but you should examine the mapping quality distribution in your BAM file and adjust accordingly.
The --minqual parameter sets the minimum base quality for reads to be used. Low-quality bases at read ends are more likely to be errors. Trimming the reads before alignment reduces the need for a strict base quality threshold, but Pilon can also filter low-quality bases internally. The --mindepth parameter sets the minimum depth of coverage required for Pilon to make a correction. If the depth at a position is below this threshold, Pilon will not change the assembly at that position. This parameter prevents Pilon from making corrections based on insufficient evidence. A value of 5 to 10 is reasonable for bacterial genomes, but you should consider the overall coverage of your data.
Memory and Runtime Considerations
Pilon loads the assembly and the alignment data into memory. The memory requirement scales with the genome size and the number of reads. A bacterial genome with 100-fold coverage of Illumina reads may require several gigabytes of memory. A human genome with similar coverage may require 32 gigabytes or more. The 2025 benchmarking study of human genome assembly pipelines reported computational cost analyses for the pipelines it evaluated, and the authors noted that polishing improved accuracy but required additional compute time. If you are polishing many genomes, consider running Pilon on a server or cluster instead of a laptop.
Runtime also depends on the genome size, coverage, and the number of changes Pilon makes. A bacterial genome can be polished in minutes to hours. A human genome may take several hours to a day. The 2020 NextPolish paper reported that NextPolish outperformed Pilon in speed and correction accuracy for human and Arabidopsis genomes, but Pilon remains a viable option when NextPolish is not available or when the user prefers Pilon's output format.
How Many Rounds of Pilon Are Needed
Diminishing Returns After the First Round
The number of Pilon rounds depends on the error rate of the initial assembly and the desired accuracy. The 2021 study of bacterial pathogen assemblies found that two or three rounds of Pilon were needed to reach genome completeness and accuracy parameters similar to reference genomes, while one round of NextPolish achieved comparable results. This does not mean that Pilon is inferior in all cases. It means that Pilon may require more rounds to converge, and the optimal number depends on the data.
After the first round of Pilon, the assembly should have fewer errors. The second round aligns the same short reads to the corrected assembly and identifies any remaining discrepancies. If the first round corrected most errors, the second round will make fewer changes. You can monitor the changes file to see how many edits Pilon makes in each round. When the number of changes drops to a small fraction of the original, additional rounds are unlikely to improve accuracy and may introduce errors.
Risk of Over-Polishing
Polishing is not a monotonic improvement process. Each round of Pilon can introduce new errors, especially in regions where the short reads are ambiguous. The 2024 benchmarking study found that using less accurate tools after more accurate ones introduced errors, which implies that the order and number of polishing steps matter. If you polish too many times, Pilon may start to correct true sequence variation that is present in the sample but not in the assembly. This is particularly problematic for genomes with high nucleotide diversity, such as bacterial populations or eukaryotic genomes with heterozygous variants.
A practical approach is to run Pilon for two or three rounds and compare the changes files. If the third round makes very few changes, stop. If the third round makes many changes, examine those changes to determine whether they are correcting real errors or introducing new ones. You can also compare the polished assembly to the unpolished assembly using k-mer-based accuracy metrics such as Merqury. If the accuracy does not improve between rounds, additional polishing is not helpful.
At a Glance
| Decision Point | Recommended Approach | Evidence Basis |
|---|---|---|
| Short-read preprocessing | Trim adapters and low-quality bases before alignment | Quality alignments require clean reads, Galaxy Training Network provides preprocessing tutorials |
| Aligner choice | Use BWA-MEM or Bowtie2 with parameters that allow expected assembly errors | Aligners must tolerate mismatches in error regions, EMBL-EBI Training documents alignment workflows |
| Pilon rounds | Start with two rounds, examine changes files, stop when changes are minimal | 2021 study found two to three Pilon rounds reached reference-like accuracy, PubMed 33716184 |
| Mapping quality filter | Set --minmq to 20 or higher to exclude multi-mapping reads | Multi-mapping reads in repeats cause persistent errors, PubMed 38978005 |
| Structural verification | Check assembly with QUAST, BUSCO, and Merqury before polishing | Polishing does not fix misjoins or collapsed repeats, PubMed 40703096 |
| Hybrid assembly alternative | Consider Unicycler hybrid assembly instead of long-read-only plus Pilon | Hybrid assembly outperformed long-read-only plus Pilon in accuracy and completeness, PubMed 31483244 |
Practical Implementation Steps
Step 1: Assess the Draft Assembly
Before running Pilon, evaluate the draft assembly to determine whether polishing is appropriate and what error types are present. Run QUAST to obtain contiguity statistics such as N50 and the number of contigs. Run BUSCO to assess completeness against conserved gene sets. Run Merqury to estimate base accuracy using k-mers from the short reads. These tools are described in the nf-core Documentation and in The Carpentries Lessons on genome assembly evaluation.
If the assembly has many contigs or large gaps, consider whether a hybrid assembler such as Unicycler would produce a better starting point. The 2019 study found that hybrid assembly with either PacBio or ONT reads facilitated high-quality genome reconstruction and was superior to the long-read assembly and polishing approach evaluated. If you must use the existing long-read assembly, proceed to polishing but document the structural limitations.
Step 2: Preprocess Short Reads
Trim adapters and low-quality bases from the short reads. Use a tool such as Trimmomatic or fastp. Check the quality reports before and after trimming to confirm that the data are clean. If the short reads were generated on an Illumina platform, verify that the read length and insert size are appropriate for the aligner you plan to use. The NCBI Data Resources provides documentation on sequence data formats and quality metrics that can help you interpret the quality reports.
Step 3: Align Short Reads to the Assembly
Align the trimmed short reads to the draft assembly using BWA-MEM or Bowtie2. Set parameters that allow mismatches and small gaps consistent with the expected error rate of the assembly. Sort the resulting BAM file by coordinate and index it with samtools. Consider marking duplicates if the library preparation produced high duplication rates. Examine the alignment statistics to confirm that a high fraction of reads mapped and that the coverage is relatively uniform across the assembly.
Step 4: Run Pilon
Run Pilon with the draft assembly and the sorted BAM file. Use the --changes flag to write a detailed report of edits. Set the memory parameter based on the genome size and coverage. Start with the default --fix all setting and a minimum mapping quality of 20. Examine the changes file after the run to understand what Pilon corrected.
Step 5: Evaluate the Polished Assembly
Compare the polished assembly to the draft assembly using QUAST, BUSCO, and Merqury. Check whether the accuracy metrics improved and whether the contiguity remained unchanged. Examine the changes file to see the types of errors that were corrected. If the accuracy improved and the changes are consistent with known error patterns in long-read assemblies, proceed to the next round or stop.
Step 6: Repeat or Stop
Run a second round of Pilon using the polished assembly as the input. Compare the changes file from the second round to the first. If the second round made few changes, stop. If the second round made many changes, run a third round and compare again. The 2021 study found that two or three rounds of Pilon were needed for bacterial genomes, but the exact number depends on the data. Do not polish indefinitely.
Records and Measurements
What to Record
Document every step of the polishing process so that the results are reproducible and auditable. Record the version of Pilon, the aligner, and all associated tools. Record the parameters used for trimming, alignment, and polishing. Record the input file names and checksums. Record the date and the person who ran the analysis. This information is essential if you need to reproduce the results or explain them to collaborators or reviewers.
The nf-core Documentation emphasizes the importance of reproducibility in bioinformatics pipelines. Following community standards for version control, containerization, and workflow management can help ensure that your polishing results are reproducible. If you are working in a laboratory that uses Bioconductor packages for downstream analysis, record the versions of those packages as well.
Metrics to Track
Track the number of changes Pilon makes in each round. The changes file lists every substitution, insertion, and deletion. Count the changes by type to understand the error profile of the assembly. Track the accuracy metrics from QUAST, BUSCO, and Merqury before and after each round. Track the runtime and memory usage for each round to plan future analyses.
The 2024 benchmarking study reported that near-perfect accuracy was defined as approximately five nucleotide errors across a 4.8 Mbp genome, excluding low confidence regions. This threshold is specific to that study and should not be treated as a universal standard. However, it illustrates the level of accuracy that is achievable with combined long-read and short-read polishing. Your target accuracy should be based on the downstream application. Phylogenetic analysis may require higher accuracy than gene annotation, and clinical applications may require higher accuracy than exploratory research.
Common Failure Patterns
Multi-Mapping Reads in Repetitive Regions
The most persistent errors after polishing are indels in homopolymers and repetitive regions where short reads cannot be uniquely mapped. The 2024 benchmarking study identified these as the most challenging errors to correct. Short reads that align to multiple locations in the assembly provide ambiguous evidence, and Pilon may not have enough information to make a confident correction. Setting a minimum mapping quality threshold can reduce the influence of multi-mapping reads, but it also reduces coverage in repetitive regions. If your assembly has many repeats, expect that some errors will remain after polishing.
Over-Polishing and Error Introduction
Each round of Pilon can introduce new errors, especially if the short reads contain systematic biases or if the assembly has regions of low complexity. The 2024 benchmarking study found that using less accurate tools after more accurate ones introduced errors. This principle applies to repeated rounds of the same tool as well. If you polish too many times, Pilon may start to correct true sequence variation or introduce errors based on ambiguous alignments. Monitor the changes file and stop when the number of changes is small.
Structural Errors That Polishing Cannot Fix
Pilon corrects base errors and small indels, but it does not fix misjoins, collapsed repeats, or chimeric contigs. If the draft assembly has structural errors, polishing will not resolve them. The 2019 study found that hybrid assembly with Unicycler was superior to long-read-only assembly with Flye followed by Pilon polishing, with respect to accuracy and completeness. If your assembly has structural problems, consider re-assembling with a hybrid assembler instead of polishing the existing assembly.
Incomplete Correction of Biological Features
Polishing does not guarantee accurate identification of all biological features. The 2021 study found that polished assemblies of Escherichia coli O157:H7, Salmonella Typhimurium, and Cronobacter sakazakii with simulated reads did not provide accurate plasmid identifications. Polishing also failed to provide an accurate antimicrobial resistance genotype for Staphylococcus aureus with real reads. These results indicate that some biological conclusions remain sensitive to assembly errors even after polishing. If your downstream analysis depends on plasmid identification or antimicrobial resistance genotyping, validate the results with additional evidence.
Limitations of Pilon and Polishing
Pilon Does Not Replace Hybrid Assembly
Pilon is a polishing tool, not an assembler. It corrects errors in an existing assembly but does not improve the assembly graph or resolve structural complexity. The 2019 study found that hybrid assembly with Unicycler was superior to long-read-only assembly followed by Pilon polishing, with respect to accuracy and completeness. If you are starting a new project, consider using a hybrid assembler that incorporates short reads during assembly instead of relying on polishing after assembly. The 2025 benchmarking study found that Flye outperformed all assemblers tested, particularly with error-corrected long reads, and that polishing with two rounds of Racon and Pilon yielded the best results. This suggests that the choice of assembler and the polishing strategy are interdependent.
Pilon Is Slower Than Some Alternatives
The 2020 NextPolish paper reported that NextPolish outperformed Pilon in speed and correction accuracy for human and Arabidopsis genomes. If you are polishing large genomes or many samples, NextPolish may be a more efficient choice. However, Pilon remains a viable option, and the 2024 benchmarking study found that Pilon performed similarly to NextPolish, Polypolish, and POLCA among short-read polishers. The choice between Pilon and NextPolish may depend on the specific data, the available compute resources, and the user's familiarity with the tools.
Polishing Does Not Guarantee Biological Accuracy
Polishing improves base accuracy, but it does not guarantee that the assembly is biologically correct. The 2021 study found that polishing did not always produce accurate plasmid identifications or antimicrobial resistance genotypes. These features depend on the assembly structure and the presence of mobile genetic elements, which are difficult to resolve even with hybrid assembly. If your downstream analysis depends on these features, use additional validation methods such as read depth analysis, PCR, or comparison to reference genomes.
Quality Controls and Verification
K-Mer-Based Accuracy Assessment
Merqury uses k-mers from the short reads to estimate the base accuracy of an assembly. It compares the k-mers in the assembly to the k-mers in the short reads and reports the fraction of assembly k-mers that are supported by the reads. This metric is useful for tracking improvement across polishing rounds. The 2025 benchmarking study used Merqury as one of three assessment tools and found that polishing improved assembly accuracy. Run Merqury before and after each round of Pilon to confirm that the accuracy is improving.
Completeness Assessment
BUSCO assesses the completeness of an assembly by searching for conserved single-copy genes. A complete assembly should contain most of these genes in full length. The 2025 benchmarking study used BUSCO as one of three assessment tools. Run BUSCO before and after polishing to confirm that polishing did not disrupt any genes. If BUSCO completeness decreases after polishing, examine the changes file to identify what Pilon modified.
Contiguity Assessment
QUAST reports contiguity statistics such as N50, the number of contigs, and the total assembly length. Polishing should not change contiguity because Pilon does not break or join contigs. The 2021 study confirmed that contiguity remained unchanged after polishing. If the contiguity changes after polishing, there may be a problem with the input files or the Pilon run. Check the logs and the output files to diagnose the issue.
Safety and Reproducibility Context
Data Management
Genome assembly and polishing generate large intermediate files, including BAM files and corrected FASTA files. Store these files in a structured directory with clear naming conventions. Record the versions of all tools and the parameters used. The nf-core Documentation describes community standards for reproducible pipelines, including version control and containerization. Following these standards can help ensure that your results are reproducible and that others can run the same analysis.
Computational Resource Planning
Pilon requires substantial memory and runtime, especially for large genomes. Plan your computational resources before running Pilon. The 2025 benchmarking study reported computational cost analyses for the pipelines it evaluated. If you are polishing many genomes, consider running the analysis on a server or cluster. The Bioconductor project provides documentation on reproducible genomic analysis, and the Galaxy Training Network offers tutorials on running bioinformatics tools in a shared environment.
Professional Escalation Criteria
If the polished assembly still has high error rates, if the changes file shows unexpected patterns, or if the downstream analysis produces inconsistent results, escalate the issue to a bioinformatics specialist or a colleague with experience in genome assembly. The EMBL-EBI Training materials provide learning pathways for bioinformatics analysis, and The Carpentries Lessons offer foundational training in computing and data analysis. Do not assume that additional rounds of polishing will fix a fundamentally flawed assembly. Investigate the root cause of the errors before proceeding.
Decision Framework for Choosing Between Pilon and Alternative Polishing Tools
Researchers often assume Pilon is the default choice for short-read polishing, but the evidence from comparative studies shows that the optimal tool depends on the specific project goals, genome characteristics, and available computational resources. This section provides a practical decision framework based on published benchmarking data, helping you determine when Pilon is the right choice and when alternatives such as NextPolish or hybrid assembly may be more appropriate.
Comparing Pilon with NextPolish
The most direct comparison in the literature comes from the 2020 NextPolish development paper, which reported that NextPolish corrected sequence errors faster and with higher correction accuracy than Pilon for human and Arabidopsis thaliana genomes. This speed advantage matters for large genomes or high-throughput projects where runtime directly affects productivity. However, the 2024 benchmarking study of Salmonella outbreak isolates found that Pilon, NextPolish, Polypolish, and POLCA performed similarly in terms of final accuracy, with NextPolish showing the highest accuracy overall. The practical implication is that Pilon remains competitive for bacterial genomes, while NextPolish may be preferable for larger genomes where runtime becomes a limiting factor.
The 2021 study comparing Pilon and NextPolish on bacterial pathogen assemblies provides additional guidance on round requirements. One round of NextPolish generated genome completeness and accuracy parameters similar to reference genomes, whereas two or three rounds of Pilon were needed to reach comparable accuracy. This difference in convergence speed does not mean Pilon produces worse final results, but it does mean Pilon requires more computational time to reach the same endpoint. If your pipeline processes many genomes, the additional rounds of Pilon may create a meaningful bottleneck.
Decision Criteria Based on Project Goals
For projects where the final accuracy target is near-perfect genome reconstruction, such as outbreak source tracking or phylogenetic analysis of closely related isolates, the 2024 study showed that combined long-read and short-read polishing pipelines achieved approximately 99.9999% accuracy, or roughly five nucleotide errors across a 4.8 Mbp genome. Pilon was part of several of the five best-performing pipelines in that study, typically used after medaka for long-read polishing. If your project requires this level of accuracy, Pilon is a viable component of the polishing pipeline, but you should plan for the full combination of long-read and short-read polishing steps instead of relying on Pilon alone.
For projects where the primary goal is gene annotation or variant discovery in non-repetitive regions, the accuracy improvement from Pilon may be sufficient after two rounds. The 2021 study found that two or three rounds of Pilon produced genome completeness and accuracy parameters similar to reference genomes for bacterial pathogens. However, the same study found that polishing did not always produce accurate plasmid identifications or antimicrobial resistance genotypes, particularly for Staphylococcus aureus with real reads. If your downstream analysis depends on mobile genetic elements or resistance gene content, you should validate the polished assembly with additional evidence regardless of which polishing tool you choose.
Decision Criteria Based on Assembly Strategy
The 2019 comparison of long-read sequencing technologies in hybrid assembly of complex bacterial genomes found that hybrid assembly with Unicycler was superior to long-read-only assembly with Flye followed by Pilon polishing, with respect to accuracy and completeness. This finding suggests that if you are starting a new project and have not yet generated the long-read assembly, you should consider whether a hybrid assembler that incorporates short reads during assembly would produce a better starting point than a long-read-only assembly that requires extensive polishing afterward.
The 2025 benchmarking study of human genome assembly pipelines found that Flye outperformed all assemblers tested, particularly with Ratatosk error-corrected long reads, and that two rounds of Racon followed by Pilon yielded the best polishing results. This study also provided a complete optimal analysis pipeline implemented on Nextflow, which enables efficient parallelization and built-in dependency management. If you are working with human or other large genomes, the choice of assembler and the polishing strategy are interdependent, and you should evaluate the full pipeline instead of selecting Pilon in isolation.
Practical Decision Matrix
Use the following criteria to decide whether Pilon is the appropriate polishing tool for your project. First, consider the genome size. For bacterial genomes under 10 Mbp, Pilon performs comparably to alternatives and the runtime difference is minimal. For eukaryotic genomes above 100 Mbp, NextPolish may be more efficient based on the 2020 benchmarking results. Second, consider the number of samples. If you are polishing more than 50 genomes, the additional rounds required for Pilon may create a significant computational burden. Third, consider the downstream application. For phylogenetic analysis requiring near-perfect accuracy, Pilon is acceptable but should be combined with long-read polishing tools such as medaka. For plasmid identification or antimicrobial resistance genotyping, no polishing tool guarantees accurate results, and you should plan for additional validation.
Record System for Tool Selection Decisions
Document the rationale for your polishing tool choice so that the decision is reproducible and auditable. Record the genome size, the number of samples, the downstream application, and the available computational resources. Record the version of each polishing tool considered and the benchmarking evidence that informed your decision. The nf-core Documentation provides standards for reproducible bioinformatics pipelines that can help you structure this documentation. The Bioconductor project also offers guidance on reproducible genomic analysis workflows.
If you switch from Pilon to NextPolish or another tool mid-project, record the reason for the change and the impact on the results. The 2024 benchmarking study found that the order of polishing tools mattered and that using less accurate tools after more accurate ones introduced errors. This principle applies to switching tools between rounds as well. If you change tools, re-evaluate the assembly accuracy with QUAST, BUSCO, and Merqury to confirm that the switch did not introduce errors.
Escalation Criteria for Tool Performance Issues
If Pilon produces unexpected results, such as a large number of changes in repetitive regions or a decrease in BUSCO completeness after polishing, escalate the issue before proceeding. Check whether the short-read alignments are of sufficient quality and whether the assembly has structural errors that Pilon cannot fix. The EMBL-EBI Training materials provide learning pathways for diagnosing assembly and polishing problems. The Galaxy Training Network offers practical tutorials that can help you troubleshoot common issues.
If the runtime or memory usage of Pilon exceeds your available resources, consider whether NextPolish would be more efficient based on the 2020 benchmarking results. If the assembly still contains errors after three rounds of Pilon, the problem may be structural instead of base-level, and you should consider re-assembling with a hybrid assembler such as Unicycler instead of continuing to polish. The 2019 study provides evidence that hybrid assembly produces more accurate and complete genomes than long-read-only assembly followed by Pilon polishing.
Frequently Asked Questions
What is the difference between Pilon and NextPolish?
Pilon and NextPolish are both short-read polishing tools that correct errors in long-read assemblies. The 2020 NextPolish paper reported that NextPolish corrected sequence errors faster and with higher accuracy than Pilon for human and Arabidopsis genomes. The 2024 benchmarking study found that NextPolish showed the highest accuracy among short-read polishers, but Pilon, Polypolish, and POLCA performed similarly. The choice between the tools may depend on the genome size, the available compute resources, and the user's familiarity with the tools. Both tools require aligned short reads and produce corrected assemblies.
How many rounds of Pilon should I run?
The optimal number of rounds depends on the error rate of the initial assembly and the desired accuracy. The 2021 study found that two or three rounds of Pilon were needed to reach genome completeness and accuracy parameters similar to reference genomes for bacterial pathogens. The 2025 benchmarking study found that two rounds of Racon followed by Pilon yielded the best results for human genome assembly. Monitor the changes file after each round. When the number of changes drops to a small fraction of the original, additional rounds are unlikely to improve accuracy and may introduce errors.
Can Pilon fix structural errors in the assembly?
No. Pilon corrects base substitutions, small insertions and deletions, and some local misassemblies, but it does not fix misjoins, collapsed repeats, or chimeric contigs. The 2019 study found that hybrid assembly with Unicycler was superior to long-read-only assembly followed by Pilon polishing, with respect to accuracy and completeness. If the draft assembly has structural errors, consider re-assembling with a hybrid assembler instead of polishing the existing assembly.
What should I do if the polished assembly still has errors?
First, examine the changes file to understand what Pilon corrected and what it did not. Run QUAST, BUSCO, and Merqury to assess the remaining errors. If the errors are concentrated in repetitive regions, consider whether additional sequencing or a different assembly strategy would help. If the errors affect downstream biological conclusions, such as plasmid identification or antimicrobial resistance genotyping, validate the results with additional evidence. The 2021 study found that polishing did not always produce accurate plasmid identifications or antimicrobial resistance genotypes.
Do I need to remove duplicates before running Pilon?
Duplicate removal is not always necessary for polishing, but it can reduce bias in regions with high duplication rates. PCR duplicates represent the same DNA fragment sequenced multiple times and can overrepresent a single fragment. If your library preparation produced high duplication rates, consider marking or removing duplicates before running Pilon. However, duplicate removal also reduces coverage, which may be problematic in low-complexity regions. The decision depends on your data and the downstream application.
What is the minimum coverage required for Pilon?
There is no universal coverage threshold for Pilon. The optimal depth depends on the genome size, the error rate of the long-read assembly, and the downstream application. Pilon has a --mindepth parameter that sets the minimum depth of coverage required for Pilon to make a correction. A value of 5 to 10 is reasonable for bacterial genomes, but you should consider the overall coverage of your data. Low coverage regions will have fewer supporting reads, and Pilon may leave errors uncorrected or introduce changes based on insufficient evidence.
Can I use Pilon to polish a human genome assembly?
Yes, Pilon can be used to polish human genome assemblies, but the memory and runtime requirements are substantial. The 2025 benchmarking study of human genome assembly pipelines found that polishing improved assembly accuracy and continuity, with two rounds of Racon and Pilon yielding the best results. The study also reported computational cost analyses and provided a complete optimal analysis pipeline implemented on Nextflow. If you are polishing a human genome, plan for sufficient memory and runtime, and consider using a workflow manager to parallelize the analysis.
How do I know if polishing improved the assembly?
Compare the assembly accuracy metrics before and after polishing. Run QUAST, BUSCO, and Merqury on both the draft and polished assemblies. The 2025 benchmarking study used these three metrics to evaluate assembly quality. If the accuracy metrics improved and the contiguity remained unchanged, polishing was successful. Examine the changes file to confirm that the changes are consistent with known error patterns in long-read assemblies. If the accuracy did not improve, additional polishing is unlikely to help.
Related Bioinformatics Guides
- Evaluating Metagenomic Assembly Tools: A Benchmarking Framework for Short-Read and Long-Read Data
- Hybrid Genome Assembly: Combining Short and Long Reads for Better Results
- Long-Read Metagenome Assembly: Overcoming Challenges with Nanopore and PacBio Data
- Long-Read Genome Assembly and Polishing Strategies
- Short-Read vs Long-Read Sequencing: Pros, Cons, and Selection Criteria
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- Benchmarking short and long read polishing tools for nanopore assemblies: achieving near-perfect genomes for outbreak isolates.. BMC genomics, 2024.
- Polishing the Oxford Nanopore long-read assemblies of bacterial pathogens with Illumina short reads to improve genomic analyses.. Genomics, 2021.
- NextPolish: a fast and efficient genome polishing tool for long-read assembly.. Bioinformatics (Oxford, England), 2020.
- Comparison of long-read sequencing technologies in the hybrid assembly of complex bacterial genomes.. Microbial genomics, 2019.
- Benchmarking of bioinformatics tools for the hybrid de novo assembly of human and non-human whole-genome sequencing data.. Computational and structural biotechnology journal, 2025.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.