Benchmarking Polishing Pipelines: How to Design a Fair Comparison of Polishing Strategies
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Standardization is paramount for fair benchmarking: Polishing tool performance is highly dependent on factors like reference genome choice, data type (e.g., ONT vs. PacBio HiFi), read depth, basecalling model, and assembler. Benchmarks must explicitly control these variables to isolate the polishing tool's impact, preventing inflated performance claims and enabling meaningful cross-study comparisons.
- Define the specific question before designing the benchmark: The objective (e.g., maximum accuracy for phylogenetics vs. practical efficiency for routine use) dictates the experimental design, data selection, and error metrics. For instance, outbreak investigations require high-resolution accuracy to distinguish closely related isolates, as demonstrated by the Salmonella study where even minor errors obscured phylogenetic relationships.
- Employ multiple, diverse error metrics and controls: Assembly quality is multidimensional; a single metric is insufficient. Report QV, mismatch rates, indel rates, and assembly contiguity (e.g., using BUSCO). Crucially, include a "no-polishing" control to quantify improvement and identify instances where polishing introduces errors, as seen in studies where tool order impacted accuracy.
- Control for confounding variables and use appropriate ground truth: Factors such as the assembler used (e.g., Flye, NextDenovo), sequencing platform (ONT vs. PacBio HiFi), and read depth (e.g., >80x for ONT, >20x for HiFi) significantly influence polishing outcomes. Complete, closed genomes serve as ideal references for unambiguous error identification, while simulated data offers controlled ground truth but may not fully capture real-world sequencing artifacts.
- Document meticulously and ensure reproducibility: Detailed records of sequencing platform, basecalling models (e.g., dorado [email protected]), assembler versions, polishing tool versions, and parameter settings are essential. Making analysis code and data publicly available (e.g., via GitHub, Zenodo) allows for verification and extension of benchmarking results.
- Consider application-specific requirements and computational efficiency: Benchmarking should align with the downstream application; for example, cgMLST-based minimum spanning trees were used to confirm concordant results between nanopore-only and Illumina assemblies in a Pseudomonas study. Computational resource usage (runtime, memory) should be reported alongside accuracy metrics, as a tool's practicality is as important as its precision.
Genome assembly polishing is the process of correcting residual base errors in draft assemblies using additional sequencing data. Researchers developing or evaluating polishing methods face a common problem: how to compare their pipeline against existing tools in a way that is fair, reproducible, and informative. A fair benchmark requires explicit decisions about reference genomes, data type, error metrics, and statistical validation before any tool is run. This framework provides concrete steps for designing such comparisons, with attention to the specific challenges of bacterial, fungal, and larger eukaryotic genomes.
Scope and Reader Context
This framework applies to researchers who have developed a new polishing tool or pipeline and need to demonstrate its performance against established methods. It also applies to laboratory professionals who must choose among existing polishing tools for a specific project. The benchmarking principles described here cover bacterial genomes, fungal genomes, and larger eukaryotic genomes, with attention to the different challenges each presents. The focus is on practical comparison design: what data to use, how to measure error, how to control for confounding variables, and how to report results transparently.
The evidence base for this framework draws from published benchmarking studies that tested multiple polishing tools across different sequencing platforms and genome types. These studies provide concrete examples of experimental design, including the number of replicates, the choice of error metrics, and the handling of difficult genomic regions. The framework also incorporates guidance from established bioinformatics training resources and community standards for reproducible analysis.
Why Polishing Benchmarks Fail Without Standardization
Polishing tools are often evaluated under conditions that favor their specific design assumptions. A tool optimized for high-coverage short-read data may perform poorly on low-coverage long-read data, but this limitation may not appear in a benchmark that only tests one data type. Conversely, a tool designed for a particular sequencing platform may show inflated performance when tested only on that platform's data. Without standardized benchmarking, published performance claims are difficult to compare across studies.
The consequences of poor benchmarking extend beyond academic disputes. In clinical and outbreak settings, polishing errors can obscure phylogenetic relationships between closely related isolates. A benchmarking study of Salmonella enterica serovar Newport isolates from a foodborne outbreak found that even small levels of error could affect high-resolution source tracking, and near-perfect accuracy was only achieved when both long-read and short-read polishing tools were combined in the correct order [<a href="#ref-1">1</a>]. This finding demonstrates that polishing performance has real-world consequences for public health investigations.
Another study of Pseudomonas aeruginosa isolates showed that an optimized nanopore-only workflow could achieve results comparable to Illumina short-read assemblies, with as few as two discordant positions per genome at the whole-genome level [<a href="#ref-2">2</a>]. However, this performance required careful optimization of basecalling models, error correction, and polishing parameters. A benchmark that did not account for these optimization steps would produce misleading results.
Core Principles of Fair Benchmarking
Define the Question Before Choosing the Tools
A benchmarking study must begin with a specific question. Are you testing whether your new tool is faster than existing tools at the same accuracy level? Are you testing whether it achieves higher accuracy on a particular data type? Are you testing whether it scales to large genomes? The question determines the experimental design, the choice of data, and the metrics used for evaluation.
For example, a study testing polishing tools for nanopore assemblies of bacterial genomes would need to specify whether the goal is maximum accuracy for phylogenetic analysis or practical efficiency for routine laboratory use. These goals may require different benchmarks. The Salmonella study tested 132 combinations of assembly and polishing tools to assess accuracy for reconstructing outbreak isolate genomes, which required a benchmark design that could distinguish between pipelines with very similar performance [<a href="#ref-1">1</a>].
Control for Confounding Variables
Polishing performance depends on multiple factors: the assembler used to create the draft assembly, the sequencing platform, the read depth, the basecalling model, and the genomic characteristics of the organism. A fair benchmark controls these factors so that differences in performance can be attributed to the polishing tools themselves.
The yeast genome benchmarking study provides an example of systematic control. The researchers performed 455 assemblies of yeast S288C using 7 assemblers with 4 polishing pipelines or no polishing across 13 subsets of different depths, and 88 assemblies using 4 assemblers with or without polishing across 11 subsets [<a href="#ref-3">3</a>]. This design allowed them to separate the effects of assembler choice, polishing tool, and sequencing depth on final assembly quality.
Use Multiple Error Metrics
Assembly quality is multidimensional. A polishing tool may improve base accuracy while introducing structural errors, or it may fix substitutions while leaving indels uncorrected. A fair benchmark reports multiple metrics that capture different aspects of error.
Common metrics include mismatch rate, indel rate, and QV (quality value), which is a logarithmic measure of base accuracy. The Salmonella study reported accuracy as a percentage and also quantified the number of nucleotide errors across the genome, which provides a more intuitive sense of error burden [<a href="#ref-1">1</a>]. The yeast study used QUAST, BUSCO, and a composite score called C_score to evaluate assembly quality from different perspectives [<a href="#ref-3">3</a>].
Include Appropriate Controls
Every benchmark should include a no-polishing control to establish the baseline error rate of the draft assembly. This control is essential for calculating the improvement attributable to polishing. It also helps identify cases where polishing introduces errors instead of correcting them.
The Salmonella study found that the order of polishing tools mattered, and using less accurate tools after more accurate ones introduced errors [<a href="#ref-1">1</a>]. This finding underscores the importance of testing pipelines as complete workflows instead of individual tools in isolation. A benchmark that only tests tools individually may miss interactions that occur when tools are combined.
At a Glance: Benchmark Design Decisions
| Decision Point | Options | Key Considerations |
|---|---|---|
| Reference genome | Complete closed genome, curated reference, or simulated genome | Complete genomes allow unambiguous error identification, simulated genomes allow known ground truth but may not capture real sequencing artifacts |
| Data type | Real sequencing data, simulated reads, or hybrid | Real data reflects actual error profiles, simulated data allows controlled depth and error rates, hybrid approaches combine both |
| Error metrics | QV, mismatch rate, indel rate, BUSCO completeness, alignment-based metrics | No single metric captures all error types, report multiple metrics and consider application-specific requirements |
| Statistical validation | Replicates, bootstrapping, paired comparisons | Single runs cannot distinguish tool performance from random variation, replicate independent assemblies or use paired statistical tests |
Choosing Reference Genomes and Ground Truth
Complete Closed Genomes as Gold Standards
The most straightforward benchmarking approach uses a complete, closed genome as the reference. For bacterial genomes, this means a single circular chromosome with no gaps or ambiguities. The Salmonella study used 15 closely related isolates from a single outbreak, which provided a challenging test case because the genomes were highly similar and small differences mattered for source tracking [<a href="#ref-1">1</a>].
Complete genomes allow unambiguous identification of every error in the polished assembly. Each position in the assembly can be compared directly to the reference, and mismatches, insertions, and deletions can be counted precisely. This approach is limited to organisms with available complete genomes, which currently includes many bacterial species and an increasing number of fungal and eukaryotic species. The NCBI provides databases for sequence data and assembly submissions that can serve as sources for reference genomes [<a href="#ref-4">4</a>].
Simulated Genomes for Controlled Experiments
Simulated reads offer the advantage of known ground truth. When reads are simulated from a reference genome, every error in the resulting assembly can be traced to its source. This allows precise measurement of polishing accuracy and enables controlled experiments that vary read depth, error rate, and error type.
However, simulated data has limitations. Simulators may not capture the complex error profiles of real sequencing platforms, particularly for Oxford Nanopore Technologies (ONT) sequencing, where error patterns depend on base context, methylation, and other factors. A tool that performs well on simulated data may perform differently on real data. The yeast benchmarking study used real ONT and PacBio HiFi datasets, which provided realistic error profiles but required careful handling of the reference genome to avoid bias [<a href="#ref-3">3</a>].
Hybrid Approaches
Some benchmarking studies combine simulated and real data. Simulated data can be used for initial testing and parameter optimization, while real data provides the final validation. This approach balances the control of simulation with the realism of actual sequencing data.
The ntEdit study provides an example of this approach. The researchers tested ntEdit on controlled Escherichia coli and Caenorhabditis elegans sequence data, then validated performance on a sub-20x coverage human genome dataset and on spruce genome data [<a href="#ref-5">5</a>]. The controlled experiments allowed precise measurement of error correction rates, while the real datasets demonstrated scalability and practical utility.
Selecting Sequencing Data for Benchmarking
Read Depth Considerations
Read depth is one of the most important variables in polishing performance. Most polishing tools require sufficient coverage to distinguish true variants from sequencing errors. Low coverage can lead to missed corrections, while very high coverage may not improve performance and can increase computational cost.
The ntEdit study found that the tool performed well at low sequence depths below 20x, fixing the majority of base substitutions and indels, with performance largely constant as coverage increased [<a href="#ref-5">5</a>]. This finding contrasts with tools that require high coverage to work effectively. A fair benchmark should test polishing tools across a range of depths to identify their coverage requirements.
The yeast study provided specific depth recommendations: more than 80x coverage for ONT data and more than 20x for HiFi data were required for high-quality genome construction [<a href="#ref-3">3</a>]. These thresholds reflect the different error rates of the two sequencing platforms and should be considered when designing benchmarks.
Platform-Specific Error Profiles
ONT and PacBio HiFi sequencing have different error profiles, and polishing tools may be optimized for one platform or the other. ONT reads have higher error rates but can be very long, while HiFi reads have lower error rates but are shorter. A fair benchmark tests tools on data from both platforms when possible.
The yeast study tested both ONT and HiFi datasets and found that different assemblers and polishing tools performed best on each platform [<a href="#ref-3">3</a>]. For ONT data, Flye was superior to other assemblers, and polishing with Pilon and Medaka improved accuracy and continuity, respectively. For HiFi data, Flye and NextDenovo performed better, and polishing was still necessary.
Basecalling Model Selection
For ONT data, the basecalling model has a major impact on downstream polishing performance. The Pseudomonas study found that optimal performance required dorado [email protected] basecalling with the inclusion of dorado error correction and dorado polish with its bacterial model [<a href="#ref-2">2</a>]. Different basecalling models produce different error profiles, and polishing tools may perform differently depending on which model was used.
A fair benchmark should specify the basecalling model used and test whether polishing tools are sensitive to this choice. If a polishing tool performs well only with a specific basecalling model, this limitation should be reported.
Error Metrics and Quality Assessment
QV and Accuracy Calculations
QV is a logarithmic measure of base accuracy, calculated as QV = -10 * log10(error_rate). A QV of 50 corresponds to an error rate of 1 in 100,000 bases, or 99.999% accuracy. The Salmonella study reported near-perfect accuracy of 99.9999%, which corresponds to approximately 5 nucleotide errors across a 4.8 Mbp genome [<a href="#ref-1">1</a>].
When reporting QV or accuracy, specify the calculation method. Some methods exclude low-confidence regions, which can inflate reported accuracy. The Salmonella study explicitly noted that their accuracy calculation excluded low-confidence regions [<a href="#ref-1">1</a>], and this caveat should be reported in any benchmark.
Mismatch Rate and Indel Rate
Mismatch rate and indel rate should be reported separately because polishing tools may have different effects on each error type. Some tools are better at correcting substitutions, while others are better at correcting insertions and deletions. The ntEdit study reported substitution and indel rates separately and found that the tool fixed the majority of both error types at low coverage [<a href="#ref-5">5</a>].
Indels in homopolymers and repetitive regions are particularly challenging to correct. The Salmonella study found that indels in these regions, where short reads could not be uniquely mapped, remained the most challenging errors to correct [<a href="#ref-1">1</a>]. A benchmark should report error rates in these difficult regions separately from genome-wide error rates.
Completeness and Contiguity Metrics
Polishing should not come at the expense of assembly completeness or contiguity. BUSCO (Benchmarking Universal Single-Copy Orthologs) scores measure the presence of conserved genes, providing an assessment of assembly completeness. The yeast study used BUSCO in addition to QUAST and C_score to evaluate assembly quality [<a href="#ref-3">3</a>].
A polishing tool that introduces structural errors or breaks contigs may improve base accuracy while degrading overall assembly quality. A fair benchmark reports both base-level accuracy and assembly-level metrics.
Application-Specific Metrics
Some applications require specific accuracy thresholds. For outbreak investigations using core genome multilocus sequence typing (cgMLST), the relevant metric is whether polishing errors affect the assignment of isolates to sequence types. The Pseudomonas study used cgMLST-based minimum spanning trees to compare nanopore-only assemblies with Illumina short-read assemblies and found fully concordant results [<a href="#ref-2">2</a>].
When designing a benchmark, consider the downstream application and include metrics that reflect its requirements. A polishing pipeline that achieves high genome-wide accuracy but introduces errors in specific genes may be unsuitable for applications that depend on those genes.
Designing the Benchmark Workflow
Step 1: Select Reference Genomes
Choose reference genomes that represent the range of challenges your polishing tool is designed to address. For bacterial benchmarks, include genomes with different GC content, repeat content, and homopolymer density. The Pseudomonas study specifically noted that library preparation incorporated temperature ramps to improve performance for high-GC content genomes [<a href="#ref-2">2</a>], indicating that GC content is a relevant variable.
For eukaryotic benchmarks, consider genome size and complexity. The ntEdit study tested performance on human chromosomes and whole genomes, as well as the 20 Gb interior and white spruce genomes [<a href="#ref-5">5</a>]. These tests demonstrated scalability to large genomes but required substantial computational resources.
Step 2: Generate or Obtain Sequencing Data
Decide whether to use real or simulated data, or both. If using real data, ensure that sequencing was performed with appropriate protocols and that basecalling was done with current models. If using simulated data, document the simulator and parameters used.
For benchmarking studies that compare multiple tools, use the same sequencing data for all tools. This controls for variation in sequencing quality and ensures that performance differences are attributable to the tools themselves.
Step 3: Create Draft Assemblies
The choice of assembler affects polishing performance. The yeast study found that the assembler plays an essential role in genome construction, especially for low-depth datasets [<a href="#ref-3">3</a>]. A fair benchmark either uses a single assembler for all polishing tools or tests multiple assemblers in combination with polishing tools.
If the goal is to evaluate polishing tools independently of assembler choice, use a single assembler and report its parameters. If the goal is to evaluate complete pipelines, test multiple assembler-polisher combinations, as was done in the Salmonella study [<a href="#ref-1">1</a>].
Step 4: Run Polishing Tools
Run each polishing tool according to its documentation, using recommended parameters. If a tool requires parameter tuning, document the tuning process and the final parameters used. Avoid using parameters that were optimized on the test data, as this introduces bias.
For tools that require training data, ensure that training and test data are separated. The Pseudomonas study used a set of unrelated clinical isolates for workflow optimization and separate retrospective hospital outbreak isolate sets for validation [<a href="#ref-2">2</a>], which is an appropriate design.
Step 5: Evaluate Results
Align polished assemblies to the reference genome and calculate error metrics. Use the same alignment and evaluation pipeline for all tools to ensure comparability. Report results for each replicate and each metric.
For statistical validation, perform multiple independent assemblies or use bootstrapping to estimate confidence intervals. The yeast study performed hundreds of assemblies across different conditions [<a href="#ref-3">3</a>], which provided robust statistical power for comparing tools.
Records and Measurements for Benchmarking
Documentation Requirements
Every benchmarking study should maintain detailed records that allow others to reproduce the results. This includes:
- Sequencing platform and basecalling model versions
- Read depth and quality metrics for each dataset
- Assembler versions and parameters
- Polishing tool versions and parameters
- Reference genome versions and accessions
- Alignment and evaluation pipeline versions
The nf-core documentation emphasizes the importance of reproducible workflow standards, including version tracking and configuration management [<a href="#ref-6">6</a>]. Following these standards ensures that benchmarking results can be verified and extended by other researchers.
Data Management
Store raw sequencing data, draft assemblies, polished assemblies, and evaluation results in organized directories with clear naming conventions. Use version control for analysis scripts and parameter files. The Carpentries lessons provide foundational training in shell, Git, and programming practices that support reproducible analysis [<a href="#ref-7">7</a>].
For large datasets, consider using established data repositories. The NCBI provides databases for sequence data and assembly submissions [<a href="#ref-4">4</a>], and the EMBL-EBI offers training on data-resource usage and practical analysis education [<a href="#ref-8">8</a>]. Depositing benchmarking data in public repositories allows other researchers to verify results and perform additional analyses.
Computational Resource Tracking
Record computational resource usage for each polishing tool, including runtime, memory, and CPU usage. The ntEdit study reported execution times for each test case, showing that the pipeline executed in less than 14 seconds for E. coli and less than 3 minutes for C. elegans on a single CPU [<a href="#ref-5">5</a>]. This information is valuable for researchers who need to estimate computational requirements for their own projects.
Resource usage should be reported alongside accuracy metrics. A tool that achieves slightly higher accuracy but requires substantially more computational resources may not be the best choice for routine use.
Common Failure Patterns in Polishing Benchmarks
Overfitting to Test Data
A polishing tool that is optimized on the same data used for evaluation will show inflated performance. This problem is particularly acute for tools that use machine learning or require parameter tuning. To avoid overfitting, use separate datasets for training and evaluation, or use cross-validation approaches.
The Pseudomonas study used unrelated clinical isolates for workflow optimization and separate outbreak isolate sets for validation [<a href="#ref-2">2</a>]. This separation ensured that the optimized parameters were not tuned to the validation data.
Ignoring Tool Interactions
Polishing tools are often combined in pipelines, and the order of tools matters. The Salmonella study found that using less accurate tools after more accurate ones introduced errors [<a href="#ref-1">1</a>]. A benchmark that tests tools individually may miss these interactions.
When designing a benchmark, test complete pipelines as well as individual tools. The Salmonella study tested 132 combinations of assembly and polishing tools [<a href="#ref-1">1</a>], which allowed identification of the best-performing pipeline combinations.
Incomplete Error Reporting
Reporting only genome-wide accuracy can hide important limitations. A tool may achieve high overall accuracy while leaving errors in specific regions, such as homopolymers or repetitive elements. The Salmonella study found that indels in homopolymers and repetitive regions remained the most challenging errors to correct [<a href="#ref-1">1</a>].
Report error rates stratified by genomic context, including coding regions, repetitive regions, and homopolymers. This information helps users understand where a polishing tool is likely to succeed and where it may fail.
Inadequate Statistical Power
Comparing polishing tools based on a single assembly per tool cannot distinguish real performance differences from random variation. The yeast study performed hundreds of assemblies across different conditions [<a href="#ref-3">3</a>], which provided the statistical power needed to identify significant differences.
For smaller studies, use paired comparisons or bootstrapping to estimate confidence intervals. Report the number of replicates and the statistical methods used.
Ignoring Computational Efficiency
Accuracy is not the only relevant metric. A tool that achieves high accuracy but requires excessive computational resources may be impractical for routine use. The ntEdit study emphasized scalability, showing that the tool executed in less than 2 hours 20 minutes on human genome assemblies using high-coverage Illumina data [<a href="#ref-5">5</a>].
Report runtime, memory usage, and CPU requirements for each tool. Consider whether the tool scales linearly with genome size or shows superlinear scaling.
Limitations of Benchmarking Approaches
Reference Bias
Using a complete reference genome as ground truth can introduce bias if the reference genome has errors or if the test strain differs from the reference strain. The Salmonella study used 15 closely related isolates from a single outbreak [<a href="#ref-1">1</a>], which minimized reference bias because the isolates were highly similar.
For organisms with high genetic diversity, consider using a reference-free evaluation approach or using multiple reference genomes. The yeast study used the S288C reference strain [<a href="#ref-3">3</a>], which is well-characterized and appropriate for benchmarking.
Platform-Specific Findings
Benchmarking results are specific to the sequencing platform, basecalling model, and read depth used in the study. Findings from ONT data may not generalize to HiFi data, and vice versa. The yeast study found that different tools performed best on each platform [<a href="#ref-3">3</a>], highlighting the need for platform-specific benchmarking.
When reporting benchmarking results, clearly state the platform and conditions used. Avoid making general claims about tool performance that extend beyond the tested conditions.
Rapid Tool Development
Polishing tools are under active development, and performance can change substantially between versions. A benchmark that compares version 1.0 of a tool with version 2.0 of another tool may not reflect current performance. The Pseudomonas study specified tool versions, including dorado v0.9.1 and specific basecalling models [<a href="#ref-2">2</a>], which allows others to understand the context of the findings.
When publishing benchmarking results, report tool versions and dates. Consider repeating benchmarks periodically to account for tool updates.
Computational Resource Constraints
Large-scale benchmarking studies require substantial computational resources. The yeast study performed 455 assemblies for ONT data and 88 assemblies for HiFi data [<a href="#ref-3">3</a>], which required significant computing time. Smaller laboratories may not have the resources to replicate such extensive benchmarks.
For resource-constrained settings, consider using smaller genomes or reduced read depths for initial benchmarking, then validate key findings on larger datasets. The ntEdit study used E. coli and C. elegans for controlled experiments before testing on human and spruce genomes [<a href="#ref-5">5</a>], which is a practical approach.
Statistical Validation and Interpretation
Replicate Assemblies
Independent assemblies from the same sequencing data can vary due to stochastic elements in assembly and polishing algorithms. Performing multiple replicates allows estimation of this variation and provides more reliable comparisons between tools.
The yeast study performed multiple assemblies for each condition, which allowed statistical comparison of tool performance [<a href="#ref-3">3</a>]. For smaller studies, at least three replicates per condition is a reasonable minimum.
Paired Comparisons
When comparing two polishing tools, use the same draft assembly as input for both tools. This paired design controls for variation in the draft assembly and isolates the effect of the polishing tool. The Salmonella study used this approach by applying different polishing tools to the same assemblies [<a href="#ref-1">1</a>].
Confidence Intervals and Significance Testing
Report confidence intervals for accuracy metrics and use appropriate statistical tests to compare tools. For paired comparisons, use paired tests such as the Wilcoxon signed-rank test. For multiple comparisons, apply corrections for multiple testing.
The yeast study used multiple quality metrics and statistical comparisons to identify significant differences between tools [<a href="#ref-3">3</a>]. This approach provides more reliable conclusions than comparing point estimates without uncertainty measures.
Interpreting Negative Results
A finding that a new tool does not outperform existing tools is still valuable. It suggests that the new tool does not provide a meaningful improvement and that resources should be directed elsewhere. Report negative results with the same rigor as positive results.
Reporting Benchmarking Results
Structured Reporting
Report benchmarking results in a structured format that includes:
- Study question and objectives
- Reference genomes and sequencing data used
- Tools and versions compared
- Experimental design and replicates
- Error metrics and evaluation methods
- Results for each tool and condition
- Statistical analysis and interpretation
- Limitations and caveats
The Galaxy Training Network provides tutorials on accessible workflow analysis and reproducibility [<a href="#ref-9">9</a>], which can help researchers develop structured reporting practices.
Data and Code Availability
Make benchmarking data and analysis code publicly available whenever possible. The Plassembler study provides an example of this practice, with the full benchmarking pipeline available on GitHub and benchmarking input and output files deposited in Zenodo [<a href="#ref-10">10</a>].
Public availability of benchmarking materials allows other researchers to verify results, extend the analysis, and apply the benchmark to new tools.
Version Documentation
Document all software versions, including the operating system, programming language versions, and tool versions. The nf-core documentation emphasizes the importance of version tracking for reproducible workflows [<a href="#ref-6">6</a>]. This documentation is essential for others who want to reproduce the benchmark.
Professional Escalation Criteria
When to Seek Additional Expertise
Some benchmarking situations require specialized expertise beyond what is described here. Consider consulting with bioinformatics specialists or statisticians when:
- The benchmark involves complex statistical designs or analyses
- The organism has unusual genomic features that complicate error assessment
- The downstream application has regulatory or clinical implications
- The benchmark results will be used for high-stakes decisions
The EMBL-EBI training resources provide pathways for developing bioinformatics skills [<a href="#ref-8">8</a>], and the Bioconductor project offers documentation for reproducible genomic analysis [<a href="#ref-11">11</a>]. These resources can help researchers build the expertise needed for rigorous benchmarking.
When to Repeat a Benchmark
Repeat a benchmark when:
- A new version of a polishing tool is released with substantial changes
- A new sequencing platform or basecalling model becomes available
- The benchmark results will be used for a different application than originally intended
- Questions arise about the validity of the original benchmark
The rapid pace of tool development means that benchmarks can become outdated quickly. The Pseudomonas study specified tool versions and basecalling models [<a href="#ref-2">2</a>], which allows others to understand when the results may no longer reflect current performance.
A Practical Decision Framework for Benchmarking Polishing Pipelines
Define the Benchmark Tier Before Selecting Tools
A common failure in polishing benchmarks is treating all evaluation questions as if they require the same experimental design. A decision framework with three tiers helps researchers match the depth of benchmarking to the specific claim they need to support. Tier 1 is a screening benchmark that answers whether a new tool or pipeline is worth pursuing at all. Tier 2 is a comparative benchmark that positions a tool against established methods on a defined dataset. Tier 3 is a deployment benchmark that validates performance for a specific application, such as outbreak investigation or clinical reporting.
The tier determines the number of replicates, the choice of reference genomes, the breadth of tools compared, and the statistical rigor required. A Tier 1 screening benchmark might use a single bacterial genome with simulated reads and a small set of comparison tools. A Tier 3 deployment benchmark requires multiple real datasets, application-specific metrics, and enough replicates to support confidence intervals. The Salmonella outbreak study exemplifies Tier 3 benchmarking, testing 132 combinations of assembly and polishing tools across 15 closely related isolates to support source tracking conclusions [<a href="#ref-1">1</a>]. The Pseudomonas study similarly represents Tier 3 validation, using four retrospective hospital outbreak isolate sets to confirm that the optimized nanopore-only workflow produced cgMLST results concordant with Illumina short-read assemblies [<a href="#ref-2">2</a>].
Select Comparison Tools Based on the Benchmark Tier
The choice of comparison tools should follow from the benchmark tier and the specific claim being tested. For a Tier 1 screening benchmark, include two or three established tools that represent different polishing strategies. For bacterial genomes, this might include one long-read polisher such as Medaka and one short-read polisher such as NextPolish or Pilon. The Salmonella study found that Medaka was a more accurate and efficient long-read polisher than Racon, and among short-read polishers, NextPolish showed the highest accuracy while Pilon, Polypolish, and POLCA performed similarly [<a href="#ref-1">1</a>]. These findings provide a baseline expectation for tool performance that a new tool must exceed.
For a Tier 2 comparative benchmark, include tools that represent the full range of approaches relevant to your data type. The yeast benchmarking study compared 7 assemblers with 4 polishing pipelines across 13 depth subsets for ONT data and 4 assemblers with or without polishing across 11 depth subsets for HiFi data [<a href="#ref-3">3</a>]. This breadth allowed the researchers to identify which tools performed best under which conditions. For a Tier 3 deployment benchmark, include the specific tools that your target users would realistically employ, and test them in the combinations and order that users would apply.
Establish the Decision Criteria Before Running the Benchmark
Define the criteria for declaring one pipeline superior to another before any tools are run. This prevents post hoc rationalization of results and ensures that the benchmark answers the question it was designed to address. Decision criteria should specify the primary metric, the minimum difference that is considered meaningful, and the statistical test that will be used.
For example, a benchmark might declare that a new pipeline is superior if it achieves a QV improvement of at least 2 over the best existing tool on the same draft assembly, with a confidence interval that does not include zero. The Salmonella study reported near-perfect accuracy of 99.9999 percent, corresponding to approximately 5 nucleotide errors across a 4.8 Mbp genome, excluding low-confidence regions [<a href="#ref-1">1</a>]. A new tool claiming improvement over this level would need to demonstrate a meaningful reduction in the remaining error burden.
Decision criteria should also address trade-offs between accuracy and computational efficiency. The ntEdit study demonstrated that a tool could fix the majority of base substitutions and indels at low sequence depths below 20x, with performance largely constant as coverage increased, while executing in less than 14 seconds for E. coli and less than 3 minutes for C. elegans on a single CPU [<a href="#ref-5">5</a>]. A benchmark that only compares accuracy would miss the substantial efficiency advantage that ntEdit offers for large genomes.
Create a Benchmark Decision Record
Maintain a structured record of every decision made during benchmark design. This record serves as documentation for publications and as a template for future benchmarks. The record should include the benchmark tier, the specific question being addressed, the reference genomes selected and the rationale for each, the sequencing data sources and accessions, the tools and versions compared, the parameters used for each tool, the error metrics calculated, the statistical methods applied, and the decision criteria established before running the benchmark.
The nf-core documentation emphasizes the importance of reproducible workflow standards, including version tracking and configuration management [<a href="#ref-6">6</a>]. A benchmark decision record extends this principle to the experimental design itself. The Carpentries lessons provide foundational training in shell, Git, and programming practices that support reproducible analysis [<a href="#ref-7">7</a>], which are directly applicable to maintaining benchmark records.
Record Tool Versions and Parameters Systematically
Tool version and parameter documentation is essential for interpreting benchmark results and for reproducing the benchmark at a later date. The Pseudomonas study specified tool versions, including dorado v0.9.1 and the [email protected] basecalling model, along with the specific error correction and polishing algorithms used [<a href="#ref-2">2</a>]. This level of detail allows other researchers to understand the context of the findings and to repeat the benchmark with updated tools.
For each tool in the benchmark, record the exact version, the command used, all non-default parameters, and the source of the parameter values. If parameters were optimized on a training dataset, document the optimization process and the final values. The Plassembler study provides an example of comprehensive documentation, with the full benchmarking pipeline available on GitHub and benchmarking input and output files deposited in Zenodo [<a href="#ref-10">10</a>].
Track Computational Resources as a Decision Variable
Computational resource usage should be recorded for every tool and pipeline tested, beyond reported as supplementary information. Runtime, memory usage, and CPU requirements are decision variables that affect whether a pipeline is practical for routine use. The ntEdit study reported execution times for each test case, showing that the pipeline executed in less than 14 seconds for E. coli and less than 3 minutes for C. elegans on a single CPU [<a href="#ref-5">5</a>]. This information is directly relevant to laboratories that need to process many samples.
The yeast study provided depth recommendations of more than 80x for ONT data and more than 20x for HiFi data for high-quality genome construction [<a href="#ref-3">3</a>]. These thresholds have direct implications for sequencing cost, which is a practical constraint for many laboratories. A benchmark that reports accuracy without considering the sequencing depth required to achieve that accuracy provides an incomplete picture for decision making.
Apply the Framework to a Worked Example
Consider a laboratory that has developed a new polishing tool for bacterial nanopore assemblies and wants to benchmark it against existing methods. Using the tier framework, the laboratory would first conduct a Tier 1 screening benchmark using a single well-characterized bacterial genome with a complete reference, such as a closed chromosome from NCBI [<a href="#ref-4">4</a>]. Simulated reads at multiple depths would provide controlled error profiles, and the new tool would be compared against Medaka and NextPolish, which represent the best-performing long-read and short-read polishers from the Salmonella study [<a href="#ref-1">1</a>].
If the Tier 1 results justify further investigation, the laboratory would proceed to a Tier 2 comparative benchmark using real sequencing data from multiple bacterial species with different GC content and repeat structure. The Pseudomonas study noted that library preparation incorporated temperature ramps to improve performance for high-GC content genomes [<a href="#ref-2">2</a>], indicating that GC content is a relevant variable for benchmarking. The benchmark would include multiple assemblers and polishing pipelines, following the design of the yeast study [<a href="#ref-3">3</a>].
For a Tier 3 deployment benchmark, the laboratory would validate the pipeline on datasets relevant to the intended application. If the application is outbreak investigation, the benchmark would use closely related isolates and evaluate whether the polished assemblies support correct cgMLST assignments, as demonstrated in the Pseudomonas study [<a href="#ref-2">2</a>]. The decision criteria would include concordance with short-read assemblies and the number of discordant positions per genome.
Common Failure Patterns in Benchmark Decision Making
Several failure patterns recur when researchers apply benchmarking frameworks. The first is selecting comparison tools that are known to perform poorly on the test data, which inflates the apparent advantage of the new tool. The second is optimizing parameters on the test data and then reporting performance on the same data, which constitutes overfitting. The third is reporting only genome-wide accuracy while ignoring errors in difficult regions such as homopolymers and repetitive elements, which the Salmonella study identified as the most challenging errors to correct [<a href="#ref-1">1</a>].
A fourth failure pattern is comparing pipelines that use different assemblers without controlling for assembler effects. The yeast study found that the assembler plays an essential role in genome construction, especially for low-depth datasets [<a href="#ref-3">3</a>]. A benchmark that compares a new polishing tool applied to one assembler's output against an existing tool applied to a different assembler's output cannot attribute performance differences to the polishing tools.
A fifth failure pattern is drawing conclusions from a single replicate. The yeast study performed 455 assemblies for ONT data and 88 assemblies for HiFi data [<a href="#ref-3">3</a>], which provided the statistical power needed to identify significant differences. Single-replicate benchmarks cannot distinguish real performance differences from stochastic variation in assembly and polishing algorithms.
Professional Escalation Criteria for Benchmark Design
Some benchmarking situations require expertise beyond what a typical laboratory can provide. Consider consulting with bioinformatics specialists or statisticians when the benchmark involves complex statistical designs, when the organism has unusual genomic features that complicate error assessment, or when the downstream application has regulatory or clinical implications. The EMBL-EBI training resources provide pathways for developing bioinformatics skills [<a href="#ref-8">8</a>], and the Bioconductor project offers documentation for reproducible genomic analysis [<a href="#ref-11">11</a>].
Consultation is also appropriate when benchmark results will be used for high-stakes decisions, such as selecting a pipeline for clinical outbreak investigation or for regulatory submissions. The Galaxy Training Network provides tutorials on accessible workflow analysis and reproducibility [<a href="#ref-9">9</a>], which can help laboratories build the skills needed for rigorous benchmarking. When questions arise about the validity of an existing benchmark, repeating the benchmark with updated tools and documented parameters is often the most direct path to resolution.
Frequently Asked Questions
What is the minimum number of replicates needed for a fair polishing benchmark?
At least three independent replicates per condition is a reasonable minimum for detecting large performance differences. The yeast benchmarking study performed hundreds of assemblies across different conditions [<a href="#ref-3">3</a>], which provided robust statistical power. For smaller studies, more replicates may be needed to detect smaller differences. The appropriate number depends on the variability of the assembly and polishing process and the size of the effect you want to detect.
How do I choose between simulated and real sequencing data for benchmarking?
Simulated data allows controlled experiments with known ground truth, which is valuable for initial testing and parameter optimization. Real data reflects actual sequencing error profiles, which is essential for final validation. A hybrid approach that uses both types of data is often the most informative. The ntEdit study used controlled E. coli and C. elegans data for initial testing and real human and spruce data for validation [<a href="#ref-5">5</a>].
What error metrics should I report in a polishing benchmark?
Report multiple metrics that capture different aspects of error, including QV or accuracy, mismatch rate, indel rate, and BUSCO completeness. Report error rates in difficult regions such as homopolymers and repetitive elements separately from genome-wide rates. The Salmonella study reported accuracy as a percentage and the number of nucleotide errors across the genome [<a href="#ref-1">1</a>], which provided an intuitive sense of error burden.
How do I account for different sequencing platforms in a benchmark?
Test polishing tools on data from each platform you intend to support. The yeast study found that different tools performed best on ONT and HiFi data [<a href="#ref-3">3</a>], so results from one platform may not generalize to another. Report results separately for each platform and specify the basecalling model used.
What should I do if my new tool does not outperform existing tools?
Report the negative result with the same rigor as a positive result. A finding that a new tool does not provide meaningful improvement is valuable information for the community. It suggests that the new tool should not be adopted and that resources should be directed elsewhere. The benchmarking framework described here provides a structure for documenting such findings transparently.
How do I handle low-confidence regions in accuracy calculations?
Report whether low-confidence regions were excluded from accuracy calculations. The Salmonella study explicitly noted that their accuracy calculation excluded low-confidence regions [<a href="#ref-1">1</a>]. If low-confidence regions are excluded, report this clearly and consider also reporting accuracy including all regions. Different studies may use different conventions, so transparency is essential for comparability.
What computational resources should I report in a benchmark?
Report runtime, memory usage, and CPU requirements for each tool. The ntEdit study reported execution times for each test case [<a href="#ref-5">5</a>], which allowed readers to assess practical utility. Resource usage should be reported alongside accuracy metrics, as a tool that achieves slightly higher accuracy but requires substantially more resources may not be the best choice for routine use.
How often should I repeat a benchmarking study?
Repeat benchmarks when new versions of polishing tools are released with substantial changes, when new sequencing platforms or basecalling models become available, or when the benchmark results will be used for a different application than originally intended. The rapid pace of tool development means that benchmarks can become outdated quickly. The Pseudomonas study specified tool versions and basecalling models [<a href="#ref-2">2</a>], which allows others to understand when the results may no longer reflect current performance.
Related Bioinformatics Guides
- Evaluating Metagenomic Assembly Tools: A Benchmarking Framework for Short-Read and Long-Read Data
- Metagenomics Tools: A Practical Guide to Software and Pipelines
- Genomic Data Analysis Tools: A Comparative Guide for Researchers
- Explainable AI for Bioinformatics: Methods, Tools, and Applications
- Metagenomic Binning Tools Benchmark: How to Evaluate and Choose
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
[1] [Benchmarking short and long read polishing tools for nanopore assemblies: achieving near-perfect genomes for outbreak isolates.](https://pubmed.ncbi.nlm.nih.gov/38978005). BMC genomics, 2024. [2] [An open-source nanopore-only sequencing workflow for analysis of clonal outbreaks delivers short-read level accuracy.](https://pubmed.ncbi.nlm.nih.gov/40679848). Journal of clinical microbiology, 2025. [3] [Benchmarking of long-read sequencing, assemblers and polishers for yeast genome.](https://pubmed.ncbi.nlm.nih.gov/35511110). Briefings in bioinformatics, 2022. [4] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information. [5] [ntEdit: scalable genome sequence polishing.](https://pubmed.ncbi.nlm.nih.gov/31095290). Bioinformatics (Oxford, England), 2019. [6] [nf-core Documentation](https://nf-co.re/docs). nf-core. [7] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries. [8] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute. [9] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project. [10] [Plassembler: an automated bacterial plasmid assembly tool.](https://pubmed.ncbi.nlm.nih.gov/37369026). Bioinformatics (Oxford, England), 2023. [11] [Bioconductor](https://bioconductor.org/). Bioconductor Project.This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.