# Polishing in the Cloud: Scaling Up Error Correction for Large Genomes with Parallel Workflows


## Key Takeaways

- Scaling genome polishing for large genomes (e.g., livestock, crops) necessitates parallel cloud workflows to overcome computational bottlenecks, as polishing cost increases with genome size, read depth, and iterations.
- Parallel polishing strategies involve chunking the draft assembly into manageable segments (1-5 Mb with 10-50 kb overlaps) to enable simultaneous processing by multiple workers, with careful overlap handling to ensure seamless merging.
- Cloud platforms offer elastic compute (e.g., AWS Spot instances, GCP preemptible instances) to match bursty polishing demands, significantly reducing costs by up to 90% for fault-tolerant workflows that can retry interrupted tasks.
- Workflow managers like Nextflow and Snakemake are crucial for orchestrating parallel tasks, managing dependencies, and ensuring reproducibility by defining discrete steps with explicit inputs and outputs.
- Quality assessment post-polishing is critical, involving reference-based alignment to identify remaining single nucleotide polymorphisms (SNPs) and indels, and completeness checks to ensure gene integrity, with metrics like median SNPs/indels per genome guiding further refinement.
- Common failure patterns include insufficient memory for large genomes, read data access bottlenecks, spot instance interruptions, inadequate chunk boundary handling, and software version inconsistencies, all of which require specific workflow design or instance configuration to mitigate.

---

Genome assembly polishing is the computational process of correcting base-level errors in draft assemblies using raw sequencing reads. For large genomes, such as those of livestock, crops, or complex pathogens, polishing becomes a compute-intensive bottleneck that can stall an entire project. Cloud computing offers a practical path to scale polishing across many parallel workers, but only when the workflow is designed with cost controls, data management, and reproducibility in mind. This article explains how to plan, execute, and evaluate cloud-based polishing for large genomes, with attention to the specific decisions that determine success or failure.

## The Polishing Bottleneck in Large Genome Projects

Draft genome assemblies produced by long-read assemblers contain systematic errors that must be corrected before downstream analysis. These errors include insertions and deletions in homopolymer regions, base substitutions, and misassemblies at repetitive elements. Polishing tools use the original sequencing reads to identify and correct these errors by aligning reads back to the assembly and computing a consensus sequence.

The computational cost of polishing scales with genome size, read depth, and the number of polishing iterations. A bacterial genome of 5 megabases can be polished on a laptop in minutes. A livestock genome of 2.5 to 3 gigabases requires hours to days of compute time on a single server. A plant genome exceeding 10 gigabases can take weeks. This scaling problem is the primary reason cloud computing has become central to large genome projects.

The [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/) provide access to reference genomes and sequence databases that are essential for validating assembly quality and comparing polished results against known standards. Researchers working with large genomes should use these resources to establish quality baselines before and after polishing.

Cloud platforms address the scaling problem through elastic compute allocation. Instead of purchasing and maintaining a fixed cluster, a researcher can provision hundreds of virtual machines for a few hours, run the polishing workflow in parallel, and release the resources when the job completes. This model matches the bursty compute demands of genome polishing, where intense computation is followed by periods of analysis and validation.

## Core Principles of Parallel Polishing

Parallel polishing requires dividing the work into independent units that can be processed simultaneously. The most common strategy is to split the draft assembly into chunks, polish each chunk independently, and then merge the polished chunks back into a complete genome. This approach works because polishing tools operate locally on sequence windows and do not require global context for most error types.

### Chunking Strategies

The choice of chunk size affects both parallelism and accuracy. Smaller chunks allow more parallel workers but increase the risk of errors at chunk boundaries. Larger chunks reduce boundary artifacts but limit parallelism. A practical approach is to use chunks of 1 to 5 megabases for large genomes, with overlapping regions of 10 to 50 kilobases at each boundary to ensure continuity.

The overlap regions are critical. After polishing, the overlapping segments are compared and trimmed to produce a seamless junction. If the overlap is too small, the polished chunks may not merge cleanly. If the overlap is too large, compute time is wasted on redundant work.

### Read Alignment and Consensus

Each polishing chunk requires access to the sequencing reads that cover that genomic region. For large genomes with high read depth, the read set can be substantial. A 3 gigabase genome at 50x coverage with Oxford Nanopore reads produces roughly 150 gigabases of sequence data. Managing this data across parallel workers requires a shared storage system or a strategy for distributing read subsets to each worker.

The [Galaxy Training Network](https://training.galaxyproject.org/) offers practical tutorials on assembly and polishing workflows that demonstrate how to structure these steps in a reproducible manner. These training materials are useful for laboratory teams that are new to parallel computing and need a structured introduction to the concepts.

### Workflow Managers

Workflow managers such as Nextflow and Snakemake provide the orchestration layer for parallel polishing. These tools track dependencies between steps, manage input and output files, and handle retries when individual tasks fail. The [nf-core Documentation](https://nf-co.re/docs) describes community standards for pipeline development that emphasize reproducibility and portability across computing environments.

A well-designed workflow defines each polishing step as a discrete task with explicit inputs and outputs. This design allows the workflow manager to schedule tasks across available compute resources and to resume from the last completed step if a failure occurs. For large genomes, the ability to resume is essential because a single failed task should not require restarting the entire polishing run.

## At a Glance: Cloud Polishing Decision Framework

The following table summarizes the key decisions that determine the success of a cloud-based polishing project. Use this framework during the planning phase to align compute choices with project goals.

| Decision Point | Option A | Option B | Consideration |
| --- | --- | --- | --- |
| Compute instance type | On-demand instances | Spot or preemptible instances | On-demand provides guaranteed availability but higher cost. Spot instances reduce compute costs substantially but require fault-tolerant workflows that can retry interrupted tasks. |
| Polishing tool selection | Long-read polisher only (Medaka) | Hybrid approach with short-read polisher (Polypolish or Pypolca) | Long-read polishing alone may suffice for many assemblies. Adding short-read polishing improves base-level accuracy but increases compute time and cost. |
| Data storage strategy | Central object storage (S3 or Google Cloud Storage) | Distributed or replicated storage | Central object storage is simpler and cheaper. Distributed storage reduces read bottlenecks when many workers access the same files concurrently. |
| Workflow orchestration | Nextflow or Snakemake on raw cloud instances | Managed platform such as Terra.bio | Direct workflow managers offer maximum flexibility and portability. Managed platforms provide user-friendly interfaces and integrated workflows but reduce configuration control. |
| Chunk size for parallelization | 1 to 5 megabase chunks with 10 to 50 kilobase overlaps | Larger chunks with reduced parallelism | Smaller chunks increase parallelism but require careful boundary handling. Larger chunks reduce boundary artifacts but limit the number of concurrent workers. |

## Cloud Platform Options for Polishing Workloads

Cloud providers offer different services that can support polishing workflows. The choice of platform depends on the scale of the project, the available budget, and the technical expertise of the research team.

### Amazon Web Services

AWS provides a range of compute options, including on-demand instances, spot instances, and dedicated hosts. Spot instances offer significant cost savings for fault-tolerant workloads because they can be interrupted with short notice. Polishing workflows are well suited to spot instances because individual tasks can be retried without losing the entire project.

A study on tick genomics demonstrated the practical use of AWS Spot instances for genome assembly and polishing. The research team assembled and polished cattle tick genomes using Oxford Nanopore sequencing data and finalized the assemblies on the AWS cloud platform, taking advantage of up to 90% discounted Spot instances. The resulting genomes were comparable to published tick genomes that had been assembled using multiple sequencing technologies and more costly bioinformatic resources. This example shows that cloud-based polishing can produce high-quality results at a fraction of the cost of traditional high-performance computing approaches.

The [EMBL-EBI Training](https://www.ebi.ac.uk/training) portal provides learning pathways that cover cloud-based bioinformatics analysis, including practical guidance on managing compute resources and data storage. These resources are valuable for teams that need to build cloud skills from a foundation.

### Google Cloud Platform

Google Cloud offers similar compute options, including preemptible instances that provide cost savings for interruptible workloads. Google Cloud also provides integration with common bioinformatics tools and workflows through its marketplace and preconfigured images.

The choice between AWS and Google Cloud often comes down to institutional preferences, existing infrastructure, and the availability of support. Both platforms can run Nextflow and Snakemake workflows effectively. The key is to design the workflow so that it is portable and does not depend on platform-specific features.

### Terra.bio

Terra.bio is a cloud-based platform that provides a managed environment for bioinformatics analysis. It offers a graphical interface for running workflows and managing data, which can be useful for teams that do not have dedicated bioinformatics staff. An evaluation of tools for decentralized antimicrobial resistance surveillance in Africa identified Terra.bio as a leading candidate for cloud-based scalability and interoperability. The platform supports integration with standard workflow languages and provides a user-friendly interface for non-specialists.

For polishing large genomes, Terra.bio can be a practical option when the research team needs a managed environment and does not want to configure cloud infrastructure directly. The tradeoff is reduced flexibility compared to running workflows directly on cloud compute instances.

## Designing a Cloud-Based Polishing Workflow

A cloud-based polishing workflow requires careful planning to balance cost, speed, and accuracy. The following steps outline a practical approach.

### Step 1: Assess the Assembly and Read Data

Before designing the polishing workflow, document the assembly size, the number of contigs, the read type and depth, and the expected error profile. This information determines the compute requirements and the polishing strategy.

For large genomes, the assembly file can be several gigabytes. The read data can be tens to hundreds of gigabytes. These files must be staged in cloud storage that is accessible to the compute instances. Object storage services such as AWS S3 or Google Cloud Storage are appropriate for this purpose because they provide high throughput and durability.

### Step 2: Select the Polishing Tools

The choice of polishing tools depends on the sequencing technology used. Oxford Nanopore reads are commonly polished with Medaka, which uses a neural network model trained on Nanopore data. Short-read polishing tools such as Polypolish and Pypolca use Illumina reads to correct errors that remain after long-read polishing.

A study of clinical Enterobacterales genomes compared multiple polishing approaches for Nanopore long-read assemblies. The research found that Medaka polishing with un-subsampled long reads produced the most accurate assemblies, with quality scores comparable to hybrid assemblies that combined long and short reads. The study also found that Medaka polishing improved indel errors but did not significantly change single nucleotide polymorphism errors. These findings indicate that the choice of polishing tool and the decision to subsample reads have measurable effects on final assembly quality.

For large genomes, the polishing strategy may combine multiple tools in sequence. A common approach is to polish with Medaka using long reads, then polish with Polypolish or Pypolca using short reads if available. Each polishing round adds compute time, so the number of rounds should be justified by quality assessments.

### Step 3: Configure the Cloud Environment

The cloud environment must be configured to support the polishing workflow. This configuration includes selecting the instance type, setting up storage, and installing the required software.

Instance selection is a critical decision. Polishing tools are memory intensive because they load the assembly and read alignments into memory. A 3 gigabase genome may require 64 to 128 gigabytes of RAM for efficient polishing. Compute-optimized instances with high CPU counts reduce wall time but increase cost per hour. The optimal configuration balances these factors based on the project budget and timeline.

The [Bioconductor](https://bioconductor.org/) project provides documentation on reproducible genomic analysis workflows, including guidance on managing software environments and dependencies. While Bioconductor focuses on R-based analysis, the principles of environment management and reproducibility apply to cloud-based polishing workflows.

### Step 4: Implement the Parallel Workflow

The workflow manager orchestrates the parallel polishing tasks. Each task processes one chunk of the assembly and produces a polished chunk as output. The workflow manager tracks the status of each task and handles retries for failed tasks.

The workflow should include a validation step after polishing to check that the polished chunks merge correctly and that the final assembly has no gaps or unexpected duplications. This validation step is essential for catching errors that may have been introduced during chunking or merging.

### Step 5: Monitor Costs and Progress

Cloud costs accrue from compute time, storage, and data transfer. Monitoring these costs is essential for keeping the project within budget. Cloud providers offer billing dashboards and cost alerts that can be configured to notify the research team when spending exceeds thresholds.

Progress monitoring is equally important. The workflow manager should provide logs that show the status of each task, including completion times and error messages. These logs are valuable for diagnosing failures and for estimating the remaining time to completion.

## Cost Considerations for Cloud Polishing

The cost of cloud-based polishing depends on several factors: the size of the genome, the read depth, the number of polishing rounds, the instance type, and the use of discounted instances.

### Compute Costs

Compute costs are the largest component of a cloud polishing budget. On-demand instances are the most expensive but provide guaranteed availability. Spot instances can reduce compute costs by up to 90% but may be interrupted, requiring the workflow to handle retries.

The tick genomics study demonstrated the cost savings potential of spot instances. The research team used AWS Spot instances to assemble and polish tick genomes, achieving results comparable to projects that used more expensive computing resources. This example illustrates that cost-conscious cloud design can make large genome projects accessible to laboratories with limited budgets.

### Storage Costs

Storage costs accrue from the assembly files, read data, intermediate files, and final outputs. Object storage is relatively inexpensive, but the volume of data for large genomes can be substantial. A 3 gigabase genome with 50x read coverage requires approximately 150 gigabytes of read storage and additional space for intermediate files.

Data lifecycle policies can reduce storage costs by automatically moving older files to cheaper storage tiers or deleting temporary files after the project completes. These policies should be configured before the polishing run begins to avoid unexpected storage charges.

### Data Transfer Costs

Data transfer costs apply when moving data into and out of the cloud. Uploading read data to cloud storage incurs transfer costs, as does downloading the final polished assembly. These costs are typically small compared to compute costs but should be included in the budget.

## Cost Comparison Across Cloud Strategies

The following table provides a practical comparison of cost-related factors across common cloud deployment strategies for polishing workloads. Actual prices vary by region and over time, so use this table as a planning framework instead of a price list.

| Strategy | Compute Cost Profile | Interruption Risk | Best Use Case | Management Overhead |
| --- | --- | --- | --- | --- |
| On-demand instances | Highest per hour, predictable total | None | Critical tasks that cannot be retried, final validation steps | Moderate, requires manual instance provisioning |
| Spot or preemptible instances | Up to 90% lower per hour, variable total | High, tasks may be terminated | Bulk polishing chunks that can be retried | Low, workflow manager handles retries |
| Managed platform such as Terra.bio | Moderate, bundled pricing | Low | Teams without dedicated cloud engineering staff | Low, platform handles infrastructure |
| Hybrid spot and on-demand mix | Moderate, balances cost and reliability | Moderate | Large projects where some tasks are critical and others are interruptible | High, requires workflow logic to assign task priority |

## Quality Assessment After Polishing

Polishing is not complete until the assembly quality has been assessed. Quality assessment involves comparing the polished assembly to reference genomes, checking for completeness, and evaluating error rates.

### Reference-Based Assessment

When a closely related reference genome is available, the polished assembly can be aligned to the reference to identify remaining errors. The [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/) provide access to reference genomes for many species, including livestock, crops, and pathogens. Alignment tools can identify single nucleotide differences, insertions, deletions, and structural variations between the polished assembly and the reference.

The interpretation of these differences requires caution. Not all differences are errors. Some differences reflect genuine biological variation between the sequenced individual and the reference. The distinction between errors and variation requires careful analysis, often involving comparison of the polished assembly to the raw reads at the positions of interest.

### Completeness Assessment

Completeness assessment checks whether the polished assembly contains all expected genes and genomic features. The [EMBL-EBI Training](https://www.ebi.ac.uk/training) resources provide guidance on using completeness assessment tools and interpreting their results. For large genomes, completeness assessment is essential because polishing errors can disrupt genes and create false frameshifts.

### Quality Metrics

Standard quality metrics for polished assemblies include the number of single nucleotide polymorphisms and indels compared to a reference, the quality value score, and the number of complete genes. The Enterobacterales study used these metrics to compare polishing approaches and found that the best long-read only polishing approach produced assemblies with zero median single nucleotide polymorphisms and zero median indels per genome. These metrics provide a quantitative basis for deciding whether additional polishing rounds are needed.

## Common Failure Patterns in Cloud Polishing

Cloud-based polishing can fail in predictable ways. Recognizing these failure patterns helps research teams diagnose problems quickly and avoid wasting compute time.

### Insufficient Memory

Polishing tools can exhaust memory on large genomes, causing tasks to fail or the instance to crash. This failure pattern is common when the chunk size is too large or when the instance type has insufficient RAM. The solution is to reduce the chunk size or select an instance with more memory.

### Read Data Access Bottlenecks

When many parallel workers attempt to read the same data files simultaneously, the storage system can become a bottleneck. This failure pattern manifests as slow task execution and increased wall time. The solution is to distribute read data across multiple storage locations or to use a parallel file system that supports high concurrent read throughput.

### Spot Instance Interruptions

Spot instances can be terminated with short notice, causing tasks to fail. The workflow manager must handle these interruptions gracefully by retrying failed tasks on new instances. If interruptions become frequent, the workflow may make little progress. In this case, the research team should consider using on-demand instances for critical tasks or adjusting the spot instance strategy.

### Chunk Boundary Artifacts

Errors at chunk boundaries can persist after polishing if the overlap regions are not handled correctly. This failure pattern produces assemblies with local errors that are difficult to detect without careful validation. The solution is to use adequate overlap regions and to validate the merged assembly at chunk boundaries.

### Software Version Inconsistencies

Polishing tools are under active development, and different versions can produce different results. If the workflow uses inconsistent software versions across tasks, the polished chunks may not be comparable. The solution is to use containerized software environments that pin specific tool versions.

## Reproducibility and Record Keeping

Reproducibility is a core requirement for scientific research. Cloud-based polishing workflows must be documented so that other researchers can understand what was done and repeat the analysis if needed.

### Workflow Documentation

The workflow definition file, whether written in Nextflow, Snakemake, or another language, serves as the primary documentation of the polishing process. This file should specify the tools, versions, parameters, and input data for each step. The [nf-core Documentation](https://nf-co.re/docs) emphasizes the importance of standardized pipeline documentation for reproducibility.

### Environment Capture

The software environment used for polishing should be captured in a container image or an environment definition file. This capture ensures that the same tool versions and dependencies are used across all tasks and that the analysis can be repeated in the future.

### Data Provenance

Data provenance records the origin and processing history of each data file. For polishing workflows, provenance includes the assembly file version, the read data version, the polishing tool versions, and the parameters used. Cloud storage systems can store metadata that captures this provenance information.

### Logs and Metrics

The workflow manager generates logs that record the execution of each task. These logs should be preserved as part of the project record. Performance metrics, such as task completion times and resource usage, are also valuable for planning future polishing runs.

## Limitations of Cloud-Based Polishing

Cloud-based polishing has limitations that research teams should understand before committing to this approach.

### Network Dependence

Cloud computing requires reliable network connectivity. Laboratories in regions with limited internet infrastructure may find it difficult to upload large read datasets and download polished assemblies. The evaluation of tools for decentralized antimicrobial resistance surveillance in Africa noted that offline capability was a key consideration for some tools. This limitation is relevant for any research team working in a low-connectivity environment.

### Cost Uncertainty

Cloud costs can be unpredictable, especially when spot instance interruptions cause tasks to be retried multiple times. Research teams should build cost monitoring into their workflow and set budget alerts to avoid unexpected charges.

### Data Security and Privacy

Genome data may be subject to data protection regulations, especially when the data comes from human subjects or commercially valuable livestock. Research teams must ensure that their cloud configuration complies with applicable regulations and that data is encrypted in transit and at rest.

### Dependency on External Providers

Cloud-based polishing creates a dependency on external service providers. If the provider experiences an outage or changes its pricing structure, the project may be affected. The tick genomics study noted that traditional genome assembly approaches create a dependency on external service providers, and the same consideration applies to cloud-based approaches.

## Professional Escalation Criteria

Research teams should know when to escalate polishing problems to more experienced colleagues or external experts. The following situations warrant escalation.

### Persistent Task Failures

If polishing tasks fail repeatedly despite adjusting chunk sizes, instance types, and software versions, the problem may be in the assembly itself or in the read data. An experienced bioinformatician should review the assembly and read quality before further polishing attempts.

### Unexpected Quality Degradation

If quality assessment shows that the polished assembly is worse than the draft assembly, the polishing process may be introducing errors. This situation requires expert review of the polishing parameters and the read data.

### Budget Overruns

If cloud costs exceed the project budget by a significant margin, the polishing strategy should be reviewed. An expert may identify more cost-effective approaches, such as using different instance types or reducing the number of polishing rounds.

### Reproducibility Concerns

If the polishing workflow cannot be reproduced because of missing documentation, inconsistent software versions, or lost data, the project may need to be repeated. An expert can help reconstruct the workflow and document it properly.

## A Practical Decision Framework for Selecting Polishing Tools and Cloud Configurations

Choosing the right polishing tools and cloud configuration for a large genome project requires a structured approach that accounts for the specific characteristics of the sequencing data, the assembly quality, and the project budget. Many research teams select tools based on familiarity or convenience instead of evidence, which can lead to suboptimal results and wasted compute spending. This section provides a practical decision framework that integrates tool selection, cloud configuration, and quality assessment into a single workflow.

### Step 1: Characterize the Input Data Before Selecting Tools

The first decision point is not which polisher to use but what data is available and what errors need correction. Document the following characteristics before any cloud resources are provisioned:

**Sequencing technology and basecalling quality.** Oxford Nanopore reads basecalled with different models have different error profiles. The [study of clinical Enterobacterales isolates](https://doi.org/10.1101/2025.09.15.676237) used Dorado v5.0.0 super-high-accuracy basecalling for R10.4.1 flow cells, which produced reads accurate enough that long-read only polishing with Medaka achieved quality comparable to hybrid assemblies. Older basecalling models or earlier flow cell versions will produce reads with more systematic errors that may require additional polishing rounds or hybrid approaches.

**Read depth and coverage uniformity.** The same Enterobacterales study found that Medaka polishing with un-subsampled long reads produced small improvements in indels but not single nucleotide polymorphisms compared to subsampled reads. For large genomes, the decision to subsample reads for polishing has direct cost implications because processing the full read set requires more compute time and storage. If the assembly is already accurate at the single nucleotide level, subsampling may be acceptable. If indel errors are the primary concern, un-subsampled reads are justified.

**Assembly error profile.** Run a quick quality assessment on the draft assembly before polishing to identify the dominant error types. If the assembly has many indels in homopolymer regions, a long-read polisher such as Medaka is appropriate. If the assembly has scattered base substitutions, short-read polishing with Polypolish or Pypolca may be more effective. The [review of artificial intelligence in metagenome-assembled genome reconstruction](https://pubmed.ncbi.nlm.nih.gov/41506577) notes that AI-based polishing tools can improve base-level precision, but these tools may require additional computational resources and should be evaluated against simpler alternatives for large genomes.

**Availability of short reads.** Hybrid polishing, which combines long-read polishing with short-read correction, generally produces the most accurate assemblies. The [study of infectious laryngotracheitis virus vaccine strains](https://doi.org/10.3390/vaccines14030245) used a hybrid strategy combining Oxford Nanopore long reads for assembly contiguity with Illumina short reads for single-nucleotide level polishing. However, generating short reads adds sequencing cost. If short reads are not available, the decision framework must account for the accuracy ceiling of long-read only polishing.

### Step 2: Match Polishing Tools to Error Profiles

The evidence from the Enterobacterales study provides a clear ranking of polishing approaches for Nanopore long-read assemblies. Autocycler combined with Medaka using un-subsampled long reads was the most accurate long-read only combination, with zero median single nucleotide polymorphisms and zero median indels per genome. This combination was comparable to hybrid assemblies that used both long and short reads.

For large genomes, this finding has practical implications. If the assembly was produced with a recent assembler and the reads were basecalled with a current high-accuracy model, a single round of Medaka polishing with un-subsampled reads may be sufficient. The study found that only 4 of 92 genomes had more than 10 single nucleotide polymorphisms or indels after this approach.

The decision to add short-read polishing should be based on the quality assessment after long-read polishing. If the polished assembly meets the project quality thresholds, additional polishing rounds add cost without proportional benefit. If the assembly still has errors, short-read polishing with Polypolish or Pypolca is the next step.

The [tick genomics study](https://pubmed.ncbi.nlm.nih.gov/40597572) demonstrates that this approach can work for large and complex genomes. The research team assembled cattle tick genomes using only Oxford Nanopore data and polished them on the AWS cloud platform. The resulting genomes were comparable to published tick genomes that used multiple sequencing technologies. This example shows that a carefully designed long-read only workflow can produce high-quality results for large genomes without the cost of hybrid sequencing.

### Step 3: Select Cloud Instances Based on Task Characteristics

Not all polishing tasks have the same compute requirements or the same tolerance for interruption. The decision framework should assign each task type to an appropriate instance category.

**Memory-intensive tasks.** Polishing tools load the assembly and read alignments into memory. A 3 gigabase genome may require 64 to 128 gigabytes of RAM for efficient polishing. These tasks should run on memory-optimized instances with sufficient RAM to avoid out-of-memory failures. The cost per hour for these instances is higher, but the tasks complete in fewer wall-clock hours.

**CPU-intensive tasks.** Read alignment and consensus computation are CPU-intensive. Compute-optimized instances with high CPU counts reduce wall time but increase cost per hour. The optimal configuration depends on whether the workflow is memory-bound or CPU-bound. Monitor resource utilization during a small test run to determine which resource is the bottleneck.

**Interruptible tasks.** Polishing chunks that can be retried without losing significant progress are good candidates for spot or preemptible instances. The [tick genomics study](https://pubmed.ncbi.nlm.nih.gov/40597572) capitalized on up to 90% discounted Spot instances for assembly and polishing. The workflow must handle interruptions gracefully by retrying failed tasks on new instances. The workflow manager tracks task status and can resubmit interrupted tasks automatically.

**Critical tasks.** Final validation steps and merging operations should run on on-demand instances to avoid interruptions. These tasks are typically short but essential. The cost difference between spot and on-demand instances for these tasks is small because the tasks complete quickly.

### Step 4: Design the Chunking Strategy Based on Genome Structure

The chunking strategy for parallel polishing should account for the structural features of the genome being assembled. Genomes with long repeats, such as the internal inverted repeats in the infectious laryngotracheitis virus genome, require careful handling at chunk boundaries.

The [ILTV vaccine study](https://doi.org/10.3390/vaccines14030245) found that the viral genome contains long internal inverted repeats that can give rise to genomic isomers, complicating short-read assembly and accurate resolution of genome structure. For genomes with similar structural complexity, chunk boundaries should be placed in regions of unique sequence instead of in repetitive regions. This placement reduces the risk of misassembly at chunk junctions.

For most large genomes, chunks of 1 to 5 megabases with overlaps of 10 to 50 kilobases provide a practical balance between parallelism and boundary accuracy. The overlap regions should be validated after merging to ensure that the polished chunks join seamlessly. If the genome has known structural features, such as long repeats or inverted regions, the chunk boundaries should be adjusted to avoid these features.

### Step 5: Establish Quality Thresholds Before Polishing

Define the quality thresholds that the polished assembly must meet before the polishing run begins. These thresholds should be based on the project requirements and the intended downstream analysis. The [Enterobacterales study](https://doi.org/10.1101/2025.09.15.676237) used quality value scores, with the best long-read only approach achieving a median quality value of 100. This score corresponds to zero errors per genome.

For large genomes, the quality thresholds may be less stringent because the cost of achieving perfect accuracy is higher. A plant genome of 10 gigabases will have more residual errors than a bacterial genome of 5 megabases, even with the same polishing approach. The thresholds should reflect the acceptable error rate for the downstream analysis.

The quality assessment should include both reference-based metrics, when a reference is available, and completeness metrics. The [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/) provide access to reference genomes for many species, which can be used to establish quality baselines. The [EMBL-EBI Training](https://www.ebi.ac.uk/training) portal provides guidance on using completeness assessment tools and interpreting their results.

### Step 6: Implement a Cost Monitoring and Control System

Cloud costs can escalate quickly if the polishing workflow is not monitored. Implement a cost monitoring system that tracks spending by task type, instance type, and data storage. The system should alert the research team when spending exceeds predefined thresholds.

**Compute cost tracking.** Cloud providers offer billing dashboards that show compute costs by instance type and time period. Configure these dashboards to show costs by project tag so that the polishing workflow costs are separated from other cloud usage.

**Storage cost tracking.** Storage costs accrue from the assembly files, read data, intermediate files, and final outputs. Implement data lifecycle policies that automatically move older files to cheaper storage tiers or delete temporary files after the project completes.

**Data transfer cost tracking.** Data transfer costs apply when moving data into and out of the cloud. These costs are typically small compared to compute costs but should be included in the budget.

**Budget alerts.** Set budget alerts that notify the research team when spending exceeds thresholds. The alerts should be configured at the project level and at the workflow level so that unexpected cost increases are detected early.

### Step 7: Document Decisions and Outcomes for Reproducibility

The decision framework should produce a record of the choices made and the rationale for each choice. This record is essential for reproducibility and for planning future polishing runs.

**Tool versions and parameters.** Record the exact versions of the polishing tools and the parameters used for each run. The [nf-core Documentation](https://nf-co.re/docs) emphasizes the importance of standardized pipeline documentation for reproducibility. Containerized software environments that pin specific tool versions ensure that the same tools are used across all tasks.

**Cloud configuration.** Record the instance types, the number of workers, the chunk sizes, and the storage configuration. This information is valuable for estimating the cost and time of future polishing runs on similar genomes.

**Quality assessment results.** Record the quality metrics before and after polishing. These metrics provide a quantitative basis for deciding whether additional polishing rounds are needed and for comparing the effectiveness of different polishing approaches.

**Cost and time data.** Record the total compute time, the total cost, and the cost by task type. This data is essential for budget planning for future projects.

### Common Failure Patterns in Tool and Configuration Selection

Several failure patterns recur when research teams select polishing tools and cloud configurations without a structured decision framework.

**Over-polishing with multiple tools.** Running multiple polishing tools in sequence without quality assessment between rounds can introduce errors and waste compute time. Each polishing round should be justified by quality assessment results.

**Under-provisioning memory.** Selecting instances with insufficient RAM causes out-of-memory failures and wasted compute time. The memory requirements should be estimated from the genome size and the polishing tool documentation before provisioning instances.

**Ignoring spot instance interruptions.** Workflows that do not handle spot instance interruptions gracefully can make little progress and waste money on retries. The workflow manager must be configured to retry failed tasks automatically.

**Inconsistent software versions across tasks.** If different workers use different versions of the polishing tools, the polished chunks may not be comparable. Containerized environments that pin specific tool versions prevent this problem.

**Missing quality gates.** Without quality assessment between polishing rounds, the workflow may continue polishing after the assembly has reached the quality threshold, wasting compute time and money.

### Professional Escalation Criteria for Tool and Configuration Decisions

Research teams should escalate tool and configuration decisions to experienced bioinformaticians or cloud architects in the following situations.

**Persistent quality problems after multiple polishing approaches.** If the polished assembly does not meet quality thresholds after trying multiple polishing tools and configurations, the problem may be in the assembly itself or in the read data. An experienced bioinformatician should review the assembly and read quality before further polishing attempts.

**Unexpected cost escalation.** If cloud costs exceed the budget by a significant margin despite cost monitoring, the polishing strategy should be reviewed. An expert may identify more cost-effective approaches, such as using different instance types or reducing the number of polishing rounds.

**Structural complexity that resists standard chunking.** If the genome has structural features that cause persistent chunk boundary artifacts, an expert should review the chunking strategy and may recommend alternative approaches.

**Reproducibility failures.** If the polishing workflow cannot be reproduced because of missing documentation or inconsistent software versions, an expert can help reconstruct the workflow and document it properly.

### Practical Implementation Steps

The following steps summarize the practical implementation of the decision framework.

1. Characterize the input data, including sequencing technology, basecalling quality, read depth, and assembly error profile.
2. Select polishing tools based on the error profile and the availability of short reads.
3. Assign task types to appropriate cloud instance categories, using spot instances for interruptible tasks and on-demand instances for critical tasks.
4. Design the chunking strategy based on genome structure, placing boundaries in unique sequence regions.
5. Establish quality thresholds before polishing and assess quality after each round.
6. Implement cost monitoring with budget alerts and data lifecycle policies.
7. Document all decisions, tool versions, cloud configurations, and quality results for reproducibility.

This framework provides a structured approach to selecting polishing tools and cloud configurations for large genome projects. The evidence from the Enterobacterales study, the tick genomics study, and the ILTV vaccine study demonstrates that careful tool selection and cloud configuration can produce high-quality assemblies at a fraction of the cost of traditional approaches. The framework is designed to be adaptable to different genome sizes, sequencing technologies, and project budgets.

## Frequently Asked Questions

### What is the difference between polishing and assembly?

Assembly is the process of constructing contiguous sequences from raw reads. Polishing is the subsequent process of correcting errors in those sequences by aligning reads back to the assembly and computing a consensus. Assembly produces the draft genome structure, while polishing improves the base-level accuracy of that structure.

### How many polishing rounds are needed for a large genome?

The number of polishing rounds depends on the assembly quality, the read type, and the desired accuracy. Some projects achieve sufficient accuracy with one round of long-read polishing, while others require additional rounds with short reads. Quality assessment after each round determines whether further polishing is needed.

### Can polishing be done without cloud computing?

Polishing can be done on local servers or workstations, but the compute time for large genomes can be substantial. Cloud computing offers the advantage of elastic scaling, allowing many polishing tasks to run in parallel and complete in a fraction of the wall time required on a single machine.

### What is the role of spot instances in cloud polishing?

Spot instances are discounted compute resources that can be interrupted with short notice. They are well suited to polishing workflows because individual tasks can be retried without losing the entire project. Spot instances can reduce compute costs by up to 90%, making large genome projects more affordable.

### How do I choose between Medaka and other polishing tools?

The choice of polishing tool depends on the sequencing technology and the error profile of the reads. Medaka is designed for Oxford Nanopore reads and uses a neural network model. Other tools, such as Polypolish and Pypolca, use short reads for polishing. The best approach may combine multiple tools in sequence.

### What quality metrics should I report after polishing?

Report the number of single nucleotide polymorphisms and indels compared to a reference, the quality value score, and the number of complete genes. These metrics provide a quantitative basis for comparing polishing approaches and for documenting the final assembly quality.

### How do I handle chunk boundary errors in parallel polishing?

Use overlapping regions of 10 to 50 kilobases at chunk boundaries and validate the merged assembly at these junctions. If boundary errors persist, increase the overlap size or use a different chunking strategy.

### What should I do if cloud costs exceed my budget?

Review the polishing strategy to identify cost reduction opportunities. Consider using spot instances, reducing the number of polishing rounds, or selecting more cost-effective instance types. Set budget alerts to catch cost overruns early.

## Related Bioinformatics Guides

- [Evaluating Genome Assembly Quality: Metrics and Tools](/knowledge/bioinformatics/evaluating-genome-assembly-quality-metrics-and-tools)
- [Metagenomics Assembly: Strategies for Reconstructing Microbial Genomes](/knowledge/bioinformatics/metagenomics-assembly-strategies-for-reconstructing-microbial-genomes)
- [Long-Read Genome Assembly and Polishing Strategies](/knowledge/bioinformatics/long-read-genome-assembly-and-polishing-strategies)
- [De Novo Genome Assembly with Long Reads: A Practical Workflow](/knowledge/bioinformatics/de-novo-genome-assembly-with-long-reads-a-practical-workflow)
- [Metagenomic Assembly and Binning: A Practical Workflow for Recovering Genomes from Complex Microbial Communities](/knowledge/bioinformatics/metagenomic-assembly-and-binning-a-practical-workflow-for-recovering-genomes-from-complex-microb)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)


## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [Tick genomics through a Nanopore: a low-cost approach for tick genomics.](https://pubmed.ncbi.nlm.nih.gov/40597572). BMC genomics, 2025.
- [Artificial intelligence in metagenome-assembled genome reconstruction: Tools, pipelines, and future directions.](https://pubmed.ncbi.nlm.nih.gov/41506577). Journal of microbiological methods, 2026.
- [Bridging the bioinformatics gap: tool selection for decentralized AMR genomic surveillance in Africa.](https://doi.org/10.3389/fpubh.2026.1756324). 2026.
- [Nanopore long-read only genome assembly of clinical Enterobacterales isolates is complete and accurate](https://doi.org/10.1101/2025.09.15.676237). 2025.
- [Hybrid Whole-Genome Sequencing for Genetic Stability Assessment of Infectious Laryngotracheitis Virus Vaccine Strains.](https://doi.org/10.3390/vaccines14030245). 2026.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.