# Polishing with Methylation-Aware Tools: Avoiding Errors in Epigenetically Modified Genomes


## Key Takeaways

- Methylated bases (e.g., 5mC, m6A) in genomes cause systematic signal changes during nanopore sequencing, which standard polishing tools misinterpret as sequencing errors, leading to false base corrections and corrupted assemblies.
- Methylation-induced errors are not random; they occur at specific sites and are not resolved by increasing sequencing depth, necessitating methylation-aware alignment or polishing strategies.
- Nanopore basecallers can output modification tags (MM/ML) that, when utilized by compatible polishing tools, allow for the protection of methylated sites from erroneous correction.
- Validation of polished assemblies from methylated genomes must be independent of the polishing process, utilizing data like Illumina short reads or reference genomes to detect systematic errors, particularly within coding sequences where 81% of errors were observed in one study.
- Species-specific optimization is critical, as polishing strategies and basecalling model performance can vary significantly, with older models sometimes yielding higher accuracy for certain bacterial species due to differential handling of modification signals.

---

## Direct Answer and Reader Context

Researchers assembling genomes from organisms with heavy DNA methylation, including many plants, bacteria, and fungi, face a specific problem during the polishing stage of genome assembly. Polishing tools compare raw sequencing reads to an initial assembly and correct base errors, but these tools assume every base mismatch between a read and the assembly represents a sequencing error. When a genome carries methylated bases such as 5-methylcytosine (5mC) or N6-methyladenine (m6A), the sequencing platform may produce systematic signal changes at those positions, and the polishing tool may interpret those changes as errors and introduce false corrections. The practical consequence is an assembly that looks polished but contains introduced mutations at methylation sites, which can corrupt gene annotations, variant calls, and downstream biological interpretation.

This article addresses that problem directly. It explains how methylation interferes with polishing, describes methylation-aware alignment and polishing strategies, and provides concrete workflow decisions for researchers working with methylated genomes. The guidance applies to genome assembly projects using long-read sequencing, particularly Oxford Nanopore Technologies (ONT) data, where base modification signals are embedded in the raw signal. The content is written for biology students, researchers, laboratory professionals, and life-science practitioners who need practical decisions instead of abstract theory.

The scope covers data inputs, workflow choices, controls, quality checks, reproducibility, interpretation limits, reporting, and practical decision criteria. The evidence base includes peer-reviewed studies on bacterial genome assembly accuracy, nanopore signal modeling for modification detection, polyploid plant genome assembly, and official documentation from bioinformatics training and workflow resources.

## Why Methylation Creates Polishing Errors

### The Mechanistic Basis of Methylation-Induced Mismatches

DNA methylation is a covalent modification where a methyl group is added to specific bases, most commonly cytosine at the C5 position (5mC) or adenine at the N6 position (m6A). These modifications are biologically abundant. In plants, methylation occurs across transposable elements, gene bodies, and intergenic regions. In bacteria, methylation is part of restriction-modification systems and can be widespread across the genome. The potato genome assembly study demonstrates that methylation divergence across haplotypes drives substantial allelic expression differentiation, indicating that methylation is not a rare event but a genome-wide feature in many organisms.

Nanopore sequencing detects bases by measuring ionic current changes as DNA passes through a protein pore. The current signal depends on the chemical identity of the bases in the pore, including modified bases. Methylated bases produce current signatures that differ from their unmodified counterparts. The unsupervised reference modeling study confirms that nanopore ionic current signals are sensitive to chemical modifications in DNA and RNA molecules, and that modified bases produce distinguishable signal patterns.

The problem for polishing arises because polishing tools compare base-called reads to the assembly sequence. Base calling converts raw electrical signals into nucleotide sequences, and the base caller may call a methylated cytosine as thymine or another base because the modified base produces a current signature that does not match the canonical cytosine model. When the polisher then aligns that read to the assembly, it sees a mismatch at the methylation site and assumes the assembly has an error. The polisher corrects the assembly to match the read, thereby introducing an error at a position where the assembly was actually correct.

### Evidence from Bacterial Genome Assembly Studies

The bacterial genome assembly study provides direct quantitative evidence for this phenomenon. The study sequenced six reference strains of highly pathogenic bacteria using ONT R10.4.1 chemistry and Illumina, then evaluated different assembly strategies against publicly available RefSeq assemblies as ground truth. The researchers found that methylation caused 6.5% of the observed errors in ONT assemblies. This is a substantial fraction when considering that the study also found 81% of observed errors were located within coding sequences, meaning methylation-induced errors can directly affect gene sequences.

The study also found that long-read polishing mainly improves assembly quality with only one round needed, but that polishing may also degrade assembly quality. This degradation is consistent with the mechanism described above: polishing tools that are not methylation-aware will systematically introduce errors at methylated positions. The study's finding that older basecalling models produced higher accuracy for certain species such as Brucella abortus suggests that newer models may be more sensitive to modification signals, which can be either beneficial or problematic depending on whether the downstream tools account for those signals.

### Why Standard Polishing Tools Fail on Methylated Genomes

Standard polishing tools operate on the assumption that the sequencing process is unbiased with respect to base identity. They model sequencing errors as random events with specific error profiles, and they use the depth of coverage and base quality scores to distinguish true errors from sequencing noise. Methylation breaks this assumption because the error is not random. It is systematic and position-specific, occurring at every site that carries a particular modification.

The systematic nature of methylation-induced errors means that increasing sequencing depth does not solve the problem. If every read at a methylated position produces the same miscall, then the polisher will see unanimous support for the wrong base. The error will be introduced with high confidence, and standard quality metrics will not flag it. This is why methylation-aware approaches are necessary instead of optional for organisms with heavy methylation.

## Methylation-Aware Alignment and Polishing Strategies

### Overview of Available Approaches

Several strategies exist for avoiding methylation-induced polishing errors. The choice depends on the sequencing platform, the availability of methylation information, and the specific polishing tool being used. The main approaches are:

1. Methylation-aware alignment that accounts for modified bases during read-to-assembly comparison
2. Masking or filtering modified positions before polishing
3. Using polishing tools that accept modification information as input
4. Using basecallers that output modification calls alongside canonical base calls
5. Validating polished assemblies against independent data such as Illumina short reads

The bacterial genome study used a combination of strategies and found that results varied by species. For Bacillus anthracis, an almost perfect assembly was achieved, while Brucella species assemblies contained five to 46 different nucleotides compared to Sanger-sequenced references. This variation highlights that no single strategy works universally, and researchers must evaluate their specific organism and data.

### Methylation-Aware Aligners

Methylation-aware aligners modify the alignment scoring to treat certain base substitutions as equivalent. For example, an aligner that knows a position is methylated may treat a C-to-T mismatch as a modification event instead of a sequencing error. This prevents the aligner from reporting a mismatch that would trigger a false correction during polishing.

The practical implementation depends on the aligner. Some aligners accept a methylation profile or modification BAM file that marks known modified positions. Others use signal-level information from nanopore data to identify modified bases before alignment. The unsupervised reference modeling study demonstrates that site-level anomaly score profiles can exhibit peak-like patterns that correspond to known modification-enriched regions, suggesting that modification detection can be performed without labeled training data.

### Modification-Aware Basecalling

Nanopore basecallers can be configured to output modification information alongside canonical base calls. This produces a BAM file with modified base tags (MM and ML tags in the SAM specification) that record the position and probability of modifications. Polishing tools that read these tags can avoid correcting positions where the modification signal explains the observed base call discrepancy.

The bacterial genome study's finding that enhanced basecalling models have generally improved assembly accuracy, but that older models produced higher accuracy for certain species, indicates that the interaction between basecalling and modification detection is complex. Researchers should test multiple basecalling models and compare assembly accuracy instead of assuming the newest model is always best.

### Polishing Tool Configuration

Some polishing tools accept parameters that control how they handle modified bases. These parameters may include options to ignore certain base substitutions, to use modification tags from the input BAM, or to skip positions with high modification probability. The specific options depend on the tool version and documentation.

The nf-core documentation provides guidance on configuring reproducible bioinformatics workflows, including polishing steps. Researchers using nf-core pipelines should check whether the pipeline version includes methylation-aware polishing options and how to enable them. The Galaxy Training Network offers accessible workflow training that can help researchers understand how to configure polishing tools within a graphical interface.

## Practical Workflow for Methylated Genomes

### Step 1: Assess Methylation Status Before Polishing

The first decision is whether the organism under study has significant methylation. This assessment can be based on literature, prior knowledge of the species, or preliminary analysis of the sequencing data. Organisms with known heavy methylation include many plants, bacteria with active restriction-modification systems, and some fungi.

For nanopore data, the raw signal can be examined for modification signatures. The unsupervised reference modeling study shows that models trained on unmodified sequences can be used to score candidate nucleotides using reconstruction error, and read-level signals can be aggregated to produce site-level modification evidence. This approach can identify modification-enriched regions without requiring labeled training data.

For organisms where methylation status is unknown, a pilot analysis of a small genomic region can reveal whether modification signals are present. If the pilot shows substantial modification evidence, the full assembly should use methylation-aware polishing.

### Step 2: Choose the Basecalling Model

The basecalling model determines whether modification information is available in the output. Models that output modified base tags provide the data needed for methylation-aware polishing. However, the bacterial genome study found that newer models do not always produce the most accurate assemblies, so model choice requires empirical testing.

The testing procedure should compare assembly accuracy across basecalling models using a common polishing strategy. For organisms with a reference genome, the polished assemblies can be compared directly to the reference. For organisms without a reference, the comparison can use assembly quality metrics such as completeness, contiguity, and the number of polishing-induced changes.

### Step 3: Select the Polishing Strategy

The polishing strategy should be selected based on the available modification information and the polishing tool. The options are:

- Standard polishing without modification awareness, followed by validation and manual correction of methylation sites
- Polishing with modification-aware alignment that treats modified positions appropriately
- Polishing with modification masking, where known modified positions are excluded from correction
- Polishing with modification-aware tools that accept MM and ML tags

The bacterial genome study's finding that one round of long-read polishing is sufficient suggests that multiple polishing rounds are unnecessary and may increase the risk of introducing errors. Researchers should plan for a single polishing round with careful parameter selection instead of iterative polishing.

### Step 4: Validate the Polished Assembly

Validation is essential for detecting methylation-induced errors. The validation strategy should include:

1. Comparison to independent data such as Illumina short reads
2. Examination of error distribution relative to known or predicted methylation sites
3. Assessment of coding sequence integrity, since the bacterial study found 81% of errors in coding sequences
4. Comparison to reference genomes when available

The bacterial genome study used publicly available RefSeq assemblies as ground truth for validation. Researchers without a reference can use the assembly itself to check for systematic errors, such as unexpected base composition at specific sequence contexts that match methylation motifs.

### Step 5: Document and Report Methylation-Aware Decisions

Reproducibility requires documenting the polishing strategy, including the basecalling model, polishing tool version, parameters, and modification handling approach. The nf-core documentation emphasizes reproducible workflow standards, and the Bioconductor project provides guidance on reproducible genomic analysis. The Carpentries lessons offer foundational training on version control and reproducible computing practices.

The documentation should record which positions were identified as methylated, how that information was used during polishing, and what validation was performed. This documentation allows other researchers to understand the assembly's limitations and to reproduce or extend the analysis.

## At a Glance: Polishing Decisions for Methylated Genomes

| Decision Point | Standard Approach | Methylation-Aware Approach | When to Choose Methylation-Aware |
|---|---|---|---|
| Basecalling model | Default model without modification output | Model with modified base tags (MM/ML) | Organism has known or suspected heavy methylation |
| Polishing rounds | Multiple rounds until convergence | Single round with careful parameter selection | Bacterial study shows one round is sufficient and more rounds may degrade quality |
| Alignment scoring | Treat all mismatches as errors | Treat C-to-T and A-to-G mismatches at modified positions as modification events | Modification calls are available from basecalling |
| Validation | Assembly quality metrics only | Additional comparison to independent short reads and methylation site examination | Methylation-induced errors are suspected or detected |
| Error correction policy | Correct all mismatches | Exclude or downweight positions with high modification probability | Modification probability data is available in BAM tags |

## Core Principles of Methylation-Aware Polishing

### Principle 1: Understand What the Polisher Sees

A polisher sees base-called reads aligned to an assembly. It does not see the raw signal or the modification state of the bases. If the base caller has already converted a methylated cytosine to thymine in the read sequence, the polisher has no way to know that the mismatch is biologically meaningful instead of a sequencing error. The only way to prevent false correction is to provide the modification information to the polisher or to prevent the polisher from correcting those positions.

### Principle 2: Modification Information Must Be Preserved Through the Pipeline

Modification information is lost if the pipeline does not carry it forward. If the basecaller outputs modification tags but the alignment step strips those tags, the polisher cannot use them. Researchers must verify that modification tags survive each pipeline step. The SAM specification defines MM and ML tags for this purpose, and tools that support these tags should preserve them through alignment and sorting.

### Principle 3: Validation Must Be Independent of the Polishing Process

Validation using the same data that was used for polishing cannot detect polishing-induced errors because the errors are supported by the data. Independent validation requires either different sequencing data, such as Illumina short reads, or biological knowledge about expected sequence features. The bacterial genome study's use of Sanger-sequenced references for Brucella species represents the gold standard for validation, but this is not available for most organisms.

### Principle 4: Error Distribution Matters More Than Total Error Count

The bacterial genome study found that 81% of errors in ONT assemblies were located within coding sequences. This distribution is particularly damaging because coding sequence errors can change predicted protein sequences and affect downstream functional analysis. When evaluating assembly quality, researchers should examine where errors occur, beyond how many errors exist. A small number of errors in coding sequences can be more damaging than a larger number in intergenic regions.

### Principle 5: Species-Specific Optimization Is Required

The bacterial genome study found that results varied by species, with some species achieving perfect genomes and others containing dozens of errors. This variation means that a polishing strategy that works for one organism may fail for another. Researchers should not assume that a published protocol for a related species will work without testing. The study's finding that older basecalling models produced higher accuracy for Brucella abortus demonstrates that newer is not always better.

## Data Inputs and Their Role in Polishing Decisions

### Long-Read Sequencing Data

Long-read data from Oxford Nanopore Technologies provides the primary input for assembly and polishing. The raw signal contains modification information that can be accessed through appropriate basecalling models. The veterinary pathogen detection review notes that nanopore sequencing enables near-complete genome assembly and identification of plasmid-borne antimicrobial resistance genes, demonstrating the utility of long reads for bacterial genomics.

The choice of sequencing chemistry affects modification detection. The bacterial genome study used R10.4.1 chemistry, which represents a recent version with improved accuracy. Researchers should be aware that different chemistry versions may have different modification signal characteristics and that basecalling models are typically developed for specific chemistry versions.

### Basecalled Reads with Modification Tags

Basecalled reads with MM and ML tags provide the modification information needed for methylation-aware polishing. The basecaller assigns a probability to each modified base call, and this probability can be used to decide whether a position should be protected from polishing correction.

The unsupervised reference modeling study demonstrates that modification detection can be performed without labeled training data by learning reference signal distributions from unmodified sequences. This approach is valuable for organisms where modification information is not available from the basecaller or where the basecaller's modification model is not trained for the specific organism.

### Reference Genomes for Validation

Reference genomes provide ground truth for validating polished assemblies. The bacterial genome study used publicly available RefSeq assemblies as ground truth, and the NCBI Data Resources provide access to these reference sequences. When a reference genome is available, researchers can directly measure polishing accuracy by comparing the polished assembly to the reference.

For organisms without a reference genome, validation is more challenging. Researchers can use assembly quality metrics, comparison to related species, or targeted validation of specific genomic regions using PCR and Sanger sequencing.

### Illumina Short Reads for Independent Validation

Illumina short reads provide an independent data source for validating polished assemblies. Because Illumina sequencing has different error characteristics than nanopore sequencing, errors introduced during nanopore polishing can be detected by aligning Illumina reads to the polished assembly and examining mismatches.

The bacterial genome study used both ONT and Illumina data, allowing comparison of assembly strategies. Researchers with access to both data types should use Illumina data for validation even if the assembly is based entirely on long reads.

## Workflow Options and Tradeoffs

### Option 1: Standard Polishing with Post-Hoc Methylation Correction

This approach uses standard polishing tools without modification awareness, then identifies and corrects methylation-induced errors after polishing. The advantage is simplicity and compatibility with existing pipelines. The disadvantage is that errors are introduced before they are corrected, and identifying all methylation-induced errors post hoc is difficult.

The post-hoc correction requires knowing which positions are methylated. This information can come from modification detection tools applied to the raw signal or from knowledge of methylation motifs in the organism. The correction process must distinguish true methylation-induced errors from genuine sequencing errors, which requires careful analysis.

### Option 2: Modification Masking Before Polishing

This approach identifies modified positions before polishing and masks them so the polisher does not correct them. Masking can be implemented by setting base qualities to zero at modified positions or by providing the polisher with a list of positions to exclude.

The advantage is that the polisher never sees the modified positions as errors, so false corrections are prevented. The disadvantage is that masking also prevents correction of genuine sequencing errors at those positions. If a modified position also has a true sequencing error, that error will remain in the assembly.

### Option 3: Modification-Aware Polishing Tools

This approach uses polishing tools that accept modification information and adjust their error model accordingly. These tools may treat a C-to-T mismatch at a modified position as a modification event instead of an error, or they may use the modification probability to weight the evidence for correction.

The advantage is that the polisher can still correct genuine errors at modified positions while avoiding false corrections. The disadvantage is that modification-aware polishing tools are less widely available and may have less mature implementations than standard tools.

### Option 4: Signal-Level Polishing

This approach uses the raw signal instead of base-called reads for polishing. Signal-level polishing can distinguish modified bases from sequencing errors because the signal characteristics are different. The unsupervised reference modeling study demonstrates that signal-level analysis can detect modifications, and this capability can be extended to polishing.

The advantage is the most accurate distinction between modifications and errors. The disadvantage is the computational cost and complexity of signal-level analysis, which requires specialized tools and substantial computing resources.

### Tradeoff Summary

| Approach | Error Prevention | Error Correction at Modified Sites | Tool Availability | Computational Cost |
|---|---|---|---|---|
| Standard polishing with post-hoc correction | Low | High after correction | High | Low |
| Modification masking | High | None at masked sites | Medium | Low |
| Modification-aware polishing | High | Medium | Low | Medium |
| Signal-level polishing | High | High | Low | High |

## Observations and Measurements for Polishing Quality

### Measuring Polishing-Induced Errors

The primary measurement for polishing quality is the number and location of differences between the polished assembly and the true genome sequence. When a reference genome is available, this measurement is straightforward. The bacterial genome study measured differences as the number of nucleotides that differed from Sanger-sequenced references, finding five to 46 differences for Brucella species.

When no reference is available, indirect measurements must be used. These include:

1. Read support for the polished sequence versus the unpolished sequence
2. Consistency between independent data sources
3. Biological plausibility of the polished sequence, such as absence of unexpected stop codons in coding sequences
4. Comparison to related species' genomes

### Tracking Modification Evidence

The modification evidence at each position should be tracked through the polishing process. This tracking includes the modification probability from the basecaller, the number of reads supporting the modification, and the consistency of modification calls across reads.

The unsupervised reference modeling study shows that site-level anomaly score profiles can exhibit peak-like patterns corresponding to known modification-enriched regions. These profiles can be used to identify regions where modification evidence is strong and where polishing errors are more likely.

### Recording Polishing Parameters

All polishing parameters should be recorded for reproducibility. These parameters include the polishing tool version, the alignment tool and parameters, the basecalling model, and any modification-related settings. The nf-core documentation provides standards for reproducible workflow configuration, and the Bioconductor project offers guidance on reproducible genomic analysis.

The record should also include the rationale for parameter choices. For example, if a researcher chose to mask modified positions, the record should explain why masking was preferred over modification-aware polishing and how the masked positions were identified.

## Records and Documentation Requirements

### Assembly Records

The assembly record should include the raw sequencing data accession numbers, the basecalling model and version, the assembly tool and parameters, and the polishing tool and parameters. The NCBI Data Resources provide systems for depositing and accessing sequencing data and assemblies, and researchers should follow the submission guidelines for their data types.

### Modification Records

The modification record should include the method used to detect modifications, the modification calls with their probabilities, and the genomic locations of modified positions. This record allows other researchers to understand which positions were protected from polishing correction and to evaluate whether the protection was appropriate.

### Validation Records

The validation record should include the validation data used, the validation method, and the results. For reference-based validation, the record should include the reference accession and the comparison method. For independent data validation, the record should include the data accession and the alignment parameters.

### Pipeline Configuration Records

The pipeline configuration record should include the exact commands used for each step, the software versions, and the environment in which the pipeline was run. The nf-core documentation provides standards for pipeline configuration, and the Carpentries lessons offer training on reproducible computing practices including version control.

## Common Failure Patterns in Polishing Methylated Genomes

### Failure Pattern 1: Systematic C-to-T Errors at Methylated CpG Sites

In organisms with CpG methylation, the most common failure pattern is systematic C-to-T errors at methylated cytosines. The polisher sees consistent support for thymine across all reads and introduces the error with high confidence. This pattern is detectable by examining the sequence context of errors, which will show enrichment at CpG dinucleotides.

### Failure Pattern 2: Errors Concentrated in Coding Sequences

The bacterial genome study found that 81% of errors were located within coding sequences. This concentration likely reflects the higher information content of coding sequences and the biological importance of these regions. Researchers should specifically examine coding sequences for polishing-induced errors, particularly at positions that change the predicted amino acid sequence.

### Failure Pattern 3: Species-Specific Polishing Failure

The bacterial genome study found that polishing results varied by species, with some species achieving perfect genomes and others containing dozens of errors. This variation means that a polishing strategy validated on one species cannot be assumed to work on another. Researchers should validate polishing quality for each new species instead of relying on prior experience with related species.

### Failure Pattern 4: Degradation from Multiple Polishing Rounds

The bacterial genome study found that one round of polishing was sufficient and that additional rounds could degrade assembly quality. This degradation likely occurs because each polishing round introduces new errors at modified positions, and these errors become fixed in the assembly. Researchers should avoid iterative polishing unless there is evidence that additional rounds improve quality.

### Failure Pattern 5: Newer Basecalling Models Producing Worse Results

The bacterial genome study found that older basecalling models produced higher accuracy for certain species such as Brucella abortus. This counterintuitive result may reflect differences in how models handle modification signals. Researchers should test multiple basecalling models instead of assuming the newest model is best.

### Failure Pattern 6: Modification Information Lost During Pipeline Processing

Modification tags can be lost during alignment, sorting, or filtering steps. If the polisher does not receive modification information, it will treat modified positions as errors. Researchers should verify that modification tags survive each pipeline step by inspecting the BAM files at each stage.

## Limitations of Methylation-Aware Polishing

### Limitation 1: Modification Detection Is Not Perfect

Modification detection from nanopore data has limited accuracy, particularly in complex biological samples. The unsupervised reference modeling study found that performance decreased in cell line samples when models trained on unmodified whole-genome-amplified DNA and in vitro-transcribed RNA were evaluated on natively modified data, reflecting the impacts of biological noise and heterogeneity. This means that some modified positions will be missed, and polishing errors at those positions will not be prevented.

### Limitation 2: Modification Status Can Vary Across Cells and Tissues

Methylation is not static. It varies across cell types, developmental stages, and environmental conditions. The potato genome study demonstrated that methylation divergence across haplotypes drives substantial allelic expression differentiation, indicating that methylation patterns are dynamic and haplotype-specific. A polishing strategy based on modification calls from one sample may not be appropriate for another sample from the same species.

### Limitation 3: Reference-Based Validation Is Not Always Available

The bacterial genome study used reference genomes for validation, but many organisms lack reference genomes. Without a reference, detecting polishing-induced errors requires indirect methods that are less sensitive. Researchers working with non-model organisms should be aware that their validation is less comprehensive than what is possible with a reference genome.

### Limitation 4: Computational Cost of Signal-Level Analysis

Signal-level analysis for modification detection and polishing is computationally expensive. The unsupervised reference modeling study used a CNN-Transformer variational autoencoder trained on large-scale data via streaming sampling and k-mer-aware soft balancing, which requires substantial computing resources. Researchers with limited computing resources may need to use base-called data with modification tags instead of signal-level analysis.

### Limitation 5: Tool Maturity and Documentation Gaps

Methylation-aware polishing tools are less mature than standard polishing tools, and documentation may be incomplete. The Bioconductor project and Galaxy Training Network provide training resources, but these may not cover the latest methylation-aware tools. Researchers should expect to invest time in understanding tool behavior and testing parameters.

## Safety and Regulatory Context

### Data Integrity and Reproducibility

The safety context for polishing methylated genomes is primarily about data integrity and reproducibility. Polishing-induced errors can propagate through downstream analyses, affecting gene annotations, variant calls, and biological interpretations. The nf-core documentation emphasizes reproducible workflow standards, and the Bioconductor project provides guidance on reproducible genomic analysis. Researchers should follow these standards to ensure that their results can be reproduced and verified.

### Clinical and Diagnostic Implications

For organisms with clinical or diagnostic relevance, polishing errors can have serious consequences. The bacterial genome study focused on highly pathogenic bacteria, including Bacillus anthracis, Brucella species, and Mycobacterium tuberculosis. Errors in assemblies of pathogenic organisms can affect genotyping, detection of genetic markers, and outbreak analysis. The veterinary pathogen detection review notes that nanopore sequencing is used for rapid whole-genome sequencing and outbreak tracing in field settings, where accuracy is critical.

The MGMT promoter methylation study in gliomas demonstrates that methylation status has clinical significance in human health, and errors in methylation detection can affect treatment decisions. While this study is not directly about genome assembly polishing, it illustrates the broader importance of accurate methylation information.

### Professional Escalation Criteria

Researchers should escalate to more specialized expertise when they encounter the following situations:

1. Polishing errors persist after applying methylation-aware strategies
2. The organism has unusual methylation patterns that are not handled by standard tools
3. The assembly is intended for clinical or regulatory use where accuracy requirements are stringent
4. The researcher lacks the computational resources or expertise to implement signal-level analysis
5. Validation reveals systematic errors that cannot be explained by known methylation patterns

In these situations, consultation with bioinformatics specialists, the tool developers, or the broader research community may be necessary. The Galaxy Training Network and EMBL-EBI Training provide pathways for developing the necessary skills, and the nf-core community offers support for pipeline configuration and troubleshooting.

## Professional Escalation Criteria

### When to Seek Specialized Assistance

The decision to escalate should be based on the impact of polishing errors on the research goals. For exploratory research where small numbers of errors are acceptable, standard approaches with validation may be sufficient. For research where accuracy is critical, such as clinical diagnostics or comparative genomics of closely related strains, specialized assistance may be required.

Specific escalation triggers include:

1. Detection of systematic errors at methylation motifs that persist after applying methylation-aware strategies
2. Inability to validate the polished assembly due to lack of independent data
3. Evidence that polishing has introduced errors in genes of biological interest
4. Plans to deposit the assembly in public databases where errors will be visible to the community
5. Use of the assembly for downstream applications where accuracy is critical, such as variant calling or phylogenetic analysis

### Resources for Escalation

The NCBI Data Resources provide access to reference genomes and validation data. The EMBL-EBI Training offers courses on bioinformatics analysis, including genome assembly and polishing. The Galaxy Training Network provides accessible workflow training that can help researchers understand and implement methylation-aware polishing. The nf-core documentation provides standards for reproducible workflow configuration. The Bioconductor project offers packages for genomic analysis, including modification detection and assembly validation.

## Frequently Asked Questions

### What is the difference between methylation-aware alignment and methylation-aware polishing?

Methylation-aware alignment modifies how reads are aligned to the assembly so that mismatches at modified positions are not treated as errors. Methylation-aware polishing modifies how the polishing tool uses aligned reads to correct the assembly. Alignment is the first step, and polishing is the second step. Both must be methylation-aware for the full benefit. If alignment treats modified positions as mismatches, the polisher will see those mismatches and may introduce false corrections. If alignment is methylation-aware but polishing is not, the polisher may still correct modified positions based on the aligned reads.

### How do I know if my organism has enough methylation to require methylation-aware polishing?

The bacterial genome study found that methylation caused 6.5% of errors in ONT assemblies, which is substantial for organisms with heavy methylation. If your organism is a plant, many bacteria, or some fungi, you should assume methylation is present and test for it. The test can be done by examining the raw nanopore signal for modification signatures or by comparing assemblies polished with and without methylation awareness. If the assemblies differ at positions that match known methylation motifs, methylation-aware polishing is needed.

### Can I use Illumina data to correct methylation-induced polishing errors?

Illumina data can be used to validate the polished assembly and identify methylation-induced errors. Because Illumina sequencing has different error characteristics than nanopore sequencing, errors introduced during nanopore polishing can be detected by aligning Illumina reads to the polished assembly and examining mismatches. However, Illumina data cannot be used directly to correct the assembly without careful analysis, because the Illumina reads may also have biases at methylated positions.

### What is the role of basecalling models in methylation-aware polishing?

The basecalling model determines whether modification information is available in the output. Models that output modified base tags provide the data needed for methylation-aware polishing. The bacterial genome study found that enhanced basecalling models have generally improved assembly accuracy, but that older models produced higher accuracy for certain species. This means the choice of basecalling model requires empirical testing for each organism.

### How many polishing rounds should I perform?

The bacterial genome study found that one round of long-read polishing was sufficient and that additional rounds could degrade assembly quality. This degradation likely occurs because each polishing round introduces new errors at modified positions. Researchers should plan for a single polishing round with careful parameter selection instead of iterative polishing.

### What validation should I perform after polishing a methylated genome?

Validation should include comparison to independent data such as Illumina short reads, examination of error distribution relative to known or predicted methylation sites, and assessment of coding sequence integrity. The bacterial genome study found that 81% of errors were located within coding sequences, so coding sequences should be specifically examined. When a reference genome is available, direct comparison to the reference provides the most reliable validation.

### Are there standard pipelines that include methylation-aware polishing?

The nf-core documentation describes community pipelines that follow reproducible workflow standards, and some of these pipelines may include methylation-aware polishing options. The Galaxy Training Network provides accessible workflow training that can help researchers understand how to configure polishing tools. Researchers should check whether the pipeline version includes methylation-aware options and how to enable them.

### What should I report about methylation-aware polishing in my methods section?

The methods section should report the basecalling model and version, the polishing tool and version, the alignment tool and parameters, and how modification information was used. The record should include which positions were identified as methylated, how that information was used during polishing, and what validation was performed. This documentation allows other researchers to understand the assembly's limitations and to reproduce or extend the analysis.

## Related Bioinformatics Guides

- [Genomic Data Analysis Tools: A Comparative Guide for Researchers](/knowledge/bioinformatics/genomic-data-analysis-tools-a-comparative-guide-for-researchers)
- [Binning in Metagenomics: From Contigs to Genomes](/knowledge/bioinformatics/binning-in-metagenomics-from-contigs-to-genomes)
- [Metagenomics Tools: A Practical Guide to Software and Pipelines](/knowledge/bioinformatics/metagenomics-tools-a-practical-guide-to-software-and-pipelines)
- [Explainable AI for Bioinformatics: Methods, Tools, and Applications](/knowledge/bioinformatics/explainable-ai-for-bioinformatics-methods-tools-and-applications)
- [Metagenomic Binning Tools Benchmark: How to Evaluate and Choose](/knowledge/bioinformatics/metagenomic-binning-tools-benchmark-how-to-evaluate-and-choose)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)


## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [Accurately assembling nanopore sequencing data of highly pathogenic bacteria.](https://pubmed.ncbi.nlm.nih.gov/40877756). BMC genomics, 2025.
- [MGMT promoter methylation across glioma subtypes: biological relevance, treatment response, and survival outcomes.](https://pubmed.ncbi.nlm.nih.gov/42313202). Molecular biology reports, 2026.
- [Unsupervised Reference Modeling of Nanopore Signals for DNA/RNA Modification Detection.](https://doi.org/10.3390/genes17050525). 2026.
- [Nanopore Sequencing in Veterinary Pathogen Detection: A Review of Technologies and Applications.](https://doi.org/10.3390/vetsci13030216). 2026.
- [Haplotype-resolved and near telomere-to-telomere assembly of the autotetraploid potato genome.](https://doi.org/10.1186/s13059-026-03980-9). 2026.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.