Polishing for Complete Genomes: Strategies for Telomere-to-Telomere Assembly Accuracy
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Standard genome polishing tools fail in complex repetitive regions (centromeres, telomeres, segmental duplications) due to ambiguous read alignments, leading to overcorrection or failure to correct errors. Repeat-aware strategies, like those used for the human CHM13 assembly, are crucial for accurate correction in these challenging genomic loci, improving Quality Value (QV) by fixing a significant percentage of existing errors.
- Achieving Telomere-to-Telomere (T2T) assembly accuracy necessitates a multi-platform sequencing approach, integrating data from technologies like Oxford Nanopore (ONT) ultra-long reads for spanning repeats, PacBio High-Fidelity (HiFi) reads for accuracy, and Illumina short reads for fine-tuning. This synergy corrects platform-specific error signatures and biases, which are critical for gapless genome reconstruction.
- Variant-aware polishing tools are essential for diploid or polyploid genomes to distinguish genuine allelic differences from sequencing errors, preventing haplotype switching. Tools like hypo-short and hypo-hybrid are designed to maintain distinct paternal and maternal haplotypes, thereby avoiding chimeric sequences and improving the accuracy of phased assemblies.
- Independent validation using data not involved in the polishing process is paramount for an unbiased assessment of assembly accuracy. Metrics such as mapping rates of raw reads to the final assembly and BUSCO completeness scores provide complementary measures of overall quality and gene content, respectively.
- Ultra-long reads (exceeding 100 kb) are critical for T2T assembly as they can span entire repeat arrays and flanking unique regions, enabling unambiguous read anchoring. Early planning for ultra-long read generation is advised due to specialized DNA extraction and sequencing protocols required to preserve high molecular weight DNA.
Telomere-to-telomere (T2T) genome assembly aims to produce chromosome-level sequences with no gaps, including complex repetitive regions such as centromeres, telomeres, and segmental duplications. Standard polishing approaches that work well for draft assemblies often fail in these regions, introducing errors or failing to correct existing ones. This article addresses the specific problem of achieving accurate polishing in repetitive and complex genomic regions, presenting strategies that combine multiple sequencing technologies, variant-aware tools, and targeted validation approaches. The content is intended for biology students, researchers, laboratory professionals, and life-science practitioners who are generating or refining complete genome assemblies.
The Polishing Problem in Complete Genome Assembly
Why Standard Polishing Falls Short in Repetitive Regions
Draft genome assemblies generated from Oxford Nanopore Technologies (ONT) long reads are known to have a higher error rate than assemblies produced from other platforms. Existing genome polishers can enhance their quality, but the error rate, including mismatches, indels, and switching errors between paternal and maternal haplotypes, can remain significant. The challenge intensifies when the goal is a complete, gapless assembly because repetitive regions present unique difficulties for error correction algorithms.
Standard polishing tools typically align reads back to the assembly and use the consensus of those alignments to identify and correct errors. In unique regions of the genome, this approach works well because reads map unambiguously. In repetitive regions, however, reads from different copies of a repeat can align to the wrong location, creating a false consensus that either fails to correct real errors or introduces new ones. This phenomenon, known as overcorrection, is a particular risk when polishing tools are applied to large tandem repeats and segmental duplications.
The first telomere-to-telomere human genome assembly, which resolved complex segmental duplications and large tandem repeats including centromeric satellite arrays, demonstrated this problem directly. Although derived from highly accurate sequences, evaluation revealed evidence of small errors and structural misassemblies in the initial draft assembly. The research team designed a new repeat-aware polishing strategy that made accurate assembly corrections in large repeats without overcorrection, ultimately fixing 51% of the existing errors and improving the assembly quality value from 70.2 to 73.9 measured from PacBio high-fidelity and Illumina k-mers [<a href="#ref-1">1</a>].
The Quality Value Metric and What It Means
Assembly quality value (QV) is a standard metric used to estimate the accuracy of a genome assembly. A QV of 50 corresponds to an estimated error rate of 1 in 100,000 bases, while a QV of 70 corresponds to 1 error in 10 million bases. For T2T assemblies, researchers typically aim for QV values above 50, with some projects achieving values exceeding 70.
The improvement from QV 70.2 to 73.9 in the human CHM13 assembly may appear modest numerically, but it represents a substantial reduction in the number of errors across a 3-gigabase genome. Each incremental improvement in QV requires increasingly sophisticated polishing strategies because the remaining errors are concentrated in the most difficult regions [<a href="#ref-1">1</a>].
The Gap Between Draft and Complete Assemblies
Draft assemblies often contain hundreds or thousands of gaps, particularly in repetitive regions that cannot be resolved with short reads alone. The transition from a draft to a T2T assembly requires also filling these gaps but also ensuring that the sequences filling them are accurate. Polishing plays a critical role in this transition because gap-filled sequences often contain errors introduced during the assembly process.
The Aspergillus fumigatus project illustrates this gap. Draft genome sequences of two clinical isolates lacked coverage of centromeres, an accurate sequence for ribosomal repeats, and a comprehensive annotation of chromosomal rearrangements such as translocations and inversions. The researchers used PacBio Single Molecule Real-Time (SMRT), Oxford Nanopore, and Illumina HiSeq sequencing for de novo genome assembly and polishing of two laboratory reference strains, generating full-length chromosome assemblies with comprehensive T2T coverage [<a href="#ref-2">2</a>].
Core Principles of Repeat-Aware Polishing
Understanding Sequencing Technology Biases
Different sequencing platforms produce characteristic error profiles, and these biases directly affect polishing outcomes. PacBio high-fidelity (HiFi) reads and Oxford Nanopore Technologies reads each have distinct error patterns that cause signature assembly errors. A diverse panel of sequencing technologies can correct these platform-specific biases [<a href="#ref-1">1</a>].
PacBio HiFi reads provide high accuracy with median read accuracy above 99.9%, but they have systematic errors in homopolymer regions and specific sequence contexts. ONT reads offer ultra-long read lengths that span complex repeats, but they have higher per-base error rates. Illumina short reads provide very high per-base accuracy but cannot span large repeats, limiting their utility for localizing errors in those regions.
The practical implication is that polishing strategies should incorporate data from multiple platforms instead of relying on a single technology. The Chinese soft-shelled turtle genome assembly used PacBio HiFi, Oxford Nanopore ultra-long reads, and Hi-C data to achieve a near-T2T chromosome-level assembly with no ambiguous bases and a QV of 41.19. Mapping rates for BGI short reads (99.4%), ONT long reads (100%), and HiFi reads (99.98%) confirmed high accuracy across platforms [<a href="#ref-3">3</a>].
The Role of Ultra-Long Reads in Polishing
Ultra-long reads, typically defined as reads exceeding 100 kilobases, provide the ability to span entire repeat arrays and unique flanking regions. This spanning capability allows polishers to anchor reads uniquely even when the repeat content itself is ambiguous. The Chinese soft-shelled turtle project demonstrated this approach by combining Oxford Nanopore ultra-long reads with PacBio HiFi data to resolve telomeres for 61 out of 68 chromosomal ends [<a href="#ref-3">3</a>].
For researchers planning a T2T assembly project, the decision to generate ultra-long reads should be made early because library preparation and sequencing protocols differ from standard long-read approaches. Ultra-long read generation requires careful DNA extraction to preserve high molecular weight DNA, and sequencing runs may need to be optimized for read length instead of throughput.
Variant-Aware Polishing Tools
Standard polishing tools treat all differences between reads and the assembly as potential errors to be corrected. Variant-aware tools recognize that some differences represent genuine biological variation, particularly in diploid or polyploid genomes where haplotypes differ. This distinction is critical for avoiding the introduction of switching errors, where the assembly incorrectly switches between paternal and maternal haplotypes.
The hypo-short and hypo-hybrid polishers were developed specifically to address this issue. Hypo-short utilizes Illumina short reads to polish an ONT-based draft assembly, resulting in a high-quality assembly with low error rates and switching errors. Hypo-hybrid incorporates ONT long reads to further refine the assembly into a diploid representation. The hypo-assembler pipeline automates the generation of highly accurate, contiguous, and nearly complete diploid assemblies using ONT long reads, Illumina short reads, and optionally Hi-C reads [<a href="#ref-4">4</a>].
For researchers working with diploid organisms, the choice between haploid and diploid polishing approaches has substantial consequences. Haploid polishing assumes a single consensus sequence and may collapse haplotype differences into errors. Diploid polishing maintains two haplotypes and requires tools that can distinguish true errors from allelic differences.
The Principle of Independent Validation
A fundamental principle in T2T polishing is that validation should use data not employed in the polishing process. This independent assessment provides an unbiased estimate of assembly accuracy. The Chinese soft-shelled turtle project validated their assembly by mapping BGI short reads (99.4% mapping rate), ONT long reads (100% mapping rate), and HiFi reads (99.98% mapping rate) to the final assembly [<a href="#ref-3">3</a>].
BUSCO analysis provides a complementary measure of completeness by assessing the presence of conserved single-copy orthologs. The Chinese soft-shelled turtle assembly achieved 97.9% BUSCO completeness, indicating that the assembly contains nearly all expected genes [<a href="#ref-3">3</a>].
At a Glance: Polishing Strategy Comparison
| Strategy | Input Data | Best Use Case | Known Limitations | Example Outcome |
|---|---|---|---|---|
| Standard short-read polishing | Illumina short reads | Draft assemblies with low repeat content | Overcorrects in large repeats, cannot resolve haplotype switching | Limited improvement in repetitive regions |
| Repeat-aware polishing | PacBio HiFi and Illumina k-mers | T2T assemblies with complex segmental duplications and centromeric arrays | Requires careful parameter tuning to avoid overcorrection | Fixed 51% of errors, QV improved from 70.2 to 73.9 in human CHM13 [<a href="#ref-1">1</a>] |
| Hybrid diploid polishing | ONT long reads plus Illumina short reads | Diploid genomes requiring haplotype resolution | Requires sufficient coverage of both data types | Produced fully phased T2T diploid genome of HG00733 with QV exceeding 50 [<a href="#ref-4">4</a>] |
| Multi-platform polishing | PacBio HiFi, ONT ultra-long, Hi-C, short reads | Near-T2T assemblies in non-model organisms | Higher cost and complexity, requires bioinformatics expertise | Chinese soft-shelled turtle assembly with no ambiguous bases and QV 41.19 [<a href="#ref-3">3</a>] |
Practical Workflow for T2T Polishing
Step 1: Assess the Starting Assembly
Before beginning polishing, evaluate the draft assembly to identify problem regions and establish baseline quality metrics. Generate a comprehensive quality report that includes contig statistics, QV estimates, and identification of regions with low read coverage or high repeat content. The NCBI provides official descriptions of databases and analysis services that can support this assessment, including tools for sequence alignment and variant calling [<a href="#ref-5">5</a>].
For researchers new to genome assembly quality assessment, the EMBL-EBI training materials offer structured learning pathways for bioinformatics data resources and practical analysis education [<a href="#ref-6">6</a>]. These resources can help laboratory professionals understand the metrics and tools used in assembly evaluation.
Step 2: Select Polishing Tools Based on Genome Characteristics
The choice of polishing tools depends on several factors: genome size, ploidy, repeat content, available sequencing data, and computational resources. For haploid or highly inbred organisms, standard polishers may suffice. For diploid organisms with significant heterozygosity, variant-aware tools such as hypo-short and hypo-hybrid are more appropriate [<a href="#ref-4">4</a>].
The Aspergillus fumigatus T2T assembly provides an example of a haploid organism with challenging repeat content. The researchers used PacBio Single Molecule Real-Time (SMRT), Oxford Nanopore, and Illumina HiSeq sequencing for de novo genome assembly and polishing of two laboratory reference strains. They generated full-length chromosome assemblies with comprehensive T2T coverage, including ribosomal repeats and centromere sequences composed of long transposon elements [<a href="#ref-2">2</a>].
Step 3: Run Initial Polishing Rounds
Begin with short-read polishing using Illumina data if available. This step corrects the majority of small errors, including single nucleotide polymorphisms and short indels. After each polishing round, reassess quality metrics to determine whether additional rounds are needed.
The Ogataea parapolymorpha DL-1 genome project illustrates the iterative nature of polishing. The researchers used long-read sequencing technology to achieve a T2T genome assembly, obtained high-quality reads, assembled de novo, and then performed polishing to enhance accuracy. The genome was analyzed to identify coding genes, telomeric motifs, rRNA genes, and methylation patterns [<a href="#ref-7">7</a>].
Step 4: Apply Repeat-Aware Polishing for Complex Regions
After initial polishing rounds, focus on regions that remain problematic. These typically include centromeres, telomeres, ribosomal DNA arrays, and segmental duplications. Repeat-aware polishing strategies that use k-mer-based validation can identify overcorrection and restore correct sequences.
The human CHM13 project demonstrated that comparing polishing results to standard automated tools reveals common polishing errors. The researchers offered practical suggestions for genome projects with limited resources, emphasizing that a diverse panel of sequencing technologies can correct platform-specific biases [<a href="#ref-1">1</a>].
Step 5: Validate with Independent Data
Final validation should use data not employed in the polishing process. This independent assessment provides an unbiased estimate of assembly accuracy. The Chinese soft-shelled turtle project validated their assembly by mapping BGI short reads (99.4% mapping rate), ONT long reads (100% mapping rate), and HiFi reads (99.98% mapping rate) to the final assembly [<a href="#ref-3">3</a>].
BUSCO analysis provides a complementary measure of completeness by assessing the presence of conserved single-copy orthologs. The Chinese soft-shelled turtle assembly achieved 97.9% BUSCO completeness, indicating that the assembly contains nearly all expected genes [<a href="#ref-3">3</a>].
Data Inputs and Quality Controls
Minimum Data Requirements for T2T Polishing
Successful T2T polishing requires sufficient coverage from multiple sequencing platforms. While exact coverage requirements vary by genome size and complexity, the following considerations apply:
For ONT-based assemblies, ultra-long reads provide the backbone for contiguity. The hypo-assembler pipeline demonstrated that ONT long reads combined with Illumina short reads can produce fully phased T2T diploid genomes with additional manual steps. The proof-of-concept assembly of HG00733 achieved a quality value exceeding 50 [<a href="#ref-4">4</a>].
For PacBio-based assemblies, HiFi reads provide high accuracy but may not span the largest repeat arrays. The Chinese soft-shelled turtle project combined PacBio HiFi with Oxford Nanopore ultra-long reads to achieve a contig N50 of 132.52 Mb and resolve telomeres for 61 out of 68 chromosomal ends [<a href="#ref-3">3</a>].
Hi-C data serves a dual purpose in T2T projects. It provides scaffolding information that guides chromosome-level assembly, and it can validate the correctness of long-range arrangements. The Chinese soft-shelled turtle project used Hi-C contact maps to guide scaffolding and curation, anchoring 40 contigs to 34 chromosomes [<a href="#ref-3">3</a>].
Quality Control Metrics to Track
Track the following metrics throughout the polishing process:
QV estimates from k-mer analysis provide a global accuracy measure. The human CHM13 project used PacBio high-fidelity and Illumina k-mers to measure QV improvements from 70.2 to 73.9 [<a href="#ref-1">1</a>].
Mapping rates indicate how well sequencing reads align to the assembly. Low mapping rates suggest structural errors or contamination. The Chinese soft-shelled turtle project reported mapping rates above 99.4% across all platforms [<a href="#ref-3">3</a>].
BUSCO completeness scores assess gene content completeness. The Chinese soft-shelled turtle assembly achieved 97.9% completeness, indicating that the assembly contains nearly all expected genes [<a href="#ref-3">3</a>].
Telomere resolution status indicates whether chromosome ends are fully assembled. The Chinese soft-shelled turtle project resolved telomeres for 61 out of 68 chromosomal ends [<a href="#ref-3">3</a>], while the Aspergillus fumigatus project achieved comprehensive T2T coverage [<a href="#ref-2">2</a>].
Record Keeping for Polishing Iterations
Maintain detailed records of each polishing iteration, including the tools used, parameters applied, input data versions, and resulting quality metrics. This documentation supports reproducibility and allows researchers to identify which steps contributed most to quality improvements.
The nf-core documentation provides standards for community pipeline usage, configuration, and reproducible workflow context [<a href="#ref-8">8</a>]. Adopting these standards can help laboratories maintain consistent polishing protocols across projects.
The Galaxy Training Network offers accessible workflow training and analysis tutorials that can help laboratory professionals implement reproducible polishing workflows [<a href="#ref-9">9</a>]. These resources provide practical guidance for researchers who may not have extensive command-line experience.
Common Failure Patterns in T2T Polishing
Overcorrection in Repetitive Regions
Overcorrection occurs when polishing tools introduce errors by incorrectly "correcting" sequences that were already accurate. This problem is particularly acute in large tandem repeats and segmental duplications where reads from different repeat copies align ambiguously. The human CHM13 project specifically designed a repeat-aware polishing strategy to make accurate assembly corrections in large repeats without overcorrection [<a href="#ref-1">1</a>].
Signs of overcorrection include decreased QV after polishing, reduced mapping rates, and discrepancies between k-mer-based and read-based accuracy estimates. If overcorrection is suspected, compare the polished assembly to the pre-polishing version in the affected regions to identify introduced changes.
Haplotype Switching Errors
In diploid genomes, polishing tools that do not account for haplotype variation may switch between paternal and maternal haplotypes within a single contig. These switching errors create chimeric sequences that do not match either parental genome. The hypo-short and hypo-hybrid polishers were specifically designed to minimize switching errors [<a href="#ref-4">4</a>].
Detection of switching errors requires comparison to known parental genotypes or phased variant data. In the absence of such data, examining read support across heterozygous sites can reveal abrupt transitions between haplotypes.
Platform-Specific Error Signatures
Each sequencing platform produces characteristic errors that can persist through polishing if only one data type is used. The human CHM13 project showed that sequencing biases in both high-fidelity and Oxford Nanopore Technologies reads cause signature assembly errors that can be corrected with a diverse panel of sequencing technologies [<a href="#ref-1">1</a>].
For example, ONT reads have systematic errors in homopolymer regions, while PacBio HiFi reads may have errors in specific sequence contexts. Using only one platform for polishing leaves these platform-specific errors uncorrected.
Incomplete Telomere Resolution
Telomeres present unique challenges because they consist of highly repetitive sequences that are difficult to assemble and polish. The Chinese soft-shelled turtle project resolved telomeres for 61 out of 68 chromosomal ends, leaving 7 ends unresolved [<a href="#ref-3">3</a>]. This outcome illustrates that even successful T2T projects may not achieve complete telomere resolution.
Strategies for improving telomere resolution include generating additional ultra-long reads targeting chromosome ends, using telomere-specific enrichment approaches, and applying targeted polishing to telomeric regions.
The Risk of Polishing Artifacts in Low-Complexity Sequence
Low-complexity regions, including homopolymers and simple sequence repeats, present a distinct challenge for polishing. These regions are prone to errors in both the assembly and the reads used for polishing. Polishing tools may introduce or perpetuate errors in these regions because the signal from correct bases is diluted by the repetitive nature of the sequence.
The Ogataea parapolymorpha DL-1 project encountered unusual telomere regulation based on the addition of non-telomeric dT, highlighting how low-complexity regions can harbor unexpected sequence features that complicate polishing [<a href="#ref-7">7</a>]. Researchers should examine low-complexity regions carefully after polishing to ensure that introduced changes are supported by multiple independent read alignments.
Advanced Polishing Strategies
Iterative Multi-Round Polishing
Single rounds of polishing rarely achieve optimal results. The human CHM13 project demonstrated that multiple rounds of polishing with different tools and data types progressively improved assembly quality [<a href="#ref-1">1</a>]. Each round should be followed by quality assessment to determine whether additional polishing is beneficial or whether overcorrection is occurring.
A practical approach is to alternate between short-read and long-read polishing, using each data type to correct errors that the other may miss. The hypo-assembler pipeline automates this process by using Illumina short reads for initial polishing and ONT long reads for diploid refinement [<a href="#ref-4">4</a>].
Targeted Polishing of Problem Regions
Instead of applying polishing uniformly across the genome, targeted approaches focus on regions that remain problematic after initial rounds. These regions can be identified through low mapping rates, high local error rates, or incomplete BUSCO matches.
The Aspergillus fumigatus project demonstrated the value of targeted approaches for specific genomic features. The researchers generated full-length chromosome assemblies with comprehensive T2T coverage, including ribosomal repeats and centromere sequences composed of long transposon elements [<a href="#ref-2">2</a>]. These regions required specialized assembly and polishing strategies.
Incorporating Epigenetic Data
Epigenetic modifications can provide additional information for polishing and validation. The Ogataea parapolymorpha DL-1 project used long-read sequencing to detect 5mC and 6mA modifications, which were validated through liquid chromatography-mass spectrometry. The study found an absence of 5mC DNA modification and the presence of 6mA in the genome [<a href="#ref-7">7</a>].
For researchers studying organisms with unusual epigenetic features, incorporating methylation data into the polishing process can improve accuracy. Methylation patterns can also help validate assembly correctness by confirming that modified bases are located in expected genomic contexts.
Manual Curation and Finishing
Despite advances in automated polishing, some regions require manual curation. The hypo-assembler pipeline allows for the production of T2T diploid genomes with additional manual steps [<a href="#ref-4">4</a>]. Manual curation involves examining read alignments in problem regions, resolving ambiguous mappings, and making targeted corrections.
The human CHM13 project's repeat-aware polishing strategy represents a semi-automated approach that combines algorithmic correction with human oversight. This strategy fixed 51% of existing errors while avoiding overcorrection in large repeats [<a href="#ref-1">1</a>].
Leveraging K-mer-Based Validation
K-mer-based validation provides a powerful tool for assessing polishing accuracy independent of read alignment. By comparing the k-mer spectrum of the assembly to that of the raw reads, researchers can identify regions where the assembly contains sequences not supported by the read data.
The human CHM13 project used PacBio high-fidelity and Illumina k-mers to measure QV improvements [<a href="#ref-1">1</a>]. This approach provides a quantitative measure of accuracy that complements read-based validation methods.
Tools and Resources for Polishing Workflows
Reproducible Workflow Management
Reproducibility is essential for polishing workflows, particularly when multiple tools and parameters are involved. The nf-core documentation provides standards for community pipeline usage, configuration, and reproducible workflow context [<a href="#ref-8">8</a>]. These standards help ensure that polishing results can be replicated across laboratories and projects.
The Galaxy Training Network offers accessible workflow training and analysis tutorials that can help researchers implement reproducible polishing workflows without extensive command-line experience [<a href="#ref-9">9</a>]. These resources provide practical guidance for laboratory professionals.
Learning Resources for Assembly and Polishing
The EMBL-EBI training materials offer structured learning pathways for bioinformatics data resources and practical analysis education [<a href="#ref-6">6</a>]. These resources can help researchers understand the underlying principles of genome assembly and polishing.
The Carpentries lessons provide foundational computing, data, shell, Git, and programming training context [<a href="#ref-10">10</a>]. These skills are essential for researchers who need to implement and customize polishing workflows.
The Bioconductor project provides official package, workflow, installation, and reproducible genomic-analysis documentation [<a href="#ref-11">11</a>]. Many polishing and validation tools are available as Bioconductor packages, and the documentation supports their proper use.
Database Resources for Validation
The NCBI provides official descriptions of databases, search systems, sequence resources, and analysis services [<a href="#ref-5">5</a>]. These resources support validation of polished assemblies through comparison to reference sequences and identification of remaining errors.
For researchers working with non-model organisms, NCBI databases can provide comparative data from related species that help identify assembly errors. The Chinese soft-shelled turtle project used known sex-linked SNP markers to validate sex chromosome inference, demonstrating the value of external validation data [<a href="#ref-3">3</a>].
Records and Measurements for Polishing Projects
Essential Records to Maintain
Maintain the following records throughout a T2T polishing project:
Assembly versions with corresponding quality metrics at each stage. This includes contig statistics, QV estimates, and BUSCO scores.
Sequencing data inventory, including platform, read length distribution, coverage, and quality scores for each data set used in polishing.
Tool versions and parameters for each polishing round. This information is essential for reproducing results and troubleshooting problems.
Validation results from independent data sets, including mapping rates and variant concordance.
Manual curation decisions and their rationale. This documentation supports future improvements and helps other researchers understand the assembly process.
Measurement Protocols for Quality Assessment
Establish consistent protocols for measuring assembly quality throughout the polishing process. Use the same metrics and methods at each stage to enable direct comparison of results.
For QV estimation, use k-mer-based methods that compare the assembly to sequencing reads. The human CHM13 project used PacBio high-fidelity and Illumina k-mers to measure QV improvements [<a href="#ref-1">1</a>].
For completeness assessment, use BUSCO analysis with appropriate lineage data sets. The Chinese soft-shelled turtle assembly achieved 97.9% BUSCO completeness, indicating that the assembly contains nearly all expected genes [<a href="#ref-3">3</a>].
For mapping rate assessment, align reads from each platform to the assembly and calculate the proportion that map successfully. The Chinese soft-shelled turtle project reported mapping rates of 99.4% for BGI short reads, 100% for ONT long reads, and 99.98% for HiFi reads [<a href="#ref-3">3</a>].
Benchmarking Against Reference Data
When a reference genome is available for the same or a closely related species, benchmarking the polished assembly against this reference provides a direct measure of accuracy. This comparison can identify errors that k-mer-based methods may miss, particularly in repetitive regions.
The Aspergillus fumigatus project generated full-length chromosome assemblies for two laboratory reference strains, CEA10 and A1160, providing a benchmark for future studies of this pathogen [<a href="#ref-2">2</a>]. Researchers working on other species should consider whether similar reference data are available for validation.
Limitations and Interpretation Constraints
What Polishing Cannot Fix
Polishing cannot correct structural misassemblies where sequences are incorrectly ordered or oriented. These errors require reassembly or manual curation instead of polishing. The human CHM13 project identified structural misassemblies in the initial draft assembly that required targeted correction [<a href="#ref-1">1</a>].
Polishing cannot recover sequences that are missing from the assembly. If a region was not assembled due to insufficient read coverage or extreme repeat content, polishing will not generate the missing sequence. Additional sequencing or targeted assembly approaches are required.
Polishing cannot resolve haplotype ambiguity in the absence of sufficient variant information. In diploid genomes, distinguishing true errors from allelic differences requires data that supports haplotype assignment.
Interpretation Limits of Quality Metrics
QV estimates provide a global measure of accuracy but do not identify specific error locations. A high QV does not guarantee that all regions are error-free, particularly in repetitive regions where errors may be concentrated.
BUSCO completeness scores assess gene content but do not evaluate non-coding regions. A high BUSCO score can coexist with substantial errors in intergenic and repetitive regions.
Mapping rates indicate how well reads align to the assembly but do not distinguish between correct and incorrect alignments in repetitive regions. High mapping rates can mask local misassembly.
The Challenge of Validating Repetitive Regions
Repetitive regions present a fundamental validation challenge because reads from different repeat copies are nearly identical. This makes it difficult to determine whether a polishing correction in a repeat region is correct or whether it has introduced an error by copying sequence from a different repeat copy.
The human CHM13 project addressed this challenge by designing a repeat-aware polishing strategy that made accurate assembly corrections in large repeats without overcorrection [<a href="#ref-1">1</a>]. This approach required careful attention to read alignment in repetitive regions and validation using multiple independent data types.
When to Escalate to Professional Support
Consider escalating to professional support or specialized services when:
Multiple polishing rounds fail to improve QV beyond a plateau. This situation may indicate structural misassemblies or systematic errors that require specialized approaches.
Telomere resolution remains incomplete after multiple attempts. The Chinese soft-shelled turtle project resolved 61 out of 68 chromosomal ends, and the remaining 7 ends may require specialized protocols [<a href="#ref-3">3</a>].
Haplotype switching errors persist despite using variant-aware tools. This situation may require additional sequencing data or manual curation.
The assembly is intended for clinical or regulatory applications where accuracy requirements exceed what standard polishing can achieve.
Safety and Regulatory Context
Data Management and Privacy Considerations
Genome assembly projects involving human data must comply with privacy and data protection regulations. The HG00733 genome used in the hypo-assembler proof-of-concept is part of the Human Genome Structural Variation Consortium and is publicly available [<a href="#ref-4">4</a>], but researchers working with human samples must ensure appropriate consent and data handling procedures.
For non-human organisms, consider whether the species is subject to any regulatory requirements. The Chinese soft-shelled turtle has economic and biomedical value, and its genome assembly may have implications for conservation or breeding programs [<a href="#ref-3">3</a>].
Reproducibility Standards
Adopting reproducible workflow standards supports scientific integrity and enables other researchers to verify results. The nf-core documentation provides standards for community pipeline usage, configuration, and reproducible workflow context [<a href="#ref-8">8</a>]. The Galaxy Training Network offers accessible workflow training that supports reproducible analysis [<a href="#ref-9">9</a>].
The Carpentries lessons provide foundational computing, data, shell, Git, and programming training context [<a href="#ref-10">10</a>]. These skills support the implementation of reproducible polishing workflows.
Ethical Considerations for Pathogen Genomics
For pathogen genomes such as Aspergillus fumigatus, the resulting reference genomes have implications for understanding pathogenicity and virulence. The T2T assembly of this fungus provides a fundamental resource for studying fungal biology and developing more effective treatments [<a href="#ref-2">2</a>]. Researchers working on pathogen genomes should consider the potential applications and ensure that their work aligns with ethical guidelines for pathogen research.
Professional Escalation Criteria
Indicators That Automated Polishing Is Insufficient
Seek specialized assistance when the following indicators are present:
QV remains below 40 after multiple polishing rounds with diverse data types. This level of accuracy may be insufficient for downstream applications such as comparative genomics or variant discovery.
BUSCO completeness remains below 95% despite adequate sequencing coverage. This result suggests either missing sequences or assembly errors that polishing cannot correct.
Mapping rates for high-quality reads remain below 95%. Low mapping rates indicate structural errors or contamination that require investigation.
Telomere resolution is achieved for fewer than 80% of chromosome ends. This outcome may indicate insufficient ultra-long read coverage or technical challenges with telomere assembly.
Situations Requiring Manual Curation
Manual curation is appropriate when automated polishing has plateaued but specific problem regions can be identified. The hypo-assembler pipeline allows for the production of T2T diploid genomes with additional manual steps [<a href="#ref-4">4</a>], suggesting that manual curation is expected for the highest-quality assemblies.
The human CHM13 project's repeat-aware polishing strategy involved human oversight of the correction process. This approach fixed 51% of existing errors while avoiding overcorrection in large repeats [<a href="#ref-1">1</a>].
Collaboration with Sequencing Facilities
When polishing challenges persist, collaboration with the sequencing facility that generated the data can be valuable. Sequencing facilities often have detailed knowledge of their platform's error profiles and can provide guidance on optimal polishing strategies.
The Chinese soft-shelled turtle project benefited from the combination of PacBio HiFi, Oxford Nanopore ultra-long reads, and Hi-C data [<a href="#ref-3">3</a>]. This multi-platform approach required coordination across sequencing facilities and demonstrates the value of collaborative data generation.
Frequently Asked Questions
What is the difference between standard polishing and repeat-aware polishing?
Standard polishing tools align reads to the assembly and use the consensus to correct errors. Repeat-aware polishing incorporates additional information to avoid overcorrection in repetitive regions where reads may align ambiguously. The human CHM13 project demonstrated that repeat-aware polishing fixed 51% of existing errors while avoiding overcorrection in large repeats, improving QV from 70.2 to 73.9 [<a href="#ref-1">1</a>].
How many rounds of polishing are typically needed for a T2T assembly?
The number of polishing rounds depends on the starting assembly quality and the complexity of the genome. The human CHM13 project required multiple rounds with different tools and data types to achieve optimal quality [<a href="#ref-1">1</a>]. Each round should be followed by quality assessment to determine whether additional polishing is beneficial or whether overcorrection is occurring.
Why do ONT-based assemblies need more polishing than PacBio-based assemblies?
Draft genomes generated from Oxford Nanopore Technologies long reads are known to have a higher error rate than assemblies from other platforms. Although existing genome polishers can enhance their quality, the error rate, including mismatches, indels, and switching errors between paternal and maternal haplotypes, can remain significant. The hypo-short and hypo-hybrid polishers were developed specifically to address this issue [<a href="#ref-4">4</a>].
What is a switching error in diploid genome polishing?
A switching error occurs when the assembly incorrectly switches between paternal and maternal haplotypes within a single contig. This creates a chimeric sequence that does not match either parental genome. The hypo-short and hypo-hybrid polishers were designed to minimize switching errors by using Illumina short reads for initial polishing and ONT long reads for diploid refinement [<a href="#ref-4">4</a>].
How can I tell if my polishing is causing overcorrection?
Signs of overcorrection include decreased QV after polishing, reduced mapping rates, and discrepancies between k-mer-based and read-based accuracy estimates. If overcorrection is suspected, compare the polished assembly to the pre-polishing version in the affected regions to identify introduced changes. The human CHM13 project specifically designed a repeat-aware polishing strategy to avoid overcorrection in large repeats [<a href="#ref-1">1</a>].
What sequencing data do I need for T2T polishing?
Successful T2T polishing requires sufficient coverage from multiple sequencing platforms. The hypo-assembler pipeline uses ONT long reads, Illumina short reads, and optionally Hi-C reads [<a href="#ref-4">4</a>]. The Chinese soft-shelled turtle project used PacBio HiFi, Oxford Nanopore ultra-long reads, and Hi-C data [<a href="#ref-3">3</a>]. A diverse panel of sequencing technologies can correct platform-specific biases [<a href="#ref-1">1</a>].
Can polishing fix structural misassemblies?
No, polishing cannot correct structural misassemblies where sequences are incorrectly ordered or oriented. These errors require reassembly or manual curation instead of polishing. The human CHM13 project identified structural misassemblies in the initial draft assembly that required targeted correction [<a href="#ref-1">1</a>].
What QV should I aim for in a T2T assembly?
The hypo-assembler proof-of-concept achieved a quality value exceeding 50 for a fully phased T2T diploid genome [<a href="#ref-4">4</a>]. The human CHM13 project improved QV from 70.2 to 73.9 through repeat-aware polishing [<a href="#ref-1">1</a>]. The Chinese soft-shelled turtle assembly achieved QV 41.19 [<a href="#ref-3">3</a>]. Target QV should be determined by the intended use of the assembly, with higher QV required for applications such as clinical variant discovery.
Related Bioinformatics Guides
- Metagenomics Assembly: Strategies for Reconstructing Microbial Genomes
- Metagenomic Assembly and Binning: A Practical Workflow for Recovering Genomes from Complex Microbial Communities
- Long-Read Sequencing for De Novo Assembly of Complex Genomes: Case Studies and Best Practices
- Long-Read Genome Assembly and Polishing Strategies
- Metagenome Co-Assembly: Strategies for Multi-Sample Data
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
[1] [Chasing perfection: validation and polishing strategies for telomere-to-telomere genome assemblies.](https://pubmed.ncbi.nlm.nih.gov/35361931). Nature methods, 2022. [2] [Telomere-to-telomere genome sequence of the model mould pathogen Aspergillus fumigatus.](https://pubmed.ncbi.nlm.nih.gov/36104328). Nature communications, 2022. [3] [A near-telomere-to-telomere genome assembly of the Chinese soft-shelled turtle (Pelodiscus sinensis).](https://pubmed.ncbi.nlm.nih.gov/41495059). Scientific data, 2026. [4] [Constructing telomere-to-telomere diploid genome by polishing haploid nanopore-based assembly.](https://pubmed.ncbi.nlm.nih.gov/38459383). Nature methods, 2024. [5] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information. [6] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute. [7] [Long-read sequencing reveals absence of 5mC in Ogataea parapolymorpha DL-1 genome and introduces telomere-to-telomere assembly.](https://pubmed.ncbi.nlm.nih.gov/40417237). Frontiers in genetics, 2025. [8] [nf-core Documentation](https://nf-co.re/docs). nf-core. [9] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project. [10] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries. [11] [Bioconductor](https://bioconductor.org/). Bioconductor Project.This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.