Telomere-to-Telomere Assembly of Repeats: How to Resolve Satellite DNA and Ribosomal RNA Arrays
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Resolving repetitive regions like telomeres, centromeres, and ribosomal RNA (rRNA) arrays necessitates ultra-long sequencing reads (e.g., >100 kb from Oxford Nanopore Technology or PacBio HiFi) to span multiple repeat units, overcoming the limitations of standard assemblers that fail on near-identical sequence blocks.
- Deep coverage (e.g., >30x) is critical for achieving high base accuracy (>99.99%) in repetitive regions, enabling the assembler to distinguish true sequence variation from sequencing errors and accurately reconstruct complex repeat structures.
- Specialized bioinformatics tools, such as Telomerecat for telomeric repeats and CentroMiner for centromeric satellites, are essential for identifying and resolving specific repeat types, but must be integrated into a broader assembly workflow for optimal results.
- A phased approach involving preliminary repeat landscape assessment, generation of appropriate ultra-long and high-fidelity sequencing data, selection of repeat-aware assemblers, and rigorous polishing/validation steps is required for successful telomere-to-telomere genome assembly.
- Complete resolution of repetitive regions yields substantial biological returns, including accurate studies of chromosome segregation (centromeres), cellular aging (telomeres), gene dosage (rRNA arrays), and the discovery of novel genes and structural variations previously hidden in draft assemblies.
Researchers attempting to assemble telomeric and centromeric regions frequently encounter persistent gaps in their genome assemblies because long tandem repeats confound standard assembly algorithms. This article provides a practical framework for resolving satellite DNA and ribosomal RNA arrays using ultra-long sequencing reads and specialized analysis tools, with concrete troubleshooting steps and quality control measures. The guidance applies to laboratory professionals, bioinformatics researchers, and biology students working with plant, animal, or human genome assemblies who need to move beyond fragmented draft assemblies toward complete telomere-to-telomere representations.
The Problem of Repetitive Regions in Genome Assembly
Genome assembly projects have historically treated repetitive regions as intractable obstacles. Telomeres, centromeres, and nucleolar organizer regions consist of long arrays of near-identical sequence units that confuse assembly algorithms designed for unique sequence. When a read contains multiple copies of the same repeat unit, the assembler cannot determine where that read belongs relative to other reads, producing collapsed or fragmented assemblies with gaps at precisely the regions that carry important biological information.
The consequences of these gaps extend beyond missing sequence. Without complete centromere assemblies, researchers cannot study the structural variation that influences chromosome segregation. Without complete telomere assemblies, studies of cellular aging and chromosome stability lose critical information. Without complete ribosomal RNA arrays, investigations of gene dosage and nucleolar organization remain incomplete. The shift toward telomere-to-telomere assembly represents a fundamental change in what researchers expect from a reference genome.
Recent projects demonstrate what becomes possible when these regions are resolved. A complete assembly of the maize genome traversed every chromosome in a single contig and revealed super-long simple-sequence-repeat arrays with consecutive thymine-adenine-guanine trinucleotide repeats extending up to 235 kilobases, along with the complete nucleolar organizer region containing 2,974 copies of the 45S ribosomal DNA [<a href="#ref-1">1</a>]. A gap-free sheep genome added 220.05 megabases of previously unresolved regions and 754 new genes to the existing reference assembly, correcting structural errors and improving variant detection in repetitive sequences [<a href="#ref-2">2</a>]. A pig telomere-to-telomere assembly uncovered 194.42 megabases of previously unresolved regions and added 1,189 new genes to the reference genome [<a href="#ref-3">3</a>]. These outcomes demonstrate that the effort required to resolve repeats produces substantial biological returns.
Core Principles of Repeat Resolution
Read Length Determines Assembly Success
The fundamental constraint in assembling tandem repeats is read length relative to repeat array length. If reads are shorter than the repeat units or cannot span multiple copies, the assembler cannot phase the copies correctly. Ultra-long reads from Oxford Nanopore Technology and high-fidelity reads from PacBio provide the read lengths needed to traverse repetitive arrays. The maize telomere-to-telomere project generated deep coverage with both ultralong Oxford Nanopore Technology reads and PacBio HiFi reads to achieve complete chromosome traversal [<a href="#ref-1">1</a>]. The cotton assembly similarly leveraged PacBio HiFi, Oxford Nanopore Technology ultralong reads, and Hi-C data to resolve 26 centromeric and 52 telomeric regions [<a href="#ref-4">4</a>].
Read length alone does not solve the problem. Coverage depth matters because repetitive regions require sufficient overlapping reads to distinguish true sequence variation from sequencing error. The maize project used deep coverage to achieve a base accuracy exceeding 99.99% [<a href="#ref-1">1</a>], while the sheep and pig assemblies reported accuracies above 99.999% [<a href="#ref-2">2</a>][<a href="#ref-3">3</a>]. These accuracy levels require both appropriate sequencing depth and careful polishing.
Assembly Strategy Must Account for Repeat Structure
Different repeat families present different assembly challenges. Centromeric satellite DNA consists of tandemly arrayed monomer units that can vary in copy number and sequence between individuals. Ribosomal RNA arrays contain many copies of the transcription unit separated by intergenic spacers, with copies arranged in tandem. Telomeric repeats are short hexamer or heptamer units repeated thousands of times. Each structure requires a slightly different assembly approach.
The sheep genome assembly identified four types of repeat units in centromeric regions, designated SatI, SatII, SatIII, and CenY [<a href="#ref-2">2</a>]. The pig assembly revealed pig-specific centromeric satellite repeat units and a unique structure in telocentric chromosomes with the arrangement telomere-SAT1B-satellite-SAT3-q_arm [<a href="#ref-3">3</a>]. The human telomere-to-telomere assembly characterized pericentromeric and centromeric repeats constituting 6.2% of the genome, or 189.9 megabases, and found multimegabase structural rearrangements including within active centromeric repeat arrays [<a href="#ref-5">5</a>]. These findings show that repeat composition varies substantially across species, so assembly strategies must be tailored to the specific repeat landscape of the organism under study.
Specialized Tools Address Specific Repeat Types
General-purpose assemblers often fail on repetitive regions because they cannot distinguish between different copies of the same repeat unit. Specialized tools target specific repeat structures. Telomerecat and similar tools identify telomeric repeat arrays by searching for the characteristic telomeric motif at read ends. CentroMiner and related approaches identify centromeric satellites by their tandem organization and copy number. These tools work best when integrated into a broader assembly workflow instead of used in isolation.
The choice of tool depends on the repeat type being targeted and the sequencing platform used. Tools designed for Oxford Nanopore Technology reads may not perform optimally on PacBio HiFi data and vice versa. Researchers should validate tool performance on their specific data before committing to a full assembly run.
At a Glance
| Assembly Challenge | Recommended Approach | Expected Outcome | Key Limitation |
|---|---|---|---|
| Telomeric repeat arrays | Ultra-long reads spanning multiple telomeric repeats, specialized telomere identification tools | Complete telomere sequences at chromosome ends | Telomere length varies between cells and individuals, so a single assembly represents one haplotype |
| Centromeric satellite arrays | Deep coverage with ultralong reads, satellite-specific identification tools, Hi-C for scaffolding | Complete centromere sequence with satellite copy number and arrangement | Centromere position and satellite composition can vary within species |
| Ribosomal RNA arrays | Ultra-long reads spanning multiple rDNA copies, copy number estimation, integration with nucleolar organizer region mapping | Complete rDNA array with copy number and arrangement | High sequence identity between copies complicates variant calling |
Practical Workflow for Repeat Resolution
Step 1: Assess the Repeat Landscape Before Sequencing
Before generating new sequencing data, examine what is already known about the repeat content of the target organism. Published assemblies, repeat databases, and cytogenetic maps provide information about the types and approximate sizes of repeat arrays. The NCBI maintains sequence databases and analysis services that can be used to examine existing assemblies and identify known repeat families [<a href="#ref-6">6</a>]. This preliminary assessment informs sequencing strategy, including read length targets and coverage depth.
For organisms with large centromeric satellite arrays, plan for ultralong reads that can span multiple satellite monomers. For organisms with moderate repeat content, standard long-read sequencing may suffice. The maize genome required deep coverage with both ultralong Oxford Nanopore Technology and PacBio HiFi reads because of its large genome size and extensive repetitive content [<a href="#ref-1">1</a>]. The cotton genome required a combination of PacBio HiFi, Oxford Nanopore Technology ultralong reads, and Hi-C to resolve its centromeric and telomeric regions [<a href="#ref-4">4</a>].
Step 2: Generate Appropriate Sequencing Data
Sequencing strategy directly determines assembly outcome. For telomere-to-telomere assembly, generate ultra-long reads with maximum read lengths exceeding the length of the largest repeat array. Oxford Nanopore Technology ultralong reads can exceed 100 kilobases, while PacBio HiFi reads provide high accuracy at moderate read lengths. The combination of both platforms provides complementary strengths: ultralong reads span repetitive arrays, while high-fidelity reads provide accurate base calls for polishing.
Coverage depth should be sufficient to ensure that every region of the genome, including repetitive regions, is covered by multiple independent reads. The maize project used deep coverage to achieve complete chromosome traversal [<a href="#ref-1">1</a>]. The sheep and pig assemblies achieved base accuracy above 99.999% [<a href="#ref-2">2</a>][<a href="#ref-3">3</a>], which requires substantial coverage for error correction.
Step 3: Select and Configure Assembly Software
Choose an assembler that supports the read types generated and can handle repetitive regions. Many modern assemblers incorporate specific strategies for repetitive sequence, including graph-based approaches that preserve repeat structure. Configure the assembler with parameters appropriate for the genome size and repeat content of the target organism.
The Galaxy Training Network provides accessible workflow training and analysis tutorials that can help researchers learn assembly procedures and understand how to configure tools appropriately [<a href="#ref-7">7</a>]. The nf-core documentation describes community pipeline standards and usage patterns that support reproducible assembly workflows [<a href="#ref-8">8</a>]. These resources help researchers move from generic assembly approaches to repeat-aware strategies.
Step 4: Identify and Resolve Telomeric Repeats
Telomeric repeats consist of short motifs repeated hundreds to thousands of times at chromosome ends. Standard assemblers often fail to extend contigs through these arrays because the reads cannot be uniquely placed. Specialized telomere identification tools search for the characteristic telomeric motif at contig ends and extend assemblies through the repeat array.
After initial assembly, check every contig for telomeric motifs at both ends. Contigs with telomeric sequence at one end represent chromosome ends. Contigs with telomeric sequence at both ends may represent complete chromosomes or may indicate assembly errors. The maize telomere-to-telomere assembly achieved complete chromosome traversal with each chromosome in a single contig [<a href="#ref-1">1</a>], demonstrating that complete telomere resolution is achievable with appropriate data and methods.
Step 5: Identify and Resolve Centromeric Repeats
Centromeric regions contain species-specific satellite DNA arrays that can extend for megabases. The human telomere-to-telomere assembly found that pericentromeric and centromeric repeats constitute 6.2% of the genome [<a href="#ref-5">5</a>]. The sheep assembly identified four types of centromeric repeat units [<a href="#ref-2">2</a>], while the pig assembly revealed pig-specific centromeric satellite repeats [<a href="#ref-3">3</a>]. These species-specific differences mean that centromere identification tools must be configured for the target organism.
Centromere resolution typically requires a combination of approaches. Satellite identification tools detect tandemly arrayed monomers. Hi-C data can place centromeric contigs within chromosome scaffolds. Comparative analysis with related species can confirm centromere identity. The cotton assembly discovered a phylogenetically recent centromere repositioning on chromosome D08, involving deactivation of an ancestral centromere and formation of a new satellite repeat-based centromere [<a href="#ref-4">4</a>], demonstrating that centromere position is not static across evolutionary time.
Step 6: Resolve Ribosomal RNA Arrays
Ribosomal RNA genes exist in tandem arrays that can span hundreds of kilobases. The maize assembly resolved the complete nucleolar organizer region of 26.8 megabases with 2,974 copies of 45S ribosomal DNA, revealing complex patterns of duplications and transposon insertions [<a href="#ref-1">1</a>]. The cotton assembly resolved 5S ribosomal DNA clusters and nucleolar organizer regions [<a href="#ref-4">4</a>].
Ribosomal RNA array assembly requires reads that span multiple copies of the transcription unit. The high sequence identity between copies means that assembly errors can easily occur if reads are too short. Copy number estimation provides a quality check: the assembled array length should be consistent with the expected copy number based on other evidence.
Step 7: Polish and Validate the Assembly
After initial assembly, polishing improves base accuracy. The maize assembly achieved base accuracy over 99.99% [<a href="#ref-1">1</a>], while the sheep and pig assemblies exceeded 99.999% [<a href="#ref-2">2</a>][<a href="#ref-3">3</a>]. Polishing involves aligning reads back to the assembly and correcting errors, typically using both long reads and short reads for complementary error correction.
Validation requires multiple approaches. Check that all expected chromosomes are represented as complete contigs. Verify that telomeric motifs appear at chromosome ends. Confirm that centromeric satellites show the expected organization. Compare assembly statistics against published assemblies for the same or related species. The sheep assembly corrected several structural errors in previous reference assemblies [<a href="#ref-2">2</a>], demonstrating that validation can identify and fix problems in earlier work.
Options and Tradeoffs in Assembly Approaches
Sequencing Platform Selection
Oxford Nanopore Technology ultralong reads provide the longest read lengths, which are essential for spanning the largest repeat arrays. The maize assembly used deep coverage ultralong Oxford Nanopore Technology reads to achieve complete chromosome traversal [<a href="#ref-1">1</a>]. PacBio HiFi reads provide higher per-base accuracy, which simplifies polishing and improves base-level accuracy. The cotton assembly used both platforms [<a href="#ref-4">4</a>], and the maize assembly also combined both [<a href="#ref-1">1</a>].
The tradeoff between platforms involves cost, throughput, and accuracy. Ultralong reads require specialized library preparation and can have lower throughput. HiFi reads provide high accuracy but shorter read lengths. Many telomere-to-telomere projects use both platforms to combine the strengths of each.
Assembly Software Selection
Different assemblers implement different strategies for handling repetitive regions. Some use overlap-layout-consensus approaches that work well with long reads. Others use de Bruijn graphs that work well with short reads but struggle with long repeats. The choice of assembler should be based on the read types available and the repeat content of the target genome.
The Bioconductor project provides official package, workflow, and installation documentation for reproducible genomic analysis [<a href="#ref-9">9</a>], which can help researchers select and configure appropriate tools. The nf-core documentation describes community pipeline standards that support reproducible assembly workflows [<a href="#ref-8">8</a>].
Coverage Depth Considerations
Higher coverage depth improves assembly quality but increases cost. The maize project used deep coverage to achieve complete chromosome traversal [<a href="#ref-1">1</a>]. The sheep and pig assemblies achieved base accuracy above 99.999% [<a href="#ref-2">2</a>][<a href="#ref-3">3</a>], which requires substantial coverage for error correction.
For repetitive regions, coverage depth matters more than for unique sequence because the assembler needs multiple independent reads spanning each repeat junction. If coverage is too low, the assembler cannot distinguish true repeat structure from sequencing error.
Observations and Measurements for Quality Assessment
Assembly Statistics
Standard assembly statistics provide a first indication of assembly quality. Contig N50 measures the contig length at which half the assembly is in contigs of that length or longer. For telomere-to-telomere assemblies, the goal is chromosome-scale contigs, with each chromosome represented as a single contig. The maize assembly achieved this with each chromosome entirely traversed in a single contig [<a href="#ref-1">1</a>].
Base accuracy measures the proportion of correctly called bases. The maize assembly achieved over 99.99% accuracy [<a href="#ref-1">1</a>], while the sheep and pig assemblies exceeded 99.999% [<a href="#ref-2">2</a>][<a href="#ref-3">3</a>]. These accuracy levels require both appropriate sequencing depth and careful polishing.
Completeness Assessment
Completeness measures the proportion of the genome represented in the assembly. The sheep assembly added 220.05 megabases of previously unresolved regions to the existing reference [<a href="#ref-2">2</a>]. The pig assembly uncovered 194.42 megabases of previously unresolved regions [<a href="#ref-3">3</a>]. These additions represent sequence that was missing from earlier assemblies.
Completeness can be assessed by comparing the assembly against expected genome size, by checking for the presence of conserved single-copy genes, and by verifying that all expected chromosomes are represented. The human telomere-to-telomere assembly enabled comprehensive characterization of pericentromeric and centromeric repeats constituting 6.2% of the genome [<a href="#ref-5">5</a>], demonstrating the value of complete assemblies for repeat characterization.
Repeat Array Measurements
For telomeric repeats, measure the length of the telomeric array at each chromosome end. Telomere length varies between cells and individuals, so the assembly represents one haplotype. For centromeric satellites, measure the copy number and arrangement of satellite monomers. The human assembly found multimegabase structural rearrangements in centromeric regions [<a href="#ref-5">5</a>]. For ribosomal RNA arrays, measure the copy number and arrangement of transcription units. The maize assembly found 2,974 copies of 45S ribosomal DNA in the nucleolar organizer region [<a href="#ref-1">1</a>].
Records and Documentation for Reproducibility
Data Management
Document all sequencing data, including platform, read length, coverage depth, and quality metrics. Store raw reads in standard formats with appropriate metadata. The NCBI provides sequence databases and analysis services for storing and accessing genomic data [<a href="#ref-6">6</a>]. The EMBL-EBI Training program provides learning pathways for data-resource training and practical analysis education [<a href="#ref-10">10</a>].
Workflow Documentation
Document every step of the assembly workflow, including software versions, parameters, and intermediate files. The nf-core documentation describes community pipeline standards that support reproducible workflow configuration [<a href="#ref-8">8</a>]. The Galaxy Training Network provides accessible workflow training and analysis tutorials [<a href="#ref-7">7</a>]. The Carpentries lessons provide foundational computing, data, shell, Git, and programming training that supports reproducible research practices [<a href="#ref-11">11</a>].
Version Control
Use version control for all scripts and configuration files. Record software versions and parameter settings. The Carpentries lessons include Git training that supports version control for research software [<a href="#ref-11">11</a>]. Reproducibility requires that another researcher can repeat the assembly workflow and obtain the same result.
Common Failure Patterns and Troubleshooting
Failure Pattern 1: Contigs End at Repeat Boundaries
When contigs consistently end at the boundaries of repeat arrays, the assembler cannot extend through the repeats. This pattern indicates that reads are too short to span the repeat array or that coverage is insufficient. Solutions include generating longer reads, increasing coverage depth, or using specialized repeat-resolution tools.
The maize assembly required deep coverage ultralong Oxford Nanopore Technology reads to traverse repetitive regions [<a href="#ref-1">1</a>]. If your assembly shows contig breaks at repeat boundaries, examine the read length distribution and coverage depth in those regions.
Failure Pattern 2: Collapsed Repeats
When the assembler merges multiple copies of a repeat into a single copy, the assembly underestimates repeat copy number. This pattern produces an assembly that is shorter than expected and missing sequence. Collapsed repeats are difficult to detect without independent evidence of repeat copy number.
The sheep assembly corrected several structural errors in previous reference assemblies [<a href="#ref-2">2</a>], demonstrating that earlier assemblies contained collapsed or misassembled repeats. Comparison with independent estimates of repeat content can identify collapsed regions.
Failure Pattern 3: Misplaced Repeats
When repeats are placed at incorrect genomic locations, the assembly has structural errors. This pattern can be detected by comparing the assembly against genetic maps, Hi-C data, or other independent evidence of genome structure. The cotton assembly discovered a centromere repositioning on chromosome D08 [<a href="#ref-4">4</a>], demonstrating that centromere position can differ from expectations based on related species.
Failure Pattern 4: Base Errors in Repeat Arrays
When base errors accumulate in repetitive regions, the assembly has lower accuracy in those regions. This pattern can be detected by comparing the assembly against reads that span the repeat array. The maize assembly achieved over 99.99% base accuracy [<a href="#ref-1">1</a>], while the sheep and pig assemblies exceeded 99.999% [<a href="#ref-2">2</a>][<a href="#ref-3">3</a>]. If your assembly has lower accuracy in repeat regions, additional polishing may be needed.
Limitations and Interpretation Boundaries
Single Haplotype Representation
A telomere-to-telomere assembly represents one haplotype of one individual. Telomere length varies between cells and individuals. Centromere position and satellite composition can vary within species. The human assembly found high degrees of structural, epigenetic, and sequence variation in centromeric regions across individuals [<a href="#ref-5">5</a>]. Researchers should not assume that the assembly represents all individuals in a species.
Technical Limitations
Ultra-long read sequencing requires specialized equipment and expertise. Not all sequencing facilities offer ultra-long read protocols. The cost of deep coverage with multiple platforms can be substantial. Researchers should assess whether telomere-to-telomere assembly is necessary for their research question or whether a draft assembly suffices.
Interpretation Limitations
Complete assemblies enable new analyses but do not automatically resolve all biological questions. The pig assembly identified new genes and variants in previously unresolved regions [<a href="#ref-3">3</a>], but functional validation requires additional experiments. The cotton assembly identified an early-maturing haplotype associated with an 11 megabase pericentric inversion [<a href="#ref-4">4</a>], but the functional significance requires experimental testing.
Safety and Regulatory Context
Data Sharing and Privacy
Genome assemblies from human samples raise privacy concerns. The human telomere-to-telomere assembly used a CHM13 cell line [<a href="#ref-5">5</a>], which avoids some privacy issues but does not represent a typical human genome. Researchers working with human samples should follow applicable data sharing and privacy regulations. The NCBI provides guidance on data submission and access [<a href="#ref-6">6</a>].
Ethical Considerations
Genome assembly of agricultural species has implications for breeding and genetic improvement. The sheep assembly identified variants associated with wool fineness [<a href="#ref-2">2</a>], and the pig assembly identified genes associated with body stature [<a href="#ref-3">3</a>]. These findings could inform breeding programs but should be validated before application.
Biosafety
Sequencing and assembly work typically involves standard molecular biology safety practices. No specific biosafety concerns apply to genome assembly beyond those associated with the source organism and sequencing protocols.
Professional Escalation Criteria
When to Seek Additional Expertise
If your assembly consistently fails to resolve repeat arrays despite appropriate data and methods, consider consulting with researchers who have completed telomere-to-telomere assemblies. The published telomere-to-telomere projects [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>][<a href="#ref-5">5</a>][<a href="#ref-3">3</a>][<a href="#ref-4">4</a>] represent successful approaches that can inform your strategy.
When to Reconsider the Approach
If the cost of telomere-to-telomere assembly exceeds the research budget, consider whether a draft assembly with targeted repeat resolution might suffice. Some research questions require complete assemblies, while others can be addressed with targeted sequencing of specific repeat regions.
When to Validate with Independent Methods
If assembly results have major implications for downstream analyses, validate with independent methods. The sheep assembly corrected structural errors in previous reference assemblies [<a href="#ref-2">2</a>], demonstrating that even established assemblies can contain errors. Independent validation methods include genetic mapping, cytogenetic analysis, and comparison with related species.
Decision Framework for Selecting Repeat Resolution Strategies
Step 1: Classify the Repeat Array by Structure and Scale
Before choosing tools or generating additional data, classify each unresolved region according to its repeat unit size, array length, and sequence divergence. This classification determines which resolution strategy will succeed. The published telomere-to-telomere projects demonstrate that repeat structures vary substantially across species, so a strategy that worked for one organism may fail for another.
Measure the repeat unit length from available draft assemblies or published repeat databases. Telomeric repeats in vertebrates use the hexamer TTAGGG, while plant telomeres often use the heptamer TTTAGGG. Centromeric satellites range from 150 to 180 base pair monomers in humans [<a href="#ref-5">5</a>] to larger satellite units in sheep and pigs [<a href="#ref-2">2</a>][<a href="#ref-3">3</a>]. Ribosomal RNA arrays contain transcription units that can exceed 10 kilobases, as seen in the maize 45S ribosomal DNA array [<a href="#ref-1">1</a>].
Estimate the total array length using pulsed-field gel electrophoresis, quantitative PCR, or read depth analysis from existing sequencing data. The maize nucleolar organizer region spans 26.8 megabases with 2,974 copies of 45S ribosomal DNA [<a href="#ref-1">1</a>]. The human centromeric and pericentromeric repeats constitute 189.9 megabases [<a href="#ref-5">5</a>]. Arrays of this scale require different strategies than arrays spanning only a few kilobases.
Record the sequence divergence between repeat copies. High-identity arrays with less than 1% divergence between copies present the greatest assembly challenge because reads cannot be uniquely placed. Arrays with higher divergence allow assemblers to distinguish individual copies. The sheep assembly identified four distinct repeat unit types in centromeric regions [<a href="#ref-2">2</a>], suggesting that some arrays contain enough variation for copy-specific placement.
Step 2: Match Strategy to Array Classification
For short arrays under 50 kilobases with moderate divergence, standard long-read assembly with PacBio HiFi or Oxford Nanopore Technology reads may suffice. The cotton assembly used PacBio HiFi and Oxford Nanopore Technology ultralong reads to resolve 26 centromeric and 52 telomeric regions [<a href="#ref-4">4</a>], demonstrating that these platforms handle a range of array sizes.
For arrays between 50 and 500 kilobases with high sequence identity, targeted read extension strategies work best. Generate ultralong reads that span the entire array, then use the assembly graph to extend contigs through the repeat. The maize project generated deep coverage ultralong Oxford Nanopore Technology reads to traverse repetitive regions [<a href="#ref-1">1</a>]. If reads cannot span the full array, use graph-based assemblers that preserve repeat structure through branching paths.
For arrays exceeding 500 kilobases with high sequence identity, combine multiple evidence types. The maize assembly required both ultralong Oxford Nanopore Technology and PacBio HiFi reads to resolve the 26.8 megabase nucleolar organizer region [<a href="#ref-1">1</a>]. The cotton assembly added Hi-C data to resolve centromeric regions [<a href="#ref-4">4</a>]. Hi-C provides long-range proximity information that places contigs within chromosomes even when sequence-based assembly fails.
For arrays with known species-specific satellite structure, use targeted satellite identification tools. The sheep assembly identified SatI, SatII, SatIII, and CenY repeat units in centromeric regions [<a href="#ref-2">2</a>]. The pig assembly revealed pig-specific centromeric satellite repeat units [<a href="#ref-3">3</a>]. Tools that search for known satellite monomers can identify and extend arrays that general assemblers collapse.
Step 3: Apply the Decision Matrix for Tool Selection
Use the following decision matrix to select tools based on array characteristics. This matrix consolidates the practical experience from published telomere-to-telomere projects [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>][<a href="#ref-5">5</a>][<a href="#ref-3">3</a>][<a href="#ref-4">4</a>] into a repeatable selection process.
| Array Characteristic | Recommended Approach | Evidence Basis | Primary Risk |
|---|---|---|---|
| Short array, moderate divergence | Standard long-read assembly with polishing | Cotton assembly resolved multiple region types with HiFi and ONT [<a href="#ref-4">4</a>] | Insufficient coverage at array boundaries |
| Long array, high identity | Ultralong reads spanning full array, graph-based assembly | Maize assembly used deep coverage ultralong ONT reads [<a href="#ref-1">1</a>] | Reads cannot span entire array |
| Megabase array, satellite structure | Satellite identification tools plus Hi-C scaffolding | Human assembly characterized 189.9 Mb of centromeric repeats [<a href="#ref-5">5</a>] | Satellite copy collapse in assembly graph |
| Ribosomal DNA array | Ultralong reads spanning multiple transcription units | Maize resolved 2,974 copies of 45S rDNA [<a href="#ref-1">1</a>] | Copy number underestimation |
| Telocentric chromosome centromere | Species-specific satellite analysis with structural validation | Pig assembly found telomere-SAT1B-satellite-SAT3 structure [<a href="#ref-3">3</a>] | Unusual centromere organization missed by generic tools |
Step 4: Establish Stop Criteria for Each Resolution Attempt
Define objective criteria for when a resolution strategy has succeeded or failed before starting. These criteria prevent wasted compute time and provide clear documentation for troubleshooting.
A resolution attempt succeeds when the assembled array meets three conditions. First, the array length falls within the expected range based on independent measurements such as pulsed-field gel electrophoresis or quantitative PCR. Second, the repeat copy number matches estimates from read depth analysis or cytogenetic data. Third, the array boundaries connect to flanking unique sequence without ambiguity.
A resolution attempt fails when any of these conditions cannot be met after two rounds of parameter adjustment. Document the failure mode using the categories in the troubleshooting section of this article. Common failure modes include contig breaks at array boundaries, collapsed copy numbers, and misplaced arrays.
The maize assembly achieved complete chromosome traversal with each chromosome in a single contig [<a href="#ref-1">1</a>], which represents the success standard for chromosome-scale resolution. The sheep assembly added 220.05 megabases of previously unresolved regions [<a href="#ref-2">2</a>], and the pig assembly added 194.42 megabases [<a href="#ref-3">3</a>]. These outcomes provide benchmarks for what complete resolution adds to a genome project.
Step 5: Document Decisions in a Repeat Resolution Log
Maintain a structured log for each repeat array targeted for resolution. This log serves as the record system for the assembly project and provides the data needed for troubleshooting and publication.
For each array, record the following fields. Array identifier and chromosome location. Repeat unit type and length. Estimated array length from independent measurements. Sequencing platforms and read length distributions used. Assembly software and version. Parameter settings for each assembly attempt. Outcome of each attempt including contig lengths and copy numbers. Failure mode if the attempt did not succeed. Time and compute resources consumed.
The nf-core documentation describes community pipeline standards that support reproducible workflow configuration [<a href="#ref-8">8</a>]. Apply these standards to the repeat resolution log so that another researcher can repeat the process. The Carpentries lessons provide Git training that supports version control for research software [<a href="#ref-11">11</a>], which applies to the scripts and configuration files used in repeat resolution.
Step 6: Compare Outcomes Against Published Benchmarks
Published telomere-to-telomere projects provide benchmarks for evaluating resolution success. The maize assembly achieved base accuracy over 99.99% [<a href="#ref-1">1</a>]. The sheep and pig assemblies exceeded 99.999% base accuracy [<a href="#ref-2">2</a>][<a href="#ref-3">3</a>]. These accuracy levels represent the quality standard for complete assemblies.
Compare the proportion of previously unresolved sequence added by your assembly against published outcomes. The sheep assembly added 220.05 megabases to the existing reference [<a href="#ref-2">2</a>]. The pig assembly added 194.42 megabases [<a href="#ref-3">3</a>]. The human assembly characterized 189.9 megabases of centromeric and pericentromeric repeats [<a href="#ref-5">5</a>]. If your assembly adds substantially less unresolved sequence than expected for the genome size, investigate whether repeats remain collapsed or missing.
Compare the number of new genes identified in previously unresolved regions. The sheep assembly added 754 new genes [<a href="#ref-2">2</a>]. The pig assembly added 1,189 new genes [<a href="#ref-3">3</a>]. These gene discoveries demonstrate the biological value of complete repeat resolution.
Step 7: Validate with Independent Biological Evidence
Sequence-based validation alone cannot confirm that an assembly correctly represents the repeat array. Independent biological evidence provides confirmation that the assembly reflects the actual genome structure.
Cytogenetic analysis with fluorescent in situ hybridization can confirm the location and approximate size of repeat arrays. The cotton assembly discovered a phylogenetically recent centromere repositioning on chromosome D08 [<a href="#ref-4">4</a>], a finding that would benefit from cytogenetic confirmation. The pig assembly revealed a unique telomere-SAT1B-satellite-SAT3 structure in telocentric chromosomes [<a href="#ref-3">3</a>], which could be validated with fiber-fluorescence in situ hybridization.
Genetic mapping data can confirm that assembled repeat arrays are placed at the correct chromosomal locations. The sheep assembly corrected several structural errors in previous reference assemblies [<a href="#ref-2">2</a>], demonstrating that even established assemblies can contain placement errors. Compare your assembly against genetic maps to verify repeat array positions.
Comparative analysis with related species provides additional validation. The human assembly found a strong relationship between centromere position and the evolution of surrounding DNA through layered repeat expansions [<a href="#ref-5">5</a>]. If your assembly shows centromere positions that differ dramatically from related species, investigate whether the difference reflects biological variation or assembly error.
Step 8: Determine When to Escalate to Collaborative Resolution
Some repeat arrays resist resolution despite appropriate data and methods. Published telomere-to-telomere projects represent collaborative efforts with substantial expertise and compute resources. The maize assembly required deep coverage with both ultralong Oxford Nanopore Technology and PacBio HiFi reads [<a href="#ref-1">1</a>]. The cotton assembly combined PacBio HiFi, Oxford Nanopore Technology ultralong reads, and Hi-C [<a href="#ref-4">4</a>].
Escalate to collaborative resolution when the repeat array exceeds 1 megabase with high sequence identity and your local resources cannot generate sufficient ultralong read coverage. The human centromeric repeats span 189.9 megabases [<a href="#ref-5">5</a>], a scale that required the coordinated effort of the Telomere-to-Telomere consortium. Individual laboratories may not have the sequencing capacity for arrays of this scale.
Escalate when the repeat array shows unusual structure that generic tools do not recognize. The pig assembly found a unique telomere-SAT1B-satellite-SAT3 structure in telocentric chromosomes [<a href="#ref-3">3</a>]. The cotton assembly discovered a centromere repositioning specific to Gossypium hirsutum [<a href="#ref-4">4</a>]. These unusual structures required specialized analysis that may not be available in standard assembly pipelines.
Escalate when the repeat array has major biological significance for the research program. The sheep assembly identified variants associated with wool fineness [<a href="#ref-2">2</a>]. The pig assembly identified genes associated with body stature [<a href="#ref-3">3</a>]. If the unresolved repeat array carries traits central to the research program, collaborative resolution may be justified despite the cost.
Step 9: Apply the Decision Framework to Ribosomal RNA Arrays
Ribosomal RNA arrays present specific challenges that require a tailored decision process. The maize assembly resolved the complete nucleolar organizer region of 26.8 megabases with 2,974 copies of 45S ribosomal DNA [<a href="#ref-1">1</a>]. The cotton assembly resolved 5S ribosomal DNA clusters and nucleolar organizer regions [<a href="#ref-4">4</a>].
First, determine whether the array contains 45S ribosomal DNA, 5S ribosomal DNA, or both. These arrays have different unit structures and require different resolution approaches. The maize assembly focused on 45S ribosomal DNA [<a href="#ref-1">1</a>], while the cotton assembly resolved both 5S and 45S arrays [<a href="#ref-4">4</a>].
Second, estimate copy number using read depth analysis. The maize assembly found 2,974 copies of 45S ribosomal DNA [<a href="#ref-1">1</a>]. Copy number estimates from read depth provide a target for assembly validation. If the assembled array contains substantially fewer copies than the read depth estimate, the assembly has collapsed the array.
Third, examine the intergenic spacer sequences between transcription units. The maize assembly revealed complex patterns of ribosomal DNA duplications and transposon insertions [<a href="#ref-1">1</a>]. These structural features affect assembly strategy because they introduce sequence variation that can help or hinder copy-specific placement.
Fourth, determine whether the array contains transposon insertions that disrupt the tandem structure. The maize assembly found transposon insertions within the ribosomal DNA array [<a href="#ref-1">1</a>]. These insertions create unique sequence anchors that can help resolve the array structure.
Step 10: Apply the Decision Framework to Centromeric Satellites
Centromeric satellite arrays require a decision process that accounts for species-specific satellite structure. The sheep assembly identified four types of repeat units in centromeric regions [<a href="#ref-2">2</a>]. The pig assembly revealed pig-specific centromeric satellite repeat units [<a href="#ref-3">3</a>]. The human assembly found multimegabase structural rearrangements in centromeric regions [<a href="#ref-5">5</a>].
First, identify the satellite monomer sequence from published databases or de novo repeat identification. The sheep assembly identified SatI, SatII, SatIII, and CenY repeat units [<a href="#ref-2">2</a>]. The pig assembly identified pig-specific satellites [<a href="#ref-3">3</a>]. Without the correct monomer sequence, satellite identification tools cannot find the array.
Second, determine whether the centromere contains a single satellite type or multiple types. The sheep assembly found four types of repeat units in centromeric regions [<a href="#ref-2">2</a>]. The human assembly found layered repeat expansions in centromeric regions [<a href="#ref-5">5</a>]. Multiple satellite types require separate analysis for each type.
Third, assess whether the centromere is satellite-rich or satellite-poor. The maize assembly precisely dissected the repeat compositions of both CentC-rich and CentC-poor centromeres [<a href="#ref-1">1</a>]. Satellite-poor centromeres may require different resolution strategies than satellite-rich centromeres.
Fourth, examine whether the centromere position matches expectations from related species. The cotton assembly discovered a phylogenetically recent centromere repositioning on chromosome D08 [<a href="#ref-4">4</a>]. The pig assembly found a unique structure in telocentric chromosomes [<a href="#ref-3">3</a>]. Unexpected centromere positions may indicate either biological variation or assembly error.
Step 11: Apply the Decision Framework to Telomeric Repeats
Telomeric repeats present a simpler decision process than centromeric or ribosomal arrays because the repeat unit is short and highly conserved. The maize assembly achieved complete chromosome traversal with each chromosome in a single contig [<a href="#ref-1">1</a>], which requires telomere resolution at every chromosome end.
First, confirm the telomeric motif for the target species. Vertebrates use TTAGGG, while plants often use TTTAGGG. The maize assembly resolved telomeric regions as part of complete chromosome traversal [<a href="#ref-1">1</a>]. The cotton assembly resolved 52 telomeric regions [<a href="#ref-4">4</a>].
Second, determine whether telomeric repeats appear at contig ends in the initial assembly. Contigs with telomeric sequence at one end represent chromosome ends. Contigs with telomeric sequence at both ends may represent complete chromosomes or may indicate assembly errors.
Third, assess whether telomeric repeats are internal to contigs. Internal telomeric repeats can indicate chromosome fusions or assembly errors. The pig assembly found a unique telomere-SAT1B-satellite-SAT3 structure in telocentric chromosomes [<a href="#ref-3">3</a>], demonstrating that telomeric sequence can appear in unexpected locations.
Fourth, measure telomere length at each chromosome end. Telomere length varies between cells and individuals, so the assembly represents one haplotype. The human assembly found high degrees of structural variation in centromeric regions [<a href="#ref-5">5</a>], and similar variation likely applies to telomeric regions.
Step 12: Review and Adjust the Decision Framework
The decision framework should be reviewed after each assembly project to incorporate lessons learned. Published telomere-to-telomere projects continue to reveal new repeat structures. The pig assembly found a unique telomere-SAT1B-satellite-SAT3 structure [<a href="#ref-3">3</a>]. The cotton assembly discovered a centromere repositioning [<a href="#ref-4">4</a>]. These discoveries expand the range of structures that researchers should expect.
The EMBL-EBI Training program provides learning pathways for data-resource training and practical analysis education [<a href="#ref-10">10</a>]. The Galaxy Training Network provides accessible workflow training and analysis tutorials [<a href="#ref-7">7</a>]. These resources help researchers stay current with new tools and approaches for repeat resolution.
The Bioconductor project provides official package, workflow, and installation documentation for reproducible genomic analysis [<a href="#ref-9">9</a>]. The nf-core documentation describes community pipeline standards [<a href="#ref-8">8</a>]. These resources support the reproducible application of the decision framework across projects.
The NCBI maintains sequence databases and analysis services that can be used to examine existing assemblies and identify known repeat families [<a href="#ref-6">6</a>]. Regular review of new assemblies in the NCBI databases can reveal repeat structures that were previously unknown.
Frequently Asked Questions
What read length is needed for telomere-to-telomere assembly?
Read length requirements depend on the size of the largest repeat array in the target genome. The maize assembly used ultralong Oxford Nanopore Technology reads to traverse repetitive regions including a 26.8 megabase nucleolar organizer region [<a href="#ref-1">1</a>]. In practice, aim for maximum read lengths exceeding the expected size of the largest repeat array, with deep coverage to ensure multiple reads span each repeat junction.
How much sequencing coverage is required for repeat resolution?
Coverage requirements depend on genome size, repeat content, and sequencing platform. The maize project used deep coverage with both ultralong Oxford Nanopore Technology and PacBio HiFi reads to achieve complete chromosome traversal [<a href="#ref-1">1</a>]. The sheep and pig assemblies achieved base accuracy above 99.999% [<a href="#ref-2">2</a>][<a href="#ref-3">3</a>]. Higher coverage improves assembly quality but increases cost, so coverage should be planned based on the specific requirements of the project.
What is the difference between telomere-to-telomere assembly and standard genome assembly?
Standard genome assembly produces contigs and scaffolds with gaps at repetitive regions. Telomere-to-telomere assembly produces complete chromosomes with no gaps, including telomeres, centromeres, and ribosomal RNA arrays. The maize assembly achieved complete chromosome traversal with each chromosome in a single contig [<a href="#ref-1">1</a>]. The sheep assembly added 220.05 megabases of previously unresolved regions [<a href="#ref-2">2</a>], and the pig assembly added 194.42 megabases [<a href="#ref-3">3</a>].
How do I know if my assembly has collapsed repeats?
Collapsed repeats produce an assembly that is shorter than expected and missing sequence. Comparison with independent estimates of repeat content can identify collapsed regions. The sheep assembly corrected several structural errors in previous reference assemblies [<a href="#ref-2">2</a>], demonstrating that earlier assemblies contained collapsed or misassembled repeats. Check assembly size against expected genome size and examine repeat copy numbers against independent estimates.
Can I resolve centromeres without ultra-long reads?
Centromere resolution typically requires ultra-long reads because centromeric satellite arrays can extend for megabases. The human telomere-to-telomere assembly found that pericentromeric and centromeric repeats constitute 6.2% of the genome [<a href="#ref-5">5</a>]. The cotton assembly used PacBio HiFi, Oxford Nanopore Technology ultralong reads, and Hi-C to resolve centromeric regions [<a href="#ref-4">4</a>]. Without ultra-long reads, centromeric regions will likely remain as gaps.
What tools are available for repeat identification and resolution?
Specialized tools include Telomerecat for telomeric repeats and CentroMiner for centromeric satellites. General assembly tools with repeat-aware strategies are also available. The Bioconductor project provides official package and workflow documentation for genomic analysis [<a href="#ref-9">9</a>]. The Galaxy Training Network provides accessible workflow training and analysis tutorials [<a href="#ref-7">7</a>]. The nf-core documentation describes community pipeline standards [<a href="#ref-8">8</a>].
How do I validate a telomere-to-telomere assembly?
Validation requires multiple approaches. Check that all expected chromosomes are represented as complete contigs. Verify that telomeric motifs appear at chromosome ends. Confirm that centromeric satellites show the expected organization. Compare assembly statistics against published assemblies for the same or related species. The sheep assembly corrected several structural errors in previous reference assemblies [<a href="#ref-2">2</a>], demonstrating that validation can identify and fix problems in earlier work.
What are the limitations of telomere-to-telomere assemblies?
Telomere-to-telomere assemblies represent one haplotype of one individual. Telomere length varies between cells and individuals. Centromere position and satellite composition can vary within species. The human assembly found high degrees of structural, epigenetic, and sequence variation in centromeric regions across individuals [<a href="#ref-5">5</a>]. Researchers should not assume that the assembly represents all individuals in a species.
Related Bioinformatics Guides
- Long-Read Sequencing for De Novo Assembly of Complex Genomes: Case Studies and Best Practices
- Metagenomics Assembly: Strategies for Reconstructing Microbial Genomes
- RNA-Seq vs DNA-Seq: Key Differences and Applications
- De Novo Genome Assembly with Long Reads: A Practical Workflow
- Long-Read Metagenome Assembly: Overcoming Challenges with Nanopore and PacBio Data
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
[1] [A complete telomere-to-telomere assembly of the maize genome.](https://pubmed.ncbi.nlm.nih.gov/37322109). Nature genetics, 2023. [2] [Telomere-to-telomere sheep genome assembly identifies variants associated with wool fineness.](https://pubmed.ncbi.nlm.nih.gov/39779954). Nature genetics, 2025. [3] [Telomere-to-telomere genome assembly of a male pig provides insight into population structure and selection for body stature.](https://pubmed.ncbi.nlm.nih.gov/41387626). Nature genetics, 2026. [4] [A telomere-to-telomere genome assembly of cotton provides insights into centromere evolution and short-season adaptation.](https://pubmed.ncbi.nlm.nih.gov/40097785). Nature genetics, 2025. [5] [Complete genomic and epigenetic maps of human centromeres.](https://pubmed.ncbi.nlm.nih.gov/35357911). Science (New York, N.Y.), 2022. [6] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information. [7] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project. [8] [nf-core Documentation](https://nf-co.re/docs). nf-core. [9] [Bioconductor](https://bioconductor.org/). Bioconductor Project. [10] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute. [11] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.