Telomere-to-Telomere Assembly Validation: Metrics and Methods for Complete Genomes
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Telomere-to-Telomere (T2T) genome assemblies require validation beyond standard metrics like N50 and BUSCO completeness, as these fail to assess challenging regions such as telomeres, centromeres, and large tandem repeats. Specialized metrics like consensus quality value (QV) from k-mer analysis, Long Terminal Repeat Assembly Index (LAI), and repeat-aware polishing are crucial for evaluating base accuracy and repeat structure.
- Base-level accuracy validation must account for the limitations of k-mer-based QV in repetitive regions; discrepancies between QV calculated from different sequencing technologies (e.g., PacBio HiFi vs. Illumina) or poor read mapping in repeats necessitate further investigation. Optical genome mapping provides an independent physical map to confirm chromosome length and structural integrity, complementing sequence-based validation.
- Centromere validation is critical and involves confirming the presence, correct copy number, and order/orientation of satellite DNA repeat units, often requiring ultra-long reads to span entire arrays. Telomere validation necessitates empirical identification of species-specific repeat motifs and confirmation of their presence at all chromosome termini.
- Common failure modes in T2T validation include overestimation of QV in repeats, missing telomeres, collapsed/expanded centromeric arrays, and misjoins disrupting long-range order, often exacerbated by overcorrection during polishing with non-repeat-aware tools. Validation using diverse sequencing technologies is essential to mitigate technology-specific assembly errors.
- A structured triage framework with clear criteria for "accept," "document," or "escalate" outcomes is vital for prioritizing validation findings, distinguishing critical errors from acceptable limitations, and guiding decisions on whether to proceed, record caveats, or seek additional data/expertise.
Telomere-to-telomere (T2T) genome assemblies aim to reconstruct every chromosome from the first telomeric repeat at one end through the centromere to the telomeric repeat at the other end, with no unexplained gaps. Standard assembly quality metrics such as N50, BUSCO completeness, and base-level accuracy remain necessary but are insufficient for T2T validation because they do not verify the presence, orientation, or sequence correctness of the most challenging genomic regions: telomeres, centromeres, ribosomal DNA arrays, and other large tandem repeats. This article provides a practical framework for researchers who have produced or received a putative T2T assembly and need to validate it with specialized metrics and methods beyond those used for draft genomes. The focus is on concrete validation decisions, measurable quality thresholds, record-keeping practices, and escalation criteria when validation fails.
The Validation Problem Specific to T2T Assemblies
A draft genome assembly can be considered useful when it captures the majority of coding sequence and provides contiguity sufficient for gene-level analysis. A T2T assembly makes a stronger claim: every base pair of every chromosome is accounted for in the correct order and orientation, including regions that have historically resisted assembly because they consist of near-identical repeats spanning millions of base pairs. The human X chromosome T2T assembly demonstrated this by reconstructing the centromeric satellite DNA array of approximately 3.1 Mb and closing 29 gaps that remained in the GRCh38 reference, including sequences from pseudoautosomal regions and cancer-testis ampliconic gene families [<a href="#ref-1">1</a>]. This achievement required high-coverage ultra-long-read nanopore sequencing combined with complementary technologies for quality improvement and validation [<a href="#ref-1">1</a>].
The validation burden increases accordingly. Standard metrics like scaffold N50 can appear excellent even when entire chromosome arms are missing or when telomeric repeats are absent from sequence ends. BUSCO completeness scores measure the presence of conserved single-copy orthologs, but these genes are rarely located in centromeric or telomeric regions, so a high BUSCO score does not confirm that repetitive genomic structures were assembled correctly. The Bactrocera dorsalis T2T assembly illustrates the challenge: dipteran insects have small body sizes, poorly conserved telomere and centromere structures, high heterozygosity, and complex genetic backgrounds, all of which complicate assembly and validation [<a href="#ref-2">2</a>]. The successful assembly of this 596 Mb genome required a low-input HiFi CCS library from a single male individual plus Oxford Nanopore Technology (ONT) sequence from pooled inbred individuals, and the validation had to confirm complete structural organization of centromeres and telomeres [<a href="#ref-2">2</a>].
Researchers must therefore adopt a validation workflow that tests the specific claims of a T2T assembly: telomere presence at every chromosome end, centromere reconstruction with correct repeat arrays, absence of collapsed or expanded repeat copies, and base-level accuracy in regions where standard mapping tools perform poorly. Each of these requires distinct methods and metrics.
At a Glance: T2T Validation Metrics and Their Roles
| Validation Layer | Primary Metric or Method | What It Confirms | Common Failure Mode |
|---|---|---|---|
| Base accuracy | Consensus quality value (QV) from k-mer analysis | Per-base error rate across the assembly | Overestimation of QV in repetitive regions where k-mers are absent |
| Completeness | BUSCO complete alignments | Presence of conserved single-copy genes | High score despite missing telomeres or centromeres |
| Repeat structure | Long terminal repeat assembly index (LAI) | Correct assembly of LTR retrotransposon arrays | Low LAI indicating collapsed or misjoined repeat copies |
| Telomere validation | Telomeric repeat detection at sequence ends | Presence of telomere motifs at all chromosome termini | Missing telomeres on one or more chromosome ends |
| Centromere validation | Satellite array reconstruction and alignment | Correct order and copy number of centromeric repeats | Collapsed arrays or chimeric joins within the centromere |
| Structural correctness | Optical genome mapping or genetic map comparison | Agreement between assembly and independent physical maps | Misjoins that preserve local sequence but disrupt long-range order |
The table above summarizes the six validation layers that a T2T assembly must pass. Each layer uses different data types and tools, and each has distinct failure modes that require different corrective actions. The following sections describe each layer in detail, including practical implementation steps and interpretation guidance.
Base-Level Accuracy: QV and Its Limitations in Repetitive Regions
The consensus quality value (QV) is the standard metric for base-level accuracy in genome assemblies. QV is typically calculated from k-mer analysis by comparing the frequency of k-mers in the assembly to the frequency expected from raw sequencing reads. A QV of 40 corresponds to one error per 10,000 bases, and a QV of 50 corresponds to one error per 100,000 bases. The Acorus tatarinowii T2T assembly reported a consensus QV of 45.41, indicating approximately one error per 34,700 bases [<a href="#ref-3">3</a>]. This value was supported by a long terminal repeat assembly index (LAI) of 10.04 and BUSCO analysis showing 97.6% complete alignments [<a href="#ref-3">3</a>].
The limitation of QV in T2T validation is that k-mer-based methods depend on the presence of unique or low-copy k-mers. In centromeric satellite arrays and other large tandem repeats, the same k-mer may appear thousands of times, and errors in these regions do not create novel k-mers that can be detected by frequency comparison. A QV calculated from whole-genome k-mer analysis therefore reflects accuracy primarily in unique and low-copy regions, while the most challenging repetitive regions may harbor undetected errors.
The CHM13 human genome polishing study directly addressed this problem. Evaluation of the initial T2T draft assembly revealed evidence of small errors and structural misassemblies despite the assembly being derived from highly accurate sequences [<a href="#ref-4">4</a>]. The researchers designed a repeat-aware polishing strategy that made accurate assembly corrections in large repeats without overcorrection, ultimately fixing 51% of existing errors and improving the assembly quality value from 70.2 to 73.9 as measured from PacBio high-fidelity and Illumina k-mers [<a href="#ref-4">4</a>]. This result demonstrates two important points for validation: first, even high-quality T2T assemblies contain errors that standard polishing misses, and second, the choice of polishing and validation data matters because sequencing biases in both high-fidelity and ONT reads cause signature assembly errors that can be corrected with a diverse panel of sequencing technologies [<a href="#ref-4">4</a>].
Practical implementation of base-level validation should therefore include multiple independent measurements. Compute QV from at least two sequencing technologies when available, such as PacBio HiFi and Illumina short reads. Compare the QV values and investigate any substantial discrepancy. Additionally, map raw reads back to the assembly and examine coverage and mismatch patterns specifically in repetitive regions, because uniform coverage with low mismatch rates provides supporting evidence that repeats were not collapsed or expanded. Record the QV calculation method, the k-mer size used, the read set used for validation, and the date of the analysis so that future comparisons are possible.
Completeness Assessment: BUSCO and Its Blind Spots
BUSCO (Benchmarking Universal Single-Copy Orthologs) is the standard tool for assessing assembly completeness by searching for a set of conserved single-copy genes expected to be present in the target lineage. The Acorus tatarinowii assembly reported 97.6% complete BUSCO alignments for the genome and 97.2% completeness for the annotated gene models [<a href="#ref-3">3</a>]. These values indicate that the assembly captured nearly all conserved coding content.
The blind spot of BUSCO for T2T validation is that the conserved genes used by BUSCO are distributed across the genome but are rarely located within centromeric satellite arrays, telomeric repeats, or ribosomal DNA clusters. A genome assembly could theoretically achieve a perfect BUSCO score while missing entire centromeres or lacking telomeres on every chromosome. BUSCO therefore confirms that the euchromatic gene space is complete, but it cannot confirm that the heterochromatic and repetitive components of the genome were assembled.
For T2T validation, BUSCO should be treated as a necessary but insufficient check. Run BUSCO with the appropriate lineage dataset for the target species and record the percentage of complete, fragmented, and missing BUSCO groups. A complete BUSCO score below 95% should trigger investigation of whether the assembly failed to capture conserved genes or whether the lineage dataset is inappropriate. A complete BUSCO score above 95% provides confidence in gene space completeness but does not reduce the need for telomere, centromere, and repeat-structure validation.
The Bactrocera dorsalis study illustrates why completeness assessment must go beyond BUSCO for T2T claims. The assembly included complete structural organization information for centromeres and telomeres, which enabled comparative genomic analysis revealing the polyphyletic origin of sex chromosomes across Diptera [<a href="#ref-2">2</a>]. This level of structural resolution required validation methods specific to those repeat classes, beyond gene-space completeness metrics.
Repeat Structure Validation: LAI and Repeat-Aware Methods
The long terminal repeat assembly index (LAI) measures the correctness of LTR retrotransposon assembly by comparing the number of intact LTR elements to the number of truncated or solo LTRs. A higher LAI indicates that LTR retrotransposon arrays were assembled without excessive collapse or misjoining. The Acorus tatarinowii assembly reported an LAI of 10.04, which is considered reference-grade for plant genomes [<a href="#ref-3">3</a>]. Repetitive sequences comprised 57.51% of that genome, making repeat structure validation essential [<a href="#ref-3">3</a>].
LAI is a useful metric but it covers only one class of repeats. T2T validation requires additional checks for other repeat classes, particularly centromeric satellites, telomeric repeats, and ribosomal DNA arrays. The CHM13 polishing study noted that standard automated polishing tools made errors in large repeats, and the researchers designed a repeat-aware strategy specifically to avoid overcorrection [<a href="#ref-4">4</a>]. This finding has direct implications for validation: tools that work well for unique regions may introduce errors when applied to repetitive regions, and validation methods must be repeat-aware in the same way.
Practical repeat validation should include the following steps. First, annotate repeat content using a combination of de novo repeat identification and known repeat databases for the target lineage. Second, calculate LAI and record the value along with the repeat annotation method and version. Third, examine the length distribution of intact LTR elements and compare it to expectations for the lineage. Fourth, for any repeat family that shows an unusual length distribution or an unexpected number of copies, investigate whether the assembly collapsed multiple copies into one or expanded a single copy into multiple. This investigation typically requires mapping long reads across the repeat array and checking whether read depth is consistent with the assembled copy number.
Telomere Validation: Detecting and Confirming Chromosome Ends
Telomere validation is a defining feature of T2T assembly assessment. A T2T assembly must have telomeric repeats at both ends of every chromosome, and the sequence of those repeats must match the expected telomere motif for the species. The human X chromosome T2T assembly closed gaps that included sequences from the pseudoautosomal regions, and the complete chromosome X allowed mapping of methylation patterns across complex tandem repeats and satellite arrays [<a href="#ref-1">1</a>]. This achievement required assembling the telomeres and validating that the assembled telomere sequences were correct and complete.
The Bactrocera dorsalis study highlighted that dipteran insects have poorly conserved telomere and centromere structures, which makes telomere validation particularly challenging for such species [<a href="#ref-2">2</a>]. Researchers working on non-model organisms cannot assume that the canonical vertebrate telomere repeat (TTAGGG) will be present. The telomere motif must be identified empirically from the sequencing data or from published information about the target lineage.
Practical telomere validation involves several steps. First, identify the expected telomere repeat motif for the target species from the literature or by searching raw reads for tandem repeat motifs at high copy number. Second, extract the terminal sequences of each assembled chromosome and search for the telomere motif. Third, confirm that the telomere motif is present at the outermost end of the sequence, not internal to other sequence. Fourth, check that every chromosome has telomeric sequence at both ends. Fifth, if any chromosome end lacks telomeric sequence, determine whether the assembly is incomplete at that end or whether the chromosome naturally lacks telomeres, which is rare but possible in some species.
Record the telomere motif used, the number of chromosome ends with confirmed telomeres, the number of chromosome ends without telomeres, and the method used for detection. If telomeres are missing from one or more chromosome ends, escalate to a review of the assembly graph and the raw read data to determine whether the chromosome end was not assembled or whether the telomere was lost during polishing.
Centromere Validation: Satellite Arrays and Structural Organization
Centromere validation is the most technically demanding component of T2T assembly assessment. Centromeres consist of large arrays of tandemly repeated satellite DNA that can span millions of base pairs, and the repeat units are often nearly identical to each other. The human X chromosome T2T assembly reconstructed the centromeric satellite DNA array of approximately 3.1 Mb [<a href="#ref-1">1</a>]. This reconstruction required high-coverage ultra-long-read nanopore sequencing because the array length exceeded the read length of other sequencing technologies [<a href="#ref-1">1</a>].
The Bactrocera dorsalis assembly included complete structural organization information for centromeres, which provided insights into the evolution of chromosome structure in insects [<a href="#ref-2">2</a>]. This finding demonstrates that centromere validation goes beyond confirming the presence of satellite repeats and includes determining the order and orientation of repeat units within the array.
Practical centromere validation requires multiple lines of evidence. First, identify the centromeric satellite repeat family for the target species. Second, confirm that the assembled centromere contains the expected repeat family at the expected copy number, which can be estimated from read depth. Third, examine the order and orientation of repeat units within the array to confirm that the assembly did not collapse or expand the array. Fourth, if a reference genome or genetic map is available for the species, confirm that the centromere is located at the expected position on each chromosome. Fifth, use an independent technology such as optical genome mapping to confirm the length and structure of the centromeric array.
The resolution of ring chromosomes and Robertsonian translocations from long-read sequencing and T2T assembly demonstrated that multiple breakpoints were localized to genomic regions previously recalcitrant to sequencing, including acrocentric p-arms, ribosomal DNA arrays, and telomeric repeats [<a href="#ref-5">5</a>]. This finding underscores that centromere validation must consider the full structural context of the chromosome, beyond the satellite array itself.
Structural Validation: Optical Mapping and Independent Confirmation
Optical genome mapping provides an independent physical measurement of genome structure that can validate T2T assemblies. The study that resolved ring chromosomes and Robertsonian translocations used optical genome mapping for validation after long-read sequencing and T2T assembly [<a href="#ref-5">5</a>]. The analyses resolved 10 of 13 cases, including a Robertsonian translocation and all ring chromosomes [<a href="#ref-5">5</a>]. This result demonstrates that optical mapping can confirm structural rearrangements that are difficult to validate from sequencing data alone.
For T2T assembly validation, optical mapping serves several purposes. It provides an independent estimate of chromosome length that can be compared to the assembled sequence length. It generates a restriction map or nick-label pattern that can be compared to the in silico digest of the assembly. It can confirm the presence and location of large structural variants, including those that involve repetitive regions. It can also detect misjoins that preserve local sequence order but disrupt long-range order.
The CHM13 polishing study noted that sequencing biases in both high-fidelity and ONT reads cause signature assembly errors that can be corrected with a diverse panel of sequencing technologies [<a href="#ref-4">4</a>]. This finding supports the use of multiple independent validation methods, including optical mapping, because any single sequencing technology may introduce characteristic errors that are invisible when validating with the same technology used for assembly.
Practical structural validation should include the following steps. First, generate an optical map or obtain a genetic map for the target species. Second, compare the map to the assembly to identify any discrepancies in chromosome length, marker order, or restriction fragment pattern. Third, investigate any discrepancies to determine whether they reflect assembly errors or map errors. Fourth, if the assembly contains structural variants such as inversions or translocations, confirm those variants with an independent method. Fifth, record the validation method, the version of the map used, and the outcome of the comparison.
Polishing and Error Correction: Repeat-Aware Strategies
Polishing is the process of correcting errors in an assembly using additional sequencing data. Standard polishing tools align reads to the assembly and identify positions where the reads disagree with the assembly consensus. These tools work well for unique regions but can introduce errors in repetitive regions, where reads from different repeat copies may align to the same location and create false consensus.
The CHM13 polishing study directly compared a repeat-aware polishing strategy to standard automated polishing tools and outlined common polishing errors [<a href="#ref-4">4</a>]. The repeat-aware strategy made accurate assembly corrections in large repeats without overcorrection, fixing 51% of existing errors and improving the assembly quality value from 70.2 to 73.9 [<a href="#ref-4">4</a>]. This result demonstrates that polishing is not a neutral process: the choice of polishing strategy affects the final assembly quality, particularly in repetitive regions.
For T2T validation, polishing should be treated as an iterative process with validation at each step. The workflow should proceed as follows. First, assemble the genome using the primary assembly strategy. Second, validate the assembly using the metrics described in this article. Third, identify errors that require correction. Fourth, apply polishing with a repeat-aware strategy, not a standard tool that may overcorrect repeats. Fifth, revalidate the assembly to confirm that polishing corrected the intended errors without introducing new ones. Sixth, if polishing introduced new errors, consider whether the polishing data or strategy needs to be changed.
The Bactrocera dorsalis assembly used a strategy with a low-input HiFi CCS library from a single male individual and ONT sequence from pooled inbred individuals [<a href="#ref-2">2</a>]. This combination of sequencing technologies provided complementary data for assembly and validation. The Acorus tatarinowii assembly reported a consensus QV of 45.41 and an LAI of 10.04, indicating that the polishing strategy produced a reference-grade assembly [<a href="#ref-3">3</a>].
Records and Measurements: What to Document for Reproducible Validation
Reproducible T2T validation requires careful record-keeping. The following records should be maintained for each assembly validation project.
First, document the assembly version and the exact command used to generate it, including all parameters and the versions of all software tools. Second, document the sequencing data used for validation, including the sequencing technology, the read length distribution, the coverage, and the basecaller version. Third, document the validation metrics and the methods used to calculate them, including the QV calculation method, the k-mer size, the BUSCO lineage dataset, and the LAI calculation method. Fourth, document the telomere motif used and the method for detecting telomeres at chromosome ends. Fifth, document the centromere validation method, including the satellite repeat family identified and the evidence for correct array structure. Sixth, document any structural validation performed, including the optical mapping or genetic map used and the outcome of the comparison.
The Galaxy Training Network provides accessible workflow training and analysis tutorials that can help researchers implement reproducible validation workflows [<a href="#ref-6">6</a>]. The nf-core documentation describes community pipeline standards, usage, configuration, and reproducible workflow context that can be applied to T2T validation pipelines [<a href="#ref-7">7</a>]. The Carpentries lessons provide foundational computing, data, shell, Git, and programming training that supports reproducible research practices [<a href="#ref-8">8</a>]. Bioconductor provides official package, workflow, installation, and reproducible genomic-analysis documentation that can support the statistical analysis of validation metrics [<a href="#ref-9">9</a>]. The EMBL-EBI Training program offers bioinformatics learning pathways and data-resource training that can help researchers build the skills needed for T2T validation [<a href="#ref-10">10</a>]. The NCBI provides official descriptions of databases, search systems, sequence resources, and analysis services that can support validation workflows [<a href="#ref-11">11</a>].
Common Failure Patterns in T2T Validation
Several failure patterns recur in T2T validation projects. Recognizing these patterns early can save substantial time and resources.
The first common failure is overestimation of base accuracy in repetitive regions. A QV calculated from whole-genome k-mer analysis may be high even when centromeric or telomeric regions contain errors, because those regions contribute few unique k-mers to the analysis. This failure is detected by comparing QV values calculated from different sequencing technologies or by examining read mapping quality specifically in repetitive regions.
The second common failure is missing telomeres on one or more chromosome ends. This failure is detected by the telomere validation step and may indicate that the assembly is incomplete at those ends. The cause may be insufficient read coverage at chromosome ends, difficulty in sequencing through telomeric repeats, or loss of telomeric sequence during assembly or polishing.
The third common failure is collapsed or expanded centromeric satellite arrays. This failure is detected by comparing read depth across the centromere to the expected depth based on genome coverage. A collapsed array shows reduced read depth, while an expanded array shows increased read depth. The cause may be the inability of the assembler to distinguish between near-identical repeat copies.
The fourth common failure is misjoins that preserve local sequence order but disrupt long-range order. This failure is detected by comparing the assembly to an independent physical map. The cause may be the presence of large duplicated regions that confuse the assembler.
The fifth common failure is overcorrection during polishing. This failure is detected by comparing the assembly before and after polishing and identifying new errors introduced by the polishing step. The cause is the use of standard polishing tools that are not repeat-aware.
The sixth common failure is the use of inappropriate validation data. If the validation reads come from the same sequencing technology and library used for assembly, systematic sequencing biases may be invisible to the validation. This failure is detected by validating with an independent sequencing technology or with optical mapping.
Limitations of T2T Validation Methods
T2T validation methods have inherent limitations that researchers should understand before interpreting results.
QV from k-mer analysis cannot detect errors in regions where k-mers are not unique. This limitation is fundamental to the method and cannot be fully overcome by increasing coverage or read length. Repeat-aware polishing can correct some errors in these regions, but the validation of those corrections requires methods that do not depend on k-mer uniqueness.
BUSCO completeness cannot detect missing repetitive regions. A genome assembly can achieve a perfect BUSCO score while lacking entire centromeres or telomeres. This limitation means that BUSCO should never be used as the sole completeness metric for a T2T assembly.
LAI covers only LTR retrotransposons. Other repeat classes, including centromeric satellites, telomeric repeats, and ribosomal DNA, require separate validation methods. A high LAI does not confirm that all repeat classes were assembled correctly.
Optical mapping provides independent structural validation but has limited resolution. The technique can confirm chromosome length and large-scale structure but cannot validate base-level accuracy or the fine structure of satellite arrays.
The CHM13 polishing study noted that sequencing biases in both high-fidelity and ONT reads cause signature assembly errors [<a href="#ref-4">4</a>]. This finding implies that validation using the same sequencing technology as assembly may miss technology-specific errors. The study recommended using a diverse panel of sequencing technologies for correction and validation [<a href="#ref-4">4</a>].
Professional Escalation Criteria
Researchers should escalate to additional expertise or resources when validation reveals problems that cannot be resolved with available tools and data. The following criteria indicate that escalation is appropriate.
First, escalate when telomeres are missing from multiple chromosome ends and the cause cannot be identified from the assembly graph or raw read data. This situation may require additional sequencing with a different technology or library preparation method.
Second, escalate when centromeric satellite arrays show evidence of collapse or expansion that cannot be resolved by polishing. This situation may require ultra-long reads that span the full array or optical mapping to determine the true array length.
Third, escalate when the assembly contains structural variants that cannot be confirmed by an independent method. This situation may require collaboration with a cytogenetics laboratory or a specialized structural variant analysis group.
Fourth, escalate when polishing introduces new errors that cannot be corrected with available tools. This situation may require the development of a repeat-aware polishing strategy or collaboration with the developers of polishing tools.
Fifth, escalate when validation metrics are inconsistent across methods, such as a high QV but a low LAI or a high BUSCO score but missing telomeres. This situation may indicate a systematic problem with the assembly or the validation approach.
The study that resolved ring chromosomes and Robertsonian translocations demonstrated that some complex structural variants remain unresolved even with advanced methods. The analyses resolved 10 of 13 cases, leaving 3 cases unresolved [<a href="#ref-5">5</a>]. This result indicates that some genomes or structural variants may not be resolvable with current technology, and researchers should be prepared to document unresolved cases and report them honestly.
A Decision Framework for Triage and Escalation in T2T Validation
The validation layers described above generate a large amount of data, and researchers must decide which problems demand immediate correction, which can be documented as acceptable limitations, and which require additional sequencing or collaboration. A structured triage framework prevents two common outcomes: spending excessive resources on minor errors that do not affect the biological conclusions, or publishing a T2T assembly with a critical structural defect that invalidates downstream analysis. This section provides a practical decision framework organized by validation layer, with concrete criteria for accept, document, or escalate outcomes.
Triage Categories and Decision Rules
The triage framework uses three outcomes for each validation finding. Accept means the finding meets the quality threshold and requires no further action. Document means the finding falls short of the ideal threshold but does not invalidate the T2T claim, and the limitation must be recorded in the assembly report. Escalate means the finding indicates a probable assembly error that requires additional data, tools, or expertise before the assembly can be considered T2T.
For base accuracy, accept a QV of 40 or higher when calculated from at least two independent sequencing technologies. Document a QV between 30 and 40 when the assembly otherwise passes all structural validation, noting that the lower QV may reflect difficult repetitive regions instead of uniform error. Escalate a QV below 30 or a discrepancy of more than 5 QV units between two independent calculations, because this inconsistency suggests a systematic problem with the assembly or the validation data instead of random error.
For BUSCO completeness, accept a complete score of 95% or higher with the appropriate lineage dataset. Document a complete score between 90% and 95% when the missing BUSCO groups are concentrated in a single genomic region that also shows other validation problems, because this pattern may indicate a local assembly failure instead of a genome-wide problem. Escalate a complete score below 90% or a high proportion of fragmented BUSCO groups, because both indicate that the assembly failed to capture conserved gene space.
For telomere validation, accept an assembly where every chromosome end has the expected telomere motif. Document an assembly where one chromosome end lacks telomeric sequence but the assembly graph shows that the end is complete and simply lacks telomeric repeats, which occurs in some species. Escalate when multiple chromosome ends lack telomeres or when the telomere motif is present internally instead of at the terminal position, because both patterns indicate incomplete assembly or misassembly.
For centromere validation, accept an assembly where the centromeric satellite array shows read depth consistent with the expected copy number and the repeat unit order matches the species-specific pattern. Document an assembly where the array length is uncertain by less than 10% but the repeat unit sequence is correct, because small length differences may reflect genuine variation between individuals. Escalate when read depth across the centromere is less than half or more than double the genome average, because this pattern indicates collapse or expansion of the array.
For structural validation, accept an assembly where the optical map or genetic map agrees with the assembly across all chromosomes. Document an assembly where the map agrees except for one small region that also shows high repeat content, because the discrepancy may reflect map resolution limits instead of assembly error. Escalate when the map disagrees with the assembly across multiple chromosomes or when a structural variant cannot be confirmed by any independent method.
Implementing the Triage Framework
The triage framework should be applied systematically after each validation round, not informally during analysis. Create a validation tracking table with one row per chromosome and one column per validation layer. Record the outcome for each cell as accept, document, or escalate, along with the specific metric value and the date of the assessment. This table serves as the primary record for the assembly validation and provides the evidence needed for the assembly report and for any future revisions.
The Bactrocera dorsalis assembly provides an example of how triage decisions interact with biological questions. The study assembled a 596 Mb T2T genome for a fruit crop pest and used the complete structural organization information for centromeres and telomeres to investigate the evolution of chromosome structure in insects [<a href="#ref-2">2</a>]. The validation had to confirm also that the assembly was complete but also that the centromere and telomere structures were accurate enough to support comparative genomic analysis revealing the polyphyletic origin of sex chromosomes across Diptera [<a href="#ref-2">2</a>]. If the centromere validation had revealed a collapsed array, the comparative analysis would have been compromised, and the triage framework would have escalated the problem before downstream analysis.
The Acorus tatarinowii assembly demonstrates the value of documenting limitations even when the assembly passes all primary thresholds. The study reported a consensus QV of 45.41, an LAI of 10.04, and BUSCO completeness of 97.6% for the genome and 97.2% for gene models [<a href="#ref-3">3</a>]. These values all meet the accept criteria in the triage framework. However, the assembly spanned 359.36 Mb with a scaffold N50 of 32.54 Mb, and repetitive sequences comprised 57.51% of the genome [<a href="#ref-3">3</a>]. The documentation of these values in the assembly report allows other researchers to assess whether the assembly meets their specific needs.
Common Triage Mistakes
Several mistakes recur when researchers apply triage decisions to T2T validation. The first mistake is treating all validation failures as equivalent. A missing telomere on one chromosome end is a different problem from a collapsed centromere on the same chromosome, and the two problems require different corrective actions. The triage framework forces researchers to classify each finding separately and to apply the appropriate response for each classification.
The second mistake is accepting a validation result without understanding the method limitations. The CHM13 polishing study demonstrated that standard automated polishing tools make errors in large repeats, and the researchers designed a repeat-aware strategy to avoid overcorrection [<a href="#ref-4">4</a>]. A researcher who accepts a QV calculated from a single sequencing technology without understanding that the QV does not reflect repetitive regions may publish an assembly with undetected errors in centromeric or telomeric sequence.
The third mistake is escalating problems that do not require escalation. A BUSCO complete score of 93% with a lineage dataset that is not well matched to the target species may reflect dataset limitations instead of assembly problems. The triage framework requires researchers to investigate the cause of a low score before deciding whether to escalate, instead of escalating automatically.
The fourth mistake is failing to document decisions. The triage framework produces a record of what was accepted, what was documented, and what was escalated, along with the evidence for each decision. This record is essential for reproducibility and for defending the assembly against later criticism.
Records and Measurements for Triage Decisions
The triage framework requires specific records for each validation layer. For base accuracy, record the QV value, the calculation method, the k-mer size, the sequencing technologies used, and the date of the calculation. For completeness, record the BUSCO version, the lineage dataset, the percentage of complete, fragmented, and missing BUSCO groups, and the date of the analysis. For telomere validation, record the telomere motif used, the number of chromosome ends with confirmed telomeres, the number without telomeres, and the detection method. For centromere validation, record the satellite repeat family, the expected and observed copy number, the read depth across the array, and the method used to determine array structure. For structural validation, record the optical map or genetic map version, the comparison method, and the outcome for each chromosome.
The Galaxy Training Network provides accessible workflow training and analysis tutorials that can help researchers implement reproducible validation workflows [<a href="#ref-6">6</a>]. The nf-core documentation describes community pipeline standards, usage, configuration, and reproducible workflow context that can be applied to T2T validation pipelines [<a href="#ref-7">7</a>]. The Carpentries lessons provide foundational computing, data, shell, Git, and programming training that supports reproducible research practices [<a href="#ref-8">8</a>]. Bioconductor provides official package, workflow, installation, and reproducible genomic-analysis documentation that can support the statistical analysis of validation metrics [<a href="#ref-9">9</a>]. The EMBL-EBI Training program offers bioinformatics learning pathways and data-resource training that can help researchers build the skills needed for T2T validation [<a href="#ref-10">10</a>]. The NCBI provides official descriptions of databases, search systems, sequence resources, and analysis services that can support validation workflows [<a href="#ref-11">11</a>].
Applying the Framework to Polishing Decisions
The triage framework also guides polishing decisions. When validation identifies errors, the researcher must decide whether to polish, how to polish, and when to stop polishing. The CHM13 study provides the key evidence for this decision: a repeat-aware polishing strategy fixed 51% of existing errors and improved the assembly quality value from 70.2 to 73.9, while standard automated polishing tools made errors in large repeats [<a href="#ref-4">4</a>]. The study also showed that sequencing biases in both high-fidelity and ONT reads cause signature assembly errors that can be corrected with a diverse panel of sequencing technologies [<a href="#ref-4">4</a>].
Apply the following polishing decision rules. First, polish only when validation identifies specific errors that can be corrected. Do not polish an assembly that passes all validation thresholds, because polishing risks introducing new errors. Second, use a repeat-aware polishing strategy when the assembly contains large repeats, and validate the polishing result with the same metrics used for the initial validation. Third, stop polishing when a validation round shows no improvement or when polishing introduces new errors. Fourth, document the polishing strategy, the data used, and the before and after validation metrics.
The ring chromosome study demonstrated that some structural variants remain unresolved even with advanced methods. The analyses resolved 10 of 13 cases, including a Robertsonian translocation and all ring chromosomes, but left 3 cases unresolved [<a href="#ref-5">5</a>]. The triage framework would classify the unresolved cases as escalate, requiring additional data or collaboration, and would document the limitations of the analysis for the resolved cases.
Professional Escalation Criteria in the Triage Context
The triage framework provides specific escalation criteria that complement the general criteria described earlier. Escalate when the validation tracking table shows escalate outcomes in multiple layers for the same chromosome, because this pattern indicates a regional assembly failure that requires investigation. Escalate when a documented limitation affects the primary biological question of the study, because the limitation may compromise the conclusions. Escalate when the assembly is intended as a reference for a community or a species with high economic or scientific importance, because the standard for validation should be higher for such assemblies.
The Bactrocera dorsalis assembly illustrates the importance of escalation for species with agricultural significance. The study noted that dipteran insects include numerous harmful species that cause significant agricultural damage, and assembly of genomes for species in this order has been difficult due to small body size, poor conservation of telomere and centromere structures, high heterozygosity, and complex genetic backgrounds [<a href="#ref-2">2</a>]. The assembly of this pest species required a low-input HiFi CCS library from a single male individual and ONT sequence from pooled inbred individuals, and the validation had to confirm complete structural organization information for centromeres and telomeres [<a href="#ref-2">2</a>]. A researcher working on a pest species with economic impact should escalate any validation problem that could affect the utility of the assembly for genetic control strategies.
The Acorus tatarinowii assembly demonstrates the value of the triage framework for species with phylogenetic importance. The study noted that this medicinal herb holds a key position as the sister lineage to all other monocots, underscoring the need for a high-quality genome assembly to facilitate in-depth evolutionary and functional studies [<a href="#ref-3">3</a>]. The assembly reported reference-grade quality with a QV of 45.41, an LAI of 10.04, and BUSCO completeness of 97.6% [<a href="#ref-3">3</a>]. These values all meet the accept criteria, and the triage framework would classify the assembly as ready for publication without escalation.
Frequently Asked Questions
What is the difference between standard assembly validation and T2T validation?
Standard assembly validation focuses on contiguity metrics like N50, completeness metrics like BUSCO, and base accuracy metrics like QV. T2T validation adds specialized checks for telomere presence at chromosome ends, centromere satellite array structure, and correct assembly of large tandem repeats. Standard metrics cannot detect missing telomeres or collapsed centromeres, so T2T validation requires additional methods.
Why is QV not sufficient to validate a T2T assembly?
QV from k-mer analysis measures base accuracy primarily in unique and low-copy regions. In centromeric satellite arrays and other large tandem repeats, the same k-mer appears many times, and errors do not create novel k-mers that can be detected by frequency comparison. The CHM13 polishing study found evidence of small errors and structural misassemblies in the initial draft assembly even though it was derived from highly accurate sequences [<a href="#ref-4">4</a>].
How do I confirm that telomeres are present at all chromosome ends?
Identify the expected telomere repeat motif for the target species, extract the terminal sequences of each assembled chromosome, and search for the motif at the outermost end of each sequence. Confirm that every chromosome has telomeric sequence at both ends. For species with poorly conserved telomere structures, such as dipteran insects, the motif must be identified empirically from the sequencing data [<a href="#ref-2">2</a>].
What does LAI measure and why is it important for T2T validation?
LAI measures the correctness of LTR retrotransposon assembly by comparing the number of intact LTR elements to the number of truncated or solo LTRs. A higher LAI indicates that LTR retrotransposon arrays were assembled without excessive collapse or misjoining. The Acorus tatarinowii assembly reported an LAI of 10.04, which is considered reference-grade for plant genomes [<a href="#ref-3">3</a>].
What should I do if polishing introduces new errors in repetitive regions?
Stop the polishing process and evaluate whether the polishing tool is repeat-aware. Standard automated polishing tools can make errors in large repeats, and the CHM13 study designed a repeat-aware polishing strategy specifically to avoid overcorrection [<a href="#ref-4">4</a>]. Consider using a diverse panel of sequencing technologies for polishing and validation to correct technology-specific errors [<a href="#ref-4">4</a>].
How can optical mapping help validate a T2T assembly?
Optical mapping provides an independent physical measurement of genome structure. It can confirm chromosome length, detect misjoins that preserve local sequence order but disrupt long-range order, and validate structural variants. The ring chromosome study used optical genome mapping for validation after long-read sequencing and T2T assembly [<a href="#ref-5">5</a>].
What records should I keep for reproducible T2T validation?
Document the assembly version and commands, the sequencing data used for validation, the validation metrics and calculation methods, the telomere motif and detection method, the centromere validation method, and any structural validation performed. Use reproducible workflow tools and training resources to support consistent documentation [<a href="#ref-6">6</a>][<a href="#ref-7">7</a>][<a href="#ref-8">8</a>].
When should I escalate a T2T validation problem to additional expertise?
Escalate when telomeres are missing from multiple chromosome ends, when centromeric arrays show collapse or expansion that cannot be resolved, when structural variants cannot be confirmed independently, when polishing introduces new errors, or when validation metrics are inconsistent across methods. Some complex structural variants remain unresolved even with advanced methods, and documenting unresolved cases is appropriate [<a href="#ref-5">5</a>].
Related Bioinformatics Guides
- Metagenomics Assembly: Strategies for Reconstructing Microbial Genomes
- Evaluating Genome Assembly Quality: Metrics and Tools
- RNA-Seq Normalization Methods: TPM, RPKM, and Beyond
- Metagenomic Assembly and Binning: A Practical Workflow for Recovering Genomes from Complex Microbial Communities
- Long-Read Sequencing for De Novo Assembly of Complex Genomes: Case Studies and Best Practices
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
[1] [Telomere-to-telomere assembly of a complete human X chromosome.](https://pubmed.ncbi.nlm.nih.gov/32663838). Nature, 2020. [2] [Telomere-to-telomere genome assembly of the Dipteran Bactrocera dorsalis from a single individual.](https://pubmed.ncbi.nlm.nih.gov/41339345). Nature communications, 2025. [3] [Telomere-to-telomere gapless genome assembly of Acorus tatarinowii.](https://pubmed.ncbi.nlm.nih.gov/41083479). Scientific data, 2025. [4] [Chasing perfection: validation and polishing strategies for telomere-to-telomere genome assemblies.](https://pubmed.ncbi.nlm.nih.gov/35361931). Nature methods, 2022. [5] [Resolution of ring chromosomes, Robertsonian translocations, and complex structural variants from long-read sequencing and telomere-to-telomere assembly.](https://pubmed.ncbi.nlm.nih.gov/39520989). American journal of human genetics, 2024. [6] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project. [7] [nf-core Documentation](https://nf-co.re/docs). nf-core. [8] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries. [9] [Bioconductor](https://bioconductor.org/). Bioconductor Project. [10] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute. [11] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.