Tandem Repeat Analysis with Long Reads: How to Accurately Estimate Repeat Lengths and Identify Expansions
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Long-read sequencing (ONT, PacBio) is essential for accurate tandem repeat length estimation and identification of pathogenic expansions, as reads can span entire repeat arrays, overcoming the limitations of short-read sequencing which often results in indirect and error-prone measurements.
- Library preparation choice is critical: amplification-free methods are recommended to preserve native repeat length and methylation status, as PCR amplification can introduce errors and erase epigenetic marks.
- Repeat genotyping tools like TREAT with otter offer cross-platform compatibility and state-of-the-art accuracy through targeted local assembly, which is superior to alignment-based methods for complex or ambiguous repeat regions.
- Adequate coverage depth, ideally at least 15-fold for genome-wide detection, is crucial to prevent dropout in repeat regions and ensure accurate genotyping; increased depth may be necessary for known pathogenic loci.
- Validation of automatic calls is mandatory, involving visual inspection of aligned reads and orthogonal methods (e.g., Southern blotting) for discordant or borderline results, as automated pipelines can miss variants or misclassify complex repeats.
Tandem repeat expansions are a class of genomic variation where short sequence motifs are repeated in tandem arrays, and abnormal expansion of these repeats is a known cause of dozens of genetic diseases. Short-read sequencing platforms generate reads that are typically too short to span large repeat arrays, which makes repeat length estimation indirect and error-prone. Long-read sequencing from Oxford Nanopore Technologies (ONT) and Pacific Biosciences (PacBio) produces reads that can cover entire repeat regions, enabling direct measurement of repeat length and sequence content. This article provides a practical workflow for researchers and laboratory professionals who need to estimate tandem repeat lengths from long-read data, identify pathogenic expansions, and validate results in a clinical research context. The focus is on concrete decisions: which tools to use, how to prepare libraries, how to assess data quality, and how to interpret results within the limits of current methods.
At a Glance
The table below summarizes the main workflow decisions for tandem repeat analysis with long reads. Each row addresses a core question that arises when designing or executing a repeat expansion study.
| Workflow Decision | Primary Consideration | Recommended Approach |
|---|---|---|
| Sequencing platform | Read length, base accuracy, and throughput differ between ONT and PacBio | Choose ONT for targeted amplification-free sequencing of known repeat loci, choose PacBio HiFi when genome-wide variant detection including SNVs and structural variants is needed |
| Library preparation | PCR amplification can introduce errors in repeat regions and alter methylation signals | Use PCR-free or amplification-free library preparation when repeat length and methylation status must be preserved |
| Repeat genotyping tool | Tools vary in assembly strategy, motif characterization, and cross-platform compatibility | Use a targeted local assembler such as otter within the TREAT workflow for cross-platform accuracy, consider RepeatHMM or TRiC for specific use cases |
| Coverage depth | Low coverage can cause dropout in repeat regions and lead to misgenotyping | Target at least 15-fold coverage for genome-wide detection, increase depth for known pathogenic loci to ensure full repeat coverage |
| Validation | Automatic calls may miss a subset of variants or misclassify complex repeats | Perform visual inspection of aligned reads and use orthogonal methods for discordant or borderline calls |
Context for Tandem Repeat Analysis
Tandem repeats are arrays of short DNA motifs repeated in a head-to-tail fashion. They are distributed throughout the human genome and contribute to both normal variation and disease risk. Abnormal expansion or shortening of tandem repeats can cause a variety of genetic diseases, and the repeat length is often the key diagnostic parameter. The use of long DNA reads has facilitated the analysis of disease-causing repeats in the human genome because long read sequencers can cover whole repeats and therefore allow direct analysis of repeat length and sequence content. This capability is considered suitable for the analysis of long tandem repeats that would be difficult or impossible to resolve with short reads.
The clinical relevance of tandem repeat expansions is well established. For example, the gene C9orf72 harbors a non-coding hexanucleotide repeat expansion known to cause amyotrophic lateral sclerosis and frontotemporal dementia. Previous studies estimated the length of this repeat expansion in multiple tissues, but technological limitations impeded exploration of additional features such as methylation levels. Targeted amplification-free long-read sequencing has since enabled measurement of repeat length, purity, and methylation status in a single assay. This example illustrates why researchers need methods that go beyond simple length estimation.
Long-read sequencing technologies have matured to the point where they can serve as a first-tier test for rare diseases. Clinical short-read exome and genome sequencing approaches have positively impacted diagnostic testing for rare diseases, yet technical limitations associated with short reads challenge their use for the detection of disease-associated variation in complex regions of the genome. Long-read sequencing technologies may overcome these challenges, potentially qualifying as a first-tier test for all rare diseases. In a study of 100 samples with 145 known clinically relevant germline variants that are difficult to detect using short-read sequencing, long-read sequencing at 30-fold high-fidelity coverage re-identified the majority of variants, including about 90 percent of structural variants, single nucleotide variants, insertions or deletions in homologous sequences, and expansions of short tandem repeats. Another 10 percent of variants were visually apparent in the data but not automatically detected, and systematic challenges remained for about 7 percent of variants, such as the detection of AG-rich repeat expansions. Titration analysis showed that 90 percent of all automatically called variants could also be identified using 15-fold coverage. Long-read genomes thus identified 93 percent of challenging pathogenic variants in that dataset.
Core Principles of Repeat Length Estimation
Accurate repeat length estimation from long reads depends on several underlying principles that shape every downstream decision.
Read Length Must Exceed Repeat Length
The fundamental advantage of long-read sequencing for tandem repeat analysis is that a single read can span the entire repeat array plus flanking unique sequence. When a read covers the full repeat, the repeat length can be measured directly from the alignment. If the repeat is longer than the read, the read will start or end within the repeat, and the length estimate becomes partial or ambiguous. This is why the choice of sequencing platform and library preparation method matters: the read length distribution determines the maximum repeat size that can be fully resolved.
Base Accuracy Affects Motif Counting
Repeat length estimation requires counting motif copies. If the base caller introduces errors within the repeat, the motif count can be inflated or deflated. High-fidelity reads from PacBio HiFi sequencing have lower error rates than traditional continuous long reads, which improves the accuracy of motif characterization. ONT reads have improved in accuracy over time, but error profiles still differ from PacBio. The choice of tool must account for the error model of the sequencing platform.
Assembly Versus Alignment
Two broad strategies exist for repeat analysis from long reads. The first is alignment-based: reads are mapped to a reference genome, and repeat lengths are inferred from the alignment. The second is assembly-based: reads spanning the repeat are assembled locally into a consensus sequence, and the repeat is characterized from the assembly. Targeted local assembly can resolve repeats that are difficult to align because the repeat sequence itself may be repetitive and cause ambiguous mapping. Tools such as otter implement fast targeted local assembly and are cross-compatible across different sequencing platforms.
Motif Purity and Interruptions
Repeat arrays are not always pure. They can contain interruptions, where the canonical motif is replaced by a variant motif. These interruptions affect both the biological interpretation and the computational analysis. For example, in the C9orf72 repeat expansion, researchers detected repeat lengths up to 4,088 repeats and found that the expansion contains few interruptions in the blood. The presence of interruptions can complicate motif counting because the repeat is no longer a simple tandem array. Tools that characterize motif composition, such as TREAT, provide information about purity and interruptions that is relevant for clinical interpretation.
Practical Workflow for Repeat Expansion Analysis
The following workflow outlines the steps from sample preparation to validated repeat length estimates. Each step includes concrete decisions and quality checks.
Step 1: Define the Target Loci and Repeat Motif
Before sequencing, define which repeat loci are of interest. This decision determines the library preparation method and the analysis tools. For a targeted study of known disease-associated repeats, such as C9orf72 or RFC1, amplification-free targeted sequencing can enrich the loci of interest without introducing PCR artifacts. For a genome-wide screen, whole-genome long-read sequencing is appropriate, but coverage and cost considerations differ.
Document the expected repeat motif, the genomic coordinates, and the disease-associated threshold for each locus. This information guides the interpretation of results and the design of validation experiments.
Step 2: Choose the Sequencing Platform and Library Preparation
The choice between ONT and PacBio depends on the research question. ONT offers long reads and is well suited for targeted amplification-free sequencing of known repeat loci. PacBio HiFi offers high accuracy and is well suited for genome-wide detection of all variant types, including single nucleotide variants, structural variants, and repeat expansions.
For targeted repeat analysis, use amplification-free library preparation to preserve the native repeat length and methylation status. PCR amplification can introduce errors in repeat regions and can bias the representation of long repeats. The targeted amplification-free long-read sequencing method used for C9orf72 analysis enabled measurement of repeat length, purity, and methylation in a single assay.
For genome-wide analysis, consider the coverage depth. Titration analysis has shown that 90 percent of automatically called challenging variants can be identified using 15-fold coverage, which suggests that reduced coverage may be sufficient for many applications. However, coverage drops can occur in repeat regions, and these drops can lead to misgenotyping. Increasing coverage at known pathogenic loci may be necessary to ensure full repeat coverage.
Step 3: Generate and Demultiplex Sequencing Data
Follow the manufacturer instructions for the sequencing platform. After sequencing, demultiplex the data if multiple samples were pooled. Document the total yield, read length distribution, and quality scores for each sample. These metrics inform downstream analysis and provide a record of data quality.
Step 4: Perform Quality Control on Raw Reads
Quality control is a mandatory step before any repeat analysis. Assess read length distribution, base quality, and coverage across target loci. For targeted sequencing, verify that the on-target rate is sufficient. For whole-genome sequencing, verify that the genome-wide coverage meets the planned depth.
Use established bioinformatics training resources to ensure that quality control steps are performed correctly. The Galaxy Training Network provides accessible workflow training and analysis tutorials that cover quality control and reproducibility. The Carpentries Lessons offer foundational computing and data skills that are useful for researchers who need to build their own analysis pipelines.
Step 5: Align Reads or Perform Targeted Assembly
The choice between alignment and assembly depends on the tool and the repeat characteristics. For alignment-based approaches, map reads to the reference genome using a long-read aligner. For assembly-based approaches, use a targeted local assembler such as otter, which is integrated in the TREAT workflow.
TREAT is an end-to-end workflow for tandem repeat characterization, visualization, and analysis across multiple genomes. It is cross-compatible across different sequencing platforms and achieves state-of-the-art genotyping and motif characterization accuracy when applied to long-read sequencing data from both ONT and PacBio platforms. The otter assembler is fast and targeted, which makes it suitable for analyzing specific repeat loci without assembling the entire genome.
Step 6: Genotype Repeats and Characterize Motifs
Run the repeat genotyping tool to estimate repeat length and motif composition for each locus. The output should include the repeat length for each allele, the motif sequence, and information about interruptions or purity. For clinical research applications, compare the estimated repeat length to the disease-associated threshold.
For the C9orf72 repeat expansion, targeted amplification-free long-read sequencing enabled measurement of repeat length and purity. In a study of 27 individuals with an expanded C9orf72 repeat, researchers obtained 7,765 on-target reads, including 1,612 fully covering the expanded allele. Repeat lengths up to 4,088 repeats were detected, and the expansion contained few interruptions in the blood.
Step 7: Assess Methylation Status When Relevant
If the research question includes epigenetic features, such as methylation of the repeat expansion, use a tool that can quantify methylation from the sequencing data. For the C9orf72 repeat expansion, targeted amplification-free long-read sequencing revealed that the expansion itself is methylated, with great variability in total methylation levels observed, as represented by the proportion of methylated CpGs ranging from 13 to 66 percent. The expanded allele was more highly methylated than the wild-type allele, and increased methylation levels were observed in longer repeat expansions. Methylation levels also correlated with age at collection and age at disease onset.
Methylation analysis requires that the library preparation preserve the native methylation status. PCR amplification can erase methylation marks, so amplification-free preparation is essential for this type of analysis.
Step 8: Validate Results and Document Limitations
Validation is a critical step in clinical research applications. Automatic variant callers may miss a subset of variants or misclassify complex repeats. In the study of 145 challenging pathogenic variants, about 10 percent of variants were visually apparent in the data but not automatically detected. Visual inspection of aligned reads can identify these cases.
For discordant or borderline calls, use orthogonal methods to confirm the repeat length. This may include Southern blotting, repeat-primed PCR, or an alternative sequencing approach. Document the validation method and the outcome for each sample.
Tools for Tandem Repeat Analysis
Several tools are available for tandem repeat analysis from long reads. The choice of tool depends on the sequencing platform, the repeat characteristics, and the research question.
TREAT and otter
TREAT is an end-to-end workflow for tandem repeat characterization, visualization, and analysis across multiple genomes. It integrates otter, a fast targeted local assembler that is cross-compatible across different sequencing platforms. In a comparison with existing tools based on long-read sequencing data from ONT Simplex and Duplex and PacBio Sequel II and Revio, otter and TREAT achieved state-of-the-art genotyping and motif characterization accuracy. Applied to clinically relevant tandem repeats, TREAT and otter significantly identified individuals with pathogenic repeat expansions. In a case-control setting, the workflow replicated previously reported associations of tandem repeats with Alzheimer's disease, including those near or within APOC1, SPI1, and ABCA7 genes.
TREAT and otter can also evaluate potential biases when genotyping tandem repeats using diverse ONT and PacBio long-read sequencing datasets. In rare cases, long-read sequencing shows coverage drops in tandem repeats, including disease-associated repeats in ABCA7 and RFC1 genes. These coverage drops can lead to tandem repeat misgenotyping, hampering the accurate characterization of repeat alleles.
RepeatHMM
RepeatHMM is a tool designed for repeat analysis using hidden Markov models. It can estimate repeat length and characterize motif composition from long-read data. RepeatHMM is appropriate when the repeat structure is complex or when the error profile of the sequencing platform requires a probabilistic model.
TRiC
TRiC is another tool for tandem repeat analysis. It is designed to characterize repeat expansions from long-read data and can be used in targeted or genome-wide analyses. The choice between TRiC and other tools should be based on the specific repeat loci and the sequencing platform.
Tool Selection Criteria
When selecting a tool, consider the following criteria:
- Cross-platform compatibility: Does the tool support both ONT and PacBio data?
- Assembly strategy: Does the tool use targeted local assembly, which can resolve repeats that are difficult to align?
- Motif characterization: Does the tool report motif purity and interruptions?
- Visualization: Does the tool provide visualization of aligned reads and repeat structures?
- Reproducibility: Does the tool support reproducible workflows, such as those provided by nf-core or Bioconductor?
The nf-core Documentation describes community pipeline standards, usage, configuration, and reproducible workflow context. Bioconductor provides official package, workflow, installation, and reproducible genomic-analysis documentation. These resources can help researchers implement reproducible repeat analysis pipelines.
Data Inputs and Quality Metrics
The quality of repeat length estimates depends on the quality of the input data. The following metrics should be recorded for each sequencing run and each sample.
Read Length Distribution
The read length distribution determines the maximum repeat size that can be fully covered by a single read. For targeted sequencing of known repeat loci, verify that the read length exceeds the expected repeat length. For genome-wide analysis, the read length distribution affects the overall ability to resolve long repeats.
Base Quality Scores
Base quality scores affect the accuracy of motif counting. High error rates within the repeat can lead to incorrect motif counts. For PacBio HiFi data, the high accuracy reduces this concern. For ONT data, the error profile should be considered when interpreting results.
Coverage Depth
Coverage depth is a critical metric for repeat analysis. Low coverage can cause dropout in repeat regions, leading to misgenotyping. In rare cases, long-read sequencing shows coverage drops in tandem repeats, including disease-associated repeats in ABCA7 and RFC1 genes. These coverage drops can lead to tandem repeat misgenotyping.
For genome-wide analysis, 15-fold coverage may be sufficient to identify the majority of automatically called challenging variants. However, increasing coverage at known pathogenic loci may be necessary to ensure full repeat coverage and accurate genotyping.
On-Target Rate
For targeted sequencing, the on-target rate measures the proportion of reads that map to the target loci. A low on-target rate reduces the effective coverage and may require additional sequencing.
Methylation Signal
If methylation analysis is planned, verify that the library preparation preserved the native methylation status. The methylation signal can be quantified from the sequencing data, and the proportion of methylated CpGs can be reported for each repeat allele.
Records and Measurements
Maintain detailed records for each sample and each analysis run. The following records support reproducibility and interpretation.
Sample Metadata
Record the sample identifier, tissue type, DNA extraction method, and any relevant clinical information. For repeat expansion disorders, the tissue type can affect repeat length and methylation status. For example, repeat length and methylation can differ between blood and other tissues.
Sequencing Run Metadata
Record the sequencing platform, flow cell type, chemistry version, and base calling software. These details affect the error profile and the interpretation of results.
Library Preparation Records
Record the library preparation method, including whether amplification-free or PCR-based methods were used. This information is essential for interpreting methylation data and for assessing the risk of PCR artifacts.
Analysis Pipeline Records
Record the software versions, parameters, and reference genome version used for each analysis. Reproducibility requires that the analysis pipeline is documented and version-controlled. The nf-core Documentation provides community pipeline standards for reproducible workflow context, and Bioconductor provides official package and workflow documentation.
Quality Control Records
Record the quality metrics for each sample, including read length distribution, base quality scores, coverage depth, and on-target rate. These records support the interpretation of repeat length estimates and provide a basis for troubleshooting.
Validation Records
Record the validation method and outcome for each sample. This includes visual inspection of aligned reads, orthogonal confirmation methods, and any discordant results.
Common Failure Patterns
Several failure patterns recur in tandem repeat analysis from long reads. Recognizing these patterns helps researchers troubleshoot and interpret results.
Coverage Drops in Repeat Regions
Coverage drops can occur in tandem repeats, including disease-associated repeats in ABCA7 and RFC1 genes. These drops can lead to tandem repeat misgenotyping, hampering the accurate characterization of repeat alleles. If coverage drops are observed, consider increasing sequencing depth or using a targeted enrichment approach.
Automatic Call Failures
Automatic variant callers may miss a subset of variants. In a study of 145 challenging pathogenic variants, about 10 percent of variants were visually apparent in the data but not automatically detected. Visual inspection of aligned reads is necessary to identify these cases.
Systematic Challenges for Specific Repeat Types
Some repeat types are systematically difficult to detect. In the same study, about 7 percent of variants presented systematic challenges, such as the detection of AG-rich repeat expansions. These repeats may require specialized tools or manual curation.
PCR Artifacts in Repeat Regions
PCR amplification can introduce errors in repeat regions and can bias the representation of long repeats. If PCR-based library preparation is used, the repeat length estimates may be inaccurate. Amplification-free library preparation is recommended for repeat analysis.
Methylation Signal Loss
If methylation analysis is planned, PCR amplification can erase methylation marks. The methylation signal will be lost or distorted, leading to incorrect conclusions about methylation status. Amplification-free library preparation is essential for methylation analysis.
Misalignment Due to Repetitive Sequence
Repeat sequences can cause ambiguous alignment, especially when the repeat is longer than the read or when the flanking sequence is not unique. Targeted local assembly can resolve these cases by assembling the repeat region without relying on alignment to a reference.
Limitations of Current Methods
Current methods for tandem repeat analysis from long reads have several limitations that researchers should understand.
Coverage Drops and Misgenotyping
Long-read sequencing can show coverage drops in tandem repeats, including disease-associated repeats in ABCA7 and RFC1 genes. These coverage drops can lead to tandem repeat misgenotyping, hampering the accurate characterization of repeat alleles. The frequency of these coverage drops is low, but they can affect specific loci.
Systematic Detection Challenges
Some repeat types are systematically difficult to detect. AG-rich repeat expansions present challenges for automatic detection, and these cases may require manual curation or specialized tools.
Validation Burden
Automatic variant callers may miss a subset of variants, and visual inspection is necessary to identify these cases. This validation burden can be significant for large studies.
Platform-Specific Biases
Different sequencing platforms have different error profiles and biases. TREAT and otter can evaluate potential biases when genotyping tandem repeats using diverse ONT and PacBio long-read sequencing datasets. Researchers should be aware of platform-specific biases when comparing results across platforms.
Interpretation Limits
Repeat length estimates are only one component of clinical interpretation. The presence of interruptions, methylation status, and other features can affect the biological significance of a repeat expansion. Researchers should interpret repeat length estimates in the context of the specific disease and the available clinical evidence.
Safety and Regulatory Context
Tandem repeat analysis from long reads is a research tool, and the results should be interpreted within the applicable regulatory framework. For clinical research applications, the laboratory should follow established quality standards and document all procedures.
Data Management and Reproducibility
Reproducibility is a core requirement for research and clinical applications. The Galaxy Training Network provides accessible workflow training and analysis tutorials that cover reproducibility. The Carpentries Lessons offer foundational computing and data skills that support reproducible analysis. The nf-core Documentation describes community pipeline standards for reproducible workflow context.
Data Resources
The National Center for Biotechnology Information provides official descriptions of databases, search systems, sequence resources, and analysis services. The European Bioinformatics Institute provides bioinformatics learning pathways, data-resource training, and practical analysis education. These resources support researchers in managing and analyzing sequencing data.
Professional Escalation Criteria
Researchers should escalate results to clinical professionals when repeat length estimates exceed disease-associated thresholds or when the results have implications for patient care. The following criteria indicate when professional consultation is appropriate:
- Repeat length exceeds the established pathogenic threshold for the specific locus
- Discordant results between automatic calls and visual inspection
- Coverage drops or other quality issues that prevent confident genotyping
- Methylation status that may affect disease interpretation
- Results that will be used for clinical decision-making
Decision Framework for Selecting Repeat Analysis Tools and Parameters
Selecting the right tool and parameter set for tandem repeat analysis from long reads is a decision that directly affects the accuracy of repeat length estimates and the ability to identify pathogenic expansions. The choice is not a single binary decision but a sequence of decisions that depend on the sequencing platform, the repeat characteristics, the research question, and the available computational resources. This section provides a practical decision framework that researchers can apply before starting an analysis, along with a record system for documenting decisions and a troubleshooting method for common failures.
Decision Point 1: Define the Repeat Locus Characteristics
The first decision is to document the characteristics of the repeat locus itself. This information determines which tools can be applied and which parameters are appropriate. Record the following for each target locus:
- Repeat motif sequence and length
- Expected repeat length range in normal and expanded alleles
- GC content of the motif
- Presence of known interruptions or variant motifs
- Genomic context, including the uniqueness of flanking sequences
- Disease-associated threshold for the specific locus
For example, the C9orf72 hexanucleotide repeat expansion is a non-coding repeat known to cause amyotrophic lateral sclerosis and frontotemporal dementia. Researchers studying this locus need to measure repeat length, purity, and methylation status in a single assay. The repeat can reach lengths up to 4,088 repeats, which corresponds to approximately 25 kilobases. This length exceeds the read length of many sequencing runs, so the choice of platform and library preparation must account for the need to cover the full expansion.
The GC content of the motif affects sequencing performance and base calling accuracy. GC-rich repeats can cause polymerase stalling during library preparation and reduced coverage during sequencing. AG-rich repeat expansions have been identified as systematically challenging for automatic detection, and these loci may require specialized tools or manual curation. Documenting the motif characteristics before analysis helps researchers anticipate these challenges.
Decision Point 2: Match the Tool to the Sequencing Platform
The second decision is to select a genotyping tool that is compatible with the sequencing platform used for the study. Tools vary in their error models, assembly strategies, and cross-platform compatibility. The TREAT workflow with the otter assembler is cross-compatible across different sequencing platforms and achieves state-of-the-art genotyping and motif characterization accuracy when applied to long-read sequencing data from both ONT and PacBio platforms. This cross-platform compatibility is valuable for studies that compare data across sequencing technologies or that need to reanalyze data generated on different platforms.
For ONT data, consider whether the data were generated with Simplex or Duplex sequencing. These two modes have different error profiles, and the choice of tool parameters should reflect these differences. For PacBio data, consider whether the data were generated on Sequel II or Revio instruments, as these platforms have different throughput and accuracy characteristics.
The decision framework should also consider whether the tool uses alignment-based or assembly-based methods. Alignment-based methods map reads to a reference genome and infer repeat length from the alignment. Assembly-based methods, such as the targeted local assembler otter, assemble the repeat region without relying on alignment to a reference. Targeted local assembly can resolve repeats that are difficult to align because the repeat sequence itself may be repetitive and cause ambiguous mapping. For loci with complex repeat structures or non-unique flanking sequences, assembly-based methods are often more reliable.
Decision Point 3: Determine Coverage Requirements
The third decision is to determine the coverage depth needed for the specific research question. Coverage depth affects the ability to detect repeat expansions and the confidence in repeat length estimates. Titration analysis has shown that 90 percent of automatically called challenging variants can be identified using 15-fold coverage. This finding suggests that reduced coverage may be sufficient for many applications, which can reduce sequencing costs and enable larger studies.
However, coverage drops can occur in tandem repeats, including disease-associated repeats in ABCA7 and RFC1 genes. These coverage drops can lead to tandem repeat misgenotyping, hampering the accurate characterization of repeat alleles. The frequency of these coverage drops is low, but they can affect specific loci. For known pathogenic loci, increasing coverage beyond the genome-wide average may be necessary to ensure full repeat coverage and accurate genotyping.
The decision framework should include a coverage plan that specifies the target depth for genome-wide analysis and the additional depth for known pathogenic loci. This plan should be documented before sequencing begins and revisited after the first quality control pass.
Decision Point 4: Choose Library Preparation Based on Downstream Analyses
The fourth decision is to choose the library preparation method based on the downstream analyses planned for the study. PCR amplification can introduce errors in repeat regions and can bias the representation of long repeats. For methylation analysis, PCR amplification can erase methylation marks, making it impossible to quantify methylation status from the sequencing data.
Targeted amplification-free long-read sequencing is the preferred method when repeat length, purity, and methylation status must be measured in a single assay. This method was used to quantify methylation of the C9orf72 repeat expansion in a study of 27 individuals with an expanded repeat. The researchers obtained 7,765 on-target reads, including 1,612 fully covering the expanded allele, and were able to measure repeat length, purity, and methylation status from the same data.
For genome-wide studies that do not require methylation analysis, PCR-based library preparation may be acceptable, but the risk of PCR artifacts in repeat regions should be documented and considered when interpreting results. The decision framework should include a justification for the chosen library preparation method and a record of the potential limitations.
Decision Point 5: Plan the Validation Strategy
The fifth decision is to plan the validation strategy before the analysis begins. Automatic variant callers may miss a subset of variants or misclassify complex repeats. In a study of 145 known clinically relevant germline variants that are challenging to detect using short-read sequencing, long-read sequencing at 30-fold high-fidelity coverage re-identified the majority of variants, including about 90 percent of structural variants, single nucleotide variants, insertions or deletions in homologous sequences, and expansions of short tandem repeats. Another 10 percent of variants were visually apparent in the data but not automatically detected.
The validation strategy should include visual inspection of aligned reads for all samples, with a focus on loci where automatic calls are discordant with expectations or where coverage is low. For discordant or borderline calls, orthogonal methods such as Southern blotting, repeat-primed PCR, or an alternative sequencing approach should be used to confirm the repeat length. The validation plan should specify which samples require orthogonal confirmation and which can be confirmed by visual inspection alone.
Record System for Analysis Decisions
A structured record system supports reproducibility and enables troubleshooting when results are unexpected. The following records should be maintained for each analysis run:
Locus Definition Records
Record the repeat motif, genomic coordinates, expected repeat length range, and disease-associated threshold for each target locus. Include the source of this information, such as a published reference or a clinical database. This record supports the interpretation of results and provides context for discordant calls.
Platform and Library Records
Record the sequencing platform, flow cell type, chemistry version, base calling software, and library preparation method for each sample. Include whether amplification-free or PCR-based methods were used and whether methylation analysis was planned. These details affect the error profile and the interpretation of results.
Tool and Parameter Records
Record the software versions, parameters, and reference genome version used for each analysis. Include the tool selection rationale, such as cross-platform compatibility or assembly strategy. The nf-core Documentation describes community pipeline standards, usage, configuration, and reproducible workflow context that can support this documentation. Bioconductor provides official package, workflow, installation, and reproducible genomic-analysis documentation.
Quality Control Records
Record the quality metrics for each sample, including read length distribution, base quality scores, coverage depth, and on-target rate. These records support the interpretation of repeat length estimates and provide a basis for troubleshooting.
Validation Records
Record the validation method and outcome for each sample. This includes visual inspection of aligned reads, orthogonal confirmation methods, and any discordant results. The validation record should specify whether the automatic call was confirmed, revised, or rejected.
Troubleshooting Method for Discordant Results
When repeat length estimates are discordant with expectations or when automatic calls fail, a structured troubleshooting method can identify the cause and guide corrective action.
Step 1: Check Coverage at the Locus
Coverage drops can occur in tandem repeats, including disease-associated repeats in ABCA7 and RFC1 genes. These drops can lead to tandem repeat misgenotyping. Examine the coverage profile across the repeat locus and the flanking regions. If coverage is low or absent, consider whether the library preparation or sequencing run failed to capture the locus. Increasing sequencing depth or using a targeted enrichment approach may be necessary.
Step 2: Inspect Aligned Reads Visually
Automatic variant callers may miss a subset of variants that are visually apparent in the data. Visual inspection of aligned reads can identify these cases. Examine the read alignments across the repeat locus, looking for reads that span the full repeat and reads that start or end within the repeat. The pattern of read starts and ends can indicate whether the repeat is longer than the read length or whether the alignment is ambiguous.
Step 3: Evaluate Motif Purity and Interruptions
Repeat arrays can contain interruptions, where the canonical motif is replaced by a variant motif. These interruptions affect both the biological interpretation and the computational analysis. For example, in the C9orf72 repeat expansion, researchers found that the expansion contains few interruptions in the blood. If the repeat contains interruptions, the motif counting algorithm may underestimate or overestimate the repeat length. Tools that characterize motif composition, such as TREAT, provide information about purity and interruptions that is relevant for clinical interpretation.
Step 4: Assess Platform-Specific Biases
Different sequencing platforms have different error profiles and biases. TREAT and otter can evaluate potential biases when genotyping tandem repeats using diverse ONT and PacBio long-read sequencing datasets. If results differ between platforms, consider whether the difference is due to platform-specific biases or to a genuine biological difference.
Step 5: Review Library Preparation Records
PCR amplification can introduce errors in repeat regions and can bias the representation of long repeats. If PCR-based library preparation was used, review the records to determine whether the repeat length estimates may be affected by PCR artifacts. For methylation analysis, PCR amplification can erase methylation marks, so the methylation signal may be lost or distorted.
Step 6: Consult Training and Support Resources
When troubleshooting is inconclusive, consult established training and support resources. The Galaxy Training Network provides accessible workflow training and analysis tutorials that cover quality control and reproducibility. The Carpentries Lessons offer foundational computing and data skills that support reproducible analysis. The European Bioinformatics Institute provides bioinformatics learning pathways, data-resource training, and practical analysis education.
Common Failure Patterns and Corrective Actions
The following table summarizes common failure patterns in tandem repeat analysis from long reads and the corrective actions that can be taken.
| Failure Pattern | Likely Cause | Corrective Action |
|---|---|---|
| Low or absent coverage at the repeat locus | Coverage drop in the repeat region, failed enrichment, or insufficient sequencing depth | Increase sequencing depth, use targeted enrichment, or verify library preparation |
| Automatic call misses a variant that is visually apparent | Tool limitation or parameter mismatch | Perform visual inspection of aligned reads and manually curate the call |
| Repeat length estimate is shorter than expected | Read length is shorter than the repeat, or the repeat contains interruptions | Use a platform with longer reads, use targeted local assembly, or characterize motif interruptions |
| Repeat length estimate is longer than expected | Base calling errors inflate the motif count | Use higher accuracy data, adjust tool parameters, or use a tool with a more appropriate error model |
| Methylation signal is absent or distorted | PCR amplification erased methylation marks | Use amplification-free library preparation for methylation analysis |
| Discordant results between platforms | Platform-specific biases in error profile or coverage | Evaluate platform-specific biases with tools such as TREAT and otter, and interpret results in the context of the platform |
Professional Escalation Criteria
Researchers should escalate results to clinical professionals when repeat length estimates exceed disease-associated thresholds or when the results have implications for patient care. The following criteria indicate when professional consultation is appropriate:
- Repeat length exceeds the established pathogenic threshold for the specific locus
- Discordant results between automatic calls and visual inspection that cannot be resolved
- Coverage drops or other quality issues that prevent confident genotyping
- Methylation status that may affect disease interpretation
- Results that will be used for clinical decision-making
The decision framework, record system, and troubleshooting method described in this section provide a structured approach to tandem repeat analysis from long reads. By documenting decisions, maintaining detailed records, and following a systematic troubleshooting method, researchers can improve the accuracy and reproducibility of their repeat length estimates and reduce the risk of misgenotyping.
Frequently Asked Questions
What is the main advantage of long-read sequencing for tandem repeat analysis?
Long-read sequencing produces reads that can cover entire repeat regions, enabling direct measurement of repeat length and sequence content. Short-read sequencing generates reads that are typically too short to span large repeat arrays, which makes repeat length estimation indirect and error-prone. Long read sequencers enable direct analysis of repeat length and sequence content by covering whole repeats, and they are therefore considered suitable for the analysis of long tandem repeats.
How do ONT and PacBio compare for repeat expansion analysis?
ONT offers long reads and is well suited for targeted amplification-free sequencing of known repeat loci. PacBio HiFi offers high accuracy and is well suited for genome-wide detection of all variant types. Tools such as TREAT and otter are cross-compatible across both platforms and achieve state-of-the-art genotyping and motif characterization accuracy when applied to data from ONT Simplex and Duplex and PacBio Sequel II and Revio.
Why is amplification-free library preparation important for repeat analysis?
PCR amplification can introduce errors in repeat regions and can bias the representation of long repeats. For methylation analysis, PCR amplification can erase methylation marks. Amplification-free library preparation preserves the native repeat length and methylation status, which is essential for accurate repeat length estimation and methylation quantification.
What coverage depth is needed for repeat expansion detection?
Titration analysis has shown that 90 percent of automatically called challenging variants can be identified using 15-fold coverage. However, coverage drops can occur in tandem repeats, including disease-associated repeats in ABCA7 and RFC1 genes, and these drops can lead to misgenotyping. Increasing coverage at known pathogenic loci may be necessary to ensure full repeat coverage.
What tools are available for tandem repeat genotyping from long reads?
TREAT is an end-to-end workflow for tandem repeat characterization, visualization, and analysis across multiple genomes. It integrates otter, a fast targeted local assembler that is cross-compatible across different sequencing platforms. Other tools include RepeatHMM and TRiC. The choice of tool depends on the sequencing platform, the repeat characteristics, and the research question.
How can I validate repeat expansion calls?
Automatic variant callers may miss a subset of variants, and visual inspection of aligned reads is necessary to identify these cases. For discordant or borderline calls, use orthogonal methods to confirm the repeat length, such as Southern blotting, repeat-primed PCR, or an alternative sequencing approach. Document the validation method and the outcome for each sample.
Can methylation status be measured from long-read repeat analysis?
Yes, targeted amplification-free long-read sequencing can quantify methylation of repeat expansions. For the C9orf72 repeat expansion, researchers measured the proportion of methylated CpGs and found great variability in total methylation levels, ranging from 13 to 66 percent. The expanded allele was more highly methylated than the wild-type allele, and increased methylation levels were observed in longer repeat expansions.
What are the limitations of current repeat analysis methods?
Current methods have several limitations, including coverage drops in tandem repeats that can lead to misgenotyping, systematic detection challenges for specific repeat types such as AG-rich repeats, and the need for visual inspection to identify variants missed by automatic callers. Platform-specific biases can also affect results, and researchers should evaluate these biases when comparing data across platforms.
Related Bioinformatics Guides
- Detecting Structural Variants with Long-Read Sequencing: Methods and Considerations
- Genomic Data Analysis Tools: A Comparative Guide for Researchers
- Longitudinal Microbiome Data Analysis: Methods and Best Practices
- Long-Read Sequencing Cost and Market: What to Expect
- Long-Read Sequencing for Isoform Quantification: Challenges and Solutions
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- Analysis of Tandem Repeat Expansions Using Long DNA Reads.. Methods in molecular biology (Clifton, N.J.), 2023.
- Characterizing tandem repeat complexities across long-read sequencing platforms with TREAT and otter.. Genome research, 2024.
- Targeted long-read sequencing to quantify methylation of the C9orf72 repeat expansion.. Molecular neurodegeneration, 2024.
- Tandem repeats in the long-read sequencing era.. Nature reviews. Genetics, 2024.
- HiFi long-read genomes for difficult-to-detect, clinically relevant variants.. American journal of human genetics, 2025.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.