Zubair Khalid

Virologist/Molecular Biologist | Veterinarian | Bioinformatician

Conventional & Molecular Virology • Vaccine Development • Computational Biology

Dr. Zubair Khalid is a veterinarian and virologist specializing in conventional and molecular virology, vaccine development, and computational biology. Dedicated to advancing animal health through innovative research and multi-omics approaches.

Dr. Zubair Khalid - Veterinarian, Virologist, and Vaccine Development Researcher specializing in Computational Biology, Multi-omics, Animal Health, and Infectious Disease Research

Category: Guides

Repeat Expansion Sequencing

Repeat expansion sequencing is the set of laboratory and computational methods used to detect and quantify abnormally long repetitive DNA sequences that are linked to many genetic disorders. This guide is written for molecular biologists, clinical geneticists, bioinformaticians, and research scientists who need a practical, source bounded framework for designing, executing, and interpreting repeat expansion experiments. It will help you decide which sequencing strategy fits your repeat of interest, avoid common pitfalls, and recognize the limits of what the data can tell you. The foundational concepts and technical workflows described here are anchored in authoritative resources such as the NCBI Bookshelf NCBI Bookshelf and EMBL EBI Training materials EMBL EBI Training.

At a Glance

Aspect Details
Definition Sequencing approaches tailored to detect long tandem repeats that exceed typical read lengths (e.g., 5 100+ bp).
Primary technologies Long read sequencing (PacBio, Oxford Nanopore), targeted PCR based enrichment, or hybrid bioinformatics calling from short reads.
Key applications Diagnosis of repeat expansion disorders (Huntington disease, Friedreich ataxia, myotonic dystrophy), research on regulatory repeats, and forensic analysis.
Main bioinformatics tools ExpansionHunter, STRling, GangSTR, RepeatMasker, and Tandem Repeats Finder.
Critical quality metrics Read depth across repeat region, repeat unit consistency, phasing of expansions, and control for GC bias.
Common pitfalls Using short reads for very large repeats, ignoring somatic mosaicism, and misinterpreting sequencing errors as expansions.
Limits of interpretation Cannot resolve expansions larger than read length without specialized protocols, unknown normal ranges for many loci.

Core Concepts of Repeat Expansion Sequencing

Tandem repeats are stretches of DNA where a short motif (2 6 base pairs for microsatellites, or longer for minisatellites) is repeated in series. Most human genomes contain thousands of these repeats, and their length can vary between individuals. When the repeat count crosses a pathogenic threshold, it can cause disease by altering gene expression, producing toxic RNA or protein, or interfering with DNA replication. [2] The EMBL EBI Training module on repeat analysis provides a useful overview of repeat biology and its clinical impact EMBL EBI Training.

Traditional PCR based genotyping fails when the repeat becomes too long to amplify or when GC content is high. Standard short read whole genome sequencing (WGS) also misses large expansions because the reads (typically 150 bp) cannot span the entire repeat. Repeat expansion sequencing overcomes this by using two main strategies. The first uses long read technology (PacBio HiFi or Oxford Nanopore) to produce reads that are long enough to cover the entire repeat. The second uses targeted capture or PCR of the repeat locus combined with either long reads or specialized short read callers that detect paired end reads containing the repeat.

A crucial concept is that repeat expansions are often not homogeneous. Somatic mosaicism, where different cells carry different repeat lengths, is common in disorders like Huntington disease and myotonic dystrophy. [10] A recent bioRxiv study using iPSC models of congenital myotonic dystrophy highlighted aberrant differentiation linked to such mosaicism Aberrant neuronal differentiation and splicing defects in Congenital Myotonic Dystrophy (DM1) iPSC models. This means that a single measurement from bulk tissue may not represent the full spectrum of repeat sizes in an individual.

Decision Criteria: When to Use Repeat Expansion Sequencing

Deciding whether to use repeat expansion sequencing depends on the repeat locus, the expected size of the expansion, and the available resources. The Galaxy Training Network offers decision trees for selecting sequencing approaches in genomics Galaxy Training Network. Here are key criteria:

  • Repeat size relative to read length If the expected pathogenic repeat count produces a locus longer than 150 300 bp, short read WGS alone cannot reliably call the expansion. You need either long read sequencing or a targeted approach that captures the repeat in fragments that can be spanned.

  • GC content and sequence complexity Repeats with high GC content (e.g., the GAA repeat in Friedreich ataxia) are prone to PCR bias and underrepresentation in short read libraries. Long read sequencing or nanopore direct sequencing can mitigate this. CRISPR Cas9 based therapies for Friedreich ataxia are an active area of research, underscoring the clinical importance of accurate expansion measurement CRISPR-Cas9-based therapies for Huntington's disease and Friedreich's ataxia: mechanisms, advances, and future perspectives.

  • Need for phasing If you need to know whether both alleles are expanded (e.g., in recessive disorders), short reads often cannot phase the repeats. Long reads or linked read technologies are necessary.

  • Somatic mosaicism characterization If your research question involves measuring variation across tissues or cell types, single cell or ultra deep long read sequencing may be required. Bulk short read methods average out mosaic lengths.

  • Cost and throughput Targeted PCR or capture based enrichment followed by short read sequencing is cheaper per sample than whole genome long read sequencing. For rare variant discovery in large cohorts, short read based callers like ExpansionHunter are efficient first passes.

Practical Workflow (Implementation Sequence)

The workflow for repeat expansion sequencing can be broken into five stages. The Bioconductor project provides R packages for each step, including quality assessment and visualization Bioconductor.

Stage 1: Sample preparation and nucleic acid extraction Use high molecular weight (HMW) DNA when planning long read sequencing. Avoid shearing to keep fragments long. For targeted PCR based methods, design primers that flank the repeat and include a GC rich region if needed.

Stage 2: Library construction

  • Option A Long read library: PacBio HiFi libraries require 10 20 kb inserts, Oxford Nanopore ligation based libraries work with 5 20 kb fragments. Use size selection if only large expansions are of interest.
  • Option B Targeted enrichment: Hybrid capture with biotinylated probes covering the repeat and flanking regions. Alternatively, long range PCR with proofreading polymerase (e.g., TaKaRa LA Taq) for amplicons up to 10 kb.
  • Option C Short read library for bioinformatics calling: Standard Illumina libraries suffice, the bioinformatics tool will detect paired end reads where one mate contains the repeat.

Stage 3: Sequencing

  • PacBio HiFi: 15 20 kb insert size, Q20 base quality.
  • Oxford Nanopore: R10.4 flow cells, high accuracy (Q20+) basecalling.
  • Illumina: paired end 150 bp, 30x or deeper coverage for detection of mosaic expansions.

Stage 4: Bioinformatics analysis

  • Align reads to a reference genome that includes known repeat annotations. Use minimap2 for long reads, BWA MEM for short reads.
  • Run a repeat caller. Popular tools:
    • ExpansionHunter (short reads, targeted via a variant catalog)
    • STRling (short reads, de novo detection)
    • GangSTR (short reads, genotype likelihood)
    • Tandem Repeats Finder (long reads)
  • For long reads, inspect alignments manually in IGV to verify the repeat spans.
  • Estimate repeat length. For long reads, the repeat length is directly inferred from the alignment. For short reads, the caller provides a posterior estimate with confidence intervals.

Stage 5: Quality checks and validation

  • Confirm coverage depth at the repeat flank is at least 20x for short read callers.
  • Check for read pile up that suggests alignment errors or collapsed repeats.
  • Validate a subset of samples with an orthogonal method, such as Southern blot or repeat primed PCR. The NCBI Sequence Read Archive (SRA) can be used to deposit and share raw data NCBI Sequence Read Archive.

Quality Checks and Validation

High quality repeat expansion sequencing data must pass several internal checks before you trust the called repeat size.

First, measure the depth of coverage across the repeat locus and its immediate flanks. Low depth (below 10x for long reads, below 30x for short reads) increases the chance of missing an expansion or underestimating its size. For short reads, ExpansionHunter outputs a number of supporting read pairs. Fewer than 10 supporting pairs should raise a flag.

Second, inspect read level evidence for GC bias. In a marine derived fungal genome study, researchers found AT rich isochores with putative regulatory functions, highlighting how GC content can affect sequencing representation A marine-derived fungal genome of Annulohypoxylon annulatoides reveals AT-rich isochores with putative regulatory functions. If your repeat is AT rich, expect lower coverage in short read data, and adjust your confidence accordingly.

Third, confirm that the called repeat length is consistent across multiple reads for long read data. A single read with an unusually long repeat could be a chimeric alignment. Look for at least five independent long reads that support the same length.

Fourth, verify that the repeat motif matches the expected pattern. Tools like Tandem Repeats Finder report the consensus motif. A mismatch may indicate a different repeat type or a sequencing error.

Finally, for clinical applications, always validate with a second method. Repeat primed PCR and Southern blot remain gold standards for large expansions.

Common Mistakes

  • Using short reads alone for large expansions A 150 bp read cannot span a 500 bp repeat. Short read callers can only infer the presence of an expansion based on read pairs that straddle the flank, but they cannot measure the exact length beyond a few hundred base pairs. This leads to ambiguous calls.

  • Ignoring the normal range Not every long repeat is pathogenic. Reference population databases like gnomAD STR provide frequency data. A repeat that is long but within the normal range should not be interpreted as disease causing.

  • Misattributing sequencing errors to expansions Nanopore homopolymer errors can create apparent insertions that mimic repeat motifs. Use Q20+ basecalling and filter reads with low quality in the repeat tract.

  • Assuming homogeneity Somatic mosaicism is the rule, not the exception, for many repeat disorders. Reporting a single repeat length from bulk tissue may misrepresent the true disease course. Consider reporting the range or the most common length using multiple reads.

  • Overlooking the need for phasing If you genotype a heterozygous repeat but do not know which allele is expanded, you lose critical diagnostic information for dominant disorders.

Limits of Interpretation

Repeat expansion sequencing, even with the best methods, has inherent limits that you must acknowledge.

The most fundamental limit is that repeat length estimates from short reads are probabilistic, not exact. Short read callers estimate the repeat size using the insert size distribution. This works well for small to moderate expansions (up to about 200 bp beyond read length), but confidence intervals widen for larger expansions. [7] A recent paper on TREPP (Tandem Repeat Expansion Pathogenicity Prediction) uses stacked CatBoost models to predict pathogenicity, which is an alternative to direct sizing TREPP: Tandem Repeat Expansion Pathogenicity Prediction via Stacked CatBoost and Context-Aware Sequence Features.

Long reads can directly measure the repeat length, but they struggle with very long expansions (thousands of base pairs) because the reads may not fully span them. Even with 20 kb reads, a 10 kb expansion might be missed if the flanking sequence is too short to align uniquely.

Another limit is that some repeats are located in extremely repetitive regions of the genome, such as centromeres or subtelomeres. Mapping reads from such regions is ambiguous, and no current tool can confidently assign a repeat length to a specific copy.

Furthermore, the evolutionary dynamics of repeats vary across species. For example, the mitochondrial control region in vertebrates shows very high substitution rates that can confound repeat annotation Evolutionary patterns of the mitochondrial control region in vertebrates: A large-scale comparative analysis. When applying repeat expansion sequencing in non model organisms, expect higher uncertainty.

Finally, the biological meaning of an expansion may not be clear. A repeat that is long in one individual may be benign, while the same length in another individual may be pathogenic due to sequence composition or genetic background. [11] Paralogs and gene families, such as the TLO family in Candida albicans, can contain repeats that are functionally interconnected with incomplete redundancy Paralogs of the Candida albicans TLO gene family form interconnected functional networks with incomplete redundancy. This reminds us that repeat length alone is not always the sole determinant of phenotype.

Frequently Asked Questions

1. What is the minimum repeat size that repeat expansion sequencing can detect?

The detection limit depends on the method. Short read callers like ExpansionHunter typically require at least five repeats beyond the reference to distinguish an expansion from noise. For a 3 bp motif, that means expansions of 15 bp or more are detectable. Long reads can detect the exact size of any repeat that is fully covered by a single read, starting from a few repeats upward.

2. Can repeat expansion sequencing distinguish somatic expansions from germline ones?

Yes and no. If you sequence DNA from a bulk tissue sample, you will see a distribution of repeat lengths. The modal length is often germline, while the tail of longer repeats indicates somatic expansion. However, to definitively assign a repeat to a specific cell, single cell sequencing is required. Software like STRling reports the distribution, but interpreting the germline length still requires caution.

3. How do I handle repeats with compound motifs (e.g., CAG/CAA)?

Most tools expect a single repeat unit. For compound motifs, you can either treat the entire composite as one repeat or analyze each motif separately. The Galaxy Training Network has a tutorial on handling complex repeats that discusses this issue Galaxy Training Network. In practice, manual examination of the aligned reads is often needed.

4. Is whole genome long read sequencing cost effective for repeat expansion screening?

For large cohorts, no. Targeted sequencing of known repeat loci is far cheaper. For discovery of novel expansion loci in rare disease families, whole genome long read sequencing is justified. The average cost per genome for PacBio HiFi is around $1000 $2000, which is falling.

References and Further Reading

Related Articles