Variant Calling in Repetitive Elements: How to Handle Alu, LINE, and Satellite Repeats
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Repetitive elements like Alu, LINE, and satellite DNA constitute approximately half of the human genome, posing significant challenges for accurate variant calling due to multi-mapping short sequencing reads. These reads, often shorter than repeat units, can align to multiple genomic locations, leading to false positive variant calls that are artifactually supported by inflated read depth.
- A robust strategy for repeat-aware variant calling involves a multi-layered approach: masking or annotating repeat regions in the reference genome prior to alignment, employing aligners (e.g., BWA-MEM, minimap2) that explicitly handle multi-mapping reads and report mapping quality, and implementing post-calling filters to remove variants supported solely by ambiguous read placements.
- Standard variant calling workflows fail in repeat regions because alignment quality scores are unreliable, variant callers can be misled by multi-mapping reads inflating read depth, and small insertions/deletions are particularly problematic due to alignment ambiguity and potential misplacement of reads containing indels within repeat units.
- Long-read sequencing technologies (e.g., Oxford Nanopore, PacBio HiFi) offer a partial solution by producing reads long enough to span repeat units, thereby reducing multi-mapping and enabling more accurate variant localization. However, long-read platforms have their own error profiles, and indel calling in repeats remains challenging, necessitating confirmation of critical variants with these methods rather than as a complete replacement for short-read analysis.
- Satellite repeats, characterized by long tandem arrays of very short repeat units (e.g., alpha satellites in centromeres), present the most intractable variant calling problem, often requiring complete exclusion from analysis and masking of centromeric/pericentromeric regions due to the near impossibility of resolving alignment ambiguity with current short-read technologies.
Repetitive elements such as Alu sequences, LINE elements, and satellite repeats create a specific and well-documented problem in variant calling: short sequencing reads that originate from different copies of a repeat can align to the same genomic location, producing multi-mapping reads and misalignments that generate false variant calls. The practical solution combines three layers of control: masking or annotation of repeat regions before alignment, use of aligners and callers that handle multi-mapping reads explicitly, and post-calling filtering that removes variants supported only by ambiguous read placements. This article provides a workflow-oriented approach for biology students, researchers, laboratory professionals, and life-science practitioners who need to produce reliable variant calls from genomes rich in repetitive content.
The Scale and Nature of the Repetitive Element Problem
Repetitive DNA sequences compose roughly half of the human genome, and transposable elements such as Alu, SVA, HERV, and L1 elements are prominent contributors to this repeat content. These elements can cause disease through disrupting genes, causing frameshift mutations, or altering splicing patterns, which makes accurate variant detection in these regions clinically relevant. The challenge is that short-read sequencing technologies produce reads that are often shorter than the repeat units themselves, so a read originating from one copy of an Alu element can align equally well to many other Alu copies scattered across the genome.
The core technical issue is multi-mapping. When a read aligns to multiple genomic locations with identical or nearly identical alignment scores, the aligner must decide where to place it. Many aligners place the read at one location arbitrarily or report it as multi-mapped. Variant callers that receive these ambiguous alignments may then support a variant call with reads that actually originate from a different genomic location. The result is a false positive variant that looks well-supported by read depth but is actually an artifact of repeat structure.
Satellite repeats present an additional challenge because they are arranged in long tandem arrays, often in centromeric and pericentromeric regions. The repetitive unit can be very short, and the array can extend for megabases. Short reads from these arrays produce alignment patterns that are difficult to resolve even with repeat-aware methods. Long-read sequencing provides a partial solution because reads can span entire repeat units and sometimes entire array substructures, but the error profiles of long-read platforms introduce their own variant calling complications.
Why Standard Variant Calling Workflows Fail in Repeat Regions
Standard variant calling workflows typically follow a linear path: raw reads are aligned to a reference genome, alignments are sorted and deduplicated, and a variant caller identifies positions where the aligned reads disagree with the reference. This path works well in unique regions of the genome but breaks down in repetitive regions for several reasons.
First, read alignment quality scores in repeat regions are unreliable. A read that aligns to a repeat copy with a few mismatches may receive a high mapping quality because the aligner does not account for the possibility that the read could align to another repeat copy with the same score. The mapping quality score is supposed to reflect the probability that the read is placed correctly, but many aligners underestimate the ambiguity in repetitive regions.
Second, variant callers that use read depth as evidence for a variant can be misled by multi-mapping reads. If a caller counts all reads that align to a position, including multi-mapped reads, the depth at that position may be inflated. Conversely, if the caller filters multi-mapped reads, the depth may be too low to make a confident call. Both scenarios produce incorrect variant calls.
Third, small insertions and deletions are particularly problematic in repeat regions. A read that contains an indel within a repeat unit may align better to a different repeat copy without the indel, causing the aligner to place the read incorrectly. The variant caller then misses the true indel or calls a false one at the wrong location.
Fourth, structural variant calling in repeat regions is confounded by the same alignment ambiguity. A structural variant that involves a repeat element may be represented in the reads as a cluster of soft-clipped bases or discordant read pairs, but distinguishing a true structural variant from alignment artifacts in repeats requires specialized methods.
At a Glance: Repeat-Aware Variant Calling Strategy
| Workflow Stage | Primary Action | Key Tool or Method | Expected Outcome |
|---|---|---|---|
| Reference preparation | Mask or annotate known repeat regions before alignment | RepeatMasker output, UCSC RepeatMasker track, or curated repeat annotations | A reference genome with repeat regions flagged for downstream filtering |
| Read alignment | Use an aligner that reports mapping quality and handles multi-mapping reads explicitly | BWA-MEM with proper multi-mapping flags, minimap2 for long reads | Alignments with mapping quality values that reflect repeat ambiguity |
| Variant calling | Use a caller that filters low mapping quality and multi-mapping reads | GATK HaplotypeCaller with repeat filtering, DeepVariant with careful thresholding | Variant calls that exclude most repeat-induced artifacts |
| Post-calling filter | Remove variants in repeat regions unless supported by unique evidence | BCFtools filtering, custom region-based filters, repeat annotation overlap | A final variant set with repeat-region variants flagged or removed |
| Long-read confirmation | Validate suspected variants in repeats with long-read sequencing | Sniffles2 for structural variants, sTELLeR for transposable element detection | Confirmed or refuted variant calls in previously ambiguous regions |
This table summarizes the five-stage approach that this article develops in detail. Each stage is described with concrete decision criteria in the sections that follow.
Reference Preparation: Masking and Annotation
The first decision in a repeat-aware variant calling workflow is whether to mask repeats in the reference genome before alignment. Masking replaces repeat bases with N characters or converts them to lowercase, which prevents reads from aligning to those regions. The tradeoff is that masking removes the possibility of calling variants in those regions entirely. For many clinical and research applications, this is acceptable because the regions are difficult to interpret even with the best methods.
An alternative to hard masking is soft masking, where repeat bases are converted to lowercase but not replaced with N. Some aligners treat lowercase bases as repeat regions and reduce the alignment score for reads that align there, but still allow alignment. This approach preserves the possibility of variant calling in repeats while discouraging spurious alignments.
A third approach is to leave the reference unmasked but annotate repeat regions separately. The annotation is then used as a filter after variant calling. This approach preserves all information but requires careful downstream filtering and increases the risk of false positives if the filter is not applied correctly.
The choice between masking and annotation depends on the research question. If the goal is to identify variants in unique regions with high confidence, hard masking is the safest choice. If the goal is to explore variants in or near repeats, annotation-based filtering is more appropriate. If the goal is clinical variant reporting, the decision should follow the validated pipeline used by the testing laboratory, because changing masking strategy can change the set of reportable variants.
For reference preparation, the NCBI Data Resources provide access to reference genomes, repeat annotations, and variation databases that support repeat-aware analysis. The EMBL-EBI Training materials include practical guidance on using reference annotations in bioinformatics workflows.
Read Alignment: Managing Multi-Mapping Reads
The aligner is the first computational step where repeat ambiguity can be managed. The key output from the aligner is the mapping quality score, which estimates the probability that a read is placed at the correct genomic location. In repeat regions, mapping quality scores are often low because the aligner detects that the read could align to multiple locations.
For short-read data, BWA-MEM is a common choice because it reports mapping quality and provides flags for multi-mapping reads. Reads that align to multiple locations are marked with the XT flag or receive a low mapping quality score. The variant caller can then use these flags to exclude multi-mapping reads from evidence.
For long-read data, minimap2 is the standard aligner. Long reads can span repeat units, which reduces but does not eliminate multi-mapping. A long read that spans an entire Alu element plus flanking unique sequence can be placed uniquely, but a read that is entirely within a repeat will still be ambiguous.
The practical decision is the mapping quality threshold. A common threshold is to exclude reads with mapping quality below 20, which corresponds to a 1 percent probability of incorrect placement. However, in repeat regions, this threshold may be too permissive. Some workflows use a threshold of 30 or higher for repeat regions, or they exclude multi-mapping reads entirely regardless of mapping quality.
The Galaxy Training Network provides accessible tutorials on read alignment and mapping quality interpretation that are useful for practitioners who are new to these concepts. The nf-core Documentation describes how community-standard pipelines handle alignment and mapping quality in production workflows.
Variant Calling: Repeat-Aware Callers and Filters
Variant callers differ substantially in how they handle reads in repetitive regions. The most important distinction is whether the caller uses local reassembly or relies on direct read comparison to the reference.
Local reassembly callers, such as GATK HaplotypeCaller, reconstruct the local haplotype in regions where reads disagree with the reference. This approach can resolve variants in repeats if the reads provide enough coverage to assemble the repeat unit correctly. However, if the reads are multi-mapped, the assembly may include reads from different repeat copies, producing a chimeric haplotype that does not exist in the sample.
Direct comparison callers, such as freebayes, count alleles at each position based on the aligned reads. These callers are more susceptible to multi-mapping artifacts because they do not attempt to resolve the local sequence context.
For somatic variant calling, the problem is compounded by the need to distinguish true somatic variants from germline variants and from artifacts. Somatic callers often use paired tumor-normal samples and apply statistical models to identify variants that are present only in the tumor. In repeat regions, the artifact rate is higher, so the statistical models must be calibrated with repeat-aware priors.
The Bioconductor project provides R packages for variant calling and filtering that support repeat-aware analysis. These packages allow practitioners to build custom filtering pipelines that incorporate repeat annotations.
Post-Calling Filtering: Removing Repeat Artifacts
After variant calling, the most important step is filtering the raw variant calls to remove artifacts from repetitive regions. The filter should be applied at multiple levels.
The first level is mapping quality filtering. Variants supported by reads with low mapping quality should be removed or flagged. The threshold depends on the aligner and the variant caller, but a common approach is to require a minimum mapping quality of 20 for all reads supporting a variant.
The second level is repeat annotation overlap. Variants that fall within annotated repeat regions should be flagged. The decision to remove or retain these variants depends on the research question. For clinical reporting, variants in repeats are typically confirmed with an orthogonal method before being reported.
The third level is read placement bias. In a true variant, the supporting reads should align to the variant location with their full length. In a repeat artifact, the supporting reads may be clipped at the repeat boundary or may show unusual alignment patterns. Tools that examine the alignment structure of supporting reads can identify these artifacts.
The fourth level is strand bias and read position bias. True variants are typically supported by reads from both strands and from various positions within the reads. Artifacts in repeats often show strong bias because the artifact is created by a specific alignment pattern.
The The Carpentries Lessons provide foundational training in data analysis workflows that are useful for implementing reproducible filtering pipelines. The Galaxy Training Network includes tutorials on variant filtering that demonstrate practical filtering steps.
Long-Read Sequencing as a Resolution Strategy
Long-read sequencing platforms from Oxford Nanopore Technologies and Pacific Biosciences produce reads that are thousands to tens of thousands of bases long. These reads can span entire Alu elements, LINE elements, and even portions of satellite repeat arrays. The ability to span repeats changes the variant calling problem fundamentally.
A long read that spans a repeat element and its flanking unique sequence can be placed uniquely in the genome. The variant caller can then use the sequence within the repeat to identify variants that are specific to that copy. This approach resolves the multi-mapping problem for many repeat types.
However, long-read sequencing has its own error profile. Oxford Nanopore reads have higher error rates than short reads, particularly for homopolymers and small indels. Pacific Biosciences HiFi reads have lower error rates but are shorter than Nanopore reads. The variant caller must account for these error profiles to avoid false variant calls.
A study of long-read sequencing for genomic profiling of myeloid cancers compared Oxford Nanopore and Pacific Biosciences data to standard short-read whole-genome sequencing. The study found more than 96 percent recall and 91 percent precision for single nucleotide variants on both long-read platforms. Performance was lower for insertions and deletions, with 66 percent recall and 42 percent precision, especially in regions with few phased reads. The long-read platforms were 95 percent accurate for copy number calls and detected all recurrent structural variants with no false-positive findings. Importantly, long reads correctly identified intronic insertions near repetitive elements that were incorrectly identified as interchromosomal structural rearrangements by standard short-read whole-genome sequencing.
This study demonstrates that long-read sequencing can resolve some repeat-associated artifacts, but it also shows that indel calling in repeats remains challenging even with long reads. The practical implication is that long-read confirmation should be used for variants in repeats that are clinically or biologically important, instead of as a replacement for short-read variant calling.
For structural variant calling, Sniffles2 implements a repeat-aware clustering approach coupled with consensus sequence generation and coverage-adaptive filtering. The method is faster and more accurate than earlier structural variant callers across different coverages, sequencing technologies, and structural variant types. Sniffles2 can call structural variants at the family level and population level, producing fully genotyped VCF files. The method identified causative structural variants around the MECP2 gene, including highly complex alleles with three overlapping structural variants, and detected mosaic structural variants in bulk long-read data.
For transposable element detection specifically, the sTELLeR tool was developed for accurate, fast, and effective detection of transposable elements in long-read genomes. The tool shows higher precision and sensitivity for calling Alu elements than similar tools, is 5 to 48 times faster, and uses less than 2 percent of the CPU hours compared to competitive callers. sTELLeR is haplotype aware and outputs results in VCF format, enabling compatibility with other variant callers and downstream analysis.
Structural Variant Calling in Repetitive Regions
Structural variants, including deletions, duplications, inversions, and translocations, are particularly difficult to call in repetitive regions. The difficulty arises because structural variant breakpoints often occur within or near repeat elements, and the reads that span the breakpoints may align ambiguously.
Short-read structural variant callers use discordant read pairs and split reads to identify breakpoints. In repeat regions, discordant read pairs may be produced by alignment artifacts instead of true structural variants. Split reads may show soft-clipped bases that represent alignment to a different repeat copy instead of a true breakpoint.
Long-read structural variant calling is more accurate because a single read can span an entire structural variant breakpoint. The read sequence can be compared to the reference to identify the exact breakpoint location. However, long-read structural variant calling in repeats still requires repeat-aware methods because the breakpoint may fall within a repeat unit.
Sniffles2 addresses this challenge with repeat-aware clustering. The method clusters reads that support the same structural variant, using the repeat structure to inform the clustering. This approach reduces false positives from reads that align to different repeat copies.
The practical decision for structural variant calling is whether to use short reads, long reads, or both. For research applications, long-read structural variant calling is preferred when the budget allows. For clinical applications, the decision depends on the validated pipeline of the testing laboratory.
Satellite Repeats: The Hardest Case
Satellite repeats present the most difficult variant calling problem because they are arranged in long tandem arrays with very short repeat units. The human genome contains several types of satellite repeats, including alpha satellites in centromeres and beta satellites in pericentromeric regions.
The repeat unit for alpha satellite DNA is approximately 171 base pairs, and the arrays can extend for megabases. Short reads from these arrays align to many positions within the array, and the alignment scores are nearly identical for all positions. The result is that short-read variant calling in satellite arrays is essentially impossible.
Long reads can span multiple satellite repeat units, which provides some resolution. However, the error rate of long-read sequencing is higher than the divergence between satellite repeat units, so distinguishing true variants from sequencing errors is difficult.
The practical approach for satellite repeats is to exclude them from variant calling entirely. Most clinical and research pipelines mask centromeric and pericentromeric regions before analysis. Variants in these regions are not reported because they cannot be validated with current technology.
The NCBI Data Resources provide access to reference genome annotations that include satellite repeat locations, which can be used to create exclusion lists for variant calling.
Ribosomal DNA Repeats: A Special Case
Ribosomal DNA repeats are a special case of repetitive elements that have received attention because of their association with human health and disease. The human genome contains hundreds of copies of ribosomal DNA, and it was unknown whether these copies possess sequence variations that form different types of ribosomes.
A study developed an algorithm for long-read variant calling in ribosomal DNA, termed RGA, which revealed that variations in human ribosomal DNA loci are predominantly insertion-deletion variants. The study developed full-length rRNA sequencing and in situ sequencing methods that showed translating ribosomes possess variation in rRNA. Over 1,000 variants are lowly expressed, but tens of variants are abundant and form distinct rRNA subtypes with different structures near the indels.
The study found that rRNA subtypes show differential expression in endoderm-derived and ectoderm-derived tissues, and in cancer, low-abundance rRNA variants can become highly expressed. This finding demonstrates that variants in repetitive elements can have biological significance, which argues for developing better methods to call and interpret these variants.
The practical implication for variant calling workflows is that ribosomal DNA repeats should be treated as a special case. Standard short-read variant calling will not produce reliable results in these regions. Long-read variant calling with repeat-aware methods is required, and the interpretation of variants in these regions requires specialized knowledge.
Practical Workflow: Step-by-Step Implementation
The following workflow provides a concrete implementation of repeat-aware variant calling. The workflow assumes a germline variant calling scenario with short-read data, with options for long-read confirmation.
Step 1: Acquire Reference and Repeat Annotations
Download the reference genome and repeat annotations from a reliable source such as NCBI Data Resources. The repeat annotations should include the genomic coordinates of Alu elements, LINE elements, satellite repeats, and other repetitive sequences.
Step 2: Decide on Masking Strategy
Choose between hard masking, soft masking, and annotation-based filtering. For most applications, annotation-based filtering is recommended because it preserves the ability to examine variants in repeats if needed. Create a BED file of repeat regions from the annotation.
Step 3: Align Reads
Align reads to the reference genome using an aligner that reports mapping quality. For short reads, BWA-MEM is a common choice. For long reads, minimap2 is recommended. Record the mapping quality distribution for reads in repeat regions versus unique regions to understand the alignment quality in your data.
Step 4: Assess Alignment Quality in Repeats
Calculate the fraction of reads that align to repeat regions with mapping quality below 20. If this fraction is high, consider whether the reference or the aligner parameters need adjustment. Some aligners have parameters that control the penalty for multi-mapping reads.
Step 5: Call Variants
Run the variant caller with parameters that account for repeat regions. If using GATK HaplotypeCaller, consider using the repeat annotation to adjust the confidence thresholds in repeat regions. If using a somatic caller, ensure that the tumor-normal comparison accounts for the higher artifact rate in repeats.
Step 6: Filter Variants
Apply the following filters in order:
- Remove variants with low mapping quality support.
- Remove variants that overlap repeat annotations unless they are supported by unique evidence.
- Remove variants with strand bias or read position bias.
- Remove variants in satellite repeat regions unless they are confirmed by long reads.
Step 7: Confirm Important Variants with Long Reads
For variants in repeat regions that are clinically or biologically important, confirm with long-read sequencing. Use a repeat-aware structural variant caller such as Sniffles2 for structural variants and a transposable element caller such as sTELLeR for transposable element variants.
Step 8: Document the Workflow
Record the reference version, repeat annotation version, aligner version and parameters, variant caller version and parameters, and filtering thresholds. This documentation is essential for reproducibility and for interpreting variant calls in the context of the workflow.
Records and Measurements for Quality Control
The following measurements should be recorded for every variant calling run to enable quality assessment and troubleshooting.
Alignment Statistics
Record the total number of reads, the number and fraction of reads aligned, the number and fraction of reads multi-mapped, and the mapping quality distribution. Compare these statistics between repeat regions and unique regions.
Coverage Statistics
Record the mean and median coverage in repeat regions and unique regions. Low coverage in repeat regions may indicate that the aligner is failing to place reads there. High coverage may indicate multi-mapping inflation.
Variant Call Statistics
Record the number of variants called before and after filtering. Record the number of variants in repeat regions and the number removed by each filter. This information helps identify which filters are most effective for your data.
False Positive Rate Estimation
If possible, estimate the false positive rate in repeat regions by comparing variant calls to known variants in the sample or by validating a subset of variants with an orthogonal method.
Common Failure Patterns and Troubleshooting
The following failure patterns are commonly observed in variant calling in repetitive elements. Each pattern has specific diagnostic features and corrective actions.
Pattern 1: Excess Variants in Alu Regions
Diagnostic feature: The variant density in Alu regions is much higher than in unique regions, and the variants are predominantly single nucleotide variants with high read depth.
Cause: Multi-mapping reads from different Alu copies are aligning to one Alu location, creating false variant support.
Corrective action: Filter variants in Alu regions that are supported by multi-mapping reads. Increase the mapping quality threshold for Alu regions. Consider masking Alu regions if the research question does not require variants there.
Pattern 2: Missing Variants in LINE Regions
Diagnostic feature: The variant density in LINE regions is much lower than in unique regions, and known variants in LINE regions are not called.
Cause: The aligner is placing reads from LINE regions at incorrect locations, or the variant caller is filtering variants in LINE regions because of low mapping quality.
Corrective action: Examine the alignment of reads in LINE regions. If reads are being placed incorrectly, consider using a repeat-aware aligner or adjusting alignment parameters. If the variant caller is filtering variants, adjust the filtering thresholds for LINE regions.
Pattern 3: False Indels at Repeat Boundaries
Diagnostic feature: Insertions and deletions are called at the boundaries of repeat elements, and the indels are not present in orthogonal validation data.
Cause: Reads that span the repeat boundary may be soft-clipped or misaligned, creating false indel evidence.
Corrective action: Filter indels at repeat boundaries that are supported by soft-clipped reads. Examine the alignment structure of supporting reads to identify the artifact.
Pattern 4: Structural Variant False Positives Near Repeats
Diagnostic feature: Structural variants are called with breakpoints near repeat elements, and the structural variants are not confirmed by long-read sequencing.
Cause: Discordant read pairs and split reads from repeat regions are being interpreted as structural variant evidence.
Corrective action: Use a repeat-aware structural variant caller such as Sniffles2. Confirm structural variants near repeats with long-read sequencing before reporting.
Pattern 5: Low Concordance Between Replicates
Diagnostic feature: Variant calls from replicate samples show low concordance in repeat regions, while concordance in unique regions is high.
Cause: The stochastic placement of multi-mapping reads creates different false variant calls in each replicate.
Corrective action: Filter variants in repeat regions that are not consistently called across replicates. Consider using a consensus approach that requires variant support from multiple replicates.
Limitations of Current Methods
The methods described in this article have important limitations that should be understood before applying them.
First, no combination of masking, alignment, and filtering can completely eliminate false variant calls in repetitive regions. The best that can be achieved is a reduction in false positives to a level that is acceptable for the specific application.
Second, the sensitivity of variant calling in repeat regions is inherently lower than in unique regions. True variants in repeats may be missed even with the best methods. This limitation is particularly important for clinical applications where a missed variant could have diagnostic significance.
Third, the performance of repeat-aware methods depends on the quality of the repeat annotation. If the annotation is incomplete or inaccurate, variants in unannotated repeats will not be filtered.
Fourth, long-read sequencing improves but does not solve the problem. The error profiles of long-read platforms create their own artifacts, and the interpretation of variants in repeats requires specialized knowledge.
Fifth, the computational cost of repeat-aware methods is higher than standard methods. The additional steps of annotation, filtering, and validation require time and expertise.
Safety and Regulatory Context
Variant calling in repetitive elements has implications for clinical testing and reporting. Laboratories that report variants in repeat regions must ensure that their methods are validated and that the limitations of the methods are documented.
The NCBI Data Resources provide access to clinical variation databases that can be used to assess whether a variant in a repeat region has known clinical significance. The EMBL-EBI Training materials include guidance on clinical variant interpretation.
For laboratories developing clinical tests, the validation should include a repeat-aware assessment. The validation should demonstrate that the test can detect known variants in repeat regions and that the false positive rate in repeat regions is acceptable. The validation should also document the limitations of the test in repeat regions.
For research applications, the safety context is less formal but still important. Researchers should document the methods used to handle repeats and should be transparent about the limitations of their variant calls in repeat regions.
Professional Escalation Criteria
The following criteria indicate when a variant call in a repetitive region should be escalated to a specialist or confirmed with an orthogonal method.
Criterion 1: Clinically Significant Variant in a Repeat Region
If a variant in a repeat region has potential clinical significance, the variant should be confirmed with an orthogonal method before clinical action. The confirmation should use a method that does not rely on the same alignment and calling approach.
Criterion 2: Structural Variant Near a Repeat Element
If a structural variant has a breakpoint near a repeat element, the structural variant should be confirmed with long-read sequencing before reporting. Short-read structural variant calls near repeats have a high false positive rate.
Criterion 3: Variant with Unusual Read Support Patterns
If a variant is supported by reads with unusual alignment patterns, such as extensive soft clipping or inconsistent insert sizes, the variant should be examined manually before reporting.
Criterion 4: Variant in a Satellite Repeat Region
If a variant is called in a satellite repeat region, the variant should be treated with extreme caution. Satellite repeat regions are the most difficult to call, and most variants in these regions are artifacts.
Criterion 5: Discordant Results Between Replicates
If replicate samples show discordant variant calls in a repeat region, the discordance should be investigated before reporting. The investigation should determine whether the discordance is due to technical artifacts or biological variation.
Frequently Asked Questions
What is the main cause of false variant calls in repetitive elements?
The main cause is multi-mapping reads. When a short read originates from one copy of a repeat element, it can align equally well to many other copies of the same repeat element across the genome. The aligner may place the read at an incorrect location, and the variant caller then uses that read as evidence for a variant at the wrong position. This produces false positive variant calls that appear well-supported by read depth.
Should I mask repeats before variant calling?
The decision depends on your research question. If you need high-confidence variants in unique regions and do not need variants in repeats, hard masking is the safest choice. If you need to examine variants in or near repeats, use annotation-based filtering instead of masking. For clinical applications, follow the validated pipeline of your testing laboratory.
What mapping quality threshold should I use for reads in repeat regions?
A common threshold is 20, which corresponds to a 1 percent probability of incorrect placement. However, in repeat regions, this threshold may be too permissive. Some workflows use a threshold of 30 or higher for repeat regions, or exclude multi-mapping reads entirely. The optimal threshold depends on your aligner, your data, and your tolerance for false positives.
Can long-read sequencing solve the repeat problem?
Long-read sequencing substantially improves variant calling in repeats because long reads can span entire repeat units and their flanking unique sequence. However, long-read sequencing does not completely solve the problem. The error profiles of long-read platforms create their own artifacts, and indel calling in repeats remains challenging even with long reads. A study of long-read sequencing for myeloid cancer profiling found 66 percent recall and 42 percent precision for indels, with lower performance in regions with few phased reads.
How do I validate a variant in a repetitive region?
The most reliable validation is long-read sequencing. A long read that spans the variant and its flanking unique sequence can confirm that the variant is present at a specific genomic location. For structural variants, use a repeat-aware caller such as Sniffles2. For transposable element variants, use a tool such as sTELLeR. If long-read sequencing is not available, consider PCR amplification and Sanger sequencing across the variant site.
What are the most difficult repeat regions for variant calling?
Satellite repeats are the most difficult because they are arranged in long tandem arrays with very short repeat units. Short reads from these arrays align to many positions within the array, and the alignment scores are nearly identical. Most clinical and research pipelines exclude centromeric and pericentromeric satellite regions from variant calling entirely.
How should I report variants in repeat regions?
Report the variant with a clear description of the methods used and the limitations of those methods. Include the repeat annotation version, the filtering thresholds, and the validation method. If the variant was confirmed with long-read sequencing, state that clearly. If the variant was not confirmed, state that the variant is a candidate that requires confirmation.
What is the role of repeat annotations in variant filtering?
Repeat annotations provide the genomic coordinates of known repeat elements. After variant calling, you can compare your variant calls to the repeat annotation and flag or remove variants that fall within repeat regions. The quality of the repeat annotation directly affects the quality of the filtering. Use the most current repeat annotation available for your reference genome.
Related Bioinformatics Guides
- Detecting Structural Variants with Long-Read Sequencing: Methods and Considerations
- From Raw Reads to Variants: A Diagnostic Blueprint for Next-Generation Sequencing (NGS) Workflows
- Metagenome Co-Assembly: Strategies for Multi-Sample Data
- Single-Cell Multi-Omics Integration: Methods and Applications
- Metagenomics Pipeline: From Raw Reads to Taxonomic and Functional Profiles
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- Detection of mosaic and population-level structural variants with Sniffles2.. Nature biotechnology, 2024.
- Diversity of ribosomes at the level of rRNA variation associated with human health and disease.. Cell genomics, 2024.
- Human genome meeting 2016 : Houston, TX, USA. 28 February - 2 March 2016.. Human genomics, 2016.
- Evaluation of Long-Read Genome Sequencing for Genomic Profiling of Myeloid Cancers.. The Journal of molecular diagnostics : JMD, 2025.
- Detecting transposable elements in long-read genomes using sTELLeR.. Bioinformatics (Oxford, England), 2024.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.