Low-Complexity Regions in Variant Calling: How to Filter Homopolymer and Microsatellite Artifacts

By Dr. Zubair Khalid, DVM, MS, PhD ·

Low-Complexity Regions in Variant Calling: How to Filter Homopolymer and Microsatellite Artifacts

Key Takeaways

  • Low-complexity regions (LCRs), comprising only 1.2% of the human genome, are disproportionately responsible for a significant majority of variant calling errors (up to 91.3% for structural variants in long-read data), necessitating targeted filtering strategies.
  • Polymerase slippage during PCR amplification on repetitive templates (homopolymers and microsatellites) generates artificial insertions and deletions, leading to strand bias and skewed allele fractions that deviate from true germline genotypes.
  • Alignment ambiguity in short-read sequencing, where reads spanning long repeats can map to multiple positions with similar scores, is a primary mechanism for erroneous realignment and pseudo-multiallelic noise in LCRs.
  • Effective LCR artifact filtering involves a multi-pronged approach: reference-based masking using tools like TRF, read-based detection of signatures like strand bias and allele balance, and careful calibration of variant caller settings (e.g., GATK's --max-assembly-region-size).
  • A tiered decision framework, stratifying LCRs by repeat length (e.g., Tier 1: short repeats, Tier 3: long repeats), is crucial for balancing sensitivity and specificity, with Tier 3 regions often requiring orthogonal validation or masking to mitigate extremely high error rates.
  • Quality control metrics such as the Ti/Tv ratio, insertion/deletion profile, and variant allele fraction distribution are vital for assessing the impact of LCR filtering, with deviations indicating potential over- or under-filtering of true biological variants.

Low-complexity regions (LCRs) such as homopolymers and microsatellites produce a disproportionate share of false-positive variant calls in both germline and somatic sequencing workflows. These genomic stretches, where a single nucleotide or short motif repeats consecutively, challenge the assumptions of short-read alignment and variant calling algorithms. This article explains why LCRs generate artifacts, how to identify them in your data, and which filtering strategies and caller settings reduce false positives without sacrificing sensitivity in flanking regions. The guidance applies to researchers, laboratory professionals, and bioinformatics practitioners managing variant calling pipelines for research or clinical genomics applications.

The Scale of the Problem: Why 1.2 Percent of the Genome Creates Most Artifacts

Low-complexity regions occupy a small fraction of the human reference genome but account for a strikingly large share of variant calling errors. A 2025 analysis of the GRCh38 reference genome identified 35.4 Mb of low-complexity regions, covering only 1.2 percent of the genome. Despite this small footprint, these regions contain 69.1 percent of confident structural variants in the benchmark sample HG002. Across long-read structural variant callers, 77.3 to 91.3 percent of erroneous calls occur within LCRs, with error rates increasing as LCR length grows [<a href="#ref-1">1</a>].

The implications extend beyond structural variants. Sequencing error rates vary substantially by genomic context, and variant calling performance decreases markedly in low-complexity regions compared to unique or mappable sequence [<a href="#ref-2">2</a>]. A 2014 study using a haploid human genome identified erroneous realignment in low-complexity regions as one of the two major sources of false variant calls, alongside an incomplete reference genome relative to the sample. The raw genotype call error rate in that study was approximately 1 in 10 to 15 kb, but post-filtered calls improved to 1 in 100 to 200 kb without significant sensitivity loss [<a href="#ref-3">3</a>].

For practical variant calling, this means that a pipeline that performs well on unique sequence can generate hundreds or thousands of spurious calls in LCRs across a whole-genome sample. The artifact burden is not uniform: homopolymers of the same nucleotide repeated five or more times, dinucleotide microsatellites, and longer tandem repeats each present distinct challenges for alignment and variant calling.

Why Homopolymers and Microsatellites Break Standard Assumptions

Polymerase Slippage and Sequencing Artifacts

During library preparation and sequencing, DNA polymerases can slip on repetitive templates, producing PCR products with insertions or deletions of one or more repeat units. This slippage creates a mixed population of molecules with different repeat lengths in the final sequencing library. When these molecules are sequenced, the resulting reads contain apparent indels that do not reflect the true germline or somatic genotype of the sample.

The error profile differs by repeat type. Homopolymers, which consist of a single repeated nucleotide, are prone to single-base insertion and deletion errors. Microsatellites with di-, tri-, or tetranucleotide motifs generate errors in multiples of the motif length. Both error types are strand-biased in some sequencing platforms, which can be exploited for filtering.

Alignment Ambiguity

Short-read aligners map reads to the reference genome by finding the best scoring alignment. In a homopolymer of length 10, a read containing 9 or 11 copies of the same nucleotide can align in multiple ways, each with similar alignment scores. The aligner may place the indel at different positions within the repeat, creating apparent variant calls at multiple sites from a single true biological event.

This alignment ambiguity is the mechanism behind the erroneous realignment identified as a major artifact source in high-coverage whole-genome sequencing [<a href="#ref-3">3</a>]. The problem worsens with longer repeats because the number of equally plausible alignments increases.

Read Length and Coverage Limitations

Standard short-read sequencing produces reads of 100 to 150 bases. When a read spans a long homopolymer, the repetitive portion may consume a substantial fraction of the read, leaving little unique sequence for confident placement. Reads that fall entirely within a long repeat may map to multiple genomic locations, and aligners typically assign them to one location with low mapping quality or mark them as multimapping.

Coverage in LCRs can also be non-uniform. PCR amplification bias during library preparation can over- or under-represent repetitive regions relative to unique sequence. This coverage distortion affects allele frequency estimates and genotype likelihood calculations.

Identifying Low-Complexity Regions in Your Data

Reference-Based Masking Approaches

The most direct way to address LCR artifacts is to identify these regions in the reference genome before variant calling. RepeatMasker and similar tools annotate repetitive elements, but they do not specifically flag homopolymers and short microsatellites. Dedicated tools such as TRF (Tandem Repeats Finder) and the Genome Browser's simple repeat tracks provide this annotation.

A 2026 study evaluated a targeted spatial masking strategy that suppresses deterministic artifacts in short-read sequencing data while preserving clinically actionable variants outside LCRs. The protocol removed thousands of sequencing and alignment artifacts while maintaining the retained biological callset, with negligible disease-associated diagnostic variants detected in the excluded artifact fraction. The masking approach preserved physiological transition-to-transversion ratios and insertion/deletion profiles in retained calls, resolved pseudo-multiallelic noise, and distinguished excluded artifact calls by distorted mutational and variant allele frequency signatures [<a href="#ref-4">4</a>].

The practical implementation involves generating a BED file of LCR coordinates and applying it as an exclusion mask during variant filtering or as an input to the variant caller. The choice of masking stringency depends on your application. Clinical germline testing may warrant aggressive masking of all LCRs, while research applications that study repetitive regions may need a more permissive approach.

Read-Based Artifact Detection

Beyond reference-based masking, you can detect LCR artifacts directly from the sequencing data. Several signatures distinguish artifact calls from true variants:

  • Strand bias: True heterozygous variants should appear on both forward and reverse strand reads at similar frequencies. Polymerase slippage artifacts often show strong strand bias because the error mechanism is strand-specific.
  • Allele balance: True heterozygous germline variants typically show allele fractions near 0.5. LCR artifacts frequently show skewed allele fractions due to differential amplification of repeat alleles.
  • Read position bias: Artifacts may cluster at specific read positions, such as read ends where base quality declines.
  • Multiallelic calls: True biallelic variants rarely produce multiple alternate alleles at the same position. LCRs frequently generate pseudo-multiallelic noise where multiple indel alleles appear supported by different reads [<a href="#ref-4">4</a>].

Quality Score Calibration

Base quality scores in LCRs are often inaccurate. The sequencing instrument assigns quality scores based on signal intensity, but the signal from a homopolymer run degrades as the run length increases. Many variant callers incorporate base quality scores into genotype likelihood calculations, so inflated quality scores in LCRs can lead to false-positive calls with high confidence.

Some pipelines apply quality score recalibration using known variant sites, but this approach can fail in LCRs because the known sites themselves may be unreliable. Family-based error rate estimation offers an alternative: using Mendelian errors in family sequencing data to produce per-sample estimates of precision and recall for any set of variant calls, regardless of sequencing platform or calling methodology [<a href="#ref-2">2</a>].

Variant Caller Settings for Low-Complexity Regions

Germline Callers

GATK HaplotypeCaller and DeepVariant are the most widely used germline variant callers. Both have parameters that affect LCR performance, though the optimal settings depend on your specific data and application.

For GATK HaplotypeCaller, the --max-assembly-region-size parameter controls the maximum size of assembly regions. Larger values allow the caller to assemble longer haplotypes, which can improve indel calling in repetitive regions but increases runtime. The --dont-use-soft-clipped-bases flag prevents soft-clipped bases from contributing to assembly, which can reduce artifacts from reads that partially align to LCRs.

DeepVariant uses a deep learning model trained on labeled examples. Its performance in LCRs depends on the training data representation. The model generally handles homopolymers better than traditional callers because it learns context-specific error patterns, but it still produces false positives in long repeats.

Somatic Callers

Somatic variant callers such as Mutect2 and Strelka2 face additional challenges in LCRs because tumor samples often have lower coverage and higher noise than germline samples. Mutect2 applies a panel of normals to filter recurrent artifacts, which is particularly important for LCRs where the same artifact positions recur across samples.

The read orientation bias filter in Mutect2 is specifically designed to remove artifacts that appear predominantly on one read orientation. This filter is effective for LCR artifacts because polymerase slippage often produces orientation-biased errors.

Multi-Platform Integration

Combining data from multiple sequencing platforms can improve variant calling in difficult genomic regions. A 2023 study demonstrated that integrating Oxford Nanopore and Illumina data improved variant calling performance, with the improvement concentrated in difficult genomic regions such as large low-complexity regions and segmental duplication regions. The deep learning-based caller Clair3-MP achieved this improvement by leveraging the complementary error profiles of the two platforms [<a href="#ref-5">5</a>].

For laboratories with access to both short-read and long-read sequencing, this multi-platform approach offers a path to higher confidence calls in LCRs. The cost and complexity are substantial, but for clinical applications where LCR variants are clinically relevant, the investment may be justified.

Practical Filtering Workflow

Step 1: Generate an LCR Annotation Track

Create a BED file of low-complexity regions for your reference genome build. Use RepeatMasker output, Tandem Repeats Finder results, or a curated LCR database. For GRCh38, the 35.4 Mb of LCRs identified in the 2025 analysis provides a reference point, but you should generate annotations specific to your reference version [<a href="#ref-1">1</a>].

Step 2: Apply Hard Filters

Apply hard filters to remove obvious artifacts before downstream analysis. Common filters include:

  • Depth filters: Remove calls with abnormally high or low depth relative to the sample mean. LCRs often show extreme depth due to PCR bias.
  • Allele balance filters: For germline heterozygous calls, require allele fraction between 0.2 and 0.8. Adjust thresholds based on your sequencing platform and coverage.
  • Strand bias filters: Remove calls where the alternate allele appears almost exclusively on one strand. Use Fisher's exact test or a similar statistical test.
  • Mapping quality filters: Remove calls with low mapping quality, which indicates ambiguous read placement.

Step 3: Apply Caller-Specific Filters

Use the filtering recommendations from your variant caller. GATK provides the VariantFiltration tool with recommended thresholds for quality by depth, Fisher strand bias, mapping quality, and read position rank sum. These thresholds are calibrated for whole-genome and whole-exome data and may need adjustment for targeted panels.

Step 4: Apply LCR-Specific Masking

Apply your LCR annotation as a mask to remove calls in these regions. The 2026 masking study demonstrated that this approach removes thousands of artifacts while preserving the biological callset, with negligible loss of disease-associated variants [<a href="#ref-4">4</a>]. The key decision is whether to mask all LCRs or only those above a length threshold.

Step 5: Validate with Independent Methods

For calls in or near LCRs that pass filtering, validate with an independent method. Sanger sequencing is the traditional validation approach, but it can also fail in repetitive regions due to primer design challenges. Long-read sequencing provides a more reliable validation for LCR variants because the longer reads span the repeat and provide unambiguous alignment.

At a Glance: LCR Artifact Sources and Filtering Responses

Artifact SourceDetection SignatureRecommended FilterExpected Outcome
Polymerase slippage during PCR amplificationStrand-biased indels, allele fractions deviating from expected genotypeStrand bias filter, allele balance filterRemoves most slippage artifacts while retaining true heterozygous calls
Erroneous read realignment in repetitive sequenceMultiple apparent variants at adjacent positions, low mapping qualityMapping quality filter, LCR maskingEliminates spurious multi-allelic calls and positional ambiguity
Base quality score inflation in homopolymer runsHigh-confidence calls with distorted mutational signaturesQuality score recalibration, family-based error estimationReduces false-positive rate without compromising sensitivity
Coverage non-uniformity from PCR biasExtreme depth values, allele fraction distortionDepth filter, copy number aware filteringNormalizes allele frequency estimates and genotype likelihoods

Common Failure Patterns in LCR Filtering

Overfiltering and Loss of True Variants

The most common failure is applying filters so aggressively that true variants in LCRs are removed. This is particularly problematic for clinically relevant genes containing homopolymers or microsatellites. The 2026 masking study found that the excluded artifact fraction contained negligible disease-associated diagnostic variants, but this finding applies to the specific cohorts and variant types studied [<a href="#ref-4">4</a>]. For other genes or variant types, the balance may shift.

To avoid overfiltering, evaluate the Ti/Tv ratio and insertion/deletion profile of your retained calls. The masking study showed that LCR masking preserved physiological Ti/Tv and Ins/Del profiles in retained calls [<a href="#ref-4">4</a>]. If your filtered callset shows distorted ratios, you may be removing true variants.

Underfiltering and Residual Artifacts

The opposite failure is leaving artifacts in the callset. This is common when filters are calibrated on unique sequence and applied uniformly across the genome. The 2014 study found that raw genotype calls had an error rate of 1 in 10 to 15 kb, which would produce thousands of false calls in a whole-genome sample [<a href="#ref-3">3</a>]. Post-filtered calls improved to 1 in 100 to 200 kb, but this still leaves hundreds of false calls.

Residual artifacts are particularly problematic for somatic variant calling, where low allele fraction variants are clinically relevant. A true somatic variant at 5 percent allele fraction can be indistinguishable from a polymerase slippage artifact at the same allele fraction.

Inconsistent Filtering Across Samples

When filters are applied with fixed thresholds across samples with different coverage or quality, the results can be inconsistent. The family-based error estimation study demonstrated that sequencing error rates between samples in the same dataset can vary by over an order of magnitude [<a href="#ref-2">2</a>]. This variability means that a filter threshold that works for one sample may be too permissive or too stringent for another.

Masking Without Context

Applying a static LCR mask without considering the variant type or clinical context can remove important variants. Some LCRs contain pathogenic variants, and masking them entirely would produce false-negative results. The 2026 study found that the masking approach preserved clinically actionable variants outside LCRs, but the study did not evaluate variants within LCRs [<a href="#ref-4">4</a>].

Records and Measurements for Quality Control

Ti/Tv Ratio

The transition-to-transversion ratio is a standard quality metric for variant calls. Whole-genome data typically shows a Ti/Tv ratio around 2.0 to 2.1, while whole-exome data shows a higher ratio around 2.8 to 3.0. LCR artifacts often distort this ratio because they are enriched for specific error types. The masking study showed that LCR masking preserved physiological Ti/Tv ratios in retained calls [<a href="#ref-4">4</a>].

Insertion/Deletion Profile

The ratio of insertions to deletions and the size distribution of indels provide quality signals. True germline indels are predominantly 1 to 3 base pairs in length, with more deletions than insertions in most genomic contexts. LCR artifacts often show distorted indel profiles, with an excess of larger indels or an unusual insertion-to-deletion ratio.

Variant Allele Fraction Distribution

For germline samples, the allele fraction distribution should show a peak near 0.5 for heterozygous variants and near 1.0 for homozygous variants. LCR artifacts often produce calls with allele fractions between these peaks or at extreme values. The masking study distinguished excluded artifact calls by distorted variant allele frequency signatures [<a href="#ref-4">4</a>].

dbSNP Annotation Rate

The fraction of calls present in dbSNP provides a rough quality indicator, but the 2026 study found cohort-dependent behavior in this metric. In the TCGA-BRCA cohort, excluded calls showed higher dbSNP annotation than retained calls, while the AURORA cohort showed the opposite direction. This finding demonstrates the vulnerability of one-dimensional database annotation for variant authentication [<a href="#ref-4">4</a>].

Mendelian Error Rate

For family samples, the Mendelian error rate provides a direct measure of variant calling accuracy. The family-based error estimation method uses Mendelian errors to produce per-sample estimates of precision and recall for any set of variant calls, regardless of sequencing platform or calling methodology [<a href="#ref-2">2</a>]. This approach is particularly valuable for LCR filtering because it provides sample-specific error estimates.

Professional Escalation Criteria

When to Consult a Bioinformatics Specialist

Escalate to a bioinformatics specialist or computational genomics expert when:

  • Your filtered callset still contains an unexplained excess of variants in LCRs after applying standard filters
  • You observe sample-to-sample variability in error rates exceeding an order of magnitude, which may indicate pipeline or sample-specific issues [<a href="#ref-2">2</a>]
  • You need to validate clinically significant variants located within LCRs and require guidance on orthogonal validation methods
  • You are considering a multi-platform sequencing strategy and need to evaluate the cost-benefit tradeoff for your specific application [<a href="#ref-5">5</a>]

When to Reconsider the Pipeline

Reconsider your variant calling pipeline when:

  • The Ti/Tv ratio or indel profile of your filtered calls deviates substantially from expected values for your sequencing type
  • You observe pseudo-multiallelic noise at multiple loci, which indicates alignment or assembly issues in repetitive regions [<a href="#ref-4">4</a>]
  • Your validation rate for LCR variants is consistently below your validation rate for unique sequence variants
  • You are changing sequencing platforms or library preparation methods, which can alter the LCR error profile

When to Involve Clinical Expertise

For clinical applications, involve clinical genomics expertise when:

  • A variant in an LCR has potential diagnostic or therapeutic implications
  • You need to determine whether a variant in a homopolymer or microsatellite is a true pathogenic variant or an artifact
  • You are considering reporting a variant that would have been excluded by standard LCR masking

Limitations of Current Approaches

Short-Read Limitations

Short-read sequencing has inherent limitations in LCRs that no amount of filtering can fully overcome. The alignment ambiguity and polymerase slippage problems are intrinsic to the technology. Even with optimal filtering, some true variants in LCRs will be missed, and some artifacts will remain.

Reference Genome Completeness

The 2014 study identified incomplete reference genome relative to the sample as a major source of variant calling errors [<a href="#ref-3">3</a>]. If the reference genome has errors or gaps in LCRs, reads from the sample may not align correctly, producing false variant calls. This problem is particularly acute for highly polymorphic repeat regions where the reference represents only one allele.

Masking Tradeoffs

Targeted spatial masking removes LCR artifacts but also removes any true variants in those regions. The 2026 study found that the excluded artifact fraction contained negligible disease-associated diagnostic variants in the cohorts studied, but this finding may not generalize to all genes and variant types [<a href="#ref-4">4</a>]. For applications where LCR variants are clinically relevant, masking is not an acceptable solution.

Database Annotation Limitations

The 2026 study demonstrated that dbSNP annotation can behave differently in excluded artifact calls depending on the cohort, with TCGA-BRCA showing higher dbSNP annotation in excluded calls and AURORA showing the opposite [<a href="#ref-4">4</a>]. This finding cautions against relying on database annotation alone to distinguish true variants from artifacts.

Safety and Regulatory Context

Clinical Genomics Standards

For clinical genomics applications, variant calling pipelines must meet regulatory standards for accuracy and reproducibility. The presence of LCR artifacts in a reported callset can lead to false-positive results with clinical consequences. Laboratories should validate their LCR filtering approach using reference standards and demonstrate acceptable performance before clinical implementation.

Reproducibility Requirements

Reproducible variant calling requires documented pipelines and version-controlled software. The nf-core documentation provides standards for community pipeline development, including usage and configuration documentation [<a href="#ref-6">6</a>]. The Galaxy Training Network offers accessible workflow training and analysis tutorials that emphasize reproducibility [<a href="#ref-7">7</a>]. Bioconductor provides official package and workflow documentation for reproducible genomic analysis [<a href="#ref-8">8</a>].

Data Management

Variant calling generates large intermediate files that require careful management. The NCBI provides official descriptions of databases, search systems, sequence resources, and analysis services that support variant calling workflows [<a href="#ref-9">9</a>]. The EMBL-EBI Training program offers bioinformatics learning pathways and data-resource training [<a href="#ref-10">10</a>]. The Carpentries provides foundational computing, data, shell, Git, and programming training that supports reproducible analysis practices [<a href="#ref-11">11</a>].

A Decision Framework for LCR Filtering Based on Repeat Length and Variant Type

Standard filtering workflows apply uniform thresholds across all low-complexity regions, but this approach ignores the substantial differences in error behavior between short homopolymers, long homopolymers, and microsatellites of varying motif sizes. A length-stratified decision framework gives you a practical structure for choosing between masking, hard filtering, and targeted review for each variant call in an LCR. This section provides that framework, along with a record system for tracking LCR filtering decisions and a troubleshooting method for persistent artifact problems.

The Length-Error Relationship as Your Primary Decision Input

The 2025 structural variant analysis demonstrated that error rates in LCRs increase with LCR length [<a href="#ref-1">1</a>]. This relationship is not linear. Short repeats produce occasional errors that resemble true indels, while long repeats generate such dense artifact clusters that individual variant calls become meaningless without orthogonal validation. Your filtering strategy should scale with this error gradient.

For practical implementation, divide LCRs into three tiers based on repeat unit length and total tract length. Tier 1 includes homopolymers of 4 to 7 bases and microsatellites with 2 to 4 repeat units. These regions produce moderate artifact rates that standard hard filters can usually manage. Tier 2 includes homopolymers of 8 to 15 bases and microsatellites with 5 to 10 repeat units. These regions require caller-specific filters and often benefit from targeted masking. Tier 3 includes homopolymers longer than 15 bases and microsatellites with more than 10 repeat units. These regions produce error rates so high that masking or orthogonal validation becomes necessary for any variant call you intend to report.

The thresholds for these tiers are starting points, not universal constants. Your sequencing platform, library preparation method, and coverage level all shift the boundaries. The family-based error estimation method provides a way to calibrate these thresholds for your specific pipeline by producing per-sample estimates of precision and recall for any set of variant calls [<a href="#ref-2">2</a>]. If your error rate in Tier 1 regions exceeds your acceptable threshold, move the boundary to include more regions in Tier 2.

Decision Matrix for LCR Variant Calls

For each variant call that falls within an annotated LCR, apply the following decision matrix. This matrix assumes you have already generated an LCR annotation track and applied the standard hard filters described in the practical filtering workflow.

Step 1: Classify the repeat context. Determine whether the variant falls in a homopolymer or microsatellite, the repeat unit length, and the total tract length. Record this information for every LCR variant call.

Step 2: Assess the variant type. Single nucleotide variants in LCRs behave differently from indels. Homopolymer indels are the most common artifact type because polymerase slippage produces single-base insertions and deletions. Single nucleotide substitutions in homopolymers are less common but still occur due to base incorporation errors during sequencing. Microsatellite indels occur in multiples of the motif length, so a CA repeat produces deletions of 2, 4, or 6 bases instead of single-base indels.

Step 3: Apply the tier-specific action. For Tier 1 variants, apply standard hard filters and retain calls that pass. For Tier 2 variants, apply caller-specific filters and require additional evidence such as allele balance within the expected range and absence of strand bias. For Tier 3 variants, do not report calls without orthogonal validation. The 2026 masking study demonstrated that targeted spatial masking removes thousands of sequencing and alignment artifacts while preserving the retained biological callset [<a href="#ref-4">4</a>]. For Tier 3 regions, masking is the default action unless you have a specific reason to investigate a variant in that region.

Step 4: Document the decision. Record the repeat classification, the variant type, the filters applied, and the final decision for every LCR variant call. This record system becomes your audit trail for quality control and troubleshooting.

A Record System for LCR Filtering Decisions

A structured record system lets you track filtering decisions across samples and identify patterns in your artifact burden. Create a table with the following fields for each LCR variant call:

  • Sample identifier
  • Genomic position and reference allele
  • Repeat type (homopolymer or microsatellite)
  • Repeat unit and tract length
  • Tier classification
  • Variant type (SNV or indel)
  • Allele fraction
  • Strand bias statistic
  • Mapping quality
  • Filter applied
  • Final decision (pass, fail, or require validation)
  • Validation result if performed

This record system serves multiple purposes. It lets you calculate your LCR artifact rate per sample and compare it across samples in the same batch. It identifies which repeat types and lengths produce the most artifacts in your specific pipeline. It provides the data needed to calibrate your tier boundaries over time. And it creates the documentation needed for clinical reporting or regulatory review.

The family-based error estimation method provides a complementary approach to per-sample record keeping. By using Mendelian errors in family sequencing data, you can produce genome-wide error estimates for each sample without relying on a truth set [<a href="#ref-2">2</a>]. This method is particularly valuable for LCR filtering because it gives you sample-specific error rates that reflect the actual performance of your pipeline on your samples.

Troubleshooting Persistent LCR Artifact Problems

When your standard filtering workflow leaves residual artifacts in LCRs, work through the following troubleshooting sequence. This method isolates the source of the problem and identifies the specific adjustment needed.

Problem 1: Residual artifacts in Tier 1 regions. If you see an excess of false-positive calls in short homopolymers and microsatellites after applying standard hard filters, the issue is likely filter threshold calibration. Check your strand bias and allele balance thresholds against the actual distribution of true variants in your data. The 2014 study found that post-filtered calls reduced the error rate from 1 in 10 to 15 kb to 1 in 100 to 200 kb without significant sensitivity loss [<a href="#ref-3">3</a>]. If your post-filtered error rate is higher, your thresholds are too permissive. Tighten the thresholds and re-evaluate using the Ti/Tv ratio and indel profile of your retained calls.

Problem 2: Pseudo-multiallelic noise in Tier 2 regions. Pseudo-multiallelic noise occurs when alignment ambiguity places the same biological event at multiple positions within a repeat, creating the appearance of multiple alternate alleles at adjacent sites. The 2026 masking study found that LCR masking resolved pseudo-multiallelic noise in retained calls [<a href="#ref-4">4</a>]. If you observe this pattern, apply a filter that requires a single dominant alternate allele at each position. Alternatively, apply the LCR mask to these regions and re-run the variant caller.

Problem 3: Sample-to-sample variability in LCR error rates. The family-based error estimation study demonstrated that sequencing error rates between samples in the same dataset can vary by over an order of magnitude [<a href="#ref-2">2</a>]. If you see this variability, do not apply fixed filter thresholds across all samples. Instead, use per-sample error estimates to calibrate thresholds for each sample. This approach is more labor-intensive but produces more consistent results across heterogeneous sample sets.

Problem 4: Distorted mutational signatures in retained calls. The 2026 masking study distinguished excluded artifact calls by distorted mutational and variant allele frequency signatures [<a href="#ref-4">4</a>]. If your retained calls show distorted Ti/Tv ratios or unusual indel size distributions, you may be retaining artifacts that passed your filters. Compare the mutational signature of your retained LCR calls against your retained non-LCR calls. Substantial differences indicate residual artifact burden in LCRs.

Problem 5: Validation failures concentrated in specific repeat types. If your orthogonal validation shows that certain repeat types consistently fail, adjust your tier boundaries to move those repeat types into a more aggressive filtering tier. For example, if dinucleotide microsatellites with 6 repeat units show a 40 percent validation failure rate, move that repeat class from Tier 2 to Tier 3 and require orthogonal validation for all calls in those regions.

Integrating the Decision Framework with Multi-Platform Data

The decision framework described above assumes a single sequencing platform. If you have access to multiple platforms, the framework extends naturally. The 2023 Clair3-MP study demonstrated that integrating Oxford Nanopore and Illumina data improved variant calling performance, with the improvement concentrated in difficult genomic regions such as large low-complexity regions and segmental duplication regions [<a href="#ref-5">5</a>].

For multi-platform data, apply the tier classification to each platform separately, then compare calls across platforms. A variant that passes filters in both platforms independently has substantially higher confidence than a variant that passes in only one. For Tier 3 regions, concordance between platforms provides the orthogonal validation that would otherwise require a separate validation experiment.

The multi-platform approach does not eliminate the need for the decision framework. Even with multi-platform data, you still need to classify repeat context, assess variant type, and document decisions. The framework provides the structure for comparing calls across platforms and deciding when concordance is sufficient for reporting.

Calibrating the Framework with Reference Standards

Before implementing this decision framework in production, calibrate it using reference standards. The 2026 masking study used analytical reference standards including EQA and NA12878 to evaluate the masking protocol [<a href="#ref-4">4</a>]. You should follow a similar approach.

Run your pipeline on a reference standard with known true variants. Apply the tier classification and decision matrix to the resulting calls. Calculate sensitivity and precision separately for each tier. Adjust your tier boundaries and filter thresholds until you achieve acceptable performance in each tier. Document the calibration results as part of your pipeline validation.

This calibration step is essential for clinical applications. The presence of LCR artifacts in a reported callset can lead to false-positive results with clinical consequences. A calibrated decision framework provides the evidence needed to demonstrate that your LCR filtering approach meets accuracy standards.

Common Failure Patterns in the Decision Framework

Failure pattern 1: Treating all LCRs the same. Applying uniform filtering to all LCRs ignores the length-error relationship documented in the 2025 structural variant analysis [<a href="#ref-1">1</a>]. This pattern produces either excessive false positives in long repeats or excessive false negatives in short repeats. The tier classification directly addresses this failure.

Failure pattern 2: Masking without documentation. Applying an LCR mask without recording which regions were masked and why creates a documentation gap. If a clinically relevant variant falls in a masked region, you need to know why that region was masked and whether the masking decision was appropriate. The record system described above provides this documentation.

Failure pattern 3: Ignoring sample-specific variability. The family-based error estimation study demonstrated that sequencing error rates between samples in the same dataset can vary by over an order of magnitude [<a href="#ref-2">2</a>]. Applying fixed thresholds across samples with different error rates produces inconsistent results. Per-sample calibration addresses this failure.

Failure pattern 4: Relying on database annotation for LCR variants. The 2026 masking study found cohort-dependent behavior in dbSNP annotation, with TCGA-BRCA showing higher dbSNP annotation in excluded calls and AURORA showing the opposite direction [<a href="#ref-4">4</a>]. This finding demonstrates the vulnerability of one-dimensional database annotation for variant authentication. Do not use dbSNP presence or absence as a primary filter for LCR variants.

Professional Escalation Criteria for the Decision Framework

Escalate to a bioinformatics specialist when you encounter the following situations:

  • Your Tier 1 regions show artifact rates that do not respond to threshold adjustment, which may indicate an alignment or reference genome issue instead of a filtering issue
  • Your Tier 2 regions produce pseudo-multiallelic noise that persists after applying the LCR mask, which may indicate a caller-specific assembly problem
  • Your Tier 3 regions contain variants that appear clinically relevant and require orthogonal validation, which may require long-read sequencing or other specialized approaches
  • Your sample-to-sample variability in LCR error rates exceeds an order of magnitude, which may indicate sample quality issues or batch effects [<a href="#ref-2">2</a>]
  • You are considering a multi-platform sequencing strategy and need to evaluate whether the cost is justified for your specific application [<a href="#ref-5">5</a>]

Limitations of the Decision Framework

The tier boundaries and filter thresholds in this framework are starting points that require calibration for your specific pipeline. The 2025 analysis identified 35.4 Mb of LCRs in GRCh38, but your reference genome version and build may differ [<a href="#ref-1">1</a>]. The framework does not eliminate the need for orthogonal validation of clinically significant variants in LCRs. It reduces the number of variants requiring validation by removing obvious artifacts, but it cannot distinguish a true variant from an artifact when both produce similar read patterns.

The framework also assumes that your LCR annotation is accurate and complete. If your annotation misses LCRs, variants in those regions will not receive the tier-specific filtering. Validate your LCR annotation against known repeat databases and update it when new reference genome versions are released.

For clinical applications, the framework must be validated using reference standards and documented as part of your pipeline validation. The 2026 masking study provides a model for this validation, using analytical reference standards and clinical cohorts to evaluate the masking protocol [<a href="#ref-4">4</a>]. Your validation should follow a similar approach, with sensitivity and precision calculated separately for each tier.

Frequently Asked Questions

What is the difference between a homopolymer and a microsatellite in variant calling?

A homopolymer is a stretch of the same nucleotide repeated consecutively, such as AAAAA or CCCCC. A microsatellite is a short motif of two to six nucleotides repeated in tandem, such as CA repeats or GATA repeats. In variant calling, homopolymers generate single-base insertion and deletion artifacts from polymerase slippage, while microsatellites generate indels in multiples of the motif length. The filtering approaches differ because the expected error patterns differ.

How do I generate a low-complexity region mask for my reference genome?

Use Tandem Repeats Finder or RepeatMasker to annotate repetitive regions in your reference genome, then filter the output to retain short tandem repeats and homopolymers. Alternatively, use the simple repeat track from the UCSC Genome Browser or a curated LCR database. The 2025 analysis identified 35.4 Mb of LCRs in GRCh38, which can serve as a reference point for validating your annotation [<a href="#ref-1">1</a>]. Generate the mask as a BED file and apply it during variant filtering or as an input to the variant caller.

What are the most effective filters for removing homopolymer artifacts?

The most effective filters combine strand bias, allele balance, and mapping quality. Polymerase slippage artifacts often show strong strand bias because the error mechanism is strand-specific. True heterozygous variants should show allele fractions near 0.5, while artifacts frequently show skewed fractions. Low mapping quality indicates ambiguous read placement in repetitive sequence. The 2014 study found that post-filtered calls reduced the error rate from 1 in 10 to 15 kb to 1 in 100 to 200 kb without significant sensitivity loss [<a href="#ref-3">3</a>].

Should I mask all low-complexity regions or only those above a length threshold?

The decision depends on your application. The 2026 masking study found that the excluded artifact fraction contained negligible disease-associated diagnostic variants in the cohorts studied, but this finding may not generalize [<a href="#ref-4">4</a>]. For clinical applications where LCR variants are relevant, use a length threshold that removes the most problematic regions while retaining shorter repeats that may contain true variants. For research applications, masking all LCRs is simpler and more conservative.

How does multi-platform sequencing improve variant calling in low-complexity regions?

Multi-platform sequencing combines the strengths of different technologies. The 2023 Clair3-MP study demonstrated that integrating Oxford Nanopore and Illumina data improved variant calling performance, with the improvement concentrated in difficult genomic regions such as large low-complexity regions and segmental duplication regions [<a href="#ref-5">5</a>]. The complementary error profiles of the two platforms allow the caller to distinguish true variants from platform-specific artifacts.

How can I estimate sequencing error rates for my specific pipeline?

Family-based error estimation uses Mendelian errors in family sequencing data to produce per-sample estimates of precision and recall for any set of variant calls, regardless of sequencing platform or calling methodology. This method demonstrated that sequencing error rates between samples in the same dataset can vary by over an order of magnitude and that variant calling performance decreases substantially in low-complexity regions [<a href="#ref-2">2</a>]. This approach provides sample-specific error estimates that are more reliable than generic quality metrics.

What is pseudo-multiallelic noise and how does it affect variant calling?

Pseudo-multiallelic noise occurs when multiple alternate alleles appear supported at the same position, creating the appearance of a multiallelic variant when no true multiallelic variant exists. This artifact is common in LCRs where alignment ambiguity places indels at different positions within a repeat. The 2026 masking study found that LCR masking resolved pseudo-multiallelic noise in retained calls [<a href="#ref-4">4</a>]. Filtering approaches that require a single dominant alternate allele can remove this artifact.

When should I use long-read sequencing for variant calling in low-complexity regions?

Use long-read sequencing when LCR variants are clinically relevant and short-read results are ambiguous. Long reads span the full repeat and provide unambiguous alignment, eliminating the alignment ambiguity that plagues short-read data. The 2025 structural variant analysis found that even long-read callers have error rates of 77.3 to 91.3 percent in LCRs, so long-read data alone does not fully solve the problem [<a href="#ref-1">1</a>]. For the highest confidence, combine long-read and short-read data using a multi-platform approach [<a href="#ref-5">5</a>].

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

[1] [Challenges in structural variant calling in low-complexity regions.](https://pubmed.ncbi.nlm.nih.gov/41384802). GigaScience, 2025. [2] [Estimating sequencing error rates using families.](https://pubmed.ncbi.nlm.nih.gov/33892748). BioData mining, 2021. [3] [Toward better understanding of artifacts in variant calling from high-coverage samples.](https://pubmed.ncbi.nlm.nih.gov/24974202). Bioinformatics (Oxford, England), 2014. [4] [Targeted Genomic Region Masking Supports Accurate Variant Calling While Suppressing Low-Complexity Sequencing Artifacts.](https://pubmed.ncbi.nlm.nih.gov/42510812). Genes, 2026. [5] [Boosting variant-calling performance with multi-platform sequencing data using Clair3-MP.](https://pubmed.ncbi.nlm.nih.gov/37537536). BMC bioinformatics, 2023. [6] [nf-core Documentation](https://nf-co.re/docs). nf-core. [7] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project. [8] [Bioconductor](https://bioconductor.org/). Bioconductor Project. [9] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information. [10] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute. [11] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.