Context-Specific Errors in Variant Calling: How to Filter Artifacts in Repetitive Regions and Homopolymers

By Dr. Zubair Khalid, DVM, MS, PhD ·

Context-Specific Errors in Variant Calling: How to Filter Artifacts in Repetitive Regions and Homopolymers

Key Takeaways

  • Standard variant quality filters are insufficient for repetitive regions and homopolymers because they assume random error distribution, which is violated by systematic sequencing chemistry limitations and alignment ambiguities.
  • Polymerase slippage in homopolymer runs leads to bidirectional insertion/deletion errors, necessitating context-aware quality recalibration using methods like molecular barcodes or local reassembly to distinguish true variants from artifacts.
  • Repetitive and variant-dense regions challenge read alignment, leading to spurious variant candidates; deep learning-based callers and hybrid short-long read approaches offer improved accuracy by mitigating misalignment issues.
  • Platform-specific error profiles, such as those observed in GC homopolymers for Element AVITI sequencing or historical ONT data, require tailored filtering strategies rather than generic approaches.
  • Practical filtering strategies involve characterizing repetitive content, applying molecular barcodes where feasible, utilizing local reassembly for ambiguous regions, and employing repeat-masking for annotation or exclusion.
  • Measuring filtering performance requires a truth set (e.g., Genome in a Bottle benchmarks) to calculate precision and recall, and quality control metrics should be stratified by genomic context (unique regions, homopolymers, microsatellites) to identify context-specific issues.

Variant calling in repetitive regions and homopolymers produces a distinct class of errors that standard quality filters often miss. These errors arise from sequencing chemistry limitations, alignment ambiguity, and polymerase slippage during library preparation. This article defines the error modes, explains why generic filtering fails, and provides concrete filtering strategies using local reassembly, repeat-masking, and context-aware quality metrics. The guidance applies to germline and somatic variant calling workflows using short-read, long-read, or hybrid sequencing data.

The Problem: Why Repetitive Regions Generate False Variants

Repetitive regions and homopolymers create systematic sequencing artifacts that mimic genuine biological variation. A homopolymer is a sequence of identical nucleotides repeated consecutively, such as a run of ten adenine bases. Microsatellites are short tandem repeats of one to six base pair motifs. Both sequence contexts challenge the base calling algorithms, the alignment step, and the variant caller itself.

Sequencing platforms produce characteristic error profiles in these regions. Short-read platforms using sequencing-by-synthesis chemistry experience phasing errors when the polymerase slips during incorporation of identical nucleotides. This slippage causes the instrument to lose synchronization across the cluster, producing insertion and deletion errors that appear as false indels in homopolymer runs. The error rate increases with homopolymer length, and the direction of the error, insertion versus deletion, depends on the specific platform chemistry.

Long-read platforms have different but equally problematic error profiles. Oxford Nanopore Technologies sequencing historically showed elevated error rates in homopolymers, although newer basecalling models have reduced this problem substantially. A 2024 benchmarking study in eLife demonstrated that deep learning-based variant callers applied to ONT super-high accuracy data mitigated traditional ONT errors in homopolymers, while also overcoming Illumina errors that arise from difficulties in aligning reads in repetitive and variant-dense genomic regions 7. The same study showed that 10x depth of ONT super-accuracy data achieved precision and recall comparable to full-depth Illumina sequencing 7.

The practical consequence is that a researcher examining a variant call in a homopolymer cannot distinguish a true variant from a sequencing artifact using the variant caller's quality score alone. The quality scores are calibrated on the assumption that errors are randomly distributed across the genome, an assumption that fails in low-complexity regions.

At a Glance: Error Modes and Filtering Strategies

Error ContextPrimary Error ModeRecommended Filtering StrategyEvidence Basis
Homopolymer runs (3+ identical bases)Polymerase slippage causes insertion/deletion errors during sequencingApply context-aware quality score recalibration using molecular barcodes or local reassemblyMolecular barcoded read correction removed all false positives in homopolymer regions while retaining true positives 8
Microsatellite and short tandem repeatsLength estimation noise from bidirectional insertion and deletion errorsUse specialized genotyping tools with error bias estimation, such as discretized Gaussian mixture modelsGenoTan genotyped 94.9% of microsatellite loci accurately from simulated 40x data 10
High-GC regions and GC homopolymersPlatform-specific coverage bias and elevated error ratesStratify variant calls by genomic context and apply platform-specific filtersAvidity sequencing showed superior coverage in high-GC regions but inferior performance in GC homopolymers 9
Repetitive and variant-dense regionsRead misalignment causes spurious variant candidatesUse deep learning-based callers or hybrid short-long read approachesDeep learning callers outperformed traditional methods and exceeded Illumina accuracy on repetitive regions 7

Core Principles of Context-Specific Variant Filtering

Generic Quality Filters Are Insufficient

Standard variant filtering relies on metrics such as read depth, mapping quality, base quality, and variant quality score. These metrics work well for variants in unique genomic regions where errors are approximately random. In repetitive regions, the assumptions break down.

Read depth in repetitive regions is inflated by multi-mapping reads. A read that aligns equally well to multiple copies of a repeat will be counted multiple times or assigned to one location arbitrarily. This inflation creates false confidence in variant calls supported by reads that actually originate from different genomic locations.

Mapping quality scores reflect the confidence that a read is placed correctly. In repetitive regions, mapping quality is appropriately low for multi-mapping reads, but variant callers handle these reads inconsistently. Some callers discard low mapping quality reads entirely, reducing sensitivity for true variants in repeats. Other callers include them, increasing false positives.

Base quality scores from the sequencer do not account for context-specific error modes. A base call in a homopolymer may receive a high quality score from the instrument while being systematically wrong due to phasing errors. The quality score reflects the fluorescence signal intensity, not the probability that the base is correct in context.

The Bidirectional Nature of Indel Errors

Insertion and deletion errors in homopolymer runs are bidirectional. A sequencing platform may systematically insert an extra base in a homopolymer run, or it may delete a base, depending on the chemistry and the specific sequence context. This bidirectionality complicates filtering because a simple rule such as "filter all indels in homopolymers" removes true variants along with artifacts.

A 2014 study in Bioinformatics introduced a homopolymer decomposition method that estimates error bias toward insertion or deletion in homopolymer sequence runs 10. The authors developed GenoTan, a program using a discretized Gaussian mixture model combined with a rules-based approach, to distinguish length variants from noise in microsatellite loci 10. The key insight is that error bias can be estimated from the data itself, allowing the filter to account for the specific platform and sequence context.

For a researcher, this means that filtering decisions should be informed by the observed error direction in control samples or known variant sites. If the platform systematically inserts bases in adenine homopolymers, then a deletion call in an adenine homopolymer deserves more scrutiny than an insertion call.

Platform-Specific Error Profiles

Different sequencing platforms produce different error patterns in repetitive regions. A 2026 study in NAR Genomics and Bioinformatics compared Element Biosciences AVITI avidity sequencing with Illumina NovaSeq X Plus on human tumor cell lines 9. The study found that AVITI showed lower duplication rates and higher base qualities, which contributed to improved mapping confidence and fewer spurious variant candidates 9. However, stratifying by genomic context revealed that AVITI genome coverage and variant calls were superior in high-GC regions while being inferior in GC homopolymers 9. AVITI also showed increased error rates on read 2 related to short fragments and sensitivity to G-quadruplex motifs 9.

The practical implication is that a filtering strategy developed for Illumina data may not transfer directly to data from another platform. Researchers should characterize the error profile of their specific platform and sequencing chemistry using control samples or publicly available benchmark data before applying context-specific filters.

Practical Workflow for Filtering Variants in Repetitive Regions

Step 1: Characterize the Repetitive Content of Your Target Regions

Before filtering, identify which genomic regions are prone to context-specific errors. This characterization should happen at the experimental design stage, not after variant calling.

For targeted sequencing panels, examine the bait design or amplicon design for homopolymer runs and repetitive motifs. For whole-genome sequencing, generate a repeat mask using standard tools that identify low-complexity regions, simple repeats, and interspersed repeats. The repeat mask should be applied to the reference genome before variant calling so that the variant caller can annotate variants falling in these regions.

The NCBI data resources provide sequence resources and analysis services that can support this characterization. Researchers can use NCBI databases to access reference genomes, repeat annotations, and variation databases for comparison with their own calls.

Step 2: Apply Molecular Barcodes Where Possible

Molecular barcodes, also known as unique molecular identifiers or UMIs, provide a powerful mechanism for correcting sequencing errors in repetitive regions. A 2019 study in the Journal of Computational Biology demonstrated a workflow that applies base score correction to molecular barcoded sequencing reads 8. The workflow reduced the false-positive rate of variant calls in homopolymer and repetitive regions where the sequencer commonly encounters phasing errors 8.

The approach works by grouping reads that share the same molecular barcode, indicating they originated from the same original DNA molecule. Errors introduced during PCR amplification or sequencing will appear in only some reads within the group, while true variants will appear in all reads. Base score correction adjusts the quality scores of bases that disagree across the barcode group, effectively suppressing PCR and sequencing errors.

The study applied this workflow to a custom QIAseq targeted DNA panel of 220 genes and found that base correction removed all false positives identified without the correction method while retaining the true positive call 8. This approach is particularly valuable for clinical validation studies where false positives in homopolymer regions can lead to incorrect variant interpretations.

Step 3: Use Local Reassembly for Ambiguous Regions

Local reassembly is a strategy that reconstructs the sequence of a region from raw reads without relying on alignment to the reference. This approach is valuable in repetitive regions where alignment is ambiguous.

The principle is straightforward. Instead of asking where each read aligns to the reference, the reassembly algorithm builds a graph of all possible sequences supported by the reads in the region. The graph is then resolved to produce the most likely sequence. Variants are called by comparing the reassembled sequence to the reference.

Local reassembly is particularly effective for indels in homopolymers because it can determine the exact length of the homopolymer run from the read data. Alignment-based approaches struggle with this task because a read with a deletion in a homopolymer can align equally well with the deletion placed at any position within the run.

Several variant callers incorporate local reassembly as part of their algorithm. The choice of caller depends on the sequencing platform and the specific application. Deep learning-based callers such as Clair3 and DeepVariant have shown superior performance on repetitive regions in benchmarking studies 7.

Step 4: Apply Repeat-Masking to Filter or Annotate

Repeat-masking serves two purposes in variant filtering. First, it can be used to exclude variants in repetitive regions from downstream analysis entirely. This approach is appropriate when the research question does not require variants in these regions and the risk of false positives is unacceptable.

Second, repeat-masking can be used to annotate variants that fall in repetitive regions, allowing for separate filtering thresholds. A variant in a homopolymer with a quality score of 500 might be treated with suspicion, while the same quality score in a unique region would be accepted.

The choice between exclusion and annotation depends on the application. For clinical reporting, exclusion of variants in repetitive regions may be appropriate if the laboratory cannot validate them with orthogonal methods. For research applications, annotation with separate thresholds preserves sensitivity while flagging potentially problematic calls.

Step 5: Apply Context-Aware Quality Metrics

Context-aware quality metrics adjust variant quality scores based on the local sequence context. These metrics can be computed during variant calling or applied as post-processing filters.

One approach is to recalibrate base qualities using known error modes. The molecular barcode correction method described above is one example. Another approach is to use machine learning models trained on known true and false variant calls to learn context-specific error patterns.

The Galaxy Training Network provides accessible workflow training and analysis tutorials that cover variant calling and filtering. Researchers can use these resources to build reproducible workflows that incorporate context-aware filtering steps.

Options and Tradeoffs in Filtering Strategies

Hard Filtering Versus Probabilistic Filtering

Hard filtering applies a binary threshold: a variant passes or fails based on specific criteria. For example, a filter might remove all variants with a quality score below 500 or all variants in homopolymers longer than five bases. Hard filtering is simple to implement and explain, but it can remove true variants that happen to fall below the threshold.

Probabilistic filtering assigns a probability that each variant is real, allowing for more nuanced decisions. Variant callers that use machine learning models, such as DeepVariant, produce probabilities that incorporate context-specific information. These probabilities can be used to rank variants for manual review instead of applying a hard threshold.

The tradeoff is between interpretability and accuracy. Hard filters are easier to document and validate for clinical use. Probabilistic filters are more accurate but harder to explain and may require extensive validation before regulatory acceptance.

Exclusion Versus Orthogonal Validation

Excluding variants in repetitive regions is the safest approach for avoiding false positives, but it sacrifices sensitivity. True disease-causing variants in homopolymers and repeats will be missed.

Orthogonal validation uses a different technology to confirm or refute a variant call. Sanger sequencing is the traditional orthogonal method, but it also struggles with homopolymers. Long-read sequencing can validate variants in repetitive regions because long reads span the entire repeat and can determine the exact length.

The choice between exclusion and orthogonal validation depends on the clinical or research context. For a research study where false positives are acceptable but false negatives are not, orthogonal validation of repetitive region variants is appropriate. For a clinical test where false positives could lead to unnecessary medical intervention, exclusion may be safer.

Single Platform Versus Hybrid Approaches

Hybrid approaches combine short-read and long-read sequencing data from the same sample. A 2025 study in Frontiers in Bioinformatics benchmarked the DNAscope Hybrid pipeline, which combines Illumina short reads with PacBio HiFi long reads 11. The study found that the hybrid approach significantly improved SNP and indel calling accuracy, particularly in complex genomic regions 11. At lower long-read depths of 5x to 10x, the hybrid approach outperformed stand-alone short-read or long-read approaches 11.

The tradeoff is cost and complexity. Hybrid sequencing requires two sequencing runs and a more complex bioinformatics pipeline. The nf-core documentation provides community pipeline standards and usage guidance that can support reproducible hybrid workflows.

For researchers who cannot afford hybrid sequencing, the alternative is to use long-read sequencing only for targeted validation of variants in repetitive regions identified by short-read sequencing.

Records and Measurements for Filtering Decisions

What to Record

Documenting filtering decisions is essential for reproducibility and for understanding why specific variants were or were not reported. The following records should be maintained for each variant calling run:

The sequencing platform, chemistry version, and basecalling model. Platform-specific error profiles change with chemistry updates, so this information is essential for interpreting filtering decisions.

The variant caller version and parameters. Variant caller algorithms change between versions, and parameter choices affect the balance between sensitivity and specificity.

The repeat mask version and parameters. Repeat annotations differ between genome builds and annotation sources.

The filtering thresholds applied and the rationale for each threshold. This documentation supports future optimization and regulatory review.

The number of variants removed by each filtering step. This measurement helps identify whether a filter is too aggressive or too lenient.

How to Measure Filtering Performance

Measuring filtering performance requires a truth set of known variants. The Genome in a Bottle consortium provides benchmark variant calls for several human cell lines, including HG002, HG003, and HG004. These benchmarks can be used to calculate precision and recall for a variant calling and filtering pipeline.

Precision is the proportion of called variants that are true positives. Recall, also called sensitivity, is the proportion of true variants that are called. A good filtering strategy maximizes both, but there is always a tradeoff. Increasing filtering stringency improves precision at the cost of recall.

For targeted panels, the truth set can be generated from orthogonal sequencing of the same samples. Long-read sequencing of the panel regions provides a high-quality truth set for evaluating short-read variant calling performance.

Quality Control Metrics for Repetitive Regions

Standard sequencing quality control metrics should be stratified by genomic context. Mean coverage, percentage of bases above a depth threshold, and uniformity of coverage should be reported separately for unique regions, homopolymers, microsatellites, and high-GC regions.

The 2026 AVITI study demonstrated that platform performance varies by genomic context 9. Reporting aggregate quality metrics can hide significant problems in specific contexts. A run with excellent overall coverage might have poor coverage in high-GC regions, leading to false negative variant calls in those regions.

Common Failure Patterns in Variant Filtering

Overly Aggressive Filtering of Homopolymer Variants

A common failure is applying a blanket filter that removes all variants in homopolymers or repetitive regions. This approach eliminates false positives but also removes true variants. Pathogenic variants in homopolymers are well documented in clinical genetics, and removing them entirely would cause false negative results.

The solution is to use context-aware filtering that distinguishes between error-prone and reliable calls. Molecular barcode correction, local reassembly, and deep learning-based callers can all improve the reliability of variant calls in these regions.

Ignoring Platform-Specific Error Profiles

Another failure is applying filtering thresholds developed for one platform to data from another platform. The error profiles of Illumina, Element AVITI, Oxford Nanopore, and PacBio differ substantially in repetitive regions. A threshold that works well for Illumina data may be too aggressive or too lenient for AVITI data.

The solution is to characterize the error profile of each platform and chemistry using control samples or benchmark data before applying filters. The eLife 2024 study provides a model for this approach, benchmarking multiple variant callers across multiple ONT basecalling models and read types 7.

Using Read Depth as a Proxy for Confidence in Repetitive Regions

Read depth is inflated in repetitive regions by multi-mapping reads. A variant call supported by 100 reads in a repetitive region may be less reliable than a variant call supported by 20 reads in a unique region, because many of the 100 reads may not actually originate from the variant site.

The solution is to use mapping quality filtered depth, which counts only reads with high mapping quality, and to interpret depth in the context of the local repeat structure.

Failing to Validate Filtering Decisions

A filtering strategy that has not been validated against a truth set is a guess. The performance of a filter can only be measured by applying it to data with known true and false variants.

The solution is to validate filtering decisions using benchmark samples before applying them to research or clinical samples. The Genome in a Bottle benchmarks provide a standard truth set for human samples.

Limitations of Current Filtering Approaches

Residual Errors After Filtering

Even the best filtering strategies do not eliminate all errors in repetitive regions. The eLife 2024 study found that deep learning-based callers mitigated ONT errors in homopolymers, but did not eliminate them entirely 7. Residual errors require manual review or orthogonal validation.

Reference Genome Limitations

The reference genome itself contains errors and ambiguities in repetitive regions. These reference errors can cause spurious variant calls when the sample sequence matches the true sequence but differs from the erroneous reference. Reference genome improvements, such as the telomere-to-telomere consortium assemblies, reduce but do not eliminate this problem.

Computational Cost

Some filtering approaches, particularly local reassembly and deep learning-based calling, require substantial computational resources. The GenoTan study noted that other programs required 5 to 30 times more computational time than GenoTan for microsatellite genotyping 10. Researchers with limited computational resources may need to balance accuracy against cost.

Transferability Across Species and Sample Types

Filtering strategies developed for human samples may not transfer directly to other species. Repeat content, GC content, and genome complexity vary across species. The eLife 2024 study evaluated variant calling across 14 bacterial species, highlighting the importance of species-specific validation 7.

Safety and Regulatory Context for Clinical Applications

Validation Requirements for Clinical Variant Calling

Clinical laboratories that report variants in repetitive regions must validate their filtering strategies according to regulatory standards. The validation should demonstrate that the filtering strategy achieves acceptable sensitivity and specificity for the intended clinical use.

The 2019 molecular barcode study was motivated by a clinical validation study that observed high false-positive rates in specific regions of a targeted panel 8. The workflow reduced false positives while retaining true positives, demonstrating the importance of validation in clinical contexts.

Reporting Variants in Repetitive Regions

When reporting variants in repetitive regions, laboratories should indicate the limitations of the testing method. A variant call in a homopolymer that has not been confirmed by an orthogonal method should be reported with appropriate caveats.

The NCBI databases provide resources that support clinical variant interpretation. ClinVar and other NCBI resources can be used to check whether a variant in a repetitive region has been previously reported and classified.

Professional Escalation Criteria

Laboratory professionals should escalate variant calls in repetitive regions for additional review when specific criteria are met. These criteria include variants in homopolymers longer than a validated threshold, variants in regions with known platform-specific error profiles, and variants that would change clinical management if confirmed.

The escalation process should include review by a molecular pathologist or clinical geneticist, orthogonal validation if available, and consultation with the ordering clinician when the variant has potential clinical significance.

Training and Reproducibility Considerations

Building Reproducible Filtering Workflows

Reproducibility in variant filtering requires version control of both the analysis code and the reference data. The Carpentries lessons provide foundational training in shell, Git, and programming that supports reproducible bioinformatics analysis. Researchers who document their filtering steps as version-controlled scripts can revisit and audit their decisions more effectively than those who rely on manual point-and-click operations.

Workflow managers provide another layer of reproducibility. The nf-core documentation describes community standards for pipeline development, usage, and configuration. These standards help ensure that a filtering workflow runs identically across different computing environments and that parameter changes are tracked systematically.

Training Resources for Context-Specific Filtering

The EMBL-EBI training portal offers learning pathways that cover sequence analysis, variant calling, and data resource usage. These materials help researchers understand the assumptions built into different filtering approaches and how to adapt them to specific data types.

The Bioconductor project provides packages for genomic analysis that include functions for repeat annotation, variant filtering, and quality control. The official package documentation describes installation procedures and reproducible analysis workflows that can be adapted for context-specific filtering.

The Galaxy Training Network offers accessible tutorials for variant calling and filtering that run on public infrastructure. These tutorials are useful for researchers who want to test different filtering strategies without installing software locally.

Decision Framework for Filtering Strategy Selection

Factors That Determine Filtering Approach

The choice of filtering strategy depends on several factors that should be evaluated before the analysis begins. The sequencing platform determines the error profile and the available correction methods. The research question determines whether sensitivity or specificity is more important. The available computational resources determine whether deep learning-based callers or local reassembly are feasible. The regulatory context determines the validation burden.

The following table summarizes the decision framework:

Analysis ContextRecommended ApproachKey Consideration
Germline research study with short-read dataApply repeat-masking annotation, use deep learning-based caller, validate with truth setBalance sensitivity and specificity for the specific research question
Somatic variant detection with molecular barcodesApply base score correction using barcode informationPreserve low-frequency variant detection while suppressing PCR errors
Clinical validation of targeted panelUse molecular barcode correction plus orthogonal validation of repetitive region callsDocument validation evidence for regulatory review
Bacterial genomics with nanopore dataUse deep learning-based callers with high-accuracy basecalling modelsEvaluate performance across species with different repeat content
Complex genomic regions with hybrid dataCombine short and long reads using integrated pipelinesAssess cost against accuracy improvement for the specific regions of interest

Matching Filtering Stringency to Research Objectives

Filtering stringency should match the downstream interpretation requirements. A discovery study that will validate candidate variants with orthogonal methods can tolerate higher false-positive rates and use lenient filters. A clinical report that directly influences patient management requires stringent filters and documented validation.

The same raw variant call set can be filtered at different stringencies for different purposes. A lenient filter set can be used for research discovery, while a stringent filter set is applied for clinical reporting. This dual filtering approach preserves the ability to revisit variants that fail stringent filters when new evidence becomes available.

A Practical Decision Framework for Context-Specific Variant Filtering

Establishing a Tiered Filtering System Based on Repeat Class and Error Risk

A single filtering threshold applied uniformly across all repetitive regions will either retain excessive artifacts or discard genuine variants. The solution is a tiered filtering system that assigns each variant to a risk class based on the specific repeat context and the observed error behavior of the sequencing platform. This framework converts the general principles of context-aware filtering into concrete, repeatable decisions that can be documented and validated.

The first tier contains variants in unique sequence with no adjacent repetitive elements. These variants can be filtered using standard quality metrics without additional context-specific adjustments. The second tier contains variants in or near homopolymers of three to ten bases. These require context-aware quality assessment because polymerase slippage errors scale with homopolymer length. The third tier contains variants in homopolymers longer than ten bases, microsatellites, and other low-complexity regions. These require the most stringent filtering, often including orthogonal validation or exclusion from clinical reporting.

The tier assignment should be automated using a repeat annotation file generated from the reference genome. The NCBI data resources provide reference genome sequences and repeat annotations that can support this classification. The repeat annotation should be generated before variant calling and applied consistently across all samples in a study.

Assigning Confidence Scores by Repeat Context

Instead of applying a binary pass or fail filter, assign each variant a context-adjusted confidence score that combines the variant caller quality score with a penalty based on the repeat context. The penalty should reflect the measured error rate for the specific platform and repeat type.

For example, if validation data shows that the platform produces a 5 percent false-positive rate for indels in five-base homopolymers and a 20 percent false-positive rate for indels in ten-base homopolymers, the confidence score should reflect this gradient. A variant in a ten-base homopolymer requires substantially higher raw quality scores to achieve the same confidence as a variant in a unique region.

The confidence score calculation should be documented in the analysis protocol and applied consistently. The Bioconductor project provides packages for genomic analysis that can implement these scoring functions within reproducible workflows. The official package documentation describes installation procedures and analysis workflows that can be adapted for context-specific scoring.

Building a Platform-Specific Error Profile From Control Data

A tiered filtering system is only as reliable as the error profile that informs it. The error profile should be measured from control data instead of assumed from published reports, because error rates vary with chemistry version, basecalling model, and laboratory-specific factors.

The control data should include a sample with known variants in repetitive regions. The Genome in a Bottle benchmarks provide this resource for human samples. For non-human species, a control sample can be sequenced with both short-read and long-read platforms, with the long-read data serving as the truth set for evaluating short-read error rates.

The error profile should record the false-positive rate for each repeat class and variant type. This profile should be updated whenever the sequencing chemistry or basecalling model changes. The 2026 AVITI study demonstrated that platform performance varies by genomic context, with AVITI showing superior coverage in high-GC regions but inferior performance in GC homopolymers 9. A laboratory using AVITI data would need a different error profile than a laboratory using Illumina data.

Implementing the Tiered Filtering Workflow

The tiered filtering workflow proceeds through five distinct stages, each with specific inputs, outputs, and decision points.

Stage one is repeat annotation. Generate a repeat mask for the reference genome and assign each genomic position to a repeat class. This annotation should be stored as a standard BED or VCF annotation file and version-controlled.

Stage two is variant calling with full output. Run the variant caller without aggressive filtering so that all candidate variants, including those in repetitive regions, are retained in the output. The Galaxy Training Network provides accessible workflow training that covers variant calling and filtering steps that can be adapted for this purpose.

Stage three is tier assignment. Annotate each variant with its repeat class using the repeat mask. Assign each variant to tier one, two, or three based on the repeat class and the distance to the nearest repetitive element.

Stage four is context-specific filtering. Apply standard quality filters to tier one variants. Apply additional context-aware filters to tier two variants, such as requiring higher mapping quality or using molecular barcode correction if available. Apply the most stringent filters to tier three variants, which may include exclusion from downstream analysis or mandatory orthogonal validation.

Stage five is documentation and review. Record the number of variants in each tier, the number removed by each filter, and the rationale for each filtering decision. This documentation supports reproducibility and regulatory review.

Recording Filtering Decisions and Outcomes

A standardized record system captures the information needed to audit filtering decisions and improve the workflow over time. The record should include the sample identifier, the sequencing platform and chemistry version, the variant caller version and parameters, the repeat annotation version, and the filtering thresholds applied.

For each variant removed by filtering, the record should note the tier assignment, the specific filter that removed the variant, and the quality metrics that triggered the filter. This information allows the laboratory to identify filters that are removing excessive numbers of variants and to adjust thresholds accordingly.

The record should also include the number of variants retained in each tier and the proportion of variants in repetitive regions that passed filtering. A sudden change in this proportion may indicate a chemistry change, a reference annotation error, or a problem with the filtering workflow.

The Carpentries lessons provide foundational training in shell, Git, and programming that supports the version control and documentation practices needed for this record system. Researchers who document their filtering steps as version-controlled scripts can revisit and audit their decisions more effectively than those who rely on manual operations.

Troubleshooting Unexpected Variant Patterns in Repetitive Regions

When the filtering workflow produces unexpected results, a systematic troubleshooting approach identifies whether the problem lies in the data, the annotation, or the filtering parameters.

The first troubleshooting step is to examine the raw read data in the affected region. Visualize the aligned reads using a genome browser or alignment viewer. Look for evidence of polymerase slippage, such as reads with variable homopolymer lengths, and for alignment artifacts, such as reads with soft-clipped ends at repeat boundaries.

The second step is to verify the repeat annotation. The repeat mask may be outdated or may not match the reference genome build used for alignment. Re-generate the repeat annotation using the same reference genome build and compare the results.

The third step is to test the filtering parameters on a control sample with known variants. If the control sample produces unexpected filtering outcomes, the parameters require adjustment. If the control sample performs correctly but the test sample does not, the problem likely lies in the test sample data quality.

The fourth step is to compare results across platforms or chemistries. The 2024 eLife study benchmarked variant calling across multiple ONT basecalling models and read types, demonstrating that performance varies substantially with these factors 7. A variant that passes filtering in one chemistry version may fail in another, and the error profile should be updated accordingly.

Common Failure Patterns in Tiered Filtering Implementation

The most common failure in implementing a tiered filtering system is applying the same thresholds to all tiers. This mistake negates the purpose of tiering and reproduces the problems of uniform filtering. Each tier requires thresholds calibrated to its specific error rate.

A second failure is using a repeat annotation that does not match the reference genome build. If the alignment uses genome build GRCh38 but the repeat annotation was generated for GRCh37, the tier assignments will be incorrect and filtering decisions will be based on wrong information.

A third failure is neglecting to update the error profile when the sequencing chemistry changes. The 2026 AVITI study showed that different platforms and chemistries produce different error patterns in repetitive regions 9. A laboratory that updates its sequencing platform without updating its error profile will apply inappropriate filtering thresholds.

A fourth failure is treating the tiered filtering system as a one-time setup instead of a continuously validated process. The error profile should be re-measured periodically, and the filtering thresholds should be adjusted based on accumulated validation data.

Measuring the Effectiveness of the Tiered Filtering System

The effectiveness of the tiered filtering system should be measured using precision and recall calculated against a truth set. Precision measures the proportion of retained variants that are true positives. Recall measures the proportion of true variants that are retained.

The measurement should be stratified by tier. A filtering system that achieves high precision and recall in tier one but poor performance in tier three has a specific problem that requires targeted adjustment. The stratification identifies whether the problem lies in the error profile, the filtering thresholds, or the repeat annotation.

The nf-core documentation describes community standards for pipeline development and usage that support reproducible benchmarking. Researchers can use these standards to build validation workflows that measure filtering performance consistently across samples and time points.

Professional Escalation Criteria for Filtering Decisions

Certain filtering decisions warrant escalation to a supervisor, bioinformatics specialist, or clinical geneticist. These escalation criteria should be defined in advance and documented in the laboratory protocol.

Escalate when a variant in a repetitive region has potential clinical significance but fails the tier three filtering thresholds. The decision to exclude or report this variant requires professional judgment that cannot be automated.

Escalate when the filtering workflow removes an unexpectedly high proportion of variants in a specific repeat class. This pattern may indicate a systematic problem with the sequencing chemistry, the repeat annotation, or the filtering parameters.

Escalate when a variant in a repetitive region is confirmed by orthogonal validation but was removed by the filtering workflow. This outcome indicates that the filtering thresholds are too aggressive for that repeat class and require adjustment.

Escalate when the error profile derived from control data does not match the observed error patterns in test samples. This discrepancy may indicate sample-specific issues such as degradation, contamination, or unusual repeat content.

Integrating the Tiered Filtering System With Existing Workflows

The tiered filtering system can be integrated with existing variant calling workflows without requiring a complete pipeline redesign. The repeat annotation and tier assignment can be added as post-processing steps after variant calling. The context-specific filtering can replace or supplement existing hard filters.

The integration should be tested on control samples before application to research or clinical samples. The EMBL-EBI training portal offers learning pathways that cover sequence analysis and variant calling, providing context for how the tiered filtering system fits within broader analysis workflows.

The tiered filtering system should be documented in the laboratory protocol with clear instructions for implementation, validation, and troubleshooting. This documentation ensures that the system is applied consistently across operators and time points, and that the rationale for filtering decisions can be reconstructed during audits or regulatory review.

Frequently Asked Questions

Why do homopolymers cause so many false variant calls?

Homopolymers cause false variant calls because the sequencing polymerase slips during incorporation of identical nucleotides. This slippage produces insertion and deletion errors that appear as false indels. The error rate increases with homopolymer length, and the direction of the error depends on the platform chemistry. Standard quality scores do not account for this context-specific error mode.

Can I simply filter out all variants in repetitive regions?

Filtering out all variants in repetitive regions eliminates false positives but also removes true variants. Pathogenic variants in homopolymers and microsatellites are well documented. A better approach is to use context-aware filtering that distinguishes reliable from unreliable calls, or to validate repetitive region variants with an orthogonal method.

What is the difference between hard filtering and probabilistic filtering?

Hard filtering applies a binary threshold, removing all variants that fail specific criteria. Probabilistic filtering assigns a probability that each variant is real, allowing for more nuanced decisions. Hard filters are easier to document and validate for clinical use. Probabilistic filters are more accurate but harder to explain and may require extensive validation.

How do molecular barcodes help with variant calling in repetitive regions?

Molecular barcodes, also known as unique molecular identifiers, group reads that originated from the same original DNA molecule. Errors introduced during PCR amplification or sequencing appear in only some reads within the group, while true variants appear in all reads. Base score correction using barcode information suppresses errors in homopolymer and repetitive regions 8.

Should I use long-read sequencing for repetitive regions?

Long-read sequencing can resolve repetitive regions because long reads span the entire repeat and can determine the exact length. However, long-read platforms have their own error profiles, particularly for homopolymers. Deep learning-based variant callers applied to high-accuracy long-read data can mitigate these errors 7. Hybrid approaches combining short and long reads provide the best accuracy but at higher cost 11.

How do I validate my filtering strategy?

Validate your filtering strategy using a truth set of known variants. The Genome in a Bottle consortium provides benchmark variant calls for several human cell lines. For targeted panels, generate a truth set using orthogonal sequencing of the same samples. Calculate precision and recall for your pipeline and adjust filtering thresholds to achieve acceptable performance.

What records should I keep for variant filtering decisions?

Keep records of the sequencing platform and chemistry version, the variant caller version and parameters, the repeat mask version and parameters, the filtering thresholds applied, and the number of variants removed by each filtering step. This documentation supports reproducibility and regulatory review.

When should I escalate a variant call for additional review?

Escalate variant calls in repetitive regions for additional review when they fall in homopolymers longer than a validated threshold, when they are in regions with known platform-specific error profiles, or when they would change clinical management if confirmed. The escalation process should include review by a qualified professional and orthogonal validation if available.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.