Tandem Repeats and Variant Calling: How to Avoid Expansion and Contraction Artifacts in STR Regions
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- STR artifacts, manifesting as false insertions and deletions, arise from PCR slippage during library preparation and ambiguous read alignment in repetitive genomic regions, leading to misinterpretation of true repeat counts.
- Repeat-aware aligners and STR-specific genotyping tools like ExpansionHunter are critical for accurate variant calling by modeling stutter patterns and providing allele-specific repeat count estimates with confidence intervals.
- Long-read sequencing offers direct resolution of repeat arrays, mitigating alignment ambiguity, but still faces challenges with small indel calling in homopolymers and tandem repeats due to platform-specific error profiles.
- A practical workflow involves assessing data inputs (platform, read length, coverage), selecting repeat-aware alignment strategies, running STR-specific genotyping, applying stringent quality filters, and meticulously documenting and reporting results.
- Quality control for STR calls includes analyzing stutter patterns, ensuring adequate coverage and uniform read depth across repeat arrays, and assessing allele balance for heterozygous loci.
- Benchmarking against pedigree-based truth sets and performing platform concordance checks are essential for validating STR genotyping pipelines and optimizing parameters for improved accuracy.
Short tandem repeats (STRs) are genomic regions composed of repeated copies of a 1 to 6 base pair motif, distributed throughout the human genome with over one million variable STR loci known. Some of these loci regulate gene expression and influence complex traits such as height, while variants in at least 60 STR loci cause genetic disorders including Huntington disease and fragile X syndrome. The core problem for variant calling is that STR regions produce alignment errors and PCR slippage artifacts that manifest as false insertion and deletion calls. This article explains how repeat-aware aligners and STR-specific callers such as ExpansionHunter improve accuracy, and it provides a practical workflow for researchers and laboratory professionals who need to distinguish true STR expansions and contractions from technical artifacts.
The intended reader is a biology student, researcher, or laboratory professional who already runs standard short-read variant calling pipelines and has observed suspicious indel calls in repetitive regions. The scope covers data inputs, workflow choices, quality controls, reproducibility, interpretation limits, and reporting decisions. The guidance applies to germline variant calling, somatic variant calling, and the integration of STR genotyping into existing bioinformatics pipelines.
The Biological and Technical Basis of STR Artifacts
Why STR Regions Produce False Indel Calls
STR loci are composed of repeated copies of a short motif, and the number of repeats can vary between individuals and between the two alleles within an individual. When sequencing libraries are prepared, PCR amplification across these repetitive regions can cause the polymerase to slip, producing amplicons with different repeat counts than the original template. This PCR slippage generates a stutter pattern where the observed read lengths cluster around the true repeat count but include shorter and longer products.
During alignment, short reads that originate from STR regions often map ambiguously because the same sequence motif appears multiple times. A read that spans part of a repeat array can align to multiple positions with equal or nearly equal scores. The aligner must decide where to place the read, and this decision directly affects whether an insertion or deletion is called at that locus. If the aligner places the read one repeat unit shorter than the true position, the downstream variant caller may interpret the discrepancy as a deletion. If the aligner places the read one repeat unit longer, the caller may interpret it as an insertion.
The result is that standard variant calling pipelines produce a high rate of false indel calls in STR regions. These false calls appear as expansions or contractions that do not reflect the true genotype. The problem is compounded in homopolymers, which are runs of a single nucleotide, because the alignment ambiguity is maximal and the PCR slippage rate is high.
The Diagnostic and Research Consequences
The clinical impact of STR artifacts is substantial. In rare disease diagnostics, short-read sequencing workflows are used to detect small variants, copy number variants, and short tandem repeats. A study of 1000 cases with 1271 known clinically relevant variants found that overall detection was 95 percent, but detection rates differed by variant category with small variants detected in 96 percent of cases and large variants in 93 percent. The study authors noted that variant calling format files were queried per variant to determine workflow-specific true positive rates, and they used a threshold of 98 percent true positive rate to decide whether genome sequencing could replace existing workflows. This means that even in a well-characterized diagnostic cohort, a meaningful fraction of variants, including STRs, are missed or miscalled by standard approaches.
Long-read sequencing offers a partial solution. A study of 98 samples from 41 families with suspected rare monogenic diseases used nanopore sequencing at approximately 36-fold average coverage with a 32 kb read N50. The long-read pipeline detected additional rare variants including structural variants and tandem repeats that were not covered by short-read sequencing. Diagnostic variants were established in 11 probands with diverse underlying genetic causes. However, the same study noted that small indel calling remains difficult within homopolymers and tandem repeats even with long-read data, though concordance to Illumina indel calls was good elsewhere.
The practical conclusion is that no single sequencing platform eliminates STR artifacts entirely. The choice of aligner, variant caller, and quality control strategy determines whether STR variants are reported accurately or whether the pipeline produces a misleading set of expansion and contraction calls.
Core Principles for STR-Aware Variant Calling
Repeat-Aware Alignment
The first line of defense against STR artifacts is the aligner. Standard short-read aligners use a seed-and-extend strategy that works well for unique genomic regions but struggles with repetitive sequences. When a read originates from an STR, the seed may match multiple locations, and the aligner may choose the wrong location or introduce gaps to force a unique alignment.
Repeat-aware aligners incorporate information about known repetitive regions into the alignment process. These aligners can soft-clip reads that extend beyond the repeat array, or they can use a repeat mask to prevent reads from being forced into incorrect alignments. The choice of aligner parameters also matters. Increasing the penalty for gap opening and extension can reduce spurious indels in repetitive regions, but it can also reduce sensitivity for true indels elsewhere in the genome.
For long-read data, the alignment problem is different. Long reads can span entire STR arrays, providing direct evidence of the repeat count. However, long-read aligners must still handle the error profile of the sequencing platform. Nanopore sequencing has a higher base-level error rate than short-read sequencing, and this error rate is concentrated in homopolymers and tandem repeats. The aligner must distinguish sequencing errors from true repeat count differences.
STR-Specific Genotyping
Repeat-aware alignment reduces but does not eliminate STR artifacts. The second line of defense is to use an STR-specific genotyping tool such as ExpansionHunter. These tools are designed to estimate the repeat count at known STR loci by analyzing the distribution of read lengths and the alignment patterns across the repeat array.
STR-specific callers use a different model than standard variant callers. Instead of calling a single consensus genotype, they estimate the repeat count for each allele and provide a confidence interval. This is important because STRs are often heterozygous with different repeat counts on the two alleles, and the stutter pattern from PCR can obscure the true allele sizes.
The output of an STR-specific caller is a repeat count estimate for each allele, along with quality metrics that indicate the confidence of the call. These calls can be integrated into the variant calling workflow as an additional annotation layer. The STR genotype can be used to validate or override indel calls from the standard variant caller in STR regions.
The Role of Long-Read Sequencing
Long-read sequencing provides a fundamentally different view of STR regions. Because long reads can span the entire repeat array, the repeat count can be read directly from the sequence. This eliminates the need to infer repeat counts from the alignment of short reads across the repeat boundary.
A study that developed a scalable nanopore sequencing protocol for human genomes found that the approach could detect single nucleotide polymorphisms with F1-scores comparable to Illumina short-read sequencing. The same study noted that small indel calling remains difficult within homopolymers and tandem repeats, but achieves good concordance to Illumina indel calls elsewhere. This suggests that long-read sequencing improves STR genotyping but does not completely solve the problem.
The decision to use long-read sequencing depends on the research or clinical question. For population-scale studies where cost is a primary constraint, short-read sequencing with STR-specific callers may be the practical choice. For rare disease diagnostics where a missed STR expansion could mean a missed diagnosis, long-read sequencing may be justified despite the higher cost.
At a Glance: STR Artifact Control Strategies
| Strategy | Primary Benefit | Key Limitation | Best Use Case |
|---|---|---|---|
| Repeat-aware short-read alignment | Reduces spurious indels from ambiguous read placement | Does not resolve PCR stutter patterns | Standard germline and somatic variant calling pipelines |
| STR-specific genotyping with ExpansionHunter | Provides allele-specific repeat count estimates with confidence metrics | Requires known STR loci and may miss novel expansions | Clinical and research workflows targeting known STR disorders |
| Long-read sequencing with repeat-aware assembly | Directly reads repeat arrays and detects large expansions | Higher cost and lower throughput than short-read sequencing | Rare disease diagnostics and validation of ambiguous short-read calls |
Practical Workflow for STR-Aware Variant Calling
Step 1: Assess Your Data Inputs
Before modifying your variant calling pipeline, document the characteristics of your sequencing data. Record the sequencing platform, read length, coverage depth, and library preparation method. PCR-free library preparation reduces but does not eliminate stutter artifacts. PCR-amplified libraries have higher stutter rates, and the stutter pattern varies by motif type and repeat length.
For short-read data, record the read length and insert size. Paired-end reads with longer inserts provide more information across the repeat array, but the alignment ambiguity remains. For long-read data, record the read N50 and the base-level error rate. These metrics determine the confidence you can place in repeat count estimates.
Step 2: Choose Your Alignment Strategy
Select an aligner that handles repetitive regions appropriately. For short-read data, consider using an aligner with repeat-aware features or adjust the parameters of your existing aligner to reduce spurious indels. Test the aligner on a control sample with known STR genotypes to establish a baseline error rate.
For long-read data, choose an aligner that is designed for the error profile of your sequencing platform. The alignment parameters should be tuned to minimize false indels in homopolymers and tandem repeats while maintaining sensitivity for true variants.
Step 3: Run STR-Specific Genotyping
Run an STR-specific caller such as ExpansionHunter on your aligned reads. This tool requires a catalog of known STR loci and their motif sequences. The output provides repeat count estimates for each allele at each locus, along with quality metrics.
Compare the STR genotype calls to the indel calls from your standard variant caller. In STR regions, the STR-specific caller should be considered the authoritative source. If the standard variant caller reports an indel that conflicts with the STR genotype, investigate the discrepancy before accepting the indel call.
Step 4: Apply Quality Filters
Apply quality filters to the STR genotype calls. The confidence interval for each allele should be narrow enough to support the biological interpretation. If the confidence interval spans more than a few repeat units, the call should be flagged as low confidence.
For clinical applications, confirm low-confidence or clinically significant STR calls with an orthogonal method. This could be long-read sequencing, PCR-based fragment analysis, or a second STR-specific caller with a different algorithmic approach.
Step 5: Document and Report
Document the STR genotyping results in your variant calling output. Include the repeat count for each allele, the confidence interval, and the quality metrics. For clinical reporting, include the STR genotype in the variant annotation and note the method used to generate the call.
When reporting STR variants, distinguish between confirmed genotypes and calls that require validation. A variant that is reported with high confidence from a single STR-specific caller may be sufficient for research purposes, but clinical reporting should follow the laboratory's validation protocols.
Options and Tradeoffs in STR Genotyping
Short-Read Sequencing with STR-Specific Callers
Short-read sequencing remains the most cost-effective approach for population-scale studies. The combination of short-read data with an STR-specific caller can genotype known STR loci with reasonable accuracy. The main limitation is that short reads cannot span very long repeat arrays, so large expansions may be missed or underestimated.
The tradeoff is between cost and completeness. For studies that require genotyping of known STR loci across many samples, short-read sequencing with an STR-specific caller is the practical choice. For studies that need to discover novel STR expansions or genotype very long repeat arrays, long-read sequencing is necessary.
Long-Read Sequencing for STR Resolution
Long-read sequencing provides direct evidence of repeat counts and can detect expansions that are invisible to short-read sequencing. The tradeoff is cost and throughput. Long-read sequencing is more expensive per sample and requires more computational resources for base calling and alignment.
The decision to use long-read sequencing should be based on the expected variant spectrum. If the study or clinical case involves known STR disorders with large expansions, long-read sequencing may be justified. If the goal is to genotype common STR polymorphisms across many samples, short-read sequencing with an STR-specific caller is more efficient.
Hybrid Approaches
A hybrid approach uses short-read sequencing for genome-wide variant calling and long-read sequencing for targeted validation of STR regions. This approach balances cost and completeness. The short-read data provides genome-wide coverage, and the long-read data resolves ambiguous STR calls.
The hybrid approach is particularly useful for clinical diagnostics where a missed STR expansion could have serious consequences. The long-read validation can be applied to a panel of known STR loci or to regions where the short-read data produces low-confidence calls.
Observations and Measurements for STR Quality Control
Stutter Pattern Analysis
The stutter pattern in PCR-amplified libraries provides a quality metric for STR genotyping. In a clean sample, the read length distribution at an STR locus should show a dominant peak at the true repeat count with smaller peaks at adjacent repeat counts. A broad stutter pattern with many peaks indicates that the PCR amplification introduced significant slippage, and the STR genotype call should be treated with caution.
Record the stutter ratio for each STR locus. The stutter ratio is the proportion of reads that differ from the dominant allele by one or more repeat units. A high stutter ratio reduces the confidence in the genotype call and may require orthogonal validation.
Coverage and Read Depth
The coverage depth at an STR locus affects the confidence of the genotype call. Low coverage produces noisy repeat count estimates, while very high coverage can amplify the effects of PCR duplicates. Record the coverage at each STR locus and flag loci with coverage below your established threshold.
For short-read data, the coverage across the repeat array should be uniform. If the coverage drops sharply at the repeat boundary, the alignment may be biased against reads that span the repeat, and the genotype call may be unreliable.
Allele Balance
For heterozygous STR loci, the two alleles should have similar read support. If one allele has substantially more reads than the other, the genotype call may be biased by PCR amplification or alignment artifacts. Record the allele balance for each heterozygous STR call and flag calls with extreme allele imbalance.
The allele balance metric is particularly important for somatic variant calling, where the variant allele fraction can vary across samples. In somatic tissues, a true STR expansion may be present in only a fraction of cells, and the allele balance reflects the mosaic fraction.
Records and Measurements for STR Variant Calling
Sample-Level Records
Maintain a record for each sample that includes the sequencing platform, library preparation method, coverage depth, and read length. These metadata are essential for interpreting STR genotype calls and for comparing results across samples.
Record the STR genotype calls for each sample in a structured format that includes the locus identifier, motif sequence, repeat count for each allele, confidence interval, and quality metrics. This record should be linked to the variant calling output and the raw sequencing data.
Locus-Level Records
For each STR locus in your catalog, record the motif sequence, the reference repeat count, and the range of repeat counts observed in your samples. This information is used to interpret new genotype calls and to identify outliers that may represent true expansions or contractions.
Record the stutter pattern and coverage metrics for each locus across your sample set. These metrics provide a baseline for quality control and help identify loci that are difficult to genotype with your current pipeline.
Pipeline-Level Records
Document the version of each tool in your variant calling pipeline, including the aligner, the STR-specific caller, and the standard variant caller. Record the parameters used for each tool and any modifications made during the analysis.
Reproducibility requires that the pipeline can be rerun with the same inputs and parameters to produce the same results. Use a workflow management system to track the pipeline execution and to ensure that all steps are recorded. Community standards for reproducible workflows are documented by the nf-core project, which provides guidelines for pipeline structure, configuration, and version tracking. The Galaxy Training Network offers accessible tutorials on reproducible analysis workflows that can be adapted for STR genotyping pipelines. For foundational skills in data management and version control, The Carpentries lessons cover shell, Git, and reproducible computing practices that apply to variant calling projects.
Common Failure Patterns in STR Variant Calling
False Expansions from PCR Slippage
The most common failure pattern is a false expansion call caused by PCR slippage. The polymerase adds or removes repeat units during amplification, and the resulting reads support a repeat count that differs from the true genotype. This pattern is more common in PCR-amplified libraries and in loci with long repeat arrays.
The solution is to use an STR-specific caller that models the stutter pattern and provides a confidence interval for the repeat count. If the confidence interval is wide, the call should be flagged as low confidence and validated with an orthogonal method.
False Contractions from Alignment Bias
A second failure pattern is a false contraction call caused by alignment bias. The aligner may fail to place reads that span the full repeat array, resulting in an underestimate of the repeat count. This pattern is more common in short-read data where the read length is shorter than the repeat array.
The solution is to use a repeat-aware aligner that can soft-clip reads at the repeat boundary and to use an STR-specific caller that accounts for the alignment bias. Long-read sequencing can resolve this issue by providing reads that span the entire repeat array.
Missed Expansions in Low-Complexity Regions
A third failure pattern is a missed expansion in a low-complexity region. The aligner may map reads from an expanded allele to the reference allele, resulting in a homozygous reference call when the sample is actually heterozygous for an expansion. This pattern is more common in loci with very long repeat arrays or with motifs that are similar to nearby sequences.
The solution is to use an STR-specific caller that is designed to detect expansions and to validate negative calls in clinically significant loci with an orthogonal method.
Limitations of Current STR Genotyping Approaches
Short-Read Limitations
Short-read sequencing cannot span very long repeat arrays, and the alignment ambiguity in repetitive regions limits the accuracy of repeat count estimates. The stutter pattern from PCR amplification further complicates the interpretation. These limitations are inherent to the technology and cannot be fully overcome by computational methods.
The practical implication is that short-read sequencing with an STR-specific caller is suitable for genotyping common STR polymorphisms but may miss large expansions. For clinical applications where large expansions are expected, long-read sequencing or orthogonal validation is required.
Long-Read Limitations
Long-read sequencing provides direct evidence of repeat counts but has a higher base-level error rate than short-read sequencing. The error rate is concentrated in homopolymers and tandem repeats, which are exactly the regions where STR genotyping is needed. The result is that long-read data can resolve large expansions but may have difficulty with small repeat count differences.
The practical implication is that long-read sequencing should be used in combination with short-read data or with careful quality control. The long-read data provides the overall structure of the repeat array, and the short-read data provides the base-level accuracy.
Computational Limitations
STR-specific callers require a catalog of known STR loci. Novel expansions at uncharacterized loci will not be detected by these tools. The catalog must be updated as new STR loci are discovered and characterized.
The computational resources required for STR genotyping are modest compared to the alignment and variant calling steps, but the analysis adds complexity to the pipeline. The STR genotype calls must be integrated with the standard variant calls, and the quality metrics must be interpreted in the context of the sample and the sequencing platform.
Benchmarking and Validation of STR Calls
Using Pedigree-Based Truth Sets
The evaluation of STR variant calling performance requires comprehensive truth sets that include tandem repeat variants. A large pedigree study that combined PacBio high-fidelity, Illumina, and Oxford Nanopore Technologies platforms generated a variant map with over 4.7 million single-nucleotide variants, 767,795 insertions and deletions, 537,486 tandem repeats, and 24,315 structural variants covering 2.77 Gb of the GRCh38 genome. This work added approximately 200 Mb of high-confidence regions and introduced the first tandem repeat and structural variant truth sets for NA12878 and her family.
The availability of these truth sets allows researchers to benchmark their STR genotyping pipelines against a comprehensive standard. When validating a new pipeline, compare the STR calls from your pipeline to the truth set and calculate sensitivity and precision separately for tandem repeat variants. This comparison will reveal whether your pipeline has a systematic bias toward false expansions or false contractions.
Platform Concordance Checks
A practical validation approach is to compare STR calls across sequencing platforms. The pedigree study demonstrated that combining data from multiple platforms produces a more complete variant map than any single platform alone. For your own samples, consider sequencing a small number of samples on both short-read and long-read platforms and comparing the STR genotype calls.
Platform concordance is a useful quality metric. If the short-read and long-read calls agree at a high rate across known STR loci, the confidence in both platforms is increased. If the calls disagree, investigate the source of the discrepancy. The discrepancy may be due to alignment bias in the short-read data, base-level errors in the long-read data, or a true difference in the sample.
Retraining and Parameter Optimization
The pedigree-based truth set has been used to retrain variant calling models. One study reported that retraining DeepVariant using the expanded truth set reduced genotyping errors by approximately 34 percent. This result demonstrates that the quality of the training data directly affects the accuracy of the variant caller.
For your own pipeline, consider whether the variant caller you use has been trained on data that includes STR regions. If the caller was trained primarily on unique regions, it may have a systematic bias in repetitive regions. Retraining the caller with STR-inclusive truth sets or adjusting the parameters to account for repetitive regions can improve accuracy.
Integration with Standard Variant Calling Workflows
Adding STR Calls as an Annotation Layer
The STR genotype calls from an STR-specific caller should be integrated into the standard variant calling workflow as an additional annotation layer. The STR calls provide repeat count estimates for known loci, and these estimates can be used to interpret or override indel calls from the standard variant caller.
The integration requires a structured format for the STR calls that can be merged with the variant calling format output. The STR calls should include the locus identifier, the repeat count for each allele, the confidence interval, and the quality metrics. This information can be added to the variant annotation for downstream analysis.
Handling Conflicts Between Callers
When the standard variant caller and the STR-specific caller disagree, the discrepancy must be investigated. The STR-specific caller is generally more reliable in STR regions because it models the stutter pattern and the alignment ambiguity. However, the STR-specific caller may miss variants that the standard caller detects, particularly if the variant is outside the catalog of known STR loci.
The investigation should include a review of the raw reads at the locus, the alignment pattern, and the quality metrics from both callers. If the standard caller reports an indel that the STR-specific caller does not support, examine the read support for the indel. If the reads supporting the indel are clustered at the repeat boundary, the call is likely an artifact.
Workflow Management and Reproducibility
The integration of STR genotyping into a variant calling workflow requires careful management of the pipeline steps and parameters. A workflow management system can track the execution of each step and ensure that the pipeline is reproducible. The nf-core documentation provides standards for community pipelines that include version tracking, configuration management, and containerization. The Bioconductor project offers packages for genomic analysis that can be integrated into reproducible workflows, and the Galaxy Training Network provides tutorials for building and running analysis workflows.
The pipeline should be tested on a control sample with known STR genotypes before it is applied to research or clinical samples. The control sample provides a baseline for the expected STR calls and allows you to verify that the pipeline is working correctly.
Safety and Regulatory Context for STR Variant Reporting
Clinical Reporting Standards
For clinical applications, STR variant reporting must follow the laboratory's validation protocols and the applicable regulatory requirements. The laboratory must validate the STR genotyping method before using it for clinical reporting, and the validation must demonstrate the accuracy and precision of the method.
The clinical report should include the STR genotype, the method used to generate the call, and the confidence in the call. If the call is based on a single method, the report should note the limitations of that method and any recommendations for confirmatory testing.
Data Sharing and Reproducibility
The reproducibility of STR variant calls depends on the availability of the raw data, the pipeline parameters, and the reference genome version. Data sharing should follow the applicable data use agreements and privacy regulations. The NCBI data resources provide repositories for raw sequencing data and processed variant calls that support data sharing and reproducibility. The EMBL-EBI training resources offer guidance on data management and sharing practices for genomic data.
For research applications, the STR genotype calls should be reported with sufficient detail to allow other researchers to reproduce the analysis. This includes the tool versions, the parameters, and the quality metrics.
Professional Escalation Criteria
A laboratory professional should escalate an STR variant call for review when the call is clinically significant and the confidence is low, when the call conflicts with the standard variant caller, or when the call is unexpected based on the clinical presentation.
The escalation process should include a review of the raw data, the alignment, and the genotype call. The review should determine whether the call is a true variant or an artifact and whether additional testing is needed.
A Decision Framework for Choosing Between STR Genotyping Strategies
Defining the Decision Context
The choice between short-read sequencing with STR-specific callers, long-read sequencing, or a hybrid approach depends on factors that can be assessed before sequencing begins. A structured decision framework helps laboratory professionals match the genotyping strategy to the biological question, the sample characteristics, and the available resources. The framework presented here uses four decision gates: expected variant spectrum, sample throughput, cost constraints, and validation requirements.
Decision Gate 1: Expected Variant Spectrum
The first decision gate asks what types of STR variants are expected in the sample set. If the study targets known STR loci with established disease associations, such as the 60 loci known to cause genetic disorders including Huntington disease and fragile X syndrome, short-read sequencing combined with an STR-specific caller can provide adequate genotyping accuracy. These loci have well-characterized motif sequences and reference repeat counts, which allows the STR-specific caller to model the expected stutter pattern and alignment ambiguity.
If the study aims to discover novel STR expansions or genotype very long repeat arrays that exceed the span of short reads, long-read sequencing becomes necessary. A study of rare disease diagnostics using nanopore sequencing demonstrated that long-read data detected additional tandem repeats that were not covered by short-read sequencing, and diagnostic variants were established in 11 probands with diverse underlying genetic causes. The long-read approach provided access to variants that were invisible to the short-read workflow.
For studies that combine genome-wide variant discovery with targeted STR validation, a hybrid approach is appropriate. The short-read data provides genome-wide coverage for single nucleotide variants and small indels, while the long-read data resolves STR regions that produce low-confidence calls in the short-read pipeline.
Decision Gate 2: Sample Throughput and Scale
The second decision gate considers the number of samples and the scale of the project. Population-scale studies that genotype common STR polymorphisms across hundreds or thousands of samples require a cost-effective approach. Short-read sequencing with an STR-specific caller is the practical choice because the per-sample cost is substantially lower than long-read sequencing.
A scalable nanopore sequencing protocol demonstrated that long-read sequencing can be applied to population-scale projects, but the study authors noted that the approach required careful optimization of the wet lab and computational protocol. The protocol achieved single nucleotide polymorphism detection with F1-scores comparable to Illumina short-read sequencing, but small indel calling remained difficult within homopolymers and tandem repeats. This suggests that even optimized long-read workflows face challenges in STR regions.
For clinical diagnostics where the number of samples is smaller but the consequences of a missed variant are higher, the cost per sample is less important than the completeness of the variant detection. Long-read sequencing or a hybrid approach may be justified even at higher cost.
Decision Gate 3: Cost and Resource Constraints
The third decision gate evaluates the financial and computational resources available for the project. Short-read sequencing remains the most cost-effective option for most applications. The computational resources required for STR genotyping with an STR-specific caller are modest compared to the alignment and variant calling steps.
Long-read sequencing requires higher per-sample costs for sequencing and additional computational resources for base calling and alignment. The error profile of long-read data also requires careful quality control, particularly in homopolymers and tandem repeats where the base-level error rate is concentrated.
The hybrid approach balances cost and completeness by using short-read data for genome-wide coverage and long-read data for targeted validation. This approach is particularly useful when the budget cannot support long-read sequencing for all samples but the study includes clinically significant STR loci that require confirmation.
Decision Gate 4: Validation and Reporting Requirements
The fourth decision gate considers the validation and reporting standards that apply to the project. Research applications may accept STR genotype calls from a single method with appropriate quality metrics. Clinical applications require validation of the genotyping method before it is used for reporting, and the validation must demonstrate the accuracy and precision of the method.
A study that evaluated genome sequencing as a diagnostic strategy for rare disease used a threshold of 98 percent true positive rate to decide whether genome sequencing could replace existing workflows. The study found that overall detection was 95 percent across 1271 known clinically relevant variants, with detection rates differing by variant category. This result demonstrates that even well-validated workflows miss a meaningful fraction of variants, and the validation requirements should reflect the clinical consequences of a missed variant.
For clinical reporting, low-confidence or clinically significant STR calls should be confirmed with an orthogonal method. The orthogonal method could be long-read sequencing, PCR-based fragment analysis, or a second STR-specific caller with a different algorithmic approach.
Implementing the Decision Framework
To apply the decision framework, document the answers to the four decision gates for each project or sample set. Record the expected variant spectrum, the sample throughput, the cost constraints, and the validation requirements. Use these records to select the genotyping strategy and to justify the selection in the project documentation.
The decision framework should be revisited when new information becomes available. If the short-read pipeline produces a high rate of low-confidence STR calls in a clinically significant locus, the decision may shift toward long-read validation. If the long-read pipeline produces concordant calls with the short-read pipeline across a validation set, the confidence in the short-read approach increases.
Recording the Decision and Outcomes
Maintain a decision record for each project that includes the rationale for the genotyping strategy, the expected variant spectrum, the sample throughput, the cost constraints, and the validation requirements. Record the outcomes of the genotyping, including the number of STR loci genotyped, the number of low-confidence calls, and the results of any orthogonal validation.
The decision record should be linked to the sample-level records and the pipeline-level records. This linkage allows the laboratory to evaluate the performance of the decision framework over time and to identify patterns that suggest a different strategy might be more appropriate for certain sample types or variant classes.
Common Decision Errors
A common error is selecting a short-read approach without assessing the expected variant spectrum. If the sample set includes individuals with suspected large STR expansions, the short-read approach will produce false negative calls that could have clinical consequences. The decision framework should be applied before sequencing begins, not after the data has been generated.
Another common error is selecting a long-read approach for all samples without considering the cost and throughput implications. The long-read approach provides additional information but at a higher cost, and the additional information may not be necessary for all samples. The decision framework should match the strategy to the specific requirements of each project.
A third error is failing to document the decision and the outcomes. Without a decision record, the laboratory cannot evaluate whether the chosen strategy was appropriate or whether a different strategy would have produced better results. The decision record is an essential component of the quality management system.
Frequently Asked Questions
Why do standard variant callers produce false indels in STR regions?
Standard variant callers assume that the reference genome is accurate and that reads align uniquely to the genome. In STR regions, reads can align to multiple positions with equal scores, and the aligner may introduce gaps to force a unique alignment. The variant caller then interprets these gaps as insertions or deletions. PCR slippage during library preparation adds another layer of error by producing reads with different repeat counts than the template. The combination of alignment ambiguity and PCR stutter produces false indel calls that do not reflect the true genotype.
What is the difference between an STR-specific caller and a standard variant caller?
A standard variant caller produces a single consensus genotype for each genomic position and reports variants relative to the reference. An STR-specific caller such as ExpansionHunter estimates the repeat count for each allele at known STR loci and provides a confidence interval for each estimate. The STR-specific caller models the stutter pattern from PCR and the alignment ambiguity in repetitive regions, which allows it to distinguish true repeat count differences from technical artifacts.
When should I use long-read sequencing for STR genotyping?
Long-read sequencing should be used when the repeat array is too long to be spanned by short reads, when the clinical presentation suggests a large expansion that short-read sequencing might miss, or when short-read STR calls are low confidence and require validation. Long-read sequencing is also useful for discovering novel STR expansions at uncharacterized loci. The higher cost and lower throughput of long-read sequencing should be weighed against the diagnostic or research value of the additional information.
How do I validate a low-confidence STR call?
A low-confidence STR call should be validated with an orthogonal method. This could be long-read sequencing, PCR-based fragment analysis, or a second STR-specific caller with a different algorithmic approach. The validation method should provide an independent estimate of the repeat count, and the results should be compared to the original call. If the validation method confirms the call, the confidence can be upgraded. If the validation method disagrees, the discrepancy should be investigated.
What quality metrics should I record for STR genotype calls?
Record the repeat count for each allele, the confidence interval for each estimate, the coverage depth at the locus, the stutter ratio, and the allele balance. These metrics provide the information needed to assess the reliability of the call and to compare calls across samples. For clinical applications, record the method used to generate the call and any validation results.
Can somatic STR variants be called with the same tools as germline variants?
The same STR-specific callers can be used for somatic and germline variant calling, but the interpretation differs. In somatic tissues, an STR expansion may be present in only a fraction of cells, and the allele balance reflects the mosaic fraction. The confidence interval for the repeat count estimate must account for the lower variant allele fraction. Somatic STR calls should be interpreted in the context of the sample type and the expected variant spectrum.
How do I integrate STR genotype calls into my existing variant calling workflow?
Run the STR-specific caller on the aligned reads from your existing pipeline. The output provides repeat count estimates for known STR loci. Compare these estimates to the indel calls from your standard variant caller and use the STR genotype as the authoritative source in STR regions. Add the STR genotype calls to your variant annotation and include them in your reporting. The integration requires a catalog of known STR loci and a workflow that can combine the STR calls with the standard variant calls.
What are the limitations of STR-specific callers?
STR-specific callers require a catalog of known STR loci and cannot detect novel expansions at uncharacterized loci. The accuracy of the repeat count estimate depends on the coverage depth, the stutter pattern, and the alignment quality. Very long repeat arrays may not be fully resolved by short-read data, and the confidence interval may be wide. The caller provides an estimate with a confidence interval, but the estimate should be validated with an orthogonal method when the call is clinically significant or when the confidence is low.
Related Bioinformatics Guides
- Detecting Structural Variants with Long-Read Sequencing: Methods and Considerations
- RNA-Seq Alignment: Choosing the Right Tool and Parameters
- RNA-Seq Alignment Tools: STAR, HISAT2, and Beyond
- Oxford Nanopore Sequencing: From Sample to Base Calls
- Short-Read vs Long-Read Sequencing: Pros, Cons, and Selection Criteria
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- Sequencing and characterizing short tandem repeats in the human genome.. Nature reviews. Genetics, 2024.
- Advancing long-read nanopore genome assembly and accurate variant calling for rare disease detection.. American journal of human genetics, 2025.
- Scalable Nanopore sequencing of human genomes provides a comprehensive view of haplotype-resolved variation and methylation.. Nature methods, 2023.
- Genome sequencing as a generic diagnostic strategy for rare disease.. Genome medicine, 2024.
- The Platinum Pedigree: a long-read benchmark for genetic variants.. Nature methods, 2025.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.