Somatic Variant Calling in Low-Purity or Low-Coverage Tumors: Strategies to Maximize Sensitivity and Specificity

By Dr. Zubair Khalid, DVM, MS, PhD ·

Somatic Variant Calling in Low-Purity or Low-Coverage Tumors: Strategies to Maximize Sensitivity and Specificity

Key Takeaways

  • Low tumor purity (<20% neoplastic content) necessitates substantially increased sequencing depth or error-correcting sequencing approaches to improve detection of low variant allele fraction (VAF) mutations, though this increases cost and may miss subclonal events.
  • For low coverage (<50X) samples, specialized variant callers designed for low-depth data or tumor-only approaches with robust germline filtering are recommended to maintain sensitivity, but increased false positive rates require careful downstream filtering.
  • Absence of a matched normal sample requires tumor-only callers with effective germline filtering and population frequency filters to distinguish somatic mutations from rare germline variants, though this limits discrimination.
  • High sequencing error regions, particularly low-mappability and repetitive areas, benefit from machine learning callers trained on error-prone genomic contexts to improve F1 scores, but these demand substantial computational resources.
  • Long-read sequencing data requires specific callers to leverage improved phasing and low VAF variant detection; while offering advantages, fewer established tools and higher per-base costs are limitations.
  • Effective somatic variant calling in challenging samples hinges on a tiered filtering strategy, starting with technical filters (depth, allele count, strand bias) and progressing to context-specific, population frequency, and validation-focused filters to balance sensitivity and specificity.

Researchers analyzing tumor samples with low purity or shallow sequencing depth face a distinct problem: true somatic variants hide among sequencing artifacts, germline polymorphisms, and allelic dropout. Standard variant callers tuned for high-quality, high-depth data often miss low allele fraction mutations or report false positives that waste validation resources. This article provides practical strategies for improving somatic variant detection in challenging samples, covering sequencing design, caller selection, filtering approaches, and quality control measures that balance sensitivity with specificity.

At a Glance

The table below summarizes key strategies for somatic variant calling under different challenging conditions. Each approach addresses a specific limitation and should be selected based on sample characteristics and research goals.

ChallengePrimary StrategyExpected BenefitKey Limitation
Low tumor purity (below 20% neoplastic content)Increase sequencing depth substantially, use error-correcting sequencing approachesImproved detection of low variant allele fraction mutationsHigher cost per sample, may still miss subclonal events
Low coverage (below 50X)Use specialized callers designed for low-depth data, consider tumor-only approaches with robust filteringMaintains reasonable sensitivity without excessive costIncreased false positive rate requires careful filtering
No matched normal sampleUse tumor-only callers with germline filtering, apply population frequency filtersEnables analysis of archival or single-sample cohortsCannot distinguish rare germline variants from somatic mutations
High sequencing error regionsApply machine learning callers trained on error-prone genomic contextsBetter F1 scores in low-mappability and repetitive regionsRequires substantial computational resources
Long-read sequencing dataUse long-read specific callers instead of short-read toolsImproved phasing and detection of variants at low allele fractionsFewer established tools, higher per-base cost

Understanding the Challenge of Low-Purity and Low-Coverage Samples

Somatic variant calling relies on distinguishing true mutations present in tumor cells from technical artifacts and germline variation. When tumor purity drops, the variant allele fraction (VAF) of true somatic mutations decreases proportionally. A mutation present in all tumor cells at 50% purity appears at approximately 25% VAF, while the same mutation in a 10% pure sample appears at only 5% VAF. At these low fractions, sequencing errors and alignment artifacts become proportionally more significant, making accurate detection difficult.

Low sequencing coverage compounds this problem. Each base position receives fewer independent read observations, reducing statistical power to distinguish true variants from noise. The combination of low purity and low coverage creates a particularly difficult scenario where true variants may be represented by only a handful of reads, and any single sequencing error can mimic a genuine mutation.

The practical consequence is a tradeoff between sensitivity and specificity. Researchers can lower detection thresholds to catch more true variants, but this inevitably increases false positives. Conversely, strict thresholds reduce false discoveries but miss clinically or biologically relevant mutations. Understanding this tradeoff and implementing strategies to shift the balance favorably is the core challenge addressed in this article.

Core Principles of Somatic Variant Detection

The Role of Matched Normal Samples

The gold standard for somatic variant calling uses a matched normal sample from the same individual. Comparing tumor and normal sequencing data allows the caller to subtract germline variants and focus on mutations unique to the tumor. This approach dramatically reduces the false positive burden because common polymorphisms and inherited variants are identified and excluded.

However, matched normal samples are frequently unavailable in clinical diagnostics, retrospective analyses of archival tumor samples in biobanks, and many research settings. In these situations, variant calling accuracy suffers because distinguishing somatic mutations from germline variants or sequencing artifacts becomes substantially more difficult. Recent developments in tumor-only calling methods have addressed this gap, with deep learning approaches demonstrating meaningful performance improvements over earlier methods when no matched normal is available [<a href="#ref-1">1</a>].

Variant Allele Fraction and Detection Limits

The variant allele fraction represents the proportion of sequencing reads supporting the alternative allele at a given position. For clonal mutations in pure tumor samples, VAF approaches 50% for heterozygous events and 100% for homozygous events. As purity decreases, these values drop proportionally.

Detection limits depend on both depth and error rate. At 100X coverage, a true variant at 5% VAF is supported by approximately five reads. The probability of observing this variant depends on the sequencing error rate at that position. Error-correcting sequencing approaches and specialized callers can reduce the impact of background noise, effectively lowering the minimum detectable VAF.

Sequencing Errors and Artifacts

Sequencing platforms introduce characteristic error patterns. Base substitution errors occur at varying rates depending on sequence context, while indels in homopolymer and repetitive regions are particularly problematic. Formalin-fixed paraffin-embedded (FFPE) samples introduce additional artifacts, including deamination of cytosine to uracil, which manifests as C to T transitions.

Machine learning approaches have proven particularly effective at learning these error patterns and distinguishing them from true variants. Callers trained on experimentally confirmed variant data can identify somatic mutations even in genomic regions characterized by high sequencing error rates, achieving better performance than traditional statistical and heuristic methods in these challenging contexts [<a href="#ref-2">2</a>].

Sequencing Design Considerations

Increasing Sequencing Depth

The most direct strategy for improving sensitivity in low-purity samples is increasing sequencing depth. Higher depth provides more observations at each position, improving statistical power to detect low VAF variants. Clinical whole-exome sequencing assays validated for somatic variant detection typically require substantial depth to achieve acceptable sensitivity [<a href="#ref-3">3</a>].

For samples with expected low purity, researchers should consider depth targets substantially higher than standard germline sequencing. The relationship between depth, purity, and detection sensitivity is not linear, and the marginal benefit of additional depth diminishes at very high coverage. Cost considerations and the law of diminishing returns should guide depth decisions based on the specific research or clinical question.

Error-Correcting Sequencing Approaches

Error-correcting sequencing methods use unique molecular identifiers or similar strategies to distinguish true variants from polymerase and sequencing errors. By grouping reads derived from the same original DNA molecule, these approaches can identify and correct errors introduced during library preparation and sequencing.

These methods are particularly valuable for low-purity samples where true VAFs approach the error rate of standard sequencing. The error correction effectively lowers the noise floor, allowing detection of variants at allele fractions that would otherwise be indistinguishable from artifacts.

Long-Read Sequencing Considerations

Long-read sequencing platforms offer advantages for somatic variant detection, particularly improved read phasing that enables more accurate single-nucleotide variation detection at low variant allele fractions [<a href="#ref-4">4</a>]. The ability to span longer genomic regions allows better alignment in repetitive and structurally complex regions.

However, long-read data requires specialized variant callers. Most existing somatic variant callers were designed for short-read sequencing data and do not perform well with long-read data [<a href="#ref-4">4</a>]. Deep learning methods developed specifically for long-read tumor-normal pairs have demonstrated robust performance across varied coverage, purity, and contamination levels, making them suitable for challenging samples [<a href="#ref-4">4</a>].

Target Enrichment and Panel Design

For targeted sequencing approaches, panel design directly impacts variant detection capability. Clinical whole-exome sequencing workflows have been optimized to detect somatic variants with high sensitivity and low false positive rates [<a href="#ref-3">3</a>]. These validated assays demonstrate that careful attention to target-enrichment strategy and downstream analysis can achieve clinical-grade performance [<a href="#ref-3">3</a>].

When designing panels for low-purity samples, consider including genomic regions with known relevance to the research question while maintaining sufficient depth across all targets. Overly broad panels may dilute sequencing capacity, reducing depth at clinically important loci.

Variant Calling Workflow Options

Matched Tumor-Normal Calling

When matched normal samples are available, paired analysis remains the preferred approach. Tumor-normal callers compare allele frequencies and genotypes between samples to identify variants specific to the tumor. This comparison effectively removes germline polymorphisms from consideration, reducing the false positive burden.

Statistical and heuristic methods have traditionally dominated this space, but machine learning approaches have demonstrated improved sensitivity, particularly in challenging genomic regions. Callers trained on experimentally confirmed variant data can achieve higher sensitivity than traditional methods while maintaining comparable or better precision [<a href="#ref-2">2</a>].

Tumor-Only Calling

Tumor-only calling is necessary when matched normal samples are unavailable. This approach must distinguish somatic mutations from both germline variants and sequencing artifacts without the benefit of a direct comparison. Population frequency databases help filter common polymorphisms, but rare germline variants remain problematic.

Deep learning frameworks designed for tumor-only analysis have demonstrated significant performance improvements over existing methods [<a href="#ref-1">1</a>]. These approaches are trained using large numbers of high-confidence variants and can accurately identify somatic variants from aligned tumor reads without a matched normal sample [<a href="#ref-1">1</a>]. The improved accuracy has practical implications for tumor mutation burden estimation and patient selection for immunotherapy [<a href="#ref-1">1</a>].

Long-Read Specific Callers

Long-read somatic variant calling requires tools designed for the error profiles and read characteristics of platforms such as Oxford Nanopore and PacBio. Deep learning methods trained on synthetic somatic variants with diverse coverages and variant allele fractions can accurately detect a wide range of somatic variants in long-read data [<a href="#ref-4">4</a>].

Tumor-only long-read callers use ensembles of neural networks trained for opposite tasks, assessing how likely or unlikely a candidate is a somatic variant [<a href="#ref-5">5</a>]. These approaches outperform short-read callers applied to long-read data and are also applicable to short-read data, providing flexibility across sequencing platforms [<a href="#ref-5">5</a>].

Machine Learning Approaches

Machine learning has transformed somatic variant calling, particularly for challenging samples. Deep learning models can learn complex patterns distinguishing true variants from artifacts, incorporating sequence context, read characteristics, and genomic features that traditional methods handle poorly [<a href="#ref-2">2</a>].

The key advantage of machine learning approaches is their ability to integrate diverse information sources and learn from experimentally confirmed variants [<a href="#ref-2">2</a>]. Active learning strategies, where predicted variants are experimentally confirmed and incorporated into training data, can continuously improve model performance [<a href="#ref-2">2</a>]. This combination of computational prediction with experimental validation creates a virtuous cycle that improves accuracy over time.

Practical Implementation Steps

Step 1: Assess Sample Characteristics

Before selecting a variant calling strategy, evaluate the expected purity and available sequencing depth. Review pathology reports for estimated tumor content, consider the sample type (fresh frozen versus FFPE), and assess DNA quality. These factors directly influence the appropriate calling approach and depth requirements.

For samples with unknown purity, consider performing a quick assessment using copy number analysis or allele-specific methods that can estimate tumor content from sequencing data itself. This information guides subsequent analysis decisions.

Step 2: Select Appropriate Sequencing Strategy

Based on sample characteristics and research goals, determine whether additional sequencing is needed before analysis. Low-purity samples may benefit from increased depth, while samples already sequenced at adequate depth may proceed directly to analysis.

Consider whether error-correcting approaches would provide meaningful benefit. For samples where true VAFs are expected to approach the sequencing error rate, the additional cost of error correction may be justified. For higher purity samples, standard sequencing may suffice.

Step 3: Choose Variant Caller Based on Data Type and Normal Availability

Match the caller to the data characteristics. Paired tumor-normal short-read data should use established paired callers or machine learning approaches trained on matched samples. Tumor-only data requires callers specifically designed for this scenario [<a href="#ref-1">1</a>]. Long-read data requires long-read specific tools [<a href="#ref-4">4</a>][<a href="#ref-5">5</a>].

Consider running multiple callers and intersecting or combining results. Different callers have different strengths and weaknesses, and consensus approaches can improve confidence in detected variants. However, this increases computational requirements and may reduce sensitivity if strict intersection is applied.

Step 4: Apply Appropriate Filters

Filtering is essential for managing false positives, particularly in low-purity or low-coverage samples. Standard filters include minimum depth, minimum alternative allele count, strand bias, and mapping quality. Population frequency filters help remove common germline variants in tumor-only analyses.

Machine learning callers often incorporate filtering into their models, but additional post-hoc filtering may still be necessary. Validate filter settings on known positive and negative controls when available.

Step 5: Validate Critical Variants

For variants that will drive downstream decisions, validation is essential. Targeted deep sequencing can confirm predicted variants and distinguish true mutations from artifacts. This approach is particularly important for low VAF variants where the confidence of initial calling may be limited [<a href="#ref-2">2</a>].

Experimental confirmation also provides valuable feedback for improving calling algorithms. Incorporating validated results into training data can improve performance for future analyses [<a href="#ref-2">2</a>].

Records and Measurements

Tracking Variant Calling Performance

Maintain detailed records of variant calling performance across samples. Track sensitivity and precision metrics when validation data are available. Document the number of variants called, the VAF distribution, and the proportion of variants passing each filtering step.

These records help identify systematic issues with specific sample types or sequencing batches. A sudden change in variant yield or VAF distribution may indicate technical problems requiring investigation.

Documenting Sample Metadata

Comprehensive sample metadata is essential for interpreting variant calling results. Record tumor purity estimates, sample type, DNA quality metrics, sequencing platform, depth, and any known technical issues. This information contextualizes variant calls and helps identify samples where results may be unreliable.

For clinical samples, maintain documentation of the validation status of the assay and any relevant quality control metrics. Clinical whole-exome sequencing assays validated for somatic variant detection should have documented sensitivity and false positive rates across variant allele fraction and depth ranges [<a href="#ref-3">3</a>].

Quality Control Metrics

Track standard sequencing quality metrics including mean depth, uniformity of coverage, and error rates. For targeted panels, monitor on-target rates and the proportion of bases meeting minimum depth thresholds. These metrics directly impact variant calling sensitivity and should be reviewed before proceeding with analysis.

For low-purity samples, consider the effective depth at variant sites. A variant at 5% VAF in a sample with 100X mean depth may be supported by only a few reads, and the confidence of the call depends on the actual depth at that specific position instead of the mean across the sample.

Common Failure Patterns

Excessive False Positives from Low-Complexity Regions

Low-complexity and repetitive genomic regions generate alignment artifacts that mimic true variants. These regions are particularly problematic for short-read sequencing where reads may map to multiple locations. Machine learning callers trained on experimentally confirmed data have demonstrated improved performance in these regions, but no method is perfect [<a href="#ref-2">2</a>].

When reviewing variant calls, pay particular attention to variants in homopolymer runs, microsatellites, and segmental duplications. These are frequently artifacts, and validation should be considered before investing resources in follow-up studies.

Missed Variants Due to Allelic Dropout

Allelic dropout occurs when one allele fails to amplify or sequence efficiently, leading to loss of true variants. This is particularly problematic in FFPE samples where DNA damage can prevent amplification of specific alleles. The result is false negative calls that are difficult to detect without validation.

Samples with low DNA quality should be analyzed with awareness of this limitation. Consider whether the research question can tolerate potential false negatives, or whether additional sequencing approaches are needed to overcome dropout.

Contamination Confounding Results

Sample contamination introduces alleles from other individuals, creating false positive somatic calls. This is particularly problematic for tumor-only analyses where contamination cannot be distinguished from true somatic variation without additional information.

Assess contamination levels using available tools and exclude or flag samples with significant contamination. For low-purity samples, the impact of contamination is amplified because the signal from true variants is already weak.

Overly Aggressive Filtering Reducing Sensitivity

In an effort to reduce false positives, researchers may apply filters so stringent that true variants are lost. This is particularly problematic for low VAF variants that may be filtered by minimum allele count or quality thresholds.

Balance filtering stringency against the research goals. If the goal is discovery of all potential variants, accept a higher false positive rate and validate candidates. If the goal is high-confidence calls for clinical decision-making, prioritize specificity even at the cost of sensitivity.

Limitations and Interpretation

Variant Calling Is Probabilistic

Somatic variant calls represent statistical inferences instead of definitive determinations. The confidence in any call depends on depth, VAF, sequencing quality, and genomic context. Low-purity and low-coverage samples produce calls with inherently lower confidence, and this uncertainty should be reflected in interpretation.

For research applications, consider reporting confidence scores and the evidence supporting each call. For clinical applications, ensure that the assay validation supports the intended use and that limitations are clearly documented [<a href="#ref-3">3</a>].

Tumor Heterogeneity Complicates Interpretation

Tumors are heterogeneous, containing multiple subclones with different mutations. Low VAF variants may represent subclonal mutations instead of artifacts, and the distinction has biological significance. The inability to detect very low VAF variants means that some subclonal mutations will be missed.

When interpreting results, consider whether the research question requires detection of subclonal mutations or only clonal events. This distinction guides depth requirements and filtering strategies.

Platform-Specific Considerations

Different sequencing platforms have different error profiles, and variant callers are often optimized for specific platforms. Long-read data requires different approaches than short-read data, and applying inappropriate tools reduces accuracy [<a href="#ref-4">4</a>][<a href="#ref-5">5</a>].

When changing platforms, validate the variant calling workflow on appropriate controls. The performance characteristics documented for one platform may not transfer to another.

Quality Control and Validation

Using Reference Materials

Reference materials with known variants provide essential quality control for somatic variant calling workflows. Cell line mixtures containing large numbers of simulated variants can benchmark assay performance across VAF ranges and genomic contexts [<a href="#ref-3">3</a>].

Regular analysis of reference materials helps detect workflow drift and ensures consistent performance over time. Document performance metrics and investigate any significant changes.

Orthogonal Validation

For critical variants, orthogonal validation using an independent method provides the highest confidence. Targeted deep sequencing of candidate variants can confirm or refute calls made from initial sequencing data [<a href="#ref-2">2</a>].

The validation approach should be appropriate for the VAF of the variant. Low VAF variants require validation methods with sufficient sensitivity to detect the variant at the expected allele fraction.

Clinical Validation Requirements

Clinical somatic variant calling requires documented validation demonstrating acceptable sensitivity and false positive rates. Validated assays should have established performance characteristics across the range of sample types and variant types for which they will be used [<a href="#ref-3">3</a>].

For clinical whole-exome sequencing, validation typically includes assessment of sensitivity at defined VAF thresholds and depth requirements, as well as false positive rates per megabase [<a href="#ref-3">3</a>]. These metrics guide appropriate use and interpretation of results.

Professional Escalation Criteria

When to Seek Additional Expertise

Researchers should consider consulting bioinformatics specialists or clinical genomics professionals when facing specific challenging situations. These include samples with extremely low purity where standard approaches fail, unusual variant patterns suggesting technical issues, or clinical decisions that depend on low-confidence variant calls.

Bioinformatics training resources from established organizations can help build the skills needed to address challenging samples. Foundational training in computing, data analysis, and genomic methods provides the basis for understanding and troubleshooting variant calling issues [<a href="#ref-6">6</a>][<a href="#ref-7">7</a>].

When to Consider Re-sequencing

Re-sequencing may be appropriate when initial data quality is inadequate for the research question. Indicators include mean depth substantially below targets, poor uniformity of coverage, high duplication rates, or evidence of sample degradation.

The decision to re-sequence should weigh the cost of additional sequencing against the value of improved data. For samples where the research question cannot be answered with existing data quality, re-sequencing may be the most efficient path forward.

When to Report Limitations

Research publications and clinical reports should clearly document the limitations of variant calling for low-purity or low-coverage samples. This includes the sensitivity and specificity of the assay at relevant VAF ranges, the depth achieved, and any known issues with specific genomic regions.

Transparent reporting allows readers and clinicians to appropriately interpret results and understand the confidence that can be placed in individual variant calls.

Building a Sample-Specific Decision Framework for Low-Purity and Low-Coverage Variant Calling

Selecting the right variant calling strategy for challenging samples requires more than choosing a tool from a list. Researchers need a structured decision process that accounts for the specific characteristics of each sample, the available computational resources, and the downstream research or clinical question. This section provides a practical framework for making these decisions, along with a record system for tracking performance and a troubleshooting method for common problems.

Step 1: Characterize the Sample Before Choosing a Caller

The first decision point occurs before any variant calling begins. Document the following sample characteristics and use them to guide all subsequent choices:

Estimated tumor purity. Review pathology reports for the neoplastic cell content estimate. If this information is unavailable or unreliable, consider estimating purity from the sequencing data itself using copy number analysis or allele-specific methods. Record whether the purity estimate comes from pathology review or computational inference, because these methods have different error profiles.

Sample type and DNA quality. Fresh frozen tissue generally produces higher quality data than formalin-fixed paraffin-embedded (FFPE) samples. FFPE samples introduce characteristic artifacts including cytosine deamination that manifests as C to T transitions. Record the sample type and any available DNA quality metrics such as DIN (DNA Integrity Number) or fragment size distribution.

Expected variant allele fraction range. For a clonal mutation in a sample with 20% purity, the expected VAF is approximately 10% for heterozygous events. Subclonal mutations will appear at even lower fractions. Estimate the minimum VAF that is biologically or clinically relevant for the research question, because this directly determines the required sequencing depth and the appropriate caller.

Sequencing platform and depth. Document the platform, the mean depth across the target region, and the uniformity of coverage. For targeted panels, record the on-target rate and the proportion of bases meeting minimum depth thresholds. These metrics directly impact variant calling sensitivity and should be reviewed before proceeding with analysis [<a href="#ref-3">3</a>].

Matched normal availability. Determine whether a matched normal sample exists and whether its sequencing data are of sufficient quality for paired analysis. If no matched normal is available, tumor-only calling approaches are required, and the analysis must account for the difficulty of distinguishing somatic mutations from germline variants or sequencing artifacts [<a href="#ref-1">1</a>].

Step 2: Match the Caller to the Data Characteristics

Once sample characteristics are documented, select the variant caller based on three primary factors: whether a matched normal is available, the sequencing platform, and the expected VAF range.

Paired tumor-normal short-read data. When matched normal samples are available, paired analysis remains the preferred approach because comparing tumor and normal sequencing data allows the caller to subtract germline variants and focus on mutations unique to the tumor. Statistical and heuristic methods have traditionally dominated this space, but machine learning approaches trained on experimentally confirmed variant data have demonstrated improved sensitivity, particularly in challenging genomic regions [<a href="#ref-2">2</a>].

Tumor-only short-read data. When matched normal samples are unavailable, use callers specifically designed for this scenario. Deep learning frameworks trained using large numbers of high-confidence variants can accurately identify somatic variants from aligned tumor reads without a matched normal sample [<a href="#ref-1">1</a>]. These approaches have demonstrated significant performance improvements over existing methods, with some showing substantially higher accuracy in tumor mutation burden classification [<a href="#ref-1">1</a>].

Long-read data. Long-read sequencing requires specialized variant callers because most existing somatic variant callers were designed for short-read sequencing data and do not perform well with long-read data [<a href="#ref-4">4</a>]. Deep learning methods developed specifically for long-read tumor-normal pairs have demonstrated robust performance across varied coverage, purity, and contamination levels [<a href="#ref-4">4</a>]. For tumor-only long-read analysis, methods using ensembles of neural networks trained for opposite tasks have outperformed short-read callers applied to long-read data [<a href="#ref-5">5</a>].

Expected VAF below 10%. For samples where true variants are expected at very low allele fractions, prioritize callers with demonstrated sensitivity at low VAF. Machine learning approaches trained on synthetic somatic variants with diverse coverages and variant allele fractions can accurately detect a wide range of somatic variants [<a href="#ref-4">4</a>]. Some callers show particularly pronounced performance in genomic regions characterized by high sequencing error rates, achieving higher F1 scores than traditional callers in these contexts [<a href="#ref-2">2</a>].

Step 3: Determine Whether Additional Sequencing Is Needed

Before proceeding with analysis, assess whether the existing data are sufficient for the research question. This decision should be based on the relationship between expected VAF, sequencing depth, and the error rate of the sequencing platform.

Calculate the expected read support. For a variant at VAF of 5% in a sample sequenced to 100X mean depth, the expected number of supporting reads is approximately five. The confidence in this call depends on the actual depth at that specific position instead of the mean across the sample. If the expected read support is below the detection limit of the chosen caller, additional sequencing may be necessary.

Consider error-correcting approaches. For samples where true VAFs approach the sequencing error rate, error-correcting sequencing methods using unique molecular identifiers can distinguish true variants from polymerase and sequencing errors. These approaches effectively lower the noise floor, enabling detection of variants at allele fractions that would otherwise be indistinguishable from artifacts.

Weigh cost against value. The decision to re-sequence should weigh the cost of additional sequencing against the value of improved data. For samples where the research question cannot be answered with existing data quality, re-sequencing may be the most efficient path forward. Clinical whole-exome sequencing assays validated for somatic variant detection typically require substantial depth to achieve acceptable sensitivity, with documented performance characteristics across defined VAF thresholds and depth requirements [<a href="#ref-3">3</a>].

Step 4: Apply a Tiered Filtering Strategy

Filtering is essential for managing false positives, particularly in low-purity or low-coverage samples. A tiered approach balances sensitivity and specificity based on the intended use of the results.

Tier 1: Technical filters. Apply standard filters including minimum depth, minimum alternative allele count, strand bias, and mapping quality. These filters remove obvious artifacts and should be applied consistently across all samples. Document the specific thresholds used and the rationale for each choice.

Tier 2: Context-specific filters. Apply filters tailored to the sample type and genomic context. For FFPE samples, consider filters that address deamination artifacts. For low-complexity and repetitive regions, apply additional scrutiny because these regions generate alignment artifacts that mimic true variants [<a href="#ref-2">2</a>]. Machine learning callers often incorporate these considerations into their models, but additional post-hoc filtering may still be necessary.

Tier 3: Population frequency filters. For tumor-only analyses, apply population frequency filters to remove common germline variants. These filters use databases of known polymorphisms to exclude variants that are unlikely to be somatic. However, rare germline variants remain problematic, and this limitation should be documented.

Tier 4: Validation-focused filtering. For variants that will drive downstream decisions, apply the most stringent filters and consider orthogonal validation. Targeted deep sequencing can confirm predicted variants and distinguish true mutations from artifacts [<a href="#ref-2">2</a>]. The validation approach should be appropriate for the VAF of the variant, with sufficient sensitivity to detect the variant at the expected allele fraction.

Step 5: Document Decisions and Outcomes in a Structured Record

Maintain a structured record for each sample that captures the decision framework inputs, the analysis choices, and the outcomes. This record serves multiple purposes: it supports reproducibility, enables troubleshooting when problems arise, and provides documentation for publications or clinical reports.

Sample identification and metadata. Record the sample identifier, tumor type, tissue source, and any relevant clinical information. Document the estimated tumor purity and the method used to obtain this estimate.

Sequencing information. Record the platform, sequencing chemistry, mean depth, uniformity metrics, and any known technical issues. For targeted panels, document the panel design and the on-target rate.

Analysis decisions. Document the variant caller or callers used, the version number, and the specific parameters or thresholds applied. Record whether a matched normal was used and the quality of that sample. Note any deviations from standard workflows and the rationale for those deviations.

Quality control metrics. Track the number of variants called, the VAF distribution, and the proportion of variants passing each filtering step. Document any validation results and the concordance between initial calls and validation outcomes.

Interpretation notes. Record any observations about unusual variant patterns, problematic genomic regions, or sample-specific issues that may affect interpretation. These notes provide context for future analyses and help identify systematic issues with specific sample types or sequencing batches.

Troubleshooting Method for Unexpected Variant Calling Results

When variant calling results deviate from expectations, use a systematic troubleshooting approach to identify the cause. The following method addresses the most common failure patterns in low-purity and low-coverage samples.

Step A: Verify data quality metrics. Before investigating variant calls, confirm that the underlying sequencing data meet quality standards. Check mean depth, uniformity of coverage, duplication rate, and error rates. A sudden change in variant yield or VAF distribution may indicate technical problems requiring investigation.

Step B: Examine the VAF distribution. Plot the VAF distribution of called variants. True somatic mutations in a clonal tumor population should cluster around the expected VAF based on purity. An unexpected distribution may indicate contamination, tumor heterogeneity, or technical artifacts. For tumor-only analyses, a large number of variants at approximately 50% VAF may indicate germline variants that were not properly filtered.

Step C: Assess contamination levels. Sample contamination introduces alleles from other individuals, creating false positive somatic calls. This is particularly problematic for tumor-only analyses where contamination cannot be distinguished from true somatic variation without additional information. Assess contamination levels using available tools and exclude or flag samples with significant contamination.

Step D: Review variants in problematic genomic regions. Pay particular attention to variants in homopolymer runs, microsatellites, and segmental duplications. These are frequently artifacts, and validation should be considered before investing resources in follow-up studies [<a href="#ref-2">2</a>].

Step E: Compare across callers. If multiple callers were used, compare the variant calls and examine the discordant set. Variants called by only one caller may represent artifacts or may be genuine variants missed by the other caller due to different algorithmic strengths. Consensus approaches can improve confidence in detected variants, but strict intersection may reduce sensitivity.

Step F: Validate representative variants. For troubleshooting purposes, validate a small set of representative variants using an independent method. This provides information about the overall false positive rate and helps determine whether the calling strategy is appropriate for the sample type.

When to Escalate to Specialized Expertise

Researchers should consider consulting bioinformatics specialists or clinical genomics professionals when facing specific challenging situations. These include samples with extremely low purity where standard approaches fail, unusual variant patterns suggesting technical issues, or clinical decisions that depend on low-confidence variant calls.

Bioinformatics training resources from established organizations can help build the skills needed to address challenging samples. Foundational training in computing, data analysis, and genomic methods provides the basis for understanding and troubleshooting variant calling issues [<a href="#ref-6">6</a>][<a href="#ref-7">7</a>]. The Galaxy Training Network offers accessible workflow training and analysis tutorials that can help researchers develop practical skills [<a href="#ref-8">8</a>]. The Carpentries provides foundational lessons in computing, data, shell, Git, and programming that support reproducible analysis practices [<a href="#ref-7">7</a>].

For researchers using community pipelines, the nf-core documentation provides standards for pipeline usage, configuration, and reproducible workflow context [<a href="#ref-9">9</a>]. Bioconductor offers official package, workflow, installation, and reproducible genomic-analysis documentation that can support custom analysis development [<a href="#ref-10">10</a>]. The EMBL-EBI Training program provides bioinformatics learning pathways and data-resource training for practical analysis education [<a href="#ref-6">6</a>].

Integrating the Decision Framework with Existing Workflows

The decision framework described in this section is designed to complement existing variant calling workflows instead of replace them. The framework provides a structured approach to decision-making that can be adapted to local protocols and computational resources.

For laboratories with established workflows, the framework can be used as a checklist to ensure that all relevant factors are considered before analysis. For researchers developing new workflows, the framework provides a logical structure for selecting tools and parameters based on sample characteristics.

The record system described in this section supports reproducibility and troubleshooting. By documenting decisions and outcomes, researchers can identify systematic issues and refine their approaches over time. This continuous improvement process is particularly valuable for challenging samples where standard approaches may not perform optimally.

The troubleshooting method provides a practical approach to diagnosing problems when variant calling results deviate from expectations. By following the steps in order, researchers can identify the most likely causes of unexpected results and take appropriate corrective action.

Frequently Asked Questions

What is the minimum tumor purity for reliable somatic variant calling?

There is no universal minimum purity threshold because detection capability depends on sequencing depth, error rate, and the specific caller used. Higher depth can compensate for lower purity by providing more observations of low VAF variants. Machine learning callers trained on diverse purity levels can detect variants in samples with very low tumor content, but sensitivity decreases as purity drops [<a href="#ref-4">4</a>][<a href="#ref-5">5</a>]. For clinical applications, validated assays typically document sensitivity across defined purity and VAF ranges, and samples below these thresholds should be interpreted with caution [<a href="#ref-3">3</a>].

How does sequencing depth affect the detection of low VAF variants?

Sequencing depth directly determines the number of observations available at each position. Higher depth provides more statistical power to distinguish true variants from sequencing errors. At 50X depth, a 5% VAF variant is supported by approximately two to three reads, which is difficult to distinguish from error. At 200X depth, the same variant is supported by approximately ten reads, providing much stronger evidence. The relationship between depth and detection sensitivity is nonlinear, and the marginal benefit of additional depth diminishes at very high coverage.

Can tumor-only sequencing reliably identify somatic variants?

Tumor-only sequencing can identify somatic variants, but with important limitations. Without a matched normal sample, distinguishing somatic mutations from rare germline variants is challenging. Deep learning methods designed for tumor-only analysis have demonstrated significant performance improvements, with some approaches showing substantially higher accuracy in tumor mutation burden classification compared to earlier methods [<a href="#ref-1">1</a>]. However, tumor-only calling remains less accurate than paired analysis, and validation of critical variants is recommended.

What are the advantages of long-read sequencing for somatic variant calling?

Long-read sequencing provides improved read phasing, which enables more accurate detection of single-nucleotide variations at low variant allele fractions [<a href="#ref-4">4</a>]. The ability to span longer genomic regions improves alignment in repetitive and structurally complex regions. Deep learning callers designed for long-read data have demonstrated robust performance across varied coverage, purity, and contamination levels [<a href="#ref-4">4</a>][<a href="#ref-5">5</a>]. However, long-read sequencing requires specialized analysis tools, and the per-base cost is generally higher than short-read sequencing.

How do machine learning callers improve variant detection in challenging samples?

Machine learning callers learn complex patterns that distinguish true variants from artifacts by training on experimentally confirmed variant data [<a href="#ref-2">2</a>]. These models can incorporate sequence context, read characteristics, and genomic features that traditional statistical methods handle poorly. In genomic regions with high sequencing error rates, machine learning approaches have demonstrated higher F1 scores than traditional callers [<a href="#ref-2">2</a>]. Active learning strategies that incorporate experimental validation into training data can continuously improve performance [<a href="#ref-2">2</a>].

What is the role of error-correcting sequencing in low-purity samples?

Error-correcting sequencing approaches use molecular barcoding or similar strategies to distinguish true variants from errors introduced during library preparation and sequencing. By grouping reads derived from the same original DNA molecule, these methods can identify and correct technical errors. This effectively lowers the noise floor, enabling detection of variants at allele fractions that would otherwise be indistinguishable from artifacts. Error correction is particularly valuable for low-purity samples where true VAFs approach the sequencing error rate.

How should researchers validate variants detected at low VAF?

Variants detected at low VAF should be validated using an independent method with sufficient sensitivity to detect the variant at the expected allele fraction. Targeted deep sequencing of candidate variants is a common approach, providing high depth at specific loci to confirm or refute initial calls [<a href="#ref-2">2</a>]. The validation approach should account for the VAF of the variant, and the depth of validation sequencing should be sufficient to detect the variant with confidence.

What quality control metrics are most important for somatic variant calling?

Key quality control metrics include mean depth, uniformity of coverage, on-target rate for targeted panels, duplication rate, and sequencing error rates. For low-purity samples, the effective depth at variant sites is more relevant than mean depth across the sample. Sample-specific metrics such as estimated tumor purity and contamination levels also influence interpretation. Regular analysis of reference materials with known variants helps ensure consistent performance over time [<a href="#ref-3">3</a>].

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

[1] [Improved tumor-only variant calling and mutation burden estimation with VarNet-T.](https://doi.org/10.1038/s41467-026-71705-4). 2026. [2] [VariantMedium: sensitive and generalizable somatic point mutation calling with 3D DenseNets trained and evaluated on experimental data.](https://doi.org/10.1186/s13073-026-01675-1). 2026. [3] [Clinical validation of a high-performance somatic exome sequencing assay: from target-enrichment strategy to variant calling.](https://doi.org/10.1038/s41525-026-00569-w). 2026. [4] [ClairS: a deep-learning method for long-read tumor-normal pair somatic small variant calling.](https://pubmed.ncbi.nlm.nih.gov/42387002). Nature methods, 2026. [5] [ClairS-TO: a deep-learning method for long-read tumor-only somatic small variant calling.](https://doi.org/10.1038/s41467-025-64547-z). 2025. [6] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute. [7] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries. [8] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project. [9] [nf-core Documentation](https://nf-co.re/docs). nf-core. [10] [Bioconductor](https://bioconductor.org/). Bioconductor Project.

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.