# Filtering Somatic Variants: How to Reduce False Positives in Low-VAF Calls


## Key Takeaways

- Low-variant allele frequency (VAF) somatic variant calling (<5%) is challenged by technical artifacts (sequencing errors, alignment mismapping, PCR duplication, oxidative damage) that mimic true biological subclones, necessitating structured filtering.
- Matched normal sample subtraction is the most powerful filter, removing germline variants and systematic artifacts by comparing tumor and normal sequencing data; a panel of normals serves as an artifact baseline when matched normals are unavailable.
- Orientation bias detection, by assessing strand-specific read distribution, effectively removes artifacts arising from PCR errors and oxidative damage (e.g., G>T transversions), which often exhibit strand asymmetry.
- Tumor purity adjustment is critical as low purity compresses VAF, pushing true low-frequency variants into the artifact-prone zone; copy number alterations also influence expected VAF and require incorporation into filtering logic.
- Coverage and read quality thresholds, including minimum variant-supporting reads, total depth, base quality, and mapping quality, are essential to ensure calls are supported by sufficient, high-quality evidence.
- Annotation-based filtering, using functional impact predictions and population frequency databases, prioritizes or deprioritizes real variants but should be applied *after* technical filters to avoid removing true somatic mutations.

---

Somatic variant calling from next-generation sequencing data frequently produces false positive calls, particularly when variant allele frequency (VAF) falls below 5%. This article presents a structured filtering approach that combines panel of normals subtraction, orientation bias detection, tumor purity adjustment, and coverage-based quality thresholds to reduce spurious calls while preserving true low-frequency variants. The workflow applies to targeted sequencing panels, whole exome sequencing, and whole genome sequencing data from tumor samples with or without matched normals.

## The False Positive Problem in Low-VAF Somatic Calling

Low-VAF variants represent true biological subclones, but they also represent the zone where technical artifacts concentrate. Sequencing errors, alignment mismapping, PCR duplication, and oxidative damage during library preparation all generate noise that mimics genuine mutations at allele frequencies below 5%. The challenge is that true subclonal mutations and technical artifacts occupy the same VAF range, making separation difficult without deliberate filtering strategies.

Targeted sequencing approaches offer higher depth that improves low-frequency variant detection, yet the transition to routine clinical use has been slow because technical and analytical obstacles remain unresolved. Gold-standard procedures and pipelines are urgently needed to accelerate this transition. The core problem is not sequencing depth alone but the analytical filters applied after alignment and initial variant calling.

False positives in somatic calling arise from several distinct mechanisms. Sequencing platform errors occur at characteristic rates per base. Alignment artifacts emerge when reads map to multiple genomic locations or when repetitive regions create ambiguous placements. PCR amplification introduces errors during library construction, and these errors become amplified alongside true variants. Formalin-fixed paraffin-embedded samples contribute deamination artifacts that convert cytosine to thymine. Each error type requires a different filtering strategy.

The practical consequence of inadequate filtering is wasted validation effort. Researchers who confirm candidate variants by orthogonal methods such as Sanger sequencing or targeted deep sequencing spend time and resources chasing artifacts. In clinical contexts, false positives can lead to incorrect treatment decisions. The filtering workflow described here reduces this burden by removing likely artifacts before validation.

## Core Principles of Somatic Variant Filtering

### Variant Allele Frequency as a Confidence Signal

VAF alone does not distinguish true from false variants. A true somatic variant at 2% VAF and a sequencing artifact at 2% VAF are indistinguishable by frequency alone. However, VAF interacts with depth to establish a confidence boundary. At a given depth, the probability that a variant call represents a real biological allele instead of random sequencing error depends on the number of independent reads supporting the alternate allele.

High depth is the foundation of low-VAF calling. Targeted sequencing achieves greater sequencing depth with reduced costs and data burden compared to whole genome or whole exome approaches, which allows targeted sequencing to identify low frequency variants in targeted regions with high confidence. This depth advantage makes targeted panels suitable for profiling low-quality and fragmented clinical DNA samples. The practical implication is that filtering thresholds must be adjusted according to the sequencing platform and depth profile of each sample.

### The Role of Matched Normal Samples

A matched normal sample from the same patient provides the most powerful filter for somatic variant calling. Germline variants appear in both tumor and normal samples. Somatic variants appear only in the tumor. Subtracting variants present in the normal sample removes germline polymorphisms and inherited variants that would otherwise appear as false positive somatic calls.

The matched normal also controls for sequencing artifacts that occur systematically across samples. If a variant appears in both tumor and normal at similar allele frequencies, it likely represents a sequencing or alignment artifact instead of a true somatic event. This is particularly important for low-VAF calls where the distinction between germline heterozygosity, germline mosaicism, and somatic mutation becomes blurred.

When matched normals are unavailable, filtering becomes more difficult. Population databases can remove common germline polymorphisms, but rare germline variants and private polymorphisms remain. The panel of normals approach partially addresses this limitation by aggregating sequencing artifacts across many normal samples processed through the same pipeline.

### Panel of Normals as an Artifact Baseline

A panel of normals is a collection of sequencing data from non-tumor samples processed through the same library preparation, sequencing, and alignment pipeline as the tumor samples being analyzed. The panel captures the systematic error profile of the entire workflow. Variants that appear recurrently in the panel are likely artifacts instead of true somatic events.

The panel of normals serves two functions. First, it identifies sites where the sequencing or alignment pipeline consistently produces false calls. These sites can be masked or filtered in tumor samples. Second, it establishes a background noise model that can be used to calculate the probability that a variant call in a tumor sample represents real biology instead of pipeline noise.

Building a useful panel of normals requires attention to sample size and diversity. The panel should include enough samples to capture recurrent artifacts. It should represent the same tissue types, library preparation methods, and sequencing platforms as the tumor samples. A panel built from one tissue type may not adequately control for artifacts in another tissue type.

## The Multi-Step Filtering Workflow

### Step 1: Initial Variant Calling and Raw Candidate Generation

The filtering workflow begins with an initial variant calling step. The choice of caller depends on the sequencing platform and analysis goals. Statistical and heuristic methods have been proposed for somatic single nucleotide variant calling from matched tumor-normal data, but they suffer from low sensitivity especially in certain genomic regions. Machine learning approaches have been developed to address these sensitivity gaps.

For targeted panel data, the variant caller should be configured to output all candidate variants without aggressive filtering. The goal at this stage is recall instead of precision. Aggressive filtering at the caller level removes true low-VAF variants before downstream filters have a chance to evaluate them. The caller should output variant calls with supporting evidence including read counts, base qualities, mapping qualities, and strand information.

The initial call set will contain a mixture of true somatic variants, germline variants, and technical artifacts. The proportion of artifacts increases as the VAF threshold decreases. At VAF above 10%, most calls are likely true. Below 5%, artifacts may outnumber true variants. The filtering steps that follow are designed to separate these categories.

### Step 2: Matched Normal Subtraction

When a matched normal is available, the first filtering step is subtraction of variants present in the normal sample. This step removes germline variants and systematic artifacts. The comparison should consider also presence or absence but also allele frequency. A variant present in the tumor at 30% VAF and in the normal at 30% VAF is clearly germline. A variant present in the tumor at 30% VAF and in the normal at 1% VAF requires careful interpretation.

The normal sample may contain low-level artifacts that coincidentally match a true somatic variant in the tumor. The filtering threshold for normal subtraction should account for this possibility. A variant present in the normal at any detectable level may warrant exclusion, or a threshold such as 2% or 3% VAF in the normal may be used. The choice depends on the depth of the normal sample and the error profile of the sequencing platform.

Matched normal subtraction also addresses clonal hematopoiesis of indeterminate potential. Older patients frequently harbor somatic mutations in blood cells that appear in normal blood samples. These mutations can be mistaken for tumor-specific variants if the normal sample is not carefully evaluated. The presence of a variant in the normal at low VAF may indicate clonal hematopoiesis instead of a tumor-specific event.

### Step 3: Panel of Normals Filtering

After matched normal subtraction, the remaining candidate variants are filtered against the panel of normals. This step removes artifacts that recur across multiple normal samples. The filtering logic is that a variant appearing in multiple normal samples at similar allele frequencies is unlikely to represent a true somatic event in the tumor.

The panel of normals filter can be applied in several ways. A simple approach is to exclude any variant that appears in more than a specified number of panel samples. A more sophisticated approach uses the panel to model the expected error rate at each site and calculates a probability that the tumor variant represents real biology.

The panel of normals is particularly valuable for targeted sequencing panels. The same genomic regions are sequenced in every sample, so the panel accumulates substantial evidence about which sites are prone to artifacts. This information is specific to the panel design, library preparation method, and sequencing platform. A panel of normals built for one assay may not transfer to another assay.

### Step 4: Orientation Bias and Strand Artifact Filtering

Orientation bias refers to the tendency of certain artifacts to appear predominantly on one sequencing strand. True somatic variants should be supported by reads from both the forward and reverse strands. Artifacts from PCR errors, oxidative damage, and other sources often show strand asymmetry.

The orientation bias filter evaluates the distribution of supporting reads across strands. A variant supported by 20 reads all on the forward strand is suspicious. A variant supported by 10 forward and 10 reverse reads is more credible. The filter can be implemented as a Fisher exact test comparing the strand distribution of variant-supporting reads to the strand distribution of reference-supporting reads.

Oxidative damage during library preparation creates a specific artifact pattern. Guanine oxidation produces 8-oxoguanine, which mispairs with adenine during PCR. This creates G to T transversions that appear predominantly on one strand. The orientation bias filter is particularly effective at removing these artifacts because the damage occurs before library amplification and therefore appears on only one strand.

### Step 5: Tumor Purity and Copy Number Adjustment

Tumor purity, the proportion of tumor cells in the sequenced sample, directly affects the expected VAF of true somatic variants. A heterozygous somatic mutation in a pure tumor sample has an expected VAF of 50%. The same mutation in a sample with 20% tumor purity has an expected VAF of 10%. Low tumor purity compresses the VAF range and pushes true variants into the low-VAF zone where artifacts concentrate.

Tumor purity adjustment involves estimating the purity of each sample and interpreting VAF in that context. A variant at 5% VAF in a sample with 10% purity may represent a clonal event present in most tumor cells. The same variant at 5% VAF in a sample with 80% purity represents a minor subclone. The filtering threshold should be adjusted accordingly.

Copy number alterations also affect expected VAF. A heterozygous mutation in a region with copy number loss has a different expected VAF than the same mutation in a region with copy number gain. Segmentation-based detection of copy number alterations from whole exome sequencing data enables accurate determination of these effects. The expected VAF calculation should incorporate local copy number state.

### Step 6: Coverage and Read Quality Thresholds

Minimum depth thresholds remove calls supported by insufficient evidence. A variant call at 1% VAF supported by 3 reads out of 300 total reads is more credible than the same VAF supported by 1 read out of 100. The filtering threshold should require a minimum number of variant-supporting reads and a minimum total depth at the site.

Base quality filtering removes calls supported by low-quality base calls. The variant-supporting reads should have adequate base quality scores at the variant position. Mapping quality filtering removes calls from reads that align ambiguously. Reads with low mapping quality may originate from elsewhere in the genome and support a false variant at the current location.

The specific thresholds depend on the sequencing platform and the desired sensitivity-specificity balance. Lower thresholds increase sensitivity but also increase false positives. Higher thresholds reduce false positives but may miss true low-VAF variants. The thresholds should be validated against known positive and negative controls.

### Step 7: Annotation-Based Filtering

Functional annotation can prioritize or filter variants based on their predicted impact. Variants in exonic regions, splice sites, or known cancer genes may be prioritized for further analysis. Variants in intergenic regions or known repetitive elements may be deprioritized. This filtering is not a substitute for technical filters but can help focus validation efforts.

Population frequency databases provide additional filtering information. Variants that appear at high frequency in population databases are likely germline polymorphisms instead of somatic mutations. However, rare germline variants and population-specific polymorphisms may not be captured in these databases. The absence of a variant from population databases does not guarantee that it is somatic.

The annotation step should be applied after technical filtering. Applying annotation filters first can remove true somatic variants that happen to fall in regions with low functional annotation confidence. The technical filters address the question of whether a variant is real. The annotation filters address the question of whether a real variant is biologically relevant.

## At a Glance

| Filter Step | Primary Artifact Removed | Input Required | Key Limitation |
| --- | --- | --- | --- |
| Matched Normal Subtraction | Germline variants and systematic artifacts | Matched normal sequencing data | Normal may contain clonal hematopoiesis artifacts |
| Panel of Normals Filtering | Recurrent pipeline artifacts | Panel of normal samples from same assay | Panel must match tumor assay conditions |
| Orientation Bias Filtering | PCR and oxidative damage artifacts | Strand-specific read counts | May remove true variants in low-complexity regions |
| Tumor Purity Adjustment | VAF compression from normal cell contamination | Tumor purity estimate and copy number data | Purity estimation introduces its own uncertainty |

## Practical Implementation Steps

### Building the Filtering Pipeline

The filtering pipeline should be implemented as a reproducible workflow. Community pipeline standards provide structured approaches for building and configuring reproducible analysis workflows. The workflow should accept aligned tumor and normal BAM files as input and produce a filtered variant call set as output.

The pipeline should be version controlled and documented. Each filtering step should be configurable through a parameter file. The parameters should be recorded in the output so that downstream users know exactly which filters were applied. This transparency is essential for interpreting results and for comparing results across samples or studies.

Containerization or virtual environments ensure that the pipeline runs consistently across different computing environments. The software dependencies should be pinned to specific versions. Changes to the pipeline should be tracked through version control, and the version used for each analysis should be recorded.

### Selecting Filter Thresholds

Filter thresholds should be selected based on the specific assay and research question. A discovery study looking for novel mutations may use permissive thresholds to maximize sensitivity. A clinical validation study may use stringent thresholds to maximize specificity. The thresholds should be documented and justified.

Validation using known positive and negative controls is essential. Positive controls are samples with known somatic variants at various allele frequencies. Negative controls are samples without known somatic variants or samples where the expected variants are absent. The filtering thresholds should be adjusted until the pipeline correctly identifies the positive controls and rejects the negative controls.

The validation set should include variants across the VAF range of interest. A pipeline that performs well at 10% VAF may perform poorly at 2% VAF. The validation should specifically test the low-VAF range where filtering is most challenging.

### Recording Filtering Decisions

Each filtering step should record the number of variants removed and the number retained. This information provides a quality metric for the overall pipeline. A pipeline that removes 99% of initial calls may be overly aggressive. A pipeline that removes only 10% may be insufficiently stringent.

The filtering decisions should be traceable. For each variant that is removed, the pipeline should record which filter removed it and the evidence supporting that decision. This traceability is essential for troubleshooting and for understanding why specific variants were excluded.

The final filtered call set should include the evidence supporting each retained variant. This evidence includes read counts, allele frequencies, strand distributions, and the results of each filtering step. This information allows downstream users to evaluate the confidence of each call.

## Options and Tradeoffs in Filtering Strategy

### Matched Normal Versus Panel of Normals Only

The matched normal approach provides the strongest filtering but requires additional sequencing costs and may not be feasible for all samples. The panel of normals approach is less powerful but can be applied to samples without matched normals. Some studies use both approaches, applying matched normal subtraction when available and panel of normals filtering to all samples.

The choice between matched normal and panel of normals depends on the study design and available resources. Clinical studies often require matched normals for accurate somatic calling. Research studies with limited budgets may rely on panel of normals filtering alone. The tradeoff is between filtering power and sequencing cost.

### Stringent Versus Permissive Thresholds

Stringent thresholds reduce false positives but also reduce sensitivity for true low-VAF variants. Permissive thresholds increase sensitivity but require more downstream validation effort. The optimal balance depends on the downstream application.

For clinical applications where false positives have direct patient consequences, stringent thresholds are appropriate. For discovery research where the goal is to identify candidate variants for further study, permissive thresholds may be more appropriate. The validation capacity of the laboratory should inform the threshold selection.

### Single Caller Versus Ensemble Approaches

Different variant callers have different strengths and weaknesses. Some callers perform well in specific genomic regions while others perform poorly. Ensemble approaches that combine calls from multiple callers can improve overall performance.

Machine learning approaches trained on experimentally confirmed variant data have shown improved sensitivity compared to traditional callers, particularly in genomic regions characterized by high sequencing error rates. These approaches combine features from multiple sources to distinguish true variants from artifacts. The training data quality is critical for the performance of these methods.

The choice between a single caller and an ensemble approach depends on the computational resources available and the required accuracy. Ensemble approaches require more computation but may provide more reliable results. The performance of each approach should be validated on the specific assay and sample type being analyzed.

## Observations and Measurements

### Concordance Across Laboratories and Platforms

Multicenter evaluations of targeted sequencing panels have shown that concordance between laboratories decreases for low-VAF variants. In one European multicenter evaluation of three amplicon-based assays for chronic lymphocytic leukemia, the panels achieved 90% to 97.7% concordance at variant allele frequency above 0.5% after filtering for variants present in the paired normal sample and removal of PCR and sequencing artifacts. However, 6 of 8 variants undetected by a single center concerned minor subclonal mutations with VAF below 5%.

This observation has direct implications for filtering strategy. The filtering steps that remove PCR and sequencing artifacts are essential for achieving acceptable concordance. Without these filters, the concordance between laboratories would be substantially lower. The study also found that a high-sensitivity assay containing unique molecular identifiers confirmed the presence of several minor subclonal mutations, suggesting that some low-VAF variants are real but require more sensitive methods for reliable detection.

### Whole Exome Sequencing Variability

A multicenter pilot study of clinical whole exome sequencing for cancer patients found that calling of somatic variants was highly concordant with a positive percentage agreement between 91% and 95% and a positive predictive value between 82% and 95% compared with a three-institution consensus. Full agreement was achieved for 16 of 17 druggable targets. Explanations for deviations included low VAF or coverage, differing annotations, and different filter protocols.

The finding that different filter protocols contribute to inter-institution variability is important. Even when laboratories use the same sequencing platform and similar analysis pipelines, differences in filtering thresholds and criteria produce different results. This observation supports the need for standardized filtering protocols and transparent documentation of filtering decisions.

### Long-Read Sequencing Considerations

Long-read sequencing platforms have different error profiles than short-read platforms. A variant caller developed for Nanopore panel sequencing data uses allele probability distributions and several other filters to robustly separate true from false positive calls. This caller effectively calls single nucleotide variants and insertions or deletions with variant allele frequencies as low as 1% and 5% respectively, producing only few low-frequency false-positive calls.

The filtering principles described in this article apply to long-read data, but the specific thresholds and implementations differ. The error profile of long-read platforms is different from short-read platforms, and the filters must be adjusted accordingly. The orientation bias filter may be less relevant for long-read data where the error mechanisms differ.

## Records and Measurements

### Essential Records for Filtering Validation

The following records should be maintained for each somatic variant calling analysis:

| Record Type | Content | Purpose |
| --- | --- | --- |
| Pipeline Version | Software versions and parameter settings | Reproducibility and troubleshooting |
| Filtering Log | Variant counts before and after each filter | Quality assessment and optimization |
| Validation Results | Performance on positive and negative controls | Threshold selection and quality assurance |
| Sample Metadata | Tumor purity, coverage, sequencing platform | Context for interpreting filtering decisions |

### Quality Metrics to Track

The transition of next-generation sequencing from research to clinical routine requires rigorous quality assessment. The comparability of different technologies for mutation profiling remains an open question, and quality metrics provide the basis for comparison.

The primary quality metric for somatic variant calling is the balance between sensitivity and specificity. Sensitivity measures the proportion of true variants detected. Specificity measures the proportion of called variants that are true. Both metrics should be evaluated across the VAF range of interest.

The positive predictive value is particularly important for clinical applications. A high positive predictive value means that most called variants are true, reducing the need for confirmatory testing. The positive predictive value should be evaluated at different VAF thresholds to understand how filtering stringency affects the reliability of calls.

## Common Failure Patterns

### Overly Aggressive Filtering

The most common failure pattern is applying filters that are too stringent, removing true low-VAF variants along with artifacts. This failure is particularly problematic because it is invisible. The researcher does not know what was missed. The filtered call set looks clean, but it lacks true biological signal.

Overly aggressive filtering often results from applying thresholds validated on one dataset to another dataset with different characteristics. A threshold that works well for high-purity fresh frozen samples may remove true variants from low-purity formalin-fixed samples. The filtering thresholds should be validated on the specific sample type being analyzed.

### Under-Filtering and Validation Overload

The opposite failure pattern is insufficient filtering, which produces a large call set dominated by artifacts. The downstream validation effort becomes overwhelming. Researchers spend time and resources confirming variants that are technical artifacts instead of biological events.

Under-filtering is common when researchers are concerned about missing true variants. The concern is legitimate, but the solution is not to eliminate filtering. The solution is to apply appropriate filters and to validate the filtering performance on known controls.

### Ignoring Tumor Purity

Failure to account for tumor purity is a common source of filtering errors. A sample with low tumor purity has compressed VAF values. True clonal variants appear at low VAF and may be filtered as artifacts. The filtering thresholds should be adjusted based on the tumor purity of each sample.

Tumor purity estimation itself introduces uncertainty. Different estimation methods may produce different purity values for the same sample. The filtering pipeline should be robust to reasonable variation in purity estimates.

### Applying Filters in the Wrong Order

The order of filtering steps matters. Applying annotation-based filters before technical filters can remove true variants that would have passed the technical filters. Applying the panel of normals filter before matched normal subtraction can remove variants that would have been retained after normal subtraction.

The recommended order is to apply technical filters first, then annotation filters. The technical filters address the question of whether a variant is real. The annotation filters address the question of whether a real variant is biologically relevant. This order preserves the maximum number of true variants while removing artifacts.

## Limitations of Filtering Approaches

### Residual False Positives

No filtering approach eliminates all false positives. Even with matched normal subtraction, panel of normals filtering, orientation bias filtering, and tumor purity adjustment, some artifacts will pass all filters. The residual false positive rate depends on the sequencing platform, sample quality, and the specific filters applied.

The residual false positive rate should be estimated using negative controls. A negative control is a sample known to lack somatic variants in the targeted regions. Any variants called in the negative control represent false positives. The false positive rate should be measured at different VAF thresholds to understand the reliability of low-VAF calls.

### Loss of True Low-VAF Variants

The filtering steps that remove artifacts also remove some true variants. This loss is unavoidable. The goal is to minimize the loss of true variants while maximizing the removal of artifacts. The balance between sensitivity and specificity should be explicitly evaluated and documented.

The loss of true low-VAF variants is particularly concerning for clinical applications where the presence of a specific mutation guides treatment decisions. A false negative result can lead to inappropriate treatment. The filtering thresholds should be validated to ensure that clinically significant variants are reliably detected.

### Platform-Specific Filter Requirements

The filtering approach described in this article is platform-agnostic in principle but requires platform-specific implementation. The error profile of each sequencing platform is different. The filters must be calibrated for the specific platform, library preparation method, and analysis pipeline.

The panel of normals is particularly platform-specific. A panel built from data generated on one sequencing platform may not be appropriate for data generated on another platform. The panel should be built from data generated using the same protocols as the tumor samples being analyzed.

### The Challenge of Low-Quality Samples

Clinical samples are often of lower quality than research samples. Formalin-fixed paraffin-embedded tissue samples have fragmented DNA and characteristic deamination artifacts. The filtering approach must account for these sample-specific artifacts.

The targeted sequencing approach is suitable for profiling low-quality and fragmented clinical DNA samples because the targeted regions are sequenced at high depth. However, the filtering thresholds may need to be adjusted for low-quality samples. The validation should include low-quality samples to ensure that the filtering approach performs adequately on the samples most likely to be encountered in clinical practice.

## Safety and Regulatory Context

### Clinical Validation Requirements

The use of somatic variant calling for clinical decision making requires rigorous validation. The validation should demonstrate that the filtering approach reliably detects clinically significant variants while minimizing false positives. The validation should include samples with known variants across the VAF range of interest.

Clinical whole exome sequencing studies have shown that inter-institution variability in variant calling can be attributed to low VAF or coverage, differing annotations, and different filter protocols. This finding supports the need for standardized filtering protocols in clinical applications. The standardization should include also the filtering steps but also the thresholds and the documentation requirements.

### Reporting Filtering Decisions

Clinical reports should include information about the filtering approach used. This information allows clinicians to interpret the results appropriately. A variant that passed stringent filters has different clinical significance than a variant that passed permissive filters.

The report should include the VAF of each reported variant, the depth at the variant site, and the results of key filtering steps. This information allows the clinician to assess the confidence of each call. The report should also include the tumor purity estimate and the expected VAF range for clonal variants.

### Professional Escalation Criteria

Certain findings warrant escalation to a molecular tumor board or other expert review. These include variants at the boundary of the validated VAF range, variants in genes with direct treatment implications, and variants where the filtering evidence is ambiguous.

The escalation criteria should be defined before the analysis begins. The criteria should be based on the clinical context and the validated performance of the filtering pipeline. The escalation process should be documented and followed consistently.

## Frequently Asked Questions

### What is the minimum VAF that can be reliably detected with targeted sequencing?

The reliable detection limit depends on sequencing depth, sample quality, and the filtering approach. Targeted sequencing approaches that achieve high depth can identify low frequency variants with high confidence. Multicenter evaluations have shown that amplicon-based approaches can be adopted for somatic mutation detection with VAF above 5% after rigorous validation. Detection below 5% VAF may require unique molecular identifiers or other high-sensitivity approaches to reach acceptable reliability.

### How many samples should be included in a panel of normals?

The panel of normals should include enough samples to capture recurrent artifacts in the assay. The required number depends on the artifact rate of the sequencing platform and the desired sensitivity for artifact detection. The panel should be built from samples processed through the same library preparation, sequencing, and alignment pipeline as the tumor samples. The panel should be evaluated to ensure that it adequately captures the artifact profile of the assay.

### Can somatic variants be called without a matched normal sample?

Somatic variants can be called without a matched normal sample, but the filtering approach must be adjusted. The panel of normals becomes the primary filter for removing germline variants and systematic artifacts. Population frequency databases can remove common germline polymorphisms. However, rare germline variants and private polymorphisms may be misclassified as somatic. The limitations of unmatched calling should be acknowledged in the interpretation of results.

### How does tumor purity affect the filtering thresholds?

Tumor purity determines the expected VAF of true somatic variants. A heterozygous mutation in a pure tumor sample has an expected VAF of 50%. The same mutation in a sample with 20% purity has an expected VAF of 10%. Low purity compresses the VAF range and pushes true variants into the low-VAF zone where artifacts concentrate. The filtering thresholds should be adjusted based on the tumor purity estimate for each sample.

### What is orientation bias and why does it matter for low-VAF calls?

Orientation bias refers to the tendency of certain artifacts to appear predominantly on one sequencing strand. True somatic variants should be supported by reads from both strands. Artifacts from PCR errors and oxidative damage often show strand asymmetry. The orientation bias filter evaluates the strand distribution of variant-supporting reads and removes calls with significant strand imbalance. This filter is particularly important for low-VAF calls where artifacts are more common.

### How should filtering thresholds be validated?

Filtering thresholds should be validated using known positive and negative controls. Positive controls are samples with known somatic variants at various allele frequencies. Negative controls are samples without known somatic variants. The validation should include variants across the VAF range of interest, with particular attention to the low-VAF range where filtering is most challenging. The validation results should be documented and used to justify the selected thresholds.

### What are unique molecular identifiers and when are they needed?

Unique molecular identifiers are random barcode sequences attached to individual DNA molecules before amplification. They allow the detection of PCR duplicates and the correction of PCR errors. Unique molecular identifiers may be necessary to reach higher sensitivity for low-frequency mutations. Multicenter evaluations have shown that assays containing unique molecular identifiers can confirm the presence of minor subclonal mutations that are missed by standard amplicon-based approaches.

### How should filtering decisions be documented for clinical reporting?

Filtering decisions should be documented in a way that allows downstream users to assess the confidence of each call. The documentation should include the pipeline version, the filtering parameters, and the variant counts before and after each filter. For each reported variant, the documentation should include the VAF, the depth at the variant site, and the results of key filtering steps. This transparency is essential for clinical interpretation and for comparing results across laboratories.

## Related Bioinformatics Guides

- [Detecting Structural Variants with Long-Read Sequencing: Methods and Considerations](/knowledge/bioinformatics/detecting-structural-variants-with-long-read-sequencing-methods-and-considerations)
- [From Raw Reads to Variants: A Diagnostic Blueprint for Next-Generation Sequencing (NGS) Workflows](/knowledge/bioinformatics/ngs-raw-reads-variant-calling-blueprint)
- [Genomic Data Analysis Tools: A Comparative Guide for Researchers](/knowledge/bioinformatics/genomic-data-analysis-tools-a-comparative-guide-for-researchers)
- [Metagenomics Assembly: Strategies for Reconstructing Microbial Genomes](/knowledge/bioinformatics/metagenomics-assembly-strategies-for-reconstructing-microbial-genomes)
- [Research Data Stewardship: Benefits and Implementation Strategies](/knowledge/bioinformatics/research-data-stewardship-benefits-and-implementation-strategies)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)


## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [Applications and analysis of targeted genomic sequencing in cancer studies.](https://pubmed.ncbi.nlm.nih.gov/31762958). Computational and structural biotechnology journal, 2019.
- [Nanopanel2 calls phased low-frequency variants in Nanopore panel sequencing data.](https://pubmed.ncbi.nlm.nih.gov/34270680). Bioinformatics (Oxford, England), 2021.
- [Comparative analysis of targeted next-generation sequencing panels for the detection of gene mutations in chronic lymphocytic leukemia: an ERIC multi-center study.](https://pubmed.ncbi.nlm.nih.gov/32273480). Haematologica, 2021.
- [Multicentric pilot study to standardize clinical whole exome sequencing (WES) for cancer patients.](https://pubmed.ncbi.nlm.nih.gov/37864096). NPJ precision oncology, 2023.
- [VariantMedium: sensitive and generalizable somatic point mutation calling with 3D DenseNets trained and evaluated on experimental data.](https://doi.org/10.1186/s13073-026-01675-1). 2026.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.