# Variant Allele Frequency (VAF) in Somatic Calling: What It Tells You About Tumor Purity, Ploidy, and Clonality


## Key Takeaways

- Variant Allele Frequency (VAF) quantifies the proportion of sequencing reads supporting a variant allele, serving as a critical metric in somatic variant calling by encoding information about tumor purity, copy number state, and clonal evolution.
- Tumor purity, the fraction of tumor cells within a sample, directly scales expected VAFs; a clonal heterozygous variant in a pure diploid tumor is expected at 50% VAF, while in a 50% pure sample, it drops to 25%, necessitating purity-adjusted interpretation of variant filtering thresholds.
- Ploidy and copy number alterations significantly impact VAF; a clonal variant at a locus with copy number gain (e.g., 3 copies, variant on 1) will yield a lower VAF (33%) than in a diploid state, requiring allele-specific copy number analysis to accurately determine the cancer cell fraction.
- Clonality is inferred from VAF distributions, where clonal variants cluster at high cancer cell fractions (near 1.0 after purity and copy number correction), while subclonal variants appear as lower frequency events reflecting later acquisition during tumor evolution.
- Reliable VAF interpretation mandates accounting for sequencing depth, variant caller specifics, and crucially, requires matched normal tissue to distinguish somatic from germline variants, as germline heterozygous variants can mimic somatic variants in tumor VAF distributions.
- The clinical utility of VAF as a predictive biomarker is limited by a lack of validated thresholds and assay standardization, though it holds prognostic relevance in specific contexts like JAK2 V617F VAF in polycythemia vera.

---

Variant allele frequency (VAF) is the proportion of sequencing reads that support a variant allele at a given genomic locus. In somatic variant calling, VAF is also a quality metric. It is a quantitative readout that encodes information about the cellular composition of the sample, the copy number state of the locus, and the evolutionary structure of the tumor cell population. For researchers interpreting targeted panel, exome, or whole-genome sequencing data, VAF provides the raw material for estimating tumor purity, detecting copy number alterations, and reconstructing clonal architecture. This article explains how to calculate VAF, how to adjust raw VAF values for purity and ploidy, and how to use VAF distributions to distinguish clonal from subclonal events. It also covers the practical decisions involved in variant filtering, the limitations of VAF-based inference, and the reporting considerations that apply when VAF informs clinical or research conclusions.

## The Biological Basis of VAF in Somatic Tissues

Somatic mutations arise during the lifetime of an organism and accumulate through cell divisions. Unlike germline variants, which are present in essentially every cell of the body, somatic mutations exist in only a subset of cells. The fraction of cells carrying a given somatic mutation depends on when the mutation occurred during development or tumor evolution and on the subsequent expansion of the mutant cell population. Early mutations that occur in a common progenitor will be shared by many descendant cells, while late mutations will be restricted to smaller branches of the cell lineage tree. This principle underlies the use of somatic mutations as endogenous barcodes for reconstructing developmental lineages and clonal dynamics, as described in a 2025 review of lineage tracing approaches that compare variant allele frequencies across bulk samples to infer lineage proximity [<a href="#ref-1">1</a>].

In cancer sequencing, the tumor sample is a mixture of tumor cells and non-tumor cells, including stromal cells, immune infiltrates, and normal epithelial cells. The observed VAF of a somatic variant therefore depends on three factors: the fraction of tumor cells in the sample (purity), the copy number state of the genomic region containing the variant, and the fraction of tumor cells that actually carry the variant (the cancer cell fraction). A variant present in all tumor cells at a diploid locus in a pure tumor sample would be expected to have a VAF of 50 percent, because each tumor cell contributes one variant allele and one reference allele. If the sample contains 50 percent normal cells, the expected VAF drops to 25 percent, because the normal cells contribute only reference alleles. If the variant is present in only a subset of tumor cells, the VAF drops further. These relationships are the foundation for using VAF to infer tumor purity and clonality.

The clinical relevance of VAF extends beyond basic research. In precision oncology, VAF has been evaluated as a predictive biomarker for targeted therapy selection, with the rationale that targeting the dominant cancer cell population, identified by its high VAF, may produce a more durable response. However, a 2023 review in Trends in Cancer notes that the absence of validated VAF thresholds and a lack of standardization between sequencing assays currently limit the clinical utility of VAF as a decision-making tool [<a href="#ref-2">2</a>]. This distinction between the biological information content of VAF and its validated clinical utility is important for researchers who may be tempted to apply VAF cutoffs in patient management decisions without appropriate analytical and clinical validation.

## Calculating VAF from Sequencing Data

VAF is calculated as the number of reads supporting the variant allele divided by the total number of reads covering the locus. If a locus has 80 reads total and 20 of those reads carry the alternate allele, the VAF is 0.25, or 25 percent. This calculation is straightforward in principle, but several technical factors affect the accuracy of the numerator and denominator.

Read depth is the primary determinant of VAF precision. At low depth, the confidence interval around a VAF estimate is wide. A VAF of 10 percent based on 20 total reads means only 2 variant reads were observed, and the true VAF could plausibly be much higher or lower. At higher depth, the estimate becomes more precise. Targeted sequencing panels designed for somatic mutation detection often aim for high depth, sometimes exceeding 1000-fold, to enable detection of low-VAF variants. A study of focal cortical dysplasia used targeted gene sequencing at 2000-fold or greater read depth to search for low-allele frequency variants in mTOR pathway genes, successfully identifying somatic variants that would have been missed at standard exome sequencing depth [<a href="#ref-3">3</a>]. This example illustrates the direct relationship between sequencing depth and the ability to detect and accurately quantify low-VAF somatic variants.

The choice of variant caller also influences VAF estimates. Different algorithms apply different base quality filters, mapping quality filters, and local realignment strategies, which can shift the ratio of variant to reference reads at a given locus. Researchers should be aware that VAF values are caller-dependent to some degree and should avoid over-interpreting small differences in VAF between variants called by different tools. Reproducible analysis pipelines, such as those provided by the nf-core community, standardize the variant calling workflow and reduce this source of variability [<a href="#ref-4">4</a>]. The Galaxy Training Network similarly offers accessible workflow training that emphasizes reproducible analysis practices for genomic data [<a href="#ref-5">5</a>].

## Tumor Purity and Its Effect on VAF

Tumor purity, also called tumor cellularity, is the proportion of cells in a sample that are tumor cells. Purity directly scales the expected VAF of somatic variants. For a clonal variant present in all tumor cells at a diploid locus, the expected VAF equals half the purity. If purity is 80 percent, the expected VAF is 40 percent. If purity is 30 percent, the expected VAF is 15 percent.

This scaling has practical consequences for variant filtering. A common filtering approach is to apply a minimum VAF threshold to distinguish true somatic variants from sequencing artifacts. However, a fixed VAF threshold will behave differently depending on sample purity. In a low-purity sample, a true clonal variant may fall below a threshold that would be appropriate for a high-purity sample. Conversely, in a high-purity sample, sequencing artifacts with VAFs of a few percent may be more easily distinguished from true variants. Researchers should therefore interpret VAF thresholds in the context of the estimated purity of each sample instead of applying a universal cutoff.

Several computational methods estimate tumor purity from sequencing data. Some approaches use the overall distribution of VAFs across all somatic variants, fitting a model that accounts for purity, ploidy, and clonality. Other approaches use copy number data or allele-specific copy number states to infer purity. The choice of method depends on the data available and the assumptions that are reasonable for the tumor type. For samples with many somatic variants, the VAF distribution approach can be robust. For samples with few variants, purity estimation becomes more uncertain, and the confidence intervals around purity estimates should be reported alongside the point estimate.

The relationship between VAF and purity also matters for germline variant calling. Germline variants are expected to have VAFs near 50 percent for heterozygous variants or near 100 percent for homozygous variants in normal tissue. In tumor samples, germline variants will also be present, and their VAFs will be influenced by copy number alterations in the tumor cells. A germline heterozygous variant in a region of copy number loss in the tumor will have a VAF that deviates from 50 percent, because the tumor cells may have lost one allele. Distinguishing germline from somatic variants requires matched normal tissue, typically from blood or adjacent normal tissue, to compare VAFs and genotypes between tumor and normal samples. The NCBI maintains databases and search systems that support this type of comparative analysis, including resources for accessing sequence data and variant annotations [<a href="#ref-6">6</a>].

## Ploidy and Copy Number Effects on VAF

Ploidy refers to the number of complete sets of chromosomes in a cell. Normal human somatic cells are diploid, with two copies of each autosome. Tumor cells frequently deviate from diploidy, exhibiting copy number gains, losses, and whole-genome duplication events. These copy number alterations change the expected VAF for a given variant, even when purity and clonality are held constant.

Consider a clonal variant present in all tumor cells at a locus with three copies of the chromosome, where the variant is present on one copy. Each tumor cell contributes one variant allele and two reference alleles, so the expected VAF in a pure sample is one-third, approximately 33 percent. If the locus has four copies with the variant on one copy, the expected VAF is 25 percent. If the variant is on two of four copies, the expected VAF is 50 percent. These examples show that VAF alone cannot distinguish between a clonal variant at a locus with copy number gain and a subclonal variant at a diploid locus. Copy number information is required to resolve this ambiguity.

Allele-specific copy number analysis provides the necessary context. By determining the copy number of each allele at a locus, researchers can calculate the expected VAF for a clonal variant and then compare the observed VAF to this expectation. The ratio of observed to expected VAF gives the cancer cell fraction, which is the proportion of tumor cells carrying the variant. A cancer cell fraction near 1.0 indicates a clonal variant present in all tumor cells. A cancer cell fraction substantially below 1.0 indicates a subclonal variant present in only a fraction of the tumor population.

Whole-genome duplication is a common event in cancer evolution and has a major effect on VAF interpretation. After whole-genome duplication, a diploid tumor becomes tetraploid. A clonal variant that was present on one of two alleles before duplication will be present on two of four alleles after duplication, maintaining a VAF of 50 percent in a pure sample. However, the copy number state has changed, and the expected VAF for a clonal variant at a tetraploid locus with the variant on one allele is 25 percent. Researchers analyzing tumors with evidence of whole-genome duplication must account for this shift when interpreting VAF distributions.

## Clonality and Subclonal Architecture

Clonality describes the fraction of tumor cells that carry a particular variant. A clonal variant is present in all tumor cells and represents an early event in tumor evolution. A subclonal variant is present in only a subset of tumor cells and represents a later event that occurred after the tumor population diversified. The distribution of VAFs across all somatic variants in a tumor provides a snapshot of its clonal architecture.

In a pure diploid tumor with no copy number alterations, clonal variants cluster around a VAF of 50 percent. Subclonal variants form a tail of lower VAFs extending downward from this cluster. The shape of this tail reflects the branching structure of the tumor phylogeny. A tumor with a dominant clone and many small subclones will show a dense cluster at 50 percent with a long sparse tail. A tumor with several large subclones will show multiple peaks in the VAF distribution, each corresponding to a different clone.

Copy number alterations and purity changes distort this simple picture. A clonal variant at a locus with copy number gain may have a VAF well below 50 percent, while a subclonal variant at a locus with copy number loss may have a VAF close to 50 percent. Without copy number correction, the VAF distribution alone cannot reliably distinguish clonal from subclonal variants. Computational tools that jointly model purity, ploidy, copy number, and clonality can reconstruct the underlying clonal architecture from the observed VAF distribution, but these tools require sufficient numbers of somatic variants and make assumptions about the mutation process that may not hold for all tumors.

The clinical significance of clonality is an active area of investigation. In polycythemia vera, the JAK2 V617F variant allele frequency is a key determinant of outcomes, including thrombosis and progression to myelofibrosis, and the dynamics of JAK2 V617F VAF over time and in response to treatment are clinically relevant [<a href="#ref-7">7</a>]. This example demonstrates that VAF can carry prognostic information in specific disease contexts, even when the general clinical utility of VAF as a predictive biomarker remains unvalidated [<a href="#ref-2">2</a>]. Researchers studying diseases where VAF has established biological relevance should be aware of the disease-specific literature and should not assume that VAF thresholds validated in one disease context transfer to another.

## At a Glance: VAF Interpretation by Sample Context

The following table summarizes how VAF should be interpreted under different sample conditions. These values assume a clonal variant present in all tumor cells and no subclonality.

| Sample Condition | Purity | Copy Number State | Expected VAF | Interpretation |
| --- | --- | --- | --- | --- |
| Pure diploid tumor | 100% | Diploid, variant on 1 of 2 alleles | 50% | Clonal heterozygous variant |
| Impure diploid tumor | 50% | Diploid, variant on 1 of 2 alleles | 25% | Clonal variant diluted by normal cells |
| Pure tumor with copy number gain | 100% | 3 copies, variant on 1 of 3 alleles | 33% | Clonal variant at gained locus |
| Pure tumor with copy number loss | 100% | 1 copy, variant on the remaining allele | 100% | Clonal variant with loss of reference allele |
| Pure tumor after whole-genome duplication | 100% | 4 copies, variant on 2 of 4 alleles | 50% | Clonal variant, pre-duplication event |
| Pure tumor after whole-genome duplication | 100% | 4 copies, variant on 1 of 4 alleles | 25% | Clonal variant, post-duplication event |

These expected values provide a reference frame for interpreting observed VAFs. When the observed VAF for a variant falls below the expected value for a clonal variant given the estimated purity and copy number state, the variant is likely subclonal. When the observed VAF matches the expected value, the variant is likely clonal. When the observed VAF exceeds the expected value, the copy number estimate or purity estimate may be inaccurate, or the variant may be present on multiple alleles.

## Practical Workflow for VAF-Based Analysis

A practical workflow for VAF-based analysis of somatic variants proceeds through several stages. Each stage involves specific decisions that affect the final interpretation.

### Step 1: Assess Sample Quality and Sequencing Depth

Before interpreting VAF values, confirm that the sequencing data meet quality standards. Check the mean depth across the targeted regions, the uniformity of coverage, and the fraction of reads passing quality filters. Low-quality samples with degraded DNA or high duplication rates will produce unreliable VAF estimates. The EMBL-EBI offers training materials on data quality assessment and analysis practices that are useful for establishing these quality checks [<a href="#ref-8">8</a>]. The Carpentries lessons provide foundational computing skills, including shell and data handling, that support reproducible quality assessment workflows [<a href="#ref-9">9</a>].

### Step 2: Call Variants with a Validated Pipeline

Use a somatic variant calling pipeline that has been validated for the intended application. Community pipelines such as those from nf-core provide standardized, reproducible workflows that incorporate best practices for alignment, variant calling, and annotation [<a href="#ref-4">4</a>]. Document the pipeline version, reference genome version, and parameter settings so that VAF calculations can be reproduced. Different pipeline versions may produce different VAF estimates for the same variant, so version control is essential.

### Step 3: Estimate Tumor Purity

Estimate tumor purity using a method appropriate for the data type. For samples with sufficient somatic variants, use the VAF distribution to estimate purity. For samples with matched copy number data, use allele-specific copy number states to refine the purity estimate. Record the purity estimate and its confidence interval. If purity cannot be estimated reliably, note this limitation and interpret VAF values with caution.

### Step 4: Determine Copy Number State at Each Variant Locus

For each somatic variant, determine the copy number state of the genomic region containing the variant. This requires copy number segmentation data, which can be derived from the same sequencing data or from a separate assay. Record the total copy number and the allele-specific copy number, if available. Without this information, VAF values cannot be converted to cancer cell fractions.

### Step 5: Calculate Cancer Cell Fraction

For each variant, calculate the expected VAF for a clonal variant given the purity and copy number state. Divide the observed VAF by the expected VAF to obtain the cancer cell fraction. Variants with cancer cell fractions near 1.0 are clonal. Variants with cancer cell fractions substantially below 1.0 are subclonal. The precision of the cancer cell fraction estimate depends on the read depth at the locus and the accuracy of the purity and copy number estimates.

### Step 6: Cluster Variants by Cancer Cell Fraction

Group variants by their cancer cell fractions to identify clonal and subclonal populations. Variants that cluster at a cancer cell fraction near 1.0 represent the founding clone. Variants that cluster at lower cancer cell fractions represent subclones. The number and size of subclonal clusters describe the tumor's evolutionary structure. This clustering step is sensitive to errors in purity and copy number estimates, so validate the clustering results using multiple methods when possible.

### Step 7: Filter Variants for Downstream Analysis

Apply variant filtering criteria that account for the biological context. For driver gene analysis, consider whether subclonal variants in known driver genes should be retained even if they fall below a VAF threshold. For clonality analysis, retain all variants with reliable VAF estimates regardless of VAF magnitude. For clinical reporting, apply thresholds that have been validated for the specific assay and clinical context, recognizing that validated VAF thresholds are not available for most applications [<a href="#ref-2">2</a>].

## Records and Measurements for VAF Analysis

Maintaining detailed records of VAF analysis is essential for reproducibility and for interpreting results in the context of sample quality and biological variability. The following measurements should be recorded for each sample and each variant.

For each sample, record the estimated purity, the method used to estimate purity, the confidence interval around the purity estimate, the ploidy estimate, and the evidence for whole-genome duplication if present. Record the sequencing platform, the target capture design, the mean depth, and the fraction of the target covered at the depth required for reliable VAF estimation.

For each variant, record the genomic position, the reference and alternate alleles, the total read depth, the variant read count, the raw VAF, the copy number state at the locus, the expected VAF for a clonal variant, the cancer cell fraction, and the confidence interval around the cancer cell fraction. Record the variant caller version and parameters, the annotation source, and any filtering decisions applied to the variant.

These records enable downstream users to assess the reliability of VAF-based conclusions and to re-analyze the data if improved methods become available. The Bioconductor project provides packages and workflows for reproducible genomic analysis, including tools for copy number analysis and clonality inference that can be incorporated into a documented analysis pipeline [<a href="#ref-10">10</a>].

## Common Failure Patterns in VAF Interpretation

Several recurring errors undermine VAF-based analysis. Recognizing these patterns helps researchers avoid incorrect conclusions.

### Applying a Fixed VAF Threshold Without Purity Adjustment

A fixed VAF threshold, such as 5 percent or 10 percent, will misclassify variants in samples with different purities. In a high-purity sample, a clonal variant may have a VAF of 40 percent, while in a low-purity sample, the same clonal variant may have a VAF of 10 percent. A threshold that works for one sample will not work for another. Always interpret VAF thresholds in the context of the sample purity.

### Ignoring Copy Number State

VAF values cannot be interpreted without copy number information. A variant with a VAF of 30 percent could be clonal at a locus with three copies, subclonal at a diploid locus, or clonal at a locus with copy number loss and an additional reference allele. Copy number data are essential for converting VAF to cancer cell fraction.

### Confusing Germline and Somatic Variants

Germline variants are present in all cells, including tumor cells. In a tumor sample, a germline heterozygous variant at a diploid locus will have a VAF near 50 percent, which is indistinguishable from a clonal somatic variant without matched normal data. Matched normal tissue is required to distinguish germline from somatic variants. The NCBI provides resources for accessing and comparing sequence data across samples, which supports this distinction [<a href="#ref-6">6</a>].

### Overinterpreting Low-VAF Variants

Low-VAF variants are prone to false positives from sequencing artifacts and to imprecise VAF estimates from low variant read counts. A variant with a VAF of 2 percent based on 100 total reads has only 2 variant reads, and the confidence interval around this estimate is wide. High-depth sequencing is required to reliably detect and quantify low-VAF variants, as demonstrated by the 2000-fold or greater depth used in the focal cortical dysplasia study [<a href="#ref-3">3</a>].

### Assuming VAF Thresholds Are Clinically Validated

The clinical utility of VAF as a predictive biomarker for therapy selection has not been established, and validated VAF thresholds are lacking [<a href="#ref-2">2</a>]. Researchers should not apply VAF cutoffs to clinical decisions without appropriate analytical and clinical validation for the specific assay and disease context. In diseases where VAF has established prognostic relevance, such as JAK2 V617F in polycythemia vera, the disease-specific literature should guide interpretation [<a href="#ref-7">7</a>].

## Limitations of VAF-Based Inference

VAF-based inference has inherent limitations that should be acknowledged in any analysis.

### Uncertainty in Purity and Copy Number Estimates

Purity and copy number estimates are themselves derived from data and carry uncertainty. Errors in these estimates propagate to cancer cell fraction calculations and can lead to incorrect clonality assignments. For samples with low purity or few somatic variants, the uncertainty is larger, and clonality conclusions should be expressed with appropriate confidence intervals.

### Inability to Resolve Subclonal Structure at Low Resolution

Bulk sequencing provides an average VAF across all cells in the sample. Subclones with similar cancer cell fractions cannot be resolved as distinct populations. Single-cell approaches or microdissection followed by sequencing of enriched cell populations can provide higher resolution, as demonstrated in the focal cortical dysplasia study where microdissected cells were analyzed to confirm that dysmorphic neurons and balloon cells carried the pathogenic variants [<a href="#ref-3">3</a>]. These approaches are more labor-intensive and are not feasible for all samples.

### Assumptions About the Mutation Process

Clonality inference methods assume that somatic mutations accumulate at a relatively constant rate and that the observed VAF distribution reflects the underlying clonal structure. Tumors with complex evolutionary histories, including chromothripsis, kataegis, or extensive copy number heterogeneity, may violate these assumptions. The resulting clonality estimates may be inaccurate.

### Technical Artifacts

Sequencing artifacts, including oxidative damage, PCR errors, and mapping errors, can produce false variants with low VAFs. These artifacts are more common in FFPE-derived DNA and in samples with low input amounts. Artifact filtering methods can reduce but not eliminate this problem. The VAF of a true variant can also be distorted by allelic bias in amplification or capture, where the variant allele is preferentially amplified or lost relative to the reference allele.

### Lack of Standardization Across Assays

Different sequencing assays have different error profiles, depth distributions, and variant calling pipelines. VAF values are not directly comparable across assays without careful calibration. The lack of standardization between sequencing assays is a recognized barrier to the clinical implementation of VAF as a decision-making tool [<a href="#ref-2">2</a>].

## Quality Controls and Reproducibility

Quality controls for VAF-based analysis should be applied at multiple levels.

### Sample-Level Controls

Include a normal control sample processed through the same pipeline to establish the baseline error rate and to distinguish germline from somatic variants. Include a positive control with known variants at defined VAFs to verify that the pipeline accurately estimates VAF across the range of interest. Include a negative control to assess contamination and artifact rates.

### Variant-Level Controls

For each variant, assess the read-level evidence, including the base quality scores, mapping quality scores, and the strand balance of variant reads. Variants supported predominantly by reads on one strand are more likely to be artifacts. Assess the local sequence context for homopolymer runs or repetitive elements that may cause mapping errors.

### Pipeline-Level Controls

Document the pipeline version, reference genome version, and all parameter settings. Use a pipeline management system that tracks these details automatically. The nf-core documentation provides guidance on pipeline configuration and reproducibility standards [<a href="#ref-4">4</a>]. The Galaxy Training Network offers tutorials on reproducible analysis workflows that can be adapted for somatic variant calling [<a href="#ref-5">5</a>].

### Analysis-Level Controls

Record the purity estimate, copy number calls, and clonality assignments for each sample. Re-run the analysis with different parameter settings or different methods to assess the robustness of the conclusions. If the clonality assignments change substantially with reasonable parameter variations, the conclusions should be reported as uncertain.

## Reporting VAF Results

When reporting VAF results, distinguish between the observed VAF, the purity-adjusted expected VAF, and the cancer cell fraction. These values answer different questions and should not be conflated.

The observed VAF is the direct measurement from sequencing data. It is the proportion of variant reads at the locus. Report the read depth and variant read count alongside the VAF so that readers can assess the precision of the estimate.

The purity-adjusted expected VAF is the VAF that would be expected for a clonal variant given the estimated purity and copy number state. This value provides the reference frame for interpreting the observed VAF.

The cancer cell fraction is the proportion of tumor cells carrying the variant. This value is the most biologically meaningful measure of clonality, but it depends on the accuracy of the purity and copy number estimates.

Report the confidence intervals around purity, copy number, and cancer cell fraction estimates. A cancer cell fraction of 0.8 with a wide confidence interval that includes 1.0 is not strong evidence of subclonality. A cancer cell fraction of 0.8 with a narrow confidence interval that excludes 1.0 is stronger evidence.

For clinical reporting, follow the reporting standards of the laboratory and the relevant regulatory framework. Recognize that VAF thresholds for clinical decision-making are not validated for most applications [<a href="#ref-2">2</a>]. In diseases where VAF has established clinical relevance, such as JAK2 V617F in polycythemia vera, report VAF in the context of the disease-specific guidelines and literature [<a href="#ref-7">7</a>].

## Professional Escalation Criteria

Researchers and laboratory professionals should escalate VAF interpretation questions to appropriate experts in several situations.

### When Purity Cannot Be Estimated Reliably

If the sample has too few somatic variants for VAF-based purity estimation and no independent purity estimate is available, escalate to a computational biologist or bioinformatician with experience in tumor purity estimation. Proceeding with clonality analysis without a reliable purity estimate will produce misleading results.

### When Copy Number Data Are Unavailable or Contradictory

If copy number data are not available for the sample, or if the copy number calls conflict with the VAF distribution, escalate to a cytogeneticist or computational biologist who can assess the copy number evidence. VAF interpretation without copy number context is unreliable.

### When Clonality Assignments Are Unstable

If clonality assignments change substantially when the analysis is repeated with different parameter settings or methods, escalate to a statistical geneticist or bioinformatician to assess the sources of instability. Unstable clonality assignments should not be used to support biological conclusions.

### When VAF Results Will Inform Clinical Decisions

If VAF results will be used to guide treatment decisions, escalate to the clinical team and the laboratory director to ensure that the results are interpreted within the validated clinical context. The absence of validated VAF thresholds for most clinical applications [<a href="#ref-2">2</a>] means that VAF-based treatment decisions require careful clinical judgment and should be documented as such.

### When Germline and Somatic Variants Cannot Be Distinguished

If matched normal tissue is not available and germline variants cannot be excluded, escalate to a molecular geneticist or genetic counselor to assess the implications of the ambiguity. Germline variants in cancer predisposition genes have implications for the patient and family members, as demonstrated by the identification of pathogenic germline variants in 8 percent of cancer cases in a large study of 10,389 adult cancers [<a href="#ref-11">11</a>].

## Frequently Asked Questions

### What is the difference between VAF and cancer cell fraction?

VAF is the proportion of sequencing reads supporting the variant allele at a locus. Cancer cell fraction is the proportion of tumor cells carrying the variant. VAF is a direct measurement from sequencing data. Cancer cell fraction is derived from VAF by adjusting for tumor purity and copy number state. A VAF of 25 percent in a sample with 50 percent purity and diploid copy number corresponds to a cancer cell fraction of 1.0, meaning all tumor cells carry the variant. The same VAF in a sample with 100 percent purity and diploid copy number corresponds to a cancer cell fraction of 0.5, meaning half of the tumor cells carry the variant.

### Why does tumor purity affect VAF interpretation?

Tumor purity determines the fraction of cells in the sample that are tumor cells. Normal cells contribute reference alleles at somatic variant loci, diluting the observed VAF. A clonal variant present in all tumor cells at a diploid locus has an expected VAF of half the purity. If purity is 80 percent, the expected VAF is 40 percent. If purity is 20 percent, the expected VAF is 10 percent. Without accounting for purity, a clonal variant in a low-purity sample can be mistaken for a subclonal variant.

### How does copy number alteration change the expected VAF?

Copy number alterations change the number of reference and variant alleles per tumor cell. At a locus with three copies and the variant on one copy, each tumor cell contributes one variant allele and two reference alleles, giving an expected VAF of 33 percent in a pure sample. At a locus with one copy and the variant on that copy, each tumor cell contributes one variant allele and no reference alleles, giving an expected VAF of 100 percent. Copy number information is required to convert VAF to cancer cell fraction.

### Can VAF distinguish clonal from subclonal variants?

VAF can distinguish clonal from subclonal variants when purity and copy number state are known. A variant with a cancer cell fraction near 1.0 is clonal. A variant with a cancer cell fraction substantially below 1.0 is subclonal. Without purity and copy number information, VAF alone cannot make this distinction, because a clonal variant at a locus with copy number gain can have the same VAF as a subclonal variant at a diploid locus.

### What read depth is needed for reliable VAF estimation?

The required read depth depends on the lowest VAF that needs to be detected and the precision required for the VAF estimate. Detecting a variant with a true VAF of 5 percent requires enough depth to observe multiple variant reads. A depth of 100 reads would give an expected 5 variant reads, which is marginal. A depth of 500 reads would give an expected 25 variant reads, which is more reliable. Studies that need to detect low-VAF somatic variants have used depths of 2000-fold or greater [<a href="#ref-3">3</a>]. Higher depth improves precision but increases cost.

### How do germline variants appear in tumor VAF distributions?

Germline heterozygous variants have a VAF near 50 percent in normal tissue. In tumor samples, germline variants are present in both tumor and normal cells, and their VAFs are influenced by copy number alterations in the tumor cells. A germline variant in a region of copy number loss in the tumor will have a VAF that deviates from 50 percent. Matched normal tissue is required to distinguish germline from somatic variants. Germline variants in cancer predisposition genes were found in 8 percent of cases in a large adult cancer cohort [<a href="#ref-11">11</a>].

### Is VAF a validated clinical biomarker?

VAF has been evaluated as a predictive biomarker for targeted therapy selection, but validated VAF thresholds and standardization between sequencing assays are lacking [<a href="#ref-2">2</a>]. In specific diseases, VAF has established clinical relevance. For example, JAK2 V617F VAF is a key determinant of outcomes in polycythemia vera, including thrombosis and progression to myelofibrosis [<a href="#ref-7">7</a>]. Researchers should consult the disease-specific literature before applying VAF thresholds to clinical decisions.

### What should I do if my VAF-based clonality results are unstable?

If clonality assignments change when the analysis is repeated with different parameter settings or methods, the results should be reported as uncertain. Escalate to a computational biologist or statistical geneticist to assess the sources of instability. Common causes include low sample purity, few somatic variants, inaccurate copy number calls, and violations of the assumptions of the clonality inference method. Consider whether higher-depth sequencing or single-cell approaches are needed to resolve the clonal structure.

## Related Bioinformatics Guides

- [Detecting Structural Variants with Long-Read Sequencing: Methods and Considerations](/knowledge/bioinformatics/detecting-structural-variants-with-long-read-sequencing-methods-and-considerations)
- [Genomic Data Analysis Tools: A Comparative Guide for Researchers](/knowledge/bioinformatics/genomic-data-analysis-tools-a-comparative-guide-for-researchers)
- [Understanding UMI in Single-Cell Sequencing: What It Is and Why It Matters](/knowledge/bioinformatics/understanding-umi-in-single-cell-sequencing-what-it-is-and-why-it-matters)
- [Genomic Data Processing: From Raw Sequencing to Analysis-Ready Files](/knowledge/bioinformatics/genomic-data-processing-from-raw-sequencing-to-analysis-ready-files)
- [RNA Sequencing Data Analysis: From Raw Reads to Differential Expression](/knowledge/bioinformatics/rna-sequencing-data-analysis-from-raw-reads-to-differential-expression)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)

## References and Further Reading

<a id="ref-1"></a>[<a href="#ref-1">1</a>] [Reconstructing developmental lineages: a retrospective approach using somatic mutations and variant allele frequency.](https://pubmed.ncbi.nlm.nih.gov/41584927). Frontiers in genetics, 2025.

<a id="ref-2"></a>[<a href="#ref-2">2</a>] [Variant allele frequency: a decision-making tool in precision oncology?](https://pubmed.ncbi.nlm.nih.gov/37704501). Trends in cancer, 2023.

<a id="ref-3"></a>[<a href="#ref-3">3</a>] [Dissecting the genetic basis of focal cortical dysplasia: a large cohort study.](https://pubmed.ncbi.nlm.nih.gov/31444548). Acta neuropathologica, 2019.

<a id="ref-4"></a>[<a href="#ref-4">4</a>] [nf-core Documentation](https://nf-co.re/docs). nf-core.

<a id="ref-5"></a>[<a href="#ref-5">5</a>] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.

<a id="ref-6"></a>[<a href="#ref-6">6</a>] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.

<a id="ref-7"></a>[<a href="#ref-7">7</a>] [JAK2 V617F allele burden in polycythemia vera: burden of proof.](https://pubmed.ncbi.nlm.nih.gov/36745865). Blood, 2023.

<a id="ref-8"></a>[<a href="#ref-8">8</a>] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.

<a id="ref-9"></a>[<a href="#ref-9">9</a>] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.

<a id="ref-10"></a>[<a href="#ref-10">10</a>] [Bioconductor](https://bioconductor.org/). Bioconductor Project.

<a id="ref-11"></a>[<a href="#ref-11">11</a>] [Pathogenic Germline Variants in 10,389 Adult Cancers.](https://pubmed.ncbi.nlm.nih.gov/29625052). Cell, 2018.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.