The Role of Tumor Purity and Ploidy in Somatic Variant Calling: How to Estimate and Adjust for Accurate Results

By Dr. Zubair Khalid, DVM, MS, PhD ·

The Role of Tumor Purity and Ploidy in Somatic Variant Calling: How to Estimate and Adjust for Accurate Results

Key Takeaways

  • Accurate somatic variant calling necessitates precise estimation of tumor purity (proportion of neoplastic cells) and ploidy (average DNA content per cell), as these parameters directly influence the expected variant allele fraction (VAF) of mutations. Without these estimates, filters based on VAF can misclassify true somatic variants as germline, subclonal, or artifactual.
  • Tools like PureCN, ASCAT, Sequenza, and Canvas estimate purity and ploidy by jointly modeling genome-wide B-allele frequency (BAF) and log R ratio (LRR) profiles, which reflect allelic imbalance at heterozygous germline SNPs and total copy number, respectively. These models resolve the ambiguity arising from the interaction between low purity and high ploidy.
  • The VAF of a somatic mutation is mathematically defined by purity ($p$), total copy number ($C_t$), and mutant copy number ($C_m$) via the formula VAF = $(p \times C_m) / (p \times C_t + 2 \times (1 - p))$, highlighting how dilution by normal cells and copy number alterations compress observed VAFs.
  • Matched normal samples simplify germline variant identification and BAF reference, improving purity and ploidy estimation accuracy, but tumor-only analysis is feasible using germline SNPs from tumor DNA and robust filtering against population databases.
  • Quality control for purity and ploidy estimation involves inspecting BAF/LRR segment profiles for model fit, evaluating solution scores for convergence, and troubleshooting common failure patterns such as sample contamination, low purity, or polyclonality, which can lead to inaccurate variant classification.

Somatic variant calling from tumor sequencing data requires accurate knowledge of two tumor-intrinsic parameters: purity, the proportion of neoplastic cells in the sequenced sample, and ploidy, the average DNA content of those cells. These values directly determine the expected variant allele fraction (VAF) for any somatic mutation, and without them, filters based on allelic fraction will misclassify variants as germline, subclonal, or artifact. This article explains how to estimate purity and ploidy using established tools such as PureCN, ASCAT, Sequenza, and Canvas, and how to incorporate those estimates into somatic variant calling and filtering decisions. The practical outcome for researchers is a defensible variant list with calibrated VAF thresholds, correct somatic versus germline classification, and interpretable copy number and clonality calls.

The Problem: Why Purity and Ploidy Matter in Variant Calling

A tumor biopsy is never pure tumor. It contains stromal cells, infiltrating immune cells, endothelial cells, and normal epithelium. The fraction of cancer cells in the sample is the purity, often called cellularity or tumor content. A sample with 40 percent purity means that for every 100 cells sequenced, 40 are neoplastic and 60 are normal. This dilution directly compresses the observed VAF of somatic mutations. A clonal heterozygous somatic mutation present in every tumor cell would be expected at 50 percent VAF in a pure sample, but at 20 percent VAF in a 40 percent pure sample. If the caller or analyst applies a germline filter that removes variants below 30 percent VAF, that somatic mutation will be discarded.

Ploidy complicates the calculation further. Many cancers are not diploid. Whole-genome duplication events, focal amplifications, and chromosome arm losses change the average copy number state. A tetraploid tumor with a heterozygous mutation on one allele of a four-copy chromosome has an expected VAF of 25 percent, not 50 percent. The same mutation in a diploid tumor has an expected VAF of 50 percent. Without a ploidy estimate, the analyst cannot distinguish a low-VAF clonal mutation in a high-ploidy tumor from a subclonal mutation in a diploid tumor.

The interaction between purity and ploidy creates a fundamental ambiguity in tumor sequencing data. A sample with low purity and high ploidy can produce VAF distributions that resemble a sample with higher purity and lower ploidy. Tools that estimate these parameters jointly, instead of independently, resolve this ambiguity by fitting the observed B-allele frequency and total copy number data to a grid of possible purity and ploidy combinations. The NCBI maintains the sequence databases and reference genomes that underpin these analyses, and the EMBL-EBI training portal offers structured learning pathways for the data resources used in tumor sequencing workflows.

Core Principles: Allele-Specific Copy Number and Variant Allele Fractions

The Relationship Between VAF, Purity, and Copy Number

The expected VAF of a somatic variant depends on three quantities: the purity p, the total copy number of the locus in tumor cells Ct, and the number of mutant copies Cm. The formula is:

VAF = (p × Cm) / (p × Ct + 2 × (1 - p))

The denominator represents the total number of alleles in the sample, combining tumor alleles and normal diploid alleles. For a clonal heterozygous mutation in a diploid tumor (Ct = 2, Cm = 1) with 100 percent purity, the VAF is 0.5. At 50 percent purity, the VAF drops to 0.33. At 25 percent purity, it drops to 0.2. For a homozygous mutation (Cm = 2) in the same diploid tumor, the VAF at 50 percent purity is 0.5, identical to the heterozygous case at 100 percent purity. This degeneracy is why purity estimation cannot rely on VAF alone.

Copy number alterations shift the expected VAF in predictable ways. A heterozygous mutation on a chromosome with a single copy in the tumor (loss of heterozygosity, Ct = 1, Cm = 1) has an expected VAF of 1.0 in a pure sample, because every tumor allele carries the mutation. At 50 percent purity, the VAF is 0.5. A mutation on one allele of a three-copy chromosome (Ct = 3, Cm = 1) has an expected VAF of 0.33 in a pure sample. These calculations form the basis for classifying variants as clonal or subclonal after purity and ploidy are known.

B-Allele Frequency and Log R Ratio as Inputs

Purity and ploidy estimation tools use two genome-wide measurements derived from sequencing data. The B-allele frequency (BAF) tracks the allelic imbalance at heterozygous germline SNP positions. In a pure diploid sample, germline heterozygotes show BAF values near 0.5. In an impure tumor with copy number alterations, BAF values shift away from 0.5 in regions of allelic imbalance, and the magnitude of the shift depends on purity and the copy number state. The log R ratio (LRR) measures the total copy number by comparing observed read depth to a neutral baseline. Regions with copy number gain show positive LRR values, and regions with loss show negative values.

The joint distribution of BAF and LRR across the genome provides the signal for purity and ploidy estimation. A tool that fits this distribution to a model of tumor cell fractions and integer copy numbers can recover purity and ploidy simultaneously. The PureCN package, distributed through Bioconductor, implements this approach for targeted and whole-exome sequencing data, with support for both matched and unmatched tumor samples. PureCN estimates purity, ploidy, copy number, loss of heterozygosity, and contamination, then uses those estimates to classify single nucleotide variants by somatic status and clonality.

At a Glance: Purity and Ploidy Estimation Tools and Their Inputs

ToolInput DataKey OutputsMatched Normal RequiredBest Use Case
PureCNTargeted panel or WES BAM files, VCF, SNP BAFPurity, ploidy, integer copy number, LOH, SNV classificationNo, supports unmatchedClinical targeted panels, tumor-only WES, integration with variant callers
ASCATSNP array or sequencing-derived BAF and LRRPurity, ploidy, allele-specific copy number segmentsNo, but benefits from germline SNP dataGenome-wide copy number profiling, WES and WGS
SequenzaWES or WGS BAM filesPurity, ploidy, allele-specific copy number, cellular prevalenceNo, uses germline SNPs from the tumorWES and WGS with matched or unmatched normals
CanvasWGS or WES BAM filesCopy number variants, ploidy, purity, heterogeneityOptional, supports matched and unmatchedLarge-scale WGS studies, diverse sample types

The choice of tool depends on the sequencing assay, the availability of a matched normal, and the downstream analysis goals. The Galaxy Training Network provides accessible workflow tutorials for copy number and variant analysis that can help researchers implement these tools in reproducible pipelines. The nf-core documentation describes community standards for pipeline configuration and reproducibility that apply when integrating purity and ploidy estimation into larger analysis workflows.

Estimating Purity and Ploidy: Practical Workflow

Step 1: Prepare Input Data with Quality Controls

Purity and ploidy estimation requires high-quality alignment and coverage data. Start with BAM files that have been aligned to the reference genome and processed with base quality score recalibration. For targeted panels, ensure that the design includes a sufficient number of heterozygous germline SNP positions across the genome or across each chromosome arm. A panel that covers only a few hundred kilobases of exonic sequence may have too few informative SNPs for reliable BAF-based estimation. The NCBI provides reference genome assemblies and dbSNP resources that are used to identify germline heterozygous positions for BAF calculation.

For whole-exome sequencing, coverage uniformity matters. Regions with extreme coverage outliers distort LRR calculations. Apply the same coverage thresholds used for variant calling, and exclude known problematic regions such as segmental duplications and low-complexity sequences. The EMBL-EBI training materials on sequence data quality can guide the selection of appropriate quality metrics.

Step 2: Generate BAF and LRR Profiles

Most purity and ploidy tools accept either raw BAM files or precomputed BAF and LRR values. For PureCN, the workflow begins with a VCF file of germline SNPs from the tumor sample, which is used to compute BAF at heterozygous positions. The tool then segments the genome using the LRR from the coverage data. For targeted panels, PureCN uses the on-target coverage to compute LRR and the SNP positions within the panel for BAF. The PureCN publication describes the methodology for estimating purity and ploidy from targeted short-read sequencing data, including the handling of unmatched tumor samples.

For whole-genome data, Canvas computes BAF and LRR across the entire genome and can infer ploidy, purity, and heterogeneity as part of its copy number calling workflow. The Canvas publication describes its application to whole-genome matched tumor-normal, small pedigree, and single-sample normal resequencing, as well as whole-exome matched and unmatched tumor-normal studies.

Step 3: Run Purity and Ploidy Estimation

Run the chosen tool with default parameters first, then inspect the fit. PureCN reports the optimal purity and ploidy solution along with a score that reflects how well the model explains the observed BAF and LRR data. ASCAT and Sequenza produce similar output, typically including a plot of the segmented BAF and LRR with the fitted model overlaid. Examine these plots for systematic deviations between the model and the data. A poor fit may indicate sample contamination, polyclonal tumor populations, or an incorrect assumption about the fraction of the genome that is aberrant.

The TOSCA workflow integrates purity and ploidy estimation into a tumor-only somatic calling pipeline that runs from raw read files through quality checks, alignment, variant calling, functional annotation, database filtering, and variant classification. TOSCA is implemented in Snakemake and is freely available, providing an end-to-end example of how purity and ploidy estimates feed into the final variant classification step.

Step 4: Incorporate Estimates into Variant Calling and Filtering

Once purity and ploidy are known, the expected VAF for a clonal heterozygous mutation can be calculated for any copy number state. Use this expected VAF to set mutation-specific filtering thresholds. Variants with observed VAFs far below the expected clonal VAF are candidates for subclonal mutations or artifacts. Variants with VAFs consistent with the expected germline heterozygous range, after adjusting for purity and copy number, should be classified as germline.

PureCN performs this classification automatically. It assigns each SNV a somatic posterior probability and a clonality estimate, using the purity and ploidy solution to compute expected VAFs for each copy number state. The PureCN documentation on Bioconductor provides installation and usage instructions, and the package integrates with standard somatic variant detection pipelines.

Options and Tradeoffs: Matched Normal versus Tumor-Only Analysis

Paired Tumor-Normal Analysis

The gold standard for somatic variant calling uses a matched normal sample from the same patient. The normal sample provides a direct measurement of germline variants, allowing the caller to subtract inherited polymorphisms from the tumor variant list. Purity and ploidy estimation still requires tumor-specific modeling, but the matched normal simplifies the identification of germline heterozygous SNPs used for BAF calculation. The deep whole-genome sequencing study of cancer cell lines used matched tumor-normal pairs to create an artificial purity ladder for benchmarking purity and ploidy estimation methods, demonstrating that matched normal data improves the accuracy of these estimates.

Tumor-Only Analysis

A matched normal sample is not always available. Retrospective analyses of clinical samples, archival tissues, and many research cohorts lack paired normals. In these cases, the analyst must distinguish somatic from germline variants using population databases, allele frequency filters, and tumor-specific features. The TOSCA publication notes that in silico screening against public or private databases and other filtering approaches are used when a paired normal is absent, but that difficulties in achieving sufficient accuracy have limited clinical applications. TOSCA addresses this by automating the tumor-only workflow, including purity and ploidy estimation, and produces somatic and germline variant estimates consistent with paired analyses.

The PureCN workflow for clinical tumor-only WES benchmarks allele-specific copy number analysis against gold standard SNP6 microarray data and matched normal WES data. The workflow produces purity and ploidy estimates that are highly concordant with the gold standard, and it classifies SNVs by somatic status, infers mutational signatures, and computes tumor mutational burden. This demonstrates that tumor-only WES can support accurate allele-specific copy number analysis when the right tools are used.

Tradeoff Summary

Matched normal analysis provides cleaner germline subtraction and more reliable BAF estimation, but it doubles sequencing cost and requires a normal tissue sample. Tumor-only analysis is cheaper and applicable to more samples, but it requires careful filtering against population databases and relies more heavily on purity and ploidy estimates to distinguish somatic from germline variants. The choice depends on the research question, the sample availability, and the tolerance for misclassification. For clinical applications where a germline variant could have hereditary implications, matched normal analysis is strongly preferred.

Observations and Measurements: What the Data Should Look Like

Expected VAF Distributions

After purity and ploidy estimation, inspect the VAF distribution of called somatic variants. In a clean clonal tumor, somatic mutations should cluster near the expected VAF for the dominant copy number state. A broad smear of VAFs below the expected clonal value suggests subclonal populations, which may be biologically relevant or may indicate poor purity estimation. A VAF distribution that peaks near 0.5 across all copy number states, regardless of the purity estimate, may indicate that the purity solution is wrong or that the sample is contaminated with a second tumor population.

BAF and LRR Segment Profiles

The segmented BAF profile should show discrete bands corresponding to integer copy number states. In a diploid tumor with purity above 60 percent, heterozygous germline SNPs in copy-neutral regions should have BAF values near 0.5, with deviations in regions of LOH or allelic imbalance. The LRR profile should show clear separation between copy number states, with the spacing between adjacent states proportional to the purity. If the LRR bands are compressed or overlapping, the purity may be too low for reliable copy number calling, or the sample may contain multiple tumor clones with different copy number profiles.

Purity and Ploidy Solution Scores

Most tools report a score or likelihood for each candidate purity and ploidy solution. A strong solution has a clear score peak, with the optimal solution substantially better than the alternatives. A flat score surface, where many purity and ploidy combinations fit equally well, indicates that the data lack sufficient information to resolve the parameters. This occurs in low-purity samples, in samples with few copy number alterations, and in targeted panels with sparse SNP coverage. In these cases, report the range of plausible solutions and treat downstream variant classifications as provisional.

Records and Measurements: Documenting Purity and Ploidy Estimates

What to Record

For every sample, record the following in the analysis metadata: the tool and version used for purity and ploidy estimation, the input data type (targeted panel, WES, WGS), the matched normal status, the estimated purity and ploidy, the solution score, and the confidence interval or alternative solutions. Record the number of informative SNPs used for BAF calculation and the fraction of the genome or panel covered by copy number segments. These records allow downstream analysts to assess the reliability of the estimates and to reproduce the analysis.

How to Use the Records

Purity and ploidy estimates should be stored in the sample metadata and passed to downstream analysis steps, including variant filtering, mutational signature analysis, and tumor mutational burden calculation. The tumor-only WES workflow demonstrates how purity and ploidy estimates feed into SNV classification, mutational signature inference, and TMB calculation, and it shows high concordance between tumor-only and matched normal pipelines when these estimates are used correctly.

Version Control and Reproducibility

Purity and ploidy estimation tools change between versions, and results can differ across releases. Record the exact software versions and parameter settings. Use workflow management systems such as Snakemake, as implemented in TOSCA, or Nextflow, as documented in the nf-core documentation, to ensure that the analysis is reproducible. The Carpentries lessons provide foundational training in version control with Git and reproducible computing practices that apply to managing analysis code and metadata.

Quality Controls and Troubleshooting

Sample Contamination

Cross-sample contamination shifts BAF values toward 0.5 and compresses LRR values, mimicking low purity. PureCN includes a contamination estimate that can be used to flag samples with excessive contamination. If the contamination estimate exceeds the tool's recommended threshold, the purity and ploidy estimates should be interpreted with caution. The PureCN publication describes the contamination estimation methodology and its integration with purity and ploidy calling.

Low Purity Samples

Samples with purity below 20 percent produce weak BAF and LRR signals. The expected VAF for a clonal heterozygous mutation in a diploid tumor at 20 percent purity is 0.11, which falls below the detection threshold of many variant callers. Purity and ploidy estimates for these samples have wide confidence intervals, and somatic variant calling may miss a substantial fraction of true mutations. If the purity estimate is below 20 percent, consider whether the sample is suitable for the intended analysis or whether a different tissue block or enrichment method would provide higher tumor content.

Polyclonal Tumors

Tumors with multiple subclones carrying different copy number profiles produce BAF and LRR patterns that no single purity and ploidy solution explains well. The solution score will be low, and the segmented profiles will show intermediate BAF values that do not match integer copy number states. Canvas explicitly models heterogeneity in cancer samples, as described in the Canvas publication, and can report the presence of multiple tumor populations. For other tools, a poor fit may indicate polyclonality, and the purity and ploidy estimates should be treated as representing the dominant clone.

Germline Contamination from the Same Patient

In tumor-only analysis, the sample contains germline SNPs from the patient's normal cells. These SNPs are expected and are used for BAF calculation. However, if the tumor sample is contaminated with DNA from a different individual, the BAF profile will show additional heterozygous positions that do not fit the expected germline pattern. This is distinct from the normal-cell contamination that purity estimation accounts for, and it requires a different correction. Check the BAF profile for unexpected heterozygous positions in regions of copy number loss, which may indicate cross-individual contamination.

Common Failure Patterns and Their Causes

Failure Pattern 1: Purity Estimate Too High

A purity estimate that exceeds the pathologist's visual estimate of tumor content by a large margin may indicate that the tool is fitting a high-purity, high-ploidy solution to data that actually represent a lower-purity, diploid tumor. This occurs when the BAF and LRR data are sparse or noisy. Compare the tool's solution to the pathology report and to the observed VAF distribution of known somatic mutations. If the estimated purity implies expected VAFs that do not match the observed data, rerun the estimation with different parameters or a different tool.

Failure Pattern 2: Purity Estimate Too Low

A purity estimate that is much lower than expected may result from a tumor with extensive copy number loss, which reduces the LRR signal and compresses the BAF deviations. It may also result from a tumor with a large fraction of the genome in copy-neutral LOH, which produces no LRR change and only subtle BAF shifts. Check whether the copy number segments show the expected patterns for the tumor type. If the sample is known to be high-purity from pathology, a low purity estimate may indicate that the tool is misinterpreting the data.

Failure Pattern 3: Ploidy Estimate Ambiguous

Many tumors have a near-diploid or near-tetraploid genome, and the data may support both solutions. The tool will report two solutions with similar scores. This ambiguity directly affects VAF interpretation, because the expected VAF for a clonal mutation differs between diploid and tetraploid solutions. When the ploidy is ambiguous, report both solutions and show how the variant classification changes under each. The deep sequencing benchmark study provides a purity ladder that can be used to evaluate how well a given tool resolves ploidy at different purity levels.

Failure Pattern 4: No Convergent Solution

Some samples produce no purity and ploidy solution that fits the data. This occurs with severe contamination, extreme polyclonality, or technical artifacts in the sequencing data. Do not force a solution. Instead, document the failure, check the input data quality, and consider whether the sample is analyzable. The Galaxy Training Network offers tutorials on quality control and data cleaning that can help identify the source of the problem.

Limitations and Interpretation Boundaries

Purity and Ploidy Are Estimates, Not Measurements

Purity and ploidy estimates are model fits to noisy data. They carry uncertainty, and that uncertainty propagates into every downstream variant classification. A variant classified as clonal under one purity solution may be subclonal under another. Report the confidence in the purity and ploidy estimates alongside the variant calls, and avoid overinterpreting small differences in clonality or VAF.

Targeted Panels Have Limited Power

Targeted panels covering a few hundred genes provide limited genome-wide information for purity and ploidy estimation. The number of informative heterozygous SNPs within the panel may be small, and the LRR signal is restricted to the targeted regions. PureCN is optimized for this setting, as described in the PureCN publication, but the confidence intervals on purity and ploidy will be wider than for WES or WGS. For research applications where accurate purity and ploidy are critical, consider WES or WGS.

Tumor-Only Analysis Cannot Definitively Classify Germline Variants

Even with accurate purity and ploidy estimates, tumor-only analysis cannot prove that a variant is somatic. A rare germline variant that is not present in population databases may have a VAF consistent with a somatic mutation after purity adjustment. The TOSCA publication notes that tumor-only calling relies on database filtering and other approaches to separate private germline mutations from somatic variants, and that these methods have limitations. For clinical decisions with hereditary implications, confirm germline status with a matched normal sample.

Purity and Ploidy Do Not Capture All Tumor Heterogeneity

A single purity and ploidy estimate represents the dominant clone in the sample. Subclonal populations with different copy number profiles will produce VAFs that deviate from the expected values for the dominant clone. Tools such as Canvas that model heterogeneity, as described in the Canvas publication, can identify multiple populations, but the resolution depends on the sequencing depth and the fraction of each subclone. Variants assigned to subclones based on VAF alone should be interpreted with caution.

Safety and Regulatory Context

Clinical Reporting Requirements

In clinical settings, purity and ploidy estimates may be used to support variant classification and treatment decisions. Regulatory frameworks require that laboratory-developed tests be validated with documented performance characteristics. The tumor-only WES workflow provides a benchmarked example of how purity and ploidy estimation can be validated against gold standard data, and it demonstrates high concordance with matched normal pipelines. Laboratories adopting tumor-only workflows should perform similar validation studies before clinical use.

Data Privacy and Sample Tracking

Tumor sequencing data are sensitive patient information. Purity and ploidy estimation does not change the privacy requirements, but the analysis metadata should be managed under the same data governance policies as the sequencing data. The NCBI provides guidance on data submission and access for controlled-access datasets, and the EMBL-EBI training materials cover data management best practices for bioinformatics projects.

Professional Escalation Criteria

Escalate to a senior bioinformatician or clinical molecular pathologist when any of the following occur: the purity and ploidy solution fails to converge, the estimated purity is below 20 percent in a sample with high pathology-estimated tumor content, the ploidy estimate is ambiguous and the variant classification changes under alternative solutions, the contamination estimate exceeds the tool's recommended threshold, or the sample shows evidence of polyclonality that affects clinical variant interpretation. Document the reason for escalation and the actions taken.

A Practical Decision Framework for Selecting Purity and Ploidy Estimation Methods

Choosing the right purity and ploidy estimation approach requires a structured evaluation of the sequencing assay, sample characteristics, and downstream analytical requirements. Researchers often default to a single tool without considering how the choice affects variant classification accuracy. This section provides a decision framework that maps common experimental scenarios to appropriate estimation strategies, along with a record system for tracking estimation quality and a troubleshooting method for resolving discrepant results.

Decision Point 1: Assess the Sequencing Assay and Available Input Data

The first decision separates whole-genome sequencing (WGS), whole-exome sequencing (WES), and targeted panel data because each assay type provides different amounts of information for purity and ploidy estimation. WGS provides genome-wide coverage of heterozygous germline SNPs and copy number signal, making it the most informative assay for these estimates. The Canvas publication describes how Canvas leverages whole-genome data to infer ploidy, purity, and heterogeneity, and it supports both matched and unmatched tumor-normal studies. For WGS data, tools that use genome-wide BAF and LRR profiles, such as ASCAT, Sequenza, or Canvas, are appropriate.

WES data provide exonic coverage with uneven distribution across the genome. The number of informative heterozygous SNPs is lower than in WGS, but still sufficient for reliable estimation in most cases. The tumor-only WES workflow demonstrates that allele-specific copy number analysis from WES without matched normals produces purity and ploidy estimates highly concordant with gold standard SNP6 microarray data. For WES, PureCN and Sequenza are well-suited because they handle the coverage non-uniformity inherent to exome capture.

Targeted panel data present the greatest challenge. Panels covering a few hundred genes may contain only dozens to hundreds of informative heterozygous SNPs, and the LRR signal is restricted to targeted regions. The PureCN publication describes a methodology optimized for targeted short-read sequencing data, with support for matched and unmatched tumor samples. For targeted panels, PureCN is the preferred choice because it was designed for this setting and integrates with standard somatic variant detection pipelines.

Decision Point 2: Determine Matched Normal Availability

The availability of a matched normal sample changes the estimation strategy and the confidence in downstream variant classification. With a matched normal, the analyst can directly identify germline heterozygous SNPs from the normal sample, providing a clean reference for BAF calculation. The deep whole-genome sequencing study of cancer cell lines used matched tumor-normal pairs to create an artificial purity ladder for benchmarking purity and ploidy estimation methods, demonstrating that matched normal data improves the accuracy of these estimates. When a matched normal is available, any of the tools can be used, and the matched normal simplifies the identification of germline SNPs.

Without a matched normal, the analyst must use germline SNPs identified from the tumor sample itself. This is feasible because normal cell contamination in the tumor sample provides germline heterozygous positions. The TOSCA workflow performs tumor-only somatic calling from raw read files through quality checks, alignment, variant calling, functional annotation, database filtering, tumor purity and ploidy estimation, and variant classification. TOSCA is implemented in Snakemake and freely available, providing an end-to-end example of tumor-only analysis. The PureCN workflow for clinical tumor-only WES demonstrates that tumor-only WES can produce purity and ploidy estimates highly concordant with matched normal pipelines, and it classifies SNVs by somatic status, infers mutational signatures, and computes tumor mutational burden.

Decision Point 3: Evaluate Sample Characteristics That Affect Estimation Reliability

Three sample characteristics determine whether purity and ploidy estimation will be reliable: expected purity, expected ploidy, and clonal heterogeneity. Samples with expected purity below 20 percent produce weak BAF and LRR signals. The expected VAF for a clonal heterozygous mutation in a diploid tumor at 20 percent purity is 0.11, which falls below the detection threshold of many variant callers. For low-purity samples, the confidence intervals on purity and ploidy estimates widen, and the estimates should be treated as provisional.

Samples with expected high ploidy, such as those from tumor types known to undergo whole-genome duplication, require tools that explicitly model multiple ploidy states. Canvas models cancer ploidy, purity, and heterogeneity, as described in the Canvas publication, and can report the presence of multiple tumor populations. For samples with suspected polyclonality, tools that model heterogeneity are preferred because a single purity and ploidy solution will not explain the observed data well.

Decision Point 4: Match the Tool to the Downstream Analysis Goal

The downstream analysis goal determines which tool output is most important. If the goal is somatic variant classification with clonality assessment, PureCN is the appropriate choice because it integrates purity and ploidy estimation with SNV classification by somatic status and clonality. The PureCN publication describes how the package estimates tumor purity, copy number, loss of heterozygosity, and contamination, and then uses those estimates to classify single nucleotide variants.

If the goal is genome-wide copy number profiling with purity and ploidy as secondary outputs, ASCAT or Sequenza are appropriate. If the goal is a comprehensive copy number variant call set with ploidy and purity inference, Canvas is appropriate. The TOSCA workflow provides an integrated solution that combines purity and ploidy estimation with the full somatic calling pipeline, making it suitable for researchers who want an end-to-end tumor-only analysis without assembling multiple tools.

A Record System for Purity and Ploidy Estimation Quality

Maintaining a structured record of purity and ploidy estimates allows downstream analysts to assess the reliability of variant classifications and to reproduce the analysis. For every sample, record the following fields in the analysis metadata:

FieldExample ValuePurpose
Sample identifierTUMOR_001Links the estimate to the sequencing data
Tool and versionPureCN 2.4.0Tracks software changes that affect results
Input data typeTargeted panel, WES, WGSDocuments the information content of the assay
Matched normal statusMatched, unmatchedDocuments the germline subtraction strategy
Estimated purity0.65The primary tumor content estimate
Estimated ploidy2.1The average DNA content of tumor cells
Solution score0.92Measures how well the model fits the data
Alternative solutionsPurity 0.55, ploidy 3.9Documents ambiguity when multiple solutions fit
Number of informative SNPs1,240Measures the BAF signal strength
Contamination estimate0.03Flags cross-sample contamination
Pathology tumor content70 percentProvides an independent reference for comparison

Store these records in the sample metadata file and pass them to downstream analysis steps, including variant filtering, mutational signature analysis, and tumor mutational burden calculation. The tumor-only WES workflow demonstrates how purity and ploidy estimates feed into SNV classification, mutational signature inference, and TMB calculation, and it shows high concordance between tumor-only and matched normal pipelines when these estimates are used correctly.

Troubleshooting Method for Discrepant Purity and Ploidy Estimates

When the tool's purity estimate differs substantially from the pathology report or from the observed VAF distribution, use the following structured troubleshooting method to identify the cause.

Step 1: Compare the tool's purity estimate to the pathology-estimated tumor content. A large discrepancy, such as a tool estimate of 30 percent purity when pathology reports 80 percent tumor content, indicates a potential model misfit. Check whether the tool's solution score is low or whether multiple solutions have similar scores. A flat score surface indicates that the data lack sufficient information to resolve the parameters.

Step 2: Examine the BAF and LRR segment profiles. In a clean clonal tumor, the segmented BAF profile should show discrete bands corresponding to integer copy number states. If the BAF bands are compressed toward 0.5, the sample may have low purity or cross-sample contamination. If the LRR bands are compressed or overlapping, the purity may be too low for reliable copy number calling, or the sample may contain multiple tumor clones with different copy number profiles.

Step 3: Check the VAF distribution of called somatic variants. After purity and ploidy estimation, somatic mutations should cluster near the expected VAF for the dominant copy number state. A broad smear of VAFs below the expected clonal value suggests subclonal populations, which may be biologically relevant or may indicate poor purity estimation. A VAF distribution that peaks near 0.5 across all copy number states, regardless of the purity estimate, may indicate that the purity solution is wrong or that the sample is contaminated with a second tumor population.

Step 4: Rerun the estimation with a different tool. If PureCN and Sequenza produce substantially different purity and ploidy estimates for the same sample, the data may be ambiguous or the tools may be making different assumptions about the fraction of the genome that is aberrant. The deep sequencing benchmark study provides a purity ladder that can be used to evaluate how well a given tool resolves purity and ploidy at different purity levels.

Step 5: Document the discrepancy and its resolution. Record the initial estimate, the alternative estimate, the reason for the discrepancy, and the final estimate used for downstream analysis. This record allows downstream analysts to understand the uncertainty in the purity and ploidy estimates and to interpret variant classifications accordingly.

Escalation Criteria for Problematic Estimates

Escalate to a senior bioinformatician or clinical molecular pathologist when any of the following occur: the purity and ploidy solution fails to converge, the estimated purity is below 20 percent in a sample with high pathology-estimated tumor content, the ploidy estimate is ambiguous and the variant classification changes under alternative solutions, the contamination estimate exceeds the tool's recommended threshold, or the sample shows evidence of polyclonality that affects clinical variant interpretation. Document the reason for escalation and the actions taken.

The Galaxy Training Network provides accessible workflow tutorials for copy number and variant analysis that can help researchers implement these tools in reproducible pipelines. The nf-core documentation describes community standards for pipeline configuration and reproducibility that apply when integrating purity and ploidy estimation into larger analysis workflows. The Carpentries lessons provide foundational training in version control with Git and reproducible computing practices that apply to managing analysis code and metadata. The EMBL-EBI training portal offers structured learning pathways for the data resources used in tumor sequencing workflows, and the NCBI maintains the sequence databases and reference genomes that underpin these analyses.

Frequently Asked Questions

What is tumor purity and why does it matter for variant calling?

Tumor purity is the fraction of neoplastic cells in the sequenced sample. It matters because it determines the expected variant allele fraction of somatic mutations. A clonal heterozygous mutation at 50 percent VAF in a pure sample appears at 25 percent VAF in a 50 percent pure sample. Without a purity estimate, filters based on VAF will misclassify somatic mutations as germline or artifact.

What is ploidy and how does it differ from purity?

Ploidy is the average DNA content of the tumor cells, expressed as the number of chromosome copies. A diploid tumor has two copies of each chromosome, while a tetraploid tumor has four. Ploidy affects the expected VAF of somatic mutations because the denominator in the VAF calculation includes the total copy number. Purity and ploidy must be estimated jointly because they interact in the VAF calculation.

How do ASCAT, Sequenza, PureCN, and Canvas differ?

ASCAT and Sequenza are general-purpose tools for allele-specific copy number analysis that estimate purity and ploidy from BAF and LRR data. PureCN is an R/Bioconductor package optimized for targeted short-read sequencing data that integrates purity and ploidy estimation with SNV classification. Canvas is a copy number variant caller that also infers ploidy, purity, and heterogeneity, and it supports diverse sample types including matched and unmatched tumor-normal studies.

Can I estimate purity and ploidy without a matched normal sample?

Yes. Tools such as PureCN, ASCAT, Sequenza, and Canvas support tumor-only analysis. They use germline SNPs from the tumor sample itself to compute BAF and use coverage data for LRR. The tumor-only WES workflow demonstrates that purity and ploidy estimates from tumor-only data are highly concordant with estimates from matched normal data.

How do I use purity and ploidy to filter somatic variants?

Once purity and ploidy are known, calculate the expected VAF for a clonal heterozygous mutation at each copy number state. Use this expected VAF as the center of the somatic variant distribution. Variants with VAFs far below the expected clonal value are candidates for subclonal mutations or artifacts. Variants with VAFs consistent with the germline heterozygous range, after purity adjustment, should be classified as germline. PureCN automates this classification.

What should I do if the purity estimate is very low?

A purity estimate below 20 percent indicates that the sample has limited tumor content. The expected VAF for clonal somatic mutations will be below 11 percent for a diploid tumor, which is near the detection limit of many variant callers. Consider whether the sample is suitable for the intended analysis, or whether a different tissue block or enrichment method would provide higher tumor content. Document the low purity and its impact on variant detection sensitivity.

How do I know if my purity and ploidy estimates are reliable?

Check the solution score reported by the tool, the fit of the model to the BAF and LRR data, and the consistency of the estimates with pathology reports and known VAF distributions. A strong solution has a clear score peak and a good visual fit. Compare the estimated purity to the pathologist's estimate of tumor content. If the estimates are inconsistent, rerun the analysis with different parameters or a different tool.

What are the limitations of tumor-only somatic variant calling?

Tumor-only calling cannot definitively classify a variant as somatic because a rare germline variant may have a VAF consistent with a somatic mutation after purity adjustment. The TOSCA publication notes that database filtering and other approaches are used in the absence of a paired normal, but these have limitations. For clinical decisions with hereditary implications, confirm germline status with a matched normal sample.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.