Subclonal Reconstruction from Somatic Variants: How to Infer Tumor Evolution and Heterogeneity

By Dr. Zubair Khalid, DVM, MS, PhD ·

Subclonal Reconstruction from Somatic Variants: How to Infer Tumor Evolution and Heterogeneity

Key Takeaways

  • Subclonal reconstruction infers distinct tumor cell populations by clustering somatic variants based on their Variant Allele Frequencies (VAFs), which are then converted to Cancer Cell Fractions (CCFs) by accounting for tumor purity and local copy number alterations.
  • The accuracy of subclonal reconstruction is critically dependent on upstream variant calling, with choices of aligners, variant callers (e.g., PyClone, SciClone), and filtering strategies significantly influencing downstream conclusions about tumor evolution and heterogeneity.
  • Sequencing depth directly impacts the ability to resolve low-frequency subclones; for instance, 100-fold coverage has a standard error of ~3% for a 10% VAF, while 500-fold coverage reduces this to ~1.3%, improving cluster separation.
  • Tumor purity is a crucial parameter, as low purity (<20-30%) compresses VAF distributions, making subclonal populations difficult to distinguish and potentially leading to unreliable reconstruction.
  • Phylogenetic tree construction relies on the infinite sites assumption, inferring evolutionary relationships from shared and private mutations across subclones, with violations like loss of heterozygosity requiring careful handling.
  • Essential records for reproducibility include sequencing metadata, alignment and variant calling parameters, copy number estimates, reconstruction tool parameters, and output files, alongside tracking quality metrics like cluster stability and convergence diagnostics.

Tumor subclonal reconstruction is the computational process of inferring distinct cell populations within a cancer sample from somatic variant data, typically variant allele frequencies (VAFs), to estimate cancer cell fractions and build phylogenetic trees that describe tumor evolution. For biology students, researchers, and laboratory professionals working with bulk or single-cell sequencing data, this article provides a practical framework for designing, executing, and interpreting subclonal reconstruction analyses using tools such as PyClone and SciClone, while acknowledging the substantial influence of upstream variant calling choices on downstream conclusions.

The Biological Rationale for Subclonal Reconstruction

Most cancers arise from a single founder cell that accumulates somatic mutations through successive clonal expansions. As these expansions proceed, multiple subclones can coexist within a tumor, each sharing subsets of mutations that define their common ancestry and distinct evolutionary trajectories. The reconstruction of this subclonal architecture from sequencing data allows researchers to identify subclonal driver mutations, detect patterns of parallel evolution, compare mutational signatures between cellular populations, and characterize mechanisms of therapy resistance, spread, and metastasis.

The clinical relevance of subclonal reconstruction extends beyond basic tumor biology. Studies of colorectal cancer have demonstrated that lymphatic and distant metastases can arise from independent subclones within the primary tumor in a substantial proportion of cases, challenging the traditional view that lymph node metastases seed distant sites. In one analysis of 213 archival biopsy samples from 17 patients, researchers used somatic variants in hypermutable DNA regions to reconstruct high-confidence phylogenetic trees, finding that in 65 percent of cases lymphatic and distant metastases arose from independent subclones in the primary tumor, while in 35 percent of cases they shared a common subclonal origin. These findings have direct implications for staging systems and surgical decision-making, as they suggest that lymph node status may not always predict the route of distant spread.

For researchers planning subclonal reconstruction experiments, the first decision point is whether the biological question requires bulk sequencing, multi-region sampling, or single-cell approaches. Bulk sequencing of a single tumor sample provides a snapshot of the overall subclonal composition but systematically underestimates intra-tumoral heterogeneity. Multi-region sampling captures spatial heterogeneity and provides more complete evolutionary histories. Single-cell sequencing offers the highest resolution but introduces distinct technical challenges including dropout, amplification bias, and higher per-cell costs. The choice among these approaches should be driven by the specific research question, available sample material, and budget constraints.

Core Principles of Subclonal Reconstruction

Subclonal reconstruction rests on several foundational principles that govern how somatic variant data are transformed into estimates of tumor cell populations. Understanding these principles is essential for interpreting the output of any reconstruction algorithm and for troubleshooting unexpected results.

Variant Allele Frequencies and Cancer Cell Fractions

The variant allele frequency (VAF) of a somatic mutation represents the proportion of sequenced reads that carry the alternative allele at a given genomic position. For a clonal mutation present in every tumor cell, the expected VAF depends on the purity of the tumor sample, the local copy number state, and the ploidy of the genome. For a heterozygous mutation in a diploid region of a pure tumor sample, the expected VAF is 0.5. As tumor purity decreases, the observed VAF decreases proportionally because normal cell contamination dilutes the variant reads.

The cancer cell fraction (CCF) is the proportion of tumor cells that carry a particular mutation. Converting VAF to CCF requires accounting for tumor purity, local copy number, and the multiplicity of the mutation, which is the number of copies of the mutant allele present in each tumor cell. This conversion is mathematically straightforward for simple cases but becomes complex in regions with copy number alterations, loss of heterozygosity, or whole-genome duplication events.

For a mutation in a diploid region with copy number neutral status, the relationship between VAF, purity, and CCF can be expressed as follows. If p is the tumor purity, the observed VAF for a mutation present in fraction f of tumor cells with multiplicity m is approximately p multiplied by m multiplied by f, divided by the total copy number contribution from tumor and normal cells. In practice, most reconstruction algorithms perform this calculation internally, but researchers should understand the underlying logic to interpret outputs correctly and to identify cases where assumptions may be violated.

Clustering of Variant Allele Frequencies

The central computational task in subclonal reconstruction is clustering somatic mutations into groups that share similar cancer cell fractions. Mutations that arose in the same subclone and were inherited by all descendant cells will have similar CCFs, forming distinct clusters in the distribution of CCF values. Clonal mutations present in all tumor cells will cluster at CCF values near 1.0, while subclonal mutations will cluster at lower values corresponding to the fraction of tumor cells that carry them.

Several algorithms implement this clustering approach with different statistical frameworks. PyClone uses a Bayesian Dirichlet process clustering model that incorporates copy number information and estimates the cellular prevalence of each mutation cluster. SciClone uses a variational Bayesian mixture model that clusters VAFs directly, with optional correction for copy number alterations. Both tools require careful input preparation and produce outputs that must be interpreted in the context of the underlying assumptions.

The number of clusters detected depends on the depth of sequencing, the number of mutations available for analysis, and the true subclonal architecture of the tumor. Low sequencing depth reduces the precision of VAF estimates and can obscure clusters that are close together. Sparse mutation counts limit the statistical power to resolve distinct clusters. Tumors with many small subclones may produce VAF distributions that are difficult to separate into discrete clusters, requiring more sophisticated approaches or additional data.

Phylogenetic Tree Construction

Once mutations are assigned to clusters representing subclones, the next step is to reconstruct the evolutionary relationships among these subclones. Phylogenetic trees for tumors are rooted at the normal cell population and branch at points where new subclones arise from ancestral populations. The tree structure encodes the order of mutation acquisition and the lineage relationships among subclones.

The infinite sites assumption underlies most phylogenetic reconstruction methods. This assumption states that each genomic position mutates at most once during tumor evolution, so mutations are never lost or reverted. Under this assumption, mutations shared between two subclones must have been present in their common ancestor, and the tree can be reconstructed from the pattern of shared and private mutations. Violations of the infinite sites assumption can occur through loss of heterozygosity, gene conversion, or back mutations, and these events must be handled carefully during tree construction.

For bulk sequencing data, phylogenetic trees are typically constructed from the cluster assignments and the hierarchical relationships implied by the CCF values. A subclone with a higher CCF that shares mutations with a subclone at lower CCF is inferred to be ancestral. For multi-region or single-cell data, more sophisticated tree-building algorithms can be applied that use the full pattern of mutation presence and absence across samples or cells.

Data Inputs and Quality Requirements

The quality of subclonal reconstruction depends critically on the quality of the input data. Researchers must understand the requirements for sequencing depth, sample purity, and variant calling accuracy before initiating a reconstruction analysis.

Sequencing Depth and Coverage Considerations

Subclonal reconstruction requires sufficient sequencing depth to detect and accurately estimate the VAFs of mutations present in small fractions of tumor cells. For bulk whole-genome sequencing, typical depths of 30 to 60-fold coverage provide adequate power for detecting clonal mutations but may miss subclones present at very low frequencies. Deeper sequencing, such as 100 to 500-fold coverage with targeted panels or whole-exome capture, improves sensitivity for low-frequency variants but increases cost and may introduce artifacts from PCR duplication or sequencing errors.

The relationship between sequencing depth and the ability to resolve subclones is governed by binomial sampling statistics. At a given depth d, the standard error of a VAF estimate is approximately the square root of VAF multiplied by (1 minus VAF) divided by d. For a mutation present at 10 percent VAF with 100-fold coverage, the standard error is approximately 3 percent, which may be too large to distinguish this mutation from a neighboring cluster at 15 percent VAF. Increasing depth to 500-fold reduces the standard error to approximately 1.3 percent, improving cluster resolution.

Researchers should also consider the number of mutations available for clustering. Algorithms such as PyClone and SciClone require a minimum number of mutations to produce stable cluster assignments. In practice, at least 50 to 100 high-confidence somatic mutations are recommended for reliable reconstruction, with more mutations providing better resolution. Tumors with low mutation burden, such as pediatric cancers or some hematologic malignancies, may not provide enough mutations for robust clustering, and alternative approaches may be needed.

Tumor Purity and Normal Contamination

Tumor purity, defined as the proportion of cells in the sample that are tumor cells, directly affects the observed VAF distribution and the ability to detect subclones. Low purity samples have compressed VAF distributions because normal cell contamination dilutes all variant reads proportionally. This compression reduces the separation between clusters and can make subclonal populations difficult to distinguish.

Estimating tumor purity is an essential preprocessing step for subclonal reconstruction. Several methods exist for purity estimation, including computational approaches that analyze the overall VAF distribution, copy number profiles, or allele-specific copy number states. Pathological assessment of tumor content from stained sections provides an independent estimate but may not reflect the molecular composition of the sequenced sample. Discrepancies between computational and pathological purity estimates should be investigated before proceeding with reconstruction.

For samples with very low purity, typically below 20 to 30 percent tumor content, subclonal reconstruction becomes increasingly unreliable. The VAF distribution becomes compressed into a narrow range, and the signal from minor subclones may be indistinguishable from noise. In such cases, researchers should consider enrichment strategies such as macrodissection, laser capture microdissection, or cell sorting to increase tumor content before sequencing.

Germline Variant Filtering

Accurate germline variant filtering is essential for subclonal reconstruction because germline polymorphisms can be mistaken for somatic mutations and distort the inferred subclonal architecture. Germline variants are present in all cells of the body, including normal cells, and their VAFs reflect the genotype instead of the tumor subclonal composition.

The standard approach to germline filtering is to sequence a matched normal sample from the same individual, typically from blood or adjacent normal tissue. Somatic variant callers compare the tumor and normal samples and retain variants that are present in the tumor but absent from the normal sample. This approach requires careful quality control to ensure that the normal sample is truly normal and that the comparison is not confounded by contamination, sample mix-ups, or clonal hematopoiesis.

When a matched normal sample is unavailable, researchers must rely on computational filtering strategies. These strategies include removing variants present in population databases such as gnomAD, filtering variants at known germline polymorphism positions, and applying statistical tests to distinguish somatic mutations from germline heterozygotes based on VAF distributions. These approaches are less reliable than matched normal filtering and should be used with caution, particularly for variants in repetitive regions or genes with pseudogenes.

The NCBI provides access to databases and search systems that support variant annotation and filtering, including reference genomes, population variant databases, and sequence analysis tools. Researchers should consult these resources when designing germline filtering strategies and when interpreting variants that fall in regions of the genome with complex architecture.

Variant Calling Workflow for Subclonal Reconstruction

The variant calling workflow is the foundation upon which subclonal reconstruction is built. Errors introduced during variant calling propagate through the reconstruction analysis and can substantially alter the inferred subclonal architecture.

Alignment and Preprocessing

The first step in the variant calling workflow is aligning sequencing reads to a reference genome. The choice of aligner, reference genome version, and alignment parameters affects the accuracy of variant detection, particularly in regions with repetitive sequence, structural variation, or high sequence divergence from the reference.

After alignment, several preprocessing steps are typically applied before variant calling. These include marking or removing PCR duplicates, which can inflate the apparent depth at specific positions and bias VAF estimates. Base quality score recalibration adjusts quality scores based on empirical error patterns, improving the accuracy of variant calling. Indel realignment, where supported by the aligner or downstream tools, corrects misalignments around insertions and deletions that can create false variant calls.

The choice of reference genome version is important for reproducibility and for downstream annotation. Different reference versions can produce different variant calls at the same genomic position, particularly in regions that have been updated or corrected between versions. Researchers should document the reference version used and ensure that all downstream analyses use the same version.

Somatic Variant Calling Strategies

Somatic variant callers compare tumor and normal samples to identify variants that are present in the tumor but absent from the normal sample. Different callers use different statistical models and heuristics, and the choice of caller can substantially affect the results of subclonal reconstruction.

An evaluation of sixteen pipelines for reconstructing the evolutionary histories of 293 localized prostate cancers from single samples found that predictions of subclonal architecture and timing of somatic mutations varied extensively across pipelines. Pipelines incorporating certain variant callers preferentially predicted homogeneous cancer cell populations, while those using other callers tended to predict multiple populations of cancer cells. These biases suggest that researchers should be cautious in interpreting specific architectures and subclonal variants, and should consider running multiple variant callers and comparing results.

For subclonal reconstruction, the key requirements for a somatic variant caller are accurate VAF estimation, high sensitivity for subclonal mutations, and low false-positive rates. Callers that use tumor-normal pairs with matched normal samples generally provide better specificity than callers that analyze tumor samples alone. However, the sensitivity for low-VAF mutations varies among callers, and some callers may systematically underestimate or overestimate VAFs for certain mutation types or genomic contexts.

Variant Filtering and Quality Control

After variant calling, a series of filtering steps is applied to remove false positives and retain high-confidence somatic mutations for subclonal reconstruction. These filters typically include thresholds for read depth, variant allele frequency, mapping quality, base quality, and strand bias. Additional filters may remove variants in repetitive regions, variants near indels, or variants with evidence of sequencing artifacts.

The choice of filtering thresholds involves a trade-off between sensitivity and specificity. Stringent filters remove more false positives but may also remove true subclonal mutations, reducing the power to detect minor subclones. Relaxed filters retain more mutations but increase the risk that false positives distort the inferred subclonal architecture. Researchers should evaluate the impact of filtering choices on the final reconstruction and document the filtering strategy used.

For long-read single-cell RNA sequencing data, a computational workflow called LongSom has been developed to call de novo somatic single-nucleotide variants, including mitochondrial variants, copy-number alterations, and gene fusions, without the need for matched normal samples. This workflow applies an extensive set of hard filters and statistical tests to distinguish somatic variants from noise and germline polymorphisms. In human ovarian cancer samples, LongSom detected clinically relevant somatic variants that were validated against matched DNA samples, and leveraging these variants and fusions identified subclones with different predicted treatment outcomes. This approach demonstrates the potential for subclonal reconstruction from RNA sequencing data, though it requires careful validation and has distinct limitations compared to DNA-based approaches.

Tools and Algorithms for Subclonal Reconstruction

Several computational tools are available for subclonal reconstruction, each with distinct strengths, limitations, and input requirements. The choice of tool should be guided by the data type, the biological question, and the computational resources available.

PyClone

PyClone is a Bayesian clustering tool that estimates the cellular prevalence of somatic mutations and groups mutations into clusters representing subclonal populations. The tool requires input files containing mutation information, including chromosome, position, reference and alternate alleles, read counts, and copy number information for each mutation.

The PyClone model incorporates copy number alterations and tumor purity to convert VAFs into cellular prevalence estimates. The Dirichlet process clustering allows the number of clusters to be inferred from the data instead of specified in advance. PyClone outputs a set of clusters with estimated cellular prevalence values and posterior probabilities for each cluster assignment.

PyClone is particularly well-suited for whole-genome or whole-exome sequencing data with matched normal samples and accurate copy number calls. The tool requires careful input preparation, including accurate copy number segmentation and allele-specific copy number estimates. Errors in copy number calling can propagate through the PyClone model and produce incorrect cluster assignments.

SciClone

SciClone is a variational Bayesian mixture model that clusters somatic mutations based on their VAFs. The tool is designed for single-nucleotide variants and can optionally incorporate copy number information to correct VAFs for local copy number alterations.

SciClone requires input files containing chromosome, position, and read counts for the reference and alternate alleles. The tool performs clustering using a beta-binomial mixture model that accounts for overdispersion in VAF estimates. SciClone outputs cluster assignments, cluster VAFs, and visualizations of the VAF distribution with cluster assignments overlaid.

SciClone is computationally efficient and can handle large numbers of mutations. The tool is well-suited for targeted sequencing data and for analyses where copy number information is limited or unavailable. However, the lack of explicit copy number correction in the default mode may lead to inaccurate cluster assignments for mutations in regions with copy number alterations.

Additional Tools and Approaches

Beyond PyClone and SciClone, numerous other tools have been developed for subclonal reconstruction, each with different modeling assumptions and output formats. Some tools focus on integrating multiple data types, such as copy number alterations and structural variants, while others specialize in multi-region or single-cell data.

The evaluation of simulation methods for tumor subclonal reconstruction highlights the complexity of benchmarking these algorithms. Most tumors originate from a single cell, and their evolution can be traced through lineages characterized by common alterations including small somatic mutations, copy number alterations, structural variants, and aneuploidies. Due to the complexity of these alterations and the errors introduced by sequencing protocols and calling algorithms, subclonal reconstruction algorithms are necessary to recapitulate the DNA sequence composition and tumor evolution in silico. With a growing number of algorithms available, consistent and comprehensive benchmarking is needed, which relies on realistic tumor sequencing generated by simulation tools.

For researchers selecting a reconstruction tool, the following considerations are important. First, the tool should be compatible with the data type and variant calling pipeline used. Second, the tool should be well-documented and maintained, with clear input and output formats. Third, the tool should be validated on data similar to the researcher's own data, either through published benchmarks or through internal validation with simulated or well-characterized samples.

Practical Workflow for Subclonal Reconstruction

The following workflow outlines the practical steps for performing subclonal reconstruction from somatic variant data. This workflow assumes bulk sequencing data with a matched normal sample, which is the most common scenario for initial reconstruction analyses.

Step 1: Data Preparation and Quality Assessment

Before initiating subclonal reconstruction, researchers should assess the quality of the sequencing data and the variant calls. This assessment includes reviewing alignment statistics, coverage distributions, and variant call quality metrics. Samples with low coverage, high duplication rates, or evidence of contamination should be flagged for further investigation.

The variant call format (VCF) file containing somatic mutations should be filtered to retain high-confidence variants. The filtering strategy should be documented and applied consistently across samples. For each retained variant, the following information should be available: chromosome, position, reference allele, alternate allele, read depth, alternate allele count, and quality scores.

Step 2: Copy Number Analysis

Copy number alterations affect the relationship between VAF and cancer cell fraction, so accurate copy number information is essential for subclonal reconstruction. Copy number analysis should be performed using tools that provide allele-specific copy number estimates, including total copy number and the number of copies of each parental allele.

For whole-genome sequencing data, copy number can be inferred from read depth and allele balance across the genome. For whole-exome or targeted sequencing data, copy number analysis is more challenging due to the non-uniform coverage and the limited number of informative positions. In such cases, researchers may need to rely on external copy number information or use tools that can accommodate uncertainty in copy number estimates.

Step 3: Input File Preparation for Reconstruction Tools

Each reconstruction tool has specific input format requirements. For PyClone, the input file should contain mutation information including read counts and copy number states. For SciClone, the input file should contain read counts for the reference and alternate alleles.

Researchers should carefully review the documentation for their chosen tool to ensure that input files are formatted correctly. Errors in input file formatting are a common cause of tool failures and incorrect results. The Bioconductor project provides documentation and workflows for reproducible genomic analysis, including packages that support subclonal reconstruction and related analyses.

Step 4: Running the Reconstruction Analysis

The reconstruction analysis should be run with parameters appropriate for the data and the biological question. For PyClone, the number of iterations and the burn-in period should be set to ensure convergence of the Markov chain Monte Carlo sampler. For SciClone, the maximum number of clusters should be specified, and the minimum cluster size should be set to avoid spurious clusters.

After running the analysis, researchers should examine the output for evidence of convergence and stability. Multiple runs with different random seeds should produce similar results. If the results vary substantially between runs, the analysis may not be stable, and the input data or parameters should be reviewed.

Step 5: Interpretation and Visualization

The output of subclonal reconstruction should be interpreted in the context of the biological question and the limitations of the data. Cluster assignments should be reviewed for biological plausibility, and mutations within each cluster should be examined for known driver genes or functional categories.

Visualization of the VAF distribution with cluster assignments overlaid provides a useful check on the reconstruction. Clusters should be visually distinct, and the cluster VAFs should be consistent with the underlying data. Phylogenetic trees should be examined for consistency with the cluster assignments and with known biology of the tumor type.

Step 6: Validation and Sensitivity Analysis

Validation of subclonal reconstruction results is essential for drawing reliable conclusions. Several validation approaches are available, including comparison with multi-region sampling data, comparison with single-cell sequencing data, and analysis of simulated data with known subclonal architecture.

Sensitivity analysis should be performed to assess the impact of key parameters and assumptions on the reconstruction. This analysis includes varying the variant calling thresholds, the copy number estimates, and the reconstruction tool parameters. If the inferred subclonal architecture changes substantially with reasonable parameter variations, the results should be interpreted with caution.

At a Glance: Key Decisions in Subclonal Reconstruction

Decision PointPrimary OptionsKey ConsiderationsCommon Pitfalls
Sequencing strategyBulk single sample, multi-region, single-cellSingle-sample underestimates heterogeneity, multi-region captures spatial diversity, single-cell provides highest resolution but higher cost and technical noiseAssuming single-sample reconstruction captures full heterogeneity
Variant caller selectionMultiple callers available with different biasesSome callers bias toward homogeneous populations, others toward multiple populations, consider running multiple callersRelying on a single caller without cross-validation
Copy number inputAllele-specific from WGS, estimated from exome, external dataCopy number errors propagate through reconstruction, allele-specific estimates preferredUsing total copy number without allele specificity
Reconstruction toolPyClone, SciClone, othersPyClone incorporates copy number and purity, SciClone is efficient for targeted dataChoosing a tool without considering input requirements and assumptions
Validation approachMulti-region comparison, single-cell validation, simulationMulti-region sampling identifies more populations than single-sample reconstructionInterpreting single-sample results as complete architecture

Records and Measurements for Subclonal Reconstruction

Systematic record-keeping is essential for reproducible subclonal reconstruction. Researchers should maintain detailed records of all analysis steps, parameters, and quality metrics to enable replication and troubleshooting.

Essential Records

The following records should be maintained for each subclonal reconstruction analysis. First, the sequencing metadata including sample identifiers, sequencing platform, read length, and coverage. Second, the alignment and preprocessing parameters including reference genome version, aligner version, and preprocessing steps applied. Third, the variant calling parameters including caller version, filtering thresholds, and the number of variants retained at each filtering stage. Fourth, the copy number analysis parameters and the copy number estimates used for reconstruction. Fifth, the reconstruction tool parameters including version, clustering parameters, and convergence diagnostics. Sixth, the output files including cluster assignments, cluster VAFs or cellular prevalence estimates, and phylogenetic trees.

Quality Metrics to Track

Several quality metrics should be tracked throughout the analysis to identify potential problems. These metrics include the number of somatic mutations passing each filtering stage, the distribution of VAFs before and after correction, the number of clusters identified, the cluster sizes, and the stability of cluster assignments across runs. For PyClone, convergence diagnostics should be recorded, including trace plots and effective sample sizes. For SciClone, the model evidence and cluster separation metrics should be documented.

Reproducibility Practices

Reproducibility requires more than recording parameters. Researchers should use version control for analysis scripts and document the computational environment, including software versions and dependencies. Workflow management systems can help ensure that analyses are run consistently and that results are traceable to specific inputs and parameters.

The Galaxy Training Network provides accessible workflow training and analysis tutorials that emphasize reproducibility in genomic analysis. The nf-core documentation describes community pipeline standards for reproducible workflows, including configuration and usage guidelines. The Carpentries lessons provide foundational training in computing, data, shell, Git, and programming that supports reproducible research practices. Researchers should consult these resources when establishing their analysis workflows.

Common Failure Patterns and Troubleshooting

Subclonal reconstruction analyses frequently encounter specific failure patterns. Recognizing these patterns and understanding their causes can help researchers troubleshoot problems and interpret results correctly.

Failure Pattern 1: Excessive Number of Clusters

When a reconstruction analysis produces many more clusters than expected based on the biological context, several causes are possible. First, the variant calling may have retained false positives that form spurious clusters. Second, copy number errors may have produced incorrect VAF corrections, splitting true clusters into multiple apparent clusters. Third, the clustering algorithm may be overfitting the data, particularly when the number of mutations is small relative to the number of clusters.

Troubleshooting steps include reviewing the variant calls for evidence of artifacts, examining the copy number estimates for regions containing mutations in small clusters, and adjusting the clustering parameters to be more conservative. If the excessive clusters persist, the analysis may need to be repeated with more stringent variant filtering or with a different reconstruction tool.

Failure Pattern 2: Single Cluster with All Mutations

The opposite failure pattern occurs when all mutations are assigned to a single cluster, suggesting a homogeneous tumor with no subclonal structure. This pattern can arise when the tumor is truly homogeneous, but it can also result from technical issues. Low sequencing depth reduces VAF precision and can obscure true subclonal structure. Low tumor purity compresses the VAF distribution and makes clusters difficult to separate. Variant calling biases that preferentially detect clonal mutations can also produce an apparent single cluster.

Troubleshooting steps include examining the VAF distribution for evidence of multiple modes, assessing the sequencing depth and purity, and considering whether the variant calling pipeline may have missed subclonal mutations. If the tumor is expected to be heterogeneous based on biological knowledge, the analysis may need to be repeated with deeper sequencing or with a more sensitive variant caller.

Failure Pattern 3: Unstable Cluster Assignments Across Runs

When cluster assignments vary substantially across runs with different random seeds, the reconstruction is not stable. This instability often indicates that the data do not strongly support a unique clustering solution. Possible causes include insufficient mutations for reliable clustering, clusters that are too close together to be resolved, or model misspecification.

Troubleshooting steps include increasing the number of iterations for Bayesian methods, examining the posterior distribution of cluster assignments, and assessing whether the data support the number of clusters identified. If the instability persists, the analysis may need to be simplified by using fewer clusters or by combining data from multiple samples.

Failure Pattern 4: Phylogenetic Trees with Contradictory Relationships

Phylogenetic trees that contain contradictory relationships, such as subclones that appear to be both ancestral and descendant of each other, indicate problems with the cluster assignments or the tree construction. These contradictions can arise when mutations are incorrectly assigned to clusters, when copy number errors distort CCF estimates, or when the infinite sites assumption is violated.

Troubleshooting steps include reviewing the cluster assignments for individual mutations, examining the CCF estimates for consistency with the tree structure, and considering whether copy number alterations or loss of heterozygosity events may have affected the analysis. In some cases, the tree construction algorithm may need to be adjusted to accommodate violations of the infinite sites assumption.

Limitations and Interpretation Boundaries

Subclonal reconstruction has inherent limitations that researchers must understand to interpret results correctly and avoid overinterpretation.

Single-Sample Reconstruction Underestimates Heterogeneity

A fundamental limitation of single-sample reconstruction is that it systematically underestimates intra-tumoral heterogeneity. An evaluation of eighteen pipelines for reconstructing the evolutionary histories of ten tumors with multi-region sampling confirmed that single-sample reconstructions predict on average fewer than half of the cancer cell populations identified by multi-region sequencing. This limitation means that the absence of detected subclones in a single-sample analysis does not demonstrate that the tumor is homogeneous.

Researchers should consider whether their biological question requires multi-region sampling. If the goal is to understand the full extent of intra-tumoral heterogeneity, single-sample reconstruction is insufficient. If the goal is to identify major subclonal populations or to compare heterogeneity across tumors, single-sample reconstruction may be adequate, provided the limitations are acknowledged.

Variant Calling Biases Propagate to Reconstruction

The evaluation of sixteen pipelines for reconstructing the evolutionary histories of 293 localized prostate cancers found that predictions of subclonal architecture and timing of somatic mutations vary extensively across pipelines. Pipelines show consistent types of biases, with some preferentially predicting homogeneous cancer cell populations and others tending to predict multiple populations of cancer cells. These biases suggest caution in interpreting specific architectures and subclonal variants.

Researchers should consider running multiple variant calling pipelines and comparing the resulting reconstructions. If the subclonal architecture is consistent across pipelines, confidence in the results is increased. If the architecture varies substantially, the results should be interpreted as pipeline-dependent and the biological conclusions should be tempered accordingly.

Copy Number Complexity Limits Accuracy

Tumors with complex copy number landscapes present substantial challenges for subclonal reconstruction. Regions with high-level amplifications, homozygous deletions, or complex rearrangements produce VAF distributions that are difficult to interpret. The copy number estimates used for VAF correction may be inaccurate in these regions, leading to incorrect CCF estimates and cluster assignments.

For tumors with extensive copy number alterations, researchers should consider using reconstruction tools that explicitly model copy number states and their uncertainties. Validation of the reconstruction using independent approaches, such as fluorescence in situ hybridization or single-cell sequencing, may be necessary to confirm the inferred subclonal architecture.

Timing Estimates Have Wide Uncertainty

Some reconstruction methods provide estimates of the timing of somatic mutations relative to tumor evolution, such as whether a mutation occurred early or late in the tumor's history. These timing estimates are based on the fraction of tumor cells carrying the mutation and the assumption that mutations present in more cells occurred earlier. However, these estimates have wide uncertainty and can be biased by the same factors that affect CCF estimation.

Researchers should treat timing estimates as approximate and should not overinterpret small differences in estimated timing between mutations. The timing of mutations should be validated using independent approaches where possible, such as analysis of mutation signatures or integration with clinical data.

Safety and Ethical Context

Subclonal reconstruction research involves human samples and genetic data, raising important ethical and regulatory considerations. Researchers must ensure that their studies comply with applicable regulations and institutional requirements.

Informed Consent and Data Protection

Human tissue samples used for subclonal reconstruction must be obtained with appropriate informed consent. The consent process should cover the intended research uses, including genomic analysis and data sharing. Researchers should verify that the consent documents for their samples permit the planned analyses and that any restrictions on data sharing are respected.

Genetic data derived from human samples are subject to data protection regulations in many jurisdictions. Researchers should ensure that their data storage, processing, and sharing practices comply with applicable laws and institutional policies. De-identification of samples and data should be performed according to established standards, and researchers should be aware that genomic data can potentially re-identify individuals even after de-identification.

Clinical Translation Considerations

When subclonal reconstruction results have potential clinical implications, such as the identification of subclonal driver mutations or the prediction of treatment resistance, researchers must be careful about the claims they make. The evaluation of subclonal reconstruction methods has shown substantial variability across pipelines, and the clinical significance of specific subclonal architectures is still being established.

Researchers should not make clinical recommendations based solely on subclonal reconstruction results without appropriate validation and clinical context. If the research has potential clinical implications, consultation with clinical colleagues and appropriate oversight bodies is recommended.

Professional Escalation Criteria

Researchers should seek additional expertise or escalate concerns in the following situations. First, if the subclonal reconstruction results are being used to guide clinical decisions, the results should be reviewed by qualified clinical personnel. Second, if the analysis reveals unexpected findings with potential clinical significance, such as germline variants in cancer predisposition genes, the findings should be handled according to institutional policies for incidental findings. Third, if the reconstruction results are inconsistent with established biological knowledge or with other experimental data, the analysis should be reviewed before drawing conclusions.

Frequently Asked Questions

What is the difference between variant allele frequency and cancer cell fraction?

Variant allele frequency (VAF) is the proportion of sequenced reads that carry the alternative allele at a genomic position. Cancer cell fraction (CCF) is the proportion of tumor cells that carry a particular mutation. The conversion from VAF to CCF requires accounting for tumor purity, local copy number, and the multiplicity of the mutation. For a heterozygous mutation in a diploid region of a pure tumor sample, the expected VAF is 0.5, corresponding to a CCF of 1.0. As purity decreases or copy number changes, the relationship between VAF and CCF becomes more complex.

How many somatic mutations are needed for reliable subclonal reconstruction?

The number of mutations required depends on the subclonal architecture of the tumor and the sequencing depth. In practice, at least 50 to 100 high-confidence somatic mutations are recommended for reliable clustering with tools such as PyClone and SciClone. Tumors with low mutation burden may not provide enough mutations for robust reconstruction, and the results should be interpreted with caution. More mutations generally provide better resolution of subclonal populations, but the quality of the mutations matters as much as the quantity.

Can subclonal reconstruction be performed without a matched normal sample?

Subclonal reconstruction can be performed without a matched normal sample, but the analysis is more challenging and less reliable. Without a matched normal, researchers must rely on computational filtering to distinguish somatic mutations from germline polymorphisms. This filtering typically uses population databases and statistical tests based on VAF distributions. The risk of retaining germline variants or removing true somatic mutations is higher without a matched normal. Some tools, such as LongSom for long-read single-cell RNA sequencing data, have been developed to call somatic variants without normal samples, but these approaches require careful validation.

What is the impact of tumor purity on subclonal reconstruction?

Tumor purity directly affects the observed VAF distribution and the ability to detect subclones. Low purity samples have compressed VAF distributions because normal cell contamination dilutes all variant reads proportionally. This compression reduces the separation between clusters and can make subclonal populations difficult to distinguish. For samples with very low purity, typically below 20 to 30 percent tumor content, subclonal reconstruction becomes increasingly unreliable. Researchers should estimate tumor purity before proceeding with reconstruction and consider enrichment strategies for low-purity samples.

How do copy number alterations affect subclonal reconstruction?

Copy number alterations affect the relationship between VAF and cancer cell fraction. In regions with copy number gains, the expected VAF for a clonal mutation depends on the number of copies of the mutant allele and the total copy number. In regions with loss of heterozygosity, the expected VAF for a clonal mutation may be 1.0 instead of 0.5. Accurate allele-specific copy number estimates are essential for converting VAFs to CCFs correctly. Errors in copy number calling can propagate through the reconstruction and produce incorrect cluster assignments.

Should I use PyClone or SciClone for my analysis?

The choice between PyClone and SciClone depends on the data type and the biological question. PyClone incorporates copy number information and tumor purity to estimate cellular prevalence, making it well-suited for whole-genome or whole-exome sequencing data with accurate copy number calls. SciClone uses a variational Bayesian mixture model to cluster VAFs directly and is computationally efficient, making it well-suited for targeted sequencing data and for analyses where copy number information is limited. Researchers should consider running both tools and comparing the results to assess the robustness of the inferred subclonal architecture.

How does multi-region sampling improve subclonal reconstruction?

Multi-region sampling captures spatial heterogeneity within a tumor and provides a more complete picture of the subclonal architecture. Single-sample reconstruction systematically underestimates intra-tumoral heterogeneity, predicting on average fewer than half of the cancer cell populations identified by multi-region sequencing. Multi-region sampling allows the detection of subclones that are present in only part of the tumor and provides information about the spatial distribution of subclones. However, multi-region sampling increases the cost and complexity of the analysis and requires careful coordination of sample collection and processing.

What validation approaches are recommended for subclonal reconstruction?

Several validation approaches are recommended for subclonal reconstruction. Comparison with multi-region sampling data can confirm whether single-sample reconstruction captured the major subclonal populations. Comparison with single-cell sequencing data provides the highest resolution validation but is expensive and technically challenging. Analysis of simulated data with known subclonal architecture can assess the accuracy of the reconstruction pipeline. Sensitivity analysis, varying key parameters and assessing the impact on the reconstruction, is essential for understanding the robustness of the results.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.