Detecting Somatic Structural Variants in Cancer Genomes: Challenges and Best Practices for Tumor-Normal Analysis

By Dr. Zubair Khalid, DVM, MS, PhD ·

Detecting Somatic Structural Variants in Cancer Genomes: Challenges and Best Practices for Tumor-Normal Analysis

Key Takeaways

  • Detecting somatic structural variants (SVs) requires robust tumor-normal analysis to distinguish acquired genomic rearrangements from inherited germline variations and sequencing artifacts.
  • Long-read sequencing platforms offer superior detection of complex SVs and repetitive regions compared to short-read technologies, necessitating specialized aligners (e.g., minimap2) and SV callers (e.g., SAVANA, Severus).
  • Somatic-specific SV callers, designed to analyze tumor-normal pairs and model tumor characteristics like aneuploidy and purity, achieve higher recall for true somatic events than generic germline callers.
  • Tumor purity significantly impacts sensitivity; low purity dilutes somatic variant allele frequencies, potentially masking true events and necessitating careful interpretation of negative results.
  • Multi-caller consensus approaches enhance precision and validation by leveraging complementary strengths of different SV detection algorithms, thereby increasing confidence in identified somatic SVs.
  • Reproducibility in somatic SV analysis is paramount, requiring meticulous record-keeping of sample metadata, pipeline configurations, quality metrics, and post-processing steps, often facilitated by workflow management systems.

Somatic structural variants (SVs) are large genomic rearrangements, including deletions, duplications, inversions, insertions, and translocations, that arise in tumor tissue and are not present in the germline. Identifying these variants from tumor-normal sequencing data requires a dedicated analytical workflow that subtracts inherited variation, accounts for tumor-specific properties such as aneuploidy and purity, and distinguishes true somatic events from sequencing and alignment artifacts. This article provides a practical framework for researchers and laboratory professionals who need to build, validate, and troubleshoot somatic SV detection pipelines using paired tumor-normal samples.

The core problem is straightforward to state but difficult to solve in practice. A tumor genome contains both germline variants inherited from the patient and somatic variants acquired during cancer development. The analytical goal is to isolate the somatic fraction with high confidence while avoiding two types of error: reporting germline variants as somatic and reporting artifacts as genuine rearrangements. The approaches described here cover data preparation, caller selection, tumor-normal subtraction, quality control, and interpretation of results within the limits of current technology.

At a Glance

The table below summarizes the key decisions and considerations for a somatic SV detection workflow. Each row addresses a distinct stage of the pipeline and the primary factors that influence outcome.

Workflow StagePrimary DecisionKey Consideration
Sequencing platformShort-read versus long-readLong-read platforms detect more SVs and resolve complex rearrangements, but require different aligners and callers
AlignmentRead mapper selectionAligner choice affects SV caller performance, long-read data requires long-read-specific aligners
SV callingGeneric versus somatic callerSomatic callers use tumor-normal pairs directly and achieve higher recall for somatic events
Tumor-normal subtractionPaired calling versus separate calling with subtractionSeparate calling in tumor and normal followed by subtraction is a common framework, but joint approaches reduce false positives
ValidationMulti-caller consensusCombining multiple callers improves precision and helps confirm candidate somatic SVs
Quality assessmentPurity and ploidy estimationLow tumor purity reduces sensitivity, aneuploidy complicates copy-number-based filtering

Understanding Somatic Structural Variants in Cancer Genomes

Structural variants are a major class of genomic alteration in cancer. Unlike single nucleotide variants and small indels, SVs can span thousands to millions of base pairs and can rearrange the architecture of entire chromosomes. These large-scale changes can disrupt tumor suppressor genes, create fusion genes that drive oncogenic signaling, place oncogenes under the control of new regulatory elements, and contribute to the genomic instability that characterizes many cancers.

The biological significance of SVs in cancer is well established. Lung cancer is a heterogeneous disease and the primary cause of cancer-related mortality worldwide, and somatic mutations including large structural variants serve as important biomarkers for selecting targeted therapy. Genomic studies of lung cancer have historically relied on short-read sequencing, but emerging long-read technologies offer a promising alternative for studying somatic SVs because they can span repetitive regions and resolve complex rearrangements that short reads cannot fully assemble.

The challenge for bioinformaticians is that SVs are inherently more difficult to detect than smaller variants. A deletion of 50 base pairs may be detectable through paired-end read analysis, but a balanced translocation that exchanges large chromosome arms requires evidence from reads that span the breakpoint junction. The accuracy of SV detection depends on sequencing depth, read length, library insert size, alignment quality, and the algorithmic approach used by the variant caller.

Core Principles of Tumor-Normal Subtraction

The fundamental principle of somatic variant calling is subtraction. A tumor sample contains a mixture of somatic and germline variation. A matched normal sample from the same individual contains the germline baseline. By comparing the two, the analyst can identify variants that are present only in the tumor and therefore likely somatic.

This subtraction can be implemented in two ways. The first approach calls variants independently in the tumor and normal samples, then removes any variant that appears in both. The second approach uses a joint calling algorithm that analyzes the tumor-normal pair simultaneously and assigns a somatic status to each candidate variant based on the evidence in both samples.

The choice between these approaches has practical consequences. Independent calling followed by subtraction is conceptually simple and allows the analyst to use any SV caller, including those designed for germline analysis. However, this approach can miss somatic variants that are present at low allele frequency in the tumor or that are obscured by noise in the normal sample. Joint calling methods are designed to model the tumor-normal relationship explicitly and can achieve higher sensitivity, but they require callers that support paired analysis.

A study of long-read somatic SV calling in lung cancer compared generic callers used in germline mode with dedicated somatic callers. The somatic callers achieved a mean recall of 72 percent for somatic events, while generic callers achieved only 53 percent. This difference demonstrates that the choice of calling strategy has a substantial impact on the number of true somatic variants identified.

Sequencing Platforms and Their Impact on SV Detection

Short-Read Sequencing

Short-read sequencing generates reads of 150 to 300 base pairs and has been the workhorse of cancer genomics for over a decade. The advantages of short-read platforms include high throughput, low per-base cost, and mature analytical tools. For SV detection, short reads provide evidence through paired-end mapping, where the distance and orientation of the two reads in a pair reveal the presence of a structural rearrangement, and through split-read analysis, where a single read maps to two distant genomic locations.

The limitations of short-read sequencing for SV detection are well documented. Repetitive regions of the genome, including segmental duplications and centromeric sequences, produce ambiguous alignments that confound SV callers. Large insertions and complex rearrangements may not be fully resolved because the breakpoint junction is longer than the read length. Despite these limitations, short-read sequencing remains a practical and widely used approach, and many somatic SV callers were originally developed for this data type.

Long-Read Sequencing

Long-read sequencing platforms produce reads of 10 kilobases or more, allowing reads to span entire repetitive elements and complex rearrangement junctions. This capability makes long-read sequencing particularly valuable for SV discovery. Studies comparing long-read and short-read approaches have found that long-read sequencing identifies more somatic SV events than short-read sequencing when both are applied to the same samples.

The adoption of long-read sequencing for somatic SV calling requires new analytical considerations. The choice of aligner has a significant influence on downstream variant detection. A benchmarking study of long-read SV callers in cancer genomes found that different combinations of aligners and variant callers influenced somatic SV detection in terms of variant type, size, sensitivity, and accuracy. The study used the minimap2 aligner with several somatic callers and found that SAVANA and Severus achieved the highest recall at 79.5 percent and 79.25 percent respectively, followed by nanomonsv at 72.5 percent.

Long-read data also enables new calling strategies. One recent approach, colorSV, uses a co-assembly method that builds a joint assembly graph from matched tumor and normal samples. By examining the local topology of this graph, the method identifies long-range SVs that are present in the tumor but absent from the normal sample. This co-assembly approach demonstrated near-perfect precision and sensitivity for calling translocations on the COLO829 cell line, outperforming four existing somatic SV callers in both metrics.

Selecting an SV Caller for Somatic Analysis

Generic Callers

Generic SV callers are designed to identify structural variants in any sample, regardless of whether the sample is germline or somatic. These tools typically analyze a single sample and report variants based on read-pair, split-read, and assembly evidence. When applied to tumor-normal analysis, generic callers are run separately on each sample, and the results are subtracted to identify somatic candidates.

The advantage of generic callers is their maturity and broad applicability. Tools such as DELLY can be run in a generic mode and are widely used in germline and somatic studies. However, the benchmarking study of long-read SV calling found that generic callers achieved lower recall for somatic events than dedicated somatic callers. This lower recall likely reflects the fact that generic callers are optimized for diploid germline samples and do not model the complexities of tumor genomes, including aneuploidy, subclonal variation, and contamination with normal cells.

Somatic Callers

Somatic SV callers are specifically designed to analyze tumor-normal pairs. These tools model the expected allele frequencies in both samples and assign a somatic likelihood to each candidate variant. Some somatic callers, such as Manta, were developed for rapid analysis of cancer sequencing data and can call structural variants, medium-sized indels, and large insertions on standard compute hardware in less than a tenth of the time required by comparable methods.

Manta is notable for its speed and its ability to assemble a high fraction of its calls to base-pair resolution. This base-pair resolution is important for downstream annotation and analysis of clinical significance, because it allows the exact breakpoint sequence to be determined. Manta was released as an open-source community resource to facilitate routine structural variant analysis in clinical and research sequencing scenarios.

The benchmarking study of long-read somatic SV callers found that somatic callers achieved higher recall than generic callers, with SAVANA and Severus performing best when used with the minimap2 aligner. These results suggest that researchers working with long-read data should prioritize somatic callers over generic tools when the goal is to maximize detection of true somatic events.

Multi-Caller Consensus

Given the variability among SV callers, a common strategy is to combine the results of multiple tools. A study that evaluated eight widely used SV callers on paired tumor and normal samples from lung cancer and melanoma cell lines found that combining multiple tools and testing different combinations significantly enhanced the validation of somatic alterations. The study used a VCF merging procedure followed by a subtraction method to identify candidate somatic SVs and explored different combinations of tools to enhance accuracy.

The rationale for multi-caller consensus is that different callers use different algorithmic strategies and therefore have complementary strengths and weaknesses. A variant detected by multiple independent callers is more likely to be genuine than a variant detected by a single caller. However, consensus approaches can also reduce sensitivity if the callers share systematic biases, and the computational cost of running multiple callers must be weighed against the benefit of increased confidence.

Handling Aneuploidy and Tumor Purity

Tumor Purity

Tumor purity refers to the proportion of cancer cells in a tumor sample, with the remainder consisting of stromal cells, immune cells, and other normal tissue. Purity has a direct impact on somatic variant detection because the allele frequency of a somatic variant is diluted by the presence of normal cells. A homozygous somatic deletion in a tumor with 50 percent purity will be present in only 50 percent of the sequencing reads, making it difficult to distinguish from a heterozygous germline variant.

Low tumor purity reduces the sensitivity of somatic SV calling because the evidence for a variant may fall below the detection threshold of the caller. Researchers should estimate tumor purity before running somatic SV callers and interpret negative results with caution when purity is low. Purity can be estimated from histology, from the allele frequencies of somatic single nucleotide variants, or from copy-number analysis.

Aneuploidy

Aneuploidy, the presence of an abnormal number of chromosomes, is a common feature of cancer genomes. Aneuploidy complicates SV calling in several ways. First, copy-number changes can create apparent breakpoints that are not true structural rearrangements. Second, the expected allele frequencies for somatic variants are altered by chromosome-level gains and losses. Third, aneuploidy can affect the alignment of reads in regions of segmental duplication, creating false positive SV calls.

Somatic SV callers that model copy number and allele frequency can account for aneuploidy to some extent, but the analyst should be aware of the limitations. When aneuploidy is extensive, manual review of candidate variants may be necessary to distinguish true somatic SVs from copy-number artifacts.

Practical Workflow for Somatic SV Detection

Step 1: Data Preparation and Quality Control

The first step in any somatic SV analysis is to verify the quality of the sequencing data. For each sample, the analyst should check sequencing depth, read length, base quality scores, and alignment statistics. Low-quality samples should be flagged before proceeding to variant calling, because poor data quality will produce unreliable results regardless of the caller used.

For tumor-normal analysis, the normal sample should be from the same individual as the tumor sample. A matched normal sample is essential for distinguishing somatic from germline variants. Unmatched normal samples, or normal samples from different individuals, will produce incorrect somatic calls because germline variation differs between individuals.

Step 2: Alignment

The choice of aligner depends on the sequencing platform. For short-read data, standard aligners such as BWA-MEM are widely used. For long-read data, the aligner choice has a significant influence on somatic SV detection. The benchmarking study of long-read SV calling used minimap2 and found that it performed well with several somatic callers. Researchers should test aligner-caller combinations on their own data or use published benchmarks to guide their choice.

Step 3: Variant Calling

Run the selected SV caller or callers on the tumor-normal pair. For somatic callers, provide both samples as input. For generic callers, run each sample separately and prepare to subtract the results. Record the parameters used, including minimum read support, minimum variant size, and any filters applied.

Step 4: Tumor-Normal Subtraction

If using separate calling, subtract variants found in the normal sample from those found in the tumor sample. This subtraction can be performed by genomic coordinates and variant type. Variants that overlap between tumor and normal are considered germline and removed. Variants present only in the tumor are candidate somatic SVs.

If using a joint somatic caller, the caller will assign a somatic status to each variant. Review the somatic calls and verify that the normal sample does not support the variant.

Step 5: Filtering and Annotation

Apply additional filters to remove artifacts. Common filters include removing variants in repetitive regions, variants with low read support, and variants that are not supported by both paired-end and split-read evidence. Annotate the remaining variants with gene information, regulatory elements, and known cancer genes to prioritize candidates for further analysis.

Step 6: Validation

For high-confidence reporting, validate candidate somatic SVs using an orthogonal method. This validation can include PCR amplification across the breakpoint, targeted sequencing, or comparison with an independent SV caller. The multi-caller consensus approach described earlier provides a computational validation that is faster and less expensive than wet-lab validation.

Records and Measurements for Reproducible Analysis

Reproducibility is a core requirement for somatic SV analysis, particularly in clinical and translational research settings. The following records should be maintained for every analysis:

Record TypeContentPurpose
Sample metadataPatient identifier, tissue type, tumor purity, sequencing platform, coverageEnables interpretation of results and comparison across samples
Pipeline configurationAligner version and parameters, caller version and parameters, reference genome versionEnsures that analyses can be reproduced exactly
Quality metricsAlignment rate, duplication rate, insert size distribution, coverage uniformityIdentifies samples that may produce unreliable SV calls
Caller outputVCF files from each caller, including filter status and quality scoresProvides the raw evidence for somatic SV calls
Post-processing stepsSubtraction method, filtering thresholds, annotation versionDocuments the decisions that produced the final variant list

The use of workflow management systems can improve reproducibility. Community standards for pipeline development, such as those promoted by the nf-core project, provide guidelines for creating portable, version-controlled analysis pipelines. These standards help ensure that the same analysis performed on different systems produces the same results.

Training in reproducible analysis practices is available through multiple sources. The Galaxy Training Network offers accessible workflow training and analysis tutorials that cover variant calling and related topics. The Carpentries provides foundational lessons in computing, data analysis, shell, Git, and programming that are useful for researchers who need to build their own pipelines. The European Bioinformatics Institute offers training pathways for bioinformatics data resources and practical analysis education.

Common Failure Patterns in Somatic SV Detection

Failure Pattern 1: Germline Contamination

The most common failure in somatic SV analysis is the misclassification of germline variants as somatic. This failure occurs when the normal sample has low coverage, when the normal sample is contaminated with tumor DNA, or when the subtraction step is performed incorrectly. The result is an inflated somatic variant count that includes inherited polymorphisms.

Prevention: Ensure adequate coverage in the normal sample, verify that the normal sample is not contaminated, and use a joint calling approach when possible. Review the allele frequency of somatic calls in the normal sample to confirm that the variant is truly absent.

Failure Pattern 2: Low Sensitivity in Impure Tumors

Low tumor purity reduces the allele frequency of somatic variants and can cause true somatic SVs to be missed. This failure is particularly problematic for deletions and other copy-number-loss events, where the evidence for the variant is reduced by the presence of normal cells.

Prevention: Estimate tumor purity before analysis and interpret negative results with caution when purity is low. Consider using callers that model tumor purity explicitly or that can detect subclonal variants.

Failure Pattern 3: Artifacts from Repetitive Regions

Repetitive regions of the genome produce ambiguous alignments that generate false positive SV calls. These artifacts are more common in short-read data but can also appear in long-read data when the repeat is longer than the read.

Prevention: Filter variants in known repetitive regions, require support from multiple read pairs or split reads, and validate candidate variants with an orthogonal method.

Failure Pattern 4: Aligner-Caller Mismatch

The performance of an SV caller depends on the aligner used to map the reads. A caller that performs well with one aligner may perform poorly with another. This mismatch can produce both false positives and false negatives.

Prevention: Use published benchmarks to select aligner-caller combinations, or test multiple combinations on a validation sample before running the full analysis.

Failure Pattern 5: Overly Aggressive Filtering

Filters that are too stringent can remove true somatic SVs along with artifacts. This failure reduces sensitivity and may cause clinically relevant variants to be missed.

Prevention: Apply filters incrementally and track the number of variants removed at each step. Compare the final variant list with known somatic SVs in the sample or in similar cancer types to assess sensitivity.

Limitations of Current Somatic SV Calling Methods

Despite advances in sequencing technology and analytical methods, somatic SV calling remains a challenging problem. The benchmarking study of long-read SV callers found that the choice of caller had a significant influence on somatic SV detection in terms of variant type, size, sensitivity, and accuracy. No single caller performed best across all variant types and sizes, and the optimal caller depended on the specific characteristics of the data.

The co-assembly approach implemented in colorSV represents a novel strategy for somatic SV detection, but it is currently limited to certain variant types. The method demonstrated near-perfect precision and sensitivity for translocations on the COLO829 cell line, but its performance on other variant types and in other cancer types requires further evaluation.

Single-cell DNA sequencing data presents additional challenges for somatic variant calling. Single-cell data is highly error-prone due to technical biases arising from uneven sequencing coverage, allelic dropout, and amplification error. These artifacts make the identification of somatic genomic variants a challenging task, and single-cell variant callers implement distinct strategies that typically result in many discordant calls when applied to real data. Researchers working with single-cell data should be aware that somatic SV calling from this data type is even less mature than from bulk sequencing.

Safety and Regulatory Context

Somatic SV analysis in clinical settings is subject to regulatory oversight that varies by jurisdiction. Laboratories performing clinical sequencing must validate their analytical workflows and demonstrate that they meet established performance standards. The specific requirements depend on whether the laboratory is operating under CLIA, ISO, or other regulatory frameworks.

For research applications, the primary concern is data privacy and patient consent. Tumor and normal samples are derived from human subjects, and the genomic data generated from these samples must be handled in accordance with applicable privacy regulations. Researchers should ensure that their data management practices comply with institutional review board requirements and applicable laws.

The NCBI provides data resources and search systems that support the deposition and retrieval of genomic data, including cancer genomics data. Researchers should follow the data submission guidelines for the relevant databases and ensure that data is deposited in a manner that protects patient privacy while enabling scientific reproducibility.

Professional Escalation Criteria

Researchers and laboratory professionals should escalate somatic SV analysis issues to a supervisor, bioinformatics specialist, or clinical consultant under the following circumstances:

  1. The somatic SV caller produces an unexpectedly high number of calls that cannot be validated by orthogonal methods.
  2. The tumor purity is below the threshold required for reliable somatic SV detection, and the analysis is intended to inform clinical decisions.
  3. A candidate somatic SV involves a known cancer gene and may have clinical significance, but the evidence is insufficient for confident reporting.
  4. The normal sample shows evidence of tumor contamination, which compromises the validity of the somatic subtraction.
  5. The analysis pipeline produces inconsistent results across replicate runs, indicating a reproducibility problem.
  6. The laboratory lacks the expertise to interpret the results of a novel or unfamiliar SV caller.

Building a Somatic SV Decision Framework for Tumor-Normal Analysis

A recurring problem in somatic structural variant detection is that researchers often select callers, filters, and validation strategies without a structured method for matching those choices to their specific sample characteristics and study goals. The result is either an excess of false positives that consume validation resources or a loss of true somatic events that undermines the biological conclusions. A practical decision framework helps the analyst move from raw sequencing data to a defensible somatic SV list by making explicit the tradeoffs at each stage and providing criteria for when to escalate or revise the approach.

Defining the Analysis Objective Before Calling Variants

The first decision in the framework is to define what the somatic SV calls will be used for. This objective determines the balance between sensitivity and precision throughout the pipeline. A discovery study aimed at finding novel fusion genes in a rare tumor type may accept lower precision in exchange for higher recall, because the goal is to generate hypotheses that will be validated later. A clinical reporting workflow, by contrast, requires high precision because false positive calls can lead to incorrect treatment decisions.

The analysis objective also determines which variant types matter most. Some studies focus on translocations that create oncogenic fusions, while others need to characterize deletions that remove tumor suppressor genes. The benchmarking study of long-read SV callers found that the choice of caller had a significant influence on somatic SV detection in terms of variant type, size, sensitivity, and accuracy. A caller that performs well for deletions may perform poorly for inversions or insertions. The analyst should therefore document the variant types of interest before selecting tools and should evaluate caller performance separately for each type instead of relying on aggregate metrics.

The objective also influences the required confidence level for each call. A candidate SV that will be followed up with PCR validation requires less computational confidence than one that will be reported directly. The decision framework should specify the minimum evidence required for each reporting category, such as research-grade candidates, validated candidates, and clinically reportable variants.

Assessing Sample Characteristics That Constrain the Analysis

Before running any SV caller, the analyst should assess three sample characteristics that fundamentally constrain what can be detected: tumor purity, ploidy status, and sequencing depth. These characteristics should be estimated from the data or from pathology records and recorded in the analysis metadata.

Tumor purity directly determines the maximum observable allele frequency of a somatic variant. A homozygous deletion in a pure tumor sample will be supported by nearly all reads at the locus, but the same deletion in a sample with 30 percent purity will be supported by only 30 percent of reads. The benchmarking study of long-read somatic SV calling in lung cancer used matched tumor and non-tumour samples and found that somatic callers achieved higher recall than generic callers, but even the best somatic callers did not detect all events. Low purity is a primary cause of missed somatic SVs, and the analyst should document the purity estimate and its method of derivation before interpreting negative results.

Ploidy status affects the expected copy number at each locus and therefore the interpretation of read-depth evidence. A diploid genome has two copies of each autosome, and a heterozygous deletion removes one copy. An aneuploid tumor may have three, four, or more copies of a chromosome, and a deletion that removes one copy from a tetraploid region may not produce a detectable change in read depth. The analyst should estimate ploidy from copy-number analysis or from the allele frequencies of germline heterozygous variants and should record whether the tumor is near-diploid, polyploid, or highly aneuploid.

Sequencing depth determines the statistical power to detect variants at a given allele frequency. Higher depth provides more reads supporting each breakpoint junction and allows detection of subclonal events. The required depth depends on the expected allele frequency, which is a function of purity and clonality. A somatic SV present in 80 percent of tumor cells in a sample with 50 percent purity will have an allele frequency of approximately 40 percent, requiring moderate depth for confident detection. A subclonal SV present in 10 percent of tumor cells in the same sample will have an allele frequency of approximately 5 percent, requiring much higher depth or long reads that span the junction.

Selecting Callers Based on Data Type and Sample Properties

The caller selection step should be guided by the sequencing platform, the sample characteristics, and the analysis objective. For short-read data, the analyst should choose callers that are optimized for the specific read length and insert size of the library. For long-read data, the choice of aligner and caller has a substantial impact on results. The benchmarking study of long-read SV callers in cancer genomes found that different combinations of aligners and variant callers influenced somatic SV detection and that the choice of caller had a significant influence in terms of variant type, size, sensitivity, and accuracy.

For long-read data, the analyst should prioritize somatic callers over generic callers when the goal is to maximize recall of true somatic events. The lung cancer benchmarking study found that somatic callers achieved a mean recall of 72 percent for somatic events, while generic callers achieved only 53 percent. Among the somatic callers tested with the minimap2 aligner, SAVANA and Severus achieved the highest recall at 79.5 percent and 79.25 percent respectively, followed by nanomonsv at 72.5 percent. These results provide a starting point for caller selection, but the analyst should verify performance on their own data because the optimal caller may differ depending on tumor type, coverage, and variant spectrum.

The decision framework should also consider whether to use a single caller or multiple callers. A study that evaluated eight widely used SV callers on paired tumor and normal samples found that combining multiple tools and testing different combinations significantly enhanced the validation of somatic alterations. The study used a VCF merging procedure followed by a subtraction method to identify candidate somatic SVs. The framework should specify the minimum number of callers required for each reporting category. A single caller may be sufficient for research-grade candidates, while two or more independent callers should be required for validated candidates.

Implementing a Tiered Evidence Classification System

A practical decision framework assigns each candidate somatic SV to a confidence tier based on the strength and diversity of evidence. This tiered system allows the analyst to prioritize validation efforts and to report results with appropriate caveats.

Tier 1 candidates are supported by multiple independent callers and by multiple lines of evidence within each caller. These candidates have high confidence and are suitable for downstream functional validation or clinical consideration. The evidence should include both paired-end and split-read support where applicable, and the breakpoint should be assembled to base-pair resolution where possible. Manta, for example, consistently assembles a higher fraction of its calls to base-pair resolution, which improves downstream annotation and analysis of clinical significance.

Tier 2 candidates are supported by a single caller or by multiple callers with weaker evidence. These candidates are suitable for research follow-up but should not be reported as definitive somatic events without additional validation. The analyst should document the specific evidence supporting each Tier 2 call and should note any conflicting evidence from other callers.

Tier 3 candidates are those that pass initial filters but have limited support or are in difficult genomic regions. These candidates should be flagged for manual review or excluded from downstream analysis unless they involve a known cancer gene. The decision framework should specify the criteria for promoting a Tier 3 candidate to a higher tier, such as the identification of a disrupted gene with known oncogenic function.

The tier assignment should be recorded in the output VCF or in a separate tracking table. This record allows the analyst to audit the decision process and to revise tier assignments if new evidence becomes available.

Establishing Filtering Thresholds Based on Data Quality

Filtering thresholds should be established before running the analysis and should be based on the observed data quality instead of on arbitrary values. The analyst should examine the distribution of variant quality scores, read support, and allele frequencies in the initial caller output and should set thresholds that separate the main distribution of true variants from the tail of likely artifacts.

For read support, the threshold should account for sequencing depth and tumor purity. A minimum of two supporting read pairs or split reads is often used, but this threshold may need to be higher in high-depth data where artifacts are more common. The analyst should also consider the insert size distribution of the library, because aberrant insert sizes can create false positive calls.

For allele frequency, the threshold should be set relative to the expected allele frequency of true somatic variants given the tumor purity. A somatic variant in a sample with 50 percent purity and 100 percent clonality will have an allele frequency near 50 percent for a heterozygous event or near 100 percent for a homozygous event. Subclonal variants will have lower allele frequencies. The analyst should set a minimum allele frequency that balances sensitivity for subclonal events against the risk of including artifacts.

The filtering thresholds should be documented in the pipeline configuration and should be reviewed when the analysis is applied to a new sample or a new data type. Thresholds that work well for high-purity, high-depth samples may be inappropriate for low-purity or low-depth samples.

Recording Decisions and Outcomes in a Structured Format

The decision framework requires a structured record of every analysis decision and its outcome. This record serves multiple purposes: it allows the analysis to be reproduced, it provides an audit trail for regulatory or publication review, and it enables the analyst to identify systematic problems in the pipeline.

The record should include the sample metadata, the pipeline configuration, the quality metrics, the caller output, and the post-processing steps. The sample metadata should include the patient identifier, tissue type, tumor purity estimate, sequencing platform, and coverage. The pipeline configuration should include the aligner version and parameters, the caller version and parameters, and the reference genome version. The quality metrics should include alignment rate, duplication rate, insert size distribution, and coverage uniformity.

The post-processing record should document the subtraction method, the filtering thresholds, the tier assignments, and the validation results. This record should be maintained in a version-controlled format so that changes to the pipeline can be tracked over time. The nf-core documentation provides community standards for pipeline development that emphasize portability and version control, and these standards can be adapted to the somatic SV analysis workflow.

The record should also include the outcome of any validation experiments. For each candidate that was validated by PCR, targeted sequencing, or an orthogonal caller, the result should be recorded as confirmed, refuted, or inconclusive. This validation record provides the ground truth needed to assess the performance of the pipeline and to refine the decision framework over time.

Troubleshooting When Results Deviate from Expectations

The decision framework should include a troubleshooting procedure for when the analysis produces unexpected results. The most common deviations are an unexpectedly high number of somatic calls, an unexpectedly low number of somatic calls, and a high proportion of calls that fail validation.

An unexpectedly high number of somatic calls often indicates germline contamination or a subtraction failure. The analyst should first check the normal sample for evidence of tumor contamination by examining the allele frequencies of known germline variants. If the normal sample shows tumor contamination, the subtraction step will fail to remove germline variants, and the somatic call count will be inflated. The analyst should also verify that the subtraction step used the correct genomic coordinates and variant types and that the normal sample had adequate coverage.

An unexpectedly low number of somatic calls often indicates low tumor purity, overly aggressive filtering, or a caller that is poorly suited to the data type. The analyst should review the purity estimate and consider whether the expected allele frequencies are below the detection threshold of the caller. The analyst should also review the filtering thresholds to determine whether true variants were removed by the filters.

A high proportion of calls that fail validation indicates a systematic problem with the calling strategy. The analyst should review the evidence supporting the failed calls to identify the common cause. If the failed calls are concentrated in repetitive regions, the analyst should add a repeat-masking filter. If the failed calls are concentrated in regions with high copy number, the analyst should review the ploidy estimate and consider whether aneuploidy is creating artifacts.

The troubleshooting procedure should specify when to escalate to a supervisor or bioinformatics specialist. Escalation is appropriate when the analyst cannot identify the cause of the deviation, when the deviation affects a clinically significant result, or when the pipeline produces inconsistent results across replicate runs.

Applying the Framework to Long-Read and Single-Cell Data

The decision framework applies to both short-read and long-read data, but the specific parameters and caller choices differ. For long-read data, the analyst should use long-read-specific aligners and callers and should evaluate caller performance using published benchmarks. The benchmarking study of long-read SV callers found that the choice of caller had a significant influence on somatic SV detection and that combining multiple tools enhanced validation. The framework should specify the minimum number of long-read callers required for each confidence tier.

For single-cell DNA sequencing data, the framework requires additional caution. Single-cell data is highly error-prone due to technical biases arising from uneven sequencing coverage, allelic dropout, and amplification error. These artifacts make the identification of somatic genomic variants a challenging task, and single-cell variant callers implement distinct strategies that typically result in many discordant calls when applied to real data. The analyst should apply more stringent filtering thresholds for single-cell data and should require validation of any candidate that will be used for downstream analysis.

The framework should also specify the training and skill requirements for the analyst. The Galaxy Training Network offers accessible workflow training and analysis tutorials that cover variant calling and related topics. The Carpentries provides foundational lessons in computing, data analysis, shell, Git, and programming that are useful for researchers who need to build their own pipelines. The European Bioinformatics Institute offers training pathways for bioinformatics data resources and practical analysis education. The analyst should have completed relevant training before applying the framework to clinical or translational samples.

Reviewing and Revising the Framework Periodically

The decision framework is not a static document. It should be reviewed and revised periodically as new callers become available, as the laboratory acquires experience with new data types, and as validation results accumulate. The review should examine the performance of the pipeline against the validation record and should identify opportunities to improve sensitivity or precision.

The review should also consider new methods that may change the optimal analysis strategy. The co-assembly approach implemented in colorSV represents a novel strategy for somatic SV detection that examines the local topology of joint assembly graphs from matched tumor-normal samples. This method demonstrated near-perfect precision and sensitivity for calling translocations on the COLO829 cell line, outperforming four existing somatic SV callers in both metrics. As new methods like colorSV become available, the framework should be updated to include them and to specify the conditions under which they should be used.

The review process should be documented and should include input from all stakeholders, including the analysts who run the pipeline, the researchers who use the results, and the laboratory leadership who are responsible for quality. The review should result in a revised version of the framework that is version-controlled and that includes a summary of the changes and the rationale for each change.

Frequently Asked Questions

What is the difference between germline and somatic structural variant calling?

Germline variant calling identifies variants that are inherited and present in all cells of an individual. Somatic variant calling identifies variants that are acquired during cancer development and present only in tumor cells. Somatic calling requires a matched normal sample to subtract germline variation and to confirm that the variant is absent from the normal tissue.

Why is a matched normal sample required for somatic SV detection?

A matched normal sample provides the germline baseline for the individual patient. Without a matched normal, the analyst cannot distinguish somatic variants from inherited polymorphisms. Using an unmatched normal or a normal from a different individual will produce incorrect somatic calls because germline variation differs between individuals.

How does tumor purity affect somatic SV detection?

Tumor purity is the proportion of cancer cells in a tumor sample. Low purity dilutes the allele frequency of somatic variants, making them harder to detect. A somatic variant in a tumor with 50 percent purity will be present in only half of the sequencing reads, which may fall below the detection threshold of the caller. Researchers should estimate purity before analysis and interpret negative results with caution when purity is low.

What is the advantage of long-read sequencing for somatic SV detection?

Long-read sequencing produces reads of 10 kilobases or more, which can span repetitive elements and complex rearrangement junctions. This capability allows long-read sequencing to identify more somatic SVs than short-read sequencing and to resolve rearrangements that short reads cannot fully assemble. However, long-read data requires different aligners and callers, and the choice of these tools significantly affects results.

Should I use a generic or somatic SV caller for tumor-normal analysis?

Somatic SV callers are specifically designed to analyze tumor-normal pairs and achieve higher recall for somatic events than generic callers. A benchmarking study found that somatic callers achieved a mean recall of 72 percent for somatic events, while generic callers achieved only 53 percent. For long-read data, somatic callers such as SAVANA and Severus performed best. Generic callers can be used with subtraction, but this approach is less sensitive.

How can I validate somatic SV calls without wet-lab experiments?

Multi-caller consensus is a computational validation approach that combines the results of multiple SV callers. A variant detected by multiple independent callers is more likely to be genuine than a variant detected by a single caller. This approach is faster and less expensive than wet-lab validation, but it cannot replace orthogonal validation for clinical reporting.

What causes false positive somatic SV calls?

False positive calls can result from germline contamination, artifacts in repetitive regions, aligner-caller mismatches, and overly permissive filters. Germline contamination occurs when the normal sample contains tumor DNA or when the subtraction step is performed incorrectly. Repetitive regions produce ambiguous alignments that generate spurious calls. The choice of aligner can also affect caller performance, and mismatched combinations produce unreliable results.

How do I handle aneuploidy in somatic SV analysis?

Aneuploidy complicates SV calling because copy-number changes can create apparent breakpoints and alter expected allele frequencies. Somatic SV callers that model copy number and allele frequency can account for aneuploidy to some extent. When aneuploidy is extensive, manual review of candidate variants may be necessary to distinguish true somatic SVs from copy-number artifacts.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.