How to Choose a Somatic Variant Caller: MuTect2, VarScan2, Strelka2, or Lancet?

By Dr. Zubair Khalid, DVM, MS, PhD ·

How to Choose a Somatic Variant Caller: MuTect2, VarScan2, Strelka2, or Lancet?

Key Takeaways

  • MuTect2 excels at SNV detection in moderate to high purity tumor-normal pairs, leveraging the GATK ecosystem and a panel of normals for artifact filtering, but requires careful post-calling filtration for germline contamination and sequencing artifacts. Its computational intensity necessitates parallelization for large datasets.
  • VarScan2 offers a low computational footprint and handles variable tumor purity and ploidy, making it suitable for low-purity samples or targeted panels, though its indel accuracy is lower and it requires a separate mpileup preprocessing step. It is sensitive to sequencing errors at low allele frequencies.
  • Strelka2 provides efficient runtime and good sensitivity for both SNVs and indels in whole-genome and whole-exome data with standard purity, but requires a candidate indel file generated by tools like Manta for optimal indel detection. It uses a mixture model accounting for tumor purity and subclonal heterogeneity.
  • Lancet is highly effective for indel detection, particularly in repetitive genomic regions, due to its local assembly approach, but exhibits slower runtime and higher memory usage. It is well-suited for indel-focused studies and validation of candidate variants.
  • Caller selection is critically dependent on data characteristics, including sequencing depth, tumor purity, and variant types of interest (SNVs vs. indels), with performance varying significantly across these parameters. Benchmarking studies highlight that regions with high sequencing error rates demand specific caller considerations.
  • A multi-caller ensemble strategy, comparing outputs and prioritizing variants called by multiple tools, can enhance confidence and improve overall sensitivity and specificity. Variants called by only one tool necessitate orthogonal validation.

Somatic variant calling identifies mutations present in tumor tissue but absent from matched normal tissue, and the choice of caller directly affects which variants you detect and report. MuTect2, VarScan2, Strelka2, and Lancet each use different statistical models, require different inputs, and perform differently across sequencing depths, tumor purity, and variant types. This article provides a systematic comparison of these four callers, including sensitivity, specificity, runtime, and input requirements, with recommendations for different research scenarios.

The practical problem is straightforward: researchers need to select a caller that matches their data type, research question, and available computational resources. A caller optimized for high-depth whole-exome data may perform poorly on low-purity whole-genome data. A caller that excels at single nucleotide variants may miss indels. The decision affects downstream filtering, annotation, and biological interpretation.

This comparison draws on published benchmarking studies and official documentation from the National Center for Biotechnology Information, the European Bioinformatics Institute, Bioconductor, the Galaxy Training Network, nf-core, and The Carpentries. The evidence base includes recent benchmarking of somatic callers on experimentally confirmed variant data, circulating tumor DNA dilution series, and structural variant detection pipelines.

At a Glance

The table below summarizes the key characteristics of the four callers covered in this article. Use it as a starting point for selection, then consult the detailed sections that follow for workflow-specific guidance.

CallerInput RequirementsStrengthsLimitationsBest Use Case
MuTect2Tumor-normal BAM files, reference genome, known sites resourceHigh sensitivity for SNVs, integrates with GATK ecosystem, widely used in cancer genomicsRequires careful filtering of germline contamination and sequencing artifacts, computationally intensiveWhole-exome and whole-genome tumor-normal pairs with moderate to high purity
VarScan2Tumor-normal BAM files, reference genome, mpileup outputHandles variable purity and ploidy, detects SNVs and indels, low computational footprintRequires separate mpileup step, sensitive to sequencing errors at low allele frequencies, less accurate for indelsLow-purity samples, targeted panels, limited computational resources
Strelka2Tumor-normal BAM files, reference genome, candidate indels fileHigh indel accuracy, efficient runtime, good sensitivity for SNVs and indelsRequires candidate indel file for optimal performance, less flexible parameter tuningWhole-genome and whole-exome data with standard purity, balanced SNV and indel detection
LancetTumor-normal BAM files, reference genome, window size parameterExcellent for indels, uses local assembly, handles repetitive regions wellSlower runtime, higher memory usage, requires parameter tuning for different data typesIndel-focused studies, repetitive genomic regions, validation of candidate variants

Understanding Somatic Variant Calling

Somatic variant calling differs fundamentally from germline variant calling. Germline variants are present in every cell of an individual and are inherited. Somatic variants arise during an individual's lifetime, typically in the context of cancer, and are present only in the tumor tissue. The matched normal sample provides the baseline for distinguishing somatic mutations from inherited polymorphisms.

The core challenge is that tumor samples are heterogeneous. A tumor biopsy contains a mixture of cancer cells and normal cells, and the proportion of cancer cells, called tumor purity, varies between samples and even within a single tumor. Variant allele frequency, the proportion of sequencing reads supporting the variant, depends on both tumor purity and the number of copies of the mutated allele. A variant present in all cancer cells of a pure tumor sample may have an allele frequency near 50 percent for heterozygous mutations, but the same variant in a low-purity sample may have an allele frequency below 5 percent.

Sequencing errors compound this problem. Base-calling errors, mapping artifacts, and PCR duplicates can produce false positive calls that mimic true somatic mutations. The four callers described here use different statistical approaches to distinguish true mutations from noise, and these approaches have different strengths and weaknesses.

The National Center for Biotechnology Information provides access to the reference genomes, sequence databases, and variation resources needed for somatic variant calling workflows. The European Bioinformatics Institute offers training materials on data resources and practical analysis education. The Galaxy Training Network provides accessible workflow training and analysis tutorials that can help researchers build and validate somatic variant calling pipelines. The Carpentries lessons cover foundational computing, data, shell, Git, and programming training that supports reproducible analysis.

Core Principles of Caller Selection

Data Type Determines Caller Suitability

Whole-genome sequencing, whole-exome sequencing, and targeted panels produce different data characteristics that affect caller performance. Whole-genome data has uniform coverage across the genome but typically lower depth, often 30x to 60x for tumor samples. Whole-exome data has higher depth, often 100x to 200x, but coverage is restricted to coding regions and capture efficiency varies across the genome. Targeted panels can achieve very high depth, often 500x to 2000x, but cover only a limited set of genes.

The benchmarking study of circulating tumor DNA detection used deep whole-genome sequencing at 150x and exome sequencing at 2000x to define a reference set of approximately 37,000 single nucleotide variants and 58,000 indels. This study benchmarked nine somatic variant callers across varying circulating tumor DNA levels and sequencing depths, providing practical guidance for selecting calling methods in liquid biopsy applications. The results showed that caller performance varies substantially with sequencing depth and variant allele frequency, reinforcing the need to match caller choice to data characteristics.

Tumor Purity and Heterogeneity

Tumor purity directly affects variant allele frequency and therefore caller sensitivity. A caller that performs well on high-purity samples may miss variants in low-purity samples because the variant allele frequency falls below the caller's detection threshold. Conversely, a caller that is too sensitive may produce excessive false positives in high-purity samples because it cannot distinguish true variants from sequencing artifacts.

The VariantMedium study trained and evaluated a somatic variant caller on experimentally confirmed variant data, using 336,839 variants from 2,956 samples with whole-exome or whole-genome sequencing for training and validation, and 118,887 variants from two independent studies with deep sequencing data for evaluation and benchmarking. The study found that VariantMedium showed the highest sensitivity among benchmarked callers and achieved similar or better F1 scores in SNV calling. Its performance was particularly pronounced in genomic regions characterized by high sequencing error rates, where it achieved higher F1 scores than MuTect2 and Strelka2. This finding highlights that caller performance is not uniform across the genome and that regions with high sequencing error rates require special consideration.

Variant Types and Detection Limits

Single nucleotide variants, small insertions and deletions, and structural variants require different detection strategies. The four callers covered in this article focus on SNVs and small indels. Structural variants, which include large deletions, duplications, inversions, and translocations, require specialized callers and are beyond the scope of this comparison.

The benchmarking of somatic structural variant callers on the HG008 genome evaluated four leading somatic SV detection tools: Sniffles2, Nanomonsv, Savana, and Severus. The study noted that most existing SV detection algorithms were originally developed for germline variants and are not well-suited to addressing the high heterogeneity of somatic mutations. The study established a multi-tool ensemble strategy for SV detection to achieve more accurate and comprehensive identification of somatic SVs. This finding underscores that different variant types require different tools and that no single caller handles all variant classes optimally.

Practical Workflow for Caller Selection

Step 1: Define Your Research Question and Data Characteristics

Before selecting a caller, document the following parameters:

  • Sequencing platform and library preparation method
  • Sequencing depth and coverage uniformity
  • Tumor purity and sample heterogeneity
  • Variant types of interest (SNVs, indels, or both)
  • Available computational resources
  • Downstream analysis requirements

Record these parameters in your laboratory notebook or electronic lab notebook. They will guide caller selection and interpretation of results.

Step 2: Assess Input Data Quality

All four callers require aligned sequencing reads in BAM format, a reference genome, and a matched normal sample. Before running any caller, verify the following:

  • Alignment was performed with a splice-aware aligner appropriate for the data type
  • Duplicate reads were marked or removed
  • Base quality scores were recalibrated
  • The reference genome version is consistent across all samples and resources

The Galaxy Training Network provides accessible workflow training and analysis tutorials that cover quality assessment and preprocessing steps. The nf-core documentation describes community pipeline standards, usage, and configuration for reproducible workflow context. These resources can help you build a robust preprocessing pipeline before variant calling.

Step 3: Select Candidate Callers Based on Data Type

For whole-exome data with moderate to high purity, MuTect2 and Strelka2 are reasonable starting points. Both callers have been widely used in cancer genomics and have documented performance characteristics.

For whole-genome data, Strelka2 offers efficient runtime and good sensitivity for both SNVs and indels. MuTect2 can also be used but may require more computational resources.

For targeted panels with high depth, VarScan2 can handle the high allele frequencies and low error rates typical of amplicon-based approaches. Lancet may be useful for indel detection in repetitive regions.

For low-purity samples or circulating tumor DNA, the benchmarking study of mutation calling in circulating tumor DNA provides practical guidance. The study used longitudinal patient-matched samples from individuals with colorectal and breast cancer, combining samples with high and ultra-low levels of tumor-derived DNA into controlled dilution series. The study also explored machine learning-based tuning of individual callers and identified features that improve accuracy in cell-free DNA. This resource clarifies the detection limits of current approaches and provides practical guidance for selecting somatic variant calling methods in liquid biopsy applications.

Step 4: Run Multiple Callers and Compare Results

Running multiple callers and comparing their outputs can improve sensitivity and specificity. The structural variant benchmarking study established a multi-tool ensemble strategy for SV detection to achieve more accurate and comprehensive identification of somatic SVs. A similar approach can be applied to SNV and indel calling.

When comparing caller outputs, focus on variants called by multiple callers as high-confidence candidates. Variants called by only one caller require additional scrutiny and may need orthogonal validation.

Step 5: Validate and Filter Candidate Variants

The tutorial for variant interrogation in tumor samples presents a practical framework that guides researchers through the critical steps of variant interrogation. The framework is broken into four phases: planning, gathering resources, filtering and validation, and dissemination and storage. The filtering and validation phase emphasizes executing a systematic approach to prioritize meaningful variants.

Filtering criteria may include:

  • Variant allele frequency thresholds
  • Read depth at the variant position
  • Mapping quality and base quality scores
  • Presence in population databases
  • Predicted functional impact

The National Center for Biotechnology Information provides access to databases that can be used for filtering, including dbSNP and gnomAD. The European Bioinformatics Institute offers training on data resources that can support variant annotation and interpretation.

Detailed Comparison of the Four Callers

MuTect2

MuTect2 is part of the Genome Analysis Toolkit and is designed for somatic SNV and indel detection in tumor-normal pairs. It uses a Bayesian approach that incorporates panel of normals information to filter recurrent sequencing artifacts and germline contamination.

MuTect2 requires a tumor BAM file, a matched normal BAM file, a reference genome, and a known sites resource such as dbSNP. The caller can also use a panel of normals to improve specificity by modeling site-specific error rates.

The VariantMedium benchmarking study found that MuTect2 achieved lower F1 scores than VariantMedium in genomic regions characterized by high sequencing error rates. This finding suggests that MuTect2 may require additional filtering or complementary approaches in difficult genomic regions.

MuTect2 is computationally intensive, particularly for whole-genome data. Runtime scales with the number of reads and the number of candidate variant sites. For large cohorts, consider using the GATK best practices workflow and parallelizing across genomic intervals.

Practical considerations for MuTect2:

  • Use the panel of normals feature when possible to improve specificity
  • Apply the recommended filtering steps after calling, including removal of variants with low tumor read depth and variants present in the normal sample
  • Consider using the tumor-only mode only when a matched normal is unavailable, and interpret results with caution

VarScan2

VarScan2 uses a heuristic approach based on read counts and allele frequencies to identify somatic variants. It requires an mpileup file generated from the tumor and normal BAM files, which adds a preprocessing step but reduces the computational burden during variant calling.

VarScan2 can handle variable tumor purity and ploidy, making it useful for samples with unusual copy number profiles. It detects both SNVs and indels, but its indel detection is generally less accurate than its SNV detection.

The computational footprint of VarScan2 is lower than MuTect2 and Strelka2, making it suitable for environments with limited resources. However, the mpileup step can be memory-intensive for whole-genome data.

Practical considerations for VarScan2:

  • Generate mpileup files with appropriate base and mapping quality filters
  • Set the minimum variant allele frequency threshold based on expected tumor purity
  • Use the Fisher exact test option to filter variants with strand bias
  • Validate indels with an independent method, as VarScan2 indel calls may have higher false positive rates

Strelka2

Strelka2 uses a mixture model approach that accounts for tumor purity and subclonal heterogeneity. It is designed for somatic SNV and indel detection in tumor-normal pairs and has been optimized for efficiency.

Strelka2 requires a candidate indel file for optimal indel detection. This file can be generated using the Manta structural variant caller, which is part of the same software suite. Without the candidate indel file, Strelka2 will only detect SNVs.

The VariantMedium benchmarking study found that Strelka2 achieved lower F1 scores than VariantMedium in genomic regions characterized by high sequencing error rates. However, Strelka2 generally performs well across the genome and has a favorable runtime compared to MuTect2.

Practical considerations for Strelka2:

  • Run Manta to generate the candidate indel file before running Strelka2
  • Use the default parameters for standard whole-exome and whole-genome data
  • Adjust the somatic prior probability for samples with unusual purity or clonality
  • Apply the recommended filtering steps, including removal of variants with low quality scores and variants present in the normal sample

Lancet

Lancet uses a local assembly approach to detect somatic variants, which makes it particularly effective for indels and variants in repetitive regions. It requires a tumor BAM file, a matched normal BAM file, a reference genome, and a window size parameter that controls the genomic region analyzed.

Lancet is computationally intensive and has higher memory usage than the other three callers. Runtime increases with window size and read depth. For whole-genome data, Lancet may require substantial computational resources and time.

The local assembly approach used by Lancet can detect variants that read-alignment-based methods miss, particularly in regions with low mappability or high sequence complexity. However, the increased sensitivity may come at the cost of additional false positives that require careful filtering.

Practical considerations for Lancet:

  • Optimize the window size parameter for your data type and computational resources
  • Use Lancet as a complementary caller to validate variants detected by other methods
  • Expect longer runtime and higher memory usage compared to MuTect2, VarScan2, and Strelka2
  • Apply stringent filtering to remove assembly artifacts

Observations and Measurements

Sensitivity and Specificity Tradeoffs

The benchmarking studies cited in this article provide quantitative evidence of caller performance differences. The VariantMedium study used experimentally confirmed variant data and found that VariantMedium showed the highest sensitivity among benchmarked callers and achieved similar or better F1 scores in SNV calling. The study also found that performance differences were most pronounced in genomic regions with high sequencing error rates.

The circulating tumor DNA benchmarking study defined a reference set of approximately 37,000 single nucleotide variants and 58,000 indels to benchmark nine somatic variant callers across varying circulating tumor DNA levels and sequencing depths. This study provides practical guidance for selecting calling methods in liquid biopsy applications, where variant allele frequencies are often very low.

These findings have direct implications for caller selection. If your study focuses on regions with high sequencing error rates, such as GC-rich regions or repetitive elements, you may need a caller with demonstrated performance in these regions or a complementary approach such as local assembly.

Runtime and Resource Requirements

Runtime and resource requirements vary substantially among the four callers. VarScan2 has the lowest computational footprint but requires a separate mpileup step. Strelka2 offers efficient runtime and is suitable for large cohorts. MuTect2 is computationally intensive but integrates with the GATK ecosystem. Lancet has the highest runtime and memory requirements due to its local assembly approach.

Record runtime and resource usage for each caller on your data. These measurements will help you plan computational resource allocation for future analyses and identify bottlenecks in your workflow.

Reproducibility Considerations

Reproducibility is a critical concern in somatic variant calling. Small changes in parameters, reference genome version, or preprocessing steps can produce different variant calls. The nf-core documentation describes community pipeline standards, usage, and configuration for reproducible workflow context. The Galaxy Training Network provides accessible workflow training and analysis tutorials that emphasize reproducibility.

To ensure reproducibility:

  • Document the exact version of each tool and reference genome used
  • Record all parameters and filtering thresholds
  • Use containerized environments or workflow managers to standardize the analysis environment
  • Store raw data, intermediate files, and final variant calls in versioned storage

The Carpentries lessons cover foundational computing, data, shell, Git, and programming training that supports reproducible analysis practices.

Records and Measurements

What to Record

Maintain detailed records of the following for each somatic variant calling run:

  • Sample identifiers and tumor-normal pair information
  • Sequencing platform, library preparation method, and sequencing depth
  • Alignment and preprocessing steps, including tool versions and parameters
  • Variant caller name, version, and all parameters used
  • Reference genome version and known sites resource versions
  • Filtering criteria and the number of variants removed at each step
  • Runtime and computational resource usage
  • Output file locations and checksums

These records support reproducibility, troubleshooting, and publication requirements.

Quality Metrics to Monitor

Monitor the following quality metrics during and after variant calling:

  • Transition to transversion ratio, which should be approximately 2 for whole-genome data
  • Variant allele frequency distribution, which should show expected patterns based on tumor purity
  • Number of variants called per sample, which should be consistent across similar samples
  • Fraction of variants in coding regions for exome data
  • TiTv ratio for known versus novel variants

Deviations from expected values may indicate problems with data quality, parameter settings, or sample contamination.

Common Failure Patterns

Failure Pattern 1: Excessive False Positives

Excessive false positives can result from sequencing artifacts, mapping errors, or germline contamination. Common causes include:

  • Insufficient filtering of low-quality reads or bases
  • Failure to remove duplicate reads
  • Contamination of the tumor sample with normal cells
  • Presence of sequencing artifacts in GC-rich or repetitive regions

If you observe an unusually high number of variant calls, review your preprocessing steps and filtering criteria. Consider using a panel of normals to model site-specific error rates.

Failure Pattern 2: Low Sensitivity

Low sensitivity can result from:

  • Low tumor purity or high subclonality
  • Insufficient sequencing depth
  • Conservative filtering thresholds
  • Caller parameters that are not optimized for your data type

If you observe fewer variants than expected, review your tumor purity estimates and consider adjusting caller parameters. The circulating tumor DNA benchmarking study provides guidance on detection limits at varying sequencing depths and variant allele frequencies.

Failure Pattern 3: Inconsistent Results Across Callers

Different callers may produce different variant calls for the same sample. This inconsistency can result from:

  • Different statistical models and detection thresholds
  • Different handling of indels and complex variants
  • Different filtering criteria applied after calling

When callers disagree, treat variants called by only one caller as lower confidence and consider orthogonal validation. The structural variant benchmarking study established a multi-tool ensemble strategy to achieve more accurate and comprehensive identification of somatic variants.

Failure Pattern 4: Runtime or Resource Exhaustion

Lancet and MuTect2 can exhaust computational resources on large datasets. Common causes include:

  • Running whole-genome data without parallelization
  • Using excessive window sizes in Lancet
  • Insufficient memory allocation for the Java virtual machine in MuTect2

If you encounter resource exhaustion, consider parallelizing across genomic intervals, reducing window sizes, or using a caller with lower resource requirements.

Limitations and Interpretation Constraints

Detection Limits at Low Variant Allele Frequencies

All four callers have detection limits at low variant allele frequencies. The circulating tumor DNA benchmarking study clarified the detection limits of current approaches and provided practical guidance for selecting somatic variant calling methods in liquid biopsy applications. At very low variant allele frequencies, below approximately 1 percent, sensitivity decreases substantially and false positive rates increase.

If your study requires detection of variants at very low allele frequencies, consider using deeper sequencing, targeted validation, or machine learning-based approaches that have been trained on experimentally confirmed variant data.

Structural Variants Require Specialized Tools

The four callers covered in this article detect SNVs and small indels. Structural variants require specialized tools. The benchmarking of somatic structural variant callers on the HG008 genome evaluated Sniffles2, Nanomonsv, Savana, and Severus and noted that most existing SV detection algorithms were originally developed for germline variants and are not well-suited to addressing the high heterogeneity of somatic mutations.

If your study requires structural variant detection, select specialized tools and validate their performance on your data type. The SV-AUTOPILOT study provides an automated, standardized pipeline for SV prediction and benchmarking, which can be useful for evaluating SV callers on non-human genomes.

Machine Learning Approaches Are Emerging

Machine learning-based somatic variant callers, such as VariantMedium, have shown promising performance in benchmarking studies. VariantMedium combines a tree-based classifier with a 3D densely connected convolutional network architecture and was trained and evaluated on experimentally confirmed variant data. The study found that VariantMedium showed the highest sensitivity among benchmarked callers and achieved similar or better F1 scores in SNV calling.

These approaches may offer advantages in difficult genomic regions, but they require careful validation and may not be suitable for all data types. Consider evaluating machine learning-based callers as complementary tools to the four callers covered in this article.

Safety and Regulatory Context

Data Privacy and Security

Somatic variant calling involves sensitive patient data. Ensure compliance with applicable data protection regulations and institutional policies. Store genomic data on secure servers with access controls and encryption. The National Center for Biotechnology Information provides guidance on data submission and access policies for genomic data.

Clinical Reporting Considerations

If variant calls will be used for clinical reporting, additional validation and quality control measures are required. The tutorial for variant interrogation in tumor samples emphasizes accessibility, reproducibility, and clinical relevance. The framework guides researchers through planning, gathering resources, filtering and validation, and dissemination and storage, ensuring findings are reproducible and accessible through transparent reporting and data sharing.

Clinical reporting requires:

  • Validation of variant calls using an orthogonal method
  • Documentation of all analysis steps and parameters
  • Adherence to applicable regulations and guidelines
  • Clear communication of limitations and uncertainties

Professional Escalation Criteria

Escalate to a bioinformatics specialist or clinical genomics professional when:

  • Variant calls will be used for clinical decision-making
  • You encounter unexpected patterns in variant calls that suggest data quality problems
  • You need to validate variants using an orthogonal method
  • You are uncertain about the interpretation of variant calls in the context of tumor heterogeneity

Building a Somatic Variant Calling Decision Matrix for Your Laboratory

Selecting a somatic variant caller from published benchmarks is only the first step. The harder problem is translating those benchmarks into a defensible decision for your specific samples, sequencing platform, and research question. A structured decision matrix turns subjective impressions into a scored, documented process that you can revisit when data characteristics change or when a new caller version is released. This section provides a practical framework for building that matrix, recording the evidence behind each choice, and troubleshooting the mismatches that commonly arise between caller assumptions and real laboratory data.

Step 1: Define Your Sample and Data Profile Before Any Tool Selection

Before you open a caller manual or run a test command, write down the characteristics of the data you will actually analyze. The tutorial for variant interrogation in tumor samples emphasizes that planning is the foundation of thoughtful experimental design and a clear understanding of sequencing outputs. Use the following checklist as your starting point and record each item in your laboratory notebook or electronic lab notebook:

  • Sequencing platform and chemistry version
  • Library preparation method, including whether it is PCR-based or PCR-free
  • Mean target depth and the distribution of depth across your regions of interest
  • Tumor purity estimate from pathology review or computational tools
  • Whether you have a matched normal sample and what tissue type it came from
  • Whether your study targets SNVs, indels, or both
  • Whether your regions of interest include repetitive elements, GC-rich areas, or known pseudogenes
  • Available compute hours, memory per node, and storage capacity

The circulating tumor DNA benchmarking study demonstrated that caller performance varies substantially with sequencing depth and variant allele frequency. A caller that performs well at 150x whole-genome depth may not be appropriate for a targeted panel at 2000x depth, and vice versa. Documenting these parameters before selection prevents the common error of choosing a caller based on a benchmark that does not match your data profile.

Step 2: Score Each Caller Against Your Data Profile

Create a scoring table with rows for each caller you are considering and columns for the characteristics that matter most for your study. Use a simple scale of 1 to 3 for each criterion, where 1 means poor fit, 2 means acceptable fit, and 3 means strong fit. The criteria should reflect the evidence from published benchmarks and the documented behavior of each caller.

For example, if you are working with low-purity samples or circulating tumor DNA, score VarScan2 higher for its documented ability to handle variable purity and ploidy. If your study focuses on indels in repetitive regions, score Lancet higher because its local assembly approach is designed for those regions. If you have limited computational resources, score VarScan2 and Strelka2 higher for their lower footprints, and score Lancet lower because of its higher runtime and memory requirements.

The VariantMedium benchmarking study provides additional evidence for scoring. It found that performance differences among callers were most pronounced in genomic regions with high sequencing error rates, where VariantMedium achieved higher F1 scores than MuTect2 and Strelka2. If your regions of interest include such areas, factor this evidence into your scoring instead of relying on whole-genome averages.

Record the rationale for each score in a notes column. This documentation is essential for defending your choice in publications, grant applications, and internal reviews. It also provides a baseline for reassessment when you add new sample types or when caller versions change.

Step 3: Run a Pilot Comparison on a Representative Subset

Do not commit to a single caller for your entire cohort before testing on a small, representative subset. Select five to ten samples that reflect the range of purity, depth, and variant burden you expect in your full dataset. Run at least two candidate callers on these samples and compare their outputs systematically.

The structural variant benchmarking study on the HG008 genome established a multi-tool ensemble strategy for SV detection to achieve more accurate and comprehensive identification of somatic variants. A similar approach applies to SNV and indel calling. Running multiple callers on a pilot subset allows you to quantify agreement and disagreement before scaling to your full cohort.

For each pilot sample, record the following measurements:

  • Total number of variants called by each caller
  • Number of variants called by all callers
  • Number of variants called by only one caller
  • Variant allele frequency distribution for each caller
  • Transition to transversion ratio for each caller
  • Runtime and peak memory usage for each caller

The circulating tumor DNA benchmarking study defined a reference set of approximately 37,000 single nucleotide variants and 58,000 indels to benchmark nine somatic variant callers. While you may not have a reference set for your own samples, you can use cross-caller agreement as a proxy for confidence. Variants called by multiple callers are higher confidence candidates, while variants called by only one caller require additional scrutiny.

Step 4: Weight Sensitivity and Specificity According to Your Research Question

The relative importance of sensitivity and specificity depends on your downstream application. If you are discovering novel candidate variants for functional validation, you may prioritize sensitivity and accept a higher false positive rate. If you are reporting variants for clinical or translational purposes, you may prioritize specificity to avoid reporting false positives.

The tutorial for variant interrogation in tumor samples emphasizes executing a systematic approach to prioritize meaningful variants. Your decision matrix should reflect this prioritization. Assign weights to sensitivity and specificity based on your research question, then multiply those weights by your caller scores to produce a weighted total.

For example, if sensitivity is twice as important as specificity for your discovery study, multiply the sensitivity score by 2 and the specificity score by 1 before summing. Document these weights and the rationale behind them. This transparency helps reviewers understand why you chose a particular caller and allows you to adjust the weights if your research question evolves.

Step 5: Document the Decision and Set a Review Trigger

Once you have completed the scoring and pilot comparison, record the final decision in your laboratory documentation. Include the following elements:

  • The caller selected and its exact version number
  • The parameters used and the rationale for each parameter
  • The pilot samples used for comparison and the measurements obtained
  • The weighted scores for each caller and the rationale for the weights
  • The date of the decision and the names of the people involved

Set a review trigger for reassessment. Common triggers include:

  • A new major version of your selected caller
  • A new benchmarking study that directly compares callers on data similar to yours
  • A change in your sequencing platform or library preparation method
  • A change in your research question that alters the sensitivity-specificity balance
  • An unexpected pattern in your variant calls that suggests caller-data mismatch

The nf-core documentation describes community pipeline standards, usage, and configuration for reproducible workflow context. Incorporating your decision matrix into a version-controlled workflow ensures that the rationale behind your caller choice is preserved alongside the analysis code.

Common Failure Patterns in Caller Selection

Failure Pattern 1: Choosing a Caller Based on a Benchmark That Does Not Match Your Data

Published benchmarks are valuable, but they are only informative if your data resembles the benchmark data. The VariantMedium study used 336,839 variants from 2,956 samples with whole-exome or whole-genome sequencing for training and validation. The circulating tumor DNA study used deep whole-genome sequencing at 150x and exome sequencing at 2000x. If your data has different depth, purity, or library preparation characteristics, the benchmark results may not transfer.

The remedy is to run a pilot comparison on your own data, as described in Step 3. If you cannot run a pilot, at minimum compare the characteristics of your data to the benchmark data and document any differences that could affect caller performance.

Failure Pattern 2: Ignoring the Matched Normal Requirement

All four callers covered in this article require a matched normal sample for reliable somatic variant calling. Some callers offer tumor-only modes, but these should be used with caution. Without a matched normal, you cannot reliably distinguish somatic variants from germline polymorphisms, and the risk of false positive calls increases substantially.

If you do not have a matched normal, document this limitation and consider whether your research question can be answered with tumor-only calling. If you proceed, interpret results with appropriate caveats and consider orthogonal validation of candidate variants.

Failure Pattern 3: Treating Caller Outputs as Final Results

Somatic variant callers produce candidate variant lists, not validated biological findings. The tutorial for variant interrogation in tumor samples emphasizes that filtering and validation are critical steps in the workflow. Apply systematic filtering criteria after calling, including variant allele frequency thresholds, read depth at the variant position, mapping quality and base quality scores, presence in population databases, and predicted functional impact.

The National Center for Biotechnology Information provides access to databases that can be used for filtering, including dbSNP and gnomAD. The European Bioinformatics Institute offers training on data resources that can support variant annotation and interpretation. Incorporate these resources into your post-calling workflow.

Failure Pattern 4: Not Recording the Decision Rationale

The most common failure in caller selection is not documenting why a particular caller was chosen. Without documentation, you cannot defend your choice in publications or grant applications, and you cannot reassess the decision when conditions change. The decision matrix described in this section provides a structured format for recording the evidence, scores, and rationale behind your selection.

Records and Measurements for Caller Selection

Maintain a dedicated section in your laboratory documentation for caller selection records. Include the following for each decision:

  • The data profile checklist from Step 1
  • The scoring table from Step 2 with notes for each score
  • The pilot comparison measurements from Step 3
  • The weighted sensitivity and specificity scores from Step 4
  • The final decision and review trigger from Step 5

The Carpentries lessons cover foundational computing, data, shell, Git, and programming training that supports reproducible analysis practices. Version control your decision matrix alongside your analysis code so that the rationale is preserved with the workflow.

Professional Escalation Criteria

Escalate to a bioinformatics specialist or clinical genomics professional when:

  • Your pilot comparison shows substantial disagreement between callers that you cannot explain
  • Your variant calls show unexpected patterns that suggest data quality problems
  • You need to validate variants using an orthogonal method
  • Your variant calls will be used for clinical decision-making
  • You are uncertain about the interpretation of variant calls in the context of tumor heterogeneity

The Galaxy Training Network provides accessible workflow training and analysis tutorials that can help you build and validate somatic variant calling pipelines. The nf-core documentation describes community pipeline standards that can support reproducible workflow context. Use these resources to strengthen your analysis before escalating to a specialist.

Frequently Asked Questions

What is the difference between somatic and germline variant calling?

Somatic variant calling identifies mutations present in tumor tissue but absent from matched normal tissue. Germline variant calling identifies inherited variants present in all cells of an individual. Somatic variant calling requires a matched normal sample to distinguish true somatic mutations from inherited polymorphisms and sequencing artifacts. The statistical models and filtering criteria differ between the two approaches because somatic variants are often present at low allele frequencies due to tumor heterogeneity and contamination with normal cells.

Which somatic variant caller is best for low-purity tumor samples?

VarScan2 can handle variable tumor purity and ploidy, making it a reasonable choice for low-purity samples. However, the circulating tumor DNA benchmarking study provides practical guidance for selecting calling methods at varying variant allele frequencies and sequencing depths. For very low variant allele frequencies, consider using deeper sequencing or machine learning-based approaches that have been trained on experimentally confirmed variant data. No single caller performs optimally across all purity levels, so validate your choice on your specific data.

Do I need a matched normal sample for somatic variant calling?

Yes, a matched normal sample is required for reliable somatic variant calling. The matched normal provides the baseline for distinguishing somatic mutations from germline polymorphisms. Without a matched normal, you cannot reliably distinguish somatic variants from inherited variants, and the risk of false positive calls increases substantially. Some callers offer tumor-only modes, but these should be used with caution and results should be interpreted with appropriate caveats.

How do I choose between MuTect2 and Strelka2?

MuTect2 and Strelka2 are both widely used for somatic SNV and indel detection. MuTect2 integrates with the GATK ecosystem and offers a panel of normals feature that can improve specificity. Strelka2 offers efficient runtime and good sensitivity for both SNVs and indels, but requires a candidate indel file for optimal indel detection. The VariantMedium benchmarking study found that both callers achieved lower F1 scores than VariantMedium in genomic regions with high sequencing error rates. Consider your data type, computational resources, and downstream analysis requirements when choosing between these two callers.

What is the role of local assembly in somatic variant calling?

Local assembly reconstructs the sequence of a genomic region from the sequencing reads, which can detect variants that read-alignment-based methods miss. Lancet uses local assembly and is particularly effective for indels and variants in repetitive regions. However, local assembly is computationally intensive and has higher memory usage than alignment-based methods. Use Lancet as a complementary caller to validate variants detected by other methods, particularly in difficult genomic regions.

How do I validate somatic variant calls?

Validation of somatic variant calls typically involves an orthogonal method such as targeted deep sequencing, droplet digital PCR, or Sanger sequencing. The tutorial for variant interrogation in tumor samples provides a practical framework for filtering and validation. Prioritize variants for validation based on their potential biological significance, allele frequency, and consistency across multiple callers. Document all validation results and integrate them into your final variant call set.

What are the limitations of somatic variant calling in circulating tumor DNA?

Circulating tumor DNA is challenging due to low variant allele frequencies and extensive DNA degradation. The circulating tumor DNA benchmarking study defined a reference set of approximately 37,000 single nucleotide variants and 58,000 indels to benchmark nine somatic variant callers across varying circulating tumor DNA levels and sequencing depths. Detection limits vary with sequencing depth and variant allele frequency, and sensitivity decreases substantially at very low allele frequencies. Consider using deeper sequencing and machine learning-based tuning to improve accuracy in liquid biopsy applications.

How do I ensure reproducibility in somatic variant calling?

Reproducibility requires documentation of all analysis steps, tool versions, parameters, and reference genome versions. Use containerized environments or workflow managers to standardize the analysis environment. The nf-core documentation describes community pipeline standards, usage, and configuration for reproducible workflow context. The Galaxy Training Network provides accessible workflow training and analysis tutorials that emphasize reproducibility. The Carpentries lessons cover foundational computing, data, shell, Git, and programming training that supports reproducible analysis practices.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.