# Manta vs. Lumpy vs. Delly: A Head-to-Head Comparison of Structural Variant Callers for Germline and Somatic Data


## Key Takeaways

- Manta demonstrates superior performance for deletion and insertion detection in short-read germline data, exhibiting high genotype concordance and efficient computational resource utilization, making it ideal for large-scale whole-genome sequencing projects.
- Lumpy's probabilistic framework allows for the integration of multiple SV evidence types, offering flexibility for germline analysis, though it requires careful parameter tuning and shows poor sensitivity for insertions in short-read data.
- Delly, while effective for deletions and some insertions in germline studies, demands substantial computational resources for massive datasets and performs less optimally for duplications and inversions compared to deletions.
- Short-read sequencing exhibits low sensitivity for insertion detection (22% in one study), necessitating consideration of long-read sequencing or complementary callers when insertions are a primary research focus.
- Duplications and inversions remain challenging for Manta, Lumpy, and Delly; specialized callers like Canvas, CNVnator (for duplications), or Sniffles2 and Severus (for complex inversions) may be required for comprehensive detection.
- Alignment choice significantly impacts SV calling, particularly for inversions in repetitive regions, underscoring the importance of selecting and documenting the aligner (e.g., BWA-MEM for short reads, Minimap2 for long reads) for reproducible results.

---

Structural variant (SV) calling is a core step in genomic analysis, yet the choice of caller can change the variants you report. This article compares Manta, Lumpy, and Delly for germline and somatic data, using published benchmarks and official documentation to guide your selection based on sensitivity, precision, and computational cost. You will learn how each caller handles deletions, insertions, duplications, and inversions, how they perform on short-read and long-read data, and which workflow decisions matter most for reproducible results.

## Scope and Reader Context

This comparison is written for bioinformatics students, researchers, laboratory professionals, and life-science practitioners who need to select an SV caller for a specific project. The three callers examined here, Manta, Lumpy, and Delly, are widely used for whole-genome sequencing data. The evidence base includes a 2024 benchmark of 11 SV callers, a 2024 comparison of short-read and nanopore sequencing against optical genome mapping, a 2019 long-read sequencing study, a 2024 study on read order sensitivity, and a 2025 preprint on inversion benchmarks. Official training and documentation resources from NCBI, EMBL-EBI, Bioconductor, Galaxy, nf-core, and The Carpentries provide context for workflow design and reproducibility.

The practical outcome of this article is a decision framework. You will learn which caller to use for germline deletions, which caller handles insertions with better precision, and how to combine callers when sensitivity matters more than speed. You will also learn the limitations of each approach, including the impact of read order, alignment choice, and sequencing depth on your results.

## At a Glance

The table below summarizes the key characteristics of Manta, Lumpy, and Delly based on published benchmark data and official documentation. Use this table for a quick comparison before reading the detailed sections.

| Caller | Best Performing Variant Types | Notable Strength | Known Limitation | Computational Profile |
|--------|-------------------------------|------------------|------------------|------------------------|
| Manta | Deletions, insertions | Highest concordance for deletions and insertions in genotype verification | Lower performance for duplications and inversions | Efficient computing resources, good for large datasets |
| Lumpy | Deletions (with paired-end and split-read evidence) | Probabilistic framework for SV evidence integration | Poor sensitivity for insertions in short-read data | Moderate memory usage, requires careful parameter tuning |
| Delly | Deletions, some insertions | Integrates paired-end and split-read signals, widely used in population studies | Better for deletions than duplications, inversions, and insertions | Substantial computational resources for massive datasets |

The 2024 BMC Genomics benchmark compared 11 SV callers including Delly, Manta, Lumpy, and others on massive whole-genome sequencing data. The study evaluated accuracy, sequence depth, running time, and memory usage. Manta identified deletion SVs with better performance and efficient computing resources, and both Manta and MELT demonstrated relatively good precision for insertions. The copy number variation callers Canvas and CNVnator performed better for long duplications because they use a read-depth approach. Genotype verification using a phased long-read assembly dataset showed Manta had the highest concordance for deletions and insertions. See the [BMC Genomics benchmark](https://pubmed.ncbi.nlm.nih.gov/38549092) for the full methodology and results.

## Structural Variant Types and Detection Challenges

Structural variants include deletions, duplications, insertions, inversions, and translocations. Each type presents different detection challenges, and no single caller performs equally well across all types. Understanding these challenges helps you interpret why Manta, Lumpy, and Delly differ in their results.

### Deletions

Deletions are the most reliably detected SV type in short-read data. The 2024 Genes study using optical genome mapping as a benchmark found that sensitivity for deletions in Illumina data was high at 86 percent, with 115 of 134 true deletions identified. This high sensitivity reflects the strong signal that deletions produce in paired-end and split-read data. Manta showed the best performance for deletions among the 11 callers tested in the BMC Genomics benchmark. See the [Genes comparison study](https://pubmed.ncbi.nlm.nih.gov/39062704) for details on deletion sensitivity across platforms.

### Insertions

Insertions are more difficult to detect than deletions in short-read data. The same Genes study found that sensitivity for insertions in Illumina data was poor at 22 percent, with only 13 of 58 true insertions identified. Manta and MELT demonstrated relatively good precision for insertions in the BMC Genomics benchmark. If your project focuses on insertions, you should consider whether short-read sequencing is appropriate or whether long-read sequencing would provide better sensitivity.

### Duplications and Inversions

Duplications and inversions remain challenging for all three callers. The BMC Genomics benchmark noted that several callers exhibited better calling performance for deletions than for duplications, inversions, and insertions. The 2025 inversion benchmark preprint found that simple inversions were recovered with high sensitivity at sufficient coverage, while complex and heterozygous events remained difficult. Sniffles2 and Severus achieved the strongest recall for complex inversions, but these callers are outside the scope of this comparison. See the [inversion benchmark preprint](https://pubmed.ncbi.nlm.nih.gov/41377520) for details on inversion detection across platforms.

## Core Principles of Structural Variant Calling

Structural variant callers use different evidence types to identify variants. The three main evidence types are paired-end reads, split reads, and read depth. Each caller combines these signals differently, which explains their performance differences.

### Paired-End and Split-Read Evidence

Paired-end reads provide information about the distance and orientation between two reads that originate from the same DNA fragment. When a structural variant is present, the insert size or orientation of the pair changes. Split reads provide evidence when a single read maps to two different genomic locations, indicating a breakpoint. Manta, Lumpy, and Delly all use paired-end and split-read evidence, but they weight these signals differently.

### Read Depth and Coverage

Read depth approaches detect copy number changes by comparing the number of reads in a region to the expected coverage. Canvas and CNVnator use this approach and performed better for long duplications in the BMC Genomics benchmark. Manta, Lumpy, and Delly primarily use breakpoint evidence instead of read depth, which explains their lower performance for duplications.

### Alignment Choice and Its Impact

The choice of aligner affects SV calling results. The 2025 inversion benchmark demonstrated that mapper choice has a substantial impact on inversion detection in repetitive regions. The study compared Minimap2 and VACmap for long-read data and found that the mapper choice changed results. For short-read data, the choice between BWA-MEM and other aligners can similarly affect SV calling. The [EMBL-EBI training resources](https://www.ebi.ac.uk/training) provide guidance on alignment and variant calling workflows.

## Manta: Assembly-Based and Paired-End Approach

Manta is a structural variant caller that uses a combination of paired-end and split-read evidence with local assembly to refine breakpoints. The BMC Genomics benchmark found that Manta identified deletion SVs with better performance and efficient computing resources. Manta also showed the highest concordance for deletions and insertions when genotypes were verified using a phased long-read assembly dataset.

### Strengths of Manta

Manta performs well for deletions and insertions in short-read data. The efficient computing resource usage makes it suitable for large whole-genome sequencing datasets. The high genotype concordance for deletions and insertions means that variants identified by Manta are more likely to be confirmed by orthogonal methods.

### Limitations of Manta

Manta does not perform as well for duplications and inversions. The BMC Genomics benchmark noted that several callers exhibited better calling performance for deletions than for duplications, inversions, and insertions. If your project focuses on these variant types, you should consider supplementing Manta with a read-depth based caller or a long-read approach.

### When to Use Manta

Use Manta for germline deletion discovery in short-read whole-genome sequencing data. Use Manta when you need efficient computing resource usage for large datasets. Use Manta when you need high genotype concordance for downstream analysis.

## Lumpy: Probabilistic Framework for SV Evidence

Lumpy uses a probabilistic framework to integrate multiple sources of SV evidence, including paired-end reads, split reads, and read depth. This approach allows Lumpy to identify variants that might be missed by callers using a single evidence type.

### Strengths of Lumpy

Lumpy provides a flexible framework for integrating different evidence types. The probabilistic approach allows for the combination of signals from multiple sources, which can improve sensitivity for certain variant types. Lumpy is widely used in population-scale studies and has been included in multiple benchmark comparisons.

### Limitations of Lumpy

Lumpy requires careful parameter tuning to achieve optimal performance. The sensitivity for insertions in short-read data is poor, consistent with the general limitation of short-read sequencing for insertion detection. The BMC Genomics benchmark included Lumpy in the comparison of 11 callers and found that several callers exhibited better calling performance for deletions than for duplications, inversions, and insertions.

### When to Use Lumpy

Use Lumpy when you want to integrate multiple evidence types in a single probabilistic framework. Use Lumpy when you have time to tune parameters for your specific dataset. Use Lumpy when you are working with germline data and need a caller that can be combined with other tools for improved sensitivity.

## Delly: Integrated Paired-End and Split-Read Calling

Delly integrates paired-end and split-read signals to identify structural variants. The caller is widely used in population studies and has been applied to large-scale whole-genome sequencing datasets.

### Strengths of Delly

Delly provides a well-established approach for germline SV calling. The integration of paired-end and split-read signals allows for the detection of deletions and some insertions. Delly has been included in multiple benchmark studies, providing a basis for comparing its performance with other callers.

### Limitations of Delly

The BMC Genomics benchmark found that Delly exhibited better calling performance for deletions than for duplications, inversions, and insertions. The computational resources required for massive whole-genome sequencing datasets are substantial. If your project requires efficient computing resource usage, Delly may not be the optimal choice.

### When to Use Delly

Use Delly for germline deletion discovery in population-scale studies. Use Delly when you need a caller with a track record of application to large datasets. Use Delly when you can allocate sufficient computational resources for the analysis.

## Germline Variant Calling Workflow

Germline variant calling identifies variants that are present in the germline genome, typically from a blood or tissue sample. The workflow involves alignment, SV calling, and filtering. The choice of caller depends on your project goals and available resources.

### Sample Preparation and Sequencing

The quality of your sequencing data affects SV calling results. The 2024 Genes study prepared high-quality DNA from 9 parent-child trios for comparison of short-read and nanopore sequencing. High-quality DNA extraction and library preparation are essential for reliable SV calling. The [NCBI data resources](https://www.ncbi.nlm.nih.gov/) provide access to reference genomes and sequence data for comparison.

### Alignment Strategy

The choice of aligner affects SV calling results. For short-read data, BWA-MEM is commonly used. For long-read data, Minimap2 and NGMLR are options. The 2019 Genome Research study found that Sniffles after NGMLR or minimap2 alignment provided the most accurate results for long-read data. The [Galaxy training resources](https://training.galaxyproject.org/) provide tutorials on alignment and variant calling workflows.

### Running the Caller

Each caller has specific command-line options and input requirements. Manta requires a configuration step followed by the run step. Lumpy requires the creation of a lumpy express configuration file. Delly requires a single command for calling. The [nf-core documentation](https://nf-co.re/docs) provides guidance on running SV callers within reproducible pipelines.

### Filtering and Annotation

After SV calling, you need to filter variants based on quality scores, read support, and population frequency. The [Bioconductor project](https://bioconductor.org/) provides packages for SV annotation and filtering. The [EMBL-EBI training resources](https://www.ebi.ac.uk/training) offer courses on variant annotation and interpretation.

## Somatic Variant Calling Workflow

Somatic variant calling identifies variants that are present in tumor tissue but not in normal tissue. The workflow differs from germline calling because you need to compare tumor and normal samples to distinguish somatic variants from germline variants.

### Tumor-Normal Comparison

Somatic SV calling requires a matched normal sample to filter out germline variants. Manta provides a somatic mode that takes tumor and normal samples as input. The somatic mode identifies variants present in the tumor but not in the normal sample. The [Galaxy training resources](https://training.galaxyproject.org/) provide tutorials on somatic variant calling workflows.

### Sensitivity and Precision Tradeoffs

Somatic SV calling involves a tradeoff between sensitivity and precision. High sensitivity ensures that you identify true somatic variants, but it may also increase false positives. High precision reduces false positives but may miss true variants. The BMC Genomics benchmark provides performance data that can inform your choice of caller for somatic analysis.

### Computational Considerations

Somatic SV calling requires more computational resources than germline calling because you need to analyze tumor and normal samples. The BMC Genomics benchmark evaluated running time and memory usage for 11 SV callers. Manta demonstrated efficient computing resources, making it a suitable choice for somatic analysis.

## Options and Tradeoffs in Caller Selection

The choice of SV caller involves tradeoffs between sensitivity, precision, and computational cost. The following sections describe the key tradeoffs you need to consider.

### Sensitivity vs. Precision

Sensitivity measures the proportion of true variants that are identified. Precision measures the proportion of identified variants that are true. The BMC Genomics benchmark found that Manta identified deletion SVs with better performance and efficient computing resources. The 2024 Genes study found that sensitivity varied according to variant type, being high for deletions but poor for insertions in short-read data.

### Computational Cost

The computational cost of SV calling includes running time and memory usage. The BMC Genomics benchmark evaluated these metrics for 11 callers. Manta demonstrated efficient computing resources. Delly required substantial computational resources for massive whole-genome sequencing datasets. If computational cost is a constraint, Manta may be the preferred choice.

### Combining Multiple Callers

The 2019 Genome Research study found that additional confidence or sensitivity can be obtained by a combination of multiple variant callers. Combining Manta, Lumpy, and Delly can improve sensitivity but also increases computational cost. The study described a scalable workflow for identification, annotation, and characterization of tens of thousands of structural variants from long-read genome sequencing.

## Observations and Measurements for Caller Performance

Published benchmarks provide quantitative data on caller performance. The following sections describe the key measurements you should consider when evaluating SV callers.

### Accuracy by Variant Type

The BMC Genomics benchmark evaluated accuracy for deletions, duplications, inversions, and insertions. Several callers exhibited better calling performance for deletions than for duplications, inversions, and insertions. Manta identified deletion SVs with better performance. Manta and MELT demonstrated relatively good precision for insertions.

### Genotype Concordance

The BMC Genomics benchmark verified genotypes inferred from each SV caller using a phased long-read assembly dataset. Manta showed the highest concordance for deletions and insertions. This finding suggests that Manta genotypes are more reliable for downstream analysis.

### Sensitivity by Sequencing Platform

The 2024 Genes study compared sensitivity across Illumina and Oxford Nanopore platforms. In the Illumina dataset, sensitivity was high for deletions at 86 percent but poor for insertions at 22 percent. In the ONT dataset, sensitivity was generally poor using the original Sniffles caller at 48 percent overall but improved substantially with Sniffles2. The study concluded that the precision of optical genome mapping is very high and that Sniffles2 with ONT data outperforms Illumina for most SV types.

### Read Order Sensitivity

The 2024 PeerJ study found that the order of input data affected the SVs predicted by each caller. The study used PacBio data from 15 Caenorhabditis elegans strains and four Arabidopsis thaliana ecotypes. The pbsv caller was highly sensitive to the order of input data, especially at the highest depths where over 70 percent of SV calls were in disagreement. The SAMtools alignment sorting algorithm was also identified as a source of variability. These findings have implications for the replication of SV studies and the development of consistent SV calling protocols. See the [PeerJ read order study](https://pubmed.ncbi.nlm.nih.gov/38500526) for details.

## Records and Measurements for Your Own Analysis

When you run SV callers on your own data, you should keep records of the parameters and metrics that affect your results. The following sections describe the records you should maintain.

### Sequencing Depth and Coverage

Record the sequencing depth and coverage for each sample. The BMC Genomics benchmark evaluated the effect of sequence depth on caller performance. Higher depth generally improves sensitivity but increases computational cost. The 2024 Genes study used high-quality DNA from parent-child trios for their comparison.

### Caller Version and Parameters

Record the version of each caller and the parameters you used. Caller versions change over time, and parameter choices affect results. The [nf-core documentation](https://nf-co.re/docs) provides guidance on version control and parameter management in reproducible pipelines.

### Alignment Metrics

Record the alignment metrics for each sample, including the aligner version, alignment parameters, and mapping quality. The 2025 inversion benchmark demonstrated that mapper choice has a substantial impact on inversion detection in repetitive regions. The [EMBL-EBI training resources](https://www.ebi.ac.uk/training) provide guidance on alignment quality assessment.

### Filtering Criteria

Record the filtering criteria you applied to the SV calls. Filtering based on quality scores, read support, and population frequency affects the final variant set. The [Bioconductor project](https://bioconductor.org/) provides packages for SV filtering and annotation.

## Quality Controls and Reproducibility

Reproducibility is a core concern in bioinformatics. The following sections describe quality controls and practices that improve the reproducibility of your SV calling results.

### Version Control and Containerization

Use version control for your analysis scripts and containerization for your software environment. The [nf-core documentation](https://nf-co.re/docs) provides guidance on reproducible pipeline standards. The [Carpentries lessons](https://carpentries.org/lessons) offer training on version control with Git and reproducible computing practices.

### Benchmarking Against Truth Sets

Benchmark your SV calling results against truth sets when available. The 2024 Genes study used optical genome mapping as a benchmark to establish a truth dataset. The 2025 inversion benchmark provides a multi-genome benchmark for simple and complex inversions. The [NCBI data resources](https://www.ncbi.nlm.nih.gov/) provide access to reference materials and benchmark datasets.

### Replicate Analysis

Run your analysis with different read orders to assess the stability of your results. The 2024 PeerJ study found that read order affected SV calling results. The study suggested that researchers should be aware of this sensitivity when developing consistent SV calling protocols.

## Common Failure Patterns in SV Calling

Understanding common failure patterns helps you troubleshoot your SV calling results. The following sections describe patterns observed in published benchmarks.

### Low Sensitivity for Insertions

Short-read data has poor sensitivity for insertions. The 2024 Genes study found that sensitivity for insertions in Illumina data was 22 percent. If your project focuses on insertions, consider long-read sequencing or a combination of callers.

### Poor Performance for Complex Inversions

The 2025 inversion benchmark found that complex and heterozygous events remained difficult to detect. Simple inversions were recovered with high sensitivity at sufficient coverage, but complex events required specialized callers. If your project focuses on inversions, consider using Sniffles2 or Severus in addition to Manta, Lumpy, or Delly.

### Read Order Sensitivity

The 2024 PeerJ study found that read order affected SV calling results. The pbsv caller was highly sensitive to read order, and the SAMtools alignment sorting algorithm was a source of variability. If your results are unstable across runs, check whether read order is affecting your analysis.

### Alignment Choice Effects

The 2025 inversion benchmark found that mapper choice has a substantial impact on inversion detection in repetitive regions. If your results differ from published benchmarks, check whether your alignment strategy matches the benchmark conditions.

## Limitations of the Evidence Base

The evidence base for this comparison has several limitations that you should consider when interpreting the results.

### Benchmark Data Types

The BMC Genomics benchmark used whole-genome sequencing data, which may not reflect performance on targeted or exome sequencing data. The 2024 Genes study used data from the Genomics England 100,000 Genomes Project, which may not represent all populations.

### Caller Version Differences

The benchmarks used specific versions of each caller. Newer versions may have improved performance. You should check the documentation for each caller to understand the version-specific changes.

### Platform Differences

The benchmarks used specific sequencing platforms. The 2024 Genes study compared Illumina and Oxford Nanopore platforms. The 2019 Genome Research study used Oxford Nanopore PromethION. Performance on other platforms may differ.

### Preprint Status

The 2025 inversion benchmark is a preprint and has not undergone peer review. You should interpret its findings with caution and check for updates to the manuscript.

## Safety and Regulatory Context

Structural variant calling has applications in clinical genomics and genetic research. The following sections describe the safety and regulatory context you should consider.

### Clinical Applications

If your SV calling results are used for clinical decision-making, you should follow applicable regulations and guidelines. The [NCBI data resources](https://www.ncbi.nlm.nih.gov/) provide access to clinical databases and resources. The [EMBL-EBI training resources](https://www.ebi.ac.uk/training) offer courses on clinical genomics and variant interpretation.

### Data Privacy and Security

Genomic data is sensitive and should be handled according to applicable privacy regulations. The [Carpentries lessons](https://carpentries.org/lessons) provide training on responsible data handling. The [nf-core documentation](https://nf-co.re/docs) provides guidance on secure data processing in reproducible pipelines.

### Professional Escalation Criteria

If your SV calling results are unexpected or inconsistent with clinical findings, you should escalate the issue to a qualified professional. The following criteria indicate when escalation is appropriate:

- Results that contradict established clinical findings
- Results that are inconsistent across replicate analyses
- Results that suggest a variant with known clinical significance
- Results that are affected by known technical artifacts

## Practical Decision Framework for SV Caller Selection

Selecting between Manta, Lumpy, and Delly requires a structured approach that accounts for your specific project goals, data characteristics, and available computational resources. Published benchmarks provide useful performance comparisons, but they do not replace a systematic evaluation on your own data. This section presents a decision framework you can apply before committing to a single caller, along with a record system for tracking performance metrics and a troubleshooting method for common failures.

### Step 1: Define Your Primary Variant Types and Priorities

The first decision point is identifying which structural variant types matter most for your biological question. The 2024 BMC Genomics benchmark demonstrated that no single caller performs equally well across all variant types. Several callers exhibited better calling performance for deletions than for duplications, inversions, and insertions. Manta identified deletion SVs with better performance and efficient computing resources, while Manta and MELT demonstrated relatively good precision for insertions. See the [BMC Genomics benchmark](https://pubmed.ncbi.nlm.nih.gov/38549092) for the full performance comparison.

For each project, write down the variant types you must detect with high sensitivity and the types where moderate sensitivity is acceptable. This prioritization directly determines your caller choice. If deletions are your primary target, Manta is the strongest candidate based on the benchmark evidence. If insertions matter, Manta remains the best short-read option among the three callers, but you should also consider whether long-read sequencing would better serve your project. The 2024 Genes study found that sensitivity for insertions in Illumina data was poor at 22 percent, with only 13 of 58 true insertions identified. See the [Genes comparison study](https://pubmed.ncbi.nlm.nih.gov/39062704) for platform-specific sensitivity data.

If duplications or inversions are central to your research question, none of the three callers in this comparison is ideal. The BMC Genomics benchmark found that copy number variation callers Canvas and CNVnator exhibited better performance in identifying long duplications because they employ the read-depth approach. The 2025 inversion benchmark preprint found that Sniffles2 and Severus achieved the strongest recall for complex inversions. See the [inversion benchmark preprint](https://pubmed.ncbi.nlm.nih.gov/41377520) for details on inversion detection across platforms. In these cases, you should plan to supplement Manta, Lumpy, or Delly with additional tools instead of relying on a single caller.

### Step 2: Assess Your Data Characteristics

Your sequencing platform, depth, and sample type influence caller performance in ways that published benchmarks may not fully capture. The 2024 Genes study compared short-read Illumina data with Oxford Nanopore long-read data using optical genome mapping as a benchmark. In the Illumina dataset, sensitivity varied according to variant type, being high for deletions at 86 percent but poor for insertions at 22 percent. In the ONT dataset, sensitivity was generally poor using the original Sniffles variant caller at 48 percent overall but improved substantially with Sniffles2. See the [Genes comparison study](https://pubmed.ncbi.nlm.nih.gov/39062704) for the full platform comparison.

For short-read data, record your sequencing depth and coverage. The BMC Genomics benchmark evaluated the effect of sequence depth on caller performance. Higher depth generally improves sensitivity but increases computational cost. For long-read data, the 2019 Genome Research study found that Sniffles after NGMLR or minimap2 alignment provided the most accurate results, but additional confidence or sensitivity can be obtained by a combination of multiple variant callers. See the [Genome Research long-read study](https://pubmed.ncbi.nlm.nih.gov/31186302) for the workflow description.

Your sample type also matters. Germline calling uses a single sample or family cohort, while somatic calling requires tumor-normal comparison. Manta provides a somatic mode that takes tumor and normal samples as input. Lumpy and Delly are primarily designed for germline analysis. If your project involves somatic SV calling, Manta is the most direct choice among the three callers.

### Step 3: Evaluate Computational Constraints

The BMC Genomics benchmark evaluated running time and memory usage for 11 SV callers on massive whole-genome sequencing datasets. Manta demonstrated efficient computing resources, making it suitable for large datasets. Delly required substantial computational resources for massive datasets. Lumpy requires moderate memory usage but needs careful parameter tuning. See the [BMC Genomics benchmark](https://pubmed.ncbi.nlm.nih.gov/38549092) for the computational performance comparison.

Before selecting a caller, estimate your total computational burden. Multiply the number of samples by the expected running time per sample for each caller. Consider whether your computing environment can handle the peak memory usage. If you are processing hundreds of whole-genome samples, the efficiency differences between callers become substantial. Manta is the most computationally efficient choice among the three callers based on the benchmark evidence.

If you are working within a shared computing cluster, also consider queue time and storage requirements. The [nf-core documentation](https://nf-co.re/docs) provides guidance on configuring reproducible pipelines for high-throughput analysis. The [Galaxy training resources](https://training.galaxyproject.org/) offer tutorials on running SV callers in accessible computing environments.

### Step 4: Run a Pilot Comparison on a Representative Subset

Before committing to a single caller for your full dataset, run a pilot comparison on a representative subset of samples. This pilot should include at least one sample with known or suspected structural variants, if available. The goal is to assess how each caller performs on your specific data instead of relying solely on published benchmarks.

For the pilot, run Manta, Lumpy, and Delly on the same aligned data using their recommended parameters. Record the number of variants called by each tool, the variant type distribution, and the computational resources used. If you have a truth set or orthogonal validation data, such as optical genome mapping results, compare the calls against this reference. The 2024 Genes study used Bionano optical genome mapping to establish a truth dataset and found that OGM calls have high precision with a positive predictive value of 95 percent. See the [Genes comparison study](https://pubmed.ncbi.nlm.nih.gov/39062704) for the truth set methodology.

The [NCBI data resources](https://www.ncbi.nlm.nih.gov/) provide access to reference genomes and benchmark datasets that can support your pilot evaluation. The [EMBL-EBI training resources](https://www.ebi.ac.uk/training) offer guidance on designing and interpreting benchmark comparisons.

### Step 5: Apply a Weighted Scoring System

To make your caller selection transparent and reproducible, apply a weighted scoring system based on your project priorities. Create a table with rows for each caller and columns for each evaluation criterion. Assign weights to each criterion based on your project goals. Common criteria include sensitivity for your primary variant types, precision, genotype concordance, running time, memory usage, and ease of parameter tuning.

For each criterion, assign a score from 1 to 5 based on the published benchmark evidence and your pilot results. Multiply each score by the criterion weight and sum the totals. The caller with the highest weighted total is your primary choice. Document this scoring process in your analysis records so that other researchers can understand your selection rationale.

The BMC Genomics benchmark provides quantitative data you can use for scoring. Manta showed the highest concordance for deletions and insertions when genotypes were verified using a phased long-read assembly dataset. Manta and MELT demonstrated relatively good precision for insertions. See the [BMC Genomics benchmark](https://pubmed.ncbi.nlm.nih.gov/38549092) for the genotype concordance results.

### Step 6: Plan for Caller Combination When Sensitivity Matters

If your project requires maximum sensitivity, plan to combine multiple callers instead of relying on a single tool. The 2019 Genome Research study found that additional confidence or sensitivity can be obtained by a combination of multiple variant callers. The study described a scalable workflow for identification, annotation, and characterization of tens of thousands of structural variants from long-read genome sequencing. See the [Genome Research long-read study](https://pubmed.ncbi.nlm.nih.gov/31186302) for the combination workflow.

When combining callers, decide whether you will take the union of calls, the intersection, or a consensus approach. The union maximizes sensitivity but increases false positives. The intersection maximizes precision but may miss true variants. A consensus approach that requires support from at least two callers provides a balance between sensitivity and precision. The [Bioconductor project](https://bioconductor.org/) provides packages for comparing and combining variant calls from multiple sources.

### Record System for SV Calling Projects

Maintain a structured record for each SV calling project to ensure reproducibility and to support troubleshooting. The following records should be kept for every sample and analysis run.

#### Sample and Sequencing Records

Record the sample identifier, tissue type, and DNA extraction method. Record the sequencing platform, read length, insert size, and sequencing depth. The 2024 Genes study prepared high-quality DNA from 9 parent-child trios for their platform comparison, demonstrating the importance of sample quality for reliable SV calling. See the [Genes comparison study](https://pubmed.ncbi.nlm.nih.gov/39062704) for the sample preparation methodology.

#### Alignment Records

Record the aligner version, alignment parameters, and reference genome version. The 2025 inversion benchmark demonstrated that mapper choice has a substantial impact on inversion detection in repetitive regions. The study compared Minimap2 and VACmap for long-read data and found that the mapper choice changed results. See the [inversion benchmark preprint](https://pubmed.ncbi.nlm.nih.gov/41377520) for the mapper comparison. For short-read data, record whether you used BWA-MEM or another aligner and note any non-default parameters.

#### Caller Records

Record the caller version, command-line parameters, and configuration files for each run. Caller versions change over time, and parameter choices affect results. The [nf-core documentation](https://nf-co.re/docs) provides guidance on version control and parameter management in reproducible pipelines. Record the date of each run and the computing environment used.

#### Output Records

Record the number of variants called by type, the quality score distributions, and the filtering criteria applied. Record the number of variants passing each filtering step. The [Bioconductor project](https://bioconductor.org/) provides packages for SV filtering and annotation that can support this record keeping.

#### Validation Records

If you validate your SV calls using orthogonal methods, record the validation results. The 2024 Genes study used optical genome mapping and visual inspection with the Integrative Genomics Viewer to verify SV calls. See the [Genes comparison study](https://pubmed.ncbi.nlm.nih.gov/39062704) for the validation methodology. Record which variants were confirmed, which were rejected, and the evidence used for each decision.

### Troubleshooting Method for SV Calling Failures

When SV calling results are unexpected or inconsistent, use a systematic troubleshooting method to identify the cause. The following steps address the most common failure patterns documented in published benchmarks.

#### Check Read Order Sensitivity

The 2024 PeerJ study found that the order of input data affected the SVs predicted by each caller. The study used PacBio data from 15 Caenorhabditis elegans strains and four Arabidopsis thaliana ecotypes. The pbsv caller was highly sensitive to the order of input data, especially at the highest depths where over 70 percent of SV calls generated from pairs of differently ordered FASTQ files were in disagreement. The SAMtools alignment sorting algorithm was also identified as a source of variability following read order randomization. See the [PeerJ read order study](https://pubmed.ncbi.nlm.nih.gov/38500526) for the full analysis.

If your results are unstable across runs, check whether read order is affecting your analysis. Run the same analysis on the original and permuted FASTQ files and compare the resulting VCF files. If the calls differ substantially, consider standardizing your read order or using a more robust caller.

#### Verify Alignment Quality

Alignment errors propagate into SV calling errors. Check your alignment metrics, including mapping quality distributions, insert size distributions, and coverage uniformity. The 2025 inversion benchmark found that mapper choice has a substantial impact on inversion detection in repetitive regions. See the [inversion benchmark preprint](https://pubmed.ncbi.nlm.nih.gov/41377520) for the mapper comparison. If your alignment quality is poor, consider realigning with different parameters or a different aligner.

#### Compare Against Orthogonal Data

If you have access to orthogonal validation data, such as optical genome mapping or long-read sequencing, compare your SV calls against this reference. The 2024 Genes study found that Bionano OGM calls have high precision with a positive predictive value of 95 percent. See the [Genes comparison study](https://pubmed.ncbi.nlm.nih.gov/39062704) for the validation results. If your calls do not match the orthogonal data, the issue may be in your alignment, caller parameters, or filtering criteria.

#### Examine Variant Type Distributions

Compare your variant type distribution against published expectations. The BMC Genomics benchmark found that several callers exhibited better calling performance for deletions than for duplications, inversions, and insertions. See the [BMC Genomics benchmark](https://pubmed.ncbi.nlm.nih.gov/38549092) for the variant type performance comparison. If your deletion calls are much lower than expected, your sensitivity may be compromised. If your duplication calls are unexpectedly high, you may be seeing false positives from read-depth artifacts.

#### Review Filtering Criteria

Overly aggressive filtering can remove true variants, while insufficient filtering can leave false positives. Review your filtering criteria against the quality score distributions in your output. The [Bioconductor project](https://bioconductor.org/) provides packages for examining and refining SV filtering criteria. The [EMBL-EBI training resources](https://www.ebi.ac.uk/training) offer courses on variant filtering and quality control.

### Professional Escalation Criteria

Some SV calling problems require escalation to a specialist or a different analytical approach. Escalate your analysis when you observe any of the following conditions:

- Results that contradict established findings for well-characterized samples
- Results that are inconsistent across replicate analyses despite identical parameters
- Results that suggest a variant with known clinical significance but cannot be validated
- Results that are affected by known technical artifacts that you cannot resolve
- Results that require interpretation beyond your expertise, particularly in clinical contexts

When escalating, document the problem clearly with your records, including the caller versions, parameters, and validation results. The [NCBI data resources](https://www.ncbi.nlm.nih.gov/) provide access to clinical databases and reference materials that can support interpretation. The [Carpentries lessons](https://carpentries.org/lessons) offer training on responsible data handling and professional communication in research contexts.

### Decision Matrix for Common Project Types

The following decision matrix summarizes the recommended caller choices for common project types based on the published benchmark evidence.

#### Population-Scale Germline Deletion Discovery

For projects that screen large cohorts for germline deletions, Manta is the primary recommendation. The BMC Genomics benchmark found that Manta identified deletion SVs with better performance and efficient computing resources. Manta also showed the highest concordance for deletions when genotypes were verified using a phased long-read assembly dataset. See the [BMC Genomics benchmark](https://pubmed.ncbi.nlm.nih.gov/38549092) for the deletion performance data. Use Delly as a secondary caller if you need to cross-validate findings or if your pipeline already includes Delly.

#### Clinical Germline Analysis

For clinical germline analysis where precision matters more than raw sensitivity, use Manta with careful filtering. The high genotype concordance for deletions and insertions supports downstream interpretation. Validate any clinically significant findings using orthogonal methods. The [NCBI data resources](https://www.ncbi.nlm.nih.gov/) provide access to clinical databases for variant interpretation.

#### Somatic Analysis with Tumor-Normal Pairs

For somatic SV calling, use Manta in its somatic mode. Manta provides a somatic mode that takes tumor and normal samples as input. The efficient computing resource usage is particularly valuable when processing multiple tumor-normal pairs. See the [Galaxy training resources](https://training.galaxyproject.org/) for tutorials on somatic variant calling workflows.

#### Insertion-Focused Studies

For projects focused on insertions, Manta is the best short-read option among the three callers. Manta and MELT demonstrated relatively good precision for insertions in the BMC Genomics benchmark. See the [BMC Genomics benchmark](https://pubmed.ncbi.nlm.nih.gov/38549092) for the insertion precision data. However, the 2024 Genes study found that short-read sensitivity for insertions is poor at 22 percent. See the [Genes comparison study](https://pubmed.ncbi.nlm.nih.gov/39062704) for the sensitivity data. If insertions are your primary focus, seriously consider long-read sequencing or optical genome mapping.

#### Multi-Caller Integration Projects

For projects that require maximum sensitivity, combine Manta with Lumpy or Delly. The 2019 Genome Research study found that additional confidence or sensitivity can be obtained by a combination of multiple variant callers. See the [Genome Research long-read study](https://pubmed.ncbi.nlm.nih.gov/31186302) for the combination workflow. Use a consensus approach that requires support from at least two callers to balance sensitivity and precision.

### Implementation Checklist

Use the following checklist when implementing your SV calling workflow:

1. Define your primary variant types and sensitivity priorities
2. Assess your data characteristics, including sequencing platform and depth
3. Evaluate your computational constraints and estimate resource requirements
4. Run a pilot comparison on a representative subset of samples
5. Apply a weighted scoring system to select your primary caller
6. Plan for caller combination if maximum sensitivity is required
7. Establish a record system for sample, alignment, caller, output, and validation data
8. Implement a troubleshooting method for common failure patterns
9. Define professional escalation criteria for unexpected results
10. Document your selection rationale and analysis parameters for reproducibility

The [nf-core documentation](https://nf-co.re/docs) provides guidance on implementing reproducible pipelines that support this checklist. The [Carpentries lessons](https://carpentries.org/lessons) offer training on the computing skills needed to implement and document your workflow effectively.

## Frequently Asked Questions

### What is the best SV caller for germline deletion detection?

Manta demonstrated the best performance for deletion SVs in the BMC Genomics benchmark, with efficient computing resources and the highest genotype concordance for deletions and insertions. Use Manta for germline deletion discovery in short-read whole-genome sequencing data. See the [BMC Genomics benchmark](https://pubmed.ncbi.nlm.nih.gov/38549092) for the full comparison.

### How do Manta, Lumpy, and Delly compare for insertion detection?

Manta and MELT demonstrated relatively good precision for insertions in the BMC Genomics benchmark. Lumpy and Delly showed lower performance for insertions. Short-read data has poor sensitivity for insertions overall, with the 2024 Genes study finding 22 percent sensitivity for insertions in Illumina data. Consider long-read sequencing if insertions are a primary focus.

### Should I use long-read sequencing instead of short-read for SV calling?

Long-read sequencing can improve sensitivity for certain SV types. The 2024 Genes study found that Sniffles2 with ONT data outperformed Illumina for most SV types. The 2019 Genome Research study found that Sniffles after NGMLR or minimap2 alignment provided the most accurate results for long-read data. However, long-read sequencing is more expensive and requires different analysis workflows.

### How does read order affect SV calling results?

The 2024 PeerJ study found that the order of input data affected the SVs predicted by each caller. The pbsv caller was highly sensitive to read order, and the SAMtools alignment sorting algorithm was a source of variability. These findings have implications for the replication of SV studies. See the [PeerJ read order study](https://pubmed.ncbi.nlm.nih.gov/38500526) for details.

### What is the role of optical genome mapping in SV benchmarking?

Optical genome mapping provides a high-precision benchmark for SV calling. The 2024 Genes study found that Bionano OGM calls have high precision with a positive predictive value of 95 percent. OGM can be used to establish truth datasets for evaluating SV callers.

### How should I combine multiple SV callers for better results?

The 2019 Genome Research study found that additional confidence or sensitivity can be obtained by a combination of multiple variant callers. Combining Manta, Lumpy, and Delly can improve sensitivity but increases computational cost. The study described a scalable workflow for identification, annotation, and characterization of tens of thousands of structural variants.

### What computational resources do Manta, Lumpy, and Delly require?

The BMC Genomics benchmark evaluated running time and memory usage for 11 SV callers. Manta demonstrated efficient computing resources. Delly required substantial computational resources for massive whole-genome sequencing datasets. Lumpy requires moderate memory usage but needs careful parameter tuning.

### How do I benchmark SV callers for my own data?

Benchmark SV callers against truth sets when available. The 2024 Genes study used optical genome mapping as a benchmark. The 2025 inversion benchmark provides a multi-genome benchmark for simple and complex inversions. The [NCBI data resources](https://www.ncbi.nlm.nih.gov/) provide access to reference materials and benchmark datasets.

## Related Bioinformatics Guides

- [Detecting Structural Variants with Long-Read Sequencing: Methods and Considerations](/knowledge/bioinformatics/detecting-structural-variants-with-long-read-sequencing-methods-and-considerations)
- [Variant Calling GATK: Structural Analysis and Computational Methodologies in Bioinformatics](/knowledge/bioinformatics/variant-calling-gatk)
- [Data Stewardship vs Data Governance: What's the Difference?](/knowledge/bioinformatics/data-stewardship-vs-data-governance-what-s-the-difference)
- [RNA-Seq vs qPCR: Validation and Comparison](/knowledge/bioinformatics/rna-seq-vs-qpcr-validation-and-comparison)
- [Metabolomics Data Analysis in R: A Practical Workflow](/knowledge/bioinformatics/metabolomics-data-analysis-in-r-a-practical-workflow)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)


## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [Comparison of structural variant callers for massive whole-genome sequence data.](https://pubmed.ncbi.nlm.nih.gov/38549092). BMC genomics, 2024.
- [A Comparison of Structural Variant Calling from Short-Read and Nanopore-Based Whole-Genome Sequencing Using Optical Genome Mapping as a Benchmark.](https://pubmed.ncbi.nlm.nih.gov/39062704). Genes, 2024.
- [Structural variants identified by Oxford Nanopore PromethION sequencing of the human genome.](https://pubmed.ncbi.nlm.nih.gov/31186302). Genome research, 2019.
- [The impact of FASTQ and alignment read order on structural variant calling from long-read sequencing data.](https://pubmed.ncbi.nlm.nih.gov/38500526). PeerJ, 2024.
- [Benchmark for simple and complex genome inversions.](https://pubmed.ncbi.nlm.nih.gov/41377520). bioRxiv : the preprint server for biology, 2025.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.