# How to Call Structural Variants from Whole-Genome Sequencing Data: A Step-by-Step Workflow with Delly and Lumpy


## Key Takeaways

- Structural variant (SV) calling from whole-genome sequencing (WGS) data necessitates a reproducible pipeline, utilizing tools like Delly and Lumpy, to convert aligned reads (BAM files) into filtered, annotated variant call format (VCF) files for downstream interpretation.
- Delly detects deletions, duplications, inversions, and translocations primarily through discordant paired-end reads and split-read analysis, while Lumpy employs a probabilistic framework integrating discordant pairs, split reads, and read depth for breakpoint detection.
- Accurate SV calling is contingent on high-quality read alignment, proper read group handling, and the selection of an appropriate reference genome, with quality control metrics such as mapping rate, duplicate rate, and insert size distribution being critical checkpoints.
- Combining calls from multiple SV callers, such as Delly and Lumpy, through intersection, union, or minimum support strategies, enhances both sensitivity and precision, with breakpoint reconciliation and reciprocal overlap being key for merging disparate calls.
- Filtering strategies for SV calls involve quality metrics (QUAL, genotype quality), supporting read counts, and population frequency, alongside methods to remove false positives by excluding problematic genomic regions and requiring multi-signal support.
- The clinical and biological significance of SVs is substantial, impacting rare disease diagnostics and contributing to genomic diversity, though short-read WGS has inherent limitations in sensitivity and breakpoint resolution, particularly in complex genomic regions.

---

Structural variant (SV) calling from whole-genome sequencing (WGS) data requires a reproducible pipeline that converts aligned reads into a filtered variant call format (VCF) file suitable for downstream interpretation. This workflow covers the complete process from aligned BAM files to annotated VCF using Delly for deletions, duplications, inversions, and translocations, and Lumpy for breakpoint detection across multiple signal types. The practical outcome is a working SV calling pipeline with concrete command examples, parameter recommendations, quality control checkpoints, and interpretation limits that you can adapt to your computing environment and research question.

## Scope and Reader Context

This article serves biology students, researchers, laboratory professionals, and life-science practitioners who need a practical, reproducible method for detecting SVs from WGS data. The workflow assumes you have aligned sequencing reads in BAM format and basic familiarity with command-line operations. If you need to build foundational command-line skills, The Carpentries offers structured lessons on shell, Git, and data handling that support reproducible bioinformatics practice. For those new to genomic data resources, the NCBI Data Resources provide official documentation on sequence databases and search systems that help you locate reference genomes and public WGS datasets for testing your pipeline.

The workflow described here uses Delly and Lumpy as the primary SV callers, but the principles of quality control, filtering, and interpretation apply broadly across SV detection tools. The Galaxy Training Network provides accessible tutorials on variant analysis workflows that complement the command-line approach presented here, and the nf-core documentation describes community standards for pipeline configuration and reproducibility that you can apply when scaling this workflow to production use.

## Structural Variants and Their Clinical and Biological Significance

Structural variants are genomic alterations typically larger than 50 base pairs that include deletions, duplications, insertions, inversions, and translocations. These variants contribute substantially to genomic diversity and disease predisposition across species. Their accurate characterization has remained challenging because of their size and complexity, which often exceed the detection limits of short-read sequencing approaches designed primarily for single nucleotide variants and small indels.

The clinical relevance of SVs is well documented. In rare disease diagnostics, more than half of families with suspected rare monogenic diseases remain unsolved after whole-genome analysis by short-read sequencing. Long-read sequencing can help bridge this diagnostic gap by capturing variants inaccessible to short-read approaches, facilitating long-range mapping and phasing, and providing haplotype-resolved methylation profiling. In a rare-disease cohort of 98 samples from 41 families sequenced with nanopore technology at approximately 36-fold average coverage, long-read sequencing detected additional rare, functionally annotated variants including SVs and tandem repeats, and established diagnostic variants in 11 probands with diverse underlying genetic causes including large-scale SVs. This evidence demonstrates that SV detection directly impacts diagnostic yield for rare monogenic diseases.

Genome sequencing as a generic diagnostic strategy has been evaluated against existing workflows for rare disease diagnosis. In a study of 1000 cases with 1271 known clinically relevant variants, short-read genome sequencing at 37-fold mean coverage detected 95% of variants overall, with detection rates of 96% for small variants, 93% for large variants including copy number variants and short tandem repeats, and 87% for other variants including structural variants and aneuploidies. The true positive rates varied between workflows from 79% to 100%, with 7 of 10 existing workflows being replaceable by genome sequencing. These findings indicate that while SV detection from short-read WGS is feasible, it does not achieve complete sensitivity, and you should interpret negative SV calls with appropriate caution.

## Core Principles of Structural Variant Calling

### Read Alignment and Its Role in SV Detection

SV calling begins with read alignment, and the quality of your alignment directly determines the quality of your SV calls. Delly and Lumpy both operate on BAM files that have been aligned to a reference genome. The alignment process places each sequencing read at its most likely genomic position, and the patterns of discordant read pairs, split reads, and read depth variations provide the evidence from which SV callers infer structural changes.

For short-read WGS data, paired-end sequencing produces reads from both ends of a DNA fragment. When a structural variant is present, the alignment patterns deviate from expectation in characteristic ways. A deletion causes reads spanning the breakpoint to align with an unexpected insert size, or one read of a pair may fail to align in the expected orientation. A duplication produces increased read depth in the duplicated region. An inversion causes read pairs to align in the wrong orientation relative to each other. Translocations produce read pairs that map to different chromosomes.

Delly uses a combination of paired-end and split-read analysis to detect deletions, duplications, inversions, and translocations. Lumpy takes a probabilistic approach that integrates multiple signals including discordant read pairs, split reads, and read depth to identify breakpoints. Both tools require properly formatted BAM files with read groups and sorted coordinates.

### Signal Types Used by SV Callers

The three primary signal types used by SV callers are discordant read pairs, split reads, and read depth. Discordant read pairs are pairs whose alignment properties deviate from the expected insert size distribution or orientation. Split reads are individual reads that align to two distinct genomic locations, indicating a breakpoint within the read. Read depth refers to the number of reads mapping to a genomic interval, which increases for duplications and decreases for deletions.

Delly primarily relies on discordant paired-end reads to identify candidate SV regions and then uses split-read analysis to refine breakpoints at base-pair resolution. Lumpy uses a probabilistic framework that combines evidence from discordant pairs, split reads, and read depth to estimate the most likely breakpoint locations. The complementary strengths of these two approaches make them a useful combination for SV detection.

### Reference Genomes and Their Limitations

The choice of reference genome affects SV calling in several ways. The reference genome provides the coordinate system against which all variants are defined. Regions of the reference that are repetitive, GC-rich, or otherwise difficult to sequence and align will have reduced mappability, leading to lower sensitivity for SV detection in those regions. The NCBI Data Resources provide official reference genome assemblies and documentation on their structure and limitations.

When selecting a reference genome, consider the species, the assembly version, and the availability of associated resources such as gene annotation tracks and population frequency databases. Different assembly versions of the same species can produce different SV calls because of improvements in the reference sequence itself. Document the reference genome version in your analysis records and report it alongside your SV calls.

## Preparing Your Data for SV Calling

### Input Requirements and File Formats

The input to the SV calling workflow is a coordinate-sorted BAM file with an associated index file. The BAM file must contain read group information that identifies the sample and library, and the reads must be aligned to the same reference genome version that you will use for SV calling. The index file, typically named with a .bai suffix, allows tools to access specific genomic regions efficiently.

Before running SV callers, verify that your BAM file meets these requirements. Use `samtools quickcheck` to confirm the file is not truncated, `samtools view -H` to inspect the header and confirm the reference genome, and `samtools flagstat` to check mapping statistics. The EMBL-EBI Training provides practical guidance on working with sequencing data files and quality assessment that supports these verification steps.

### Quality Control of Aligned Reads

Quality control of aligned reads is a critical step that prevents downstream errors. Check the following metrics before proceeding to SV calling:

- Mapping rate: the proportion of reads that aligned to the reference genome. Low mapping rates may indicate contamination, adapter contamination, or reference mismatch.
- Duplicate rate: the proportion of reads that are PCR or optical duplicates. High duplicate rates reduce effective coverage and can distort read depth signals used by SV callers.
- Insert size distribution: the distribution of fragment sizes for properly paired reads. SV callers use this distribution to identify discordant pairs, so an accurate estimate is essential.
- Coverage uniformity: the consistency of read depth across the genome. Extreme coverage fluctuations can create false SV signals.

The Galaxy Training Network provides tutorials on quality control workflows that demonstrate these metrics in practice. Document these quality metrics in your analysis records and compare them against established thresholds for your sequencing platform and library preparation method.

### Read Group Handling and Sample Identity

Read groups in the BAM header identify the sample, library, and sequencing run for each read. SV callers use read group information to separate samples in multi-sample analyses and to model sequencing platform-specific error profiles. Ensure that each sample has a unique read group identifier and that the sample name is consistent across all files in your analysis.

Sample identity errors are a common source of analysis failure. Verify that the sample name in the BAM header matches the sample name in your analysis records and that no sample mixing occurred during library preparation or sequencing. The NCBI Data Resources provide guidance on sequence data submission standards that include sample naming conventions.

## Setting Up the SV Calling Environment

### Software Installation and Dependency Management

Delly and Lumpy require specific software dependencies that must be installed before the workflow can run. Delly requires a C++ compiler, the Boost libraries, and zlib. Lumpy requires Python, the pysam library, and the lumpy-sv package that includes the core breakpoint detection algorithm.

The Bioconductor project provides official documentation on reproducible genomic-analysis workflows and package installation that can help you manage dependencies in R-based environments. For command-line tools, consider using a package manager such as Conda or a container platform such as Docker or Singularity to create reproducible environments. The nf-core documentation describes community standards for containerized pipeline execution that you can adapt for your SV calling workflow.

### Reference Genome Indexing

Before running SV callers, you must index the reference genome. Delly requires a dictionary file for the reference, created with `samtools dict`, and Lumpy requires a fasta index, created with `samtools faidx`. These index files allow the tools to access reference sequences efficiently during breakpoint refinement.

The reference genome file itself must be in FASTA format with a consistent naming convention. Chromosome names in the reference must match chromosome names in the BAM file. Mismatches between reference and BAM chromosome names are a common cause of SV calling failure.

### Computing Resources and Runtime Expectations

SV calling from WGS data is computationally intensive. The runtime depends on the genome size, sequencing depth, number of samples, and available computing resources. For a human genome at 30-fold coverage, Delly typically requires several hours of compute time, and Lumpy requires similar resources. The NanoVar protocol for long-read SV detection reports that users can identify and analyze SVs in a typical human dataset with a conventional computational setup in approximately 2 to 5 hours after read mapping, providing a reference point for workflow expectations.

Plan your computing resources based on the size of your dataset. For large cohorts, consider using a high-performance computing cluster or cloud computing environment. The Carpentries lessons on shell and data management provide foundational skills for working with large files and automating repetitive tasks in these environments.

## Running Delly for Structural Variant Calling

### Delly Command Structure and Options

Delly uses a subcommand structure where the first argument specifies the operation. The main operations are `call` for SV detection, `merge` for combining calls across samples, and `filter` for applying quality filters. The basic call command has the following structure:

```
delly call -g reference.fa -o output.bcf sample.bam
```

The `-g` option specifies the reference genome, and the `-o` option specifies the output file. Delly produces output in BCF format, which is a binary version of VCF that requires compression and indexing.

Key parameters that affect Delly performance include:

- `-q`: minimum mapping quality for read inclusion, with a default of 20
- `-s`: minimum insert size, which should be set based on your library preparation
- `-x`: exclude regions from the analysis, useful for excluding known problematic regions
- `-t`: SV type to call, with options for deletions, duplications, inversions, and translocations

### Calling Deletions, Duplications, Inversions, and Translocations

Delly can call four SV types: deletions, duplications, inversions, and translocations. By default, Delly calls all four types, but you can restrict the analysis to specific types using the `-t` option. This is useful when you have a specific hypothesis about the types of SVs present in your samples.

For deletion calling, Delly identifies regions where the insert size distribution is larger than expected, indicating that a deletion has removed sequence between the paired reads. For duplication calling, Delly identifies regions with increased read depth. For inversion calling, Delly identifies read pairs with unexpected orientation. For translocation calling, Delly identifies read pairs that map to different chromosomes or to distant locations on the same chromosome.

### Multi-Sample Calling with Delly

Delly supports multi-sample calling, which improves sensitivity for rare variants and provides population context for interpreting variant frequency. The multi-sample workflow involves three steps:

1. Call SVs in each sample individually to produce per-sample BCF files
2. Merge the per-sample calls using `delly merge` to create a unified site list
3. Re-genotype all samples at the merged sites using `delly call -g reference.fa -v merged.bcf -o genotyped.bcf sample.bam`

The merge step identifies SVs that are shared across samples and creates a non-redundant list of variant sites. The re-genotyping step then determines the genotype of each sample at each site. This approach reduces the computational burden of calling all samples jointly while maintaining the benefits of multi-sample analysis.

### Delly Output and Quality Scores

Delly output in BCF format contains the SV calls with genotype information and quality scores. The quality score, stored in the QUAL field of the VCF, reflects the confidence in the variant call. Higher quality scores indicate greater confidence. The INFO field contains additional annotations including the SV type, the breakpoint confidence interval, and the number of supporting read pairs.

Convert the BCF output to VCF format using `bcftools view` for downstream analysis and visualization. The VCF file can then be filtered based on quality scores, genotype quality, and supporting evidence.

## Running Lumpy for Structural Variant Calling

### Lumpy Workflow Overview

Lumpy uses a multi-step workflow that extracts evidence from the BAM file, identifies candidate breakpoints, and genotypes the resulting variants. The workflow requires several preprocessing steps before the core Lumpy algorithm runs.

The Lumpy workflow consists of the following steps:

1. Extract discordant read pairs using `samtools view`
2. Extract split reads using `lumpy_filter` or an equivalent tool
3. Run the Lumpy breakpoint detection algorithm on the combined evidence
4. Genotype the resulting breakpoints using `svtyper` or a similar tool

### Extracting Discordant Read Pairs and Split Reads

Discordant read pairs are extracted from the BAM file by selecting pairs whose alignment properties deviate from the expected insert size distribution. The `samtools view` command with appropriate filters can extract these reads. The specific filters depend on the insert size distribution of your library, which you should estimate from the BAM file before setting thresholds.

Split reads are extracted using tools that identify reads with soft-clipped alignments, which indicate that part of the read aligns to one location and part aligns to another. The `lumpy_filter` tool provides this functionality and produces a BAM file containing candidate split reads.

### Running the Lumpy Breakpoint Detection Algorithm

The core Lumpy algorithm takes the discordant read pairs and split reads as input and identifies breakpoints where structural variants are present. The algorithm uses a probabilistic model that integrates evidence from multiple signals to estimate the most likely breakpoint locations.

The basic Lumpy command has the following structure:

```
lumpy -mw 4 -tt 0 -x exclude.bed -b bamlist.txt -o output.vcf
```

The `-mw` option sets the maximum window size for breakpoint clustering, the `-tt` option sets the trim threshold for split read alignment, and the `-x` option specifies regions to exclude. The `-b` option specifies a file containing the paths to the discordant and split read BAM files.

### Genotyping Lumpy Calls with svtyper

Lumpy produces breakpoint calls without genotype information. The `svtyper` tool adds genotype information by examining read depth and breakpoint support in the original BAM files. This step is essential for downstream analysis that requires genotype information, such as association studies or family-based analyses.

The svtyper command has the following structure:

```
svtyper -B sample.bam -i lumpy.vcf -o genotyped.vcf
```

The `-B` option specifies the BAM file, and the `-i` option specifies the input VCF file from Lumpy. The output VCF contains genotype calls with genotype quality scores and supporting evidence counts.

## Combining Delly and Lumpy Calls

### Rationale for Multi-Caller Approaches

Different SV callers have complementary strengths and weaknesses. Delly excels at detecting deletions and other SVs with clear paired-end signatures, while Lumpy integrates multiple signal types and can detect SVs that individual signals might miss. Combining calls from multiple tools increases sensitivity and provides a basis for assessing call confidence.

The SV-JIM pipeline demonstrates the value of multi-caller approaches in a systematic way. SV-JIM streamlines the numerous SV calling stages into a single process and evaluates the multiple SV sets produced using the Jaccard index measure to identify those with the highest consistency among the included SV callers. SV-JIM then produces aggregated SV results based on how many callers supported the reported SVs. Case studies on the Homo sapiens genome and two plant genomes identified a significant amount of inter-caller variance that varied by tens of thousands of results on the larger Brassica nigra and Homo sapiens genomes. Aggregating the SV sets helped simplify better retention of the less frequently occurring SV types by requiring a level of minimum support instead of from a specific SV caller combination.

### Intersection and Union Approaches

Two common strategies for combining SV calls are intersection and union. The intersection approach retains only SVs detected by both callers, which increases precision at the cost of sensitivity. The union approach retains SVs detected by either caller, which increases sensitivity at the cost of precision.

The choice between intersection and union depends on your research question. For clinical applications where false positives are costly, the intersection approach may be preferred. For discovery applications where sensitivity is paramount, the union approach may be more appropriate. The SV-JIM approach of requiring minimum support from multiple callers provides a middle ground that balances sensitivity and precision.

### Breakpoint Reconciliation and Reciprocal Overlap

Combining SV calls requires reconciling breakpoints between callers. Two SVs from different callers may represent the same biological variant but have slightly different breakpoint coordinates. The standard approach is to consider two SVs as the same variant if their breakpoints are within a specified distance of each other, typically 100 to 1000 base pairs depending on the application.

For deletions and duplications, reciprocal overlap is a common criterion. Two SVs are considered the same if the proportion of the shorter SV that overlaps the longer SV exceeds a threshold, typically 50% to 80%. For inversions and translocations, breakpoint proximity is the primary criterion because the variant spans may be large.

### Tools for Combining SV Calls

Several tools facilitate the combination of SV calls from multiple callers. The `survivor` tool provides functions for merging, comparing, and filtering SV calls. The `svimmer` tool combines calls from multiple callers using a reciprocal overlap approach. The `SV-JIM` pipeline provides an integrated workflow for multi-caller SV analysis with aggregation based on minimum support.

When using these tools, document the parameters used for breakpoint reconciliation and reciprocal overlap. These parameters directly affect the number of combined calls and should be reported alongside your results.

## Filtering and Quality Control of SV Calls

### Quality Metrics for SV Calls

SV calls come with several quality metrics that you should examine before downstream analysis. The QUAL field in the VCF provides an overall quality score. The genotype quality, stored in the GQ field, reflects confidence in the genotype call. The number of supporting read pairs, stored in the INFO field, indicates the strength of evidence for the variant.

Additional quality metrics include the breakpoint confidence interval, which indicates the uncertainty in the breakpoint coordinates, and the read depth ratio, which compares the observed read depth in the variant region to the expected depth. These metrics help distinguish true SVs from alignment artifacts.

### Filtering Strategies Based on Quality Scores

A common filtering strategy retains SVs with quality scores above a threshold, genotype quality above a threshold, and a minimum number of supporting read pairs. The specific thresholds depend on your sequencing depth and the balance between sensitivity and precision that your application requires.

For germline SV calling from WGS data at 30-fold coverage, typical filters include:

- QUAL greater than 100
- Genotype quality greater than 20
- At least 3 supporting read pairs for deletions and duplications
- At least 2 supporting read pairs for inversions and translocations

These thresholds are starting points that you should adjust based on your data quality and validation results. The Galaxy Training Network provides tutorials on variant filtering that demonstrate these concepts in practice.

### Population Frequency Filtering

Population frequency filtering removes SVs that are common in the general population, which are less likely to be pathogenic or biologically interesting. Population frequency is estimated from large cohorts of sequenced individuals, such as those available through the NCBI Data Resources.

For rare disease studies, a common filter retains SVs with population frequency below 1%. For cancer studies, the threshold may be higher because somatic SVs can be present at low frequency in the population. The appropriate threshold depends on the inheritance model and the expected frequency of disease-associated variants.

### Removing False Positive SVs

False positive SVs arise from several sources, including alignment errors in repetitive regions, GC bias in sequencing, and artifacts from library preparation. Strategies for removing false positives include:

- Excluding SVs in known problematic regions such as segmental duplications and centromeres
- Requiring support from multiple independent signals
- Visualizing candidate SVs in a genome browser to confirm the evidence
- Validating a subset of calls using an orthogonal method such as PCR or long-read sequencing

The SV-JIM case studies identified a potential for inflated precision reporting that can occur during evaluation, highlighting the importance of careful validation when assessing SV calling performance.

## Annotating and Interpreting SV Calls

### Gene and Regulatory Region Annotation

Annotating SV calls with gene and regulatory region information helps prioritize variants for downstream analysis. Annotation tools such as ANNOVAR, SnpEff, and VEP can add gene names, variant consequences, and regulatory annotations to SV calls.

The annotation process maps SV breakpoints to genomic features and determines the potential functional impact. A deletion that removes an exon has a different consequence than a deletion in an intergenic region. An inversion that disrupts a gene promoter has a different consequence than an inversion in a gene desert.

### SV-Derived Neoantigen Prediction

In cancer research, SVs can generate novel chimeric sequences and neopeptides that serve as targets for personalized vaccines and T-cell therapies. The SVNeoPP workflow takes WGS and RNA-seq as inputs, performs SV calling and annotation, and reconstructs altered transcripts and coding sequences in a traceable, isoform-aware manner to generate candidate peptides. Candidates are prescreened by integrating antigen-processing features with HLA binding prediction, and then hierarchically filtered and prioritized based on transcript expression, proteomics evidence, immunogenicity predictions, and sequence similarity to experimentally validated neoantigen databases.

This workflow demonstrates the downstream utility of SV calls beyond basic variant annotation. When your research question involves immunology or cancer biology, consider whether SV-derived neoantigen analysis is relevant to your study.

### Clinical Interpretation of SV Calls

Clinical interpretation of SV calls requires comparison against databases of known pathogenic variants and careful assessment of the evidence for pathogenicity. The American College of Medical Genetics and Genomics provides guidelines for variant interpretation that apply to SVs as well as smaller variants.

For clinical applications, consider the following factors:

- The size and type of the SV
- The genes and regulatory elements affected
- The inheritance pattern of the condition
- The population frequency of the variant
- The consistency of the variant with the clinical presentation

The evidence from long-read sequencing studies demonstrates that SVs contribute to rare disease diagnoses, with large-scale SVs established as diagnostic variants in a subset of probands. This evidence supports the clinical relevance of accurate SV calling.

## Reproducibility and Workflow Management

### Documenting Your Workflow

Reproducibility requires complete documentation of your workflow, including software versions, parameters, reference genome versions, and input file paths. Record the following information for each analysis:

- Software names and versions, including Delly, Lumpy, and all dependencies
- Reference genome version and source
- Alignment software and parameters
- SV calling parameters for each tool
- Filtering thresholds and criteria
- Date of analysis and analyst name

The nf-core documentation describes community standards for pipeline documentation and configuration that you can adapt for your SV calling workflow. The Bioconductor project provides guidance on reproducible genomic-analysis workflows that emphasize documentation and version control.

### Using Workflow Management Systems

Workflow management systems such as Snakemake and Nextflow automate the execution of multi-step pipelines and track the dependencies between steps. These systems provide several benefits for SV calling:

- Automatic restart from failed steps
- Parallel execution of independent steps
- Consistent parameter handling across samples
- Built-in logging and reporting

The SV-JIM pipeline uses the Snakemake workflow management system, and the SVNeoPP workflow is implemented in Snakemake to enable modular extension, checkpoint-based restarts, and end-to-end reproducibility. These examples demonstrate the value of workflow management systems for complex SV analysis pipelines.

### Containerization for Reproducibility

Containerization packages software and its dependencies into a single image that can be run consistently across different computing environments. Docker and Singularity are the most commonly used container platforms in bioinformatics.

Containers address the problem of software version drift, where updates to dependencies change the behavior of your analysis. By freezing the software environment at the time of analysis, containers ensure that the same analysis can be reproduced months or years later. The nf-core documentation provides guidance on container usage in bioinformatics pipelines.

### Version Control for Analysis Code

Version control systems such as Git track changes to analysis code and documentation over time. The Carpentries lessons on Git provide foundational training for using version control in research contexts.

For SV calling workflows, version control is essential for tracking changes to analysis scripts, parameter files, and documentation. Each analysis should be associated with a specific commit of the analysis code, allowing you to reproduce the exact analysis at any point in the future.

## Common Failure Patterns and Troubleshooting

### Low Mapping Quality and Its Causes

Low mapping quality in the BAM file leads to poor SV calling because SV callers rely on read alignment to identify discordant pairs and split reads. Common causes of low mapping quality include:

- Sequencing errors that prevent accurate alignment
- Reference genome mismatches due to sample diversity
- Adapter contamination that causes reads to align incorrectly
- Repetitive regions where reads map to multiple locations

Check mapping quality distributions in your BAM file before running SV callers. If a substantial proportion of reads have mapping quality below 20, investigate the cause before proceeding.

### Insert Size Distribution Problems

SV callers use the insert size distribution to identify discordant read pairs. If the insert size distribution is bimodal or has a long tail, the SV caller may misclassify normal pairs as discordant or miss true discordant pairs.

Estimate the insert size distribution from properly mapped read pairs before running SV callers. The mean and standard deviation of the insert size distribution should be recorded and used to set parameters for discordant pair detection. If the insert size distribution is abnormal, investigate library preparation issues before proceeding.

### Reference Genome Mismatches

Mismatches between the reference genome used for alignment and the reference genome used for SV calling cause systematic errors. SV callers may fail to identify variants in regions where the reference sequences differ, or they may identify false variants at positions of reference mismatch.

Verify that the reference genome version used for alignment matches the reference genome version used for SV calling. The chromosome names and sequences should be identical. If you need to use a different reference version, realign the reads to the new reference before SV calling.

### Memory and Runtime Failures

SV calling from WGS data requires substantial memory and runtime. Delly and Lumpy can fail with out-of-memory errors on large genomes or high-depth samples. Common solutions include:

- Increasing available memory
- Processing chromosomes separately
- Reducing the number of samples in multi-sample analyses
- Using a computing cluster with appropriate resource allocation

Monitor resource usage during SV calling and adjust your approach based on observed requirements. The EMBL-EBI Training provides guidance on working with large genomic datasets that includes resource management considerations.

## Limitations of Short-Read SV Calling

### Sensitivity Limitations in Complex Regions

Short-read SV calling has well-documented limitations in complex genomic regions. Repetitive regions, segmental duplications, and regions with high GC content are difficult to align uniquely, leading to reduced sensitivity for SV detection in these regions. The evidence from long-read sequencing studies demonstrates that long-read approaches capture variants inaccessible to short-read sequencing, including SVs and tandem repeats.

When interpreting negative SV calls from short-read data, consider the mappability of the region. A negative call in a low-mappability region is less informative than a negative call in a uniquely mappable region. Tools that provide mappability tracks can help you assess the confidence of negative calls.

### Breakpoint Resolution Limits

Short-read SV callers typically resolve breakpoints to within a few base pairs, but the resolution depends on the availability of split reads that span the breakpoint. In regions where split reads are absent, breakpoint coordinates may be imprecise, with confidence intervals spanning hundreds of base pairs.

The breakpoint confidence interval, reported in the VCF, provides a measure of this uncertainty. When breakpoint precision is important for your application, consider validating candidate SVs with long-read sequencing or targeted PCR approaches.

### Detection Limits for Small SVs

The lower size limit for SV detection by short-read approaches is approximately 50 base pairs, but sensitivity decreases for smaller SVs. Variants smaller than 50 base pairs are typically classified as indels and detected by small variant callers instead of SV callers.

The genome sequencing study that evaluated detection of clinically relevant variants found that detection rates differed by variant category, with small variants detected in 96% of cases, large variants in 93%, and other variants including structural variants in 87%. These detection rates provide a reference for the sensitivity of short-read genome sequencing for different variant types.

### The Role of Long-Read Sequencing

Long-read sequencing addresses many of the limitations of short-read SV calling. Long reads span repetitive regions and provide contiguous sequence information that enables detection of SVs inaccessible to short-read approaches. The NanoVar protocol provides a comprehensive workflow for SV detection in long-read sequencing data, and the evidence from rare disease studies demonstrates the additional diagnostic yield from long-read approaches.

When short-read SV calling is insufficient for your research question, consider whether long-read sequencing is appropriate. The additional cost and computational requirements may be justified by the improved sensitivity for complex SVs.

## Records and Measurements for Your SV Calling Analysis

### Essential Records to Maintain

Maintain the following records for each SV calling analysis:

- Sample identifiers and their corresponding BAM file paths
- Reference genome version and source
- Software versions for all tools in the workflow
- Parameter values for each tool
- Quality control metrics for input BAM files
- SV call counts before and after filtering
- Filtering thresholds and the rationale for each threshold

These records support reproducibility and provide the basis for troubleshooting when results are unexpected. The Carpentries lessons on data management provide guidance on organizing and documenting research data.

### Measurements to Track Across Samples

Track the following measurements across samples to identify systematic issues:

- Number of SVs called per sample
- Distribution of SV types
- SV size distribution
- Genotype quality distributions
- Transition to transversion ratio for SNVs as a quality proxy

Samples with substantially different SV counts may indicate sample quality issues or batch effects. The SV-JIM case studies identified significant inter-caller variance in SV counts, highlighting the importance of understanding expected variation in SV calling results.

### Comparing Calls Across Methods

When you have SV calls from multiple methods or multiple callers, compare the calls to assess consistency. The Jaccard index, which measures the overlap between two sets of calls, provides a quantitative measure of consistency. The SV-JIM pipeline uses the Jaccard index measure to identify SV sets with the highest consistency among included SV callers.

For individual samples, compare SV calls from Delly and Lumpy to identify calls supported by both tools. These high-confidence calls are more likely to be true positives than calls supported by only one tool.

## Professional Escalation Criteria

### When to Seek Expert Assistance

Seek expert assistance when you encounter any of the following situations:

- SV calling produces an unexpectedly high or low number of calls
- Quality control metrics indicate problems with input data
- You need to interpret SV calls for clinical decision-making
- You are uncertain about the appropriate filtering thresholds for your application
- You need to validate SV calls using orthogonal methods

Bioinformatics core facilities and collaborators with SV calling expertise can provide guidance on troubleshooting and interpretation. The EMBL-EBI Training provides learning pathways that can help you build the skills needed to address these challenges independently.

### When to Consider Alternative Approaches

Consider alternative approaches to short-read SV calling when:

- Your research question requires detection of SVs in complex genomic regions
- Breakpoint precision is critical for your application
- Short-read SV calling produces insufficient numbers of high-confidence calls
- You need to detect SVs that are inaccessible to short-read sequencing

Long-read sequencing approaches, such as those described in the NanoVar protocol, provide an alternative for these applications. The evidence from rare disease studies demonstrates that long-read sequencing can detect additional diagnostic variants in cases where short-read approaches are uninformative.

### When to Validate SV Calls

Validate SV calls using orthogonal methods when:

- The SV is a candidate for clinical reporting
- The SV is central to your research hypothesis
- The SV has borderline quality metrics
- The SV is in a region with known alignment difficulties

Validation methods include PCR amplification across breakpoints, long-read sequencing of the region, and fluorescence in situ hybridization for large SVs. The choice of validation method depends on the SV type and the resources available.

## At a Glance: SV Calling Workflow Summary

| Workflow Step | Primary Tool | Key Input | Key Output | Critical Parameters |
|---|---|---|---|---|
| Read alignment verification | samtools | BAM file | Quality metrics | Mapping rate, duplicate rate, insert size distribution |
| Deletion, duplication, inversion, translocation calling | Delly | Aligned BAM, reference FASTA | BCF file with SV calls | Mapping quality threshold, insert size threshold, SV types |
| Breakpoint detection with multi-signal integration | Lumpy | Discordant pairs, split reads, reference FASTA | VCF file with breakpoint calls | Window size, trim threshold, exclusion regions |
| Genotype refinement | svtyper | Lumpy VCF, original BAM | VCF with genotype calls | Read depth model, breakpoint support threshold |
| Multi-caller combination | SURVIVOR or SV-JIM | Delly VCF, Lumpy VCF | Combined VCF | Reciprocal overlap threshold, minimum caller support |
| Filtering and quality control | bcftools | Combined VCF | Filtered VCF | QUAL threshold, genotype quality threshold, supporting read count |
| Annotation | ANNOVAR, SnpEff, VEP | Filtered VCF, gene annotation | Annotated VCF | Gene annotation version, regulatory annotation tracks |

## Comparison of Delly and Lumpy for SV Detection

| Feature | Delly | Lumpy |
|---|---|---|
| Primary signal types | Paired-end discordance, split reads | Paired-end discordance, split reads, read depth |
| SV types detected | Deletions, duplications, inversions, translocations | Deletions, duplications, inversions, translocations |
| Multi-sample support | Yes, with merge and re-genotype workflow | Limited, primarily single-sample analysis |
| Output format | BCF, convertible to VCF | VCF |
| Genotype calling | Included in the call step | Requires separate genotyping with svtyper |
| Computational requirements | Moderate memory, multi-threaded support | Moderate memory, multi-threaded support |
| Strengths | Robust deletion detection, multi-sample workflow | Probabilistic integration of multiple signals |
| Limitations | May miss SVs with weak paired-end signals | Requires careful preprocessing of discordant and split reads |

## Practical Implementation Steps

### Step 1: Verify Input Data Quality

Before running SV callers, verify the quality of your aligned BAM files. Run `samtools quickcheck` to confirm file integrity, `samtools flagstat` to check mapping statistics, and `samtools view -H` to inspect the header. Record the mapping rate, duplicate rate, and insert size distribution for each sample.

### Step 2: Prepare the Reference Genome

Download the reference genome in FASTA format from the NCBI Data Resources or another official source. Create the required index files using `samtools faidx` for the FASTA index and `samtools dict` for the dictionary file. Verify that chromosome names in the reference match chromosome names in your BAM files.

### Step 3: Run Delly for SV Calling

Run Delly on each sample using the appropriate parameters for your data. Start with default parameters and adjust based on your insert size distribution and coverage. Record the command and parameters for each sample.

### Step 4: Run Lumpy for SV Calling

Extract discordant read pairs and split reads from each BAM file, then run the Lumpy breakpoint detection algorithm. Run svtyper to add genotype information to the Lumpy calls. Record the commands and parameters for each step.

### Step 5: Combine Calls from Multiple Callers

Use SURVIVOR or SV-JIM to combine the Delly and Lumpy calls. Choose the reciprocal overlap threshold and minimum caller support based on your application. Record the parameters used for combination.

### Step 6: Filter and Annotate the Combined Calls

Apply quality filters to the combined calls, then annotate with gene and regulatory region information. Record the filtering thresholds and the number of calls retained at each stage.

### Step 7: Document and Archive the Analysis

Document the complete workflow including software versions, parameters, and input file paths. Archive the analysis code and configuration files using version control. Store the final VCF files and supporting documentation in a secure location.

## Common Failure Patterns and Their Solutions

### Failure Pattern: Delly Produces No Calls

If Delly produces no SV calls, check the following:

- Verify that the BAM file contains properly paired reads with read groups
- Check that the insert size distribution is reasonable and that the insert size threshold is appropriate
- Confirm that the reference genome version matches the alignment reference
- Check that the sample name in the BAM header matches the expected sample name

### Failure Pattern: Lumpy Produces Excessive Calls

If Lumpy produces an excessive number of calls, check the following:

- Verify that discordant read pair extraction used appropriate insert size thresholds
- Check that split read extraction used appropriate soft-clip thresholds
- Confirm that exclusion regions are specified for known problematic regions
- Examine the quality score distribution to identify low-confidence calls

### Failure Pattern: Delly and Lumpy Calls Show Poor Concordance

If Delly and Lumpy calls show poor concordance, check the following:

- Verify that both tools used the same reference genome version
- Check that the breakpoint reconciliation parameters are appropriate
- Examine the SV type distribution for each tool to identify systematic differences
- Consider whether one tool is detecting SVs that the other misses due to methodological differences

### Failure Pattern: Filtering Removes Too Many Calls

If filtering removes too many calls, check the following:

- Verify that the quality score distributions are appropriate for your sequencing depth
- Check that the supporting read count thresholds are appropriate for your coverage
- Examine whether specific SV types are disproportionately affected by filtering
- Consider whether the population frequency threshold is too stringent for your application

## Safety and Ethical Considerations

### Data Privacy and Security

WGS data contains sensitive genetic information that requires careful handling. Follow institutional and regulatory requirements for data storage, access control, and sharing. The NCBI Data Resources provide guidance on data submission and access standards that include privacy considerations.

When working with human WGS data, ensure that appropriate consent and ethical approval are in place for your research. De-identify samples before data sharing and use secure data transfer methods.

### Responsible Interpretation of SV Calls

SV calls from WGS data have limitations that should be communicated to collaborators and research participants. Negative SV calls do not exclude the presence of SVs, particularly in complex genomic regions. Positive SV calls require validation before clinical action.

The evidence from genome sequencing studies demonstrates that detection rates vary by variant category, with structural variants detected at lower rates than small variants. Communicate these limitations when reporting SV calling results.

### Reproducibility as an Ethical Obligation

Reproducible analysis is an ethical obligation in research because it enables verification and builds trust in scientific findings. Document your SV calling workflow completely and share analysis code and parameters when publishing results. The nf-core documentation and Bioconductor project provide guidance on reproducible analysis practices.

## Frequently Asked Questions

### What is the minimum sequencing depth required for SV calling with Delly and Lumpy?

The minimum sequencing depth depends on the SV type and size, the genome complexity, and the balance between sensitivity and precision that your application requires. For human WGS data, 30-fold coverage is a common standard that supports detection of germline SVs with reasonable sensitivity. Lower coverage reduces sensitivity, particularly for smaller SVs and SVs in complex regions. Higher coverage increases sensitivity but also increases computational requirements and may introduce artifacts from sequencing errors. The genome sequencing study that evaluated detection of clinically relevant variants used 37-fold mean coverage and detected 87% of structural variants, providing a reference point for expected sensitivity at this coverage level.

### How do I choose between Delly and Lumpy for my SV calling analysis?

The choice between Delly and Lumpy depends on your research question and data characteristics. Delly provides a robust multi-sample workflow with integrated genotyping, making it suitable for cohort studies and germline variant analysis. Lumpy integrates multiple signal types including read depth, which may improve detection of SVs with weak paired-end signals. For most applications, running both tools and combining the results provides the best balance of sensitivity and precision. The SV-JIM pipeline demonstrates the value of multi-caller approaches and provides a framework for aggregating results based on minimum caller support.

### What is the difference between germline and somatic SV calling?

Germline SV calling detects SVs present in the germline genome, which are inherited or arise de novo in the germline. Somatic SV calling detects SVs present in somatic tissues, such as tumors, which arise during development or disease progression. Germline SV calling typically uses a single sample or family samples and expects heterozygous or homozygous variants. Somatic SV calling typically compares tumor and normal samples from the same individual and expects variants present only in the tumor. The bioinformatics workflows differ in the calling strategy, filtering criteria, and interpretation approach.

### How do I filter SV calls to reduce false positives?

Filtering SV calls to reduce false positives involves multiple criteria. Start with quality score thresholds, retaining calls with QUAL above a threshold appropriate for your sequencing depth. Apply genotype quality thresholds to ensure confident genotype calls. Require a minimum number of supporting read pairs, with higher thresholds for larger SVs. Exclude SVs in known problematic regions such as segmental duplications and centromeres. Consider population frequency filtering to remove common variants. Validate a subset of calls using orthogonal methods to assess the false positive rate and adjust thresholds accordingly.

### What are the limitations of short-read SV calling?

Short-read SV calling has several limitations. Sensitivity is reduced in repetitive regions, segmental duplications, and regions with extreme GC content where reads cannot be uniquely aligned. Breakpoint resolution is limited by the availability of split reads spanning the breakpoint. Small SVs below approximately 50 base pairs are typically not detected by SV callers. The evidence from long-read sequencing studies demonstrates that long-read approaches capture variants inaccessible to short-read sequencing, including SVs and tandem repeats. When these limitations are relevant to your research question, consider whether long-read sequencing is appropriate.

### How do I validate SV calls from Delly and Lumpy?

Validation of SV calls uses orthogonal methods that provide independent evidence for the variant. PCR amplification across the breakpoint confirms the presence of a deletion or inversion breakpoint. Long-read sequencing of the region provides sequence-level confirmation of the SV structure. Fluorescence in situ hybridization can confirm large SVs at the cytogenetic level. The choice of validation method depends on the SV type, the resources available, and the level of evidence required for your application. For clinical applications, validation is typically required before reporting.

### What is the role of long-read sequencing in SV detection?

Long-read sequencing addresses many limitations of short-read SV calling by providing reads that span repetitive regions and provide contiguous sequence information. The NanoVar protocol provides a comprehensive workflow for SV detection in long-read sequencing data, and the evidence from rare disease studies demonstrates the additional diagnostic yield from long-read approaches. Long-read sequencing detected additional rare, functionally annotated variants including SVs and tandem repeats, and established diagnostic variants in a subset of probands with rare monogenic diseases. When short-read SV calling is insufficient for your research question, long-read sequencing may provide the additional sensitivity needed.

### How do I report SV calling results in a publication?

Reporting SV calling results requires complete documentation of the workflow and parameters. Report the software versions for all tools, the reference genome version, the alignment and SV calling parameters, and the filtering thresholds. Report the number of SVs called before and after filtering, the distribution of SV types and sizes, and the validation results for a subset of calls. The nf-core documentation and Bioconductor project provide guidance on reproducible analysis reporting. Include the analysis code and configuration files as supplementary materials to enable others to reproduce your analysis.

## Related Bioinformatics Guides

- [Detecting Structural Variants with Long-Read Sequencing: Methods and Considerations](/knowledge/bioinformatics/detecting-structural-variants-with-long-read-sequencing-methods-and-considerations)
- [Metabolomics Data Analysis in R: A Practical Workflow](/knowledge/bioinformatics/metabolomics-data-analysis-in-r-a-practical-workflow)
- [Whole Slide Image Analysis: A Practical Workflow for Pathologists](/knowledge/bioinformatics/whole-slide-image-analysis-a-practical-workflow-for-pathologists)
- [From Raw Reads to Variants: A Diagnostic Blueprint for Next-Generation Sequencing (NGS) Workflows](/knowledge/bioinformatics/ngs-raw-reads-variant-calling-blueprint)
- [De Novo Genome Assembly with Long Reads: A Practical Workflow](/knowledge/bioinformatics/de-novo-genome-assembly-with-long-reads-a-practical-workflow)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)


## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [SV-JIM, detailed pairwise structural variant calling using long-reads and genome assemblies.](https://pubmed.ncbi.nlm.nih.gov/39826659). Methods (San Diego, Calif.), 2025.
- [Advancing long-read nanopore genome assembly and accurate variant calling for rare disease detection.](https://pubmed.ncbi.nlm.nih.gov/39862869). American journal of human genetics, 2025.
- [NanoVar: a comprehensive workflow for structural variant detection to uncover the genome's hidden patterns.](https://pubmed.ncbi.nlm.nih.gov/41034474). Nature protocols, 2026.
- [SVNeoPP: A Workflow for Structural-Variant-Derived Neoantigen Prediction and Prioritization Using Multi-Omics Data.](https://pubmed.ncbi.nlm.nih.gov/41892252). Biology, 2026.
- [Genome sequencing as a generic diagnostic strategy for rare disease.](https://pubmed.ncbi.nlm.nih.gov/38355605). Genome medicine, 2024.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.