A Comprehensive Guide to Structural Variant Calling: Read-Pair, Split-Read, Read-Depth, and Assembly-Based Methods Compared
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Structural variant (SV) calling, crucial for understanding genetic diseases and cancer, relies on four primary paradigms: read-pair, split-read, read-depth, and assembly-based methods, each exploiting distinct signals in sequencing data with varying sensitivities and precisions.
- Read-pair analysis detects SVs via discordant insert sizes or orientations of paired reads, offering computational efficiency for initial screening but limited breakpoint resolution, while split-read analysis achieves base-pair resolution by identifying reads spanning breakpoints, crucial for refining calls.
- Read-depth analysis, by measuring changes in sequencing coverage, is the sole method capable of detecting copy-number variants (CNVs) in exome data, demonstrating high sensitivity comparable to microarrays, but it cannot identify balanced events like inversions or translocations.
- Assembly-based methods de novo assemble reads into contigs, excelling at detecting novel insertions and complex SVs, particularly in repetitive regions, but demand significant computational resources, making them ideal for long-read datasets.
- Short-read sequencing, while cost-effective, exhibits low sensitivity for insertions (22%) compared to deletions (86%), whereas long-read sequencing significantly improves insertion detection (up to 74% with Sniffles2) and can phase variants across long distances, albeit with higher costs and error rates.
- Integrating results from multiple SV callers, as demonstrated by machine learning approaches in the Deciphering Developmental Disorders project, is essential for improving overall dataset quality and mitigating the common failure patterns of low insertion sensitivity and false positives in repetitive regions.
Structural variant (SV) calling is the process of identifying genomic alterations larger than 50 base pairs, including deletions, insertions, duplications, inversions, and translocations. These variants are implicated in rare genetic diseases, cancer progression, and population diversity, yet they remain more difficult to detect reliably than single-nucleotide variants or small indels. This article compares four core detection paradigms: read-pair analysis, split-read analysis, read-depth analysis, and assembly-based methods. Each approach exploits a different signal in sequencing data, and each carries distinct strengths and limitations that affect sensitivity, precision, and computational cost. The practical outcome for researchers is a decision framework for selecting the appropriate SV calling strategy based on sequencing platform, sample type, variant classes of interest, and available compute resources.
The Problem of Structural Variant Detection
Structural variants are a major source of genetic variation in human genomes and other organisms. Unlike single-nucleotide variants, which alter one base, SVs can span thousands to millions of base pairs and disrupt gene structure, regulatory elements, or chromosomal architecture. The clinical relevance of SVs is well established. In rare disease diagnostics, structural variants such as multiexon deletions and duplications are an important cause of disease but are often overlooked in standard exome or genome sequencing analysis. A 2024 study from the Deciphering Developmental Disorders project found that integrating calls from multiple exome-based CNV algorithms using random forest machine learning generated a higher quality dataset than using individual algorithms, and that exome-based CNV calling had higher sensitivity than low-resolution chromosomal microarrays currently in clinical use. This finding underscores that SV detection is also a technical exercise but a diagnostic necessity.
The challenge in SV calling stems from the nature of the variants themselves. Short-read sequencing platforms produce reads of 100 to 150 base pairs, which are far shorter than most structural variants. When a read spans a deletion breakpoint, the alignment software must decide whether the read maps to a novel junction or represents a misalignment. When a read falls entirely within a duplicated region, the alignment may be ambiguous. These ambiguities create systematic errors that vary by variant type and genomic context.
The choice of detection method therefore determines which variants are found and which are missed. A 2024 comparison of short-read and nanopore-based whole-genome sequencing using optical genome mapping as a benchmark found that in Illumina short-read data, sensitivity varied according to variant type, being high for deletions (115 of 134, 86 percent) but poor for insertions (13 of 58, 22 percent). In nanopore long-read data, sensitivity was generally poor using the original Sniffles variant caller (48 percent overall) but improved substantially with Sniffles2, reaching 90 percent for deletions and 74 percent for insertions. These numbers illustrate a central truth: no single method detects all SV types equally well, and the choice of caller and platform materially affects the results.
At a Glance: Four SV Detection Paradigms Compared
The table below summarizes the core principles, strengths, limitations, and typical use cases for each of the four SV detection approaches covered in this article.
| Method | Detection Signal | Strengths | Limitations | Typical Use Case |
|---|---|---|---|---|
| Read-Pair | Discordant insert sizes or orientations of paired reads | Detects deletions, insertions, inversions, and translocations, computationally efficient | Low breakpoint resolution, affected by insert size variability and repetitive regions | Initial screening in germline and somatic short-read datasets |
| Split-Read | Reads that align to two non-contiguous genomic locations | Base-pair breakpoint resolution, validates read-pair calls | Requires reads spanning the breakpoint, limited in repetitive regions | Confirming and refining breakpoints from read-pair or assembly calls |
| Read-Depth | Changes in read coverage along the genome | Detects copy-number variants including deletions and duplications, works with exome data | Cannot detect balanced events like inversions or translocations, resolution limited by bin size | Copy-number variant screening in germline and somatic datasets |
| Assembly-Based | De novo assembly of reads into contigs followed by comparison to reference | Detects novel insertions and complex SVs, captures variants in repetitive regions | High computational cost, sensitive to sequencing errors and coverage depth | Long-read datasets and reference-free discovery of novel insertions |
Core Principles of Structural Variant Detection
Read-Pair Analysis
Read-pair analysis, also called the paired-end mapping approach, relies on the expected distance and orientation between two reads from the same DNA fragment. During library preparation, DNA is fragmented to a target size, typically 300 to 500 base pairs for short-read sequencing. Both ends of each fragment are sequenced, producing read pairs that should align to the reference genome at a distance consistent with the fragment size and in a convergent orientation.
When a structural variant is present, the alignment of read pairs deviates from this expectation. A deletion in the reference genome causes read pairs that span the deletion to align farther apart than expected, because the deleted sequence is absent from the reference. An insertion in the sample genome causes read pairs that flank the insertion to align closer together than expected. An inversion causes read pairs to align in the wrong orientation. A translocation causes read pairs to align to different chromosomes.
The Manta caller, described in a 2016 paper in Bioinformatics, is optimized for rapid germline and somatic analysis and discovers structural variants and indels based on supporting paired and split-read evidence. Manta can analyze NA12878 at 50x genomic coverage in less than 20 minutes, a speed that makes it practical for routine clinical and research sequencing scenarios. The scoring models are optimized for germline analysis of diploid individuals and somatic analysis of tumor-normal sample pairs.
The primary limitation of read-pair analysis is breakpoint resolution. Because the signal is the distance between paired reads, the exact breakpoint can only be localized to within the fragment size, typically several hundred base pairs. This imprecision complicates downstream annotation and clinical interpretation. Read-pair analysis also struggles in repetitive regions, where reads may map to multiple locations and produce spurious discordant signals.
Split-Read Analysis
Split-read analysis identifies structural variants by finding reads that align to two non-contiguous locations in the reference genome. When a read spans a deletion breakpoint, the portion of the read upstream of the deletion aligns to one location and the portion downstream aligns to another location, with a gap corresponding to the deleted sequence. The alignment software must split the read into two or more segments and determine the optimal alignment for each segment.
The strength of split-read analysis is breakpoint resolution. Because the split point within the read corresponds to the exact breakpoint, this method can localize breakpoints to base-pair resolution. This precision is valuable for downstream annotation, primer design for validation, and clinical interpretation.
The limitation of split-read analysis is that it requires reads that actually span the breakpoint. In short-read data, the probability of a read spanning a breakpoint decreases as the variant size increases. For large deletions, most reads will fall entirely within the deleted region or entirely outside it, and only a small fraction will span the junction. This limitation explains why split-read analysis alone has low sensitivity for large variants and is often combined with read-pair analysis.
Manta combines both read-pair and split-read evidence, and the 2016 paper reports that it consistently assembles a higher fraction of its calls to base-pair resolution, allowing for improved downstream annotation and analysis of clinical significance. This integration of signals is a common design pattern in modern SV callers.
Read-Depth Analysis
Read-depth analysis, also called copy-number analysis, detects structural variants by measuring changes in sequencing coverage along the genome. The underlying assumption is that the number of reads mapping to a region is proportional to the copy number of that region in the sample genome. A deletion reduces coverage to approximately half of the expected level in a diploid genome, while a duplication increases coverage to approximately 1.5 times the expected level.
Read-depth analysis is the only method among the four that can detect copy-number variants in exome sequencing data. The 2024 DDD study demonstrated that exome-based CNV calling, when integrating multiple algorithms, achieved the same sensitivity of 89 percent as exon-resolution array comparative genomic hybridization and detected the same number of unique pathogenic CNVs not called by the other approach. This finding is significant because exome sequencing is widely used in clinical diagnostics, and the ability to extract CNV information from existing data avoids the need for additional microarray testing.
The limitations of read-depth analysis are substantial. It cannot detect balanced structural variants such as inversions or translocations, because these events do not change copy number. The resolution is limited by the bin size used to count reads, typically 100 to 1,000 base pairs for whole-genome data and one exon for exome data. Coverage biases from GC content, mappability, and library preparation can create false signals. In the DDD study, the authors noted that integrating calls from multiple ES-based CNV algorithms using random forest machine learning generated a higher quality dataset than using individual algorithms, implying that no single read-depth caller is sufficiently reliable on its own.
Assembly-Based Analysis
Assembly-based analysis constructs contiguous sequences, or contigs, from the sequencing reads without relying on a reference genome. The assembled contigs are then compared to the reference genome to identify structural variants. This approach can detect novel insertions that are absent from the reference, complex variants that involve multiple breakpoints, and variants in repetitive regions that are difficult to map with short reads.
The computational cost of assembly is the primary limitation. De novo assembly requires substantial memory and processing time, and the quality of the assembly depends on sequencing depth, read length, and error rate. Short-read assemblies are fragmented and often fail to span large repetitive regions, limiting their utility for SV detection. Long-read assemblies, particularly those from nanopore or PacBio sequencing, produce longer contigs and can resolve complex variants that are invisible to short-read methods.
A 2025 study in the American Journal of Human Genetics evaluated nanopore long-read sequencing in a rare-disease cohort of 98 samples from 41 families, achieving approximately 36x average coverage and a 32-kilobase read N50 from a single flow cell. The Napu pipeline generated assemblies, phased variants, and methylation calls. Long-read sequencing covered coding exons in approximately 280 genes and about 5 known Mendelian disease-associated genes that were not covered by short-read sequencing. The study established diagnostic variants by long-read sequencing in 11 probands, with diverse underlying genetic causes including de novo and compound heterozygous variants, large-scale SVs, and epigenetic modifications.
The assembly-based approach is particularly valuable for detecting insertions, which are poorly detected by short-read methods. The 2024 comparison study found that short-read sensitivity for insertions was only 22 percent, while nanopore long-read sequencing with Sniffles2 achieved 74 percent sensitivity for insertions. This difference reflects the fundamental limitation of short reads: an insertion that is longer than the read length cannot be fully sequenced, and its presence can only be inferred from the gap between flanking reads.
Sequencing Platforms and Their Impact on SV Calling
Short-Read Sequencing
Short-read sequencing, dominated by Illumina platforms, remains the most widely used technology for genomic analysis. The advantages are well established: high throughput, low cost per base, and low error rates. For SV calling, the short-read platform supports read-pair, split-read, and read-depth approaches, but the read length of 100 to 150 base pairs imposes fundamental limits.
The 2024 comparison study using optical genome mapping as a benchmark provides quantitative evidence of these limits. In the Illumina dataset, sensitivity was high for deletions (86 percent) but poor for insertions (22 percent). This asymmetry is expected: deletions create a detectable gap in read-pair distances, while insertions require reads to span the inserted sequence, which is impossible when the insertion is longer than the read length.
Short-read data also present challenges for repetitive regions. Reads from repetitive elements map to multiple locations, creating ambiguous alignments that generate false SV calls or obscure true variants. The 2025 long-read study noted that long-read sequencing covered coding exons in approximately 280 genes and about 5 known Mendelian disease-associated genes that were not covered by short-read sequencing, indicating that short-read data have systematic blind spots in difficult-to-map regions.
Long-Read Sequencing
Long-read sequencing, including nanopore and PacBio platforms, produces reads of 10 to 100 kilobases or more. These reads can span entire structural variants, allowing direct observation of breakpoints and inserted sequences. The 2025 study achieved a 32-kilobase read N50 from a single flow cell, meaning that half of the sequenced bases were in reads of at least 32 kilobases.
The advantages of long reads for SV calling are substantial. Long reads can span repetitive regions, resolve complex variants, and phase variants across long distances. The 2025 study reported that long-read sequencing completely phased 87 percent of protein-coding genes, a capability that is important for determining whether variants are on the same or different chromosomes.
The limitations of long-read sequencing include higher cost per base, higher error rates (particularly for nanopore data), and lower throughput compared to short-read platforms. The 2024 comparison study found that the original Sniffles variant caller achieved only 48 percent overall sensitivity on nanopore data, although Sniffles2 improved this to 90 percent for deletions and 74 percent for insertions. This improvement demonstrates that the accuracy of long-read SV calling depends heavily on the caller software, beyond the sequencing platform.
Optical Genome Mapping
Optical genome mapping is a distinct technology that images long DNA molecules labeled at specific sequence motifs. It does not sequence the DNA but produces a genome-wide map of label positions that can be compared to a reference map to detect structural variants. The 2024 comparison study used Bionano optical genome mapping to establish a truth dataset, finding that 222 of 234 rare proband SV calls were verified by visualization with the Integrative Genomics Viewer, indicating a positive predictive value of 95 percent.
Optical genome mapping is valuable as a benchmark and validation tool because it provides an independent, sequence-independent view of structural variation. However, it is not a sequencing technology and does not provide base-level sequence information. Its role in SV calling is primarily as a reference standard and a complement to sequencing-based methods.
Practical Workflow for Structural Variant Calling
Step 1: Define the Biological Question and Variant Classes of Interest
The first decision in any SV calling project is to define which variant classes are relevant to the biological or clinical question. A researcher studying rare disease may prioritize deletions, duplications, and insertions that disrupt coding genes. A cancer researcher may need to detect translocations and copy-number changes across the genome. A population geneticist may be interested in all SV classes at the highest possible sensitivity.
This decision determines the choice of detection method. Read-depth analysis is the only method that detects copy-number variants in exome data, making it the default choice for exome-based studies. Read-pair and split-read analysis detect balanced and unbalanced events but have limited sensitivity for insertions. Assembly-based analysis is the only method that reliably detects novel insertions and complex variants.
Step 2: Select the Sequencing Platform and Coverage
The sequencing platform determines which detection methods are feasible. Short-read data support read-pair, split-read, and read-depth analysis but have limited sensitivity for insertions and variants in repetitive regions. Long-read data support assembly-based analysis and direct breakpoint detection but require higher coverage and more compute resources.
Coverage depth affects SV calling sensitivity. Higher coverage increases the number of reads supporting each variant call, improving confidence. However, the relationship between coverage and sensitivity is not linear, and the optimal coverage depends on the variant class and detection method. The 2025 long-read study used approximately 36x average coverage, which was sufficient to detect diagnostic variants in 11 of 41 families.
Step 3: Choose the SV Calling Software
The choice of software is as important as the choice of platform. The 2024 comparison study demonstrated that the original Sniffles caller achieved only 48 percent overall sensitivity on nanopore data, while Sniffles2 achieved 90 percent for deletions and 74 percent for insertions. This difference is larger than the difference between platforms, emphasizing that caller selection is a critical decision.
For short-read data, Manta is a well-established choice for germline and somatic analysis. The 2016 paper reports that Manta is optimized for rapid analysis, calling structural variants, medium-sized indels, and large insertions in less than a tenth of the time that comparable methods require. Manta discovers and scores variants based on supporting paired and split-read evidence, with scoring models optimized for germline analysis of diploid individuals and somatic analysis of tumor-normal sample pairs.
For long-read data, Sniffles2 is a current standard, but other callers are available. The choice should be informed by the specific variant classes of interest and the characteristics of the dataset.
Step 4: Run Quality Control on Input Data
Quality control is essential before running SV callers. Poor-quality reads, adapter contamination, and sequencing errors can generate false SV calls. The Galaxy Training Network provides accessible workflow training and analysis tutorials that cover quality control and reproducible analysis practices. The Carpentries lessons provide foundational computing and data skills that are useful for implementing quality control workflows.
Key quality metrics to assess include read depth, read length distribution, base quality scores, and mapping rates. For paired-end data, the insert size distribution should be checked because read-pair analysis depends on the expected fragment size. Anomalous insert sizes can indicate library preparation problems that will confound SV calling.
Step 5: Run Multiple Callers and Integrate Results
The DDD study demonstrated that integrating calls from multiple exome-based CNV algorithms using random forest machine learning generated a higher quality dataset than using individual algorithms. This finding supports a general principle: no single SV caller is sufficiently reliable on its own, and integrating results from multiple callers improves precision and recall.
Integration can be performed at several levels. Simple intersection of calls from multiple callers reduces false positives but may also reduce sensitivity. Union of calls increases sensitivity but introduces false positives. Machine learning approaches, as used in the DDD study, can learn the characteristics of true and false calls and produce a higher quality integrated dataset.
The nf-core documentation describes community pipeline standards for reproducible workflow configuration and usage. These pipelines often include multiple SV callers and integration steps, providing a practical starting point for researchers who do not want to build their own workflows.
Step 6: Validate and Annotate Candidate Variants
Validation of SV calls is essential, particularly for clinical applications. Validation methods include PCR amplification across breakpoints, Sanger sequencing, optical genome mapping, and visualization of aligned reads in a genome browser. The 2024 comparison study used visualization with the Integrative Genomics Viewer to verify optical genome mapping calls, demonstrating that manual inspection of raw sequence data remains a valuable validation step.
Annotation of validated variants involves determining their genomic context, overlapping genes, regulatory elements, and potential functional impact. The NCBI Data Resources provide official descriptions of databases, search systems, sequence resources, and analysis services that support variant annotation. The EMBL-EBI Training program offers bioinformatics learning pathways and data-resource training that cover variant annotation and interpretation.
Step 7: Document and Report Results
Reproducibility requires documentation of every step in the SV calling workflow, including software versions, parameters, reference genome version, and quality metrics. The nf-core documentation emphasizes community pipeline standards for reproducible workflow configuration and usage. The Galaxy Training Network provides accessible workflow training and analysis tutorials that emphasize reproducibility.
Reporting should include the number of variants detected by class, the validation rate, and the proportion of variants with potential functional impact. For clinical applications, reporting should follow established guidelines for variant interpretation and include the evidence supporting each reported variant.
Options and Trade-offs in SV Calling
Sensitivity versus Precision
The fundamental trade-off in SV calling is between sensitivity and precision. Sensitivity is the proportion of true variants that are detected. Precision is the proportion of detected variants that are true. Increasing sensitivity by relaxing thresholds typically decreases precision, and vice versa.
The 2024 comparison study provides quantitative evidence of this trade-off. Optical genome mapping achieved a positive predictive value of 95 percent, meaning that 95 percent of its calls were true. Short-read sequencing achieved 86 percent sensitivity for deletions but only 22 percent for insertions. Long-read sequencing with Sniffles2 achieved 90 percent sensitivity for deletions and 74 percent for insertions. These numbers illustrate that the optimal balance between sensitivity and precision depends on the variant class and the downstream use of the calls.
For clinical applications, precision is often prioritized to avoid reporting false variants. For discovery research, sensitivity may be prioritized to avoid missing true variants. The choice should be documented and justified in the analysis plan.
Computational Cost
The computational cost of SV calling varies widely by method. Read-pair and read-depth analysis are computationally efficient and can process a whole-genome dataset in minutes to hours. Split-read analysis adds moderate computational cost. Assembly-based analysis is the most expensive, requiring substantial memory and processing time.
Manta was designed to address the computational bottleneck in SV calling. The 2016 paper reports that Manta analyzes NA12878 at 50x genomic coverage in less than 20 minutes, a speed that makes it practical for routine clinical and research sequencing scenarios. This efficiency is achieved through optimized algorithms that focus on regions with evidence of structural variation instead of scanning the entire genome.
The choice of method should consider available compute resources. Researchers with limited compute may need to prioritize read-pair and read-depth methods over assembly-based methods. The Carpentries lessons provide foundational computing and data skills that are useful for managing computational workflows.
Variant Classes Detected
Each detection method is sensitive to different variant classes. Read-depth analysis detects copy-number variants but cannot detect balanced events. Read-pair analysis detects deletions, insertions, inversions, and translocations but has limited breakpoint resolution. Split-read analysis provides breakpoint resolution but requires reads spanning the breakpoint. Assembly-based analysis detects novel insertions and complex variants but is computationally expensive.
The 2024 comparison study found that short-read sensitivity was high for deletions (86 percent) but poor for insertions (22 percent). This asymmetry is a fundamental limitation of short-read sequencing. Long-read sequencing with Sniffles2 achieved 90 percent sensitivity for deletions and 74 percent for insertions, demonstrating that long reads are superior for insertion detection.
Germline versus Somatic Calling
Germline SV calling analyzes constitutional DNA and assumes that variants are present in all cells. Somatic SV calling analyzes tumor DNA and must distinguish somatic variants from germline variants and from sequencing artifacts. The 2016 Manta paper describes scoring models optimized for germline analysis of diploid individuals and somatic analysis of tumor-normal sample pairs.
Somatic SV calling requires matched normal samples to filter germline variants. The tumor-normal comparison also helps to identify somatic variants that are present at low allele fractions due to tumor heterogeneity. The choice of caller and parameters should reflect whether the analysis is germline or somatic.
Exome versus Whole-Genome Sequencing
Exome sequencing targets the protein-coding regions of the genome, which constitute approximately 1 to 2 percent of the genome. SV calling from exome data is limited to read-depth analysis because read-pair and split-read signals are sparse in targeted data. The 2024 DDD study demonstrated that exome-based CNV calling, when integrating multiple algorithms, achieved the same sensitivity of 89 percent as exon-resolution array comparative genomic hybridization.
Whole-genome sequencing provides genome-wide coverage and supports all four SV detection methods. The trade-off is cost: whole-genome sequencing is more expensive than exome sequencing, and the data analysis is more computationally intensive. The choice between exome and whole-genome sequencing should consider the variant classes of interest and the available budget.
Observations and Measurements in SV Calling
Measuring Sensitivity and Precision
Sensitivity and precision are the primary metrics for evaluating SV calling performance. Sensitivity is calculated as the number of true variants detected divided by the total number of true variants. Precision is calculated as the number of true variants detected divided by the total number of variants detected.
Measuring these metrics requires a truth dataset, which can be obtained from orthogonal technologies such as optical genome mapping, from validated variant sets, or from simulated data. The 2024 comparison study used optical genome mapping to establish a truth dataset, finding that 222 of 234 rare proband SV calls were verified, indicating a positive predictive value of 95 percent.
Measuring Breakpoint Resolution
Breakpoint resolution is the precision with which a variant call localizes the exact breakpoint. Read-pair analysis typically achieves breakpoint resolution of several hundred base pairs, limited by fragment size. Split-read analysis achieves base-pair resolution when reads span the breakpoint. Assembly-based analysis achieves base-pair resolution for assembled contigs.
Breakpoint resolution matters for downstream applications. Base-pair resolution is required for PCR primer design, for determining whether a breakpoint disrupts a gene, and for clinical interpretation. The 2016 Manta paper reports that Manta consistently assembles a higher fraction of its calls to base-pair resolution, allowing for improved downstream annotation and analysis of clinical significance.
Measuring Computational Performance
Computational performance is measured by runtime, memory usage, and disk usage. The 2016 Manta paper reports that Manta analyzes NA12878 at 50x genomic coverage in less than 20 minutes, a speed that makes it practical for routine clinical and research sequencing scenarios. This performance is achieved through optimized algorithms that focus on regions with evidence of structural variation.
Computational performance should be measured on the specific hardware and dataset that will be used for the analysis. Performance can vary substantially with coverage depth, read length, and genome complexity. The nf-core documentation provides community pipeline standards for reproducible workflow configuration and usage, which can help standardize performance measurement.
Recording Analysis Parameters
Reproducibility requires recording all analysis parameters, including software versions, reference genome version, alignment parameters, and SV calling thresholds. The nf-core documentation emphasizes community pipeline standards for reproducible workflow configuration and usage. The Galaxy Training Network provides accessible workflow training and analysis tutorials that emphasize reproducibility.
Analysis parameters should be recorded in a structured format that can be shared with collaborators and included in publications. The Carpentries lessons provide foundational computing and data skills, including version control with Git, that support reproducible analysis.
Common Failure Patterns in SV Calling
Failure Pattern 1: Low Sensitivity for Insertions in Short-Read Data
The most common failure pattern in SV calling is low sensitivity for insertions in short-read data. The 2024 comparison study found that short-read sensitivity for insertions was only 22 percent, compared to 86 percent for deletions. This asymmetry is a fundamental limitation of short-read sequencing: an insertion that is longer than the read length cannot be fully sequenced, and its presence can only be inferred from the gap between flanking reads.
Mitigation strategies include using long-read sequencing for insertion detection, using assembly-based methods that can reconstruct inserted sequences, and acknowledging the limitation in the analysis plan. Researchers who need comprehensive insertion detection should consider long-read sequencing despite the higher cost.
Failure Pattern 2: False Positives in Repetitive Regions
Repetitive regions of the genome, including segmental duplications, transposable elements, and tandem repeats, are prone to false SV calls. Reads from repetitive elements map to multiple locations, creating ambiguous alignments that generate spurious discordant signals. The 2025 long-read study noted that long-read sequencing covered coding exons in approximately 280 genes and about 5 known Mendelian disease-associated genes that were not covered by short-read sequencing, indicating that short-read data have systematic blind spots in difficult-to-map regions.
Mitigation strategies include filtering calls in repetitive regions, using mappability tracks to identify regions with ambiguous alignments, and validating calls with orthogonal methods. Long-read sequencing can resolve many repetitive regions that are inaccessible to short reads.
Failure Pattern 3: Batch Effects and Coverage Bias in Read-Depth Analysis
Read-depth analysis is sensitive to coverage biases from GC content, mappability, and library preparation. These biases can create false copy-number signals or obscure true variants. The DDD study addressed this challenge by integrating calls from multiple exome-based CNV algorithms using random forest machine learning, which generated a higher quality dataset than using individual algorithms.
Mitigation strategies include using multiple read-depth callers and integrating results, normalizing coverage for GC bias and other known confounders, and validating calls with orthogonal methods. The choice of bin size also affects read-depth analysis: smaller bins provide higher resolution but more noise, while larger bins provide lower resolution but more stability.
Failure Pattern 4: Inconsistent Results Across Callers
Different SV callers often produce inconsistent results for the same dataset. This inconsistency reflects differences in algorithms, parameters, and underlying assumptions. The DDD study found that integrating calls from multiple algorithms generated a higher quality dataset than using individual algorithms, implying that no single caller is sufficiently reliable on its own.
Mitigation strategies include running multiple callers and integrating results, using machine learning approaches for integration, and validating calls with orthogonal methods. The nf-core documentation provides community pipeline standards that often include multiple SV callers and integration steps.
Failure Pattern 5: Poor Breakpoint Resolution in Read-Pair Analysis
Read-pair analysis localizes breakpoints to within the fragment size, typically several hundred base pairs. This imprecision complicates downstream annotation and clinical interpretation. The 2016 Manta paper reports that Manta consistently assembles a higher fraction of its calls to base-pair resolution, allowing for improved downstream annotation and analysis of clinical significance.
Mitigation strategies include using callers that assemble breakpoint regions, such as Manta, and validating breakpoints with split-read analysis or PCR amplification. For clinical applications, base-pair resolution is often required for variant interpretation.
Limitations of Current SV Calling Methods
Short-Read Limitations
Short-read sequencing has fundamental limitations for SV detection. The read length of 100 to 150 base pairs limits the detection of insertions, which cannot be fully sequenced when longer than the read length. The 2024 comparison study found that short-read sensitivity for insertions was only 22 percent. Short reads also struggle in repetitive regions, where ambiguous alignments generate false calls or obscure true variants.
These limitations are not fully addressable by improved algorithms. The information content of short reads is insufficient to resolve certain variant classes and genomic contexts. Researchers who need comprehensive SV detection should consider long-read sequencing despite the higher cost.
Long-Read Limitations
Long-read sequencing addresses many short-read limitations but introduces new challenges. The error rate of nanopore sequencing is higher than that of short-read sequencing, which can generate false variant calls. The 2024 comparison study found that the original Sniffles variant caller achieved only 48 percent overall sensitivity on nanopore data, although Sniffles2 improved this substantially.
Long-read sequencing also has higher cost per base and lower throughput than short-read sequencing. The 2025 study achieved approximately 36x average coverage from a single flow cell, which is lower than typical short-read coverage of 30 to 50x. The choice between short-read and long-read sequencing should consider the variant classes of interest and the available budget.
Exome Sequencing Limitations
Exome sequencing targets only the protein-coding regions of the genome, which constitute approximately 1 to 2 percent of the genome. SV calling from exome data is limited to read-depth analysis because read-pair and split-read signals are sparse in targeted data. The DDD study demonstrated that exome-based CNV calling achieved the same sensitivity of 89 percent as exon-resolution array comparative genomic hybridization, but this sensitivity applies only to copy-number variants, not to balanced events or variants in non-coding regions.
Researchers who need comprehensive SV detection should consider whole-genome sequencing. The trade-off is cost: whole-genome sequencing is more expensive than exome sequencing, and the data analysis is more computationally intensive.
Validation Limitations
Validation of SV calls is essential but can be challenging. PCR amplification across breakpoints requires base-pair breakpoint resolution, which is not always available. Optical genome mapping provides an independent view of structural variation but is not widely available and does not provide base-level sequence information. The 2024 comparison study used optical genome mapping to establish a truth dataset, but this approach is not practical for routine validation in most laboratories.
Visualization of aligned reads in a genome browser is a valuable validation step but is time-consuming and subjective. The 2024 comparison study used visualization with the Integrative Genomics Viewer to verify optical genome mapping calls, demonstrating that manual inspection remains a valuable validation method.
Quality Controls and Reproducibility
Input Data Quality Control
Quality control of input data is essential before running SV callers. Poor-quality reads, adapter contamination, and sequencing errors can generate false SV calls. The Galaxy Training Network provides accessible workflow training and analysis tutorials that cover quality control and reproducible analysis practices.
Key quality metrics to assess include read depth, read length distribution, base quality scores, and mapping rates. For paired-end data, the insert size distribution should be checked because read-pair analysis depends on the expected fragment size. Anomalous insert sizes can indicate library preparation problems that will confound SV calling.
Alignment Quality Control
The quality of read alignment directly affects SV calling accuracy. Poorly aligned reads generate spurious discordant signals and false split-read calls. Alignment quality should be assessed by mapping rate, mapping quality distribution, and the proportion of reads with discordant alignments.
The choice of aligner and alignment parameters affects SV calling. Different aligners have different strengths and weaknesses for SV detection. The nf-core documentation provides community pipeline standards for reproducible workflow configuration and usage, which can help standardize alignment and SV calling.
Caller Parameter Tuning
SV callers have parameters that control sensitivity and precision. Default parameters are often appropriate for standard datasets, but tuning may be necessary for specific applications. The 2024 comparison study found that the original Sniffles variant caller achieved only 48 percent overall sensitivity on nanopore data, while Sniffles2 achieved 90 percent for deletions and 74 percent for insertions. This difference reflects improvements in the caller algorithm, not parameter tuning.
Parameter tuning should be documented and justified. The Carpentries lessons provide foundational computing and data skills, including version control with Git, that support reproducible analysis.
Reproducibility Standards
Reproducibility requires documentation of every step in the SV calling workflow, including software versions, parameters, reference genome version, and quality metrics. The nf-core documentation emphasizes community pipeline standards for reproducible workflow configuration and usage. The Galaxy Training Network provides accessible workflow training and analysis tutorials that emphasize reproducibility.
Containerization and workflow management systems can improve reproducibility by capturing the software environment. The nf-core documentation provides guidance on using containers and workflow management systems for reproducible analysis.
Safety and Regulatory Context
Clinical Diagnostic Applications
SV calling in clinical diagnostic applications is subject to regulatory oversight and professional guidelines. The 2025 long-read study demonstrated that long-read sequencing can enhance diagnostic yield for rare monogenic diseases, implying utility in future clinical genomics workflows. The 2021 study of an artificial intelligence-based clinical decision support tool found that it ranked over 90 percent of causal genes among the top or second candidate and prioritized for review a median of 3 candidate genes per case.
Clinical SV calling requires validation of all reported variants, documentation of evidence, and adherence to professional guidelines for variant interpretation. The NCBI Data Resources provide official descriptions of databases, search systems, sequence resources, and analysis services that support clinical variant interpretation.
Data Privacy and Security
Genomic data are sensitive personal information that require appropriate safeguards. Researchers and clinicians must comply with applicable regulations for data privacy and security. The NCBI Data Resources provide official descriptions of databases and search systems that include data access controls and security measures.
Data sharing for research purposes should follow established guidelines and obtain appropriate consent. The EMBL-EBI Training program offers bioinformatics learning pathways and data-resource training that cover data sharing and ethical considerations.
Professional Escalation Criteria
Researchers and clinicians should escalate SV calling results to appropriate professionals when certain criteria are met. These criteria include:
- Detection of a variant in a gene with established clinical significance
- Detection of a variant that disrupts a known disease-associated gene
- Detection of a variant with potential reproductive implications
- Inability to validate a variant that has clinical implications
- Discrepancies between SV calling results and other diagnostic tests
The 2021 study of an artificial intelligence-based clinical decision support tool found that it identified causal SVs as the top candidate in 17 of 20 cases with diagnostic SVs and within the top five in 19 of 20 cases. This finding supports the use of decision support tools to prioritize variants for clinical review, but final interpretation should be performed by qualified professionals.
Frequently Asked Questions
What is the difference between read-pair and split-read SV calling?
Read-pair SV calling uses the expected distance and orientation between paired reads to detect structural variants. When a deletion is present, read pairs that span the deletion align farther apart than expected. When an insertion is present, read pairs that flank the insertion align closer together than expected. Split-read SV calling identifies reads that align to two non-contiguous locations in the reference genome, with the split point corresponding to the breakpoint. Read-pair analysis detects variants across a wide size range but has limited breakpoint resolution. Split-read analysis provides base-pair breakpoint resolution but requires reads that actually span the breakpoint. Many modern callers, including Manta, combine both signals to improve sensitivity and breakpoint resolution.
Why is insertion detection so poor in short-read sequencing data?
Insertion detection is poor in short-read data because an insertion that is longer than the read length cannot be fully sequenced. The presence of the insertion can only be inferred from the gap between flanking reads, which provides limited information about the inserted sequence. The 2024 comparison study found that short-read sensitivity for insertions was only 22 percent, compared to 86 percent for deletions. Long-read sequencing can span entire insertions, allowing direct observation of the inserted sequence. The same study found that nanopore long-read sequencing with Sniffles2 achieved 74 percent sensitivity for insertions.
Can structural variants be detected from exome sequencing data?
Yes, copy-number variants can be detected from exome sequencing data using read-depth analysis. The 2024 DDD study demonstrated that exome-based CNV calling, when integrating multiple algorithms, achieved the same sensitivity of 89 percent as exon-resolution array comparative genomic hybridization. However, exome sequencing cannot detect balanced structural variants such as inversions or translocations, and read-pair and split-read signals are sparse in targeted data. Researchers who need comprehensive SV detection should consider whole-genome sequencing.
What is the role of optical genome mapping in structural variant calling?
Optical genome mapping is a technology that images long DNA molecules labeled at specific sequence motifs, producing a genome-wide map of label positions that can be compared to a reference map to detect structural variants. The 2024 comparison study used Bionano optical genome mapping to establish a truth dataset, finding that 222 of 234 rare proband SV calls were verified, indicating a positive predictive value of 95 percent. Optical genome mapping is valuable as a benchmark and validation tool because it provides an independent, sequence-independent view of structural variation. However, it is not a sequencing technology and does not provide base-level sequence information.
How should I choose between short-read and long-read sequencing for SV detection?
The choice between short-read and long-read sequencing depends on the variant classes of interest, the available budget, and the computational resources. Short-read sequencing is less expensive and has lower error rates but has limited sensitivity for insertions and variants in repetitive regions. The 2024 comparison study found that short-read sensitivity was high for deletions (86 percent) but poor for insertions (22 percent). Long-read sequencing can span entire structural variants and resolve repetitive regions but has higher cost per base and higher error rates. The 2025 long-read study demonstrated that long-read sequencing can enhance diagnostic yield for rare monogenic diseases, implying utility in future clinical genomics workflows.
What is the best strategy for improving SV calling accuracy?
The best strategy for improving SV calling accuracy is to run multiple callers and integrate the results. The DDD study demonstrated that integrating calls from multiple exome-based CNV algorithms using random forest machine learning generated a higher quality dataset than using individual algorithms. Integration can be performed at several levels, from simple intersection or union of calls to machine learning approaches that learn the characteristics of true and false calls. Validation of candidate variants with orthogonal methods, such as PCR amplification or optical genome mapping, is also essential for clinical applications.
How do somatic SV calling and germline SV calling differ?
Somatic SV calling analyzes tumor DNA and must distinguish somatic variants from germline variants and from sequencing artifacts. Germline SV calling analyzes constitutional DNA and assumes that variants are present in all cells. The 2016 Manta paper describes scoring models optimized for germline analysis of diploid individuals and somatic analysis of tumor-normal sample pairs. Somatic SV calling requires matched normal samples to filter germline variants and to identify somatic variants that are present at low allele fractions due to tumor heterogeneity.
What quality metrics should I report for SV calling results?
Quality metrics for SV calling results should include the number of variants detected by class, the validation rate, the proportion of variants with potential functional impact, and the computational performance of the analysis. Sensitivity and precision should be reported when a truth dataset is available. Breakpoint resolution should be reported for variants that will be used for downstream applications. All analysis parameters, including software versions, reference genome version, and SV calling thresholds, should be documented for reproducibility.
Related Bioinformatics Guides
- Detecting Structural Variants with Long-Read Sequencing: Methods and Considerations
- Metagenome Co-Assembly: Strategies for Multi-Sample Data
- Long-Read Metagenome Assembly: Overcoming Challenges with Nanopore and PacBio Data
- Evaluating Metagenomic Assembly Tools: A Benchmarking Framework for Short-Read and Long-Read Data
- Metagenomics Assembly: Strategies for Reconstructing Microbial Genomes
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- A Comparison of Structural Variant Calling from Short-Read and Nanopore-Based Whole-Genome Sequencing Using Optical Genome Mapping as a Benchmark.. Genes, 2024.
- Manta: rapid detection of structural variants and indels for germline and cancer sequencing applications.. Bioinformatics (Oxford, England), 2016.
- Artificial intelligence enables comprehensive genome interpretation and nomination of candidate diagnoses for rare genetic diseases.. Genome medicine, 2021.
- Detection and characterization of copy-number variants from exome sequencing in the DDD study.. Genetics in medicine open, 2024.
- Advancing long-read nanopore genome assembly and accurate variant calling for rare disease detection.. American journal of human genetics, 2025.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.