Resolving Structural Variant Breakpoints: How to Use Split-Read and Assembly-Based Methods for Base-Pair Resolution

By Dr. Zubair Khalid, DVM, MS, PhD ·

Resolving Structural Variant Breakpoints: How to Use Split-Read and Assembly-Based Methods for Base-Pair Resolution

Key Takeaways

  • Split-read methods detect structural variant (SV) breakpoints by identifying reads that align to non-contiguous genomic locations, providing direct evidence of a junction. This approach is computationally efficient and suitable for simple SVs in unique genomic regions, but its resolution is limited by read length and can be ambiguous in repetitive sequences.
  • Assembly-based methods, particularly local assembly, reconstruct contiguous sequences (contigs) spanning breakpoints from overlapping reads. This allows for base-pair resolution of SVs that no single read spans, including those in repetitive regions or complex rearrangements, by anchoring the contig to unique flanking sequences.
  • Integrated callers like GRIDSS combine split-read, discordant read-pair, and assembly evidence within a probabilistic model to enhance SV detection and breakpoint resolution accuracy, offering a unified approach for short-read data.
  • Assembly-centric callers like SvABA prioritize local assembly for detecting and resolving SVs and small indels, particularly effective for complex cancer genomes with multiple rearrangements and copy number changes.
  • Short-read limitations in resolving breakpoints within highly repetitive regions or complex rearrangements necessitate escalation to long-read sequencing or optical genome mapping for definitive base-pair resolution and comprehensive structural characterization.
  • Reproducibility and validation are critical, requiring meticulous record-keeping of software versions, parameters, reference genomes, and quality metrics, with orthogonal validation methods like PCR, long-read sequencing, or optical mapping for high-confidence breakpoint calls.

Structural variant (SV) breakpoint resolution is the process of determining the exact nucleotide position where a genomic rearrangement begins and ends. Split-read and assembly-based methods are the two primary computational strategies for achieving base-pair resolution, and both are required in a complete variant calling workflow because each captures different classes of rearrangements. This article explains the principles behind both approaches, provides practical guidance for implementing tools such as GRIDSS and SvABA, and outlines the quality checks, records, and interpretation limits that determine whether a resolved breakpoint is reliable enough for downstream functional analysis.

The Breakpoint Resolution Problem in Structural Variant Analysis

Structural variants include deletions, duplications, inversions, insertions, translocations, and complex rearrangements that combine multiple events. Unlike single-nucleotide variants, SVs can span thousands to millions of base pairs, and their functional impact often depends on the precise breakpoint sequence. A deletion that removes a promoter has a different consequence than a deletion that fuses two genes, and only base-pair resolution can distinguish these outcomes.

The core difficulty is that short-read sequencing produces fragments of 150 to 300 base pairs. When a read spans a breakpoint, only a portion of that read aligns to the reference genome on one side of the junction, and the remainder aligns to a distant location on the other side. This split-read signature is the raw material for breakpoint detection, but it is incomplete. Reads that do not span the junction, repetitive regions that cause ambiguous alignment, and complex rearrangements with multiple junctions all create gaps in the evidence.

Assembly-based methods address these gaps by reconstructing the sequence around the breakpoint from overlapping reads. Local assembly builds a de novo contig from reads in a candidate region, then aligns that contig back to the reference to identify the exact junction. This approach can resolve breakpoints that no single read spans, including those in repetitive or complex regions.

The choice between split-read and assembly-based methods is not either-or. Production pipelines use both, often in a single tool. GRIDSS integrates split-read, discordant read-pair, and assembly evidence into a unified model. SvABA uses local assembly as its primary mechanism for detecting and resolving SVs. Understanding how each method works, what evidence it requires, and where it fails is essential for interpreting results and deciding when to escalate to long-read sequencing or optical genome mapping.

Core Principles of Split-Read Breakpoint Detection

Split-read detection identifies breakpoints by finding reads that align to two or more non-contiguous locations in the reference genome. When a read spans a deletion breakpoint, the first portion aligns to the reference position upstream of the deletion, and the second portion aligns to the position downstream. The alignment software reports this as a split alignment, with a gap or soft-clipped segment representing the unaligned portion.

How Split-Read Alignments Are Generated

Modern aligners such as BWA-MEM and minimap2 produce split alignments as part of their standard output. The aligner first maps the read to its best overall position, then attempts to align the unmapped or soft-clipped portion to another location. The resulting SAM or BAM record contains multiple alignment segments, each with its own chromosome, position, and strand.

For a simple deletion, the split-read signature is straightforward. The two aligned segments flank the deleted region, and the gap between them corresponds to the deletion size. For an inversion, the two segments align to opposite strands. For a translocation, the segments align to different chromosomes. The orientation and distance between the aligned segments define the SV type and approximate breakpoint location.

Limitations of Split-Read Evidence Alone

Split-read evidence has three significant limitations. First, the breakpoint can only be resolved to the point where the read alignment ends. If the read does not extend far enough past the junction, the exact nucleotide position remains ambiguous. Second, reads that span breakpoints in repetitive regions often align ambiguously, and the aligner may place the split at an incorrect location. Third, complex rearrangements with multiple junctions may produce split-read patterns that are consistent with several different underlying structures.

These limitations mean that split-read calls are hypotheses about breakpoint location, not confirmations. The confidence in a split-read call depends on the number of supporting reads, the length of the aligned segments, and the uniqueness of the flanking sequence. A single split read with short aligned segments in a repetitive region is weak evidence. Multiple split reads with long, uniquely mapping segments provide strong evidence.

Assembly-Based Breakpoint Resolution

Assembly-based methods reconstruct the sequence spanning a breakpoint from overlapping reads, then align the assembled contig to the reference to identify the exact junction. This approach does not require any single read to span the breakpoint. Instead, it uses the redundancy of overlapping reads to build a contiguous sequence that crosses the junction.

Local Assembly Versus Whole-Genome Assembly

Local assembly, also called targeted assembly, restricts the assembly process to candidate regions identified by preliminary evidence such as discordant read pairs or split reads. This approach is computationally efficient because it only assembles the regions most likely to contain breakpoints. GRIDSS and SvABA both use local assembly as a refinement step after initial SV detection.

Whole-genome assembly reconstructs the entire genome from all reads, then compares the assembled contigs to the reference to identify structural differences. This approach is more computationally intensive but can detect SVs that local assembly misses, particularly those in regions with no preliminary evidence. Long-read sequencing projects increasingly use whole-genome assembly, and telomere-to-telomere assemblies have resolved SVs in regions that were previously intractable.

The Assembly and Alignment Workflow

The local assembly workflow follows a standard sequence. First, the tool collects reads that map to a candidate breakpoint region, including reads that map to either side of the junction and reads that map nowhere nearby. Second, the tool assembles these reads into contigs using a de Bruijn graph or overlap-layout-consensus algorithm. Third, the tool aligns the assembled contigs back to the reference genome. Fourth, the tool identifies the breakpoint by finding where the contig alignment switches from one reference location to another.

The key advantage of assembly is that the contig provides a continuous sequence across the junction. This sequence can be examined directly for microhomology, inserted bases, or other features that reveal the mechanism of rearrangement. The contig also provides a template for validating the breakpoint by aligning individual reads back to it.

What Assembly Resolves That Split Reads Cannot

Assembly resolves breakpoints in three situations where split reads fail. First, when the breakpoint falls in a region where reads are too short to span the junction, assembly can bridge the gap using overlapping reads. Second, when the breakpoint is in a repetitive region, assembly can use the unique flanking sequence to anchor the contig and determine the correct location. Third, when the rearrangement is complex with multiple junctions, assembly can reconstruct the entire rearranged segment and reveal the order and orientation of the component pieces.

GRIDSS: Integrated Split-Read and Assembly Calling

GRIDSS is a structural variant caller that integrates split-read, discordant read-pair, and assembly evidence into a single probabilistic model. It is designed for short-read whole-genome sequencing data and is widely used in both germline and somatic variant calling workflows.

GRIDSS Evidence Collection and Processing

GRIDSS begins by collecting all reads that provide evidence of a potential structural variant. This includes reads with split alignments, read pairs with unexpected insert sizes or orientations, and reads that fail to map to the reference. The tool then performs local assembly on the reads in each candidate region, producing contigs that span potential breakpoints.

The assembled contigs are aligned back to the reference, and the alignment is analyzed to identify breakpoint junctions. GRIDSS uses a probabilistic model to assign a quality score to each breakpoint based on the supporting evidence. The model accounts for the number of supporting reads, the quality of the assembly, and the complexity of the rearrangement.

GRIDSS Output and Interpretation

GRIDSS produces a VCF file with breakpoint calls that include the exact genomic coordinates of each junction, the orientation of the joined segments, and the inserted sequence at the breakpoint. The quality score, called QUAL, reflects the confidence in the call. Higher scores indicate stronger evidence.

For germline variant calling, GRIDSS is typically run on a single sample with matched normal data for somatic calling. The tool can be run in tumor-normal mode to identify somatic SVs, and it provides filters for common artifacts such as reads that map to multiple locations or reads with low mapping quality.

Practical Considerations for GRIDSS

GRIDSS requires a reference genome, a BAM file with aligned reads, and sufficient computational resources for the assembly step. The tool is available through Bioconductor, and the official package documentation provides installation and usage instructions. The assembly step is the most computationally intensive part of the pipeline, and runtime scales with the number of candidate regions.

For reproducible workflows, GRIDSS can be integrated into pipeline frameworks such as nf-core, which provides standardized configuration and execution options. The nf-core documentation describes how to configure pipelines for different compute environments and how to ensure reproducibility across runs.

SvABA: Assembly-Based Structural Variant Detection

SvABA is a structural variant caller that uses local assembly as its primary detection mechanism. It is designed to identify SVs and small indels from short-read sequencing data, with a focus on somatic variants in cancer genomes.

SvABA Assembly Strategy

SvABA assembles reads from candidate regions into contigs, then aligns the contigs to the reference genome to identify structural differences. The tool uses a multi-step approach that first identifies regions with abnormal read depth or discordant read pairs, then assembles the reads in those regions, and finally analyzes the assembled contigs for breakpoints.

The assembly step in SvABA is designed to handle the complexity of cancer genomes, which often contain multiple rearrangements, copy number changes, and regions of high sequence similarity. The tool uses a graph-based assembly approach that can resolve complex junctions and identify the precise breakpoint sequence.

SvABA Output and Variant Representation

SvABA reports breakpoints with exact genomic coordinates and the inserted sequence at each junction. The output includes both simple SVs such as deletions and duplications, and complex SVs with multiple breakpoints. The tool also reports small indels that are detected during the assembly process.

For somatic variant calling, SvABA can be run with matched normal data to filter germline variants. The tool provides a variant quality score and annotations that indicate whether the variant is supported by assembly evidence, read-pair evidence, or both.

Practical Considerations for SvABA

SvABA requires a reference genome and aligned BAM files. The tool is computationally intensive because assembly is performed for every candidate region. Runtime can be reduced by providing a list of candidate regions or by adjusting the parameters that control the assembly process.

SvABA is available as a standalone tool and can be integrated into larger pipelines. The tool documentation provides detailed instructions for installation, configuration, and interpretation of output.

Choosing Between Split-Read and Assembly Methods

The choice between split-read and assembly-based methods depends on the research question, the data available, and the computational resources. Both approaches have strengths and weaknesses, and production pipelines typically use both.

When Split-Read Methods Are Sufficient

Split-read methods are sufficient for detecting simple SVs with strong read support in unique regions of the genome. A deletion with multiple split reads spanning the junction, all with long aligned segments and unique flanking sequence, can be confidently resolved without assembly. Split-read methods are also faster and require fewer computational resources than assembly.

For large cohort studies where the goal is to identify common SVs, split-read methods provide a practical first pass. The calls can be filtered by read support and mapping quality to retain high-confidence breakpoints.

When Assembly Is Required

Assembly is required when split-read evidence is insufficient or ambiguous. This includes breakpoints in repetitive regions, complex rearrangements with multiple junctions, and SVs where no single read spans the breakpoint. Assembly is also required when the goal is to determine the exact breakpoint sequence, including microhomology and inserted bases, because the assembled contig provides a continuous sequence across the junction.

Studies of complex structural variants in Mendelian disorders have shown that breakpoint resolution can be critical for interpreting pathogenicity. In one cohort of undiagnosed rare disease patients, short-read whole-genome sequencing identified complex SVs affecting known disease genes, and long-read sequencing was used to resolve the precise configuration of one rearrangement. The breakpoint analysis revealed microhomology and repetitive elements that suggested replication-based mechanisms of formation 11.

The Complementary Role of Long-Read Sequencing

Long-read sequencing technologies such as Oxford Nanopore and Pacific Biosciences produce reads that span entire SVs, providing direct evidence of breakpoint junctions. Long-read data can resolve breakpoints in regions that are intractable to short-read methods, including highly repetitive regions and complex rearrangements.

A study of complex duplication variants in autism spectrum disorder used Oxford Nanopore long-read sequencing to characterize rearrangements in five families. The study resolved all breakpoint junctions at nucleotide resolution and identified potential fusion genes formed through duplication rearrangements. The long-read data also allowed direct assessment of methylation status across the rearranged regions 10.

Long-read sequencing is not always necessary, but it is the escalation path when short-read methods cannot resolve a breakpoint. The decision to escalate depends on the clinical or research significance of the variant and the cost of additional sequencing.

The Role of Optical Genome Mapping in Breakpoint Validation

Optical genome mapping is a complementary technology that provides genome-wide detection of structural variants without sequencing. The method uses ultra-high-molecular-weight DNA labeled at specific sequence motifs and imaged in nanochannel arrays. The resulting maps are compared to a reference to identify structural differences.

Optical Genome Mapping Capabilities

Optical genome mapping can detect all major classes of chromosomal aberrations, including aneuploidies, deletions, duplications, translocations, inversions, insertions, isochromosomes, ring chromosomes, and complex rearrangements. A proof-of-principle study of 85 blood or cultured cell samples achieved 100 percent concordance with standard assays for all aberrations with non-centromeric breakpoints 7.

The resolution of optical genome mapping is lower than sequencing-based methods, but it provides a genome-wide view that can identify variants missed by targeted approaches. The method is particularly useful for detecting balanced SVs and for determining the genomic localization and orientation of duplicated segments.

Using Optical Genome Mapping for Validation

Optical genome mapping can validate breakpoints identified by sequencing and can resolve cases where sequencing is ambiguous. In a study of ring chromosomes and Robertsonian translocations, optical genome mapping was used to validate the structures resolved by long-read sequencing. The combination of long-read sequencing and optical genome mapping provided confidence in the breakpoint calls 8.

For clinical applications, optical genome mapping offers a cost-effective approach for comprehensive detection of chromosomal aberrations. The method can be used as a first-line test, with sequencing reserved for cases that require base-pair resolution.

Practical Workflow for Breakpoint Resolution

A complete workflow for breakpoint resolution involves multiple stages, from initial variant detection to final validation. Each stage has specific inputs, outputs, and quality checks.

Stage 1: Initial Variant Detection

The first stage is to identify candidate structural variants from aligned sequencing data. This can be done using a caller such as GRIDSS or SvABA, or using a combination of tools. The input is a BAM file with aligned reads, and the output is a VCF file with candidate breakpoints.

Quality checks at this stage include assessing the number of supporting reads, the mapping quality of those reads, and the complexity of the candidate region. Variants with weak support or ambiguous mapping should be flagged for further analysis.

Stage 2: Breakpoint Refinement

The second stage is to refine the breakpoint coordinates using assembly. This can be done within the variant caller, as with GRIDSS and SvABA, or as a separate step using a dedicated assembly tool. The input is the candidate variant list and the aligned reads, and the output is a refined VCF with base-pair resolution breakpoints.

Quality checks at this stage include examining the assembled contig for the breakpoint junction, assessing the alignment of the contig to the reference, and verifying that the breakpoint coordinates are consistent with the read evidence.

Stage 3: Validation and Interpretation

The third stage is to validate the breakpoint and interpret its functional impact. Validation can involve visual inspection of the reads in a genome browser, PCR amplification across the breakpoint, or orthogonal validation using a different technology such as optical genome mapping or long-read sequencing.

Interpretation involves determining whether the breakpoint disrupts a gene, creates a fusion gene, or affects a regulatory region. This requires annotation of the breakpoint coordinates against gene models and regulatory elements.

Stage 4: Reporting and Archiving

The final stage is to report the results and archive the data. The report should include the breakpoint coordinates, the supporting evidence, the validation status, and the functional interpretation. The data should be archived in a format that allows reproduction of the analysis.

Reproducibility requires documenting the software versions, parameters, and reference genome used in the analysis. Pipeline frameworks such as nf-core provide standardized approaches for ensuring reproducibility.

At a Glance: Breakpoint Resolution Methods

MethodEvidence TypeResolutionStrengthsLimitationsTypical Use
Split-readReads spanning the junctionBase-pair when reads extend past junctionFast, computationally efficient, direct evidenceFails in repetitive regions, requires reads spanning junctionInitial detection, simple SVs in unique regions
Local assemblyOverlapping reads assembled into contigsBase-pair from contig alignmentResolves breakpoints no single read spans, provides junction sequenceComputationally intensive, requires sufficient read depthRefinement of candidate breakpoints, complex SVs
Long-read sequencingSingle reads spanning the entire SVBase-pair from read alignmentResolves repetitive and complex regions, provides methylation dataHigher cost, lower throughput, requires specialized equipmentEscalation for unresolved breakpoints, complex rearrangements
Optical genome mappingGenome-wide maps of labeled DNAKilobase to megabaseDetects all SV classes, cost-effective, no sequencing requiredLower resolution than sequencing, cannot provide base-pair sequenceGenome-wide screening, validation of sequencing results

Records and Measurements for Breakpoint Analysis

Accurate record keeping is essential for breakpoint analysis. The records should capture the data inputs, the analysis parameters, and the quality metrics for each variant call.

Essential Records

The essential records for each breakpoint call include the sample identifier, the reference genome version, the aligner and version, the variant caller and version, the breakpoint coordinates, the SV type, the supporting read count, the mapping quality, and the assembly quality. These records allow the analysis to be reproduced and the calls to be compared across samples.

The reference genome version is particularly important because breakpoint coordinates are only meaningful relative to a specific reference. Different reference versions can have different coordinates for the same variant, and comparisons across studies require consistent reference usage.

Quality Metrics to Track

Quality metrics for breakpoint calls include the number of supporting reads, the fraction of reads with high mapping quality, the length of the assembled contig, the alignment identity of the contig to the reference, and the complexity of the breakpoint region. These metrics provide a basis for filtering calls and for prioritizing variants for validation.

For somatic variant calling, additional metrics include the variant allele fraction, the coverage at the breakpoint, and the presence of the variant in matched normal data. Somatic calls with low allele fraction or evidence in normal data should be treated with caution.

Data Management and Archiving

The raw sequencing data, aligned reads, and variant calls should be archived in a format that allows reanalysis. The NCBI provides databases for archiving sequencing data and variant calls, and the official documentation describes the submission process. Archived data should include the metadata necessary to understand the analysis, including sample information, experimental protocols, and analysis parameters.

Common Failure Patterns in Breakpoint Resolution

Understanding common failure patterns helps researchers identify problems early and avoid wasted effort on unreliable calls.

Failure Pattern 1: Insufficient Read Support

The most common failure pattern is insufficient read support for a breakpoint call. This occurs when the read depth at the breakpoint is too low to provide confidence in the call, or when the supporting reads have low mapping quality. Breakpoint calls with fewer than three supporting reads should be treated as tentative, and calls with reads that map to multiple locations should be flagged.

The solution is to increase sequencing depth, use a different library preparation method, or escalate to long-read sequencing. For germline variants, increasing depth to 30x or higher typically provides sufficient support for most breakpoints.

Failure Pattern 2: Repetitive Region Ambiguity

Breakpoints in repetitive regions often produce ambiguous alignments. The split reads may align to multiple locations, and the assembled contig may not have a unique alignment to the reference. This failure pattern is common in centromeric, telomeric, and ribosomal DNA regions.

The solution is to use long-read sequencing, which can span the entire repetitive region, or to use optical genome mapping, which does not rely on sequence alignment. A study of ring chromosomes and Robertsonian translocations used long-read sequencing and telomere-to-telomere assembly to resolve breakpoints in acrocentric p-arms, ribosomal DNA arrays, and telomeric repeats 8.

Failure Pattern 3: Complex Rearrangement Misassembly

Complex rearrangements with multiple junctions can be misassembled, producing contigs that do not accurately represent the true structure. This occurs when the assembly algorithm cannot determine the correct order and orientation of the component segments.

The solution is to use multiple assembly approaches and to validate the structure with orthogonal methods. Long-read sequencing can resolve complex rearrangements by providing reads that span multiple junctions.

Failure Pattern 4: Reference Bias

Reference bias occurs when reads that differ from the reference genome are less likely to align, causing variants to be missed or mischaracterized. This is a particular problem for SVs in regions where the reference genome is incomplete or incorrect.

The solution is to use a reference genome that is appropriate for the sample population and to be aware of the limitations of the reference. Telomere-to-telomere assemblies provide a more complete reference for human samples.

Limitations of Short-Read Breakpoint Resolution

Short-read breakpoint resolution has inherent limitations that cannot be overcome by improved algorithms or increased depth. These limitations should be considered when designing studies and interpreting results.

Read Length Constraints

Short reads of 150 to 300 base pairs cannot span large SVs. A deletion of 10 kilobases cannot be detected by a single read, and the breakpoint must be inferred from split reads or assembly. This inference is reliable for simple SVs but becomes less reliable for complex rearrangements.

Repetitive Sequence Complexity

The human genome contains extensive repetitive sequence, including transposable elements, segmental duplications, and tandem repeats. These regions produce ambiguous alignments and assembly graphs that are difficult to resolve. Short-read methods systematically miss SVs in these regions.

Structural Variant Size Limits

Short-read methods have limited sensitivity for SVs at the extremes of the size distribution. Very small SVs, such as indels of a few base pairs, are difficult to distinguish from sequencing errors. Very large SVs, such as chromosomal rearrangements, may not be detected if the breakpoints fall in regions with low coverage.

The Case for Long-Read Sequencing

Long-read sequencing addresses these limitations by providing reads that span entire SVs. A study of human de novo mutation rates used five complementary sequencing technologies to phase and assemble more than 95 percent of each diploid genome in a four-generation family. The study estimated 98 to 206 de novo mutations per transmission and demonstrated that the mutation rate varies by an order of magnitude depending on repeat content, length, and sequence identity 9.

Long-read sequencing also provides methylation data directly from the sequencing reads, allowing assessment of the functional impact of rearrangements on gene expression. A study of complex duplication variants in autism spectrum disorder used methylation analysis from long-read data to identify aberrant methylation in carriers across a rearrangement affecting the CREBBP locus 10.

Quality Controls and Reproducibility

Quality controls are essential for ensuring that breakpoint calls are reliable and that analyses can be reproduced.

Alignment Quality Controls

The alignment step should be assessed for mapping rate, insert size distribution, and coverage uniformity. Low mapping rates may indicate contamination or poor library quality. Abnormal insert size distributions may indicate problems with the library preparation. Coverage uniformity affects the sensitivity of variant detection.

Variant Calling Quality Controls

Variant calling should be assessed for the number of calls, the distribution of variant types, and the transition-transversion ratio for small variants. An unexpectedly high number of calls may indicate systematic artifacts. An abnormal distribution of variant types may indicate a problem with the calling algorithm.

Reproducibility Controls

Reproducibility requires documenting the software versions, parameters, and reference genome used in the analysis. Pipeline frameworks such as nf-core provide standardized approaches for ensuring reproducibility. The nf-core documentation describes how to configure pipelines and how to track the parameters used in each run.

Training in reproducible analysis practices is available from The Carpentries, which provides lessons on shell, Git, and data analysis. The EMBL-EBI Training program offers courses on bioinformatics data resources and analysis methods. The Galaxy Training Network provides accessible workflow training and analysis tutorials for reproducible genomics.

Safety and Regulatory Context

Breakpoint resolution has implications for clinical diagnosis and genetic counseling. The results of breakpoint analysis can determine whether a variant is pathogenic, and the interpretation can affect patient management.

Clinical Interpretation Standards

Clinical interpretation of structural variants requires consideration of the breakpoint location, the genes affected, and the mechanism of formation. Breakpoints that disrupt a gene, create a fusion gene, or affect a regulatory region are more likely to be pathogenic. Breakpoints that fall in intergenic regions with no known function are less likely to be pathogenic.

A study of complex structural variants in Mendelian disorders identified three pathogenic complex SVs in a cohort of 1324 undiagnosed rare disease patients. The study recommended that complex SVs be considered during clinical investigations and showed that resolution of breakpoints can be critical for interpreting pathogenicity 11.

Reporting and Escalation Criteria

Breakpoint calls that are clinically significant should be validated by an orthogonal method before being reported. Validation can involve PCR amplification across the breakpoint, optical genome mapping, or long-read sequencing. Calls that cannot be validated should be reported as tentative.

Professional escalation is appropriate when a breakpoint cannot be resolved by short-read methods and the variant is potentially clinically significant. Escalation to long-read sequencing or optical genome mapping should be considered in these cases.

Data Sharing and Privacy

Variant data from human samples should be shared in accordance with applicable regulations and ethical guidelines. The NCBI provides databases for sharing genomic data, and the official documentation describes the submission process and the privacy protections in place.

Professional Escalation Criteria

Knowing when to escalate from short-read to long-read methods is important for efficient use of resources. The following criteria indicate that escalation should be considered.

Unresolved Breakpoints in Clinically Significant Genes

If a breakpoint falls in or near a gene of clinical significance and cannot be resolved to base-pair resolution, escalation to long-read sequencing should be considered. The precise breakpoint may determine whether the gene is disrupted, fused, or unaffected.

Complex Rearrangements with Ambiguous Structure

If a rearrangement involves multiple breakpoints and the structure cannot be determined from short-read data, escalation should be considered. Long-read sequencing can resolve the order and orientation of the component segments.

Breakpoints in Repetitive Regions

If a breakpoint falls in a repetitive region and the exact location cannot be determined, escalation should be considered. Long-read sequencing can span the repetitive region and provide a unique alignment.

Discrepancies Between Methods

If different analysis methods produce inconsistent breakpoint calls, escalation should be considered. The discrepancy may indicate a complex rearrangement that requires longer reads to resolve.

Decision Framework for Selecting Breakpoint Resolution Methods

Choosing between split-read, assembly-based, long-read, and optical genome mapping approaches requires a structured decision process that accounts for variant characteristics, data availability, and downstream interpretation needs. A practical framework helps researchers avoid over-sequencing simple cases and under-resolving complex ones.

Tier 1: Initial Triage Based on Variant Class

The first decision point is the variant class identified during initial detection. Simple deletions and duplications with strong split-read support in unique regions can proceed directly to validation without assembly refinement. These variants typically have multiple supporting reads with long aligned segments and unambiguous mapping locations.

Complex variants require immediate escalation in analytical approach. A variant with multiple breakpoint junctions, inconsistent read-pair orientations, or evidence of both deletion and duplication in the same region should be routed to assembly-based refinement regardless of the initial call confidence. The presence of microhomology at breakpoints, detected through preliminary sequence examination, also warrants assembly-based confirmation because microhomology can indicate replication-based mechanisms that produce complex structures 11.

Tier 2: Region Complexity Assessment

The genomic context of the breakpoint determines whether short-read methods can achieve base-pair resolution. Breakpoints in unique sequence with moderate GC content and no segmental duplications are amenable to split-read resolution. Breakpoints near or within repetitive elements, including Alu elements, LINE-1 sequences, or tandem repeats, require assembly-based methods because split-read alignments in these regions are ambiguous.

For breakpoints in highly repetitive regions such as centromeres, telomeres, acrocentric p-arms, or ribosomal DNA arrays, short-read methods are unlikely to provide reliable resolution regardless of depth. A study of ring chromosomes and Robertsonian translocations demonstrated that these regions required long-read sequencing and telomere-to-telomere assembly for resolution 8. The decision to escalate to long-read sequencing should be made early for variants in these regions instead of after failed short-read attempts.

Tier 3: Evidence Strength Evaluation

Before committing to a resolution strategy, evaluate the available evidence using three quantitative measures. First, count the number of reads supporting the breakpoint junction. Fewer than three supporting reads indicates insufficient evidence for confident resolution by any short-read method. Second, assess the mapping quality of supporting reads. Reads with mapping quality below the threshold recommended by the aligner documentation should be excluded from consideration. Third, measure the distance between the outermost supporting reads and the predicted breakpoint. If no read extends more than 50 base pairs past the junction, the breakpoint coordinate is uncertain.

These measures can be recorded in a standardized format for each variant. The records should include the supporting read count, the median mapping quality, the maximum read extension past the junction, and the assembly contig length if assembly was performed. This documentation allows comparison across variants and provides a basis for escalation decisions.

Tier 4: Downstream Analysis Requirements

The intended use of the breakpoint determines the required resolution level. If the goal is to determine whether a breakpoint disrupts a coding exon, breakpoint coordinates within a few hundred base pairs may be sufficient. If the goal is to identify fusion genes, the exact breakpoint sequence is required to determine the reading frame and whether the fusion produces a functional transcript.

If the goal is to identify the mechanism of rearrangement, including microhomology-mediated end joining or non-allelic homologous recombination, base-pair resolution is essential. These mechanisms are distinguished by the sequence features at the breakpoint junction, including the length of microhomology and the presence of inserted bases 11. Methylation analysis from long-read data can also reveal position effects when a breakpoint relocates a gene near heterochromatin 8.

Tier 5: Resource Allocation Decision

The final decision point balances the cost of additional sequencing or analysis against the value of the information gained. For research studies with large cohorts, a tiered approach is efficient. Run split-read and assembly-based calling on all samples, then escalate to long-read sequencing only for variants that meet clinical significance criteria or that remain unresolved after short-read analysis.

For clinical cases, the escalation threshold should be lower. A study of complex structural variants in Mendelian disorders recommended that complex SVs be considered during clinical investigations and demonstrated that breakpoint resolution was critical for interpreting pathogenicity 11. In these cases, the cost of long-read sequencing is justified by the diagnostic value of a resolved breakpoint.

Implementing the Framework in Practice

The framework can be implemented as a decision tree with documented criteria at each node. The tree should be versioned and stored with the analysis records. Each variant call should be annotated with the tier at which it was resolved and the evidence that supported the decision.

For reproducible implementation, the decision criteria can be encoded in workflow configuration files. Pipeline frameworks such as nf-core provide standardized approaches for configuring analysis steps and documenting parameters. The nf-core documentation describes how to create reproducible workflows with versioned parameters and containerized software.

Training in the computational skills needed to implement this framework is available from multiple sources. The Galaxy Training Network provides accessible workflow training and analysis tutorials for reproducible genomics. The EMBL-EBI Training program offers courses on bioinformatics data resources and analysis methods. The Carpentries provides foundational lessons on shell, Git, and data analysis that support reproducible research practices.

Recording Framework Decisions

Each variant should have a decision record that documents the tier at which it was resolved, the evidence that supported the resolution, and the escalation path if resolution was not achieved. The record should include the date, the analyst, the software versions, and the reference genome version. These records support audit and reanalysis when new methods become available.

The decision record should also note any discrepancies between methods. If split-read and assembly-based methods produce different breakpoint coordinates, the discrepancy should be documented and investigated. Discrepancies may indicate a complex rearrangement that requires long-read sequencing for resolution.

Common Decision Errors

A common error is escalating to long-read sequencing before exhausting short-read assembly options. Local assembly with GRIDSS or SvABA can resolve many breakpoints that split-read methods miss, and this approach is substantially less expensive than long-read sequencing. The decision framework should require documented evidence that assembly-based methods were attempted before escalation.

Another common error is failing to escalate when the evidence clearly indicates the need. Breakpoints in centromeric or telomeric regions, complex rearrangements with more than three junctions, and variants with no assembly contig spanning the junction all indicate that short-read methods have reached their limit. Continuing to increase short-read depth in these cases wastes resources and delays resolution.

A third error is treating optical genome mapping as a replacement for sequencing-based resolution. Optical genome mapping provides genome-wide detection and can validate breakpoints, but it does not provide base-pair sequence 7. The method is complementary to sequencing and should be used in combination with, not instead of, sequencing-based approaches.

Frequently Asked Questions

What is the difference between split-read and assembly-based breakpoint resolution?

Split-read resolution identifies breakpoints by finding reads that align to two non-contiguous locations in the reference genome. The breakpoint is placed where the read alignment switches from one location to the other. Assembly-based resolution reconstructs the sequence spanning the breakpoint from overlapping reads, then aligns the assembled contig to the reference. Assembly can resolve breakpoints that no single read spans, including those in repetitive or complex regions.

When should I use GRIDSS versus SvABA?

GRIDSS integrates split-read, discordant read-pair, and assembly evidence into a unified model and is well suited for germline and somatic variant calling. SvABA uses local assembly as its primary detection mechanism and is designed for somatic variants in cancer genomes. The choice depends on the research question and the data available. Both tools can be used in the same workflow, with one providing initial detection and the other providing refinement.

How many supporting reads are needed for a confident breakpoint call?

The number of supporting reads needed depends on the complexity of the region and the quality of the reads. In unique regions with high mapping quality, three or more supporting reads may be sufficient. In repetitive regions or regions with low mapping quality, more reads are needed. The variant caller provides a quality score that reflects the strength of the evidence, and this score should be used to filter calls.

What is the role of optical genome mapping in breakpoint resolution?

Optical genome mapping provides genome-wide detection of structural variants without sequencing. The method can detect all major classes of chromosomal aberrations and can validate breakpoints identified by sequencing 7. The resolution is lower than sequencing-based methods, but the method is cost-effective and can be used as a first-line test.

When should I escalate to long-read sequencing?

Escalation to long-read sequencing should be considered when a breakpoint cannot be resolved by short-read methods and the variant is potentially clinically significant. This includes breakpoints in repetitive regions, complex rearrangements with ambiguous structure, and breakpoints in or near genes of clinical significance.

Can assembly-based methods resolve breakpoints in repetitive regions?

Assembly-based methods can resolve some breakpoints in repetitive regions by using unique flanking sequence to anchor the contig. However, highly repetitive regions such as centromeres, telomeres, and ribosomal DNA arrays may require long-read sequencing. A study of ring chromosomes and Robertsonian translocations used long-read sequencing and telomere-to-telomere assembly to resolve breakpoints in these regions 8.

What records should I keep for breakpoint analysis?

Records should include the sample identifier, reference genome version, aligner and version, variant caller and version, breakpoint coordinates, SV type, supporting read count, mapping quality, and assembly quality. These records allow the analysis to be reproduced and the calls to be compared across samples.

How do I validate a breakpoint call?

Validation can involve visual inspection of the reads in a genome browser, PCR amplification across the breakpoint, or orthogonal validation using a different technology such as optical genome mapping or long-read sequencing. Validation is particularly important for clinically significant variants and for variants with weak read support.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.