Resolving Breakpoints with Long Reads: How to Pinpoint Exact Insertion and Deletion Junctions in Structural Variants

By Dr. Zubair Khalid, DVM, MS, PhD ·

Resolving Breakpoints with Long Reads: How to Pinpoint Exact Insertion and Deletion Junctions in Structural Variants

Key Takeaways

  • Long-read sequencing (e.g., Oxford Nanopore Technologies, PacBio HiFi) offers base-pair resolution of structural variant (SV) breakpoints by generating reads that span entire genomic rearrangements, enabling direct observation of junction sequences, unlike short-read methods that rely on indirect inference.
  • Specialized alignment tools (e.g., Minimap2) and SV callers (e.g., Sniffles, Picky) are crucial for processing long-read data, identifying breakpoint signatures such as split or soft-clipped alignments, and generating variant calls in VCF format.
  • Visual inspection using tools like IGV is essential for confirming breakpoint support from multiple reads and assessing alignment patterns, while custom scripts are used to extract precise junction sequences for mechanistic interpretation and primer design.
  • Validation of breakpoints is critical and can be achieved through orthogonal methods like optical genome mapping for large SVs or PCR amplification followed by Sanger sequencing for base-pair confirmation of junction sequences.
  • Challenges in breakpoint resolution include mapping ambiguity in highly repetitive genomic regions, which can be mitigated by using complete reference genomes (e.g., T2T assemblies), and potential alignment errors at homopolymers, particularly with ONT data.
  • For clinical applications, rigorous diagnostic validation, standardized variant interpretation according to HGVS nomenclature, and adherence to data privacy regulations are paramount when utilizing long-read breakpoint analysis.

Structural variants (SVs) are large genomic rearrangements including deletions, insertions, inversions, duplications, and translocations that span typically 50 base pairs or more. Short-read sequencing platforms generate reads of 150 to 300 base pairs, which often cannot span the full length of these variants or uniquely map across repetitive elements flanking the rearrangement. Long-read sequencing platforms, including Oxford Nanopore Technologies (ONT) and Pacific Biosciences (PacBio), produce reads from 10 kilobases to over 100 kilobases, enabling reads to span entire SV events and align across the breakpoint junction. This article explains how long reads achieve base-pair resolution of SV breakpoints, how to extract and validate junction sequences from alignment files using tools such as the Integrative Genomics Viewer (IGV) and custom scripts, and the limitations encountered in repetitive genomic regions. The intended readers are biology students, researchers, laboratory professionals, and life-science practitioners who need a practical workflow for pinpointing exact insertion and deletion junctions in their own sequencing data.

At a Glance

The table below summarizes the key decisions a researcher faces when using long reads for SV breakpoint resolution. Each row presents a common scenario, the recommended approach, and the evidence basis for that choice.

ScenarioRecommended ApproachEvidence Basis
Detecting SVs in cancer genomes with complex rearrangementsUse long-read pipelines such as Picky that exploit nanopore reads for genome-wide breakpoint detection at nucleotide resolutionPicky identified the full spectrum of SVs with superior specificity and sensitivity relative to short-read analyses in a breast cancer model [<a href="#ref-1">1</a>]
Resolving SVs in highly repetitive regions such as acrocentric p-arms, ribosomal DNA arrays, and telomeric repeatsCombine long-read sequencing with telomere-to-telomere assembly and tools like BigClipper for breakpoint discoveryLong-read sequencing with T2T assembly resolved ring chromosomes and Robertsonian translocations previously intractable by short reads [<a href="#ref-2">2</a>]
Validating large SVs detected by optical genome mappingUse nanopore sequencing as a complementary high-resolution method to refine breakpoint positionsNanopore sequencing detected 94,400 SVs compared to 49,677 by optical mapping, and provided base-pair refinement of breakpoints [<a href="#ref-3">3</a>]
Diagnosing rare monogenic disorders with suspected structural variantsApply full-genome analysis combining long-range assembly and whole-genome sequencing for breakpoint-resolved SV detectionFull-genome analysis identified SVs missed by short reads, including non-coding duplications, with a 40% diagnostic yield in 50 cases [<a href="#ref-4">4</a>]

Understanding Structural Variant Breakpoints

A breakpoint is the precise genomic coordinate where a structural rearrangement interrupts the reference sequence. For a simple deletion, two breakpoints exist: one at the start of the deleted segment and one at the end. The junction sequence is the novel sequence formed when the two sides of the breakpoint are joined. For insertions, the breakpoints flank the inserted sequence, which may originate from elsewhere in the genome or from a mobile element. For inversions and translocations, the breakpoints define the orientation and location of the rearranged segments.

Short-read sequencing detects SVs primarily through discordant read pairs and split reads. Discordant pairs indicate that two reads map farther apart than expected or in an unexpected orientation, but they do not reveal the exact junction sequence. Split reads, where a single read maps to two different genomic locations, can provide breakpoint information, but the short read length limits how much flanking sequence is available for unique alignment. In repetitive regions, short reads often fail to map uniquely, leaving breakpoints unresolved.

Long reads overcome these limitations by spanning the entire SV event. A single long read can cover a deletion of several kilobases, including the flanking unique sequence on both sides of the junction. When this read aligns to the reference genome, the alignment shows a gap or a split at the exact breakpoint position. The base-pair resolution comes from the read sequence itself: the nucleotides at the junction are directly observed, not inferred from paired-end distances or read depth.

The practical outcome for researchers is the ability to determine the exact sequence context of a rearrangement. This information supports mechanistic interpretation, such as identifying microhomology-mediated end joining or non-homologous end joining as the likely formation mechanism [<a href="#ref-2">2</a>]. It also enables PCR primer design for validation, functional assessment of disrupted genes, and accurate reporting of variant coordinates for clinical or research databases.

Core Principles of Long-Read Breakpoint Resolution

Read Length and Spanning Capacity

The fundamental advantage of long reads is their ability to span entire SV events. A deletion of 5 kilobases requires a read of at least 5 kilobases plus flanking unique sequence to align confidently on both sides. ONT reads commonly exceed 10 kilobases, and PacBio HiFi reads typically range from 10 to 25 kilobases. For larger SVs, such as the 195-kilobase intergenic deletion near ITPR1 detected in a Parkinson's disease study, no single read can span the entire event [<a href="#ref-3">3</a>]. In these cases, breakpoint resolution relies on reads that span the junction itself, even if they do not span the full variant.

The junction-spanning read aligns partially to one side of the breakpoint and partially to the other side. The alignment software must recognize the split or soft-clipped portion of the read. Soft-clipped bases are those at the ends of a read that do not align to the reference. When a read spans a deletion junction, the bases on one side of the junction align to the reference upstream of the deletion, and the bases on the other side align downstream. The point where the alignment switches from one location to the other defines the breakpoint.

Alignment Signatures at Breakpoints

Three alignment signatures indicate a breakpoint in long-read data. The first is a split alignment, where a single read produces two or more alignment segments to different genomic locations. The second is a soft-clipped alignment, where the unaligned portion of the read contains the junction sequence. The third is an insertion in the read relative to the reference, where the read contains extra bases not present in the reference at the breakpoint.

For a deletion, the read aligns continuously to the reference until the deletion start, then skips the deleted sequence, and resumes alignment at the deletion end. The alignment shows a gap in the reference coordinates. For an insertion, the read contains additional bases that do not align to the reference at that position. These inserted bases may be a novel sequence, a duplicated segment from elsewhere, or a mobile element insertion.

The distinction between these signatures matters for interpretation. A deletion junction shows reference bases on both sides with a gap between them. An insertion junction shows reference bases flanking a novel sequence. A complex rearrangement may show combinations of these signatures, such as a deletion-inversion where the read aligns in reverse orientation on one side of the breakpoint [<a href="#ref-2">2</a>].

Base-Pair Resolution and Junction Sequence Extraction

Base-pair resolution means that the exact nucleotide where the reference sequence is interrupted can be identified. This resolution is achieved when a read spans the junction and the alignment software reports the precise coordinate where the split occurs. The junction sequence itself is the novel sequence formed by the rearrangement. For a simple deletion, the junction sequence is the concatenation of the reference bases immediately upstream and downstream of the deleted segment. For an insertion, the junction sequence includes the inserted bases.

Extracting the junction sequence requires examining the read alignment at the breakpoint. The Integrative Genomics Viewer (IGV) displays read alignments graphically, allowing visual inspection of split reads and soft-clipped bases. Custom scripts can parse alignment files in SAM or BAM format to extract the sequence at the breakpoint. The extracted sequence can then be compared to the reference to confirm the rearrangement and identify any micro-insertions or microhomology at the junction.

Micro-insertions are small inserted sequences, often a few base pairs, found at breakpoint junctions. A study using the Picky pipeline found micro-insertions to be common structural features associated with SVs in cancer genomes [<a href="#ref-1">1</a>]. These micro-insertions may be templated from nearby sequence or may be non-templated additions. Their presence provides clues about the DNA repair mechanism that created the rearrangement.

Practical Workflow for Breakpoint Resolution

Step 1: Generate or Obtain Long-Read Sequencing Data

Long-read data can be generated on ONT or PacBio platforms. ONT offers real-time sequencing and can produce ultra-long reads exceeding 100 kilobases. PacBio HiFi sequencing produces highly accurate reads of 10 to 25 kilobases with base quality scores comparable to short reads. The choice of platform depends on the research question, available equipment, and budget.

For researchers without access to a sequencing facility, public data repositories provide long-read datasets. The NCBI Sequence Read Archive (SRA) hosts raw sequencing data from numerous long-read projects [<a href="#ref-5">5</a>]. The European Bioinformatics Institute also provides training on accessing and using these data resources [<a href="#ref-6">6</a>]. Downloading existing datasets allows researchers to practice breakpoint analysis without generating new sequencing data.

Step 2: Align Long Reads to a Reference Genome

Long-read alignment requires specialized aligners designed for the error profiles of ONT and PacBio data. Minimap2 is a widely used aligner that handles both platforms. The output is a SAM file, typically converted to a sorted BAM file with an index for efficient access.

Alignment parameters affect breakpoint detection. For ONT data, the aligner must tolerate the higher error rate, particularly in homopolymer regions. For PacBio HiFi data, the aligner can use more stringent parameters because the base quality is higher. The choice of reference genome also matters. A telomere-to-telomere reference, such as the T2T-CHM13 assembly, includes previously unresolved repetitive regions and improves breakpoint mapping in these areas [<a href="#ref-2">2</a>].

Step 3: Call Structural Variants

SV callers for long-read data identify breakpoints from alignment signatures. Tools such as Sniffles, cuteSV, and Picky analyze split reads, soft-clipped bases, and read depth to identify SVs. The output is a VCF file with breakpoint coordinates and supporting read information.

The Picky pipeline was specifically designed for nanopore long reads and demonstrated superior specificity and sensitivity relative to short-read analyses in a breast cancer model [<a href="#ref-1">1</a>]. It identified the full spectrum of SV architectures and uncovered repetitive DNA as a major source of variation. For researchers studying cancer genomes, Picky provides a validated approach for comprehensive SV detection.

Step 4: Visualize Breakpoints in IGV

IGV provides a graphical interface for examining read alignments at candidate breakpoints. Load the BAM file and the reference genome, then navigate to the breakpoint coordinates from the VCF file. The display shows reads aligned to the reference, with split reads and soft-clipped bases visible as colored segments.

Visual inspection confirms that the breakpoint is supported by multiple reads and that the alignment signature matches the expected pattern. For a deletion, multiple reads should show the same gap in reference coordinates. For an insertion, multiple reads should show the same inserted sequence. The number of supporting reads and the consistency of the breakpoint coordinates provide confidence in the call.

Step 5: Extract Junction Sequences with Custom Scripts

Custom scripts can extract the exact junction sequence from the alignment file. The script parses the BAM file, identifies reads that span the breakpoint, and extracts the sequence at the junction. For a deletion, the script identifies the reference coordinates where the read alignment splits and extracts the reference bases on both sides. For an insertion, the script extracts the inserted bases from the read sequence.

The extracted junction sequence can be written to a FASTA file for further analysis. This sequence can be compared to the reference genome using BLAST or similar tools to identify the origin of inserted sequences. The NCBI BLAST service provides this comparison capability [<a href="#ref-5">5</a>]. The junction sequence can also be used to design PCR primers for experimental validation.

Step 6: Validate Breakpoints

Validation confirms that the breakpoint is real and not an alignment artifact. Multiple lines of evidence support validation. First, the breakpoint should be supported by multiple independent reads. Second, the breakpoint coordinates should be consistent across reads. Third, the junction sequence should be reproducible when extracted by different methods.

Optical genome mapping provides an orthogonal validation method for large SVs. A comparison of ONT and optical mapping in Parkinson's disease patients found that both methods identified SVs larger than 50 kilobases, but optical mapping detected fewer total SVs while detecting six times more in the 50 to 80 kilobase range [<a href="#ref-3">3</a>]. The study concluded that optical mapping can be a powerful first-line method for detecting large SVs, but requires a high-resolution method such as nanopore sequencing to refine breakpoint positions [<a href="#ref-3">3</a>]. This complementary approach provides strong validation for large rearrangements.

Options and Tradeoffs in Long-Read Platforms

Oxford Nanopore Technologies

ONT sequencing offers several advantages for breakpoint analysis. The platform can produce ultra-long reads exceeding 100 kilobases, which can span very large SVs. The real-time sequencing allows researchers to stop data collection once sufficient coverage is achieved. The portable devices make sequencing accessible in diverse laboratory settings.

The tradeoff is lower per-base accuracy compared to PacBio HiFi. ONT reads have higher error rates, particularly in homopolymer regions and areas with complex repeats. This error profile can complicate breakpoint identification if the junction falls within a low-complexity region. However, the study of Parkinson's disease patients found that nanopore sequencing was highly capable of detecting large variants independently [<a href="#ref-3">3</a>].

PacBio HiFi

PacBio HiFi sequencing produces reads with accuracy exceeding 99.9 percent, comparable to short-read platforms. The high accuracy simplifies breakpoint identification because the junction sequence is read with high confidence. The circular consensus sequencing approach generates multiple passes over each molecule, producing a consensus read with minimal errors.

The tradeoff is shorter read lengths compared to ONT, typically 10 to 25 kilobases. This length is sufficient for many SVs but may not span the largest rearrangements. The cost per gigabase is also higher than ONT, which may limit coverage depth for large genomes.

Choosing Between Platforms

The choice between ONT and PacBio depends on the specific research question. For detecting the full spectrum of SVs in a cancer genome, ONT with a pipeline like Picky provides comprehensive detection [<a href="#ref-1">1</a>]. For resolving breakpoints in repetitive regions where base accuracy is critical, PacBio HiFi may be preferable. For very large SVs exceeding 50 kilobases, combining long-read sequencing with optical genome mapping provides complementary information [<a href="#ref-3">3</a>].

Researchers should also consider the bioinformatics infrastructure available. Both platforms generate data that can be analyzed with the same alignment and variant calling tools, but the parameter settings differ. The Galaxy Training Network provides accessible workflows for long-read analysis that can be adapted to either platform [<a href="#ref-7">7</a>]. The nf-core community maintains standardized pipelines for reproducible long-read analysis [<a href="#ref-8">8</a>].

Records and Measurements for Breakpoint Analysis

Documenting Breakpoint Coordinates

Accurate record keeping is essential for reproducible breakpoint analysis. For each SV, record the chromosome, start coordinate, end coordinate, and variant type. Record the reference genome build used for alignment, as coordinates differ between builds. Record the number of supporting reads and the read depth at the breakpoint.

The VCF file from the SV caller provides a structured format for these records. Each variant entry includes the chromosome, position, reference allele, alternate allele, and quality score. Additional fields may include the number of supporting reads, the read names, and the breakpoint confidence interval.

Measuring Breakpoint Support

Breakpoint support is measured by the number of reads that span the junction and the consistency of their alignment. A high-confidence breakpoint is supported by multiple reads with identical or nearly identical breakpoint coordinates. A low-confidence breakpoint may be supported by a single read or by reads with inconsistent coordinates.

The read depth at the breakpoint provides additional context. For a deletion, the read depth drops to zero across the deleted region. For an insertion, the read depth increases across the inserted region if the inserted sequence is present in multiple copies. These depth changes can be visualized in IGV and quantified with depth-of-coverage tools.

Recording Junction Sequences

The junction sequence should be recorded for each breakpoint. This sequence is the novel sequence formed by the rearrangement and is the key evidence for the variant. Record the length of the junction sequence, the presence of any micro-insertions, and the sequence context on both sides.

For clinical or publication purposes, the junction sequence may be reported in the variant description. The Human Genome Variation Society (HGVS) nomenclature provides a standard format for describing sequence variants, including those with breakpoint resolution. Following this standard ensures that the variant description is unambiguous and machine-readable.

Common Failure Patterns in Breakpoint Resolution

Repetitive Regions and Mapping Ambiguity

The most common failure pattern is the inability to map reads uniquely in repetitive regions. Long reads that span a breakpoint in a repetitive region may align to multiple locations in the reference, producing ambiguous breakpoint coordinates. This problem is particularly acute in acrocentric p-arms, ribosomal DNA arrays, and telomeric repeats, which were previously recalcitrant to sequencing [<a href="#ref-2">2</a>].

The telomere-to-telomere assembly addressed this problem by providing a complete reference that includes these repetitive regions [<a href="#ref-2">2</a>]. When using an older reference with gaps, breakpoints in these regions cannot be resolved. Researchers should use the most complete reference available for their organism and be aware of the limitations of older references.

Low Coverage at the Breakpoint

Breakpoints may be missed if coverage is insufficient. Long-read sequencing has higher per-read cost than short-read sequencing, so coverage is often lower. If only one or two reads span a breakpoint, the call may be unreliable. Increasing coverage improves breakpoint detection but increases cost.

The required coverage depends on the variant type and the genome complexity. For diploid genomes, each allele should be covered by multiple reads. For somatic variants in cancer, the variant allele fraction may be low, requiring higher coverage to detect the breakpoint. The Picky study identified breakpoints in a breast cancer model, demonstrating that long-read analysis can detect somatic SVs when coverage is adequate [<a href="#ref-1">1</a>].

Alignment Errors at Homopolymers and Low-Complexity Sequence

ONT reads have higher error rates in homopolymer regions, where the same nucleotide is repeated many times. These errors can shift the apparent breakpoint position by one or more base pairs. If the breakpoint falls within or near a homopolymer, the exact junction may be ambiguous.

PacBio HiFi reads have lower error rates in these regions, but the shorter read length may not span the full SV. Combining platforms or using optical genome mapping for validation can resolve this ambiguity [<a href="#ref-3">3</a>]. The optical mapping comparison study found that nanopore sequencing could refine breakpoint positions identified by optical mapping, providing a complementary approach [<a href="#ref-3">3</a>].

Complex Rearrangements with Multiple Breakpoints

Complex SVs may involve multiple breakpoints that are difficult to resolve individually. A deletion-inversion, for example, has two breakpoints where the sequence orientation changes. An interchromosomal rearrangement has breakpoints on two different chromosomes. These complex events require careful analysis of the alignment signatures at each breakpoint.

The study of ring chromosomes and Robertsonian translocations found that complex SVs could be resolved using long-read sequencing with T2T assembly and the BigClipper tool [<a href="#ref-2">2</a>]. The analysis resolved 10 of 13 cases, including all ring chromosomes and a Robertsonian translocation [<a href="#ref-2">2</a>]. The breakpoint sequences suggested mechanisms of SV formation including microhomology-mediated end joining, non-homologous end joining, and non-allelic homologous recombination [<a href="#ref-2">2</a>].

Limitations of Long-Read Breakpoint Resolution

Read Length Constraints for Very Large SVs

Even long reads have length limits. A deletion of 500 kilobases cannot be spanned by a single ONT read, even with ultra-long sequencing. In these cases, breakpoint resolution relies on reads that span the junction, not the full variant. The junction-spanning reads provide the breakpoint coordinates, but the full extent of the deletion must be inferred from read depth or optical mapping.

The comparison of ONT and optical mapping in Parkinson's disease found that both methods identified SVs larger than 50 kilobases, but optical mapping detected significantly larger deletions and insertions [<a href="#ref-3">3</a>]. For very large SVs, optical mapping provides a genome-wide view that complements the base-pair resolution of long-read sequencing [<a href="#ref-3">3</a>].

Error Rates and Base Accuracy

The error rate of the sequencing platform affects the confidence in the junction sequence. ONT reads have higher error rates, which can introduce errors in the extracted junction sequence. These errors may be misinterpreted as micro-insertions or micro-deletions at the breakpoint.

PacBio HiFi reads have lower error rates, providing higher confidence in the junction sequence. However, the shorter read length may not span the full SV. Researchers should consider the tradeoff between read length and base accuracy when choosing a platform for breakpoint analysis.

Reference Genome Completeness

The reference genome used for alignment determines which breakpoints can be resolved. A reference with gaps in repetitive regions cannot support breakpoint mapping in those regions. The T2T assembly provided a gapless reference for the human genome, enabling breakpoint resolution in previously intractable regions [<a href="#ref-2">2</a>].

For non-human organisms, the reference genome may have more gaps and errors. Researchers should assess the quality of their reference genome and consider whether a new assembly is needed for their research question. The NCBI provides resources for accessing and evaluating reference genomes [<a href="#ref-5">5</a>].

Computational Requirements

Long-read alignment and variant calling require substantial computational resources. The alignment of long reads is computationally intensive, and the BAM files are large. SV callers may require significant memory and processing time. Researchers without access to high-performance computing may need to use cloud resources or smaller datasets.

The Galaxy Training Network provides accessible workflows that can be run on public servers [<a href="#ref-7">7</a>]. The nf-core community maintains standardized pipelines that can be deployed on various computing infrastructures [<a href="#ref-8">8</a>]. The Carpentries offers foundational training in shell, Git, and programming that supports reproducible bioinformatics analysis [<a href="#ref-9">9</a>].

Quality Controls and Reproducibility

Alignment Quality Metrics

Quality control begins with the alignment. The SAM file contains mapping quality scores for each read alignment. Reads with low mapping quality should be filtered or examined carefully. The alignment statistics, such as the percentage of reads mapped and the median read length, provide an overview of data quality.

The BAM file should be sorted and indexed for efficient access. Duplicate reads, which may arise from PCR amplification or sequencing artifacts, should be marked or removed. For long-read data, duplicates are less common than in short-read data, but they can still occur.

Variant Call Quality Filters

SV callers assign quality scores to each variant call. These scores reflect the number of supporting reads, the consistency of the breakpoint coordinates, and the alignment quality. Researchers should apply quality filters to remove low-confidence calls.

The specific filters depend on the variant caller and the research question. For high-confidence breakpoint resolution, require multiple supporting reads with consistent breakpoint coordinates. For discovery purposes, more permissive filters may be appropriate, with validation of candidate variants by visual inspection or orthogonal methods.

Reproducibility Through Workflow Management

Reproducible analysis requires documenting the exact commands, parameters, and software versions used. Workflow management systems such as nf-core provide standardized pipelines that ensure reproducibility across runs and users [<a href="#ref-8">8</a>]. The nf-core documentation describes how to configure and run these pipelines [<a href="#ref-8">8</a>].

The Galaxy Training Network provides tutorials that teach reproducible analysis practices [<a href="#ref-7">7</a>]. The Carpentries lessons cover foundational skills in shell scripting, version control with Git, and data management that support reproducible research [<a href="#ref-9">9</a>]. The EMBL-EBI Training program offers courses on bioinformatics data resources and analysis [<a href="#ref-6">6</a>].

Validation by Independent Methods

Validation by an independent method provides the strongest evidence for a breakpoint. Optical genome mapping can validate large SVs detected by long-read sequencing [<a href="#ref-3">3</a>]. PCR amplification across the breakpoint followed by Sanger sequencing provides validation at base-pair resolution. The PCR primers are designed from the junction sequence, and the resulting amplicon is sequenced to confirm the junction.

The comparison study of ONT and optical mapping found that both methods detected a benign intergenic deletion near ITPR1, and optical mapping validated a previously published 7-megabase PRKN inversion [<a href="#ref-3">3</a>]. This cross-platform validation demonstrates the value of using multiple methods for high-confidence breakpoint resolution [<a href="#ref-3">3</a>].

Safety and Regulatory Context for Clinical Applications

Diagnostic Validation Requirements

When long-read breakpoint analysis is used for clinical diagnosis, the methods must meet regulatory standards for diagnostic testing. The full-genome analysis study demonstrated a 40% diagnostic yield in 50 cases of rare monogenic disorders, including 35% in exome-negative cases [<a href="#ref-4">4</a>]. The study identified structural variants missed by short reads, including non-coding duplications, and phased variants across distances of more than 180 kilobases [<a href="#ref-4">4</a>].

Clinical validation requires demonstrating that the test accurately and reliably detects the variants it is designed to detect. This validation includes analytical sensitivity, analytical specificity, and reproducibility studies. The breakpoint resolution provided by long reads must be confirmed by orthogonal methods before clinical reporting.

Variant Interpretation and Reporting

The interpretation of structural variants in a clinical context requires assessing the potential functional impact. A deletion that disrupts a gene coding sequence is more likely to be pathogenic than an intergenic deletion. The breakpoint sequence can reveal whether the rearrangement disrupts regulatory elements, creates fusion genes, or alters gene dosage.

The full-genome analysis study found that structural variants could be prioritized using a variant prioritization pipeline [<a href="#ref-4">4</a>]. The study suggested that longer DNA technologies could replace multiple tests for monogenic disorders and expand the range of variants detected [<a href="#ref-4">4</a>]. This finding has implications for the future of clinical genetic testing.

Data Sharing and Privacy

Clinical genomic data are subject to privacy regulations that vary by jurisdiction. Researchers and clinicians must ensure that patient data are handled in compliance with applicable laws. Data sharing for research purposes requires appropriate consent and de-identification.

The NCBI provides databases for depositing and accessing genomic data, including the Sequence Read Archive and dbGaP for controlled-access data [<a href="#ref-5">5</a>]. Researchers should follow the data sharing policies of their funding agencies and institutions.

Professional Escalation Criteria

When to Seek Additional Expertise

Breakpoint analysis can be challenging, particularly for complex rearrangements or repetitive regions. Researchers should seek additional expertise when they encounter the following situations:

  • Breakpoints in highly repetitive regions that cannot be resolved with the available reference
  • Complex rearrangements with multiple breakpoints or interchromosomal involvement
  • Inconsistent breakpoint coordinates across supporting reads
  • Discrepancies between long-read and optical mapping results
  • Clinical cases where the breakpoint interpretation affects patient management

Consulting Bioinformatics Specialists

Bioinformatics specialists can provide guidance on alignment parameters, variant calling strategies, and custom script development. They can also help with computational infrastructure and workflow optimization. The EMBL-EBI Training program offers courses that build bioinformatics skills [<a href="#ref-6">6</a>]. The Bioconductor project provides documentation and support for genomic analysis packages [<a href="#ref-10">10</a>].

Referring to Reference Laboratories

For clinical cases, reference laboratories with expertise in long-read sequencing and structural variant analysis can provide confirmatory testing. These laboratories have validated protocols and can provide interpretation in the context of clinical guidelines. The full-genome analysis study demonstrated the clinical utility of long-read approaches for rare disease diagnosis [<a href="#ref-4">4</a>].

Decision Framework for Selecting Breakpoint Validation Methods

Choosing the correct validation strategy for a resolved breakpoint depends on the variant size, the genomic context, the available budget, and the downstream use of the result. A structured decision framework prevents both over-validation, which wastes resources, and under-validation, which risks reporting artifacts as real rearrangements. The framework below organizes the choice by variant characteristics and the strength of evidence already obtained from the long-read alignment.

Step 1: Classify the Variant by Size and Complexity

Begin by recording the approximate size of the variant and the number of breakpoints involved. Simple deletions and insertions below 10 kilobases typically have two breakpoints that can be confirmed by PCR amplification across the junction. Variants between 10 and 50 kilobases may require a combination of PCR and optical genome mapping if the breakpoint falls in a region with low mappability. Variants larger than 50 kilobases benefit from optical genome mapping as a first-line confirmation because PCR across the full event is impractical and the breakpoint-spanning reads may be few [<a href="#ref-3">3</a>].

Complex rearrangements with more than two breakpoints, such as deletion-inversions or interchromosomal translocations, require validation at each junction independently. The study of ring chromosomes and Robertsonian translocations demonstrated that complex SVs resolved by long reads needed orthogonal confirmation by optical genome mapping to establish the full architecture [<a href="#ref-2">2</a>]. For these cases, validate each breakpoint separately instead of assuming that confirmation of one junction validates the entire event.

Step 2: Assess the Supporting Read Evidence

Before selecting a validation method, quantify the support for the breakpoint from the alignment file. Count the number of reads that span the junction and record the consistency of their breakpoint coordinates. A breakpoint supported by five or more reads with identical coordinates in a diploid sample provides strong evidence that the variant is real. A breakpoint supported by one or two reads requires validation regardless of the variant size.

Record the mapping quality of the supporting reads. Reads with mapping quality below the threshold used by your aligner may align ambiguously, particularly in repetitive regions. The Picky pipeline demonstrated that nanopore long reads could identify breakpoints at nucleotide resolution in cancer genomes, but the study also found repetitive DNA to be a major source of variation [<a href="#ref-1">1</a>]. Low mapping quality at a breakpoint should trigger additional validation instead of acceptance of the call.

Step 3: Match the Validation Method to the Variant Class

For simple deletions and insertions below 10 kilobases with strong read support, PCR amplification across the junction followed by Sanger sequencing provides base-pair confirmation. Design primers in the unique flanking sequence on both sides of the junction. The amplicon should be the size predicted by the junction sequence. Sanger sequencing of the amplicon confirms the exact nucleotide sequence at the junction.

For variants between 10 and 50 kilobases, PCR remains feasible if the junction-spanning sequence is unique. However, if the breakpoint falls within a repetitive element, PCR primers may not amplify specifically. In these cases, optical genome mapping provides an orthogonal view of the variant without requiring sequence-specific amplification. The comparison of ONT and optical mapping in Parkinson's disease found that optical mapping detected six times more variants in the 50 to 80 kilobase range than nanopore sequencing, suggesting that optical mapping is particularly valuable for confirming mid-sized variants [<a href="#ref-3">3</a>].

For variants larger than 50 kilobases, optical genome mapping is the preferred first-line validation. The same study found that both methods identified SVs larger than 50 kilobases, but optical mapping detected significantly larger deletions and insertions [<a href="#ref-3">3</a>]. Optical mapping provides a genome-wide view that confirms the presence and approximate location of the variant, while the long-read data provides the base-pair breakpoint resolution [<a href="#ref-3">3</a>].

Step 4: Apply the Escalation Criteria for Ambiguous Cases

Escalate to additional validation methods when the initial validation is inconclusive. If PCR amplification fails but the read support is strong, the failure may indicate that the breakpoint sequence contains a repetitive element that interferes with primer binding. In this case, attempt optical genome mapping or use a different PCR strategy with primers further from the junction.

If optical genome mapping and long-read sequencing disagree on the breakpoint location, resolve the discrepancy by examining the raw alignment in IGV. The discrepancy may arise from a misassembly in the reference genome or from a complex rearrangement that both methods detect partially. The full-genome analysis study demonstrated that long-range assembly combined with whole-genome sequencing could detect and map structural variants missed by short reads, including non-coding duplications [<a href="#ref-4">4</a>]. For clinical cases where the discrepancy affects patient management, refer to a laboratory with expertise in structural variant analysis.

Step 5: Document the Validation Decision

Record the validation method used for each breakpoint and the outcome. This record supports reproducibility and provides evidence for the confidence level assigned to each variant. For publication or clinical reporting, state the validation method explicitly and describe any limitations of that method for the specific variant class.

The nf-core documentation emphasizes the importance of standardized workflows for reproducible genomic analysis [<a href="#ref-8">8</a>]. Incorporate the validation decision into your analysis records so that the rationale for each choice is available to collaborators or reviewers. The Galaxy Training Network provides tutorials on documenting and sharing analysis workflows [<a href="#ref-7">7</a>].

Record System for Breakpoint Validation Tracking

A structured record system tracks each breakpoint from initial detection through validation. This system prevents loss of information and supports audit for clinical or publication purposes. The record should capture the variant identifier, the breakpoint coordinates, the supporting read count, the validation method, and the validation outcome.

Required Fields for Each Breakpoint Record

Create a table with the following fields for each breakpoint: variant identifier, chromosome, breakpoint start coordinate, breakpoint end coordinate, variant type, reference genome build, number of supporting reads, mapping quality of supporting reads, validation method, validation date, validation outcome, and notes. The reference genome build is essential because coordinates differ between builds, and reporting a breakpoint without the build creates ambiguity.

Record the junction sequence in a separate field or as an attachment. The junction sequence is the key evidence for the variant and should be preserved in the record. For insertions, record the inserted sequence and its origin if identified by comparison to the reference. The NCBI provides BLAST services for identifying the origin of inserted sequences [<a href="#ref-5">5</a>].

Tracking Validation Status

Assign a status to each breakpoint: pending validation, validated by one method, validated by two methods, or failed validation. A breakpoint validated by two independent methods, such as long-read sequencing and optical genome mapping, receives the highest confidence. A breakpoint that fails validation should be flagged and the reason recorded.

The comparison of ONT and optical mapping in Parkinson's disease provides an example of cross-platform validation. Both methods detected a benign intergenic deletion near ITPR1, and optical mapping validated a previously published 7-megabase PRKN inversion [<a href="#ref-3">3</a>]. Recording both the long-read and optical mapping results for each variant creates a complete validation history.

Periodic Review of Validation Records

Review the validation records periodically to identify patterns in validation failures. If a particular genomic region consistently fails PCR validation, the region may contain repetitive elements that interfere with primer binding. If optical mapping consistently disagrees with long-read calls in a specific region, the reference genome may have an error in that region.

The Bioconductor project provides packages for genomic data management and visualization that support systematic review of validation records [<a href="#ref-10">10</a>]. The Carpentries lessons teach data management practices that support organized record keeping [<a href="#ref-9">9</a>]. The EMBL-EBI Training program offers courses on data resources and analysis that include record-keeping best practices [<a href="#ref-6">6</a>].

Troubleshooting Breakpoint Validation Failures

PCR Amplification Fails Across the Junction

When PCR fails to amplify across a predicted junction, first verify that the primers are specific to the flanking sequence. Use a primer design tool that checks for off-target binding. If the primers are specific, the failure may indicate that the breakpoint coordinates are incorrect or that the junction sequence differs from the prediction.

Re-examine the alignment in IGV to confirm the breakpoint coordinates. If the coordinates are correct, the junction sequence may contain a secondary structure that inhibits amplification. Try a different polymerase or add a GC-rich enhancer to the reaction. If amplification continues to fail, use optical genome mapping as an alternative validation method.

Optical Genome Mapping Does Not Detect the Variant

Optical genome mapping may fail to detect a variant that is clearly present in the long-read data. This failure can occur if the variant is smaller than the detection limit of the optical mapping platform or if the variant falls in a region with low label density. The comparison study found that optical mapping detected fewer total SVs than nanopore sequencing, detecting 49,677 compared to 94,400 [<a href="#ref-3">3</a>].

If optical mapping does not detect the variant, review the size distribution of the variants that optical mapping detects reliably. The study found that optical mapping detected six times more variants in the 50 to 80 kilobase range than nanopore sequencing [<a href="#ref-3">3</a>]. If the variant is smaller than this range, optical mapping may not be the appropriate validation method.

Breakpoint Coordinates Differ Between Methods

When long-read sequencing and optical mapping report different breakpoint coordinates for the same variant, the discrepancy may arise from the different resolution of the two methods. Optical mapping provides approximate breakpoint locations based on label patterns, while long-read sequencing provides base-pair resolution. The study concluded that optical mapping requires a high-resolution method to refine breakpoint positions [<a href="#ref-3">3</a>].

Resolve the discrepancy by examining the long-read alignment at the optical mapping breakpoint location. If the long-read data supports a breakpoint at the optical mapping location, the discrepancy may be a reporting difference instead of a true disagreement. If the long-read data does not support the optical mapping location, the optical mapping call may be an artifact.

Validation Fails in Repetitive Regions

Breakpoints in repetitive regions present the most difficult validation challenges. The study of ring chromosomes and Robertsonian translocations found that breakpoints localized to acrocentric p-arms, ribosomal DNA arrays, and telomeric repeats were previously recalcitrant to sequencing [<a href="#ref-2">2</a>]. The telomere-to-telomere assembly enabled resolution of these regions, but validation remains challenging because PCR primers may not be specific in repetitive sequence.

For breakpoints in repetitive regions, use optical genome mapping as the primary validation method. If optical genome mapping is not available, validate the breakpoint by examining the consistency of the supporting reads and the reproducibility of the junction sequence across multiple reads. The full-genome analysis study demonstrated that long-range assembly could resolve structural variants in challenging regions [<a href="#ref-4">4</a>].

Welfare and Safety Context for Clinical Reporting

Avoiding False Positive Clinical Reports

The clinical impact of a false positive structural variant report is substantial. A patient may receive unnecessary follow-up testing, surveillance, or intervention based on a variant that is not present. The full-genome analysis study demonstrated a 40% diagnostic yield in 50 cases of rare monogenic disorders, but this yield depends on accurate variant detection [<a href="#ref-4">4</a>].

For clinical reporting, require validation by at least one independent method before reporting a breakpoint as confirmed. The validation method should be appropriate for the variant size and genomic context. For variants that affect patient management, consider validation by two independent methods.

Communicating Uncertainty in Reports

Clinical reports should communicate the confidence level of each breakpoint call. A breakpoint supported by multiple reads and validated by an orthogonal method receives high confidence. A breakpoint supported by a single read or with inconsistent coordinates receives low confidence and should be reported as unconfirmed.

The full-genome analysis study found that structural variants could be prioritized using a variant prioritization pipeline [<a href="#ref-4">4</a>]. This prioritization supports clinical interpretation by focusing attention on variants most likely to be pathogenic. The report should state the validation status and the evidence supporting each call.

Data Protection for Patient Samples

Clinical breakpoint analysis involves patient samples that are subject to privacy regulations. Ensure that patient data are handled in compliance with applicable laws and institutional policies. The NCBI provides databases for depositing and accessing genomic data, including controlled-access databases for sensitive data [<a href="#ref-5">5</a>].

Researchers and clinicians should follow the data sharing policies of their funding agencies and institutions. The EMBL-EBI Training program provides guidance on responsible data management [<a href="#ref-6">6</a>]. The Carpentries lessons teach data management practices that support privacy and security [<a href="#ref-9">9</a>].

Frequently Asked Questions

What is the difference between breakpoint resolution and structural variant detection?

Structural variant detection identifies the presence and approximate location of a rearrangement. Breakpoint resolution determines the exact nucleotide where the rearrangement interrupts the reference sequence. Long reads enable breakpoint resolution because they span the junction and provide the sequence context on both sides. Short reads can detect SVs through discordant read pairs but often cannot resolve the exact breakpoint, particularly in repetitive regions.

How many long reads are needed to confirm a breakpoint?

The number of supporting reads needed depends on the variant type, the genome ploidy, and the required confidence. For a heterozygous variant in a diploid genome, multiple reads supporting the breakpoint on each allele provide high confidence. For somatic variants with low allele fraction, more reads are needed to distinguish the variant from sequencing errors. Visual inspection in IGV and consistency of breakpoint coordinates across reads provide additional confidence.

Can long reads resolve breakpoints in all repetitive regions?

Long reads resolve breakpoints in many repetitive regions that are intractable to short reads, but not all. The telomere-to-telomere assembly enabled breakpoint resolution in acrocentric p-arms, ribosomal DNA arrays, and telomeric repeats [<a href="#ref-2">2</a>]. However, some regions may still be ambiguous if the flanking sequence is not unique. The completeness of the reference genome and the read length determine which regions can be resolved.

What is the role of optical genome mapping in breakpoint validation?

Optical genome mapping provides a genome-wide view of large structural variants and can validate variants detected by long-read sequencing. A comparison study found that optical mapping detected fewer total SVs than nanopore sequencing but detected larger variants and validated a previously published inversion [<a href="#ref-3">3</a>]. The study concluded that optical mapping is a powerful first-line method for large SVs but requires high-resolution methods to refine breakpoint positions [<a href="#ref-3">3</a>].

How do micro-insertions at breakpoints inform the mechanism of SV formation?

Micro-insertions are small inserted sequences found at breakpoint junctions. A study using the Picky pipeline found micro-insertions to be common structural features associated with SVs in cancer genomes [<a href="#ref-1">1</a>]. The presence and sequence of micro-insertions can indicate the DNA repair mechanism, such as microhomology-mediated end joining or non-homologous end joining [<a href="#ref-2">2</a>]. Breakpoint sequences from resolved ring chromosomes and translocations suggested these mechanisms of SV formation [<a href="#ref-2">2</a>].

What are the limitations of using short reads for breakpoint resolution?

Short reads of 150 to 300 base pairs often cannot span the full length of structural variants or map uniquely across repetitive elements. Discordant read pairs indicate the presence of a variant but do not reveal the exact junction sequence. Split reads can provide breakpoint information, but the short read length limits the flanking sequence available for unique alignment. Long reads overcome these limitations by spanning the junction and providing base-pair resolution.

How should breakpoint coordinates be reported for clinical or publication purposes?

Breakpoint coordinates should be reported with the reference genome build, the chromosome, and the exact nucleotide positions. The Human Genome Variation Society (HGVS) nomenclature provides a standard format for describing sequence variants. The junction sequence should be reported when available. For clinical reporting, the variant should be validated by an orthogonal method and interpreted in the context of the patient's phenotype.

What training resources are available for learning long-read breakpoint analysis?

The Galaxy Training Network provides accessible workflows and tutorials for long-read analysis [<a href="#ref-7">7</a>]. The nf-core documentation describes standardized pipelines for reproducible analysis [<a href="#ref-8">8</a>]. The Carpentries offers foundational training in shell, Git, and programming [<a href="#ref-9">9</a>]. The EMBL-EBI Training program provides courses on bioinformatics data resources and analysis [<a href="#ref-6">6</a>]. The Bioconductor project documents genomic analysis packages and workflows [<a href="#ref-10">10</a>].

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

[1] [Picky comprehensively detects high-resolution structural variants in nanopore long reads.](https://pubmed.ncbi.nlm.nih.gov/29713081). Nature methods, 2018. [2] [Resolution of ring chromosomes, Robertsonian translocations, and complex structural variants from long-read sequencing and telomere-to-telomere assembly.](https://pubmed.ncbi.nlm.nih.gov/39520989). American journal of human genetics, 2024. [3] [Complementarity of Long-Reads and Optical Mapping in Parkinson's Disease for Structural Variants.](https://pubmed.ncbi.nlm.nih.gov/41653029). Annals of clinical and translational neurology, 2026. [4] [Application of full-genome analysis to diagnose rare monogenic disorders.](https://pubmed.ncbi.nlm.nih.gov/34556655). NPJ genomic medicine, 2021. [5] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information. [6] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute. [7] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project. [8] [nf-core Documentation](https://nf-co.re/docs). nf-core. [9] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries. [10] [Bioconductor](https://bioconductor.org/). Bioconductor Project.

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.