# How to Detect Structural Variants from RNA-Seq Data: Challenges and Computational Approaches


## Key Takeaways

- RNA-seq detects structural variants (SVs) using split-read junctions, read-depth changes, and discordant paired-end mappings, mirroring DNA-based methods but applied to transcribed sequences. The primary challenge is distinguishing true SVs from normal splicing events, particularly for split-read analysis where canonical and alternative splicing can mimic deletion breakpoints.

- Detection power is intrinsically linked to transcript expression levels; highly expressed genes provide robust signals for read-depth and split-read analysis, while lowly expressed genes yield insufficient coverage for reliable SV calls, limiting RNA-seq's utility for genes with low tissue-specific expression.

- Isoform diversity complicates read-depth interpretation, as variations in exon inclusion or skipping can create coverage patterns resembling deletions or duplications, necessitating isoform-aware models or careful manual review for accurate copy number inference.

- RNA-seq SV calling is best employed as a complementary approach to DNA-based methods, adding diagnostic value when relevant genes are expressed in the sampled tissue, but contributing little when they are not, as demonstrated in studies of inherited diseases where blood expression levels were limiting.

- Practical workflows require careful consideration of sample type, library preparation (e.g., strand-specific for fusion detection), splice-aware alignment (e.g., STAR, HISAT2), and appropriate SV calling tools (e.g., STAR-Fusion, Arriba), followed by rigorous filtering and orthogonal validation with DNA-based methods.

---

Structural variant detection from RNA sequencing data is possible through split-read and read-depth signals, but the approach carries distinct limitations tied to splicing complexity and transcript expression levels. Researchers can use RNA-seq to identify deletions, duplications, inversions, and fusions that affect gene expression, yet the accuracy of these calls depends on careful alignment strategies, coverage thresholds, and orthogonal validation. This article explains the computational workflow for RNA-seq based SV detection, the biological and technical challenges that shape results, and the practical decisions researchers must make when interpreting variant calls from transcriptome data.

## At a Glance

RNA-seq based structural variant detection uses the same underlying signals as DNA-based approaches but applies them to transcribed sequences. Split reads indicate junctions between non-contiguous genomic regions, read depth reflects copy number changes across expressed loci, and discordant paired-end mappings reveal rearrangements. The critical difference from DNA-seq is that RNA-seq only samples expressed transcripts, so detection power varies by gene expression level and isoform structure.

| Detection Signal | What It Captures | Primary Challenge in RNA-Seq | Best Use Case |
| --- | --- | --- | --- |
| Split-read analysis | Exact breakpoint junctions from reads spanning two genomic locations | Distinguishing true SVs from normal splicing events | Fusions and deletions with expressed breakpoints |
| Read-depth analysis | Copy number changes inferred from coverage across exons | Coverage variation from expression differences between genes | Large deletions and duplications in highly expressed genes |
| Discordant paired-end mapping | Paired reads mapping farther apart or in wrong orientation than expected | Library insert size variation and transcript structure complexity | Insertions, inversions, and translocations involving expressed regions |

The practical implication is that RNA-seq SV calling works best as a complementary approach to DNA-based methods. In a diagnostic study of 1000 patients with inherited eye diseases, complementary whole-blood RNA-seq supported classification of two unclear variants, but the authors noted that the added diagnostic value was limited by low expression levels of the major disease genes in blood. This pattern repeats across applications: RNA-seq adds value when the relevant genes are expressed in the sampled tissue, and it contributes little when they are not.

## Why RNA-Seq Presents Unique Challenges for SV Detection

Structural variant calling from DNA sequencing has matured through dedicated algorithms that leverage read depth, split reads, and assembly. RNA-seq introduces a different set of complications because the input material is transcribed RNA, not genomic DNA. The transcriptome is a filtered and processed version of the genome, and that processing directly interferes with the signals SV callers rely on.

### Splicing Confounds Split-Read Analysis

The most direct challenge is that RNA-seq reads span exon-exon junctions as a normal feature of gene expression. A read that aligns to two distant genomic locations is the primary signal for a deletion breakpoint in DNA-based calling. In RNA-seq, that same alignment pattern appears constantly for every multi-exon transcript. The computational problem is distinguishing a true structural variant breakpoint from a canonical or alternative splicing event.

Splice-aware aligners such as STAR and HISAT2 handle this by allowing reads to split across annotated or novel junctions. The output contains junction information that can be repurposed for SV detection, but the interpretation requires additional filtering. A split read supporting a fusion between two genes must be evaluated against the possibility that it represents a read-through transcript or an alignment artifact. The distinction matters clinically because fusion detection from RNA-seq is a standard approach in cancer genomics, and false positives from spurious junctions create reporting problems.

### Expression Level Determines Detection Power

DNA sequencing provides relatively uniform coverage across the genome, aside from GC bias and mappability issues. RNA-seq coverage is proportional to transcript abundance, which varies by several orders of magnitude between genes. A gene expressed at one transcript per million reads will have sparse coverage, making split-read and read-depth detection unreliable. A highly expressed gene may have hundreds of reads covering each exon, providing robust signal for copy number changes.

This expression dependence means that RNA-seq SV calling has an inherent detection bias toward highly expressed transcripts. Lowly expressed genes, including many disease-relevant transcription factors and developmental regulators, will have insufficient coverage for confident SV calls. The inherited eye disease study demonstrated this limitation directly: RNA-seq added diagnostic value only for genes with adequate blood expression, while the major disease genes were not expressed at useful levels in that tissue.

### Isoform Diversity Complicates Read-Depth Interpretation

Read-depth based copy number analysis assumes that coverage across a genomic interval reflects the number of copies present. In RNA-seq, coverage across exons within a single gene varies according to isoform usage. A cassette exon that is skipped in the predominant isoform will show reduced coverage relative to constitutive exons, creating a pattern that resembles a deletion. Conversely, alternative promoter usage or 3-prime end processing can create coverage gradients that mimic partial gene deletions.

Long-read RNA sequencing adds another layer of isoform information but also introduces its own challenges. Single-cell long-read targeted sequencing of ovarian cancer revealed extensive transcriptional variation, including associations between expressed single nucleotide variants and alternative transcript structures. The authors used hybridization capture to improve on-target transcript yield by 29-fold, demonstrating that targeted approaches can address throughput limitations. However, the complexity of isoform diversity means that read-depth based SV inference from RNA-seq requires isoform-aware models or careful manual review.

## Core Principles of RNA-Seq SV Detection

The computational approaches for RNA-seq SV detection build on principles established for DNA-based calling but adapt them to transcriptome-specific signals. Understanding these principles helps researchers choose appropriate tools and interpret results correctly.

### Split-Read Principle Applied to Transcripts

Split-read detection identifies reads where one portion aligns to one genomic location and another portion aligns to a distant location. In DNA-seq, this directly indicates a breakpoint. In RNA-seq, the split may represent a splicing event or a structural variant. The distinguishing features include whether the junction is annotated, whether the two sides have consistent orientation, and whether the junction is supported by multiple independent reads.

Fusion detection tools such as STAR-Fusion and Arriba implement this principle by first aligning reads with a splice-aware aligner, then clustering chimeric alignments that suggest intergenic junctions. The output includes the genomic coordinates of both partners, the number of supporting reads, and annotation status. These tools are widely used in cancer genomics where gene fusions are actionable biomarkers.

### Read-Depth Principle Applied to Expressed Regions

Read-depth analysis for RNA-seq measures coverage across exonic intervals and compares it to expected levels based on gene expression. The challenge is that expected coverage varies by gene, so the analysis must normalize within genes or use external controls. Some approaches compare coverage between exons of the same gene to detect partial deletions. Others compare expression levels across samples to detect copy number changes that affect overall transcript abundance.

The sensitivity of read-depth approaches depends on the magnitude of the copy number change and the baseline expression level. A homozygous deletion of a highly expressed gene produces a clear drop to zero coverage. A heterozygous deletion reduces coverage by roughly half, which may be detectable with sufficient read depth but requires careful statistical modeling. Amplifications increase coverage proportionally to copy number, but the relationship can be distorted by feedback regulation of gene expression.

### Discordant Paired-End Principle in Transcript Context

Paired-end sequencing generates reads from both ends of a fragment. When the fragment length is known, the expected distance and orientation between the two reads are predictable. Discordant pairs indicate rearrangements. In RNA-seq, the fragment is derived from cDNA, and the insert size distribution depends on the library preparation method. Transcript structure adds complexity because a pair may span an exon-exon junction, causing the reads to map with unexpected distance even in the absence of a structural variant.

Despite these complications, discordant paired-end analysis can identify large-scale rearrangements involving expressed regions. The approach works best when combined with split-read analysis, as discordant pairs provide evidence for the presence of a rearrangement while split reads localize the breakpoint.

## Practical Workflow for RNA-Seq SV Detection

A practical workflow for RNA-seq SV detection follows a sequence of decisions from experimental design through variant interpretation. Each step involves tradeoffs between sensitivity, specificity, and computational cost.

### Step 1: Define the Biological Question and Sample Type

The first decision is whether RNA-seq is the appropriate technology for the SV detection question. If the goal is comprehensive genome-wide SV detection, DNA sequencing is the appropriate primary approach. RNA-seq adds value when the question concerns expressed structural variants, particularly fusions that create chimeric transcripts or SVs that affect expression levels.

Sample type matters because expression is tissue-specific. Blood is convenient for clinical studies but may not express the genes of interest. Tumor samples express cancer-specific transcripts but have variable cellularity and RNA quality. The inherited eye disease study illustrates the tissue limitation: blood RNA-seq could not support classification of variants in genes that are not expressed in blood. Researchers should verify that the genes of interest are expressed in the available sample type before investing in RNA-seq for SV detection.

### Step 2: Choose Library Preparation and Sequencing Strategy

Standard short-read RNA-seq with paired-end 150 base pair reads is sufficient for many fusion detection applications. The paired-end information helps resolve complex rearrangements, and the read length supports split-read alignment across exon-exon junctions. Strand-specific library preparation provides information about transcript orientation, which helps distinguish genuine fusions from read-through artifacts.

Long-read RNA sequencing provides full-length transcript information and can resolve isoform structure that short reads cannot. The scTaILoR-seq method demonstrated that targeted long-read approaches can achieve high on-target yield, making them feasible for focused studies. However, long-read RNA-seq has lower throughput and higher cost per transcript, so the choice depends on whether isoform resolution is essential for the biological question.

### Step 3: Align Reads with Splice-Aware Aligners

Splice-aware alignment is a prerequisite for RNA-seq SV detection. Aligners such as STAR, HISAT2, and minimap2 for long reads handle exon-exon junctions during alignment. The choice of aligner affects downstream SV calling because each aligner has different junction detection sensitivity and error profiles.

Alignment parameters matter for SV detection. Allowing novel junction discovery increases sensitivity for fusion detection but also increases false positives from spurious alignments. Restricting to annotated junctions reduces false positives but misses novel fusions. The balance depends on whether the study prioritizes discovery or precision.

### Step 4: Run SV Calling Tools with Appropriate Parameters

Multiple tools accept RNA-seq alignments and call structural variants. Fusion callers such as STAR-Fusion and Arriba specialize in intergenic junctions. General SV callers designed for DNA-seq can also process RNA-seq alignments, but their assumptions about coverage uniformity and insert size distributions may not hold.

The choice of tool should match the SV class of interest. Deletions and duplications within expressed genes are best detected with read-depth approaches. Fusions and translocations require split-read and discordant paired-end analysis. Inversions involving expressed regions may be detectable through split reads at the breakpoints, but the internal rearrangement is not visible in transcript data.

### Step 5: Filter and Annotate Candidate Variants

Raw SV calls from RNA-seq contain a high proportion of false positives. Filtering strategies include requiring minimum read support, checking for annotation status, evaluating breakpoint sequence context, and comparing against known artifacts. Read-through transcripts from adjacent genes are a common source of false fusion calls. Alignment artifacts from repetitive regions create spurious split reads.

Annotation of candidate variants should include the affected genes, the predicted functional consequence, and the expression level of the involved transcripts. Variants in lowly expressed genes may be technically valid but biologically irrelevant if the transcript is not present at meaningful levels.

### Step 6: Validate with Orthogonal Methods

RNA-seq SV calls require validation before clinical reporting or publication. The standard approach is DNA-based confirmation using whole-genome sequencing, whole-exome sequencing, or targeted breakpoint PCR. The validation strategy depends on the variant class and the available samples.

In the Open Pediatric Cancer Project, harmonized multiomic data from over 6000 pediatric cancer patients included structural variants and fusions called from whole-genome sequencing and RNA-seq. The project integrated these data types to deliver reproducible workflows, demonstrating that multi-platform validation is feasible at scale. Researchers should follow similar principles: RNA-seq findings should be cross-referenced against DNA-level evidence when available.

## Computational Tools and Resources for RNA-Seq SV Analysis

The bioinformatics ecosystem provides multiple entry points for researchers who need to implement RNA-seq SV detection. Official training resources and community pipelines reduce the burden of building workflows from scratch.

### Training Resources for Building Analysis Skills

The Galaxy Training Network offers accessible tutorials for RNA-seq analysis, including alignment, quantification, and differential expression. These tutorials provide step-by-step instructions that are useful for researchers who are new to command-line analysis or who prefer graphical interfaces. The training materials cover quality control, read alignment, and interpretation of results, which are foundational skills for SV detection.

The Carpentries lessons provide foundational computing skills including shell, Git, and programming in Python or R. These skills are prerequisites for reproducible bioinformatics analysis. Researchers who can navigate the command line, manage scripts, and use version control will find it easier to implement and document SV calling workflows.

The EMBL-EBI Training portal offers learning pathways for bioinformatics data resources and practical analysis. These materials cover the use of major databases and analysis services, which helps researchers understand the data resources available for variant annotation and interpretation.

### Community Pipelines for Reproducible Workflows

The nf-core documentation describes community-developed pipelines that follow best practices for reproducibility and configuration. These pipelines use the Nextflow workflow manager and containerized software, making them portable across computing environments. For RNA-seq analysis, nf-core provides pipelines for RNA-seq alignment and quantification, and the framework supports customization for SV detection.

Bioconductor provides official documentation for R packages used in genomic analysis. Packages for RNA-seq analysis, variant annotation, and visualization are available through Bioconductor, and the documentation includes installation instructions and workflow examples. Researchers who prefer R for downstream analysis will find these resources valuable.

### Database Resources for Variant Interpretation

The NCBI Data Resources provide access to sequence databases, variation databases, and analysis services. These resources support variant interpretation by providing reference sequences, gene annotations, and population frequency data. Researchers should use these resources to evaluate whether candidate SVs overlap known genes or regulatory elements.

## Options and Tradeoffs in RNA-Seq SV Detection

Different experimental and computational choices produce different tradeoffs between sensitivity, specificity, cost, and interpretability. Researchers should select approaches based on their specific biological questions and constraints.

### Short-Read versus Long-Read RNA Sequencing

Short-read RNA-seq remains the most common approach because of cost, throughput, and established analysis tools. Paired-end short reads provide sufficient information for most fusion detection and for read-depth analysis of highly expressed genes. The main limitation is that short reads cannot resolve complex isoform structures or phase variants across long distances.

Long-read RNA-seq provides full-length transcript information, enabling direct observation of isoform structure and variant phasing. The longcallR tool demonstrated joint SNP calling, haplotype phasing, and allele-specific analysis from long RNA-seq reads, identifying significant allele-specific splicing events per sample. This capability is valuable for understanding how structural variants affect transcript structure and allele-specific expression. The tradeoff is lower throughput and higher cost, which limits the number of cells or samples that can be analyzed.

### Genome-Guided versus De Novo Transcriptome Analysis

Genome-guided analysis aligns reads to a reference genome and uses gene annotations to interpret results. This approach leverages existing knowledge and is generally more sensitive for known genes. De novo transcriptome assembly reconstructs transcripts without a reference, which can discover novel transcripts but requires more computational resources and produces more fragmented results.

For SV detection, genome-guided analysis is the standard approach because SV breakpoints need to be mapped to genomic coordinates. De novo assembly may complement genome-guided analysis by revealing transcripts that do not align well to the reference, but it is not a primary SV detection strategy.

### Targeted versus Whole-Transcriptome Approaches

Whole-transcriptome RNA-seq captures all expressed genes and supports discovery of unexpected SVs. Targeted approaches use hybridization capture or amplicon-based methods to focus sequencing on genes of interest. The scTaILoR-seq method demonstrated that targeted capture can improve on-target transcript yield by 29-fold, making long-read sequencing more efficient for focused studies.

The choice between targeted and whole-transcriptome approaches depends on whether the study aims to discover novel SVs or to characterize known candidates. Targeted approaches provide deeper coverage of specific genes, improving detection sensitivity for lowly expressed targets. Whole-transcriptome approaches provide broader coverage but with lower depth per gene.

## Observations and Measurements for Quality Control

Quality control is essential for RNA-seq SV detection because technical artifacts can masquerade as structural variants. Researchers should track specific metrics throughout the analysis and establish thresholds for data quality.

### RNA Quality and Quantity Metrics

RNA integrity affects the proportion of reads that map to exons and the uniformity of coverage across transcripts. Degraded RNA produces reads that are biased toward the 5-prime end of transcripts, distorting read-depth signals. The RNA integrity number or equivalent metric should be recorded for each sample, and samples with poor integrity should be flagged for cautious interpretation.

Library preparation artifacts can introduce biases that mimic structural variants. Adapter contamination, PCR duplicates, and index hopping create spurious signals. Quality control reports from tools such as FastQC and MultiQC should be reviewed for these issues before proceeding with SV calling.

### Alignment and Coverage Statistics

The proportion of reads that map to the genome, the proportion that map to exons, and the uniformity of coverage across transcripts provide important quality signals. Low mapping rates suggest contamination or alignment issues. Low exonic mapping rates suggest genomic DNA contamination or poor RNA quality. Coverage uniformity across transcripts affects the reliability of read-depth based SV calls.

For SV detection specifically, the number of chimeric or split alignments should be tracked. An unusually high number of chimeric alignments may indicate library preparation artifacts or alignment errors. An unusually low number may indicate that the aligner parameters are too restrictive for the data.

### Reproducibility Metrics

Replicate samples provide the most direct assessment of reproducibility. Variants called in multiple replicates are more likely to be genuine. The concordance between replicates should be reported for SV calls, particularly for low-confidence variants. Technical replicates control for library preparation and sequencing variability, while biological replicates control for biological variability.

## Records and Documentation for Reproducible Analysis

Reproducible SV calling requires documentation of every analysis decision, from software versions to parameter settings. The nf-core documentation emphasizes the importance of versioned pipelines and containerized software for reproducibility. Researchers should adopt similar practices even when not using nf-core pipelines.

### Version Control for Code and Configuration

All analysis scripts and configuration files should be tracked in a version control system such as Git. The Carpentries lessons provide training in Git and GitHub that is directly applicable to bioinformatics projects. Version control enables researchers to reproduce previous analyses, compare results across versions, and collaborate effectively.

### Analysis Logs and Parameter Records

Each SV calling run should produce a log that records the software version, command-line parameters, reference genome version, and input file checksums. These logs enable exact reproduction of the analysis and provide evidence for audit purposes. The Open Pediatric Cancer Project released dockerized workflows and versioned processed data, demonstrating a model for reproducible multiomic analysis.

### Data Management and Storage

Raw sequencing data, alignment files, and variant calls should be stored according to a defined data management plan. The NCBI Data Resources provide repositories for raw sequencing data and processed results. Researchers should deposit data in appropriate repositories to enable verification and reuse by the scientific community.

## Common Failure Patterns in RNA-Seq SV Detection

Understanding common failure modes helps researchers troubleshoot their analyses and interpret results appropriately. The following patterns appear frequently in RNA-seq SV detection projects.

### False Fusions from Read-Through Transcription

Read-through transcription produces chimeric transcripts that span adjacent genes without a genomic rearrangement. These transcripts are common in cancer cells and can be mistaken for gene fusions. The distinguishing feature is that read-through fusions involve genes in the same orientation on the same chromosome with no evidence of a genomic breakpoint. Filtering against this pattern requires comparing RNA-seq fusion calls with DNA-level evidence.

### Coverage Artifacts from Sequence Composition

GC-rich and repetitive regions produce coverage artifacts that mimic copy number changes. PCR amplification bias during library preparation creates uneven coverage that is not related to genomic copy number. Read-depth based SV callers may interpret these artifacts as deletions or duplications. Normalization methods that account for sequence composition can reduce these false positives.

### Alignment Errors in Repetitive Regions

Reads from repetitive regions often align to multiple genomic locations, and aligners may assign them arbitrarily. This creates spurious split reads and discordant pairs that suggest structural variants. Filtering reads with low mapping quality or multi-mapping status reduces these artifacts but may also remove genuine variants in repetitive regions.

### Expression-Dependent Detection Bias

The most fundamental limitation is that RNA-seq cannot detect SVs in genes that are not expressed in the sampled tissue. This is not a technical artifact but a biological property of the data. Researchers must interpret negative results cautiously: failure to detect an SV in RNA-seq does not mean the SV is absent, only that it is not detectable in the transcriptome of the sampled tissue.

## Limitations and Interpretation Boundaries

RNA-seq SV detection has inherent limitations that shape how results should be interpreted and reported. Researchers should communicate these limitations clearly in publications and clinical reports.

### Incomplete Coverage of the Genome

RNA-seq only samples transcribed regions, which constitute a small fraction of the human genome. SVs in non-coding regions, regulatory elements, and genes that are not expressed in the sampled tissue are invisible to RNA-seq. This limitation is particularly important for clinical applications where a comprehensive genomic assessment is needed.

### Inability to Determine Genomic Breakpoints Precisely

RNA-seq provides evidence about transcript structure but does not directly reveal genomic breakpoints. A fusion transcript may result from a genomic rearrangement, but the exact genomic coordinates require DNA-level analysis. Similarly, a deletion that removes an exon may be detected through reduced coverage, but the genomic extent of the deletion requires DNA sequencing to define.

### Confounding of Expression Changes and Copy Number Changes

Changes in transcript abundance can result from copy number changes, transcriptional regulation, or RNA stability. Read-depth analysis of RNA-seq cannot distinguish these mechanisms. A gene with reduced expression due to transcriptional repression will show the same coverage pattern as a gene with a heterozygous deletion. This confounding limits the interpretation of read-depth based SV calls from RNA-seq.

### Limited Power for Lowly Expressed Genes

The inherited eye disease study demonstrated that RNA-seq has limited diagnostic value when disease genes are not expressed in the sampled tissue. This principle extends to all applications: the sensitivity of RNA-seq SV detection is directly proportional to expression level. Researchers should estimate the expected expression of target genes in the sampled tissue before designing RNA-seq based SV detection studies.

## Safety and Regulatory Context for Clinical Applications

When RNA-seq SV detection is used for clinical diagnosis or treatment decisions, additional considerations apply. The regulatory landscape for genomic testing varies by jurisdiction, and researchers should be aware of the requirements in their setting.

### Validation Requirements for Clinical Testing

Clinical laboratories must validate their SV detection methods before offering them as clinical tests. Validation typically involves testing samples with known variants, establishing sensitivity and specificity, and documenting the analytical performance. RNA-seq based SV detection requires additional validation because of the expression-dependent detection bias.

### Reporting Standards for Incidental Findings

RNA-seq analysis may reveal variants in genes unrelated to the clinical question. Laboratories need policies for handling incidental findings, including whether to report them and how to communicate results to patients. These policies should be established before clinical testing begins.

### Data Sharing and Privacy Considerations

Genomic data are sensitive personal information. Researchers and laboratories must comply with applicable privacy regulations when storing, sharing, and publishing RNA-seq data. The NCBI Data Resources provide controlled-access repositories for human genomic data that balance data sharing with privacy protection.

## Professional Escalation Criteria

Researchers should escalate RNA-seq SV findings to appropriate professionals when specific conditions are met. The following criteria help determine when additional expertise or validation is needed.

### When to Consult a Clinical Geneticist

RNA-seq SV findings that may explain a patient's phenotype should be discussed with a clinical geneticist before any clinical action. The geneticist can assess the medical significance of the finding, recommend confirmatory testing, and coordinate clinical management. This is particularly important for variants in genes associated with inherited diseases.

### When to Request DNA-Based Confirmation

Any RNA-seq SV call that will be used for clinical decision-making should be confirmed with DNA-based methods. Whole-genome sequencing provides the most comprehensive confirmation, but targeted breakpoint analysis may be sufficient for specific variants. The confirmation should be performed in a clinical laboratory with appropriate validation.

### When to Seek Bioinformatics Consultation

Researchers who encounter unexpected patterns in their SV calls, such as an unusually high number of fusions or poor reproducibility between replicates, should consult with a bioinformatics specialist. The specialist can review the analysis pipeline, identify potential artifacts, and recommend alternative approaches.

## Decision Framework for Selecting RNA-Seq SV Detection Approaches

Choosing the right strategy for RNA-seq based structural variant detection requires a structured evaluation of biological context, technical constraints, and downstream interpretation needs. A decision framework helps researchers avoid the common error of applying a single tool or workflow to every question. The framework below organizes the key decisions into sequential checkpoints that can be documented and revisited as project goals evolve.

### Checkpoint 1: Confirm Expression Evidence Before Committing to RNA-Seq

The first decision point is whether the genes of interest are expressed in the available sample type at levels sufficient for SV detection. This determination should be made before library preparation and sequencing, not after data collection. The inherited eye disease study demonstrated the consequence of skipping this step: blood RNA-seq could not support classification of variants in genes that are not expressed in blood, and the added diagnostic value was limited by low expression levels of the major disease genes in that tissue.

To assess expression feasibility, researchers should consult existing expression databases such as the Genotype-Tissue Expression project or The Cancer Genome Atlas for the tissue type under study. These resources provide median transcript per million values across tissues and can indicate whether target genes are likely to have adequate coverage. A practical threshold is to require that the gene of interest has a median expression level that would produce at least 10 to 20 reads across the breakpoint region given the planned sequencing depth. If expression evidence is absent or marginal, researchers should either select a different sample type, increase sequencing depth substantially, or use DNA-based methods as the primary detection strategy.

### Checkpoint 2: Match Detection Signal to Variant Class

Different structural variant classes require different detection signals, and the choice of signal determines which tools are appropriate. The Genome biology review of structural variant calling approaches emphasizes that multiple methods have emerged to target various SV classes, zygosities, and size ranges, and that no single approach works across the full spectrum of variation.

For fusion detection, split-read analysis is the primary signal because fusions create chimeric transcripts with junctions that can be identified in aligned reads. Read-depth analysis is not suitable for fusion detection because the copy number of the involved loci may be unchanged. For deletions and duplications that affect expressed exons, read-depth analysis is the appropriate signal because these variants change the number of copies of the transcribed region. Discordant paired-end mapping provides supporting evidence for rearrangements but is rarely sufficient on its own because transcript structure creates discordant mappings as a normal feature.

The decision framework should specify which variant classes are of interest before selecting tools. A study focused on gene fusions in cancer should prioritize split-read based fusion callers. A study investigating copy number changes in expressed genes should prioritize read-depth based approaches. A study attempting to detect multiple variant classes should plan to integrate signals from multiple tools, which increases complexity and requires a clear strategy for reconciling conflicting evidence.

### Checkpoint 3: Evaluate the Tradeoff Between Discovery and Precision

The balance between discovering novel variants and avoiding false positives depends on the intended use of the results. Discovery-oriented studies, such as those characterizing the landscape of fusions in a new cancer type, may accept higher false positive rates in exchange for sensitivity. Precision-oriented studies, such as those generating variants for clinical reporting, require stringent filtering and orthogonal validation.

This tradeoff is operationalized through alignment parameters and filtering thresholds. Allowing novel junction discovery during alignment increases sensitivity for unannotated fusions but also increases false positives from spurious alignments. Restricting to annotated junctions reduces false positives but misses novel events. The longcallR study identified significant allele-specific splicing events per sample, of which 46 percent involved unannotated junctions, demonstrating that novel junction discovery can reveal biologically meaningful events that would be missed by annotation-restricted approaches.

Researchers should document their position on this tradeoff explicitly in the analysis plan. The documentation should specify the minimum read support threshold, whether novel junctions are allowed, and the criteria for distinguishing genuine variants from artifacts. This documentation enables reproducibility and provides a basis for interpreting the confidence of individual calls.

### Checkpoint 4: Determine Whether Isoform Resolution Is Required

The need to resolve isoform structure influences the choice between short-read and long-read RNA sequencing. Short-read RNA-seq provides sufficient information for most fusion detection and for read-depth analysis of highly expressed genes. However, short reads cannot resolve complex isoform structures or phase variants across long distances within a transcript.

Long-read RNA sequencing provides full-length transcript information that enables direct observation of isoform structure and variant phasing. The scTaILoR-seq method demonstrated that targeted long-read approaches can achieve high on-target yield, improving the median number of on-target transcripts per cell by 29-fold. This capability is valuable when the biological question concerns how structural variants affect transcript structure, such as whether a deletion removes specific exons from all isoforms or only from a subset.

The decision framework should include a question about whether isoform resolution is essential for the biological question. If the goal is simply to identify the presence of a fusion or copy number change, short-read sequencing is sufficient and more cost-effective. If the goal is to understand how a variant alters transcript structure or to phase variants across transcripts, long-read sequencing is warranted despite the higher cost and lower throughput.

### Checkpoint 5: Plan the Validation Strategy Before Data Collection

Validation is not an afterthought but a planned component of the study design. The validation strategy should be specified before sequencing begins because it affects sample requirements and data collection decisions. DNA-based confirmation using whole-genome sequencing, whole-exome sequencing, or targeted breakpoint PCR is the standard approach for confirming RNA-seq SV calls.

The Open Pediatric Cancer Project demonstrated that multi-platform validation is feasible at scale, integrating structural variants and fusions called from whole-genome sequencing and RNA-seq across more than 6000 pediatric cancer patients. The project released reproducible dockerized workflows and versioned processed data, providing a model for how multi-platform integration can be operationalized.

The validation plan should specify which variants will be selected for confirmation, the method of confirmation, and the criteria for considering a variant confirmed. A common approach is to prioritize variants that are clinically actionable, that affect genes of known biological significance, or that have strong read support but uncertain annotation status. The plan should also specify how discordant results between RNA-seq and DNA-based methods will be resolved.

### Checkpoint 6: Establish Quality Thresholds and Documentation Standards

Quality thresholds should be established before analysis begins, not derived from the data after the fact. This prevents the common failure pattern of adjusting thresholds until the results look reasonable, which inflates false positive rates and undermines reproducibility.

Key quality metrics to track include RNA integrity scores, mapping rates, exonic mapping rates, coverage uniformity across transcripts, and the number of chimeric alignments. Samples with poor RNA integrity should be flagged for cautious interpretation because degraded RNA produces reads biased toward the 5-prime end of transcripts, distorting read-depth signals. Low exonic mapping rates suggest genomic DNA contamination or poor RNA quality. An unusually high number of chimeric alignments may indicate library preparation artifacts or alignment errors.

Documentation standards should follow the principles emphasized in the nf-core documentation, which describes community pipelines that follow best practices for reproducibility and configuration. Each analysis run should record the software version, command-line parameters, reference genome version, and input file checksums. The Carpentries lessons provide foundational training in Git and version control that is directly applicable to documenting analysis workflows.

### Checkpoint 7: Interpret Results Within the Boundaries of RNA-Seq Evidence

The final checkpoint is interpretation. RNA-seq SV calls must be interpreted within the boundaries of what transcriptome data can and cannot demonstrate. RNA-seq provides evidence about transcript structure but does not directly reveal genomic breakpoints. A fusion transcript may result from a genomic rearrangement, but the exact genomic coordinates require DNA-level analysis. A deletion that removes an exon may be detected through reduced coverage, but the genomic extent of the deletion requires DNA sequencing to define.

Changes in transcript abundance can result from copy number changes, transcriptional regulation, or RNA stability. Read-depth analysis of RNA-seq cannot distinguish these mechanisms. A gene with reduced expression due to transcriptional repression will show the same coverage pattern as a gene with a heterozygous deletion. This confounding limits the interpretation of read-depth based SV calls from RNA-seq.

The interpretation should also account for the expression-dependent detection bias. Failure to detect an SV in RNA-seq does not mean the SV is absent, only that it is not detectable in the transcriptome of the sampled tissue. This limitation should be communicated clearly in publications and clinical reports.

### Implementing the Framework in Practice

The decision framework can be implemented as a structured document that is completed before data collection and updated as the analysis progresses. The document should record the biological question, the sample type and expression evidence, the variant classes of interest, the discovery-precision tradeoff, the isoform resolution requirement, the validation strategy, the quality thresholds, and the interpretation boundaries.

The Galaxy Training Network provides accessible tutorials for RNA-seq analysis that can help researchers build the foundational skills needed to implement this framework. The EMBL-EBI Training portal offers learning pathways for bioinformatics data resources and practical analysis. The Bioconductor documentation provides official guidance for R packages used in genomic analysis. These resources support the skill development needed to execute each checkpoint effectively.

The framework is not a rigid protocol but a decision aid that helps researchers make explicit choices that are often left implicit. By documenting each decision, researchers create a record that supports reproducibility, facilitates troubleshooting, and provides a basis for interpreting the confidence of individual variant calls. This documentation is particularly important when RNA-seq SV findings are used for clinical decision-making, where the analytical and interpretive decisions must be transparent and defensible.

## Frequently Asked Questions

### What types of structural variants can be detected from RNA-seq data?

RNA-seq can detect deletions, duplications, inversions, and translocations that affect expressed transcripts. Split-read analysis identifies breakpoints at exon-exon junctions, read-depth analysis reveals copy number changes across expressed exons, and discordant paired-end mapping detects rearrangements. The detection is limited to regions that are transcribed in the sampled tissue, and the sensitivity depends on expression level.

### How does splicing interfere with structural variant detection from RNA-seq?

Splicing creates split reads that span exon-exon junctions, which is the same signal used to detect structural variant breakpoints. The computational challenge is distinguishing true breakpoints from normal splicing events. Fusion detection tools address this by evaluating whether the two sides of a split read come from different genes or distant regions of the same gene, and by requiring multiple supporting reads.

### Why is expression level important for RNA-seq SV detection?

RNA-seq coverage is proportional to transcript abundance, so genes with low expression have sparse coverage that limits detection sensitivity. A gene expressed at very low levels may have no reads covering a breakpoint, making the SV invisible to split-read analysis. Read-depth analysis similarly requires sufficient coverage to distinguish copy number changes from noise.

### What is the difference between RNA-seq and DNA-seq for SV detection?

DNA-seq provides uniform coverage across the genome and can detect SVs in coding and non-coding regions. RNA-seq only samples expressed transcripts, so it cannot detect SVs in non-expressed genes or regulatory regions. RNA-seq provides information about the functional consequences of SVs on transcript structure, which DNA-seq cannot provide directly.

### How should RNA-seq SV calls be validated?

RNA-seq SV calls should be validated with DNA-based methods before clinical use. Whole-genome sequencing provides the most comprehensive confirmation, while targeted breakpoint PCR can confirm specific variants. The validation approach depends on the variant class and the intended use of the result.

### What are the main sources of false positive SV calls from RNA-seq?

Read-through transcription between adjacent genes is a common source of false fusion calls. Alignment errors in repetitive regions create spurious split reads. PCR amplification bias during library preparation creates coverage artifacts that mimic copy number changes. Filtering strategies should address each of these sources.

### Can long-read RNA-seq improve SV detection compared to short-read?

Long-read RNA-seq provides full-length transcript information that can resolve isoform structure and phase variants across transcripts. This enables detection of complex SVs that affect transcript structure in ways that short reads cannot resolve. The tradeoff is lower throughput and higher cost, which limits the number of samples that can be analyzed.

### When should RNA-seq be used for SV detection instead of DNA-seq?

RNA-seq should be used when the biological question concerns expressed structural variants, particularly fusions that create chimeric transcripts. RNA-seq is also valuable for understanding how SVs affect transcript structure and expression. For comprehensive genome-wide SV detection, DNA-seq is the appropriate primary approach, with RNA-seq as a complementary method.

## Related Bioinformatics Guides

- [RNA-Seq vs ChIP-Seq: Complementary Approaches for Gene Regulation](/knowledge/bioinformatics/rna-seq-vs-chip-seq-complementary-approaches-for-gene-regulation)
- [RNA-Seq vs Microarray: Choosing the Right Gene Expression Profiling Platform](/knowledge/bioinformatics/rna-seq-vs-microarray-choosing-the-right-gene-expression-profiling-platform)
- [Detecting Structural Variants with Long-Read Sequencing: Methods and Considerations](/knowledge/bioinformatics/detecting-structural-variants-with-long-read-sequencing-methods-and-considerations)
- [RNA-Seq Databases: Accessing and Using Public RNA-Seq Data](/knowledge/bioinformatics/rna-seq-databases-accessing-and-using-public-rna-seq-data)
- [RNA-Seq Data Analysis in Galaxy: A User-Friendly Platform](/knowledge/bioinformatics/rna-seq-data-analysis-in-galaxy-a-user-friendly-platform)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)


## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [Diagnostic genome sequencing improves diagnostic yield: a prospective single-centre study in 1000 patients with inherited eye diseases.](https://pubmed.ncbi.nlm.nih.gov/37734845). Journal of medical genetics, 2024.
- [Single-cell long-read targeted sequencing reveals transcriptional variation in ovarian cancer.](https://pubmed.ncbi.nlm.nih.gov/39134520). Nature communications, 2024.
- [Structural variant calling: the long and the short of it.](https://pubmed.ncbi.nlm.nih.gov/31747936). Genome biology, 2019.
- [The Open Pediatric Cancer Project.](https://pubmed.ncbi.nlm.nih.gov/40891528). GigaScience, 2025.
- [SNP calling, haplotype phasing and allele-specific analysis with long RNA-seq reads.](https://pubmed.ncbi.nlm.nih.gov/41912802). Nature methods, 2026.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.