How to Visualize Alternative Splicing Events: A Guide to Sashimi Plots and Other Visualization Tools for RNA-seq
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Sashimi plots visualize RNA-seq junction reads, with arc thickness proportional to read count, enabling direct comparison of alternative splicing events across conditions.
- Accurate Sashimi plot generation necessitates high-quality RNA-seq alignments, validated by metrics like mapping rate and exonic read proportion, and requires appropriate GTF/GFF annotations for genomic context.
- Tools like IGV offer interactive exploration, while ggsashimi (R package) and Gviz (Bioconductor) provide programmatic control for generating publication-quality, customizable figures.
- Data preparation involves quality control of raw reads (e.g., FastQC), proper BAM file indexing (
samtools index), and selection of gene regions guided by differential splicing analysis (e.g., rMATS, MISO). - Common pitfalls include misalignment in repetitive regions, annotation mismatches, insufficient sequencing depth leading to false negatives, and overinterpretation of minor isoforms without validation.
Alternative splicing is a central mechanism through which a single gene produces multiple mRNA isoforms, and visualizing these events is essential for interpreting RNA-seq data in both basic research and clinical contexts. This article provides a practical workflow for generating Sashimi plots and other splicing visualizations using established bioinformatics tools, with attention to data preparation, customization, interpretation, and common pitfalls. The guidance is intended for biology students, researchers, laboratory professionals, and life-science practitioners who need publication-quality figures that accurately represent splicing events from their RNA-seq experiments.
Understanding Alternative Splicing and Why Visualization Matters
Alternative splicing generates transcriptomic diversity by differentially including or excluding exons, retaining introns, or selecting alternative splice sites. The functional consequences of these events are substantial. Research in C. elegans using RNA-seq and bioinformatic analysis has categorized non-triplet alternative splicing into three classes: NMD-sensitive isoforms, alternative C-terminal length isoforms, and dual-coding isoforms, with hundreds of events identified across these categories. Analysis of human transcriptomes reveals broadly similar patterns and distributions, indicating that non-triplet splicing is a large but underappreciated class of alternative splicing that regulates gene expression and generates protein-coding diversity (Pervasive non-triplet alternative splicing drives functional isoform diversity).
The clinical relevance of splicing is equally important. Germline genetic variants that impact splicing are a frequent cause of disease, yet predicting the nature and abundance of aberrant splicing remains challenging. Paired DNA and RNA testing has revealed unexpected splice events that alter variant interpretation, including variants at consensus donor nucleotide positions lacking a splice impact, mid-exonic missense variants creating novel donor sites, and variants causing splicing impacts through pyrimidine tract optimization (A Plot Twist: When RNA Yields Unexpected Findings in Paired DNA-RNA Germline Genetic Testing). These findings underscore why researchers need robust visualization methods to detect and present splicing alterations accurately.
Visualization serves multiple purposes in splicing analysis. It allows researchers to confirm that a splicing event detected by computational tools is real and not an artifact of alignment or quantification. It provides a means to compare splicing patterns across conditions, such as control versus treatment or wild-type versus knockdown. It also enables the communication of complex splicing biology to collaborators, reviewers, and clinical audiences. A well-constructed Sashimi plot can reveal the structure of isoforms, the relative abundance of junction reads, and the differences between conditions in a single figure.
Core Principles of Splicing Visualization
The Relationship Between Read Alignment and Splicing Inference
RNA-seq reads that span exon-exon junctions are the primary evidence for splicing events. When reads align to the genome, those that cross a splice junction are split into two or more segments, with the alignment software recording the genomic coordinates of each segment. The number of reads supporting a particular junction is a quantitative measure of how frequently that splice event occurs in the sample. Visualization tools display these junction reads as arcs or curves connecting the exons they span, with the thickness or height of the arc proportional to the number of supporting reads.
The accuracy of splicing inference depends on the quality of the underlying alignment. Spliced aligners such as STAR or HISAT2 are designed to handle reads that cross splice junctions, but they require appropriate parameters and a reference genome or transcriptome. Errors in alignment can produce false junction calls, particularly in regions of high sequence similarity or repetitive content. Before generating any visualization, researchers should verify that their alignments meet quality standards, including mapping rates, the distribution of reads across genes, and the proportion of reads mapping to exonic versus intronic regions.
The Role of Annotations in Splicing Visualization
Gene annotations provide the genomic coordinates of exons, introns, and transcripts, which are necessary for interpreting splicing events in context. Most visualization tools accept annotations in GTF or GFF format, typically derived from Ensembl, UCSC, or NCBI. The choice of annotation version can affect the appearance of Sashimi plots, as different annotation builds may include or exclude particular isoforms. Researchers should document the annotation version used in their analysis and consider whether the annotation adequately represents the transcripts relevant to their study.
Annotations also help distinguish known splicing events from novel ones. A junction that is present in the annotation is likely a constitutive or annotated alternative event, while a junction that is absent may represent a novel splicing event. This distinction is important for interpretation, as novel junctions may be biologically meaningful or may result from alignment artifacts. Visualization tools can color or label junctions differently based on whether they are annotated, facilitating this assessment.
Quantitative Interpretation of Splicing Events
Sashimi plots provide a visual representation of junction read counts, but they do not directly quantify splicing ratios or percent-spliced-in (PSI) values. PSI is the proportion of transcripts that include a particular exon or splice site, and it is the standard metric for comparing splicing between conditions. Tools such as rMATS, MISO, or SUPPA2 calculate PSI values and statistical significance for differential splicing. The Sashimi plot complements these quantitative analyses by showing the actual read support for the junctions that drive the PSI calculation.
When interpreting Sashimi plots, researchers should consider the depth of coverage at the relevant exons and junctions. Low coverage can make a splicing event appear absent when it is merely underdetected. Conversely, high coverage can reveal minor isoforms that may or may not be biologically significant. The visualization should be accompanied by the quantitative metrics from splicing analysis tools to provide a complete picture.
At a Glance: Sashimi Plot Tools and Their Characteristics
The following table summarizes the primary tools for generating Sashimi plots, their input requirements, and their key features. This comparison helps researchers select the appropriate tool for their specific needs.
| Tool | Input Requirements | Key Features | Best For |
|---|---|---|---|
| IGV (Integrative Genomics Viewer) | BAM files, index files, reference genome, optional GTF annotation | Interactive browsing, built-in Sashimi plot generation, no programming required, export as SVG or PNG | Quick exploration, single-sample or small-scale comparison, researchers unfamiliar with command-line tools |
| ggsashimi (R package) | BAM files, GTF annotation, tab-delimited coordinates file | Highly customizable publication-quality plots, supports multiple conditions, color and transparency controls, junction count labels | Producing final figures for manuscripts, complex multi-sample comparisons, users comfortable with R |
| Gviz (Bioconductor) | BAM files, GTF annotation, R environment | Programmatic control over all plot elements, integration with other Bioconductor packages, supports multiple tracks | Automated figure generation within analysis pipelines, users already working in R/Bioconductor |
Bioconductor provides official documentation for R-based genomic analysis packages, including those used for splicing visualization and reproducible workflows. The Galaxy Training Network offers accessible tutorials for RNA-seq analysis and visualization that can be run without command-line expertise, making these tools available to a broader range of researchers.
Preparing Your Data for Splicing Visualization
Quality Control of RNA-seq Alignments
Before generating Sashimi plots, researchers must ensure that their BAM files are properly aligned and quality controlled. The RNA-seq workflow begins with raw sequencing reads, which undergo quality assessment, trimming, and alignment to a reference genome. Each step affects the reliability of downstream splicing visualization.
Quality control of raw reads typically involves checking per-base quality scores, GC content, adapter contamination, and duplication rates. Tools such as FastQC provide these metrics, and the EMBL-EBI Training resources offer practical guidance on interpreting quality reports and making decisions about trimming and filtering. Reads with poor quality at their ends may align incorrectly or fail to span splice junctions, reducing the sensitivity of splicing detection.
After alignment, additional quality metrics should be examined. These include the overall mapping rate, the proportion of reads mapping to exonic regions, the number of reads spanning splice junctions, and the consistency of coverage across genes. The nf-core documentation describes community standards for RNA-seq pipeline outputs, including quality control metrics that should be reviewed before proceeding to downstream analysis. If alignment quality is poor, the resulting Sashimi plots will be unreliable, and the underlying data issues should be addressed before visualization.
Indexing and Processing BAM Files
Sashimi plot tools require indexed BAM files to efficiently retrieve reads from genomic regions of interest. Indexing is performed with samtools index, which creates a .bai file that allows random access to the alignment data. The BAM file must be sorted by genomic coordinates before indexing, which is typically done during the alignment or with samtools sort.
For experiments with multiple samples or conditions, researchers may need to merge or compare BAM files. Some tools, such as ggsashimi, accept multiple BAM files and display them as separate tracks, allowing direct visual comparison of splicing between conditions. In this case, each BAM file must be individually indexed, and the tool will use the read counts from each file to draw the corresponding junction arcs.
Coverage normalization is an important consideration when comparing samples with different sequencing depths. Sashimi plots can display raw read counts or normalized values, and the choice affects interpretation. If one sample has twice the sequencing depth of another, the junction arcs will appear thicker even if the splicing ratio is identical. Some tools allow normalization by library size or by the coverage of a reference gene, and researchers should apply appropriate normalization when comparing across samples.
Selecting Genomic Regions for Visualization
Sashimi plots are typically generated for specific genes or genomic regions of interest instead of for entire chromosomes. The selection of regions is guided by the results of differential splicing analysis, which identifies genes with statistically significant splicing changes between conditions. Researchers should prioritize genes with strong evidence of differential splicing, as indicated by adjusted p-values and effect sizes, and should examine the specific exons or junctions that drive the signal.
The genomic coordinates for visualization can be obtained from the splicing analysis output or from genome browsers such as NCBI. The NCBI Data Resources provide search systems and sequence resources that allow researchers to locate genes and obtain their genomic coordinates. When defining the region for a Sashimi plot, researchers should include sufficient flanking sequence to show the full context of the splicing event, including the exons that are alternatively included or skipped.
Generating Sashimi Plots with IGV
Installing and Loading Data in IGV
The Integrative Genomics Viewer (IGV) is a widely used genome browser that includes built-in Sashimi plot functionality. IGV is available as a desktop application and does not require programming skills, making it accessible to researchers at all levels of bioinformatics experience. The tool accepts BAM files, alignment indexes, and annotation files, and it displays coverage tracks and junction arcs in an interactive interface.
To begin, researchers load their indexed BAM files into IGV, along with the reference genome and an annotation file in GTF or GFF format. IGV can load annotations from local files or from remote sources such as Ensembl or UCSC. Once the data is loaded, researchers navigate to the gene or region of interest by entering the gene name or genomic coordinates in the search box.
The Sashimi plot is generated by right-clicking on a BAM track and selecting the Sashimi Plot option. IGV will display the coverage profile for the region and draw arcs representing splice junctions, with the height of each arc proportional to the number of reads supporting that junction. The plot can be customized by adjusting the minimum number of reads required to display a junction, the color of the arcs, and the display of junction read counts.
Customizing IGV Sashimi Plots
IGV provides several options for customizing Sashimi plots to meet publication requirements. The minimum junction read count threshold determines which junctions are displayed, allowing researchers to filter out low-support junctions that may represent noise. Setting this threshold too high may hide genuine splicing events in low-coverage regions, while setting it too low may produce cluttered plots with many spurious arcs.
The plot can be exported as an image file in SVG or PNG format. SVG is preferred for publication because it is a vector format that scales without loss of quality, allowing editors and reviewers to zoom into details. Researchers should ensure that the exported image includes all relevant labels, such as the gene name, genomic coordinates, and the sample or condition represented by each track.
IGV also allows the display of multiple BAM tracks simultaneously, with each track representing a different sample or condition. This feature enables direct visual comparison of splicing between conditions, as the junction arcs from each sample are drawn in separate panels. The color of each track can be customized to distinguish conditions, and the coverage tracks provide context for the junction arcs.
Limitations of IGV for Publication Figures
While IGV is excellent for interactive exploration, its Sashimi plots have limitations for final publication figures. The default styling may not match journal requirements, and the level of customization is limited compared to programmatic tools. The plot layout is determined by the IGV interface, and researchers have limited control over the placement of labels, the size of panels, and the overall composition of the figure.
For these reasons, many researchers use IGV for initial exploration and then generate final figures with ggsashimi or Gviz, which offer greater control over plot aesthetics. The IGV Sashimi plot serves as a quick check that the splicing event is visually supported by the data, while the programmatic tools produce the polished figure for the manuscript.
Generating Publication-Quality Sashimi Plots with ggsashimi
Installation and Setup
ggsashimi is an R package that generates Sashimi plots using the ggplot2 graphics system, providing extensive customization options for publication-quality figures. The package is available from GitHub and can be installed using the devtools or remotes package in R. Installation requires R and the dependencies specified in the package documentation.
The Bioconductor project provides official documentation for R-based genomic analysis packages and workflows, and researchers working in R will find resources for managing packages and ensuring reproducible analyses. The The Carpentries Lessons offer foundational training in R programming, data manipulation, and reproducible research practices that are useful for researchers new to programmatic figure generation.
Input File Preparation for ggsashimi
ggsashimi requires three main inputs: BAM files, a GTF annotation file, and a coordinates file specifying the genomic regions to plot. The coordinates file is a tab-delimited text file with columns for the chromosome, start position, end position, and a name for each region. Multiple regions can be specified in a single file, and ggsashimi will generate a separate plot for each.
The BAM files must be indexed and sorted, as described previously. ggsashimi accepts multiple BAM files and can group them by condition, with each group displayed as a separate track. The grouping is specified in the command line or in a configuration file, allowing flexible comparisons between conditions.
The GTF annotation file provides the exon structure for the plotted region. ggsashimi uses the annotation to draw the exon blocks at the bottom of the plot, with the junction arcs connecting the exons. The annotation should match the reference genome version used for alignment to ensure consistency between the read alignments and the displayed gene structure.
Customizing ggsashimi Plots
ggsashimi offers extensive customization options that allow researchers to produce figures tailored to their specific needs. The color of each condition track can be specified, and the transparency of the junction arcs can be adjusted to reduce visual clutter when many junctions are present. The minimum read count for displaying a junction can be set, and the font size of junction count labels can be controlled.
The plot can be customized to show or hide the coverage track, to adjust the height of the coverage and junction panels, and to modify the appearance of the exon blocks. Researchers can also add a legend to distinguish conditions and can adjust the overall dimensions of the output image to fit journal specifications.
The output format is controlled by the graphics device, with PDF and PNG being the most common choices for publication. PDF is preferred for vector graphics, while PNG is suitable for raster images at high resolution. Researchers should generate figures at the resolution and dimensions required by their target journal, typically 300 dpi or higher for raster images.
Interpreting ggsashimi Output
The ggsashimi plot displays the coverage profile for each condition at the top, with the junction arcs below and the gene annotation at the bottom. The height of each arc represents the number of reads supporting that junction, and the number is displayed above the arc. Comparing the arcs between conditions reveals differences in splicing, such as increased inclusion of a particular exon or the appearance of a novel junction in one condition.
Researchers should examine the plot for consistency with the quantitative splicing analysis. If a gene is identified as differentially spliced with a particular exon skipped in the treatment condition, the Sashimi plot should show reduced junction reads spanning that exon in the treatment track. Discrepancies between the quantitative analysis and the visualization may indicate issues with the analysis, such as misalignment or incorrect annotation, and should be investigated before proceeding.
Using Gviz for Programmatic Splicing Visualization
Integration with Bioconductor Workflows
Gviz is a Bioconductor package that provides programmatic control over genomic visualization, including the ability to create Sashimi-like plots within larger analysis workflows. The package is designed to work with other Bioconductor packages, allowing researchers to integrate visualization into their existing RNA-seq analysis pipelines. The Bioconductor website provides official documentation for package installation, usage, and workflow integration.
Gviz uses a track-based system, where each track represents a different type of genomic data. For splicing visualization, researchers create a track for the gene annotation, tracks for the coverage of each sample, and tracks for the junction reads. The tracks are combined into a single plot, with the layout controlled by the researcher.
Creating Sashimi-Style Plots with Gviz
To create a Sashimi-style plot with Gviz, researchers first load the BAM files and annotation into R. The AlignmentsTrack function creates a track from a BAM file, displaying the coverage and read alignments for a specified genomic region. The GeneRegionTrack function creates a track from a GTF annotation, displaying the exon-intron structure of genes in the region.
Junction reads can be displayed using the SashimiTrack function or by customizing the alignment track to show split reads. The appearance of each track can be customized, including colors, heights, and labels. The tracks are combined using the plotTracks function, which arranges them in a vertical stack.
Gviz provides greater flexibility than IGV for integrating splicing visualization into automated workflows. Researchers can write R scripts that generate plots for multiple genes or regions in a loop, producing consistent figures for all differentially spliced genes identified in the analysis. This automation is valuable for generating supplementary figures for manuscripts or for creating standardized reports for large datasets.
Reproducibility Considerations
Programmatic visualization with Gviz or ggsashimi supports reproducibility by encoding the figure generation in scripts that can be rerun with updated data or parameters. The nf-core documentation emphasizes the importance of reproducible workflows in bioinformatics, and the The Carpentries Lessons provide training in version control and reproducible research practices that complement programmatic visualization.
Researchers should document the versions of all software and packages used for visualization, including R, Bioconductor, and the specific visualization packages. This documentation ensures that figures can be regenerated with the same parameters and that the analysis is transparent to reviewers and collaborators.
Comparing Sashimi Plots with Other Splicing Visualization Approaches
Coverage Plots and Exon Usage Tracks
Sashimi plots are one of several approaches to visualizing splicing events. Coverage plots display the read depth across a genomic region, with exons showing higher coverage than introns. While coverage plots can reveal differences in exon usage between conditions, they do not directly show the junction reads that connect exons, making it difficult to distinguish between different isoforms that share exons.
Exon usage tracks, such as those generated by DEXSeq or similar tools, display the relative usage of individual exons across conditions. These plots are useful for identifying exons that are differentially included or excluded, but they do not show the connectivity between exons. Sashimi plots complement these approaches by showing the actual junction reads that define isoform structure.
Transcript Structure Diagrams
Transcript structure diagrams, such as those generated by Ensembl or UCSC, display the exon-intron structure of annotated isoforms. These diagrams are useful for understanding the known isoforms of a gene and for identifying which exons are alternative. However, they do not show the expression level of each isoform or the splicing changes between conditions.
Sashimi plots provide quantitative information that transcript structure diagrams lack, showing the read support for each junction and enabling comparison between conditions. For a complete picture, researchers may combine a transcript structure diagram with a Sashimi plot, using the diagram to identify the isoforms and the Sashimi plot to show their expression.
Heatmaps and Clustering of Splicing Events
For large-scale analyses, heatmaps and clustering can summarize splicing patterns across many genes and samples. Tools such as PSI values across samples can be displayed as heatmaps, with rows representing splicing events and columns representing samples. Clustering can reveal groups of co-regulated splicing events or samples with similar splicing profiles.
These approaches are useful for identifying patterns in large datasets, but they lack the detail of Sashimi plots for individual events. Researchers typically use heatmaps for overview and Sashimi plots for detailed examination of specific genes of interest.
Practical Workflow for Splicing Visualization
Step 1: Run Differential Splicing Analysis
The first step in the workflow is to identify genes with differential splicing between conditions using a dedicated tool such as rMATS, MISO, or SUPPA2. These tools take aligned BAM files or count tables as input and produce lists of splicing events with statistical significance and effect sizes. The output includes the genomic coordinates of the events, the PSI values for each condition, and the read counts supporting inclusion and exclusion.
Researchers should review the results to identify the most significant and biologically relevant events. The number of events to visualize depends on the scope of the study, but a typical manuscript may include Sashimi plots for several representative genes, with additional plots in supplementary materials.
Step 2: Select Genes and Regions for Visualization
Based on the differential splicing results, researchers select genes for visualization. The selection should prioritize events with strong statistical support, large effect sizes, and biological relevance to the study question. For each gene, the genomic region for the Sashimi plot should encompass the alternatively spliced exons and sufficient flanking sequence to show the full context.
The coordinates for visualization can be obtained from the splicing analysis output or from genome browsers. The NCBI Data Resources provide search systems for locating genes and obtaining genomic coordinates, and the EMBL-EBI Training resources offer guidance on navigating genomic databases.
Step 3: Generate Initial Plots with IGV
For initial exploration, researchers should load the BAM files into IGV and generate Sashimi plots for the selected regions. This step allows quick visual confirmation that the splicing event is supported by the data and that the region is correctly specified. Researchers can adjust the minimum read threshold and examine the junction arcs to ensure they match the quantitative analysis.
The IGV exploration also helps identify potential issues, such as low coverage in the region, alignment artifacts, or annotation discrepancies. If problems are detected, they should be addressed before generating final figures.
Step 4: Generate Publication Figures with ggsashimi or Gviz
Once the regions are confirmed, researchers generate publication-quality figures using ggsashimi or Gviz. The figures should be customized to meet journal requirements, including appropriate colors, font sizes, and image dimensions. The plots should clearly show the splicing differences between conditions, with junction read counts labeled and conditions distinguished by color.
Researchers should generate figures for all selected genes and review them for consistency with the quantitative analysis. Any discrepancies should be investigated, and the figures should be revised as needed.
Step 5: Document the Visualization Parameters
For reproducibility, researchers should document all parameters used for visualization, including the software versions, the minimum read thresholds, the normalization method, and the annotation version. This documentation should be included in the methods section of the manuscript or in a supplementary file.
The nf-core documentation provides guidance on documenting bioinformatics workflows, and the The Carpentries Lessons offer training in reproducible research practices that support thorough documentation.
Records and Measurements for Splicing Visualization
Tracking Junction Read Counts
The primary measurement in Sashimi plots is the number of reads supporting each splice junction. These counts should be recorded for each junction in each condition, along with the total number of reads spanning the region. The junction counts provide the quantitative basis for comparing splicing between conditions and for calculating PSI values.
Researchers should maintain a record of the junction counts for all visualized events, including the genomic coordinates of each junction, the number of supporting reads in each condition, and the PSI values calculated from these counts. This record allows verification of the visualization against the quantitative analysis and provides the data needed for supplementary tables.
Documenting Coverage Depth
Coverage depth is an important contextual measurement for interpreting Sashimi plots. Low coverage can make splicing events difficult to detect, while high coverage provides confidence in the observed junctions. Researchers should record the average coverage across the visualized region for each condition, as well as the coverage at the specific exons involved in the alternative splicing event.
The coverage information helps readers assess the reliability of the Sashimi plot. If a junction appears absent in one condition, the coverage at that region should be sufficient to detect it if it were present. Reporting coverage depth alongside Sashimi plots strengthens the evidence for differential splicing.
Recording Analysis Parameters
All parameters used in the analysis and visualization should be recorded, including the alignment software and version, the reference genome and annotation version, the splicing analysis tool and parameters, and the visualization tool and settings. This documentation ensures that the analysis can be reproduced and that the results can be interpreted in the context of the specific methods used.
The Galaxy Training Network provides tutorials that emphasize the importance of recording analysis parameters and maintaining reproducible workflows. Researchers should follow these practices to ensure the integrity of their splicing visualization.
Common Failure Patterns in Splicing Visualization
Misalignment in Repetitive or Homologous Regions
One of the most common problems in splicing visualization is misalignment of reads in repetitive or homologous regions. Reads that map to multiple locations in the genome may be assigned to the wrong location, producing false junction calls or incorrect coverage patterns. This problem is particularly acute in genes with paralogs or in regions with transposable elements.
Researchers should examine the mapping quality of reads in the visualized region and consider filtering low-quality alignments. Tools that report mapping quality scores can help identify reads that are ambiguously mapped. If misalignment is suspected, researchers may need to use more stringent alignment parameters or exclude multi-mapping reads from the analysis.
Annotation Mismatches
Discrepancies between the annotation used for visualization and the actual transcripts expressed in the sample can lead to misinterpretation of Sashimi plots. If the annotation lacks a particular isoform, the Sashimi plot may show junctions that appear novel when they are actually known. Conversely, if the annotation includes isoforms that are not expressed, the plot may suggest splicing events that are not supported by the data.
Researchers should verify that the annotation version is appropriate for their organism and cell type and should consider whether the annotation adequately represents the transcripts relevant to their study. In some cases, it may be necessary to use a custom annotation derived from transcript assembly of the RNA-seq data.
Low Coverage Leading to False Negative Results
Low sequencing depth can result in insufficient reads to detect splicing events, particularly for rare isoforms or in genes with low expression. A Sashimi plot from a low-coverage sample may show no junction reads for an event that is actually present, leading to a false negative conclusion.
Researchers should assess the coverage in the visualized region and consider whether the sequencing depth is adequate for detecting the splicing events of interest. If coverage is insufficient, additional sequencing may be required, or the analysis may need to focus on more highly expressed genes.
Overinterpretation of Minor Isoforms
Sashimi plots can reveal minor isoforms that are present at low levels, and researchers may be tempted to interpret these as biologically significant. However, low-level splicing events may represent noise, splicing errors, or non-functional isoforms. The Pervasive non-triplet alternative splicing drives functional isoform diversity study notes that non-triplet alternative splicing is sometimes considered evidence of splicing errors or noise, although some examples are functionally important.
Researchers should apply appropriate thresholds for junction read counts and should consider the biological context when interpreting minor isoforms. Validation with orthogonal methods, such as RT-PCR, may be necessary to confirm the existence and functional significance of low-abundance splicing events.
Limitations of Sashimi Plots and Splicing Visualization
Inability to Resolve Full-Length Isoforms
Sashimi plots display the read support for individual junctions but do not directly show the full-length isoforms that are expressed. A gene with multiple alternative exons can produce many different isoforms, and the Sashimi plot shows the junctions but not the combinations of junctions that define each isoform. Determining which isoforms are actually expressed requires transcript assembly or long-read sequencing.
Researchers should be cautious about inferring isoform structure from Sashimi plots alone. The plot shows the evidence for individual splicing events, but the combination of events into full-length isoforms requires additional analysis.
Limited Quantitative Precision
The height of the junction arcs in a Sashimi plot is proportional to the number of supporting reads, but the visual representation is not a precise quantitative measure. Small differences in junction read counts may be difficult to discern visually, and the plot does not provide confidence intervals or statistical significance.
Researchers should rely on the quantitative splicing analysis for precise measurements and use the Sashimi plot for visual confirmation and communication. The plot should be accompanied by the PSI values and statistical results from the splicing analysis.
Dependence on Read Length and Sequencing Platform
The ability to detect splicing events depends on the read length and the sequencing platform. Short reads may not span the full length of a junction, particularly for junctions between distant exons. Long-read sequencing can resolve full-length isoforms but has different error profiles and throughput characteristics.
Researchers should consider the limitations of their sequencing data when interpreting Sashimi plots. Junctions that are not detected may be absent or may be undetectable given the read length and sequencing depth.
Safety and Regulatory Context for Splicing Visualization
Clinical Interpretation of Splicing Variants
In clinical settings, splicing visualization can inform the interpretation of genetic variants that may affect splicing. The A Plot Twist: When RNA Yields Unexpected Findings in Paired DNA-RNA Germline Genetic Testing study describes cases where RNA testing revealed unexpected splice events that altered variant interpretation. These findings highlight the complexity of splicing mechanisms and the importance of careful interpretation in clinical contexts.
Researchers and clinicians should be aware that splicing predictions from DNA sequence alone may not accurately reflect the actual splicing outcomes. RNA testing and visualization provide direct evidence of splicing that can complement or correct predictions. However, the interpretation of splicing data in clinical contexts requires expertise and should follow established guidelines.
Data Privacy and Sharing Considerations
RNA-seq data may contain sensitive information, particularly if derived from human subjects. Researchers should follow applicable regulations and institutional policies for data storage, sharing, and publication. The NCBI Data Resources provide repositories for sequence data with controlled access options for sensitive data.
When publishing Sashimi plots, researchers should ensure that the data sharing complies with consent agreements and privacy regulations. In some cases, it may be necessary to aggregate data or remove identifying information before publication.
Reproducibility and Transparency Requirements
Many journals and funding agencies require that bioinformatics analyses be reproducible and transparent. The nf-core documentation describes community standards for reproducible workflows, and the The Carpentries Lessons provide training in reproducible research practices. Researchers should ensure that their splicing visualization is documented sufficiently for others to reproduce the figures.
The Galaxy Training Network offers accessible training in reproducible analysis workflows, and the EMBL-EBI Training provides resources for bioinformatics education. Researchers should take advantage of these resources to ensure their analyses meet reproducibility standards.
Professional Escalation Criteria for Splicing Visualization
When to Seek Bioinformatics Support
Researchers should seek bioinformatics support when they encounter issues that exceed their expertise. These situations include persistent alignment problems, unexpected discrepancies between quantitative analysis and visualization, or the need for advanced customization of figures. Bioinformatics specialists can help troubleshoot issues and ensure that the analysis is technically sound.
The Bioconductor support forum and the Galaxy Training Network provide resources for getting help with bioinformatics problems. Researchers should document their issues clearly and provide the relevant files and parameters when seeking support.
When to Consult Clinical Genetics Experts
In clinical contexts, splicing visualization that informs variant interpretation should be reviewed by clinical genetics experts. The A Plot Twist: When RNA Yields Unexpected Findings in Paired DNA-RNA Germline Genetic Testing study demonstrates that splicing impacts can be complex and unexpected, and expert interpretation is essential for accurate clinical decisions.
Researchers who are not clinical genetics experts should escalate findings with potential clinical implications to qualified professionals. The interpretation of splicing data in the context of genetic testing should follow established guidelines and consider the full clinical picture.
When to Validate with Orthogonal Methods
Splicing events that are central to the conclusions of a study should be validated with orthogonal methods, such as RT-PCR, quantitative PCR, or long-read sequencing. Validation is particularly important for novel splicing events, low-abundance isoforms, or events with potential clinical significance.
The KHSRP has oncogenic functions and regulates the expression and alternative splicing of DNA repair genes in breast cancer MDA-MB-231 cells study exemplifies the use of multiple approaches to characterize splicing regulation. Researchers should consider the level of evidence required for their conclusions and validate accordingly.
Frequently Asked Questions
What is the minimum read count needed to display a junction in a Sashimi plot?
The minimum read count is a user-defined threshold that filters out low-support junctions. A common default is 5 reads, but the appropriate threshold depends on the sequencing depth and the purpose of the visualization. For exploratory analysis, a lower threshold may be used to detect rare splicing events, while for publication figures, a higher threshold may be preferred to reduce clutter and focus on robust junctions. Researchers should report the threshold used and consider the coverage depth when interpreting the plot.
How do I compare splicing between multiple conditions in a Sashimi plot?
Tools like ggsashimi and IGV allow multiple BAM files to be displayed as separate tracks, with each track representing a condition. The junction arcs from each condition are drawn in separate panels, allowing direct visual comparison. The conditions should be distinguished by color, and the plot should include a legend to identify each condition. Normalization for sequencing depth should be applied when comparing conditions with different coverage.
Can Sashimi plots show novel splicing events that are not in the annotation?
Yes, Sashimi plots display junction reads regardless of whether the junction is annotated. Novel junctions will appear as arcs connecting exons that are not connected in the annotation. Researchers should verify that novel junctions are not alignment artifacts and should consider whether they represent biologically meaningful splicing events. Validation with orthogonal methods may be necessary for novel events.
What is the difference between a Sashimi plot and a coverage plot?
A coverage plot displays the read depth across a genomic region, showing which regions are expressed. A Sashimi plot additionally displays the junction reads that connect exons, providing information about splicing. Coverage plots can show differences in exon usage but cannot distinguish between isoforms that share exons. Sashimi plots provide the junction-level evidence needed to understand isoform structure.
How do I choose between IGV, ggsashimi, and Gviz for Sashimi plots?
The choice depends on the researcher's needs and expertise. IGV is interactive and requires no programming, making it suitable for exploration and quick checks. ggsashimi produces highly customizable publication figures and requires R programming skills. Gviz offers programmatic control within Bioconductor workflows and is suitable for automated figure generation. Researchers may use IGV for exploration and ggsashimi or Gviz for final figures.
What normalization should I apply before generating Sashimi plots?
Normalization is important when comparing samples with different sequencing depths. Options include normalizing by library size, by the total number of reads in the region, or by the coverage of a reference gene. The choice of normalization affects the visual comparison of junction read counts between conditions. Researchers should apply appropriate normalization and document the method used.
How do I interpret the numbers displayed above the junction arcs?
The numbers above the junction arcs represent the count of reads supporting that junction. A higher number indicates more reads spanning the junction, suggesting that the splicing event is more frequent. Comparing the numbers between conditions reveals differences in splicing. The counts should be interpreted in the context of the total coverage in the region.
Can Sashimi plots be used for clinical variant interpretation?
Sashimi plots can provide evidence for the splicing impact of genetic variants, but their interpretation in clinical contexts requires expertise. The A Plot Twist: When RNA Yields Unexpected Findings in Paired DNA-RNA Germline Genetic Testing study demonstrates that splicing impacts can be complex and unexpected. Clinical interpretation should be performed by qualified professionals following established guidelines.
Related Bioinformatics Guides
- RNA-Seq Visualization: Volcano Plots, Heatmaps, and PCA
- RNA-Seq Quality Control: Essential Checks and Tools
- Alternative Splicing Analysis from RNA-Seq Data
- RNA-Seq Alignment Tools: STAR, HISAT2, and Beyond
- Volcano Plot Proteomics: How to Create and Interpret Them Effectively
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- Integrative RNA-seq and iRIP-seq analysis links SNRPA overexpression to transcriptomic and splicing alterations in hepatocellular carcinoma cells.. 2026.
- Emerging roles for SNRPB2 in governing the cell cycle and steering tumor immune modulation in breast cancer.. 2026.
- A Plot Twist: When RNA Yields Unexpected Findings in Paired DNA-RNA Germline Genetic Testing.. 2025.
- Pervasive non-triplet alternative splicing drives functional isoform diversity.. 2026.
- Correction of the molecular phenotype of X-linked Dystonia-Parkinsonism reveals a non-canonical function of BRD4.. 2026.
- Author Correction: KHSRP has oncogenic functions and regulates the expression and alternative splicing of DNA repair genes in breast cancer MDA-MB-231 cells (Scientific Reports, (2024), 14, 1, (14694), 10.1038/s41598-024-64687-0). Scientific Reports, 2024.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.