How to Choose the Right Tool for Differential Splicing Analysis: A Decision Guide for RNA-seq Researchers
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- The choice of differential splicing analysis tool is critically dependent on the specific biological question, with event-based tools (e.g., rMATS, SUPPA2) suited for known splicing events, exon-based tools (e.g., DEXSeq, limma) for broader exon/junction usage changes, and isoform-based tools (e.g., NBSplice) for differential transcript usage.
- Tool performance varies significantly with the number of biological replicates; methods borrowing information across genes (e.g., limma) are preferred for small sample sizes (≤3 replicates), while event-based tools (e.g., rMATS, MAJIQ) require adequate replication (≥5 replicates) for reliable estimation of splicing event variability.
- For studies with covariates (e.g., sex, age, clinical attributes), specialized tools like MntJULiP or frameworks like saseR are essential to avoid spurious findings, as standard tools often lack native covariate handling capabilities.
- Detecting novel splicing events necessitates tools that can process de novo identified junctions (e.g., DJExpress) or do not rely on annotations (e.g., LeafCutter), as most event-based tools are limited to annotated gene models.
- Input data requirements differ substantially, with some tools accepting aligned reads (BAM files) for junction counting (e.g., rMATS, MAJIQ), while others require pre-computed transcript quantifications (e.g., SUPPA2) or exon/junction count matrices (e.g., DEXSeq, JunctionSeq).
- Benchmarking studies highlight that no single tool is universally superior, and results can differ dramatically across datasets and sample sizes, suggesting a practical approach of running multiple tools and comparing consensus findings to improve robustness.
Differential splicing analysis asks whether the relative usage of transcript isoforms changes between experimental conditions. The choice of computational tool determines which splicing events you can detect, how you must prepare your input data, and whether your results will hold up under replication. This guide provides a decision framework for researchers who need to select among rMATS, DEXSeq, MISO, SUPPA2, LeafCutter, MAJIQ, and related tools based on experimental design, biological questions, and computational resources.
RNA sequencing has become an indispensable tool for transcriptome-wide analysis of differential gene expression and differential splicing of mRNAs [<a href="#ref-1">1</a>]. As sequencing technologies have developed, so too has the range of analytical approaches available to researchers [<a href="#ref-1">1</a>]. The practical problem is not a shortage of tools but a surplus of options with different assumptions, input requirements, and output formats. A 2023 benchmarking study evaluated 21 different tools on simulated RNA-seq data with known differential splicing events and observed immense discrepancies among them [<a href="#ref-2">2</a>]. The same study found that SUPPA, DARTS, rMATS, and LeafCutter outperformed other event-based tools in their evaluation [<a href="#ref-2">2</a>]. A separate 2020 systematic evaluation of 10 tools found that exon-based methods generally performed better than isoform-based and event-based methods, but that different tools performed strikingly differently across different datasets or numbers of samples [<a href="#ref-3">3</a>].
These findings point to a central reality: there is no single best tool for all differential splicing questions. The right choice depends on your experimental design, the type of splicing events you care about, your tolerance for false positives, and your computational environment. This article walks through the decision process step by step.
At a Glance
The table below summarizes the main tool categories and their practical characteristics. Use it as a starting point before reading the detailed sections that follow.
| Tool Category | Representative Tools | Input Requirements | Event Types Detected | Statistical Approach | Computational Demand |
|---|---|---|---|---|---|
| Exon-based | DEXSeq, limma, edgeR, JunctionSeq | Count data at exon or junction level from aligned reads | Differential exon usage, differential junction usage | Generalized linear models with count distributions | Moderate, runs in R/Bioconductor |
| Event-based | rMATS, SUPPA2, MAJIQ, LeafCutter, MISO | Splice junction counts or transcript quantifications | Exon skipping, alternative 5' and 3' splice sites, intron retention, mutually exclusive exons | Various, including hierarchical models and Bayesian approaches | Variable, some require substantial memory |
| Isoform-based | Cuffdiff2, DiffSplice, NBSplice | Transcript-level quantifications | Differential transcript usage and expression | Statistical tests on estimated isoform abundances | Moderate to high |
| Covariate-aware | MntJULiP, Jutils, saseR | Junction counts with covariate information | Intron-level splicing differences, aberrant splicing | Bayesian mixture models, offset-based approaches | Scalable to thousands of samples |
Understanding the Biological Question Before Choosing a Tool
The first decision is not about software. It is about what you need to learn from your data. Differential splicing analysis can answer several distinct questions, and each question maps to a different class of tools.
If your question is whether specific known splicing events change between conditions, event-based tools such as rMATS or SUPPA2 are appropriate. These tools test for changes in the inclusion level of defined splicing events, such as cassette exons or alternative splice sites [<a href="#ref-4">4</a>]. If your question is broader, such as whether any exon within a gene is used differently between conditions, exon-based tools such as DEXSeq or limma are better suited [<a href="#ref-5">5</a>][<a href="#ref-3">3</a>]. If your question concerns which full-length isoforms change in relative abundance, isoform-based tools such as NBSplice may be necessary [<a href="#ref-6">6</a>].
The distinction matters because tools in different categories do not always agree. The 2020 systematic evaluation found that exon-based methods and two event-based methods, MAJIQ and rMATS, scored well on consistency, reproducibility, precision, recall, and false discovery rate control [<a href="#ref-3">3</a>]. However, the same study noted that the different tools performed strikingly differently across different datasets or numbers of samples [<a href="#ref-3">3</a>]. This means that your choice of tool can change your biological conclusions, beyond your list of significant events.
Consider also whether you need to discover novel splicing events. Most event-based tools are unsuitable for discovering novel splice sites because they rely on annotated gene models [<a href="#ref-2">2</a>]. If novel isoform discovery is a goal, you may need tools that can assemble transcripts from scratch or that accept de novo identified splice junctions. DJExpress, for example, can handle both annotated and de novo identified splice junctions, allowing quantification of novel splice events [<a href="#ref-7">7</a>].
Core Principles of Differential Splicing Analysis
Differential splicing analysis rests on several core principles that apply regardless of which tool you choose. Understanding these principles helps you evaluate whether a tool is appropriate for your data.
Splicing Is Measured Relative to Gene Expression
Differential splicing is distinct from differential gene expression. A gene can be expressed at the same total level between conditions while the proportions of its isoforms change. Conversely, a gene can change in overall expression while its splicing pattern remains constant. Tools for differential splicing must therefore measure the relative usage of exons, junctions, or isoforms, beyond their absolute abundance.
Exon-based tools such as DEXSeq and limma model the counts of individual exons or junctions while accounting for the overall expression level of the gene [<a href="#ref-5">5</a>][<a href="#ref-3">3</a>]. Event-based tools such as rMATS and SUPPA2 compute inclusion levels for specific splicing events, such as the percent spliced in (PSI) value for a cassette exon [<a href="#ref-4">4</a>]. Isoform-based tools estimate the relative expression of each transcript isoform within a gene [<a href="#ref-6">6</a>].
Count Data Requires Appropriate Statistical Models
RNA-seq data consist of counts of reads or fragments mapped to genomic features. These counts follow discrete distributions, typically negative binomial or Poisson, and their variance is related to their mean. Tools that ignore this count nature, or that apply models designed for continuous data without appropriate transformation, risk inflated false positive rates.
The limma package, originally developed for microarray data, has been extended to handle RNA-seq count data through the voom transformation, which estimates the mean-variance relationship and applies linear modeling [<a href="#ref-5">5</a>]. This approach allows limma to perform both differential expression and differential splicing analyses of RNA-seq data [<a href="#ref-5">5</a>]. Other tools use generalized linear models with negative binomial distributions, such as DEXSeq and edgeR [<a href="#ref-3">3</a>].
Sample Size Affects Tool Performance
The number of biological replicates in your experiment affects which tools will perform reliably. The 2020 benchmarking study found that tool performance varied with the number of samples [<a href="#ref-3">3</a>]. Tools that borrow information across genes, such as limma, are designed to overcome the problem of small sample sizes [<a href="#ref-5">5</a>]. However, the optimal tool for a two-versus-two comparison may not be the optimal tool for a ten-versus-ten comparison.
If your experiment includes covariates such as sex, age, ethnicity, or clinical attributes, you need tools that can account for these factors. MntJULiP detects intron-level differences in alternative splicing using a Bayesian mixture model and can process thousands of samples within hours [<a href="#ref-8">8</a>]. The same tool was applied to GTEx brain RNA-seq samples to deconvolute the effects of sex and age at death on splicing patterns [<a href="#ref-8">8</a>]. saseR similarly provides a framework for differential usage and aberrant splicing that scales to large datasets [<a href="#ref-9">9</a>][<a href="#ref-10">10</a>].
Input Data Requirements and Preparation
Each tool category has specific input requirements. Preparing your data correctly is often the most time-consuming part of the analysis and the most common source of errors.
Aligned Reads or Count Matrices
Most differential splicing tools require either aligned reads in BAM format or precomputed count matrices. Exon-based tools such as DEXSeq and limma typically accept count matrices derived from aligned reads using tools such as HTSeq or featureCounts [<a href="#ref-5">5</a>][<a href="#ref-3">3</a>]. Event-based tools such as rMATS and MAJIQ require BAM files aligned to the reference genome, from which they extract junction counts [<a href="#ref-4">4</a>][<a href="#ref-11">11</a>].
The choice between BAM files and count matrices affects your computational requirements. BAM files are large and require substantial storage and memory to process. Count matrices are smaller and faster to analyze but require you to decide in advance which genomic features to count.
Annotations and Reference Genomes
Most tools require a gene annotation file in GTF or GFF format. The quality and completeness of your annotation directly affect your results. Tools that rely on annotated splice junctions cannot detect novel splicing events [<a href="#ref-2">2</a>]. If you are working with a less well-annotated organism, or if you suspect that novel isoforms are biologically important, you may need tools that can assemble transcripts de novo or that accept custom junction definitions.
The NCBI provides official descriptions of databases, search systems, sequence resources, and analysis services that can help you obtain reference genomes and annotations [<a href="#ref-12">12</a>]. The EMBL-EBI Training portal offers learning pathways for bioinformatics data resources and practical analysis education [<a href="#ref-13">13</a>].
Long-Read Sequencing Data
Long-read RNA sequencing technologies, such as Oxford Nanopore Technologies, produce full-length transcripts that can reveal isoform structure directly. These data require different analysis approaches than short-read data. A 2026 study used ONT long-read RNA sequencing on blood samples from 145 individuals and performed differential gene expression, differential transcript expression, differential transcript usage, and alternative splicing analyses using ONT-aware tools such as DRIMSeq, DEXSeq, and stageR [<a href="#ref-14">14</a>]. The study identified extensive isoform-level dysregulation, including novel transcript variants [<a href="#ref-14">14</a>].
If you are using long-read data, you need tools that can handle the characteristics of these data, including higher error rates and different count distributions. saseR provides a single framework for various short- and long-read applications by replacing normalization offsets [<a href="#ref-9">9</a>][<a href="#ref-10">10</a>].
Event-Based Tools: rMATS, SUPPA2, MAJIQ, and LeafCutter
Event-based tools are the most commonly used category for differential splicing analysis. They define specific splicing events, quantify their inclusion levels, and test for differences between conditions.
rMATS
rMATS detects differential splicing of cassette exons, alternative 5' splice sites, alternative 3' splice sites, mutually exclusive exons, and retained introns. It uses a hierarchical model to estimate the inclusion level of each event and test for differences between conditions [<a href="#ref-4">4</a>].
In benchmarking studies, rMATS has performed well. The 2023 study found that rMATS outperformed most other event-based tools [<a href="#ref-2">2</a>]. The 2020 study similarly found that rMATS scored well on precision, recall, and false discovery rate control [<a href="#ref-3">3</a>]. In a targeted RNA long-amplicon sequencing evaluation, rMATS was able to detect all exon-skipping events tested, although it failed to detect other types of events besides exon skipping [<a href="#ref-11">11</a>].
rMATS requires BAM files aligned to the reference genome and a GTF annotation file. It is computationally intensive but can be run on a standard server. The output includes inclusion levels for each event and statistical tests for differential splicing.
SUPPA2
SUPPA2 computes the percent spliced in value for each splicing event from transcript quantifications. It is fast and can handle large datasets [<a href="#ref-4">4</a>]. The 2023 benchmarking study found that SUPPA outperformed most other event-based tools [<a href="#ref-2">2</a>].
SUPPA2 requires transcript-level quantifications, typically from tools such as Salmon or kallisto, instead of BAM files. This makes it computationally efficient and suitable for datasets with many samples. However, because it relies on transcript quantifications, its accuracy depends on the quality of the transcript quantification step.
MAJIQ
MAJIQ detects and quantifies splicing variations using a Bayesian approach. It can detect complex splicing events that other tools miss, including combinations of alternative splice sites [<a href="#ref-4">4</a>]. In the targeted long-amplicon sequencing evaluation, MAJIQ was able to detect all types of splicing events tested, including exon skipping, multiple-exon skipping, alternative 5' splicing, and alternative 3' splicing [<a href="#ref-11">11</a>].
MAJIQ requires BAM files and an annotation file. It is more computationally demanding than SUPPA2 but provides more detailed information about splicing complexity. The 2020 benchmarking study found that MAJIQ scored well on the selected measures [<a href="#ref-3">3</a>].
LeafCutter
LeafCutter focuses on intron usage instead of exon inclusion. It identifies differential splicing by comparing the usage of introns between conditions [<a href="#ref-4">4</a>]. The 2023 benchmarking study found that LeafCutter outperformed most other event-based tools [<a href="#ref-2">2</a>].
LeafCutter requires BAM files and does not require an annotation file, which makes it useful for discovering novel splicing events. However, its output is in terms of intron usage, which may be less intuitive to interpret than exon inclusion levels.
Choosing Among Event-Based Tools
The choice among event-based tools depends on your specific needs. If you are primarily interested in exon skipping, rMATS is a strong choice [<a href="#ref-11">11</a>]. If you need to detect a wide range of event types, including complex combinations, MAJIQ may be better [<a href="#ref-11">11</a>]. If you have many samples and need speed, SUPPA2 is efficient [<a href="#ref-4">4</a>]. If you want to discover novel intron usage patterns without relying on annotations, LeafCutter is appropriate [<a href="#ref-2">2</a>][<a href="#ref-4">4</a>].
A practical approach is to run two tools and compare their results. The 2023 benchmarking study proposed a tool-pair combination approach to improve overall performance [<a href="#ref-2">2</a>]. Running both rMATS and SUPPA2, for example, and focusing on events identified by both tools can reduce false positives.
Exon-Based Tools: DEXSeq, limma, and JunctionSeq
Exon-based tools test whether individual exons or junctions are used differently between conditions, accounting for the overall expression level of the gene. These tools are well suited for detecting differential splicing when you do not have a specific hypothesis about which splicing events are involved.
DEXSeq
DEXSeq models exon counts using a generalized linear model with a negative binomial distribution. It tests whether the relative usage of each exon differs between conditions [<a href="#ref-3">3</a>]. DEXSeq is available through Bioconductor, which provides official package, workflow, installation, and reproducible genomic-analysis documentation [<a href="#ref-15">15</a>].
DEXSeq requires count matrices for each exon, which can be generated from aligned reads using the DEXSeq counting script. It is moderately computationally demanding and can handle complex experimental designs.
limma
The limma package provides an integrated solution for analyzing data from gene expression experiments and can perform both differential expression and differential splicing analyses of RNA-seq data [<a href="#ref-5">5</a>]. It uses the voom transformation to model count data and borrows information across genes to overcome the problem of small sample sizes [<a href="#ref-5">5</a>].
limma is particularly useful for experiments with complex designs or small numbers of replicates. Its facilities for reading, normalizing, and exploring data are well developed [<a href="#ref-5">5</a>]. The package is available through Bioconductor [<a href="#ref-15">15</a>][<a href="#ref-5">5</a>].
JunctionSeq
JunctionSeq tests for differential usage of splice junctions, complementing exon-level analyses. It is an exon-based method that scored well in the 2020 benchmarking study [<a href="#ref-3">3</a>]. JunctionSeq requires count data at the junction level and can detect both annotated and novel junctions.
Advantages and Limitations of Exon-Based Tools
The 2020 systematic evaluation found that exon-based methods generally performed better than isoform-based and event-based methods across the selected measures [<a href="#ref-3">3</a>]. However, exon-based tools do not directly identify which isoforms are affected. They tell you that an exon or junction is used differently, but not which full-length transcripts are involved [<a href="#ref-6">6</a>].
If you need to know the identity of affected isoforms, you may need to combine exon-based analysis with isoform quantification or use an isoform-based tool such as NBSplice [<a href="#ref-6">6</a>].
Isoform-Based Tools: Cuffdiff2, DiffSplice, and NBSplice
Isoform-based tools estimate the abundance of each transcript isoform and test whether the relative expression of isoforms changes between conditions. These tools provide the most direct information about which isoforms are affected but are generally less accurate than exon-based or event-based tools [<a href="#ref-3">3</a>].
Cuffdiff2 and DiffSplice
Cuffdiff2 and DiffSplice are isoform-based tools that were included in the 2020 benchmarking study [<a href="#ref-3">3</a>]. The study found that isoform-based methods performed generally worse than exon-based and event-based methods [<a href="#ref-3">3</a>]. These tools require transcript-level quantifications and are computationally demanding.
NBSplice
NBSplice is an R package for differential splicing analysis using isoform expression data [<a href="#ref-6">6</a>]. It estimates differences in relative expression of gene transcripts between experimental conditions to infer changes in alternative splicing patterns [<a href="#ref-6">6</a>]. In an evaluation using a synthetic RNA-seq dataset with controlled differential splicing, NBSplice accurately predicted differential splicing occurrence, outperforming current methods in terms of accuracy, sensitivity, F-score, and false discovery rate control [<a href="#ref-6">6</a>].
NBSplice was also applied to a real cancer dataset, revealing new differentially spliced genes that could be studied for colorectal cancer biomarker discovery [<a href="#ref-6">6</a>]. The package is available through Bioconductor [<a href="#ref-15">15</a>][<a href="#ref-6">6</a>].
When to Use Isoform-Based Tools
Isoform-based tools are appropriate when the biological question concerns specific isoforms, such as which protein isoforms are produced from a gene. However, because isoform quantification from short-read data is inherently uncertain, results from isoform-based tools should be interpreted with caution and validated experimentally.
Tools for Covariates and Large Datasets
Emerging large and complex RNA-seq datasets from disease and population studies include multiple confounders such as sex, age, ethnicity, and clinical attributes [<a href="#ref-8">8</a>]. Few methods are equipped to handle these challenges [<a href="#ref-8">8</a>]. If your study includes covariates, you need tools that can account for them.
MntJULiP and Jutils
MntJULiP detects intron-level differences in alternative splicing using a Bayesian mixture model and takes covariates into account [<a href="#ref-8">8</a>]. Jutils visualizes alternative splicing variation with heatmaps, PCA, sashimi plots, and Venn diagrams [<a href="#ref-8">8</a>]. These tools are scalable and can process thousands of samples within hours [<a href="#ref-8">8</a>].
The developers applied these methods to GTEx brain RNA-seq samples to deconvolute the effects of sex and age at death on splicing patterns [<a href="#ref-8">8</a>]. Clustering of covariate-adjusted data identified a subgroup of individuals undergoing a distinct splicing program during aging [<a href="#ref-8">8</a>].
saseR
saseR provides a framework for prioritizing expression and usage outliers that is much faster than state-of-the-art methods and significantly outperforms these for aberrant splicing detection [<a href="#ref-9">9</a>][<a href="#ref-10">10</a>]. It works by replacing normalization offsets, which unlocks bulk RNA-seq tools for differential usage and aberrant splicing [<a href="#ref-9">9</a>][<a href="#ref-10">10</a>].
saseR is suitable for rare disease diagnosis through splicing and expression outliers and can handle both short- and long-read applications [<a href="#ref-9">9</a>][<a href="#ref-10">10</a>].
Practical Workflow for Tool Selection
The following workflow provides a structured approach to selecting and applying a differential splicing tool.
Step 1: Define Your Biological Question
Write down the specific question you need to answer. Are you looking for changes in known splicing events, novel splicing events, or isoform-level changes? Do you need to account for covariates? This decision determines the tool category.
Step 2: Assess Your Data
Check your number of biological replicates, whether you have covariates, whether you are using short-read or long-read sequencing, and the quality of your genome annotation. The 2020 benchmarking study found that tool performance varied with the number of samples [<a href="#ref-3">3</a>], so your sample size matters.
Step 3: Select a Primary Tool
Based on your question and data, select a primary tool from the appropriate category. If you are unsure, consider running two tools from different categories and comparing results [<a href="#ref-2">2</a>].
Step 4: Prepare Input Data
Follow the tool documentation to prepare your input data. This typically involves aligning reads to the reference genome, generating count matrices, and formatting annotation files. The Galaxy Training Network provides accessible workflow training and analysis tutorials that can help with these steps [<a href="#ref-16">16</a>].
Step 5: Run the Analysis and Assess Quality
Run the tool and examine the output for quality indicators, such as the distribution of p-values and the number of significant events. If the results look anomalous, check your input data for errors.
Step 6: Validate and Interpret Results
Validate a subset of significant events using an independent method, such as RT-PCR or visualization of read alignments. Interpret your results in the context of your biological question.
Records and Measurements to Keep
Good record keeping is essential for reproducible differential splicing analysis. The Carpentries Lessons provide foundational training in computing, data, shell, Git, and programming that supports reproducible workflows [<a href="#ref-17">17</a>]. The nf-core Documentation describes community pipeline standards, usage, configuration, and reproducible workflow context [<a href="#ref-18">18</a>].
For each analysis, record the following:
- Tool name and version
- Reference genome and annotation versions
- Alignment tool and parameters
- Count generation method and parameters
- Statistical model and thresholds
- Number of samples and experimental design
- Covariates included in the model
- Date of analysis and person who performed it
These records allow you to reproduce your analysis and to compare results across tools or datasets.
Common Failure Patterns and How to Avoid Them
Several common errors lead to failed or misleading differential splicing analyses.
Using the Wrong Input Format
Each tool has specific input requirements. Feeding a count matrix to a tool that expects BAM files, or vice versa, will cause errors or incorrect results. Read the tool documentation carefully before preparing your data.
Ignoring Annotation Quality
Tools that rely on annotated splice junctions cannot detect novel splicing events [<a href="#ref-2">2</a>]. If your organism has poor annotation, or if you suspect novel isoforms are important, use tools that can handle de novo junctions or that do not require annotations.
Insufficient Replicates
Differential splicing analysis requires biological replicates to estimate variability. The 2020 benchmarking study found that tool performance varied with the number of samples [<a href="#ref-3">3</a>]. With too few replicates, results will be unreliable regardless of which tool you use.
Overlooking Covariates
If your study includes covariates such as sex, age, or clinical attributes, failing to account for them can produce spurious results. Use tools such as MntJULiP or saseR that can handle covariates [<a href="#ref-8">8</a>][<a href="#ref-9">9</a>][<a href="#ref-10">10</a>].
Running Only One Tool
Different tools can produce different results on the same data [<a href="#ref-2">2</a>][<a href="#ref-3">3</a>]. Running a single tool may miss events that other tools would detect or report false positives that other tools would filter. Consider running two tools and focusing on events identified by both [<a href="#ref-2">2</a>].
Limitations and Interpretation Boundaries
Differential splicing tools have inherent limitations that affect interpretation of results.
Event-Based Tools Cannot Discover Novel Splice Sites
Most event-based tools are unsuitable for discovering novel splice sites because they rely on annotated gene models [<a href="#ref-2">2</a>]. If novel splice site discovery is a goal, you need tools that can assemble transcripts or accept de novo junctions.
RNA-Level Results Do Not Guarantee Protein-Level Changes
Alternative splicing affects 95% of multi-exon genes, generating protein isoforms with distinct functions [<a href="#ref-19">19</a>]. However, changes in RNA splicing do not always translate to changes in protein isoforms. Tools such as IsoPepTracker can help analyze and visualize differential peptides across canonical and novel isoforms that are theoretically detectable by shotgun mass spectrometry-based proteomics [<a href="#ref-19">19</a>].
Short-Read Tools Have Limited Isoform Resolution
Short-read sequencing cannot always determine which combinations of exons are present in a single transcript. Isoform-based tools estimate isoform abundances but with inherent uncertainty. Long-read sequencing provides full-length transcripts and can reveal isoform structure directly [<a href="#ref-14">14</a>].
Benchmarking Results Depend on Data Characteristics
Benchmarking studies use simulated or specific real datasets, and their results may not generalize to your data. The 2020 study found that different tools performed strikingly differently across different datasets or numbers of samples [<a href="#ref-3">3</a>]. Use benchmarking results as guidance, not as definitive rankings.
Professional Escalation Criteria
Some situations warrant consulting a bioinformatics specialist or collaborating with an expert.
When to Seek Help
- You are unsure which tool category is appropriate for your biological question
- Your experimental design includes complex covariates or batch effects
- You are working with a non-model organism with poor annotation
- You are using long-read sequencing data and are unsure which tools are appropriate
- Your results are inconsistent across tools or datasets
- You need to discover novel splicing events and are unsure which tools can do this
Where to Find Help
The EMBL-EBI Training portal offers bioinformatics learning pathways and practical analysis education [<a href="#ref-13">13</a>]. The Galaxy Training Network provides accessible workflow training and analysis tutorials [<a href="#ref-16">16</a>]. Bioconductor provides official package, workflow, installation, and reproducible genomic-analysis documentation [<a href="#ref-15">15</a>]. The Carpentries Lessons offer foundational computing and data training [<a href="#ref-17">17</a>].
A Practical Decision Framework for Matching Tools to Experimental Constraints
Beyond the general categories of exon-based, event-based, and isoform-based tools, researchers need a concrete decision framework that maps specific experimental constraints to tool choices. The benchmarking literature consistently shows that tool performance varies dramatically with data characteristics, so a fixed recommendation for all situations will fail in practice [<a href="#ref-2">2</a>][<a href="#ref-3">3</a>]. This section provides a structured framework based on five experimental constraints that determine which tool class will perform reliably for your specific dataset.
Constraint 1: Number of Biological Replicates
The number of biological replicates in your experiment is the first constraint that should narrow your tool choices. The 2020 systematic evaluation found that different tools performed strikingly differently across different datasets or numbers of samples [<a href="#ref-3">3</a>]. This means that a tool that works well with eight replicates per condition may produce unreliable results with only three.
For experiments with three or fewer replicates per condition, tools that borrow information across genes are preferable. The limma package was specifically designed to overcome the problem of small sample sizes through information borrowing [<a href="#ref-5">5</a>]. Its voom transformation estimates the mean-variance relationship and applies linear modeling in a way that stabilizes estimates when replication is limited [<a href="#ref-5">5</a>]. Exon-based methods in this category, including DEXSeq and JunctionSeq, also scored well in benchmarking studies and can handle modest replication [<a href="#ref-3">3</a>].
For experiments with five or more replicates per condition, event-based tools such as rMATS and MAJIQ become more reliable. These tools estimate inclusion levels for individual splicing events and require sufficient replication to estimate variability accurately. The 2023 benchmarking study found that rMATS and LeafCutter outperformed other event-based tools in their evaluation, but this performance was established on datasets with adequate replication [<a href="#ref-2">2</a>].
For experiments with dozens or hundreds of samples, scalability becomes the dominant concern. MntJULiP can process thousands of samples within hours using a Bayesian mixture model [<a href="#ref-8">8</a>]. saseR similarly provides a framework that scales to large datasets and is much faster than state-of-the-art methods for prioritizing expression and usage outliers [<a href="#ref-9">9</a>][<a href="#ref-10">10</a>]. If your study includes hundreds of samples from population cohorts, these tools are more appropriate than tools designed for small controlled experiments.
Constraint 2: Covariate Structure
Emerging large and complex RNA-seq datasets from disease and population studies include multiple confounders such as sex, age, ethnicity, and clinical attributes [<a href="#ref-8">8</a>]. Few methods are equipped to handle these challenges [<a href="#ref-8">8</a>]. If your experimental design includes covariates that you must account for, your tool choices are substantially restricted.
Standard event-based tools such as rMATS and SUPPA2 do not natively incorporate covariate information into their statistical models. Running these tools on data with strong batch effects or demographic confounders risks producing spurious differential splicing calls. The 2020 benchmarking study did not evaluate covariate handling because the tools tested were not designed for this purpose [<a href="#ref-3">3</a>].
For covariate-aware analysis, MntJULiP detects intron-level differences in alternative splicing using a Bayesian mixture model that takes covariates into account [<a href="#ref-8">8</a>]. The developers applied this method to GTEx brain RNA-seq samples to deconvolute the effects of sex and age at death on splicing patterns [<a href="#ref-8">8</a>]. This application demonstrates the practical value of covariate adjustment: clustering of covariate-adjusted data identified a subgroup of individuals undergoing a distinct splicing program during aging that would have been obscured without adjustment [<a href="#ref-8">8</a>].
saseR offers an alternative approach by replacing normalization offsets to unlock bulk RNA-seq tools for differential usage and aberrant splicing [<a href="#ref-9">9</a>][<a href="#ref-10">10</a>]. This offset-based strategy allows existing tools to accommodate covariates and other technical factors without requiring new statistical models [<a href="#ref-9">9</a>][<a href="#ref-10">10</a>].
If your study includes covariates and you cannot use MntJULiP or saseR, consider whether you can stratify your analysis or use a two-stage approach. However, stratifying reduces power and may not fully account for continuous covariates such as age. The safer approach is to select a tool designed for covariate handling from the outset.
Constraint 3: Event Type Priorities
Different tools have different strengths for detecting specific types of splicing events. The targeted RNA long-amplicon sequencing evaluation provides a clear example of this variation. MAJIQ was able to detect all types of splicing events tested, including exon skipping, multiple-exon skipping, alternative 5' splicing, and alternative 3' splicing [<a href="#ref-11">11</a>]. However, MAJIQ was unable to detect one of the exon-skipping events in the test set [<a href="#ref-11">11</a>]. In contrast, rMATS was able to detect all exon-skipping events but failed to detect other types of events besides exon skipping [<a href="#ref-11">11</a>].
This finding has direct practical implications. If your biological question centers on exon skipping, rMATS is a strong choice because it detects this event type reliably [<a href="#ref-11">11</a>]. If you need to survey a broad range of event types, including alternative splice site usage, MAJIQ provides more comprehensive coverage [<a href="#ref-11">11</a>]. The 2023 benchmarking study similarly found that SUPPA, DARTS, rMATS, and LeafCutter outperformed other event-based tools, but the optimal choice among these depends on which event types matter for your question [<a href="#ref-2">2</a>].
Consider also whether you need to detect intron retention. LeafCutter focuses on intron usage instead of exon inclusion, making it appropriate when intron retention is biologically relevant [<a href="#ref-4">4</a>]. The 2023 benchmarking study found that LeafCutter outperformed most other event-based tools, supporting its use for intron-centric questions [<a href="#ref-2">2</a>].
Constraint 4: Novel Event Discovery Requirements
If your biological question requires discovering novel splicing events, most event-based tools are unsuitable because they rely on annotated gene models [<a href="#ref-2">2</a>]. The 2023 benchmarking study examined the abilities of tools to identify novel splicing events and found that most event-based tools were unsuitable for discovering novel splice sites [<a href="#ref-2">2</a>].
For novel event discovery, you have three practical options. First, LeafCutter does not require an annotation file and can identify novel intron usage patterns [<a href="#ref-2">2</a>][<a href="#ref-4">4</a>]. This makes it useful when you suspect unannotated splicing events are biologically important. Second, DJExpress can handle both annotated and de novo identified splice junctions, allowing quantification of novel splice events [<a href="#ref-7">7</a>]. This tool accepts raw splice junction counts as input and can process data on an average laptop computer, making it accessible to researchers without high-performance computing resources [<a href="#ref-7">7</a>]. Third, you can use transcript assembly approaches that reconstruct full-length transcripts without relying on existing annotations.
If you use a tool that relies on annotations, you will miss novel events entirely. The 2023 benchmarking study found that most event-based tools were unsuitable for discovering novel splice sites [<a href="#ref-2">2</a>]. This limitation is not a minor detail, it means that your results will be constrained by the completeness of your annotation file. For poorly annotated organisms or for discovering disease-specific splicing alterations, this constraint can be decisive.
Constraint 5: Computational Resources and User Expertise
The computational demands of differential splicing tools vary substantially, and your available resources should inform your tool choice. The 2023 benchmarking study noted that tools vary in computational requirements, ease of use, and the specificity of the splicing events they detect [<a href="#ref-4">4</a>]. Some tools require substantial memory and processing power, while others can run on standard laptops.
DJExpress was specifically designed to run on an average laptop computer while providing interactive visualization formats [<a href="#ref-7">7</a>]. This makes it accessible to researchers without access to high-performance computing clusters. The tool also offers a web-compatible graphical interface, reducing the need for command-line expertise [<a href="#ref-7">7</a>].
For researchers with limited computational training, the Galaxy Training Network provides accessible workflow training and analysis tutorials that can help with the steps required to prepare data and run analyses [<a href="#ref-16">16</a>]. The EMBL-EBI Training portal offers bioinformatics learning pathways and practical analysis education [<a href="#ref-13">13</a>]. Bioconductor provides official package, workflow, installation, and reproducible genomic-analysis documentation [<a href="#ref-15">15</a>].
For researchers comfortable with command-line tools and pipeline management, nf-core documentation describes community pipeline standards, usage, configuration, and reproducible workflow context [<a href="#ref-18">18</a>]. The Carpentries Lessons offer foundational training in computing, data, shell, Git, and programming that supports reproducible workflows [<a href="#ref-17">17</a>].
Applying the Framework: A Worked Example
Consider a researcher studying alternative splicing in a cancer cohort with 50 tumor samples and 50 normal samples, with sex and age as covariates, using short-read RNA-seq data from a well-annotated genome. The biological question is whether specific exon-skipping events differ between tumor and normal tissue.
Applying the framework: the sample size is large, so scalability matters. Covariates are present, so covariate-aware tools are required. The event type priority is exon skipping, which rMATS detects reliably [<a href="#ref-11">11</a>]. Novel event discovery is not a primary goal because the genome is well annotated. Computational resources include access to a standard university server.
The framework suggests MntJULiP or saseR as primary candidates because they handle covariates and scale to hundreds of samples [<a href="#ref-8">8</a>][<a href="#ref-9">9</a>][<a href="#ref-10">10</a>]. If the researcher prefers an event-based approach, they could run rMATS on covariate-adjusted data or use a two-stage approach. The 2023 benchmarking study proposed a tool-pair combination approach to improve overall performance [<a href="#ref-2">2</a>], so running both MntJULiP and rMATS and focusing on events identified by both would provide more confidence in the results.
Recording Your Decision Process
For reproducibility, record the rationale for your tool selection. The Carpentries Lessons provide foundational training in computing, data, shell, Git, and programming that supports reproducible workflows [<a href="#ref-17">17</a>]. The nf-core Documentation describes community pipeline standards, usage, configuration, and reproducible workflow context [<a href="#ref-18">18</a>].
For each analysis, document the following decision points:
- Number of biological replicates per condition
- Covariates included in the experimental design
- Priority splicing event types
- Whether novel event discovery is required
- Available computational resources
- Tool name and version selected
- Rationale for selection based on the constraints above
- Alternative tools considered and reasons for rejection
This record allows you to revisit your decision if results are unexpected and to compare findings across tools or datasets. The 2020 benchmarking study found that different tools performed strikingly differently across different datasets or numbers of samples [<a href="#ref-3">3</a>], so documenting your choices is essential for interpreting your results in context.
When to Reconsider Your Tool Choice
Your initial tool selection may need revision as you examine your data. If your results show an anomalous distribution of p-values, an unexpectedly large number of significant events, or poor overlap with known biology, reconsider whether your tool choice matches your experimental constraints. The 2023 benchmarking study observed immense discrepancies among 21 different tools on simulated data with known differential splicing events [<a href="#ref-2">2</a>]. If your tool produces results that conflict with validated biology, the tool may be mismatched to your data characteristics.
Similarly, if you discover that your data have stronger batch effects or more covariates than anticipated, you may need to switch to a covariate-aware tool. The 2020 benchmarking study found that tool performance varied with the number of samples [<a href="#ref-3">3</a>], so if your replication is lower than planned, consider whether your chosen tool can handle the reduced sample size.
Professional Escalation Criteria for Tool Selection
Some situations warrant consulting a bioinformatics specialist before proceeding with analysis. If your experimental design includes complex covariates or batch effects that you are unsure how to model, seek expert advice. If you are working with a non-model organism with poor annotation and need to discover novel splicing events, the tool selection is more complex and may require custom pipelines. If you are using long-read sequencing data and are unsure which tools are appropriate, consult the literature on ONT-aware tools such as DRIMSeq, DEXSeq, and stageR [<a href="#ref-14">14</a>]. If your results are inconsistent across tools or datasets, a specialist can help you diagnose whether the issue is tool choice, data quality, or experimental design.
The EMBL-EBI Training portal offers bioinformatics learning pathways and practical analysis education [<a href="#ref-13">13</a>]. The Galaxy Training Network provides accessible workflow training and analysis tutorials [<a href="#ref-16">16</a>]. Bioconductor provides official package, workflow, installation, and reproducible genomic-analysis documentation [<a href="#ref-15">15</a>]. These resources can help you build the skills needed to make informed tool choices and to troubleshoot problems when they arise.
Frequently Asked Questions
What is the difference between differential expression and differential splicing analysis?
Differential expression analysis tests whether the total abundance of a gene changes between conditions. Differential splicing analysis tests whether the relative usage of transcript isoforms changes between conditions, independent of total gene expression. A gene can be expressed at the same total level while the proportions of its isoforms change, and vice versa. Tools for differential splicing measure the relative usage of exons, junctions, or isoforms, beyond their absolute abundance [<a href="#ref-5">5</a>][<a href="#ref-3">3</a>].
Which differential splicing tool should I use for a simple two-group comparison?
For a simple two-group comparison with biological replicates, rMATS is a strong choice for detecting exon skipping and other common splicing events [<a href="#ref-2">2</a>][<a href="#ref-11">11</a>]. DEXSeq or limma are appropriate if you want to test all exons without specifying events in advance [<a href="#ref-5">5</a>][<a href="#ref-3">3</a>]. The 2020 benchmarking study found that exon-based methods and two event-based methods, MAJIQ and rMATS, scored well on the selected measures [<a href="#ref-3">3</a>].
Can I use the same tool for short-read and long-read RNA-seq data?
Some tools can handle both short- and long-read data, but not all. saseR provides a single framework for various short- and long-read applications by replacing normalization offsets [<a href="#ref-9">9</a>][<a href="#ref-10">10</a>]. Long-read data require tools that can handle higher error rates and different count distributions. A 2026 study used ONT-aware tools such as DRIMSeq, DEXSeq, and stageR for long-read data [<a href="#ref-14">14</a>].
How many biological replicates do I need for differential splicing analysis?
The number of replicates affects tool performance. The 2020 benchmarking study found that different tools performed strikingly differently across different datasets or numbers of samples [<a href="#ref-3">3</a>]. Tools such as limma are designed to borrow information across genes to overcome the problem of small sample sizes [<a href="#ref-5">5</a>]. In general, more replicates provide more reliable results, but the optimal number depends on the variability in your system and the tool you choose.
What should I do if two tools give different results on the same data?
Different tools can produce different results on the same data [<a href="#ref-2">2</a>][<a href="#ref-3">3</a>]. This is expected because tools use different statistical models, input data, and event definitions. A practical approach is to run two tools and focus on events identified by both [<a href="#ref-2">2</a>]. The 2023 benchmarking study proposed a tool-pair combination approach to improve overall performance [<a href="#ref-2">2</a>].
Can differential splicing tools detect novel splicing events?
Most event-based tools are unsuitable for discovering novel splice sites because they rely on annotated gene models [<a href="#ref-2">2</a>]. Tools that can handle de novo identified splice junctions, such as DJExpress, allow quantification of novel splice events [<a href="#ref-7">7</a>]. LeafCutter does not require an annotation file and can identify novel intron usage patterns [<a href="#ref-2">2</a>][<a href="#ref-4">4</a>].
Do RNA-level splicing changes always lead to protein-level changes?
No. Alternative splicing affects 95% of multi-exon genes, generating protein isoforms with distinct functions [<a href="#ref-19">19</a>]. However, changes in RNA splicing do not always translate to changes in protein isoforms due to mechanisms such as nonsense-mediated decay and translational regulation. Tools such as IsoPepTracker can help analyze and visualize differential peptides across canonical and novel isoforms that are theoretically detectable by shotgun mass spectrometry-based proteomics [<a href="#ref-19">19</a>].
How do I account for covariates such as sex and age in differential splicing analysis?
Few methods are equipped to handle covariates in differential splicing analysis [<a href="#ref-8">8</a>]. MntJULiP detects intron-level differences in alternative splicing using a Bayesian mixture model and takes covariates into account [<a href="#ref-8">8</a>]. saseR provides a framework for prioritizing expression and usage outliers and can handle large datasets [<a href="#ref-9">9</a>][<a href="#ref-10">10</a>]. If your study includes covariates, use tools that can account for them instead of tools that assume all samples are exchangeable.
Related Bioinformatics Guides
- Single-Cell RNA Sequencing Depth: A Cost-Benefit Analysis for Experimental Design
- RNA-Seq Alignment: Choosing the Right Tool and Parameters
- Genomic Data Analysis Tools: A Comparative Guide for Researchers
- TMT Proteomics: Experimental Design, Labeling, and Data Analysis
- RNA-Seq Alignment Tools: STAR, HISAT2, and Beyond
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
[1] [RNA sequencing: the teenage years.](https://pubmed.ncbi.nlm.nih.gov/31341269). Nature reviews. Genetics, 2019. [2] [A comprehensive benchmarking of differential splicing tools for RNA-seq analysis at the event level.](https://pubmed.ncbi.nlm.nih.gov/37020334). Briefings in bioinformatics, 2023. [3] [Systematic evaluation of differential splicing tools for RNA-seq studies.](https://pubmed.ncbi.nlm.nih.gov/31802105). Briefings in bioinformatics, 2020. [4] [Comprehensive Overview of Computational Tools for Alternative Splicing Analysis.](https://pubmed.ncbi.nlm.nih.gov/41326318). Wiley interdisciplinary reviews. RNA, 2025. [5] [limma powers differential expression analyses for RNA-sequencing and microarray studies.](https://pubmed.ncbi.nlm.nih.gov/25605792). Nucleic acids research, 2015. [6] [Differential splicing analysis based on isoforms expression with NBSplice.](https://pubmed.ncbi.nlm.nih.gov/31972288). Journal of biomedical informatics, 2020. [7] [DJExpress: An Integrated Application for Differential Splicing Analysis and Visualization](https://doi.org/10.3389/fbinf.2022.786898). Frontiers in Bioinformatics, 2022. [8] [MntJULiP and Jutils: Differential splicing analysis of RNA-seq data with covariates.](https://pubmed.ncbi.nlm.nih.gov/38260578). bioRxiv : the preprint server for biology, 2024. [9] [saseR: juggling offsets unlocks RNA-seq tools for fast and scalable differential usage, aberrant splicing and expression retrieval.](https://doi.org/10.1186/s13059-026-03973-8). 2026. [10] [saseR: Juggling offsets unlocks RNA-seq tools for fast and Scalable differential usage, Aberrant Splicing and Expression Retrieval](https://doi.org/10.1101/2023.06.29.547014). bioRxiv, 2024. [11] [Computational Comparison of Differential Splicing Tools for Targeted RNA Long-Amplicon Sequencing (rLAS).](https://pubmed.ncbi.nlm.nih.gov/40244027). International journal of molecular sciences, 2025. [12] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information. [13] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute. [14] [A whole-transcriptome analysis of differentially expressed genes, transcripts, and transcript usage in blood samples from Parkinson's disease patients.](https://doi.org/10.3389/ebm.2026.11099). 2026. [15] [Bioconductor](https://bioconductor.org/). Bioconductor Project. [16] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project. [17] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries. [18] [nf-core Documentation](https://nf-co.re/docs). nf-core. [19] [IsoPepTracker: An interactive web application for peptide-driven isoform analysis.](https://doi.org/10.1371/journal.pcbi.1014324). 2026.This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.