RSEM vs. Salmon: Which Tool Should You Use for Isoform Quantification in RNA-seq?
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- RSEM utilizes full read alignment via external aligners (e.g., STAR, Bowtie2) to generate BAM/SAM files, which are essential for downstream analyses like variant calling or genome browser visualization. This alignment-based approach is computationally more intensive but provides comprehensive mapping information.
- Salmon employs a quasi-mapping strategy, significantly accelerating quantification by avoiding full alignment and directly estimating transcript abundances. This method is ideal for large datasets or when rapid turnaround is critical, though it does not produce alignment files in its primary mode.
- Salmon offers built-in bias correction for sequence-specific, GC, and positional biases, potentially improving accuracy, especially in datasets with library preparation artifacts. RSEM's bias correction is primarily through its read generation model, offering less explicit correction.
- Integration with downstream differential expression analysis tools differs: RSEM outputs are often used directly or with minimal conversion, while Salmon results typically require the
tximportpackage for seamless import into tools like DESeq2, edgeR, and limma. - Computational resource requirements vary: Salmon generally exhibits lower memory usage due to its compact indexing, making it more suitable for systems with limited RAM compared to the larger indices required by alignment-based methods like RSEM.
Researchers planning RNA sequencing studies face a critical decision at the quantification step: whether to use RSEM or Salmon for estimating transcript and isoform abundances. This choice affects computational time, memory requirements, accuracy of isoform-level estimates, and compatibility with downstream differential expression analysis. This article provides a direct comparison of both tools, explains their methodological differences, and offers practical recommendations based on your specific research context, data characteristics, and computational resources.
Understanding the Quantification Problem in RNA-seq
RNA sequencing produces millions of short reads that must be assigned to genomic features to measure gene and transcript expression. The core challenge is that reads often map to multiple locations, particularly across isoforms that share exonic sequences. Quantification tools must resolve this ambiguity to produce accurate abundance estimates.
The quantification step sits between read preprocessing and differential expression analysis in the standard RNA-seq workflow. After quality control and trimming, reads are aligned or pseudo-aligned to a reference, and counts are generated for genes or transcripts. These counts feed into statistical models that identify differentially expressed genes between conditions.
Isoform-level quantification is more complex than gene-level quantification because transcripts from the same gene frequently share exons. A read originating from a shared exon cannot be unambiguously assigned to a single isoform without additional information. Both RSEM and Salmon address this challenge using expectation-maximization algorithms, but they differ substantially in their implementation and computational strategy.
The choice between RSEM and Salmon depends on several factors: whether you need alignment files for downstream analyses, the size of your dataset, available computational resources, and compatibility with your preferred differential expression tools. Understanding these tradeoffs helps you select the appropriate tool for your specific workflow.
Core Principles of Transcript Quantification
How RSEM Works
RSEM, which stands for RNA-Seq by Expectation-Maximization, uses an alignment-based approach to quantify transcript abundances. The tool first aligns RNA-seq reads to a reference transcriptome using an external aligner such as Bowtie, Bowtie2, or STAR. The resulting alignments are then processed through an expectation-maximization algorithm that estimates transcript abundances while accounting for multi-mapping reads.
The RSEM pipeline consists of two main stages. The first stage prepares the reference by building an index of the transcriptome, which includes information about transcript sequences, gene-transcript relationships, and read generation parameters. The second stage aligns reads to this index and estimates abundances using the expectation-maximization algorithm.
RSEM produces several output files, including gene-level and transcript-level count matrices, estimated expression values in transcripts per million (TPM), and posterior probability estimates for read assignments. These outputs integrate directly with downstream tools such as EBSeq and edgeR for differential expression analysis.
One important characteristic of RSEM is that it requires aligned reads in BAM or SAM format. This means you must run an alignment step before quantification, which adds computational time but produces alignment files that can be used for other purposes such as variant calling or visualization in genome browsers.
How Salmon Works
Salmon uses a fundamentally different approach called quasi-mapping, which avoids full alignment of reads to the reference transcriptome. Instead, Salmon identifies the likely origin of each read by matching k-mers between the read and the transcriptome index, then refines these mappings using a lightweight algorithm that considers the compatibility of the read with candidate transcripts.
This quasi-mapping approach is substantially faster than traditional alignment because it does not need to compute base-to-base alignments for every read. Salmon processes reads in a streaming fashion and uses a dual-phase inference procedure that combines quasi-mapping with an expectation-maximization algorithm to estimate transcript abundances.
Salmon also implements a technique called selective alignment, which improves accuracy by considering the full read sequence when determining compatibility with candidate transcripts. This helps distinguish between reads that originate from different isoforms or from highly similar transcripts.
The output from Salmon includes transcript-level counts and TPM values, as well as gene-level summaries when a gene-to-transcript mapping is provided. Salmon can also output equivalence classes, which describe sets of transcripts that share compatible reads, providing additional information for downstream analyses.
Key Methodological Differences
The primary difference between RSEM and Salmon lies in their read-to-transcript assignment strategies. RSEM relies on external aligners to produce alignments, while Salmon uses quasi-mapping to determine read compatibility. This difference drives the substantial speed advantage of Salmon, which can process datasets in minutes that might take hours with alignment-based approaches.
Both tools use expectation-maximization algorithms to resolve multi-mapping reads, but they differ in how they model the read generation process. RSEM uses a generative model that accounts for read length distributions and sequencing errors, while Salmon uses a more flexible model that can incorporate fragment-level information and bias correction.
Salmon includes built-in bias correction for sequence-specific biases, fragment-level GC bias, and positional biases that can arise during library preparation and sequencing. RSEM can also account for some biases through its read generation model, but it does not offer the same level of built-in bias correction as Salmon.
Another important difference is that Salmon can operate in alignment-based mode using pre-computed alignments from tools like STAR, in addition to its quasi-mapping mode. This flexibility allows researchers to use Salmon for quantification even when they need alignments for other purposes.
At a Glance: RSEM vs. Salmon Comparison
| Feature | RSEM | Salmon |
|---|---|---|
| Read assignment method | Full alignment via external aligner (Bowtie, Bowtie2, STAR) | Quasi-mapping with selective alignment |
| Computational speed | Slower due to alignment step | Fast, often 10-20 times faster than alignment-based methods |
| Memory requirements | Moderate, depends on aligner and index size | Lower, uses compact indexing |
| Bias correction | Limited, through read generation model | Built-in sequence, GC, and positional bias correction |
| Output formats | Gene and transcript counts, TPM, posterior probabilities | Transcript counts, TPM, equivalence classes |
| Downstream compatibility | EBSeq, edgeR, DESeq2 with appropriate imports | tximport for DESeq2, edgeR, limma, sleuth |
| Alignment files produced | Yes, BAM/SAM files | No alignment files in quasi-mapping mode |
| Best use case | When alignments are needed for other analyses | Large datasets, fast turnaround, integration with tximport |
Practical Workflow Considerations
Reference Preparation and Indexing
Both RSEM and Salmon require a reference transcriptome for quantification. The quality of this reference directly affects the accuracy of your results. For model organisms with well-annotated genomes, you can download transcript sequences from public databases such as NCBI or Ensembl. The NCBI provides official descriptions of its databases and sequence resources, which can help you identify appropriate reference files for your organism of interest.
RSEM requires you to prepare the reference using the rsem-prepare-reference command, which builds an index from transcript sequences. This step also generates gene-to-transcript mapping files that are used during quantification. The indexing step can take considerable time and memory for large genomes, so plan accordingly.
Salmon uses the salmon index command to build its index from transcript sequences. The indexing process creates a compact representation of the transcriptome that enables fast quasi-mapping. Salmon also supports the use of decoy sequences, which are genomic sequences that help prevent spurious mappings to transcripts that share sequence similarity with genomic regions.
For both tools, the choice of transcriptome annotation is critical. Using an incomplete or outdated annotation can lead to inaccurate quantification, particularly for genes with many isoforms. Check the annotation version and ensure it matches your experimental context.
Read Preprocessing and Quality Control
Before quantification, raw sequencing reads should undergo quality control to remove adapter sequences, low-quality bases, and contaminating sequences. The Galaxy Training Network provides accessible workflow training and analysis tutorials that cover standard RNA-seq quality control procedures, which can help you establish a reproducible preprocessing pipeline.
Quality control steps typically include:
- Assessing raw read quality using tools like FastQC
- Trimming adapter sequences and low-quality bases
- Removing reads that map to ribosomal RNA or other contaminants
- Checking for batch effects or technical artifacts
The choice of preprocessing steps can affect quantification accuracy. Overly aggressive trimming can remove informative bases, while insufficient trimming can leave adapter sequences that interfere with mapping. Test different preprocessing parameters on a subset of your data to determine the optimal approach for your specific dataset.
Running RSEM
The RSEM workflow requires two main commands. First, prepare the reference:
rsem-prepare-reference --gtf annotation.gtf transcriptome.fa rsem_index
Then align and quantify reads:
rsem-calculate-expression --paired-end --star -p 8 reads_1.fastq reads_2.fastq rsem_index output_prefix
The --star flag tells RSEM to use the STAR aligner, which is recommended for larger genomes due to its speed and accuracy. You can also use Bowtie2 with the --bowtie2 flag, which may be more appropriate for smaller transcriptomes.
RSEM produces several output files, including:
.genes.resultsfile with gene-level counts and TPM values.isoforms.resultsfile with transcript-level counts and TPM values.bamand.baifiles with alignments.statfile with alignment statistics
Review the alignment statistics to check for problems such as low overall alignment rates or excessive multi-mapping reads. These issues may indicate problems with the reference or the preprocessing steps.
Running Salmon
Salmon also requires two main commands. First, build the index:
salmon index -t transcriptome.fa -i salmon_index
Then quantify reads:
salmon quant -i salmon_index -l A -1 reads_1.fastq -2 reads_2.fastq -p 8 -o output_directory
The -l A flag tells Salmon to automatically determine the library type, which is useful when you are unsure about the strandedness of your library. You can also specify the library type explicitly if you know it from your library preparation protocol.
Salmon produces a directory with several output files, including:
quant.sffile with transcript-level counts and TPM valuesquant.genes.sffile with gene-level summariesaux_info/directory with additional information such as equivalence classes and bias parameters
The quant.sf file is the primary output for downstream analyses. It contains columns for transcript ID, length, effective length, TPM, and estimated number of reads.
Integrating Quantification Results with Downstream Tools
The output from RSEM and Salmon feeds into different downstream analysis pipelines. RSEM results can be used directly with EBSeq for differential expression analysis, or converted for use with edgeR and DESeq2. Salmon results are typically imported into R using the tximport package, which prepares count matrices for DESeq2, edgeR, limma, and other tools.
The Bioconductor project provides official package documentation and workflow guidance for many of these downstream tools. The tximport package is specifically designed to import transcript-level quantification data from tools like Salmon and RSEM into gene-level count matrices for differential expression analysis.
When using tximport, you need to provide a transcript-to-gene mapping file that tells the package which transcripts belong to which genes. This mapping is typically derived from the annotation file used to build the reference index.
The choice of differential expression tool depends on your experimental design and the assumptions you are willing to make about your data. DESeq2 and edgeR use negative binomial models that account for overdispersion in count data, while limma uses linear models with empirical Bayes moderation. Each tool has strengths and limitations, and the choice should be based on your specific research question.
Performance Benchmarks and Accuracy Considerations
Speed and Resource Requirements
The most significant practical difference between RSEM and Salmon is computational speed. Salmon's quasi-mapping approach processes reads much faster than alignment-based methods, making it particularly attractive for large datasets or when rapid turnaround is needed.
For a typical RNA-seq dataset with 20-30 million paired-end reads, Salmon can complete quantification in 10-20 minutes on a standard workstation, while RSEM with STAR alignment may take 1-2 hours. The speed advantage becomes more pronounced with larger datasets, where Salmon can process data in a fraction of the time required by alignment-based approaches.
Memory requirements also differ between the tools. Salmon uses a compact index that requires less memory than the full genome index used by STAR for RSEM. This makes Salmon more suitable for analysis on machines with limited memory, such as standard laptops or cloud instances with modest specifications.
The nf-core documentation provides community pipeline standards and configuration guidance that can help you optimize resource allocation for RNA-seq analysis workflows. Many nf-core pipelines include both RSEM and Salmon as options, allowing you to choose the appropriate tool for your computational environment.
Accuracy of Isoform-Level Estimates
Both RSEM and Salmon produce accurate isoform-level estimates for most datasets, but their performance can vary depending on the complexity of the transcriptome and the characteristics of the sequencing data.
Salmon's built-in bias correction can improve accuracy for datasets with sequence-specific biases, which are common in certain library preparation protocols. The selective alignment approach also helps distinguish between highly similar transcripts, reducing spurious assignments.
RSEM's generative model accounts for read length distributions and sequencing errors, which can be beneficial for datasets with unusual read length distributions or higher error rates. The alignment-based approach may also be more accurate for transcripts with low expression levels, where the additional information from full alignments can help resolve ambiguities.
For most standard RNA-seq datasets, the accuracy differences between RSEM and Salmon are small and unlikely to affect biological conclusions. However, for datasets with complex isoform structures or high sequence similarity between transcripts, the choice of tool may have a more substantial impact.
Compatibility with Downstream Analyses
The compatibility of quantification results with downstream tools is an important consideration. Salmon's output integrates seamlessly with tximport, which is the recommended approach for importing transcript-level quantification into DESeq2, edgeR, and limma. This workflow is well-documented and widely used in the bioinformatics community.
RSEM results can also be imported into these tools, but the process is slightly more involved. The rsem-generate-data-matrix command can create count matrices from RSEM output, which can then be used with edgeR or DESeq2. Alternatively, you can use the tximport package with RSEM output by providing the appropriate files.
The Bioconductor project provides official documentation for many of these workflows, including detailed vignettes that walk through the entire analysis pipeline from quantification to differential expression. These resources can help you implement a reproducible analysis workflow that meets your specific needs.
Options and Tradeoffs in Tool Selection
When to Choose RSEM
RSEM is the appropriate choice when you need alignment files for downstream analyses beyond quantification. If you plan to perform variant calling, visualize reads in a genome browser, or conduct other analyses that require aligned reads, RSEM provides these files as part of its standard output.
RSEM may also be preferred when you need to integrate quantification with tools that have been specifically designed to work with RSEM output. Some differential expression tools and downstream analysis pipelines have been optimized for RSEM results, and using RSEM can simplify the integration process.
For researchers who are already familiar with alignment-based workflows and have established pipelines that use RSEM, continuing with RSEM may be more practical than switching to Salmon. The learning curve for Salmon is relatively shallow, but changing tools requires updating scripts and potentially re-validating results.
When to Choose Salmon
Salmon is the preferred choice for most new RNA-seq projects, particularly those with large datasets or limited computational resources. The speed advantage of Salmon allows for faster iteration and more comprehensive analysis, which can be valuable when exploring different parameters or analyzing multiple datasets.
Salmon's built-in bias correction provides improved accuracy for datasets with known biases, and the equivalence class output enables additional analyses that are not possible with RSEM. The integration with tximport and modern differential expression tools makes Salmon a natural fit for contemporary RNA-seq workflows.
For researchers working with non-model organisms or organisms with incomplete annotations, Salmon's ability to use decoy sequences can improve quantification accuracy by reducing spurious mappings. This feature is particularly useful when the reference transcriptome does not fully represent the organism's transcript diversity.
Hybrid Approaches
Some workflows use both RSEM and Salmon in a complementary manner. For example, you might use Salmon for initial quantification and exploratory analysis, then use RSEM with STAR alignments for final quantification when you need alignment files for other purposes.
The nf-core community provides pipeline standards and configuration guidance that can help you implement hybrid workflows. Many nf-core pipelines support both tools and allow you to switch between them based on your specific requirements.
The choice between RSEM and Salmon is not permanent. You can run both tools on the same dataset and compare the results to assess the impact of the quantification method on your biological conclusions. This approach can be valuable for validating results and understanding the sensitivity of your analysis to the quantification choice.
Observations and Measurements for Quality Assessment
Alignment and Mapping Statistics
Both RSEM and Salmon produce statistics that help you assess the quality of your quantification. Review these statistics carefully to identify potential problems with your data or reference.
For RSEM, the .stat file contains information about the total number of reads, the number of aligned reads, and the number of multi-mapping reads. A low overall alignment rate may indicate problems with the reference, contamination in the sample, or issues with read preprocessing.
For Salmon, the aux_info/meta_info.json file contains similar statistics, including the number of reads processed, the mapping rate, and the number of fragments assigned to transcripts. Salmon also reports the fraction of reads that map to multiple transcripts, which can indicate ambiguity in the transcriptome.
Compare these statistics across samples in your dataset to identify outliers that may indicate sample quality issues. Samples with substantially lower mapping rates or higher multi-mapping rates may need to be excluded or re-processed.
Expression Distribution Checks
After quantification, examine the distribution of expression values across genes and transcripts. Most datasets should show a roughly bimodal distribution, with a large number of lowly expressed features and a smaller number of highly expressed features.
Check for anomalies such as an excessive number of zero-count genes or an unusual number of very highly expressed transcripts. These patterns may indicate problems with the reference annotation, library preparation, or quantification.
The EMBL-EBI Training provides bioinformatics learning pathways and practical analysis education that can help you develop skills for assessing quantification quality. These resources cover topics such as expression distribution analysis and quality control for RNA-seq data.
Reproducibility Checks
Reproducibility is a critical concern in RNA-seq analysis. Run the same quantification pipeline on technical replicates to assess the variability introduced by the quantification step. The variability between technical replicates should be substantially lower than the variability between biological replicates.
The Galaxy Training Network provides accessible workflow training and analysis tutorials that emphasize reproducibility in bioinformatics analysis. Following these practices can help you ensure that your quantification results are reproducible and reliable.
Document all parameters used in your quantification, including the reference version, tool version, and any non-default settings. This documentation is essential for reproducing your analysis and for interpreting results in the context of your specific workflow.
Records and Documentation Requirements
Maintaining Analysis Records
Proper documentation of your RNA-seq analysis is essential for reproducibility and for meeting publication requirements. Maintain detailed records of all steps in your quantification workflow, including:
- Tool versions and installation methods
- Reference transcriptome version and source
- Parameters used for indexing and quantification
- Computational resources used
- Date and time of analysis
The Carpentries Lessons provide foundational computing and data management training that emphasizes good practices for documentation and reproducibility. These skills are essential for managing complex bioinformatics workflows.
Version Control for Analysis Scripts
Use version control systems such as Git to track changes to your analysis scripts and configuration files. This practice allows you to reproduce previous analyses and to understand how changes in parameters or tools affect your results.
The Carpentries Lessons include training on Git and version control, which can help you implement these practices in your own workflow. Version control is particularly important when collaborating with other researchers or when revisiting analyses after a period of time.
Data Management and Storage
RNA-seq datasets can be large, and proper data management is essential for efficient analysis and long-term storage. Plan your data storage strategy before beginning your analysis, considering factors such as:
- Storage capacity for raw and processed data
- Backup procedures for critical files
- Data organization and naming conventions
- Metadata documentation for samples and experiments
The NCBI provides official descriptions of its data resources and submission guidelines, which can help you prepare your data for deposition in public databases. Many journals require data deposition as a condition of publication, so plan for this requirement early in your project.
Common Failure Patterns and Troubleshooting
Low Mapping Rates
Low mapping rates are a common problem in RNA-seq quantification. If a substantial fraction of reads do not map to the reference transcriptome, investigate potential causes:
- Contamination with genomic DNA or other organisms
- Adapter contamination that was not removed during preprocessing
- Reference transcriptome that does not match the organism or tissue being studied
- Sequencing errors or poor read quality
Check the quality of your raw reads and the effectiveness of your preprocessing steps. The Galaxy Training Network provides workflow training that covers quality control procedures for RNA-seq data, which can help you identify and address these issues.
Excessive Multi-Mapping Reads
A high fraction of multi-mapping reads can indicate problems with the reference annotation or with the quantification approach. Multi-mapping reads are reads that map to multiple transcripts, which creates ambiguity in the quantification.
If you observe excessive multi-mapping, consider:
- Using a more complete or updated reference annotation
- Filtering low-complexity reads that may map spuriously
- Using Salmon's selective alignment to improve discrimination between similar transcripts
The choice of reference annotation can have a substantial impact on multi-mapping rates. Compare results with different annotation versions to assess the sensitivity of your analysis to this choice.
Discrepancies Between Tools
If you run both RSEM and Salmon on the same dataset and observe substantial discrepancies in the results, investigate the source of the differences. Potential causes include:
- Differences in bias correction approaches
- Differences in how multi-mapping reads are resolved
- Differences in the handling of reads that map to unannotated regions
Compare the results at the gene level and transcript level to identify where the discrepancies occur. This analysis can help you understand the strengths and limitations of each tool for your specific dataset.
Memory and Computational Failures
Both RSEM and Salmon can fail due to insufficient memory or computational resources. If you encounter memory errors, consider:
- Reducing the number of threads used
- Using a smaller reference index
- Processing data in smaller batches
- Using a machine with more memory
The nf-core documentation provides configuration guidance that can help you optimize resource allocation for RNA-seq analysis. Many cloud computing platforms offer instances with varying memory and CPU configurations, allowing you to select the appropriate resources for your analysis.
Limitations and Interpretation Boundaries
Isoform Quantification Uncertainty
Isoform-level quantification is inherently uncertain, particularly for genes with many isoforms that share exonic sequences. Both RSEM and Salmon provide estimates of transcript abundance, but these estimates have associated uncertainty that should be considered when interpreting results.
The posterior probability estimates from RSEM and the equivalence class information from Salmon can help you assess the confidence in specific isoform assignments. For transcripts with high ambiguity, the estimated abundances may be unreliable, and conclusions based on these estimates should be interpreted with caution.
Reference Annotation Dependence
The accuracy of both RSEM and Salmon depends heavily on the quality and completeness of the reference transcriptome. Incomplete annotations can lead to inaccurate quantification, particularly for novel isoforms or genes that are not well represented in the reference.
For organisms with incomplete annotations, consider using approaches that can identify novel transcripts, such as reference-free assembly or genome-guided assembly. These approaches can complement reference-based quantification and provide a more complete picture of the transcriptome.
Batch Effects and Technical Variability
Quantification results can be affected by batch effects and technical variability that are not related to biological differences. These effects can arise from differences in library preparation, sequencing runs, or other technical factors.
Statistical methods for batch effect correction, such as those implemented in ComBat or RUVSeq, can help account for these effects in downstream analyses. The Bioconductor project provides documentation for many of these methods, which can help you implement appropriate corrections for your data.
Interpretation of TPM Values
TPM values are commonly used to compare expression levels across samples and genes. However, TPM values are relative measures that depend on the total transcriptome composition, and they should not be interpreted as absolute measures of transcript abundance.
When comparing TPM values across samples, be aware that differences in total RNA composition can affect the relative values. For differential expression analysis, count-based methods that account for library size and composition are generally preferred over direct comparisons of TPM values.
Safety and Regulatory Context
Data Privacy and Confidentiality
RNA-seq data can contain sensitive information, particularly when derived from human subjects or endangered species. Ensure that your data handling procedures comply with relevant regulations and institutional policies.
For human data, de-identification and secure storage are essential. For data from endangered species, consider whether publication of the data could pose risks to the species. The NCBI provides guidance on data submission and access controls that can help you manage these considerations.
Computational Security
Bioinformatics analyses often involve downloading software and data from the internet, which can pose security risks. Use trusted sources for software and data, and verify the integrity of downloaded files.
The Carpentries Lessons provide foundational computing training that covers security best practices for research computing. These practices include using secure connections, managing passwords, and verifying software authenticity.
Ethical Use of Computational Resources
Large-scale RNA-seq analyses can consume substantial computational resources. Be mindful of your institution's resource usage policies and consider the environmental impact of your computations.
Optimize your workflows to minimize unnecessary computation, and consider using shared computational resources efficiently. The nf-core documentation provides guidance on resource-efficient pipeline configuration that can help you reduce waste.
Professional Escalation Criteria
When to Seek Expert Assistance
Some RNA-seq analysis challenges require specialized expertise. Consider seeking assistance from bioinformatics core facilities, collaborators, or consultants when:
- You encounter persistent technical problems that you cannot resolve
- Your analysis requires specialized methods beyond standard workflows
- You need to interpret complex results that have substantial biological implications
- You are planning a large-scale study that requires careful experimental design
The EMBL-EBI Training provides bioinformatics learning pathways that can help you develop the skills needed for many RNA-seq analyses. However, some challenges require hands-on assistance from experienced practitioners.
Validating Results with Independent Methods
When quantification results have important implications, consider validating them with independent methods. Options include:
- Quantitative PCR to validate expression differences for specific genes
- Alternative quantification tools to confirm results
- Independent biological replicates to confirm reproducibility
The choice of validation approach depends on your specific research question and the resources available. For high-stakes conclusions, independent validation is strongly recommended.
Consulting Statistical Expertise
Differential expression analysis involves complex statistical considerations that may require specialized expertise. If you are uncertain about the appropriate statistical methods for your data, consult with a statistician or bioinformatics expert.
The Bioconductor project provides documentation for many statistical methods used in RNA-seq analysis, which can help you understand the assumptions and limitations of different approaches. However, expert guidance can be valuable for designing appropriate analyses and interpreting results.
A Practical Decision Framework for Selecting RSEM or Salmon
Beyond the technical comparison of alignment strategies and speed, researchers need a structured method for choosing between RSEM and Salmon that accounts for their specific project constraints. The following decision framework organizes the selection process around five assessment points that can be evaluated before committing to a quantification pipeline.
Step 1: Inventory Your Downstream Analysis Requirements
Begin by listing every analysis you plan to perform after quantification. Create a table with three columns: the analysis type, whether it requires alignment files, and the specific tool you intend to use. Common downstream analyses include differential expression, variant calling, fusion detection, and genome browser visualization.
If your list includes variant calling, fusion detection, or any analysis that requires BAM files, RSEM becomes the more practical choice because it produces alignment files as standard output. If your only downstream analysis is differential expression using DESeq2, edgeR, or limma, Salmon integrates more directly through the tximport package. The Bioconductor project provides official package documentation for tximport and the differential expression tools that accept its output, which can help you confirm compatibility before you begin.
Step 2: Assess Your Computational Environment
Evaluate the computational resources available for your analysis. Record the number of CPU cores, available memory, and the time limit for job completion on your computing cluster or workstation. For a typical dataset with 20 to 30 million paired-end reads, Salmon can complete quantification in 10 to 20 minutes, while RSEM with STAR alignment may require 1 to 2 hours. This difference becomes more pronounced with larger datasets or when processing many samples.
If you are working on a standard laptop or a shared server with limited memory, Salmon's compact index requires less memory than the full genome index used by STAR for RSEM. The nf-core documentation provides community pipeline standards and configuration guidance that can help you estimate resource requirements for both tools across different dataset sizes.
Step 3: Evaluate Reference Annotation Quality
The quality and completeness of your reference transcriptome affects both tools, but the impact differs. For organisms with well-annotated genomes, both RSEM and Salmon perform reliably. For non-model organisms or species with incomplete annotations, Salmon's ability to use decoy sequences can reduce spurious mappings to genomic regions that share sequence similarity with transcripts.
Check the annotation version and its release date. If your organism has frequent annotation updates, consider whether the latest version includes isoforms that are relevant to your experimental context. The NCBI provides official descriptions of its sequence resources and annotation databases, which can help you identify the appropriate reference files for your organism.
Step 4: Consider Dataset Size and Number of Samples
The total number of reads across all samples in your study influences the practical difference between RSEM and Salmon. For small pilot studies with a few samples, the speed advantage of Salmon may be less critical. For large studies with dozens or hundreds of samples, the cumulative time savings from Salmon can be substantial.
Calculate the estimated total processing time for each tool by multiplying the per-sample time by the number of samples. Include time for indexing and any re-runs you might need if you adjust parameters. This calculation provides a concrete basis for deciding whether the speed advantage of Salmon justifies any workflow changes.
Step 5: Document Your Decision and Validation Plan
Record your tool selection, the rationale for the choice, and the specific version numbers of all software involved. Include the reference transcriptome version, the parameters used for indexing and quantification, and the date of analysis. This documentation supports reproducibility and helps you or your collaborators understand the analysis context when revisiting the data later.
The Carpentries Lessons provide foundational training on documentation practices and version control that can help you establish a reliable record-keeping system for your bioinformatics workflows.
A Structured Comparison Record for Your Project
Use the following record template to document your tool evaluation for each RNA-seq project. This record serves as a reference for future projects and helps standardize decision-making across your research group.
| Decision Point | RSEM | Salmon | Your Project Context |
|---|---|---|---|
| Alignment files needed for other analyses | Yes, BAM files produced | No in quasi-mapping mode | List specific analyses requiring alignments |
| Estimated processing time for all samples | Hours to days depending on dataset | Minutes to hours | Calculate based on your sample count |
| Memory available on analysis machine | Moderate to high | Lower | Record available RAM |
| Reference annotation completeness | Works with standard annotations | Decoy sequences help with incomplete annotations | Note annotation version and known gaps |
| Downstream differential expression tool | EBSeq, edgeR, DESeq2 with conversion | tximport for DESeq2, edgeR, limma, sleuth | List your planned tools |
| Bias correction needs | Limited built-in correction | Built-in sequence, GC, and positional bias correction | Note library preparation method |
Complete this table before running either tool. The process of filling in the context column forces you to make explicit decisions about your analysis requirements instead of defaulting to a familiar tool.
Troubleshooting Method for Quantification Discrepancies
When results from RSEM and Salmon disagree substantially, use this systematic troubleshooting approach to identify the source of the discrepancy.
Step 1: Compare at the Gene Level First
Start by comparing gene-level summaries instead of transcript-level estimates. Gene-level comparisons are more robust because they aggregate across isoforms and reduce the impact of isoform assignment differences. Calculate the correlation between gene-level TPM values from both tools across all samples. A high correlation at the gene level with low correlation at the transcript level points to isoform assignment as the source of disagreement.
Step 2: Identify the Transcripts with Largest Differences
Rank transcripts by the absolute difference in estimated TPM between the two tools. Examine the top 50 transcripts with the largest differences. For each transcript, check the number of isoforms in its gene and the sequence similarity among those isoforms. Transcripts from genes with many highly similar isoforms are the most likely to show substantial differences between tools.
Step 3: Check Read Coverage and Multi-Mapping Rates
For the transcripts with large discrepancies, examine the read coverage and multi-mapping rates reported by each tool. RSEM reports posterior probability estimates for read assignments, while Salmon provides equivalence class information. High multi-mapping rates for specific transcripts indicate that the quantification is uncertain regardless of the tool used.
Step 4: Assess Bias Correction Effects
If the discrepancies concentrate in transcripts with extreme GC content or specific sequence features, the difference may stem from Salmon's built-in bias correction. Run Salmon with bias correction disabled using the appropriate flag and compare the results. If disabling bias correction reduces the discrepancy, the difference is likely due to bias correction instead of fundamental methodological differences.
Step 5: Validate with an Independent Method
For transcripts where the tools disagree and the biological conclusion depends on the result, validate with an independent method such as quantitative PCR. This validation is particularly important when the transcript of interest has clinical, agricultural, or conservation implications. The choice of validation approach depends on your specific research question and available resources.
Common Failure Patterns in Tool Selection
Pattern 1: Choosing RSEM for Speed-Sensitive Projects
Researchers sometimes default to RSEM because they are familiar with alignment-based workflows, only to find that the processing time creates bottlenecks for large datasets. If your project involves more than 50 samples and you do not need alignment files for other analyses, Salmon is generally the more practical choice. The time savings allow for faster iteration and more comprehensive parameter exploration.
Pattern 2: Choosing Salmon Without Checking Alignment Needs
The reverse failure occurs when researchers choose Salmon for its speed but later discover they need alignment files for variant calling or visualization. If you anticipate needing alignments for any downstream analysis, factor this into your decision from the start. Running RSEM after Salmon to generate alignments requires additional computational time and may produce different quantification results than the Salmon output you already have.
Pattern 3: Ignoring Reference Annotation Quality
Both tools depend on the quality of the reference transcriptome, but researchers sometimes overlook this dependency when comparing tools. If your organism has a poorly annotated genome, the choice of tool matters less than the choice of reference. Invest time in selecting the best available annotation and consider whether transcriptome assembly or augmentation is needed before quantification.
Pattern 4: Failing to Document the Decision
Without documentation of why a particular tool was chosen, future researchers may repeat the evaluation process or question the validity of the results. Record your decision framework, the alternatives considered, and the specific context that led to your choice. This documentation is particularly important for multi-year projects where personnel may change.
Welfare and Reproducibility Context
The choice between RSEM and Salmon has implications beyond computational efficiency. Reproducibility of RNA-seq analysis depends on transparent documentation of all analysis steps, including the quantification tool and its parameters. The Galaxy Training Network provides accessible workflow training that emphasizes reproducible analysis practices, which can help you establish a reliable pipeline regardless of the tool you select.
For studies involving agricultural species or conservation biology, the accuracy of isoform quantification can affect conclusions about gene expression patterns that inform breeding decisions or conservation strategies. The EMBL-EBI Training provides bioinformatics learning pathways that cover quality assessment and interpretation of quantification results, which can help you avoid common pitfalls in drawing biological conclusions from your data.
When your analysis informs decisions about animal management, breeding programs, or conservation actions, consider validating key findings with independent methods. The cost of validation is small compared to the potential consequences of acting on incorrect quantification results.
Frequently Asked Questions
What is the main difference between RSEM and Salmon?
RSEM uses a full alignment approach, requiring reads to be aligned to the reference transcriptome using an external aligner such as Bowtie or STAR before quantification. Salmon uses quasi-mapping, which identifies the likely origin of each read by matching k-mers without computing full alignments. This makes Salmon substantially faster while producing comparable accuracy for most datasets.
Which tool is faster for isoform quantification?
Salmon is typically 10-20 times faster than RSEM because quasi-mapping avoids the computational cost of full alignment. For a standard dataset with 20-30 million paired-end reads, Salmon can complete quantification in 10-20 minutes, while RSEM with STAR alignment may take 1-2 hours.
Can I use Salmon results with DESeq2 or edgeR?
Yes, Salmon results can be imported into R using the tximport package, which prepares count matrices for DESeq2, edgeR, limma, and other differential expression tools. The Bioconductor project provides documentation for this workflow, which is widely used in the bioinformatics community.
Does RSEM produce alignment files that can be used for other analyses?
Yes, RSEM produces BAM and SAM alignment files as part of its standard output. These files can be used for variant calling, visualization in genome browsers, and other analyses that require aligned reads. Salmon in quasi-mapping mode does not produce alignment files.
Which tool handles multi-mapping reads better?
Both tools use expectation-maximization algorithms to resolve multi-mapping reads, but they differ in their approach. Salmon's selective alignment can improve discrimination between similar transcripts, while RSEM's generative model accounts for read length distributions and sequencing errors. For most datasets, the accuracy differences are small.
Do I need to perform quality control before using RSEM or Salmon?
Yes, quality control is essential before quantification. Raw reads should be assessed for quality, and adapter sequences and low-quality bases should be removed. The Galaxy Training Network provides workflow training that covers standard quality control procedures for RNA-seq data.
Can I use both RSEM and Salmon on the same dataset?
Yes, you can run both tools on the same dataset and compare the results. This approach can help you assess the sensitivity of your analysis to the quantification method and validate your conclusions. Discrepancies between tools may indicate areas of uncertainty in the quantification.
What reference transcriptome should I use for quantification?
The choice of reference transcriptome depends on your organism and research question. For model organisms, use the latest annotation from public databases such as NCBI or Ensembl. For non-model organisms, you may need to construct a reference from transcriptome assembly or use a closely related species as a reference.
Related Bioinformatics Guides
- RNA-Seq Alignment: Choosing the Right Tool and Parameters
- RNA-Seq Alignment Tools: STAR, HISAT2, and Beyond
- RNA-Seq Quality Control: Essential Checks and Tools
- RNA-Seq vs qPCR: Validation and Comparison
- RNA-Seq Batch Effect Detection and Correction
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- Dynamic changes of endogenous hormones and transcriptome analysis reveal the regulatory mechanism of gibberellin-promoted male cone formation in Platycladus orientalis.. 2026.
- Salmonids reveal principles of regulatory evolution following autotetraploidization. 2026.
- Application of genomic tools to study and potentially improve the upper thermal tolerance of farmed Atlantic salmon (Salmo salar).. 2025.
- Comparative transcriptomic analysis of embryonic stem cells across mammalian species.. 2026.
- LncRNA and mRNA expression characteristic and bioinformatic analysis in anemic diabetic foot ulcers.. 2025.
- Commentary: a review of technical considerations for planning an RNA-Sequencing experiment.. 2025.
- Fishy business in Seattle: Salmon mislabeling fraud in sushi restaurants vs grocery stores. PLoS ONE, 2024.
- Sampling of Atlantic salmon using the Norwegian Quality cut (NQC) vs. Whole fillet, differences in contaminant and nutrient contents.. Food Chemistry, 2023.
- Effects of Yellowstripe Scad vs Salmon Intake Intervention among Overweight Malaysian Adults. 2017.
- Hydropower vs. Salmon: The Struggle of the Pacific Northwest's Anadromous Fish Resources for a Peaceful Coexistence with the Federal Columbia River Power System. 1981.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.