# Why Are My Isoform Quantification Results Unreliable? Troubleshooting Common Issues in Salmon and RSEM


## Key Takeaways

- **Reference Transcriptome Integrity is Paramount:** Unreliable isoform quantification often stems from an incomplete, outdated, or misannotated reference transcriptome. Mismatches between the annotation version and genome build, or the presence of duplicate/highly similar sequences, lead to low mapping rates, misassigned reads, and inflated counts for incorrect isoforms. Always verify and document the exact reference version and source.

- **Library Preparation and Sequencing Parameters are Critical:** Incorrectly specifying library parameters such as strandedness, read orientation, and read length to Salmon or RSEM will result in erroneous fragment length distribution estimation and skewed isoform counts. Mismatched parameters, particularly for stranded libraries, significantly increase ambiguity and reduce the ability to differentiate isoforms transcribed from opposite strands.

- **Bias Correction is Essential for Accurate Isoform Abundance:** Systematic biases introduced during library preparation and sequencing, such as GC-content and sequence-specific biases, must be corrected. Disabling or misapplying these corrections, as implemented in Salmon, leads to systematic over- or underestimation of transcript abundance, particularly impacting isoforms with differing sequence compositions.

- **Quantification Mode Consistency is Non-Negotiable:** The choice between alignment-free (Salmon) and alignment-based (Salmon, RSEM) quantification modes must be consistent throughout a study. Inconsistent modes lead to differing count matrices for the same sample, compromising downstream differential expression analysis. Alignment-free modes sacrifice positional information, potentially impacting resolution for highly similar isoforms.

- **Transcript-Level Uncertainty Must Be Quantified and Accounted For:** Ignoring transcript-level uncertainty, often represented by equivalence classes in Salmon, leads to unreliable differential expression calls. Utilizing equivalence class information, running multiple random seeds, and incorporating bootstrap estimates into downstream analyses (e.g., with sleuth or tximport) are crucial for robust interpretation and avoiding false positives.

---

Isoform quantification with Salmon or RSEM produces unreliable results when the reference transcriptome is incomplete or misannotated, when parameters do not match the library preparation and sequencing chemistry, when bias correction is disabled or misapplied, when the quantification mode is inconsistent with the downstream analysis, or when the results are interpreted without accounting for transcript-level uncertainty. This article walks through the specific failure points in a typical RNA-seq workflow, the decisions that cause them, and the records and checks that separate a reproducible quantification run from one that cannot be trusted.

The scope here covers the practical pipeline from raw FASTQ files to transcript-level count matrices. The focus is on Salmon and RSEM because they are the two most widely used lightweight quantification tools in modern RNA-seq analysis. The troubleshooting guidance applies to bulk RNA-seq of purified cell populations, sorted subpopulations, and tissue samples. It also applies to protocols that isolate specific cell types for downstream sequencing, such as the RNA FISH-based cell purification approach described for planaria, where the quality of the input material directly determines whether transcript-level estimates carry biological meaning [<a href="#ref-1">1</a>].

## At a Glance

The table below summarizes the most common causes of unreliable isoform quantification, the observable symptom in the output files, and the first action to take.

| Failure Point | Observable Symptom | First Action |
| --- | --- | --- |
| Incomplete or outdated reference transcriptome | Low mapping rate, many reads assigned to no feature, inflated counts for annotated isoforms | Rebuild the transcriptome index from the current annotation release and verify the version matches the genome build |
| Mismatched library parameters | Fragment length distribution looks wrong, mapping rate drops, counts shift between isoforms | Confirm the library type, strandedness, and read orientation against the kit documentation and set the corresponding Salmon or RSEM flags |
| Disabled or misapplied bias correction | Systematic overestimation of GC-rich or sequence-specific transcripts | Enable sequence-specific bias correction in Salmon and verify the GC bias model is active in the output logs |
| Inconsistent quantification mode | Counts differ between alignment-based and alignment-free modes for the same sample | Choose one mode for the entire study and document the exact command used for every sample |
| Ignoring transcript-level uncertainty | Differential expression calls change when the same data is rerun with a different random seed | Use equivalence class information, run multiple seeds, and report the range of estimates |
| Reference contamination or duplication | Unexpected isoforms with identical sequences, inflated counts for paralogous genes | Check the transcriptome for duplicate sequences and remove or flag them before indexing |

## The Reference Transcriptome Is the Foundation of Every Quantification Run

The reference transcriptome determines what Salmon and RSEM can detect. If a transcript is not in the reference, no read originating from that transcript will map to it. If a transcript sequence is wrong, reads will map with mismatches and the quantification algorithm will distribute those reads incorrectly. If the reference contains duplicate or near-identical sequences, reads will be assigned ambiguously and the counts will be split between the wrong isoforms.

### Annotation Version and Genome Build Must Match

The most common reference-related failure is a mismatch between the annotation version and the genome build. A transcriptome extracted from an annotation file that was built against an older genome assembly will contain sequences that do not align cleanly to the current genome. When Salmon or RSEM quantifies against this transcriptome, reads from regions that were reannotated or reassembled will map poorly or not at all.

The fix is to download the transcriptome and annotation from the same source and the same release. The National Center for Biotechnology Information maintains the major sequence databases and search systems that researchers use to obtain reference genomes and annotations [<a href="#ref-2">2</a>]. When you download a reference transcriptome from NCBI, record the assembly accession, the annotation release number, and the date of download. These three pieces of information must be stored with the quantification results.

For non-model organisms, the reference may be less complete. The planaria protocol example shows that even with a well-designed cell isolation method, the downstream sequencing analysis depends on the quality of the reference [<a href="#ref-1">1</a>]. If the reference transcriptome for a non-model organism is assembled from short reads, it will contain gaps, chimeric transcripts, and incomplete isoforms. Quantification against such a reference produces counts that reflect the assembly errors as much as the biological expression levels.

### Transcriptome Indexing Is a One-Time Decision That Affects Every Sample

Salmon and RSEM both require an index built from the reference transcriptome. The index is built once and reused for all samples in the study. This means the index building step is a single point of failure. If the index is built from a truncated or corrupted reference, every sample quantified against that index will be wrong.

The index should be built with the same version of the tool that will be used for quantification. Salmon and RSEM both change their index formats between major versions. An index built with an older version may not be readable by a newer version of the tool, or worse, it may be readable but produce subtly different results because the newer version expects a different index structure.

Record the exact command used to build the index, including the tool version, the reference file path, and any parameters such as k-mer size. The k-mer size in Salmon affects the sensitivity and specificity of the mapping step. A k-mer size that is too long will miss short reads or reads with many errors. A k-mer size that is too short will produce many spurious matches and increase the ambiguity in the equivalence classes.

### Duplicate and Highly Similar Sequences Confuse the Quantification

Paralogous genes and retained duplicate copies of the same gene produce transcripts with high sequence similarity. When these transcripts are in the reference, reads that originate from one copy will map to both copies. Salmon and RSEM handle this ambiguity by distributing the reads according to the estimated abundance of each transcript, but the estimates are only as good as the model.

If the reference contains exact duplicate transcript sequences, the quantification tools cannot distinguish between them. The reads will be split evenly or according to the prior, which is not the true biological distribution. Check the reference for duplicate sequences before building the index. Remove exact duplicates or keep only one representative transcript per gene if the biological question does not require distinguishing between identical copies.

## Library Preparation and Sequencing Chemistry Determine the Correct Parameters

The quantification tools need to know how the library was prepared and how the sequencer read the fragments. If these parameters are wrong, the fragment length distribution will be estimated incorrectly, the read orientation will be misinterpreted, and the counts will shift between isoforms.

### Strandedness Is the Most Common Parameter Error

RNA-seq libraries can be stranded or unstranded. In a stranded library, the read orientation tells you which strand the RNA came from. In an unstranded library, the read can come from either strand and the orientation carries no information.

Salmon and RSEM both have flags for library type. In Salmon, the library type is specified with the `--libType` flag. In RSEM, the library type is specified with the `--forward-prob` parameter or the `--strandedness` option. If the library is stranded but the quantification is run with the unstranded setting, the tools will treat reads from both strands as equally likely. This doubles the ambiguity for every read and reduces the ability to distinguish between isoforms that are transcribed from opposite strands.

The library type is determined by the kit used for library preparation. The kit documentation states whether the library is stranded and, if so, whether the first strand or the second strand is sequenced. Record the kit name, the lot number, and the protocol version. This information must be stored with the quantification parameters.

### Read Length and Fragment Length Distribution

The fragment length distribution is estimated by the quantification tools from the mapped reads. If the reads are shorter than expected, or if the fragments are longer than the read length, the tools will have less information to work with. The fragment length distribution is used to model the probability that a read came from a particular position on a transcript. If the distribution is wrong, the assignment of reads to isoforms will be wrong.

Salmon estimates the fragment length distribution automatically from the data. RSEM also estimates it, but the estimation is more sensitive to the initial parameters. If the library contains a wide range of fragment lengths, the estimation may converge to a local optimum that does not reflect the true distribution.

The read length is determined by the sequencing platform and the kit. A 150-base-pair paired-end read provides more information than a 75-base-pair single-end read. The choice of read length affects the ability to distinguish between isoforms that differ only in a short exon. If the isoforms of interest differ by a region shorter than the read length, the quantification will be unreliable regardless of the tool.

### Paired-End Reads Must Be Treated as Paired

Salmon and RSEM both support paired-end reads, but the pairing information must be preserved in the input files. If the FASTQ files are not properly paired, or if the read names do not match between the two files, the tools will treat the reads as single-end and lose the fragment length information.

Check the read names in the paired FASTQ files before running the quantification. The names must match exactly between the two files, and the order must be the same. If the files were processed with a tool that reordered the reads, the pairing is broken and the quantification will be unreliable.

## Bias Correction Models Are Not Optional for Accurate Isoform Counts

RNA-seq data contains systematic biases that are introduced during library preparation and sequencing. These biases affect the number of reads that map to each transcript, independent of the true expression level. Salmon was the first transcriptome-wide quantifier to correct for fragment GC-content bias, and this correction substantially improves the accuracy of abundance estimates and the sensitivity of subsequent differential expression analysis [<a href="#ref-3">3</a>].

### Sequence-Specific Bias

The sequence-specific bias model in Salmon accounts for the fact that the sequencing adapters and the library preparation enzymes have preferences for certain sequences. These preferences cause some transcripts to be sequenced more or less often than their true abundance would predict.

The sequence-specific bias model is enabled by default in Salmon, but it can be disabled. If it is disabled, the quantification will overestimate the abundance of transcripts that are favored by the bias and underestimate the abundance of transcripts that are disfavored. The effect is larger for isoforms that differ in their sequence composition, which is exactly the case for alternative isoforms of the same gene.

### GC-Content Bias

The GC-content bias model accounts for the fact that fragments with extreme GC content are amplified less efficiently during library preparation. This bias affects the number of reads that map to each transcript, and it is particularly strong for transcripts with high or low GC content.

Salmon corrects for GC-content bias using a model that is learned from the data. The model is fit during the quantification run, and the fitted parameters are reported in the output. Check the output logs to confirm that the GC bias model was fit and that the parameters are in a reasonable range. If the model did not converge, or if the parameters are extreme, the quantification may be unreliable.

### Bias Correction in RSEM

RSEM does not have the same bias correction models as Salmon. RSEM uses a different approach to handle biases, and it is less effective at correcting for GC-content bias. If the library has strong GC bias, RSEM will produce less accurate estimates than Salmon.

The choice between Salmon and RSEM should be based on the specific requirements of the analysis. Salmon is faster and has more sophisticated bias correction. RSEM provides a more detailed model of the read generation process and can be used with the EBSeq package for differential expression analysis. Both tools are valid, but they are not interchangeable. The choice must be made before the analysis starts and documented in the methods.

## Quantification Mode and Downstream Analysis Must Be Consistent

Salmon can be run in alignment-based mode or alignment-free mode. In alignment-based mode, reads are first aligned to the genome or transcriptome with an external aligner, and the alignments are passed to Salmon. In alignment-free mode, Salmon maps the reads directly to the transcriptome using its own lightweight mapping procedure.

RSEM requires alignments. RSEM uses the Bowtie or Bowtie2 aligner to map reads to the transcriptome, and then it estimates abundances from the alignments. The alignment step is a separate process that must be run before RSEM.

### Alignment-Free Mode Is Faster but Has Different Failure Modes

Salmon's alignment-free mode is fast because it does not produce full alignments. Instead, it maps reads to transcripts using a k-mer-based approach and then uses the equivalence classes to estimate abundances. The equivalence classes group reads that map to the same set of transcripts. This grouping is the key to Salmon's speed, but it also means that the information about the exact position of each read is lost.

The loss of positional information does not affect the abundance estimates for most transcripts, but it can affect the estimates for transcripts that have very similar sequences. If two isoforms differ only in a small region, the reads from that region will be in the same equivalence class, and the abundance will be distributed between the two isoforms based on the model.

### Alignment-Based Mode Provides More Information

In alignment-based mode, the aligner provides the exact position of each read on the transcriptome. This information is used by Salmon to improve the abundance estimates. The alignment-based mode is more accurate for transcripts with high sequence similarity, but it requires an additional alignment step and is slower.

The choice between alignment-free and alignment-based mode should be based on the complexity of the transcriptome and the similarity of the isoforms of interest. For a well-annotated transcriptome with distinct isoforms, alignment-free mode is sufficient. For a transcriptome with many highly similar isoforms, alignment-based mode may be necessary.

### RSEM Requires Alignments and Has Its Own Parameters

RSEM uses the alignments from Bowtie or Bowtie2 to estimate abundances. The alignment parameters affect the quality of the RSEM estimates. If the aligner is run with parameters that are too permissive, reads will map to multiple transcripts and the ambiguity will increase. If the aligner is run with parameters that are too strict, reads will not map and the mapping rate will drop.

RSEM has its own parameters for the abundance estimation step. The `--estimate-rspd` parameter enables the estimation of the read start position distribution, which improves the accuracy of the estimates. The `--calc-ci` parameter enables the calculation of confidence intervals for the abundance estimates. These parameters should be enabled for the most accurate results.

## Equivalence Classes and Transcript-Level Uncertainty

The equivalence classes are the core data structure that Salmon uses to estimate abundances. Each equivalence class contains the reads that map to the same set of transcripts. The size of the equivalence class and the number of transcripts in it determine the uncertainty of the abundance estimates.

### The Number of Transcripts in an Equivalence Class Affects the Uncertainty

If an equivalence class contains reads that map to a single transcript, the abundance of that transcript is known with high confidence. If an equivalence class contains reads that map to multiple transcripts, the abundance of each transcript is uncertain. The uncertainty is higher when the transcripts in the equivalence class have similar sequences and when the reads are distributed evenly between them.

Salmon reports the uncertainty in the form of the Gibbs sampling output or the bootstrap estimates. The bootstrap estimates provide a range of possible abundance values for each transcript. If the range is wide, the abundance estimate is unreliable.

### Running Multiple Seeds Reveals the Uncertainty

Salmon uses a random seed for the optimization and the Gibbs sampling. Different seeds produce slightly different results. If the results differ substantially between seeds, the quantification is unstable and the estimates are unreliable.

Run Salmon with multiple seeds and compare the results. The European Bioinformatics Institute provides training materials on RNA-seq analysis that cover the practical aspects of running quantification tools and interpreting the output [<a href="#ref-4">4</a>]. The training materials emphasize the importance of reproducibility and the need to document the exact parameters used for each run.

### The Bootstrap Estimates Should Be Used in Downstream Analysis

The bootstrap estimates from Salmon can be used in downstream differential expression analysis. Tools like sleuth and tximport can incorporate the bootstrap estimates to account for the uncertainty in the abundance estimates. If the bootstrap estimates are ignored, the differential expression analysis will treat the abundance estimates as if they were exact, which will produce false positives for transcripts with high uncertainty.

The number of bootstrap replicates should be set based on the complexity of the transcriptome and the desired level of confidence. More bootstrap replicates provide a more accurate estimate of the uncertainty, but they require more computational time. A typical number is 100 bootstrap replicates, but the optimal number depends on the specific analysis.

## Quality Control Checks Before and After Quantification

Quality control is not a single step in the workflow. It is a continuous process that starts with the raw sequencing data and continues through the quantification and downstream analysis. The checks described below should be performed at each stage, and the results should be recorded.

### Raw Read Quality

The raw reads should be checked for quality before quantification. The quality scores, the GC content, the adapter contamination, and the duplication rate should be examined. If the reads have low quality scores, the quantification will be unreliable because the mapping step will fail or produce incorrect alignments.

The Galaxy Training Network provides accessible workflow training that covers the quality control steps in RNA-seq analysis [<a href="#ref-5">5</a>]. The training materials show how to use FastQC and other tools to assess the quality of the raw reads and how to trim or filter reads that do not meet the quality thresholds.

### Mapping Rate

The mapping rate is the percentage of reads that map to the reference transcriptome. A low mapping rate indicates a problem with the reference, the reads, or the parameters. A mapping rate below 50 percent for a well-annotated organism suggests that the reference is incomplete or that the reads are contaminated with non-transcriptomic sequences.

The mapping rate should be recorded for each sample. If the mapping rate varies substantially between samples, the variation may indicate a problem with one of the samples, such as degradation or contamination.

### Count Distribution

The distribution of counts across transcripts should be examined after quantification. Most transcripts should have low counts, and a small number of transcripts should have high counts. If the distribution is unusual, such as many transcripts with zero counts or a few transcripts with extremely high counts, the quantification may be unreliable.

The count distribution can be examined with a histogram or a density plot. The Bioconductor project provides packages for the analysis of count data, including tools for quality control and visualization [<a href="#ref-6">6</a>]. The packages are documented with reproducible examples that show how to assess the quality of the quantification results.

## Common Failure Patterns and Their Causes

The failure patterns described below are the most common in practice. Each pattern has a specific cause and a specific fix.

### Low Mapping Rate with a Complete Reference

If the reference transcriptome is complete and the mapping rate is still low, the problem is likely in the reads or the parameters. The reads may contain adapter contamination, or the library may have been prepared with a protocol that produces reads that do not match the reference.

Check the adapter contamination in the raw reads. If the adapters are present, trim them before quantification. Check the library type and the strandedness. If the library is stranded and the quantification is run with the wrong library type, the mapping rate will drop because the reads will not match the expected orientation.

### Inflated Counts for Annotated Isoforms

If the counts for annotated isoforms are inflated, the reference may contain duplicate sequences or the bias correction may be disabled. Check the reference for duplicate sequences and remove them. Enable the bias correction models in Salmon and verify that they are active in the output logs.

### Inconsistent Results Between Runs

If the results differ between runs of the same data, the quantification is unstable. The instability may be caused by the random seed, the number of bootstrap replicates, or the parameters. Run the quantification with multiple seeds and compare the results. If the results differ substantially, increase the number of bootstrap replicates or change the parameters.

### Counts That Do Not Match the Biological Expectation

If the counts do not match the biological expectation, the problem may be in the reference annotation. The annotation may contain errors, such as incorrect exon boundaries or missing isoforms. The NCBI provides tools for examining the annotation and comparing it to the genome [<a href="#ref-2">2</a>]. If the annotation is incorrect, the quantification will be unreliable regardless of the tool.

## Records and Measurements for Reproducible Quantification

Reproducibility requires records. The records must include the exact versions of the tools, the reference files, the parameters, and the input data. The records must be stored with the results so that anyone can reproduce the analysis.

### Tool Versions

The versions of Salmon, RSEM, and any other tools used in the analysis must be recorded. The versions affect the results because the tools change their algorithms and their default parameters between versions. The version information is available from the tool's documentation or from the command line.

### Reference Files

The reference files must be recorded with their version and their source. The NCBI provides accession numbers for the reference genomes and annotations [<a href="#ref-2">2</a>]. The accession numbers must be recorded with the quantification results. The date of download must also be recorded because the reference files are updated regularly.

### Parameters

The parameters used for the quantification must be recorded. The parameters include the library type, the read length, the k-mer size, the number of bootstrap replicates, and the bias correction settings. The exact command used for the quantification must be recorded, including all flags and their values.

### Input Data

The input data must be recorded with the sample names, the file paths, and the checksums. The checksums ensure that the input data has not been corrupted or changed. The checksums can be calculated with tools like md5sum or sha256sum.

## Professional Escalation Criteria

Some problems cannot be solved by adjusting the parameters or the reference. These problems require the expertise of a bioinformatics specialist or a statistician. The escalation criteria below describe the situations where professional help is needed.

### Persistent Low Mapping Rate

If the mapping rate remains low after checking the reference, the reads, and the parameters, the problem may be in the sequencing data itself. The sequencing run may have failed, or the library may have been prepared incorrectly. A bioinformatics specialist can examine the raw data and determine the cause of the low mapping rate.

### Extreme Bias in the Count Distribution

If the count distribution is extremely biased, with a few transcripts accounting for most of the reads, the problem may be in the library preparation. The library may contain a few highly abundant transcripts that dominate the sequencing. A specialist can determine whether the bias is biological or technical.

### Inconsistent Results Across Tools

If Salmon and RSEM produce substantially different results for the same data, the problem may be in the models or the parameters. A specialist can compare the models and determine which tool is more appropriate for the specific data.

### Results That Contradict Known Biology

If the quantification results contradict known biology, such as the expression of housekeeping genes or the expected response to a treatment, the problem may be in the reference annotation or the experimental design. A specialist can examine the annotation and the experimental design to determine the cause of the contradiction.

## Limitations of Isoform Quantification

Isoform quantification has inherent limitations that cannot be overcome by adjusting the parameters or the reference. These limitations must be understood before the results can be interpreted.

### Short Reads Cannot Distinguish All Isoforms

Short reads cannot distinguish between isoforms that differ only in a region shorter than the read length. If two isoforms differ by a single exon of 50 base pairs, and the reads are 75 base pairs long, the reads that span the exon boundary will distinguish the isoforms, but the reads that fall entirely within the shared exons will not. The quantification tools use the reads that span the exon boundaries to estimate the abundance of each isoform, but the estimates are uncertain.

### The Reference Must Be Complete

The reference transcriptome must be complete for the quantification to be accurate. If the reference is missing isoforms, the reads from those isoforms will map to the closest available isoform, and the counts will be inflated for that isoform. The completeness of the reference is a particular problem for non-model organisms, where the annotation is often incomplete.

### The Quantification Is Only as Good as the Model

The quantification tools use models of the read generation process to estimate abundances. The models are approximations of the true process, and they contain assumptions that may not hold for all data. The bias correction models in Salmon improve the accuracy of the estimates, but they do not eliminate all biases.

## Safety and Regulatory Context

Isoform quantification is a computational analysis, and it does not involve the safety or regulatory concerns of wet-lab experiments. However, the results of the analysis can have regulatory implications if they are used to support a claim about a drug, a diagnostic, or a treatment.

The protocol for purifying and sequencing cell subpopulations based on RNA FISH signal in planaria describes a method for isolating specific cells for downstream RNA sequencing [<a href="#ref-1">1</a>]. The protocol emphasizes the importance of avoiding RNA-degrading formaldehyde during the dissociation and fixation steps. This attention to RNA integrity is critical for the downstream quantification because degraded RNA produces reads that do not map cleanly to the reference.

The structural biology study of insulin receptor antagonists describes a different type of analysis, but it illustrates the importance of accurate quantification in a regulatory context [<a href="#ref-7">7</a>]. The study uses cryo-EM to determine the structure of the insulin receptor bound to antagonists, and the structural insights may inform the development of next-generation treatments for congenital hyperinsulinism. If the quantification of gene expression were used to support a claim about the efficacy of a treatment, the quantification would need to be reproducible and accurate.

The protocol for genome-wide identification of transcription factor binding motifs describes a method for profiling intrinsic TF motifs using pull-down sequencing [<a href="#ref-8">8</a>]. The protocol bypasses antibody dependence and applies to any affinity-tagged DNA-binding protein. The analysis of the sequencing data from this protocol requires accurate quantification of the enriched motifs, and the same troubleshooting principles apply.

The protocol for marker-free genome editing in Saccharomyces cerevisiae describes a method for generating gene deletions using CRISPR-Cas9 [<a href="#ref-9">9</a>]. The verification of the edits by PCR is a quality control step that ensures the edits are correct. The same principle applies to isoform quantification: the results must be verified before they are used for downstream analysis.

The protocol for polysome profiling in differentiating adipocytes describes a method for isolating polysome-bound mRNA for RNA sequencing [<a href="#ref-10">10</a>]. The protocol enables the analysis of actively translating mRNAs and translational gene regulation. The quantification of the polysome-bound mRNA requires the same attention to the reference, the parameters, and the quality control as any other RNA-seq analysis.

## A Practical Decision Framework for Choosing Between Salmon and RSEM

The choice between Salmon and RSEM is often treated as a matter of personal preference, but it is a technical decision that should be driven by the structure of the reference transcriptome, the sequencing library characteristics, and the downstream analysis requirements. A structured decision framework reduces the risk of selecting a tool that cannot handle the specific features of the data. The framework below uses observable properties of the input data and the analysis goals to guide the selection.

### Step 1: Assess the Reference Transcriptome Complexity

The first decision point is the degree of sequence similarity among the isoforms in the reference. Transcriptomes with many paralogous genes, retained duplicate copies, or alternatively spliced isoforms that share long exonic regions present a greater challenge for lightweight quantification tools. The presence of these features can be assessed by examining the annotation file before building any index.

Count the number of transcripts per gene in the annotation. If the median number of transcripts per gene is above three, or if any gene has more than ten annotated transcripts, the transcriptome has high isoform complexity. Next, check for exact duplicate transcript sequences. The NCBI provides access to sequence databases and search systems that allow you to identify duplicate sequences in the reference [<a href="#ref-2">2</a>]. If exact duplicates are present, they must be removed or flagged before quantification, regardless of which tool is selected.

For transcriptomes with high isoform complexity, RSEM's alignment-based approach provides more positional information for each read. The aligner maps each read to a specific location on the transcript, and RSEM uses this information to distinguish between isoforms that share long regions of sequence identity. Salmon's alignment-free mode groups reads into equivalence classes, which loses the exact positional information. If the isoforms of interest differ only in a short exon, the equivalence class approach may not provide enough information to separate them reliably.

### Step 2: Evaluate the Library Preparation and Sequencing Parameters

The second decision point is the library preparation protocol and the sequencing chemistry. The key parameters are the strandedness, the read length, and the fragment length distribution. These parameters determine whether the bias correction models in Salmon will provide a substantial advantage over RSEM.

Salmon was the first transcriptome-wide quantifier to correct for fragment GC-content bias, and this correction substantially improves the accuracy of abundance estimates and the sensitivity of subsequent differential expression analysis [<a href="#ref-3">3</a>]. If the library preparation kit is known to produce GC bias, or if the sequencing facility has reported GC bias in previous runs, Salmon is the preferred choice. The GC bias correction model in Salmon is learned from the data during the quantification run, and the fitted parameters are reported in the output logs.

For libraries with strong sequence-specific bias, such as those prepared with certain transposase-based kits, Salmon's sequence-specific bias model provides a correction that RSEM does not offer. The sequence-specific bias model accounts for the preferences of the library preparation enzymes and the sequencing adapters for certain sequences. If the library was prepared with a kit that is known to introduce sequence-specific bias, Salmon will produce more accurate estimates.

RSEM is the preferred choice when the library is unstranded and the fragment length distribution is narrow. RSEM's model of the read generation process is more detailed than Salmon's, and it can provide more accurate estimates when the data does not have strong biases. RSEM also provides confidence intervals for the abundance estimates when the `--calc-ci` parameter is enabled, which is useful for assessing the reliability of the estimates for individual transcripts.

### Step 3: Consider the Downstream Analysis Requirements

The third decision point is the downstream analysis. The choice between Salmon and RSEM affects the available tools for differential expression analysis and the way the uncertainty in the abundance estimates is propagated.

Salmon produces bootstrap estimates that can be used with the sleuth package for differential expression analysis. The bootstrap estimates capture the uncertainty in the abundance estimates, and sleuth uses this uncertainty to improve the statistical power of the differential expression test. If the downstream analysis will use sleuth, Salmon is the required choice.

RSEM produces expected counts that can be used with the EBSeq package for differential expression analysis. EBSeq models the uncertainty in the RSEM estimates using a Bayesian approach. If the downstream analysis will use EBSeq, RSEM is the required choice.

The Bioconductor project provides packages for the analysis of count data, including tools for quality control and visualization [<a href="#ref-6">6</a>]. The choice between Salmon and RSEM should be made before the analysis starts, and the same tool must be used for all samples in the study. Mixing tools across samples introduces a systematic difference that cannot be distinguished from a biological effect.

### Step 4: Document the Decision and the Rationale

The decision framework should be applied to each study, and the outcome should be documented. The documentation must include the reference transcriptome version, the library preparation kit, the sequencing platform, the read length, the strandedness, and the downstream analysis tools. The rationale for the choice between Salmon and RSEM should be recorded with the quantification parameters.

The nf-core documentation provides community pipeline standards that emphasize the importance of reproducible workflow configuration [<a href="#ref-11">11</a>]. The documentation shows how to record the tool versions, the parameters, and the reference files in a structured format that can be shared with collaborators and reviewers. The same level of detail should be applied to the decision framework.

### Step 5: Validate the Choice with a Pilot Sample

Before running the full study, validate the choice of tool with a pilot sample. Run both Salmon and RSEM on the same sample and compare the results. The comparison should include the mapping rate, the count distribution, and the estimates for a set of known isoforms.

If the results from Salmon and RSEM are substantially different, investigate the cause. The difference may be due to the bias correction models, the equivalence class approach, or the alignment parameters. The European Bioinformatics Institute provides training materials on RNA-seq analysis that cover the practical aspects of running quantification tools and interpreting the output [<a href="#ref-4">4</a>]. The training materials emphasize the importance of understanding the assumptions of each tool before interpreting the results.

The pilot sample should also be used to assess the stability of the results. Run Salmon with multiple random seeds and compare the estimates. If the estimates vary substantially between seeds, the quantification is unstable and the results are unreliable. The instability may be due to the complexity of the transcriptome or the quality of the reads.

### Step 6: Establish a Record System for the Quantification Runs

The record system for the quantification runs must capture the decision framework outcome, the tool version, the reference file, the parameters, and the input data. The records must be stored with the results so that anyone can reproduce the analysis.

The record system should include a table with one row per sample. The columns should include the sample name, the tool, the tool version, the reference transcriptome version, the library type, the read length, the k-mer size, the number of bootstrap replicates, and the bias correction settings. The checksums of the input FASTQ files should also be recorded to ensure that the input data has not been corrupted or changed.

The Carpentries lessons provide foundational training on computing, data, shell, Git, and programming that is relevant to establishing a reproducible record system [<a href="#ref-12">12</a>]. The lessons show how to use version control to track changes to the analysis scripts and how to document the analysis steps in a structured format.

### Step 7: Escalate When the Decision Framework Does Not Resolve the Problem

If the decision framework does not resolve the reliability issues, the problem may be in the reference annotation or the experimental design. The escalation criteria include persistent low mapping rate, extreme bias in the count distribution, inconsistent results across tools, and results that contradict known biology.

A bioinformatics specialist can examine the reference annotation for errors, such as incorrect exon boundaries or missing isoforms. The NCBI provides tools for examining the annotation and comparing it to the genome [<a href="#ref-2">2</a>]. The specialist can also assess whether the experimental design is appropriate for the biological question and whether the sequencing depth is sufficient for isoform-level analysis.

The decision framework is not a substitute for expertise. It is a structured approach to selecting the appropriate tool and documenting the rationale. The framework reduces the risk of arbitrary tool selection, but it does not eliminate the need for careful examination of the results and the assumptions of the tools.

## Frequently Asked Questions

### Why does my mapping rate drop when I switch from gene-level to isoform-level quantification?

The mapping rate drops because the reference transcriptome contains more sequences than the genome. Reads that map to intronic regions or intergenic regions in the genome do not map to the transcriptome. The mapping rate for isoform-level quantification is expected to be lower than the mapping rate for gene-level quantification. If the mapping rate drops dramatically, check the reference transcriptome for completeness and the reads for adapter contamination.

### How do I know if my library is stranded?

The library preparation kit documentation states whether the library is stranded. If the documentation is not available, you can determine the strandedness empirically by examining the read orientation relative to the annotated transcripts. The Galaxy Training Network provides tutorials that show how to determine the strandedness of a library from the data [<a href="#ref-5">5</a>].

### What is the difference between Salmon and RSEM?

Salmon is a lightweight quantification tool that uses a dual-phase parallel inference algorithm and feature-rich bias models [<a href="#ref-3">3</a>]. It is faster than RSEM and corrects for fragment GC-content bias. RSEM uses alignments from Bowtie or Bowtie2 and provides a more detailed model of the read generation process. The choice between the two tools depends on the specific requirements of the analysis.

### How many bootstrap replicates should I use in Salmon?

The number of bootstrap replicates depends on the complexity of the transcriptome and the desired level of confidence. A typical number is 100, but the optimal number depends on the specific analysis. More bootstrap replicates provide a more accurate estimate of the uncertainty, but they require more computational time.

### Why do I get different results when I run Salmon with different random seeds?

Salmon uses a random seed for the optimization and the Gibbs sampling. Different seeds produce slightly different results. If the results differ substantially between seeds, the quantification is unstable and the estimates are unreliable. Run Salmon with multiple seeds and compare the results to assess the stability.

### How do I check if the reference transcriptome is complete?

The completeness of the reference transcriptome can be checked by comparing the number of annotated transcripts to the number of known genes and by examining the mapping rate. If the mapping rate is low and the reference is from a well-annotated organism, the reference may be incomplete. The NCBI provides tools for examining the annotation and comparing it to the genome [<a href="#ref-2">2</a>].

### What should I do if the counts for a gene do not match the biological expectation?

If the counts for a gene do not match the biological expectation, check the reference annotation for errors, such as incorrect exon boundaries or missing isoforms. Check the parameters used for the quantification, including the library type and the bias correction settings. If the problem persists, consult a bioinformatics specialist.

### Can I use the same index for Salmon and RSEM?

No, Salmon and RSEM use different index formats. The index must be built separately for each tool. The index building step is a one-time decision that affects every sample, so the index must be built carefully and the exact command must be recorded.

## Related Bioinformatics Guides

- [Long-Read Sequencing for Isoform Quantification: Challenges and Solutions](/knowledge/bioinformatics/long-read-sequencing-for-isoform-quantification-challenges-and-solutions)
- [How to Interpret Gene Set Enrichment Analysis Results](/knowledge/bioinformatics/how-to-interpret-gene-set-enrichment-analysis-results)
- [Genomic Data Analysis Tools: A Comparative Guide for Researchers](/knowledge/bioinformatics/genomic-data-analysis-tools-a-comparative-guide-for-researchers)
- [Understanding UMI in Single-Cell Sequencing: What It Is and Why It Matters](/knowledge/bioinformatics/understanding-umi-in-single-cell-sequencing-what-it-is-and-why-it-matters)
- [Spatial Transcriptomics Study Design: Key Considerations for Robust Results](/knowledge/bioinformatics/spatial-transcriptomics-study-design-key-considerations-for-robust-results)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)

## References and Further Reading

<a id="ref-1"></a>[<a href="#ref-1">1</a>] [Protocol for purifying and sequencing cell subpopulations based on RNA FISH signal in planaria.](https://doi.org/10.1016/j.xpro.2026.104547). 2026.

<a id="ref-2"></a>[<a href="#ref-2">2</a>] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.

<a id="ref-3"></a>[<a href="#ref-3">3</a>] [Salmon: fast and bias-aware quantification of transcript expression using dual-phase inference](https://doi.org/10.1038/nmeth.4197). Nature Methods, 2017.

<a id="ref-4"></a>[<a href="#ref-4">4</a>] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.

<a id="ref-5"></a>[<a href="#ref-5">5</a>] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.

<a id="ref-6"></a>[<a href="#ref-6">6</a>] [Bioconductor](https://bioconductor.org/). Bioconductor Project.

<a id="ref-7"></a>[<a href="#ref-7">7</a>] [Structural basis of insulin receptor antagonism by bivalent site 1-site 2 ligands S961 and Ins-AC-S2.](https://doi.org/10.1038/s41467-026-73851-1). 2026.

<a id="ref-8"></a>[<a href="#ref-8">8</a>] [Protocol for the genome-wide identification of intrinsic transcription factor binding motifs by mammalian-optimized pull-down sequencing.](https://doi.org/10.1016/j.xpro.2026.104513). 2026.

<a id="ref-9"></a>[<a href="#ref-9">9</a>] [Protocol for marker-free genome editing in Saccharomyces cerevisiae using universal donor templates and multiplexed CRISPR-Cas9.](https://doi.org/10.1016/j.xpro.2025.104280). 2026.

<a id="ref-10"></a>[<a href="#ref-10">10</a>] [Protocol to perform polysome profiling in primary differentiating murine adipocytes.](https://doi.org/10.1016/j.xpro.2025.103799). 2025.

<a id="ref-11"></a>[<a href="#ref-11">11</a>] [nf-core Documentation](https://nf-co.re/docs). nf-core.

<a id="ref-12"></a>[<a href="#ref-12">12</a>] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.