Transcript Quantification with Salmon: A Step-by-Step Guide to Isoform-Level Analysis from RNA-seq Data

By Dr. Zubair Khalid, DVM, MS, PhD ·

Transcript Quantification with Salmon: A Step-by-Step Guide to Isoform-Level Analysis from RNA-seq Data

Key Takeaways

  • Salmon quantifies RNA-seq reads at the transcript isoform level, enabling the detection of alternative splicing events and isoform switching that are obscured by gene-level analysis. This is critical for understanding biological processes like the sleep deprivation study where over 1500 genes showed isoform switching invisible at the gene level.
  • The workflow involves building a reference transcriptome index, followed by read quantification using an expectation-maximization algorithm, which avoids full alignment for computational efficiency. The choice of reference transcriptome (e.g., GENCODE, Ensembl, RefSeq) and its version is paramount for accurate isoform assignment.
  • Quantification output includes TPM (Transcripts Per Million) for normalized expression and NumReads for differential expression analysis, alongside effective transcript lengths. Quality assessment relies on mapping rates (typically >70% for well-annotated organisms) and count distributions, often aggregated using tools like MultiQC.
  • Downstream analysis can be performed at the gene level using tools like tximport or at the transcript level using salmon quantmerge, with specialized statistical methods (e.g., Sleuth, DEXSeq) required to account for isoform abundance uncertainty.
  • Handling single-cell RNA-seq data with Salmon, often via tools like Alevin, requires specific considerations due to low input material and amplification biases, though isoform-level analysis can reveal cellular subsets missed by gene-based methods.
  • Reproducibility is achieved through version control of scripts (e.g., Git), containerization (e.g., Docker), workflow management systems (e.g., Snakemake, Nextflow), and meticulous documentation of all parameters, including reference versions and bias correction settings.

RNA sequencing produces millions of short reads that must be assigned to the transcripts that generated them. Salmon is a computational tool that performs this assignment at the transcript level, allowing researchers to distinguish between isoforms of the same gene. This guide provides a complete workflow from raw FASTQ files to transcript-level count matrices, including index construction, quantification commands, output interpretation, and downstream analysis considerations for splicing research.

The primary audience for this material includes biology students entering computational research, laboratory professionals who generate sequencing data, and life-science practitioners who need to analyze transcriptomes without deep programming expertise. The workflow described here assumes access to a Unix-based computing environment, either a local server, a high-performance computing cluster, or a cloud instance. The practical outcome is a reproducible pipeline that produces transcript-level quantification files suitable for differential expression analysis and isoform switching detection.

At a Glance

Salmon operates in two main phases. First, an index is built from a reference transcriptome. Second, quantification is performed by mapping reads to the index and estimating transcript abundances using an expectation-maximization algorithm. The output includes per-sample quantification files with estimated counts, transcript per million values, and effective lengths.

Workflow StagePrimary InputKey OutputCommon Decision Point
Index buildingReference transcriptome FASTASalmon index directoryChoosing transcriptome source and version
QuantificationFASTQ files and indexQuant.sf per sampleSelecting library type and number of bootstrap samples
Quality assessmentQuant.sf files and alignment logsMapping rates, count distributionsDetermining whether samples meet inclusion criteria
Downstream analysisCombined count matrixDifferential expression or isoform switching resultsSelecting gene-level versus transcript-level analysis

The table above summarizes the four main stages covered in this guide. Each stage requires specific inputs and produces outputs that feed into the next step. The decision points listed are the most common places where researchers must make choices that affect downstream results.

Understanding Transcript-Level Quantification

Gene-level quantification assigns reads to genes, collapsing all isoforms into a single measurement. Transcript-level quantification assigns reads to specific isoforms, preserving information about alternative splicing and isoform usage. This distinction matters because many biological processes involve changes in isoform expression without changes in overall gene expression.

Isoform-level expression is biologically more informative than gene-level expression and can reveal cellular subsets and corresponding biomarkers that are not visible at the gene level. Studies of sleep deprivation in the mouse frontal cortex demonstrated isoform switching of over 1500 genes, particularly those involved in splicing and RNA binding, while global transcription was broadly repressed. These isoform-level changes would have been invisible in a gene-level analysis that collapsed all transcripts into single measurements.

The practical implication for researchers is that transcript-level quantification enables detection of biological phenomena that gene-level approaches miss. When designing an RNA-seq experiment, the choice between gene-level and transcript-level analysis should be made before sequencing, because the reference and quantification strategy affect the entire pipeline.

How Salmon Differs from Alignment-Based Methods

Traditional RNA-seq analysis pipelines use splice-aware aligners such as STAR or HISAT2 to map reads to a reference genome, followed by counting tools that assign aligned reads to genes or transcripts. Salmon uses a different approach based on lightweight algorithms that avoid full base-to-base alignment.

Salmon builds an index of the transcriptome that contains k-mers and other sequence features. During quantification, reads are compared against this index to determine which transcripts they are compatible with. The expectation-maximization algorithm then resolves ambiguities where a read could originate from multiple transcripts, particularly reads that map to regions shared by isoforms of the same gene.

This approach offers computational efficiency compared to full alignment. The tradeoff is that Salmon does not produce a traditional alignment file in BAM format by default, although it can output alignments if needed. For most downstream analyses, the quantification files contain sufficient information.

The Role of the Reference Transcriptome

The reference transcriptome is the set of known transcript sequences used to build the Salmon index. The choice of reference has a direct impact on quantification accuracy. If a transcript is missing from the reference, reads originating from that transcript will be unassigned or incorrectly assigned to similar transcripts.

For human and mouse studies, the standard references are the GENCODE and Ensembl transcript sets. For other organisms, RefSeq from the National Center for Biotechnology Information provides curated transcript sequences. The NCBI maintains databases of nucleotide sequences, genome assemblies, and annotation resources that researchers can use to obtain reference transcriptomes for a wide range of species.

When selecting a reference, consider the version and the date of the annotation. Newer versions may include additional transcripts or correct errors in previous versions. However, changing the reference version between samples in the same study will introduce artifacts, so the same reference must be used for all samples in a comparison.

Preparing the Computing Environment

Salmon is a command-line tool that runs on Linux and macOS systems. Windows users can run Salmon through the Windows Subsystem for Linux or through a virtual machine. The tool is distributed as a precompiled binary, which simplifies installation compared to tools that require compilation from source.

Before installing Salmon, verify that the computing environment meets the basic requirements. Salmon requires a 64-bit operating system and sufficient memory to load the index into RAM. The memory requirement depends on the size of the transcriptome. For human transcriptomes, several gigabytes of RAM are typically needed. For smaller transcriptomes such as bacterial or yeast, the requirement is lower.

Installing Salmon

The official Salmon documentation provides installation instructions for various operating systems. The simplest approach is to download the precompiled binary from the release page and add it to the system PATH. Alternatively, package managers such as conda can install Salmon along with its dependencies.

After installation, verify that Salmon is working by running the command salmon --version. This command should print the version number. If the command is not found, check that the installation directory is included in the PATH environment variable.

Organizing Project Directories

A well-organized project directory structure makes the workflow easier to manage and reproduce. A typical structure includes separate directories for raw data, reference files, indexes, quantification results, and scripts.

project/
  raw_data/
  reference/
  index/
  quant/
  scripts/
  results/

Place raw FASTQ files in the raw_data directory. Store the reference transcriptome FASTA file in the reference directory. Build the Salmon index into the index directory. Run quantification with output directed to the quant directory. Keep analysis scripts in the scripts directory and final results in the results directory.

This structure supports reproducibility because each step has a defined location for inputs and outputs. When sharing the analysis with collaborators or reviewers, the directory structure makes it clear where each file belongs.

Building the Salmon Index

The index is a data structure that allows Salmon to quickly determine which transcripts are compatible with each read. Building the index is a one-time step for a given reference transcriptome. The same index can be used for all samples in a study.

Obtaining the Reference Transcriptome

Download the reference transcriptome FASTA file from an appropriate source. For model organisms, the Ensembl or GENCODE websites provide transcript sequences. The NCBI provides RefSeq transcript sequences through its nucleotide database and FTP servers.

The FASTA file should contain one entry per transcript, with the transcript identifier in the header line. The identifiers should be consistent with the annotation files used for downstream analysis. If the downstream analysis will use transcript-to-gene mapping, ensure that the mapping file is available and matches the transcript identifiers in the FASTA file.

Running the Index Command

The basic command to build a Salmon index is:

salmon index -t reference_transcripts.fasta -i salmon_index -k 31

The -t flag specifies the transcript FASTA file. The -i flag specifies the output directory for the index. The -k flag sets the k-mer length used for indexing. The default k-mer length is 31, which works well for most applications.

For transcriptomes with high sequence similarity, such as those containing many closely related isoforms, a longer k-mer length may improve specificity. However, longer k-mers increase the memory requirement for the index. The Salmon documentation provides guidance on selecting appropriate k-mer lengths for different applications.

Verifying the Index

After building the index, verify that it was created successfully by listing the contents of the index directory. The index directory should contain several files, including the hash table and other data structures used by Salmon.

A quick test of the index can be performed by running quantification on a small subset of reads. If the index is corrupted or incompatible with the Salmon version, the quantification step will fail with an error message. Resolve any errors before proceeding with the full dataset.

Quantifying Transcript Abundance

Quantification is the core step where reads are assigned to transcripts and abundances are estimated. The command structure for quantification depends on whether the data are single-end or paired-end.

Single-End Quantification

For single-end reads, the basic quantification command is:

salmon quant -i salmon_index -l A -r sample.fastq.gz -o quant/sample --seqBias --gcBias

The -i flag specifies the index directory. The -l flag specifies the library type, with A indicating automatic detection. The -r flag specifies the FASTQ file containing the reads. The -o flag specifies the output directory. The --seqBias and --gcBias flags enable bias correction, which improves accuracy by accounting for sequence-specific and GC-content-specific biases in library preparation.

Paired-End Quantification

For paired-end reads, the command includes both read files:

salmon quant -i salmon_index -l A -1 sample_R1.fastq.gz -2 sample_R2.fastq.gz -o quant/sample --seqBias --gcBias

The -1 flag specifies the file containing the first read of each pair, and the -2 flag specifies the file containing the second read. The paired-end information improves the assignment of reads to transcripts because the two reads of a pair provide more sequence information than a single read.

Library Type Specification

The -l A flag tells Salmon to automatically detect the library type. This works for most standard library preparation protocols. However, for libraries prepared with unusual protocols, automatic detection may fail. In such cases, specify the library type explicitly.

The library type string describes the orientation and strandedness of the reads. Common values include ISR for inward-stranded reverse, IU for inward-unstranded, and ISF for inward-stranded forward. The Salmon documentation provides a complete list of library type codes and their meanings.

If the library type is specified incorrectly, the quantification results will be wrong. The automatic detection mode is recommended for most applications because it avoids this risk.

Bootstrap Samples for Inference

Salmon can generate bootstrap samples that estimate the uncertainty in transcript abundance estimates. Bootstrap samples are useful for downstream analyses that require confidence intervals or that use methods incorporating estimation uncertainty.

To generate bootstrap samples, add the --numBootstraps flag with the desired number of samples:

salmon quant -i salmon_index -l A -r sample.fastq.gz -o quant/sample --seqBias --gcBias --numBootstraps 100

A typical number of bootstrap samples is 100, which provides a reasonable balance between accuracy and computational cost. Bootstrap samples increase the runtime and disk usage, so consider whether the downstream analysis requires them before enabling this option.

Understanding the Quantification Output

Each quantification run produces a directory containing several files. The primary output file is quant.sf, which contains the quantification results in a tabular format.

The Quant.sf File Format

The quant.sf file is a tab-separated table with a header line and one row per transcript. The columns are:

ColumnDescription
NameTranscript identifier
LengthTranscript length in nucleotides
EffectiveLengthLength adjusted for sequence-specific biases
TPMTranscripts per million
NumReadsEstimated number of reads assigned to the transcript

The TPM value normalizes for both sequencing depth and transcript length, allowing comparison of expression levels across transcripts and samples. The NumReads value is the estimated count that can be used for differential expression analysis.

Additional Output Files

The quantification directory also contains a logs subdirectory with detailed information about the quantification run. The log files record the command used, the parameters set, and any warnings or errors encountered during the run.

The aux_info directory contains auxiliary files, including the meta_info.json file with run metadata and the fld file with fragment length distribution information. These files are useful for quality assessment and for downstream tools that require additional information.

Mapping Rate and Quality Metrics

The log file reports the mapping rate, which is the percentage of reads that were assigned to transcripts. A high mapping rate indicates that the reads are consistent with the reference transcriptome. A low mapping rate may indicate problems with the reference, the sample, or the sequencing quality.

For most well-annotated organisms, mapping rates above 70 percent are expected. Lower mapping rates warrant investigation. Possible causes include contamination, reference mismatch, or poor sequencing quality. The log file also reports the number of reads that were too short, too long, or failed other quality filters.

Quality Control and Assessment

Quality control is essential for producing reliable results. The quantification output provides several metrics that can be used to assess sample quality before proceeding to downstream analysis.

Checking Mapping Rates Across Samples

Compile the mapping rates from all samples in the study. Samples with substantially lower mapping rates than the rest of the cohort may have quality issues. Investigate such samples before including them in the analysis.

The mapping rate can be affected by the reference transcriptome completeness. If the organism has a poorly annotated genome, many reads may map to unannotated regions and be excluded from quantification. In such cases, a lower mapping rate may be expected and should be interpreted in the context of the reference quality.

Examining Count Distributions

The distribution of counts across transcripts should follow a typical pattern, with a small number of highly expressed transcripts and a long tail of lowly expressed transcripts. Samples with unusual count distributions may have technical issues.

Plot the distribution of log-transformed counts for each sample. Samples with bimodal distributions or with an excess of zero counts may warrant further investigation. The count distribution can also reveal batch effects when comparing samples processed in different batches.

Using MultiQC for Aggregated Reports

MultiQC is a tool that aggregates quality metrics from multiple samples into a single report. It can parse Salmon log files and produce summary tables and plots. This tool is particularly useful for studies with many samples, where manual inspection of individual log files is impractical.

The Galaxy Training Network provides tutorials on RNA-seq analysis that include quality control steps and demonstrate the use of tools like MultiQC. These tutorials are designed for researchers who want to learn the practical aspects of RNA-seq analysis in an accessible format.

Combining Quantification Results for Downstream Analysis

Individual quantification files must be combined into a count matrix for differential expression analysis. Several tools can perform this step, including tximport in R and the salmon quantmerge command.

Using tximport for Gene-Level Summarization

The tximport package in Bioconductor imports transcript-level quantification files and summarizes them to the gene level when needed. This is useful when the downstream analysis tool requires gene-level counts, such as DESeq2 or edgeR.

tximport requires a mapping between transcript identifiers and gene identifiers. This mapping can be obtained from the annotation files used to build the reference transcriptome. The Bioconductor project provides documentation and workflows for using tximport with Salmon output.

Using salmon quantmerge for Transcript-Level Matrices

The salmon quantmerge command combines quantification files from multiple samples into a single matrix. This command is useful when the downstream analysis will be performed at the transcript level.

salmon quantmerge --quants quant/sample1 quant/sample2 quant/sample3 -o transcript_counts.tsv

The --quants flag specifies the quantification directories to merge. The -o flag specifies the output file. The resulting matrix has transcripts as rows and samples as columns, with the estimated counts as values.

Considerations for Differential Expression Analysis

Differential expression analysis at the transcript level requires methods that account for the uncertainty in transcript abundance estimates. Tools such as Sleuth and DEXSeq incorporate this uncertainty into their statistical models.

The choice between gene-level and transcript-level differential expression depends on the biological question. If the question concerns overall gene expression changes, gene-level analysis is appropriate. If the question concerns isoform switching or alternative splicing, transcript-level analysis is required.

The study of sleep deprivation effects in the mouse frontal cortex provides an example where transcript-level analysis revealed isoform switching of over 1500 genes, a finding that would not have been apparent from gene-level analysis alone. Researchers interested in splicing regulation should plan for transcript-level analysis from the start of their study.

Handling Single-Cell RNA-Seq Data

Single-cell RNA-seq presents unique challenges for transcript quantification. The low input material and amplification steps introduce technical noise that must be accounted for in the analysis.

Isoform-Level Quantification in Single-Cell Data

Isoform-level expression in single-cell data can reveal cellular subsets and biomarkers that are not visible at the gene level. However, the strong 3-prime bias of many single-cell protocols limits the information available for isoform discrimination.

Methods such as Scasa have been developed to address this challenge by exploiting transcription clusters and isoform paralogs. These methods compare favorably against competing approaches including Alevin, Cellranger, Kallisto, Salmon, Terminus, and STARsolo at both isoform and gene level expression.

The reanalysis of a CITE-Seq dataset with isoform-based Scasa revealed a subgroup of CD14 monocytes missed by gene-based methods. This finding demonstrates the biological value of isoform-level analysis in single-cell data.

Salmon in Single-Cell Pipelines

Salmon can be used for single-cell RNA-seq quantification through tools like Alevin, which is part of the Salmon suite. Alevin processes single-cell data and produces count matrices suitable for downstream analysis with tools like Seurat or Scanpy.

The current best practices in single-cell RNA-seq analysis include pre-processing steps such as quality control, normalization, data correction, feature selection, and dimensionality reduction. These steps are detailed in tutorials and reviews that provide workflow guidance for new entrants to the field.

Integration with Multimodal Data

Recent advances in single-cell technology enable simultaneous measurement of multiple modalities, such as RNA expression and chromatin accessibility. The integration of these multimodal data can improve cell type identification and reveal rare cell types.

A systematic evaluation of multimodal single-cell integration strategies in the kidney demonstrated that horizontal integration of scRNA and snRNA-seq improves cell-type identification, while vertical integration of snRNA-seq and snATAC-seq enhances resolution in homogeneous populations and difficult-to-identify states. Global integration was especially effective in identifying adaptive states and rare cell types.

For researchers working with multimodal data, the quantification step must be coordinated across modalities to ensure that cell identities are consistent. The transcript quantification from Salmon provides the RNA expression component that is integrated with other data types.

Reproducibility and Workflow Management

Reproducibility is a core requirement for computational biology research. The Salmon workflow should be documented and versioned so that results can be reproduced by collaborators or reviewers.

Version Control for Analysis Scripts

Store analysis scripts in a version control system such as Git. The Carpentries provides lessons on version control with Git that cover the fundamentals of tracking changes, committing updates, and collaborating with others.

Version control allows researchers to track changes to analysis scripts over time and to revert to previous versions when needed. This is particularly important when the analysis evolves during the course of a study.

Containerization and Workflow Tools

Containerization tools such as Docker and Singularity package the analysis environment, including the operating system, software versions, and dependencies, into a single image. This ensures that the analysis runs identically on different computing systems.

Workflow management tools such as Snakemake and Nextflow automate the execution of analysis pipelines. The nf-core project provides community-developed pipelines that follow best practices for reproducibility and configuration. These pipelines include RNA-seq analysis workflows that use Salmon for quantification.

The nf-core documentation describes the standards for pipeline development, including versioning, testing, and documentation requirements. Using a community pipeline can save time and improve the reproducibility of the analysis.

Documenting the Analysis

Document the analysis steps, including software versions, parameter settings, and reference versions. This documentation should be sufficient for another researcher to reproduce the analysis from raw data.

The EMBL-EBI provides training materials on bioinformatics data resources and analysis methods. These materials include guidance on documenting computational analyses and on using public data resources effectively.

Common Failure Patterns and Troubleshooting

Several common problems can arise during Salmon quantification. Recognizing these patterns and knowing how to address them can save substantial time.

Low Mapping Rates

Low mapping rates can result from several causes. The reference transcriptome may be incomplete or mismatched to the sample species. The sample may be contaminated with DNA or RNA from other organisms. The sequencing quality may be poor, with many reads failing quality filters.

To diagnose low mapping rates, examine the log file for details about read classification. Check whether the reads match the expected species by running a quick alignment against a reference genome or by examining the taxonomic composition of the reads.

Memory Errors During Index Building

Building an index for a large transcriptome can require substantial memory. If the index building fails with a memory error, consider using a longer k-mer length, which reduces the memory requirement, or use a computing node with more memory.

The memory requirement also depends on the number of transcripts and the sequence diversity. Transcriptomes with many highly similar transcripts require more memory for the index.

Quantification Failures

Quantification can fail for several reasons, including corrupted index files, incompatible Salmon versions, or malformed FASTQ files. Check the error message and the log file for details.

If the index was built with a different version of Salmon, rebuild the index with the current version. If the FASTQ files are corrupted, re-download or re-generate them from the sequencing facility.

Batch Effects

Batch effects are systematic differences between samples processed in different batches. These effects can arise from differences in library preparation, sequencing runs, or reagent lots.

Batch effects can be detected by examining the count distributions and by performing principal component analysis on the count matrix. If batch effects are present, include the batch information in the differential expression model.

Limitations of Transcript-Level Quantification

Transcript-level quantification has limitations that researchers should understand when interpreting results.

Ambiguous Read Assignment

Reads that map to regions shared by multiple isoforms cannot be unambiguously assigned. The expectation-maximization algorithm distributes these reads among the compatible transcripts based on the estimated abundances. This introduces uncertainty in the abundance estimates for transcripts that share sequence regions.

The bootstrap samples generated by Salmon provide a measure of this uncertainty. Downstream analysis tools that incorporate bootstrap information can account for this uncertainty in statistical tests.

Reference Dependence

The quantification results depend entirely on the reference transcriptome. Transcripts that are not in the reference will not be quantified. Novel isoforms that are not annotated will be missed or incorrectly assigned to annotated isoforms.

For organisms with incomplete annotations, transcript-level quantification may be less reliable than gene-level quantification. Consider the annotation quality when deciding whether to perform transcript-level analysis.

Bias Correction Limitations

Salmon includes bias correction methods for sequence-specific and GC-content-specific biases. These methods improve accuracy but cannot fully eliminate all biases. The effectiveness of bias correction depends on the library preparation protocol and the sequencing platform.

The --seqBias and --gcBias flags enable bias correction. These flags are recommended for most applications, but they increase the runtime and memory usage.

Professional Escalation Criteria

Some situations require consultation with a bioinformatics specialist or a statistician. Recognizing these situations can prevent wasted effort and incorrect conclusions.

When to Seek Expert Assistance

Consult a bioinformatics specialist when the mapping rates are consistently low across all samples, when the reference transcriptome is poorly annotated, or when the analysis requires advanced statistical methods that are unfamiliar.

Consult a statistician when the experimental design is complex, such as with multiple factors, repeated measures, or batch effects that are confounded with the biological variable of interest.

Documentation for Escalation

When escalating a problem, provide the relevant documentation, including the Salmon log files, the quantification output, and the analysis scripts. This documentation allows the specialist to diagnose the problem efficiently.

The NCBI provides resources for finding bioinformatics support, including databases of tools and services. The EMBL-EBI training materials also provide guidance on when to seek expert assistance and how to prepare for consultations.

Records and Measurements for Quality Assurance

Maintaining detailed records of the analysis is essential for quality assurance and for publication requirements.

Recording Analysis Parameters

Record the Salmon version, the reference transcriptome version, the k-mer length, the library type, and all flags used for quantification. This information should be included in the methods section of any publication.

The meta_info.json file in the quantification directory contains some of this information automatically. However, the reference version and the rationale for parameter choices should be documented separately.

Tracking Sample Quality Metrics

Maintain a table of sample quality metrics, including mapping rates, total read counts, and the number of transcripts with nonzero counts. This table can be used to identify problematic samples and to report quality metrics in publications.

The table should include the sample identifier, the sequencing run, the library preparation batch, and any other relevant metadata. This information is essential for diagnosing batch effects and for reporting the analysis accurately.

Archiving Raw Data and Results

Archive the raw FASTQ files, the reference transcriptome, the index, and the quantification results. Public repositories such as the NCBI Sequence Read Archive provide long-term storage for raw sequencing data.

The NCBI provides data submission services and maintains databases that support the sharing and reuse of genomic data. Depositing data in public repositories is a requirement for many journals and funding agencies.

Safety and Ethical Considerations

RNA-seq analysis involves handling large data files and using computational resources. While there are no direct physical safety concerns, researchers should be aware of data management and ethical considerations.

Data Management

Raw sequencing data can be large, often exceeding hundreds of gigabytes per study. Ensure that data are stored securely and backed up regularly. Follow institutional policies for data storage and retention.

Patient-derived data, including single-cell data from human tissues, may be subject to privacy regulations. De-identify data before analysis and follow institutional review board requirements for data handling.

Computational Resource Use

Salmon quantification can be computationally intensive, particularly for large datasets with many samples. Be mindful of resource limits on shared computing systems and schedule jobs appropriately.

The nf-core documentation provides guidance on configuring workflows for different computing environments, including resource allocation and job scheduling.

Building a Salmon Analysis Decision Log for Isoform-Level Studies

A recurring failure in transcript-level RNA-seq projects is the loss of context around quantification decisions. Researchers often run Salmon, store the quant.sf files, and later discover that they cannot reconstruct why a particular reference version, k-mer length, or bias correction setting was chosen. This section provides a practical decision framework and record system that links each Salmon parameter to the biological question being asked, enabling you to audit your own analysis and to communicate your methods clearly to collaborators or reviewers.

The Decision Framework: Linking Parameters to Biological Questions

Before running any Salmon command, document the biological question that the quantification must support. This question determines the reference choice, the quantification mode, and the downstream analysis path. The framework below organizes these decisions into four layers that you should record before execution.

Decision LayerQuestion to AnswerSalmon Parameter or InputRecorded Output
Reference layerWhich transcript isoforms are expected in this sample?Transcriptome FASTA source and versionReference accession and download date
Assignment layerHow should multimapping reads be distributed?k-mer length, --validateMappingsIndex build log and parameter values
Correction layerWhich library biases are present in this protocol?--seqBias, --gcBias, library typeBias correction flags and library type code
Inference layerWhat uncertainty is needed for downstream statistics?--numBootstraps, --recoverOrphansBootstrap count and auxiliary file list

The reference layer is the most consequential decision. If you are studying isoform switching, the reference must contain the isoforms you expect to detect. For human and mouse, GENCODE and Ensembl provide comprehensive transcript sets. For other organisms, RefSeq from the National Center for Biotechnology Information provides curated transcript sequences. Record the exact version and the date you downloaded it, because annotation updates can change transcript identifiers and sequences between releases.

The assignment layer controls how Salmon handles reads that are compatible with multiple transcripts. The default k-mer length of 31 works for most applications, but transcriptomes with many closely related isoforms may benefit from longer k-mers. The --validateMappings flag improves accuracy by validating candidate mappings, at the cost of increased runtime. Record which flags you used and why, because this affects the comparability of results across studies.

The correction layer addresses library preparation biases. The --seqBias flag corrects for sequence-specific biases, and the --gcBias flag corrects for GC-content-dependent biases. These flags are recommended for most applications, but they increase runtime and memory usage. The library type code, such as ISR for inward-stranded reverse, must be recorded because it affects how reads are oriented relative to transcripts.

The inference layer determines whether you can use uncertainty-aware downstream tools. If you plan to use Sleuth or other methods that incorporate estimation uncertainty, you must generate bootstrap samples with --numBootstraps. Record the number of bootstraps and the rationale for that number, because this affects both runtime and the statistical power of downstream tests.

Implementing the Record System

Create a plain-text decision log for each study, stored in the scripts or results directory of your project. The log should contain one entry per sample or per batch of samples processed together. Each entry should include the following fields:

  • Sample identifier and sequencing run ID
  • Reference transcriptome source, version, and download date
  • Salmon version and index build date
  • k-mer length and validation flag
  • Library type code and whether automatic detection was used
  • Bias correction flags enabled
  • Bootstrap count and reason for that count
  • Mapping rate and total read count from the log file
  • Any warnings or errors observed during the run

This log serves two purposes. First, it allows you to audit your own analysis when results are unexpected. If a sample shows an unusual isoform pattern, you can check whether it was processed with the same reference and parameters as the rest of the cohort. Second, it provides the information needed for the methods section of a publication. Journals increasingly require detailed reporting of bioinformatics parameters, and a complete decision log makes this reporting straightforward.

The meta_info.json file in each quantification directory records some of this information automatically, including the Salmon version and the command used. However, it does not record the reference version or the rationale for parameter choices. Your decision log fills this gap.

Common Failure Patterns in Decision Documentation

Three failure patterns recur when researchers skip the decision log. The first is reference version drift. A study that spans several months may use different reference versions for different batches of samples. If the decision log does not record the reference version for each batch, the resulting count matrix will contain artifacts that are difficult to diagnose. The fix is to record the reference version at the time of index building and to verify that all samples in a comparison use the same index.

The second failure pattern is parameter inconsistency across samples. If one batch of samples is quantified with bias correction and another without, the counts are not directly comparable. The decision log should record the exact flags used for each sample, and the analysis should be rerun if inconsistencies are found.

The third failure pattern is missing bootstrap information. If you decide after quantification that you need uncertainty estimates for downstream analysis, you cannot recover them without rerunning Salmon. The decision log should record whether bootstraps were generated and how many, so that this information is available when you select downstream tools.

Using the Decision Log for Troubleshooting

When a sample fails quality checks, the decision log provides the first place to look. Compare the recorded parameters for the failing sample against those for samples that passed. Differences in reference version, k-mer length, or bias correction flags can explain systematic differences in quantification results.

The log also supports escalation to a bioinformatics specialist. When you consult an expert, provide the decision log along with the Salmon log files and the quantification output. This documentation allows the specialist to diagnose problems efficiently without re-deriving your analysis history.

The EMBL-EBI provides training materials on bioinformatics data resources and analysis methods that include guidance on documenting computational analyses. The Carpentries lessons on version control and reproducible analysis practices complement this decision log by providing the skills to track changes to analysis scripts over time.

Integrating the Decision Log with Workflow Tools

If you use a workflow management tool such as Snakemake or Nextflow, the decision log can be generated automatically from the workflow configuration. The nf-core documentation describes standards for pipeline development, including versioning and documentation requirements. Community pipelines that use Salmon for quantification typically record the Salmon version and parameters in their output files, but you should verify that the reference version is also recorded.

For researchers using the Galaxy platform, the Galaxy Training Network provides tutorials on RNA-seq analysis that include quality control steps and demonstrate the use of tools like MultiQC. These tutorials emphasize the importance of recording analysis parameters and provide templates for documenting workflows.

Records and Measurements for Quality Assurance

In addition to the decision log, maintain a summary table of sample quality metrics. This table should include the sample identifier, the mapping rate, the total read count, the number of transcripts with nonzero counts, and the median TPM value. Update this table after each batch of samples is quantified.

The summary table serves two purposes. First, it allows you to identify problematic samples early, before downstream analysis. Second, it provides the quality metrics that are typically reported in publications. The table should be stored in a plain text or CSV format that can be shared with collaborators or included as supplementary material.

When you archive the analysis, include the decision log, the summary table, the Salmon index build command, and the quantification commands for each sample. This archive should be sufficient for another researcher to reproduce the analysis from raw data. The NCBI provides data submission services and maintains databases that support the sharing and reuse of genomic data, and depositing raw sequencing data in the Sequence Read Archive is a requirement for many journals and funding agencies.

Professional Escalation Criteria for Decision Documentation

Escalate to a bioinformatics specialist when you cannot reconstruct the parameters used for a batch of samples, when the decision log reveals inconsistencies across samples that cannot be resolved by rerunning quantification, or when the reference transcriptome version cannot be verified. These situations indicate that the analysis history is incomplete, and continuing with downstream analysis risks producing incorrect biological conclusions.

When escalating, provide the decision log, the Salmon log files, the quantification output, and the raw FASTQ file paths. This documentation allows the specialist to assess whether the quantification is reliable or whether samples must be reprocessed. The NCBI provides resources for finding bioinformatics support, including databases of tools and services, and the EMBL-EBI training materials provide guidance on when to seek expert assistance.

Frequently Asked Questions

What is the difference between gene-level and transcript-level quantification?

Gene-level quantification assigns reads to genes, collapsing all isoforms into a single measurement. Transcript-level quantification assigns reads to specific isoforms, preserving information about alternative splicing and isoform usage. Transcript-level analysis can reveal biological phenomena that are invisible at the gene level, such as isoform switching where different isoforms of the same gene are expressed under different conditions.

How do I choose the reference transcriptome for Salmon?

The reference transcriptome should match the species and the annotation version used for the study. For human and mouse, GENCODE and Ensembl provide comprehensive transcript sets. For other organisms, RefSeq from NCBI provides curated transcript sequences. Use the same reference version for all samples in a study to avoid artifacts from annotation changes.

What library type should I specify for Salmon quantification?

The -l A flag enables automatic library type detection, which works for most standard library preparation protocols. If automatic detection fails, specify the library type explicitly using the codes documented in the Salmon manual. The library type describes the orientation and strandedness of the reads, and specifying it incorrectly will produce wrong results.

How many bootstrap samples should I generate?

A typical number is 100 bootstrap samples, which provides a reasonable balance between accuracy and computational cost. Bootstrap samples are needed for downstream analysis tools that incorporate estimation uncertainty, such as Sleuth. If the downstream analysis does not require uncertainty information, bootstrap samples can be omitted to save time and disk space.

Why is my mapping rate low?

Low mapping rates can result from an incomplete reference transcriptome, contamination, or poor sequencing quality. Check the log file for details about read classification. Verify that the reference matches the sample species. Consider whether the sample may be contaminated with other organisms.

Can Salmon be used for single-cell RNA-seq data?

Yes, Salmon can be used for single-cell RNA-seq through tools like Alevin, which is part of the Salmon suite. Isoform-level quantification in single-cell data can reveal cellular subsets that are not visible at the gene level, but the 3-prime bias of many single-cell protocols limits isoform discrimination. Methods such as Scasa have been developed to address this challenge.

What downstream analysis tools work with Salmon output?

The tximport package in Bioconductor imports Salmon quantification files and can summarize them to the gene level for use with DESeq2 or edgeR. For transcript-level analysis, tools such as Sleuth and DEXSeq incorporate the uncertainty in transcript abundance estimates. The Bioconductor project provides documentation and workflows for these analyses.

How do I ensure my analysis is reproducible?

Use version control for analysis scripts, document all software versions and parameters, and consider using containerization or workflow management tools. The Carpentries provides lessons on version control and reproducible analysis practices. The nf-core project provides community pipelines that follow reproducibility best practices.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.