STAR Alignment for RNA-seq: Key Parameters, Output Metrics, and Troubleshooting Common Issues
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- STAR's speed and accuracy in spliced alignment stem from its maximal mappable prefixes (MMPs) and seed-and-extend strategy, enabling it to span splice junctions without pre-annotated splice sites.
- Accurate alignment hinges on a high-quality reference genome assembly and a matching annotation file (GTF/GFF) for index generation, crucial for correctly identifying exon-exon junctions and improving mapping rates.
- Key parameters like
--outFilterMismatchNmaxand--outFilterMultimapNmaxcontrol alignment stringency, with defaults suitable for standard RNA-seq but requiring adjustment for divergent organisms or repetitive element studies. - The
Log.final.outfile provides critical mapping statistics, including the percentage of uniquely mapped reads (typically >80% for high-quality data) and multimapping reads, which are essential for quality control and identifying potential issues. - Troubleshooting low uniquely mapped rates involves checking raw read quality (FastQC), reference genome/annotation compatibility, and potential sample contamination, while high multimapping rates may indicate repetitive sequences or multi-copy genes.
RNA-seq analysis begins with a critical decision: how to map millions of short sequencing reads back to a reference genome. STAR (Spliced Transcripts Alignment to a Reference) has become one of the most widely used aligners for this task because it performs highly accurate spliced sequence alignment at ultrafast speed. The alignment step determines which reads are assigned to which genes, and errors introduced here propagate through every downstream analysis, including differential expression testing, isoform quantification, and fusion detection. This article explains the key STAR parameters that control alignment behavior, how to interpret the output metrics that indicate mapping quality, and how to diagnose and resolve common alignment failures. The practical goal is to help researchers configure STAR correctly for their specific data types and organism, interpret alignment statistics to catch problems early, and make informed decisions about when parameter adjustments are necessary versus when default settings are appropriate.
Understanding STAR Alignment Fundamentals
STAR operates by searching for maximal mappable prefixes (MMPs) in the reference genome. When a read cannot be mapped contiguously, STAR identifies the longest sequence that matches the genome, then searches for the next MMP downstream, allowing it to span splice junctions without requiring a separate splice-site annotation file. This seed-and-extend strategy is what gives STAR its speed advantage over earlier aligners that relied on exhaustive read splitting.
The algorithm's performance depends on the genome index, which must be built before alignment. The index is a compressed representation of the reference genome that enables rapid lookup of sequence matches. Building a genome index requires the reference genome FASTA file and, optionally, a gene annotation file in GTF or GFF format. When an annotation file is provided, STAR can use it to improve alignment accuracy around known splice junctions and to generate read counts per gene during alignment.
STAR's alignment algorithm can be controlled by many user-defined parameters, and understanding which parameters matter for which applications is essential for obtaining reliable results. The original description of STAR's key options and best practices for achieving maximum mapping accuracy and speed remains the foundational reference for parameter selection. Researchers should consult this guidance when deciding which parameters to adjust for their specific experimental context.
At a Glance: STAR Parameter Decision Table
| Parameter | Function | Typical Setting | When to Adjust |
|---|---|---|---|
--outSAMtype | Output file format | BAM SortedByCoordinate | Use SAM for debugging or when downstream tools require unsorted SAM input |
--quantMode | Generate gene counts during alignment | GeneCounts | Enable when using STAR for quantification, use TranscriptomeSAM for alignment to transcriptome |
--outFilterMismatchNmax | Maximum mismatches per read | 10 for 100 bp reads | Increase for longer reads or divergent organisms, decrease for stringent variant calling |
--outFilterMultimapNmax | Maximum number of genomic loci for multimapping reads | 10 | Lower to 1 for unique-mapping-only analyses, increase for repetitive element studies |
--alignIntronMax | Maximum intron length | 0 (unlimited) for default | Set to species-specific maximum for organisms with known intron size distributions |
--outFilterScoreMinOverLread | Minimum alignment score as fraction of read length | 0.66 | Increase for stringent filtering, decrease for short reads or degraded RNA |
--chimSegmentMin | Minimum length of chimeric segment for fusion detection | 12 | Lower for sensitive fusion detection, increase to reduce false positives |
--sjdbGTFfile | Annotation file for splice junction database | Path to GTF file | Required for accurate junction mapping, omit for de novo transcript discovery |
Building a STAR Genome Index
Reference Genome Selection
The choice of reference genome is the first decision that affects alignment quality. For human and mouse data, the standard references are GRCh38 and GRCm39 respectively. For other organisms, researchers should select the most recent assembly available from authoritative sources. The NCBI provides official descriptions of genome assemblies, sequence resources, and search systems that can help researchers identify the appropriate reference for their organism. Using an outdated or incorrect reference genome will produce alignment artifacts that are difficult to detect without careful inspection of mapping statistics.
When working with non-model organisms, the quality of the reference genome assembly directly impacts alignment rates. Fragmented assemblies with many small contigs will produce lower uniquely mapped rates because reads may span assembly gaps. Researchers working with such organisms should document the assembly version and consider whether a transcriptome-based reference might be more appropriate for their analysis.
Annotation File Considerations
The gene annotation file provides STAR with known transcript structures. When provided during index generation, STAR builds a splice junction database that improves alignment of reads spanning known exon-exon junctions. This is particularly important for RNA-seq data because reads spanning junctions cannot be aligned contiguously to the genome.
For organisms with well-annotated genomes, using the annotation file during index generation is strongly recommended. The annotation should match the genome assembly version. Mismatched annotation and genome versions will cause spurious alignments and reduced mapping rates. For organisms with incomplete annotations, researchers may choose to build the index without an annotation file and rely on STAR's de novo junction discovery. This approach is appropriate for transcript discovery studies but will reduce sensitivity for known splice junctions.
Index Generation Parameters
The --runMode genomeGenerate command builds the index. Key parameters include --genomeDir to specify the output directory, --genomeFastaFiles to provide the reference FASTA, and --sjdbGTFfile to provide the annotation. The --sjdbOverhang parameter specifies the length of the genomic sequence around annotated splice junctions to include in the database. The recommended value is read length minus one, so for 100 bp paired-end reads, the value should be 99.
The --genomeSAindexNbases parameter controls the size of the suffix array index. For genomes larger than approximately 3 billion bases, the default value of 14 is appropriate. For smaller genomes, STAR will issue a warning recommending a smaller value. Ignoring this warning can cause the index generation to fail or produce an index that uses excessive memory. The --genomeChrBinNbits parameter controls how the genome is divided into bins for parallel processing and may need adjustment for genomes with many small contigs.
Index generation is computationally intensive and may require substantial memory. For mammalian genomes, 30 GB of RAM is typically sufficient, but larger genomes or genomes with high repeat content may require more. The index should be generated once per genome and annotation combination and reused across all samples in a study.
Core Alignment Parameters
Output Format and Quantification Settings
The --outSAMtype parameter determines the output file format. The most common setting is BAM SortedByCoordinate, which produces a coordinate-sorted BAM file ready for downstream analysis. This setting is appropriate for most applications because sorted BAM files are required by many downstream tools, including featureCounts, HTSeq, and most variant callers. The --outSAMtype BAM Unsorted option produces an unsorted BAM file, which may be useful when the alignment order must match the input read order, such as for certain fusion detection algorithms.
The --quantMode parameter enables quantification during alignment. The GeneCounts option produces a read count per gene file that can be used directly for differential expression analysis. This option requires an annotation file in the index. The TranscriptomeSAM option produces alignments to the transcriptome instead of the genome, which is required for transcript-level quantification tools such as RSEM and Salmon. Both options can be enabled simultaneously with --quantMode TranscriptomeSAM GeneCounts.
Read Filtering Parameters
STAR applies several filters to determine which alignments are reported. The --outFilterMismatchNmax parameter sets the maximum number of mismatches allowed per read. The default value of 10 is appropriate for standard 100 bp reads from human and mouse. For reads from organisms with higher sequence divergence from the reference, such as non-model organisms or samples with high SNP density, this value may need to be increased. However, increasing this parameter also increases the chance of spurious alignments.
The --outFilterMultimapNmax parameter controls how many genomic loci a read can map to and still be reported. The default value of 10 means that reads mapping to up to 10 locations are reported with all their alignments. Reads mapping to more than 10 locations are not reported. For analyses that require uniquely mapped reads only, this parameter should be set to 1. For studies of repetitive elements or multi-copy genes, higher values may be appropriate.
The --outFilterScoreMinOverLread and --outFilterMatchNminOverLread parameters set minimum alignment scores as fractions of read length. The default value of 0.66 for score and 0 for match means that alignments must achieve at least 66% of the maximum possible score. These parameters provide a global quality threshold that can be adjusted for specific applications. For degraded RNA samples or very short reads, lowering these thresholds may increase mapping rates at the cost of alignment accuracy.
Splice Junction Parameters
The --alignIntronMax parameter sets the maximum intron length. The default value of 0 means no limit, which is appropriate for most applications. However, for organisms with known intron size distributions, setting this parameter can reduce spurious alignments across very long genomic distances. For example, the maximum intron length in humans is approximately 1 Mb, and setting --alignIntronMax 1000000 can improve alignment specificity.
The --alignMatesGapMax parameter sets the maximum distance between paired-end reads. The default value of 0 means no limit. For standard RNA-seq libraries with insert sizes of 200 to 500 bp, this parameter can be set to a value that accommodates the expected insert size distribution plus the largest expected intron. Setting this parameter too low will cause valid alignments to be rejected.
The --alignSJoverhangMin parameter sets the minimum overhang required for a splice junction alignment. The default value of 5 means that at least 5 bases must overhang the junction on each side. For short reads, this parameter may need to be lowered to detect junctions near read ends. For long reads, higher values can improve junction specificity.
Output Metrics and Their Interpretation
Log Files and Mapping Statistics
STAR produces a Log.final.out file for each sample that contains the key mapping statistics. This file should be examined for every sample before proceeding with downstream analysis. The most important metrics include the total number of input reads, the number of uniquely mapped reads, the number of reads mapped to multiple loci, and the number of reads mapped to too many loci.
The percentage of uniquely mapped reads is the most commonly cited quality metric. For high-quality RNA-seq data from well-annotated organisms, this value typically exceeds 80%. Lower values may indicate sample contamination, reference genome issues, or parameter problems. However, the interpretation of this metric depends on the organism and the analysis goals. A study of repetitive elements will necessarily have lower uniquely mapped rates than a study of protein-coding genes.
The number of reads mapped to multiple loci is reported as a percentage of total reads. High multimapping rates can indicate the presence of repetitive sequences or multi-copy genes in the sample. For differential expression analysis, multimapping reads are often excluded or distributed among their mapped locations. The choice of how to handle multimapping reads should be documented and consistent across all samples in a study.
Splice Junction Metrics
The SJ.out.tab file contains information about detected splice junctions, including their genomic coordinates, whether they are annotated or novel, and the number of reads supporting each junction. This file is useful for assessing the complexity of splicing in the sample and for identifying potential artifacts.
The number of novel splice junctions detected can indicate whether the annotation file is complete for the organism and cell type being studied. A very high proportion of novel junctions may suggest that the annotation is incomplete or that the sample contains unexpected transcript isoforms. A very low proportion of novel junctions may indicate that the alignment parameters are too stringent for detecting new junctions.
Chimeric Read Metrics
Chimeric reads are reads that map to two different genomic locations, potentially indicating structural rearrangements or gene fusions. STAR detects chimeric alignments when the --chimSegmentMin parameter is set to a value greater than 0. The default value of 0 disables chimeric alignment detection.
The number of chimeric reads in a sample depends on the biological context and the detection parameters. Cancer samples may have many chimeric reads due to genomic rearrangements, while normal samples typically have few. The interpretation of chimeric read counts requires careful consideration of the expected biology and the detection parameters used.
Practical Workflow for STAR Alignment
Step 1: Quality Control of Raw Reads
Before alignment, raw sequencing reads should be assessed for quality. Tools such as FastQC provide per-base quality scores, GC content, adapter contamination, and overrepresented sequences. The Galaxy Training Network provides accessible workflow training and analysis tutorials that cover quality control procedures for RNA-seq data. Low-quality bases and adapter sequences should be trimmed before alignment to improve mapping rates and reduce spurious alignments.
The decision to trim should be based on the quality assessment results. If adapter contamination is detected, trimming is necessary because adapter sequences will not map to the genome and will reduce the effective read length. If quality scores drop significantly toward the 3' end of reads, trimming low-quality bases can improve alignment accuracy. However, aggressive trimming reduces read length and may reduce the ability to detect splice junctions.
Step 2: Index Generation
Generate the genome index using the appropriate reference genome and annotation files. Document the genome assembly version, annotation version, and all index generation parameters. This documentation is essential for reproducibility and for interpreting results across studies.
The index generation step should be performed once per genome and annotation combination. The resulting index directory can be reused for all samples in a study. The index should be stored in a location with sufficient disk space and should be backed up to prevent loss.
Step 3: Alignment of Individual Samples
Run STAR for each sample using the appropriate parameters. The basic command includes --runMode alignReads, --genomeDir pointing to the index, --readFilesIn specifying the FASTQ files, and --outSAMtype BAM SortedByCoordinate. For paired-end data, both read files are specified. For gzipped FASTQ files, the --readFilesCommand zcat parameter is required.
The alignment step is computationally intensive and can be parallelized using the --runThreadN parameter. The number of threads should be set based on the available computational resources. STAR scales well with multiple threads, but the memory requirement also increases with thread count.
Step 4: Examine Alignment Statistics
After alignment, examine the Log.final.out file for each sample. Compare the mapping statistics across samples in the study to identify outliers. Samples with substantially lower uniquely mapped rates or higher multimapping rates than the rest of the study should be investigated.
The alignment statistics should be recorded in a sample metadata table that includes the sample identifier, the total read count, the uniquely mapped read count, the multimapping read count, and the unmapped read count. This table serves as a quality control record and should be included in the study documentation.
Step 5: Downstream Analysis
The aligned BAM files can be used for various downstream analyses, including gene quantification, differential expression analysis, isoform quantification, and fusion detection. The choice of downstream analysis determines which STAR output files are needed. For gene-level quantification, the ReadsPerGene.out.tab file produced by --quantMode GeneCounts can be used directly. For transcript-level quantification, the transcriptome BAM file produced by --quantMode TranscriptomeSAM is required.
The Bioconductor project provides official package and workflow documentation for many downstream analysis tools, including DESeq2, edgeR, and limma for differential expression analysis. These tools accept gene count matrices as input and provide statistical frameworks for identifying differentially expressed genes.
Options and Tradeoffs in STAR Configuration
Default Parameters Versus Custom Configuration
The question of whether to use default parameters or customize STAR settings depends on the specific application. Research on STAR's performance across a range of alignment parameters found that technical metrics such as fraction mapping or expression profile correlation are uninformative, capturing properties unlikely to have any role in biological discovery. The same study found that changes in alignment parameters within a wide range have little impact on both technical and biological performance. However, when performance does break, it happens in difficult regions such as X-Y paralogs and MHC genes.
This finding suggests that for most standard RNA-seq analyses, default parameters are appropriate and produce results that are robust to parameter variation. Researchers should focus their attention on ensuring that the reference genome and annotation are correct and that the sample quality is adequate, instead of on fine-tuning alignment parameters.
Parameter Sensitivity in Specific Contexts
While default parameters are appropriate for most applications, certain contexts require parameter adjustment. For organisms with high sequence divergence from the reference, such as non-model organisms or samples with high SNP density, the mismatch tolerance may need to be increased. For studies of repetitive elements or multi-copy genes, the multimapping filter may need to be adjusted.
Benchmarking studies using the Arabidopsis thaliana genome found that STAR's overall performance was superior to other aligners at the read base-level assessment, with overall accuracy reaching over 90% under different test conditions. However, junction base-level assessment produced varying results depending upon the applied algorithm. This finding highlights the importance of evaluating alignment performance in the context of the specific biological question.
Paired-End Versus Single-End Data
Paired-end sequencing provides additional information that can improve alignment accuracy, particularly for reads spanning splice junctions and for resolving ambiguous mappings. STAR uses the paired-end information to filter alignments that are inconsistent with the expected insert size and orientation.
For paired-end data, the --alignMatesGapMax parameter should be set to accommodate the expected insert size distribution. The default value of 0 means no limit, which is appropriate when the insert size distribution is unknown. However, setting this parameter can improve alignment specificity by rejecting alignments with implausible insert sizes.
Single-end data requires more careful parameter selection because there is no paired-end information to resolve ambiguous mappings. The --outFilterMultimapNmax parameter becomes more important for single-end data because multimapping reads cannot be resolved by mate information.
Records and Measurements for Quality Assessment
Sample-Level Quality Metrics
Each aligned sample should have a record of the following metrics from the Log.final.out file:
| Metric | Interpretation | Action Threshold |
|---|---|---|
| Uniquely mapped reads percentage | Proportion of reads with unique genomic location | Investigate if below 70% for well-annotated organisms |
| Multimapping reads percentage | Proportion of reads mapping to multiple locations | Investigate if above 20% for standard RNA-seq |
| Reads mapped to too many loci | Reads exceeding the multimapping limit | Investigate if above 5% for standard RNA-seq |
| Splice junction reads | Reads spanning exon-exon junctions | Expected to be 20-40% for mRNA libraries |
| Chimeric reads | Reads mapping to two genomic locations | Investigate if above 5% for non-cancer samples |
| Unmapped reads percentage | Reads not mapped to the reference | Investigate if above 15% for standard RNA-seq |
These thresholds are general guidelines and should be adjusted based on the organism, library type, and analysis goals. The key is to identify samples that are outliers relative to the rest of the study and to investigate the cause of the deviation.
Study-Level Quality Assessment
Beyond individual sample metrics, the consistency of alignment statistics across samples in a study should be assessed. Large variation in uniquely mapped rates or total mapped reads across biological replicates may indicate sample quality issues or technical batch effects.
The alignment statistics should be recorded in a machine-readable format, such as a CSV or TSV file, to facilitate automated quality assessment. This file should be included in the study's data repository to enable reproducibility and reanalysis.
Documentation of Parameters
All STAR parameters used for index generation and alignment should be documented for each study. This documentation should include the STAR version, the reference genome assembly version, the annotation file version, and all non-default parameter settings. The nf-core documentation provides community pipeline standards for usage and configuration that emphasize the importance of parameter documentation for reproducibility.
The documentation should be stored with the analysis results and should be referenced in any publications resulting from the study. This practice enables other researchers to reproduce the analysis and to assess the impact of parameter choices on the results.
Common Failure Patterns and Troubleshooting
Low Uniquely Mapped Read Percentage
When the uniquely mapped read percentage is lower than expected, several causes should be investigated. The first check is the quality of the raw reads. Low-quality reads with many sequencing errors will fail to map or will map with many mismatches. Re-examine the FastQC results and consider whether more aggressive quality trimming is needed.
The second check is the reference genome and annotation. If the reference genome is incomplete or contains errors, reads from the corresponding genomic regions will fail to map. If the annotation file is mismatched with the genome assembly, splice junction alignment will be impaired. Verify that the genome and annotation versions are compatible.
The third check is sample contamination. If the sample contains DNA from another organism, such as a pathogen or a contaminating cell line, reads from the contaminating organism will not map to the intended reference. Check for unexpected taxonomic composition using tools such as Kraken or by examining the unmapped reads for known contaminant sequences.
High Multimapping Rate
A high proportion of multimapping reads can indicate the presence of repetitive sequences or multi-copy genes in the sample. This is expected for certain sample types, such as those enriched for ribosomal RNA or containing many transposable elements. For standard mRNA-seq libraries, a high multimapping rate may indicate incomplete ribosomal RNA depletion.
The handling of multimapping reads depends on the downstream analysis. For differential expression analysis, multimapping reads are often excluded because their assignment to specific genes is ambiguous. Some analysis tools distribute multimapping reads among their mapped locations proportionally. The choice of approach should be documented and consistent across all samples.
Excessive Chimeric Reads
A high number of chimeric reads can indicate sample quality issues, such as template switching during library preparation, or biological phenomena, such as gene fusions in cancer samples. The interpretation depends on the expected biology and the detection parameters used.
For non-cancer samples, a high chimeric read count may indicate library preparation artifacts. Re-examine the library preparation protocol and consider whether the fragmentation or amplification steps could be introducing artifacts. For cancer samples, chimeric reads may indicate genuine gene fusions that are biologically relevant.
STAR Crashes or Memory Errors
STAR alignment requires substantial memory, particularly for mammalian genomes. If STAR crashes with a memory error, the --limitBAMsortRAM parameter may need to be increased. This parameter sets the maximum RAM available for sorting the BAM file. The default value is 0, which means STAR will use all available memory.
For very large genomes or genomes with high repeat content, the index generation step may require more memory than the default settings provide. The --genomeSAindexNbases parameter may need to be reduced for small genomes, and the --genomeChrBinNbits parameter may need adjustment for genomes with many small contigs.
Inconsistent Results Across Replicates
When alignment statistics vary substantially across biological replicates, the cause may be sample quality differences or technical batch effects. Examine the raw read quality metrics for each sample to identify outliers. Consider whether the samples were processed in different batches and whether batch effects could explain the variation.
The alignment statistics should be examined in the context of the experimental design. If the variation correlates with a known technical factor, such as sequencing run or library preparation batch, this factor should be included in the downstream statistical analysis.
Limitations of STAR Alignment
Multimapping Read Ambiguity
STAR cannot resolve the ambiguity of reads that map to multiple genomic locations. These reads are reported with all their alignments, and the choice of which alignment to use for quantification is left to the downstream analysis tool. This limitation is particularly relevant for genes with high sequence similarity, such as paralogs and members of gene families.
Research on STAR's performance found that when alignment performance breaks, it happens in difficult regions such as X-Y paralogs and MHC genes. These regions contain highly similar sequences that are difficult to distinguish even with careful parameter selection. Researchers studying these regions should be aware of the limitations of short-read alignment and may need to use additional approaches, such as long-read sequencing or targeted capture.
Reference Genome Dependence
STAR alignment is entirely dependent on the quality and completeness of the reference genome. Reads from regions that are absent from the reference or that contain assembly errors will not map correctly. This limitation is particularly relevant for non-model organisms with incomplete genome assemblies.
For organisms with poor genome assemblies, alternative approaches such as de novo transcriptome assembly or alignment to a closely related species' genome may be more appropriate. The choice of reference genome should be documented and justified in the study methods.
Splice Junction Detection Limits
STAR detects splice junctions by finding reads that span exon-exon boundaries. The sensitivity of junction detection depends on the read length and the alignment parameters. Short reads may not span entire junctions, and reads with very short overhangs on either side of a junction may not be detected.
The --alignSJoverhangMin parameter controls the minimum overhang required for junction detection. Lowering this parameter can increase sensitivity for junctions near read ends but may also increase false positive junction calls. The optimal setting depends on the read length and the expected junction architecture.
Computational Resource Requirements
STAR alignment requires substantial computational resources, particularly memory. For mammalian genomes, 30 GB of RAM is typically required for alignment, and index generation may require more. This requirement can be prohibitive for researchers with limited computational resources.
Alternative aligners with lower memory requirements are available, and the choice of aligner should be based on the specific requirements of the study. The Galaxy Training Network provides accessible workflow training that includes options for running STAR and other aligners on shared computational infrastructure.
Quality Control and Reproducibility Considerations
Reproducibility Standards
Reproducibility in RNA-seq analysis requires documentation of all analysis steps, including software versions, parameter settings, and reference data versions. The nf-core documentation provides community pipeline standards that emphasize reproducible workflow configuration. Following these standards ensures that the analysis can be reproduced by other researchers.
The alignment parameters and reference data versions should be recorded in a configuration file that is stored with the analysis results. This file should be referenced in the study methods and should be made available with the data.
Workflow Management
Workflow management systems can help ensure that the alignment step is performed consistently across all samples in a study. These systems track the inputs, outputs, and parameters of each analysis step and can automatically resume failed analyses. The nf-core documentation provides guidance on using workflow management systems for reproducible genomic analysis.
For researchers who prefer a graphical interface, the Galaxy Training Network provides accessible workflow training that covers RNA-seq analysis pipelines. These platforms handle the technical details of running STAR and other tools, allowing researchers to focus on the biological interpretation of results.
Data Management
The alignment output files, including BAM files and log files, should be stored in a structured directory layout that facilitates data management. The raw FASTQ files should be archived separately from the analysis outputs. The NCBI provides official descriptions of data submission procedures for sequence data, and researchers should consider depositing their raw data in appropriate repositories.
The alignment statistics should be recorded in a sample metadata table that is maintained throughout the study. This table should be updated as new samples are added and should be included in the final data submission.
Professional Escalation Criteria
When to Seek Expert Assistance
Certain alignment problems require expert assistance to resolve. If the uniquely mapped read percentage is below 50% for a well-annotated organism, or if the alignment statistics are highly inconsistent across samples in a study, consultation with a bioinformatics specialist is recommended.
The EMBL-EBI Training provides bioinformatics learning pathways and practical analysis education that can help researchers develop the skills needed to troubleshoot alignment problems. Researchers who encounter persistent alignment issues should consider seeking training or consultation.
When to Consider Alternative Tools
While STAR is appropriate for most RNA-seq applications, certain situations may require alternative tools. For transcript-level quantification, alignment-free tools such as Salmon or Kallisto may be more appropriate. For splice junction detection, tools that integrate transcriptome guidance with deep learning-based junction scoring have shown improved performance in some benchmarks.
A 2025 study comparing splice junction detection tools found that a transformer-based approach achieved the highest mean F1 score for splice junction detection, outperforming STAR and other aligners on a standard benchmark of human simulated datasets. However, this study used simulated data, and the performance on real biological data may differ. Researchers should evaluate the performance of different tools on their specific data before selecting an alignment strategy.
When to Revisit the Experimental Design
Persistent alignment problems may indicate issues with the experimental design instead of the alignment parameters. If all samples in a study have low mapping rates, the library preparation protocol or the sequencing platform may be the cause. If only specific samples have low mapping rates, sample quality or handling may be the issue.
The alignment statistics should be examined in the context of the experimental design. If the problems correlate with a specific experimental condition, the biological interpretation of the results may need to be reconsidered.
Frequently Asked Questions
What is the difference between uniquely mapped reads and multimapping reads?
Uniquely mapped reads align to exactly one location in the reference genome. Multimapping reads align to multiple locations with equal or nearly equal alignment scores. STAR reports both types of reads, but the handling of multimapping reads in downstream analysis depends on the tool and the analysis goals. For gene quantification, multimapping reads are often excluded or distributed among their mapped locations. The proportion of multimapping reads in a sample depends on the genome content and the read length.
How do I choose between --quantMode GeneCounts and --quantMode TranscriptomeSAM?
The GeneCounts option produces a read count per gene file that can be used directly for gene-level differential expression analysis. The TranscriptomeSAM option produces alignments to the transcriptome that are required for transcript-level quantification tools such as RSEM and Salmon. The choice depends on the analysis goal. For gene-level analysis, GeneCounts is sufficient. For isoform-level analysis, TranscriptomeSAM is required. Both options can be enabled simultaneously.
What does the --outFilterMultimapNmax parameter control?
This parameter sets the maximum number of genomic loci that a read can map to and still be reported. The default value of 10 means that reads mapping to up to 10 locations are reported with all their alignments. Reads mapping to more than 10 locations are not reported. Setting this parameter to 1 restricts the output to uniquely mapped reads only. Higher values are appropriate for studies of repetitive elements or multi-copy genes.
Why is my uniquely mapped read percentage lower than expected?
Several factors can cause low uniquely mapped read percentages. Low-quality reads with many sequencing errors will fail to map or will map with many mismatches. Reference genome or annotation issues can prevent reads from mapping correctly. Sample contamination with DNA from another organism will produce reads that do not map to the intended reference. The cause should be investigated by examining the raw read quality, the reference genome and annotation versions, and the unmapped reads for contaminant sequences.
How should I handle multimapping reads in my analysis?
The handling of multimapping reads depends on the downstream analysis. For differential expression analysis, multimapping reads are often excluded because their assignment to specific genes is ambiguous. Some analysis tools distribute multimapping reads among their mapped locations proportionally. The choice of approach should be documented and consistent across all samples in a study. For studies of repetitive elements or multi-copy genes, multimapping reads may be the primary signal of interest.
What is the --sjdbOverhang parameter and how do I set it?
The --sjdbOverhang parameter specifies the length of the genomic sequence around annotated splice junctions to include in the splice junction database during index generation. The recommended value is read length minus one. For 100 bp paired-end reads, the value should be 99. This parameter ensures that reads spanning annotated junctions can be aligned with the appropriate overhang on each side of the junction.
How do I detect gene fusions with STAR?
Gene fusion detection requires enabling chimeric alignment detection by setting --chimSegmentMin to a value greater than 0. The default value of 0 disables chimeric alignment detection. The chimeric reads detected by STAR can be used as input to fusion detection tools. The sensitivity and specificity of fusion detection depend on the detection parameters and the downstream analysis tools. A 2026 study comparing gene fusion detection algorithms found that detection rates varied substantially across algorithms, with some algorithms failing to detect fusions resulting from small deletions, lowly expressed fusions, and certain types of rearrangements.
What should I do if STAR crashes with a memory error?
STAR alignment requires substantial memory, particularly for mammalian genomes. If STAR crashes with a memory error, the --limitBAMsortRAM parameter may need to be increased. This parameter sets the maximum RAM available for sorting the BAM file. The default value is 0, which means STAR will use all available memory. For very large genomes, the index generation step may also require more memory than the default settings provide.
Related Bioinformatics Guides
- RNA-Seq Alignment Tools: STAR, HISAT2, and Beyond
- RNA-Seq Alignment: Choosing the Right Tool and Parameters
- RNA-Seq vs DNA-Seq: Key Differences and Applications
- RNA-Seq vs ChIP-Seq: Complementary Approaches for Gene Regulation
- RNA-Seq vs Microarray: Choosing the Right Gene Expression Profiling Platform
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- Optimizing RNA-Seq Mapping with STAR.. Methods in molecular biology (Clifton, N.J.), 2016.
- The fractured landscape of RNA-seq alignment: the default in our STARs.. Nucleic acids research, 2018.
- Benchmarking RNA-Seq Aligners at Base-Level and Junction Base-Level Resolution Using the Arabidopsis thaliana Genome.. Plants (Basel, Switzerland), 2024.
- Transcriptomic Profile of Lin(-)Sca1(+)c-kit (LSK) cells in db/db mice with long-standing diabetes.. BMC genomics, 2024.
- A comprehensive transcriptomic dataset of Sorghum bicolor seedlings under abiotic stress conditions.. 2026.
- Genome concatenation enables accurate dual RNA-seq mapping for lignocellulolytic fungi in coculture.. 2026.
- Comparison of gene fusion detection algorithms reveals frequently overlooked driver fusions in hematologic malignancies.. 2026.
- Identifying and Visualizing Novel Small Open Reading Frames (sORFs) Via Ribo-Seq, Transcriptome Assembly, and ggRibo.. 2026.
- THRAISE: An automated and reproducible web platform for RNA-seq analysis.. 2025.
- Processed nascent RNA from Hetzel et al., 2016. 2019.
- DeepSAP: improved RNA-seq alignment by integrating transcriptome guidance with transformer-based splice junction scoring. bioRxiv, 2025.
- RNA-seq Insights into Major Regulatory Networks in Pediatric Leukemia. Research journal of biotechnology, 2025.
- The phylogeny of extant starfish (Asteroidea: Echinodermata) including Xyloplax, based on comparative transcriptomics. Molecular Phylogenetics and Evolution, 2017.
- Sample size estimation for detection of splicing events in transcriptome sequencing data. International Journal of Molecular Sciences, 2017.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.