RNA-seq Quality Control Pipeline: A Complete Workflow from Raw FASTQ to Count Matrix

By Dr. Zubair Khalid, DVM, MS, PhD ·

RNA-seq Quality Control Pipeline: A Complete Workflow from Raw FASTQ to Count Matrix

Key Takeaways

  • Raw FASTQ file integrity is paramount; verify file structure (header, sequence, quality lines of equal length) and checksums against repository data to prevent corrupted downloads from mimicking biological variation.
  • Initial quality assessment with FastQC is critical, focusing on per-base quality scores (declining towards 3' end is typical for Illumina), adapter content (indicates need for trimming), and sequence duplication levels (excessive levels suggest low input RNA or PCR bias).
  • Read trimming and filtering, using tools like Trimmomatic or cutadapt, addresses adapter contamination and low-quality bases (e.g., below Q20/Q30) and removes short reads (e.g., <36 bp) to improve alignment specificity.
  • Alignment to a reference genome or transcriptome (e.g., using STAR or HISAT2) is assessed by mapping rate and unique alignment proportion; high intronic mapping can signal genomic DNA contamination, while high intergenic mapping may indicate alignment issues.
  • Quantification methods (e.g., featureCounts for counts, Salmon for transcript-level) generate a count matrix, with the number of detected genes per sample serving as a key metric for library complexity and sequencing depth.
  • Post-alignment QC includes evaluating gene body coverage (uniformity is ideal, 5' or 3' bias suggests degradation or protocol issues), assessing rRNA contamination (typically <10% of aligned reads), and performing sample clustering (e.g., PCA) to detect batch effects or sample mix-ups.

RNA sequencing produces large volumes of raw sequence data that require systematic processing before biological interpretation is possible. The path from raw FASTQ files to a usable count matrix involves multiple stages, each with distinct quality concerns that can compromise downstream results if left unaddressed. This workflow covers each step from raw data acquisition through alignment and quantification, with concrete decision points and quality thresholds that researchers can apply to their own datasets.

The workflow described here applies to bulk RNA-seq experiments, where gene expression is measured across populations of cells. Single-cell RNA-seq requires additional considerations beyond the scope of this article, though many of the foundational quality control principles transfer directly. Researchers working with single-cell data should consult specialized protocols that address cell-level quality metrics, ambient RNA contamination, and doublet detection [<a href="#ref-1">1</a>].

At a Glance

The table below summarizes the major pipeline stages, the primary quality concerns at each stage, and the tools commonly used for assessment.

Pipeline StagePrimary Quality ConcernsCommon Assessment Tools
Raw data acquisitionFile integrity, format correctness, base call qualitymd5sum, NCBI SRA tools, FastQC
Read trimming and filteringAdapter contamination, low-quality bases, short read removalTrimmomatic, cutadapt, fastp
Alignment to referenceMapping rate, unique alignment proportion, rRNA contaminationSTAR, HISAT2, Salmon, samtools flagstat
QuantificationGene-level counts, transcript-level estimates, detected gene numbersfeatureCounts, HTSeq, Salmon, RSEM
Post-alignment QCSample clustering, batch effects, outlier detection, gene body coverageMultiQC, RSeQC, edgeR, DESeq2

Understanding RNA-seq Data Inputs and File Formats

RNA-seq experiments produce raw data in FASTQ format, which contains nucleotide sequence reads along with associated quality scores for each base. The quality scores are encoded as ASCII characters and represent the probability of a base call being incorrect. Understanding the structure of FASTQ files is essential for troubleshooting quality issues at the earliest stage of the pipeline.

Most sequencing facilities deliver data as compressed FASTQ files, often with one file pair per sample for paired-end sequencing. Paired-end reads provide two reads from each fragment, one from each end, which improves alignment accuracy and enables detection of splice junctions. The choice between single-end and paired-end sequencing affects downstream analysis options, with paired-end data generally providing better performance for transcript-level quantification and isoform detection [<a href="#ref-2">2</a>].

Raw sequencing data is frequently deposited in public repositories such as the NCBI Sequence Read Archive, which provides standardized formats for data submission and retrieval [<a href="#ref-3">3</a>]. When downloading public datasets, researchers should verify file integrity using checksums provided by the repository. Corrupted downloads can introduce errors that mimic biological variation and are difficult to detect after alignment.

The European Bioinformatics Institute provides training materials on data resources and analysis approaches that cover file formats and data retrieval methods [<a href="#ref-4">4</a>]. Familiarity with these resources helps researchers navigate the practical aspects of data acquisition and format conversion.

File Format Verification Steps

Before beginning any analysis, verify that each FASTQ file meets basic structural requirements. Each read should occupy exactly four lines in the file: a header line beginning with the @ symbol, the nucleotide sequence, a separator line beginning with the + symbol, and the quality score string. The sequence and quality lines must have equal length. Files that fail these checks may be truncated or corrupted during transfer.

Record the file size and line count for each sample. Unexpected discrepancies between paired-end files from the same sample may indicate incomplete transfer or sequencing problems. The checksums provided by the sequencing facility or repository should match the downloaded files exactly. Any mismatch requires re-downloading before proceeding.

Read Length and Sequencing Platform Considerations

Read length varies by sequencing platform and configuration. Illumina platforms commonly produce reads of 50, 75, 100, or 150 bases. Longer reads improve alignment specificity and enable better detection of splice junctions, but they also increase sequencing cost. The optimal read length depends on the research question and the complexity of the transcriptome being studied.

For paired-end data, the insert size distribution affects the ability to detect splice junctions and isoforms. Very short inserts may result in overlapping reads, while very long inserts may reduce the efficiency of paired-end alignment. The insert size distribution can be assessed after alignment using tools that examine the distance between paired reads.

Initial Quality Assessment with FastQC

FastQC is the standard first step in RNA-seq quality assessment. It generates a report containing multiple quality modules, including per-base quality scores, per-sequence quality scores, GC content distribution, adapter content, and sequence duplication levels. Each module provides a pass, warning, or failure status based on built-in thresholds.

Per-base quality scores are displayed as box plots across read positions. For Illumina sequencing, quality scores typically decline toward the 3' end of reads. A sharp drop in quality at specific positions may indicate technical issues with the sequencing run. The warning threshold for per-base quality is typically triggered when any base position falls below a quality score of 20 in more than 15% of reads, though researchers should examine the actual plots instead of relying solely on automated flags.

GC content distribution should approximate a normal distribution for most genomes, though RNA-seq libraries can show deviations due to transcript composition. A narrow abnormal peak in GC content may indicate adapter contamination or PCR amplification bias. The presence of overrepresented sequences, particularly adapter sequences, signals the need for trimming before alignment.

FastQC reports should be generated for every sample in the experiment before any processing steps. Aggregating FastQC reports across all samples using MultiQC provides a convenient overview and helps identify systematic issues affecting multiple samples simultaneously. This aggregation step is particularly valuable for large experiments where examining individual reports becomes impractical.

Interpreting FastQC Module Results

The per-base sequence quality module displays the range of quality scores at each position across all reads. The background of the plot is colored green for high quality, yellow for moderate quality, and red for low quality. Reads from Illumina platforms typically show high quality in the central portion with declining quality toward the 3' end. A sudden drop in quality at a specific position may indicate a problem with the sequencing chemistry or a specific cycle in the run.

The per-sequence quality scores module shows the distribution of average quality scores across all reads. Most reads should have high average quality. A substantial population of reads with low average quality may indicate problems with library preparation or sequencing. These reads may need to be filtered before alignment.

The adapter content module shows the cumulative percentage of reads containing adapter sequences at each position. Adapter contamination appears as an increasing curve toward the 3' end of reads. The presence of adapter sequences at high levels indicates that the insert fragments are shorter than the read length, and trimming is required.

The sequence duplication levels module shows the distribution of duplicate reads in the library. Some duplication is expected in RNA-seq due to high expression of abundant transcripts. However, excessive duplication across the entire library may indicate low input RNA or PCR amplification bias.

Recording Initial Quality Observations

For each sample, record the following observations from the initial FastQC report: the number of reads, the read length, the per-base quality pattern, the GC content distribution, the adapter content level, and any modules that fail. This record provides a baseline for comparison after trimming and filtering. Samples with unusual initial quality profiles may require additional attention at later stages.

The MultiQC aggregation report should be saved as part of the project records. This report provides a single view of all samples and enables quick identification of samples that deviate from the overall pattern. Systematic issues affecting multiple samples are often more visible in the aggregated view than in individual reports.

Read Trimming and Adapter Removal

Adapter contamination occurs when the sequencing read extends through the entire insert fragment and into the adapter sequence. This is common when fragment sizes are shorter than the read length. Adapter sequences interfere with alignment because they do not match the reference genome and can cause reads to align incorrectly or fail to align entirely.

Trimming tools such as Trimmomatic and cutadapt identify and remove adapter sequences, low-quality bases, and short reads. The specific parameters depend on the library preparation kit and sequencing platform used. Most tools support automatic adapter detection, though specifying the expected adapter sequences improves accuracy.

Quality trimming removes low-confidence bases from read ends. A common approach is to trim bases from the 3' end until the quality score exceeds a threshold, typically Q20 or Q30. The choice of threshold balances read retention against data quality. Aggressive trimming removes more low-quality bases but also shortens reads, which can reduce alignment specificity.

After trimming, reads that fall below a minimum length threshold should be removed. Very short reads are unlikely to align uniquely and can introduce noise into quantification. A minimum length of 36 bases is commonly used, though the optimal threshold depends on read length and the reference genome complexity.

The decision to trim should be guided by the FastQC results. If adapter content is minimal and quality scores are uniformly high, trimming may be unnecessary. However, most real-world datasets benefit from at least light quality trimming. The pipeline should record trimming statistics, including the number of reads retained, the number removed, and the reasons for removal, to enable quality assessment at later stages.

Trimming Parameter Selection

The choice of trimming parameters requires balancing several considerations. Adapter trimming should be performed when the FastQC adapter content module shows significant contamination. The adapter sequences used by the library preparation kit should be specified explicitly when known. Automatic adapter detection can identify unexpected adapters but may also remove legitimate sequence in some cases.

Quality trimming thresholds should be selected based on the overall quality profile of the dataset. For datasets with generally high quality, a threshold of Q20 removes only the lowest quality bases. For datasets with lower quality, a threshold of Q30 may be appropriate to remove more low-quality sequence. The sliding window approach used by Trimmomatic evaluates quality across a window of bases and trims when the average quality falls below the threshold.

The minimum read length threshold determines which reads are retained after trimming. Reads shorter than the threshold are discarded because they are unlikely to align uniquely. The optimal threshold depends on the read length and the complexity of the reference. For 150-base reads, a minimum length of 36 bases retains reads that cover at least one exon junction. For shorter reads, the minimum length should be adjusted accordingly.

Evaluating Trimming Effectiveness

After trimming, run FastQC again on the processed reads to confirm that adapter contamination has been removed and quality scores have improved. Compare the before and after reports to document the effect of trimming. The adapter content should be near zero, and the per-base quality should be consistently high across the retained read positions.

Record the proportion of reads retained after trimming. A very low retention rate may indicate severe quality problems or an issue with the trimming parameters. A very high retention rate suggests that trimming had minimal effect, which may be appropriate if the data was already high quality. The retention rate should be consistent across samples from the same experiment.

Alignment to Reference Genome or Transcriptome

Read alignment maps sequencing reads to a reference genome or transcriptome. The choice of alignment strategy depends on the research question and the availability of a high-quality reference. Genome alignment is preferred when studying splice variants, novel transcripts, or genomic features beyond simple gene expression. Transcriptome alignment is faster and can be more accurate for quantification when the transcript annotation is complete [<a href="#ref-2">2</a>].

STAR is a widely used splice-aware aligner that can handle large genomes and complex splice junctions. It requires a genome index built from the reference genome and annotation file. The index building step is computationally intensive but only needs to be performed once per genome version. STAR produces aligned reads in BAM format along with alignment statistics that include mapping rates and the proportion of uniquely mapped reads.

HISAT2 is an alternative splice-aware aligner that uses a hierarchical indexing strategy. It is generally faster and requires less memory than STAR, making it suitable for computing environments with limited resources. The choice between STAR and HISAT2 often depends on personal preference and the specific requirements of the analysis.

Alignment quality is assessed using several metrics. The overall mapping rate indicates what proportion of reads aligned to the reference. Low mapping rates may indicate contamination, adapter problems, or a mismatch between the sample species and the reference genome. The proportion of uniquely mapped reads is particularly important for quantification, as reads that map to multiple locations cannot be confidently assigned to a specific gene.

For RNA-seq data, the proportion of reads mapping to exonic regions versus intronic and intergenic regions provides insight into library quality. High intronic mapping may indicate genomic DNA contamination, while high intergenic mapping may indicate alignment problems or annotation issues. The Sequencing Quality Control Consortium found that measurement performance depends on both the platform and the data analysis pipeline, with variation being particularly large for transcript-level profiling [<a href="#ref-5">5</a>].

Reference Genome Selection and Index Building

The reference genome version and annotation file must be selected carefully. Different versions of the same genome may produce different alignment results, particularly in regions with structural variation or annotation updates. The reference genome and annotation versions should be recorded for reproducibility.

Genome indexes should be built with the same reference and annotation files that will be used for the actual alignment. Mismatches between the index and the reference can cause subtle alignment errors. The index building process should be documented, including the software version and parameters used.

For species with well-annotated genomes, the primary assembly is typically used. For species with multiple assemblies, the choice of assembly should be based on the quality of the annotation and the relevance to the research question. The NCBI provides access to reference genomes and annotations for many species [<a href="#ref-3">3</a>].

Alignment Statistics and Interpretation

The alignment summary produced by STAR or HISAT2 includes the total number of reads, the number and percentage of reads mapped to multiple loci, the number and percentage of reads mapped to too many loci, and the number and percentage of reads unmapped. These statistics provide the first indication of alignment quality.

The percentage of uniquely mapped reads is a key quality metric. For high-quality RNA-seq data, this percentage is typically above 70%. Lower values may indicate repetitive sequence content, alignment parameter issues, or problems with the reference genome. The percentage of reads mapped to multiple loci should be reported, as these reads are typically excluded from quantification.

The percentage of reads unmapped for various reasons should be examined. Reads that are unmapped because they are too short may indicate trimming was too aggressive. Reads that are unmapped because they contain too many mismatches may indicate contamination or reference mismatch. Reads that are unmapped because they are chimeric may indicate structural rearrangements or alignment artifacts.

Exonic, Intronic, and Intergenic Mapping Distribution

The distribution of aligned reads across genomic features provides insight into library quality. Tools such as RSeQC can calculate the percentage of reads mapping to exons, introns, and intergenic regions. For a typical mRNA library, the majority of reads should map to exons.

High intronic mapping may indicate genomic DNA contamination in the library. This can occur when DNase treatment is incomplete during library preparation. High intergenic mapping may indicate alignment problems, annotation issues, or the presence of unannotated transcripts. The expected distribution depends on the library preparation protocol and the annotation completeness.

The gene body coverage profile shows the read coverage across the length of transcripts. This profile can reveal 5' or 3' bias that may indicate RNA degradation or library preparation issues. The Area Under the Gene Body Coverage Curve has been proposed as a metric for assessing sample quality, with lower values associated with degraded samples [<a href="#ref-6">6</a>].

Quantification of Gene and Transcript Expression

Quantification converts aligned reads into expression measurements. The two main approaches are count-based methods and transcript-level quantification methods. Count-based methods, such as featureCounts and HTSeq, assign reads to genomic features based on the alignment and annotation. Transcript-level methods, such as Salmon and RSEM, estimate transcript abundances using probabilistic models.

Count-based methods produce a count matrix where each row represents a gene and each column represents a sample. The counts represent the number of reads assigned to each gene. This matrix serves as the input for differential expression analysis tools such as DESeq2 and edgeR. The choice of counting method and annotation version can significantly affect the results, so these parameters should be recorded and reported.

Transcript-level quantification methods estimate expression at the transcript isoform level. These methods can handle reads that map to multiple transcripts from the same gene by using expectation-maximization algorithms to allocate reads probabilistically. Salmon and other alignment-free methods can quantify expression directly from raw reads without a separate alignment step, which reduces computational requirements.

The number of detected genes per sample is a useful quality metric. Samples with unusually low gene detection may indicate RNA degradation, low sequencing depth, or technical problems. The integration of multiple quality metrics, including uniquely aligned reads, rRNA reads, and detected genes, provides a more reliable assessment of sample quality than any single metric alone [<a href="#ref-6">6</a>].

Count-Based Quantification Parameters

FeatureCounts and HTSeq require an annotation file in GTF or GFF format. The annotation file must be compatible with the reference genome used for alignment. The choice of annotation version can affect the number of reads assigned to each gene, particularly for genes with multiple isoforms.

The counting mode determines how reads that overlap multiple features are handled. The default mode for featureCounts assigns each read to a single feature, prioritizing features based on the order in the annotation file. The multi-mapping mode assigns reads that map to multiple locations to all overlapping features, which may be appropriate for some analyses.

The strand-specificity setting must match the library preparation protocol. For strand-specific libraries, reads should be assigned only to features on the correct strand. Incorrect strand settings can result in reads being assigned to the wrong genes, particularly for genes with overlapping antisense transcription.

Transcript-Level Quantification Considerations

Transcript-level quantification methods estimate expression at the isoform level using probabilistic models. These methods can resolve ambiguities in read assignment by allocating reads to transcripts based on their compatibility with the transcript sequences. The resulting transcript-level estimates can be summarized to the gene level for differential expression analysis.

Salmon and other alignment-free methods quantify expression directly from raw reads without a separate alignment step. These methods use a transcriptome index and can be faster than alignment-based approaches. However, they may be less sensitive to novel splice junctions and structural variants.

The choice between count-based and transcript-level quantification depends on the research question. Count-based methods are simpler and more widely used for gene-level differential expression analysis. Transcript-level methods provide additional information about isoform usage but require more complex downstream analysis.

Detected Gene Counts and Library Complexity

The number of detected genes per sample provides a measure of library complexity and sequencing depth. For mammalian genomes, a typical RNA-seq library detects 10,000 to 20,000 genes. Samples with substantially fewer detected genes may have degraded RNA, low sequencing depth, or technical problems.

The relationship between sequencing depth and detected genes is nonlinear. As sequencing depth increases, the number of detected genes increases but with diminishing returns. The Sequencing Quality Control Consortium found that measurement performance depends on sequencing depth and that deeper sequencing improves detection of lowly expressed genes [<a href="#ref-5">5</a>].

The proportion of reads assigned to genes is another useful metric. This proportion represents the fraction of aligned reads that overlap annotated genes. Low values may indicate annotation issues, alignment problems, or the presence of unannotated transcripts.

Post-Alignment Quality Control

Post-alignment QC examines the aligned data for issues that may not be apparent from alignment statistics alone. This includes checking gene body coverage, assessing rRNA contamination, evaluating strand specificity, and examining sample relationships.

Gene body coverage plots show the read coverage across the length of transcripts. Ideally, coverage should be relatively uniform across the gene body. A 5' bias may indicate RNA degradation, while a 3' bias may indicate issues with library preparation. The Area Under the Gene Body Coverage Curve has been proposed as a metric for assessing sample quality, with lower values associated with degraded samples [<a href="#ref-6">6</a>].

rRNA contamination is a common problem in RNA-seq libraries. Ribosomal RNA constitutes the majority of cellular RNA, and incomplete depletion during library preparation results in a large proportion of reads mapping to rRNA genes. High rRNA content reduces the effective sequencing depth for messenger RNA and can indicate problems with the library preparation protocol.

Strand-specific RNA-seq preserves information about which strand of the DNA the RNA was transcribed from. This information is useful for resolving overlapping genes and detecting antisense transcription. The alignment should be checked to confirm that the expected strand specificity is maintained. Incorrect strand settings during alignment can result in reads being assigned to the wrong genes.

Sample-level quality assessment examines relationships between samples using methods such as principal component analysis and hierarchical clustering. Samples from the same experimental group should cluster together, while samples from different groups should separate. Outlier samples that cluster with the wrong group may indicate sample mix-ups, contamination, or technical problems. These analyses can be performed using the count matrix and are often integrated into the early stages of differential expression analysis.

Gene Body Coverage Assessment

Gene body coverage plots are generated by dividing each transcript into a fixed number of bins and calculating the average read coverage in each bin. The resulting profile shows the coverage across the transcript from the 5' end to the 3' end. A uniform profile indicates that all regions of the transcript are equally represented.

A 5' bias, where coverage is higher at the 5' end of transcripts, may indicate RNA degradation. Degraded RNA has reduced coverage toward the 3' end because reverse transcription is less efficient on fragmented RNA. A 3' bias may indicate that the library preparation protocol enriches for the 3' end of transcripts, which is common in some protocols.

The Area Under the Gene Body Coverage Curve provides a quantitative summary of the coverage profile. This metric has been proposed as a quality indicator, with lower values associated with degraded samples [<a href="#ref-6">6</a>]. The metric can be compared across samples to identify samples with unusual coverage profiles.

rRNA Contamination Assessment

rRNA contamination can be assessed by counting the number of reads that align to rRNA genes. The rRNA content is typically reported as a percentage of total aligned reads. For well-prepared libraries, this percentage should be low, typically below 10%.

High rRNA content reduces the effective sequencing depth for messenger RNA. If rRNA contamination is detected, the rRNA reads can be filtered before quantification. However, this filtering reduces the total read count and may affect the sensitivity of downstream analysis.

For future experiments, the library preparation protocol should be reviewed to improve rRNA depletion efficiency. The choice of rRNA depletion method can significantly affect the level of rRNA contamination. Some methods are more effective for specific sample types or species.

Sample Clustering and Outlier Detection

Principal component analysis of the count matrix provides a low-dimensional view of sample relationships. The first few principal components typically capture the largest sources of variation, which may include biological differences, batch effects, and technical artifacts. Samples from the same experimental group should cluster together in the principal component plot.

Hierarchical clustering with a distance metric based on expression similarity provides another view of sample relationships. The resulting dendrogram shows the similarity between samples. Samples that cluster with the wrong group may indicate sample mix-ups or contamination.

Outlier samples should be examined carefully before exclusion. The decision to exclude an outlier should be based on evidence of technical problems instead of simply because the sample does not fit expectations. The integration of multiple quality metrics provides a more reliable basis for identifying low-quality samples [<a href="#ref-6">6</a>].

Reproducibility and Pipeline Management

RNA-seq analysis pipelines should be reproducible, meaning that the same inputs produce the same outputs regardless of when or where the analysis is performed. Reproducibility requires version control for both the analysis code and the software tools used. Container technologies such as Docker and Singularity can encapsulate the analysis environment, ensuring that software versions remain consistent.

Workflow management systems such as Snakemake and Nextflow provide structured frameworks for building analysis pipelines. These systems track the dependencies between steps, manage input and output files, and enable parallel execution on computing clusters. The nf-core project provides a collection of community-curated pipelines built with Nextflow that follow standardized practices for quality control and analysis [<a href="#ref-7">7</a>].

The Galaxy platform offers a web-based interface for running bioinformatics analyses without requiring command-line expertise. The Galaxy Training Network provides tutorials covering RNA-seq analysis workflows, including quality control, alignment, and differential expression [<a href="#ref-8">8</a>]. These resources are valuable for researchers who are new to bioinformatics or who prefer a graphical interface.

Version control for analysis scripts using Git enables tracking of changes and facilitates collaboration. The Carpentries provides lessons on foundational computing skills, including shell scripting, version control with Git, and data management practices [<a href="#ref-9">9</a>]. These skills are essential for building reproducible analysis pipelines.

Workflow Management System Selection

The choice of workflow management system depends on the computing environment and the complexity of the analysis. Snakemake uses a Python-based syntax and can run on a single machine or a computing cluster. Nextflow uses a Groovy-based syntax and provides built-in support for container execution and cloud computing.

The nf-core project provides a collection of community-curated pipelines built with Nextflow that follow standardized practices for quality control and analysis [<a href="#ref-7">7</a>]. These pipelines include RNA-seq analysis workflows that cover quality control, alignment, quantification, and differential expression. Using a community-curated pipeline can reduce the effort required to build and maintain an analysis pipeline.

The Compi RNA-seq pipeline provides a customizable framework that supports multiple counting strategies and downstream analysis tools [<a href="#ref-10">10</a>]. This pipeline is designed to be user-friendly while maintaining flexibility for customization. The pipeline includes tasks for FASTQ quality control, two alternative read-counting strategies, and tools for differential expression and pathway analysis.

Containerization and Environment Management

Container technologies such as Docker and Singularity encapsulate the analysis environment, including the operating system, software, and dependencies. Containers ensure that the analysis runs with the same software versions regardless of the host system. This is particularly important for long-term reproducibility, as software versions change over time.

Workflow management systems can automatically pull and use containers for each step of the pipeline. This ensures that each step runs with the specified software version. The container images should be recorded in the pipeline documentation to enable reproduction of the analysis.

Bioconductor provides a collection of R packages for genomic analysis, including packages for RNA-seq quality control, normalization, and differential expression [<a href="#ref-11">11</a>]. These packages are versioned and can be installed in a reproducible manner using the Bioconductor release system.

Documentation and Record Keeping

Detailed documentation of the analysis pipeline is essential for reproducibility and for troubleshooting problems that arise during analysis. The following information should be recorded for each sample:

  • Raw data file names and checksums
  • Sequencing platform and run parameters
  • FastQC results for raw and processed data
  • Trimming parameters and statistics
  • Alignment software version and parameters
  • Reference genome and annotation versions
  • Quantification method and parameters
  • Quality metrics at each stage

This information should be stored alongside the analysis scripts and results. Many workflow management systems automatically generate logs and reports that capture this information. The MultiQC report provides a convenient summary of quality metrics across all samples and pipeline stages.

For publications, the specific software versions and parameters used should be reported to enable others to reproduce the analysis. Public data repositories such as the NCBI provide mechanisms for depositing both raw data and processed results [<a href="#ref-3">3</a>]. The EMBL-EBI training resources provide guidance on data submission and reporting standards [<a href="#ref-4">4</a>].

Common Failure Patterns and Troubleshooting

Several quality problems recur frequently in RNA-seq datasets. Recognizing these patterns and understanding their causes enables researchers to make informed decisions about whether to proceed with analysis, apply corrective measures, or exclude problematic samples.

Low mapping rates are often caused by adapter contamination, sample contamination with other species, or a mismatch between the sample and the reference genome. Checking the FastQC reports for adapter content and examining the taxonomic composition of unmapped reads can help identify the cause. If the sample is contaminated with another species, the reads from the contaminating species may need to be filtered before alignment.

High duplication rates can indicate PCR amplification bias, particularly in libraries with low input RNA. Duplicate reads are reads that start at the same position and have the same sequence. While some duplication is expected in RNA-seq, excessive duplication reduces the effective library complexity and can bias expression measurements. Marking or removing duplicates may be appropriate for some analyses, though this is less common for RNA-seq than for DNA-seq.

GC bias refers to the tendency of sequencing to over- or under-represent regions with extreme GC content. This bias can affect both library preparation and sequencing. The FastQC GC content plot can reveal deviations from the expected distribution. Some quantification methods include corrections for GC bias, though these corrections are not universally applied.

Batch effects are systematic technical differences between groups of samples processed at different times or in different batches. These effects can confound biological differences if not properly accounted for. Including batch information in the experimental design and using statistical methods that model batch effects can help mitigate this problem. The eQTLQC pipeline provides automated quality control and normalization approaches for multi-source heterogeneous data, which is particularly relevant for studies combining data from multiple batches [<a href="#ref-12">12</a>].

Troubleshooting Low Mapping Rates

When the mapping rate is unexpectedly low, the first step is to examine the FastQC reports for the raw and trimmed reads. Adapter contamination can cause reads to fail alignment because the adapter sequence does not match the reference genome. If adapter contamination is present, additional trimming may be required.

The taxonomic composition of unmapped reads can be assessed using tools that classify reads by species. If a substantial proportion of unmapped reads match a different species, the sample may be contaminated. This can occur during sample collection, library preparation, or sequencing. The source of contamination should be identified and addressed.

A mismatch between the sample species and the reference genome can also cause low mapping rates. This can occur when samples are mislabeled or when the reference genome is not appropriate for the sample. The sample identity should be verified using genetic markers or other methods.

Addressing High Duplication Rates

High duplication rates can be identified from the FastQC sequence duplication levels module or from alignment statistics. For RNA-seq, some duplication is expected due to the high expression of abundant transcripts. However, excessive duplication across the entire library may indicate low input RNA or PCR amplification bias.

The decision to remove duplicates depends on the analysis goals. For differential expression analysis, duplicate removal is less common for RNA-seq than for DNA-seq because the duplication pattern is influenced by expression levels. However, for some applications, such as variant calling, duplicate removal may be appropriate.

If high duplication rates are observed, the library preparation protocol should be reviewed. Increasing the input RNA amount, reducing the number of PCR cycles, or using a library preparation method with lower amplification bias may reduce duplication rates.

Managing Batch Effects

Batch effects can be identified by examining sample clustering patterns. If samples cluster by batch instead of by experimental group, batch effects may be present. The experimental design should include batch information, and the analysis should account for batch effects using statistical methods.

The eQTLQC pipeline provides automated quality control and normalization approaches for multi-source heterogeneous data [<a href="#ref-12">12</a>]. This pipeline is designed for eQTL analysis but includes general quality control and normalization procedures that can be applied to other RNA-seq analyses.

When combining data from multiple batches, the batch information should be included in the statistical model. Methods such as ComBat can adjust for batch effects, though the effectiveness of these methods depends on the experimental design and the severity of the batch effects.

Limitations and Interpretation Considerations

RNA-seq quality control identifies technical problems but cannot fully correct for all sources of variation. Even with rigorous quality control, RNA-seq measurements contain biases that affect absolute expression levels. The Sequencing Quality Control Consortium found that RNA-seq does not provide accurate absolute measurements and that gene-specific biases are observed across platforms [<a href="#ref-5">5</a>]. Relative expression measurements between conditions are generally more reliable than absolute measurements.

The choice of analysis pipeline can significantly affect results. Different alignment and quantification methods can produce different expression estimates, particularly at the transcript level. Researchers should document their pipeline choices and consider how these choices might affect their conclusions. The Compi RNA-seq pipeline provides a customizable framework that supports multiple counting strategies and downstream analysis tools, allowing researchers to compare results across methods [<a href="#ref-10">10</a>].

Quality control thresholds should be applied with consideration of the experimental context. A sample with moderate quality issues may still be usable if the biological signal is strong and the issues are consistent across samples. Conversely, a sample with excellent technical metrics may still be problematic if it is an outlier in terms of biological content. The integration of multiple quality metrics provides a more robust assessment than any single metric [<a href="#ref-6">6</a>].

Understanding Technical Versus Biological Variation

RNA-seq measurements reflect both technical variation, which arises from the measurement process, and biological variation, which reflects genuine differences between samples. Quality control aims to minimize technical variation, but some technical variation is unavoidable. The relative contribution of technical and biological variation can be assessed using replicate samples.

Technical replicates, which are multiple libraries prepared from the same RNA sample, provide a measure of technical variation. Biological replicates, which are libraries prepared from different biological samples, provide a measure of total variation. The comparison of technical and biological variation can inform the design of future experiments.

The Sequencing Quality Control Consortium found that measurements of relative expression are accurate and reproducible across sites and platforms if specific filters are used [<a href="#ref-5">5</a>]. However, absolute measurements are less reliable, and gene-specific biases are observed across platforms. These findings highlight the importance of using appropriate normalization and statistical methods.

Pipeline Choice and Result Variability

Different analysis pipelines can produce different results from the same raw data. The choice of alignment method, quantification method, and annotation version can all affect the final count matrix. Researchers should document their pipeline choices and consider how these choices might affect their conclusions.

The Compi RNA-seq pipeline supports two alternative read-counting strategies and includes recently developed tools for differential expression and pathway analysis [<a href="#ref-10">10</a>]. This flexibility allows researchers to compare results across methods and assess the robustness of their findings.

For transcript-level analysis, the variation between methods is particularly large. The Sequencing Quality Control Consortium found that variation is large for transcript-level profiling [<a href="#ref-5">5</a>]. Researchers should be cautious when interpreting transcript-level results and should validate findings using independent methods.

Context-Dependent Quality Thresholds

Quality control thresholds should be applied with consideration of the experimental context. A sample with moderate quality issues may still be usable if the biological signal is strong and the issues are consistent across samples. The decision to exclude a sample should be based on the specific analysis goals and the potential impact of the quality issues.

The integration of multiple quality metrics provides a more robust assessment than any single metric [<a href="#ref-6">6</a>]. A sample that fails one quality metric may be acceptable if other metrics are within normal ranges. Conversely, a sample that passes all individual metrics may still be problematic if the metrics are inconsistent with each other.

The Quality Control Diagnostic Renderer (QC-DR) software was developed to simultaneously visualize a comprehensive panel of quality metrics and flag samples with aberrant values compared to a reference dataset [<a href="#ref-6">6</a>]. This approach supports the conclusion that any individual quality metric is limited in its predictive value and suggests approaches based on the integration of multiple metrics with quality thresholds.

Records and Documentation

Maintaining detailed records of the analysis pipeline is essential for reproducibility and for troubleshooting problems that arise during analysis. The following information should be recorded for each sample:

  • Raw data file names and checksums
  • Sequencing platform and run parameters
  • FastQC results for raw and processed data
  • Trimming parameters and statistics
  • Alignment software version and parameters
  • Reference genome and annotation versions
  • Quantification method and parameters
  • Quality metrics at each stage

This information should be stored alongside the analysis scripts and results. Many workflow management systems automatically generate logs and reports that capture this information. The MultiQC report provides a convenient summary of quality metrics across all samples and pipeline stages.

For publications, the specific software versions and parameters used should be reported to enable others to reproduce the analysis. Public data repositories such as the NCBI provide mechanisms for depositing both raw data and processed results [<a href="#ref-3">3</a>]. The EMBL-EBI training resources provide guidance on data submission and reporting standards [<a href="#ref-4">4</a>].

Quality Metrics Summary Table

The table below summarizes the key quality metrics to record at each pipeline stage, along with the typical interpretation of unusual values.

Quality MetricPipeline StageTypical Interpretation of Unusual Values
Per-base quality scoresRaw dataLow scores may indicate sequencing problems or instrument issues
Adapter contentRaw dataHigh content indicates short inserts and the need for trimming
Mapping rateAlignmentLow rate may indicate contamination, adapter problems, or reference mismatch
Uniquely mapped readsAlignmentLow proportion may indicate repetitive content or alignment issues
rRNA contentPost-alignmentHigh content indicates incomplete rRNA depletion
Detected genesQuantificationLow count may indicate degradation, low depth, or technical problems
Gene body coveragePost-alignment5' or 3' bias may indicate degradation or library preparation issues
Sample clusteringPost-alignmentUnexpected clustering may indicate batch effects or sample mix-ups

Reporting Standards for Publications

Publications should report the software versions and parameters used for each pipeline step, along with key quality metrics such as mapping rates, the number of detected genes, and the proportion of reads assigned to genes. The MultiQC report provides a convenient summary of these metrics. Reporting these details enables readers to assess the reliability of the results and to reproduce the analysis.

The specific quality metrics to report depend on the journal and the analysis goals. Some journals require detailed reporting of quality control steps, while others accept a summary. The key is to provide enough information for readers to assess the reliability of the results.

The NCBI provides mechanisms for depositing both raw data and processed results [<a href="#ref-3">3</a>]. Raw data should be deposited in the Sequence Read Archive, and processed data can be deposited in the Gene Expression Omnibus. The accession numbers should be reported in the publication.

Professional Escalation Criteria

Some quality problems require consultation with sequencing facility staff, bioinformatics support, or statistical experts. The following situations warrant escalation:

  • Complete sequencing failure or extremely low yield across all samples
  • Systematic quality problems affecting multiple samples that cannot be resolved with standard trimming and filtering
  • Evidence of sample contamination or sample mix-ups
  • Unexpected clustering patterns that suggest batch effects or technical artifacts
  • Discrepancies between technical quality metrics and expected biological patterns

When escalating, provide the relevant quality reports, pipeline logs, and experimental metadata to enable efficient troubleshooting. Document the steps already taken to address the problem and the specific questions that need resolution.

When to Contact the Sequencing Facility

The sequencing facility should be contacted when quality problems appear to originate from the sequencing run itself. This includes complete sequencing failure, extremely low yield, or systematic quality problems affecting multiple samples. The facility may be able to identify instrument issues or provide guidance on troubleshooting.

When contacting the sequencing facility, provide the run parameters, the quality reports, and the specific observations that suggest a sequencing problem. The facility may request additional information or may re-run the samples if the problem is identified.

When to Consult Bioinformatics Support

Bioinformatics support should be consulted when quality problems cannot be resolved with standard processing steps. This includes persistent low mapping rates, unexpected clustering patterns, or discrepancies between quality metrics and expected biological patterns. Bioinformatics experts can help identify the cause of the problem and recommend appropriate solutions.

When consulting bioinformatics support, provide the pipeline logs, the quality reports, and the specific questions that need resolution. The support team may request access to the raw data or the intermediate files to perform additional analysis.

When to Consult Statistical Experts

Statistical experts should be consulted when quality problems affect the statistical analysis of the data. This includes batch effects that cannot be adequately modeled, outlier samples that are difficult to classify, or questions about the appropriate normalization method. Statistical experts can help design the analysis to account for technical variation.

When consulting statistical experts, provide the experimental design, the quality metrics, and the specific statistical questions that need resolution. The experts may request the count matrix and the sample metadata to perform additional analysis.

Frequently Asked Questions

What is the minimum sequencing depth required for RNA-seq?

The required sequencing depth depends on the research question and the expression levels of the genes of interest. For differential expression analysis of highly expressed genes, 10 to 20 million reads per sample may be sufficient. Detecting lowly expressed genes or subtle expression changes requires greater depth. The Sequencing Quality Control Consortium found that measurement performance depends on sequencing depth and that deeper sequencing improves detection of lowly expressed genes [<a href="#ref-5">5</a>]. Researchers should consider their specific goals and consult published guidelines for their experimental context.

How do I choose between genome alignment and transcriptome alignment?

Genome alignment is necessary when studying splice variants, novel transcripts, or genomic features. Transcriptome alignment is faster and can be more accurate for gene-level quantification when the annotation is complete. The choice also depends on the downstream analyses planned. A survey of RNA-seq best practices notes that no single analysis pipeline works for all cases and that the choice of alignment strategy should be guided by the research question [<a href="#ref-2">2</a>].

What should I do if my samples have high rRNA contamination?

High rRNA content reduces the effective sequencing depth for messenger RNA. If rRNA contamination is detected, the first step is to confirm the finding using alignment statistics and rRNA-specific metrics. For the current dataset, rRNA reads can be filtered before quantification, though this reduces the total read count. For future experiments, the library preparation protocol should be reviewed to improve rRNA depletion efficiency.

How many biological replicates do I need?

The number of biological replicates required depends on the expected effect size, the variability between samples, and the desired statistical power. More replicates improve the reliability of differential expression results and enable detection of smaller effects. The best practices survey emphasizes that experimental design, including replication, is a critical first step in RNA-seq analysis [<a href="#ref-2">2</a>].

Can I combine data from different sequencing batches?

Combining data from different batches is possible but requires careful attention to batch effects. Batch information should be included in the statistical model, and the analysis should check for systematic differences between batches. The eQTLQC pipeline provides automated quality control and normalization approaches for multi-source heterogeneous data [<a href="#ref-12">12</a>]. If batch effects are severe, combining datasets may not be appropriate.

What is the difference between raw counts and normalized expression values?

Raw counts represent the number of reads assigned to each gene and are used as input for differential expression tools. Normalized expression values account for differences in sequencing depth and library composition between samples. The choice of normalization method can affect results, and different methods are appropriate for different experimental designs. The RNA-seq analysis pipeline described in recent methods chapters includes normalization as a distinct step between quantification and differential expression analysis [<a href="#ref-13">13</a>].

How do I identify and handle outlier samples?

Outlier samples can be identified using principal component analysis, hierarchical clustering, and quality metrics. Samples that cluster with the wrong experimental group or show extreme values for multiple quality metrics should be examined carefully. The decision to exclude an outlier should be based on evidence of technical problems instead of simply because the sample does not fit expectations. The integration of multiple quality metrics provides a more reliable basis for identifying low-quality samples [<a href="#ref-6">6</a>].

What quality metrics should I report in my publication?

Publications should report the software versions and parameters used for each pipeline step, along with key quality metrics such as mapping rates, the number of detected genes, and the proportion of reads assigned to genes. The MultiQC report provides a convenient summary of these metrics. Reporting these details enables readers to assess the reliability of the results and to reproduce the analysis.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

[1] [Current best practices in single-cell RNA-seq analysis: a tutorial.](https://pubmed.ncbi.nlm.nih.gov/31217225). Molecular systems biology, 2019. [2] [A survey of best practices for RNA-seq data analysis.](https://pubmed.ncbi.nlm.nih.gov/26813401). Genome biology, 2016. [3] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information. [4] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute. [5] [A comprehensive assessment of RNA-seq accuracy, reproducibility and information content by the Sequencing Quality Control Consortium.](https://pubmed.ncbi.nlm.nih.gov/25150838). Nature biotechnology, 2014. [6] [Integration of Bulk RNA-seq Pipeline Metrics for Assessing Low-Quality Samples.](https://pubmed.ncbi.nlm.nih.gov/40630536). Research square, 2025. [7] [nf-core Documentation](https://nf-co.re/docs). nf-core. [8] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project. [9] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries. [10] [Updates and validation of the Compi RNA-seq pipeline with a case study in Alzheimer's disease.](https://doi.org/10.1515/jib-2025-0064). 2026. [11] [Bioconductor](https://bioconductor.org/). Bioconductor Project. [12] [A pipeline for RNA-seq based eQTL analysis with automated quality control procedures.](https://pubmed.ncbi.nlm.nih.gov/34433407). BMC bioinformatics, 2021. [13] [An RNA-Seq Data Analysis Pipeline.](https://pubmed.ncbi.nlm.nih.gov/39068354). Methods in molecular biology (Clifton, N.J.), 2024.

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.