RNA-seq Quality Control: A Step-by-Step Guide to Running FastQC and Interpreting the Results
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- FastQC is a foundational tool for RNA-seq quality control, assessing raw data through modules like Per Base Sequence Quality, GC Content, Overrepresented Sequences, Adapter Content, and Duplication Levels to identify potential issues before downstream analysis.
- Per Base Sequence Quality flags declining Phred scores towards read ends, necessitating trimming of low-quality bases, while Per Sequence GC Content can reveal adapter contamination or organismal contamination through deviations from expected distributions.
- Overrepresented Sequences often indicate adapter dimers or rRNA contamination, requiring trimming or filtering, respectively, whereas Adapter Content directly identifies ligated adapter sequences that must be removed.
- Sequence Duplication Levels are inherently higher in RNA-seq due to highly expressed genes but excessive duplication can signal PCR bias or low input RNA, requiring careful interpretation against biological expectations.
- RNA-seq specific considerations, such as high rRNA content (often >80% of total RNA) and potential GC bias, necessitate careful interpretation of FastQC warnings and may require specialized downstream correction methods beyond basic trimming.
- A robust quality control workflow involves running FastQC on raw data, performing targeted trimming (e.g., with Cutadapt or Trimmomatic) based on identified issues, and re-running FastQC to validate improvements before proceeding to alignment and quantification.
RNA sequencing has become a standard method for exploring gene expression differences between experimental conditions or cell types, and the resulting data can inform further hypotheses about biological processes. While the laboratory protocols required to generate sequencing libraries can be performed in most research facilities, the computational analysis that follows is often an area where researchers have limited experience. Quality control is the first and most consequential step in this computational workflow, because decisions made at this stage determine whether downstream alignment, quantification, and differential expression analyses rest on reliable foundations. This article provides a practical walkthrough of running FastQC on raw RNA-seq data and interpreting its output modules, with specific attention to RNA-seq-specific quality issues and the decision thresholds that justify trimming, filtering, or re-sequencing.
FastQC is a widely used and well-documented tool for assessing sequencing data quality, and it appears as the first step in numerous published RNA-seq workflows. A user-friendly bioinformatics workflow that takes raw RNA-seq data to interpretable results applies FastQC for data quality assessment and Cutadapt for read trimming before alignment with STAR and quantification with featureCounts. Similarly, a transcriptomic dataset from beef heifers used FastQC and MultiQC for quality control before STAR alignment and DESeq2 differential expression analysis. A prostate cancer RNA-seq study applied FastQC for quality controls, Trimmomatic for trimming, and Kallisto for pseudoalignment. These examples illustrate that FastQC is the common entry point across diverse RNA-seq pipelines, regardless of the organism, sequencing platform, or downstream analysis strategy.
The purpose of this guide is to help biology students, researchers, and laboratory professionals run FastQC correctly on raw RNA-seq data and interpret the output metrics to identify quality issues before preprocessing. The scope covers the standard FastQC modules that matter most for RNA-seq, including per-base sequence quality, GC content, overrepresented sequences, adapter contamination, and duplication levels. The guide also addresses RNA-seq-specific considerations such as the expected differences between RNA-seq and DNA-seq quality profiles, the role of library preparation methods, and the practical thresholds that indicate when trimming is necessary or when a sample should be flagged for re-sequencing.
At a Glance
The table below summarizes the key FastQC modules, what each module measures, the common RNA-seq quality issues they reveal, and the practical actions to take when a module flags a warning or failure.
| FastQC Module | What It Measures | Common RNA-seq Issue Detected | Practical Action When Flagged |
|---|---|---|---|
| Per Base Sequence Quality | Phred quality scores across read positions | Declining quality toward read ends, sequencing run deterioration | Trim low-quality bases from read ends, then re-run FastQC to confirm improvement |
| Per Sequence GC Content | GC distribution across all reads | Adapter contamination, PCR bias, contamination from other organisms | Check for overrepresented sequences, verify library preparation, consider contamination screening |
| Overrepresented Sequences | Sequences appearing more often than expected | Adapter dimers, rRNA contamination, highly abundant transcripts | Identify the sequence source, trim adapters, remove rRNA reads, or filter specific sequences |
| Adapter Content | Presence of adapter sequences in reads | Adapter contamination from short inserts or over-sequencing | Trim adapters with Cutadapt or Trimmomatic, then re-run FastQC |
| Sequence Duplication Levels | Degree of duplicate reads in the library | PCR amplification bias, low input RNA, over-amplification | Assess whether duplication is biological (highly expressed genes) or technical, consider deduplication only for certain analyses |
| Per Base N Content | Proportion of undetermined bases at each position | Sequencing errors, poor base calling | Flag sample for review, check sequencing run metrics, consider re-sequencing if N content is high |
Understanding RNA-seq Data and Quality Metrics
RNA-seq generates millions of short sequencing reads that represent fragments of the transcriptome. The quality of these reads directly affects every downstream analysis step, from alignment to differential expression calling. Low-quality reads can introduce spurious alignments, bias expression estimates, and reduce the statistical power to detect genuine biological differences between conditions.
The raw data produced by sequencing platforms are stored in FASTQ format, which contains both the nucleotide sequence and a corresponding quality score for each base. These quality scores, known as Phred scores, are logarithmically related to the probability of an incorrect base call. A Phred score of 20 indicates a 1 in 100 chance of an incorrect base call, while a Phred score of 30 indicates a 1 in 1000 chance. Higher scores represent more reliable base calls.
FastQC systematically evaluates these quality scores and other sequence characteristics across the entire dataset, generating a report that includes multiple modules. Each module is assigned a status of pass, warning, or failure based on built-in thresholds. A warning indicates that something unusual may be present, while a failure indicates that the data are likely to have a significant problem. It is important to understand that these thresholds are general guidelines instead of absolute rules, and the biological context of the experiment should inform the final interpretation.
RNA-seq data have quality profiles that differ from DNA-seq data in several important ways. The transcriptome is dominated by a relatively small number of highly expressed genes, which means that duplication levels are naturally higher than in whole-genome sequencing. The GC content distribution can also be skewed by the underlying transcriptome composition of the organism being studied. These RNA-seq-specific characteristics mean that some FastQC warnings are expected and do not necessarily indicate a problem with the sequencing run.
The National Center for Biotechnology Information provides access to sequence databases and analysis services that support the broader context of RNA-seq data management and quality assessment. Researchers can use these resources to understand the standards for sequence data deposition and retrieval, which are relevant when preparing raw data for public repositories such as the Gene Expression Omnibus or the European Nucleotide Archive.
Preparing Your Environment for FastQC
Before running FastQC, you need a working computational environment with the necessary software installed. FastQC is a Java-based application that runs on Windows, macOS, and Linux systems. It can be downloaded directly from the Babraham Bioinformatics website or installed through package managers such as conda, apt, or Homebrew.
For researchers who prefer not to install software locally, several web-based platforms provide access to FastQC through graphical interfaces. The Galaxy Training Network offers accessible workflow training and analysis tutorials that include FastQC as part of RNA-seq analysis pipelines. These platforms are particularly useful for researchers who are new to command-line computing or who lack access to high-performance computing resources.
The Carpentries provides foundational lessons in computing, data handling, shell, Git, and programming that are valuable for researchers who need to build the skills required for command-line bioinformatics. These lessons cover the basics of navigating the file system, running commands, and managing data files, which are prerequisites for working effectively with FastQC and other bioinformatics tools.
Bioconductor provides official package, workflow, installation, and reproducible genomic-analysis documentation that can help researchers integrate quality control into broader analysis pipelines. While FastQC itself is not a Bioconductor package, the Bioconductor ecosystem includes tools for visualizing and processing quality control results, and its documentation standards are useful for understanding how to structure reproducible analyses.
The European Bioinformatics Institute offers training resources for bioinformatics data resources and practical analysis education. These training materials cover the principles of sequence data quality assessment and provide hands-on exercises that complement the guidance in this article.
For researchers working within structured pipeline frameworks, the nf-core documentation describes community pipeline standards, usage, configuration, and reproducible workflow context. Many nf-core pipelines include FastQC as a mandatory first step, and understanding how these pipelines handle quality control can help researchers adapt their own workflows.
Running FastQC on Raw RNA-seq Data
The basic command to run FastQC on one or more FASTQ files is straightforward. Navigate to the directory containing your raw sequencing files and run the following command:
fastqc sample1.fastq.gz sample2.fastq.gz
For paired-end data, both read files should be included in the same FastQC run so that the quality of both mates can be assessed together:
fastqc sample1_R1.fastq.gz sample1_R2.fastq.gz
When working with large numbers of samples, it is often convenient to run FastQC on all FASTQ files in a directory using a wildcard:
fastqc *.fastq.gz
FastQC generates an HTML report file for each input FASTQ file, along with a zip archive containing the same data in a machine-readable format. The HTML report can be opened in any web browser and provides a visual summary of all quality modules. The zip file is useful for downstream processing, such as aggregating results across multiple samples with MultiQC.
Several command-line options are useful for RNA-seq quality control. The --outdir option specifies the directory where output files should be written, which helps keep results organized. The --threads option specifies the number of threads to use for processing, which can speed up analysis when multiple samples are being processed. The --extract option automatically extracts the zip archive so that the underlying data files are immediately accessible.
For RNA-seq data, it is often helpful to run FastQC both before and after trimming. The pre-trimming run provides a baseline assessment of the raw data quality, while the post-trimming run confirms that the trimming step successfully addressed the identified issues. This before-and-after comparison is a standard practice in published RNA-seq workflows and provides documentation that quality control was performed appropriately.
Interpreting the Per Base Sequence Quality Module
The Per Base Sequence Quality module displays a box-and-whisker plot showing the distribution of quality scores at each position across all reads. The x-axis represents the position within the read, and the y-axis represents the Phred quality score. The plot includes the mean quality score at each position, the interquartile range, and the 10th and 90th percentile values.
For most sequencing platforms, quality scores are highest in the middle of the read and decline toward the 3-prime end. This decline is a normal characteristic of sequencing-by-synthesis chemistry and does not necessarily indicate a problem. The FastQC module flags a warning if the lower quartile of quality scores falls below 28 at any position, and it flags a failure if the lower quartile falls below 20 at any position.
In RNA-seq data, the per-base quality profile can be affected by the library preparation method. Some protocols produce reads with lower quality at the 5-prime end due to the presence of random hexamer priming sites or other technical artifacts. These patterns are often consistent across all samples in an experiment and should be interpreted in the context of the specific library preparation kit used.
When the per-base quality module flags a warning or failure, the practical response depends on the severity and location of the low-quality bases. If quality declines only in the final few bases of the read, trimming those bases with a tool such as Cutadapt or Trimmomatic is usually sufficient. If quality is poor across a substantial portion of the read, more aggressive trimming may be necessary, or the sample may need to be re-sequenced.
A published RNA-seq workflow for endothelial cells used FastQC for data quality assessment and Cutadapt for read trimming, demonstrating that this combination is effective for addressing quality issues identified in the per-base quality module. The workflow proceeded to STAR alignment and featureCounts quantification after trimming, indicating that the quality control step successfully prepared the data for downstream analysis.
Interpreting the Per Sequence GC Content Module
The Per Sequence GC Content module displays the distribution of GC content across all reads in the dataset. The plot shows the observed GC content distribution as a blue line and the theoretical normal distribution as a red line. The theoretical distribution is calculated from the mean GC content of the dataset and represents what would be expected if GC content were randomly distributed across reads.
For most RNA-seq datasets, the observed GC content distribution should approximate a normal distribution centered on the average GC content of the transcriptome. The FastQC module flags a warning if the observed distribution deviates significantly from the theoretical distribution, and it flags a failure if the deviation is substantial.
A shifted GC content distribution can indicate several types of problems. Adapter contamination typically produces a secondary peak in the GC distribution, because adapter sequences have a distinct GC content that differs from the transcriptome. Contamination from other organisms, such as bacterial or fungal sequences, can also shift the GC distribution. PCR bias during library amplification can distort the GC distribution by preferentially amplifying fragments with certain GC contents.
In RNA-seq data, the GC content distribution can be influenced by the organism being studied. Different organisms have different average GC contents in their transcriptomes, and some organisms have bimodal GC distributions due to the presence of distinct gene families with different base compositions. These biological characteristics should be considered when interpreting the GC content module.
A metatranscriptomic reanalysis of Alzheimer's brain samples illustrates the importance of considering contamination when interpreting GC content and other quality metrics. The study screened for non-human transcripts in ribosomal-depleted RNA-seq data and identified low-biomass microbial signals, including enrichment of Acinetobacter radioresistens in the Alzheimer's disease group. The authors emphasized the technical challenges of inferring microbial signals from post-mortem brain RNA-seq data, including contamination risk, low microbial biomass, and overwhelming host background. These findings highlight the need for careful quality control when contamination is a concern.
Interpreting the Overrepresented Sequences Module
The Overrepresented Sequences module identifies sequences that appear in the dataset more frequently than would be expected by chance. FastQC calculates the expected frequency of each sequence based on the total number of reads and the read length, and it flags sequences that exceed this threshold. The module reports the sequence, its length, its count, its percentage of the total reads, and its possible source.
In RNA-seq data, overrepresented sequences can arise from several sources. Adapter dimers are a common artifact of library preparation, where adapter sequences ligate to each other instead of to RNA fragments. Ribosomal RNA contamination occurs when rRNA is not completely depleted during library preparation, resulting in a small number of rRNA sequences dominating the dataset. Highly expressed genes can also produce overrepresented sequences, particularly in organisms with compact transcriptomes where a few genes account for a large proportion of total RNA.
The practical response to overrepresented sequences depends on their source. Adapter dimers should be removed by trimming adapters with Cutadapt or Trimmomatic. Ribosomal RNA contamination can be addressed by removing rRNA reads during the alignment step or by using a more effective rRNA depletion method in the library preparation. Overrepresented sequences from highly expressed genes are generally not a problem and can be left in the dataset, although they may affect the sensitivity of detecting lowly expressed genes.
A study of canine testicular Leydig cell tumors applied quality control using FastQC and Trimmomatic before differential expression analysis. The workflow identified 1500 transcripts, including 982 upregulated and 168 downregulated genes, and the quality control step was essential for ensuring that these expression differences were not artifacts of poor-quality data. This example demonstrates the importance of addressing overrepresented sequences before proceeding with biological interpretation.
Interpreting the Adapter Content Module
The Adapter Content module detects the presence of adapter sequences within the reads. Adapters are short oligonucleotide sequences that are ligated to RNA fragments during library preparation to enable sequencing. When the insert size is shorter than the read length, the sequencing reaction reads through the insert and into the adapter sequence, producing reads that contain adapter contamination at the 3-prime end.
The Adapter Content module plots the cumulative percentage of reads containing each adapter sequence at each position. The module flags a warning if any adapter sequence is present in more than 5 percent of reads, and it flags a failure if any adapter sequence is present in more than 10 percent of reads.
Adapter contamination is a common issue in RNA-seq data, particularly when the RNA fragments are short or when the library is over-sequenced. The presence of adapter sequences in reads can cause alignment problems, because the adapter sequence does not match the reference genome or transcriptome. Adapter trimming is therefore an essential preprocessing step for RNA-seq data.
Cutadapt is a widely used tool for adapter trimming, and it was applied in the endothelial cell RNA-seq workflow alongside FastQC. Trimmomatic is another popular option that was used in the prostate cancer RNA-seq study and the canine Leydig cell tumor study. Both tools can remove adapter sequences from reads and can also perform quality trimming in the same step.
After trimming adapters, it is important to re-run FastQC to confirm that the adapter content module now passes. This verification step ensures that the trimming was effective and that no adapter sequences remain in the data.
Interpreting the Sequence Duplication Levels Module
The Sequence Duplication Levels module displays the distribution of duplication levels across all sequences in the dataset. The plot shows the percentage of sequences with each duplication level, from unique sequences that appear only once to sequences that appear many times. The module also reports the total number of sequences and the percentage of sequences that are duplicates.
In RNA-seq data, duplication levels are naturally higher than in DNA-seq data because the transcriptome is dominated by highly expressed genes. A small number of genes can account for a large proportion of the total RNA, and the corresponding reads will appear many times in the dataset. This biological duplication is expected and should not be interpreted as a technical artifact.
However, high duplication levels can also indicate technical problems. PCR amplification bias during library preparation can artificially inflate the number of duplicate reads, particularly when the input RNA amount is low or when too many PCR cycles are used. Over-sequencing can also produce excessive duplication, because the same fragments are sequenced multiple times.
The FastQC module flags a warning if more than 20 percent of sequences are duplicated, and it flags a failure if more than 50 percent of sequences are duplicated. For RNA-seq data, these thresholds should be interpreted with caution, because biological duplication from highly expressed genes can trigger warnings even in high-quality datasets.
A study that integrated bulk RNA-seq pipeline metrics for assessing low-quality samples found that no individual quality control metric is sufficient on its own to identify low-quality samples. The study developed the Quality Control Diagnostic Renderer (QC-DR), software designed to simultaneously visualize a comprehensive panel of quality control metrics and flag samples with aberrant values compared to a reference dataset. Among the most highly correlated pipeline quality control metrics were the percentage and number of uniquely aligned reads, the percentage of rRNA reads, the number of detected genes, and a newly developed metric of Area Under the Gene Body Coverage Curve. The study concluded that approaches based on the integration of multiple metrics with quality control thresholds are more reliable than any single metric.
RNA-seq-Specific Quality Considerations
RNA-seq data have several characteristics that require special attention during quality control. These characteristics arise from the biology of the transcriptome and the technical aspects of RNA-seq library preparation, and they can affect the interpretation of FastQC modules.
Ribosomal RNA Contamination
Ribosomal RNA comprises the majority of total RNA in most cells, typically accounting for more than 80 percent of the RNA content. Library preparation protocols use poly-A selection or ribosomal RNA depletion to remove rRNA before sequencing, but these methods are not always completely effective. Residual rRNA contamination can consume a significant proportion of sequencing reads, reducing the effective depth for messenger RNA and other RNA species.
The presence of rRNA contamination is often visible in the overrepresented sequences module, where a small number of rRNA sequences appear at high frequency. The percentage of rRNA reads is also a useful quality control metric, as identified in the study of bulk RNA-seq pipeline metrics. High rRNA content can be addressed by more aggressive rRNA depletion during library preparation or by filtering rRNA reads during alignment.
GC Bias
GC bias refers to the tendency of sequencing and library preparation methods to preferentially amplify or sequence fragments with certain GC contents. This bias can distort the relationship between the number of reads mapped to a gene and the actual expression level of that gene. GC bias is particularly problematic for differential expression analysis, because it can create false differences between samples with different overall GC compositions.
The per-sequence GC content module in FastQC can reveal GC bias, particularly when the observed distribution deviates from the expected normal distribution. However, GC bias is often subtle and may not be visible in the FastQC report. More sophisticated methods for detecting and correcting GC bias are available in downstream analysis tools, such as the GC correction functions in DESeq2 and edgeR.
Transcript Length Bias
RNA-seq reads are typically generated from fragments of transcripts, and the number of reads mapping to a gene depends on both the expression level and the length of the transcript. Longer transcripts produce more fragments and therefore more reads than shorter transcripts at the same expression level. This transcript length bias is a well-known characteristic of RNA-seq data and is addressed by normalization methods that account for transcript length, such as transcripts per million (TPM).
FastQC does not directly assess transcript length bias, but the duplication levels module can provide indirect information. Highly expressed long transcripts produce many duplicate reads, which can inflate the duplication levels metric. Understanding this relationship helps interpret duplication warnings in the context of the transcriptome composition.
Strand-Specificity
RNA-seq library preparation protocols can be strand-specific or non-strand-specific. Strand-specific protocols preserve the information about which strand of the DNA the RNA was transcribed from, while non-strand-specific protocols lose this information. The choice of protocol affects the interpretation of alignment results and the quantification of gene expression.
FastQC does not directly assess strand-specificity, but the overrepresented sequences module can provide clues. In non-strand-specific libraries, reads from both strands of highly expressed genes appear in the dataset, while in strand-specific libraries, reads appear predominantly from one strand. This difference can affect the interpretation of overrepresented sequences and the choice of downstream analysis tools.
Practical Workflow for RNA-seq Quality Control
The following workflow provides a step-by-step approach to running FastQC and interpreting the results for RNA-seq data. This workflow is designed to be practical and reproducible, and it can be adapted to different computational environments and analysis pipelines.
Step 1: Organize Your Data
Before running FastQC, organize your raw sequencing data in a consistent directory structure. Create a directory for each sample or each experiment, and ensure that all FASTQ files are named consistently. A common convention is to include the sample identifier, the read pair (R1 or R2 for paired-end data), and the file extension in the filename.
Document the metadata for each sample, including the sample identifier, the biological condition, the library preparation protocol, the sequencing platform, and the sequencing date. This metadata is essential for interpreting quality control results and for reproducing the analysis at a later time.
Step 2: Run FastQC on All Raw Samples
Run FastQC on all raw FASTQ files in the experiment. Use the --outdir option to specify an output directory for the reports, and use the --threads option to speed up processing when multiple samples are being analyzed.
For large experiments with many samples, consider using MultiQC to aggregate the FastQC reports into a single summary report. MultiQC scans the FastQC output files and generates a combined report that allows quick comparison of quality metrics across all samples. This comparison is useful for identifying samples that deviate from the overall pattern.
Step 3: Review the FastQC Reports
Open the FastQC HTML reports and review each module systematically. Start with the per-base sequence quality module to assess the overall quality of the sequencing run. Then review the per-sequence GC content module to check for contamination or bias. Next, review the overrepresented sequences and adapter content modules to identify adapter contamination or rRNA contamination. Finally, review the sequence duplication levels module to assess the complexity of the library.
For each module, note whether the status is pass, warning, or failure, and record any observations in a quality control log. This log provides documentation of the quality control process and supports decisions about trimming, filtering, or re-sequencing.
Step 4: Decide on Trimming and Filtering
Based on the FastQC results, decide whether trimming or filtering is necessary. If the per-base sequence quality declines at the read ends, trim low-quality bases with Cutadapt or Trimmomatic. If adapter contamination is present, trim adapter sequences with the same tools. If rRNA contamination is substantial, consider filtering rRNA reads during alignment.
Record the trimming parameters used, including the quality threshold, the adapter sequence, and the minimum read length after trimming. These parameters should be consistent across all samples in the experiment to avoid introducing bias.
Step 5: Re-run FastQC After Trimming
After trimming and filtering, re-run FastQC on the processed FASTQ files. Compare the post-trimming reports to the pre-trimming reports to confirm that the identified issues were resolved. The per-base sequence quality should improve at the read ends, the adapter content should be reduced, and the overrepresented sequences should be less prominent.
Document the before-and-after comparison in the quality control log. This documentation is valuable for publications, for data repository submissions, and for troubleshooting if downstream analyses reveal unexpected results.
Step 6: Proceed to Downstream Analysis
After confirming that the quality control is satisfactory, proceed to the downstream analysis steps of the RNA-seq workflow. These steps typically include alignment to a reference genome or transcriptome, quantification of gene expression, and differential expression analysis.
The choice of alignment and quantification tools depends on the specific research question and the available computational resources. STAR is a widely used aligner that was applied in the endothelial cell workflow, the beef heifer dataset, and the canine Leydig cell tumor study. Kallisto is a pseudoaligner that was used in the prostate cancer study. Salmon is another popular option that was used in the end-to-end reproducible RNA-seq workflow.
Common Failure Patterns and Troubleshooting
Several quality issues recur frequently in RNA-seq data, and recognizing these patterns helps researchers diagnose problems quickly and take appropriate action.
Pattern 1: Quality Drops at Read Ends
A common pattern is high quality in the middle of the read with declining quality toward the 3-prime end. This pattern is normal for sequencing-by-synthesis platforms and becomes more pronounced with longer read lengths. The practical response is to trim low-quality bases from the read ends, using a quality threshold that balances the need to remove errors with the need to retain useful sequence.
Pattern 2: Adapter Contamination in Short Inserts
When RNA fragments are shorter than the read length, the sequencing reaction reads into the adapter sequence. This produces reads with adapter sequences at the 3-prime end, which are detected by the adapter content module. The practical response is to trim adapter sequences, which also removes the low-quality bases that often accompany adapter contamination.
Pattern 3: rRNA Contamination
Incomplete rRNA depletion during library preparation produces datasets with a high proportion of rRNA reads. These reads appear as overrepresented sequences in the FastQC report and consume sequencing depth that could be used for messenger RNA. The practical response is to filter rRNA reads during alignment or to improve the rRNA depletion method in the library preparation protocol.
Pattern 4: PCR Duplicates
Excessive PCR amplification during library preparation produces high duplication levels in the FastQC report. This pattern is more common with low input RNA amounts or excessive PCR cycles. The practical response is to reduce the number of PCR cycles during library preparation or to use a library preparation kit with lower amplification bias.
Pattern 5: Contamination from Other Organisms
Contamination from bacteria, fungi, or other organisms can appear in RNA-seq data, particularly for samples collected from environmental or clinical sources. This contamination can shift the GC content distribution and produce overrepresented sequences from the contaminating organism. The practical response is to screen the reads against reference databases to identify the contaminating organism and to filter the contaminating reads if necessary.
A study of nasopharyngeal carcinoma patients in Indonesia analyzed raw sequence data quality using FastQC and interpreted the results using HISAT2, HTSeq, edgeR, and PANTHER. The study identified 25493 gene transcripts, with 1956 genes significantly upregulated and 90 genes significantly downregulated in nasopharyngeal carcinoma samples compared to normal individuals. The quality control step was essential for ensuring that these expression differences were reliable, particularly given the challenges of working with clinical biopsy samples.
Records and Documentation for Quality Control
Maintaining detailed records of the quality control process is essential for reproducible research and for troubleshooting downstream analysis problems. The following records should be maintained for each RNA-seq experiment.
Quality Control Log
The quality control log records the FastQC status for each module for each sample, both before and after trimming. This log can be maintained as a spreadsheet or as a text file, and it should include the sample identifier, the module name, the status (pass, warning, or failure), and any relevant observations.
Trimming Parameters
The trimming parameters used for each sample should be recorded, including the tool version, the quality threshold, the adapter sequence, and the minimum read length after trimming. These parameters should be consistent across all samples in the experiment, and any deviations should be documented.
MultiQC Summary Report
For experiments with many samples, the MultiQC summary report provides a comprehensive overview of quality metrics across all samples. This report should be saved and included in the project documentation.
Pipeline Configuration
If the analysis is performed using a pipeline framework such as nf-core or Snakemake, the pipeline configuration files should be saved and version-controlled. These files document the exact parameters used for each analysis step and support reproducibility.
The nf-core documentation describes community pipeline standards, usage, configuration, and reproducible workflow context. Following these standards helps ensure that quality control and downstream analysis steps are performed consistently and reproducibly.
Limitations of FastQC and Quality Control Metrics
FastQC is a powerful tool for assessing sequencing data quality, but it has several limitations that researchers should understand.
FastQC Thresholds Are General Guidelines
The warning and failure thresholds in FastQC are based on general sequencing data characteristics and may not be appropriate for all RNA-seq datasets. RNA-seq data naturally have higher duplication levels and different GC content distributions than DNA-seq data, and these differences can trigger warnings that do not indicate actual problems. Researchers should interpret FastQC results in the context of their specific experiment and organism.
Individual Metrics Have Limited Predictive Value
A study that integrated bulk RNA-seq pipeline metrics for assessing low-quality samples found that any individual quality control metric is limited in its predictive value. The study recommended approaches based on the integration of multiple metrics with quality control thresholds. This finding emphasizes the importance of reviewing all FastQC modules together instead of focusing on any single metric.
FastQC Does Not Assess All Quality Issues
FastQC assesses the quality of the raw sequencing reads, but it does not assess all aspects of data quality that affect downstream analysis. For example, FastQC does not assess alignment rates, gene body coverage, or the presence of batch effects. These issues are assessed by other tools, such as Qualimap for alignment quality and principal component analysis for batch effects.
Quality Control Is Not a Substitute for Experimental Design
Quality control can identify problems with sequencing data, but it cannot fix problems with experimental design. Poor experimental design, such as inadequate biological replication or confounding of conditions with batch effects, cannot be resolved by quality control or any downstream analysis method. Researchers should invest in careful experimental design before generating sequencing data.
Professional Escalation Criteria
Some quality issues require escalation to a bioinformatics specialist, a sequencing facility, or a statistical consultant. The following criteria indicate when professional escalation is appropriate.
Consistent Failures Across Multiple Modules
If a sample fails multiple FastQC modules, this indicates a systemic problem with the library preparation or sequencing run. Escalate to the sequencing facility to investigate the cause and to determine whether re-sequencing is necessary.
Unexpected Patterns in Multiple Samples
If multiple samples show unexpected quality patterns, such as a shifted GC content distribution or high contamination levels, this may indicate a problem with the library preparation protocol or the sequencing run. Escalate to the sequencing facility or to a bioinformatics specialist to investigate the cause.
Discrepancies Between Quality Metrics and Experimental Expectations
If quality metrics are inconsistent with experimental expectations, such as high duplication levels in samples with high input RNA amounts, this may indicate a technical problem. Escalate to a bioinformatics specialist to investigate the cause and to determine the appropriate response.
Issues That Affect Downstream Analysis
If quality issues persist after trimming and filtering, or if downstream analysis results are unexpected, escalate to a bioinformatics specialist. The specialist can perform more detailed quality assessments, such as alignment quality analysis with Qualimap or gene body coverage analysis, to identify the underlying cause.
Frequently Asked Questions
What is the difference between a warning and a failure in FastQC?
A warning in FastQC indicates that a module has detected something unusual that may or may not indicate a problem with the data. A failure indicates that the module has detected a significant issue that is likely to affect downstream analysis. For RNA-seq data, some warnings are expected due to the biological characteristics of the transcriptome, such as high duplication levels from highly expressed genes. Failures should be investigated and addressed before proceeding with downstream analysis.
How many samples should I run through FastQC before deciding on trimming parameters?
FastQC should be run on all samples in the experiment, both before and after trimming. The pre-trimming run provides a baseline assessment of the raw data quality, and the post-trimming run confirms that the trimming was effective. Trimming parameters should be determined based on the quality profile of the worst samples, because all samples should be processed consistently to avoid introducing bias.
Can I use FastQC on single-cell RNA-seq data?
FastQC can be used on single-cell RNA-seq data, but the interpretation of the results differs from bulk RNA-seq data. Single-cell RNA-seq data have different quality characteristics, including higher dropout rates and different duplication patterns. The study of single-cell RNA-seq data from wild-type and fli1b mutant zebrafish embryos used the 10X Genomics Chromium platform and identified 34 distinct cell clusters. Quality control for single-cell data typically involves additional steps beyond FastQC, such as filtering cells based on the number of detected genes and the percentage of mitochondrial reads.
What should I do if my RNA-seq data shows high duplication levels?
High duplication levels in RNA-seq data can be caused by biological factors, such as highly expressed genes, or by technical factors, such as PCR amplification bias. Review the overrepresented sequences module to determine whether the duplicates come from a small number of highly expressed genes or from a broad distribution of sequences. If the duplicates are from highly expressed genes, the high duplication levels are expected and do not require action. If the duplicates are from PCR bias, consider reducing the number of PCR cycles during library preparation or using a library preparation kit with lower amplification bias.
How do I know if adapter contamination is affecting my RNA-seq data?
The adapter content module in FastQC detects the presence of adapter sequences within the reads. If the module flags a warning or failure, adapter contamination is present and should be addressed by trimming. Adapter contamination is more common in RNA-seq data with short inserts or over-sequencing. After trimming, re-run FastQC to confirm that the adapter content module now passes.
What is the role of MultiQC in RNA-seq quality control?
MultiQC aggregates the FastQC reports from multiple samples into a single summary report. This aggregation allows quick comparison of quality metrics across all samples and helps identify samples that deviate from the overall pattern. MultiQC is particularly useful for experiments with many samples, where reviewing individual FastQC reports would be time-consuming. The beef heifer transcriptomic dataset used FastQC and MultiQC for quality control before STAR alignment and DESeq2 differential expression analysis.
Should I trim my RNA-seq reads before alignment?
Trimming is recommended when the FastQC report identifies quality issues that can be addressed by trimming, such as low-quality bases at read ends or adapter contamination. Trimming is not always necessary, and over-trimming can reduce the amount of usable sequence. The decision to trim should be based on the FastQC results and should be applied consistently across all samples in the experiment. Published RNA-seq workflows commonly use Cutadapt or Trimmomatic for trimming after FastQC quality assessment.
How does the choice of genome build affect RNA-seq quality control and interpretation?
The choice of genome build can affect RNA-seq analysis at multiple stages, including alignment and quantification. A study of the effect of genome build on expression quantification and outlier detection found that 61 percent of quantified genes were not influenced by genome build, but 1492 genes had build-dependent quantification, 3377 genes had build-exclusive expression, and 9077 genes had annotation-specific expression across six routinely collected biospecimens. The study recommended that transcriptomics-guided analyses and diagnoses be cross-referenced with data on genes impacted by build choice for robustness. While FastQC itself is not affected by genome build, the downstream alignment and quantification steps are, and the choice of build should be documented in the analysis records.
Related Bioinformatics Guides
- Single-Cell RNA Sequencing Quality Control: A Practical Guide to Filtering and Metrics
- RNA-Seq Data Analysis Workflow: From Raw Reads to Insights
- RNA-Seq Quality Control: Essential Checks and Tools
- RNA-Seq Databases: Accessing and Using Public RNA-Seq Data
- Spatial Transcriptomics Data Analysis: A Guide to Preprocessing, Integration, and Interpretation
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- Endothelial Cell RNA-Seq Data: Differential Expression and Functional Enrichment Analyses to Study Phenotypic Switching.. Methods in molecular biology (Clifton, N.J.), 2022.
- Transcriptomic dataset from peripheral white blood cells of beef heifers at weaning.. Data in brief, 2023.
- Circular RNA Identification and Characterization with CircRNAFlow: A Bioinformatics Approach.. Advances in experimental medicine and biology, 2025.
- Step-by-Step Construction of Gene Co-expression Networks from High-Throughput Arabidopsis RNA Sequencing Data.. Methods in molecular biology (Clifton, N.J.), 2018.
- RNA sequencing data of human prostate cancer cells treated with androgens.. Data in brief, 2019.
- Transcriptome Profile of Next-Generation Sequencing Data Relate to Proliferation Aberration of Nasopharyngeal Carcinoma Patients in Indonesia.. Asian Pacific journal of cancer prevention : APJCP, 2020.
- Transcriptomic Profiling of Canine Testicular Leydig Cell Tumors Uncovers Key Upregulated Gene Pathways.. Animals : an open access journal from MDPI, 2026.
- An End-to-End Reproducible RNA-Seq Workflow from Raw Sequencing Reads to Differential Expression, Pathway Enrichment, and Biological Interpretation. 2026.
- RAGER: A user-friendly computational platform for integrated analysis of RNA-Seq and ATAC-seq data.. 2026.
- Single-cell RNA-seq data of wild type and fli1b mutant zebrafish embryos.. 2026.
- Metatranscriptomic Reanalysis of Alzheimer's Brains Identifies Low-Biomass Microbial Signals Including Enrichment of <,i>,Acinetobacter radioresistens<,/i>,.. 2026.
- Integration of bulk RNA-seq pipeline metrics for assessing low-quality samples.. 2026.
- Combining optical genome mapping and RNA-seq for structural variants detection and interpretation in unsolved neurodevelopmental disorders. Genome Medicine, 2024.
- RNA-seq data science: From raw data to effective interpretation. Frontiers in Genetics, 2023.
- Impact of genome build on RNA-seq interpretation and diagnostics.. American Journal of Human Genetics, 2024.
- Normalization choice drives biological interpretation in single-cell RNA-seq cancer studies: A systematic benchmarking of 465 computational pipelines. Comput. Biol. Chem., 2026.
- Integrating single-cell RNA-seq datasets with substantial batch effects. BMC Genomics, 2025.
- Generation-Specific Heterosis in Lactation, Reproduction, and Blood Transcriptomic Profiles of Chinese Simmental × Holstein Crossbred Cows. Animals, 2026.
- Subtype-resolved transcriptomic analysis reveals distinct zinc-finger regulatory hubs in breast cancer. Computational Biology and Chemistry, 2026.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.