# RNA-seq Quality Control: Key Metrics and Thresholds for Library and Sequencing Data


## Key Takeaways

- RNA Integrity Number (RIN) ≥ 7.0 is a critical pre-library preparation metric, with lower values indicating degradation that can lead to 3-prime bias and reduced transcript complexity; DV200 is a valuable alternative for degraded samples like FFPE tissues.
- Per-base Phred scores > 30 across most read positions in raw sequencing data are essential, as significant declines, particularly early in reads, signal sequencing run failures or instrument issues.
- Mapping rates > 70% for eukaryotic genomes are expected, with lower rates potentially indicating sample contamination (e.g., bacterial, rRNA) or reference genome mismatches, necessitating investigation of unmapped reads.
- Duplication rates < 30-40% in bulk RNA-seq libraries are generally acceptable; higher rates often point to over-amplification during library preparation or low input RNA, reducing effective sequencing depth.
- Gene body coverage should be relatively uniform, as significant 3-prime bias suggests RNA degradation or specific library preparation artifacts, complementing RIN assessment.
- For single-cell RNA-seq, cell-level metrics like number of genes detected per cell and mitochondrial read proportion are crucial for identifying and excluding low-quality cells or technical artifacts before downstream analysis.

---

RNA sequencing quality control is the systematic assessment of RNA integrity, library preparation success, sequencing output, and alignment performance before any biological interpretation begins. For researchers and laboratory professionals, the practical problem is knowing which metrics matter, what values indicate acceptable performance, and when to stop and troubleshoot instead of proceed to downstream analysis. This article defines the critical quality control metrics for RNA-seq libraries and raw sequencing data, provides recommended thresholds for common experimental designs, and outlines troubleshooting steps for the most frequent failure patterns. The scope covers bulk RNA-seq primarily, with notes on single-cell RNA-seq where workflows diverge. Quality control decisions made before differential expression analysis determine whether downstream results can be trusted, so understanding these metrics and their limitations is a prerequisite for reliable transcriptome research.

## The Role of Quality Control in the RNA-seq Workflow

Quality control in RNA-seq is not a single step but a continuous process that begins with sample collection and extends through alignment and quantification. The RNA-seq analysis pipeline typically proceeds from raw sequencing reads through quality assessment, preprocessing, alignment to a reference genome, and quantification of transcript abundance before differential expression analysis can be performed. Each stage generates metrics that inform whether the next stage will produce reliable results. The importance of ensuring data integrity at the outset is emphasized throughout the literature, as poor-quality input data propagates errors through every subsequent analysis step.

The core principle is that quality control decisions should be made using objective metrics instead of intuition. RNA-seq data analysis requires a thorough understanding of the entire pipeline, from quality control through preprocessing, alignment, differential expression, and functional analysis. Researchers who skip or rush quality control often discover problems only after investing substantial time in downstream analysis, at which point the cost of correction is much higher.

Quality control serves three distinct purposes in the RNA-seq workflow. First, it identifies failed or compromised samples that should be excluded or re-sequenced. Second, it characterizes technical variation that can be accounted for in statistical models. Third, it provides a record of data quality that supports reproducibility and interpretation of results. Each purpose requires different metrics and different thresholds.

## At a Glance: Key RNA-seq Quality Control Metrics and Thresholds

The table below summarizes the primary quality control metrics that should be assessed for RNA-seq libraries and raw sequencing data. Thresholds are provided as general guidance for standard bulk RNA-seq experiments using Illumina short-read technology. Values outside these ranges warrant investigation and may indicate problems requiring corrective action.

| Metric | Assessment Stage | Typical Acceptable Range | Common Failure Indication |
| --- | --- | --- | --- |
| RNA Integrity Number (RIN) | Before library preparation | 7.0 or higher for most applications | Degraded RNA, failed extraction, improper storage |
| Library concentration | After library preparation | Depends on platform and protocol, typically within expected range for the kit used | Failed amplification, excessive adapter dimer, quantification error |
| Fragment size distribution | After library preparation | Tight peak at expected insert size, no significant adapter dimer peak | Over-fragmentation, under-fragmentation, adapter contamination |
| Per-base sequence quality | Raw sequencing data | Mean Phred score above 30 for most positions | Sequencing run failure, instrument issues, reagent problems |
| GC content distribution | Raw sequencing data | Consistent with expected transcriptome composition | Contamination, PCR bias, library preparation artifacts |
| Adapter contamination | Raw sequencing data | Minimal adapter sequences in reads | Over-amplification, insufficient adapter removal, short inserts |
| Mapping rate | After alignment | Above 70% for standard eukaryotic genomes | Sample contamination, reference mismatch, rRNA contamination |
| Duplication rate | After quantification | Below 30-40% for most RNA-seq libraries | Over-amplification, low input RNA, sequencing depth saturation |
| Gene body coverage | After alignment | Relatively uniform coverage across gene length | RNA degradation, 3-prime bias, library preparation issues |
| rRNA contamination | After alignment | Below 5-10% of mapped reads | Incomplete rRNA depletion, failed poly-A selection |

These thresholds represent general guidance instead of universal standards. The acceptable range for each metric depends on the organism, tissue type, library preparation method, sequencing platform, and experimental design. Researchers should establish project-specific thresholds based on their protocols and prior experience with similar samples.

## RNA Quality Assessment Before Library Preparation

The quality of input RNA is the single most important determinant of RNA-seq data quality. Degraded RNA produces libraries with 3-prime bias, reduced complexity, and unreliable quantification of transcript abundance. Assessment of RNA quality should occur before library preparation, and samples failing quality thresholds should be re-extracted or excluded.

### RNA Integrity Number and Alternative Metrics

The RNA Integrity Number (RIN) is the most widely used metric for assessing RNA quality. RIN values range from 1 to 10, with 10 representing completely intact RNA. For most bulk RNA-seq applications, a RIN of 7.0 or higher is considered acceptable. However, the appropriate threshold depends on the research question and the tissue type. Some tissues, such as those obtained from clinical biopsies or post-mortem samples, may yield lower RIN values that still produce usable data.

RIN is calculated from the electrophoretic trace of RNA, assessing the ratio of ribosomal RNA peaks and the presence of degradation products. The metric assumes that ribosomal RNA integrity reflects the integrity of the entire transcriptome. This assumption holds for most samples but can fail in specific circumstances, such as when samples have been subjected to selective degradation or when ribosomal RNA has been partially removed.

Alternative metrics include the DV200 value, which measures the percentage of RNA fragments longer than 200 nucleotides. DV200 is particularly useful for degraded samples, such as those from formalin-fixed paraffin-embedded tissues, where RIN values are consistently low. For such samples, DV200 provides a more informative assessment of whether sufficient intact RNA remains for library preparation.

### Practical RNA Quality Assessment Protocol

The following steps outline a practical approach to RNA quality assessment before library preparation:

1. Quantify RNA concentration using a fluorometric method such as Qubit or a spectrophotometric method such as NanoDrop. Fluorometric methods are preferred because they specifically detect RNA instead of contaminating DNA or proteins.
2. Assess RNA integrity using a microfluidic electrophoresis system such as the Agilent Bioanalyzer or TapeStation. Record the RIN value and examine the electrophoretic trace for signs of degradation.
3. Check for genomic DNA contamination using a method that distinguishes DNA from RNA, such as PCR amplification without reverse transcription or analysis of the electrophoretic trace.
4. Record all quality metrics in a sample tracking sheet that accompanies the samples through library preparation and sequencing.
5. For samples with RIN values below the acceptable threshold, consider whether re-extraction is possible or whether the sample can be processed with a protocol optimized for degraded RNA.

The decision to proceed with a low-quality sample should be documented and justified. Some downstream analyses, such as detection of highly expressed genes or analysis of transcript isoforms, may tolerate moderate degradation. Other analyses, such as detection of low-abundance transcripts or precise quantification of isoform usage, require high-quality RNA.

## Library Preparation Quality Control

Library preparation converts RNA into a sequencing-ready DNA library through reverse transcription, adapter ligation, and amplification. Each step can introduce biases or artifacts that affect data quality. Quality control at this stage focuses on library concentration, fragment size distribution, and the presence of contaminants.

### Library Concentration and Quantification

Library concentration must be accurately measured to determine the appropriate amount of library to load onto the sequencer. Under-loading reduces sequencing output, while over-loading causes cluster overcrowding and reduced data quality. Quantitative PCR is the most accurate method for quantifying libraries that will be sequenced on Illumina platforms, as it measures the concentration of adapter-ligated molecules that can form clusters.

The expected library concentration depends on the library preparation kit and the quantification method used. Most kits provide expected concentration ranges in their protocols. Libraries with concentrations far outside the expected range may indicate problems with amplification, purification, or quantification.

Library concentration should be measured after the final purification step and again after any dilution steps. Records of concentration at each stage help identify where losses occur and support troubleshooting if sequencing fails.

### Fragment Size Distribution

The fragment size distribution of a library reflects the insert size and the success of adapter ligation. A typical RNA-seq library should show a tight distribution of fragment sizes centered on the expected insert size, which depends on the fragmentation method and the library preparation kit. For Illumina sequencing, fragment sizes typically range from 200 to 500 base pairs, with the exact range depending on the kit.

The fragment size distribution is assessed using a microfluidic electrophoresis system. The trace should show a single dominant peak with minimal signal at very small sizes, which would indicate adapter dimers, and minimal signal at very large sizes, which would indicate incomplete fragmentation or concatenated fragments.

Adapter dimers are a common library preparation artifact that occurs when adapters ligate to each other instead of to cDNA fragments. Adapter dimers sequence efficiently and consume sequencing capacity without producing useful data. If adapter dimers constitute a significant fraction of the library, the library should be re-purified or re-prepared.

### Library Quality Metrics and Their Interpretation

The following metrics should be recorded for each library:

1. Final library concentration in nanomolar or picomolar units
2. Mean and median fragment size in base pairs
3. Presence and proportion of adapter dimers
4. Presence of high-molecular-weight contaminants
5. Yield relative to input RNA amount

These metrics provide a baseline for troubleshooting. If a library produces poor sequencing data, comparison of these metrics with libraries that performed well can identify the source of the problem.

## Raw Sequencing Data Quality Assessment

Once libraries are sequenced, the raw data must be assessed for quality before any alignment or quantification is performed. This assessment examines per-base quality scores, base composition, GC content, adapter contamination, and sequence duplication levels.

### Per-Base Quality Scores

Per-base quality scores, expressed as Phred scores, indicate the probability of an incorrect base call at each position in the read. A Phred score of 30 corresponds to an error probability of 1 in 1000, and a Phred score of 20 corresponds to an error probability of 1 in 100. Quality scores typically decline toward the end of reads, and the rate of decline depends on the sequencing platform and run conditions.

Quality assessment tools generate per-base quality plots that show the distribution of quality scores across read positions. Acceptable data typically shows mean quality scores above 30 for most positions, with some decline at the 3-prime end of reads. Severe quality decline, particularly in early read positions, indicates problems with the sequencing run, reagents, or instrument.

Per-base quality assessment should be performed separately for read 1 and read 2 in paired-end sequencing, as the two reads often show different quality profiles. Read 2 frequently has lower quality than read 1 because of the chemistry of paired-end sequencing.

### Base Composition and GC Content

The base composition of sequencing reads should reflect the composition of the transcriptome being sequenced. For most eukaryotic transcriptomes, GC content ranges from 40 to 60 percent, with variation depending on the organism and the genes expressed in the sample. Per-base base composition plots should show relatively stable proportions of A, C, G, and T across read positions, with some expected variation at the first few positions due to random hexamer priming.

Deviations from expected base composition can indicate several problems. A consistent bias toward one base across all positions may indicate contamination or a sequencing artifact. Position-specific biases at the start of reads are common and usually result from random hexamer priming bias during reverse transcription. Severe GC bias, where the GC content of sequenced fragments differs substantially from the expected transcriptome composition, can result from PCR amplification bias during library preparation.

### Adapter Contamination

Adapter sequences appear in reads when the insert fragment is shorter than the read length. When this occurs, the sequencer continues reading into the adapter sequence at the 3-prime end of the read. Adapter contamination is more common in libraries with short inserts and in sequencing runs with longer read lengths.

Quality assessment tools detect adapter contamination by searching for known adapter sequences in the reads. The proportion of reads containing adapter sequence and the position at which adapter sequence begins should be recorded. If adapter contamination is significant, reads should be trimmed before alignment.

Adapter trimming is a preprocessing step that removes adapter sequences from reads. The trimming decision involves a tradeoff between removing adapter sequence and retaining read length. Most trimming tools identify the adapter sequence and trim only the portion of the read that matches the adapter, preserving the insert-derived sequence.

### Sequence Duplication Levels

Sequence duplication occurs when multiple reads originate from the same cDNA fragment. Some duplication is expected in RNA-seq data because highly expressed genes produce many copies of the same transcript. However, excessive duplication indicates problems such as over-amplification during library preparation or insufficient input RNA.

Duplication rates are calculated after alignment, as reads from different genes can have identical sequences by chance. The duplication rate should be interpreted in the context of sequencing depth and expression levels. Highly expressed genes will naturally show high duplication rates, while the overall duplication rate across all genes should remain moderate.

For most bulk RNA-seq libraries, an overall duplication rate below 30 to 40 percent is considered acceptable. Higher duplication rates reduce the effective sequencing depth and may indicate that the library was over-amplified. In such cases, the library may need to be re-prepared with fewer PCR cycles or with more input RNA.

## Alignment and Quantification Quality Metrics

After preprocessing, reads are aligned to a reference genome or transcriptome. Alignment quality metrics provide information about the success of the alignment and the composition of the sequenced library.

### Mapping Rate

The mapping rate is the proportion of reads that align to the reference genome or transcriptome. For standard eukaryotic samples, mapping rates above 70 percent are generally expected. Lower mapping rates may indicate sample contamination, reference genome issues, or the presence of sequences that do not originate from the expected organism.

Several factors can reduce mapping rates. Ribosomal RNA contamination reduces mapping rates if the reference does not include ribosomal RNA sequences or if the alignment tool does not handle multi-mapping reads well. Contamination with DNA from other organisms, such as bacteria or mycoplasma, produces reads that do not map to the reference genome. Sequence variants in the sample that differ from the reference genome can also reduce mapping rates, particularly for highly polymorphic organisms.

The interpretation of mapping rate depends on the alignment strategy. Spliced aligners that map reads across exon-exon junctions may achieve lower mapping rates than unspliced aligners because junction reads are more difficult to map. The choice of reference genome version and annotation also affects mapping rates.

### rRNA Contamination

Ribosomal RNA constitutes the majority of total RNA in most cells, typically 80 to 90 percent. Library preparation methods must remove ribosomal RNA through poly-A selection or ribosomal RNA depletion. Incomplete removal results in a large proportion of reads mapping to ribosomal RNA genes, which reduces the effective sequencing depth for messenger RNA.

The proportion of reads mapping to ribosomal RNA should be below 5 to 10 percent for most libraries. Higher proportions indicate incomplete ribosomal RNA removal and may require re-preparation of the library with a different depletion method. Some protocols tolerate higher ribosomal RNA contamination, particularly those designed for degraded RNA or for organisms with unusual RNA composition.

### Gene Body Coverage

Gene body coverage describes the distribution of sequencing reads across the length of transcripts. Ideally, reads should be distributed relatively uniformly across the entire transcript, with some expected variation due to fragmentation bias and sequence composition.

RNA degradation produces a characteristic 3-prime bias, where coverage is higher at the 3-prime end of transcripts than at the 5-prime end. This bias occurs because reverse transcription is primed from the 3-prime end, and degraded RNA produces cDNA fragments that are enriched for the 3-prime portion of transcripts. Gene body coverage plots can detect this bias and provide an independent assessment of RNA quality that complements RIN values.

Gene body coverage is also affected by the fragmentation method and the library preparation protocol. Some protocols produce libraries with inherent 3-prime or 5-prime bias, and the expected coverage profile depends on the specific protocol used.

### Read Distribution Across Genomic Features

The distribution of reads across genomic features provides information about library composition. In a typical mRNA-seq library, the majority of reads should map to exonic regions, with smaller proportions mapping to intronic and intergenic regions. The expected distribution depends on the library preparation method and the organism.

Poly-A selected libraries should show high exonic read proportions, typically above 70 percent. Ribosomal RNA depleted libraries may show higher intronic read proportions because they capture pre-mRNA and other non-coding transcripts. Unusually high intergenic read proportions may indicate contamination or annotation issues.

## Single-Cell RNA-seq Quality Control Considerations

Single-cell RNA-seq (scRNA-seq) presents additional quality control challenges beyond those of bulk RNA-seq. The analysis of single-cell data requires identification and removal of low-quality cells and technical artifacts before downstream analysis. The mixture of technical noise and biological variability makes separating technical artifacts from real biological variation particularly challenging.

### Cell-Level Quality Metrics

Single-cell RNA-seq quality control operates at two levels: the library level and the cell level. Library-level quality control follows the same principles as bulk RNA-seq, assessing RNA quality, library concentration, and sequencing quality. Cell-level quality control assesses the quality of each individual cell's expression profile.

Common cell-level quality metrics include the number of genes detected per cell, the total number of unique molecular identifiers (UMIs) per cell, and the proportion of mitochondrial reads per cell. Cells with very few detected genes or very low UMI counts are likely to be empty droplets or damaged cells. Cells with high mitochondrial read proportions are likely to be dying or dead, as mitochondrial RNA is more stable than cytoplasmic RNA.

The thresholds for these metrics depend on the cell type, the protocol, and the sequencing depth. Best practices recommend examining the distributions of these metrics across all cells and setting thresholds based on the data instead of using fixed values. Cells that fall outside the expected distributions should be examined individually before exclusion.

### Technical Artifact Detection

Technical artifacts in single-cell RNA-seq data can arise from multiple sources, including ambient RNA contamination, barcode swapping, and amplification bias. Ambient RNA from lysed cells can contaminate the capture solution and produce spurious expression signals in cells that did not express those genes. This contamination is particularly problematic for highly expressed genes.

Detection of technical artifacts requires integration of gene expression patterns with data quality metrics. Methods that combine multiple quality metrics can identify cells with unusual expression profiles that indicate technical problems instead of biological variation. The protocol for detecting technical artifacts should be applied consistently across all samples in an experiment.

### Integration of Quality Control into Single-Cell Workflows

Single-cell RNA-seq analysis workflows typically include quality control as the first step after generation of the count matrix. The workflow proceeds from quality control through normalization, data correction, feature selection, and dimensionality reduction before cell clustering and downstream analysis. Each step builds on the quality of the previous step, so inadequate quality control compromises all subsequent analyses.

Several software packages provide integrated quality control workflows for single-cell RNA-seq data. These packages calculate quality metrics, generate diagnostic plots, and support filtering decisions. The choice of package depends on the analysis environment and the specific requirements of the experiment.

## Tools and Platforms for Quality Control

Quality control for RNA-seq data can be performed using a variety of tools and platforms, ranging from command-line tools to web-based platforms. The choice of tool depends on the researcher's computational skills, the scale of the experiment, and the specific quality control tasks required.

### Command-Line Tools and Workflow Managers

Command-line tools provide the most flexibility for quality control analysis. FastQC is a widely used tool for assessing raw sequencing data quality, generating per-base quality plots, GC content distributions, adapter contamination reports, and sequence duplication levels. MultiQC aggregates quality control reports from multiple samples and tools into a single summary report, facilitating comparison across samples.

Workflow managers such as Snakemake and Nextflow support the automation of quality control pipelines. These tools enable reproducible analysis by defining the steps of the pipeline, the parameters used, and the expected outputs. The nf-core project provides community-curated pipelines that include quality control steps and follow standardized practices for reproducible analysis.

### Web-Based Platforms

Web-based platforms lower the barrier to RNA-seq analysis for researchers without extensive computational skills. The Galaxy platform provides a web interface for executing bioinformatics tools, including quality control tools, without requiring command-line expertise. Galaxy workflows can be saved, shared, and re-run, supporting reproducibility.

The Galaxy Training Network provides tutorials for RNA-seq analysis, including quality control steps. These tutorials guide users through the process of uploading data, running quality control tools, interpreting the results, and proceeding to downstream analysis. The training materials are designed to be accessible to researchers with limited bioinformatics experience.

### Reproducibility Considerations

Reproducibility is a central concern in RNA-seq quality control. The same dataset analyzed with different tools or parameters can produce different quality control results, affecting decisions about sample inclusion and preprocessing. To support reproducibility, researchers should document the versions of all tools used, the parameters applied, and the thresholds used for filtering decisions.

Workflow managers and containerization technologies support reproducibility by capturing the computational environment in which analysis was performed. Containers such as Docker and Singularity package software with their dependencies, ensuring that the same software version is used across different computing environments. The use of these technologies is increasingly expected in published research.

## Common Failure Patterns and Troubleshooting

Understanding common failure patterns in RNA-seq quality control helps researchers diagnose problems quickly and take corrective action. The following sections describe the most frequent failure patterns and their likely causes.

### Low Mapping Rate

A low mapping rate, typically below 70 percent for eukaryotic samples, indicates that a substantial proportion of reads do not align to the reference genome. The first step in troubleshooting is to examine the unmapped reads to determine their origin. If the unmapped reads show high similarity to ribosomal RNA sequences, ribosomal RNA contamination is likely. If the unmapped reads show similarity to bacterial or other non-target organism sequences, sample contamination is likely.

Reference genome issues can also cause low mapping rates. If the reference genome is incomplete or contains errors, reads from the corresponding genomic regions will not map. Using a different reference genome version or a more complete assembly may resolve this issue. For organisms with high genetic diversity, reads from the sample may differ substantially from the reference, requiring alignment parameters that allow more mismatches.

### Excessive Duplication

Excessive duplication, typically above 40 percent, reduces the effective sequencing depth and can bias expression quantification. The most common cause is over-amplification during library preparation, which occurs when too many PCR cycles are used or when input RNA is insufficient. Re-preparing the library with fewer PCR cycles or more input RNA usually resolves this issue.

Duplication can also result from sequencing depth that exceeds the complexity of the library. For libraries prepared from low-input RNA, the number of unique cDNA fragments may be limited, and high sequencing depth produces many duplicate reads. In this case, the duplication rate reflects the library complexity instead of a technical problem.

### Adapter Dimer Contamination

Adapter dimers appear as a small fragment peak in the library size distribution and produce reads that consist entirely of adapter sequence. These reads do not map to the reference genome and consume sequencing capacity. If adapter dimers constitute more than a few percent of the library, the library should be re-purified to remove the small fragments.

Adapter dimer contamination is more common in libraries prepared from low-input RNA, where the concentration of cDNA fragments is low relative to adapters. Optimization of the adapter-to-insert ratio and careful purification can reduce adapter dimer formation.

### RNA Degradation

RNA degradation produces libraries with 3-prime bias and reduced complexity. The first indication of RNA degradation is often a low RIN value, but degradation can also be detected in the gene body coverage plot after sequencing. If RNA degradation is detected after sequencing, the affected samples may need to be excluded or analyzed with methods that account for degradation.

For samples with moderate degradation, protocols optimized for degraded RNA may produce usable data. These protocols typically use random priming instead of oligo-dT priming and may include additional steps to maximize the capture of intact RNA fragments.

### Sample Mix-Ups and Contamination

Sample mix-ups are a serious problem in RNA-seq studies, particularly those involving multiple samples from the same individual or closely related samples. Verification of sample identity should be part of the quality control workflow. Methods that use genetic markers, such as HLA typing, can confirm that samples are correctly assigned to individuals. The RNA2HLA tool demonstrates the importance of sample identity verification by identifying wrongly assigned samples in public datasets.

Cross-contamination between samples can occur during library preparation or sequencing. The presence of reads from unexpected organisms or unexpected expression patterns may indicate contamination. Careful laboratory practices and the use of unique dual indexes can reduce the risk of cross-contamination.

## Records and Documentation for Quality Control

Maintaining detailed records of quality control metrics is essential for reproducible research and for troubleshooting problems that arise during analysis. The following records should be maintained for each RNA-seq experiment.

### Sample Tracking Records

Sample tracking records document the identity and provenance of each sample from collection through sequencing. These records should include:

1. Sample identifier and source
2. Date and method of collection
3. Storage conditions and duration
4. RNA extraction method and yield
5. RNA quality metrics including RIN or DV200 values
6. Library preparation method and kit
7. Library concentration and fragment size
8. Sequencing platform and run identifier

These records support traceability and enable identification of systematic problems that affect multiple samples.

### Quality Control Decision Logs

Quality control decision logs document the thresholds used for filtering and the decisions made for each sample. These logs should include:

1. The quality control metrics assessed
2. The thresholds applied for each metric
3. The rationale for any deviations from standard thresholds
4. The disposition of each sample, including whether it passed, failed, or was flagged for review
5. Any corrective actions taken, such as re-extraction or re-sequencing

Decision logs support transparency and reproducibility by documenting the basis for sample inclusion or exclusion.

### Analysis Parameter Records

Analysis parameter records document the software versions and parameters used for quality control and downstream analysis. These records should include:

1. Software names and versions
2. Reference genome version and annotation
3. Alignment parameters
4. Quality filtering thresholds
5. Normalization methods

These records enable replication of the analysis and support interpretation of results by other researchers.

## Limitations of Quality Control Metrics

Quality control metrics provide valuable information about data quality, but they have important limitations that researchers should understand.

### Thresholds Are Context-Dependent

The thresholds described in this article represent general guidance, not universal standards. The appropriate threshold for each metric depends on the organism, tissue type, library preparation method, sequencing platform, and experimental design. Researchers should establish thresholds based on their specific protocols and validate these thresholds using positive and negative controls.

For example, the acceptable mapping rate for a well-annotated model organism such as mouse or human may be above 80 percent, while the acceptable mapping rate for a poorly annotated non-model organism may be below 60 percent. Similarly, the acceptable duplication rate depends on the sequencing depth and the expression complexity of the sample.

### Metrics Can Be Misleading

Quality control metrics can be misleading in specific circumstances. A high RIN value does not guarantee that the RNA is suitable for all applications, as RIN measures ribosomal RNA integrity instead of messenger RNA integrity. A high mapping rate does not guarantee that the alignment is correct, as reads can map to the wrong location or map ambiguously to multiple locations.

Researchers should interpret quality control metrics in the context of the biological question and the expected characteristics of the samples. Unusual values should be investigated instead of automatically accepted or rejected.

### Quality Control Cannot Fix All Problems

Quality control can identify problems but cannot fix all of them. Samples with severe degradation, contamination, or other issues may need to be excluded from the analysis. Re-sequencing can address some problems, such as low sequencing depth or poor run quality, but cannot fix problems that originate in the sample or library.

The cost of re-sequencing should be weighed against the value of the data. For valuable samples that cannot be replaced, analysis with methods that accommodate lower quality may be preferable to exclusion. For samples that can be replaced, re-extraction and re-sequencing may be the most reliable path to high-quality data.

## Professional Escalation Criteria

Researchers should know when to escalate quality control problems to supervisors, core facility staff, or collaborators. The following situations warrant escalation:

1. Systematic quality failures affecting multiple samples or libraries, which may indicate problems with reagents, equipment, or protocols
2. Quality metrics that fall far outside expected ranges and cannot be explained by sample characteristics
3. Evidence of contamination that may affect other samples or experiments
4. Uncertainty about whether to include or exclude samples with borderline quality metrics
5. Disagreements between quality control metrics that suggest conflicting interpretations

Escalation should include a clear description of the problem, the quality control data, and the steps already taken to troubleshoot. Core facility staff and experienced collaborators can often identify problems that are not apparent to researchers with less experience.

## Frequently Asked Questions

### What is the minimum RIN value acceptable for RNA-seq library preparation?

The minimum acceptable RIN value depends on the research question and the library preparation method. For most standard bulk RNA-seq applications, a RIN of 7.0 or higher is recommended. However, protocols optimized for degraded RNA can produce usable data from samples with lower RIN values, particularly when using DV200 as an alternative metric. Researchers should establish thresholds based on their specific protocols and validate them with appropriate controls.

### How much sequencing depth is needed for reliable differential expression analysis?

The required sequencing depth depends on the organism, the number of samples, the expression complexity of the tissue, and the goals of the analysis. For standard bulk RNA-seq experiments in mammalian systems, 20 to 30 million reads per sample is commonly used for gene-level differential expression analysis. Higher depth may be needed for detecting low-abundance transcripts, analyzing isoform usage, or studying complex transcriptomes. Lower depth may suffice for highly expressed genes or for experiments with many biological replicates.

### What does a high duplication rate indicate in RNA-seq data?

A high duplication rate, typically above 30 to 40 percent, indicates that many reads originate from the same cDNA fragments. This can result from over-amplification during library preparation, insufficient input RNA, or sequencing depth that exceeds library complexity. High duplication rates reduce the effective sequencing depth and can bias expression quantification. The appropriate response depends on the cause, ranging from re-preparing the library to accepting the data with appropriate statistical adjustments.

### How should adapter contamination be handled in RNA-seq data?

Adapter contamination should be addressed by trimming adapter sequences from reads before alignment. Trimming tools identify adapter sequences and remove them from the 3-prime end of reads. The decision to trim involves a tradeoff between removing adapter sequence and retaining read length. For libraries with significant adapter contamination, trimming is essential to prevent alignment problems and inaccurate quantification.

### What is the difference between quality control for bulk RNA-seq and single-cell RNA-seq?

Bulk RNA-seq quality control focuses on library-level metrics such as RNA integrity, library concentration, fragment size, mapping rate, and duplication rate. Single-cell RNA-seq quality control includes these library-level metrics but adds cell-level metrics such as the number of genes detected per cell, UMI counts per cell, and mitochondrial read proportions. Single-cell analysis also requires identification and removal of technical artifacts such as ambient RNA contamination and empty droplets.

### How can sample mix-ups be detected in RNA-seq studies?

Sample mix-ups can be detected using genetic markers that distinguish individuals. Methods such as HLA typing can verify that samples are correctly assigned to individuals, as demonstrated by tools like RNA2HLA. Other approaches include comparing genotype calls from RNA-seq data with known genotypes or using expression of sex-specific genes to verify sample sex. Sample identity verification should be part of the quality control workflow for studies involving multiple samples from the same individual.

### What should be done when quality control metrics conflict with each other?

Conflicting quality control metrics require careful investigation. For example, a sample with a high RIN value but a low mapping rate may have contamination or reference genome issues. A sample with a low RIN value but a high mapping rate may have RNA degradation that does not affect alignment. When metrics conflict, researchers should examine the underlying data, consider the biological context, and consult with experienced colleagues or core facility staff before making decisions.

### How should quality control results be reported in publications?

Quality control results should be reported in the methods section of publications, including the metrics assessed, the thresholds applied, and the number of samples excluded. Key quality control metrics such as mapping rate and duplication rate should be reported for each sample or as summary statistics. Reporting quality control results supports reproducibility and enables readers to assess the reliability of the findings.

## Related Bioinformatics Guides

- [Single-Cell RNA Sequencing Quality Control: A Practical Guide to Filtering and Metrics](/knowledge/bioinformatics/single-cell-rna-sequencing-quality-control-a-practical-guide-to-filtering-and-metrics)
- [RNA-Seq Data Analysis Workflow: From Raw Reads to Insights](/knowledge/bioinformatics/rna-seq-data-analysis-workflow-from-raw-reads-to-insights)
- [RNA Sequencing Data Analysis: From Raw Reads to Differential Expression](/knowledge/bioinformatics/rna-sequencing-data-analysis-from-raw-reads-to-differential-expression)
- [RNA-Seq Data Analysis in Galaxy: A User-Friendly Platform](/knowledge/bioinformatics/rna-seq-data-analysis-in-galaxy-a-user-friendly-platform)
- [RNA-Seq Quality Control: Essential Checks and Tools](/knowledge/bioinformatics/rna-seq-quality-control-essential-checks-and-tools)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)


## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [Current best practices in single-cell RNA-seq analysis: a tutorial.](https://pubmed.ncbi.nlm.nih.gov/31217225). Molecular systems biology, 2019.
- [RNA-Seq Data Analysis.](https://pubmed.ncbi.nlm.nih.gov/38907924). Methods in molecular biology (Clifton, N.J.), 2024.
- [RNA-Seq Data Analysis in Galaxy.](https://pubmed.ncbi.nlm.nih.gov/33835453). Methods in molecular biology (Clifton, N.J.), 2021.
- [Quality Control of Single-Cell RNA-seq.](https://pubmed.ncbi.nlm.nih.gov/30758816). Methods in molecular biology (Clifton, N.J.), 2019.
- [RNA2HLA: HLA-based quality control of RNA-seq datasets.](https://pubmed.ncbi.nlm.nih.gov/33758920). Briefings in bioinformatics, 2021.
- [Short-Read RNA-Seq.](https://pubmed.ncbi.nlm.nih.gov/38907923). Methods in molecular biology (Clifton, N.J.), 2024.
- [Quality and composition control of complex TCM preparations through a novel "Herbs-in vivo Compounds-Targets-Pathways" network methodology: The case of Lianhuaqingwen capsules.](https://pubmed.ncbi.nlm.nih.gov/39947450). Pharmacological research, 2025.
- [Single-Cell RNA Sequencing Analysis: A Step-by-Step Overview.](https://pubmed.ncbi.nlm.nih.gov/33835452). Methods in molecular biology (Clifton, N.J.), 2021.
- [Increasing usable reads in RNA-seq protocols.](https://doi.org/10.1016/j.isci.2026.116984). 2026.
- [Integrating Multi-dimensional RNA Sequencing to Construct a Prognostic Risk Model for Bladder Cancer](https://doi.org/10.21203/rs.3.rs-9821610/v1). 2026.
- [VizR: An Interactive Web Platform for End-to-End RNA-Seq Analysis and Visualization in Plant Biology](https://doi.org/10.64898/2026.07.24.740467). 2026.
- [RAGER: A user-friendly computational platform for integrated analysis of RNA-Seq and ATAC-seq data.](https://doi.org/10.1371/journal.pone.0349941). 2026.
- [A comprehensive assessment of RNA-seq accuracy, reproducibility and information content by the Sequencing Quality Control consortium](https://doi.org/10.1038/nbt.2957). Nature Biotechnology, 2014.
- [popsicleR: A R Package for Pre-processing and Quality Control Analysis of Single Cell RNA-seq Data.](https://doi.org/10.1016/j.jmb.2022.167560). Journal of Molecular Biology, 2022.
- [Rup (RNA-seq Usability Assessment Pipeline)-Quality Control for Bulk RNA-seq Experiments in Eukaryotes](https://doi.org/10.3791/69253). Journal of Visualized Experiments, 2025.
- [Quality control of RNA-seq experiments](https://doi.org/10.1007/978-1-4939-2291-8_8). Methods in Molecular Biology, 2015.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.