RNA-seq Quality Control Metrics: A Decision Guide for When to Fail or Pass a Sample

By Dr. Zubair Khalid, DVM, MS, PhD ·

RNA-seq Quality Control Metrics: A Decision Guide for When to Fail or Pass a Sample

Key Takeaways

  • Per-base quality scores (Phred scores) are critical; median scores below 20 across most read lengths warrant concern, while scores above 30 indicate good sequencing quality, with thresholds adjusted for downstream applications like variant detection versus differential expression.
  • Excessive read duplication rates, particularly above 50-60% in bulk RNA-seq, can signal library preparation issues like low input RNA or over-amplification, whereas in single-cell RNA-seq, some duplication may be biologically valid.
  • Low alignment rates (<70%) suggest contamination, adapter issues, or reference mismatches, with uniquely aligned reads being particularly informative, as identified by the QC-DR analysis as highly correlated with overall sample quality.
  • Ribosomal RNA (rRNA) content exceeding 10% in poly-A selected libraries indicates inefficient depletion, necessitating pre-sequencing assessment via methods like 18S rRNA qPCR to prevent wasted sequencing resources.
  • Mitochondrial content thresholds in single-cell RNA-seq are tissue-specific; a universal 5% threshold is inadequate for tissues like cardiac muscle, where naturally higher mitochondrial proportions exist, necessitating context-dependent reference values.
  • Integrating multiple QC metrics, rather than relying on single indicators, provides more reliable quality assessment, with approaches like the QC-DR analysis and frameworks like scQCenrich demonstrating the value of multi-metric evaluation.

RNA sequencing quality control requires objective thresholds that separate technical artifacts from biological signal. This article provides a decision framework for evaluating per-base quality, duplication rates, alignment rates, gene detection, rRNA contamination, and mitochondrial content, with specific pass, warn, and fail criteria that researchers can apply before committing to downstream analysis. The guidance applies to bulk RNA-seq and single-cell RNA-seq workflows, with attention to the distinct failure modes of each approach.

The Quality Control Problem in RNA Sequencing

RNA-seq experiments generate millions of reads per sample, but the biological conclusions drawn from those reads depend entirely on the quality of the underlying data. A sample with degraded RNA, inefficient rRNA depletion, or sequencing artifacts can produce expression measurements that misrepresent the true transcriptome. The challenge for researchers is that no single quality metric reliably identifies all low-quality samples. Recent work on the Quality Control Diagnostic Renderer (QC-DR) applied to a clinical dataset of 252 alveolar macrophage samples found that individual QC metrics have limited predictive value when used alone, and that integrating multiple metrics with appropriate thresholds provides more reliable quality assessment [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>].

The practical consequence is that researchers need a structured decision process instead of a single pass-fail test. This article presents a checklist approach where each metric contributes evidence toward a final decision about whether to include, exclude, or re-sequence a sample. The thresholds provided here represent commonly used starting points, but they require adjustment based on library preparation method, organism, tissue type, and experimental design.

Core Quality Metrics and Their Biological Meaning

Per-Base Sequence Quality

Per-base quality scores, reported as Phred scores by sequencing instruments, indicate the probability of an incorrect base call at each position in the read. A Phred score of 30 corresponds to a 1 in 1000 error rate, while a score of 20 corresponds to a 1 in 100 error rate. Quality scores typically decline toward the 3-prime end of reads due to sequencing chemistry limitations.

For RNA-seq data, the per-base quality profile matters because low-quality bases contribute to misalignment and incorrect variant calls. The standard approach is to examine the per-base quality plot from tools such as FastQC, which is commonly taught in bioinformatics training programs through resources like the Galaxy Training Network and The Carpentries [<a href="#ref-3">3</a>][<a href="#ref-4">4</a>]. A sample with median quality scores below 20 across most of the read length warrants concern, while scores above 30 indicate good sequencing quality.

The decision threshold depends on downstream application. For differential expression analysis, moderate quality degradation at read ends may be tolerable because alignment algorithms can often still place reads correctly. For variant detection or analyses requiring high base-level accuracy, more stringent thresholds apply.

Read Duplication Rate

Duplicate reads arise when the same cDNA fragment is sequenced multiple times. In RNA-seq, some duplication is expected because highly expressed genes generate many copies of the same transcript. However, excessive duplication can indicate library preparation problems, such as low input RNA or over-amplification during PCR.

The interpretation of duplication rates differs between bulk and single-cell RNA-seq. In bulk RNA-seq, duplication rates above 50-60% often signal that the library has low complexity, meaning the same few transcripts dominate the sequencing output. This can occur when starting RNA quantity is too low or when too many PCR cycles are used during library preparation. In single-cell RNA-seq, duplicate reads can be valid biological duplicates, particularly for highly expressed genes, and some pipelines retain them for mitochondrial variant detection [<a href="#ref-5">5</a>].

The key question is whether duplication reflects biology or technical artifact. A sample with high duplication but good gene coverage and expected rRNA levels may be acceptable, while a sample with high duplication and low gene detection likely has a library preparation problem.

Alignment Rate

The alignment rate measures the proportion of reads that map to the reference genome or transcriptome. Low alignment rates indicate contamination, adapter issues, or reference mismatches. For human and mouse RNA-seq data, alignment rates above 70-80% are typical for good quality samples, though the exact expected value depends on the reference and alignment tool used.

The percentage of uniquely aligned reads is particularly informative. Reads that align to multiple locations in the genome may represent repetitive elements or homologous genes, and their inclusion or exclusion affects expression quantification. The QC-DR analysis identified the percentage and number of uniquely aligned reads as among the most highly correlated metrics with overall sample quality [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>].

Low alignment rates require investigation before deciding whether to fail a sample. Possible causes include rRNA contamination, adapter dimers, sample mix-ups, or a reference genome that does not match the sample species. Each cause has a different remedy, so the alignment rate alone does not determine the decision.

Gene Detection

The number of detected genes reflects the complexity of the RNA population captured in the library. A typical mammalian bulk RNA-seq sample detects 10,000 to 20,000 genes, while single-cell samples detect fewer genes per cell due to the limited RNA content of individual cells.

Low gene detection can indicate RNA degradation, inefficient reverse transcription, or sequencing depth that is too low to capture rare transcripts. The number of detected genes was among the most informative QC metrics in the QC-DR analysis, correlating strongly with overall sample quality [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>].

For single-cell RNA-seq, the number of detected genes per cell is a standard QC metric, but thresholds must be set based on the cell type and tissue. A fixed threshold that works for one tissue may exclude healthy cells from another tissue with different transcriptional complexity.

Ribosomal RNA Content

Ribosomal RNA comprises more than 80-90% of total cellular RNA, and efficient removal is essential for capturing the transcriptome, particularly low-abundance mRNAs [<a href="#ref-6">6</a>]. Inefficient rRNA removal during library preparation can result from variations in sample quality, preparation methods, and handling [<a href="#ref-6">6</a>].

The percentage of reads mapping to rRNA is a direct measure of depletion efficiency. For poly-A selected libraries, rRNA content should be low because the selection process targets messenger RNA with poly-A tails. For total RNA libraries with rRNA depletion, the expected rRNA percentage depends on the depletion method used.

A recent study introduced a real-time qPCR-based assay targeting human 18S rRNA to evaluate rRNA depletion efficiency before sequencing [<a href="#ref-6">6</a>]. The assay was optimized using serial dilutions of Universal Human Reference control, and Ct thresholds were established using pilot data from 644 libraries [<a href="#ref-6">6</a>]. Analysis of 1748 human total RNA libraries and 445 poly-A libraries demonstrated a strong correlation between 18S rRNA qPCR results and post-sequencing rRNA rates [<a href="#ref-6">6</a>]. This pre-sequencing assessment allows researchers to identify problematic libraries before spending sequencing resources.

Mitochondrial Content

Mitochondrial transcripts can indicate cell stress or lysis in single-cell RNA-seq, where the proportion of reads mapping to mitochondrial genes is a standard QC metric. The commonly used threshold of 5% mitochondrial transcripts originated in early single-cell studies and has been adopted as a default in many software packages [<a href="#ref-7">7</a>].

However, this threshold does not apply universally. A systematic analysis of 5,530,106 cells from 1349 annotated datasets found that average mitochondrial percentage in human tissues is significantly higher than in mouse tissues, and that the 5% threshold fails to accurately discriminate between healthy and low-quality cells in 29.5% of the 44 human tissues analyzed [<a href="#ref-7">7</a>]. For cardiac tissue, mitochondrial transcripts can comprise almost 30% of total mRNA due to high energy demands, and applying a 5% threshold causes unacceptable exclusion of cardiomyocytes and introduces bias that discriminates against pacemaker cells [<a href="#ref-8">8</a>].

The recommendation is to determine tissue-specific mitochondrial thresholds instead of relying on a universal default. For human tissues, reference values have been proposed for 44 tissues, while for mouse tissues the 5% threshold generally performs well [<a href="#ref-7">7</a>].

At a Glance: Pass, Warn, and Fail Thresholds

The following table provides starting thresholds for common RNA-seq QC metrics. These values require adjustment based on library preparation method, organism, tissue type, and experimental design. The warn category indicates that the metric should be investigated in context with other metrics before making a final decision.

MetricPassWarnFail
Per-base quality (median Phred score)Above 30 across most read positions20-30 at read endsBelow 20 across most of the read
Alignment rateAbove 80% total aligned70-80% total alignedBelow 70% total aligned
Uniquely aligned readsAbove 70% of total reads50-70% of total readsBelow 50% of total reads
rRNA content (poly-A library)Below 5% of reads5-10% of readsAbove 10% of reads
Detected genes (bulk, mammalian)Above 12,000 genes8,000-12,000 genesBelow 8,000 genes
Duplication rate (bulk)Below 40%40-60%Above 60% with low gene detection
Mitochondrial content (single-cell)Tissue-specific reference2-3 times tissue referenceAbove 3 times tissue reference

Practical Workflow for Quality Control Decisions

Step 1: Run Standard QC Tools

The first step is to generate QC reports using established tools. FastQC provides per-base quality, duplication, and adapter content information. MultiQC aggregates results across samples for comparison. Alignment-based metrics come from tools such as STAR, HISAT2, or Salmon, depending on the analysis pipeline.

Training resources from the Galaxy Training Network provide step-by-step tutorials for running these tools and interpreting their output [<a href="#ref-3">3</a>]. The nf-core documentation describes community-standard pipelines that include QC steps as part of reproducible workflows [<a href="#ref-9">9</a>]. Bioconductor packages provide R-based tools for downstream QC analysis and visualization [<a href="#ref-10">10</a>].

Step 2: Compare Samples Within the Batch

Individual QC metrics are most informative when compared across samples from the same experiment. A sample that falls outside the distribution of other samples warrants investigation even if it passes absolute thresholds. The mirnaQC webserver demonstrates this approach for miRNA-seq data, ranking quality attributes using a reference distribution obtained from over 36,000 publicly available datasets [<a href="#ref-11">11</a>].

For single-cell RNA-seq, the miQC package provides a data-driven approach that jointly models mitochondrial proportion and detected genes using mixture models, instead of applying fixed thresholds [<a href="#ref-12">12</a>]. This approach adapts to different tissue types and preserves high-quality cells that would be excluded by uniform thresholds [<a href="#ref-12">12</a>].

Step 3: Evaluate Metrics in Combination

The QC-DR analysis concluded that any individual QC metric is limited in its predictive value, and that approaches based on integration of multiple metrics with QC thresholds provide more reliable quality assessment [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>]. A sample that fails one metric but passes all others may be acceptable, while a sample that marginally fails several metrics likely has a real quality problem.

The scQCenrich framework extends this multi-metric approach for single-cell RNA-seq, integrating canonical metrics with intronic fraction, MALAT1 enrichment, and dissociation-stress features [<a href="#ref-13">13</a>]. This framework reduces over-filtering relative to conventional methods while preserving coherent cell populations [<a href="#ref-13">13</a>].

Step 4: Document the Decision

Record the QC metrics for every sample, the thresholds applied, and the rationale for inclusion or exclusion. This documentation supports reproducibility and allows reviewers to assess the impact of QC decisions on the final results. The nf-core documentation emphasizes the importance of reproducible workflows that include QC reporting [<a href="#ref-9">9</a>].

Records and Measurements for Quality Control

Maintain a laboratory notebook or electronic record that includes the following information for each RNA-seq sample:

  • RNA quantity and quality measurements from the extraction step, including RNA integrity number or equivalent metrics
  • Library preparation method and any deviations from the standard protocol
  • rRNA depletion method and efficiency measurements if pre-sequencing assessment was performed
  • Sequencing platform, read length, and sequencing depth
  • All QC metric values from the bioinformatics pipeline
  • The final pass, warn, or fail decision with justification

The pre-sequencing qPCR assay for 18S rRNA provides a cost-effective method for evaluating rRNA depletion efficiency before committing sequencing resources [<a href="#ref-6">6</a>]. Including this measurement in the records allows early identification of problematic libraries.

Common Failure Patterns and Their Causes

Pattern 1: High rRNA Content with Low Gene Detection

This pattern indicates inefficient rRNA depletion during library preparation. The library is dominated by rRNA reads, leaving fewer reads available for messenger RNA. The result is reduced gene detection and potentially biased expression measurements for low-abundance transcripts.

The pre-sequencing qPCR assay can identify this problem before sequencing, allowing the library to be remade or the depletion step to be repeated [<a href="#ref-6">6</a>]. If the problem is detected after sequencing, the sample may need to be re-sequenced with a new library.

Pattern 2: Low Alignment Rate with High Adapter Content

This pattern suggests adapter dimers or other library preparation artifacts. Adapter dimers are short fragments consisting of adapter sequences ligated to each other without an insert. They align poorly to the reference genome and consume sequencing capacity.

Adapter trimming can remove some adapter sequences, but excessive adapter dimer content indicates a library preparation problem that may require remaking the library.

Pattern 3: High Duplication Rate with Adequate Gene Detection

This pattern can occur with low input RNA or excessive PCR amplification. The library has sufficient complexity to detect many genes, but the sequencing output is dominated by a small number of highly amplified fragments. This reduces the effective sequencing depth for rare transcripts.

The decision to fail or pass depends on whether the duplication affects the specific analyses planned. For differential expression of highly expressed genes, the impact may be minimal. For detection of rare transcripts or isoform-level analysis, the impact can be substantial.

Pattern 4: Tissue-Specific Mitochondrial Content Exceeding Default Thresholds

This pattern occurs when researchers apply a universal mitochondrial threshold to a tissue with naturally high mitochondrial content. Cardiac tissue, kidney tissue, and archived tumor tissues can have mitochondrial proportions well above the 5% default [<a href="#ref-8">8</a>][<a href="#ref-12">12</a>]. Applying the default threshold excludes healthy cells and introduces biological bias.

The solution is to determine tissue-specific thresholds using reference values or data-driven approaches such as miQC [<a href="#ref-12">12</a>][<a href="#ref-7">7</a>].

Single-Cell RNA-Seq Quality Control Considerations

Cell-Level Filtering

Single-cell RNA-seq requires quality control at two levels: the sample level and the cell level. Sample-level QC follows similar principles to bulk RNA-seq, while cell-level QC involves filtering individual cells based on detected genes, UMI counts, and mitochondrial content.

The scReady pipeline automates cell-level QC, including ambient RNA removal, doublet detection, and cell and gene filtering based on mitochondrial content and customizable thresholds [<a href="#ref-14">14</a>]. This pipeline is designed for novice bioinformaticians and produces a fully processed Seurat object with diagnostic plots and a comprehensive quality control report [<a href="#ref-14">14</a>].

Adaptive Thresholds for Cell Filtering

Fixed thresholds for cell filtering are problematic because optimal values vary by tissue, cell type, and experimental condition. An intersective approach that combines adaptive QC based on median absolute deviation with manual QC based on visual inspection of histograms and violin plots has been applied to gastric cancer scRNA-seq data [<a href="#ref-15">15</a>]. This approach filtered 134,367 cells from 28 patients down to 116,666 high-quality cells by intersecting the sets of cells removed by both methods [<a href="#ref-15">15</a>].

The miQC package provides a probabilistic framework that jointly models mitochondrial proportion and detected genes, adapting to the specific characteristics of each dataset [<a href="#ref-12">12</a>]. This approach is available through Bioconductor and can be integrated into standard single-cell analysis workflows [<a href="#ref-10">10</a>][<a href="#ref-12">12</a>].

Doublet Detection

Doublets occur when two cells are captured in the same droplet or well and sequenced together. Doublets can be mistaken for novel cell types or intermediate cell states. Detection methods use the observation that doublets have higher UMI counts and express marker genes from multiple cell types.

The scReady pipeline includes doublet detection as part of its automated preprocessing [<a href="#ref-14">14</a>]. The scQCenrich framework also addresses doublet-related artifacts through its multi-metric approach [<a href="#ref-13">13</a>].

Quality Control for miRNA-Seq

MicroRNA sequencing has distinct quality control requirements because the short read length and the nature of small RNA libraries create different failure modes. The mirnaQC webserver uses 34 quality parameters to assist in miRNA-seq quality control, ranking quality attributes using a reference distribution from over 36,000 publicly available datasets [<a href="#ref-11">11</a>].

Key miRNA-seq quality issues include microRNA yield, the fraction of putative degradation products such as rRNA fragments, and the percentage of adapter dimers [<a href="#ref-11">11</a>]. These issues are difficult to assess using absolute thresholds, which is why the reference distribution approach provides more interpretable results [<a href="#ref-11">11</a>].

The mirnaQC webserver accepts FASTQ files and SRA accessions, and the results page includes sections for library preparation artifacts, sequencing issues, contamination, and yield [<a href="#ref-11">11</a>]. Principal component analysis and heatmaps help identify underlying issues across samples [<a href="#ref-11">11</a>].

Reproducibility and Workflow Standards

Containerized Pipelines

Reproducible quality control requires standardized pipelines that produce consistent results across runs and users. The nf-core documentation describes community-developed pipelines that follow best practices for reproducibility, including containerization and version control [<a href="#ref-9">9</a>]. These pipelines include QC steps as part of the standard workflow.

The scReady pipeline for single-cell RNA-seq is containerized and designed for both single machines and high-performance computing clusters [<a href="#ref-14">14</a>]. This approach ensures that the same input data produces the same output regardless of the computing environment.

Training and Skill Development

Quality control decisions require bioinformatics skills that can be developed through structured training. The EMBL-EBI Training program provides learning pathways for bioinformatics data resources and practical analysis education [<a href="#ref-16">16</a>]. The Carpentries offers foundational lessons in computing, data handling, shell, Git, and programming that support reproducible analysis practices [<a href="#ref-4">4</a>].

The Galaxy Training Network provides accessible workflow training with tutorials that cover quality control, alignment, and downstream analysis [<a href="#ref-3">3</a>]. These resources allow researchers to develop the skills needed to make informed QC decisions.

Documentation Standards

The nf-core documentation emphasizes the importance of documenting pipeline usage, configuration, and results for reproducibility [<a href="#ref-9">9</a>]. For quality control decisions, documentation should include the version of each tool used, the parameters applied, and the rationale for any threshold adjustments.

Limitations of Quality Control Metrics

Metrics Cannot Detect All Artifacts

Quality control metrics detect technical artifacts that manifest in the sequencing data, but they cannot detect all problems. Sample mix-ups, contamination with RNA from other species, and batch effects that uniformly affect all samples may not be apparent from QC metrics alone.

The QC-DR analysis found that experimental QC metrics derived from the laboratory were not significantly correlated with pipeline QC metrics [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>]. This suggests that laboratory measurements and bioinformatics metrics capture different aspects of sample quality, and both should be considered in the decision process.

Thresholds Are Context-Dependent

The thresholds provided in this article are starting points that require adjustment based on the specific experimental context. A threshold that works for one tissue, library preparation method, or sequencing platform may not work for another. The systematic analysis of mitochondrial proportions across tissues demonstrates this clearly, with the 5% threshold failing for 29.5% of human tissues analyzed [<a href="#ref-7">7</a>].

Data-driven approaches that determine thresholds from the distribution of metrics within each dataset provide more reliable results than fixed thresholds [<a href="#ref-12">12</a>][<a href="#ref-15">15</a>]. These approaches require sufficient sample sizes to estimate the distribution reliably.

Quality Control Cannot Rescue Poor Experimental Design

Quality control identifies problematic samples, but it cannot compensate for poor experimental design. Inadequate biological replication, confounding between conditions and batch, and insufficient sequencing depth are design issues that QC metrics cannot resolve. The decision to exclude a sample should be made with consideration of the overall experimental design and the impact of exclusion on statistical power.

Professional Escalation Criteria

When to Consult a Bioinformatics Specialist

Researchers should escalate quality control decisions to a bioinformatics specialist or core facility when:

  • Multiple samples in the same batch fail the same QC metric, suggesting a systematic problem with the library preparation or sequencing run
  • The cause of a QC failure cannot be identified from the available metrics
  • The sample is irreplaceable, such as a rare clinical specimen, and the decision to exclude or include has major consequences
  • The QC metrics are ambiguous, with some metrics passing and others failing
  • The downstream analysis is sensitive to small quality differences, such as variant detection or isoform-level analysis

When to Consider Re-sequencing

Re-sequencing should be considered when:

  • The sample fails multiple QC metrics and the failure is likely due to a technical problem that can be corrected
  • The library preparation method can be improved based on the observed failure pattern
  • The biological question requires high-quality data that the current sample cannot provide
  • The cost of re-sequencing is justified by the importance of the sample to the overall study

When to Exclude a Sample

Exclusion should be considered when:

  • The sample fails multiple QC metrics and the failure cannot be corrected by re-sequencing
  • The sample is one of several replicates, and exclusion does not compromise the statistical design
  • The QC failure is likely to introduce bias into the downstream analysis
  • The sample quality is so poor that even re-sequencing is unlikely to produce usable data

Safety and Ethical Context

Sample Provenance and Consent

Quality control decisions should consider the provenance of samples and the consent under which they were collected. For clinical samples, the decision to exclude or re-sequence may have implications for the use of limited patient material. Researchers should document the rationale for QC decisions in a way that supports transparency and reproducibility.

Data Sharing and Public Repositories

The NCBI provides databases and search systems for sequence data, including the Sequence Read Archive and the Gene Expression Omnibus [<a href="#ref-17">17</a>]. When sharing RNA-seq data, quality control metrics should be included as metadata to allow downstream users to assess data quality. The EMBL-EBI Training program provides guidance on data submission and the use of public data resources [<a href="#ref-16">16</a>].

Misinterpretation Risks

Poor quality control can lead to erroneous biological interpretations. The analysis of mitochondrial proportions in single-cell RNA-seq demonstrated that omitting the mitochondrial QC filter or adopting a suboptimal threshold may lead to erroneous biological interpretations [<a href="#ref-7">7</a>]. Similarly, the exclusion of pacemaker cells due to an inappropriate mitochondrial threshold would bias any analysis of cardiac tissue [<a href="#ref-8">8</a>].

Researchers should report QC decisions transparently in publications, including the thresholds applied and the number of samples excluded. This allows readers to assess the potential impact of QC decisions on the reported results.

Building a Structured QC Decision Log and Escalation Protocol

The pass, warn, and fail thresholds described in the previous sections provide the technical basis for quality decisions, but they do not by themselves create a consistent decision process. Researchers often apply thresholds inconsistently across batches, fail to document the rationale for borderline decisions, or revisit the same quality questions repeatedly when new samples arrive. A structured decision log that records every QC evaluation, the evidence considered, and the final disposition of each sample transforms quality control from an informal judgment into an auditable process. This section presents a practical framework for building that log, integrating multiple metrics into a single decision score, and establishing clear escalation criteria that prevent both premature exclusion of salvageable samples and inappropriate inclusion of failed libraries.

The Multi-Metric Decision Matrix

The QC-DR analysis of 252 clinical RNA-seq samples demonstrated that individual QC metrics have limited predictive value when used alone, and that approaches based on integration of multiple metrics with QC thresholds provide more reliable quality assessment [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>]. The metrics most highly correlated with overall sample quality were the percentage and number of uniquely aligned reads, the percentage of rRNA reads, the number of detected genes, and the area under the gene body coverage curve [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>]. Notably, experimental QC metrics derived from the laboratory were not significantly correlated with pipeline QC metrics, suggesting that laboratory measurements and bioinformatics metrics capture different aspects of sample quality [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>].

A practical way to operationalize this finding is to build a decision matrix that assigns a score to each metric based on whether it falls in the pass, warn, or fail range. The matrix approach prevents the common error of excluding a sample based on a single alarming metric while ignoring the overall pattern. It also prevents the opposite error of including a sample that marginally passes every metric but collectively shows a consistent pattern of degradation.

The following scoring system provides a starting point for building a decision matrix:

MetricPass (score 2)Warn (score 1)Fail (score 0)
Per-base quality (median Phred score)Above 30 across most read positions20-30 at read endsBelow 20 across most of the read
Total alignment rateAbove 80%70-80%Below 70%
Uniquely aligned readsAbove 70% of total reads50-70% of total readsBelow 50% of total reads
rRNA content (poly-A library)Below 5% of reads5-10% of readsAbove 10% of reads
Detected genes (bulk, mammalian)Above 12,000 genes8,000-12,000 genesBelow 8,000 genes
Duplication rate (bulk)Below 40%40-60%Above 60% with low gene detection
Gene body coverage uniformityEven coverage across gene lengthModerate 3-prime or 5-prime biasSevere coverage bias

A sample with a total score of 12 or higher across six metrics generally passes without further investigation. A score between 8 and 11 warrants review of the specific failing metrics and their potential impact on the planned analysis. A score below 8 indicates that the sample likely has a quality problem requiring either re-sequencing or exclusion. These score ranges are starting points instead of absolute rules, and they should be calibrated against the distribution of scores across the full batch of samples in an experiment.

The gene body coverage metric deserves particular attention because it captures RNA degradation patterns that other metrics miss. The area under the gene body coverage curve (AUC-GBC) was identified as one of the most informative metrics in the QC-DR analysis [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>]. A sample with degraded RNA typically shows reduced coverage toward the 5-prime end of transcripts, while a sample with 3-prime bias from library preparation shows the opposite pattern. This metric is not included in standard FastQC reports and requires alignment-based analysis, but it provides information that per-base quality and alignment rates cannot reveal.

Building the Sample Decision Record

For each sample in an RNA-seq experiment, create a decision record that captures the following information in a consistent format:

Sample identification and provenance. Record the sample identifier, tissue type, organism, and any relevant clinical or experimental metadata. This information provides context for interpreting QC metrics, particularly for tissues with known characteristics such as high mitochondrial content in cardiac tissue [<a href="#ref-8">8</a>] or high rRNA content in certain sample types.

Laboratory measurements. Record RNA quantity, RNA integrity number or equivalent quality metric, library preparation method, rRNA depletion method, and any pre-sequencing quality assessments. The 18S rRNA qPCR assay provides a cost-effective method for evaluating rRNA depletion efficiency before sequencing, with Ct thresholds established from pilot data and validated across 1748 total RNA and 445 poly-A libraries [<a href="#ref-6">6</a>]. Including this measurement in the decision record allows early identification of problematic libraries before sequencing resources are committed.

Pipeline metrics. Record all QC metrics generated by the bioinformatics pipeline, including per-base quality, duplication rate, alignment rate, uniquely aligned reads, rRNA content, detected genes, and gene body coverage. The nf-core documentation describes community-standard pipelines that include QC reporting as part of reproducible workflows [<a href="#ref-9">9</a>]. These pipelines generate consistent metric outputs that can be recorded in a standardized format.

Thresholds applied. Record the specific thresholds used for each metric and the source of those thresholds, whether from published guidelines, tissue-specific reference values, or data-driven determination from the current dataset. For mitochondrial content in single-cell RNA-seq, tissue-specific reference values should be used instead of the universal 5% threshold [<a href="#ref-7">7</a>]. The systematic analysis of 5,530,106 cells from 1349 annotated datasets provides reference values for 121 mouse tissues and 44 human tissues [<a href="#ref-7">7</a>].

Decision and rationale. Record the final disposition of the sample, whether pass, warn with conditions, re-sequence, or exclude. Document the specific reasons for the decision, including which metrics influenced the outcome and how the decision aligns with the overall experimental design.

Follow-up actions. Record any actions taken, such as re-sequencing with a new library, adjusting thresholds for subsequent samples, or consulting a bioinformatics specialist.

The decision record serves multiple purposes. It provides documentation for publications and reviewers, allowing them to assess the impact of QC decisions on reported results. It creates a reference for future experiments, allowing researchers to identify recurring quality issues and adjust protocols accordingly. It also supports reproducibility by ensuring that the same input data and thresholds produce the same decisions regardless of who performs the analysis.

Batch-Level Pattern Recognition

Individual sample decisions are necessary but not sufficient for reliable quality control. Patterns across samples within a batch often reveal systematic problems that individual metrics cannot detect. A batch in which multiple samples fail the same metric suggests a problem with the library preparation batch, the sequencing run, or the reagent lot, instead of with individual samples.

The mirnaQC webserver demonstrates the value of reference distributions for identifying aberrant samples, ranking quality attributes using a reference distribution obtained from over 36,000 publicly available miRNA-seq datasets [<a href="#ref-11">11</a>]. This approach allows researchers to identify samples that fall outside the expected range even when they pass absolute thresholds. Principal component analysis and heatmaps help identify underlying issues across samples [<a href="#ref-11">11</a>].

For batch-level analysis, create a summary table that lists all samples in the batch with their QC metric values. Look for the following patterns:

One sample deviates from the batch. A single sample with outlier values for multiple metrics likely has an individual problem such as RNA degradation, sample mix-up, or library preparation failure. This sample should be evaluated for exclusion or re-sequencing.

Multiple samples share the same deviation. When several samples from the same preparation batch or sequencing run show similar QC failures, the problem is likely systematic. This pattern requires investigation of the shared step, whether library preparation, rRNA depletion, or sequencing.

The batch shows a shift from previous batches. If the current batch has consistently lower alignment rates or higher duplication rates than previous batches using the same protocol, a reagent lot change, equipment issue, or protocol drift may be responsible.

Samples cluster by biological condition. If QC metrics correlate with the biological condition being studied, the quality differences may reflect genuine biological variation instead of technical artifacts. For example, degraded samples from clinical settings may cluster by disease severity. This pattern requires careful interpretation to avoid removing biologically meaningful variation.

The QC-DR software was designed to simultaneously visualize a comprehensive panel of QC metrics and flag samples with aberrant values when compared to a reference dataset [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>]. This approach provides a model for batch-level quality assessment that goes beyond individual sample evaluation.

The Escalation Protocol

A structured escalation protocol ensures that quality decisions receive appropriate review based on their impact on the study. The protocol should define three levels of review:

Level 1: Standard review. Samples that pass all metrics or fail only one metric with a clear explanation require no additional review. The decision is recorded in the sample decision record and the sample proceeds to downstream analysis.

Level 2: Supervised review. Samples that fail multiple metrics, have ambiguous patterns, or are irreplaceable require review by a second researcher or a bioinformatics specialist. The reviewer examines the decision record, evaluates the pattern of metrics, and makes a recommendation. This level of review is appropriate for samples that are expensive to replace, such as rare clinical specimens, or samples where the QC failure pattern does not match a known technical problem.

Level 3: Committee review. Samples that will be excluded from a study, samples that fail in ways that suggest systematic problems, or samples where the QC decision materially affects the study conclusions require review by the full research team. This review considers the impact of the decision on statistical power, the potential for bias, and the availability of alternative samples.

The escalation protocol should also define when to consult external expertise. Researchers should escalate to a bioinformatics specialist or core facility when multiple samples in the same batch fail the same QC metric, when the cause of a QC failure cannot be identified from the available metrics, when the sample is irreplaceable and the decision has major consequences, when QC metrics are ambiguous with some passing and others failing, or when the downstream analysis is sensitive to small quality differences such as variant detection or isoform-level analysis.

Re-sequencing Decision Criteria

The decision to re-sequence a sample should be based on the likelihood that the quality problem can be corrected and the cost-benefit of the re-sequencing effort. Re-sequencing should be considered when the sample fails multiple QC metrics and the failure is likely due to a technical problem that can be corrected, when the library preparation method can be improved based on the observed failure pattern, when the biological question requires high-quality data that the current sample cannot provide, and when the cost of re-sequencing is justified by the importance of the sample to the overall study.

The pre-sequencing 18S rRNA qPCR assay provides a particularly valuable tool for avoiding wasted sequencing resources [<a href="#ref-6">6</a>]. By identifying libraries with inefficient rRNA depletion before sequencing, this assay allows researchers to remake libraries or adjust depletion protocols before committing sequencing capacity. The assay demonstrated a strong correlation between 18S rRNA qPCR results and post-sequencing rRNA rates across 1748 total RNA and 445 poly-A libraries [<a href="#ref-6">6</a>].

When re-sequencing is not feasible, researchers should consider whether the sample can be included with appropriate caveats. A sample that fails one metric but passes all others may be acceptable, particularly if the failing metric has a known cause that does not affect the planned analysis. The decision should consider the pattern of metrics across the sample and the sensitivity of the downstream analysis to the specific quality issue.

Common Decision Errors and How to Avoid Them

Several recurring errors undermine quality control decisions. Recognizing these patterns helps researchers avoid them:

Threshold rigidity. Applying fixed thresholds without considering tissue type, library preparation method, or experimental context leads to inappropriate exclusions. The 5% mitochondrial threshold in single-cell RNA-seq exemplifies this problem, failing to discriminate healthy from low-quality cells in 29.5% of human tissues analyzed [<a href="#ref-7">7</a>]. Data-driven approaches such as miQC, which jointly models mitochondrial proportion and detected genes using mixture models, adapt to the specific characteristics of each dataset [<a href="#ref-12">12</a>].

Single-metric decisions. Making inclusion or exclusion decisions based on one metric while ignoring the overall pattern leads to both false inclusions and false exclusions. The QC-DR analysis concluded that any individual QC metric is limited in its predictive value [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>]. The multi-metric decision matrix addresses this error by requiring consideration of the full metric panel.

Inconsistent documentation. Applying different thresholds or decision criteria across batches without documentation makes the QC process unreproducible and difficult to defend in publication. The decision record system described above addresses this error by requiring consistent documentation for every sample.

Confusing biological variation with technical artifacts. Some QC metrics reflect genuine biological variation instead of technical problems. Cardiac tissue has mitochondrial transcripts comprising almost 30% of total mRNA due to high energy demands [<a href="#ref-8">8</a>]. Applying a threshold appropriate for other tissues would exclude healthy cardiomyocytes and introduce bias that discriminates against pacemaker cells [<a href="#ref-8">8</a>]. Similarly, high duplication rates in single-cell RNA-seq can reflect valid biological duplicates for highly expressed genes [<a href="#ref-5">5</a>].

Ignoring batch effects. Quality differences between batches can confound biological comparisons. If all samples from one condition were sequenced in one batch and all samples from another condition in a different batch, QC differences between batches become inseparable from biological differences. This design problem cannot be corrected by QC filtering and requires attention at the experimental design stage.

Implementing the Decision Framework in Practice

The following implementation steps provide a practical path for establishing a structured QC decision process in a research laboratory:

Step 1: Define the metric panel. Select the QC metrics that will be evaluated for every sample. The panel should include per-base quality, alignment rate, uniquely aligned reads, rRNA content, detected genes, duplication rate, and gene body coverage for bulk RNA-seq. Single-cell RNA-seq adds cell-level metrics including detected genes per cell, UMI counts per cell, and mitochondrial content per cell.

Step 2: Establish baseline thresholds. Use the thresholds provided in the At a Glance table as starting points. Adjust thresholds based on the specific tissue, library preparation method, and sequencing platform used in the laboratory. For mitochondrial content in single-cell RNA-seq, consult tissue-specific reference values [<a href="#ref-7">7</a>].

Step 3: Create the decision record template. Develop a standardized template for recording QC metrics, thresholds, decisions, and rationale for each sample. The template should be accessible to all members of the research team and should be used consistently across experiments.

Step 4: Run the pipeline and generate metrics. Use established pipelines that include QC reporting as part of the standard workflow. The nf-core documentation describes community-standard pipelines that follow best practices for reproducibility [<a href="#ref-9">9</a>]. The Galaxy Training Network provides accessible workflow training with tutorials covering quality control, alignment, and downstream analysis [<a href="#ref-3">3</a>].

Step 5: Apply the decision matrix. Score each metric using the pass, warn, and fail categories. Calculate the total score and apply the decision criteria. Document the score and the rationale for the final decision.

Step 6: Review batch-level patterns. Compare QC metrics across all samples in the batch. Identify outlier samples and systematic patterns. Investigate any patterns that suggest shared technical problems.

Step 7: Escalate when appropriate. Apply the escalation protocol for samples that fail multiple metrics, are irreplaceable, or have ambiguous quality patterns. Consult a bioinformatics specialist when the cause of a QC failure cannot be identified.

Step 8: Document and archive. Store all decision records in a location accessible to the research team. Include QC decisions in publications and data submissions. The NCBI provides databases for sequence data, including the Sequence Read Archive and the Gene Expression Omnibus [<a href="#ref-17">17</a>]. When sharing RNA-seq data, include QC metrics as metadata to allow downstream users to assess data quality.

The structured decision framework described in this section provides a practical method for implementing the multi-metric approach to quality control that the QC-DR analysis identified as most reliable [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>]. By combining a decision matrix, a standardized record system, batch-level pattern recognition, and a clear escalation protocol, researchers can make objective, defensible quality decisions that support reproducible and reliable RNA-seq analysis.

Frequently Asked Questions

What is the most important RNA-seq quality control metric?

No single metric reliably identifies all low-quality samples. The QC-DR analysis found that the percentage and number of uniquely aligned reads, the percentage of rRNA reads, the number of detected genes, and the area under the gene body coverage curve were among the most highly correlated metrics with sample quality [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>]. The most informative approach integrates multiple metrics instead of relying on any single one.

How should I set quality control thresholds for my experiment?

Start with commonly used thresholds from published guidelines, then adjust based on your specific tissue, library preparation method, and sequencing platform. For mitochondrial content in single-cell RNA-seq, use tissue-specific reference values instead of the universal 5% threshold [<a href="#ref-7">7</a>]. Data-driven approaches such as miQC can determine thresholds from the distribution of metrics within your dataset [<a href="#ref-12">12</a>].

Can I use a sample that fails one quality control metric?

A sample that fails one metric but passes all others may be acceptable, particularly if the failing metric has a known cause that does not affect the planned analysis. The decision should consider the pattern of metrics across the sample and the sensitivity of the downstream analysis to the specific quality issue.

What causes high duplication rates in RNA-seq?

High duplication rates can result from low input RNA, excessive PCR amplification during library preparation, or sequencing depth that exceeds the complexity of the library. In single-cell RNA-seq, duplicate reads can be valid biological duplicates for highly expressed genes [<a href="#ref-5">5</a>]. The interpretation depends on whether the duplication reflects biology or technical artifact.

How does rRNA contamination affect RNA-seq results?

Ribosomal RNA comprises more than 80-90% of total cellular RNA, and inefficient removal means that rRNA reads consume sequencing capacity that would otherwise capture messenger RNA [<a href="#ref-6">6</a>]. High rRNA content reduces gene detection and can bias expression measurements for low-abundance transcripts. Pre-sequencing qPCR assessment can identify rRNA depletion problems before sequencing [<a href="#ref-6">6</a>].

Why is the 5% mitochondrial threshold problematic for some tissues?

The 5% threshold was established in early single-cell studies and does not account for tissue-specific differences in mitochondrial content. Cardiac tissue can have mitochondrial transcripts comprising almost 30% of total mRNA [<a href="#ref-8">8</a>]. A systematic analysis found that the 5% threshold fails to discriminate healthy from low-quality cells in 29.5% of human tissues analyzed [<a href="#ref-7">7</a>].

What is the difference between quality control for bulk and single-cell RNA-seq?

Bulk RNA-seq quality control operates at the sample level, evaluating metrics such as alignment rate, duplication rate, and gene detection across the entire library. Single-cell RNA-seq requires additional cell-level filtering based on detected genes, UMI counts, and mitochondrial content per cell. Single-cell workflows also address ambient RNA and doublets, which are not relevant to bulk RNA-seq [<a href="#ref-14">14</a>].

How should I document quality control decisions for publication?

Record all QC metric values, the thresholds applied, and the rationale for inclusion or exclusion of each sample. Include the versions of all software tools and the parameters used. Report the number of samples excluded and the reasons in the methods section of your publication. This documentation supports reproducibility and allows reviewers to assess the impact of QC decisions.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

[1] [Integration of Bulk RNA-seq Pipeline Metrics for Assessing Low-Quality Samples.](https://pubmed.ncbi.nlm.nih.gov/40630536). Research square, 2025. [2] [Integration of bulk RNA-seq pipeline metrics for assessing low-quality samples.](https://pubmed.ncbi.nlm.nih.gov/41593502). BMC bioinformatics, 2026. [3] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project. [4] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries. [5] [A bioinformatics pipeline for identifying homoplasmic and heteroplasmic mitochondrial DNA SNVs in single-cell RNA-Seq datasets.](https://doi.org/10.1016/j.ygeno.2025.111122). Genomics, 2025. [6] [Pre-sequencing assessment of RNA-Seq library quality using real-time qPCR.](https://pubmed.ncbi.nlm.nih.gov/41293783). BioTechniques, 2025. [7] [Systematic determination of the mitochondrial proportion in human and mice tissues for single-cell RNA-sequencing data quality control.](https://pubmed.ncbi.nlm.nih.gov/32840568). Bioinformatics (Oxford, England), 2021. [8] [Quality control in scRNA-Seq can discriminate pacemaker cells: the mtRNA bias.](https://pubmed.ncbi.nlm.nih.gov/34427691). Cellular and molecular life sciences : CMLS, 2021. [9] [nf-core Documentation](https://nf-co.re/docs). nf-core. [10] [Bioconductor](https://bioconductor.org/). Bioconductor Project. [11] [mirnaQC: a webserver for comparative quality control of miRNA-seq data.](https://pubmed.ncbi.nlm.nih.gov/32484556). Nucleic acids research, 2020. [12] [miQC: An adaptive probabilistic framework for quality control of single-cell RNA-sequencing data.](https://pubmed.ncbi.nlm.nih.gov/34428202). PLoS computational biology, 2021. [13] [ScQCenrich enables multi-metric quality control for single-cell RNA sequencing.](https://pubmed.ncbi.nlm.nih.gov/42342867). Communications biology, 2026. [14] [scReady - an automated and accessible pipeline for single-cell RNA-Seq preprocessing: Empowering novice bioinformaticians](https://doi.org/10.12688/wellcomeopenres.25152.1). Wellcome Open Research, 2026. [15] [Enhancing Single-Cell Transcriptomic Quality in Gastric Cancer: An Intersective Approach Combining Adaptive and Manual Quality Control](https://doi.org/10.1109/IICAI70155.2026.11621043). Indian International Conference on Artificial Intelligence, 2026. [16] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute. [17] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.