The Complete Guide to Quality Control Metrics in Single-Cell RNA-Seq: From UMI Counts to Ambient RNA
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- UMI counts and gene counts are primary indicators of cellular RNA content and library complexity. UMI counts correct for PCR amplification bias, reflecting original mRNA molecules, while gene counts indicate the diversity of detected transcripts. Low UMI counts often signal empty droplets or damaged cells, whereas a low gene count relative to UMI count suggests degraded RNA.
- Mitochondrial fraction serves as a critical proxy for cell viability. Healthy cells maintain a low proportion of mitochondrial transcripts (typically 5-20%), but this fraction increases significantly as cytoplasmic mRNA leaks from compromised cell membranes, indicating cell stress or death.
- Intronic fraction is a key differentiator between whole-cell and single-nucleus RNA-seq (snRNA-seq). High intronic fractions are expected in snRNA-seq due to the capture of unspliced nuclear RNA, whereas low intronic fractions are characteristic of whole-cell preparations enriched for spliced cytoplasmic mRNA.
- Ambient RNA contamination ("soup problem") requires estimation and correction. Cell-free mRNA from lysed cells can be captured by droplets, leading to spurious expression signals, particularly for highly expressed genes. Tools like SoupX are essential for quantifying and mitigating this contamination.
- Doublet detection is crucial for removing barcodes containing multiple cells. These arise from co-encapsulation and create artificial expression profiles that can be mistaken for genuine cell types. Computational methods simulate doublets to train classifiers for their identification.
- QC metric interpretation is highly context-dependent. Thresholds for UMI counts, gene counts, and mitochondrial fraction must be tailored to specific tissue types, dissociation protocols, and sequencing depths, as fixed thresholds can lead to over- or under-filtering of genuine biological signals.
Single-cell RNA sequencing (scRNA-seq) generates gene expression profiles for thousands of individual cells in a single experiment, but the technical artifacts introduced during tissue dissociation, library preparation, and sequencing can obscure genuine biological signals. Quality control (QC) is the process of identifying and removing barcodes that represent empty droplets, damaged cells, doublets, or ambient RNA contamination before downstream analysis. This article explains each core QC metric, the biological rationale behind common filtering thresholds, and how to design a filtering strategy appropriate for your specific dataset and tissue type.
The Purpose and Scope of Quality Control in Single-Cell RNA-Seq
The central challenge in scRNA-seq analysis is distinguishing real cellular heterogeneity from technical noise. Each droplet or well in a single-cell experiment is assigned a barcode, and the sequencing reads associated with that barcode are assumed to originate from a single cell. In practice, this assumption frequently fails. Empty droplets capture cell-free RNA from the surrounding solution, damaged cells release their cytoplasmic contents, and multiple cells can be captured in the same droplet. These artifacts produce barcodes with expression profiles that do not accurately represent any single living cell.
QC metrics are quantitative features computed for each barcode that help classify it as a high-quality cell, a low-quality cell, a doublet, or empty background. The metrics themselves are straightforward to calculate from the count matrix, but interpreting them correctly requires understanding what each metric measures and how tissue type, dissociation method, and sequencing platform influence typical values. A threshold that works well for fresh peripheral blood mononuclear cells may be inappropriate for archived tumor tissue or for single-nucleus data from post-mortem brain samples.
The goal of QC is not to maximize the number of retained cells but to produce a dataset where the remaining cells faithfully represent the biological states present in the original tissue. Overly stringent filtering removes real cell populations, particularly large cells with high mitochondrial content or metabolically active cells with low total RNA. Overly lenient filtering leaves contaminating signals that can create spurious clusters and distort differential expression results. The filtering strategy must therefore be dataset-specific and informed by the biological expectations for the tissue under study.
Core Quality Control Metrics and Their Biological Basis
UMI Counts and Total Transcript Detection
Unique molecular identifiers (UMIs) are short random sequences attached to each cDNA molecule during library preparation. Counting UMIs instead of raw reads corrects for PCR amplification bias, because each unique UMI represents one original mRNA molecule. The number of UMIs per cell is a proxy for the total mRNA content of that cell, which varies substantially across cell types. Large cells such as hepatocytes and neurons contain more mRNA than small cells such as resting lymphocytes.
Low UMI counts can indicate several problems. A barcode with very few UMIs may represent an empty droplet that captured a small amount of ambient RNA, a cell that lysed before capture and lost most of its cytoplasmic mRNA, or a cell that was damaged during dissociation. However, some legitimate cell types naturally have low mRNA content, and quiescent cells in G0 phase transcribe fewer genes than actively dividing cells. The distribution of UMI counts across all barcodes typically shows a clear bimodal pattern, with a small population of low-count barcodes representing empty droplets and a larger population of higher-count barcodes representing real cells. The inflection point between these populations provides a data-driven threshold for separating cells from background.
The relationship between UMI counts and gene counts is informative. In high-quality cells, the number of detected genes increases with UMI count in a predictable saturating curve. Low-quality cells often show a lower gene count than expected for their UMI count, indicating that the captured RNA is fragmented or degraded. The singleCellTK package provides a standardized pipeline for generating these metrics and visualizing their distributions, which helps identify the appropriate thresholds for each dataset.
Gene Counts and Detection Sensitivity
The number of genes detected per cell reflects both the cellular mRNA content and the sequencing depth. A cell with 1,000 detected genes may be a low-quality cell or a deeply sequenced cell with restricted gene expression, such as a mature erythrocyte. The gene count distribution across cells in a dataset should be examined in relation to the UMI count distribution instead of in isolation.
Genes are counted as detected when at least one UMI is assigned to that gene in a given cell. Because sequencing is a sampling process, cells sequenced at greater depth will have more genes detected simply by chance. This technical relationship means that gene count thresholds must be adjusted for the sequencing depth of the experiment. A dataset sequenced to 50,000 reads per cell will have higher gene counts across all cells than a dataset sequenced to 10,000 reads per cell, even if the biological samples are identical.
The popsicleR package integrates gene count and UMI count metrics with normalization and clustering steps, allowing researchers to evaluate how different filtering thresholds affect downstream results. The package was designed to make QC accessible to users who are not expert computational biologists, and it provides interactive visualizations that show the consequences of each filtering decision.
Mitochondrial Fraction and Cell Viability
The proportion of UMIs mapping to mitochondrial genes is one of the most informative QC metrics. In healthy cells, mitochondrial transcripts constitute a small fraction of total mRNA, typically 5 to 20 percent depending on cell type and metabolic activity. When a cell is damaged or dying, cytoplasmic mRNA is lost through leakage across a compromised membrane, but mitochondrial mRNA remains trapped within the mitochondrial organelles. The mitochondrial fraction therefore increases as a cell degrades.
High mitochondrial fraction is a marker of cell stress and death, but the threshold for defining "high" depends on the tissue and dissociation protocol. Cells with high metabolic demand, such as cardiomyocytes and hepatocytes, naturally have higher mitochondrial fractions than lymphocytes. Tissues that require prolonged enzymatic dissociation show higher mitochondrial fractions across all cells because the dissociation process itself stresses cells. The miQC framework models the joint distribution of mitochondrial fraction and gene count using a mixture model, allowing the data to determine which cells are likely low-quality instead of applying a fixed threshold. This approach is particularly valuable for archived tumor tissues, where conventional thresholds remove an excessive proportion of cells.
The valiDrops method extends this concept by using clustering-based approaches to identify barcodes with distinct biological signals after initial filtering on standard metrics. The authors demonstrate that valiDrops can predict and flag dead cells with high accuracy, and that biological signals from cell types and states are more distinct after filtering compared to existing tools. This work highlights that mitochondrial fraction alone does not capture all aspects of cell quality, and that combining multiple metrics with data-driven approaches improves outcomes.
Intronic Fraction and Nuclear RNA Content
The intronic fraction represents the proportion of UMIs mapping to intronic regions of genes. In whole-cell scRNA-seq, most captured RNA is mature cytoplasmic mRNA that has been spliced, so intronic reads are relatively rare. In single-nucleus RNA-seq (snRNA-seq), the captured RNA is predominantly unspliced precursor mRNA still in the nucleus, so intronic fractions are much higher. The expected intronic fraction therefore depends on whether the experiment profiles whole cells or isolated nuclei.
The scQCenrich framework integrates intronic fraction with canonical metrics including mitochondrial RNA, gene counts, and UMIs to provide a multi-metric QC approach for whole-cell scRNA-seq. The framework also incorporates MALAT1 enrichment and dissociation-stress features. MALAT1 is a long non-coding RNA that is highly expressed in the nucleus, so its enrichment can indicate nuclear contamination in whole-cell preparations. The authors report that scQCenrich reduces over-filtering relative to conventional and model-based comparators while preserving coherent cell populations across mouse brain, heart, and lung cancer datasets.
The choice between whole-cell and single-nucleus approaches has direct consequences for QC. A study of post-mortem human brain neocortex developed a preliminary QC pipeline for snRNA-seq and found that approximately 79 percent of samples were high quality based on quantitative laboratory and data metrics. The study identified the proportion of unique or duplicate reads and the proportion of reads remaining after quality trimming as useful features for pass/fail classification. This work demonstrates that sample-level QC metrics, collected before and during sequencing, complement cell-level metrics derived from the count matrix.
Doublet Detection and Multiple Cell Capture
Doublets occur when two or more cells are captured in the same droplet or well and share a single barcode. The resulting expression profile is a mixture of the component cells, which can create artificial cell types that fall between genuine populations in clustering analyses. Doublet rates depend on the loading concentration: higher cell concentrations increase the probability of co-encapsulation. Most droplet-based platforms have doublet rates of 1 to 5 percent at standard loading densities, but the rate increases with the number of cells loaded.
Doublet detection methods work by comparing each barcode's expression profile to the profiles of other cells in the dataset. Computational approaches simulate artificial doublets by combining the expression profiles of random cell pairs, then train classifiers to identify real barcodes that resemble these simulated doublets. The singleCellTK QC pipeline includes doublet prediction as a standard step, integrating it with empty droplet detection and ambient RNA estimation in a single workflow.
Doublet detection is more challenging when the component cells are transcriptionally similar. Two T cells captured together produce a profile that resembles a single T cell with slightly higher gene counts, which is difficult to distinguish from a genuine cell. Doublet detection is most effective for identifying doublets composed of transcriptionally distinct cell types, such as a T cell and a macrophage. The practical consequence is that some doublets will always escape detection, and the filtering strategy should account for this residual contamination.
Ambient RNA and the Soup Problem
Ambient RNA refers to cell-free mRNA present in the input solution that is captured by droplets along with the RNA from intact cells. This contamination originates from cells that lyse during tissue dissociation or during the microfluidic capture process. The released mRNA forms a "soup" of transcripts that reflects the overall gene expression of the tissue, with highly expressed genes contributing more to the soup than lowly expressed genes.
SoupX demonstrated that ambient RNA contamination is ubiquitous in droplet-based scRNA-seq experiments, with experiment-specific variations in composition and magnitude. The authors showed that this contamination confounds biological interpretation and that correcting for it improves quality control metrics and downstream analysis. SoupX estimates the contamination fraction for each cell and produces background-corrected expression profiles that can be used with existing analysis tools.
Ambient RNA affects QC metrics in several ways. Empty droplets contain only ambient RNA, which is why empty droplet detection methods use the ambient profile as a reference. Cells with low mRNA content are most affected by ambient contamination because the contaminating transcripts constitute a larger fraction of their total captured RNA. Highly expressed genes in the tissue, such as hemoglobin in blood or albumin in liver, can appear in all cells regardless of cell type, creating false signals of expression. The scPerturb resource applied uniform QC pipelines across 44 publicly available perturbation datasets, highlighting that consistent handling of ambient RNA is essential for comparing results across experiments.
At a Glance: Quality Control Metrics and Typical Thresholds
| Metric | What It Measures | Typical Threshold Range | Biological Rationale | Caveats |
|---|---|---|---|---|
| UMI count | Total mRNA molecules captured per cell | Remove barcodes below the inflection point in the UMI distribution | Empty droplets and lysed cells capture few transcripts | Large cells and highly active cells naturally have more UMIs |
| Gene count | Number of genes with at least one UMI | Remove cells below the knee of the gene-UMI curve | Degraded RNA produces fewer detected genes per UMI | Sequencing depth directly affects gene detection |
| Mitochondrial fraction | Proportion of UMIs mapping to mitochondrial genes | 5 to 20 percent for whole cells, higher for stressed tissues | Damaged cells lose cytoplasmic mRNA but retain mitochondrial mRNA | Metabolically active cells have higher baseline fractions |
| Intronic fraction | Proportion of UMIs mapping to intronic regions | Low for whole cells, high for single nuclei | Nuclear RNA is unspliced precursor mRNA | Platform and protocol determine expected values |
| Doublet score | Probability that a barcode contains multiple cells | Remove the top 1 to 5 percent of predicted doublets | Multiple cells share one barcode, creating artificial profiles | Doublets of similar cell types are hard to detect |
| Ambient RNA fraction | Proportion of captured RNA originating from cell-free mRNA | Correct if contamination exceeds a few percent | Lysed cells release mRNA into the input solution | Highly expressed genes dominate the ambient profile |
Designing a Filtering Strategy for Your Dataset
Step 1: Generate and Visualize the Core Metrics
Before applying any thresholds, generate the full set of QC metrics for every barcode in the dataset. The singleCellTK pipeline provides a standardized workflow that imports data from multiple platforms and preprocessing tools, generates standard QC metrics, and produces visualizations. The Galaxy Training Network offers accessible tutorials for running these analyses without extensive programming experience, and the Bioconductor project hosts the underlying R packages with documentation for reproducible analysis.
Plot the distributions of UMI counts, gene counts, and mitochondrial fraction as histograms and scatter plots. Examine the relationship between UMI counts and gene counts, and between UMI counts and mitochondrial fraction. These visualizations reveal the structure of the data and help identify where the cell population separates from the background.
Step 2: Identify Empty Droplets
Empty droplets are barcodes that captured ambient RNA but no cell. These barcodes have low UMI counts and expression profiles that match the ambient soup. The standard approach for droplet-based data is to compare each barcode's profile to the ambient profile and remove barcodes that are not significantly different from background. The inflection point in the UMI count distribution provides a simple initial threshold, but more sophisticated methods use the full expression profile to distinguish empty droplets from cells with low mRNA content.
The valiDrops method uses data-adaptive thresholding on community-standard quality metrics followed by a clustering-based approach to identify barcodes with distinct biological signals. This two-stage approach reduces the risk of removing genuine cells that happen to have low UMI counts while still eliminating empty droplets.
Step 3: Apply Cell-Level Filters
After removing empty droplets, apply filters for gene count, mitochondrial fraction, and other metrics. The thresholds should be informed by the distributions observed in step 1 instead of applied as fixed values. The miQC framework provides a principled approach by jointly modeling mitochondrial fraction and gene count with mixture models, allowing the data to determine which cells are likely low-quality.
For whole-cell data, consider the intronic fraction and MALAT1 enrichment as additional metrics, as recommended by the scQCenrich framework. These metrics can identify cells with nuclear contamination or dissociation stress that would pass conventional filters.
Step 4: Detect and Remove Doublets
Run doublet detection on the filtered dataset. The singleCellTK pipeline includes doublet prediction as a standard step. Examine the distribution of doublet scores and remove the highest-scoring barcodes. The expected doublet rate depends on the loading concentration, so the number of doublets removed should be consistent with the experimental design.
Step 5: Estimate and Correct Ambient RNA
Run ambient RNA estimation using a tool such as SoupX to quantify the contamination fraction and generate corrected expression profiles. Compare the results of downstream analyses with and without ambient RNA correction to determine whether correction changes biological conclusions.
Step 6: Evaluate the Impact of Filtering on Downstream Results
QC is not complete until the filtered data have been carried through clustering and cell type annotation. The popsicleR package integrates QC with normalization, clustering, and annotation, allowing users to see how filtering decisions affect the final cell types identified. The scQCEA framework provides automated cell type annotation using differential gene expression patterns and a repository of reference marker genes, enabling expression-based QC that discriminates between true variation and background noise.
If filtering removes an entire expected cell type, the thresholds are likely too stringent. If clustering reveals populations that cannot be annotated to any known cell type, the thresholds may be too lenient. The filtering strategy should be adjusted iteratively until the retained cells produce biologically interpretable clusters.
Options and Tradeoffs in QC Approaches
Fixed Thresholds versus Data-Driven Methods
Fixed thresholds, such as removing all cells with mitochondrial fraction above 10 percent or gene count below 500, are simple to implement and easy to report. However, they assume that all cells in all tissues have similar QC metric distributions, which is not the case. The miQC framework demonstrates that fixed thresholds are often overly stringent, particularly for lower-quality tissues such as archived tumor samples. Data-driven methods that model the actual distributions in each dataset preserve more high-quality cells while still removing low-quality ones.
The tradeoff is complexity and reproducibility. Fixed thresholds are transparent and easy to document. Data-driven methods require more sophisticated statistical modeling and may produce different results when applied to the same data with different random seeds. The choice depends on the research question and the need for comparability across datasets.
Whole-Cell versus Single-Nucleus QC
Single-nucleus RNA-seq requires different QC thresholds than whole-cell RNA-seq because the RNA content and composition differ fundamentally. Nuclei contain predominantly unspliced precursor mRNA, so intronic fractions are high and cytoplasmic transcripts are depleted. Mitochondrial RNA is largely absent from nuclei, so mitochondrial fraction is not a useful viability metric for snRNA-seq. The study of post-mortem human brain neocortex developed a QC framework specifically for snRNA-seq, using quantitative laboratory and data metrics as features for classification models.
The choice between whole-cell and single-nucleus approaches affects also QC thresholds but also the biological questions that can be addressed. Whole-cell data capture the full transcriptome including cytoplasmic mRNA, while nuclear data reflect the nuclear transcriptome and may be more representative of the cell's regulatory state. The study of flexor digitorum brevis mouse myofibers demonstrated that whole-cell flow sorting can generate deep expression patterns from intact myofibers, with quality control metrics indicating that only a subset of sorted cells met optimal quality standards.
Dissociation Stress and Its Effects on QC Metrics
The method of tissue dissociation has a profound effect on QC metrics. A study comparing cold active protease and collagenase dissociation found that collagenase digestion at 37 degrees Celsius induces a stress response characterized by a core set of 512 heat shock and stress response genes, including FOS and JUN. This stress response was conserved across all cell types but showed cell type-specific patterns in patient tissues. Dissociation with a cold active protease at 6 degrees Celsius minimized this stress response.
The practical implication is that QC metrics reflecting cell stress, particularly mitochondrial fraction and stress gene expression, will be elevated in datasets generated with collagenase dissociation. Comparing QC metrics across datasets generated with different dissociation protocols requires caution, and the scQCenrich framework includes dissociation-stress features as part of its multi-metric QC approach.
Records and Measurements for QC Documentation
What to Record for Each Dataset
Reproducible QC requires documenting every filtering decision and its rationale. For each dataset, record the following information:
The sequencing platform and library preparation kit, including the version and any protocol modifications. The cell loading concentration and expected doublet rate. The number of barcodes before and after each filtering step. The thresholds applied for each metric and whether they were fixed or data-driven. The number and percentage of cells removed at each step. The ambient RNA contamination fraction estimated by tools such as SoupX. The doublet detection method and the number of doublets removed. The final number of cells retained and the distribution of QC metrics in the retained population.
Using QC Reports for Quality Assurance
The scQCEA framework generates interactive reports of QC metrics for comparing sets of samples and visual evaluation of quality scores. These reports allow researchers to identify batches or samples with unusual QC profiles before downstream analysis. The framework includes a repository of 2348 marker genes for 95 human and mouse cell types, enabling expression-based QC that verifies expected cell types are present in each sample.
The nf-core documentation describes community standards for reproducible bioinformatics pipelines, including requirements for QC reporting. Following these standards ensures that QC decisions are transparent and that results can be compared across studies. The EMBL-EBI training resources provide guidance on data management and analysis best practices that support reproducible QC.
Integrating Laboratory Metrics with Computational Metrics
The preliminary QC framework for snRNA-seq demonstrated that laboratory metrics collected during sample preparation and sequencing can be combined with computational metrics from the count matrix to improve QC classification. These laboratory metrics include RNA integrity numbers, library concentration, sequencing depth, and the proportion of reads mapping to the reference genome. Recording these metrics alongside the computational QC metrics provides a more complete picture of data quality and can help diagnose the source of QC failures.
Common Failure Patterns in QC
Over-Filtering and Loss of Real Cell Populations
The most common QC failure is removing genuine cells because they fall outside thresholds that were appropriate for a different tissue or protocol. Large cells with high mitochondrial content, such as hepatocytes and cardiomyocytes, are frequently removed by standard mitochondrial fraction thresholds. The miQC framework was developed specifically to address this problem, and the authors demonstrate that data-driven approaches preserve high-quality cells that would be removed by fixed thresholds.
Signs of over-filtering include the absence of expected cell types in the final clustering results, unusually low cell counts relative to the number of barcodes detected, and the loss of rare cell populations that were present in the unfiltered data. The scQCenrich framework reports reduced over-filtering relative to conventional comparators while preserving coherent cell populations, suggesting that multi-metric approaches are less prone to this failure mode.
Under-Filtering and Retention of Contaminating Signals
The opposite failure is retaining too many low-quality barcodes, which introduces noise that obscures genuine biological variation. Low-quality cells cluster together based on their shared technical artifacts instead of their biological identity, creating artificial populations that are difficult to interpret. Ambient RNA contamination can cause all cells to appear to express highly expressed genes from the tissue, masking true cell type differences.
Signs of under-filtering include clusters that cannot be annotated to any known cell type, high similarity between all clusters driven by ambient genes, and differential expression results that are dominated by stress genes or mitochondrial genes. The SoupX study demonstrated that ambient RNA contamination can produce misleading biological interpretations, and that correction improves the utility of existing and future datasets.
Inconsistent Thresholds Across Batches
When multiple samples or batches are processed together, applying different thresholds to each batch can introduce batch effects that are confounded with biological differences. The scPerturb resource applied uniform QC pipelines across 44 perturbation datasets to enable comparison and integration across experiments. Consistent QC is essential for any multi-sample study, and the thresholds should be determined using the pooled data or using a method that adapts to each batch while maintaining comparable stringency.
Ignoring the Ambient RNA Profile
Many QC pipelines remove empty droplets and low-quality cells but do not correct for ambient RNA contamination in the retained cells. This omission is particularly problematic for genes that are highly expressed in the tissue, because ambient contamination can create false expression signals in all cells. The SoupX method provides a practical solution by estimating the contamination fraction and producing corrected profiles, and the authors recommend applying this correction before downstream analysis.
Limitations of QC Metrics and Interpretation Caveats
QC Metrics Are Correlated and Redundant
The core QC metrics are not independent. UMI count, gene count, and mitochondrial fraction are correlated because they all reflect the amount and quality of captured RNA. A cell with low UMI count will generally have low gene count and may have high mitochondrial fraction if the low UMI count results from cell damage. This correlation means that applying multiple thresholds sequentially can remove more cells than any single threshold would, and the combined effect is difficult to predict without examining the joint distributions.
The miQC framework addresses this issue by jointly modeling mitochondrial fraction and gene count, and the scQCenrich framework extends this approach to additional metrics. Joint modeling avoids the cumulative over-filtering that occurs when thresholds are applied independently.
Thresholds Are Tissue-Specific and Protocol-Specific
There are no universal QC thresholds that work for all datasets. The appropriate thresholds depend on the tissue type, the dissociation protocol, the sequencing platform, and the sequencing depth. The study of flexor digitorum brevis myofibers found that only 171 of 763 sorted myofibers met optimal quality criteria, with a median read count of 239,252 and an average of 12,098 transcripts per cell. These values reflect the specific challenges of profiling large, multinucleated muscle cells and would not be appropriate for other tissues.
The dissociation study demonstrated that the method of tissue dissociation influences cell yield and transcriptome state in a tissue- and cell-type-dependent manner. QC thresholds developed for one dissociation protocol may not transfer to another, and the stress response genes identified in this study can help identify dissociation-induced artifacts in scRNA-seq experiments.
QC Cannot Fully Correct for Technical Artifacts
QC filtering removes the most obviously problematic barcodes, but it cannot eliminate all technical artifacts. Doublets of transcriptionally similar cells escape detection. Ambient RNA correction estimates the contamination fraction but cannot perfectly reconstruct the true cellular expression profiles. Dissociation stress alters gene expression in ways that cannot be fully reversed by filtering or correction.
The valiDrops study notes that isolation of cells or nuclei for sxRNA-seq releases contaminating RNA through cell damage and transcript leakage, and that identifying high-quality barcodes is a critical analytical step. The authors acknowledge that their method improves data quality but does not eliminate all artifacts. Researchers should interpret results with awareness of the residual technical noise that remains after QC.
Safety and Regulatory Context for QC Decisions
Data Reproducibility and Reporting Standards
QC decisions affect every downstream result, so they must be documented and reported transparently. The nf-core documentation describes community standards for reproducible pipelines, including requirements for version control, containerization, and reporting. Following these standards allows other researchers to understand exactly how QC was performed and to reproduce the results.
The Bioconductor project provides official documentation for the R packages used in QC analysis, including installation instructions and workflow examples. The Galaxy Training Network offers accessible tutorials that teach the practical skills needed to run QC analyses. The Carpentries lessons provide foundational training in the computing skills, including shell, Git, and programming, that support reproducible analysis.
Professional Escalation Criteria
Some QC findings indicate problems that require consultation with a bioinformatics specialist or the sequencing facility. Escalate when the proportion of barcodes passing QC is unexpectedly low, such as below 30 percent of detected barcodes, because this may indicate a problem with the library preparation or sequencing run. Escalate when QC metrics differ dramatically between batches that were expected to be similar, because this may indicate a technical issue that affects data comparability. Escalate when the ambient RNA contamination fraction is very high, because this may indicate excessive cell lysis during dissociation that requires protocol optimization. Escalate when expected cell types are completely absent from the filtered data, because this may indicate that the filtering thresholds are inappropriate or that the tissue was not adequately dissociated.
The EMBL-EBI training resources provide guidance on when to seek expert assistance and how to communicate QC findings to collaborators. The NCBI data resources offer access to public datasets and tools that can help benchmark QC decisions against established standards.
Frequently Asked Questions
What is the difference between UMI counts and gene counts in QC?
UMI counts represent the total number of unique mRNA molecules captured from a cell, while gene counts represent the number of distinct genes with at least one UMI. UMI counts reflect the total mRNA content of the cell, which varies by cell type and size. Gene counts reflect the diversity of gene expression and are influenced by both the mRNA content and the sequencing depth. A cell with high UMI counts but low gene counts may be a specialized cell expressing a restricted set of genes, while a cell with low UMI counts and proportionally low gene counts may be damaged or empty.
Why is mitochondrial fraction used as a QC metric?
Mitochondrial fraction increases in damaged cells because cytoplasmic mRNA leaks out through a compromised cell membrane while mitochondrial mRNA remains trapped inside the mitochondria. High mitochondrial fraction therefore indicates cell stress or death. However, the baseline mitochondrial fraction varies by cell type and tissue, with metabolically active cells such as cardiomyocytes and hepatocytes having higher fractions than resting lymphocytes. The miQC framework models the joint distribution of mitochondrial fraction and gene count to identify low-quality cells without applying a fixed threshold.
How do I choose the right QC thresholds for my dataset?
Start by visualizing the distributions of UMI counts, gene counts, and mitochondrial fraction for all barcodes. Identify the inflection points where the cell population separates from the background. Use data-driven methods such as miQC or valiDrops that model the actual distributions instead of applying fixed thresholds. After filtering, check that expected cell types are present in the clustering results and adjust thresholds if entire populations are missing.
What is ambient RNA and why does it matter?
Ambient RNA is cell-free mRNA present in the input solution that is captured by droplets along with RNA from intact cells. It originates from cells that lyse during tissue dissociation or microfluidic capture. SoupX demonstrated that ambient contamination is ubiquitous and experiment-specific, and that it confounds biological interpretation. Highly expressed genes in the tissue appear in all cells due to ambient contamination, creating false expression signals. Ambient RNA correction should be applied before downstream analysis.
Should I use the same QC thresholds for single-nucleus and whole-cell data?
No. Single-nucleus RNA-seq captures predominantly unspliced precursor mRNA, so intronic fractions are high and cytoplasmic transcripts are depleted. Mitochondrial RNA is largely absent from nuclei, so mitochondrial fraction is not a useful viability metric. The preliminary QC framework for snRNA-seq used laboratory and data metrics specific to nuclear preparations. The scQCenrich framework includes intronic fraction and MALAT1 enrichment as QC metrics for whole-cell data, reflecting the different RNA composition of whole cells versus nuclei.
How many doublets should I expect and remove?
The doublet rate depends on the cell loading concentration. Most droplet-based platforms have doublet rates of 1 to 5 percent at standard loading densities, with higher rates at higher concentrations. Doublet detection methods such as those included in the singleCellTK pipeline assign a doublet score to each barcode, and the highest-scoring barcodes are removed. The number removed should be consistent with the expected doublet rate for the loading concentration used.
Can QC remove real biological variation?
Yes. Overly stringent QC thresholds can remove genuine cell populations, particularly large cells with high mitochondrial content, cells with low mRNA content, and rare cell types. The miQC study demonstrated that fixed thresholds are often overly stringent for lower-quality tissues. The scQCenrich study reported reduced over-filtering relative to conventional comparators. QC should be evaluated by checking that expected cell types are present in the filtered data.
What should I do if my QC metrics look unusual?
First, verify that the metrics were calculated correctly and that the data were imported properly. Check whether the unusual metrics are consistent across all samples or confined to specific batches. If confined to specific batches, investigate whether those batches had different dissociation protocols, sequencing depths, or other technical differences. If the metrics are unusual across all samples, consider whether the tissue type or dissociation method explains the pattern. The dissociation study showed that collagenase digestion induces a conserved stress response, and the scQCenrich framework includes dissociation-stress features in its QC approach. If the metrics cannot be explained by known factors, consult a bioinformatics specialist or the sequencing facility.
Related Bioinformatics Guides
- Single-Cell RNA Sequencing Quality Control: A Practical Guide to Filtering and Metrics
- RNA-Seq Quality Control: Essential Checks and Tools
- Single-Cell RNA Sequencing Depth: A Cost-Benefit Analysis for Experimental Design
- Single-Cell Sequencing Depth: How Much Is Enough?
- Single-Cell Sequencing Analysis Pipeline: From Raw Data to Biological Insights
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- Automatic quality control of single-cell and single-nucleus RNA-seq using valiDrops.. NAR genomics and bioinformatics, 2023.
- scPerturb: harmonized single-cell perturbation data.. Nature methods, 2024.
- SoupX removes ambient RNA contamination from droplet-based single-cell RNA sequencing data.. GigaScience, 2020.
- Normalization of Single-Cell RNA-Seq Data.. Methods in molecular biology (Clifton, N.J.), 2021.
- popsicleR: A R Package for Pre-processing and Quality Control Analysis of Single Cell RNA-seq Data.. Journal of molecular biology, 2022.
- Comprehensive generation, visualization, and reporting of quality control metrics for single-cell RNA sequencing data.. Nature communications, 2022.
- Single cell RNA-seq analysis of the flexor digitorum brevis mouse myofibers.. Skeletal muscle, 2021.
- PRODUCTION OF A PRELIMINARY QUALITY CONTROL PIPELINE FOR SINGLE NUCLEI RNA-SEQ AND ITS APPLICATION IN THE ANALYSIS OF CELL TYPE DIVERSITY OF POST-MORTEM HUMAN BRAIN NEOCORTEX.. Pacific Symposium on Biocomputing. Pacific Symposium on Biocomputing, 2017.
- ScQCenrich enables multi-metric quality control for single-cell RNA sequencing.. 2026.
- Editorial: AI in single-cell biology.. 2026.
- LAIOR: a hyperbolic neural ODE variational framework for interpretable single-cell manifold learning and trajectory inference.. 2026.
- HistoMap: Reconstructing Spatially Resolved Single-Cell Profiles from Bulk RNA-Seq to Decipher the Immune-Excluded Microenvironment in Colon Cancer. 2026.
- Cross-species transcriptomic integration reveals a MIRO1-mediated macrophage-T cell axis in glioma.. 2026.
- Dissociation of solid tumor tissues with cold active protease for single-cell RNA-seq minimizes conserved collagenase-associated stress responses. Genome Biology, 2019.
- miQC: An adaptive probabilistic framework for quality control of single-cell RNA-sequencing data. bioRxiv, 2021.
- scQCEA: a framework for annotation and quality control report of single-cell RNA-sequencing data. BMC Genomics, 2023.
- dupRadar: A Bioconductor package for the assessment of PCR artifacts in RNA-Seq data. BMC Bioinformatics, 2016.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.