How to Choose the Right QC Metrics for Single-Cell RNA-Seq Data from Different Tissues and Conditions

By Dr. Zubair Khalid, DVM, MS, PhD ·

How to Choose the Right QC Metrics for Single-Cell RNA-Seq Data from Different Tissues and Conditions

Key Takeaways

  • Universal QC thresholds for single-cell RNA-seq (scRNA-seq) and single-nucleus RNA-seq (snRNA-seq) are inappropriate due to significant biological variation in mitochondrial read fraction, UMI counts, and gene detection across different tissues and cell types. For instance, cardiomyocytes inherently possess high mitochondrial content, necessitating higher thresholds than typically applied to other tissues.
  • Tissue dissociation protocols profoundly influence QC metrics by inducing transcriptional stress responses, such as the upregulation of immediate early genes. Cold-active proteases may minimize these stress responses compared to standard warm enzymatic digestion, directly impacting gene counts and stress gene expression, thus requiring protocol-specific threshold adjustments.
  • The distinction between scRNA-seq and snRNA-seq necessitates different QC strategies; nuclei preparations generally exhibit lower cytoplasmic transcript and mitochondrial read fractions, leading to lower gene detection per nucleus compared to whole-cell preparations from the same tissue.
  • Core QC metrics like UMI count and gene count are influenced by cell size and metabolic activity; quiescent cell populations (e.g., resting T cells) naturally have lower UMI counts than metabolically active cells (e.g., hepatocytes), requiring thresholds that accommodate this biological heterogeneity.
  • Ambient RNA contamination, arising from lysed cells during dissociation, can inflate gene counts and obscure true cell-type signatures, particularly in tissues with robust extracellular matrices like heart and skeletal muscle, necessitating specific bioinformatic correction methods like CellBender or SoupX.
  • Iterative QC threshold selection, informed by initial distribution analysis, expected cell composition, and evaluation of cell type recovery, is critical. Documenting the rationale for each threshold decision in a structured log ensures reproducibility and facilitates auditing of downstream biological conclusions.

Single-cell RNA sequencing (scRNA-seq) and single-nucleus RNA sequencing (snRNA-seq) generate transcriptomic profiles from individual cells or nuclei, but the quality metrics that define a usable library differ substantially across tissue types, dissociation protocols, and experimental conditions. A universal threshold for metrics such as mitochondrial read fraction, gene count, or unique molecular identifier (UMI) count will discard biologically meaningful populations in one tissue while retaining damaged or low-quality cells in another. This article provides a decision framework for selecting and tuning quality control (QC) metrics based on tissue type, dissociation protocol, and expected cell composition, with examples from published studies across muscle, brain, heart, lung, liver, adipose tissue, kidney, and tumor samples.

Why QC Thresholds Cannot Be Universal

The core problem in scRNA-seq QC is that the metrics used to identify low-quality cells are confounded with biological variation. Mitochondrial read fraction, for example, is commonly used as a proxy for cell stress or membrane damage because high mitochondrial content with low gene detection often indicates a dying or lysed cell. However, some cell types naturally contain high mitochondrial transcript proportions. Cardiomyocytes are densely packed with mitochondria to support continuous contraction, and their mitochondrial read fraction in single-cell or single-nucleus libraries can exceed levels that would trigger removal in other tissues. Similarly, hepatocytes in the liver carry high metabolic load and may show elevated mitochondrial transcripts. Applying a blanket threshold of 10% or 20% mitochondrial reads, values commonly cited in general tutorials, would eliminate a large fraction of these biologically relevant cells.

Tissue dissociation itself introduces stress responses that alter QC metrics. Enzymatic digestion of solid tissues activates immediate early genes and stress-related transcriptional programs. A study comparing dissociation methods for solid tumors found that cold-active protease treatment minimized conserved collagenase-associated stress responses compared with standard warm enzymatic digestion [<a href="#ref-1">1</a>]. This finding demonstrates that the choice of dissociation enzyme and temperature directly affects the transcriptional state captured in the library, which in turn influences QC metrics such as total gene counts and the expression of stress genes. Researchers working with tumor tissue must therefore account for dissociation-induced stress when setting thresholds, because a stressed but viable cell may look similar to a dying cell by standard metrics.

The distinction between scRNA-seq and snRNA-seq adds another layer of complexity. Single-nucleus approaches are often used for tissues that are difficult to dissociate, such as adult brain, heart, and adipose tissue. Nuclei preparations generally contain fewer cytoplasmic transcripts than whole cells, so gene detection per nucleus is lower. Mitochondrial transcripts are also less abundant in nuclear preparations because mitochondria reside in the cytoplasm. A protocol for isolating nuclei from murine cardiac tissue for single-nucleus multiomic sequencing describes mechanical homogenization, sequential filtration, sucrose cushion purification, and fluorescence-activated nuclei sorting to obtain high-quality nuclei from fresh-frozen tissue [<a href="#ref-2">2</a>]. The QC metrics appropriate for these nuclear libraries differ from those used for whole-cell libraries from the same organ.

Core QC Metrics and Their Biological Meaning

UMI Count per Cell or Nucleus

The number of unique molecular identifiers per cell reflects the amount of mRNA captured and sequenced from that cell. Low UMI counts can indicate failed capture, sequencing dropout, or a cell that was damaged during dissociation. However, UMI counts vary by cell type and tissue. Small quiescent cells such as resting T cells or naive B cells naturally contain less mRNA than large metabolically active cells such as hepatocytes or plasma cells. In lung tissue, a study analyzing normal human lung identified 20,154 high-quality cells that clustered into seven major cell types including macrophages, monocytes, CD8+ T cells, epithelial cells, endothelial cells, adipocytes, and NK cells [<a href="#ref-3">3</a>]. The epithelial cells were further divided into seven subtypes with distinct developmental trajectories. Within this dataset, the appropriate UMI threshold had to accommodate both the abundant alveolar macrophages and the rarer epithelial progenitor populations, which differ substantially in transcriptional output.

Gene Count per Cell or Nucleus

The number of detected genes per cell correlates with sequencing depth and cell complexity. Cells with very low gene counts are often empty droplets or damaged cells, but some quiescent populations legitimately express fewer genes. In skeletal muscle, a single-cell study of sarcopenia used a senescence-accelerated mouse strain and retained 13,612 tibialis anterior muscle single-cell transcriptomes after QC [<a href="#ref-4">4</a>]. Muscle tissue contains multinucleated myofibers that are difficult to capture as intact single cells, so the recovered cells were predominantly mononuclear populations such as endothelial cells, immune cells, and fibro-adipogenic progenitors. The gene count thresholds for these populations differ from those for hepatocytes or neurons, which express larger numbers of genes per cell.

Mitochondrial Read Fraction

Mitochondrial read fraction serves as a stress and damage indicator in whole-cell scRNA-seq. High mitochondrial content with low cytoplasmic gene detection suggests that cytoplasmic mRNA leaked from a compromised cell membrane while mitochondrial transcripts remained trapped. This metric is less informative for snRNA-seq because nuclei contain few mitochondrial transcripts. For heart tissue, where cardiomyocytes are inherently mitochondrial-rich, the mitochondrial threshold must be set higher or replaced with alternative metrics. A single-cell study of the healthy and diseased adult mouse heart optimized RNA amplification from digested cardiac tissue and identified cytoskeleton-associated protein 4 as a novel marker for activated fibroblasts [<a href="#ref-5">5</a>]. The authors had to account for the high mitochondrial content of cardiomyocytes when setting QC thresholds, because standard cutoffs would have removed these essential cells from the analysis.

Doublet Rate

Doublets are droplets or wells that contain two or more cells or nuclei. They generate libraries with high UMI and gene counts that can appear as distinct clusters between their component cell types. Doublet rates scale with loading density, so higher cell loads produce more doublets. Computational doublet detection methods can flag potential doublets, but they are imperfect and may remove legitimate large cells or dividing cells. In a study of childhood-onset lupus nephritis, researchers used spatial transcriptomics to generate a single-cell resolution atlas of kidney tissue from eight patients and four controls, assigning annotated cells to 30 reference cell types [<a href="#ref-6">6</a>]. The authors had to manage doublets carefully because the kidney contains both large resident cells and infiltrating immune cells, and a doublet between a proximal tubule cell and a macrophage could be misread as a novel cell type.

Ambient RNA Contamination

Ambient RNA consists of transcripts released from lysed cells during dissociation that contaminate the capture solution. This contamination appears as low-level expression of many genes across all cells, which can obscure true cell-type signatures and inflate gene counts. Tissues with tough extracellular matrices, such as heart and skeletal muscle, often produce more ambient RNA because aggressive dissociation is required to release cells. A protocol for minimized-bias profiling of liver and visceral adipose tissue in mice using integrated snRNA-seq describes tissue-specific nuclei extraction and selective enrichment to isolate high-quality nuclei while separating whole-liver and non-parenchymal cell workflows to reduce dissociation bias [<a href="#ref-7">7</a>]. The authors explicitly designed their protocol to minimize ambient contamination from damaged hepatocytes, which are abundant and easily lysed during liver processing.

Tissue-Specific QC Considerations

Brain and Neural Tissue

Brain tissue is commonly analyzed by snRNA-seq because adult neurons are fragile and have extensive processes that are damaged during enzymatic dissociation. A study integrating snRNA-seq and spatial transcriptomics in a mouse model of lipopolysaccharide-induced sepsis-associated encephalopathy collected brain tissues at 0, 12, 24, and 72 hours after injection and identified specialized subpopulations of astrocytes, microglia, and vascular cells [<a href="#ref-8">8</a>]. The authors used single-nucleus approaches to capture these populations, and their QC workflow had to accommodate the low cytoplasmic content of nuclei. For brain snRNA-seq, gene count thresholds are typically lower than for whole-cell scRNA-seq, and mitochondrial read fraction is less useful as a quality metric because nuclei contain few mitochondrial transcripts.

Guidelines for bioinformatics of single-cell sequencing data analysis in Alzheimer's disease systematically reviewed 14 major analytic directions, including quality control and normalization, and applied their recommended workflow to a large snRNA-seq dataset in Alzheimer's disease [<a href="#ref-9">9</a>]. The authors addressed challenges in using human postmortem tissue, where RNA quality varies with postmortem interval and tissue preservation. Postmortem brain tissue often shows elevated ambient RNA and reduced gene detection compared with fresh tissue, so QC thresholds must be adjusted downward to retain genuine neuronal populations.

Heart and Skeletal Muscle

Cardiac and skeletal muscle tissues present unique QC challenges because of their cellular composition. Cardiomyocytes are large, multinucleated, and mitochondrial-rich, making them difficult to capture as intact single cells. Most cardiac studies therefore use snRNA-seq. A study of human heart failure with preserved ejection fraction analyzed single-nucleus RNA sequencing in myocardial biopsies from 19 patients and 24 nonfailing controls, using genotype-based demultiplexing and recovering 48,886 nuclei after QC [<a href="#ref-10">10</a>]. The authors identified 14 cell types and found thousands of differentially expressed genes across cell types. Their QC workflow had to retain cardiomyocyte nuclei, which have lower gene detection than other cardiac cell types because of their large size and the nuclear capture method.

A protocol for isolating nuclei from murine cardiac tissue for single-nucleus multiomic sequencing describes mechanical homogenization, sequential filtration, sucrose cushion purification, and fluorescence-activated nuclei sorting [<a href="#ref-2">2</a>]. This protocol enables integrated analysis of gene expression and chromatin accessibility across cardiac cell types. The authors note that cardiac tissue requires more rigorous purification than softer tissues because of the abundance of contractile proteins and connective tissue, which can clog capture devices and increase ambient RNA.

For skeletal muscle, the sarcopenia study used a senescence-accelerated mouse strain and identified 14 cell clusters from tibialis anterior muscle, with two distinct endothelial subtypes dominant in the sarcopenia group and control groups [<a href="#ref-4">4</a>]. Muscle dissociation releases myofiber fragments and satellite cells, and the QC thresholds must distinguish between genuine mononuclear cells and damaged myofiber remnants. The authors validated their findings with immunofluorescence staining and western blotting, demonstrating that QC decisions directly affect downstream biological conclusions.

Tumor Tissue

Tumor tissue presents additional QC complications because of its heterogeneity, variable cellularity, and the presence of dead or dying cells within the tumor microenvironment. A study of ovarian cancer used scRNA-seq to characterize the tumor microenvironment in samples from immunotherapy cohorts, analyzing data from TCGA, GTEx, GEO, and spatial transcriptome databases [<a href="#ref-11">11</a>]. The authors examined the ferroptosis-related gene HMOX1 and found decreased expression in ovarian cancer epithelial cells but upregulation in macrophages. Their QC workflow had to accommodate the wide range of cell sizes and transcriptional activities present in tumor samples, from small immune cells to large malignant epithelial cells.

The dissociation method for solid tumors directly affects QC metrics. The cold-active protease study demonstrated that standard collagenase digestion induces conserved stress responses that alter the transcriptional state of captured cells [<a href="#ref-1">1</a>]. These stress responses can inflate mitochondrial read fraction and reduce gene detection, making stressed but viable tumor cells appear as low-quality cells by standard thresholds. Researchers working with tumor tissue should therefore consider the dissociation protocol when interpreting QC metrics and may need to use higher mitochondrial thresholds or alternative stress markers.

Lung Tissue

Lung tissue contains a diverse mix of epithelial, endothelial, immune, and stromal cells, and its QC thresholds must accommodate this heterogeneity. A study of lung tissue in a rat model of acute respiratory distress syndrome presented a standardized protocol for scRNA-seq analysis, including optimized tissue dissociation, rigorous quality control, doublet removal, batch effect correction, and comprehensive bioinformatic pipelines [<a href="#ref-12">12</a>]. The approach reliably identified 21 distinct cell populations and revealed six macrophage and monocyte subpopulations with unique transcriptional signatures. The authors emphasized that rigorous QC was essential for identifying these rare subpopulations, because low-quality cells could obscure genuine biological variation.

A study reconstructing the developmental trajectories of pulmonary parenchymal epithelial cells used the Seurat R package for data quality control and identified 20,154 high-quality cells from normal human lung tissue [<a href="#ref-3">3</a>]. The authors initially divided the cells into 17 clusters and identified seven cell types, then extracted 4,240 epithelial cells for further analysis. The epithelial cells were divided into seven clusters including alveolar cells, alveolar endothelial progenitors, ciliated cells, secretory cells, and ionocytes. The QC thresholds for this dataset had to retain both the abundant alveolar macrophages and the rarer epithelial progenitor populations, which differ substantially in transcriptional output.

Liver and Adipose Tissue

Liver and adipose tissue are challenging for scRNA-seq because of their lipid content and the fragility of certain cell populations. Hepatocytes are large, metabolically active cells that are easily damaged during dissociation, releasing high levels of ambient RNA. Adipocytes are lipid-laden and buoyant, making them difficult to capture in standard microfluidic devices. The protocol for minimized-bias profiling of liver and visceral adipose tissue in mice using integrated snRNA-seq describes tissue-specific nuclei extraction and selective enrichment to isolate high-quality nuclei from fresh mouse liver and visceral adipose tissue [<a href="#ref-7">7</a>]. The authors separated whole-liver and non-parenchymal cell workflows to reduce dissociation bias, and they used magnetic enrichment and BD Rhapsody capture followed by cDNA synthesis, whole-transcriptome amplification, and index PCR.

A study of adipose tissue-derived mesenchymal stem cells under chondrogenic induction used droplet-based scRNA-seq and analyzed 37,219 high-quality transcripts from control cells and cells induced for 1 and 2 weeks [<a href="#ref-13">13</a>]. The authors identified four distinct cell clusters with varying proportions across conditions, and gene ontology enrichment analysis revealed cluster-specific variations in biological processes related to ribosome biogenesis, mitochondrial oxidative metabolism, cell proliferation, and collagen fibril organization. The QC thresholds for this dataset had to accommodate the transcriptional changes induced by chondrogenic differentiation, which altered gene counts and mitochondrial content across clusters.

Kidney Tissue

Kidney tissue contains a complex mix of epithelial, endothelial, immune, and stromal cells, and its QC thresholds must accommodate the large proximal tubule cells, which have high transcriptional output, alongside smaller immune cells. The childhood-onset lupus nephritis study used spatial transcriptomics to generate a single-cell resolution atlas of kidney tissue from eight patients and four controls, assigning annotated cells to 30 reference cell types [<a href="#ref-6">6</a>]. The authors found that individual immune lineages localized to specific regions in lupus nephritis kidneys, including myeloid cells that trafficked to inflamed glomeruli and B cells that clustered within tubulointerstitial immune hotspots. Their QC workflow had to preserve the spatial information while removing low-quality cells, which required careful threshold selection for each cell type.

At a Glance: QC Metric Selection by Tissue Type

Tissue TypeRecommended ApproachKey QC MetricsCommon Pitfalls
Brain and neural tissueUse snRNA-seq preferentially, adjust gene count thresholds downward for nucleiGene count per nucleus, ambient RNA fraction, doublet rateApplying whole-cell mitochondrial thresholds to nuclear data, losing fragile neuronal populations
Heart and skeletal muscleUse snRNA-seq for cardiomyocytes, expect high mitochondrial contentNuclear gene count, doublet rate, ambient RNA from myofiber damageRemoving cardiomyocytes with standard mitochondrial cutoffs, clogging capture devices with contractile proteins
Tumor tissueAccount for dissociation-induced stress, consider cold-active proteaseMitochondrial read fraction with higher threshold, stress gene expression, doublet rateDiscarding stressed but viable tumor cells, misclassifying doublets as novel malignant populations
Lung tissueUse whole-cell scRNA-seq, accommodate diverse cell sizesUMI count, gene count, doublet rate, ambient RNA from alveolar damageLosing rare epithelial progenitor populations, retaining damaged alveolar cells
Liver and adipose tissueUse snRNA-seq to avoid lipid interference, separate parenchymal and non-parenchymal workflowsNuclear gene count, ambient RNA from hepatocyte lysis, doublet rateHigh ambient RNA obscuring true signals, losing lipid-laden adipocytes
Kidney tissueUse whole-cell or spatial approaches, accommodate large epithelial cellsUMI count, gene count, doublet rate, ambient RNA from tubular damageMisclassifying proximal tubule cells as doublets, losing infiltrating immune cells

Practical Workflow for Setting QC Thresholds

Step 1: Examine the Distribution of QC Metrics Before Filtering

Before setting any thresholds, generate violin plots and scatter plots of UMI count, gene count, mitochondrial read fraction, and doublet score across all cells or nuclei. Examine the distributions for each sample separately, because batch effects and processing differences can shift these metrics. A study of heart failure with preserved ejection fraction used genotype-based demultiplexing with souporcell and quantified gene expression with CellRanger and CellBender, demonstrating that the choice of quantification and demultiplexing tools affects the QC metric distributions [<a href="#ref-10">10</a>]. The authors assigned more than 70% of nuclei to individuals after demultiplexing pooled myocardial biopsies, and their QC thresholds were informed by the distributions of these assigned nuclei.

Step 2: Identify the Expected Cell Composition

Determine which cell types should be present in your sample based on the tissue and experimental condition. For example, a lung sample should contain alveolar macrophages, epithelial cells, endothelial cells, and immune cells. A heart sample should contain cardiomyocytes, fibroblasts, endothelial cells, pericytes, and immune cells. The hypertrophic cardiomyopathy study identified macrophages as the key cell cluster most associated with the disease and found that F13A1 expression was predominantly restricted to macrophage clusters [<a href="#ref-14">14</a>]. The authors used published scRNA-seq datasets and integrated them with bulk RNA-seq data, demonstrating that expected cell composition guides QC decisions.

Step 3: Set Initial Thresholds Based on Distribution Inflections

Use the inflection points in the QC metric distributions to set initial thresholds. For gene count, identify the point where the distribution drops sharply, which typically indicates the boundary between real cells and empty droplets or damaged cells. For mitochondrial read fraction, identify the population with elevated mitochondrial content and determine whether it represents a genuine cell type or damaged cells. In cardiac tissue, the mitochondrial-rich population includes cardiomyocytes, so the threshold must be set above the cardiomyocyte distribution. In brain snRNA-seq, mitochondrial read fraction is less informative, and gene count and ambient RNA fraction become more important.

Step 4: Evaluate the Impact of Thresholds on Cell Type Recovery

After applying initial thresholds, cluster the data and annotate cell types. Check whether any expected cell types are missing or underrepresented. If a known cell type is absent, the thresholds may be too stringent. If unexpected clusters appear with low gene counts and high mitochondrial content, the thresholds may be too lenient. The lung ARDS study identified 21 distinct cell populations and six macrophage and monocyte subpopulations after rigorous QC, demonstrating that appropriate thresholds preserve rare populations while removing low-quality cells [<a href="#ref-12">12</a>].

Step 5: Iterate and Document Threshold Decisions

QC threshold selection is an iterative process. Adjust thresholds based on the observed cell type recovery and document the final decisions, including the rationale for each threshold. This documentation is essential for reproducibility and for interpreting downstream results. The Alzheimer's disease guidelines implemented their recommended workflow for each major analytic direction and shared the scripts and data with the research community through GitHub, emphasizing the importance of transparent and reproducible QC decisions [<a href="#ref-9">9</a>].

Records and Measurements for QC Decisions

Maintain a QC log for each sample that records the following measurements:

  • Total number of cells or nuclei captured
  • Number of cells or nuclei passing QC
  • Percentage of cells or nuclei removed at each QC step
  • Median UMI count before and after filtering
  • Median gene count before and after filtering
  • Median mitochondrial read fraction before and after filtering
  • Estimated doublet rate
  • Ambient RNA contamination level
  • Dissociation protocol used and any deviations
  • Sequencing depth and platform

These records allow you to compare QC metrics across batches and experiments, identify systematic issues in sample processing, and justify threshold decisions in publications. The in vivo Perturb-seq protocol highlights where quality control checks can offer critical go-no-go points for a time- and cost-intensive method, emphasizing that QC decisions should be made at defined stages of the workflow instead of only at the end [<a href="#ref-15">15</a>].

Common Failure Patterns in QC Threshold Selection

Using Default Thresholds Without Examination

Many analysis pipelines include default QC thresholds, such as minimum gene count of 200 or maximum mitochondrial read fraction of 20%. Applying these defaults without examining the actual distributions in your data will remove biologically meaningful cells in some tissues and retain damaged cells in others. The sarcopenia study retained 13,612 tibialis anterior muscle single-cell transcriptomes after QC, and the thresholds used were specific to the muscle tissue and the senescence-accelerated mouse model [<a href="#ref-4">4</a>]. Default thresholds would not have been appropriate for this dataset.

Ignoring the Difference Between scRNA-seq and snRNA-seq

Whole-cell and nuclear preparations produce different QC metric distributions. Nuclei contain fewer transcripts and fewer mitochondrial reads, so thresholds developed for whole-cell data will remove genuine nuclei. The cardiac nuclei isolation protocol describes steps for mechanical homogenization, sequential filtration, sucrose cushion purification, and fluorescence-activated nuclei sorting, and the QC metrics for these nuclear libraries differ from those for whole-cell libraries [<a href="#ref-2">2</a>]. Researchers must use thresholds appropriate for the library type.

Applying the Same Thresholds Across Batches

Batch effects from sample processing, sequencing runs, and reagent lots can shift QC metric distributions. Applying the same thresholds across batches will remove different proportions of cells from each batch, introducing batch-specific bias. The childhood-onset lupus nephritis study used spatial transcriptomics across eight patients and four controls, and the authors had to account for batch effects when setting QC thresholds [<a href="#ref-6">6</a>]. Normalize QC metrics within batches or use methods that adjust for batch effects before filtering.

Removing Doublets Without Considering Biological Variation

Computational doublet detection methods flag cells with high UMI and gene counts, but some genuine cell types have high transcriptional output. Large cells such as hepatocytes, cardiomyocytes, and megakaryocytes can be flagged as doublets. The hypertrophic cardiomyopathy study found that F13A1 expression was predominantly restricted to macrophage clusters, and the authors had to distinguish between genuine macrophage populations and doublets involving macrophages and other cell types [<a href="#ref-14">14</a>]. Examine the expression of known cell-type markers in flagged doublets before removing them.

Overlooking Ambient RNA Contamination

Ambient RNA from lysed cells can inflate gene counts and obscure cell-type signatures. This is particularly problematic in tissues with fragile cells, such as liver and adipose tissue. The minimized-bias profiling protocol for liver and visceral adipose tissue separated whole-liver and non-parenchymal cell workflows to reduce dissociation bias, and the authors used shallow sequencing-based assessment of cell-type recovery to evaluate the effectiveness of their approach [<a href="#ref-7">7</a>]. Use methods such as CellBender or SoupX to estimate and remove ambient RNA contamination before setting gene count thresholds.

Limitations of QC Metrics and When to Escalate

QC metrics are imperfect proxies for cell quality, and their interpretation requires biological context. A cell with low gene count and high mitochondrial read fraction may be damaged, but it may also be a genuine quiescent or stressed population. The ovarian cancer study found that HMOX1 inhibition in epithelial cells could secrete TGF-beta1 to activate macrophage subtypes, demonstrating that stressed cells can have biologically meaningful signaling functions [<a href="#ref-11">11</a>]. Removing all stressed cells may eliminate important biological variation.

Professional escalation is warranted when QC metrics show systematic patterns that suggest protocol problems instead of biological variation. These patterns include:

  • A sudden drop in median UMI or gene count across all samples in a batch, suggesting a reagent or instrument failure
  • Consistently high doublet rates across samples, suggesting incorrect cell loading concentrations
  • High ambient RNA contamination across all samples, suggesting inadequate washing or excessive cell death during dissociation
  • Loss of all cells of a particular type across replicates, suggesting that the dissociation protocol is biased against that cell type

When these patterns appear, consult with the sequencing facility or a bioinformatics specialist before proceeding with analysis. The in vivo Perturb-seq protocol emphasizes that quality control checks can offer critical go-no-go points for time- and cost-intensive methods, and the same principle applies to standard scRNA-seq experiments [<a href="#ref-15">15</a>].

Integration of QC Decisions with Downstream Analysis

QC decisions directly affect downstream analyses including clustering, cell type annotation, differential expression, and trajectory inference. The immune cell annotation review highlighted that annotating immune cells based solely on transcriptomic data remains challenging because of gene expression heterogeneity and post-transcriptional regulation, which contribute to mismatches between mRNA and protein expression [<a href="#ref-16">16</a>]. These discrepancies can lead to cell misclassification and obscure functional insights, particularly in heterogeneous populations such as peripheral blood mononuclear cells. The authors emphasized that multimodal approaches, such as Cellular Indexing of Transcriptomes and Epitopes by Sequencing, can address the shortcomings of single-modality analyses.

The Alzheimer's disease guidelines covered 14 major analytic directions, including quality control and normalization, dimension reduction and feature extraction, cell clustering analysis, cell type inference and annotation, differential expression, trajectory inference, copy number variation analysis, integration of single-cell multi-omics, epigenomic analysis, gene network inference, prioritization of cell subpopulations, integrative analysis of human and mouse scRNA-seq data, spatial transcriptomics, and comparison of single-cell mouse model studies and single-cell human studies [<a href="#ref-9">9</a>]. The authors emphasized that QC decisions at the beginning of the workflow propagate through all downstream analyses, so careful threshold selection is essential for valid biological conclusions.

For studies that integrate single-cell data with bulk RNA-seq data, QC decisions affect the accuracy of cell type-specific inference. The EPIC-unmix method for cell type-specific inference from bulk RNA-sequencing data integrates single-cell reference profiles, and the quality of the single-cell reference directly affects the accuracy of the bulk deconvolution [<a href="#ref-17">17</a>]. Similarly, the periodontitis study combined scRNA-seq with bulk RNA-seq analysis to identify necroptosis-related genes as therapeutic targets, and the QC of the single-cell data determined which cell types and genes were included in the downstream analysis [<a href="#ref-18">18</a>].

Spatial Transcriptomics and QC Considerations

Spatial transcriptomics adds a spatial dimension to single-cell analysis, and QC decisions must account for the spatial context. The childhood-onset lupus nephritis study used spatial transcriptomics to demonstrate that individual immune lineages localized to specific regions in lupus nephritis kidneys, including myeloid cells that trafficked to inflamed glomeruli and B cells that clustered within tubulointerstitial immune hotspots [<a href="#ref-6">6</a>]. The authors found that gene expression varied as a function of tissue location, and they identified modules of spatially correlated gene expression with predicted roles in induction of inflammation and the development of tubulointerstitial fibrosis.

For spatial transcriptomics, QC metrics include the number of transcripts per spot, the number of genes detected per spot, and the fraction of reads mapping to the transcriptome. These metrics vary across tissue regions, and thresholds must be set to retain informative spots while removing background or damaged regions. The sepsis-associated encephalopathy study integrated snRNA-seq and spatial transcriptomics and identified a distinct region where Astro-2 and Micro-2 cells surrounded Vas-1 cells, demonstrating that spatial QC decisions affect the identification of biologically meaningful spatial domains [<a href="#ref-8">8</a>].

A Structured Decision Log for QC Threshold Selection Across Tissues and Conditions

A recurring failure in single-cell RNA-seq projects is not the absence of QC but the absence of a documented rationale for why specific thresholds were chosen. When thresholds are selected informally or copied from a previous project, the decisions cannot be revisited when downstream results reveal missing cell types or unexpected clusters. A structured decision log forces explicit reasoning at each QC step and creates a record that can be audited when biological conclusions are questioned. This section provides a practical framework for building such a log, with tissue-specific examples drawn from published studies.

The QC Decision Log Format

Create a table with one row per QC decision point and columns for the metric name, the tissue and condition, the dissociation and library preparation method, the initial threshold, the observed distribution, the final threshold, and the biological justification. Record the decision log as a plain text file or spreadsheet in the project directory before any filtering is applied. The log should be updated whenever a threshold is revised during the iterative QC process described in the practical workflow section.

The first entry in the log should always be the expected cell composition for the sample. For skeletal muscle from the sarcopenia mouse model, the expected populations include endothelial cells, immune cells, and fibro-adipogenic progenitors, because multinucleated myofibers are not captured as intact single cells [<a href="#ref-4">4</a>]. For cardiac tissue, the expected populations include cardiomyocyte nuclei, fibroblasts, endothelial cells, pericytes, and immune cells, with the recognition that cardiomyocyte nuclei show lower gene detection than other cardiac cell types [<a href="#ref-10">10</a>]. Writing down the expected composition before filtering prevents the common error of adjusting thresholds until a preconceived cluster structure appears.

Recording Metric Distributions Before Filtering

For each sample, record the median and range of UMI count, gene count, mitochondrial read fraction, and doublet score before any filtering. Also record the shape of each distribution, noting whether it is unimodal, bimodal, or skewed. These baseline distributions are the reference against which all threshold decisions are evaluated. In the lung ARDS rat model study, the authors emphasized rigorous quality control and doublet removal as essential for identifying 21 distinct cell populations and six macrophage and monocyte subpopulations [<a href="#ref-12">12</a>]. The baseline distributions in that study would have shown the heterogeneity across alveolar macrophages, epithelial cells, and immune cells that made a single universal threshold inappropriate.

The baseline distribution record also captures batch-specific variation. When samples are processed on different days or with different reagent lots, the median UMI count can shift by 20% or more without any change in biological state. Recording these shifts in the decision log prevents the misinterpretation of batch effects as biological differences. The heart failure study used genotype-based demultiplexing with souporcell and quantified expression with CellRanger and CellBender, assigning more than 70% of nuclei to individuals from pooled myocardial biopsies [<a href="#ref-10">10</a>]. The QC metric distributions for each individual were recorded after demultiplexing, allowing thresholds to be set within the context of each sample's unique distribution.

Threshold Justification Entries

For each threshold, write a justification that connects the cutoff to the expected biology of the tissue. A justification such as "mitochondrial threshold set at 20% because this is the standard value" is not acceptable. A defensible justification states the observed distribution, the cell types that would be removed at different cutoff values, and the biological reason for the final choice.

For cardiac tissue, the justification for a high mitochondrial threshold would state that cardiomyocytes are densely packed with mitochondria and that standard cutoffs would remove these essential cells. The single-cell study of the healthy and diseased adult mouse heart had to account for the high mitochondrial content of cardiomyocytes when setting QC thresholds, because standard cutoffs would have removed these cells from the analysis [<a href="#ref-5">5</a>]. The decision log entry would record the mitochondrial read fraction distribution, identify the cardiomyocyte population within that distribution, and justify the threshold as the point that retains the cardiomyocyte peak while removing cells with extreme mitochondrial content and very low gene detection.

For brain snRNA-seq, the justification for a low gene count threshold would state that nuclei contain fewer transcripts than whole cells and that the threshold is set at the inflection point of the gene count distribution. The sepsis-associated encephalopathy study used snRNA-seq to identify specialized subpopulations of astrocytes, microglia, and vascular cells in mouse brain [<a href="#ref-8">8</a>]. The decision log for that study would record that mitochondrial read fraction was not used as a primary quality metric because nuclei contain few mitochondrial transcripts, and that gene count and ambient RNA fraction were the informative metrics.

Linking QC Decisions to Dissociation Protocol Records

The decision log should include the dissociation protocol used for each sample, because the protocol directly affects QC metric distributions. The cold-active protease study demonstrated that standard collagenase digestion induces conserved stress responses in solid tumors, while cold-active protease treatment minimizes these responses [<a href="#ref-1">1</a>]. A decision log entry for a tumor sample would record the dissociation enzyme, temperature, and duration, and would note whether stress gene expression was examined as an additional quality metric.

For liver and adipose tissue, the minimized-bias profiling protocol separated whole-liver and non-parenchymal cell workflows to reduce dissociation bias, and used magnetic enrichment and BD Rhapsody capture followed by cDNA synthesis and whole-transcriptome amplification [<a href="#ref-7">7</a>]. The decision log for these tissues would record the nuclei extraction method, the enrichment strategy, and the expected impact on ambient RNA contamination from damaged hepatocytes.

Using the Decision Log for Troubleshooting

When a downstream analysis reveals a problem, the decision log is the first place to look. If a known cell type is missing, check the threshold justification for the metric that would have removed that population. If an unexpected cluster appears with low gene counts and high mitochondrial content, check whether the mitochondrial threshold was set too high for the tissue.

The in vivo Perturb-seq protocol highlights where quality control checks can offer critical go-no-go points for a time- and cost-intensive method [<a href="#ref-15">15</a>]. The decision log serves the same function for standard scRNA-seq experiments. When the log shows that a threshold was set based on a distribution inflection point, and the downstream analysis reveals that the inflection point corresponded to a genuine biological population, the log provides the evidence needed to revise the threshold and re-run the analysis.

Common Failure Patterns in Decision Documentation

The most common failure is recording only the final thresholds without the baseline distributions or the justification. This makes it impossible to determine whether a threshold was chosen for biological reasons or for convenience. A second failure is updating the log after the analysis is complete, which introduces recall bias and reduces the reliability of the record. A third failure is using different log formats for different samples, which makes comparison across batches difficult.

The Alzheimer's disease guidelines implemented their recommended workflow for each major analytic direction and shared the scripts and data with the research community through GitHub [<a href="#ref-9">9</a>]. The decision log should follow the same principle of transparency. When the log is shared with the analysis code, reviewers and collaborators can evaluate whether the QC decisions were appropriate for the tissue and condition.

Records and Measurements for the Decision Log

Maintain the following measurements in the decision log for each sample:

  • Sample identifier, tissue type, and experimental condition
  • Dissociation protocol including enzyme, temperature, and duration
  • Library preparation method and sequencing platform
  • Total cells or nuclei captured before filtering
  • Median and range of UMI count before filtering
  • Median and range of gene count before filtering
  • Median and range of mitochondrial read fraction before filtering
  • Estimated doublet rate before filtering
  • Ambient RNA contamination estimate
  • Expected cell composition with reference to published studies
  • Initial threshold for each QC metric
  • Observed distribution inflection points
  • Final threshold for each QC metric
  • Biological justification for each final threshold
  • Cells or nuclei retained after filtering
  • Cell types identified after clustering and annotation

These records allow the QC decisions to be evaluated in the context of the full experimental workflow. The hypertrophic cardiomyopathy study integrated published scRNA-seq datasets with bulk RNA-seq data and identified macrophages as the key cell cluster most associated with the disease [<a href="#ref-14">14</a>]. The decision log for that study would show how the QC thresholds were set to retain macrophage populations while removing low-quality cells, and how those decisions affected the downstream identification of F13A1 as a macrophage-specific marker.

Escalation Criteria Based on Decision Log Patterns

The decision log can reveal systematic patterns that warrant escalation to a bioinformatics specialist or sequencing facility. If the log shows that the median UMI count dropped sharply across all samples in a batch, this suggests a reagent or instrument failure. If the log shows consistently high doublet rates across samples, this suggests incorrect cell loading concentrations. If the log shows high ambient RNA contamination across all samples, this suggests inadequate washing or excessive cell death during dissociation.

The decision log also reveals whether the dissociation protocol is biased against a particular cell type. If the log shows that a specific cell type is absent across all replicates, and the threshold justifications do not explain the absence, the dissociation protocol itself may be the cause. The lung epithelial study identified alveolar epithelial progenitors as a rare population, and the QC thresholds had to be adjusted to retain these cells [<a href="#ref-3">3</a>]. The decision log for that study would show the threshold adjustments and the biological justification for each adjustment, providing a template for troubleshooting similar losses in other tissues.

Frequently Asked Questions

What is the difference between QC thresholds for scRNA-seq and snRNA-seq?

Whole-cell scRNA-seq libraries contain cytoplasmic transcripts, so mitochondrial read fraction is a useful quality metric because high mitochondrial content with low cytoplasmic gene detection indicates membrane damage. Single-nucleus libraries contain few cytoplasmic transcripts, so mitochondrial read fraction is less informative. Gene count thresholds are typically lower for snRNA-seq because nuclei contain less RNA than whole cells. The cardiac nuclei isolation protocol describes steps for obtaining high-quality nuclei from fresh-frozen tissue, and the QC metrics for these libraries differ from those for whole-cell libraries [<a href="#ref-2">2</a>].

How should I set the mitochondrial read fraction threshold for heart tissue?

Cardiomyocytes are densely packed with mitochondria, so their mitochondrial read fraction in whole-cell libraries can exceed the 10% to 20% threshold commonly used in general tutorials. For cardiac tissue, examine the distribution of mitochondrial read fraction and identify the population that corresponds to cardiomyocytes based on known markers. Set the threshold above this population to retain cardiomyocytes while removing cells with extreme mitochondrial content and very low gene detection. For snRNA-seq of heart tissue, mitochondrial read fraction is less informative because nuclei contain few mitochondrial transcripts.

Why does my tumor sample show high mitochondrial read fraction across many cells?

Tumor tissue often contains stressed and dying cells within the tumor microenvironment, and the dissociation process itself induces stress responses. A study of solid tumor dissociation found that cold-active protease treatment minimized conserved collagenase-associated stress responses compared with standard warm enzymatic digestion [<a href="#ref-1">1</a>]. High mitochondrial read fraction in tumor samples may reflect genuine stress in the tumor microenvironment instead of technical artifacts. Consider using higher mitochondrial thresholds for tumor tissue and examine stress gene expression to distinguish between stressed but viable cells and damaged cells.

How do I handle ambient RNA contamination in liver samples?

Liver tissue contains abundant hepatocytes that are easily damaged during dissociation, releasing high levels of ambient RNA. This contamination appears as low-level expression of many genes across all cells, which can inflate gene counts and obscure cell-type signatures. The minimized-bias profiling protocol for liver and visceral adipose tissue separated whole-liver and non-parenchymal cell workflows to reduce dissociation bias [<a href="#ref-7">7</a>]. Use computational methods to estimate and remove ambient RNA contamination before setting gene count thresholds, and consider using snRNA-seq to reduce the impact of cytoplasmic ambient RNA.

What should I do when QC thresholds remove an expected cell type?

If an expected cell type is absent after QC filtering, examine the QC metrics for that cell type in the unfiltered data. The cells may have been removed because they have low gene counts or high mitochondrial read fraction, which could indicate that the thresholds are too stringent for that population. Alternatively, the dissociation protocol may be biased against that cell type. The lung epithelial study identified alveolar epithelial progenitors as a rare population, and the QC thresholds had to be adjusted to retain these cells [<a href="#ref-3">3</a>]. Iterate on the thresholds and document the rationale for each adjustment.

How do batch effects influence QC threshold selection?

Batch effects from sample processing, sequencing runs, and reagent lots can shift QC metric distributions. Applying the same thresholds across batches will remove different proportions of cells from each batch, introducing batch-specific bias. Examine QC metric distributions for each sample separately and set thresholds within batches or use methods that adjust for batch effects before filtering. The heart failure study used genotype-based demultiplexing and quantified gene expression with CellRanger and CellBender, and the QC thresholds were informed by the distributions of demultiplexed nuclei across pooled samples [<a href="#ref-10">10</a>].

When should I escalate QC issues to a specialist?

Escalate when QC metrics show systematic patterns that suggest protocol problems instead of biological variation. These patterns include a sudden drop in median UMI or gene count across all samples in a batch, consistently high doublet rates across samples, high ambient RNA contamination across all samples, or loss of all cells of a particular type across replicates. The in vivo Perturb-seq protocol highlights where quality control checks can offer critical go-no-go points for time- and cost-intensive methods, and the same principle applies to standard scRNA-seq experiments [<a href="#ref-15">15</a>].

How do QC decisions affect downstream cell type annotation?

QC decisions directly affect clustering and cell type annotation because low-quality cells can form spurious clusters or obscure genuine populations. The immune cell annotation review highlighted that annotating immune cells based solely on transcriptomic data remains challenging because of gene expression heterogeneity and post-transcriptional regulation [<a href="#ref-16">16</a>]. Removing too many cells can eliminate rare populations, while retaining too many low-quality cells can create artifactual clusters. The Alzheimer's disease guidelines implemented their recommended workflow for each major analytic direction and shared the scripts and data with the research community, emphasizing the importance of transparent and reproducible QC decisions [<a href="#ref-9">9</a>].

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

[1] [Dissociation of solid tumor tissues with cold active protease for single-cell RNA-seq minimizes conserved collagenase-associated stress responses](https://doi.org/10.1186/s13059-019-1830-0). Genome Biology, 2019. [2] [Protocol for isolation of nuclei from murine cardiac tissue for single-nucleus multiomic sequencing.](https://doi.org/10.1016/j.xpro.2026.104615). 2026. [3] [Reconstructing the Developmental Trajectories of Multiple Subtypes in Pulmonary Parenchymal Epithelial Cells by Single-Cell RNA-seq](https://doi.org/10.3389/fgene.2020.573429). Frontiers in Genetics, 2020. [4] [Single-cell RNA-seq reveals interferon-induced guanylate-binding proteins are linked with sarcopenia.](https://pubmed.ncbi.nlm.nih.gov/36162807). Journal of cachexia, sarcopenia and muscle, 2022. [5] [Single-Cell Sequencing of the Healthy and Diseased Heart Reveals Cytoskeleton-Associated Protein 4 as a New Modulator of Fibroblasts Activation.](https://pubmed.ncbi.nlm.nih.gov/29386203). Circulation, 2018. [6] [Childhood-onset lupus nephritis is characterized by complex interactions between kidney stroma and infiltrating immune cells.](https://pubmed.ncbi.nlm.nih.gov/39602512). Science translational medicine, 2024. [7] [Protocol for minimized-bias profiling of liver and visceral adipose tissue in mice using integrated single-nucleus RNA sequencing](https://doi.org/10.1016/j.xpro.2026.104699). STAR Protocols, 2026. [8] [Integrating single-nucleus RNA sequencing and spatial transcriptomics to elucidate a specialized subpopulation of astrocytes, microglia and vascular cells in brains of mouse model of lipopolysaccharide-induced sepsis-associated encephalopathy.](https://pubmed.ncbi.nlm.nih.gov/38961424). Journal of neuroinflammation, 2024. [9] [Guidelines for bioinformatics of single-cell sequencing data analysis in Alzheimer's disease: review, recommendation, implementation and application.](https://pubmed.ncbi.nlm.nih.gov/35236372). Molecular neurodegeneration, 2022. [10] [Single-Cell Analysis of Human Heart Failure With Preserved Ejection Fraction.](https://pubmed.ncbi.nlm.nih.gov/42137938). Circulation research, 2026. [11] [Distinct roles of HMOX1 on tumor epithelium and macrophage for regulation of immune microenvironment in ovarian cancer.](https://pubmed.ncbi.nlm.nih.gov/40638251). International journal of surgery (London, England), 2025. [12] [Single-Cell RNA Sequencing of Lung Tissue in a Rat Model of Acute Respiratory Distress Syndrome.](https://doi.org/10.3791/70491). Journal of Visualized Experiments, 2026. [13] [Single-cell RNA sequencing reveals the heterogeneity of adipose tissue-derived mesenchymal stem cells under chondrogenic induction](https://doi.org/10.5483/BMBRep.2023-0161). BMB Reports, 2023. [14] [Integrative Transcriptomics and Machine Learning Identify Macrophage-Associated Biomarkers in Hypertrophic Cardiomyopathy.](https://doi.org/10.3390/ijms27115102). 2026. [15] [Massively parallel in vivo Perturb-seq screening.](https://pubmed.ncbi.nlm.nih.gov/39939709). Nature protocols, 2025. [16] [Immune cell annotation in the single-cell studies: technologies, challenges, and integrative solutions.](https://doi.org/10.1007/s12026-026-09780-4). 2026. [17] [Cell type-specific inference from bulk RNA-sequencing data by integrating single-cell reference profiles via EPIC-unmix](https://doi.org/10.1186/s13059-025-03847-5). Genome Biology, 2025. [18] [Single-cell RNA-seq combined with bulk RNA-seq analysis identifies necroptosis-related genes as therapeutic targets for periodontitis](https://doi.org/10.1186/s12920-025-02241-1). BMC Medical Genomics, 2025.

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.