Single-Cell vs. Single-Nucleus RNA-Seq: How Quality Control Metrics Differ and What to Adjust in Your Filtering Pipeline
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Mitochondrial fraction interpretation shifts from cell health to contamination: In single-nucleus RNA sequencing (snRNA-seq), high mitochondrial transcript percentages indicate ambient contamination or attached mitochondria, not compromised nuclear integrity, necessitating lower thresholds (1-5%) compared to single-cell RNA sequencing (scRNA-seq) (10-20%).
- Gene detection rates are inherently lower in nuclei: Nuclear RNA content is less abundant than whole-cell RNA, leading to fewer detected genes per nucleus, requiring a reduction in minimum gene count thresholds to avoid systematic removal of transcriptionally less active cell types.
- Ambient RNA profiles differ significantly and require targeted correction: Nuclear isolation buffers contain fewer cytoplasmic transcripts, but specific tissues may exhibit dominant ambient transcripts (e.g., globin genes in blood) that necessitate computational correction (e.g., CellBender) or targeted depletion strategies during library preparation.
- Doublet detection is more challenging in snRNA-seq: The lower RNA content of nuclei makes distinguishing doublets from single nuclei with high RNA content more difficult, favoring genotype-based demultiplexing (e.g., souporcell) for pooled samples over standard computational doublet detection algorithms.
- Sample condition dictates assay choice and QC strategy: Frozen tissues are primarily suitable for snRNA-seq, while fresh tissues allow for both scRNA-seq and snRNA-seq, with the latter often preferred for fragile cell types and archived samples, impacting downstream QC metric interpretation.
- Verification of cell-type retention post-filtering is critical: After applying QC thresholds, it is essential to cluster and examine the cell-type composition to ensure that specific, potentially rare or transcriptionally low, cell populations have not been disproportionately removed, which could bias downstream biological conclusions.
Researchers moving from single-cell RNA sequencing (scRNA-seq) to single-nucleus RNA sequencing (snRNA-seq) often apply familiar quality control thresholds and encounter unexpected results. The core problem is that nuclear preparations produce different distributions of key metrics, particularly mitochondrial transcript fraction, ambient RNA contamination, and gene detection rates. This article explains how QC metrics differ between whole-cell and nuclear data, what threshold adjustments are justified by current evidence, and how to build a filtering pipeline that accounts for these differences without discarding biologically meaningful nuclei.
Why Nuclear Preparations Change Your QC Assumptions
The transition from scRNA-seq to snRNA-seq is not a simple swap of input material. Whole-cell protocols capture cytoplasmic and nuclear transcripts, while nuclear protocols capture primarily nuclear RNA. This distinction changes the meaning of every standard QC metric you have been using.
Biological Basis for Metric Differences
In whole-cell preparations, the mitochondrial fraction is a standard indicator of cell health. High mitochondrial reads typically signal a compromised cell membrane that has lost cytoplasmic mRNA while retaining mitochondrial transcripts. This logic does not transfer directly to nuclear data. Nuclei do not contain mitochondria, so mitochondrial transcripts detected in snRNA-seq data come from ambient contamination or from mitochondria that remain attached to the nuclear envelope during isolation. The biological signal you are measuring is fundamentally different.
The evidence from archived blood samples illustrates this point. When researchers evaluated two nuclei-isolation techniques for snRNA-seq of PAXgene whole blood samples, they found that cell lysis produced up to two orders of magnitude higher nuclei yields than mechanical separation and also produced less biased proportions of immune cells. High ambient globin gene counts following lysis were substantially reduced by CRISPR-guided globin gene depletion of complementary DNA, resulting in more sensitive and efficient gene detection per cell. This finding shows that the isolation method itself shapes the ambient RNA profile you will need to filter.
What Changes in the Metric Distributions
Gene detection rates shift because nuclear RNA is less abundant than total cellular RNA. A nucleus contains a fraction of the mRNA present in the whole cell, so you should expect lower median genes per nucleus compared to cells from the same tissue. This is not a quality failure. It is an expected property of the assay.
Mitochondrial fraction distributions compress toward zero in nuclear data. Most nuclei will show very low mitochondrial percentages, and the threshold that worked for cells, often 10 to 20 percent, will be meaningless because almost no nuclei approach those values. The useful signal in nuclear data is often the small number of nuclei with unusually high mitochondrial reads, which may indicate damaged nuclei or excessive ambient contamination.
Ambient RNA behaves differently because nuclear isolation buffers contain fewer cytoplasmic transcripts than whole-cell dissociation buffers. However, certain tissues produce persistent ambient signals. The blood sample study showed that globin genes dominated the ambient profile after cell lysis, requiring active depletion to recover sensitivity. Your tissue of interest may have similar dominant ambient transcripts that need specific handling.
Core QC Metrics and Their Nuclear Interpretations
Standard scRNA-seq QC relies on three primary metrics: total UMI count, number of detected genes, and mitochondrial fraction. Each requires reinterpretation for snRNA-seq data.
Total UMI Count
The total UMI count reflects sequencing depth and capture efficiency. In nuclear preparations, you should expect lower total UMIs per nucleus compared to cells from the same tissue because the starting RNA pool is smaller. The practical implication is that your minimum UMI threshold should be set relative to the nuclear distribution, not carried over from your cell-based pipeline.
A useful approach is to examine the UMI distribution across your nuclei and identify the inflection point where the distribution drops off. Nuclei below this point often represent empty droplets or nuclei with very low RNA content that will not support reliable clustering. The exact threshold depends on your tissue, protocol, and sequencing depth.
Number of Detected Genes
Detected genes per nucleus follows the same logic as UMI count. Nuclear transcriptomes are less complex than whole-cell transcriptomes because nuclear RNA is enriched for intronic and nascent transcripts while lacking some cytoplasmic mRNA species. The evidence from the aldosteronoma study, which analyzed 140,742 cells from 28 samples using 10x single-nucleus RNA-seq, demonstrates that nuclear data can support detailed cell-type identification and trajectory analysis when QC is handled appropriately.
The key adjustment is to lower your minimum gene threshold and to verify that your chosen threshold does not remove specific cell types. Some cell types, particularly those with low transcriptional activity, will naturally have fewer detected genes. Setting the threshold too high will systematically remove these populations.
Mitochondrial Fraction
Mitochondrial fraction in snRNA-seq data should be interpreted as a contamination metric instead of a cell-health metric. Since nuclei lack mitochondria, any mitochondrial reads represent ambient contamination from lysed cells or mitochondria that remained attached during isolation.
The practical threshold for nuclear data is often much lower than for whole-cell data. Many snRNA-seq pipelines use thresholds between 1 and 5 percent mitochondrial fraction. However, you should examine your own distribution instead of applying a fixed value. If your isolation protocol produces consistently low mitochondrial contamination, you may be able to use a more stringent threshold without losing nuclei.
The heart failure studies provide a useful reference point. Both the preprint and the published version used snRNA-seq on myocardial biopsies with nuclei isolated from combined samples and applied quality control before clustering and annotation. The published study recovered 48,886 nuclei after QC and identified 14 cell types. This scale of analysis depends on QC thresholds that retain sufficient nuclei while removing damaged or contaminated ones.
At a Glance: QC Metric Comparison Table
| Metric | Whole-Cell scRNA-seq | Single-Nucleus snRNA-seq | Practical Adjustment |
|---|---|---|---|
| Mitochondrial fraction | High values indicate cell damage, typical threshold 10 to 20 percent | Low values expected, high values indicate ambient contamination or attached mitochondria | Lower threshold to 1 to 5 percent and examine distribution before setting cutoff |
| Median genes per cell or nucleus | Higher due to cytoplasmic mRNA capture | Lower due to nuclear RNA content | Reduce minimum gene threshold and verify cell-type retention |
| Ambient RNA profile | Cytoplasmic transcripts from dissociation | Buffer and lysis-derived transcripts, tissue-specific dominant species possible | Use ambient RNA removal tools and consider targeted depletion for dominant transcripts |
| Doublet detection | Based on combined transcriptomes of two cells | Based on combined nuclear transcriptomes, less distinct due to lower RNA content | Adjust doublet detection parameters or use genotype-based demultiplexing when pooling samples |
Building a Nuclear-Specific Filtering Pipeline
A filtering pipeline for snRNA-seq data requires deliberate choices at each stage. The following workflow reflects current practice in published snRNA-seq studies and available bioinformatics infrastructure.
Step 1: Generate Count Matrices with Appropriate Tools
The count matrix generation step determines the quality of all downstream metrics. The heart failure studies used CellRanger for gene expression quantification followed by CellBender for ambient RNA removal. This combination addresses the two main sources of technical noise in nuclear data: read alignment and ambient contamination.
CellBender is particularly relevant for nuclear data because it models and removes ambient RNA using the empty droplets in your sample. The blood sample study showed that ambient globin genes can dominate the signal after certain isolation methods, and active removal was necessary to recover gene detection sensitivity. If your tissue has a dominant ambient transcript, consider whether targeted depletion during library preparation is warranted.
Step 2: Examine Metric Distributions Before Filtering
Do not set thresholds before looking at your data. Generate violin plots and scatter plots of UMI count, gene count, and mitochondrial fraction for each sample. Look for bimodal distributions that indicate distinct populations of high-quality and low-quality nuclei.
The popsicleR package provides an interactive framework for this examination. It integrates methods from widely used pipelines for estimating quality-control metrics, filtering low-quality cells, data normalization, and removal of technical and biological biases. The package starts from either Cell Ranger output files or a feature-barcode matrix of raw counts from any scRNA-seq technology. This flexibility is valuable when you are working with snRNA-seq data from different platforms.
Step 3: Set Sample-Specific Thresholds
Each sample will have its own metric distributions based on tissue type, isolation protocol, and sequencing depth. Set thresholds per sample instead of applying global cutoffs across your entire dataset.
For mitochondrial fraction, start with a threshold between 1 and 5 percent and examine which nuclei are removed. If the removed nuclei are spread evenly across cell types, the threshold is likely appropriate. If specific cell types are disproportionately removed, the threshold may be too stringent.
For UMI and gene counts, identify the natural break in the distribution. Nuclei below this break are likely empty droplets or low-quality captures. The exact position of this break will vary between samples and should be determined empirically.
Step 4: Handle Doublets and Multiplexed Samples
Doublet detection in nuclear data is more challenging than in whole-cell data because nuclear transcriptomes are less complex, making it harder to distinguish a doublet from a single nucleus with high RNA content.
The heart failure studies used genotype-based demultiplexing with souporcell for pooled samples. This approach assigns droplets to individual patients based on genetic variants, which also enables doublet detection because a droplet containing nuclei from two individuals will show mixed genotypes. The preprint reported assigning more than 75 percent of droplets to individual patients, while the published version reported more than 70 percent of nuclei assigned. This method is particularly valuable when pooling samples to reduce cost.
If you are not pooling samples, consider using doublet detection tools that are calibrated for nuclear data. Be aware that default parameters trained on whole-cell data may not perform optimally on nuclear data.
Step 5: Assess Ambient RNA and Apply Correction
Ambient RNA correction is more critical for snRNA-seq than for scRNA-seq because nuclear isolation buffers can contain transcripts from lysed cells, and the lower RNA content of nuclei makes ambient contamination proportionally more impactful.
CellBender, used in the heart failure studies, models ambient RNA and removes it from the count matrix. The blood sample study demonstrated that ambient globin genes could be substantially reduced by CRISPR-guided depletion, which is a library preparation approach instead of a computational correction. Both strategies have value, and the choice depends on your tissue and the severity of contamination.
Step 6: Verify Cell-Type Retention After Filtering
After applying thresholds, verify that you have not systematically removed specific cell types. This is a critical step that is often skipped. Cluster your filtered nuclei and compare the cell-type composition to your expectations based on the tissue.
The Lynch Syndrome study provides an example of why this verification matters. The researchers used scRNA-seq for fresh biopsy samples and snRNA-seq for a frozen biopsy sample, then performed integrative computational analysis. The ability to integrate these data types depends on retaining comparable cell populations in both datasets. If your snRNA-seq filtering removes a cell type that is present in your scRNA-seq data, your integration will be biased.
Practical Implementation Steps
The following steps translate the pipeline concepts into concrete actions for your analysis.
Step 1: Document Your Isolation Protocol
Record the isolation method, buffer composition, and any depletion steps. The blood sample study showed that mechanical separation and cell lysis produced very different yields and cell-type proportions. This information is essential for interpreting your QC metrics and for troubleshooting if your data quality is poor.
Step 2: Run Initial QC Metric Calculation
Use a standard tool such as Seurat, Scanpy, or popsicleR to calculate QC metrics for each nucleus. Generate distributions for UMI count, gene count, and mitochondrial fraction. Save these plots for your records.
Step 3: Apply Ambient RNA Correction
Run CellBender or an equivalent tool before filtering. The heart failure studies applied CellBender after CellRanger quantification, and this order is important because ambient RNA correction should occur before you set thresholds based on gene detection.
Step 4: Set and Apply Thresholds
Based on your metric distributions, set thresholds for minimum UMI count, minimum gene count, and maximum mitochondrial fraction. Apply these thresholds and record how many nuclei are removed at each step.
Step 5: Assess Doublet Rate
If you pooled samples, run genotype-based demultiplexing and assess the doublet rate. If you did not pool, run a doublet detection tool and examine the proportion of nuclei flagged as doublets.
Step 6: Cluster and Verify
Cluster your filtered nuclei and examine the cell-type composition. Compare this to your expectations and to any matched scRNA-seq data you have. If specific cell types are missing, revisit your thresholds.
Step 7: Document All Decisions
Record every threshold you applied and the rationale for each decision. This documentation is essential for reproducibility and for interpreting your downstream results.
Records and Measurements to Maintain
Maintaining detailed records of your QC process is essential for reproducibility and for troubleshooting when results are unexpected.
Sample-Level Records
For each sample, record the tissue type, isolation method, pooling strategy, sequencing depth, and the number of nuclei captured before and after filtering. The heart failure studies provide a model for this documentation, reporting the number of patients per pool, the total nuclei recovered after QC, and the number of cell types identified.
Metric Distribution Records
Save the metric distributions for each sample before and after filtering. This includes the median and range of UMI counts, gene counts, and mitochondrial fractions. These records allow you to compare samples and to identify batch effects that may require correction.
Threshold Decision Records
Document the threshold values you applied and the evidence that supported each decision. If you adjusted a threshold because a specific cell type was being removed, record that observation. This information is valuable when you revisit the analysis or when you process additional samples.
Computational Environment Records
Record the versions of all software tools you used, including CellRanger, CellBender, souporcell, Seurat, Scanpy, and any other packages. The Bioconductor project provides official documentation for reproducible genomic analysis, and the nf-core documentation describes community standards for reproducible workflows. Following these standards ensures that your analysis can be reproduced by others.
Common Failure Patterns in Nuclear QC
Several failure patterns recur when researchers transition from scRNA-seq to snRNA-seq. Recognizing these patterns can save substantial time and prevent incorrect biological conclusions.
Pattern 1: Applying Whole-Cell Mitochondrial Thresholds
Researchers who apply a 10 or 20 percent mitochondrial threshold to nuclear data often find that almost no nuclei are removed. This creates a false sense that their data are exceptionally clean. The real issue is that the metric is not measuring what they think it is. Nuclear mitochondrial reads reflect contamination, not cell health, and the threshold should be set based on the contamination distribution.
Pattern 2: Overly Stringent Gene Count Thresholds
Setting a minimum gene count based on whole-cell expectations will remove a large fraction of nuclei, particularly those from cell types with lower transcriptional activity. The result is a filtered dataset that is biased toward highly transcribed cell types. This bias can lead to incorrect conclusions about cell-type composition.
Pattern 3: Ignoring Ambient RNA
Nuclear data from certain tissues, particularly blood and liver, can have substantial ambient RNA from dominant transcripts. The blood sample study showed that globin genes could dominate the ambient profile and reduce gene detection sensitivity. Ignoring this contamination will inflate the apparent quality of low-quality nuclei and obscure biological signal.
Pattern 4: Inadequate Doublet Handling
Doublet detection tools trained on whole-cell data may not perform well on nuclear data because the lower RNA content makes doublets harder to distinguish from single nuclei. The heart failure studies addressed this by using genotype-based demultiplexing, which provides a direct measure of whether a droplet contains nuclei from multiple individuals.
Pattern 5: Failing to Verify Cell-Type Retention
Applying thresholds without checking which cell types are removed can systematically eliminate rare or low-transcription cell types. This is particularly problematic for studies of heterogeneous tissues where rare populations are biologically important.
Limitations of Nuclear QC Metrics
Understanding the limitations of each QC metric is essential for interpreting your data correctly.
Mitochondrial Fraction Limitations
The mitochondrial fraction in nuclear data is not a reliable indicator of nuclear integrity. A nucleus with high mitochondrial reads may be contaminated by ambient RNA instead of damaged. Conversely, a damaged nucleus may not show elevated mitochondrial reads if the mitochondria were removed during isolation. This metric should be used cautiously and in combination with other indicators.
Gene Count Limitations
Gene count is influenced by sequencing depth, capture efficiency, and the transcriptional activity of the cell type. Comparing gene counts across samples with different sequencing depths is not meaningful. Within a sample, gene count can help distinguish empty droplets from nuclei, but the threshold must be set based on the distribution.
UMI Count Limitations
UMI count is similarly influenced by sequencing depth and capture efficiency. The blood sample study showed that different isolation methods produced different nuclei yields and cell-type proportions, which means that UMI distributions can vary substantially between protocols. Comparing UMI counts across protocols is not valid.
Ambient RNA Correction Limitations
Ambient RNA correction tools model the ambient profile from empty droplets, but this model may not capture all sources of contamination. The blood sample study showed that targeted depletion of dominant transcripts was necessary to recover sensitivity, suggesting that computational correction alone may be insufficient for some tissues.
Quality Controls and Verification Steps
Beyond the initial filtering thresholds, several quality controls should be applied throughout the analysis.
Verify Marker Gene Expression
After clustering, verify that known marker genes are expressed in the expected cell types. The heart failure studies identified 14 cell types based on specific marker genes, and the Lynch Syndrome study validated proposed markers using immuno-histochemical staining. This verification step confirms that your QC decisions did not remove the biological signal you are studying.
Compare to Matched Whole-Cell Data
If you have matched scRNA-seq and snRNA-seq data from the same tissue, compare the cell-type proportions and gene expression profiles. The blood sample study found that CL-derived samples maintained similar cell-type proportions and gene expression as matched peripheral blood mononuclear cell samples, while retaining granulocytes. This comparison provides confidence that your nuclear data are biologically representative.
Assess Batch Effects
If you processed samples in multiple batches, assess whether batch effects are present in your data. The heart failure studies pooled samples from multiple patients, which introduces batch structure that must be accounted for in the analysis. The nf-core documentation provides guidance on reproducible workflow configuration that can help standardize processing across batches.
Evaluate Integration Quality
If you are integrating snRNA-seq data with scRNA-seq data, evaluate the quality of the integration. The Lynch Syndrome study performed integrative computational analysis of scRNA-seq and snRNA-seq data, which requires careful handling of the technical differences between the two assays. Poor integration can create artificial cell populations or obscure real ones.
Welfare and Safety Context for Laboratory Practice
While snRNA-seq is not an animal husbandry procedure, the laboratory practices involved have safety considerations that warrant attention.
Sample Handling Safety
Nuclei isolation involves the use of lysis buffers and mechanical disruption methods. The blood sample study evaluated mechanical separation using an Acrodisc filter and cell lysis as alternative approaches. Both methods require careful handling of biological samples and chemical reagents. Follow your institutional biosafety guidelines for handling human or animal tissues.
Chemical Safety
Lysis buffers often contain detergents and other chemicals that require appropriate personal protective equipment. Consult the safety data sheets for all reagents used in your isolation protocol. The Carpentries lessons provide foundational training in data handling and reproducible practices that can help you document your protocols safely.
Data Management Safety
Single-nucleus datasets are large and require substantial computational resources. The Galaxy Training Network provides accessible workflow training that includes guidance on managing large datasets. Ensure that your data storage and backup procedures comply with your institutional requirements, particularly if you are working with human data.
Professional Escalation Criteria
Knowing when to seek help is an important part of any bioinformatics analysis. The following situations warrant consultation with a bioinformatics specialist or a statistician.
Escalate When Metric Distributions Are Unexpected
If your mitochondrial fraction distribution is bimodal with a substantial high peak, or if your gene count distribution does not show a clear separation between nuclei and empty droplets, consult a specialist. These patterns may indicate problems with your isolation protocol or sequencing run that require troubleshooting.
Escalate When Cell-Type Composition Is Biologically Implausible
If your filtered data lack a cell type that is known to be present in your tissue, or if the proportions of cell types are dramatically different from published data, escalate. This may indicate that your thresholds are too stringent or that your isolation protocol is biased against certain cell types.
Escalate When Integration Fails
If you cannot integrate your snRNA-seq data with matched scRNA-seq data, or if integration produces artificial populations, escalate. The technical differences between the two assays can be challenging to handle, and specialized methods may be required.
Escalate When Computational Resources Are Insufficient
If your dataset is too large for your available computational resources, or if your analysis is taking an unreasonable amount of time, escalate. The nf-core documentation provides guidance on configuring workflows for different computational environments, and the Bioconductor project provides documentation for reproducible genomic analysis that can help you optimize your pipeline.
A Practical Decision Framework for Choosing Between scRNA-seq and snRNA-seq Based on Your Sample Type
The decision to use single-cell or single-nucleus RNA sequencing is often driven by sample availability and research questions, but the quality control implications of that choice deserve explicit consideration before you begin. A structured decision framework helps you anticipate which QC challenges you will face and which thresholds you will need to adjust. This section provides a practical framework for matching your sample type and research question to the appropriate assay, with specific attention to how that choice shapes your QC pipeline.
Step 1: Assess Your Sample Condition and Storage History
The first decision point concerns the physical state of your tissue. Fresh tissue samples are compatible with both scRNA-seq and snRNA-seq, but frozen tissue is generally only suitable for snRNA-seq because the freeze-thaw process disrupts cell membranes while leaving nuclei intact. The Lynch Syndrome study illustrates this distinction clearly. The researchers used scRNA-seq for paired fresh biopsy samples from three patients and snRNA-seq for a frozen biopsy sample from the same patient cohort. This design allowed them to compare the two assays directly on matched tissue from the same genetic background.
Your sample storage history should drive your assay choice. If your samples have been frozen for any length of time, snRNA-seq is the appropriate choice. If you have fresh tissue, you have the option of either assay, and your choice should depend on the downstream questions and the cell types of interest. The blood sample study provides another storage consideration. The researchers worked with PAXgene whole blood samples, which are preserved for RNA analysis but were not previously considered suitable for single-cell analysis. Their work demonstrated that snRNA-seq with appropriate isolation methods can recover comprehensive cellular information from archived samples. If you have access to archived preserved samples, snRNA-seq may be your only viable option.
Step 2: Evaluate Your Cell-Type Recovery Priorities
Different cell types present different challenges for nuclear isolation and whole-cell dissociation. Some cell types are more fragile than others, and the choice of assay can systematically bias your recovered cell populations.
The blood sample study provides direct evidence for this bias. When the researchers compared mechanical separation using an Acrodisc filter to cell lysis for nuclei isolation, they found that cell lysis produced up to two orders of magnitude higher nuclei yields and less biased proportions of immune cells. This finding has direct implications for your QC pipeline. If you use a mechanical separation method, you may need to adjust your expectations for cell-type representation and may need to sequence more deeply to recover rare populations.
For tissues with fragile cell types, such as neurons or cardiomyocytes, snRNA-seq is often preferred because the isolation process is gentler on the tissue. The heart failure studies used snRNA-seq on myocardial biopsies and successfully identified 14 cell types, including cardiomyocytes, fibroblasts, endothelial cells, pericytes, and macrophages. The published study recovered 48,886 nuclei after quality control, demonstrating that snRNA-seq can capture the full cellular diversity of a complex tissue when the isolation protocol is appropriate.
Your decision framework should include a list of the cell types you need to recover and an assessment of whether each is likely to survive your chosen isolation method. If you need to recover fragile cell types, snRNA-seq is generally the safer choice. If you need to capture cytoplasmic transcripts that are lost in nuclear preparations, scRNA-seq may be necessary despite the higher risk of cell loss.
Step 3: Determine Whether Cytoplasmic Transcript Information Is Essential
The most fundamental biological difference between scRNA-seq and snRNA-seq is the RNA compartment captured. Whole-cell protocols capture both cytoplasmic and nuclear transcripts, while nuclear protocols capture primarily nuclear RNA. This difference matters for specific research questions.
If your research question depends on cytoplasmic mRNA species, such as secreted factors or signaling molecules that are rapidly exported from the nucleus, scRNA-seq may be necessary. The blood sample study noted that nuclear protocols capture only nuclear transcripts, yet the CL-derived samples maintained similar cell-type proportions and gene expression as matched peripheral blood mononuclear cell samples. This finding suggests that for many cell types, nuclear transcriptomes are representative of the overall cellular state. However, the study also found that granulocytes were retained in the nuclear data, which is a distinct advantage because granulocytes are often lost during standard whole-cell processing.
For research questions that focus on transcriptional regulation, splicing, or nascent RNA, snRNA-seq may actually be advantageous because nuclear RNA is enriched for intronic and nascent transcripts. The aldosteronoma study used snRNA-seq to analyze 140,742 cells from 28 samples and performed pseudotime trajectory analysis and snATAC-seq analysis to trace cell lineage. This type of regulatory analysis benefits from the nuclear transcriptome's enrichment for nascent RNA.
Your decision framework should include a specific list of genes or pathways that are central to your research question. If those genes are known to be enriched in the cytoplasm or rapidly exported from the nucleus, scRNA-seq is the appropriate choice. If your genes of interest are well represented in nuclear RNA, snRNA-seq offers the advantage of working with frozen or archived samples.
Step 4: Consider Your Pooling and Demultiplexing Strategy
The heart failure studies used a pooling strategy that has direct implications for your QC pipeline. Both the preprint and the published version isolated nuclei from combined samples with six patients per pool. This approach reduces cost and technical variability but requires genotype-based demultiplexing to assign nuclei to individual patients.
The preprint reported assigning more than 75 percent of droplets to individual patients using souporcell, while the published version reported more than 70 percent of nuclei assigned. This difference in assignment rates between the preprint and published version may reflect differences in the final analysis pipeline or in the stringency of the assignment criteria. For your own analysis, you should expect that a fraction of nuclei will not be assignable to individual patients and will need to be excluded from patient-level analyses.
Your decision framework should include an assessment of whether pooling is appropriate for your study. Pooling reduces cost and technical variability but introduces the need for genotype-based demultiplexing and reduces the number of nuclei available per patient. If you need high per-patient nuclei counts for rare cell-type analysis, individual sample processing may be preferable despite the higher cost.
Step 5: Match Your QC Pipeline to Your Assay Choice
Once you have chosen your assay, your QC pipeline should be tailored accordingly. The following table summarizes the key QC adjustments for each assay choice.
| Decision Point | scRNA-seq Approach | snRNA-seq Approach | Practical Implication |
|---|---|---|---|
| Sample storage | Fresh tissue required | Fresh or frozen tissue | Frozen samples force snRNA-seq choice |
| Mitochondrial fraction | Cell-health indicator, threshold 10 to 20 percent | Contamination indicator, threshold 1 to 5 percent | Different biological meaning requires different threshold |
| Gene detection | Higher median genes per cell | Lower median genes per nucleus | Lower minimum gene threshold for nuclear data |
| Ambient RNA | Cytoplasmic transcripts from dissociation | Buffer and lysis-derived transcripts | More aggressive ambient RNA correction for nuclear data |
| Doublet detection | Based on combined transcriptomes | Less distinct due to lower RNA content | Genotype-based demultiplexing preferred for pooled samples |
| Cell-type recovery | May lose fragile cell types | Better for fragile cell types | Verify cell-type retention after filtering |
Step 6: Plan for Validation Experiments
Your decision framework should include a validation plan that confirms your assay choice is appropriate for your research question. The Lynch Syndrome study provides a model for this validation. The researchers used immuno-histochemical staining to validate three proposed key markers identified through their single-cell and single-nucleus analysis. This validation step confirmed that the transcriptional signals detected in the sequencing data corresponded to protein expression in the tissue.
For your own study, plan to validate at least a subset of your key findings using an orthogonal method. This could include immuno-histochemical staining, quantitative PCR, or spatial transcriptomics. The macaque claustrum study used single-cell spatial transcriptome analysis to complement their single-nucleus data, providing spatial context for the transcriptional states they identified. If spatial information is important for your research question, consider whether spatial transcriptomics should be part of your validation plan.
Step 7: Document Your Decision Rationale
The final step in your decision framework is documentation. Record the reasons for your assay choice, the sample storage conditions, the isolation protocol, and the expected cell-type composition. This documentation serves two purposes. First, it helps you interpret unexpected QC results by providing context for why certain metric distributions look the way they do. Second, it supports reproducibility by allowing other researchers to understand the choices you made and to compare their results to yours.
The EMBL-EBI Training provides learning pathways for bioinformatics data analysis that include guidance on documenting analysis decisions. The Galaxy Training Network offers accessible workflow training that emphasizes reproducibility through documented analysis steps. These resources can help you build a documentation framework that supports your QC pipeline.
A Record System for Tracking QC Decisions Across Assay Types
Maintaining a structured record system is essential when you are working with both scRNA-seq and snRNA-seq data, particularly if you are integrating the two data types. The Lynch Syndrome study performed integrative computational analysis of scRNA-seq and snRNA-seq data, which requires careful tracking of how each dataset was processed.
Sample Tracking Table
Create a table that tracks each sample through your analysis pipeline. Include columns for sample identifier, tissue type, storage condition, assay type, isolation method, pooling status, sequencing depth, and the number of cells or nuclei captured before and after filtering. The heart failure studies provide a model for this documentation, reporting the number of patients per pool, the total nuclei recovered after QC, and the number of cell types identified.
Threshold Decision Log
For each sample, record the threshold values you applied for minimum UMI count, minimum gene count, and maximum mitochondrial fraction. Also record the rationale for each threshold decision. If you adjusted a threshold because a specific cell type was being removed, record that observation. This log is essential for troubleshooting when your downstream results are unexpected.
Ambient RNA Correction Record
Record the ambient RNA correction method you used for each sample, including the software version and any parameters you adjusted. The heart failure studies used CellBender for ambient RNA removal after CellRanger quantification. The blood sample study used CRISPR-guided globin gene depletion of complementary DNA to reduce ambient globin gene counts. Both approaches should be documented in your records.
Doublet Detection Record
Record the doublet detection method you used and the proportion of cells or nuclei flagged as doublets. If you used genotype-based demultiplexing, record the assignment rate and the number of nuclei that could not be assigned to individual patients. The heart failure studies reported assignment rates of more than 75 percent in the preprint and more than 70 percent in the published version, providing a reference range for your own expectations.
Cell-Type Retention Record
After clustering and annotation, record the cell-type composition of your filtered data. Compare this to your expectations based on the tissue type and to any published reference data. The heart failure studies identified 14 cell types in myocardial tissue, providing a reference for expected cell-type diversity in that tissue.
Troubleshooting Method for Unexpected QC Metric Distributions
When your QC metric distributions do not match your expectations, a systematic troubleshooting approach can help you identify the cause and determine whether your data are usable.
Step 1: Compare to Protocol-Specific Expectations
The first step is to determine whether your metric distributions are consistent with your isolation protocol. The blood sample study demonstrated that different isolation methods produce very different results. Mechanical separation produced lower nuclei yields and more biased immune cell proportions, while cell lysis produced higher yields and less bias. If your metric distributions are unexpected, first check whether they are consistent with your isolation method.
Step 2: Examine Sample-Level Variation
If you processed multiple samples, examine whether the unexpected distribution is present in all samples or only in specific ones. Sample-specific patterns may indicate problems with a particular isolation batch or sequencing run. The heart failure studies pooled samples from multiple patients, which introduces batch structure that must be accounted for in the analysis.
Step 3: Check for Dominant Ambient Transcripts
If your gene detection rates are lower than expected, check whether a small number of genes dominate your ambient profile. The blood sample study found that globin genes dominated the ambient profile after cell lysis, requiring active depletion to recover sensitivity. If your tissue has a dominant ambient transcript, you may need to apply targeted depletion or more aggressive computational correction.
Step 4: Verify Cell-Type Representation
If your filtered data lack a cell type that is known to be present in your tissue, this may indicate that your thresholds are too stringent or that your isolation protocol is biased against that cell type. The blood sample study found that cell lysis produced less biased proportions of immune cells than mechanical separation, suggesting that isolation method can systematically affect cell-type representation.
Step 5: Consult Published Reference Data
Compare your metric distributions to published snRNA-seq studies from the same or similar tissues. The heart failure studies provide reference distributions for myocardial tissue, and the aldosteronoma study provides reference distributions for adrenal tissue. The NCBI Data Resources provide access to published sequencing data that can serve as references for your own analysis.
Step 6: Escalate When Necessary
If your troubleshooting does not resolve the issue, escalate to a bioinformatics specialist or statistician. The Bioconductor project provides official documentation for reproducible genomic analysis, and the nf-core documentation describes community standards for reproducible workflows. These resources can help you identify specialized methods for challenging datasets.
Common Failure Patterns When Transitioning Between Assay Types
The transition from scRNA-seq to snRNA-seq introduces specific failure patterns that are distinct from those encountered when working with a single assay type.
Pattern 1: Assuming Metric Distributions Are Comparable
Researchers often assume that the same QC thresholds will work for both assay types. This assumption fails because nuclear RNA is less abundant than total cellular RNA, producing lower gene detection rates and lower UMI counts. The blood sample study demonstrated that even within snRNA-seq, different isolation methods produce substantially different yields and cell-type proportions, highlighting the need for protocol-specific thresholds.
Pattern 2: Misinterpreting Mitochondrial Fraction
Applying whole-cell mitochondrial thresholds to nuclear data produces misleading results. Since nuclei lack mitochondria, the mitochondrial fraction in snRNA-seq data reflects contamination instead of cell health. The appropriate threshold is much lower, typically between 1 and 5 percent, and should be based on the contamination distribution in your specific data.
Pattern 3: Ignoring Isolation Method Effects
The choice of isolation method has a substantial impact on your data quality and cell-type representation. The blood sample study found that cell lysis produced up to two orders of magnitude higher nuclei yields than mechanical separation and less biased immune cell proportions. If you change your isolation method between samples, you introduce batch effects that must be accounted for in your analysis.
Pattern 4: Inadequate Ambient RNA Correction
Ambient RNA correction is more critical for snRNA-seq than for scRNA-seq because nuclear isolation buffers can contain transcripts from lysed cells, and the lower RNA content of nuclei makes ambient contamination proportionally more impactful. The blood sample study demonstrated that ambient globin genes could dominate the signal and reduce gene detection sensitivity, requiring active depletion to recover.
Pattern 5: Failing to Validate Cell-Type Retention
Applying thresholds without checking which cell types are removed can systematically eliminate rare or low-transcription cell types. This is particularly problematic when integrating scRNA-seq and snRNA-seq data, as the Lynch Syndrome study did, because the two assays may retain different cell populations.
Professional Escalation Criteria for Cross-Assay Analyses
When you are working with both scRNA-seq and snRNA-seq data, specific situations warrant escalation to a specialist.
Escalate When Integration Produces Artificial Populations
If your integration of scRNA-seq and snRNA-seq data produces cell populations that do not appear in either dataset individually, this may indicate that the technical differences between the assays are not being adequately handled. The Lynch Syndrome study performed integrative computational analysis of both data types, and the success of this approach depends on appropriate handling of the technical differences.
Escalate When Cell-Type Proportions Differ Dramatically Between Assays
If your matched scRNA-seq and snRNA-seq samples show dramatically different cell-type proportions, this may indicate that one assay is biased against specific cell types. The blood sample study found that CL-derived samples maintained similar cell-type proportions as matched peripheral blood mononuclear cell samples, providing a reference for expected concordance.
Escalate When Ambient RNA Correction Is Insufficient
If computational ambient RNA correction does not adequately remove contamination, you may need to consider targeted depletion approaches. The blood sample study used CRISPR-guided globin gene depletion to reduce ambient globin counts, which is a library preparation approach that requires specialized expertise.
Escalate When You Cannot Determine Appropriate Thresholds
If your metric distributions do not show clear separation between nuclei and empty droplets, or if you cannot identify appropriate thresholds for your data, escalate to a specialist. The Carpentries lessons provide foundational training in data handling that can help you build the skills needed to make these decisions, but some situations require specialized expertise.
Frequently Asked Questions
Why is the mitochondrial fraction so low in my single-nucleus data?
Nuclei do not contain mitochondria, so mitochondrial transcripts in snRNA-seq data come from ambient contamination or from mitochondria attached to the nuclear envelope during isolation. The low mitochondrial fraction is expected and should be interpreted as a contamination metric instead of a cell-health metric. Set your threshold based on the distribution in your data instead of applying whole-cell thresholds.
Should I use the same gene count threshold for single-nucleus and single-cell data?
No. Nuclear RNA is less abundant than total cellular RNA, so you should expect lower median genes per nucleus compared to cells from the same tissue. Set your minimum gene threshold based on the nuclear distribution and verify that your threshold does not remove specific cell types with lower transcriptional activity.
How do I handle ambient RNA in single-nucleus data?
Ambient RNA correction is more critical for snRNA-seq than for scRNA-seq because nuclear isolation buffers can contain transcripts from lysed cells. The heart failure studies used CellBender for ambient RNA removal after CellRanger quantification. For tissues with dominant ambient transcripts, such as globin genes in blood, targeted depletion during library preparation may be necessary.
What is the best way to detect doublets in single-nucleus data?
Genotype-based demultiplexing is a robust approach when you pool samples from multiple individuals. The heart failure studies used souporcell for this purpose and successfully assigned more than 70 percent of nuclei to individuals. If you are not pooling samples, use doublet detection tools calibrated for nuclear data and be aware that default parameters trained on whole-cell data may not perform optimally.
Can I integrate single-nucleus and single-cell data from the same tissue?
Yes, but integration requires careful handling of the technical differences between the two assays. The Lynch Syndrome study performed integrative computational analysis of scRNA-seq and snRNA-seq data from colorectal cancer samples. Verify that your QC decisions retain comparable cell populations in both datasets before integration.
Why does my single-nucleus data have lower gene detection than my single-cell data?
Nuclear RNA is enriched for intronic and nascent transcripts while lacking some cytoplasmic mRNA species. This biological difference results in lower gene detection per nucleus compared to whole cells. This is an expected property of the assay, not a quality failure. Adjust your expectations and thresholds accordingly.
How should I set mitochondrial thresholds for single-nucleus data?
Examine the distribution of mitochondrial fraction in your data and set the threshold based on the contamination profile. Many snRNA-seq pipelines use thresholds between 1 and 5 percent, but the appropriate value depends on your isolation protocol and tissue type. Verify that your threshold does not disproportionately remove specific cell types.
What records should I keep for reproducible single-nucleus analysis?
Document your isolation protocol, sequencing depth, software versions, threshold decisions, and metric distributions for each sample. The Bioconductor project and nf-core documentation provide standards for reproducible genomic analysis. This documentation is essential for troubleshooting and for publishing your results.
Related Bioinformatics Guides
- Single-Cell RNA Sequencing Quality Control: A Practical Guide to Filtering and Metrics
- RNA-Seq Quality Control: Essential Checks and Tools
- Single-Cell Sequencing Depth: How Much Is Enough?
- Single-Cell Sequencing Analysis Pipeline: From Raw Data to Biological Insights
- Single-Cell vs Single-Nucleus RNA Sequencing: Choosing the Right Approach
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- Single cell transcriptomic analyses of human heart failure with preserved ejection fraction.. bioRxiv : the preprint server for biology, 2025.
- Single-Cell Analysis of Human Heart Failure With Preserved Ejection Fraction.. Circulation research, 2026.
- Tracing the stemness and malignant transition in a heritable colorectal cancer Lynch Syndrome by single-cell RNA-seq analysis.. 2026.
- Comprehensive cellular analysis with single-nucleus RNA-seq of archived PAXgene whole blood samples.. 2026.
- Divergent surgical outcomes for CACNA1D- and KCNJ5-mutant aldosteronomas are traceable to their cell-of-origin and most differentially expressed gene, CCM2L. 2026.
- popsicleR: A R Package for Pre-processing and Quality Control Analysis of Single Cell RNA-seq Data.. Journal of Molecular Biology, 2022.
- Research on the role of the key gene RhoJ in human limb venous malformation endothelial cells using single-nucleus RNA sequencing technology. Chinese Journal of Plastic Surgery, 2025.
- Single-cell spatial transcriptome atlas and whole-brain connectivity of the macaque claustrum. Cell, 2025.
- Alzheimer’s disease and diabetes-associated cognitive dysfunction: the microglia link?. Metabolic Brain Disease, 2025.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.