A Step-by-Step Tutorial for Quality Control and Filtering of Single-Cell RNA-Seq Data in R (Seurat)
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Single-cell RNA-sequencing (scRNA-seq) data requires rigorous quality control (QC) to mitigate technical noise, low-quality cells, and unwanted variation before downstream analysis. Key QC metrics include the number of detected genes (nFeature_RNA), total UMI counts (nCount_RNA), and the percentage of mitochondrial reads (percent.mt).
- High percent.mt values, often exceeding 10-20%, are indicative of damaged cells that have lost cytoplasmic mRNA, while low nFeature_RNA and nCount_RNA can signal empty droplets or failed library preparation.
- Doublets, where two or more cells are captured and sequenced as one, present as cells with abnormally high nCount_RNA and nFeature_RNA, and can be computationally identified using methods like scDblFinder by comparing real cells to simulated doublets.
- Threshold selection for filtering must be dataset-specific, balancing the removal of technical artifacts with the preservation of biological signal, and should be informed by visualizations such as violin plots and scatter plots of metric distributions.
- Reproducibility in scRNA-seq QC is paramount, necessitating detailed documentation of all threshold decisions, rationale, and cell counts at each filtering stage, alongside careful consideration of biological variation between cell types and experimental batches.
Single-cell RNA sequencing (scRNA-seq) generates gene expression measurements for thousands of individual cells in a single experiment, but the resulting data contain substantial technical noise, low-quality cells, and unwanted sources of variation that must be identified and removed before downstream analysis. This tutorial provides a practical, code-based workflow for performing quality control (QC) and filtering in R using the Seurat package, with detailed annotations for calculating QC metrics, visualizing distributions, selecting thresholds, and removing doublets. The workflow assumes basic R knowledge and is designed for biology students, researchers, and laboratory professionals who need reproducible decision criteria instead of arbitrary cutoff values.
The approach presented here follows the canonical analytic workflow described in the single-cell literature, which includes read mapping, quality controls, gene expression quantification, normalization, feature selection, dimensionality reduction, and cell clustering [<a href="#ref-1">1</a>]. Quality control is the first critical step after obtaining a count matrix, and the decisions made at this stage directly affect all subsequent analyses. Poor filtering can retain damaged cells that distort clustering, while overly aggressive filtering can remove legitimate cell populations, particularly rare or large cell types.
This tutorial uses Seurat, an R package widely adopted for single-cell analysis. The workflow covers loading data, calculating QC metrics, visualizing metric distributions, selecting filtering thresholds, applying filters, and addressing doublets. Each step includes code, interpretation guidance, and common failure patterns to help you make informed decisions with your own datasets.
Understanding Single-Cell RNA-Seq Data Structure and Quality Issues
What the Count Matrix Contains
The starting point for scRNA-seq analysis is a count matrix where rows represent genes and columns represent cells. Each entry indicates the number of unique molecular identifiers (UMIs) or reads mapped to a given gene in a given cell. This matrix can be generated by various pipelines, including Cell Ranger from 10X Genomics, or from any scRNA-seq technology that produces a feature-barcode matrix of raw counts [<a href="#ref-2">2</a>]. The raw count matrix contains biological signal mixed with technical artifacts, and the goal of QC is to distinguish genuine cellular states from technical failures.
Single-cell technologies measure gene expression at individual cell resolution, which creates opportunities for resolving cellular heterogeneity but also introduces high data complexity [<a href="#ref-1">1</a>]. Unlike bulk RNA-seq, where signals are averaged across many cells, scRNA-seq data are sparse, with many genes showing zero counts in individual cells due to dropout events. This sparsity, combined with technical variation in capture efficiency and sequencing depth, means that raw counts cannot be directly compared across cells without normalization and QC.
Sources of Technical Artifacts
Low-quality cells in scRNA-seq data typically arise from several sources. Cells that are damaged during dissociation may lose cytoplasmic mRNA while retaining genomic DNA, leading to low total counts and high mitochondrial read fractions. Empty droplets or wells may contain ambient RNA from the surrounding solution, producing low counts with no clear biological identity. Doublets occur when two or more cells are captured together and sequenced as a single barcode, creating cells with abnormally high counts and mixed gene expression profiles.
The presence of technical artifacts, noise, and biological biases requires researchers to first identify and eventually remove unreliable signals from low-quality cells and unwanted sources of variation [<a href="#ref-2">2</a>]. This process is laborious because it requires the manual combination of different computational strategies to quantify QC metrics and define optimal sets of preprocessing parameters [<a href="#ref-2">2</a>]. The workflow presented here systematizes this process while preserving the flexibility needed for different biological contexts.
Why Quality Control Matters for Downstream Analysis
Quality control directly influences the reliability of every downstream step, including normalization, feature selection, dimensionality reduction, clustering, and differential expression. Low-quality cells can form spurious clusters that are driven by technical artifacts instead of biological differences. Doublets can create artificial cell states that appear intermediate between two genuine populations. Cells with extreme mitochondrial content may reflect stress responses or apoptosis instead of true biological states.
The single-cell field lacks universal standardization, which reflects the immaturity of the field but can also encumber newcomers [<a href="#ref-1">1</a>]. Different studies use different filtering thresholds, and the optimal parameters depend on tissue type, dissociation protocol, sequencing platform, and biological questions. This tutorial provides a framework for making these decisions transparently and reproducibly instead of relying on arbitrary defaults.
Setting Up the R Environment and Required Packages
Installing R and RStudio
R is a free, open-source programming language for statistical computing and graphics. RStudio provides an integrated development environment that simplifies script management, visualization, and package installation. If you do not have R and RStudio installed, download them from the official R project website and RStudio website before proceeding.
The Carpentries offers foundational computing and data lessons that cover R programming basics, data structures, and reproducible analysis practices [<a href="#ref-3">3</a>]. If you are new to R, completing an introductory lesson before starting this tutorial will help you understand the code structure and avoid common syntax errors.
Installing Seurat and Supporting Packages
Seurat is available through CRAN and the Satija Lab GitHub repository. The official installation instructions recommend using the remotes package to install the development version, which includes the latest features and bug fixes. For most users, installing from CRAN provides a stable version suitable for standard analyses.
install.packages("Seurat")
Additional packages used in this workflow include tidyverse for data manipulation and visualization, ggplot2 for plotting, and DoubletFinder or scDblFinder for doublet detection. Install these packages using the standard install.packages function.
install.packages("tidyverse")
install.packages("ggplot2")
Bioconductor provides a repository for genomics-focused R packages, including many single-cell analysis tools [<a href="#ref-4">4</a>]. Some doublet detection methods and additional QC packages are available through Bioconductor and can be installed using the BiocManager package.
if (!requireNamespace("BiocManager", quietly = TRUE))
install.packages("BiocManager")
BiocManager::install("scDblFinder")
Loading the Required Libraries
After installation, load the required libraries at the beginning of your script. This ensures that all functions are available and helps document your analysis environment.
library(Seurat)
library(tidyverse)
library(ggplot2)
Loading Single-Cell RNA-Seq Data into Seurat
Supported Input Formats
Seurat can read data from multiple sources. The most common input is the output of the Cell Ranger pipeline from 10X Genomics, which produces a directory containing barcodes.tsv, features.tsv, and matrix.mtx files [<a href="#ref-2">2</a>]. Seurat provides the Read10X function to load this format directly.
counts <- Read10X(data.dir = "path/to/filtered_feature_bc_matrix")
For data from other technologies or custom pipelines, you can load a count matrix from a CSV or TSV file using standard R functions and then create a Seurat object. The key requirement is that the matrix contains raw integer counts, not normalized or log-transformed values.
counts <- read.csv("count_matrix.csv", row.names = 1)
counts <- as(as.matrix(counts), "dgCMatrix")
Creating a Seurat Object
The CreateSeuratObject function initializes a Seurat object from the count matrix. The min.cells and min.features parameters provide initial filtering by removing genes detected in fewer than a specified number of cells and cells with fewer than a specified number of genes. These parameters remove the most obvious empty droplets and undetected genes but should be set conservatively to avoid removing legitimate cell populations.
seurat_obj <- CreateSeuratObject(
counts = counts,
project = "my_project",
min.cells = 3,
min.features = 200
)
The min.cells = 3 parameter removes genes detected in fewer than three cells, which reduces noise from genes that are essentially absent. The min.features = 200 parameter removes cells with fewer than 200 detected genes, which eliminates the most extreme low-quality cells and empty droplets. These values are starting points, not final filtering thresholds.
Understanding the Seurat Object Structure
A Seurat object contains the raw count matrix, metadata for each cell, and slots for normalized data, scaled data, and dimensionality reduction results. The metadata can be accessed and modified using standard R syntax. Understanding this structure helps you track which cells remain after each filtering step.
## View cell metadata
head([email protected])
## View dimensions of the count matrix
dim(seurat_obj)
The metadata data frame contains columns for nCount_RNA (total UMI counts per cell) and nFeature_RNA (number of genes detected per cell). These are the primary metrics used for quality control.
Calculating Quality Control Metrics
Percentage of Mitochondrial Reads
The percentage of reads mapping to mitochondrial genes is a key indicator of cell quality. High mitochondrial content typically indicates that cytoplasmic mRNA has been lost from damaged cells, leaving a higher proportion of mitochondrial transcripts. This metric is calculated by identifying mitochondrial genes, which are annotated with the MT- prefix in human data and mt- prefix in mouse data.
seurat_obj[["percent.mt"]] <- PercentageFeatureSet(seurat_obj, pattern = "^MT-")
For mouse data, use the pattern "^mt-" instead. For other organisms, consult the gene annotation to identify mitochondrial genes. The NCBI provides official descriptions of gene annotation resources and sequence databases that can help you identify mitochondrial genes for your organism of interest [<a href="#ref-5">5</a>].
Total UMI Counts per Cell
The nCount_RNA metric represents the total number of UMIs detected in each cell. This metric reflects sequencing depth and capture efficiency. Cells with very low counts may be empty droplets or damaged cells, while cells with extremely high counts may be doublets or multiplets.
Number of Detected Genes per Cell
The nFeature_RNA metric represents the number of genes with at least one UMI detected in each cell. This metric reflects cellular complexity and is correlated with cell type. Large cells typically have more detected genes than small cells, and this biological variation must be considered when setting thresholds.
Additional Quality Metrics
Depending on your data and research question, you may calculate additional QC metrics. The percentage of reads mapping to ribosomal genes can indicate ribosome contamination. The percentage of reads mapping to hemoglobin genes is relevant for blood or tissue samples with red blood cell contamination. The ratio of UMIs to genes detected can help identify low-complexity cells.
seurat_obj[["percent.ribo"]] <- PercentageFeatureSet(seurat_obj, pattern = "^RP[SL]")
These additional metrics provide context for interpreting the primary QC metrics and can help identify specific technical artifacts.
Visualizing Quality Control Metrics
Violin Plots for Metric Distributions
Violin plots show the distribution of each QC metric across all cells. Seurat provides the VlnPlot function for this purpose. These plots help you identify the range of values and spot outliers before setting filtering thresholds.
VlnPlot(seurat_obj, features = c("nFeature_RNA", "nCount_RNA", "percent.mt"), ncol = 3)
Interpret the violin plots by examining the shape of each distribution. A typical dataset shows a main peak representing healthy cells, with a tail of low-quality cells at low counts and high mitochondrial percentages. The width of the violin indicates the density of cells at each value.
Scatter Plots for Metric Relationships
Scatter plots reveal relationships between QC metrics. The FeatureScatter function creates these plots and can help identify doublets, which typically show high nCount_RNA and nFeature_RNA values simultaneously.
plot1 <- FeatureScatter(seurat_obj, feature1 = "nCount_RNA", feature2 = "nFeature_RNA")
plot2 <- FeatureScatter(seurat_obj, feature1 = "nCount_RNA", feature2 = "percent.mt")
plot1 + plot2
In a healthy dataset, nCount_RNA and nFeature_RNA show a positive correlation, with cells falling along a diagonal band. Cells that deviate from this band may be problematic. Cells with high nCount_RNA but low nFeature_RNA may be low-complexity cells or contain ambient RNA. Cells with high values for both metrics may be doublets.
Histograms for Threshold Selection
Histograms provide a complementary view of metric distributions and can help you identify natural breakpoints for filtering thresholds. The ggplot2 package provides flexible histogram creation.
ggplot([email protected], aes(x = nFeature_RNA)) +
geom_histogram(bins = 100) +
geom_vline(xintercept = 500, color = "red") +
theme_minimal()
Examine the histogram for each metric and identify where the distribution drops off. The goal is to remove the tail of low-quality cells while retaining the main population. Natural breakpoints in the distribution provide more defensible thresholds than arbitrary values.
Selecting Filtering Thresholds
Principles for Threshold Selection
Threshold selection requires balancing sensitivity and specificity. Aggressive filtering removes more low-quality cells but risks removing legitimate cell populations. Conservative filtering retains more cells but leaves technical artifacts that can distort downstream analysis. The optimal thresholds depend on your tissue type, dissociation protocol, and biological questions.
The single-cell field lacks universal standardization, and different studies use different thresholds [<a href="#ref-1">1</a>]. This means you must justify your thresholds based on your data instead of relying on published defaults. The visualization steps above provide the evidence needed to make these decisions transparently.
Minimum Number of Detected Genes
The minimum number of detected genes per cell is the most common filtering threshold. Cells with very few detected genes are likely empty droplets or severely damaged cells. A common starting point is 200 to 500 genes, but the appropriate value depends on your cell types.
Small cells such as T cells and B cells naturally express fewer genes than large cells such as macrophages or neurons. If your dataset contains small cell types, a threshold of 500 genes may remove them. Examine the nFeature_RNA distribution and identify the natural breakpoint before setting this threshold.
seurat_obj <- subset(seurat_obj, subset = nFeature_RNA > 500)
Maximum Number of Detected Genes
The maximum number of detected genes helps remove doublets and multiplets. Cells with abnormally high gene counts are likely to contain two or more cells. However, some legitimate cell types have high gene counts, so this threshold must be set carefully.
A common approach is to set the maximum threshold at a value above the main distribution but below the obvious doublet population. Examine the histogram and scatter plots to identify where the doublet population begins.
seurat_obj <- subset(seurat_obj, subset = nFeature_RNA < 6000)
Percentage of Mitochondrial Reads
The maximum percentage of mitochondrial reads is a critical threshold for removing damaged cells. A common starting point is 10% to 20%, but the appropriate value depends on tissue type and dissociation protocol.
Some tissues, such as heart and skeletal muscle, naturally have high mitochondrial content. Cells from these tissues may show mitochondrial percentages above 20% even when healthy. Conversely, some dissociation protocols cause more cellular stress, leading to higher mitochondrial content across all cells.
seurat_obj <- subset(seurat_obj, subset = percent.mt < 20)
Total UMI Counts
The minimum total UMI count is sometimes used as an additional filter. Cells with very low UMI counts may have failed library preparation or sequencing. However, this metric is correlated with nFeature_RNA, and filtering on both metrics can be redundant.
If you choose to filter on nCount_RNA, set the threshold based on the distribution and consider the relationship with nFeature_RNA. Cells with low UMI counts but reasonable gene counts may be small cells instead of low-quality cells.
Documenting Threshold Decisions
Record your threshold values and the rationale for each decision. This documentation is essential for reproducibility and for defending your analysis in publications or presentations. The Galaxy Training Network emphasizes the importance of accessible workflow training and reproducibility in bioinformatics analysis [<a href="#ref-6">6</a>]. Documenting your QC decisions follows the same principle.
## Record thresholds in a text file
writeLines(c(
"QC thresholds for dataset: my_project",
paste0("Date: ", Sys.Date()),
"min.features: 500",
"max.features: 6000",
"max.percent.mt: 20"
), con = "qc_thresholds.txt")
Applying Filters and Verifying Results
Subsetting the Seurat Object
The subset function applies filtering thresholds to the Seurat object. You can combine multiple criteria in a single subset call or apply them sequentially. Sequential filtering allows you to examine the effect of each threshold on cell numbers.
seurat_obj_filtered <- subset(
seurat_obj,
subset = nFeature_RNA > 500 & nFeature_RNA < 6000 & percent.mt < 20
)
Checking Cell Counts Before and After Filtering
Compare the number of cells before and after filtering to understand the impact of your thresholds. A typical dataset retains 70% to 90% of cells after QC filtering, but this proportion varies widely depending on sample quality and threshold stringency.
cat("Cells before filtering:", ncol(seurat_obj), "\n")
cat("Cells after filtering:", ncol(seurat_obj_filtered), "\n")
cat("Percentage retained:", round(ncol(seurat_obj_filtered) / ncol(seurat_obj) * 100, 1), "%\n")
If you retain fewer than 50% of cells, your thresholds may be too aggressive. If you retain more than 95%, you may not be removing enough low-quality cells. Re-examine the distributions and adjust thresholds accordingly.
Re-plotting Metrics After Filtering
After filtering, re-plot the QC metrics to verify that the remaining cells form a clean distribution. The violin plots and scatter plots should show a more uniform population without the tails of low-quality cells.
VlnPlot(seurat_obj_filtered, features = c("nFeature_RNA", "nCount_RNA", "percent.mt"), ncol = 3)
Compare these plots to the pre-filtering plots to confirm that the filtering removed the intended cells. The distributions should be narrower and more symmetric, with fewer extreme outliers.
Examining the Effect on Gene Expression
Filtering removes cells, which changes the set of genes detected in the dataset. After filtering, re-examine the number of genes detected and the total UMI counts to ensure that the remaining data are sufficient for downstream analysis.
## Check dimensions after filtering
dim(seurat_obj_filtered)
## Check average metrics after filtering
summary([email protected]$nFeature_RNA)
summary([email protected]$percent.mt)
Normalization and Scaling After Filtering
Log Normalization
After QC filtering, the next step is normalization. Seurat provides the NormalizeData function, which applies log normalization by default. This step divides each cell's counts by the total counts, multiplies by a scale factor of 10,000, and applies a natural log transformation.
seurat_obj_filtered <- NormalizeData(seurat_obj_filtered, normalization.method = "LogNormalize", scale.factor = 10000)
Normalization is essential because raw counts are not directly comparable across cells with different sequencing depths. The log transformation stabilizes variance and makes the data more suitable for downstream statistical analysis.
Identification of Highly Variable Features
After normalization, identify genes that show high variability across cells. These genes are likely to be biologically informative and are used for dimensionality reduction and clustering. Seurat provides the FindVariableFeatures function for this purpose.
seurat_obj_filtered <- FindVariableFeatures(seurat_obj_filtered, selection.method = "vst", nfeatures = 2000)
The default selection method uses variance stabilizing transformation to identify the 2,000 most variable genes. This step reduces the dimensionality of the data and focuses downstream analysis on informative genes.
Scaling the Data
Scaling transforms the data so that each gene has a mean of zero and a variance of one across cells. This step ensures that genes with high expression do not dominate the analysis. Seurat provides the ScaleData function for this purpose.
seurat_obj_filtered <- ScaleData(seurat_obj_filtered)
Scaling is a critical step before dimensionality reduction because principal component analysis is sensitive to the scale of the input features. Without scaling, highly expressed genes would dominate the principal components.
Doublet Detection and Removal
What Are Doublets and Why Do They Matter
Doublets occur when two or more cells are captured together and sequenced under a single barcode. They create artificial cell states that can confound clustering and differential expression analysis. Doublets are particularly problematic in datasets with diverse cell types because they can appear as intermediate states between two genuine populations.
The frequency of doublets depends on the loading density of the microfluidic device or droplet generator. Higher loading densities increase the probability of doublet formation. Most protocols provide expected doublet rates based on the number of cells loaded, but the actual rate varies between experiments.
Computational Doublet Detection Methods
Several computational methods can identify doublets based on gene expression patterns. These methods simulate artificial doublets by combining the expression profiles of random cell pairs and then train a classifier to distinguish real cells from simulated doublets.
The scDblFinder package, available through Bioconductor, provides a widely used approach for doublet detection [<a href="#ref-4">4</a>]. It uses a combination of clustering and classification to identify cells that resemble simulated doublets.
library(scDblFinder)
sce <- as.SingleCellExperiment(seurat_obj_filtered)
sce <- scDblFinder(sce)
seurat_obj_filtered$doublet_score <- sce$scDblFinder.score
seurat_obj_filtered$doublet_class <- sce$scDblFinder.class
Setting Doublet Removal Thresholds
After calculating doublet scores, examine the distribution and set a threshold for removal. The expected doublet rate provides a useful reference. If you loaded 10,000 cells with an expected doublet rate of 5%, you would expect approximately 500 doublets.
## Examine doublet score distribution
ggplot([email protected], aes(x = doublet_score)) +
geom_histogram(bins = 100) +
theme_minimal()
## Remove predicted doublets
seurat_obj_filtered <- subset(seurat_obj_filtered, subset = doublet_class == "singlet")
Limitations of Doublet Detection
Computational doublet detection methods have limitations. They may miss doublets formed by two cells of the same type, which have expression profiles similar to a single cell of that type. They may also incorrectly classify large or highly complex cells as doublets. The machine learning approaches used for doublet detection require training data that is similar to the test data, and their accuracy depends on the cell type composition of the dataset [<a href="#ref-7">7</a>].
Consider doublet detection results as a guide instead of a definitive classification. If you have prior knowledge about expected cell types in your sample, use this knowledge to evaluate whether the doublet calls are biologically plausible.
Common Failure Patterns in Quality Control
Overly Aggressive Filtering
The most common mistake in scRNA-seq QC is setting thresholds too aggressively, which removes legitimate cell populations. This is particularly problematic for rare cell types or small cells with naturally low gene counts. If you remove too many cells, you may lose the biological signal you are trying to study.
Signs of overly aggressive filtering include retaining fewer than 50% of cells, losing known cell types, or seeing a sudden drop in the number of detected genes after filtering. If you observe these signs, relax your thresholds and re-examine the distributions.
Under-Filtering and Retaining Low-Quality Cells
The opposite failure is retaining too many low-quality cells. This can happen when thresholds are set too conservatively or when the distributions do not show clear breakpoints. Low-quality cells can form spurious clusters, increase noise, and distort differential expression results.
Signs of under-filtering include clusters with low gene counts and high mitochondrial percentages, cells with extreme values in scatter plots, and poor separation between clusters after dimensionality reduction. If you observe these signs, tighten your thresholds and consider additional QC metrics.
Ignoring Biological Variation in Thresholds
A common error is applying the same thresholds to all samples or cell types without considering biological variation. Different tissues and cell types have different baseline values for QC metrics. A threshold that works for peripheral blood mononuclear cells may not work for neurons or hepatocytes.
Examine QC metrics separately for each sample or condition in your dataset. If you have multiple samples, consider whether they have similar distributions before applying a common threshold. The practical guide for single-cell transcriptome data analysis in neuroscience emphasizes the importance of addressing the unique challenges of scRNA-seq data and illustrating methods for robust analysis [<a href="#ref-8">8</a>].
Failing to Document Threshold Decisions
Reproducibility requires documentation of all analysis decisions, including QC thresholds. Without documentation, you cannot defend your choices in publications or reproduce your analysis months later. The nf-core documentation emphasizes the importance of reproducible workflow standards in bioinformatics [<a href="#ref-9">9</a>].
Record the threshold values, the rationale for each choice, and the number of cells removed at each step. Include this information in your methods section or as supplementary material.
Records and Measurements for Quality Control
Maintaining an Analysis Log
Keep a detailed log of your QC analysis, including the date, software versions, input data, thresholds, and results. This log serves as a record of your decisions and supports reproducibility. The EMBL-EBI training resources emphasize the importance of structured learning pathways and practical analysis education in bioinformatics [<a href="#ref-10">10</a>].
A simple text file or R script with comments can serve as your analysis log. Include the following information:
- Date and analyst name
- R and Seurat versions
- Input data source and file paths
- Initial cell and gene counts
- QC metric distributions
- Threshold values and rationale
- Cell counts after each filtering step
- Doublet detection method and parameters
Tracking Cell Counts Through the Workflow
Create a table that tracks cell counts at each stage of the workflow. This table helps you identify where cells are lost and evaluate whether the losses are appropriate.
| Analysis Stage | Cells Remaining | Genes Detected | Median UMI Count | Median Genes per Cell | Median Percent MT |
|---|---|---|---|---|---|
| Raw data | 12,500 | 22,000 | 3,200 | 1,100 | 8.5 |
| After min.features filter | 11,800 | 21,500 | 3,400 | 1,200 | 7.8 |
| After max.features filter | 11,200 | 21,000 | 3,300 | 1,150 | 7.5 |
| After percent.mt filter | 10,400 | 20,500 | 3,500 | 1,200 | 5.2 |
| After doublet removal | 9,800 | 20,000 | 3,600 | 1,250 | 4.8 |
This table provides a clear summary of the filtering process and can be included in your methods or supplementary materials.
Comparing Samples and Batches
If your dataset contains multiple samples or batches, compare QC metrics across groups before filtering. Batch effects can cause systematic differences in QC metrics that are not related to cell quality. The Galaxy single-cell and spatial omics community has developed tools and training resources to help researchers perform and interpret their own analyses, including handling batch effects [<a href="#ref-11">11</a>].
Plot QC metrics grouped by sample or batch to identify systematic differences. If one sample shows consistently lower gene counts or higher mitochondrial percentages, this may indicate a technical problem with that sample instead of a biological difference.
At a Glance: Quality Control Decision Table
The following table summarizes the key QC metrics, their interpretation, and common threshold considerations. Use this table as a reference when making filtering decisions for your own datasets.
| QC Metric | What It Measures | Low Value Indicates | High Value Indicates | Threshold Considerations |
|---|---|---|---|---|
| nFeature_RNA | Number of genes detected per cell | Empty droplet, damaged cell, or small cell type | Doublet, multiplet, or large cell type | Set minimum based on distribution breakpoint, set maximum above main population but below doublet peak |
| nCount_RNA | Total UMI counts per cell | Low sequencing depth or failed library | Doublet or high capture efficiency | Correlated with nFeature_RNA, use as secondary filter |
| percent.mt | Percentage of reads mapping to mitochondrial genes | Healthy cell with intact cytoplasm | Damaged cell with lost cytoplasmic mRNA | Set maximum based on tissue type, 10% to 20% is common but varies |
| Doublet score | Probability that a cell is a doublet | Singlet cell | Doublet or multiplet | Use expected doublet rate as reference, examine distribution for breakpoint |
Practical Implementation Steps
Step 1: Load and Inspect Your Data
Load your count matrix and create a Seurat object. Examine the dimensions and basic metadata to understand the scale of your dataset. Check for obvious issues such as very low total counts or unusual gene annotations.
counts <- Read10X(data.dir = "path/to/filtered_feature_bc_matrix")
seurat_obj <- CreateSeuratObject(counts = counts, project = "my_project")
print(dim(seurat_obj))
Step 2: Calculate QC Metrics
Calculate the percentage of mitochondrial reads and any additional QC metrics relevant to your sample type. Add these metrics to the Seurat object metadata.
seurat_obj[["percent.mt"]] <- PercentageFeatureSet(seurat_obj, pattern = "^MT-")
Step 3: Visualize Metric Distributions
Create violin plots and scatter plots to examine the distributions of QC metrics. Identify the main population of healthy cells and the tails representing low-quality cells.
VlnPlot(seurat_obj, features = c("nFeature_RNA", "nCount_RNA", "percent.mt"), ncol = 3)
FeatureScatter(seurat_obj, feature1 = "nCount_RNA", feature2 = "nFeature_RNA")
Step 4: Select and Apply Thresholds
Based on the visualizations, select thresholds for each QC metric. Apply the filters using the subset function and record the number of cells removed.
seurat_obj_filtered <- subset(
seurat_obj,
subset = nFeature_RNA > 500 & nFeature_RNA < 6000 & percent.mt < 20
)
Step 5: Verify Filtering Results
Re-plot the QC metrics after filtering and compare the distributions to the pre-filtering plots. Confirm that the remaining cells form a clean population without extreme outliers.
VlnPlot(seurat_obj_filtered, features = c("nFeature_RNA", "nCount_RNA", "percent.mt"), ncol = 3)
Step 6: Detect and Remove Doublets
Run doublet detection using scDblFinder or another method. Examine the doublet score distribution and remove predicted doublets.
sce <- as.SingleCellExperiment(seurat_obj_filtered)
sce <- scDblFinder(sce)
seurat_obj_filtered$doublet_class <- sce$scDblFinder.class
seurat_obj_filtered <- subset(seurat_obj_filtered, subset = doublet_class == "singlet")
Step 7: Proceed to Downstream Analysis
After QC filtering and doublet removal, proceed to normalization, feature selection, scaling, dimensionality reduction, and clustering. The steps are described in the normalization section above and follow the canonical single-cell analysis workflow [<a href="#ref-1">1</a>].
Limitations and Considerations
Quality Control Is Not a Substitute for Experimental Design
Quality control cannot fix fundamental problems with sample preparation or sequencing. If your dissociation protocol damages cells, no amount of computational filtering will recover the lost biological information. The best QC strategy is to optimize sample preparation and minimize technical artifacts before sequencing.
The single-cell field continues to evolve, with new technologies and analysis methods emerging regularly [<a href="#ref-1">1</a>]. Stay informed about best practices and be willing to update your workflow as new evidence becomes available.
Thresholds Are Dataset-Specific
The thresholds used in this tutorial are starting points, not universal recommendations. The optimal thresholds for your dataset depend on tissue type, dissociation protocol, sequencing platform, and biological questions. Always examine your data distributions and justify your thresholds based on evidence.
The lack of universal standardization in scRNA-seq analysis means that different studies may use different thresholds [<a href="#ref-1">1</a>]. This is acceptable as long as you document your decisions and explain your rationale.
mRNA Abundance Does Not Always Reflect Protein Expression
Single-cell RNA-seq measures mRNA abundance, which is not always a reliable proxy for protein expression due to post-transcriptional and translational regulation [<a href="#ref-7">7</a>]. Quality control based on mRNA metrics cannot identify cells with abnormal protein expression if their mRNA profiles appear normal. This limitation is inherent to scRNA-seq technology and should be considered when interpreting results.
Differential Expression Results Can Be Sensitive to Filtering Choices
The choice of QC thresholds can affect downstream differential expression results. Different filtering decisions can lead to inconsistent findings even in high-quality samples [<a href="#ref-12">12</a>]. This sensitivity underscores the importance of documenting your QC decisions and considering how they might affect your conclusions.
Professional Escalation Criteria
When to Seek Expert Assistance
Some QC situations require expert assistance. If you encounter any of the following situations, consider consulting a bioinformatics specialist or collaborating with an experienced single-cell analyst:
- Your dataset shows unusual QC metric distributions that do not match any common failure pattern
- You are working with a novel tissue type or cell population with no published QC reference
- Your filtering removes an unexpectedly large or small proportion of cells
- You observe conflicting results between different QC metrics
- You need to integrate data from multiple batches or platforms with different QC characteristics
Resources for Further Learning
Several resources provide additional training and support for single-cell analysis. The EMBL-EBI training program offers structured learning pathways for bioinformatics analysis [<a href="#ref-10">10</a>]. The Galaxy Training Network provides accessible workflow training and analysis tutorials [<a href="#ref-6">6</a>]. The Carpentries offers foundational computing and data lessons that can strengthen your R programming skills [<a href="#ref-3">3</a>].
The nf-core documentation provides standards for reproducible workflow development that can help you structure your analysis pipelines [<a href="#ref-9">9</a>]. Bioconductor offers official package documentation and workflow guides for genomic analysis [<a href="#ref-4">4</a>].
Frequently Asked Questions
What is the difference between filtered and unfiltered feature-barcode matrices from Cell Ranger?
Cell Ranger produces two output matrices. The filtered matrix contains cells that pass Cell Ranger's internal empty droplet detection, while the unfiltered matrix contains all barcodes with detectable counts. For most analyses, start with the filtered matrix and apply additional QC filtering as described in this tutorial. The unfiltered matrix is useful if you want to examine the distribution of all barcodes or apply your own empty droplet detection.
How do I choose the minimum number of genes per cell threshold?
Examine the histogram of nFeature_RNA and identify the natural breakpoint where the distribution drops off. Cells below this breakpoint are likely empty droplets or severely damaged cells. A common starting point is 200 to 500 genes, but the appropriate value depends on your cell types. Small cells such as T cells naturally express fewer genes than large cells such as macrophages.
What percentage of mitochondrial reads is too high?
The acceptable percentage of mitochondrial reads depends on tissue type and dissociation protocol. A common threshold is 10% to 20%, but some tissues naturally have higher mitochondrial content. Examine the distribution of percent.mt in your data and identify where the main population ends. Cells with mitochondrial percentages well above the main distribution are likely damaged.
How do I know if my filtering removed too many cells?
If you retain fewer than 50% of cells after filtering, your thresholds may be too aggressive. Re-examine the QC metric distributions and consider whether you are removing legitimate cell populations. If you have prior knowledge about expected cell types in your sample, check whether known cell types are still present after filtering.
What is the expected doublet rate in scRNA-seq data?
The expected doublet rate depends on the number of cells loaded into the microfluidic device or droplet generator. Higher loading densities increase doublet rates. Most protocols provide expected doublet rates based on cell loading density. Computational doublet detection methods can identify cells that resemble simulated doublets, but they have limitations and should be interpreted with caution.
Should I filter on nCount_RNA in addition to nFeature_RNA?
Filtering on nCount_RNA is often redundant with filtering on nFeature_RNA because the two metrics are correlated. However, filtering on nCount_RNA can help remove cells with high UMI counts but low gene counts, which may indicate ambient RNA contamination. Examine the scatter plot of nCount_RNA versus nFeature_RNA to determine whether additional filtering is needed.
Can I use the same QC thresholds for single-nucleus RNA-seq data?
Single-nucleus RNA-seq (snRNA-seq) data have different QC characteristics than scRNA-seq data. Nuclei typically have lower gene counts and different mitochondrial read percentages than whole cells. The practical guide for single-cell transcriptome data analysis in neuroscience discusses the unique challenges of snRNA-seq data [<a href="#ref-8">8</a>]. Examine your data distributions and adjust thresholds accordingly.
How should I report QC decisions in my publication?
Report the number of cells before and after QC filtering, the thresholds used for each metric, and the number of cells removed at each step. Include this information in the methods section or as supplementary material. The Galaxy Training Network emphasizes the importance of reproducibility in bioinformatics analysis [<a href="#ref-6">6</a>], and documenting your QC decisions follows this principle.
Related Bioinformatics Guides
- Single-Cell RNA Sequencing Quality Control: A Practical Guide to Filtering and Metrics
- RNA-Seq Quality Control: Essential Checks and Tools
- Single-Cell Sequencing Depth: How Much Is Enough?
- Single-Cell RNA Sequencing Depth: A Cost-Benefit Analysis for Experimental Design
- Single-Cell Isolation Techniques: A Practical Comparison
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
[1] [Single-Cell RNA Sequencing Analysis: A Step-by-Step Overview.](https://pubmed.ncbi.nlm.nih.gov/33835452). Methods in molecular biology (Clifton, N.J.), 2021. [2] [popsicleR: A R Package for Pre-processing and Quality Control Analysis of Single Cell RNA-seq Data.](https://doi.org/10.1016/j.jmb.2022.167560). Journal of Molecular Biology, 2022. [3] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries. [4] [Bioconductor](https://bioconductor.org/). Bioconductor Project. [5] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information. [6] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project. [7] [Machine learning predictions surpass individual mRNAs as a proxy of single-cell protein expression.](https://doi.org/10.1186/s13059-026-04083-1). 2026. [8] [A practical guide for single-cell transcriptome data analysis in neuroscience.](https://pubmed.ncbi.nlm.nih.gov/40164433). Neuroscience research, 2025. [9] [nf-core Documentation](https://nf-co.re/docs). nf-core. [10] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute. [11] [Galaxy single-cell & spatial omics community update: Navigating new frontiers in 2025.](https://doi.org/10.1016/j.xgen.2025.101005). 2025. [12] [Differential expression analysis in single-cell and spatial RNA-seq without model assumptions.](https://doi.org/10.1016/j.crmeth.2026.101383). 2026.This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.