Pseudobulk vs. Single-Cell Level Differential Expression: A Statistical Guide for Single-Cell RNA-Seq
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Pseudobulk analysis is the statistically robust default for comparing experimental groups across biological replicates. This approach aggregates gene counts from cells within each sample and cell type, treating the biological replicate as the unit of analysis, thereby reliably controlling false positive rates by avoiding pseudoreplication.
- **Single-cell level differential expression methods are appropriate for comparing cell types or states within a single biological sample.** These methods model expression at the individual cell level and are suitable when cells from the same sample are considered independent observations, such as identifying marker genes between T cells and B cells in one individual.
- Ignoring the correlation structure between cells from the same biological replicate leads to pseudoreplication and inflated false discovery rates. This statistical error, where cells are treated as independent observations despite sharing genetic background and technical batch effects, can result in hundreds of false positive differentially expressed genes.
- The number of biological replicates is paramount; pseudobulk analysis is limited by this count, and single-cell sequencing cannot compensate for insufficient biological replication. Studies with fewer than 5 replicates per group exhibit low power, and even with more replicates, the choice of method (pseudobulk vs. mixed-effects models) depends on the replicate count to ensure stable variance estimates and appropriate statistical power.
- Batch effects must be modeled as covariates in differential expression analysis, not corrected prior to testing. Using batch-corrected data for differential expression can distort expression values and yield unreliable results; integrating batch as a fixed effect in the statistical model is the recommended strategy to account for systematic technical variation.
- Cell type composition differences between groups can confound differential expression results. Changes in the proportion of a cell type can manifest as altered gene expression in pseudobulk analysis, necessitating careful examination of cell type proportions and their potential impact on observed differential expression.
Differential expression (DE) analysis is a core step in single-cell RNA sequencing (scRNA-seq) studies, but the choice between aggregating cells into pseudobulk samples and modeling expression at the single-cell level carries major statistical consequences. The central problem is that single-cell-level methods that ignore variation between biological replicates are biased and prone to false discoveries, with widely used methods capable of finding hundreds of differentially expressed genes when no biological difference exists. Pseudobulk approaches, which sum counts across cells within a sample and then apply bulk RNA-seq DE tools, treat the biological replicate as the unit of analysis and therefore control false positive rates more reliably. This guide provides a statistical framework for deciding between these approaches, covering when each is appropriate, how to implement them, and how to interpret results without overstating biological conclusions.
The Statistical Problem in Single-Cell Differential Expression
Single-cell transcriptomics measures expression in thousands or millions of individual cells, which creates an illusion of large sample sizes. A typical experiment with 10 patients and 10 controls may yield 50,000 cells after quality control, and researchers might be tempted to treat each cell as an independent observation. This reasoning is statistically invalid because cells from the same individual are not independent. They share the donor's genetic background, environmental exposures, sample processing history, and technical batch effects. Ignoring this correlation structure inflates the effective sample size and produces p-values that are far too small.
The core issue is the distinction between biological replication and technical replication. Biological replicates are independent samples from different individuals or animals. Technical replicates are repeated measurements from the same biological sample. In scRNA-seq, each cell is a technical measurement nested within a biological sample. When a DE method treats cells as independent replicates, it commits pseudoreplication, a statistical error that systematically underestimates variance and overstates significance.
Benchmarking studies have demonstrated the practical consequences of this error. Methods that ignore variation between biological replicates are biased and prone to false discoveries, and the most widely used methods can discover hundreds of differentially expressed genes in the absence of biological differences. This finding comes from a systematic evaluation of DE methods applied to single-cell data, including a demonstration using injured mouse spinal cord data where true and false discoveries were explicitly exposed. The problem is not limited to one method or one data type. It reflects a fundamental mismatch between the statistical assumptions of the methods and the structure of single-cell data.
The four major challenges in single-cell DE analysis have been dissected as excessive zeros, normalization, donor effects, and cumulative biases. Excessive zeros arise because dropout events, where a gene is not detected in a cell despite being expressed, create a large proportion of zero counts. Normalization is complicated by differences in sequencing depth across cells and samples. Donor effects capture the biological and technical variation attributable to each individual. Cumulative biases emerge when these issues compound through the analysis pipeline. Each challenge can distort DE results, and methods that address only one or two of them may still produce unreliable conclusions.
Pseudobulk Aggregation as the Reference Approach
Pseudobulk analysis addresses the pseudoreplication problem by aggregating counts across all cells of a given cell type within each biological sample. The result is one expression profile per sample per cell type, which matches the structure of bulk RNA-seq data. Standard bulk DE tools such as limma-voom, DESeq2, and edgeR can then be applied with the biological replicate as the unit of analysis.
The statistical logic is straightforward. If a study has 8 patients and 8 controls, pseudobulk aggregation produces 8 pseudobulk profiles per group for each cell type. The DE analysis then compares these 8 versus 8 profiles, which is an honest representation of the sample size. The method does not pretend that 50,000 cells provide 50,000 independent observations. It accepts the limitation of the experimental design and models the data accordingly.
Pseudobulk methods have become the default recommendation in many single-cell analysis workflows because they control false positive rates under a wide range of conditions. The approach is particularly important when the research question concerns differences between experimental groups, such as disease versus control or treated versus untreated. In these designs, the biological replicate is the fundamental unit of comparison, and any analysis that ignores this structure risks reporting genes that are not truly differentially expressed.
The practical implementation of pseudobulk analysis requires careful attention to the aggregation step. Counts should be summed across cells within each sample and cell type, not averaged. Summing preserves the count nature of the data and allows the use of count-based statistical models. Averaging or normalizing before aggregation can distort the distributional assumptions of downstream tools. The aggregation should be performed on raw or minimally filtered counts, before any normalization that would remove the library size information needed by methods like DESeq2 or edgeR.
Single-Cell Level Models and Their Appropriate Uses
Single-cell level DE methods model expression at the level of individual cells while attempting to account for the correlation structure within samples. These methods include mixed-effects models, where the biological sample is treated as a random effect, and specialized tools designed for single-cell data that incorporate zero inflation or other distributional features.
The appropriate use of single-cell level models depends on the research question. When the comparison is between cell types within the same sample, such as asking which genes distinguish T cells from B cells in a single individual, the cells are the relevant units and single-cell level analysis is appropriate. When the comparison is between experimental conditions across individuals, such as disease versus control, the biological replicate must be the unit of analysis, and pseudobulk or mixed-effects models that respect this structure are required.
Mixed-effects models offer a compromise. They use all the cells in the analysis, which can improve power for detecting small effects, while including a random intercept for each biological sample to account for within-sample correlation. These models are more computationally intensive than pseudobulk approaches and require careful specification of the random effects structure. They also assume that the random effects are normally distributed, which may not hold for count data without appropriate transformations.
The choice between pseudobulk and single-cell level models is not always clear cut. Benchmarking studies have shown that the relative performance of DE methods depends on their ability to account for variation between biological replicates. Methods that ignore this variation are biased, but methods that model it appropriately can perform well. The key is to match the statistical model to the experimental design and to validate the assumptions of the chosen approach.
At a Glance: Choosing Between Pseudobulk and Single-Cell Level DE
| Analysis Scenario | Recommended Approach | Statistical Rationale | Key Implementation Detail |
|---|---|---|---|
| Disease versus control comparison across individuals | Pseudobulk with limma-voom, DESeq2, or edgeR | Biological replicate is the unit of analysis, controls false positive rate | Sum counts across cells within each sample and cell type before DE testing |
| Cell type comparison within a single sample | Single-cell level methods | Cells are independent within a sample, no cross-sample correlation to model | Use methods appropriate for sparse count data with zero inflation |
| Small number of biological replicates per group | Pseudobulk with empirical Bayes moderation | Borrows information across genes to stabilize variance estimates | Use limma-voom for count data with precision weights |
| Large cohort with many biological replicates | Mixed-effects models or pseudobulk | Both can control false positives, mixed models may improve power | Specify sample as a random effect, validate model convergence |
| Batch effects confounded with condition | Pseudobulk with batch covariate in the model | Models batch as a fixed effect, avoids using batch-corrected data for DE | Include batch as a covariate in the design matrix |
| Exploratory analysis without clear group structure | Single-cell level methods for clustering and visualization | DE testing is secondary to cell type discovery | Reserve DE testing for confirmatory analyses |
Workflow for Pseudobulk Differential Expression
Step 1: Define the Comparison Groups
Before any analysis, define the biological question precisely. The comparison groups should reflect the experimental design, such as disease versus control, treated versus untreated, or time points within a longitudinal study. Each group must contain biological replicates from different individuals or animals. Technical replicates, such as multiple libraries from the same sample, do not count as independent replicates.
Step 2: Perform Quality Control and Cell Type Annotation
Quality control is a prerequisite for reliable DE analysis. Filter cells based on the number of detected genes, total UMI counts, and mitochondrial read fraction. The specific thresholds depend on the tissue and protocol, but the goal is to remove low-quality cells and doublets that would distort expression profiles. After filtering, cluster the cells and annotate cell types using known marker genes. The quality of the cell type annotation directly affects the validity of the pseudobulk profiles, since each profile represents the aggregate expression of a putative cell type.
Step 3: Aggregate Counts by Sample and Cell Type
For each biological sample and each cell type, sum the UMI counts across all cells. The result is a count matrix with rows representing genes and columns representing sample-cell type combinations. This matrix is the input for downstream DE analysis. The aggregation step should be performed on the raw count matrix, before any normalization or transformation.
Step 4: Filter Lowly Expressed Genes
Genes with very low total counts across all samples provide little information and increase the multiple testing burden. Filter genes that are expressed in fewer than a minimum number of samples or that have very low total counts. The specific thresholds depend on the sequencing depth and the expected expression levels of the genes of interest. Common choices include requiring expression in at least 20 percent of samples or a minimum total count across all samples.
Step 5: Apply a Bulk DE Method
Use a bulk RNA-seq DE tool that is appropriate for count data. limma-voom transforms the counts to log2 counts per million, estimates precision weights, and fits a linear model with empirical Bayes moderation. DESeq2 models the counts with a negative binomial distribution and uses a shrinkage estimator for dispersion. edgeR uses a similar negative binomial framework with an empirical Bayes approach to dispersion estimation. All three methods are well established for bulk RNA-seq and have been adapted for pseudobulk single-cell data.
Step 6: Adjust for Multiple Testing
The number of genes tested in a single-cell DE analysis can be 20,000 or more. Multiple testing correction is essential to control the false discovery rate. The Benjamini-Hochberg procedure is the standard approach, producing adjusted p-values that can be interpreted as the expected proportion of false positives among the genes called significant. Report both the raw and adjusted p-values, and base biological conclusions on the adjusted values.
Step 7: Validate with Independent Data
DE results from a single dataset should be treated as hypotheses instead of conclusions. Validation can take several forms. External validation uses an independent cohort to test whether the same genes are differentially expressed. Internal validation might use a different statistical method or a different subset of the data. The strongest evidence comes from consistent results across multiple analytical contexts, such as pseudobulk analysis of single-cell data combined with bulk RNA-seq data from the same condition.
Practical Implementation Steps for Single-Cell Level Models
Step 1: Confirm the Experimental Design Supports Single-Cell Level Analysis
Single-cell level models are appropriate when the research question involves comparisons within samples or when the number of biological replicates is large enough to support mixed-effects modeling. For between-group comparisons with fewer than 10 replicates per group, pseudobulk analysis is generally safer. The choice should be made before the analysis, not after seeing the results.
Step 2: Specify the Mixed-Effects Model
If using a mixed-effects model, specify the fixed effects for the experimental condition and the random effects for the biological sample. The model should also include technical covariates such as sequencing batch or library preparation date if these are known. The random effects structure must reflect the nesting of cells within samples.
Step 3: Choose an Appropriate Distributional Assumption
Single-cell count data are overdispersed relative to a Poisson distribution, meaning the variance exceeds the mean. The negative binomial distribution is a common choice that accommodates overdispersion. Zero-inflated models add a separate component for excess zeros, which can improve fit for genes with high dropout rates. However, zero-inflated models can deteriorate performance for low depth data, so the choice should be guided by the data characteristics.
Step 4: Assess Model Convergence and Fit
Mixed-effects models can fail to converge, especially with sparse data or complex random effects structures. Check convergence diagnostics and examine the fitted values for obvious problems. If the model does not converge, simplify the random effects structure or switch to a pseudobulk approach.
Step 5: Compare Results with Pseudobulk Analysis
A practical validation step is to run both pseudobulk and single-cell level analyses and compare the results. Genes that are significant in both analyses are more likely to be true positives. Genes that are significant only in the single-cell level analysis may reflect inflated false positives, especially if the single-cell level method ignores biological replication.
Options and Tradeoffs in DE Methods
limma-voom
limma-voom is a widely used approach that applies linear modeling to log-transformed counts with precision weights. The voom transformation estimates the mean-variance relationship and assigns weights to each observation. The empirical Bayes moderation in limma borrows information across genes, which is particularly valuable when the number of biological replicates is small. limma-voom performs well in benchmarking studies for pseudobulk analysis of single-cell data.
DESeq2
DESeq2 models count data with a negative binomial distribution and uses a shrinkage estimator for dispersion and log fold changes. The shrinkage is stronger for genes with low counts and high dispersion, which reduces false positives for weakly expressed genes. DESeq2 is appropriate for pseudobulk analysis and has been used in single-cell studies where the comparison is between groups of biological replicates.
edgeR
edgeR also uses a negative binomial model with empirical Bayes dispersion estimation. It offers several tests for differential expression, including exact tests for simple designs and likelihood ratio tests for more complex designs. edgeR is computationally efficient and scales well to large pseudobulk matrices.
Mixed-Effects Models
Mixed-effects models such as those implemented in glmmTMB or lme4 can model single-cell data directly while accounting for within-sample correlation. These models are more flexible than pseudobulk approaches but require careful specification and are computationally intensive. They are most useful when the number of biological replicates is large and the research question requires cell-level inference.
Specialized Single-Cell Methods
Several methods have been developed specifically for single-cell DE analysis, including those based on zero-inflated models or generalized linear models with UMI counts. These methods can perform well in specific scenarios, but benchmarking studies have shown that their relative performance depends on the data characteristics. Methods that ignore biological replication remain problematic regardless of their other features.
Observations and Measurements in DE Analysis
Effect Size Estimation
The log fold change is the standard effect size measure in DE analysis. For pseudobulk data, the log fold change is calculated from the aggregated counts, which provides a stable estimate of the average expression difference between groups. For single-cell level models, the log fold change can be estimated from the model coefficients, but the interpretation is complicated by the zero inflation and normalization issues.
Dispersion Estimation
Dispersion measures the variability of expression across biological replicates. Genes with high dispersion are more variable and require larger effect sizes to reach statistical significance. Pseudobulk methods estimate dispersion from the aggregated counts, which is more stable than estimating dispersion from individual cells. The dispersion estimates are used in the statistical tests and influence the power to detect DE genes.
Zero Proportions
The proportion of cells with zero counts for a given gene is a useful diagnostic. Genes with very high zero proportions are expressed in only a small fraction of cells, and their DE status may reflect changes in the proportion of expressing cells instead of changes in expression level per cell. This distinction is biologically important and should be reported alongside the DE results.
Library Size Variation
Sequencing depth varies across cells and samples. In pseudobulk analysis, the library size of each pseudobulk sample is the sum of the library sizes of its constituent cells. Methods like DESeq2 and edgeR account for library size differences through normalization factors. limma-voom uses the log counts per million, which normalizes for library size before modeling.
Records and Documentation for Reproducibility
Analysis Code and Parameters
Record the exact code and parameters used for every step of the analysis, from quality control through DE testing. This includes the software versions, the filtering thresholds, the normalization method, and the statistical model. Reproducibility requires that another researcher can run the same analysis and obtain the same results.
Data Versions and Accession Numbers
Record the version of the reference genome, the annotation files, and the count matrix. If the data are publicly available, record the accession numbers. Public data repositories such as those maintained by the National Center for Biotechnology Information provide stable access to sequence data and analysis results. The NCBI Data Resources offer search systems and analysis services that support reproducible research.
Quality Control Metrics
Record the number of cells before and after quality control, the number of genes detected per cell, the mitochondrial read fraction, and any other metrics used for filtering. These numbers provide context for interpreting the DE results and allow readers to assess the robustness of the findings.
Analysis Environment
Record the computing environment, including the operating system, the version of R or Python, and the versions of all analysis packages. Containerized workflows and pipeline frameworks can improve reproducibility by standardizing the analysis environment. Community pipeline standards such as those documented in the nf-core Documentation provide guidance for reproducible workflow configuration.
Common Failure Patterns in DE Analysis
Pseudoreplication
The most common and most serious failure is treating cells as independent replicates when they are nested within biological samples. This error produces p-values that are far too small and can generate hundreds of false positive DE genes. The solution is to use pseudobulk analysis or mixed-effects models that account for the sample structure.
Using Batch-Corrected Data for DE Testing
Batch correction methods such as Harmony or Seurat integration are designed for visualization and clustering, not for DE testing. Using batch-corrected data for DE analysis can distort the expression values and produce unreliable results. Benchmarking studies have shown that the use of batch-corrected data rarely improves the analysis for sparse data, whereas batch covariate modeling improves the analysis for substantial batch effects. The recommended approach is to model batch as a covariate in the DE analysis instead of correcting the data beforehand.
Inappropriate Normalization
Normalization choices have a major impact on DE results. Methods that use relative abundance instead of absolute expression can introduce biases, particularly for genes with large changes in expression. The choice of normalization method should be guided by the data characteristics and validated with diagnostic plots.
Ignoring Zero Inflation
The excessive zeros in single-cell data violate the assumptions of many statistical methods. Methods that do not account for zero inflation can produce biased estimates of expression levels and dispersion. However, zero-inflated models are not always the best choice, and their performance depends on the sequencing depth and the proportion of zeros in the data.
Overinterpreting Small Effect Sizes
Statistical significance does not imply biological importance. A gene can be significantly differentially expressed with a log fold change of 0.1, which may have no biological consequence. Report effect sizes alongside p-values and interpret the results in the context of the biological question.
Limitations of Pseudobulk and Single-Cell Level Approaches
Loss of Cell-Level Information
Pseudobulk aggregation discards information about the distribution of expression across cells within a sample. Two samples with the same mean expression can have very different distributions, with one having uniform expression across all cells and the other having high expression in a small subpopulation. This distinction is lost in pseudobulk analysis. Single-cell level models can capture this information but require more complex statistical machinery.
Small Sample Sizes
Pseudobulk analysis is limited by the number of biological replicates. A study with 3 patients and 3 controls has very low power to detect DE genes, regardless of the number of cells sequenced. Increasing the number of cells per sample does not increase the effective sample size for between-group comparisons. Researchers should design studies with adequate biological replication, recognizing that single-cell sequencing cannot compensate for a small number of donors.
Cell Type Composition Effects
Pseudobulk analysis of a cell type assumes that the cell type is present in all samples and that the cells are comparable across samples. If a cell type is rare or absent in some samples, the pseudobulk profile may be based on very few cells and be highly variable. Cell type composition differences between groups can confound DE analysis, since changes in the proportion of a cell type can appear as changes in expression.
Technical Batch Effects
Batch effects are a major source of variation in single-cell data. Samples processed on different days, by different operators, or with different reagent lots can show systematic differences in expression. These batch effects can be confounded with the experimental condition if all cases are processed in one batch and all controls in another. The design should balance cases and controls across batches whenever possible.
Quality Controls and Validation Strategies
Diagnostic Plots
Examine diagnostic plots before interpreting DE results. A mean-variance plot shows whether the variance increases with the mean as expected. A p-value histogram shows whether the p-values are uniformly distributed under the null hypothesis, with an excess of small p-values indicating true signals and an excess of very large p-values indicating model misspecification. An MA plot shows the relationship between mean expression and log fold change.
Replicate Correlation
Check the correlation between pseudobulk profiles from the same group. High correlation indicates that the samples are similar and the analysis is stable. Low correlation may indicate technical problems, batch effects, or genuine biological heterogeneity. The correlation should be reported as a quality metric.
Independent Validation
The strongest validation is replication in an independent cohort. If the same genes are differentially expressed in two independent studies, the evidence for a true biological effect is much stronger. External validation can use bulk RNA-seq data, which is often available from public repositories, or a second single-cell dataset.
Sensitivity Analysis
Test whether the DE results are robust to the analysis choices. Run the analysis with different filtering thresholds, different normalization methods, and different statistical models. If the results change substantially, the conclusions should be tempered. Sensitivity analysis is particularly important for genes with borderline significance.
Safety and Regulatory Context for Research Applications
Data Privacy and Consent
Single-cell RNA-seq data from human subjects contain sensitive biological information. Researchers must ensure that the data were collected with appropriate informed consent and that the analysis complies with relevant privacy regulations. De-identified data should be used whenever possible, and access to raw data should be controlled.
Clinical Translation
DE results from single-cell studies are research findings, not clinical diagnostics. Translating these findings to clinical applications requires extensive validation, regulatory approval, and clinical testing. Researchers should be careful not to overstate the clinical implications of their findings, particularly when the sample sizes are small and the results have not been replicated.
Animal Research Compliance
Studies using animal models must comply with institutional animal care and use regulations. The number of animals should be minimized while maintaining adequate statistical power. The choice between pseudobulk and single-cell level analysis affects the required number of biological replicates, and this should be considered in the experimental design.
Professional Escalation Criteria
When to Seek Statistical Consultation
Consult a statistician or bioinformatics specialist when the experimental design is complex, when the number of biological replicates is small, when batch effects are confounded with the experimental condition, or when the DE results will be used for regulatory submissions or clinical decisions. Statistical consultation is also warranted when the analysis produces unexpected results, such as thousands of DE genes with very small effect sizes.
When to Reanalyze the Data
Reanalyze the data when the quality control metrics indicate problems, when the DE results are not robust to sensitivity analysis, when the p-value histogram shows anomalies, or when the results contradict established biological knowledge. Reanalysis should use a different approach or address the identified problems in the original analysis.
When to Collect Additional Data
Collect additional data when the power is inadequate to detect the expected effect sizes, when the batch effects cannot be resolved with the current data, or when the cell type of interest is too rare for reliable pseudobulk analysis. Additional biological replicates are more valuable than additional cells from the same samples.
A Practical Decision Framework for Pseudobulk Versus Single-Cell Level DE
Choosing between pseudobulk and single-cell level differential expression is not a one-time decision made at the start of a project. It is a series of decisions that should be revisited as the analysis progresses and as new information about the data emerges. This section provides a structured framework for making those decisions, with concrete criteria, record-keeping practices, and troubleshooting steps that can be applied to real datasets.
Step 1: Classify the Primary Research Question
The first decision point is to classify the research question into one of three categories. Category A questions compare experimental groups across biological replicates, such as disease versus control, treated versus untreated, or time points in a longitudinal study. Category B questions compare cell types or cell states within a single biological sample, such as identifying marker genes that distinguish T cells from B cells in one individual. Category C questions involve both dimensions, such as asking whether the difference between two cell types changes between disease and control groups.
For Category A questions, pseudobulk analysis should be the default approach. The biological replicate is the fundamental unit of comparison, and any method that ignores this structure risks inflated false positives. For Category B questions, single-cell level methods are appropriate because cells within a sample can be treated as independent observations. For Category C questions, the analysis typically proceeds in two stages: first, identify cell types of interest using single-cell level methods, then apply pseudobulk analysis within each cell type to test for group differences.
Step 2: Assess the Biological Replicate Count
The number of biological replicates per group is the single most important factor in choosing a DE method. Studies with fewer than 5 replicates per group have very low power for pseudobulk analysis, and the dispersion estimates from methods like DESeq2 or edgeR will be unstable. With 5 to 10 replicates per group, pseudobulk analysis with empirical Bayes moderation, as implemented in limma-voom, is generally reliable. With more than 10 replicates per group, mixed-effects models become feasible and may offer improved power by using all cells in the analysis.
The replicate count also determines whether single-cell level methods can be used at all. Methods that ignore biological replication are problematic regardless of the number of replicates, but the consequences are more severe with fewer replicates. With 3 replicates per group, a single outlier sample can dominate the pseudobulk analysis, and the results should be interpreted with caution. With 20 or more replicates per group, the analysis is more robust, and the choice between pseudobulk and mixed-effects models becomes less consequential.
Step 3: Evaluate the Cell Type Composition
Cell type composition varies across samples and can confound DE analysis. Before aggregating counts, examine the proportion of each cell type in each sample. If a cell type is present in fewer than 10 cells in any sample, the pseudobulk profile for that sample will be noisy and may not be reliable. If a cell type is absent from several samples in one group, the pseudobulk analysis for that cell type should be reconsidered.
The cell type composition should be recorded for every sample and cell type combination. This record allows the analyst to identify samples where the pseudobulk profile is based on very few cells and to assess whether composition differences between groups are driving the DE results. In studies where cell type proportions differ substantially between groups, such as the increased proportion of classical monocytes observed in fibrotic hypersensitivity pneumonitis, the DE results should be interpreted in the context of these composition changes.
Step 4: Document the Analysis Decisions
Every decision in the DE analysis should be documented in a structured format that can be reviewed and reproduced. The documentation should include the software versions, the filtering thresholds, the normalization method, the statistical model, and the rationale for each choice. This record is essential for troubleshooting and for defending the analysis in peer review.
A practical approach is to maintain an analysis log with the following entries for each decision point: the date, the decision made, the reason for the decision, the alternative options considered, and the expected impact on the results. This log serves as a reference for sensitivity analysis and allows another researcher to understand the reasoning behind each choice.
Step 5: Run the Analysis and Record the Results
Run the chosen DE analysis and record the results in a structured format. The record should include the number of genes tested, the number of significant genes at the chosen threshold, the distribution of effect sizes, and the dispersion estimates. These summary statistics provide context for interpreting the results and for comparing the performance of different methods.
For pseudobulk analysis, record the number of cells per sample and cell type, the total library size of each pseudobulk sample, and the number of genes passing the expression filter. For single-cell level analysis, record the number of cells per sample, the proportion of zeros for each gene, and the convergence status of the model. These records allow the analyst to identify potential problems, such as a sample with very few cells or a gene with an extreme zero proportion.
Step 6: Compare Results Across Methods
A practical validation step is to run both pseudobulk and single-cell level analyses and compare the results. This comparison serves two purposes. First, it identifies genes that are significant in both analyses, which are more likely to be true positives. Second, it identifies genes that are significant only in the single-cell level analysis, which may reflect inflated false positives if the method ignores biological replication.
The comparison should be recorded in a table with columns for the gene, the pseudobulk p-value, the single-cell level p-value, the effect size in each analysis, and the direction of the effect. Genes that are significant in both analyses with consistent effect sizes provide the strongest evidence for a true biological difference. Genes that are significant only in the single-cell level analysis should be examined carefully, and the analysis should be checked for pseudoreplication or other statistical problems.
Step 7: Perform Sensitivity Analysis
Sensitivity analysis tests whether the DE results are robust to the analysis choices. Run the analysis with different filtering thresholds, different normalization methods, and different statistical models. If the results change substantially, the conclusions should be tempered.
A structured sensitivity analysis should vary one factor at a time and record the impact on the number of significant genes and the identity of the top genes. Common sensitivity analyses include changing the minimum number of cells per sample for pseudobulk aggregation, changing the expression filter threshold, and comparing results from limma-voom, DESeq2, and edgeR. The results of the sensitivity analysis should be reported alongside the primary analysis.
Troubleshooting Common Problems
When the DE results show unexpected patterns, several diagnostic checks can identify the source of the problem. A p-value histogram with an excess of values near 1 indicates model misspecification, often caused by overdispersion or incorrect distributional assumptions. A p-value histogram with an excess of values near 0 may indicate true signals or may indicate pseudoreplication if the analysis ignores biological replication.
An MA plot showing the relationship between mean expression and log fold change can reveal systematic biases. If the log fold changes are correlated with the mean expression, the normalization may be inadequate. If the variance increases with the mean more than expected, the dispersion model may be misspecified.
The correlation between pseudobulk profiles from the same group is a useful quality metric. High correlation indicates that the samples are similar and the analysis is stable. Low correlation may indicate technical problems, batch effects, or genuine biological heterogeneity. The correlation should be reported as a quality metric and examined for outlier samples.
Records and Measurements for the Decision Framework
The decision framework requires a structured record of the analysis decisions and their outcomes. The following measurements should be recorded for each analysis: the number of biological replicates per group, the number of cells per sample and cell type, the total library size of each pseudobulk sample, the number of genes passing the expression filter, the dispersion estimates, the number of significant genes at the chosen threshold, and the results of the sensitivity analysis.
These records serve multiple purposes. They allow the analyst to assess the reliability of the results, to identify potential problems, and to defend the analysis in peer review. They also provide the information needed to compare results across methods and to determine whether the choice of method affected the conclusions.
Common Failure Patterns in the Decision Framework
The most common failure in applying this framework is making the decision between pseudobulk and single-cell level analysis once at the start of the project and never revisiting it. The choice should be revisited when new information about the data emerges, such as the discovery that a cell type is rare in some samples or that batch effects are more severe than expected.
Another common failure is applying the same analysis approach to all cell types without considering the cell type composition. A cell type with 500 cells per sample may support reliable pseudobulk analysis, while a cell type with 20 cells per sample may not. The analysis should be tailored to the data characteristics of each cell type.
A third failure is ignoring the results of the sensitivity analysis. If the DE results change substantially when the filtering threshold is changed, the conclusions should be tempered. The sensitivity analysis is not a formality, it is a critical check on the robustness of the findings.
When to Escalate to Statistical Consultation
Statistical consultation is warranted when the experimental design is complex, when the number of biological replicates is small, when batch effects are confounded with the experimental condition, or when the DE results will be used for regulatory submissions or clinical decisions. Consultation is also warranted when the analysis produces unexpected results, such as thousands of DE genes with very small effect sizes, or when the sensitivity analysis shows that the results are not robust.
The decision framework should be applied before the analysis begins, but it should also be applied after the analysis is complete. If the results are surprising or if the sensitivity analysis reveals instability, the analysis should be revisited with statistical guidance. The goal is not to find the analysis that produces the most significant results, but to find the analysis that produces the most reliable results.
Frequently Asked Questions
What is the main difference between pseudobulk and single-cell level DE analysis?
Pseudobulk analysis aggregates counts across all cells of a cell type within each biological sample, creating one expression profile per sample. The DE test then compares these sample-level profiles between groups, treating the biological replicate as the unit of analysis. Single-cell level analysis models expression at the level of individual cells, which requires accounting for the correlation of cells within samples. The main difference is the unit of analysis, which determines how the sample size is calculated and how the false positive rate is controlled.
Why do single-cell level methods produce false positives?
Single-cell level methods that treat cells as independent observations ignore the correlation of cells within biological samples. Cells from the same individual share genetic, environmental, and technical factors that make their expression values correlated. Treating them as independent inflates the effective sample size and produces p-values that are too small. This pseudoreplication can generate hundreds of false positive DE genes when no biological difference exists.
When should I use pseudobulk analysis?
Use pseudobulk analysis when the research question involves comparing groups of biological replicates, such as disease versus control or treated versus untreated. Pseudobulk is the recommended default for between-group comparisons because it controls the false positive rate and is straightforward to implement with established bulk RNA-seq tools. It is particularly important when the number of biological replicates is small.
When is single-cell level analysis appropriate?
Single-cell level analysis is appropriate when the comparison is between cell types within the same sample, such as asking which genes distinguish T cells from B cells in a single individual. It is also appropriate when using mixed-effects models that include the biological sample as a random effect, which accounts for within-sample correlation. For between-group comparisons, single-cell level analysis should be used with caution and validated against pseudobulk results.
How many biological replicates do I need for pseudobulk DE analysis?
The number of biological replicates depends on the effect size, the variability of expression, and the desired statistical power. Studies with fewer than 5 replicates per group have very low power and are prone to both false positives and false negatives. Increasing the number of cells per sample does not compensate for a small number of biological replicates. A minimum of 5 to 10 replicates per group is often recommended, but the specific number should be determined by a power analysis.
Should I use batch-corrected data for DE analysis?
Batch-corrected data should generally not be used for DE testing. Batch correction methods are designed for visualization and clustering, and they can distort the expression values in ways that affect DE results. The recommended approach is to model batch as a covariate in the DE analysis, using methods like limma-voom, DESeq2, or edgeR that can include batch in the design matrix. Benchmarking studies have shown that batch covariate modeling improves the analysis for substantial batch effects.
What is the role of zero inflation in single-cell DE analysis?
Zero inflation refers to the excess of zero counts in single-cell data, which arises from dropout events where a gene is not detected despite being expressed. Zero inflation violates the assumptions of standard count models and can bias DE results. Some methods use zero-inflated models to address this issue, but these methods can deteriorate performance for low depth data. The choice of method should be guided by the data characteristics and validated with diagnostic plots.
How should I validate my DE results?
Validation can take several forms. External validation uses an independent cohort to test whether the same genes are differentially expressed. Sensitivity analysis tests whether the results are robust to the analysis choices, such as filtering thresholds, normalization methods, and statistical models. Consistency across multiple analytical contexts, such as pseudobulk analysis of single-cell data and bulk RNA-seq data from the same condition, provides stronger evidence than any single analysis.
Related Bioinformatics Guides
- Single-Cell RNA Sequencing Depth: A Cost-Benefit Analysis for Experimental Design
- Single-Cell Sequencing Depth: How Much Is Enough?
- RNA Sequencing Data Analysis: From Raw Reads to Differential Expression
- Single-Cell RNA-Seq Analysis Pipelines for Veterinary Immunology
- Single-Cell RNA Sequencing Quality Control: A Practical Guide to Filtering and Metrics
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- Single cell transcriptomic analyses of human heart failure with preserved ejection fraction.. bioRxiv : the preprint server for biology, 2025.
- Monocyte Subset Predicts Clinical Outcomes in Fibrotic Hypersensitivity Pneumonitis.. bioRxiv : the preprint server for biology, 2025.
- Redistribution of peripheral blood immune cell subsets and remodeling of intercellular communication in rare Behçet's disease: a single-cell transcriptomics study.. Frontiers in medicine, 2026.
- Cellular and molecular dysregulation of the esophageal epithelium in systemic sclerosis.. JCI insight, 2026.
- Single cell transcriptome signatures and cell-cell interactions associated with sarcoidosis in lung immune cell populations.. 2026.
- Integrative single-cell and bulk transcriptomics identify an <,i>,ALK<,/i>,-associated three-gene signature predicting neuroblastoma outcomes.. 2026.
- Single-cell and spatial landscapes reveal coordinated CD8+ T-cell programs defining responder-specific immune niches in NSCLC. 2026.
- Peripheral B cell immune dysregulation genetically contributes to stage-dependent neuroinflammation and identifies priority therapeutic targets in Parkinson's disease: a computational integration of Mendelian randomization and single-cell transcriptomics.. 2026.
- A Systemic Immune State Axis Distinguishes Psoriatic Arthritis from Psoriasis.. 2026.
- Exploring and mitigating shortcomings in single-cell differential expression analysis with a new statistical paradigm. Genome Biology, 2025.
- Confronting false discoveries in single-cell differential expression. Nature Communications, 2021.
- Benchmarking integration of single-cell differential expression. Nature Communications, 2023.
- Bayesian approach to single-cell differential expression analysis. Nature Methods, 2014.
- SINGLE-CELL SEQUENCING REVEALS PATHWAYS AND PERTURBATIONS PREDICTIVE OF CENTRAL NERVOUS SYSTEM (CNS) DISSEMINATION IN B-CELL ACUTE LYMPHOBLASTIC LEUKEMIA (B-ALL). Haematologica, 2026.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.