A Practical Guide to Pseudobulk Differential Expression for Single-Cell Data: How to Aggregate Cells and Avoid Statistical Pitfalls
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Pseudobulk differential expression analysis aggregates single-cell RNA-seq counts per biological sample and cell type, enabling the application of established bulk RNA-seq tools like edgeR or DESeq2. This approach correctly treats biological samples (e.g., donors) as independent replicates, mitigating the inflated false discovery rates inherent in treating individual cells as independent observations.
- The fundamental requirement for valid pseudobulk analysis is the presence of multiple biological replicates per condition; a minimum of three is recommended to reliably estimate within-group variance and achieve adequate statistical power. Without sufficient biological replication, statistical inference is impossible, regardless of the number of cells sequenced.
- Summing raw UMI or read counts per sample and cell type is the statistically recommended aggregation method, as it preserves the count distribution necessary for negative binomial models used by tools like edgeR and DESeq2. Averaging normalized expression values requires different statistical frameworks and does not fit count-based models.
- Before aggregation, rigorous cell-level quality control is critical, including filtering low-quality cells, removing doublets, and correcting for ambient RNA contamination using methods like CellBender or SoupX. Poor quality input data will directly distort pseudobulk expression values and lead to misleading differential expression results.
- Statistical models must account for potential unequal variances between experimental groups, a common issue in pseudobulk data; methods like voomByGroup or voomWithQualityWeights are recommended over standard approaches that assume equal variances. Improper removal of unwanted variation using RUV methods can also compromise results by removing biological signal.
- Pseudobulk analysis reflects average expression within a cell type and can be confounded by differences in cell type composition between conditions. Researchers must report cell type proportions and consider these compositional effects when interpreting differential expression results, as they can drive apparent expression changes.
Pseudobulk differential expression analysis is the process of summing or averaging single-cell RNA sequencing (scRNA-seq) counts within each biological sample and cell type, then applying conventional bulk RNA-seq differential expression tools such as edgeR or DESeq2 to the aggregated data. This approach treats each donor or patient as the experimental unit instead of each individual cell, which preserves biological replication structure and controls false discovery rates more reliably than cell-level testing. Researchers working with single-cell or single-nucleus RNA-seq data should aggregate cells into pseudobulk profiles whenever they have multiple biological replicates per condition, because cell-level tests treat thousands of cells from the same donor as independent observations and produce inflated significance. This guide covers aggregation methods, statistical model selection, quality control steps, common failure patterns, and practical decision criteria for pseudobulk analysis.
Why Pseudobulk Analysis Is Necessary for Differential Expression
Single-cell RNA-seq experiments generate expression measurements for thousands to millions of individual cells, but those cells are not independent observations. Cells collected from the same biological sample share the donor's genetic background, environmental exposures, sample processing history, and technical batch effects. When researchers test for differential expression at the cell level, each cell contributes a data point, and the analysis effectively treats all cells from one patient as independent replicates. This violates the fundamental statistical assumption of independence and leads to artificially small p-values and inflated false discovery rates.
The core problem is that biological replication occurs at the donor or sample level, not at the cell level. A study with five patients in each condition has five biological replicates per group regardless of whether each patient contributes 500 or 50,000 cells. Pseudobulk analysis acknowledges this structure by collapsing all cells from one sample and one cell type into a single expression profile, creating one observation per biological replicate per cell type. Standard differential expression tools designed for bulk RNA-seq, such as edgeR and DESeq2, then operate on these aggregated profiles with the correct number of replicates.
Cell-level test methods fail in single-cell data because they do not account for the intra-donor correlation among cells. The technical literature directly addresses this failure mode, with work titled "Why do cell-level test methods for detecting differentially expressed genes fail in single-cell RNA-seq data?" examining the statistical mechanisms behind inflated false positives when cells are treated as independent replicates [<a href="#ref-1">1</a>]. The practical consequence is that researchers who skip pseudobulk aggregation and run Wilcoxon rank-sum tests or similar cell-level comparisons on their clusters will report hundreds or thousands of differentially expressed genes, most of which reflect technical noise and donor effects instead of true biological differences between conditions.
Pseudobulk analysis also aligns single-cell studies with the analytical standards of bulk transcriptomics. The same statistical frameworks that have been validated for bulk RNA-seq over decades, including negative binomial models and voom precision weighting, apply directly to aggregated single-cell data. This continuity allows researchers to leverage established tools, quality metrics, and interpretation practices from the bulk RNA-seq field. The Bioconductor project provides the primary software ecosystem for both bulk and single-cell differential expression analysis, with packages for aggregation, normalization, and statistical modeling maintained under reproducible version control [<a href="#ref-2">2</a>].
Biological Replicates and Experimental Design Requirements
The single most important requirement for valid pseudobulk differential expression is the presence of biological replicates. Pseudobulk analysis cannot create replication where the experimental design lacks it. A study with one control sample and one treated sample produces one pseudobulk profile per condition, and no statistical method can estimate within-group variance from a single observation. Researchers must design their experiments with at least three biological replicates per condition, and ideally more, to achieve reasonable statistical power.
The number of replicates needed depends on the expected effect size, the variability between donors, and the number of cells captured per sample. Studies of human disease tissue often use matched case-control designs, as illustrated by a rheumatoid arthritis study that performed single-cell RNA sequencing on peripheral blood mononuclear cells from 18 patients and 18 matched controls, with matching for age, sex, race, and ethnicity [<a href="#ref-3">3</a>]. This design controls for demographic confounders while providing balanced replication across conditions. The study identified 168 differentially expressed genes through pseudobulk analysis by cell type, demonstrating that moderate sample sizes can yield biologically meaningful results when the experimental design is sound [<a href="#ref-3">3</a>].
Single-nucleus RNA sequencing studies face additional constraints because tissue availability is often limited. A heart failure study using single-nucleus RNA sequencing on myocardial biopsies from 30 patients and 29 non-failing donor controls pooled samples from six patients per pool and used genotype-based demultiplexing to assign nuclei back to individual patients [<a href="#ref-4">4</a>]. This pooling strategy allowed the researchers to process limited biopsy material efficiently while still recovering donor-level information for pseudobulk analysis. After quality control, the study recovered 48,886 nuclei from eight pooled samples representing 19 heart failure patients and 24 controls, and identified cell type-specific differential expression patterns using limma-voom [<a href="#ref-4">4</a>].
Researchers should document their experimental design decisions before data collection begins. Key design parameters include the number of biological replicates per condition, the expected number of cells per sample, the cell types of interest, and the planned statistical model. These decisions determine whether pseudobulk analysis will have adequate power and whether the results can support the intended biological conclusions. The Galaxy Training Network offers accessible tutorials on experimental design and single-cell analysis workflows that can help researchers plan their studies before generating data [<a href="#ref-5">5</a>].
At a Glance: Pseudobulk Analysis Decision Table
| Decision Point | Recommended Approach | Key Consideration |
|---|---|---|
| Aggregation method | Sum raw UMI or read counts per sample and cell type | Preserves count distribution required for negative binomial models in edgeR and DESeq2 |
| Statistical model | edgeR, DESeq2, or limma-voom on pseudobulk counts | Match model to data characteristics and experimental design complexity |
| Unequal group variances | Use voomByGroup or voomWithQualityWeights with blocked design | Standard methods assume equal variance and can inflate false discovery rates when violated [<a href="#ref-6">6</a>] |
| Unwanted variation | Apply RUV2 or RUVIII per cell type with validated negative controls | Improper RUV implementation can remove biological signal [<a href="#ref-7">7</a>] |
| Minimum replicates | At least three biological replicates per condition | Fewer replicates cannot reliably estimate within-group variance |
| Cell type stratification | Construct separate pseudobulk profiles per annotated cell type | Prevents mixing signals from different cell populations |
| Quality control | Filter low-quality cells, doublets, and ambient RNA before aggregation | Poor quality input distorts pseudobulk expression values |
| Reporting | Document cell counts, library sizes, model parameters, and software versions | Supports reproducibility and interpretation of results |
Aggregation Methods for Constructing Pseudobulk Profiles
Pseudobulk construction requires summing or averaging expression values across all cells from the same biological sample and cell type. The choice of aggregation method affects downstream analysis results, and researchers should understand the tradeoffs between different approaches.
Summing Raw Counts
The most common and statistically recommended approach is to sum the raw UMI counts or read counts across all cells from a given sample and cell type. This produces a count matrix where each row is a gene, each column is a sample-cell type combination, and each value is the total number of transcripts detected for that gene across all cells in that group. Summing raw counts preserves the count nature of the data, which is required for negative binomial models used by edgeR and DESeq2.
Summing counts gives more weight to cell types that are more abundant in a sample. A cell type comprising 30% of the cells in a sample will contribute more total counts than a rare cell type comprising 2% of cells. This is appropriate when the research question concerns the total transcriptional output of a cell type within a tissue, but it can complicate comparisons across samples with different cell type compositions. Researchers should report cell type proportions alongside pseudobulk results so that readers can interpret expression differences in the context of cellular composition.
Averaging Normalized Expression
An alternative approach is to normalize expression within each cell first, then average the normalized values across cells from the same sample and cell type. This produces pseudobulk profiles on a continuous scale instead of a count scale. Averaging normalized values treats each cell equally regardless of its total transcript count, which can be useful when comparing cell types with very different sizes or RNA content.
However, averaged normalized expression values are not counts and do not fit the negative binomial distribution assumed by edgeR and DESeq2. Researchers who average normalized values must use statistical methods appropriate for continuous data, such as limma with voom transformation or linear models on log-transformed values. The choice between summing counts and averaging normalized values should be guided by the statistical framework planned for downstream analysis.
Cell Type Specific Aggregation
Pseudobulk profiles are typically constructed separately for each cell type. After clustering and cell type annotation, researchers subset the data to one cell type, then aggregate cells from that cell type within each biological sample. This produces one pseudobulk profile per sample per cell type, allowing cell type specific differential expression analysis.
Cell type specific aggregation requires reliable cell type annotations. If clustering or annotation is inaccurate, cells from different biological types will be mixed within pseudobulk profiles, diluting cell type specific signals. A study of diabetic foot ulcers performed marker-based cell type annotation after integration and clustering, resolving 10 major cell compartments before conducting macrophage-specific pseudobulk differential expression analysis [<a href="#ref-8">8</a>]. The quality of the annotation directly determined the interpretability of the downstream cell type specific results [<a href="#ref-8">8</a>].
Pseudoreplicates and Pooled Samples
Some experimental designs involve pooling cells from multiple donors before sequencing, as in the heart failure study that pooled six patient biopsies per sample [<a href="#ref-4">4</a>]. In these cases, genotype-based demultiplexing can assign cells back to individual donors, enabling donor-level pseudobulk analysis. When demultiplexing is not possible, researchers may need to treat each pooled sample as a pseudoreplicate, which reduces the effective number of biological replicates and limits statistical power.
The concept of pseudoreplicates extends to computational approaches that split cells from the same donor into multiple pseudobulk profiles. While pseudoreplicates can be useful for exploratory analysis or method validation, they do not provide true biological replication and should not be used as the basis for formal differential expression testing. A study on removing unwanted variation in pseudobulk analysis introduced a strategy called RUVIII PBPS that leverages pseudoreplicates to control false discovery rates when technical replicates or negative control genes are unavailable [<a href="#ref-7">7</a>]. This approach demonstrates that pseudoreplicates have analytical value, but their role is supplementary to true biological replication instead of a replacement for it [<a href="#ref-7">7</a>].
Statistical Models for Pseudobulk Differential Expression
Once pseudobulk profiles are constructed, researchers apply statistical models designed for bulk RNA-seq data. The choice of model affects error control, statistical power, and the interpretation of results.
edgeR and Negative Binomial Models
edgeR models count data using the negative binomial distribution, which accounts for both technical variation in sequencing and biological variation between replicates. The method estimates dispersion parameters from the data and uses empirical Bayes moderation to stabilize dispersion estimates for genes with low counts. edgeR is appropriate for pseudobulk data when counts are summed across cells, because the aggregated values retain count properties.
The Bioconductor project hosts edgeR and provides extensive documentation on its use for differential expression analysis [<a href="#ref-2">2</a>]. Researchers should follow the package vignettes and workflows to ensure correct model specification, including proper handling of the experimental design matrix and dispersion estimation.
DESeq2 and Moderated Estimation
DESeq2 also uses negative binomial models but employs a different approach to dispersion estimation and statistical testing. DESeq2 shrinks gene-wise dispersion estimates toward a fitted trend, which improves stability for genes with low counts. The method is widely used for bulk RNA-seq and transfers directly to pseudobulk single-cell data.
Both edgeR and DESeq2 require a design matrix that encodes the experimental groups and any covariates to be adjusted for. Common covariates include sex, age, batch, and sequencing depth. Researchers should include covariates that are biologically relevant and balanced across conditions to avoid confounding.
limma-voom and Precision Weighting
limma-voom is an alternative approach that transforms count data to log2 counts per million, estimates the mean-variance relationship, and assigns precision weights to each observation before fitting linear models. The method was originally developed for bulk RNA-seq and has been adapted for pseudobulk single-cell data. The heart failure study used limma-voom for cell type specific differential expression between heart failure patients and controls, identifying thousands of differentially expressed genes in cardiomyocytes and fibroblasts [<a href="#ref-4">4</a>].
limma-voom is particularly useful when the data do not fit the negative binomial assumptions well or when researchers prefer the flexibility of linear modeling. The method handles complex experimental designs and provides access to the full suite of limma functions for contrast testing and gene set analysis.
Accounting for Unequal Group Variances
Standard differential expression methods assume equal variance between experimental groups, but pseudobulk single-cell data frequently violate this assumption. Group heteroscedasticity, where one condition has substantially more variable expression than another, can hamper the detection of differentially expressed genes and inflate false discovery rates when ignored.
Two approaches that account for heteroscedastic groups are voomByGroup and voomWithQualityWeights using a blocked design, abbreviated as voomQWB. These methods were introduced specifically for pseudobulk single-cell data and demonstrated superior error control and statistical power compared to gold-standard methods when group variances are unequal [<a href="#ref-6">6</a>]. Researchers who observe substantial differences in variability between conditions should consider these approaches instead of assuming equal variances [<a href="#ref-6">6</a>].
Removing Unwanted Variation
Technical variation from batch effects, sequencing depth differences, and sample processing can obscure biological signals in pseudobulk data. Removing unwanted variation (RUV) methods estimate and remove factors associated with technical noise, but improper implementation can remove biological information and compromise both power and false positive control.
A systematic evaluation of RUV methods in pseudobulk single-cell data found that removing unwanted variation per cell type with RUV2 or RUVIII extracts factors associated with technical noise and controls the false discovery rate, even when the factor of interest is confounded with technical variation [<a href="#ref-7">7</a>]. The study also introduced RUVIII PBPS, a strategy that controls false discovery rates when technical replicates are missing, the factor of interest depends on the sources of unwanted variation, or plausible negative control genes are unavailable [<a href="#ref-7">7</a>]. Researchers should apply RUV methods cautiously and validate that the removed factors are technical instead of biological in origin [<a href="#ref-7">7</a>].
Quality Control Before Aggregation
The quality of pseudobulk profiles depends on the quality of the single-cell data before aggregation. Poor quality cells, ambient RNA contamination, and doublets can distort pseudobulk expression values and produce misleading differential expression results.
Cell Level Quality Filters
Before aggregating cells into pseudobulk profiles, researchers should apply standard single-cell quality control filters. Common filters include removing cells with low total UMI counts, low numbers of detected genes, high mitochondrial read fractions, and evidence of doublet contamination. These filters remove low quality cells that contribute noise to pseudobulk aggregates.
The specific thresholds for quality filters depend on the tissue type, sequencing platform, and experimental protocol. Researchers should examine the distributions of quality metrics across cells and choose thresholds that remove clear outliers while retaining the majority of cells. A study of gastric cancer analyzed an open-access single-cell dataset with 24 patients and 48 samples, defining 43 clusters grouped into 15 cell subtypes after quality control and clustering [<a href="#ref-9">9</a>]. The quality control decisions directly affected the reliability of the pseudobulk profiles generated from the retained cells [<a href="#ref-9">9</a>].
Doublet Detection
Doublets, where two cells are captured and sequenced together, produce expression profiles that are mixtures of two cell types. Doublets can distort pseudobulk aggregates, particularly for rare cell types where a small number of doublets can contribute disproportionately to the summed counts. Computational doublet detection methods should be applied before aggregation, and suspected doublets should be removed.
Ambient RNA Contamination
Ambient RNA from lysed cells can contaminate the droplet suspension and contribute counts to cells that do not actually express those genes. This contamination is particularly problematic for highly expressed genes and can create spurious expression signals in cell types that should not express them. Methods such as CellBender and SoupX estimate and remove ambient RNA contamination before downstream analysis. The heart failure study used CellBender for gene expression quantification, reflecting the importance of ambient RNA correction in single-nucleus data [<a href="#ref-4">4</a>].
Sample Level Quality Assessment
After aggregation, researchers should assess the quality of each pseudobulk profile. Samples with very few cells may produce sparse pseudobulk profiles with low total counts, which can be noisy and unreliable. Researchers should record the number of cells contributing to each pseudobulk profile and consider excluding profiles with too few cells from downstream analysis.
Library size, or the total number of counts in each pseudobulk profile, should be examined across samples. Large differences in library size can indicate technical variation that should be accounted for in the statistical model. Normalization methods within edgeR, DESeq2, and limma-voom adjust for library size differences, but extreme imbalances may warrant additional investigation.
Practical Workflow for Pseudobulk Analysis
The following workflow outlines the steps for constructing pseudobulk profiles and performing differential expression analysis. Researchers should adapt these steps to their specific data and experimental design.
Step 1: Define Biological Replicates
Identify the biological replicate variable in the metadata, typically the donor, patient, or sample identifier. Confirm that each condition has at least three biological replicates. Record the number of replicates per condition and the total number of cells per replicate.
Step 2: Apply Cell Level Quality Control
Filter cells based on quality metrics appropriate for the data type. Remove low quality cells, doublets, and cells with excessive ambient RNA contamination. Document the number of cells removed at each step and the final cell counts per sample.
Step 3: Cluster and Annotate Cell Types
Perform normalization, integration, dimensionality reduction, and clustering to identify cell populations. Annotate clusters based on canonical marker genes and biological knowledge. Record the marker genes used for each cell type annotation.
Step 4: Construct Pseudobulk Profiles
For each cell type and biological replicate, sum the raw counts across all cells. Create a matrix with genes as rows and sample-cell type combinations as columns. Record the number of cells contributing to each pseudobulk profile.
Step 5: Filter Low Expression Genes
Remove genes with very low total counts across samples. Common filters include requiring a minimum count in a minimum number of samples. The specific thresholds depend on the sequencing depth and the number of samples.
Step 6: Specify the Statistical Model
Define the design matrix with the condition of interest and any covariates. Choose the differential expression method based on the data characteristics and the need to account for unequal group variances or unwanted variation.
Step 7: Run Differential Expression Analysis
Apply the chosen method to the pseudobulk count matrix. Generate results tables with log fold changes, p-values, and adjusted p-values. Examine the distribution of p-values and the number of significant genes at the chosen false discovery rate threshold.
Step 8: Validate and Interpret Results
Check that significant genes are biologically plausible for the cell type and condition being compared. Compare results across cell types to identify shared and cell type specific signals. Consider whether cell type composition differences between conditions could explain the observed expression differences.
Records and Measurements for Pseudobulk Analysis
Maintaining detailed records is essential for reproducible pseudobulk analysis. Researchers should document the following information for each analysis:
Cell Counts and Composition
Record the number of cells per biological replicate before and after quality control. Record the number of cells per cell type per replicate. Cell type composition differences between conditions can confound pseudobulk differential expression results, and these proportions should be reported alongside expression results.
A study of diabetic foot ulcers performed sample-level cell fraction analysis and found that no cell type proportion differences remained statistically significant after false discovery rate correction [<a href="#ref-8">8</a>]. This analysis provided important context for interpreting the pseudobulk differential expression results, because cell type composition changes can drive apparent expression differences even when within-cell-type expression is unchanged [<a href="#ref-8">8</a>].
Library Sizes
Record the total number of counts in each pseudobulk profile. Library size variation across samples should be examined and accounted for in normalization. Extreme library size differences may indicate technical problems with specific samples.
Quality Metrics
Record the quality control thresholds applied and the number of cells removed at each step. Document the software versions and parameters used for clustering, annotation, and differential expression analysis. This information supports reproducibility and troubleshooting.
Analysis Parameters
Record the statistical model specification, including the design matrix, covariates, and dispersion estimation approach. Document any RUV or batch correction methods applied and the factors removed. These parameters directly affect the results and must be reported for transparency.
Common Failure Patterns in Pseudobulk Analysis
Several recurring problems undermine pseudobulk differential expression analyses. Recognizing these failure patterns helps researchers avoid them and interpret results correctly.
Treating Cells as Independent Replicates
The most common and consequential error is performing differential expression at the cell level without aggregation. This approach produces severely inflated false discovery rates because cells from the same donor are highly correlated. Researchers who observe thousands of differentially expressed genes from a study with only a handful of donors per condition should suspect cell-level testing or inadequate replication.
Insufficient Biological Replication
Pseudobulk analysis with only two replicates per condition has very low statistical power and cannot reliably estimate within-group variance. Results from such analyses should be treated as exploratory. Researchers should design studies with at least three replicates per condition and ideally more, depending on the expected effect sizes and variability.
Ignoring Unequal Group Variances
Standard differential expression methods assume equal variance between groups, but pseudobulk single-cell data often violate this assumption. When one condition is substantially more variable than another, standard methods can produce inflated false discovery rates. Methods that account for group heteroscedasticity, such as voomByGroup and voomQWB, should be used when unequal variances are present [<a href="#ref-6">6</a>].
Improper Removal of Unwanted Variation
RUV methods can remove biological signal when implemented incorrectly. Removing unwanted variation per cell type with RUV2 or RUVIII controls false discovery rates, but the choice of negative control genes and the implementation strategy matter [<a href="#ref-7">7</a>]. Researchers should validate that RUV factors correlate with technical variables instead of biological variables of interest [<a href="#ref-7">7</a>].
Aggregating Across Confounded Cell Types
If cell type annotations are inaccurate or if cell type composition differs substantially between conditions, pseudobulk profiles may mix signals from different cell populations. This mixing dilutes cell type specific effects and can produce misleading results. Researchers should examine cell type proportions and consider composition effects when interpreting pseudobulk results.
Overlooking Cell Type Specific Signals
Pseudobulk analysis that aggregates all cells from a sample without cell type stratification can miss cell type specific effects. A study of gastric cancer found that pseudobulk analysis showed no significant differences in mitochondrial conserved gene expression between cancer and control samples, while single-cell analysis revealed a cell type specific imbalance with tumor-associated epithelial cells showing elevated expression and non-epithelial cells showing reduced expression [<a href="#ref-9">9</a>]. This example demonstrates that whole-sample pseudobulk analysis can obscure compartment-specific biology that is visible at cell type resolution [<a href="#ref-9">9</a>].
Interpretation Limits and Biological Context
Pseudobulk differential expression results must be interpreted within the limits of the method and the biological context of the study.
Pseudobulk Reflects Average Expression
Pseudobulk analysis measures the average expression across all cells of a given type within a sample. It does not capture heterogeneity within cell types. Two samples with identical average expression could have very different distributions of expression across individual cells. Researchers interested in within-cell-type heterogeneity should complement pseudobulk analysis with cell-level approaches that examine expression distributions.
Cell Type Composition Confounds
Differences in cell type composition between conditions can drive apparent pseudobulk expression differences. If a disease condition has more activated macrophages and fewer resting macrophages, pseudobulk analysis of the macrophage compartment will reflect this composition shift even if no individual macrophage changes its expression. Researchers should report cell type proportions and consider composition effects when interpreting results.
Pseudobulk Results Depend on Annotation Quality
The cell type annotations used to stratify pseudobulk profiles directly determine the biological interpretation of results. Different annotation strategies can produce different cell type definitions and different pseudobulk profiles. Researchers should use well-established marker genes and validate annotations with multiple approaches when possible.
Direction of Effect Can Be Misleading
Rank-based pathway scoring methods can produce wrong-direction effects when non-pathway competitor genes saturate the top-rank scoring window. A benchmark study of pseudobulk pathway activity scoring methods identified rank-window competition as a mechanism by which rank-based methods lose biological signal and in extreme cases invert it [<a href="#ref-10">10</a>]. Controlled simulations demonstrated sign inversion in 40% of replicates under high competitor burden, with real-data confirmation from extracellular matrix remodeling in chronic kidney disease [<a href="#ref-10">10</a>]. Researchers should be cautious when interpreting pathway scores from rank-based methods in pseudobulk data [<a href="#ref-10">10</a>].
Deconvolution Is Not a Substitute for Pseudobulk
Computational deconvolution methods that infer cell type proportions and cell type specific expression from bulk RNA-seq data have limitations compared to matched single-cell data. A study comparing deconvolution of equine bronchoalveolar lavage fluid RNA-seq with matched single-cell data found that deconvolution primarily captured mRNA-derived cell type proportions instead of true cell counts, with moderate agreement for cell counts that improved after mRNA content correction [<a href="#ref-11">11</a>]. Recovery of cell type specific differential expression was inconsistent, with frequent cross-lineage signal spillover [<a href="#ref-11">11</a>]. Researchers should prefer pseudobulk analysis of matched single-cell data over deconvolution of bulk data when single-cell data are available [<a href="#ref-11">11</a>].
Safety and Reproducibility Context
Pseudobulk analysis involves computational methods that must be applied reproducibly to support valid scientific conclusions. The following practices support reproducibility and reduce the risk of analytical errors.
Version Control and Containerization
All software packages used for pseudobulk analysis should be versioned and documented. Containerization tools and workflow managers help ensure that analyses can be reproduced exactly. The nf-core documentation describes community standards for reproducible bioinformatics pipelines, including version control, containerization, and configuration management [<a href="#ref-12">12</a>]. Researchers should follow these standards when developing or using analysis workflows [<a href="#ref-12">12</a>].
Training and Foundational Skills
Pseudobulk analysis requires foundational skills in programming, data manipulation, and statistical analysis. The Carpentries lessons provide training in shell, Git, and programming that supports reproducible bioinformatics practice [<a href="#ref-13">13</a>]. The EMBL-EBI Training program offers learning pathways for bioinformatics data resources and practical analysis education [<a href="#ref-14">14</a>]. Researchers should invest in these foundational skills to conduct pseudobulk analysis correctly [<a href="#ref-14">14</a>][<a href="#ref-13">13</a>].
Data Resources and Repositories
Single-cell data used for pseudobulk analysis should be deposited in public repositories following community standards. The NCBI provides databases for sequence data, gene expression, and associated metadata that support data sharing and reuse [<a href="#ref-15">15</a>]. Researchers should deposit their processed data and analysis code to enable validation and replication of their results [<a href="#ref-15">15</a>].
Reporting Standards
Publications reporting pseudobulk differential expression results should include the number of biological replicates per condition, the number of cells per sample and cell type, the aggregation method, the statistical model, and the software versions used. This information allows readers to assess the validity of the analysis and compare results across studies.
A Decision Framework for Choosing Between Pseudobulk and Alternative Analysis Strategies
Researchers often assume pseudobulk analysis is the only valid approach for differential expression in single-cell data, but the choice between pseudobulk, cell-level testing, and deconvolution depends on the specific research question, experimental design, and data characteristics. This section provides a practical decision framework that researchers can apply before committing to a particular analytical strategy.
When Pseudobulk Analysis Is the Correct Choice
Pseudobulk analysis is the appropriate strategy when the research question concerns average expression differences between conditions within a defined cell type, and when the experimental design includes biological replicates. This scenario covers most case-control studies in disease research, where the goal is to identify genes that are consistently up or down regulated in a cell type across affected individuals compared to controls.
A rheumatoid arthritis study demonstrated this application by performing pseudobulk analysis by cell type on peripheral blood mononuclear cells from 18 patients and 18 matched controls, identifying 168 differentially expressed genes between conditions [<a href="#ref-3">3</a>]. The study design with balanced replication and cell type stratification made pseudobulk analysis the natural choice for detecting consistent disease-associated transcriptional changes [<a href="#ref-3">3</a>].
Pseudobulk analysis is also preferred when the downstream analysis requires gene-level statistics that integrate with established bulk RNA-seq tools and pathway databases. The statistical framework of edgeR, DESeq2, and limma-voom provides access to mature methods for dispersion estimation, contrast testing, and gene set enrichment that have been validated over decades of bulk transcriptomics research. The Bioconductor project maintains these tools with reproducible version control and extensive documentation [<a href="#ref-2">2</a>].
When Cell-Level Analysis May Be Appropriate
Cell-level differential expression testing is sometimes appropriate for exploratory analysis or hypothesis generation, but it should not be used for formal statistical inference about condition differences. Cell-level tests treat each cell as an independent observation, which violates the independence assumption because cells from the same donor share biological and technical factors. The technical literature directly addresses this failure mode, with work examining why cell-level test methods fail in single-cell RNA-seq data [<a href="#ref-1">1</a>].
Cell-level analysis can be useful for identifying genes with heterogeneous expression patterns within a cell type, or for discovering rare cell states that would be averaged away in pseudobulk profiles. A gastric cancer study found that pseudobulk analysis showed no significant differences in mitochondrial conserved gene expression between cancer and control samples, while single-cell analysis revealed a cell type specific imbalance with tumor-associated epithelial cells showing elevated expression and non-epithelial cells showing reduced expression [<a href="#ref-9">9</a>]. This example illustrates that cell-level resolution can reveal biology that pseudobulk averaging obscures, but the findings should be validated with pseudobulk or bulk approaches before drawing firm conclusions [<a href="#ref-9">9</a>].
When Deconvolution Is the Only Option
Computational deconvolution becomes necessary when researchers have bulk RNA-seq data without matched single-cell data, but still need cell type specific expression estimates. Deconvolution methods infer cell type proportions and cell type specific gene expression from bulk data using reference signatures. A study comparing deconvolution of equine bronchoalveolar lavage fluid RNA-seq with matched single-cell data found that deconvolution primarily captured mRNA-derived cell type proportions instead of true cell counts, with moderate agreement for cell counts that improved after mRNA content correction [<a href="#ref-11">11</a>]. Recovery of cell type specific differential expression was inconsistent, with frequent cross-lineage signal spillover [<a href="#ref-11">11</a>].
Researchers should use deconvolution only when single-cell data are unavailable, and should interpret deconvolution results with caution given the demonstrated limitations in recovering cell type specific expression [<a href="#ref-11">11</a>]. When single-cell data exist for the same samples, pseudobulk analysis of the matched single-cell data is preferred over deconvolution of bulk data [<a href="#ref-11">11</a>]. Newer deep learning approaches such as BLUE aim to improve deconvolution accuracy for cell type specific expression profiles, but these methods require validation against matched single-cell data before adoption [<a href="#ref-16">16</a>].
Decision Criteria for Method Selection
The following criteria help researchers select the appropriate analysis strategy for their specific situation.
Research Question Alignment
Define whether the question concerns average expression within a cell type, heterogeneity within a cell type, or cell type composition differences. Average expression questions align with pseudobulk analysis. Heterogeneity questions require cell-level approaches. Composition questions require cell fraction analysis instead of differential expression testing.
Replicate Structure Assessment
Count the number of biological replicates per condition. Pseudobulk analysis requires at least three replicates per condition for reliable variance estimation. Studies with fewer replicates should be treated as exploratory regardless of the analysis method chosen. A study of type 2 diabetes pancreatic islets found that cell-level differential expression analysis did not identify robust disease-associated genes after multiple-testing correction, while pseudobulk analysis showed an exploratory signal in the immune-enriched compartment with CXCL3 meeting a relaxed significance threshold [<a href="#ref-17">17</a>]. The subtle and compartment-specific nature of the transcriptional alterations required donor-aware cell type resolved analysis to detect any signal at all [<a href="#ref-17">17</a>].
Data Quality Evaluation
Assess the quality of cell type annotations, the number of cells per sample, and the presence of technical variation. Poor annotations undermine both pseudobulk and cell-level approaches. Low cell counts per sample produce noisy pseudobulk profiles. Technical variation may require RUV methods before pseudobulk analysis [<a href="#ref-7">7</a>].
Variance Structure Examination
Examine whether group variances are approximately equal across conditions. Standard pseudobulk methods assume equal variance, but group heteroscedasticity is common in single-cell data. When variances are unequal, methods such as voomByGroup and voomWithQualityWeights using a blocked design provide superior error control and power [<a href="#ref-6">6</a>].
Implementation Steps for Method Selection
Researchers should follow a structured process for selecting and documenting their analysis strategy.
Step 1: Document the Research Question
Write a clear statement of the biological question and the specific comparison being made. Record whether the question concerns average expression, heterogeneity, or composition.
Step 2: Inventory the Experimental Design
List the number of biological replicates per condition, the number of cells per sample, the cell types of interest, and any covariates such as sex, age, or batch. Record this information in the analysis notebook before any computational work begins.
Step 3: Assess Data Characteristics
Examine quality metrics, cell type proportions, and variance structure across samples. Create diagnostic plots showing library sizes, cell counts per sample, and expression variability within and between conditions.
Step 4: Select the Primary Analysis Strategy
Choose pseudobulk analysis when the question concerns average expression and replicates are sufficient. Choose cell-level analysis only for exploratory heterogeneity questions. Choose deconvolution only when single-cell data are unavailable.
Step 5: Document the Rationale
Record the reasons for the chosen strategy, including the research question, replicate structure, data quality, and variance characteristics. This documentation supports reproducibility and helps reviewers understand the analytical decisions.
Step 6: Plan Sensitivity Analyses
Identify alternative analysis strategies that could validate the primary results. For pseudobulk analysis, consider running a cell-level analysis to check for heterogeneity effects, or a composition analysis to rule out confounding by cell type proportions.
Common Decision Errors
Several recurring errors undermine method selection in single-cell differential expression analysis.
Defaulting to Cell-Level Testing
Researchers sometimes default to cell-level testing because it is the default output of popular single-cell analysis packages. This approach produces inflated false discovery rates and should be avoided for formal inference. The statistical failure of cell-level testing is well documented in the technical literature [<a href="#ref-1">1</a>].
Choosing Pseudobulk Without Checking Replicates
Pseudobulk analysis with insufficient replication produces unreliable results. Researchers should verify that at least three biological replicates exist per condition before committing to pseudobulk analysis. Studies with fewer replicates require different designs or should be treated as exploratory.
Ignoring Composition Effects
Pseudobulk analysis can be confounded by cell type composition differences between conditions. A study of diabetic foot ulcers found that no cell type proportion differences remained statistically significant after false discovery rate correction, but the analysis of cell fractions provided essential context for interpreting pseudobulk results [<a href="#ref-8">8</a>]. Researchers should always examine composition before interpreting expression differences [<a href="#ref-8">8</a>].
Applying Deconvolution When Single-Cell Data Exist
Deconvolution of bulk data is unnecessary and potentially misleading when matched single-cell data are available. The demonstrated limitations of deconvolution in recovering cell type specific expression, including cross-lineage signal spillover, argue for pseudobulk analysis of single-cell data whenever possible [<a href="#ref-11">11</a>].
Validation of Method Choice
Researchers should validate their method selection through sensitivity analyses and comparison with alternative approaches.
Sensitivity Analysis With Alternative Methods
Run the differential expression analysis with at least one alternative method to confirm that results are robust to methodological choices. For pseudobulk analysis, compare results from edgeR and limma-voom, or from standard methods and methods that account for unequal group variances [<a href="#ref-6">6</a>].
Comparison With Cell-Level Results
Compare pseudobulk results with cell-level results to identify genes that show consistent signals across both approaches. Genes significant in both analyses are more likely to reflect robust biological differences, while genes significant only in cell-level analysis may reflect inflated false positives.
Composition Adjustment
Repeat the pseudobulk analysis with cell type proportion as a covariate to assess whether composition differences drive the observed expression changes. If results change substantially after composition adjustment, the expression differences may reflect composition effects instead of within-cell-type changes.
External Validation
Validate key findings using independent datasets or experimental approaches. A study of ischemic cardiomyopathy used pseudobulk analysis to identify differentially expressed genes, then validated the expression of signature genes in an in vivo rat model [<a href="#ref-18">18</a>]. This external validation strengthened confidence in the computational findings [<a href="#ref-18">18</a>].
Records for Method Selection
Maintain a method selection record that documents the research question, experimental design, data characteristics, chosen strategy, and rationale. This record should include the following elements:
Research Question Statement
Record the exact biological question and the specific comparison being tested. Include the cell types of interest and the conditions being compared.
Design Inventory
List the number of biological replicates per condition, the number of cells per sample, and the covariates included in the analysis. Record the source of the data and any relevant experimental details.
Data Quality Summary
Document the quality control thresholds applied, the number of cells removed, and the final cell counts per sample and cell type. Include the software versions used for quality control and processing.
Method Selection Rationale
Record the reasons for choosing pseudobulk, cell-level, or deconvolution analysis. Include the specific criteria that led to the decision and any alternative strategies considered.
Sensitivity Analysis Results
Document the results of sensitivity analyses with alternative methods. Record any differences in results and the interpretation of those differences.
Troubleshooting Method Selection Problems
When results are unexpected or difficult to interpret, researchers should examine whether the method selection was appropriate.
Unexpected Lack of Significant Genes
If pseudobulk analysis identifies few or no significant genes, examine whether the study has sufficient power. Low cell counts per sample, high between-donor variability, or small effect sizes can all reduce power. Consider whether the research question requires cell-level resolution to detect heterogeneous effects [<a href="#ref-9">9</a>].
Excessive Significant Genes
If pseudobulk analysis identifies thousands of significant genes, examine whether the analysis was actually performed at the cell level, whether composition differences are driving results, or whether unwanted variation was improperly removed [<a href="#ref-7">7</a>].
Inconsistent Results Across Cell Types
If results differ dramatically across cell types, examine whether cell type annotations are accurate and whether cell type composition differs between conditions. Consider whether the research question requires a different stratification of cell types.
Disagreement With Published Literature
If results contradict published findings, compare the analytical methods used in both studies. Differences in aggregation methods, statistical models, quality control, or cell type annotations can produce different results. The Galaxy Training Network provides tutorials that can help researchers understand how methodological choices affect outcomes [<a href="#ref-5">5</a>].
Frequently Asked Questions
What is the difference between pseudobulk and cell-level differential expression?
Pseudobulk differential expression aggregates cells from the same biological sample and cell type into a single expression profile before statistical testing, treating each donor as one replicate. Cell-level differential expression tests each cell as an independent observation, which violates statistical independence assumptions because cells from the same donor are highly correlated. Cell-level tests produce inflated false discovery rates and should not be used for formal differential expression testing when biological replicates are available.
How many biological replicates do I need for pseudobulk analysis?
At least three biological replicates per condition are generally required for pseudobulk differential expression analysis. Two replicates per condition cannot reliably estimate within-group variance and produce results with very low statistical power. The optimal number of replicates depends on the expected effect size, between-donor variability, and the number of cells captured per sample. Studies with more replicates provide more reliable estimates of dispersion and greater statistical power.
Should I sum counts or average normalized values for pseudobulk construction?
Summing raw counts is the recommended approach when using edgeR or DESeq2, because these methods require count data and model the count distribution with negative binomial models. Averaging normalized values produces continuous data that require different statistical approaches such as limma with voom transformation. The choice of aggregation method should match the statistical framework planned for downstream analysis.
How do I handle samples with very few cells in pseudobulk analysis?
Samples with very few cells produce sparse pseudobulk profiles with low total counts, which can be noisy and unreliable. Researchers should record the number of cells contributing to each pseudobulk profile and consider excluding profiles with too few cells from downstream analysis. The threshold for minimum cell numbers depends on the sequencing depth and the cell type being analyzed.
Can I use pseudoreplicates for differential expression testing?
Pseudoreplicates, where cells from the same donor are split into multiple pseudobulk profiles, do not provide true biological replication and should not be used as the basis for formal differential expression testing. Pseudoreplicates can be useful for exploratory analysis or method validation, and some RUV strategies leverage pseudoreplicates to control false discovery rates when technical replicates are unavailable [<a href="#ref-7">7</a>]. However, true biological replication at the donor level remains essential for valid statistical inference.
What should I do if group variances are unequal in my pseudobulk data?
Standard differential expression methods assume equal variance between groups, but pseudobulk single-cell data frequently violate this assumption. Methods that account for group heteroscedasticity, such as voomByGroup and voomWithQualityWeights using a blocked design, provide superior error control and statistical power when group variances are unequal [<a href="#ref-6">6</a>]. Researchers should assess variance homogeneity and use these methods when unequal variances are present [<a href="#ref-6">6</a>].
How does cell type composition affect pseudobulk results?
Differences in cell type composition between conditions can drive apparent pseudobulk expression differences even when no individual cell changes its expression. Researchers should report cell type proportions alongside pseudobulk results and consider composition effects when interpreting results. Cell fraction analysis with false discovery rate correction can identify whether composition differences are statistically significant [<a href="#ref-8">8</a>].
When should I use deconvolution instead of pseudobulk analysis?
Deconvolution methods infer cell type proportions and cell type specific expression from bulk RNA-seq data when single-cell data are not available. However, deconvolution has limitations compared to matched single-cell data, including inconsistent recovery of cell type specific differential expression and cross-lineage signal spillover [<a href="#ref-11">11</a>]. When single-cell data are available, pseudobulk analysis of the matched single-cell data is preferred over deconvolution of bulk data [<a href="#ref-11">11</a>].
Related Bioinformatics Guides
- Proteomics Data Analysis in R: A Practical Workflow for Differential Expression and Visualization
- RNA Sequencing Data Analysis: From Raw Reads to Differential Expression
- Single-Cell RNA Sequencing Quality Control: A Practical Guide to Filtering and Metrics
- Single-Cell RNA Sequencing Depth: A Cost-Benefit Analysis for Experimental Design
- Single-Cell Sequencing Depth: How Much Is Enough?
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
[1] [Title: Why do cell-level test methods for detecting differentially expressed genes fail in single-cell RNA-seq data?](https://www.semanticscholar.org/paper/b65fef8b1161db7c4cc5722dfad6a092211e34ec). 2023. [2] [Bioconductor](https://bioconductor.org/). Bioconductor Project. [3] [Single-cell RNA-Seq analysis reveals cell subsets and gene signatures associated with rheumatoid arthritis disease activity.](https://pubmed.ncbi.nlm.nih.gov/38954480). JCI insight, 2024. [4] [Single cell transcriptomic analyses of human heart failure with preserved ejection fraction.](https://pubmed.ncbi.nlm.nih.gov/40161758). bioRxiv : the preprint server for biology, 2025. [5] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project. [6] [Modeling group heteroscedasticity in single-cell RNA-seq pseudo-bulk data.](https://pubmed.ncbi.nlm.nih.gov/37147723). Genome biology, 2023. [7] [Removal of unwanted variation in pseudobulk analysis of single-cell RNA sequencing data and the leveraging of pseudoreplicates.](https://pubmed.ncbi.nlm.nih.gov/41368194). NAR genomics and bioinformatics, 2025. [8] [Single-cell reanalysis identifies macrophage-associated transcriptional and intercellular communication features in diabetic foot ulcers.](https://doi.org/10.3389/fcimb.2026.1875481). 2026. [9] [Single-cell and pseudobulk analyses reveal hidden mitochondrial expression imbalance in gastric cancer.](https://doi.org/10.3389/fgene.2026.1826214). 2026. [10] [PathwayBench: a multi-criterion benchmark of pseudobulk pathway activity scoring methods reveals rank-window competition as a mechanism of biological signal loss in single-cell RNA-seq](https://doi.org/10.21203/rs.3.rs-10271634/v1). 2026. [11] [Cell-Type Deconvolution of Equine BALF RNA-Seq: A Critical Comparison with Matched Single-Cell Data.](https://doi.org/10.3390/genes17070773). 2026. [12] [nf-core Documentation](https://nf-co.re/docs). nf-core. [13] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries. [14] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute. [15] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information. [16] [Deconvolving cell-type-specific gene expression profiles from bulk RNA-seq samples.](https://doi.org/10.1371/journal.pcbi.1014101). 2026. [17] [Single-cell and pseudobulk transcriptomic profiling reveals immune-enriched disease- associated signatures in human pancreatic islets from type 2 diabetes](https://doi.org/10.21203/rs.3.rs-10040418/v1). 2026. [18] [Single-cell RNA sequencing pseudobulk analysis and machine learning identify candidate biomarkers for ischemic cardiomyopathy.](https://pubmed.ncbi.nlm.nih.gov/42591306). Frontiers in cardiovascular medicine, 2026.This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.