Why Are My Batch-Corrected RNA-seq Data Still Showing Batch Effects? Troubleshooting Common Pitfalls
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Batch correction methods assume technical variation is separable from biological variation; failure occurs when this assumption is violated, often due to confounded experimental designs where batch membership correlates with biological conditions.
- The choice of batch correction method must align with the data type: count-based methods (e.g., ComBat-seq, ComBat-ref) are essential for raw RNA-seq counts to preserve their distributional properties, while methods designed for continuous data are inappropriate for raw counts.
- Incorrect model specification, such as treating a categorical batch variable as continuous or omitting it entirely, prevents accurate estimation and removal of systematic technical variation.
- Overcorrection, where biological signal is removed along with technical noise, can be detected by examining the expression patterns of known marker genes; reducing correction strength or using methods that preserve biological variation is indicated.
- Quality control (QC) must precede batch correction to remove low-quality samples or cells that can disproportionately influence correction algorithms and distort results.
- Confounded experimental designs, where batch membership is inextricably linked to the biological variable of interest, cannot be rectified by computational batch correction and necessitate experimental redesign or cautious interpretation.
Batch correction is a computational step applied to RNA sequencing data to remove systematic technical variation that arises from processing samples in different groups, at different times, or on different instruments. When batch-corrected data still display clustering by batch, the correction has not achieved its purpose. This article explains the common reasons for failed batch correction, how to diagnose each cause, and what to do when correction does not work as expected. The guidance applies to bulk RNA-seq count data and single-cell RNA-seq data, with attention to the differences between these data types.
Understanding What Batch Correction Can and Cannot Do
Batch effects are systematic non-biological differences between groups of samples processed under different technical conditions. These differences can obscure true biological variation and compromise the reliability of differential expression analysis. A benchmark study of batch-effect correction methods for single-cell RNA sequencing data described batch-specific systematic variations as a major challenge for data integration, noting that effective batch integration requires methods that remove technical variation while preserving cell type identity (PubMed 31948481).
Batch correction methods operate on the assumption that the technical variation you want to remove is identifiable and separable from the biological variation you want to keep. When this assumption fails, correction either does nothing useful or removes biological signal along with technical noise. Understanding the limits of batch correction helps you diagnose why your corrected data still show batch structure.
Batch correction cannot fix problems that originate before the computational step. If the experimental design is confounded, if the data were processed with incompatible pipelines, or if the correction model is misspecified, no algorithm will produce clean results. The correction step is one part of a larger workflow that includes experimental design, quality control, normalization, and downstream validation.
At a Glance: Common Causes of Failed Batch Correction
| Cause | How to Detect It | Primary Fix |
|---|---|---|
| Incorrect model specification | Batch variable not included in the model or included as a continuous variable when it is categorical | Verify the model formula includes batch as a factor with the correct levels |
| Wrong data type for the method | Using a method designed for continuous data on raw counts or using a count-based method on log-transformed data | Match the method to the data type, for example use count-based methods for raw counts |
| Confounded design | Batch membership correlates with the biological condition of interest | Redesign the experiment or use methods that can separate the effects |
| Overcorrection | Biological signal disappears after correction, and known marker genes no longer distinguish cell types | Reduce correction strength or use a method that preserves biological variation |
| Correction applied before necessary quality control | Low-quality samples or cells dominate the correction and distort the result | Perform quality control before batch correction |
| Incompatible data across batches | Different gene annotations, different reference genomes, or different quantification methods | Harmonize the processing pipeline across all batches |
Core Principles of Batch Correction
The Distinction Between Technical and Biological Variation
Batch correction methods estimate the component of variation that is associated with technical factors and remove it from the data. The challenge is that technical and biological variation are not always cleanly separated. For example, if all samples from one treatment group were processed in one batch and all samples from the control group were processed in another batch, the correction method cannot distinguish between batch effects and treatment effects. This situation is called confounding, and it is a fundamental limitation of batch correction.
A study of cutaneous squamous cell carcinoma progression demonstrated the importance of this principle. The authors noted that individual RNA sequencing studies showed study-specific clustering and varied widely in their differential gene expression detection. Only after batch correction did they define a consensus set of differentially expressed genes (PubMed 38867481). This example shows that batch correction can reveal biological signal when the underlying design supports separation of technical and biological variation.
Data Type Determines Method Choice
RNA-seq data are counts. They are integer values that are skewed and over-dispersed, meaning the variance is larger than the mean. Many batch correction methods were developed for microarray data, which are continuous and approximately normally distributed after transformation. Applying these methods to raw count data can produce erroneous results because the distributional assumptions are violated.
ComBat-seq was developed specifically to address this problem. It uses a negative binomial regression model that retains the integer nature of count data, making the batch-adjusted data compatible with common differential expression software packages that require integer counts. The developers showed in realistic simulations that ComBat-seq adjusted data resulted in better statistical power and control of false positives in differential expression compared to data adjusted by other available methods (PubMed 33015620).
A refined version called ComBat-ref builds on ComBat-seq by using a pooled dispersion parameter for entire batches and preserving count data for the reference batch. This method demonstrated superior performance in simulated environments and real datasets, improving sensitivity and specificity over existing methods (PubMed 38746101).
For single-cell RNA-seq data, the choice of method depends on the analysis goals. A benchmark study compared 14 batch correction methods for single-cell data and found that Harmony, LIGER, and Seurat 3 were the recommended methods for batch integration. Harmony was recommended as the first method to try due to its significantly shorter runtime (PubMed 31948481).
The Role of Normalization
Batch correction is not the same as normalization, but the two steps interact. Normalization adjusts for differences in sequencing depth and library size across samples. Batch correction adjusts for systematic technical variation across groups of samples. If normalization is performed incorrectly, batch correction may not work because the data still contain technical artifacts that the correction method cannot identify.
For bulk RNA-seq data, normalization to a common scale such as transcripts per million (TPM) is a common practice before batch correction. One study of metabolic dysfunction-associated steatotic liver disease normalized expression matrices to TPM, log-transformed the data, and then applied ComBat for batch correction (PubMed 41185013). This workflow is common in the literature, but it is important to recognize that the choice of normalization method affects the performance of batch correction.
Practical Workflow for Diagnosing Failed Batch Correction
Step 1: Verify the Data Type and Method Match
Check what type of data you are feeding into the batch correction method. If you are using raw counts, use a count-based method such as ComBat-seq or ComBat-ref. If you are using log-transformed continuous data, methods designed for continuous data may be appropriate. Using a method designed for continuous data on raw counts will likely fail because the distributional assumptions are violated.
For single-cell RNA-seq data, the processing pipeline matters. A review of quantitative single-cell transcriptomics noted that over 50 protocols have been developed and that data processing and analysis tools are evolving fast. The review compared essential methods to process single-cell data from mapping, filtering, normalization, and batch corrections to basic differential expression analysis (PubMed 29579145). The choice of protocol and processing method affects the performance of batch correction.
Step 2: Check the Model Specification
The batch correction model must include the correct batch variable. If you have samples from three different sequencing runs, the batch variable must have three levels. If you accidentally code the batch variable as a continuous variable, the model will not correctly estimate the batch effects.
Also check whether the model includes other sources of unwanted variation. Some methods allow you to specify additional covariates such as library preparation date, sequencing lane, or technician. Including these covariates can improve correction performance if they explain meaningful variation.
Step 3: Assess Confounding
Examine the relationship between batch membership and the biological variable of interest. Create a table showing how many samples from each treatment group are in each batch. If the design is balanced, with each treatment group represented in each batch, correction can work. If the design is confounded, with all samples from one treatment group in one batch, correction cannot separate technical from biological variation.
Confounded designs require a different approach. Options include collecting additional samples to balance the design, using methods that can model the confounding structure, or acknowledging the limitation and interpreting results with caution.
Step 4: Evaluate Correction Performance
After correction, visualize the data to assess whether batch structure remains. Principal component analysis is a common tool for this purpose. Plot the first two principal components and color the points by batch. If samples from different batches are intermingled, the correction has worked. If samples still cluster by batch, the correction has failed.
For single-cell data, additional metrics are available. The benchmark study used four metrics including kBET, LISI, ASW, and ARI to evaluate batch correction performance (PubMed 31948481). These metrics quantify different aspects of batch mixing and cell type preservation. Using multiple metrics provides a more complete picture of correction performance.
Step 5: Check for Overcorrection
Overcorrection occurs when the correction removes biological signal along with technical noise. This problem is often harder to detect than undercorrection because the data look clean but the biological conclusions are wrong.
To check for overcorrection, examine known marker genes for the cell types or conditions in your study. If the marker genes no longer distinguish the groups after correction, the correction has removed too much biological variation. For example, in a study of osteoarthritis, researchers identified twelve chondrocyte clusters including ferroptosis-active homeostasis chondrocytes (PubMed 40636124). If batch correction removed the expression differences that define these clusters, the biological interpretation would be compromised.
Common Failure Patterns and Their Solutions
Failure Pattern 1: Batch Structure Persists After Correction
If you apply batch correction and the data still show batch clustering, the first thing to check is whether the method was appropriate for the data type. Count-based methods such as ComBat-seq and ComBat-ref are designed for raw count data. If you applied a method designed for continuous data to raw counts, the method may not have converged to a useful solution.
Another possibility is that the batch variable was not correctly specified. Check the levels of the batch factor and ensure that each batch is represented by enough samples. Methods generally require a minimum number of samples per batch to estimate batch effects reliably.
A third possibility is that the batch effect is nonlinear or interacts with biological variables. Some methods assume that batch effects are additive and constant across the range of expression values. If the batch effect is multiplicative or varies by expression level, the correction may be incomplete.
Failure Pattern 2: Biological Signal Disappears After Correction
Overcorrection is a common problem when the batch effect is large relative to the biological effect. The correction method cannot distinguish between the two sources of variation and removes both.
To address overcorrection, consider using a method that explicitly models biological variation. Some methods allow you to specify known biological covariates in the model, which helps the method preserve variation associated with those covariates. Alternatively, reduce the strength of the correction or use a method that is less aggressive.
For single-cell data, the choice of method matters. The benchmark study found that different methods have different tradeoffs between batch mixing and cell type preservation. Harmony was recommended as the first method to try because it balances these objectives and has a shorter runtime (PubMed 31948481).
Failure Pattern 3: Correction Works for Some Batches but Not Others
If the correction removes batch effects for some batches but not others, the problem may be that the batches differ in ways that the model does not capture. For example, if one batch was sequenced at much higher depth than the others, the correction may not fully account for this difference.
Check the quality metrics for each batch. Library size, mapping rate, and gene detection rate can vary across batches. If one batch is an outlier on these metrics, consider whether that batch should be included in the analysis at all.
Failure Pattern 4: Correction Introduces Artifacts
Sometimes batch correction creates new problems. For example, correction can produce negative expression values, which are not meaningful for count data. Or correction can create spurious differences between groups that were not present in the original data.
If you observe artifacts after correction, check whether the method is appropriate for the data type. Count-based methods preserve the integer nature of count data, which avoids the problem of negative values. If you used a method designed for continuous data on counts, consider switching to a count-based method.
Failure Pattern 5: Batch Effects Reappear in Downstream Analysis
Even if the corrected data look clean in the initial visualization, batch effects can reappear in downstream analyses. This can happen if the downstream analysis method does not account for the corrected nature of the data or if the correction was not complete.
For differential expression analysis, consider whether the analysis method expects raw counts or corrected data. Some methods require integer counts, and using corrected data that are no longer integers can cause problems. ComBat-seq was designed to produce adjusted data that are compatible with common differential expression software packages that require integer counts (PubMed 33015620).
Records and Measurements for Batch Correction
What to Record
Maintain a record of the batch correction process for reproducibility. The record should include the version of the software used, the input data files, the parameters specified, and the output files. This information allows you to reproduce the analysis and to diagnose problems if they arise.
The nf-core documentation emphasizes the importance of reproducible workflows. Community pipelines provide standards for configuration and usage that support reproducibility (nf-core Documentation). Following these standards can help ensure that your batch correction is reproducible.
Quality Metrics to Track
Track quality metrics before and after batch correction. Common metrics include the number of genes detected per sample, the distribution of expression values, and the results of principal component analysis. Comparing these metrics before and after correction helps you assess whether the correction had the intended effect.
For single-cell data, additional metrics include the number of cells per batch, the number of genes detected per cell, and the proportion of mitochondrial reads. These metrics can identify low-quality cells that may distort the correction.
Visualization for Diagnosis
Create visualizations to diagnose batch correction problems. Principal component analysis plots colored by batch and by biological condition are essential. Also create plots of known marker genes to check for overcorrection.
The Galaxy Training Network provides accessible workflow training that includes guidance on visualization and quality assessment (Galaxy Training Network). These resources can help you build the skills needed to diagnose batch correction problems.
Options and Tradeoffs in Batch Correction Methods
Count-Based Methods for Bulk RNA-seq
ComBat-seq and ComBat-ref are count-based methods designed for bulk RNA-seq data. They use negative binomial regression models that retain the integer nature of count data. The advantage of these methods is that they produce adjusted data compatible with common differential expression software packages.
ComBat-ref introduces a pooled dispersion parameter for entire batches and preserves count data for the reference batch. This approach demonstrated superior performance in simulated environments and real datasets, improving sensitivity and specificity over existing methods (PubMed 38746101).
Methods for Single-Cell RNA-seq
The benchmark study of batch correction methods for single-cell RNA-seq data compared 14 methods and recommended Harmony, LIGER, and Seurat 3 for batch integration. Harmony was recommended as the first method to try due to its shorter runtime (PubMed 31948481).
The choice of method depends on the analysis goals. Some methods prioritize batch mixing, while others prioritize cell type preservation. The benchmark study evaluated performance using four metrics including kBET, LISI, ASW, and ARI, which quantify different aspects of correction performance.
Integration Tools for Single-Cell Data
Tools such as Celline provide one-step retrieval and integrative analysis of public single-cell RNA sequencing data. Celline wraps established tools including Seurat and Scanpy for quality control and cell-type annotation, Harmony and scVI for batch correction, and Slingshot for trajectory inference. This approach standardizes the integration process and reduces the burden of manual curation (Frontiers in Bioinformatics 2025).
The developers of Celline demonstrated its utility by applying it to two mouse brain cortex datasets. The tool successfully retrieved data, standardized metadata, and enabled standard analyses that removed low-quality cells, annotated 11 major cell types, improved integration quality, and completed trajectory analysis.
Feature Selection and Its Impact
Feature selection methods affect the performance of single-cell RNA-seq data integration and querying (Nature Methods 2025). The choice of which genes to use for integration can have a substantial impact on the quality of the result. This finding highlights the importance of considering the entire analysis pipeline, beyond the batch correction step.
Multi-Omics Integration Considerations
For studies that integrate multiple data types, batch correction becomes more complex. A study of diffuse glioma used MOFA+ to derive intrinsic molecular signatures from transcriptional, methylation, and genomic profiles of 667 tumors. The researchers derived factor scores for two separate Chinese Glioma Genome Atlas batches without retraining the model, demonstrating that batch correction approaches for multi-omics data require careful consideration of how each data type contributes to the overall structure (Cancers 2026).
Common Mistakes in Batch Correction Workflows
Mistake 1: Applying Correction Before Quality Control
Batch correction should be applied after quality control has removed low-quality samples or cells. Low-quality data can distort the correction and produce misleading results. For single-cell data, quality control typically includes filtering cells based on the number of genes detected, the proportion of mitochondrial reads, and other metrics.
Mistake 2: Using Incompatible Data Across Batches
If batches were processed with different reference genomes, different gene annotations, or different quantification methods, the data may not be directly comparable. Before batch correction, harmonize the processing pipeline across all batches. This may require reprocessing some batches with the same pipeline used for the others.
Mistake 3: Ignoring the Experimental Design
The experimental design determines what batch correction can achieve. If the design is confounded, no correction method can separate technical from biological variation. Plan the experiment with batch structure in mind, distributing samples from each treatment group across batches whenever possible.
Mistake 4: Failing to Validate the Correction
Batch correction should be validated using multiple approaches. Visualize the corrected data, check known marker genes, and compare results with and without correction. If the biological conclusions change dramatically after correction, investigate whether the correction is appropriate.
Mistake 5: Treating Batch Correction as a Black Box
Batch correction methods make assumptions about the data and the nature of batch effects. Understanding these assumptions helps you choose the right method and interpret the results correctly. The Bioconductor project provides official documentation for many batch correction packages, including workflows and reproducible analysis examples (Bioconductor).
Limitations of Batch Correction
Confounded Designs Cannot Be Fixed
The most important limitation of batch correction is that it cannot fix confounded designs. If batch membership is correlated with the biological variable of interest, the correction method cannot distinguish between the two sources of variation. In this situation, the results of batch correction are unreliable, and the limitation should be acknowledged in the interpretation.
Batch Correction Can Remove Biological Signal
Overcorrection is a real risk, especially when the batch effect is large relative to the biological effect. The correction method cannot distinguish between technical and biological variation, so it may remove both. Checking known marker genes after correction helps detect this problem.
Methods Have Different Assumptions
Different batch correction methods make different assumptions about the data. Some assume continuous normally distributed data, while others are designed for count data. Some assume additive batch effects, while others can model more complex patterns. Choosing a method that matches the data type and the nature of the batch effects is essential.
Single-Cell Data Present Additional Challenges
Single-cell RNA-seq data present additional challenges for batch correction. The data are sparse, with many genes detected in only a subset of cells. The number of cells per batch can vary substantially. And the biological variation includes cell type composition, which can differ across batches for biological reasons.
A review of quantitative single-cell transcriptomics noted that over 50 protocols have been developed and that data processing and analysis tools are evolving fast (PubMed 29579145). The choice of experimental protocol and computational methods affects the performance of batch correction.
Spatial Transcriptomics Adds Another Layer of Complexity
Spatial transcriptomics technologies introduce additional technical considerations for batch correction. A study of Xenium spatial transcriptomics data systematically dissected technical noise including transcript spillover, assay specificity, panel performance, and segmentation strategies. The authors introduced SPLIT, a method that improves signal purity by resolving mixed transcriptomic signals, demonstrating that spatial data require specialized approaches beyond standard batch correction (Nature Methods 2026).
Quality Controls and Validation
Before Correction
Before applying batch correction, verify that the data are suitable for the method. Check the data type, the distribution of expression values, and the quality metrics for each sample or cell. Remove low-quality data that could distort the correction.
After Correction
After applying batch correction, validate the result using multiple approaches. Visualize the corrected data to check for remaining batch structure. Examine known marker genes to check for overcorrection. Compare the results of downstream analyses with and without correction.
Documentation
Document the batch correction process thoroughly. Record the software version, the input data, the parameters, and the output. This documentation supports reproducibility and helps diagnose problems if they arise.
Professional Escalation Criteria
Some batch correction problems require consultation with a bioinformatics specialist or statistician. Consider escalation in the following situations:
- The experimental design is confounded, and you need guidance on how to proceed.
- Batch correction produces artifacts that you cannot explain.
- The results of downstream analyses change dramatically depending on the correction method.
- You are unsure whether the correction method is appropriate for your data type.
- You need to integrate data from many different sources with different processing pipelines.
Bioinformatics training resources can help you build the skills needed to address these challenges. The EMBL-EBI Training program provides learning pathways for bioinformatics, including practical analysis education (EMBL-EBI Training). The Carpentries Lessons provide foundational computing and data skills that support reproducible analysis (The Carpentries Lessons).
A Decision Framework for Choosing Between Count-Based and Embedding-Based Batch Correction
When batch-corrected RNA-seq data still show batch structure, the problem often lies not in the execution of the correction but in the initial choice of correction strategy. Researchers frequently default to a single method without considering whether the data structure, downstream analysis goals, and biological questions align with the assumptions of that method. This section provides a practical decision framework for selecting between count-based methods such as ComBat-seq and ComBat-ref and embedding-based methods such as Harmony, with explicit criteria for when each approach is appropriate and how to document the decision.
The Core Distinction: What Each Method Family Preserves
Count-based methods operate directly on the expression matrix and produce corrected values that retain the data structure of the original measurements. ComBat-seq uses a negative binomial regression model that preserves the integer nature of count data, making the adjusted data compatible with common differential expression software packages that require integer counts (PubMed 33015620). ComBat-ref builds on this foundation by using a pooled dispersion parameter for entire batches and preserving count data for the reference batch, which improves sensitivity and specificity over existing methods (PubMed 38746101).
Embedding-based methods such as Harmony operate differently. instead of correcting the expression matrix itself, these methods learn a low-dimensional embedding that removes batch structure while preserving biological variation. The benchmark study of batch correction methods for single-cell RNA sequencing data compared 14 methods and found that Harmony, LIGER, and Seurat 3 were the recommended methods for batch integration, with Harmony recommended as the first method to try due to its significantly shorter runtime (PubMed 31948481).
The choice between these families has consequences for every downstream analysis step. Count-based methods produce corrected data that can be fed directly into differential expression tools that expect integer counts. Embedding-based methods produce corrected coordinates that are suitable for clustering, visualization, and trajectory inference but cannot be used for count-based differential expression analysis.
Decision Point 1: What Is the Primary Downstream Analysis?
The first question in the decision framework concerns the destination of the corrected data. If the primary downstream analysis is differential expression between conditions, count-based methods are the appropriate choice. The developers of ComBat-seq demonstrated in realistic simulations that the adjusted data resulted in better statistical power and control of false positives in differential expression compared to data adjusted by other available methods (PubMed 33015620). This advantage comes from preserving the integer nature of the data, which matches the assumptions of differential expression tools.
If the primary downstream analysis is clustering, cell type identification, or trajectory inference, embedding-based methods may be more appropriate. The benchmark study evaluated performance using four metrics including kBET, LISI, ASW, and ARI, which quantify different aspects of batch mixing and cell type preservation (PubMed 31948481). These metrics are designed to assess whether cells from different batches are intermingled while cells of the same type remain grouped.
For studies that require both differential expression and clustering, the decision becomes more complex. One option is to run both types of correction and use the appropriate corrected data for each downstream analysis. Another option is to use a count-based method and accept that clustering performance may be suboptimal. The choice depends on which analysis is the primary scientific question.
Decision Point 2: What Is the Batch Structure?
The number of batches and the number of samples per batch influence method choice. Count-based methods estimate batch-specific parameters, and the reliability of these estimates depends on having enough samples per batch. Methods generally require a minimum number of samples per batch to estimate batch effects reliably. If some batches have very few samples, the parameter estimates may be unstable, and the correction may be incomplete.
Embedding-based methods such as Harmony are often more robust to unbalanced batch sizes because they learn a global embedding that integrates information across all batches. The benchmark study noted that Harmony had a significantly shorter runtime than other methods, which is an advantage when working with large datasets (PubMed 31948481). For datasets with many batches or highly unbalanced batch sizes, embedding-based methods may be more practical.
The number of batches also matters. With only two batches, the batch effect is a single contrast, and count-based methods can estimate this contrast reliably if each batch has enough samples. With many batches, the number of parameters increases, and the risk of overfitting or unstable estimates grows. In this situation, embedding-based methods that share information across batches may perform better.
Decision Point 3: What Is the Biological Variation of Interest?
The nature of the biological signal affects the choice of correction method. If the biological variation of interest is a continuous gradient, such as a developmental trajectory or a disease progression score, the correction method must preserve this continuous structure. Embedding-based methods that learn a low-dimensional manifold may be better suited to preserving continuous gradients.
If the biological variation of interest is discrete, such as distinct cell types or treatment groups, the correction method must preserve the separation between these groups. The benchmark study emphasized that effective batch integration requires methods that remove technical variation while preserving cell type identity (PubMed 31948481). Count-based methods that correct each gene independently may preserve discrete group structure better than embedding-based methods that project data into a lower-dimensional space.
A study of cutaneous squamous cell carcinoma progression demonstrated the importance of preserving biological signal during batch correction. The authors noted that individual RNA sequencing studies showed study-specific clustering and varied widely in their differential gene expression detection. Only after batch correction did they define a consensus set of differentially expressed genes, including those altered in the preinvasive stages of disease development (PubMed 38867481). This example shows that the choice of correction method can determine whether biological signal is recovered or lost.
Decision Point 4: What Is the Data Type?
The data type is a non-negotiable constraint on method choice. Raw count data require count-based methods because the distributional assumptions of continuous-data methods are violated. The developers of ComBat-seq noted that many existing methods for batch effects adjustment assume the data follow a continuous, bell-shaped Gaussian distribution, but RNA-seq data are typically skewed, over-dispersed counts, so this assumption is not appropriate and may lead to erroneous results (PubMed 33015620).
For single-cell RNA-seq data, the processing pipeline matters. A review of quantitative single-cell transcriptomics noted that over 50 protocols have been developed and that data processing and analysis tools are evolving fast. The review compared essential methods to process single-cell data from mapping, filtering, normalization, and batch corrections to basic differential expression analysis (PubMed 29579145). The choice of protocol and processing method affects the performance of batch correction.
For data that have already been normalized and log-transformed, the choice is less constrained. Continuous-data methods may be appropriate, but the correction may be less effective than count-based methods applied to raw counts. One study of metabolic dysfunction-associated steatotic liver disease normalized expression matrices to TPM, log-transformed the data, and then applied ComBat for batch correction (PubMed 41185013). This workflow is common in the literature, but it is important to recognize that the choice of normalization method affects the performance of batch correction.
A Practical Decision Table
| Decision Factor | Count-Based Methods (ComBat-seq, ComBat-ref) | Embedding-Based Methods (Harmony, LIGER, Seurat 3) |
|---|---|---|
| Primary downstream analysis | Differential expression requiring integer counts | Clustering, visualization, trajectory inference |
| Batch structure | Few batches with adequate samples per batch | Many batches or unbalanced batch sizes |
| Biological variation | Discrete groups or conditions | Continuous gradients or complex cell type mixtures |
| Data type | Raw counts | Normalized or log-transformed data, single-cell embeddings |
| Output format | Corrected count matrix | Corrected embedding coordinates |
| Compatibility with differential expression tools | Direct compatibility | Requires additional steps for count-based analysis |
Implementing the Decision Framework
To apply this framework, document the decision process before running the correction. Record the primary downstream analysis, the batch structure, the biological variation of interest, and the data type. This documentation serves two purposes. First, it ensures that the choice of method is deliberate and justified. Second, it provides a record that can be revisited if the correction fails to remove batch structure.
The nf-core documentation emphasizes the importance of reproducible workflows and provides community pipeline standards for configuration and usage (nf-core Documentation). Following these standards can help ensure that the batch correction decision is documented and reproducible.
Testing the Decision with a Pilot Analysis
Before committing to a single method, run a pilot analysis with both a count-based method and an embedding-based method on a subset of the data. Compare the results using the following criteria:
- Does the corrected data still show batch structure in principal component analysis?
- Do known marker genes distinguish the biological groups of interest?
- Do the downstream analysis results differ substantially between methods?
This pilot analysis provides empirical evidence to support the method choice. The benchmark study of batch correction methods used a similar approach, comparing 14 methods across five scenarios including identical cell types with different technologies, non-identical cell types, multiple batches, big data, and simulated data (PubMed 31948481). The study design demonstrates the value of systematic comparison before committing to a single method.
When to Use Both Approaches
Some studies benefit from running both count-based and embedding-based corrections. This dual approach is appropriate when the study requires both differential expression analysis and clustering or trajectory inference. The corrected count matrix from ComBat-seq or ComBat-ref supports differential expression analysis, while the corrected embedding from Harmony supports cell type identification and trajectory analysis.
The dual approach also provides a form of validation. If both correction methods produce consistent biological conclusions, the results are more robust than if only one method was used. If the methods produce conflicting conclusions, the discrepancy indicates that the correction is sensitive to method choice, and the results should be interpreted with caution.
Documenting the Decision for Reproducibility
The decision framework should be documented as part of the analysis record. Include the following information:
- The primary downstream analysis and why it determined the method choice
- The batch structure, including the number of batches and samples per batch
- The biological variation of interest and how it was assessed
- The data type and preprocessing steps
- The results of any pilot analyses comparing methods
- The software versions and parameters used for the chosen method
The Bioconductor project provides official documentation for many batch correction packages, including workflows and reproducible analysis examples (Bioconductor). These resources can support the documentation process and help ensure that the analysis is reproducible.
Common Failure Patterns in Method Selection
Failure Pattern 1: Using a Count-Based Method on Log-Transformed Data
Count-based methods such as ComBat-seq and ComBat-ref are designed for raw count data. Applying these methods to log-transformed data violates the distributional assumptions of the negative binomial model and can produce incorrect results. If the data have already been normalized and log-transformed, use a method designed for continuous data or return to the raw counts and apply the count-based method before transformation.
Failure Pattern 2: Using an Embedding-Based Method for Differential Expression
Embedding-based methods produce corrected coordinates, not corrected expression values. These coordinates cannot be used for count-based differential expression analysis. If the primary downstream analysis is differential expression, use a count-based method that preserves the integer nature of the data.
Failure Pattern 3: Ignoring the Batch Structure
The batch structure determines what correction can achieve. If the design is confounded, with batch membership correlated with the biological variable of interest, no correction method can separate technical from biological variation. Assess the batch structure before choosing a method and before running the correction.
Failure Pattern 4: Choosing a Method Without a Pilot Analysis
The choice of correction method should be based on evidence, not habit. Run a pilot analysis with candidate methods and compare the results. The benchmark study of batch correction methods provides a model for this approach, comparing multiple methods across multiple scenarios to identify the most suitable method for batch-effect removal (PubMed 31948481).
Records and Measurements for Method Selection
Maintain a record of the method selection process. This record should include the decision factors, the candidate methods considered, the results of any pilot analyses, and the final method choice with justification. This documentation supports reproducibility and helps diagnose problems if the correction fails.
The Galaxy Training Network provides accessible workflow training that includes guidance on analysis design and method selection (Galaxy Training Network). These resources can help researchers build the skills needed to make informed method choices.
Professional Escalation Criteria for Method Selection
Some method selection problems require consultation with a bioinformatics specialist or statistician. Consider escalation in the following situations:
- The batch structure is complex, with many batches or highly unbalanced batch sizes, and you are unsure which method family is appropriate.
- The biological variation of interest is not well defined, and you need guidance on how to assess whether the correction preserves the relevant signal.
- The pilot analysis produces conflicting results between methods, and you need help interpreting the discrepancy.
- The downstream analysis requires both differential expression and clustering, and you need guidance on how to structure the analysis pipeline.
The EMBL-EBI Training program provides learning pathways for bioinformatics, including practical analysis education (EMBL-EBI Training). The Carpentries Lessons provide foundational computing and data skills that support reproducible analysis (The Carpentries Lessons). These resources can help researchers build the skills needed to address method selection challenges.
Frequently Asked Questions
Why does my data still show batch clustering after I applied ComBat?
ComBat was designed for continuous data, and applying it to raw count data can produce incomplete correction. Count-based methods such as ComBat-seq or ComBat-ref are more appropriate for RNA-seq count data because they use negative binomial regression models that match the distribution of the data (PubMed 33015620). Check whether the batch variable was correctly specified and whether the design is confounded.
What is the difference between ComBat-seq and ComBat-ref?
ComBat-seq uses a negative binomial regression model that retains the integer nature of count data. ComBat-ref builds on this approach by using a pooled dispersion parameter for entire batches and preserving count data for the reference batch. ComBat-ref demonstrated superior performance in simulated environments and real datasets, improving sensitivity and specificity over existing methods (PubMed 38746101).
How do I know if my batch correction has overcorrected the data?
Check known marker genes for the cell types or conditions in your study. If the marker genes no longer distinguish the groups after correction, the correction has removed too much biological variation. Also compare the results of downstream analyses with and without correction. If the biological conclusions change dramatically, investigate whether the correction is appropriate.
Can batch correction fix a confounded experimental design?
No. If batch membership is correlated with the biological variable of interest, the correction method cannot distinguish between technical and biological variation. The results of batch correction are unreliable in this situation. Options include collecting additional samples to balance the design or acknowledging the limitation in the interpretation.
Should I use the same batch correction method for bulk and single-cell RNA-seq data?
Not necessarily. Bulk RNA-seq data are typically analyzed with count-based methods such as ComBat-seq or ComBat-ref. Single-cell RNA-seq data are often analyzed with methods such as Harmony, LIGER, or Seurat 3, which were benchmarked for single-cell data integration (PubMed 31948481). The choice of method depends on the data type and the analysis goals.
What quality metrics should I check before batch correction?
Check the number of genes detected per sample or cell, the distribution of expression values, and the proportion of mitochondrial reads for single-cell data. Also check library size and mapping rate for bulk data. Low-quality data can distort the correction and produce misleading results.
How do I document batch correction for reproducibility?
Record the software version, the input data files, the parameters specified, and the output files. Also record the quality metrics before and after correction. This documentation supports reproducibility and helps diagnose problems if they arise. Community pipeline standards such as those described in the nf-core documentation can support reproducible workflows (nf-core Documentation).
What should I do if batch effects reappear in downstream analysis?
Check whether the downstream analysis method expects raw counts or corrected data. Some methods require integer counts, and using corrected data that are no longer integers can cause problems. Also check whether the correction was complete by visualizing the corrected data and examining known marker genes.
Related Bioinformatics Guides
- RNA-Seq Batch Effect Detection and Correction
- RNA-Seq Databases: Accessing and Using Public RNA-Seq Data
- RNA-Seq Data Analysis in Galaxy: A User-Friendly Platform
- RNA-Seq Data Analysis Workflow: From Raw Reads to Insights
- RNA-Seq vs qPCR: Validation and Comparison
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- A benchmark of batch-effect correction methods for single-cell RNA sequencing data.. Genome biology, 2020.
- Gene expression landscape of cutaneous squamous cell carcinoma progression.. The British journal of dermatology, 2024.
- ComBat-seq: batch effect adjustment for RNA-seq count data.. NAR genomics and bioinformatics, 2020.
- A novel palmitoylation-based molecular signature reveals COX6A1 as a key regulator in metabolic dysfunction-associated steatotic liver disease.. Journal of translational medicine, 2025.
- PIK3CA variants selectively initiate brain hyperactivity during gliomagenesis.. Nature, 2020.
- Single-cell transcriptome and multi-omics integration reveal ferroptosis-driven immune microenvironment remodeling in knee osteoarthritis.. Frontiers in immunology, 2025.
- Highly Effective Batch Effect Correction Method for RNA-seq Count Data.. bioRxiv : the preprint server for biology, 2024.
- Quantitative single-cell transcriptomics.. Briefings in functional genomics, 2018.
- A Vascular-Extracellular Matrix Molecular Program Identifies High-Risk Diffuse Glioma Across Independent Multi-Omics.. 2026.
- Resolving sensitivity, specificity and signal contamination in Xenium spatial transcriptomics.. 2026.
- Single-cell atlas of AML reveals age-related gene regulatory networks in t(8,21) AML.. 2026.
- YKL-40 alleviates the TNF-α-Induced chondrocyte injury in osteoarthritis in vitro.. 2026.
- Applications of AI to single-cell and spatial transcriptomics: current state-of-the-art and challenges.. 2025.
- Celline: a flexible tool for one-step retrieval and integrative analysis of public single-cell RNA sequencing data.. 2025.
- Feature selection methods affect the performance of scRNA-seq data integration and querying. Nature Methods, 2025.
- Single-cell RNA sequencing analysis of T helper cell differentiation and heterogeneity. Biochimica Et Biophysica Acta Molecular Cell Research, 2022.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.