Why Are My DE Results Different Between DESeq2 and edgeR? Troubleshooting Common Discrepancies
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Discrepancies between DESeq2 and edgeR arise from distinct normalization (median of ratios vs. TMM), dispersion estimation (empirical Bayes shrinkage differences), and statistical testing (Wald vs. exact/likelihood ratio tests) methodologies. These core differences impact gene-wise fold changes and significance calls, particularly for low-count or highly variable genes.
- Input data preparation, including the use of raw counts versus pre-normalized data and the stringency of filtering thresholds for low-expression genes, is a critical determinant of result concordance. Both tools require raw integer counts, and differing filtering criteria can exclude genes from testing in one pipeline but not the other.
- Batch effects and covariates must be handled consistently through model formula specification in both tools; differing inclusion of such terms can alter which genes reach statistical significance. Examining MA plots and comparing p-value distributions are essential diagnostic steps for identifying differences in statistical testing structure.
- Troubleshooting involves systematically comparing normalization factors, examining dispersion plots for shrinkage differences, and assessing the impact of filtering thresholds. Analyzing the raw counts for discordant genes across samples is crucial to identify potential outlier samples influencing results.
- Sample size significantly impacts concordance, with DESeq2 generally identifying more genes at smaller sample sizes, while edgeR may offer higher precision for downstream classification. Consensus approaches (genes significant in both tools) are more reliable with increasing biological replicates.
- Independent validation of a subset of differentially expressed genes using orthogonal methods like RT-qPCR is recommended, especially when results between DESeq2 and edgeR diverge for genes of biological interest. Pathway-level concordance can often be maintained even when gene-level results differ.
Analysts routinely observe that DESeq2 and edgeR produce different lists of differentially expressed genes from the same raw count matrix. These differences are expected because the two tools implement distinct normalization strategies, dispersion estimation procedures, and statistical tests. A comparative evaluation of four software tools, including DESeq2 and edgeR, found that when programs used different approaches in each of the three core steps of mapping, normalization, and statistics, only 31 to 40 differentially expressed genes were found in common between programs [<a href="#ref-1">1</a>]. When two programs shared two of the three steps, the overlap increased to 53 genes [<a href="#ref-1">1</a>]. This article explains the specific methodological differences that drive these discrepancies, provides a systematic workflow for diagnosing why your results differ, and offers practical criteria for deciding when to reconcile results or escalate to a more fundamental review of the data.
At a Glance
The table below summarizes the primary sources of discrepancy between DESeq2 and edgeR, the practical symptom you will observe, and the first diagnostic step to take.
| Source of Discrepancy | Typical Symptom in Results | First Diagnostic Step |
|---|---|---|
| Normalization method difference | Genes with moderate expression show different fold changes between tools | Compare normalized count distributions and check library size factors |
| Dispersion estimation approach | Low-count genes are called significant by one tool but not the other | Examine dispersion plots and check the number of biological replicates |
| Statistical test structure | P-value distributions differ in shape, especially for genes near the significance threshold | Generate MA plots and compare the rank order of genes by adjusted p-value |
| Input data preparation | One tool runs on raw counts while the other received normalized or filtered data | Verify that both tools received the same raw count matrix and filtering steps |
| Model formula specification | Batch effects or covariates are handled differently, changing which genes reach significance | Review the design formula and check for hidden confounders |
| Filtering thresholds applied before testing | Low-expression genes are removed by one pipeline but retained by another | Document the minimum count and minimum total read filters used |
Understanding the Core Methodological Differences
Normalization Strategies
DESeq2 uses the median of ratios method for normalization. This approach calculates a size factor for each sample based on the median of the ratios of each gene count to the geometric mean across all samples. The method assumes that most genes are not differentially expressed and that the counts follow a negative binomial distribution. The median of ratios approach is robust to the presence of a small number of highly expressed genes that could otherwise skew the normalization.
edgeR uses the trimmed mean of M-values (TMM) method. This approach selects a reference sample and computes log fold changes between each sample and the reference for a subset of genes, then trims the extreme values and averages the remaining values to produce a normalization factor. TMM is designed to be robust to genes that are truly differentially expressed between conditions.
The practical consequence of these different normalization methods is that the effective library sizes differ between the two tools. A gene with moderate expression can receive a different normalized count value from each tool, which changes the calculated fold change and the statistical significance. In the E. coli comparison study, one program using DESeq2 produced more conservative fold changes of 1.5 to 3.5 fold, while three other programs reported fold changes of 15 to 178 fold for the same comparison [<a href="#ref-1">1</a>]. This dramatic difference illustrates that normalization and statistical choices can substantially alter the magnitude of reported effects.
Dispersion Estimation
Both DESeq2 and edgeR model count data with a negative binomial distribution, which requires an estimate of the dispersion parameter for each gene. The dispersion parameter captures the variability of gene expression beyond what would be expected from Poisson sampling alone.
edgeR estimates dispersion using a conditional maximum likelihood approach with empirical Bayes moderation. The tool first estimates a common dispersion across all genes, then a trended dispersion that depends on the mean expression level, and finally a genewise dispersion that is shrunk toward the trend. The amount of shrinkage depends on the number of biological replicates and the consistency of the dispersion estimates across genes.
DESeq2 estimates dispersion using a similar empirical Bayes approach but with a different implementation. The tool estimates genewise dispersion using maximum likelihood, then fits a curve through the dispersion estimates as a function of the mean expression, and finally shrinks the genewise estimates toward the fitted curve. DESeq2 also incorporates information from the distribution of dispersions across genes to determine the strength of shrinkage.
The practical consequence is that genes with low counts and high variability can receive very different dispersion estimates from the two tools. A gene with a dispersion estimate that is strongly shrunk toward the trend in one tool but less shrunk in the other can cross the significance threshold in one analysis but not the other. This effect is most pronounced in experiments with few biological replicates, where the genewise dispersion estimates are noisy and the shrinkage prior has a larger influence.
Statistical Testing Framework
DESeq2 uses a Wald test for differential expression. The test statistic is the estimated log2 fold change divided by its standard error, and the resulting p-value is calculated from a normal distribution. DESeq2 also implements an independent filtering step that automatically removes genes with very low counts from the testing procedure, which can increase statistical power by reducing the multiple testing burden.
edgeR uses an exact test for comparing two groups, which is analogous to the Fisher exact test but adapted for negative binomial data. For designs with more than two groups or with covariates, edgeR uses a likelihood ratio test or a quasi-likelihood F-test. The quasi-likelihood approach is generally recommended because it accounts for uncertainty in the dispersion estimates.
The different test statistics produce different p-value distributions, especially for genes near the significance threshold. A gene with a moderate fold change and moderate expression may receive a p-value of 0.04 from one tool and 0.06 from the other, placing it on different sides of the significance cutoff. The comparative study of E. coli data found that the overlap between tools decreased when the programs used different statistical approaches, even when the normalization method was the same [<a href="#ref-1">1</a>].
Input Data Preparation and Its Impact on Results
Raw Count Matrix Requirements
Both DESeq2 and edgeR expect a matrix of raw integer counts, where each row represents a gene or transcript and each column represents a sample. The counts should be generated from a consistent quantification pipeline, such as featureCounts, HTSeq, or a transcript quantification tool. Mixing counts from different quantification methods or different genome annotations will introduce systematic differences that affect both tools differently.
The Galaxy Training Network provides accessible workflow training for RNA-seq analysis, including guidance on generating count matrices and performing quality control [<a href="#ref-2">2</a>]. Following a standardized workflow for read trimming, alignment, and quantification reduces the risk of introducing artifacts that could cause discrepancies between tools.
Filtering Low-Expression Genes
The decision to filter low-expression genes before differential expression testing can have a large impact on the final gene lists. Both tools recommend some form of filtering, but the specific criteria differ. DESeq2 performs independent filtering automatically as part of its workflow, while edgeR typically requires the analyst to apply a filter based on counts per million (CPM) or total counts across samples.
If you apply a more aggressive filter in one pipeline than the other, genes with low but detectable expression will be tested by one tool and excluded by the other. These genes often have high variability and can be called significant by the tool that retains them, especially if the filter removes genes that would have contributed to the multiple testing correction.
Library Preparation and Batch Effects
The presence of batch effects or other technical covariates can cause discrepancies between tools if the model formula is specified differently. Both tools can incorporate covariates into the design matrix, but the default behavior differs. If one analysis includes a batch term and the other does not, the results will diverge for genes that are affected by the batch.
The IDEAMEX web server was developed to address the challenge of integrated RNA-seq analysis, allowing users to perform differential expression analysis with or without batch effect awareness using DESeq2 and edgeR among other tools [<a href="#ref-3">3</a>]. The availability of such integrated platforms highlights the importance of consistent model specification when comparing results across tools.
Practical Workflow for Diagnosing Discrepancies
Step 1: Verify Input Data Consistency
Confirm that both tools received the same raw count matrix. Check that the row names match, the column names match, and the values are identical. A common source of error is running one tool on raw counts and the other on counts that have been normalized or transformed. Both tools must receive raw integer counts.
Check the genome annotation version and the quantification method used to generate the counts. If the counts were generated with different annotation versions, the gene identifiers may not match, and the results will differ for genes that were added or removed between annotation versions.
Step 2: Compare Normalization Factors
Extract the size factors from DESeq2 and the normalization factors from edgeR. These values should be similar but not identical. If the size factors differ substantially for a particular sample, investigate whether that sample has an unusual library composition, such as a high proportion of reads from a small number of highly expressed genes.
Plot the normalized counts from both tools against each other for a set of housekeeping genes or genes known to be stably expressed. If the normalized values diverge for these genes, the normalization methods are producing systematically different results for your dataset.
Step 3: Examine Dispersion Estimates
Generate dispersion plots from both tools. The plots show the genewise dispersion estimates as a function of the mean expression, with the fitted trend and the final shrunken estimates. Compare the amount of shrinkage applied to genes in different expression ranges.
If one tool applies substantially more shrinkage than the other, genes with high variability will be called significant by the tool with stronger shrinkage. This situation is more common in experiments with few biological replicates, where the genewise dispersion estimates are unreliable and the shrinkage prior has a larger influence.
Step 4: Compare Test Statistics and P-Value Distributions
For each gene, extract the log2 fold change and the adjusted p-value from both tools. Create a scatter plot of the log2 fold changes to identify genes where the tools disagree on the direction or magnitude of the effect. Create a histogram of the p-values from each tool to check the overall distribution.
Genes with discordant fold changes between tools should be examined individually. Check the raw counts for these genes across all samples to determine whether the discrepancy is driven by a single outlier sample or by a systematic difference in normalization.
Step 5: Assess the Impact of Filtering
Review the filtering criteria applied in each pipeline. If one pipeline filtered genes based on a minimum count threshold and the other did not, the set of genes tested will differ. Re-run the analysis with identical filtering criteria to determine whether the filtering decision explains the discrepancy.
Step 6: Document and Report the Differences
Record the number of genes called significant by each tool, the number of genes called by both tools, and the number called by only one tool. Report these numbers in your methods section along with the specific version numbers of both tools and the parameters used. This documentation allows other researchers to understand the basis for your results and to compare their own analyses.
Options and Tradeoffs for Reconciling Results
Using a Consensus Approach
One common strategy is to report genes that are called significant by both DESeq2 and edgeR. This conservative approach reduces the number of false positives but may also discard true positives that one tool failed to detect. The E. coli comparison study found that the overlap between tools ranged from 31 to 53 genes depending on the similarity of the methods used [<a href="#ref-1">1</a>]. A consensus approach would report only these overlapping genes.
The tradeoff is that the consensus list may miss genes with genuine biological relevance, particularly in experiments with small effect sizes or limited sample sizes. The comparative study of DESeq2 and edgeR found that DESeq2 generally identified more differentially expressed genes than edgeR at smaller sample sizes, while the tools became more concordant as sample size increased [<a href="#ref-4">4</a>]. This finding suggests that the consensus approach becomes more reliable as the number of biological replicates increases.
Using a Single Tool with Justification
Another strategy is to select one tool based on the characteristics of your experiment and justify the choice in your methods. DESeq2 may be preferred for experiments with small sample sizes because it identifies more genes at smaller sample sizes [<a href="#ref-4">4</a>]. edgeR may be preferred when the analysis requires the quasi-likelihood F-test for complex designs or when the analyst needs the flexibility of the TMM normalization for datasets with unusual composition.
The choice should be documented with a rationale based on the experimental design and the known properties of each tool. The Bioconductor project provides official documentation for both packages, including detailed descriptions of the statistical methods and guidance on parameter selection [<a href="#ref-5">5</a>].
Using an Integrated Platform
Several platforms integrate multiple differential expression tools and provide a unified framework for comparing results. The IDEAMEX web server performs differential expression analysis with DESeq2 and edgeR among other tools and generates integrated results using correlograms, heatmaps, Venn diagrams, and text lists [<a href="#ref-3">3</a>]. The Confidence web application performs simultaneous statistical analysis of RNA-seq count data and assigns a confidence score from 1 to 4 to each gene to aid in prioritization [<a href="#ref-6">6</a>].
These integrated platforms can reduce the effort required to compare results across tools and provide a structured approach to gene prioritization. However, they may not offer the same level of control over parameters as running the tools directly.
Using a Third Tool for Validation
Some analysts use a third tool, such as limma-voom, to validate the results from DESeq2 and edgeR. If a gene is called significant by all three tools, the confidence in that result is higher. The IDEAMEX platform includes limma-voom in its integrated analysis [<a href="#ref-3">3</a>]. The tradeoff is the additional computational time and the need to understand the assumptions of a third tool.
Observations and Measurements for Troubleshooting
Sample Size Effects
The number of biological replicates has a substantial impact on the agreement between DESeq2 and edgeR. A comparative study using real and semi-simulated bulk RNA-seq datasets found that DESeq2 generally identified more differentially expressed genes than edgeR at smaller sample sizes, while the tools became more concordant as sample size increased [<a href="#ref-4">4</a>]. This observation has practical implications for experimental design.
If your experiment has only two or three biological replicates per condition, expect larger discrepancies between the tools. The dispersion estimates will be noisy, and the shrinkage parameters will have a larger influence on which genes reach significance. Consider increasing the number of biological replicates if the experimental budget allows, or interpret the results with caution and validate a subset of genes with an independent method such as RT-qPCR.
Outlier Sensitivity
Both tools show similar responses to simulated outliers, with the Jaccard similarity between tools decreasing as more swapped samples were introduced [<a href="#ref-4">4</a>]. An outlier sample can have a large influence on the normalization factors and the dispersion estimates, causing the tools to diverge for genes that are affected by the outlier.
When you identify a sample with unusual library composition or extreme counts for a subset of genes, investigate whether the sample is a technical artifact or a genuine biological outlier. If the sample is a technical artifact, consider removing it from the analysis or applying a more robust normalization method. If the sample represents genuine biological variability, consider whether the experimental design adequately captures this variability.
Effect Size and Significance Thresholds
The choice of significance threshold and fold change cutoff affects the agreement between tools. A study of Atlantic salmon hepatic transcriptomes identified 1750 differentially expressed transcripts between the 10.5 and 16.5 degree Celsius temperature groups using both DESeq2 and edgeR with a q-value threshold of 0.05 and a fold change threshold of 1.5 [<a href="#ref-7">7</a>]. The same study found 172 and 52 differentially expressed transcripts for the 10.5 versus 13.5 and 13.5 versus 16.5 comparisons, respectively [<a href="#ref-7">7</a>].
The agreement between tools is typically higher for genes with large fold changes and strong statistical significance. Genes near the significance threshold are more likely to be called by one tool but not the other. When reporting results, consider reporting the number of genes at different significance thresholds to give readers a sense of the robustness of the findings.
Records and Documentation Standards
Version Control and Parameter Documentation
Record the exact versions of DESeq2 and edgeR used in the analysis, along with the versions of R and Bioconductor. The statistical methods in both tools have evolved over time, and results from different versions may not be directly comparable. The Bioconductor project provides release-specific documentation and versioned packages [<a href="#ref-5">5</a>].
Document all parameters used in the analysis, including the normalization method, the dispersion estimation approach, the test statistic, and any filtering thresholds. This documentation should be included in the methods section of any publication or report.
Analysis Logs
Maintain a log of all analysis steps, including the commands used to generate the count matrix, the filtering steps applied, and the differential expression analysis commands. The Carpentries lessons provide foundational training on reproducible data analysis practices, including version control with Git and documentation of analysis steps [<a href="#ref-8">8</a>]. These practices are essential for troubleshooting discrepancies and for ensuring that the analysis can be reproduced by other researchers.
Reproducibility Frameworks
Consider using a workflow management system to ensure reproducibility. The nf-core documentation describes community standards for pipeline usage and configuration, providing a framework for reproducible genomic analysis [<a href="#ref-9">9</a>]. The Galaxy Training Network offers accessible workflow training that emphasizes reproducibility [<a href="#ref-2">2</a>]. These frameworks can help ensure that the analysis is consistent across runs and that the results can be traced back to specific inputs and parameters.
Common Failure Patterns and Their Solutions
Failure Pattern 1: One Tool Reports Many More Significant Genes Than the Other
If DESeq2 reports substantially more significant genes than edgeR, the most likely cause is the difference in statistical power at small sample sizes. DESeq2 generally identifies more genes at smaller sample sizes [<a href="#ref-4">4</a>]. Another possible cause is that edgeR applied a more aggressive filter for low-expression genes, removing genes that DESeq2 retained.
Solution: Check the number of genes tested by each tool. If the numbers differ substantially, review the filtering criteria. Consider running edgeR with a less aggressive filter or running DESeq2 with the same filter applied to both tools.
Failure Pattern 2: Fold Changes Differ Substantially Between Tools
If the log2 fold changes reported by the two tools differ by more than a factor of two for a substantial number of genes, the normalization methods are producing systematically different results. This situation can occur when the library composition varies substantially between samples or when a small number of genes dominate the read counts.
Solution: Examine the size factors from DESeq2 and the normalization factors from edgeR. If the factors differ substantially for any sample, investigate the library composition of that sample. Consider whether the TMM normalization in edgeR or the median of ratios normalization in DESeq2 is more appropriate for your data.
Failure Pattern 3: Genes Are Called Significant in Opposite Directions
If a gene is called significantly upregulated by one tool and significantly downregulated by the other, the discrepancy is likely driven by a small number of samples with extreme counts. This situation can occur when a single outlier sample has a large influence on the normalization factors.
Solution: Examine the raw counts for the discordant genes across all samples. Identify the samples that are driving the difference. Determine whether these samples are technical outliers or genuine biological variation. Consider removing the outlier samples and re-running the analysis.
Failure Pattern 4: Results Differ Between Runs of the Same Tool
If the results differ between runs of the same tool with the same input data, the analysis is not reproducible. This situation can occur when the tool uses a random seed for some step in the analysis or when the input data are modified between runs.
Solution: Verify that the input count matrix is identical between runs. Check whether the tool uses a random seed and whether the seed is set consistently. Use a workflow management system to ensure that the analysis steps are identical between runs.
Failure Pattern 5: Results Differ When Using Different Versions of the Same Tool
If the results differ when using different versions of DESeq2 or edgeR, the statistical methods have changed between versions. Both tools have been updated over time to improve their statistical methods and to fix bugs.
Solution: Document the exact version of each tool used in the analysis. When comparing results across studies, verify that the same version of each tool was used. Consider re-running the analysis with the current version of the tool to confirm that the results are stable.
Limitations of Differential Expression Tools
Assumptions of the Negative Binomial Model
Both DESeq2 and edgeR assume that the count data follow a negative binomial distribution. This assumption may not hold for all datasets, particularly those with unusual library compositions or with a large proportion of genes with zero counts. When the assumption is violated, the tools may produce unreliable results.
The saseR tool was developed to address some of these limitations by replacing normalization offsets to unlock bulk RNA-seq tools for differential usage and aberrant splicing applications [<a href="#ref-10">10</a>]. This development highlights the ongoing need for methods that can handle the diversity of RNA-seq data types and applications.
Sensitivity to Experimental Design
The performance of both tools depends on the experimental design, including the number of biological replicates, the balance of the design, and the presence of confounding factors. A comparative review of differential gene expression analysis methods noted that RNA-seq data analysis comprises a multi-phase workflow including read trimming, alignment, quantification, normalization, and differential expression testing [<a href="#ref-11">11</a>]. Each phase introduces potential sources of variability that can affect the final results.
Interpretation Limits
Differential expression analysis identifies genes with statistically significant changes in expression, but statistical significance does not necessarily imply biological significance. A gene with a small fold change may be statistically significant with a large sample size but may not have a meaningful biological effect. Conversely, a gene with a large fold change may not reach statistical significance with a small sample size but may have an important biological role.
The Confidence web application was developed to address the challenge of prioritizing target genes for further investigation, incorporating a confidence score from 1 to 4 to aid in gene prioritization [<a href="#ref-6">6</a>]. This approach recognizes that the statistical results from differential expression tools are only one input into the biological interpretation process.
Quality Control and Validation Approaches
Independent Validation with RT-qPCR
Independent validation of a subset of differentially expressed genes with an orthogonal method such as RT-qPCR provides confidence in the results. A study of Atlantic salmon hepatic transcriptomes found a significant positive correlation between RNA-seq and RT-qPCR results for 8 of 9 metabolic-related transcripts tested [<a href="#ref-7">7</a>]. This validation approach is particularly important when the results from DESeq2 and edgeR differ for genes of biological interest.
Cross-Study Validation
The generalizability of tool-specific gene sets can be assessed by validating the results in independent datasets. A comparative study found that in cross-study validation using four independent SARS-CoV-2 datasets, edgeR-specific genes yielded higher AUC, precision, and recall in held-out datasets [<a href="#ref-4">4</a>]. This finding suggests that the choice of tool can affect the generalizability of the results.
Pathway-Level Concordance
Even when the gene-level results differ between tools, the pathway-level results may be concordant. The comparative study found that many contrasts retained substantial pathway-level agreement between tools, although selected contrasts still showed tool-specific enriched pathways [<a href="#ref-4">4</a>]. If the biological interpretation is based on pathway enrichment instead of individual genes, the discrepancies between tools may have less impact on the conclusions.
Welfare and Safety Context for Research Applications
Ethical Use of Animal Data
When differential expression analysis is applied to data from animal studies, the analysis should be conducted with attention to the ethical use of animal data. The number of biological replicates should be sufficient to support the statistical analysis without causing unnecessary animal use. The EMBL-EBI Training provides resources on bioinformatics learning pathways that include guidance on responsible data analysis [<a href="#ref-12">12</a>].
Data Sharing and Reproducibility
The NCBI provides data resources for sharing and accessing genomic data, including sequence data and analysis results [<a href="#ref-13">13</a>]. Depositing raw count data and analysis code in public repositories supports reproducibility and allows other researchers to verify the results. The nf-core documentation provides standards for reproducible workflow configuration [<a href="#ref-9">9</a>].
Reporting Standards
When reporting the results of differential expression analysis, include the specific versions of the tools used, the parameters applied, and the criteria for significance. This information allows other researchers to understand the basis for the results and to compare their own analyses. The Galaxy Training Network provides training on reproducible analysis practices that support transparent reporting [<a href="#ref-2">2</a>].
Professional Escalation Criteria
When to Consult a Bioinformatics Specialist
If the discrepancies between DESeq2 and edgeR cannot be resolved by the troubleshooting steps described above, consult a bioinformatics specialist. Situations that warrant escalation include:
- The results differ in the direction of effect for a substantial number of genes
- The normalization factors differ by more than a factor of two for any sample
- The dispersion estimates are highly variable and the shrinkage parameters have a large influence on the results
- The results are not reproducible between runs of the same tool
- The analysis involves a complex experimental design with multiple covariates or batch effects
When to Reconsider the Experimental Design
If the discrepancies between tools are driven by a small number of outlier samples or by high variability in the data, the experimental design may need to be reconsidered. Increasing the number of biological replicates can improve the agreement between tools [<a href="#ref-4">4</a>]. If the experimental budget does not allow additional replicates, consider whether the biological question can be addressed with the available data.
When to Seek Statistical Consultation
If the statistical methods underlying the tools are not well understood, or if the analysis involves a nonstandard experimental design, seek statistical consultation. The Bioconductor project provides documentation and support for the statistical methods implemented in DESeq2 and edgeR [<a href="#ref-5">5</a>]. The Carpentries lessons provide foundational training on data analysis practices that can support the interpretation of statistical results [<a href="#ref-8">8</a>].
Building a Decision Framework for Tool Selection and Result Reconciliation
Defining the Analysis Context Before Running Either Tool
The first practical decision occurs before you execute DESeq2 or edgeR. Define the primary purpose of your analysis because the optimal tool choice depends on whether you are prioritizing sensitivity, specificity, or cross-study generalizability. A comparative evaluation of edgeR and DESeq2 across viral infection, bacterial infection, and fibrotic condition datasets found that classification models trained on edgeR-specific genes achieved higher F1 scores in 9 of 13 contrasts and more frequently reached perfect or near-perfect precision [<a href="#ref-4">4</a>]. The same study found that DESeq2 generally identified more differentially expressed genes at smaller sample sizes [<a href="#ref-4">4</a>]. These findings support a decision rule based on your experimental priorities.
Use the following criteria to select a primary tool before running the analysis:
| Analysis Priority | Recommended Primary Tool | Rationale |
|---|---|---|
| Maximum gene discovery with limited replicates | DESeq2 | Identifies more DEGs at smaller sample sizes [<a href="#ref-4">4</a>] |
| High precision for downstream classification or validation | edgeR | Higher F1 scores in 9 of 13 contrasts tested [<a href="#ref-4">4</a>] |
| Cross-study generalizability | edgeR | Higher AUC, precision, and recall in held-out datasets [<a href="#ref-4">4</a>] |
| Complex designs with multiple groups or covariates | edgeR with quasi-likelihood F-test | Accounts for uncertainty in dispersion estimates |
| Pathway-level interpretation | Either tool with pathway concordance check | Many contrasts retain substantial pathway-level agreement [<a href="#ref-4">4</a>] |
Document this decision and the rationale in your analysis notebook before running either tool. This pre-registration prevents post hoc tool selection based on which results look more favorable.
Establishing a Concordance Baseline for Your Dataset
Before troubleshooting discrepancies, establish a baseline measure of agreement between the tools for your specific dataset. The overlap between tools depends on the similarity of the methods used and the characteristics of the data [<a href="#ref-1">1</a>]. A comparative study of E. coli transcriptomes found that when programs used different approaches in each of the three core steps of mapping, normalization, and statistics, only 31 to 40 differentially expressed genes were found in common, while programs sharing two of the three steps found 53 genes in common [<a href="#ref-1">1</a>].
Calculate the Jaccard similarity index between the significant gene sets from DESeq2 and edgeR using this formula:
Jaccard similarity = (number of genes significant in both tools) / (number of genes significant in either tool)
Record this value as your dataset-specific concordance baseline. A Jaccard similarity above 0.5 generally indicates acceptable agreement for most biological interpretations. Values below 0.3 suggest that the tools are producing substantially different results that require investigation before proceeding with biological interpretation.
Implementing a Tiered Reconciliation Strategy
When the tools produce different gene lists, apply a tiered reconciliation strategy that matches the level of evidence required for your downstream application.
Tier 1: High-confidence consensus genes. Genes called significant by both tools with the same direction of effect receive the highest confidence. These genes are suitable for immediate biological interpretation and validation. The Confidence web application assigns a confidence score from 1 to 4 to aid in gene prioritization, where 1 represents low confidence and 4 represents high confidence [<a href="#ref-6">6</a>]. Genes in the consensus set typically receive the highest confidence scores.
Tier 2: Tool-specific genes with biological support. Genes called significant by only one tool should be examined for biological plausibility before being discarded. Check whether the gene belongs to a pathway that is enriched in the consensus gene set. A study of Atlantic salmon hepatic transcriptomes found that pathway analysis using ClueGO showed impacts on amino acid and lipid metabolism that were consistent with the gene-level results [<a href="#ref-7">7</a>]. If a tool-specific gene participates in a pathway already implicated by the consensus genes, it may represent a true positive that one tool failed to detect.
Tier 3: Discordant genes requiring individual review. Genes with opposite directions of effect between tools or with large fold change discrepancies require individual examination. Extract the raw counts for these genes across all samples and identify the samples driving the difference. A single outlier sample can have a large influence on the normalization factors and cause the tools to diverge [<a href="#ref-4">4</a>].
Recording Reconciliation Decisions in an Analysis Log
Maintain a structured analysis log that records each reconciliation decision. The Carpentries lessons provide foundational training on reproducible data analysis practices, including version control with Git and documentation of analysis steps [<a href="#ref-8">8</a>]. Use the following table format for each gene or gene set that requires a reconciliation decision:
| Gene ID | DESeq2 Result | edgeR Result | Reconciliation Decision | Rationale |
|---|---|---|---|---|
| Example | Significant, log2FC = 1.8 | Not significant, log2FC = 0.9 | Retained in Tier 2 | Gene participates in enriched pathway from consensus set |
| Example | Significant, log2FC = 2.1 | Significant, log2FC = -0.3 | Excluded | Opposite direction of effect, driven by one outlier sample |
This log provides a transparent record of the analytical decisions and supports reproducibility. The nf-core documentation describes community standards for pipeline usage and configuration that support reproducible genomic analysis [<a href="#ref-9">9</a>].
Setting Thresholds for Professional Escalation
Define objective criteria for when to escalate the analysis to a bioinformatics specialist or statistical consultant. These criteria should be based on measurable values instead of subjective impressions.
Escalate the analysis when any of the following conditions are met:
- The Jaccard similarity between the significant gene sets is below 0.3
- More than 10 percent of significant genes show opposite directions of effect between tools
- The normalization factors from edgeR differ from the DESeq2 size factors by more than a factor of two for any sample
- The results are not reproducible between runs of the same tool with identical input data
- The analysis involves a complex experimental design with multiple covariates that cannot be adequately modeled by either tool
The Bioconductor project provides official documentation for both packages, including detailed descriptions of the statistical methods and guidance on parameter selection [<a href="#ref-5">5</a>]. Consult this documentation before escalating to ensure that the issue is not caused by a parameter setting that can be adjusted.
Using Integrated Platforms for Structured Comparison
Integrated analysis platforms can reduce the effort required to compare results across tools and provide a structured approach to reconciliation. The IDEAMEX web server performs differential expression analysis with DESeq2 and edgeR among other tools and generates integrated results using correlograms, heatmaps, Venn diagrams, and text lists [<a href="#ref-3">3</a>]. The Confidence web application performs simultaneous statistical analysis of RNA-seq count data and assigns a confidence score from 1 to 4 to each gene to aid in prioritization [<a href="#ref-6">6</a>].
These platforms are particularly useful when the analysis must be repeated across multiple datasets or when the analyst lacks the programming expertise to run both tools directly. The EMBL-EBI Training provides bioinformatics learning pathways that include practical analysis education for researchers who want to develop these skills [<a href="#ref-12">12</a>]. The Galaxy Training Network offers accessible workflow training for RNA-seq analysis that emphasizes reproducibility [<a href="#ref-2">2</a>].
Validating the Reconciliation Outcome
After applying the tiered reconciliation strategy, validate the outcome using independent methods. A study of Atlantic salmon hepatic transcriptomes found a significant positive correlation between RNA-seq and RT-qPCR results for 8 of 9 metabolic-related transcripts tested [<a href="#ref-7">7</a>]. Select genes from each tier for validation, prioritizing Tier 1 consensus genes and Tier 2 genes with strong biological support.
For cross-study validation, test whether the reconciled gene set generalizes to independent datasets. The comparative study found that edgeR-specific genes yielded higher AUC, precision, and recall in held-out SARS-CoV-2 datasets [<a href="#ref-4">4</a>]. If your reconciled gene set performs poorly in cross-study validation, reconsider whether the reconciliation strategy was appropriate for your data.
Reviewing the Decision Framework After Each Analysis
After completing the analysis and reconciliation, review the decision framework to identify improvements for future analyses. Record the following information:
- The Jaccard similarity achieved with the selected primary tool
- The number of genes in each reconciliation tier
- The proportion of Tier 2 and Tier 3 genes that were validated by RT-qPCR or cross-study analysis
- Any parameter adjustments that improved agreement between tools
This review process builds a dataset-specific knowledge base that improves the efficiency and reliability of future analyses. The nf-core documentation provides guidance on reproducible workflow configuration that supports this iterative improvement process [<a href="#ref-9">9</a>].
Frequently Asked Questions
Why does DESeq2 report more significant genes than edgeR in my analysis?
DESeq2 generally identifies more differentially expressed genes than edgeR at smaller sample sizes, while the tools become more concordant as sample size increases [<a href="#ref-4">4</a>]. This difference is driven by the distinct normalization methods, dispersion estimation procedures, and statistical tests implemented in each tool. If your experiment has few biological replicates per condition, expect DESeq2 to report more significant genes. The difference should decrease as you add more biological replicates.
Should I use the intersection of DESeq2 and edgeR results as my final gene list?
Using the intersection of results from both tools is a conservative approach that reduces false positives but may discard true positives that one tool failed to detect. The overlap between tools depends on the similarity of the methods used and the characteristics of the dataset [<a href="#ref-1">1</a>]. Consider the biological context of your experiment and the consequences of false positives versus false negatives when deciding whether to use the intersection or the results from a single tool.
Can I run DESeq2 and edgeR on normalized counts instead of raw counts?
No. Both DESeq2 and edgeR expect a matrix of raw integer counts. The tools perform their own normalization as part of the analysis. Running the tools on normalized counts will produce invalid results because the statistical models assume raw count data with a negative binomial distribution. Always provide the raw count matrix to both tools.
How many biological replicates do I need for DESeq2 and edgeR to produce similar results?
The agreement between DESeq2 and edgeR increases with the number of biological replicates [<a href="#ref-4">4</a>]. With two or three replicates per condition, expect substantial differences between the tools. With five or more replicates per condition, the results should be more concordant. The exact number depends on the variability of your data and the magnitude of the biological effects you are studying.
What should I do if a gene is called significant by one tool but not the other?
Examine the raw counts for that gene across all samples. Check whether the discrepancy is driven by a single outlier sample or by a systematic difference in normalization. If the gene is of biological interest, validate the result with an independent method such as RT-qPCR. A study of Atlantic salmon found significant positive correlation between RNA-seq and RT-qPCR results for 8 of 9 metabolic-related transcripts tested [<a href="#ref-7">7</a>].
Do the differences between DESeq2 and edgeR affect pathway enrichment results?
Pathway-level results can be concordant even when gene-level results differ between tools. A comparative study found that many contrasts retained substantial pathway-level agreement between tools, although selected contrasts still showed tool-specific enriched pathways [<a href="#ref-4">4</a>]. If your biological interpretation is based on pathway enrichment, the discrepancies between tools may have less impact on your conclusions.
How do I report the use of DESeq2 and edgeR in my methods section?
Report the exact versions of both tools, the parameters used, and the criteria for significance. Document the normalization method, the dispersion estimation approach, and the statistical test for each tool. Report the number of genes called significant by each tool and the number called by both tools. This documentation allows other researchers to understand the basis for your results and to compare their own analyses.
What is the best way to validate my differential expression results?
Independent validation with an orthogonal method such as RT-qPCR provides confidence in the results. Cross-study validation in independent datasets can assess the generalizability of the findings [<a href="#ref-4">4</a>]. Pathway enrichment analysis can place the results in biological context. The Confidence web application provides a structured approach to gene prioritization that can support the validation process [<a href="#ref-6">6</a>].
Related Bioinformatics Guides
- How to Interpret Gene Set Enrichment Analysis Results
- Hybrid Genome Assembly: Combining Short and Long Reads for Better Results
- Functional Metagenomics: From Gene Prediction to Pathway Reconstruction
- RNA-Seq vs ChIP-Seq: Complementary Approaches for Gene Regulation
- Genomic Data Visualization Tools: Choosing and Using Them Effectively
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
[1] [A transcriptome software comparison for the analyses of treatments expected to give subtle gene expression responses.](https://pubmed.ncbi.nlm.nih.gov/35725382). BMC genomics, 2022. [2] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project. [3] [Integrative Differential Expression Analysis for Multiple EXperiments (IDEAMEX): A Web Server Tool for Integrated RNA-Seq Data Analysis.](https://pubmed.ncbi.nlm.nih.gov/30984248). Frontiers in genetics, 2019. [4] [Tool choice matters: Evaluating edgeR vs. DESeq2 for sensitivity, robustness, and cross-study performance](https://doi.org/10.1371/journal.pone.0353788). PLoS ONE, 2026. [5] [Bioconductor](https://bioconductor.org/). Bioconductor Project. [6] [Confidence: a web app for cross-platform differential gene expression analysis, gene scoring, and enrichment analysis.](https://doi.org/10.1038/s41598-026-50527-w). 2026. [7] [RNA-Seq Analysis of the Growth Hormone Transgenic Female Triploid Atlantic Salmon (Salmo salar) Hepatic Transcriptome Reveals Broad Temperature-Mediated Effects on Metabolism and Other Biological Processes.](https://pubmed.ncbi.nlm.nih.gov/35677560). Frontiers in genetics, 2022. [8] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries. [9] [nf-core Documentation](https://nf-co.re/docs). nf-core. [10] [saseR: juggling offsets unlocks RNA-seq tools for fast and scalable differential usage, aberrant splicing and expression retrieval.](https://doi.org/10.1186/s13059-026-03973-8). 2026. [11] [From bulk RNA sequencing to spatial transcriptomics: a comparative review of differential gene expression analysis methods.](https://doi.org/10.1186/s40246-025-00884-w). 2025. [12] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute. [13] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.