Differential Abundance Testing in Metagenomics: DESeq2 vs. edgeR vs. ANCOM-BC
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- DESeq2 and edgeR employ negative binomial models with empirical Bayes dispersion estimation and normalization strategies (median-of-ratios and TMM, respectively) that assume a stable reference set of taxa, making them sensitive for moderate community shifts but potentially prone to inflated false discovery rates in metagenomics due to implicit handling of compositionality.
- ANCOM-BC explicitly models the compositional nature of metagenomic data by applying bias correction to log-transformed counts, offering superior false discovery rate control and suitability for studies expecting large community shifts, though it may exhibit reduced power for small effects and require careful handling of zero counts.
- All three tools can be challenged by extreme sparsity and library size variation; ANCOM-BC's bias correction is less stable with many zeros, while DESeq2 and edgeR may produce unstable estimates for rare taxa due to high dispersion and limited shrinkage.
- Robust differential abundance analysis necessitates thorough data preparation, including quality control and filtering of rare taxa, followed by exploratory data analysis (e.g., PCA) to identify potential confounding factors before applying statistical tests.
- Interpreting results requires caution, as differential abundance reflects relative, not absolute, changes; validation with independent methods or alternative tools with different statistical assumptions is crucial for confirming biological significance.
Researchers analyzing metagenomic data face a central decision when identifying taxa that differ between experimental groups: which differential abundance testing tool should they use. DESeq2, edgeR, and ANCOM-BC represent three widely adopted approaches, each built on different statistical assumptions about how sequencing data behave. This article compares these tools in terms of their mathematical foundations, performance on sparse and compositional data, practical usability, and suitability for specific research questions. The goal is to help researchers match their data characteristics and study design to the most appropriate tool instead of defaulting to a familiar option.
Understanding the Compositional Nature of Metagenomic Count Data
Metagenomic sequencing produces count tables where each entry represents the number of sequencing reads assigned to a particular taxon in a particular sample. These counts do not measure absolute abundances of organisms in the original environment. Instead, they reflect relative proportions constrained by the total number of reads generated for each sample, a quantity known as library size or sequencing depth [<a href="#ref-1">1</a>]. Two samples with identical microbial communities but different sequencing depths will produce different count values for the same taxa, even though the true biological abundances are identical.
This constraint has profound implications for statistical analysis. When one taxon increases in abundance, the counts of all other taxa must decrease proportionally to maintain the fixed library size, regardless of whether those other taxa actually changed in the environment. This property, known as compositionality, creates spurious negative correlations between taxa and can generate false signals of differential abundance if not properly accounted for [<a href="#ref-2">2</a>]. A researcher comparing counts between groups without addressing compositionality may conclude that a taxon differs between conditions when the difference merely reflects shifts in other taxa.
Sparsity compounds the problem. Metagenomic datasets typically contain many taxa that are present in only a subset of samples, and many taxa that appear at very low counts. Zero-inflated distributions arise because a taxon may be genuinely absent from a sample, present but below the detection limit of sequencing, or lost during bioinformatics processing. Tools that assume count data follow a particular distribution must handle these zeros appropriately, and different tools make different choices about how to model or transform them [<a href="#ref-2">2</a>].
The practical consequence is that no single statistical test can be applied naively to raw count tables. Researchers must either normalize the data to make samples comparable, transform the data to break the compositional constraint, or use a model that explicitly accounts for the relative nature of the measurements. DESeq2, edgeR, and ANCOM-BC represent three different strategies for addressing these challenges, and understanding their differences requires examining how each tool handles normalization, distributional assumptions, and compositionality.
Core Statistical Principles of DESeq2, edgeR, and ANCOM-BC
DESeq2: Negative Binomial Modeling with Median-of-Ratios Normalization
DESeq2 models raw count data using a negative binomial distribution, which accommodates the overdispersion typical of biological count data where variance exceeds the mean. The tool estimates a size factor for each sample based on the median of ratios of each taxon's count to the geometric mean across all samples. This normalization approach assumes that most taxa do not change between conditions, so the median ratio provides a robust estimate of sequencing depth that is not skewed by highly abundant or highly variable taxa [<a href="#ref-1">1</a>].
The DESeq2 model fits a generalized linear model for each taxon, with the negative binomial distribution parameterized by the mean and a dispersion parameter. The dispersion is estimated using an empirical Bayes approach that shrinks per-taxon dispersion estimates toward a fitted trend, which improves stability for taxa with low counts. Statistical significance is assessed using a Wald test or a likelihood ratio test for more complex designs [<a href="#ref-1">1</a>].
DESeq2 performs well when the assumption of a negative binomial distribution holds and when most taxa are not differentially abundant, which justifies the median-of-ratios normalization. However, the method does not explicitly account for compositionality beyond the normalization step. If a large fraction of the community shifts between conditions, the normalization assumption may be violated, and the tool can produce inflated false positive rates [<a href="#ref-2">2</a>].
edgeR: Negative Binomial Modeling with Trimmed Mean of M-Values Normalization
edgeR also uses a negative binomial model for count data but employs a different normalization strategy. The trimmed mean of M-values (TMM) approach calculates a normalization factor for each sample by comparing the log-fold-changes of a subset of taxa that are deemed not differentially abundant. The trimming removes the most extreme log-fold-changes and the most extreme absolute abundances, leaving a core set of taxa assumed to be stable between samples [<a href="#ref-1">1</a>].
Like DESeq2, edgeR estimates dispersion using empirical Bayes methods, with separate steps for estimating the common dispersion across all taxa, the trended dispersion as a function of abundance, and the tagwise dispersion for individual taxa. The tool supports a range of experimental designs through its generalized linear model framework and provides multiple testing options, including exact tests for simple two-group comparisons and likelihood ratio tests for more complex designs [<a href="#ref-1">1</a>].
The TMM normalization in edgeR is more robust than simple total-count normalization because it ignores taxa with extreme fold-changes that could skew the scaling factor. However, edgeR shares the same fundamental limitation as DESeq2: the normalization assumes that a stable reference set of taxa exists, and the model does not explicitly account for the compositional constraint. Benchmarking studies have shown that edgeR achieves high sensitivity in detecting true differences but may fail to control the false discovery rate adequately in some metagenomic contexts [<a href="#ref-2">2</a>].
ANCOM-BC: Bias Correction for Compositional Data
ANCOM-BC takes a fundamentally different approach by explicitly modeling the compositional structure of the data. The method begins with a log transformation of the observed counts and then estimates a bias term for each sample that accounts for differences in sampling fractions. The sampling fraction represents the ratio of the expected total abundance in a sample to the observed library size, and differences in sampling fractions between groups can create spurious signals of differential abundance [<a href="#ref-2">2</a>].
The bias correction is achieved through an iterative procedure that estimates the sampling fractions using the observed data, then adjusts the log-counts accordingly. After correction, the tool fits a linear regression model for each taxon, with the bias-corrected log-abundances as the response variable. Hypothesis testing is performed using standard linear model inference, and multiple testing correction is applied to control the false discovery rate [<a href="#ref-2">2</a>].
ANCOM-BC is designed specifically for compositional data and does not rely on the assumption that most taxa are unchanged between conditions. This makes it particularly attractive for metagenomic studies where large shifts in community composition are expected. The method has been shown to control the false discovery rate effectively while maintaining reasonable sensitivity, though it may be less powerful than DESeq2 or edgeR in scenarios where those tools' assumptions are met [<a href="#ref-2">2</a>].
Performance on Sparse and Compositional Data
Sensitivity and False Discovery Rate Control
The choice of differential abundance tool involves a tradeoff between sensitivity, the ability to detect true differences, and specificity, the ability to avoid false positives. Benchmarking studies using both simulated and real datasets have consistently shown that DESeq2 and edgeR achieve high sensitivity in detecting differentially abundant taxa [<a href="#ref-2">2</a>]. This means they are likely to find true biological signals when they exist. However, this sensitivity comes at a cost: both tools have been observed to inadequately control the false discovery rate in metagenomic contexts, producing more false positives than expected [<a href="#ref-2">2</a>].
ANCOM-BC, by contrast, prioritizes false discovery rate control. The bias correction procedure and the explicit handling of compositionality reduce the number of spurious findings, but this may come at the expense of reduced power to detect smaller true effects [<a href="#ref-2">2</a>]. For researchers whose primary concern is avoiding false discoveries, particularly in exploratory studies where many candidate taxa will be subjected to downstream validation, ANCOM-BC may be the safer choice.
The appropriate balance between sensitivity and specificity depends on the research question. A hypothesis-driven study testing a small number of candidate taxa may tolerate a higher false discovery rate because each finding will be individually validated. A discovery-oriented study screening thousands of taxa for potential biomarkers may require stricter false discovery rate control to avoid overwhelming downstream resources with false leads [<a href="#ref-2">2</a>].
Handling of Zero Counts and Rare Taxa
Sparse data present challenges for all three tools, but they respond differently. DESeq2 and edgeR handle zeros through their negative binomial models, which can accommodate zero counts as part of the distribution. However, taxa with very low counts across most samples have high dispersion estimates, and the empirical Bayes shrinkage may not fully stabilize their variance estimates. This can lead to unstable fold-change estimates and unreliable p-values for rare taxa [<a href="#ref-1">1</a>].
ANCOM-BC's log transformation requires handling of zeros before analysis. The tool adds a small pseudocount to zero values to enable log transformation, and the choice of pseudocount can influence results for rare taxa. The bias correction procedure also becomes less stable when many taxa have zero counts in some samples, because the sampling fraction estimates rely on the observed counts across all taxa [<a href="#ref-2">2</a>].
For datasets with extreme sparsity, where many taxa are present in fewer than 10% of samples, none of the three tools performs optimally. Researchers may need to filter rare taxa before analysis, apply a tool specifically designed for zero-inflated data, or interpret results for rare taxa with caution. The filtering threshold should be based on the study design and the minimum count required for reliable statistical inference, and this decision should be documented in the analysis protocol.
Impact of Library Size Variation
Unequal sequencing depths between samples are a common feature of metagenomic datasets, arising from differences in DNA extraction efficiency, library preparation, or sequencing runs. All three tools attempt to correct for these differences through normalization or bias correction, but their approaches differ in robustness.
DESeq2's median-of-ratios normalization is relatively robust to extreme library size differences because the median is less affected by outliers than the mean. edgeR's TMM normalization similarly trims extreme values before calculating scaling factors. ANCOM-BC's bias correction explicitly estimates sampling fractions, which should account for library size differences even when the compositional structure is complex [<a href="#ref-2">2</a>].
However, extreme variation in library sizes, such as a 100-fold difference between the smallest and largest samples, can strain all three methods. In such cases, samples with very low sequencing depth contribute noisy estimates that may dominate the analysis. Researchers should examine library size distributions before analysis and consider whether samples with extremely low depth should be excluded or whether the analysis should be stratified by sequencing batch.
Practical Workflow for Differential Abundance Testing
Step 1: Data Preparation and Quality Control
Before any differential abundance testing, researchers must ensure that the input count table is accurate and complete. The count table should be generated from a well-validated bioinformatics pipeline that includes quality trimming, adapter removal, and taxonomic classification [<a href="#ref-3">3</a>]. For 16S rRNA amplicon data, the pipeline typically involves denoising to resolve amplicon sequence variants, while shotgun metagenomic data requires assembly or direct read classification against reference databases [<a href="#ref-4">4</a>].
Quality control steps should include examining the number of reads per sample, the number of taxa detected per sample, and the distribution of counts across taxa. Samples with very low read counts may need to be excluded or re-sequenced. Taxa present in very few samples may be filtered to reduce the multiple testing burden and improve the stability of normalization estimates. The filtering criteria should be pre-specified and documented to avoid bias.
The choice of taxonomic level for analysis also affects results. Genus-level analysis reduces sparsity compared to species-level analysis but may obscure biologically relevant differences between closely related species. Researchers should consider the resolution of their taxonomic classification and the biological question being addressed when choosing the analysis level [<a href="#ref-5">5</a>].
Step 2: Exploratory Data Analysis
Before running differential abundance tests, researchers should visualize the data to understand its structure. Principal component analysis or ordination methods can reveal whether samples cluster by experimental group or by technical factors such as sequencing batch [<a href="#ref-6">6</a>]. Alpha diversity metrics describe within-sample diversity, while beta diversity metrics describe between-sample community differences [<a href="#ref-5">5</a>].
These exploratory analyses serve several purposes. They can identify outliers that may need to be removed or investigated. They can reveal unexpected clustering that suggests confounding variables. They can also provide context for interpreting differential abundance results, since a taxon that differs between groups may be part of a broader community shift instead of an isolated change [<a href="#ref-6">6</a>].
Tools such as MetaDAVis provide interactive interfaces for these exploratory analyses, allowing researchers without extensive programming experience to examine their data before proceeding to statistical testing [<a href="#ref-6">6</a>]. The R Shiny application supports taxonomic abundance distribution, alpha and beta diversity analyses, dimension reduction, correlation analysis, heatmap generation, and differential abundance analysis in a single platform [<a href="#ref-6">6</a>].
Step 3: Running Differential Abundance Tests
The three tools are implemented as R packages and are available through Bioconductor, which provides documentation, installation instructions, and workflow examples for reproducible genomic analysis [<a href="#ref-7">7</a>]. Researchers should ensure they are using current versions of the packages and should record the version numbers in their analysis documentation.
For DESeq2 and edgeR, the analysis workflow involves creating a count matrix, defining the experimental design, estimating size factors or normalization factors, estimating dispersion, and fitting the statistical model. Both tools require the user to specify the experimental design formula, which should include all relevant biological and technical variables. For ANCOM-BC, the workflow involves log transformation, bias correction, and linear model fitting, with the experimental design specified similarly.
The choice of multiple testing correction method should be made before analysis. The Benjamini-Hochberg procedure for controlling the false discovery rate is commonly used and is the default in many tools. The significance threshold should be pre-specified, typically at a false discovery rate of 0.05, though more stringent thresholds may be appropriate for studies with many taxa or when downstream validation is expensive.
Step 4: Interpreting and Validating Results
Differential abundance testing produces a list of taxa with associated fold-changes, p-values, and adjusted p-values. These results should be interpreted in the context of the study design and the exploratory analyses. A taxon that shows a large fold-change but is present in very few samples may be less reliable than a taxon with a smaller fold-change that is consistently detected across samples.
Validation of results can take several forms. Technical validation involves confirming that the differential abundance signal is reproducible using an independent method, such as quantitative PCR or a different bioinformatics pipeline. Biological validation involves assessing whether the identified taxa are consistent with known biology or with results from other studies. Computational validation involves running the analysis with different tools or parameters to assess the robustness of the findings [<a href="#ref-2">2</a>].
Researchers should also consider whether the identified taxa are biologically plausible. A taxon that is known to be rare in the environment but shows a large fold-change may warrant additional scrutiny. The functional implications of differential abundance can be explored through pathway prediction tools, which infer the metabolic capabilities of the community based on taxonomic composition [<a href="#ref-5">5</a>].
Comparing the Three Tools: A Decision Framework
At a Glance
| Feature | DESeq2 | edgeR | ANCOM-BC |
|---|---|---|---|
| Statistical model | Negative binomial with empirical Bayes dispersion | Negative binomial with empirical Bayes dispersion | Linear model on bias-corrected log-abundances |
| Normalization approach | Median-of-ratios size factors | Trimmed mean of M-values (TMM) | Explicit sampling fraction bias correction |
| Compositionality handling | Implicit through normalization assumption | Implicit through normalization assumption | Explicit through bias correction |
| Sensitivity | High | High | Moderate |
| False discovery rate control | Can be inadequate in metagenomic contexts | Can be inadequate in metagenomic contexts | Generally adequate |
| Best suited for | RNA-seq-like data, moderate community shifts | RNA-seq-like data, moderate community shifts | Metagenomic data with large community shifts |
| Handling of sparse data | Moderate, may be unstable for rare taxa | Moderate, may be unstable for rare taxa | Requires pseudocount, may be unstable with extreme sparsity |
| Ease of use | Well documented, many tutorials | Well documented, many tutorials | Less familiar to many researchers |
| Availability | Bioconductor R package | Bioconductor R package | Bioconductor R package |
When to Use DESeq2
DESeq2 is a reasonable choice when the study design resembles a typical RNA-seq experiment, with a moderate number of samples per group and the expectation that most taxa do not change between conditions. The tool is well documented, widely used, and supported by extensive training materials [<a href="#ref-7">7</a>]. Researchers who are already familiar with DESeq2 for gene expression analysis can apply the same workflow to metagenomic count data with minimal adjustment.
However, researchers should be cautious when using DESeq2 for metagenomic data with large community shifts. If a substantial fraction of the microbial community changes between groups, the median-of-ratios normalization assumption may be violated, leading to inflated false positive rates [<a href="#ref-2">2</a>]. In such cases, the results should be interpreted with caution and ideally validated with an alternative method.
When to Use edgeR
edgeR is similar to DESeq2 in its statistical approach and is appropriate in many of the same scenarios. The TMM normalization is robust to outliers and performs well when a stable reference set of taxa exists. edgeR offers flexible experimental design options and is well supported by documentation and training resources [<a href="#ref-7">7</a>].
The choice between DESeq2 and edgeR often comes down to personal preference and familiarity, as both tools perform similarly in benchmark studies [<a href="#ref-1">1</a>]. Researchers may choose to run both tools and compare results, focusing on taxa that are identified by both methods as a conservative approach to reducing false positives.
When to Use ANCOM-BC
ANCOM-BC is the preferred choice when the research question involves large community shifts or when the compositional nature of the data is a primary concern. The explicit bias correction addresses the sampling fraction differences that can confound other methods, and the tool has been shown to control the false discovery rate effectively in metagenomic contexts [<a href="#ref-2">2</a>].
ANCOM-BC is also appropriate when the study involves longitudinal designs or complex experimental structures, as the linear model framework can accommodate these designs. The tool is less familiar to many researchers than DESeq2 or edgeR, but it is available through Bioconductor with documentation and examples [<a href="#ref-7">7</a>].
Using Multiple Tools for Robustness
Given the tradeoffs between sensitivity and false discovery rate control, many researchers choose to run multiple differential abundance tools and compare results. Taxa that are identified by multiple tools with different statistical assumptions are more likely to represent true biological differences than taxa identified by only one tool [<a href="#ref-2">2</a>].
A common strategy is to run DESeq2 or edgeR for sensitivity and ANCOM-BC for false discovery rate control, then focus on the intersection of results. This approach sacrifices some sensitivity but provides greater confidence in the identified taxa. The specific combination of tools should be documented in the analysis protocol, and the results from each tool should be reported transparently.
Common Failure Patterns and How to Avoid Them
Failure Pattern 1: Ignoring Compositionality
The most common error in differential abundance analysis is treating count data as if it measured absolute abundances. Researchers who compare raw counts or simple proportions between groups without accounting for the compositional constraint will produce results that reflect library size differences and community shifts instead of true biological differences [<a href="#ref-2">2</a>].
Prevention: Use tools that explicitly address compositionality, such as ANCOM-BC, or apply appropriate normalization before analysis. Interpret results in the context of the overall community structure instead of as isolated taxon changes.
Failure Pattern 2: Inadequate False Discovery Rate Control
Researchers who rely solely on DESeq2 or edgeR may report many differentially abundant taxa that do not replicate in validation studies. The high sensitivity of these tools comes with a cost of increased false positives, particularly in datasets with high sparsity or large community shifts [<a href="#ref-2">2</a>].
Prevention: Apply multiple testing correction and consider using a more stringent significance threshold. Validate findings with an independent method or an alternative statistical tool. Report the false discovery rate control method used in publications.
Failure Pattern 3: Overlooking Library Size Variation
Samples with very different sequencing depths can dominate the analysis if normalization is inadequate. A sample with 10 times more reads than another sample will have higher counts for all taxa, creating spurious differences if the library size is not accounted for [<a href="#ref-1">1</a>].
Prevention: Examine library size distributions before analysis. Use tools with robust normalization procedures. Consider excluding samples with extremely low read counts or stratifying the analysis by sequencing batch.
Failure Pattern 4: Applying RNA-seq Tools Without Adjustment
DESeq2 and edgeR were originally developed for RNA-seq gene expression analysis, and their assumptions may not hold for metagenomic data. The compositional structure of microbial communities differs from the structure of transcriptomes, and the normalization assumptions may be violated [<a href="#ref-1">1</a>].
Prevention: Understand the assumptions of each tool and assess whether they are met for the specific dataset. Consider using tools designed for metagenomic data, such as ANCOM-BC, or validate RNA-seq tool results with a compositional method.
Failure Pattern 5: Ignoring Rare Taxa
Rare taxa present in only a few samples are difficult to model reliably, and their fold-change estimates can be unstable. Researchers may either miss true differences in rare taxa or report spurious differences driven by a single sample [<a href="#ref-2">2</a>].
Prevention: Apply pre-filtering to remove taxa with very low counts or very low prevalence. Interpret results for rare taxa with caution. Consider aggregating to a higher taxonomic level to reduce sparsity.
Failure Pattern 6: Confounding Experimental and Technical Variables
If experimental groups differ in sequencing batch, DNA extraction date, or other technical factors, differential abundance results may reflect technical artifacts instead of biological differences. This is a particular risk in studies where samples are processed in batches [<a href="#ref-6">6</a>].
Prevention: Include technical variables in the experimental design model. Use exploratory analysis to check for batch effects. Consider using tools that can accommodate complex designs, such as the generalized estimating equation approach implemented in some frameworks [<a href="#ref-2">2</a>].
Records and Measurements for Reproducible Analysis
Documentation Requirements
Reproducible differential abundance analysis requires detailed documentation of every step in the workflow. The documentation should include the version of the bioinformatics pipeline used to generate count tables, the reference database used for taxonomic classification, the filtering criteria applied, the normalization method, the statistical model, and the multiple testing correction procedure [<a href="#ref-7">7</a>].
The R environment and package versions should be recorded, as results can vary between versions. The random seed should be set and recorded if any stochastic elements are involved in the analysis. All parameters that could affect results should be documented, including pseudocount values, filtering thresholds, and significance levels.
Data Storage and Sharing
Raw sequencing data should be deposited in public repositories such as the NCBI Sequence Read Archive, which provides infrastructure for storing and sharing sequencing data [<a href="#ref-4">4</a>]. Processed count tables and analysis scripts should be shared alongside publications to enable replication and reanalysis.
The NCBI provides a range of data resources for metagenomic research, including databases for raw sequences, processed data, and metadata [<a href="#ref-4">4</a>]. Researchers should follow community standards for data deposition and metadata annotation to ensure that their data can be reused by others.
Analysis Scripts and Workflow Management
Analysis scripts should be written in a clear, commented style that allows others to understand and reproduce the analysis. Workflow management tools can help ensure that analyses are run consistently and that results are traceable to specific inputs and parameters [<a href="#ref-8">8</a>].
Community standards for reproducible workflows emphasize version control, containerization, and automated testing [<a href="#ref-8">8</a>]. Researchers should use version control for their analysis scripts and consider using containerized environments to ensure that the analysis can be reproduced in the future.
Training and Skill Development for Differential Abundance Analysis
Foundational Bioinformatics Skills
Differential abundance analysis requires a foundation in bioinformatics and statistical computing. Researchers should be comfortable working with the command line, managing files, and writing scripts in R or Python. The Carpentries provides lessons on foundational computing, data management, shell, Git, and programming that are relevant to researchers entering this field [<a href="#ref-9">9</a>].
The EMBL-EBI Training program offers learning pathways for bioinformatics data resources and practical analysis education [<a href="#ref-3">3</a>]. These resources can help researchers understand the databases and tools used in metagenomic analysis and develop the skills needed for differential abundance testing.
Tool-Specific Training
Bioconductor provides documentation, vignettes, and workflow examples for DESeq2, edgeR, and ANCOM-BC [<a href="#ref-7">7</a>]. Researchers should work through these examples with their own data to understand how the tools behave and how to interpret their output.
The Galaxy Training Network offers accessible workflow training for bioinformatics analysis, including tutorials that cover metagenomic data processing and analysis [<a href="#ref-10">10</a>]. These tutorials provide hands-on experience with the tools and workflows used in differential abundance testing.
Statistical Literacy
Understanding the statistical assumptions underlying each tool is essential for choosing the appropriate method and interpreting results. Researchers should be familiar with the negative binomial distribution, empirical Bayes methods, multiple testing correction, and the concept of compositionality [<a href="#ref-2">2</a>].
Statistical literacy also involves understanding the limitations of each approach. No tool can perfectly distinguish true biological differences from technical artifacts, and researchers should interpret results with appropriate caution. Consulting with a statistician or bioinformatics specialist may be appropriate for complex study designs or when results are unexpected.
Limitations and Interpretation Caveats
Technical Limitations of Each Tool
DESeq2 and edgeR share the limitation that their normalization procedures assume most taxa are unchanged between conditions. When this assumption is violated, the normalization factors are biased, and the resulting p-values are unreliable [<a href="#ref-2">2</a>]. The tools also do not explicitly model the compositional constraint, which can lead to spurious correlations between taxa.
ANCOM-BC addresses compositionality but has its own limitations. The bias correction procedure requires sufficient data to estimate sampling fractions reliably, and the method may be less stable with small sample sizes or extreme sparsity. The choice of pseudocount for zero values can influence results, and the method may be less powerful than DESeq2 or edgeR when the data meet those tools' assumptions [<a href="#ref-2">2</a>].
Biological Interpretation Limits
Differential abundance analysis identifies taxa whose relative abundance differs between groups, but this does not necessarily indicate that the absolute abundance of the taxon differs. A taxon may appear to increase in relative abundance simply because other taxa decreased. Researchers should be cautious about interpreting differential abundance results as evidence of changes in absolute abundance [<a href="#ref-1">1</a>].
The functional implications of differential abundance are also uncertain. A taxon that differs between groups may or may not be causally related to the phenotype of interest. Functional prediction tools can provide hypotheses about the metabolic implications of community changes, but these predictions require experimental validation [<a href="#ref-5">5</a>].
Study Design Limitations
The power of differential abundance analysis depends on the number of samples per group, the effect size, and the variability within groups. Studies with few samples per group have limited power to detect small differences, and the false discovery rate may be poorly controlled with small sample sizes [<a href="#ref-2">2</a>].
Longitudinal studies present additional challenges, as repeated measurements from the same individual are correlated. Tools that assume independent samples may produce inflated significance for longitudinal data. The metaGEENOME framework addresses this by integrating generalized estimating equations, which account for within-subject correlation [<a href="#ref-2">2</a>].
Professional Escalation Criteria
When to Consult a Bioinformatics Specialist
Researchers should consider consulting a bioinformatics specialist when their study design is complex, when they are uncertain about the appropriate analysis approach, or when their results are unexpected or difficult to interpret. Specific situations that warrant escalation include:
- Studies with longitudinal or repeated-measures designs that require specialized statistical methods [<a href="#ref-2">2</a>]
- Datasets with extreme sparsity or library size variation that challenge standard tools
- Results that differ substantially between analysis tools or that contradict biological expectations
- Studies where the choice of analysis approach could affect regulatory or clinical decisions
When to Seek Statistical Consultation
Statistical consultation is appropriate when researchers need to determine the required sample size for a study, when they are uncertain about the appropriate multiple testing correction, or when they need to interpret complex interactions between variables. Statisticians can also help with the design of validation studies and the interpretation of negative results.
When to Escalate Data Quality Issues
If exploratory analysis reveals severe batch effects, contamination, or other data quality issues, researchers should escalate the issue before proceeding with differential abundance testing. Continuing with a flawed dataset will produce unreliable results regardless of the statistical tool used. The appropriate response may involve re-sequencing samples, removing contaminated samples, or applying more sophisticated batch correction methods.
A Practical Decision Framework for Tool Selection Based on Data Characteristics
Selecting between DESeq2, edgeR, and ANCOM-BC requires a systematic assessment of the dataset before any analysis begins. The following framework translates data characteristics into concrete tool choices, helping researchers avoid the common failure pattern of defaulting to a familiar tool without evaluating whether its assumptions match the data at hand.
Step 1: Assess the Expected Magnitude of Community Shift
The first decision point concerns the anticipated proportion of the microbial community that differs between experimental groups. This assessment should be based on the biological question and prior knowledge from the literature, not on the observed data after analysis.
For studies where the intervention or condition is expected to produce modest changes affecting a minority of taxa, DESeq2 or edgeR are appropriate choices. Their normalization procedures assume that most taxa remain stable between conditions, and this assumption holds when community shifts are limited [<a href="#ref-1">1</a>]. Examples include studies comparing closely related host genotypes, subtle dietary modifications, or mild environmental gradients.
For studies where the intervention is expected to produce substantial community restructuring, ANCOM-BC is the safer choice. The bias correction procedure does not rely on the assumption of a stable reference set of taxa, making it more robust when large fractions of the community change between groups [<a href="#ref-2">2</a>]. Examples include antibiotic treatment studies, major dietary interventions, or comparisons between healthy and diseased states where the microbiome is known to differ substantially.
Researchers uncertain about the expected magnitude of community shift should run an exploratory analysis using beta diversity metrics. If ordination plots show clear separation between groups with high explained variance, the community shift is likely substantial, favoring ANCOM-BC. If separation is modest or overlapping, DESeq2 or edgeR may be adequate.
Step 2: Quantify Sparsity and Zero Inflation
The proportion of zero counts in the dataset directly affects tool performance. Before analysis, calculate the percentage of taxa present in fewer than 10% of samples and the overall proportion of zero entries in the count table.
Datasets with fewer than 30% zero entries and most taxa present in a majority of samples are suitable for all three tools. DESeq2 and edgeR will handle these data with stable dispersion estimates, and ANCOM-BC will produce reliable bias corrections [<a href="#ref-2">2</a>].
Datasets with 30% to 60% zero entries require more careful consideration. DESeq2 and edgeR may produce unstable estimates for rare taxa, while ANCOM-BC requires pseudocount addition that can influence results. In this range, researchers should apply prevalence filtering before analysis, removing taxa present in fewer than 10% to 20% of samples. The filtering threshold should be documented and justified in the analysis protocol.
Datasets with more than 60% zero entries are challenging for all three tools. None of the methods performs optimally under extreme sparsity, and researchers should consider whether the taxonomic level of analysis can be raised to reduce sparsity. Aggregating from species to genus level often substantially reduces the zero proportion while preserving biologically meaningful resolution [<a href="#ref-5">5</a>].
Step 3: Evaluate Library Size Distribution
Examine the distribution of total reads per sample before choosing a tool. Calculate the ratio of the largest to smallest library size and identify any samples that are clear outliers.
When library sizes vary by less than 10-fold, all three tools handle the variation adequately through their respective normalization or bias correction procedures. DESeq2 and edgeR use robust normalization that is not unduly influenced by a few high-depth samples, while ANCOM-BC estimates sampling fractions explicitly [<a href="#ref-2">2</a>].
When library sizes vary by 10-fold to 50-fold, the choice of tool becomes more consequential. ANCOM-BC's explicit sampling fraction estimation may provide more accurate correction than the normalization approaches used by DESeq2 and edgeR, particularly if the compositional structure is complex. Researchers should also consider whether low-depth samples should be excluded or whether the analysis should be restricted to samples meeting a minimum depth threshold.
When library sizes vary by more than 50-fold, no tool can fully compensate for the noise introduced by very low-depth samples. The appropriate response is to examine the low-depth samples for quality issues, consider excluding them, or stratify the analysis by sequencing batch. Continuing with extreme library size variation will produce unreliable results regardless of the statistical tool selected [<a href="#ref-1">1</a>].
Step 4: Consider Study Design Complexity
The experimental design influences tool selection through the statistical models each tool supports. DESeq2 and edgeR offer flexible generalized linear model frameworks that accommodate multifactorial designs, interactions, and continuous covariates. These tools are well suited for studies with multiple experimental factors or complex designs [<a href="#ref-7">7</a>].
ANCOM-BC also supports linear model frameworks but may be less familiar to researchers for complex designs. The tool has been applied successfully in cross-sectional and longitudinal contexts, and frameworks such as metaGEENOME integrate generalized estimating equations for repeated-measures data [<a href="#ref-2">2</a>].
For longitudinal studies with repeated measurements from the same subjects, none of the three tools fully accounts for within-subject correlation. The metaGEENOME framework addresses this limitation by integrating generalized estimating equations, which model the correlation structure of repeated measurements [<a href="#ref-2">2</a>]. Researchers with longitudinal designs should consider whether this approach is more appropriate than the standard implementations of DESeq2, edgeR, or ANCOM-BC.
Step 5: Document the Decision and Pre-Specify the Analysis
The tool selection decision should be made before running the analysis and documented in the study protocol. This pre-specification prevents the common failure pattern of selecting a tool based on which produces the most favorable results, a practice that inflates false discovery rates and undermines reproducibility.
The documentation should include the rationale for tool selection based on the data characteristics assessed in the previous steps. It should also specify the filtering criteria, normalization approach, statistical model, multiple testing correction method, and significance threshold. This documentation should be included in the analysis script and referenced in any publication reporting the results [<a href="#ref-7">7</a>].
Record System for Tool Selection Decisions
Maintaining a structured record of tool selection decisions enables transparency and reproducibility. The following record fields should be completed for each analysis:
| Record Field | Description |
|---|---|
| Study identifier | Unique identifier linking to the study protocol |
| Data version | Version of the count table and processing pipeline |
| Expected community shift | Qualitative assessment based on biological knowledge |
| Zero proportion | Percentage of zero entries in the filtered count table |
| Prevalence filtering threshold | Minimum percentage of samples where a taxon must be present |
| Library size ratio | Ratio of largest to smallest library size |
| Study design type | Cross-sectional, longitudinal, or complex multifactorial |
| Tool selected | DESeq2, edgeR, ANCOM-BC, or combination |
| Selection rationale | Specific data characteristics that drove the decision |
| Analysis date | Date the analysis was run |
| Package versions | Version numbers of all R packages used |
| Random seed | Seed value if stochastic elements were involved |
This record system allows researchers to track how tool selection decisions were made and to revisit those decisions if data quality issues emerge during analysis. It also facilitates comparison across studies within a research program, helping to identify whether tool selection is consistent with data characteristics or driven by habit [<a href="#ref-7">7</a>].
Troubleshooting When Results Conflict Across Tools
When running multiple tools as a robustness check, researchers may encounter situations where tools produce conflicting results. A taxon may be identified as differentially abundant by DESeq2 but not by ANCOM-BC, or the direction of change may differ between tools. These conflicts require systematic troubleshooting instead of arbitrary resolution.
First, examine the taxon's count distribution across samples. A taxon identified by DESeq2 but not ANCOM-BC may have moderate counts with a consistent difference between groups, where the negative binomial model has more power. Alternatively, the taxon may have extreme counts in a few samples that drive the DESeq2 result, while ANCOM-BC's bias correction attenuates this effect [<a href="#ref-2">2</a>].
Second, check whether the conflicting taxon is part of a broader community shift. If many taxa change between groups, DESeq2 and edgeR may produce spurious results due to violated normalization assumptions, while ANCOM-BC correctly identifies only the taxa that change beyond what is expected from the overall shift. In this case, the ANCOM-BC result is more trustworthy [<a href="#ref-2">2</a>].
Third, assess the biological plausibility of the conflicting result. A taxon with a large fold-change but very low prevalence may be a false positive driven by a single outlier sample. Visualizing the count distribution for the taxon across samples can reveal whether the signal is consistent or driven by outliers.
Fourth, consider whether the conflict arises from different handling of zeros. ANCOM-BC requires pseudocount addition, and the choice of pseudocount can affect results for rare taxa. Re-running ANCOM-BC with different pseudocount values can assess the sensitivity of the result to this parameter choice [<a href="#ref-2">2</a>].
When conflicts persist after troubleshooting, the conservative approach is to report the taxon as a candidate finding that requires validation instead of as a confirmed differentially abundant taxon. The validation can involve quantitative PCR, an independent sequencing run, or a different bioinformatics pipeline. This approach acknowledges the limitations of computational tools while providing a path forward for biologically important findings.
Escalation Criteria for Tool Selection Uncertainty
Researchers should escalate to a bioinformatics specialist or statistician when the data characteristics fall into ranges where tool performance is uncertain. Specific escalation triggers include:
- Zero proportions above 60% that persist after taxonomic aggregation
- Library size ratios above 50-fold that cannot be resolved by sample exclusion
- Longitudinal designs where within-subject correlation must be modeled [<a href="#ref-2">2</a>]
- Results that conflict substantially between tools and cannot be resolved through troubleshooting
- Studies where the differential abundance results will inform regulatory decisions or clinical recommendations
In these situations, the cost of an incorrect tool selection is high, and specialized expertise can help identify the most appropriate analytical approach or determine whether additional data collection is needed before analysis can proceed reliably.
Frequently Asked Questions
What is the main difference between DESeq2 and edgeR?
DESeq2 and edgeR both use negative binomial models for count data, but they differ in their normalization approaches. DESeq2 uses median-of-ratios size factors, while edgeR uses trimmed mean of M-values normalization. In practice, the two tools perform similarly in benchmark studies, and the choice between them often comes down to user preference and familiarity [<a href="#ref-1">1</a>].
Why does ANCOM-BC handle compositional data differently?
ANCOM-BC explicitly models the compositional structure of metagenomic data by estimating and correcting for differences in sampling fractions between samples. This bias correction addresses the constraint that counts are relative instead of absolute, which can create spurious signals in tools that do not account for compositionality [<a href="#ref-2">2</a>].
Which tool has better false discovery rate control?
ANCOM-BC has been shown to control the false discovery rate more effectively than DESeq2 or edgeR in metagenomic contexts. DESeq2 and edgeR achieve high sensitivity but may produce more false positives, particularly when large community shifts occur between groups [<a href="#ref-2">2</a>].
Should I run multiple differential abundance tools?
Running multiple tools and comparing results is a reasonable strategy for increasing confidence in findings. Taxa identified by multiple tools with different statistical assumptions are more likely to represent true biological differences. This approach sacrifices some sensitivity but reduces false positives [<a href="#ref-2">2</a>].
How should I handle zero counts in my data?
Zero counts can be handled through filtering, pseudocount addition, or the use of tools that accommodate zero-inflated distributions. The appropriate approach depends on the tool and the extent of sparsity in the data. Researchers should document their zero-handling strategy and consider its impact on results [<a href="#ref-2">2</a>].
What sample size do I need for reliable differential abundance testing?
The required sample size depends on the effect size, the variability within groups, and the desired statistical power. Studies with small sample sizes have limited power to detect small differences and may have poorly controlled false discovery rates. Researchers should consult statistical guidance for their specific study design [<a href="#ref-2">2</a>].
Can I use these tools for shotgun metagenomic data?
Yes, DESeq2, edgeR, and ANCOM-BC can be applied to count tables generated from shotgun metagenomic data. The same considerations regarding compositionality, sparsity, and normalization apply, though the specific characteristics of shotgun data may differ from 16S rRNA amplicon data [<a href="#ref-2">2</a>].
How do I report differential abundance results in my publication?
Publications should report the analysis pipeline, including the version of each tool, the normalization method, the statistical model, the multiple testing correction, and the significance threshold. Raw data should be deposited in public repositories, and analysis scripts should be shared to enable replication [<a href="#ref-4">4</a>][<a href="#ref-7">7</a>].
Related Bioinformatics Guides
- Metagenomics Data Analysis: From Raw Reads to Biological Insights
- Microbiome Data Analysis in R: A Practical Guide for Compositional Data
- Genomic Data Analysis Tools: A Comparative Guide for Researchers
- Metagenomic Binning Tools Benchmark: How to Evaluate and Choose
- Metagenomic Contamination Control: Best Practices for Clean Data
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
[1] [Benchmarking differential expression analysis tools for RNA-Seq: normalization-based vs. log-ratio transformation-based methods.](https://pubmed.ncbi.nlm.nih.gov/30021534). BMC bioinformatics, 2018. [2] [metaGEENOME: an integrated framework for differential abundance analysis of microbiome data in cross-sectional and longitudinal studies.](https://pubmed.ncbi.nlm.nih.gov/40691525). BMC bioinformatics, 2025. [3] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute. [4] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information. [5] [Association between secondhand smoke exposure and ocular microbiome changes in children.](https://doi.org/10.1016/j.crmicr.2026.100630). 2026. [6] [MetaDAVis: An R shiny application for metagenomic data analysis and visualization.](https://doi.org/10.1371/journal.pone.0319949). 2025. [7] [Bioconductor](https://bioconductor.org/). Bioconductor Project. [8] [nf-core Documentation](https://nf-co.re/docs). nf-core. [9] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries. [10] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.