# Normalization Methods for Metagenomic Count Data: A Guide to TMM, RLE, and Rarefaction


## Key Takeaways

- **Normalization is critical for metagenomic count data due to sequencing depth variation and inherent data compositionality**, which violate assumptions of standard statistical methods and can lead to false positives/negatives.
- **TMM (Trimmed Mean of M-values) and RLE (Relative Log Expression) are robust methods for differential abundance testing**, computing sample-specific scaling factors to account for technical variation, with TMM focusing on trimmed M-values and RLE using a synthetic reference sample.
- **Rarefaction subsamples reads to a common depth, primarily suitable for diversity estimation and exploratory analysis**, but it discards significant data, reducing statistical power and introducing randomness.
- **Method selection should be guided by the research question (e.g., differential abundance vs. diversity estimation) and data characteristics (e.g., sparsity, depth variation, batch effects)**, with batch correction recommended as an initial step when batch effects are present.
- **Reproducibility necessitates meticulous documentation of all preprocessing steps, including filtering thresholds, normalization method, software versions, and parameters**, with a normalization decision log serving as an auditable record of rationale and outcomes.

---

Metagenomic count data require normalization before downstream analysis because sequencing depth varies across samples and the compositional nature of the data violates assumptions of many statistical methods. This article compares three widely used normalization approaches, trimmed mean of M-values (TMM), relative log expression (RLE), and rarefaction, and provides a decision framework based on data type and analysis goals. The practical outcome is a reproducible workflow for selecting and applying normalization methods that supports valid biological interpretation.

## The Problem Normalization Solves

Raw metagenomic count tables contain the number of sequencing reads assigned to each taxon or functional feature in each sample. These raw counts are not directly comparable across samples because the total number of reads generated per sample differs due to sequencing machine capacity, library preparation efficiency, and sample DNA concentration. A sample sequenced to 10 million reads will have roughly twice the counts of an identical sample sequenced to 5 million reads, even if the microbial community composition is identical.

The compositional structure of metagenomic data compounds this problem. When one taxon increases in abundance, the relative proportions of all other taxa decrease, even if their absolute abundances remain unchanged. This interdependence means that standard statistical methods designed for independent observations can produce false positives and false negatives when applied to raw or improperly normalized count data.

Normalization methods attempt to remove technical variation while preserving biological signal. The choice of method affects every downstream step, including diversity estimation, differential abundance testing, correlation analysis, and machine learning model performance. A benchmark of 22 normalization methods for predicting quantitative phenotypes found that none demonstrated significant superiority across all scenarios, and the authors recommended batch correction methods as an initial step when batch effects are present [<a href="#ref-1">1</a>]. This finding underscores that normalization is not a one-size-fits-all procedure and that method selection requires consideration of the specific research question.

## Core Principles of Count Normalization

### Sequencing Depth Variation

Sequencing depth refers to the total number of reads generated for a sample. Metagenomic datasets routinely contain samples with 5-fold to 20-fold differences in sequencing depth. Without normalization, samples with higher depth appear to have higher diversity and more abundant taxa, creating artifacts that obscure true biological differences.

Normalization methods address depth variation by scaling counts to a common baseline. The simplest approach divides each count by the total number of reads in the sample, producing relative abundances. However, this approach assumes that the total microbial load is constant across samples, an assumption that fails when samples differ in biomass or when dominant taxa change dramatically between conditions.

### Compositionality

Compositional data are constrained to sum to a constant, typically 1 or 100 percent. This constraint creates negative correlations between features that are purely mathematical artifacts. For example, if one taxon doubles in relative abundance, all other taxa must decrease in relative abundance to maintain the constant sum, regardless of whether their absolute abundances changed.

A review of preprocessing methods in human microbiome research noted that many studies use relative abundance or normalization-based transformations that do not specifically account for the statistical attributes of microbiome data, including compositionality and sparsity [<a href="#ref-2">2</a>]. The authors raised concerns about reproducibility and comparability when preprocessing details are incomplete or missing from publications.

### Sparsity and Overdispersion

Metagenomic count tables are sparse, meaning many taxa are absent or present at very low counts in many samples. This sparsity creates zero-inflated distributions that violate normality assumptions. Overdispersion, where the variance exceeds the mean, is also common because microbial communities are influenced by many interacting factors.

Normalization methods must handle these characteristics without introducing artifacts. Methods that add pseudocounts to zero values, for example, can distort the relative relationships between features, particularly when the pseudocount is large relative to the actual counts.

## At a Glance: Normalization Method Comparison

| Method | Core Approach | Best Suited For | Key Limitation | Implementation |
|--------|---------------|-----------------|----------------|----------------|
| TMM | Computes scaling factors from trimmed mean of M-values between samples | Differential abundance testing, metagenomic count data with variable depth | Assumes most features are not differentially abundant | edgeR package in Bioconductor [<a href="#ref-3">3</a>] |
| RLE | Computes scaling factors relative to a synthetic reference sample using median fold changes | Between-sample comparisons, condition-specific modeling | Sensitive to large numbers of zeros and extreme outliers | DESeq2 package in Bioconductor [<a href="#ref-3">3</a>] |
| Rarefaction | Subsamples reads to a common depth across all samples | Diversity estimation, exploratory analysis, legacy workflows | Discards valid data, reduces statistical power, introduces randomness | phyloseq, vegan, QIIME2 |

## TMM Normalization

### How TMM Works

Trimmed mean of M-values (TMM) computes a scaling factor for each sample relative to a reference sample. The method identifies a set of features that are not differentially abundant between the sample and the reference, then calculates a weighted mean of the log fold changes for those features. The trimming removes features with extreme log fold changes and extreme absolute abundance, reducing the influence of highly variable or highly abundant features.

The resulting scaling factor is applied to the raw counts to produce normalized values. TMM assumes that most features are not differentially abundant between samples, which is a reasonable assumption for many metagenomic datasets where the majority of taxa maintain similar abundances across conditions.

### When to Use TMM

TMM is appropriate when the research question involves comparing relative abundances between groups, such as identifying differentially abundant taxa between healthy and diseased individuals. The method performs well when sequencing depth varies substantially across samples and when the dataset contains a moderate number of features.

A benchmark of RNA-seq normalization methods for genome-scale metabolic modeling found that TMM, along with RLE and GeTMM, enabled the production of condition-specific models with considerably low variability in the number of active reactions compared to within-sample normalization methods [<a href="#ref-4">4</a>]. The same study reported that TMM-normalized data captured disease-associated genes with average accuracy around 0.80 for Alzheimer's disease and 0.67 for lung adenocarcinoma, with improved accuracy after covariate adjustment [<a href="#ref-4">4</a>].

### Practical Implementation

TMM is implemented in the edgeR package, available through Bioconductor. The Bioconductor project provides official documentation for package installation, workflow execution, and reproducible genomic analysis [<a href="#ref-3">3</a>]. Users should consult the edgeR documentation for current function names and parameters, as the package undergoes regular updates.

The standard workflow involves creating a DGEList object from the count table, filtering low-count features, calculating normalization factors with the calcNormFactors function, and then proceeding to downstream analysis. The filtering step is important because features with very low counts contribute noise to the scaling factor calculation.

## RLE Normalization

### How RLE Works

Relative log expression (RLE) normalization, also called median-of-ratios normalization, computes scaling factors based on the ratio of each sample to a synthetic reference. The reference is constructed by taking the geometric mean of counts across all samples for each feature. For each sample, the ratio of observed counts to the reference is calculated, and the median of these ratios becomes the scaling factor.

RLE assumes that most features are not differentially abundant and that the median fold change between a sample and the reference represents technical variation instead of biological signal. The method is robust to outliers because it uses the median instead of the mean.

### When to Use RLE

RLE is well suited for datasets where the majority of features have similar abundances across samples and where the research question involves comparing overall community structure between conditions. The method performs well in differential expression analysis and in modeling applications where condition-specific predictions are needed.

The RNA-seq normalization benchmark found that RLE, TMM, and GeTMM produced condition-specific metabolic models with low variability in active reaction counts, and these methods captured disease-associated genes more accurately than within-sample normalization methods [<a href="#ref-4">4</a>]. The study also reported that covariate adjustment improved accuracy for all methods, highlighting the importance of accounting for factors such as age and gender [<a href="#ref-4">4</a>].

### Practical Implementation

RLE is implemented in the DESeq2 package, available through Bioconductor [<a href="#ref-3">3</a>]. The DESeq2 workflow involves creating a DESeqDataSet object from the count table and experimental design, then running the estimateSizeFactors function to compute scaling factors. The normalized counts are accessed through the counts function with normalized set to TRUE.

Users should be aware that RLE can be sensitive to datasets with a very large proportion of zeros. When most features are absent in most samples, the geometric mean reference becomes zero for many features, and the median ratio calculation can be unstable. Filtering low-count features before normalization helps mitigate this issue.

## Rarefaction

### How Rarefaction Works

Rarefaction subsamples reads from each sample without replacement to a common depth, typically the depth of the sample with the fewest reads. After rarefaction, all samples have the same total number of reads, and the counts represent a random subset of the original data.

The method is conceptually simple and has been widely used in microbial ecology for diversity estimation. Rarefaction curves, which plot the number of observed features against sequencing depth, are used to assess whether sampling depth is sufficient to capture the community diversity.

### When to Use Rarefaction

Rarefaction is appropriate for diversity estimation, particularly alpha diversity metrics such as observed species richness and Shannon diversity. The method ensures that diversity comparisons are not confounded by differences in sequencing depth.

However, rarefaction discards a substantial portion of the data. If the minimum sequencing depth is 10,000 reads and the median depth is 100,000 reads, rarefaction discards 90 percent of the reads from the median sample. This data loss reduces statistical power and can obscure rare taxa that are present at low abundance.

### Practical Implementation

Rarefaction is implemented in several tools, including the phyloseq package in R, the vegan package in R, and QIIME2. The rarefy function in vegan and the rarefy_even_depth function in phyloseq provide standard implementations.

The choice of rarefaction depth requires careful consideration. Setting the depth too low discards excessive data, while setting it too high excludes samples with low sequencing depth. Some workflows use multiple rarefaction depths and compare results across depths to assess robustness.

## Comparing Normalization Methods

### Performance in Differential Abundance Testing

A large-scale comparison of 14 differential abundance testing methods on 38 16S rRNA gene datasets found that methods identified drastically different numbers and sets of significant features, and that results depended on data preprocessing [<a href="#ref-5">5</a>]. The study reported that ALDEx2 and ANCOM-II produced the most consistent results across studies, and the authors recommended a consensus approach based on multiple methods to ensure robust biological interpretations [<a href="#ref-5">5</a>].

This finding has direct implications for normalization choice. Different normalization methods feed into different differential abundance testing frameworks, and the combination of normalization and testing method affects the final results. Researchers should not assume that results are robust to the choice of normalization method.

### Performance in Quantitative Phenotype Prediction

The evaluation of 22 normalization methods for predicting quantitative phenotypes found that none demonstrated significant superiority in reducing root mean squared error of predictions [<a href="#ref-1">1</a>]. The study examined three simulation scenarios and 31 real datasets, covering a range of microbiome data types and prediction targets [<a href="#ref-1">1</a>].

The authors strongly recommended using batch correction methods as the initial step when predicting quantitative phenotypes, particularly when batch effects are present [<a href="#ref-1">1</a>]. This recommendation reflects the observation that batch effects, which are technical variations introduced during sample processing, can dominate biological signal and undermine model reproducibility.

### Performance in Diversity Estimation

A benchmark of 13 analytical approaches using simulated microbial communities found that quantitative approaches incorporating microbial load variation performed significantly better than computational strategies designed to mitigate compositionality and sparsity [<a href="#ref-6">6</a>]. The study reported that quantitative methods improved identification of true positive associations and reduced false positive detection [<a href="#ref-6">6</a>].

For diversity estimation specifically, the benchmark found that methods correcting for sampling depth showed higher precision compared to uncorrected scaling when analyzing simulated scenarios of low microbial load dysbiosis [<a href="#ref-6">6</a>]. The authors advocated for wider adoption of experimental quantitative approaches, while acknowledging that preferred transformations exist for cases where microbial load determination is not feasible [<a href="#ref-6">6</a>].

## Decision Framework for Method Selection

### Step 1: Define the Research Question

The choice of normalization method depends on the downstream analysis goal. Differential abundance testing, diversity estimation, correlation analysis, and machine learning prediction each have different requirements and sensitivities to normalization choices.

For differential abundance testing, TMM and RLE are appropriate starting points, but the testing method should be selected in conjunction with the normalization method. The consensus approach recommended by the 38-dataset comparison suggests using multiple methods and focusing on features identified by several approaches [<a href="#ref-5">5</a>].

For diversity estimation, rarefaction or quantitative methods that account for microbial load are appropriate. The choice depends on whether absolute microbial load data are available.

For quantitative phenotype prediction, batch correction methods should be considered as an initial step, particularly when datasets are combined from multiple sources or processing batches [<a href="#ref-1">1</a>].

### Step 2: Assess Data Characteristics

Examine the distribution of sequencing depths across samples. If depths vary widely, methods that scale by total counts may be inadequate. Examine the proportion of zeros in the count table. High sparsity can destabilize RLE and TMM calculations.

Assess whether the dataset contains extreme outliers in total counts or individual feature counts. Outliers can disproportionately influence scaling factor calculations.

### Step 3: Evaluate Batch Structure

Determine whether samples were processed in multiple batches, sequenced on different runs, or collected at different times. Batch effects can introduce systematic variation that normalization methods do not fully remove.

If batch structure is present, consider batch correction methods before or in addition to normalization. The quantitative phenotype prediction study found that batch correction methods performed satisfactorily in datasets affected by batch effects [<a href="#ref-1">1</a>].

### Step 4: Validate the Choice

Apply the selected normalization method and examine the distribution of normalized values. Check whether the normalization successfully removed depth-related trends by plotting normalized counts against sequencing depth.

Compare results across multiple normalization methods. If conclusions change substantially depending on the normalization method, the findings should be interpreted with caution and reported with appropriate caveats.

## Practical Workflow for Normalization

### Data Preparation

Start with a raw count table where rows represent features and columns represent samples. The count table should contain integer counts of reads assigned to each feature in each sample. Feature identifiers should be consistent with the reference database used for taxonomic or functional assignment.

The NCBI provides official descriptions of sequence databases and analysis services that can be used for feature assignment and annotation [<a href="#ref-7">7</a>]. Researchers should document the database version and assignment parameters used.

### Quality Filtering

Filter low-count features before normalization. Features present in very few samples or with very low total counts contribute noise to normalization calculations and downstream analyses. Common filtering thresholds remove features present in fewer than 10 percent of samples or with fewer than 10 total counts across all samples.

The filtering threshold should be documented and reported. The Galaxy Training Network provides accessible workflow training and analysis tutorials that cover quality filtering and normalization procedures for metagenomic data [<a href="#ref-8">8</a>].

### Normalization Calculation

Apply the selected normalization method using the appropriate software package. For TMM, use edgeR from Bioconductor [<a href="#ref-3">3</a>]. For RLE, use DESeq2 from Bioconductor [<a href="#ref-3">3</a>]. For rarefaction, use phyloseq, vegan, or QIIME2.

Record the normalization factors or rarefaction depth used. These values are essential for reproducibility and should be included in the methods section of any publication.

### Downstream Analysis

Proceed with the intended downstream analysis using the normalized data. For differential abundance testing, use methods appropriate for the data type and normalization approach. For diversity estimation, use the rarefied or normalized data as appropriate.

The EMBL-EBI Training program provides bioinformatics learning pathways and practical analysis education that cover downstream analysis of metagenomic data [<a href="#ref-9">9</a>].

### Documentation and Reporting

Document all preprocessing steps, including filtering thresholds, normalization method, software versions, and parameters. The review of preprocessing methods in human microbiome research found that preprocessing information was incomplete or missing in many publications, leading to reproducibility concerns and comparability issues [<a href="#ref-2">2</a>].

Report normalized data characteristics, including the distribution of normalization factors and the effect of normalization on the count distribution. This information allows readers to assess the appropriateness of the normalization choice.

## Records and Measurements

### Normalization Factors

Record the normalization factor calculated for each sample. TMM and RLE produce a scaling factor per sample that indicates the relative adjustment applied to that sample's counts. Samples with factors substantially above or below 1 may indicate technical issues or extreme biological variation.

Examine the distribution of normalization factors across samples. A wide distribution suggests substantial depth variation or compositional differences between samples. Extreme factors warrant investigation before proceeding with downstream analysis.

### Sequencing Depth Metrics

Record the total number of reads per sample before and after filtering. The ratio of retained reads to total reads indicates the proportion of data that passed quality filtering. Samples with very low retention rates may have quality issues that affect downstream analysis.

For rarefaction, record the rarefaction depth and the number of samples excluded because their depth was below the threshold. Excluding samples reduces statistical power and may introduce bias if excluded samples differ systematically from retained samples.

### Feature Retention

Record the number of features before and after filtering. The proportion of features retained indicates the sparsity of the dataset and the stringency of the filtering criteria. Datasets with very low feature retention may have limited power for detecting differences in rare taxa.

## Common Failure Patterns

### Inappropriate Use of Relative Abundance

Many studies use relative abundance normalization, where each count is divided by the total counts in the sample. This approach does not account for compositionality and can produce spurious correlations between features. The review of preprocessing methods found prevalent usage of relative and normalization-based transformations that do not specifically account for the statistical attributes of microbiome data [<a href="#ref-2">2</a>].

Relative abundance is appropriate for some descriptive analyses but should not be used as the sole normalization for differential abundance testing or correlation analysis.

### Ignoring Batch Effects

Normalization methods address sequencing depth variation but do not fully remove batch effects. When samples are processed in multiple batches, batch-related variation can dominate biological signal. The quantitative phenotype prediction study found that batch correction methods performed satisfactorily in datasets affected by batch effects and recommended their use as an initial step [<a href="#ref-1">1</a>].

Researchers should assess batch structure before selecting normalization methods and consider batch correction when batch effects are present.

### Applying a Single Method Without Validation

The choice of normalization method can substantially affect results. The 38-dataset comparison of differential abundance methods found that results depended on data preprocessing and that methods identified drastically different numbers and sets of significant features [<a href="#ref-5">5</a>].

Researchers should apply multiple normalization methods and compare results. If conclusions are consistent across methods, confidence in the findings increases. If conclusions differ, the reasons for the differences should be investigated.

### Overlooking Compositionality

Compositional data analysis methods, such as centered log-ratio transformation, address the mathematical constraints of relative abundance data. However, these methods have their own limitations and are not always appropriate for metagenomic data.

The benchmark of 13 analytical approaches found that quantitative approaches incorporating microbial load variation performed better than computational strategies designed to mitigate compositionality and sparsity [<a href="#ref-6">6</a>]. This finding suggests that experimental design choices, such as measuring microbial load, can be more effective than computational corrections.

## Limitations of Normalization Methods

### No Universal Best Method

The evaluation of 22 normalization methods for quantitative phenotype prediction found that none demonstrated significant superiority across all scenarios [<a href="#ref-1">1</a>]. This finding indicates that normalization method selection requires consideration of the specific research question, data characteristics, and downstream analysis goals.

Researchers should not assume that a method that performed well in one study will perform well in their dataset. Validation on the specific dataset is essential.

### Normalization Does Not Correct Experimental Design Flaws

Normalization methods cannot compensate for poor experimental design, inadequate sample size, or uncontrolled technical variation. The quantitative phenotype prediction study emphasized that batch correction methods should be used when batch effects are present, but prevention of batch effects through careful experimental design is preferable [<a href="#ref-1">1</a>].

### Computational Methods Have Limits

The benchmark of 13 analytical approaches found that experimental quantitative approaches, such as incorporating microbial load variation, performed better than computational strategies designed to mitigate compositionality and sparsity [<a href="#ref-6">6</a>]. This finding suggests that computational normalization has inherent limits and that experimental design choices can provide more effective solutions.

When microbial load determination is not feasible, the benchmark suggested preferred transformations for specific cases, but these transformations are less effective than quantitative approaches [<a href="#ref-6">6</a>].

## Safety and Reproducibility Context

### Reproducibility Standards

Reproducibility requires complete documentation of all preprocessing steps, including normalization method, parameters, and software versions. The review of preprocessing methods in human microbiome research found that preprocessing information was incomplete or missing in many publications, leading to reproducibility concerns and comparability issues [<a href="#ref-2">2</a>].

The nf-core documentation provides community pipeline standards, usage, configuration, and reproducible workflow context that can support reproducible metagenomic analysis [<a href="#ref-10">10</a>]. The Carpentries lessons provide foundational computing, data, shell, Git, and programming training that supports reproducible research practices [<a href="#ref-11">11</a>].

### Data Availability

Deposit raw count data and normalized data in public repositories to support reproducibility and reanalysis. The NCBI provides official descriptions of sequence databases and analysis services that support data deposition and retrieval [<a href="#ref-7">7</a>].

Include normalization factors and filtering parameters in data deposition to allow other researchers to reproduce the analysis.

### Version Control

Record software versions for all tools used in the analysis pipeline. Version differences can affect normalization results and downstream analyses. The Bioconductor project provides official documentation for package installation and workflow execution, including version information [<a href="#ref-3">3</a>].

## Professional Escalation Criteria

### When to Seek Expert Consultation

Consult a bioinformatics specialist or statistician when the dataset has unusual characteristics, such as extreme sparsity, very wide depth variation, or complex batch structure. These situations require specialized expertise to select appropriate normalization methods and validate results.

Consult a domain expert when the biological interpretation of normalized data is unclear or when normalization results conflict with expected biological patterns.

### When to Reconsider the Analysis Approach

Reconsider the analysis approach when normalization factors are extreme, when results change substantially across normalization methods, or when batch effects dominate biological signal. These situations may require different normalization methods, batch correction, or changes to the experimental design.

Reconsider the analysis approach when the proportion of features retained after filtering is very low, when many samples are excluded due to low sequencing depth, or when the distribution of normalized counts remains highly skewed after normalization.

## Building a Normalization Decision Log for Audit and Reanalysis

A normalization decision log is a structured record that captures the rationale, parameters, and outcomes of every normalization choice made during a metagenomic analysis. This log serves as the bridge between the raw count table and the published results, allowing another researcher or a future version of yourself to reconstruct exactly why a particular method was selected and how the data were transformed. The review of preprocessing methods in human microbiome research found that preprocessing information was incomplete or missing in many publications, leading to reproducibility concerns, comparability issues, and questionable results [<a href="#ref-2">2</a>]. A decision log directly addresses this failure mode by forcing documentation at the time decisions are made instead of relying on retrospective reconstruction.

### Why a Decision Log Matters Beyond Standard Methods Sections

Standard methods sections describe what was done but rarely capture why alternatives were rejected or what exploratory analyses informed the final choice. This omission matters because normalization choices are conditional on data characteristics that vary between datasets. A method that performed well in one study may fail in another due to differences in sparsity, depth distribution, or batch structure. The evaluation of 22 normalization methods for predicting quantitative phenotypes found that none demonstrated significant superiority across three simulation scenarios and 31 real datasets [<a href="#ref-1">1</a>]. This finding implies that the justification for a normalization choice cannot be borrowed from the literature but must be established for each dataset individually.

A decision log records the evidence considered at each decision point. It captures the sequencing depth distribution before filtering, the proportion of zeros in the count table, the batch structure of the samples, and the results of exploratory analyses that compared multiple normalization methods. This record transforms normalization from an opaque step into an auditable process. When results are challenged or when a reanalysis is needed with updated taxonomic databases or improved statistical methods, the decision log provides the starting point for understanding what was done and why.

### Core Components of the Decision Log

The decision log should contain five distinct sections that correspond to the stages of the normalization workflow. Each section records specific measurements and observations that support the decisions made.

#### Sample and Sequencing Metadata

Record the total number of reads generated for each sample before any filtering. This information establishes the baseline sequencing depth distribution that drives many normalization decisions. Include the sequencing platform, the number of samples per sequencing run, and any sample processing batches. The NCBI provides official descriptions of sequence databases and analysis services that can be used to document the origin and processing of sequence data [<a href="#ref-7">7</a>].

Record the date of sequencing, the library preparation protocol, and the kit version used. These details matter because batch effects often correlate with processing dates and reagent lots. The quantitative phenotype prediction study found that batch correction methods performed satisfactorily in datasets affected by batch effects and strongly recommended their use as an initial step when such effects are present [<a href="#ref-1">1</a>]. Without batch metadata recorded at the time of analysis, batch correction cannot be applied or evaluated.

#### Filtering Decisions and Rationale

Record the exact filtering criteria applied to the count table before normalization. This includes the minimum number of samples a feature must be present in, the minimum total count threshold, and any abundance-based filters. Record the number of features retained and removed at each filtering step.

Document why specific thresholds were chosen. A threshold of 10 percent prevalence may be appropriate for a dataset with 100 samples but overly aggressive for a dataset with 20 samples. The rationale should reference the specific characteristics of the dataset instead of a generic default. The Galaxy Training Network provides accessible workflow training and analysis tutorials that cover quality filtering and normalization procedures for metagenomic data [<a href="#ref-8">8</a>], and these resources can inform threshold selection.

#### Normalization Method Selection

Record the primary normalization method selected and the alternative methods that were evaluated. For each alternative, record the software package, version, and parameters used. Record the date the analysis was run because software versions change over time and results may differ between versions.

Document the criteria used to select the final method. These criteria might include the distribution of normalization factors, the stability of results across methods, or the requirements of the downstream analysis. The comparison of 14 differential abundance testing methods on 38 datasets found that results depended on data preprocessing and that methods identified drastically different numbers and sets of significant features [<a href="#ref-5">5</a>]. Recording the comparison results in the decision log provides evidence that the chosen method was selected based on data-specific evaluation instead of habit or convenience.

#### Normalization Outputs and Diagnostics

Record the normalization factors calculated for each sample. For TMM and RLE, these are scaling factors applied to each sample's counts. For rarefaction, record the rarefaction depth and the number of samples excluded because their depth fell below the threshold.

Record diagnostic plots and their interpretation. This includes plots of normalized counts against sequencing depth, the distribution of normalization factors across samples, and the effect of normalization on the count distribution. The benchmark of 13 analytical approaches found that quantitative methods correcting for sampling depth showed higher precision compared to uncorrected scaling when analyzing simulated scenarios of low microbial load dysbiosis [<a href="#ref-6">6</a>]. Diagnostic plots provide the evidence needed to determine whether the normalization successfully removed depth-related trends.

#### Downstream Analysis Linkage

Record how the normalized data were used in downstream analyses. This includes the differential abundance testing method, the diversity metrics calculated, or the machine learning approach applied. Record the software versions for these downstream tools as well.

Document any sensitivity analyses that examined whether conclusions changed when alternative normalization methods were used. The 38-dataset comparison recommended that researchers use a consensus approach based on multiple differential abundance methods to help ensure robust biological interpretations [<a href="#ref-5">5</a>]. A decision log that records the results of multiple normalization methods provides the evidence needed to support a consensus interpretation.

### Implementing the Decision Log in Practice

#### Template Structure

Create the decision log as a structured document with one section per analysis project. A spreadsheet format works well for sample-level metadata and normalization factors, while a narrative format works better for documenting decision rationale. The log should be version controlled so that changes over time are tracked.

The Carpentries lessons provide foundational computing, data, shell, Git, and programming training that supports reproducible research practices [<a href="#ref-11">11</a>]. Using Git to version control the decision log ensures that the record of decisions cannot be accidentally overwritten or lost.

#### Timing of Documentation

Record decisions at the time they are made instead of at the end of the analysis. This practice captures the reasoning while it is fresh and prevents the common failure of reconstructing methods from memory or from incomplete notes. The review of preprocessing methods found that preprocessing information was incomplete or missing in many publications [<a href="#ref-2">2</a>], a problem that arises when documentation is deferred.

Schedule specific checkpoints for documentation. Record sample metadata when the count table is first generated. Record filtering decisions when filters are applied. Record normalization factors when normalization is calculated. Record downstream analysis parameters when those analyses are run.

#### Integration with Analysis Scripts

Embed decision logging into the analysis scripts themselves. Each script should write its parameters and outputs to the decision log automatically. This integration ensures that the log reflects what was actually done instead of what was intended or remembered.

The nf-core documentation provides community pipeline standards, usage, configuration, and reproducible workflow context that can support this integration [<a href="#ref-10">10</a>]. Pipelines that automatically generate log files as part of their execution reduce the burden of manual documentation and improve accuracy.

### Using the Decision Log for Troubleshooting

#### Identifying Unusual Normalization Factors

When normalization factors are extreme, the decision log provides the context needed to determine whether the extreme values reflect technical issues or genuine biological variation. Check the sample metadata recorded in the log for batch membership, processing date, or quality metrics that might explain the extreme factors.

If a sample has a normalization factor substantially above or below the range of other samples, investigate whether the sample had unusual sequencing depth, low quality scores, or contamination. The decision log should contain the information needed to trace the sample back to its original sequencing records.

#### Diagnosing Inconsistent Results Across Methods

When results differ substantially across normalization methods, the decision log provides the evidence needed to understand why. Compare the normalization factors calculated by each method and examine whether specific samples or features drive the differences.

The comparison of differential abundance methods found that for many tools the number of features identified correlated with aspects of the data such as sample size, sequencing depth, and effect size of community differences [<a href="#ref-5">5</a>]. The decision log records these data characteristics, allowing the analyst to determine whether the inconsistency between methods is explained by known data properties.

#### Supporting Reanalysis with Updated Methods

When new normalization methods become available or when updated taxonomic databases are released, the decision log provides the starting point for reanalysis. The log records the original count table, the filtering decisions, and the normalization parameters, allowing the reanalysis to proceed from the same starting point.

The EMBL-EBI Training program provides bioinformatics learning pathways and practical analysis education that cover reanalysis and updated methods [<a href="#ref-9">9</a>]. The decision log ensures that the reanalysis can be compared directly with the original analysis because the preprocessing steps are fully documented.

### Common Failure Patterns in Decision Logging

#### Incomplete Sample Metadata

A decision log that lacks batch information, processing dates, or sequencing run identifiers cannot support batch effect evaluation or correction. The quantitative phenotype prediction study found that batch correction methods performed satisfactorily in datasets affected by batch effects [<a href="#ref-1">1</a>], but this correction requires knowing which samples belong to which batch.

Record batch information at the time samples are received or sequenced. Do not rely on memory or on reconstructing batch structure from sequence data alone.

#### Vague Filtering Rationale

A decision log that records filtering thresholds without explaining why those thresholds were chosen provides limited value for troubleshooting or reanalysis. The rationale should reference the specific characteristics of the dataset, such as the distribution of sequencing depths or the proportion of zeros in the count table.

Record the exploratory analyses that informed the filtering choice. If multiple thresholds were evaluated, record the results of each evaluation and the reason for selecting the final threshold.

#### Missing Diagnostic Interpretations

A decision log that records diagnostic plots without recording their interpretation provides an incomplete record. The interpretation is the analytical content that supports the normalization decision.

Record what the diagnostic plots showed and what conclusion was drawn from them. For example, record that the plot of normalized counts against sequencing depth showed no remaining depth trend after TMM normalization, supporting the use of TMM for downstream analysis.

#### Disconnected Downstream Analysis Records

A decision log that records normalization decisions but not downstream analysis parameters creates a gap in the analytical record. The downstream analysis parameters are needed to reproduce the final results and to understand how normalization choices propagated through the analysis.

Record the downstream analysis method, software version, and parameters in the same decision log as the normalization decisions. This integration ensures that the complete analytical pipeline is documented in one place.

### Records and Measurements for the Decision Log

#### Normalization Factor Distributions

Record the mean, median, and range of normalization factors for each method evaluated. A wide range of factors indicates substantial depth variation or compositional differences between samples. The benchmark of analytical approaches found that quantitative methods correcting for sampling depth showed higher precision compared to uncorrected scaling [<a href="#ref-6">6</a>], and the normalization factor distribution provides evidence about whether depth correction was adequate.

Record the number of samples with normalization factors outside the central range. These samples may require investigation before proceeding with downstream analysis.

#### Feature Retention Statistics

Record the number of features before filtering, after filtering, and after normalization. The proportion of features retained indicates the sparsity of the dataset and the stringency of the filtering criteria.

Record the proportion of counts retained after filtering. This metric indicates how much of the sequencing data contributed to the normalized dataset.

#### Comparison Results Across Methods

Record the results of any comparisons between normalization methods. This includes the number of differentially abundant features identified with each method, the overlap between methods, and any differences in diversity estimates.

The 38-dataset comparison found that methods identified drastically different numbers and sets of significant features [<a href="#ref-5">5</a>]. Recording these differences in the decision log provides evidence about the robustness of the conclusions to normalization choices.

### Professional Escalation Criteria for Decision Logging

#### When to Seek Expert Consultation

Consult a bioinformatics specialist when the decision log reveals patterns that are difficult to interpret. This includes normalization factor distributions that are strongly skewed, filtering decisions that remove a large proportion of features, or inconsistencies between normalization methods that cannot be explained by known data characteristics.

Consult a statistician when the decision log reveals that the choice of normalization method substantially affects the conclusions of the analysis. The evaluation of 22 normalization methods found that none demonstrated significant superiority in predicting quantitative phenotypes [<a href="#ref-1">1</a>], suggesting that method choice can introduce uncertainty that requires statistical expertise to interpret.

#### When to Reconsider the Analysis Approach

Reconsider the analysis approach when the decision log reveals that batch effects were present but not addressed. The quantitative phenotype prediction study strongly recommended using batch correction methods as the initial step when batch effects are present [<a href="#ref-1">1</a>].

Reconsider the analysis approach when the decision log reveals that the normalization method was selected without comparison to alternatives. The comparison of differential abundance methods recommended using a consensus approach based on multiple methods to help ensure robust biological interpretations [<a href="#ref-5">5</a>]. A decision log that records only a single method without comparison provides insufficient evidence for robust conclusions.

Reconsider the analysis approach when the decision log reveals that diagnostic plots showed remaining depth trends after normalization. The benchmark of analytical approaches found that quantitative methods correcting for sampling depth showed higher precision compared to uncorrected scaling [<a href="#ref-6">6</a>]. If normalization did not remove depth trends, the downstream analysis may be confounded by sequencing depth variation.

## Frequently Asked Questions

### What is the difference between TMM and RLE normalization?

TMM computes scaling factors from a trimmed mean of M-values, where M-values are log fold changes between a sample and a reference sample. RLE computes scaling factors from the median of ratios between observed counts and a synthetic reference constructed from geometric means across samples. Both methods assume that most features are not differentially abundant, but they use different statistical approaches to estimate scaling factors. TMM is implemented in edgeR, while RLE is implemented in DESeq2, both available through Bioconductor [<a href="#ref-3">3</a>].

### When should I use rarefaction instead of TMM or RLE?

Rarefaction is appropriate for diversity estimation, particularly alpha diversity metrics such as observed species richness and Shannon diversity. The method ensures that diversity comparisons are not confounded by differences in sequencing depth. However, rarefaction discards valid data and reduces statistical power. TMM and RLE are preferred for differential abundance testing and other analyses that require retaining all data. The benchmark of analytical approaches found that quantitative methods correcting for sampling depth showed higher precision compared to uncorrected scaling for diversity-related analyses [<a href="#ref-6">6</a>].

### Does normalization method choice affect differential abundance testing results?

Yes, normalization method choice substantially affects differential abundance testing results. A comparison of 14 differential abundance testing methods on 38 datasets found that methods identified drastically different numbers and sets of significant features, and that results depended on data preprocessing [<a href="#ref-5">5</a>]. The study recommended using a consensus approach based on multiple differential abundance methods to help ensure robust biological interpretations [<a href="#ref-5">5</a>].

### Can I use the same normalization method for all metagenomic datasets?

No, the choice of normalization method depends on data characteristics and research questions. The evaluation of 22 normalization methods for quantitative phenotype prediction found that none demonstrated significant superiority across all scenarios [<a href="#ref-1">1</a>]. Researchers should assess their data characteristics, including sequencing depth distribution, sparsity, and batch structure, before selecting a normalization method.

### How do I account for batch effects in metagenomic data?

Batch effects should be addressed through experimental design and computational methods. The quantitative phenotype prediction study strongly recommended using batch correction methods as the initial step when predicting quantitative phenotypes, particularly when batch effects are present [<a href="#ref-1">1</a>]. Researchers should document batch structure and consider batch correction methods in addition to normalization.

### What is compositionality and why does it matter for normalization?

Compositionality refers to the constraint that relative abundances sum to a constant, typically 1 or 100 percent. This constraint creates negative correlations between features that are mathematical artifacts instead of biological signals. The review of preprocessing methods noted that many studies use relative abundance transformations that do not specifically account for compositionality [<a href="#ref-2">2</a>]. Normalization methods that address compositionality, such as centered log-ratio transformation, have different properties than TMM, RLE, or rarefaction.

### How do I validate that my normalization method is appropriate?

Validate the normalization method by examining the distribution of normalization factors, plotting normalized counts against sequencing depth, and comparing results across multiple normalization methods. If conclusions change substantially depending on the normalization method, the findings should be interpreted with caution. The benchmark of analytical approaches found that quantitative methods incorporating microbial load variation performed better than computational strategies, suggesting that experimental validation of normalization choices is valuable [<a href="#ref-6">6</a>].

### What information should I report about normalization in my methods section?

Report the normalization method, software package and version, parameters used, filtering thresholds, and the number of features and samples retained after filtering. Record normalization factors for each sample. The review of preprocessing methods found that preprocessing information was incomplete or missing in many publications, leading to reproducibility concerns [<a href="#ref-2">2</a>]. Complete documentation supports reproducibility and allows readers to assess the appropriateness of the normalization choice.

## Related Bioinformatics Guides

- [RNA Sequencing Methods: A Guide to Library Prep, Strandedness, and Sequencing Depth](/knowledge/bioinformatics/rna-sequencing-methods-a-guide-to-library-prep-strandedness-and-sequencing-depth)
- [RNA-Seq Normalization Methods: TPM, RPKM, and Beyond](/knowledge/bioinformatics/rna-seq-normalization-methods-tpm-rpkm-and-beyond)
- [Longitudinal Microbiome Data Analysis: Methods and Best Practices](/knowledge/bioinformatics/longitudinal-microbiome-data-analysis-methods-and-best-practices)
- [Single-Cell Sequencing Methods: A Comparative Overview](/knowledge/bioinformatics/single-cell-sequencing-methods-a-comparative-overview)
- [Metagenomic Binning Tools Benchmark: How to Evaluate and Choose](/knowledge/bioinformatics/metagenomic-binning-tools-benchmark-how-to-evaluate-and-choose)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)

## References and Further Reading

<a id="ref-1"></a>[<a href="#ref-1">1</a>] [Evaluation of normalization methods for predicting quantitative phenotypes in metagenomic data analysis.](https://doi.org/10.3389/fgene.2024.1369628). 2024.

<a id="ref-2"></a>[<a href="#ref-2">2</a>] [Overview of data preprocessing for machine learning applications in human microbiome research.](https://doi.org/10.3389/fmicb.2023.1250909). 2023.

<a id="ref-3"></a>[<a href="#ref-3">3</a>] [Bioconductor](https://bioconductor.org/). Bioconductor Project.

<a id="ref-4"></a>[<a href="#ref-4">4</a>] [A benchmark of RNA-seq data normalization methods for transcriptome mapping on human genome-scale metabolic networks](https://doi.org/10.1038/s41540-024-00448-z). npj Systems Biology and Applications, 2024.

<a id="ref-5"></a>[<a href="#ref-5">5</a>] [Microbiome differential abundance methods produce different results across 38 datasets.](https://doi.org/10.1038/s41467-022-28034-z). 2022.

<a id="ref-6"></a>[<a href="#ref-6">6</a>] [Benchmarking microbiome transformations favors experimental quantitative approaches to address compositionality and sampling depth biases.](https://doi.org/10.1038/s41467-021-23821-6). 2021.

<a id="ref-7"></a>[<a href="#ref-7">7</a>] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.

<a id="ref-8"></a>[<a href="#ref-8">8</a>] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.

<a id="ref-9"></a>[<a href="#ref-9">9</a>] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.

<a id="ref-10"></a>[<a href="#ref-10">10</a>] [nf-core Documentation](https://nf-co.re/docs). nf-core.

<a id="ref-11"></a>[<a href="#ref-11">11</a>] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.