TMM vs. DESeq2 Median-of-Ratios: A Side-by-Side Comparison of RNA-seq Normalization Methods
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- TMM (Trimmed Mean of M-values) normalizes by calculating a weighted mean of log fold changes relative to a reference sample, trimming extreme values to mitigate the influence of highly expressed or differentially expressed genes. This method is particularly robust to composition effects where a few genes dominate library proportions.
- Median-of-Ratios (DESeq2) normalizes by calculating the median of ratios of each gene's count to its geometric mean across all samples, using the median as a robust estimator against outliers without explicit trimming parameters. This method assumes a majority of genes are not differentially expressed.
- Both methods are designed for between-sample normalization, correcting for library size and composition differences to enable valid comparisons of gene expression across samples, unlike within-sample normalization methods such as TPM.
- The choice of normalization method can impact downstream differential expression results; TMM's explicit trimming offers more control for datasets with strong composition effects, while Median-of-Ratios is simpler and robust for moderate differences.
- Practical implementation involves quality control of raw counts, gene filtering, applying the chosen method (TMM via
edgeR::calcNormFactorsor Median-of-Ratios viaDESeq2::estimateSizeFactors), and verifying normalization through diagnostic plots before proceeding to differential expression testing. - Reproducibility is enhanced by recording raw library sizes, calculated scaling factors, normalization parameters, and software versions, and by utilizing standardized community pipelines for consistent workflow execution.
RNA-seq differential expression analysis requires a normalization step before biological comparisons can be made. The two most widely used between-sample normalization methods are Trimmed Mean of M-values (TMM) from edgeR and the median-of-ratios method (often called RLE) from DESeq2. Both methods scale raw count data to account for differences in sequencing depth and library composition, but they use different statistical approaches to calculate scaling factors. This article compares TMM and DESeq2 median-of-ratios normalization across data inputs, workflow choices, quality checks, reproducibility, and interpretation limits, with practical guidance on when each method is preferable.
The choice between TMM and median-of-ratios normalization can change which genes appear differentially expressed in your final results. Normalization approaches rescale the data, and these rescaling decisions affect downstream analysis outcomes. For analysts working with RNA-seq count data, understanding the assumptions behind each method helps prevent false discoveries and improves reproducibility across datasets.
At a Glance
The table below summarizes the key differences between TMM and DESeq2 median-of-ratios normalization for practical decision-making.
| Feature | TMM (edgeR) | Median-of-Ratios (DESeq2) |
|---|---|---|
| Core calculation | Trimmed mean of log fold changes relative to a reference sample | Median of ratios of each gene count to its geometric mean across samples |
| Reference basis | One sample selected as reference | Geometric mean across all samples |
| Trimming approach | Removes extreme log fold changes and absolute expression values | No explicit trimming, uses median as robust estimator |
| Typical implementation | edgeR package in Bioconductor | DESeq2 package in Bioconductor |
| Best suited for | Datasets with variable library sizes and composition effects | Datasets with moderate composition differences and balanced designs |
| Known limitation | Can be influenced by highly expressed genes if trimming parameters are not appropriate | Assumes most genes are not differentially expressed |
| Output scale | Normalized counts suitable for edgeR downstream tests | Median-of-ratios normalized counts with variance stabilization options |
Both methods belong to the between-sample normalization category, which is used when samples in a dataset are compared with each other. This differs from within-sample normalization methods such as TPM, which are preferred when genes within a single sample are compared. The distinction matters because the analytical goal determines which normalization approach is appropriate.
Understanding RNA-seq Count Data and the Normalization Problem
RNA-seq quantifies the abundance of transcripts within a biological sample and performs differential analysis between different conditions to reveal regulated gene signatures. The raw output from sequencing instruments is a matrix of integer counts, where each value represents the number of sequencing reads that mapped to a particular gene in a particular sample.
These raw counts are not directly comparable across samples for several reasons. First, the total number of reads generated per sample, called the library size, varies between samples due to sequencing depth differences. A sample with more total reads will have higher counts for all genes simply because more sequencing was performed. Second, the composition of the transcriptome can differ between samples. If a small number of genes are highly expressed in one condition, they consume a larger fraction of the sequencing reads, which reduces the counts for all other genes in that sample.
Normalization methods attempt to correct for these technical artifacts so that observed differences in gene expression reflect biological changes instead of sequencing artifacts. The constrained nature of count data means that the data are limited by the sequencing depth of the assay, and normalization is typically required before statistical analysis.
The normalization problem is not unique to RNA-seq. Similar challenges arise in microbiome studies where sequencing depth and sparsity can strongly affect downstream analyses. In real datasets, the underlying biological signal is unknown, which makes it difficult to determine whether a normalization method preserves true group differences or introduces distortions. Simulation-based evaluation frameworks that generate realistic datasets with known ground truth enable quantitative comparison of normalization methods.
Core Principles of TMM Normalization
TMM normalization, implemented in the edgeR package, calculates scaling factors based on the weighted mean of log fold changes between a sample and a reference sample. The method assumes that most genes are not differentially expressed between samples, so the majority of log fold changes should center around zero.
The TMM calculation proceeds through several steps. First, a reference sample is selected from the dataset. The choice of reference can affect results, though the method is designed to be relatively robust to this choice. Second, for each gene, the log fold change between the sample and the reference is calculated, along with the absolute expression level. Third, the method trims the most extreme log fold changes and the highest absolute expression values. This trimming removes genes that are likely to be truly differentially expressed or that have very high counts, which could otherwise dominate the scaling factor calculation. Fourth, the remaining genes are used to calculate a weighted mean of the log fold changes, which becomes the scaling factor for that sample.
The trimming parameters are important for TMM performance. The default settings trim 30 percent of log fold changes and 5 percent of absolute expression values, but these parameters can be adjusted. When a dataset contains a small number of very highly expressed genes that differ dramatically between conditions, the default trimming may not be sufficient to prevent these genes from influencing the scaling factors.
TMM normalization is designed to be robust to composition effects, where a few highly expressed genes differ between conditions. By trimming extreme values, the method reduces the influence of these genes on the scaling factor calculation. This property makes TMM particularly useful for datasets where strong biological differences exist between conditions.
Core Principles of DESeq2 Median-of-Ratios Normalization
The median-of-ratios method, implemented in DESeq2, takes a different approach to calculating scaling factors. For each gene, the method first calculates the geometric mean of the counts across all samples. Then, for each sample and each gene, the ratio of the gene count to the geometric mean is calculated. The scaling factor for a sample is the median of these ratios across all genes.
The geometric mean serves as a pseudo-reference sample that represents the typical expression level for each gene across the dataset. Genes with zero counts in any sample are typically excluded from the calculation because the geometric mean would be zero, making the ratio undefined.
The median is used as a robust estimator of the typical ratio. Because the median is resistant to outliers, genes that are strongly differentially expressed between conditions do not unduly influence the scaling factor, provided that they are in the minority. The method assumes that most genes are not differentially expressed, so the median ratio should reflect the true library size difference between samples.
The median-of-ratios method does not use explicit trimming parameters. Instead, the median provides inherent robustness to outliers. This simplicity is an advantage in practice because there are fewer parameters to tune, but it also means that the method has less flexibility to adapt to datasets where a large proportion of genes are differentially expressed.
Both TMM and median-of-ratios methods have been shown to perform similarly in many benchmark studies. The choice between them often depends on the specific characteristics of the dataset and the downstream analysis goals.
Comparing Method Performance in Benchmark Studies
Multiple benchmark studies have compared TMM and median-of-ratios normalization across different datasets and analysis scenarios. The results consistently show that both methods perform well in typical RNA-seq analysis workflows, but specific conditions can favor one method over the other.
A benchmark study of differential expression analysis tools compared normalization-based methods with log-ratio transformation-based methods. The study found that conventional RNA-seq methods such as edgeR and DESeq2 identify differentially expressed genes with high precision and, given sufficient sample sizes, high recall. The choice of normalization method affected performance, but both TMM and median-of-ratios approaches produced reliable results in most scenarios.
A study comparing normalization methods for Alzheimer's disease RNA-seq datasets found that covariate-adjusted TMM and covariate-adjusted DESeq2 methods performed better than other approaches in both transcriptome datasets examined. The study used two commonly used Alzheimer's disease RNA-seq datasets and mapped differentially expressed genes onto the human protein interactome to discover disease-specific subnetworks. Capturing known Alzheimer's disease genes and genes associated with disease-related functional terms served as the criteria for comparing normalization methods. The results showed that applying covariate adjustment had a positive effect on normalization by removing confounder effects.
A simulation-based benchmark of microbiome normalization methods found that model-based normalization-factor methods, particularly edgeR-TMM and in some settings DESeq2, gave the closest match to the simulated biological contrast. These methods showed better recovery of taxa-level differences while preserving sample-level separation. The study also found that sequencing-depth differences alone could create false-positive group differences for several methods, highlighting the importance of choosing an appropriate normalization approach.
A comparison of library size normalization and statistical methods for balanced two-group comparisons found that RLE, TMM, and upper quartile normalization methods performed similarly given a desired sample size. The study evaluated normalization combined with different statistical tests, including the Wald test from DESeq2 and exact test or quasi-likelihood F-test from edgeR. The results showed that the choice of statistical test mattered more than the choice among these normalization methods in many scenarios.
Mathematical Similarities and Differences
The TMM and median-of-ratios methods are mathematically related, and under certain conditions they produce identical results. A study published in Frontiers in Genetics proved properties showing when TMM, RLE, and MRN methods give exactly the same results. These properties were demonstrated mathematically and illustrated with in silico calculations on a given RNA-seq dataset.
The key insight is that both methods estimate a scaling factor that represents the effective library size for each sample. The scaling factors are used to adjust the raw counts so that samples can be compared on an equal footing. The difference lies in how the scaling factors are estimated.
TMM estimates the scaling factor using a weighted mean of log fold changes after trimming extreme values. The median-of-ratios method estimates the scaling factor using the median of ratios to a geometric mean reference. Both approaches are designed to be robust to outliers, but they achieve this robustness through different mechanisms.
The mathematical equivalence between the methods depends on the distribution of log fold changes and the trimming parameters used. When the data meet certain conditions, the trimmed mean and the median produce the same estimate. In practice, these conditions are rarely met exactly, so the methods produce slightly different scaling factors for real datasets.
Understanding these mathematical relationships helps analysts interpret why the methods sometimes produce different results. When the methods disagree, the difference typically arises from genes with extreme expression changes or from datasets where a large fraction of genes are differentially expressed.
Data Inputs and Preprocessing Requirements
Both TMM and median-of-ratios normalization require raw count data as input. The counts should represent the number of sequencing reads mapped to each gene or transcript. The choice of alignment and quantification procedure can affect the count matrix, but both normalization methods are designed to work with standard count data from common quantification tools.
Before normalization, analysts should perform quality control on the raw count data. This includes checking for samples with very low total read counts, which may indicate sequencing failures or sample quality issues. Samples with unusual library sizes or composition profiles should be examined carefully, as they can influence normalization scaling factors.
The quality of the input data affects the reliability of normalization. Spatial transcriptomic analysis has shown that different tissue preservation methods can affect gene expression profiles. A study comparing matched frozen and formalin-fixed paraffin-embedded colorectal cancer tissues found that FFPE tissue samples offered improved resolution of cellular morphology, while fresh frozen and snap frozen tissues showed high levels of cancer gene expression. These differences in sample quality can affect downstream normalization and analysis.
For standard RNA-seq analysis, the count matrix should include genes with sufficient expression levels. Genes with very low counts across all samples are often filtered before normalization to reduce noise and improve the stability of scaling factor estimates. The filtering threshold depends on the dataset and the analysis goals.
Workflow Integration with edgeR and DESeq2
TMM normalization is integrated into the edgeR package, while median-of-ratios normalization is integrated into the DESeq2 package. Both packages are available through Bioconductor, which provides official package, workflow, installation, and reproducible genomic-analysis documentation.
The edgeR workflow typically proceeds through several steps. First, the count matrix and sample metadata are loaded into a DGEList object. Second, quality control and filtering are performed. Third, TMM normalization is applied to calculate scaling factors. Fourth, the normalized data are used for differential expression testing using edgeR's statistical methods, which include the exact test and the quasi-likelihood F-test.
The DESeq2 workflow follows a similar structure. First, the count matrix and sample metadata are loaded into a DESeqDataSet object. Second, the median-of-ratios normalization is applied automatically when the DESeq function is called. Third, the normalized data are used for differential expression testing using DESeq2's Wald test or likelihood ratio test.
Both packages provide functions for extracting normalized counts for visualization and downstream analysis. The normalized counts can be used for principal component analysis, clustering, and other exploratory analyses. However, the normalized counts from different methods are not directly comparable because they use different scaling approaches.
The choice of statistical test is coupled with the choice of normalization method in practice. edgeR's statistical methods are designed to work with TMM-normalized data, while DESeq2's statistical methods are designed to work with median-of-ratios-normalized data. Mixing normalization methods across packages is possible but requires careful consideration of the statistical assumptions.
Sample Size Considerations
The number of biological replicates per condition affects the performance of both normalization methods and the statistical tests used for differential expression analysis. Benchmark studies have shown that the choice of statistical test matters more when sample sizes are small.
A study comparing normalization and statistical methods for balanced two-group comparisons found that a Wald test performs better than an exact test when the number of sample replicates is large, and that a quasi-likelihood F-test performs best given sample sizes of 5, 10, and 15 for any normalization method. The study used the MAQC RNA-seq datasets with small sample replicates and found that a recently developed normalization method combined with an exact test could achieve better performance in terms of power and specificity in some scenarios.
For experiments with very small sample sizes, such as two conditions without replicates, the choice of normalization method becomes particularly important. A study comparing TMM, RLE, and MRN normalization methods for a simple two-conditions-without-replicates experimental design highlighted the similarities between the methods and proved conditions under which they give exactly the same results. This scenario is challenging because there is no within-condition variability to inform the statistical model.
In general, larger sample sizes improve the reliability of normalization scaling factors and the power of differential expression tests. The median-of-ratios method may be more stable with small sample sizes because the median is a robust estimator, but both methods can produce unreliable results when sample sizes are very small.
Highly Expressed Genes and Composition Effects
The presence of highly expressed genes that differ between conditions is a common challenge in RNA-seq analysis. These genes consume a large fraction of sequencing reads, which reduces the counts for all other genes in the sample. This composition effect can create apparent differences in expression for genes that are not actually changing.
TMM normalization is specifically designed to handle composition effects through its trimming approach. By removing the most extreme log fold changes and highest expression values, TMM reduces the influence of highly expressed genes on the scaling factor calculation. The default trimming parameters remove 30 percent of log fold changes and 5 percent of absolute expression values, but these can be adjusted if the dataset contains an unusual number of highly expressed genes.
The median-of-ratios method also provides robustness to composition effects through the use of the median. Because the median is resistant to outliers, a small number of highly expressed genes with extreme ratios will not shift the median substantially. However, if a large fraction of genes are differentially expressed, the median may be biased.
In practice, the choice between TMM and median-of-ratios for datasets with strong composition effects depends on the proportion of genes that are affected. TMM's explicit trimming parameters provide more control in extreme cases, while the median-of-ratios method is simpler and requires no parameter tuning.
Covariate Adjustment and Confounder Control
Real-world RNA-seq datasets often include covariates such as age, sex, post-mortem interval, or batch effects that can confound the comparison between conditions. Normalization methods do not automatically account for these covariates, so additional adjustment is often needed.
A study comparing normalization methods for Alzheimer's disease datasets found that applying covariate adjustment had a positive effect on normalization by removing confounder effects. The study adjusted for gender, age of death, and post-mortem interval after normalization and found that covariate-adjusted TMM and covariate-adjusted DESeq2 methods performed better in both transcriptome datasets examined.
The workflow for covariate adjustment typically involves normalizing the count data first, then including covariates in the statistical model for differential expression analysis. Both edgeR and DESeq2 support the inclusion of covariates in their statistical models. The choice of normalization method interacts with the covariate adjustment strategy, so analysts should consider both decisions together.
For datasets with known batch effects, more sophisticated approaches such as surrogate variable analysis or ComBat may be needed in addition to normalization. These methods identify and remove hidden sources of variation that are not captured by the measured covariates.
Practical Implementation Steps
The following steps outline a practical workflow for choosing and applying TMM or median-of-ratios normalization for RNA-seq differential expression analysis.
First, assess the experimental design and data characteristics. Determine the number of biological replicates per condition, the expected magnitude of biological differences, and whether composition effects are likely to be present. This assessment informs the choice of normalization method.
Second, perform quality control on the raw count data. Check for samples with very low total read counts, unusual gene expression distributions, or evidence of sample contamination. Remove or re-sequence problematic samples before normalization.
Third, filter genes with very low expression levels. Genes with zero counts in most samples contribute noise to the normalization calculation and are unlikely to be biologically meaningful. Common filtering approaches include keeping genes with a minimum number of counts in a minimum number of samples.
Fourth, apply the chosen normalization method. For TMM, use the edgeR package and the calcNormFactors function. For median-of-ratios, use the DESeq2 package and the estimateSizeFactors function. Both functions are well-documented in their respective Bioconductor packages.
Fifth, verify the normalization results. Check that the scaling factors are reasonable and that the normalized data show expected patterns in principal component analysis or clustering. Samples that cluster by condition instead of by batch or library size indicate successful normalization.
Sixth, proceed with differential expression analysis using the statistical methods appropriate for the chosen normalization method. For TMM-normalized data, use edgeR's exact test or quasi-likelihood F-test. For median-of-ratios-normalized data, use DESeq2's Wald test or likelihood ratio test.
Records and Measurements for Normalization Assessment
Keeping detailed records of the normalization process supports reproducibility and helps identify problems when results are unexpected. The following measurements should be recorded for each analysis.
Record the raw library sizes for each sample before normalization. These values indicate the total sequencing depth and can be compared across samples to identify outliers. Samples with library sizes that differ substantially from the median may require special attention.
Record the normalization scaling factors calculated by the chosen method. The scaling factors indicate the relative adjustment applied to each sample. Scaling factors that vary widely across samples suggest substantial differences in library composition or sequencing depth.
Record the number of genes included in the normalization calculation after filtering. This number affects the stability of the scaling factor estimates. Very low gene counts after filtering may indicate that the filtering threshold was too aggressive.
Record the normalization parameters used, including any trimming parameters for TMM or filtering thresholds. These parameters affect the results and should be reported in publications for reproducibility.
Record the version numbers of the software packages used, including edgeR, DESeq2, and the R version. Software updates can change normalization behavior, so version information is essential for reproducing results.
Common Failure Patterns and Troubleshooting
Several common problems can arise during normalization that lead to poor differential expression results. Recognizing these patterns helps analysts diagnose and correct issues.
One common failure pattern is the presence of a single sample with an unusual library composition. This sample may have a very different distribution of gene expression compared to the other samples, which can skew the normalization scaling factors. Diagnostic plots such as principal component analysis or hierarchical clustering can identify such samples. If a sample is clearly an outlier, it may need to be removed from the analysis.
Another failure pattern occurs when a small number of genes dominate the sequencing reads in one condition but not the other. This composition effect can create false positives in differential expression analysis if the normalization method does not adequately account for it. TMM's trimming parameters can be adjusted to be more aggressive, or the median-of-ratios method may be more appropriate if the effect is moderate.
A third failure pattern involves genes with zero counts in some samples. Both normalization methods handle zero counts differently, and genes with many zero counts can be problematic. Filtering these genes before normalization is often necessary to obtain stable scaling factors.
A fourth failure pattern is the use of inappropriate statistical tests for the sample size. The exact test in edgeR may have low power with small sample sizes, while the Wald test in DESeq2 may have inflated false positive rates with very small sample sizes. Choosing the appropriate statistical test for the sample size is important for reliable results.
Interpretation Limits and Reporting Considerations
Normalized count data have inherent limitations that should be considered when interpreting differential expression results. The normalized values are relative measures, not absolute transcript abundances. They indicate how expression levels compare between samples, but they do not provide information about the absolute number of transcripts in a cell or tissue.
The choice of normalization method can affect which genes are identified as differentially expressed. Studies have shown that different analytical packages can report different expression patterns and false discovery rates and P-values. This variability across methods means that results should be interpreted with appropriate caution, particularly for genes with small fold changes or borderline significance.
For publication, analysts should report the normalization method used, the software versions, and the parameters applied. This information allows other researchers to reproduce the analysis and compare results across studies. The lack of standardized analytical methods leads to uncertainties in data interpretation and study reproducibility, especially with studies reporting high false discovery rates.
When reporting results, it is important to distinguish between genes with strong, consistent expression changes and those with marginal changes that may be sensitive to the normalization choice. Genes that are identified as differentially expressed by multiple normalization methods are more likely to represent true biological differences.
Quality Control and Reproducibility
Reproducibility in RNA-seq analysis requires careful attention to quality control at every step, including normalization. The Galaxy Training Network provides accessible workflow training and analysis tutorials that emphasize reproducibility in bioinformatics analysis. Similarly, nf-core documentation describes community pipeline standards for reproducible workflow configuration and usage.
Quality control for normalization should include visual inspection of the data before and after normalization. Before normalization, check the distribution of raw counts across samples and identify any samples with unusual patterns. After normalization, verify that the scaling factors are reasonable and that the normalized data show expected biological patterns.
The Carpentries lessons provide foundational computing and data skills that support reproducible analysis practices. These skills include version control, documentation, and structured project organization, which are essential for maintaining reproducible RNA-seq analysis workflows.
For large-scale or collaborative projects, using standardized pipelines can improve reproducibility. Community pipelines such as those documented by nf-core provide consistent implementations of RNA-seq analysis workflows, including normalization steps. These pipelines reduce the risk of errors from manual analysis and make it easier to share and reproduce results.
Choosing Between TMM and Median-of-Ratios
The decision between TMM and median-of-ratios normalization depends on several factors related to the dataset and analysis goals. No single method is universally best, and the choice should be informed by the specific characteristics of the data.
For datasets with strong composition effects, where a small number of genes are highly expressed and differ between conditions, TMM's explicit trimming parameters provide more control. The trimming can be adjusted to be more aggressive if the default settings are insufficient to prevent highly expressed genes from dominating the scaling factor calculation.
For datasets with moderate composition differences and balanced designs, the median-of-ratios method is a solid choice. Its simplicity and lack of tuning parameters make it easy to apply consistently, and its performance is comparable to TMM in most benchmark studies.
For datasets with small sample sizes, both methods can be unstable, but the median-of-ratios method may be more robust because the median is resistant to outliers. However, the choice of statistical test is often more important than the choice of normalization method for small sample sizes.
For datasets with known covariates or batch effects, the normalization method should be chosen in conjunction with the covariate adjustment strategy. Both TMM and median-of-ratios can be combined with covariate adjustment, but the specific implementation differs between edgeR and DESeq2.
For datasets where highly expressed genes are a concern, analysts should examine the data to determine the extent of the problem. If a small number of genes account for a large fraction of the sequencing reads, TMM with adjusted trimming parameters may be preferable. If the composition effect is moderate, either method should perform adequately.
Professional Escalation Criteria
Some analysis scenarios warrant consultation with a bioinformatics specialist or biostatistician. Recognizing these situations helps prevent incorrect conclusions from being drawn from the data.
Escalate to a specialist when the normalization scaling factors vary widely across samples and the cause is not apparent. This pattern may indicate a technical problem with the sequencing data or a biological effect that requires careful investigation.
Escalate when the choice of normalization method substantially changes the list of differentially expressed genes. If the results are highly sensitive to the normalization approach, the biological conclusions may not be robust, and additional analysis is needed to understand the source of the discrepancy.
Escalate when the dataset has a complex design, such as multiple factors, interactions, or repeated measures. These designs require careful statistical modeling that goes beyond standard normalization and differential expression workflows.
Escalate when the data come from a non-standard assay, such as spatial transcriptomics or single-cell RNA-seq. These assays have unique characteristics that require specialized normalization and analysis approaches.
Escalate when the results will be used for clinical or regulatory decisions. The stakes are higher for these applications, and the analysis must meet rigorous standards for reproducibility and validation.
A Practical Decision Framework for Selecting TMM or Median-of-Ratios Normalization
Choosing between TMM and median-of-ratios normalization requires a structured evaluation of your specific dataset instead of relying on habit or default settings. The following decision framework translates the statistical properties of each method into concrete assessment steps that can be applied before committing to an analysis path.
Step 1: Profile Library Composition Before Normalization
Begin by examining the distribution of raw counts across all samples. Calculate the total library size for each sample and identify the proportion of reads contributed by the top 50 most highly expressed genes. This profile reveals whether composition effects are likely to be a problem.
Record the following measurements for each sample:
- Total mapped read count
- Percentage of reads from the top 50 genes
- Number of genes with zero counts
- Ratio of the largest to smallest library size
If the top 50 genes account for more than 30 percent of reads in any sample and these genes differ substantially between conditions, composition effects are present. TMM with adjusted trimming parameters may be preferable because the method explicitly removes extreme log fold changes and high expression values from the scaling factor calculation. The default trimming removes 30 percent of log fold changes and 5 percent of absolute expression values, but these parameters can be increased when highly expressed genes dominate the library.
If the top genes account for a smaller proportion of reads and library sizes are relatively consistent, the median-of-ratios method provides adequate robustness through its use of the median as a resistant estimator.
Step 2: Assess the Proportion of Differentially Expressed Genes
Both methods assume that most genes are not differentially expressed between conditions. When this assumption is violated, the scaling factors become biased and downstream results are unreliable.
Estimate the likely proportion of differentially expressed genes using prior knowledge of the biological system or by running a preliminary analysis with both methods. If more than 30 percent of genes are expected to change between conditions, neither method performs optimally. In this scenario, consult a bioinformatics specialist before proceeding, because the normalization step itself may introduce systematic bias that no parameter adjustment can fully correct.
For datasets where the proportion of changing genes is moderate, the median-of-ratios method may be more stable because the median is inherently resistant to a minority of extreme ratios. TMM can also perform well, but the trimming parameters should be verified to ensure they exclude the changing genes from the scaling factor calculation.
Step 3: Evaluate Sample Size and Replication Structure
The number of biological replicates per condition influences both normalization stability and the choice of statistical test. For experiments with fewer than five replicates per condition, record the exact replication structure and consider how it affects method performance.
Benchmark evidence shows that the quasi-likelihood F-test performs best given sample sizes of 5, 10, and 15 for any normalization method. The Wald test performs better than the exact test when the number of sample replicates is large. These findings indicate that the statistical test choice often matters more than the normalization method for typical sample sizes.
For experiments with very small sample sizes, such as two conditions without replicates, both normalization methods can produce unstable scaling factors. The median-of-ratios method may be more robust because the median is resistant to outliers, but results from such experiments should be interpreted with caution regardless of the normalization approach.
Step 4: Check for Covariates and Batch Structure
Record all known covariates such as age, sex, post-mortem interval, or processing batch before choosing a normalization method. Covariate adjustment after normalization has been shown to improve performance for both TMM and DESeq2 methods in transcriptome datasets.
The workflow for covariate adjustment differs between packages. edgeR supports the inclusion of covariates in its statistical model through the design matrix. DESeq2 also supports covariates through its formula interface. The normalization method should be chosen in conjunction with the covariate adjustment strategy, because the interaction between these decisions affects the final results.
If batch effects are suspected but not measured, additional approaches such as surrogate variable analysis may be needed. These methods identify hidden sources of variation that are not captured by measured covariates and should be applied after normalization.
Step 5: Run Both Methods and Compare Scaling Factors
A practical validation step is to run both normalization methods on the same dataset and compare the resulting scaling factors. This comparison provides direct evidence about whether the choice of method matters for your specific data.
Calculate the correlation between TMM scaling factors and median-of-ratios scaling factors across samples. If the correlation is high, the choice of method is unlikely to substantially affect downstream results. If the correlation is low, investigate the source of the discrepancy by examining which genes contribute most to the difference.
Also compare the lists of differentially expressed genes produced by each method. Genes that are identified by both methods are more likely to represent true biological differences. Genes identified by only one method should be examined carefully, particularly if they have small fold changes or borderline significance.
Step 6: Document the Decision and Rationale
Record the normalization method chosen, the parameters used, and the rationale for the decision. This documentation supports reproducibility and helps other researchers understand why a particular approach was selected.
Include the following information in the analysis record:
- Raw library sizes for all samples
- Scaling factors calculated by the chosen method
- Trimming parameters used for TMM or filtering thresholds for median-of-ratios
- Software package versions including edgeR, DESeq2, and R version
- Results of the comparison between methods if both were run
This record allows the analysis to be reproduced exactly and provides a basis for troubleshooting if results are questioned during review or replication attempts.
Common Failure Patterns in Method Selection
Several recurring problems emerge when analysts choose a normalization method without adequate assessment. Recognizing these patterns helps avoid costly reanalysis.
The first pattern is using TMM with default trimming parameters when a small number of genes dominate the library composition. The default trimming may not exclude these genes from the scaling factor calculation, leading to biased normalization. Adjust the trimming parameters or switch to the median-of-ratios method if the composition effect is moderate.
The second pattern is applying the median-of-ratios method when a large fraction of genes are differentially expressed. The median becomes biased when more than half of the genes change between conditions, producing scaling factors that do not reflect true library size differences. Neither method handles this scenario well, and specialist consultation is warranted.
The third pattern is ignoring covariate structure when it is known to exist. Normalization does not automatically account for covariates, and failing to include them in the statistical model can produce false positives. Both edgeR and DESeq2 support covariate adjustment, and this step should not be skipped when covariates are measured.
The fourth pattern is choosing a normalization method based on the statistical test instead of the data characteristics. While TMM is integrated with edgeR and median-of-ratios is integrated with DESeq2, the normalization choice should be driven by the dataset properties. The statistical test can be selected independently in many cases, though the integrated workflows are generally recommended for simplicity and reliability.
When to Escalate to Specialist Consultation
Some datasets require expert input beyond standard normalization workflows. Escalate to a bioinformatics specialist or biostatistician when the scaling factors from TMM and median-of-ratios methods disagree substantially, when more than 30 percent of genes are expected to be differentially expressed, when the experimental design includes complex factors or interactions, or when the data come from non-standard assays such as spatial transcriptomics or single-cell RNA-seq.
The decision framework described here provides a systematic approach to method selection, but it cannot replace expert judgment for unusual datasets. The cost of incorrect normalization is inflated false discovery rates and unreliable biological conclusions, which are far more expensive than a consultation with a specialist.
Frequently Asked Questions
What is the main difference between TMM and DESeq2 median-of-ratios normalization?
TMM calculates scaling factors using a trimmed mean of log fold changes relative to a reference sample, while DESeq2 median-of-ratios calculates scaling factors using the median of ratios of each gene count to its geometric mean across all samples. Both methods aim to correct for sequencing depth and library composition differences, but they use different statistical approaches to achieve this correction.
When should I use TMM instead of DESeq2 median-of-ratios?
TMM is often preferred when the dataset has strong composition effects, such as a small number of highly expressed genes that differ dramatically between conditions. The trimming parameters in TMM provide explicit control over how much influence these genes have on the scaling factor calculation. TMM is also a good choice when using edgeR for downstream differential expression testing.
When should I use DESeq2 median-of-ratios instead of TMM?
DESeq2 median-of-ratios is a good choice for datasets with moderate composition differences and balanced experimental designs. The method is simple to apply with no tuning parameters, and it performs comparably to TMM in most benchmark studies. DESeq2 is also preferred when using the Wald test or likelihood ratio test for differential expression analysis.
Can I use TMM normalization with DESeq2 for differential expression testing?
It is possible to use TMM-normalized counts with DESeq2, but this requires careful consideration of the statistical assumptions. DESeq2's statistical model is designed to work with its own normalization approach, and mixing methods may produce unreliable results. It is generally recommended to use the normalization method that is integrated with the statistical testing framework.
How do highly expressed genes affect normalization?
Highly expressed genes that differ between conditions consume a larger fraction of sequencing reads, which reduces the counts for all other genes in the sample. This composition effect can create apparent differences in expression for genes that are not actually changing. Both TMM and median-of-ratios methods are designed to be robust to this effect, but extreme cases may require adjusted parameters or careful interpretation.
Does sample size affect the choice of normalization method?
Sample size affects the reliability of both normalization methods and the statistical tests used for differential expression analysis. With very small sample sizes, such as two conditions without replicates, the choice of normalization method becomes particularly important. The median-of-ratios method may be more stable with small sample sizes because the median is a robust estimator, but both methods can produce unreliable results when sample sizes are very small.
How do I know if my normalization worked correctly?
After normalization, check that the scaling factors are reasonable and that the normalized data show expected biological patterns. Principal component analysis or hierarchical clustering should show samples clustering by condition instead of by batch or library size. If samples cluster by technical factors, the normalization may not have adequately corrected for these effects.
Should I report which normalization method I used in my publication?
Yes, reporting the normalization method, software versions, and parameters is essential for reproducibility. The lack of standardized analytical methods leads to uncertainties in data interpretation and study reproducibility. Other researchers need this information to reproduce your analysis and compare results across studies.
Related Bioinformatics Guides
- RNA-Seq Normalization Methods: TPM, RPKM, and Beyond
- RNA-Seq vs qPCR: Validation and Comparison
- RNA-Seq Data Analysis in Galaxy: A User-Friendly Platform
- RNA-Seq Data Analysis Workflow: From Raw Reads to Insights
- Single-Cell RNA Sequencing Depth: A Cost-Benefit Analysis for Experimental Design
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- Confidence: a web app for cross-platform differential gene expression analysis, gene scoring, and enrichment analysis.. 2026.
- From bench to bytes: a practical guide to RNA sequencing data analysis.. 2025.
- Comparative spatial whole transcriptome analysis of matched frozen and formalin-fixed paraffin-embedded colorectal cancer tissues.. 2026.
- A urinary microRNA aging clock accurately predicts biological age.. 2025.
- Population-specific MicroRNA biomarker discovery in breast ductal carcinoma via explainable graph neural multi-omics modeling.. 2026.
- Comprehensive Transcriptomic Analysis and Biomarker Prioritization of Hydroxyprogesterone in Breast Cancer.. 2026.
- In Papyro Comparison of TMM (edgeR), RLE (DESeq2), and MRN Normalization Methods for a Simple Two-Conditions-Without-Replicates RNA-Seq Experimental Design. Frontiers in Genetics, 2016.
- Effect of RNA-Seq data normalization on protein interactome mapping for Alzheimer's disease. Comput. Biol. Chem., 2024.
- Benchmarking differential expression analysis tools for RNA-Seq: normalization-based vs. log-ratio transformation-based methods. BMC Bioinformatics, 2018.
- A realistic simulation-based benchmark of microbiome normalization in sample stratification and taxa-level analysis. Frontiers in Bioinformatics, 2026.
- Supplementary Information for Benchmarking differential expression analysis tools for RNA-Seq : normalization-based vs . log-ratio transformation-based methods. 2018.
- Choice of library size normalization and statistical methods for differential gene expression analysis in balanced two-group comparisons for RNA-seq studies. BMC Genomics, 2020.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.