RNA-seq Normalization Methods: Which One to Use for Accurate Differential Expression Analysis?

By Dr. Zubair Khalid, DVM, MS, PhD ·

RNA-seq Normalization Methods: Which One to Use for Accurate Differential Expression Analysis?

Key Takeaways

  • Raw RNA-seq counts are not directly comparable due to technical variations in sequencing depth, library size, gene length, and transcriptome composition, necessitating normalization for accurate differential expression analysis.
  • Methods like TPM and RPKM/FPKM normalize for gene length and library size but are generally unsuitable for differential expression analysis due to their susceptibility to transcriptome composition biases and their output not being raw counts.
  • Robust differential expression analysis typically relies on methods like DESeq2's median-of-ratios or edgeR's TMM, which estimate scaling factors based on the assumption that most genes remain unchanged between conditions, thereby correcting for composition effects.
  • The selection of a normalization method is critically dependent on experimental design factors including sample size, data distribution assumptions (e.g., negative binomial), and the degree of transcriptome composition variation between samples.
  • Proper normalization is a prerequisite for reliable downstream analyses such as gene set enrichment, pathway analysis, and deconvolution, and its effectiveness should be assessed via methods like PCA or hierarchical clustering to ensure samples group by biological condition.

RNA-seq Normalization Methods: Which One to Use for Accurate Differential Expression Analysis

RNA-seq normalization is the computational step that adjusts raw read counts to make gene expression measurements comparable across samples before differential expression testing. The method you choose directly affects which genes appear differentially expressed, and the wrong choice can produce false positives that waste validation resources and misdirect biological interpretation. This article compares the most common normalization approaches, explains their underlying assumptions, and provides practical selection criteria based on your experimental design and data characteristics.

The Core Problem: Why Raw Counts Cannot Be Compared Directly

RNA-seq experiments generate raw counts that reflect also the true abundance of transcripts in your samples but also technical artifacts that vary between libraries. Two samples sequenced at different depths will produce different raw counts for the same gene even if the biological expression level is identical. Similarly, samples with different total RNA compositions will show apparent differences in individual gene counts that are driven by the overall mixture instead of by genuine regulation of that specific gene.

The constrained nature of sequencing data means that raw counts do not measure absolute transcript abundances. Instead, the data are limited by the sequencing depth of the assay, and normalization is required before statistical analysis can be meaningfully performed [<a href="#ref-1">1</a>]. This constraint is fundamental to understanding why normalization choices matter. If you compare raw counts between samples, you are comparing library sizes as much as biological signals.

Normalization methods attempt to remove these technical effects while preserving the biological differences you want to detect. Each method makes different assumptions about what constitutes a technical artifact versus a biological signal, and those assumptions determine when the method will perform well and when it will fail.

At a Glance: Normalization Method Comparison

MethodUnitKey AssumptionBest Use CasePrimary Limitation
TPMTranscripts per millionGene length and total transcript content should be controlledComparing expression across genes within a sample, visualizationNot suitable as input for most differential expression tools
RPKM/FPKMReads/fragments per kilobase per millionGene length and sequencing depth are the main technical factorsLegacy compatibility, within-sample gene comparisonsBiased by total transcript composition, not recommended for between-sample comparisons
DESeq2 median-of-ratiosNormalized countsMost genes are not differentially expressed between conditionsDifferential expression with small sample sizesAssumes similar library composition across samples
TMM (edgeR)Normalization factorsMost genes are not differentially expressed, composition should be balancedDifferential expression with varied library sizesCan be influenced by highly expressed genes if not trimmed properly

Understanding the Technical Sources of Variation

Sequencing Depth and Library Size

The total number of reads generated for each sample varies due to sequencing machine output, pooling strategies, and sample preparation efficiency. A sample with 30 million reads will have roughly twice the raw counts of a sample with 15 million reads for the same gene, assuming identical biological expression. Library size normalization divides counts by the total number of reads to make samples comparable on a per-million-read basis.

The challenge is that library size alone is not always an adequate scaling factor. If one sample contains a small number of very highly expressed genes that consume a large fraction of the reads, other genes in that sample will appear artificially reduced. This composition effect is not corrected by simple total-count scaling.

Gene Length

Longer genes generate more reads than shorter genes at the same expression level because more fragments can be produced from a longer transcript. This is not a biological difference in regulation but a physical property of the transcript. Methods that correct for gene length, such as RPKM, FPKM, and TPM, divide counts by the gene length in kilobases.

Gene length correction is essential when comparing expression levels across different genes within the same sample. It is less relevant when comparing the same gene across samples, because the gene length is constant. This distinction explains why some methods include length correction and others do not.

Transcriptome Composition

The total RNA content of a cell or tissue is not constant across conditions. A treatment that globally increases transcription will dilute the relative proportion of reads from any single gene. Conversely, a treatment that reduces overall transcription will make individual genes appear more abundant. This composition effect is a major source of false positives in differential expression analysis.

Methods that assume most genes are unchanged between conditions can estimate a scaling factor that accounts for composition differences. Methods that do not make this assumption are vulnerable to composition artifacts.

Core Normalization Methods and Their Assumptions

TPM: Transcripts per Million

TPM normalizes for both gene length and sequencing depth. The calculation first divides read counts by gene length to obtain reads per kilobase, then divides by the total of these length-normalized values across all genes and multiplies by one million. The result is a value that represents the proportion of transcripts in the sample attributable to each gene.

TPM is useful for comparing expression levels across genes within a sample because it accounts for gene length. It is also the recommended unit for many visualization approaches and for methods that assume relative abundance measurements. However, TPM values are not appropriate as direct input to most differential expression tools, which expect raw integer counts and model count distributions.

The within-sample normalization methods, including TPM and FPKM, have been shown to produce higher variability in downstream analyses compared to between-sample methods. In the context of mapping RNA-seq data onto genome-scale metabolic models, TPM and FPKM normalization led to considerably higher variability in the number of active reactions compared to methods like RLE, TMM, and GeTMM [<a href="#ref-2">2</a>]. This variability suggests that within-sample normalization is less stable for analyses that depend on consistent absolute measurements across samples.

RPKM and FPKM: Reads and Fragments per Kilobase per Million

RPKM normalizes read counts by gene length and total reads in the library. FPKM is the paired-end sequencing equivalent, using fragments instead of reads. Both methods were widely used in early RNA-seq analysis but have known limitations.

The primary problem with RPKM and FPKM is that the total number of reads in the library is influenced by the transcriptome composition. If a sample has a large number of highly expressed genes, the total read count is inflated, and all other genes appear artificially reduced. This makes RPKM and FPKM unreliable for comparing expression levels across samples with different compositions.

The total read count used in the denominator is not a stable reference point when the composition of the transcriptome changes between conditions. This instability propagates through the normalization and can create apparent differential expression where none exists biologically.

DESeq2 Median-of-Ratios

The DESeq2 approach estimates a size factor for each sample by taking the median of the ratios of each gene's count to the geometric mean of that gene's counts across all samples. The underlying assumption is that most genes are not differentially expressed, so the median ratio provides a robust estimate of the sequencing depth and composition effect for that sample.

This method is implemented in the DESeq2 Bioconductor package, which provides a complete workflow for differential expression analysis [<a href="#ref-3">3</a>]. The size factors are then used to scale the raw counts before fitting the statistical model.

The median-of-ratios method performs well when the assumption of mostly unchanged genes holds. In benchmark comparisons, DESeq2 showed slightly better performance than other methods when sample sizes reached six or twelve per group, particularly for data following a negative binomial distribution [<a href="#ref-4">4</a>]. For log-normal distributed data, both DESeq and DESeq2 outperformed other methods in terms of false discovery rate control, power, and stability across all sample sizes tested [<a href="#ref-4">4</a>].

TMM: Trimmed Mean of M Values

The TMM method, implemented in the edgeR package, estimates normalization factors by comparing each sample to a reference sample. The method trims the most extreme log-fold-changes and the most extreme absolute expression values, then calculates a weighted mean of the remaining values. This trimmed mean provides a scaling factor that accounts for composition differences while being robust to outliers.

TMM is designed for situations where a minority of genes are highly expressed and could otherwise dominate the normalization. By trimming extreme values, the method reduces the influence of these genes on the scaling factor.

In the metabolic model mapping benchmark, TMM normalization enabled the production of condition-specific models with considerably lower variability in the number of active reactions compared to within-sample methods like FPKM and TPM [<a href="#ref-2">2</a>]. The accuracy of capturing disease-associated genes was approximately 0.80 for Alzheimer's disease data and 0.67 for lung adenocarcinoma data when using RLE, TMM, or GeTMM normalization [<a href="#ref-2">2</a>].

RLE: Relative Log Expression

The RLE method is used in the DESeq package and is conceptually similar to the median-of-ratios approach. It calculates the median of the log ratios between each sample and the geometric mean across samples. The method assumes that most genes are not differentially expressed and uses the median as a robust central tendency measure.

RLE normalization performed well in the metabolic model benchmark, producing condition-specific models with low variability and high accuracy for capturing disease-associated genes [<a href="#ref-2">2</a>]. The method is appropriate when the assumption of mostly unchanged genes holds.

GeTMM: Gene Length Corrected TMM

GeTMM combines gene length correction with TMM normalization. The method first corrects counts for gene length, then applies TMM scaling factors. This approach attempts to provide both the between-sample comparability of TMM and the within-sample gene length correction that is useful for certain analyses.

In the metabolic model benchmark, GeTMM performed comparably to RLE and TMM, with low variability in active reaction numbers and high accuracy for disease-associated genes [<a href="#ref-2">2</a>]. The method may be useful when you need both gene length correction and robust between-sample normalization.

Selecting a Normalization Method for Differential Expression

Sample Size Considerations

The number of biological replicates per condition is a critical factor in method selection. Benchmark studies have shown that method performance varies with sample size. With three samples per group, EBSeq performed better than other methods in terms of false discovery rate control, power, and stability for negative binomial distributed data [<a href="#ref-4">4</a>]. When sample sizes increased to six or twelve per group, DESeq2 performed slightly better than the other methods [<a href="#ref-4">4</a>].

All methods improved their performance when sample size increased to twelve per group, with the exception of DESeq [<a href="#ref-4">4</a>]. This finding suggests that increasing biological replication can compensate for some normalization method limitations, but the choice of method still matters at smaller sample sizes.

For experiments with limited replication, which is common in many research settings, the choice of normalization method can have a more pronounced effect on results. The median-of-ratios approach in DESeq2 and the TMM approach in edgeR are both designed to be robust with small sample sizes, but their assumptions about the data distribution differ.

Data Distribution Assumptions

RNA-seq count data are commonly modeled with a negative binomial distribution. Methods like edgeR and DESeq2 are built around this distribution and estimate dispersion parameters from the data. When the data actually follow a negative binomial distribution, these methods perform well [<a href="#ref-4">4</a>].

However, real data do not always conform to the assumed distribution. The benchmark study found that for log-normal distributed data, both DESeq and DESeq2 performed better than other methods across all sample sizes [<a href="#ref-4">4</a>]. This finding highlights the importance of understanding your data characteristics before selecting a method.

If you are uncertain about the distribution of your data, you can examine the mean-variance relationship in your counts. Negative binomial data show a quadratic mean-variance relationship, while log-normal data show a different pattern. Diagnostic plots can help you assess which assumption is more appropriate for your dataset.

Library Size Variation

Unequal library sizes between samples are common in practice. Some samples may produce more reads due to differences in RNA quality, library preparation efficiency, or sequencing lane allocation. The benchmark study found no significant differences in false discovery rate control, power, or stability across methods when comparing equal versus unequal library sizes [<a href="#ref-4">4</a>].

This finding suggests that the normalization methods are generally robust to library size variation. However, it is still important to check the distribution of library sizes in your dataset and to ensure that no sample has an extremely low read count that would make normalization unreliable.

Composition Effects

The most challenging scenario for normalization is when the transcriptome composition differs substantially between conditions. This can occur in comparisons between different cell types, tissues, or developmental stages where large fractions of the transcriptome are genuinely regulated.

Methods that assume most genes are unchanged will struggle in these situations because the assumption is violated. The trimmed mean approach in TMM and the median-of-ratios approach in DESeq2 are designed to be robust to composition effects by downweighting the influence of extreme values. However, if the majority of genes are differentially expressed, no method can reliably distinguish technical from biological variation.

Practical Workflow for Normalization and Differential Expression

Step 1: Quality Control of Raw Counts

Before any normalization, examine the quality of your raw count data. Check the total number of reads per sample, the number of genes with zero counts, and the distribution of counts across genes. Samples with very low total reads may need to be excluded or resequenced.

The Galaxy Training Network provides accessible tutorials for RNA-seq quality control and analysis workflows that can guide you through these steps [<a href="#ref-5">5</a>]. Similarly, the EMBL-EBI Training resources offer learning pathways for bioinformatics data analysis that cover quality assessment and normalization concepts [<a href="#ref-6">6</a>].

Step 2: Choose Your Analysis Platform

Most differential expression analysis is performed using R and Bioconductor packages. The Bioconductor project provides official documentation for package installation, workflow execution, and reproducible genomic analysis [<a href="#ref-3">3</a>]. DESeq2 and edgeR are both available through Bioconductor and have extensive documentation.

For users who prefer a graphical interface, the Galaxy platform offers web-based workflows for RNA-seq analysis [<a href="#ref-5">5</a>]. For users who work with large-scale pipelines, the nf-core community provides standardized workflows with documented usage and configuration options [<a href="#ref-7">7</a>].

Step 3: Apply the Normalization Method

For DESeq2, the median-of-ratios normalization is applied automatically when you run the DESeq function on a DESeqDataSet object. The size factors are estimated and stored in the object, and the normalized counts can be extracted for visualization or other purposes.

For edgeR, the TMM normalization factors are calculated using the calcNormFactors function. These factors are then used in the differential expression analysis. The edgeR workflow is documented in the Bioconductor package vignette [<a href="#ref-3">3</a>].

Step 4: Assess Normalization Effectiveness

After normalization, check whether the normalization successfully removed technical variation. Principal component analysis or hierarchical clustering of the normalized data should show samples clustering by biological condition instead of by sequencing depth or other technical factors.

You can also examine the size factors or normalization factors estimated by the methods. Large variation in these factors between samples indicates substantial composition or depth differences that the normalization is attempting to correct.

Step 5: Run Differential Expression Analysis

The differential expression analysis uses the normalized data to fit statistical models and test for differences between conditions. The choice of normalization method is embedded in the analysis tool, so you do not need to manually normalize the data before running DESeq2 or edgeR.

The output of the analysis includes log fold changes, p-values, and adjusted p-values for each gene. The adjusted p-values control the false discovery rate and should be used to select differentially expressed genes.

Step 6: Validate Results

Differential expression results should be validated before proceeding to downstream analyses. Check that the top differentially expressed genes make biological sense in the context of your experiment. Consider validating a subset of genes using an orthogonal method such as quantitative PCR.

The single-cell RNA-seq studies cited in this article used multiple validation approaches, including co-immunostaining and RNA in situ hybridization to confirm findings from sequencing data [<a href="#ref-8">8</a>]. Similar validation approaches can strengthen confidence in bulk RNA-seq results.

Records and Measurements for Normalization Decisions

Documenting Library Statistics

Maintain a record of the total reads, mapped reads, and uniquely mapped reads for each sample. These statistics are essential for diagnosing normalization problems and for reporting in publications. The sequence read archive at NCBI provides a repository for raw sequencing data and associated metadata [<a href="#ref-9">9</a>].

Recording Normalization Parameters

Document the normalization method used, the version of the software, and any parameters that were changed from defaults. This information is essential for reproducibility. The nf-core documentation emphasizes the importance of standardized pipeline usage and configuration for reproducible analysis [<a href="#ref-7">7</a>].

Tracking Quality Metrics

Record quality metrics before and after normalization. These include the distribution of counts, the number of genes detected per sample, and the correlation between technical replicates. Changes in these metrics can indicate problems with the normalization.

Reproducibility Records

For published work, record the exact commands used for normalization and differential expression analysis. Version numbers of all software packages should be documented. The Carpentries lessons provide foundational training in computing skills that support reproducible data analysis, including version control and documentation practices [<a href="#ref-10">10</a>].

Common Failure Patterns in Normalization

Failure Pattern 1: Using TPM or FPKM for Differential Expression

A common error is to normalize counts to TPM or FPKM and then use these values as input to differential expression tools that expect raw counts. This approach violates the distributional assumptions of the tools and can produce unreliable results.

The within-sample normalization methods like TPM and FPKM are designed for different purposes than between-sample normalization. Using them for differential expression testing can lead to inflated false positive rates because the normalization does not adequately account for composition effects.

Failure Pattern 2: Ignoring Composition Effects

When comparing conditions with very different transcriptome compositions, such as tumor versus normal tissue, simple normalization methods can produce many false positives. The highly expressed genes in one condition consume a large fraction of the reads, making other genes appear artificially reduced.

The benchmark study on normalization methods for metabolic model mapping found that within-sample methods like FPKM and TPM produced considerably higher variability in downstream analyses compared to between-sample methods like RLE, TMM, and GeTMM [<a href="#ref-2">2</a>]. This variability reflects the instability of within-sample normalization when composition differs between samples.

Failure Pattern 3: Insufficient Biological Replication

With only two or three samples per condition, the statistical power to detect differential expression is limited, and the results are sensitive to outliers. The benchmark study found that method performance varied with sample size, with EBSeq performing better at three samples per group and DESeq2 performing better at six or twelve samples per group [<a href="#ref-4">4</a>].

Increasing biological replication is the most effective way to improve the reliability of differential expression results. If replication is limited by cost or sample availability, consider using methods that are robust to small sample sizes and interpret results with appropriate caution.

Failure Pattern 4: Ignoring Batch Effects

Normalization methods do not correct for batch effects, which are systematic technical differences between groups of samples processed at different times or in different batches. Batch effects can confound biological differences and produce false positives or false negatives.

If batch effects are present, they should be modeled in the statistical analysis instead of relying on normalization to remove them. Some analysis workflows include batch correction steps, but these should be applied carefully to avoid removing biological signal.

Failure Pattern 5: Applying Single-Cell Normalization Methods to Bulk Data

Single-cell RNA-seq data have different characteristics than bulk RNA-seq data, including higher dropout rates and greater technical variability. Normalization methods developed for single-cell data, such as SCnorm, are designed for these specific challenges [<a href="#ref-11">11</a>]. Applying these methods to bulk data may not be appropriate, and applying bulk methods to single-cell data can produce poor results.

The normalization of single-cell RNA-seq data is an active area of research, and no single method outperforms all others across all datasets [<a href="#ref-12">12</a>]. Data-driven metrics can be used to rank normalization methods and select the best performers for a given dataset [<a href="#ref-12">12</a>].

Limitations of Normalization Methods

No Universal Best Method

The choice of normalization method can have a profound impact on the results, and no single method outperforms all others in all datasets [<a href="#ref-12">12</a>]. The performance of a method depends on the characteristics of the data, including the distribution of counts, the sample size, and the degree of composition difference between conditions.

This finding means that researchers should not assume that a method that worked well for one dataset will work equally well for another. It is prudent to assess the performance of different methods on your specific data and to consider how sensitive your conclusions are to the choice of normalization.

Normalization Cannot Fix Poor Experimental Design

Normalization methods can correct for technical variation, but they cannot compensate for inadequate biological replication, poor RNA quality, or confounding between biological and technical factors. The most sophisticated normalization approach will not produce reliable results from a poorly designed experiment.

Transcriptome Size Variation

Recent research has highlighted that variation in transcriptome size across cell types significantly impacts RNA-seq data normalization [<a href="#ref-13">13</a>]. The total amount of RNA per cell can vary substantially between cell types, and this variation affects the interpretation of normalized counts.

A new normalization approach called Count based on Linearized Transcriptome Size (CLTS) has been developed to correct differentially expressed genes that are typically misidentified by standard count per 10K normalization [<a href="#ref-13">13</a>]. This approach maintains transcriptome size variation and enhances the accuracy of downstream analyses such as bulk deconvolution [<a href="#ref-13">13</a>].

This finding is particularly relevant for studies comparing different cell types or tissues, where transcriptome size differences may be substantial. Standard normalization methods that assume similar total RNA content per cell may produce misleading results in these comparisons.

Gene Length Effects

Gene length effects can influence normalization and downstream analyses. The ReDeconv algorithm, which incorporates transcriptome size into normalization, also mitigates gene length effects and models expression variances, improving deconvolution outcomes particularly for rare cell types [<a href="#ref-13">13</a>].

For standard differential expression analysis, gene length effects are less problematic because the same gene is compared across samples. However, for analyses that compare expression across genes, such as gene set enrichment analysis, gene length effects should be considered.

Welfare and Safety Context in Research Applications

Ethical Use of Animal-Derived Samples

Many RNA-seq experiments use samples derived from animal tissues. Researchers must ensure that sample collection complies with institutional animal care and use regulations. The normalization and analysis of the data are separate from the sample collection, but the quality of the data depends on proper sample handling.

Data Management and Reproducibility

Reproducible analysis practices are essential for the integrity of research findings. The nf-core documentation emphasizes community pipeline standards for reproducible analysis [<a href="#ref-7">7</a>]. The Carpentries lessons provide foundational training in data management and reproducible computing practices [<a href="#ref-10">10</a>].

Reporting Standards

Publications reporting RNA-seq results should include detailed descriptions of the normalization method, software versions, and parameters used. This information allows other researchers to reproduce the analysis and to assess the robustness of the findings.

Professional Escalation Criteria

When to Seek Expert Consultation

Consider consulting a bioinformatics specialist or biostatistician when your data present any of the following characteristics:

  • Extreme composition differences between conditions, such as comparing different tissue types or cell populations
  • Very small sample sizes with high variability
  • Evidence of strong batch effects that cannot be easily modeled
  • Unusual count distributions that do not fit standard models
  • The need to integrate RNA-seq data with other data types, such as single-cell data or proteomics data

When to Reconsider Your Analysis Approach

You should reconsider your normalization and analysis approach if you observe any of the following:

  • Principal component analysis shows samples clustering by technical factors instead of biological conditions
  • The number of differentially expressed genes is implausibly high or low
  • Known positive or negative control genes do not show the expected expression patterns
  • Results are highly sensitive to the choice of normalization method

When to Seek Additional Resources

The EMBL-EBI Training portal provides learning pathways for bioinformatics data analysis that can help you build the skills needed for RNA-seq analysis [<a href="#ref-6">6</a>]. The Galaxy Training Network offers hands-on tutorials for RNA-seq analysis workflows [<a href="#ref-5">5</a>]. The Bioconductor project provides official documentation for the packages used in differential expression analysis [<a href="#ref-3">3</a>].

Integration with Downstream Analyses

Differential Expression Results as Input to Other Analyses

The choice of normalization method affects also the list of differentially expressed genes but also downstream analyses that use these results. Gene set enrichment analysis, pathway analysis, and network analysis all depend on the quality of the differential expression results.

In the context of genome-scale metabolic modeling, the normalization method of choice affects the model content produced by algorithms like iMAT and INIT and their predictive accuracy [<a href="#ref-2">2</a>]. The benchmark study found that RLE, TMM, and GeTMM normalization enabled the production of condition-specific metabolic models with considerably lower variability in the number of active reactions compared to within-sample methods [<a href="#ref-2">2</a>].

Integration with Single-Cell Data

Many research projects integrate bulk RNA-seq data with single-cell RNA-seq data. The normalization choices for each data type can affect the integration. The single-cell studies cited in this article used various approaches to integrate data types, including using bulk RNA-seq data to confirm findings from single-cell analyses [<a href="#ref-14">14</a>][<a href="#ref-15">15</a>].

The dilated cardiomyopathy study integrated single-cell RNA-seq data from 70,958 single cells with bulk RNA-seq data to confirm the roles of specific cell subpopulations and genes [<a href="#ref-14">14</a>]. The non-small cell lung cancer study used single-cell RNA-seq to identify mixed-lineage tumor cells and then validated the findings with orthogonal methods [<a href="#ref-8">8</a>].

Deconvolution Analyses

Bulk RNA-seq deconvolution aims to estimate the cell type composition of bulk samples using reference signatures from single-cell data. The normalization of both the single-cell reference data and the bulk data affects the accuracy of deconvolution.

The ReDeconv algorithm incorporates transcriptome size into both single-cell normalization and bulk deconvolution, improving the accuracy of cell type percentage predictions [<a href="#ref-13">13</a>]. This approach addresses the issue that standard normalization methods can distort the relative proportions of cell types when transcriptome sizes differ.

Case Examples from the Literature

Dilated Cardiomyopathy Study

A study of dilated cardiomyopathy integrated single-cell RNA-seq data from 7 DCM samples and 3 normal heart tissue samples, totaling 70,958 single-cell data points [<a href="#ref-14">14</a>]. The analysis identified 9 distinct cell subtypes and used machine learning methods to quantify bulk RNA-seq data, finding significant differences in fibroblasts, T cells, and macrophages between DCM and normal samples [<a href="#ref-14">14</a>].

The study demonstrated the importance of appropriate normalization for identifying cell type differences. The integration of single-cell and bulk data required consistent normalization approaches across data types to produce reliable results.

Non-Small Cell Lung Cancer Study

A study of non-small cell lung cancer performed single-cell RNA-seq on 7,364 individual cells from tumor tissues and matched normal tissues from 19 primary lung cancer patients and 1 pulmonary chondroid hamartoma patient [<a href="#ref-8">8</a>]. The analysis identified mixed-lineage tumor cells that simultaneously expressed marker genes for two or three histologic subtypes of NSCLC [<a href="#ref-8">8</a>].

The study used copy number variation patterns and mitochondrial mutations to show that mixed-lineage and single-lineage tumor cells from the same patient had common tumor ancestors [<a href="#ref-8">8</a>]. These findings depended on accurate normalization to distinguish genuine biological differences from technical artifacts.

Hepatocellular Carcinoma Study

A study of hepatocellular carcinoma used single-cell RNA-seq data from the GEO database and bulk RNA-seq data from TCGA to explore cancer-associated fibroblast characteristics [<a href="#ref-15">15</a>]. The analysis identified 4 CAF clusters, 3 of which were associated with prognosis, and constructed a risk signature with 6 genes [<a href="#ref-15">15</a>].

The study demonstrated the integration of multiple data sources and the importance of consistent normalization for cross-dataset comparisons. The risk signature was significantly associated with stromal and immune scores, as well as some immune cells [<a href="#ref-15">15</a>].

Major Depressive Disorder Study

A study of major depressive disorder performed single-cell RNA-seq and ATAC-seq on peripheral blood mononuclear cells from 8 healthy individuals and 8 MDD patients before or after 12 weeks of antidepressant treatment [<a href="#ref-16">16</a>]. The analysis revealed a lower proportion of naive T cells in MDD patients compared to controls [<a href="#ref-16">16</a>].

The study used flow cytometry data from an independent cohort of 35 patients and 40 healthy individuals to confirm the findings [<a href="#ref-16">16</a>]. This validation approach demonstrates the importance of confirming sequencing-based findings with orthogonal methods.

Cutaneous Squamous Cell Carcinoma Study

A study of cutaneous squamous cell carcinoma combined single-cell RNA sequencing with spatial transcriptomics and multiplexed ion beam imaging [<a href="#ref-17">17</a>]. The analysis identified four tumor subpopulations and a tumor-specific keratinocyte population unique to cancer [<a href="#ref-17">17</a>].

The study mapped ligand-receptor networks to specific cell types and identified the tumor-specific keratinocyte population as a hub for intercellular communication [<a href="#ref-17">17</a>]. The integration of multiple data types required careful normalization to ensure that comparisons across data modalities were valid.

Frequently Asked Questions

What is the difference between TPM and RPKM normalization?

TPM and RPKM both normalize for gene length and sequencing depth, but they calculate the normalization differently. RPKM first divides counts by gene length and then by the total number of reads in the library. TPM first divides counts by gene length, then divides by the total of these length-normalized values across all genes, and multiplies by one million. The key difference is that TPM values sum to one million across all genes, making them interpretable as proportions of the transcriptome. RPKM values do not have this property, and the total RPKM across genes varies between samples depending on the transcriptome composition. For comparing expression across samples, TPM is generally preferred over RPKM because the sum-to-one-million property makes the values more comparable.

Should I use TPM values for differential expression analysis?

TPM values are not appropriate as direct input to most differential expression tools. Tools like DESeq2 and edgeR expect raw integer counts and model the count distribution using statistical distributions such as the negative binomial. If you provide TPM values, the tools will treat them as counts, which violates the distributional assumptions and can produce unreliable results. TPM values are useful for visualization, for comparing expression across genes within a sample, and for certain downstream analyses that expect normalized expression values. For differential expression testing, use the normalization methods built into the analysis tools, such as the median-of-ratios method in DESeq2 or the TMM method in edgeR.

How do I choose between DESeq2 and edgeR for differential expression?

DESeq2 and edgeR use different normalization approaches and statistical models. DESeq2 uses the median-of-ratios method, while edgeR uses the TMM method. Both methods assume that most genes are not differentially expressed and estimate scaling factors to account for composition differences. The choice between them depends on your data characteristics. Benchmark studies have shown that DESeq2 performs slightly better than other methods when sample sizes reach six or twelve per group for negative binomial distributed data [<a href="#ref-4">4</a>]. For log-normal distributed data, both DESeq and DESeq2 performed better than other methods [<a href="#ref-4">4</a>]. In practice, many researchers run both methods and compare the results to identify robust differentially expressed genes.

What sample size do I need for reliable differential expression results?

The benchmark study found that method performance varied with sample size. With three samples per group, EBSeq performed better than other methods in terms of false discovery rate control, power, and stability for negative binomial distributed data [<a href="#ref-4">4</a>]. When sample sizes increased to six or twelve per group, DESeq2 performed slightly better than the other methods [<a href="#ref-4">4</a>]. All methods improved their performance when sample size increased to twelve per group, except DESeq [<a href="#ref-4">4</a>]. These findings suggest that at least three biological replicates per condition are needed for meaningful analysis, and six or more replicates provide more reliable results. The optimal sample size depends on the biological variability in your system and the magnitude of the expression differences you want to detect.

What is the median-of-ratios normalization method in DESeq2?

The median-of-ratios method estimates a size factor for each sample by calculating the median of the ratios of each gene's count to the geometric mean of that gene's counts across all samples. The underlying assumption is that most genes are not differentially expressed, so the median ratio provides a robust estimate of the sequencing depth and composition effect for that sample. The size factors are then used to scale the raw counts before fitting the statistical model. This method is robust to outliers because the median is less sensitive to extreme values than the mean. The method is implemented in the DESeq2 Bioconductor package, which provides a complete workflow for differential expression analysis [<a href="#ref-3">3</a>].

What is the TMM normalization method in edgeR?

The TMM method, which stands for Trimmed Mean of M values, estimates normalization factors by comparing each sample to a reference sample. The method trims the most extreme log-fold-changes and the most extreme absolute expression values, then calculates a weighted mean of the remaining values. This trimmed mean provides a scaling factor that accounts for composition differences while being robust to outliers. TMM is designed for situations where a minority of genes are highly expressed and could otherwise dominate the normalization. By trimming extreme values, the method reduces the influence of these genes on the scaling factor. The method is implemented in the edgeR Bioconductor package.

How does transcriptome size affect normalization?

Transcriptome size variation across cell types significantly impacts RNA-seq data normalization [<a href="#ref-13">13</a>]. The total amount of RNA per cell can vary substantially between cell types, and this variation affects the interpretation of normalized counts. Standard normalization methods that assume similar total RNA content per cell may produce misleading results when comparing cell types with different transcriptome sizes. A new normalization approach called Count based on Linearized Transcriptome Size (CLTS) has been developed to correct differentially expressed genes that are typically misidentified by standard count per 10K normalization [<a href="#ref-13">13</a>]. This approach maintains transcriptome size variation and enhances the accuracy of downstream analyses such as bulk deconvolution [<a href="#ref-13">13</a>].

Can I use the same normalization method for bulk and single-cell RNA-seq data?

Bulk and single-cell RNA-seq data have different characteristics, and normalization methods developed for one data type may not be appropriate for the other. Single-cell data have higher dropout rates, greater technical variability, and often show different count distributions than bulk data. Normalization methods developed for single-cell data, such as SCnorm, are designed for these specific challenges [<a href="#ref-11">11</a>]. The normalization of single-cell RNA-seq data is an active area of research, and no single method outperforms all others across all datasets [<a href="#ref-12">12</a>]. If you are integrating bulk and single-cell data, you should consider the normalization choices for each data type and how they affect the integration.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

[1] [Benchmarking differential expression analysis tools for RNA-Seq: normalization-based vs. log-ratio transformation-based methods](https://doi.org/10.1186/s12859-018-2261-8). BMC Bioinformatics, 2018. [2] [A benchmark of RNA-seq data normalization methods for transcriptome mapping on human genome-scale metabolic networks](https://doi.org/10.1038/s41540-024-00448-z). npj Systems Biology and Applications, 2024. [3] [Bioconductor](https://bioconductor.org/). Bioconductor Project. [4] [An evaluation of RNA-seq differential analysis methods.](https://pubmed.ncbi.nlm.nih.gov/36112652). PloS one, 2022. [5] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project. [6] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute. [7] [nf-core Documentation](https://nf-co.re/docs). nf-core. [8] [Molecular profiling of human non-small cell lung cancer by single-cell RNA-seq.](https://pubmed.ncbi.nlm.nih.gov/35962452). Genome medicine, 2022. [9] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information. [10] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries. [11] [SCnorm: Robust normalization of single-cell RNA-seq data](https://doi.org/10.1038/nmeth.4263). Nature Methods, 2017. [12] [Normalization of Single-Cell RNA-Seq Data.](https://pubmed.ncbi.nlm.nih.gov/33835450). Methods in molecular biology (Clifton, N.J.), 2021. [13] [Transcriptome size matters for single-cell RNA-seq normalization and bulk deconvolution](https://doi.org/10.1038/s41467-025-56623-1). Nature Communications, 2025. [14] [Comprehensive analysis of scRNA-seq and bulk RNA-seq reveals the non-cardiomyocytes heterogeneity and novel cell populations in dilated cardiomyopathy.](https://pubmed.ncbi.nlm.nih.gov/39762897). Journal of translational medicine, 2025. [15] [Characterization of cancer-related fibroblasts (CAF) in hepatocellular carcinoma and construction of CAF-based risk signature based on single-cell RNA-seq and bulk RNA-seq data.](https://pubmed.ncbi.nlm.nih.gov/36211448). Frontiers in immunology, 2022. [16] [Integrated Single-Cell RNA-seq and ATAC-seq Reveals Heterogeneous Differentiation of CD4(+) Naive T Cell Subsets is Associated with Response to Antidepressant Treatment in Major Depressive Disorder.](https://pubmed.ncbi.nlm.nih.gov/38867657). Advanced science (Weinheim, Baden-Wurttemberg, Germany), 2024. [17] [Multimodal Analysis of Composition and Spatial Architecture in Human Squamous Cell Carcinoma.](https://pubmed.ncbi.nlm.nih.gov/32579974). Cell, 2020.

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.