Heatmaps for RNA-seq: Best Practices for Clustering, Scaling, and Visualization with pheatmap

By Dr. Zubair Khalid, DVM, MS, PhD ·

Heatmaps for RNA-seq: Best Practices for Clustering, Scaling, and Visualization with pheatmap

Key Takeaways

  • Data Preparation is Paramount: Raw RNA-seq counts require normalization (e.g., median-of-ratios, VST) to account for library size and composition effects before heatmap visualization; log-transformed normalized values are generally preferred over raw counts for reduced dynamic range and improved visual accessibility.
  • Gene Selection Drives Interpretation: Genome-wide heatmaps are rarely informative; focus on biologically relevant subsets such as differentially expressed genes (DEGs) identified via statistical testing, genes within specific pathways, or the most variable genes to highlight key biological signals.
  • Z-score Scaling is Standard for Relative Patterns: Row-wise z-score transformation is crucial for comparing relative expression changes across samples for genes with different absolute expression levels, where a positive z-score indicates expression above the gene's mean and a negative z-score indicates expression below its mean.
  • Hierarchical Clustering Reveals Co-expression and Sample Similarity: Default hierarchical clustering with Euclidean distance and Ward linkage groups genes with similar expression profiles and samples with similar overall transcriptomes, aiding in the identification of co-regulated gene sets and experimental condition effects.
  • Color Scales and Annotations Enhance Interpretability: Diverging color scales (e.g., blue-white-red) are optimal for z-scored data, with carefully chosen breaks and limits to ensure appropriate contrast; sample and gene annotations are essential for contextualizing observed patterns with experimental metadata or functional categories.

RNA sequencing produces genome-wide expression measurements that require careful visualization for interpretation. Heatmaps are among the most common graphical outputs in transcriptomics research, yet their construction involves decisions that materially affect what readers can conclude from the figure. This article addresses the specific problem of creating informative heatmaps from gene expression data using the R package pheatmap. The focus is on practical choices for data transformation, clustering, color scales, and annotation that determine whether a heatmap communicates biological signal or obscures it. The intended readers are biology students, researchers, laboratory professionals, and life-science practitioners who have count matrices or normalized expression tables and need to produce defensible visualizations for analysis, reporting, or publication.

The Role of Heatmaps in RNA-seq Analysis

Heatmaps serve a distinct function within the broader RNA-seq analysis workflow. They provide a compact visual summary of expression patterns across many genes and many samples simultaneously. This makes them useful for exploratory analysis, quality assessment, and communicating results to readers who need to grasp overall trends without inspecting individual values.

The RNA-seq analysis pipeline typically proceeds from raw sequencing reads through quality control, alignment or pseudoalignment, quantification, normalization, and differential expression testing. Each stage produces intermediate outputs that can be visualized. Heatmaps commonly appear at two points in this workflow. First, they can display sample-level patterns such as library composition, batch effects, or outlier samples. Second, they can display gene-level patterns such as clusters of co-expressed genes, expression differences between experimental groups, or the behavior of a curated gene set of interest.

The practical guide by From bench to bytes: a practical guide to RNA sequencing data analysis emphasizes that RNA-seq analysis demands proficiency with computational and statistical approaches to manage technical issues and large data sizes. The same principle applies to visualization. A heatmap is a statistical graphic whose appearance depends on choices about scaling, clustering, and color mapping. Different choices can lead different viewers to different conclusions about the same underlying data.

The Bioconductor project provides the official documentation and package infrastructure for many R-based genomic analysis tools, including pheatmap and the packages that prepare data for it. Researchers should treat pheatmap as one component within a reproducible workflow that includes documented data processing steps and version-controlled code.

Data Inputs and Preparation for pheatmap

Before any heatmap can be drawn, the data must be in a suitable format. pheatmap accepts a numeric matrix where rows represent genes or genomic features and columns represent samples. The values in the matrix should be expression measurements that have already undergone appropriate processing.

Count Matrices versus Normalized Expression Values

Raw count matrices contain integer values representing the number of sequencing reads that mapped to each gene in each sample. These values are not directly suitable for heatmap visualization because they are confounded by library size, gene length, and composition effects. A sample sequenced to greater depth will have higher counts across all genes, which would make the sample appear artificially distinct in a heatmap.

Normalization is required before visualization. Common approaches include library size normalization, trimmed mean of M-values, median-of-ratios, and variance stabilizing transformation. The choice of normalization method depends on the downstream analysis goals and the assumptions that are reasonable for the data. The Galaxy Training Network provides accessible workflow training that covers normalization choices within complete analysis pipelines, which is useful for researchers who want to see how normalization fits into the broader context.

For heatmap purposes, the key requirement is that the values being plotted are comparable across samples. This means that technical artifacts such as sequencing depth differences have been removed or accounted for. Normalized log-scale values are typically preferred over raw counts because they reduce the dynamic range and make patterns more visually accessible.

Selecting Genes for Visualization

Genome-wide heatmaps that include all expressed genes are rarely informative. With tens of thousands of rows, individual patterns become invisible and the figure becomes a solid block of color. Practical heatmaps typically display a subset of genes selected according to the biological question.

Common selection strategies include differentially expressed genes from a statistical test, genes belonging to a pathway or gene set of interest, the most variable genes across samples, or genes from a published signature. Each strategy serves a different purpose. Differentially expressed genes highlight the response to the experimental condition. Highly variable genes reveal the dominant sources of variation in the dataset. Curated gene sets test a specific hypothesis about pathway activity.

The TCC-GUI application includes heatmap with hierarchical clustering as part of its exploratory analysis tools, demonstrating that gene selection and heatmap construction are often integrated into differential expression workflows. Researchers using pheatmap directly should make the gene selection step explicit and documented so that readers understand why particular genes appear in the figure.

Missing Values and Filtering

RNA-seq matrices frequently contain missing values or zero counts. pheatmap requires a complete numeric matrix without missing entries. Genes with zero counts across all samples provide no information and should be removed. Genes with missing values in some samples require a decision about imputation or removal.

The safest approach is to filter genes based on expression level before visualization. A common filter requires a minimum count in a minimum number of samples. The specific thresholds depend on the dataset and the analysis goals. Filtering reduces noise from lowly expressed genes and improves the stability of clustering results. The filtering decisions should be recorded in the analysis documentation so that the visualization can be reproduced.

Scaling and Transformation Choices

The choice of scaling determines what the colors in a heatmap represent. This is the most consequential decision in heatmap construction because it directly controls the visual message.

Z-Score Transformation

The most common scaling approach for gene expression heatmaps is the z-score, also called the row-standardized score. For each gene, the mean expression across samples is subtracted and the result is divided by the standard deviation across samples. This produces values centered at zero with a standard deviation of one.

Z-score scaling makes genes with different absolute expression levels comparable. A highly expressed gene and a lowly expressed gene can both show the same color when they deviate from their own means by the same number of standard deviations. This is appropriate when the research question concerns relative changes in expression pattern instead of absolute expression level.

The interpretation of a z-score heatmap requires care. A red cell does not mean the gene is highly expressed in an absolute sense. It means the gene is expressed above its own average across the samples shown in the heatmap. This distinction matters when communicating results to readers who may not be familiar with the scaling.

Row Scaling versus Column Scaling

pheatmap allows scaling by row, by column, or not at all. Row scaling is the standard choice for gene expression heatmaps because it normalizes each gene across samples. Column scaling normalizes each sample across genes, which can be useful for comparing sample-level expression profiles but is less common for gene-focused visualizations.

No scaling means the raw values are plotted directly. This is appropriate only when the values are already on a comparable scale, such as log-transformed normalized counts from samples with similar library sizes. Without scaling, genes with higher absolute expression will dominate the color scale and lowly expressed genes will appear uniformly as background color.

Log Transformation

Expression data are typically right-skewed, with a small number of highly expressed genes and a long tail of lowly expressed genes. Log transformation compresses the dynamic range and makes the distribution more symmetric. Most RNA-seq visualization workflows apply a log2 transformation to normalized counts before further processing.

The choice of log base affects the interpretation of color differences. A log2 scale means that a difference of one unit corresponds to a twofold change in expression. This is a natural scale for biologists who think in terms of fold changes. The EMBL-EBI Training resources cover transformation choices within the context of practical analysis education, helping researchers understand when log transformation is appropriate.

Variance Stabilizing Transformation

Some normalization methods produce values that are already transformed to stabilize variance across the expression range. The variance stabilizing transformation from DESeq2 and the regularized log transformation are examples. These transformations produce values that can be plotted without additional scaling, although row scaling is still often applied for visualization purposes.

The choice between z-score scaling of log-transformed counts and using variance stabilized values directly depends on the downstream analysis. For heatmaps intended to accompany differential expression results, using the same transformation as the statistical analysis maintains consistency.

Clustering Methods and Their Consequences

Clustering arranges rows and columns so that similar expression patterns appear adjacent to each other. The clustering method determines what similarity means and therefore what patterns the eye perceives.

Hierarchical Clustering in pheatmap

pheatmap uses hierarchical clustering by default. The algorithm builds a tree of similarities, then cuts the tree to produce an ordering. The user controls the distance metric and the linkage method.

The distance metric defines how similarity between two expression profiles is measured. Common choices include Euclidean distance, correlation-based distance, and Manhattan distance. Euclidean distance measures the straight-line distance between expression vectors. Correlation-based distance measures the similarity of patterns regardless of absolute magnitude. For z-scored data, Euclidean distance and correlation distance often produce similar results because the scaling has already removed magnitude differences.

The linkage method defines how clusters are merged. Ward's method minimizes the total within-cluster variance and tends to produce compact clusters. Complete linkage uses the maximum distance between members of two clusters and tends to produce more spread-out clusters. Average linkage uses the mean distance and represents a compromise. The choice of linkage affects the tree structure and therefore the final ordering of rows and columns.

Clustering Rows versus Columns

Row clustering orders genes so that genes with similar expression patterns are adjacent. This reveals groups of co-expressed genes that may share regulatory mechanisms or participate in common pathways. Column clustering orders samples so that samples with similar overall expression profiles are adjacent. This reveals sample groups that may correspond to experimental conditions, tissue types, or disease states.

Both row and column clustering can be applied simultaneously, which is the default in pheatmap. This produces the characteristic block structure of many published heatmaps, where groups of genes show coordinated patterns across groups of samples.

When to Disable Clustering

There are situations where clustering should be disabled. If the rows represent genes in a defined order, such as a pathway diagram or a genomic position, clustering would destroy the meaningful order. If the columns represent time points or dose levels, clustering would disrupt the natural ordering and make trends harder to see.

pheatmap allows the user to set cluster_rows = FALSE or cluster_cols = FALSE to preserve a specified order. This is often the right choice for time series data, dose-response experiments, or gene sets with a defined biological order.

Assessing Cluster Quality

Clustering always produces clusters, even when the data contain no real structure. Researchers should assess whether the observed clusters are biologically meaningful or artifacts of the algorithm. The CONCORDEX approach provides a quantitative assessment of single-cell RNA-seq clustering, and similar principles apply to bulk RNA-seq heatmaps. The degree of separation between compared groups can be measured, and the stability of clusters can be assessed by resampling or by testing different clustering parameters.

A practical check is to run the clustering with different distance metrics and linkage methods. If the major clusters remain consistent across methods, the structure is likely robust. If the clusters change dramatically with method choice, the apparent structure may not be reliable.

Color Scales and Visual Perception

The color scale maps numeric values to colors. This mapping determines what the viewer sees and therefore what conclusions the viewer draws.

Sequential versus Diverging Color Scales

Gene expression heatmaps typically use a diverging color scale, where a central color represents the midpoint and two different colors represent values above and below the midpoint. The classic choice is blue for low values, white for the midpoint, and red for high values. This scheme works well with z-scored data where the midpoint of zero has a clear interpretation.

Sequential color scales, where a single hue progresses from light to dark, are appropriate when the values have a natural zero and only positive values are meaningful. These are less common for gene expression heatmaps but may be useful for displaying raw counts or other non-negative quantities.

The FIt-SNE heatmap-style visualization demonstrates that color choices matter for simultaneously visualizing expression patterns from thousands of genes. The authors implemented a heatmap-style visualization for single-cell RNA-seq based on one-dimensional t-SNE, showing that thoughtful color mapping can make large-scale expression patterns visible.

Setting Breaks and Limits

The color scale has limits that define the minimum and maximum values mapped to the extreme colors. Values beyond these limits are clipped to the extreme colors. The choice of limits determines the contrast in the figure.

If the limits are too narrow, the extreme colors will be oversaturated and subtle differences will be invisible. If the limits are too wide, most cells will appear as intermediate colors and the figure will lack contrast. A common approach is to set symmetric limits around zero for z-scored data, such as negative two to positive two, and to clip extreme values.

pheatmap allows the user to specify breaks that define the boundaries between color intervals. The breaks should be chosen to match the data distribution. For z-scored data, breaks at integer values from negative two to positive two provide a natural scale where each color step corresponds to one standard deviation.

Color Blindness Considerations

A substantial fraction of the population has some form of color vision deficiency. Red-green color scales are problematic for these readers. Alternatives include blue-yellow scales, purple-yellow scales, or scales that use color and brightness differences that are distinguishable by most viewers.

The choice of color scale is a scientific decision, also an aesthetic one. The scale should be perceptually uniform, meaning that equal steps in data values produce equal steps in perceived color difference. Many modern color scales, such as viridis and magma, are designed to be perceptually uniform and color-blind safe.

Annotation and Figure Organization

Annotations add context to a heatmap by displaying additional information about rows or columns. This information helps readers interpret the patterns in the main heatmap body.

Sample Annotation

Sample annotation displays metadata about each column, such as treatment group, tissue type, time point, or batch. pheatmap supports annotation data frames that are displayed as colored bars above the heatmap. The annotation colors should be chosen to be distinguishable and consistent across figures in the same study.

Sample annotation is essential for interpreting whether the clustering in the heatmap corresponds to experimental groups. If samples cluster by batch instead of by treatment, this is visible in the annotation and signals a potential batch effect that requires attention.

Gene Annotation

Gene annotation displays information about each row, such as membership in a pathway, chromosome location, or gene type. This can help readers see whether genes with similar expression patterns share functional characteristics.

Gene annotation is particularly useful when the heatmap displays a curated gene set. The annotation can show which sub-pathway or functional category each gene belongs to, allowing readers to see whether the expression patterns align with functional groupings.

Row and Column Labels

Labels identify individual genes and samples. In large heatmaps, labels may overlap or become unreadable. pheatmap provides options for font size, label angle, and whether labels are shown at all.

For heatmaps with many rows, showing all gene labels is impractical. Options include showing labels only for selected genes, using a smaller font size, or omitting labels and providing the gene list in supplementary material. The decision should balance readability against the need for readers to identify specific genes.

Figure Dimensions and Resolution

The physical dimensions of the heatmap affect its readability. A heatmap with many rows requires a tall figure. A heatmap with many columns requires a wide figure. The font sizes for labels and annotations should scale with the figure dimensions.

For publication, heatmaps should be saved at sufficient resolution. Vector formats such as PDF or SVG preserve detail at any size. Raster formats such as PNG should be saved at high resolution, typically 300 dots per inch or higher for print.

Practical Workflow for Creating Heatmaps with pheatmap

The following workflow provides a structured approach to creating heatmaps with pheatmap. Each step includes decisions that should be documented for reproducibility.

Step 1: Prepare the Expression Matrix

Load the normalized expression data into R. The matrix should have genes as rows and samples as columns. Verify that the row names are unique gene identifiers and the column names are unique sample identifiers. Check for missing values and decide how to handle them.

Filter the matrix to the genes of interest. This may be based on differential expression results, a gene set, or variability across samples. Record the filtering criteria and the number of genes retained.

Step 2: Apply Scaling

Decide whether to scale by row, by column, or not at all. For most gene expression heatmaps, row scaling with z-scores is appropriate. Apply the scaling and verify that the resulting values have the expected distribution.

If the data are not already log-transformed, apply the transformation before scaling. The choice of transformation should match the downstream analysis.

Step 3: Configure Clustering

Choose the distance metric and linkage method for row and column clustering. Consider whether clustering is appropriate for both dimensions or whether one dimension should preserve a defined order.

Test the stability of the clustering by varying the parameters. If the major clusters are consistent across parameter choices, the structure is likely robust.

Step 4: Select the Color Scale

Choose a diverging color scale for z-scored data. Set the limits and breaks to match the data distribution. Consider color blindness when selecting the color palette.

Verify that the color scale produces a figure with appropriate contrast. If the figure appears saturated or washed out, adjust the limits.

Step 5: Add Annotations

Create annotation data frames for samples and genes. Choose colors for the annotation categories that are distinguishable and consistent with other figures in the study.

Verify that the annotation colors do not conflict with the heatmap color scale.

Step 6: Generate and Inspect the Figure

Create the heatmap and inspect it carefully. Check whether the clustering reveals meaningful patterns. Look for unexpected groupings that may indicate technical artifacts.

Save the figure in a vector format for publication. Record the code and parameters used to generate the figure.

Step 7: Document the Parameters

Record all parameters used in the heatmap construction, including the data version, filtering criteria, scaling method, clustering parameters, color scale, and annotation definitions. This documentation enables reproduction of the figure and supports the integrity of the analysis.

The nf-core documentation emphasizes reproducibility standards for community pipelines, and the same principles apply to individual visualization steps. Reproducible figures require documented code and parameters.

At a Glance: Key Decisions for pheatmap Heatmaps

The table below summarizes the primary decisions researchers face when constructing heatmaps with pheatmap, along with recommended defaults and situations where alternatives should be considered.

Decision PointRecommended DefaultWhen to Choose an Alternative
Data inputNormalized log-transformed countsVariance stabilized values when they match the statistical analysis
Gene selectionDifferentially expressed genes or curated gene setMost variable genes for unsupervised exploration
ScalingRow z-scoresNo scaling when values are already comparable across samples
ClusteringHierarchical with Euclidean distance and Ward linkageDisable clustering for time series or defined biological order
Color scaleDiverging blue-white-red for z-scoresPerceptually uniform scales such as viridis for color-blind safety
AnnotationSample metadata including treatment and batchGene annotation for pathway membership or functional categories

Common Failure Patterns and How to Avoid Them

Several recurring problems undermine the usefulness of heatmaps in RNA-seq analysis. Recognizing these patterns helps researchers avoid them in their own figures.

Misleading Color Scales

The most common failure is a color scale that does not match the data distribution. If the limits are too narrow, most cells appear as extreme colors and the figure suggests more uniformity than exists. If the limits are too wide, most cells appear as intermediate colors and the figure hides real differences.

The solution is to examine the distribution of the values being plotted and choose limits that capture the relevant range. For z-scored data, limits around negative two to positive two are often appropriate, but the actual distribution should guide the choice.

Clustering Artifacts

Clustering can produce apparent structure that is not biologically meaningful. This is especially problematic when the number of genes is large relative to the number of samples, because the clustering algorithm can always find patterns in noise.

The solution is to assess cluster stability and to interpret clusters cautiously. Clusters that appear consistently across different parameter choices are more likely to reflect real structure.

Overcrowded Figures

Heatmaps with too many rows or columns become unreadable. Labels overlap, individual cells become invisible, and the figure conveys little information.

The solution is to limit the number of rows to a manageable number, typically fewer than a few hundred for publication figures. If more genes must be shown, consider splitting the figure into multiple panels or providing the full heatmap as supplementary material.

Inconsistent Scaling Across Figures

When multiple heatmaps appear in the same study, they should use consistent scaling and color choices. Inconsistent scales make comparisons across figures impossible and can mislead readers.

The solution is to define a standard set of visualization parameters for the study and apply them consistently across all figures.

Ignoring Batch Effects

Heatmaps can reveal batch effects when samples cluster by technical factors instead of biological ones. Ignoring these patterns can lead to incorrect biological conclusions.

The solution is to include batch information in the sample annotation and to examine whether clustering corresponds to batch. If batch effects are present, they should be addressed in the statistical analysis before visualization.

Quality Control and Reproducibility

Heatmaps serve as quality control tools as well as presentation tools. Examining a heatmap of the most variable genes across all samples can reveal outlier samples, batch effects, or other technical problems.

Using Heatmaps for Sample-Level Quality Control

A heatmap of sample-level correlations or of the top variable genes can show whether biological replicates cluster together and whether technical replicates are more similar to each other than to biological replicates. Samples that cluster unexpectedly may have technical problems that require investigation.

The RAGER platform integrates popular bioinformatics tools in an automated workflow for joint analysis of RNA-seq and ATAC-seq data, demonstrating that quality assessment is a standard component of transcriptomic analysis pipelines. Heatmaps are part of this quality assessment toolkit.

Recording Analysis Parameters

Reproducibility requires that the analysis parameters be recorded. For heatmaps, this includes the data version, the normalization method, the gene filtering criteria, the scaling method, the clustering parameters, and the color scale.

The The Carpentries lessons provide foundational training in reproducible computing practices, including version control and documentation. These practices apply to visualization code as well as to statistical analysis code.

Version Control for Analysis Code

The R script that generates a heatmap should be under version control. This allows the analysis to be reproduced exactly and allows changes to be tracked over time. The script should be self-contained, reading data from files and writing figures to files, so that the entire process can be rerun.

Session Information

Recording the R session information, including package versions, is important because package updates can change default behaviors. The sessionInfo() function in R provides this information and should be saved with the analysis documentation.

Interpretation Limits and Reporting Standards

Heatmaps are visual summaries, not statistical tests. They should be interpreted with appropriate caution and reported with appropriate context.

What a Heatmap Can and Cannot Show

A heatmap can show patterns of expression across genes and samples. It can reveal clusters of co-expressed genes and groups of samples with similar profiles. It cannot establish statistical significance, prove causality, or quantify the strength of associations.

The transcriptomic signatures study in prostate cancer and the diabetes-related gene signature study in breast cancer both use heatmaps as part of larger analyses that include statistical testing and survival modeling. The heatmaps illustrate patterns that are established by other methods.

Reporting Heatmap Parameters in Publications

Publications should report the key parameters used to construct heatmaps. This includes the data transformation, the scaling method, the clustering method, and the color scale. This information allows readers to interpret the figure correctly and to compare it with other figures.

Avoiding Overinterpretation

The visual impression of a heatmap can be compelling, but it should not be overinterpreted. Clusters in a heatmap are the result of algorithmic choices as well as biological signal. The same data can produce different-looking heatmaps with different parameters.

Researchers should be cautious about drawing strong conclusions from visual patterns alone. Statistical analysis should support any claims about differential expression or sample grouping.

Advanced Considerations for Large Datasets

Single-cell RNA-seq datasets present additional challenges for heatmap visualization because of their size. The FIt-SNE method was developed to address the scalability problems of t-SNE for large single-cell datasets and includes a heatmap-style visualization for simultaneously viewing expression patterns from thousands of genes.

Downsampling and Aggregation

For very large datasets, downsampling or aggregation may be necessary. Cells can be grouped by cluster, and the mean expression per cluster can be displayed in a heatmap. This reduces the number of columns to a manageable size while preserving the overall patterns.

Interactive Visualization

Interactive heatmaps allow readers to explore the data by hovering over cells, zooming into regions, and selecting subsets. These tools are useful for exploratory analysis but are less common in static publications.

Memory and Computation Considerations

Large expression matrices require substantial memory. pheatmap may struggle with matrices containing tens of thousands of rows. Filtering genes before visualization reduces the memory requirement and improves performance.

Integration with Differential Expression Analysis

Heatmaps frequently accompany differential expression results. The genes displayed in the heatmap are often the statistically significant genes from a differential expression test.

Selecting Genes from Differential Expression Results

The TCC-GUI application integrates differential expression analysis with visualization tools including heatmap with hierarchical clustering. This integration allows researchers to move from statistical results to visual summaries within a single workflow.

When selecting genes for a heatmap from differential expression results, the thresholds for significance and fold change should be reported. The heatmap then shows the expression patterns of the genes that passed these thresholds.

Coordinating Heatmap Colors with Statistical Results

The color scale in a heatmap should be consistent with the statistical results. If the differential expression analysis uses log2 fold changes, the heatmap should use a scale that reflects this. If the analysis uses z-scores, the heatmap should use a diverging scale centered at zero.

Presenting Heatmaps Alongside Other Visualizations

Heatmaps are often presented alongside volcano plots, MA plots, or principal component analysis plots. These visualizations complement each other, with the heatmap showing detailed gene-level patterns and the other plots showing overall distributions or sample relationships.

Case Examples from Published Studies

Published studies demonstrate the range of heatmap applications in RNA-seq analysis.

Whole-Blood Transcriptomics in Renal Cell Carcinoma

The study of immune-related adverse events in metastatic renal cell carcinoma used whole-blood transcriptomic analysis to characterize immune pathways associated with treatment outcomes. The study identified candidate genes that may contribute to the immune features associated with adverse event development. Heatmaps in such studies typically display the expression of key genes across patient groups, with clustering revealing coordinated patterns.

Glioblastoma and Drug Resistance

The study of IL-37 in glioblastoma used RNA sequencing to reveal altered transcriptomic profiles characterized by activation of MAPK pathways and senescence-related pathways. Heatmaps displaying pathway-related genes can show how the expression of these genes changes under different treatment conditions.

Iron Availability and Cell Fate

The multi-omic study of iron availability examined how nutrient availability affects transcriptional programs in hepatocytes. The study found that iron availability influences the fidelity of hepatocyte cell fate. Heatmaps displaying the affected transcriptional programs would show coordinated changes in gene expression under different culture conditions.

Wheat Cultivar Comparison

The comparison of transcriptome profiles between two wheat cultivars with different antioxidant activities demonstrates the application of RNA-seq to non-model organisms. Heatmaps in such studies display cultivar-specific expression patterns that may relate to the phenotypic differences.

These examples illustrate that heatmaps are used across diverse biological contexts. The principles of scaling, clustering, and color choice apply regardless of the organism or the specific research question.

Professional Escalation Criteria

Some situations require consultation with a bioinformatics specialist or statistician instead of continued self-directed troubleshooting.

Persistent Batch Effects

If samples consistently cluster by batch or by technical factors despite normalization, the problem may require specialized batch correction methods. A bioinformatics specialist can advise on appropriate approaches.

Unexpected Clustering Patterns

If the clustering in a heatmap reveals patterns that contradict the experimental design in unexpected ways, the data may have technical problems. A specialist can help determine whether the pattern reflects a real biological signal or a technical artifact.

Reproducibility Failures

If the heatmap cannot be reproduced from the documented code and parameters, the analysis pipeline may have undocumented dependencies or version issues. A specialist can help establish a reproducible workflow.

Statistical Questions

If the interpretation of a heatmap requires statistical methods beyond the researcher's expertise, such as assessing cluster significance or comparing clustering results, a statistician should be consulted.

Frequently Asked Questions

What is the difference between scaling by row and scaling by column in pheatmap?

Row scaling normalizes each gene across samples, so each row has a mean of zero and a standard deviation of one. This makes genes with different absolute expression levels comparable and is the standard choice for gene expression heatmaps. Column scaling normalizes each sample across genes, which can be useful for comparing sample-level expression profiles but is less common for gene-focused visualizations. The choice depends on whether the research question concerns gene patterns across samples or sample patterns across genes.

How do I choose the number of genes to display in a heatmap?

The number of genes should be small enough that individual patterns are visible and labels are readable. For publication figures, a few hundred genes is often the practical maximum. The genes should be selected according to the biological question, such as differentially expressed genes, genes in a pathway of interest, or the most variable genes across samples. If more genes must be shown, consider splitting the figure into multiple panels or providing the full heatmap as supplementary material.

Why does my heatmap look different when I change the clustering method?

Different clustering methods define similarity and cluster merging differently, so they can produce different orderings of rows and columns. This is expected behavior. If the major clusters remain consistent across different methods, the structure is likely robust. If the clusters change dramatically with method choice, the apparent structure may not be reliable and should be interpreted cautiously.

What color scale should I use for a gene expression heatmap?

A diverging color scale is appropriate for z-scored data, where a central color represents the midpoint of zero and two different colors represent values above and below the midpoint. The classic blue-white-red scale works well for many readers, but red-green scales should be avoided because of color blindness. Modern perceptually uniform scales such as viridis are good alternatives. The limits and breaks should match the data distribution.

Should I cluster both rows and columns in my heatmap?

Clustering both rows and columns is the default in pheatmap and is appropriate when the goal is to reveal coordinated patterns between gene groups and sample groups. However, clustering should be disabled for one dimension when the order is meaningful, such as time points, dose levels, or genomic position. The choice depends on the biological question and the structure of the data.

How do I add annotations to my pheatmap?

pheatmap accepts annotation data frames that are displayed as colored bars above the heatmap for columns and beside the heatmap for rows. The annotation data frame should have row names that match the column names or row names of the expression matrix. The annotation colors should be chosen to be distinguishable and consistent across figures in the same study.

What should I do if my samples cluster by batch instead of by treatment?

Samples clustering by batch indicates a batch effect that may confound the biological signal. The batch information should be included in the sample annotation so that the pattern is visible. The batch effect should be addressed in the statistical analysis, potentially using batch correction methods or including batch as a covariate in the model. A bioinformatics specialist can advise on the appropriate approach for the specific data.

How do I make my heatmap reproducible?

Save the R script that generates the heatmap under version control. The script should read data from files and write figures to files so that the entire process can be rerun. Record the R session information, including package versions. Document all parameters used in the heatmap construction, including the data version, filtering criteria, scaling method, clustering parameters, and color scale.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.