t-SNE vs. UMAP for Single-Cell Visualization: A Comparative Guide to Choosing the Right Algorithm

By Dr. Zubair Khalid, DVM, MS, PhD ·

t-SNE vs. UMAP for Single-Cell Visualization: A Comparative Guide to Choosing the Right Algorithm

Key Takeaways

  • UMAP generally outperforms t-SNE in single-cell visualization by preserving more global structure and exhibiting greater stability across random initializations, making it the preferred choice for exploratory analysis of large datasets.
  • Both t-SNE and UMAP are sensitive to hyperparameter choices (perplexity for t-SNE, n_neighbors and min_dist for UMAP), which critically influence local versus global structure preservation and can lead to misinterpretation of biological relationships if not carefully optimized and validated.
  • The input representation significantly impacts embedding quality; applying t-SNE or UMAP to a pre-processed, PCA-reduced latent space derived from normalized and feature-selected scRNA-seq data yields more robust and biologically meaningful visualizations than direct application to raw, sparse expression matrices.
  • Quantitative validation methods, such as scDEED, are essential for assessing embedding reliability by comparing neighbors in the low-dimensional space to those in the pre-embedding space, thereby identifying and minimizing dubious cell embeddings.
  • UMAP's faster runtime and better scalability to millions of cells, compared to t-SNE's poor scaling beyond ~100,000 cells, enable more extensive parameter exploration and iterative refinement of visualizations in large-scale single-cell projects.
  • Overinterpretation of cluster separation, fragmentation of continuous biological processes, and batch effects masquerading as biology are common failure patterns in single-cell embeddings that necessitate rigorous validation and appropriate preprocessing steps like batch correction.

Researchers analyzing single-cell RNA sequencing data face a practical decision early in their workflow: which nonlinear dimensionality reduction method should produce the two-dimensional embedding that drives cluster interpretation and biological discovery. The two dominant choices, t-distributed stochastic neighbor embedding (t-SNE) and uniform manifold approximation and projection (UMAP), both place cells in a low-dimensional space for visualization, but they rest on different mathematical foundations, respond differently to hyperparameter changes, and produce embeddings with different structural meanings. This article compares the algorithmic principles, practical trade-offs in speed and structure preservation, parameter sensitivity, and interpretability of t-SNE and UMAP for typical single-cell datasets, with concrete guidance for choosing between them and for assessing whether an embedding is trustworthy.

The decision matters because the embedding is also a picture. It shapes how clusters are perceived, which marker genes are explored, and what biological conclusions are drawn. A visualization that distorts global relationships can lead a researcher to overinterpret separation between cell types or miss meaningful continuities. Understanding what each algorithm optimizes, how it handles the sparsity and noise of single-cell data, and how to validate the resulting embedding is essential for responsible analysis.

The Role of Dimensionality Reduction in Single-Cell Analysis

Single-cell RNA sequencing (scRNA-seq) produces expression measurements for thousands of genes across thousands to millions of individual cells. The resulting data matrix is high-dimensional, sparse, and noisy. Most genes are not detected in most cells, technical dropout creates zeros that do not reflect true absence of expression, and batch effects can introduce systematic variation unrelated to biology. Directly visualizing this raw matrix is impossible, and even principal component analysis (PCA), which captures global variance, often fails to separate subtle cell populations that differ in complex, nonlinear ways.

Nonlinear dimensionality reduction methods address this limitation by learning a low-dimensional representation that preserves certain properties of the high-dimensional data. For single-cell analysis, the goal is typically to place cells with similar expression profiles near each other in the embedding, allowing clusters of cells to become visually apparent. Both t-SNE and UMAP accomplish this goal, but they define similarity and preserve structure differently.

The standard workflow in single-cell analysis proceeds through quality control, normalization, feature selection, PCA, and then nonlinear embedding. Quality control typically involves filtering cells based on the number of detected genes and the proportion of mitochondrial reads, as described in the unified pipeline used in the scInfoMaxVAE study, which applied cell- and gene-level quality control with minimum detected gene thresholds before library-size normalization and log transformation [<a href="#ref-1">1</a>]. After normalization, highly variable genes are selected, and PCA reduces the dimensionality to a tractable number of principal components, typically between 10 and 50. The nonlinear embedding method then operates on this PCA-reduced representation instead of on the raw gene expression matrix.

This preprocessing context matters for comparing t-SNE and UMAP. Both methods are sensitive to the number of PCA components used as input, and both perform differently when applied directly to raw, sparse expression data versus a denoised latent representation. A comparative study of autoencoder-based dimensionality reduction found that direct application of t-SNE or UMAP to the raw, sparse expression matrix often yields unstable, poorly separated clusters, while applying these methods to a compact latent representation learned by an autoencoder produces more robust visualizations with enhanced cluster consistency [<a href="#ref-2">2</a>]. This finding underscores that the choice of embedding method cannot be separated from the choice of input representation.

Algorithmic Foundations of t-SNE

t-SNE, introduced as an improvement over stochastic neighbor embedding, constructs a probability distribution over pairs of high-dimensional points such that similar points have a high probability of being neighbors. In the high-dimensional space, the similarity between two points is modeled as a Gaussian distribution centered on each point, with a perplexity parameter controlling the effective number of neighbors. In the low-dimensional embedding, the similarity is modeled using a Student t-distribution with one degree of freedom, which has heavier tails than a Gaussian. The algorithm minimizes the Kullback-Leibler divergence between the high-dimensional and low-dimensional probability distributions through gradient descent.

The heavy-tailed t-distribution in the low-dimensional space addresses the crowding problem, where moderate distances in high-dimensional space become too large in low-dimensional space. By allowing moderate distances in the embedding to be represented as larger distances, t-SNE creates more space for local structure. However, this choice has a consequence: the cost function is dominated by local similarities, and the global arrangement of clusters is largely determined by the random initialization and the optimization trajectory instead of by any explicit preservation of global distances.

The perplexity parameter is the most influential hyperparameter in t-SNE. It can be interpreted as a smooth measure of the effective number of neighbors for each point. Low perplexity values emphasize very local structure, potentially fragmenting clusters, while high perplexity values incorporate more global information but can obscure fine-grained distinctions. The scDEED study demonstrated that the hyperparameter setting of the embedding method has a major impact on the reliability of the resulting two-dimensional visualization, and the authors developed a statistical approach to detect dubious cell embeddings and optimize hyperparameters for both t-SNE and UMAP [<a href="#ref-3">3</a>].

t-SNE is stochastic. Each run with a different random seed can produce a different embedding, particularly in the global arrangement of clusters. The relative positions of distant clusters in a t-SNE plot should not be interpreted as meaningful distances, because the algorithm does not preserve global structure. This limitation is well recognized in the field and has motivated the development of alternative methods.

Algorithmic Foundations of UMAP

UMAP is built on the mathematical framework of manifold learning and topological data analysis. The algorithm constructs a fuzzy simplicial complex representation of the high-dimensional data, where each point is connected to its nearest neighbors with weights that decay with distance. It then optimizes the low-dimensional representation to minimize the cross-entropy between the high-dimensional fuzzy graph and the low-dimensional fuzzy graph.

The key difference from t-SNE lies in the objective function. UMAP uses cross-entropy instead of Kullback-Leibler divergence, which means it balances the preservation of local structure with the preservation of global structure. The repulsive forces in UMAP are designed to push apart points that are not connected in the high-dimensional graph, which helps maintain some global arrangement. As a result, UMAP embeddings tend to preserve more of the global topology of the data, and the relative positions of clusters are more meaningful than in t-SNE.

UMAP has two primary hyperparameters: n_neighbors, which controls the size of the local neighborhood used to construct the high-dimensional graph, and min_dist, which controls how tightly points are allowed to pack together in the low-dimensional embedding. The n_neighbors parameter plays a role analogous to perplexity in t-SNE, balancing local versus global structure. Smaller values emphasize local detail, while larger values capture broader relationships. The min_dist parameter affects the visual appearance of clusters, with smaller values producing tighter, more separated clusters and larger values producing more diffuse arrangements.

UMAP is also stochastic but is generally considered more stable than t-SNE across random initializations. The optimization is typically faster, and the algorithm scales better to large datasets. A comparative study using autoencoder-derived latent spaces found that UMAP outperformed t-SNE in cluster cohesion, global structure preservation, robustness to initialization and data perturbation, and computational cost [<a href="#ref-2">2</a>]. These practical advantages have made UMAP the default choice in many single-cell analysis pipelines.

At a Glance: t-SNE versus UMAP

The following table summarizes the key differences between t-SNE and UMAP for single-cell visualization. These comparisons reflect typical behavior on scRNA-seq datasets and should be interpreted with the understanding that specific results depend on the dataset and hyperparameter choices.

Featuret-SNEUMAP
Mathematical basisProbability distribution matching with Kullback-Leibler divergenceFuzzy topological graph with cross-entropy optimization
Primary hyperparametersPerplexity, learning rate, iterationsn_neighbors, min_dist, learning rate
Local structure preservationStrong, with emphasis on nearest neighborsStrong, with tunable neighborhood size
Global structure preservationWeak, cluster positions are not meaningfulModerate, some global topology is retained
Runtime on large datasetsSlower, scales poorly beyond ~100,000 cellsFaster, scales better to millions of cells
Stability across runsVariable, global arrangement changes with random seedMore stable, but still stochastic
Typical use in single-cellLegacy analyses, small datasets, confirmatory visualizationDefault choice for exploratory analysis of large datasets
Interpretability of distancesInter-cluster distances not interpretableInter-cluster distances partially interpretable

The choice between t-SNE and UMAP depends on the specific goals of the analysis. For exploratory analysis of large single-cell datasets, UMAP is generally preferred due to its speed, scalability, and better preservation of global structure. For small datasets where the researcher wants to compare results with published t-SNE embeddings, or for confirmatory visualization alongside other methods, t-SNE remains a valid choice.

Practical Workflow for Embedding Single-Cell Data

A reproducible embedding workflow requires attention to each stage of the pipeline. The following steps outline a practical approach for generating and validating t-SNE or UMAP embeddings of single-cell data.

Step 1: Quality Control and Preprocessing

Before any embedding is computed, the data must pass through quality control. Filter cells based on the number of detected genes, the total number of unique molecular identifiers (UMIs), and the proportion of mitochondrial reads. The scInfoMaxVAE study applied cell- and gene-level quality control with minimum detected gene thresholds as part of its unified pipeline [<a href="#ref-1">1</a>]. A typical glioma scRNA-seq analysis reported median gene counts of 5,437 and UMI counts of 14,207 per cell after quality control, with 4,753 highly variable genes identified for downstream analysis [<a href="#ref-4">4</a>]. These numbers provide a reference point for expected data quality, though thresholds should be adjusted based on tissue type and protocol.

Step 2: Normalization and Feature Selection

Normalize the expression matrix to account for differences in sequencing depth across cells. Library-size normalization followed by log transformation is a standard approach [<a href="#ref-1">1</a>]. Identify highly variable genes, typically using the variance-stabilizing transformation or the dispersion-based method implemented in Seurat. The number of highly variable genes typically ranges from 2,000 to 5,000 for scRNA-seq datasets.

Step 3: PCA for Initial Dimensionality Reduction

Run PCA on the scaled expression matrix of highly variable genes. The number of principal components to retain is often determined by inspecting the elbow in the variance explained curve or by using a statistical test such as JackStraw. For most datasets, retaining between 10 and 50 principal components is appropriate. The PCA-reduced representation serves as the input to t-SNE or UMAP.

Step 4: Compute the Embedding

For UMAP, set the n_neighbors parameter based on the expected cluster structure. A value between 5 and 50 is typical, with smaller values for datasets with many rare populations and larger values for datasets with broad, continuous transitions. Set min_dist between 0.01 and 0.5, with smaller values for tighter clusters. For t-SNE, set perplexity between 5 and 50, with the caveat that the algorithm is sensitive to this parameter and the optimal value depends on the dataset size and structure [<a href="#ref-3">3</a>].

Step 5: Validate the Embedding

The reliability of a two-dimensional embedding cannot be assumed. The scDEED method provides a statistical approach for detecting dubious cell embeddings by comparing the neighbors of each cell in the two-dimensional embedding with its neighbors in the pre-embedding space [<a href="#ref-3">3</a>]. Cells with low reliability scores are identified as dubious, and the method can guide hyperparameter optimization by minimizing the number of dubious embeddings. Applying such validation is particularly important when the embedding will be used to draw biological conclusions.

Step 6: Annotate and Interpret

Cluster the cells in the embedding using a graph-based method such as the Louvain algorithm, which operates on the shared nearest neighbor graph instead of on the two-dimensional coordinates directly. Annotate clusters based on canonical marker genes. For example, a glioma study identified seven major populations including astrocytes, oligodendrocytes, microglia, neural stem cells, OPC/immature neurons, pericytes, and T cells using marker-based annotation [<a href="#ref-4">4</a>]. The embedding provides the visual context for these annotations, but the clustering should be performed in the high-dimensional space.

Parameter Sensitivity and Its Consequences

Both t-SNE and UMAP are sensitive to their hyperparameters, and the consequences of poor parameter choices differ between the two methods. Understanding this sensitivity is critical for producing trustworthy visualizations.

Perplexity in t-SNE

The perplexity parameter in t-SNE controls the effective number of neighbors considered for each point. Low perplexity values, such as 5 or 10, emphasize very local structure and can cause clusters to fragment into multiple pieces. High perplexity values, such as 50 or 100, incorporate more global information but can obscure rare populations and create artificial bridges between clusters. The scDEED study demonstrated that the hyperparameter setting of t-SNE has a major impact on the reliability of the resulting embedding, and the authors recommend optimizing perplexity by minimizing the number of dubious cell embeddings [<a href="#ref-3">3</a>].

For datasets with thousands of cells, a perplexity between 15 and 50 is commonly used. For larger datasets, higher perplexity values may be necessary, but the computational cost increases. The stochastic nature of t-SNE means that different perplexity values can produce qualitatively different global arrangements, and there is no universally optimal value.

n_neighbors and min_dist in UMAP

The n_neighbors parameter in UMAP controls the size of the local neighborhood used to construct the high-dimensional fuzzy graph. Small values, such as 5, emphasize local structure and can reveal fine-grained subpopulations but may fragment continuous transitions. Large values, such as 50 or 100, capture broader relationships and produce more connected embeddings but can obscure rare populations.

The min_dist parameter controls the minimum distance between points in the low-dimensional embedding. Small values, such as 0.01, produce tight, well-separated clusters that are visually appealing but may exaggerate separation. Large values, such as 0.5, produce more diffuse embeddings that better represent continuous transitions but may obscure cluster boundaries.

The interaction between n_neighbors and min_dist is important. A small n_neighbors with a small min_dist produces a highly fragmented embedding with many small clusters. A large n_neighbors with a large min_dist produces a diffuse embedding with few distinct clusters. The optimal combination depends on the biological question and the expected structure of the data.

Consequences of Poor Parameter Choices

Poor parameter choices can lead to embeddings that misrepresent the underlying biology. A t-SNE embedding with too low perplexity may split a homogeneous cell population into artificial subclusters, leading to spurious marker gene analyses. A UMAP embedding with too small n_neighbors may fragment a continuous developmental trajectory, obscuring the progression from one cell state to another. Conversely, too large a neighborhood may merge distinct populations, hiding meaningful heterogeneity.

The scDEED approach addresses this problem by providing a quantitative measure of embedding reliability. By calculating a reliability score for every cell embedding based on the similarity between the cell's two-dimensional embedding neighbors and its pre-embedding neighbors, the method identifies cells whose embedding positions are not trustworthy [<a href="#ref-3">3</a>]. This validation step should be a routine part of any single-cell analysis workflow.

Speed and Scalability Considerations

The computational cost of t-SNE and UMAP differs substantially, and this difference becomes more pronounced as dataset size increases. For typical single-cell datasets ranging from tens of thousands to millions of cells, the choice of algorithm can affect also the time required for analysis but also the feasibility of iterative exploration.

Runtime Benchmarks

UMAP is generally faster than t-SNE for datasets of comparable size. The comparative study using autoencoder-derived latent spaces reported that UMAP achieved lower computational cost than t-SNE while producing better cluster cohesion and structural preservation [<a href="#ref-2">2</a>]. The speed advantage of UMAP stems from its optimization strategy, which uses stochastic gradient descent and negative sampling, and from its ability to leverage approximate nearest neighbor search algorithms.

For datasets with 100,000 cells, t-SNE can take hours to converge, while UMAP typically completes in minutes. For datasets with more than one million cells, t-SNE becomes impractical on standard hardware, while UMAP remains feasible with appropriate settings. This scalability difference has driven the widespread adoption of UMAP in large-scale single-cell projects, including the Human Cell Atlas and other consortium efforts.

Memory Usage

Both methods require storing the high-dimensional representation of the data, but their memory footprints differ. t-SNE requires computing pairwise similarities, which can be memory-intensive for large datasets. UMAP constructs a sparse nearest neighbor graph, which is more memory-efficient. For datasets that exceed available memory, both methods offer approximate implementations, but UMAP's approximate approach is generally more robust.

Practical Implications

The speed advantage of UMAP enables more extensive parameter exploration. A researcher can compute multiple UMAP embeddings with different n_neighbors and min_dist values in the time it would take to compute a single t-SNE embedding. This capability supports the validation approach recommended by scDEED, which involves optimizing hyperparameters by minimizing the number of dubious cell embeddings [<a href="#ref-3">3</a>].

However, speed should not be the sole criterion for method selection. For small datasets where t-SNE is computationally feasible, the choice should be based on the biological question and the interpretability of the resulting embedding. The comparative study of autoencoder-based pipelines found that UMAP outperformed t-SNE in multiple metrics, but both methods produced meaningful visualizations when applied to a well-constructed latent space [<a href="#ref-2">2</a>].

Global Structure Preservation and Interpretability

The most consequential difference between t-SNE and UMAP lies in how they handle global structure. This difference affects also the visual appearance of the embedding but also the biological interpretations that can be drawn from it.

t-SNE and the Limits of Local Structure

t-SNE is explicitly designed to preserve local structure. The Kullback-Leibler divergence objective penalizes errors in local neighborhoods heavily while largely ignoring the global arrangement of distant points. As a result, the relative positions of clusters in a t-SNE plot are not meaningful. Two clusters that appear far apart in the embedding may be close in the high-dimensional space, and two clusters that appear adjacent may be distant.

This property has important consequences for interpretation. Distances between clusters in a t-SNE plot should never be used to infer similarity or developmental relationships. The size of a cluster in a t-SNE plot is also not meaningful, because the algorithm tends to produce clusters of roughly equal size regardless of the actual number of cells in each population. A rare population may appear as a small, tight cluster, but a common population may also appear small if it is homogeneous.

The scDEED study emphasized that t-SNE and UMAP embeddings might not reliably inform the similarities among cell clusters [<a href="#ref-3">3</a>]. This caution applies with particular force to t-SNE, where the global arrangement is essentially arbitrary. Researchers should treat t-SNE plots as visual aids for identifying candidate populations, not as quantitative representations of biological relationships.

UMAP and Partial Global Structure

UMAP's cross-entropy objective balances local and global structure preservation. The algorithm explicitly attempts to arrange clusters in a way that reflects the global topology of the high-dimensional data. As a result, the relative positions of clusters in a UMAP embedding are more meaningful than in t-SNE, and distances between clusters provide some information about their similarity.

However, UMAP does not preserve global distances exactly. The min_dist parameter introduces a floor on inter-point distances, and the optimization prioritizes local structure. The global arrangement is approximate, and the degree of global structure preservation depends on the n_neighbors parameter. Larger values of n_neighbors capture more global structure but may obscure local detail.

The practical implication is that UMAP embeddings can support certain types of biological interpretation that t-SNE embeddings cannot. For example, if a researcher observes a continuous gradient of cells in a UMAP embedding, this gradient may reflect a genuine developmental trajectory or a continuous transition between cell states. In a t-SNE embedding, such a gradient would be more difficult to interpret, because the global arrangement is not meaningful.

The Local-Global Trade-off

The fundamental challenge in dimensionality reduction for single-cell data is the local-global trade-off. Methods optimized for local neighborhood preservation distort global topology, while those emphasizing global coherence obscure fine-grained cell states. This trade-off is explicitly recognized in the development of newer methods, such as the Lorentz-regularized variational autoencoder (LiVAE), which was designed to balance local fidelity with global coherence [<a href="#ref-5">5</a>]. The same trade-off applies to the choice between t-SNE and UMAP, and researchers should be aware of what each method sacrifices.

For most single-cell analyses, the primary goal is to identify and characterize cell populations, which requires strong local structure preservation. Both t-SNE and UMAP achieve this goal, but UMAP provides the additional benefit of partial global structure preservation at no cost to local fidelity. This combination makes UMAP the more interpretable choice for most applications.

Embedding Quality Assessment and Validation

The reliability of a two-dimensional embedding cannot be assumed, and several approaches exist for assessing embedding quality. These approaches range from visual inspection to quantitative validation methods.

Visual Inspection Criteria

A trustworthy embedding should show clear separation between known cell populations and continuity along known developmental trajectories. Clusters should be compact and well separated, with minimal fragmentation of homogeneous populations. The embedding should be stable across multiple runs with different random seeds, and the results should be consistent with known biology.

However, visual inspection has limitations. The human eye is adept at finding patterns, including patterns that are not biologically meaningful. An embedding that looks clean and well separated may still misrepresent the underlying data, particularly if the hyperparameters were chosen to produce an appealing visualization instead of a faithful one.

Quantitative Validation with scDEED

The scDEED method provides a statistical approach for detecting dubious cell embeddings. The method calculates a reliability score for every cell embedding based on the similarity between the cell's two-dimensional embedding neighbors and its pre-embedding neighbors. Cells with low reliability scores are identified as dubious, and the method can guide hyperparameter optimization by minimizing the number of dubious cell embeddings [<a href="#ref-3">3</a>].

This approach is particularly valuable for comparing t-SNE and UMAP on a specific dataset. By computing the proportion of dubious cell embeddings for each method across a range of hyperparameter values, a researcher can determine which method and which parameter settings produce the most trustworthy visualization for their data.

Cross-Method Consistency

Another validation approach is to compare embeddings produced by different methods. If t-SNE and UMAP produce broadly consistent cluster structures, the results are more likely to reflect genuine biology. If the methods disagree substantially, the researcher should investigate the source of the discrepancy, which may indicate that the data do not support clear cluster structure or that the hyperparameters are poorly chosen.

The scInfoMaxVAE study compared its method against t-SNE and UMAP embeddings and reported that the variational autoencoder approach achieved competitive clustering and structure preservation across all datasets [<a href="#ref-1">1</a>]. This comparison highlights that t-SNE and UMAP are not the only options for single-cell visualization, and that alternative methods may provide better performance for specific datasets.

Reproducibility Considerations

Both t-SNE and UMAP are stochastic, meaning that different runs with different random seeds produce different embeddings. This stochasticity has implications for reproducibility. A researcher who reports a t-SNE or UMAP embedding should specify the random seed, the hyperparameters, and the software version used to generate the embedding. The nf-core documentation emphasizes the importance of reproducibility in bioinformatics workflows, and this principle applies to dimensionality reduction as much as to any other analysis step [<a href="#ref-6">6</a>].

For publications, it is good practice to report the embedding parameters in the methods section and to make the code available. The Galaxy Training Network provides tutorials on reproducible single-cell analysis workflows, and the Bioconductor project offers packages for single-cell analysis with documented best practices [<a href="#ref-7">7</a>][<a href="#ref-8">8</a>].

Common Failure Patterns in Single-Cell Embeddings

Several recurring problems can compromise the quality of t-SNE and UMAP embeddings. Recognizing these failure patterns is the first step toward avoiding them.

Overinterpretation of Cluster Separation

The most common failure pattern is overinterpreting the visual separation between clusters in an embedding. Both t-SNE and UMAP can produce visually distinct clusters even when the underlying populations are not well separated in the high-dimensional space. The min_dist parameter in UMAP, in particular, can create artificial separation by forcing points apart in the low-dimensional embedding.

Researchers should validate cluster separation using quantitative methods, such as differential expression analysis between clusters or silhouette scores computed in the high-dimensional space. A cluster that appears distinct in the embedding but lacks differentially expressed marker genes may be an artifact of the embedding instead of a genuine biological population.

Fragmentation of Continuous Populations

Both t-SNE and UMAP can fragment continuous populations into multiple clusters. This problem is particularly acute for developmental trajectories, where cells transition gradually from one state to another. Low perplexity in t-SNE and small n_neighbors in UMAP can cause a continuous trajectory to appear as a series of discrete clusters, leading to spurious interpretations of discrete cell states.

The LAIOR study noted that existing approaches typically achieve only one of two goals: classical methods emphasize either local neighborhoods or global variance, and deep generative models cluster cell types well but often fracture trajectory continuity [<a href="#ref-9">9</a>]. This observation applies to t-SNE and UMAP as well, and researchers studying developmental processes should be cautious about interpreting cluster boundaries in embeddings.

Batch Effects Masquerading as Biology

Technical variation between samples, sequencing batches, or experimental conditions can dominate the embedding and create clusters that reflect technical artifacts instead of biological differences. Both t-SNE and UMAP will faithfully represent whatever variation is present in the input data, including batch effects.

The LiVAE study emphasized that its method balances local fidelity with global coherence without requiring specialized batch-correction procedures [<a href="#ref-5">5</a>]. For t-SNE and UMAP, batch correction is typically performed before embedding, using methods such as Harmony, Seurat integration, or scVI. Researchers should check whether clusters in the embedding correspond to known batches and should apply batch correction if necessary.

Hyperparameter Overfitting

The ability to tune hyperparameters to produce an appealing visualization carries the risk of overfitting. A researcher may choose perplexity, n_neighbors, or min_dist values that produce the cleanest-looking clusters, but these values may not produce the most faithful representation of the data. The scDEED method addresses this problem by providing a quantitative criterion for hyperparameter selection, minimizing the number of dubious cell embeddings instead of maximizing visual appeal [<a href="#ref-3">3</a>].

Ignoring the Input Representation

The quality of the embedding depends on the quality of the input representation. Applying t-SNE or UMAP directly to the raw, sparse expression matrix often yields unstable, poorly separated clusters [<a href="#ref-2">2</a>]. The choice of preprocessing steps, including normalization, feature selection, and PCA, has a major impact on the resulting embedding. Researchers should not assume that a poor embedding reflects a limitation of t-SNE or UMAP when the input representation may be the problem.

Alternative and Complementary Methods

While t-SNE and UMAP are the most widely used methods for single-cell visualization, they are not the only options. Several alternative methods address specific limitations of these approaches, and researchers should be aware of these options.

Variational Autoencoders

Variational autoencoders (VAEs) provide a probabilistic approach to dimensionality reduction that can capture nonlinear structure while providing a generative model of the data. The scInfoMaxVAE method uses a mutual-information-maximizing variational autoencoder with a zero-inflated count likelihood tailored for scRNA-seq, designed for dimensionality reduction and cell-type classification [<a href="#ref-1">1</a>]. In a comparison against t-SNE and UMAP, scInfoMaxVAE achieved competitive clustering and structure preservation across 12 public datasets, with normalized mutual information of 0.94, matching the VASC method and exceeding t-SNE's 0.66 [<a href="#ref-1">1</a>].

The advantage of VAE-based methods is their ability to model the count distribution of scRNA-seq data explicitly, including zero inflation. The scInfoMaxVAE study attributed its performance to information-theoretic training and explicit modeling of zero inflation [<a href="#ref-1">1</a>]. However, the study also noted limitations, including sensitivity to hyperparameters and modest run-to-run variance, suggesting benefits from automated tuning [<a href="#ref-1">1</a>].

Hyperbolic and Lorentz-Regularized Methods

Hyperbolic geometry provides a natural framework for representing hierarchical structure, which is common in biological data. The LAIOR method combines Lorentz geometric regularization, a dual-path information bottleneck, and neural ODE regularization to learn embeddings that preserve local cell-state structure, global hierarchy, and smooth developmental trajectories [<a href="#ref-9">9</a>]. The LiVAE method applies hyperbolic geometry as soft regularization over standard Euclidean latent spaces, balancing local fidelity with global coherence [<a href="#ref-5">5</a>].

These methods are more complex than t-SNE or UMAP and require more computational resources, but they may provide better representations for datasets with strong hierarchical structure, such as developmental atlases or tumor microenvironments.

Optimal Transport Approaches

Optimal transport (OT) provides a framework for comparing distributions of cellular states, which is useful for longitudinal studies and treatment comparisons. The OT approach described in the immuno-oncology study uses the Sinkhorn algorithm to compute distances between marker expression distributions, providing reproducible measures of change in high-dimensional space [<a href="#ref-10">10</a>]. This approach is complementary to t-SNE and UMAP, which are designed for visualization instead of for quantitative comparison between samples.

When to Use Alternatives

Alternative methods should be considered when t-SNE or UMAP embeddings are unsatisfactory or when the biological question requires properties that these methods do not provide. For example, if a researcher needs to compare cell populations across multiple samples or time points, an OT-based approach may be more appropriate than a single embedding [<a href="#ref-10">10</a>]. If the data have strong hierarchical structure, a hyperbolic method may provide a better representation [<a href="#ref-9">9</a>][<a href="#ref-5">5</a>].

However, alternative methods have their own limitations. The scInfoMaxVAE study noted sensitivity to hyperparameters and run-to-run variance [<a href="#ref-1">1</a>]. The LAIOR study acknowledged that hyperbolic embeddings are numerically fragile in practice [<a href="#ref-9">9</a>]. Researchers should validate any embedding method, whether classical or modern, using the same criteria of local structure preservation, global coherence, and biological interpretability.

Records and Measurements for Embedding Quality

Maintaining records of embedding parameters and quality metrics is essential for reproducible and interpretable single-cell analysis. The following measurements should be recorded for each embedding.

Hyperparameter Records

For t-SNE, record the perplexity, learning rate, number of iterations, and random seed. For UMAP, record the n_neighbors, min_dist, learning rate, and random seed. Also record the software version and the input representation, including the number of PCA components and the number of highly variable genes.

Quality Metrics

Record the proportion of dubious cell embeddings as determined by scDEED or a similar method [<a href="#ref-3">3</a>]. Record the normalized mutual information or adjusted Rand index for clustering results, if available. Record the runtime and memory usage for the embedding computation.

Validation Records

Record the results of cross-method consistency checks, comparing the t-SNE and UMAP embeddings for the same dataset. Record the results of cluster validation, including differential expression analysis and marker gene annotation. Record any batch correction applied before embedding.

Reproducibility Records

Record the exact commands used to generate the embedding, including all parameter values. The nf-core documentation provides guidance on reproducible workflow configuration, and the Bioconductor project offers tools for reproducible genomic analysis [<a href="#ref-7">7</a>][<a href="#ref-6">6</a>]. The Carpentries lessons provide foundational training in version control and reproducible computing practices [<a href="#ref-11">11</a>].

Professional Escalation Criteria

Certain situations warrant consultation with a bioinformatics specialist or computational biologist. The following criteria indicate that the embedding analysis may require expert input.

Persistent Instability Across Runs

If t-SNE or UMAP embeddings vary substantially across runs with different random seeds, the data may not support clear cluster structure, or the hyperparameters may be poorly chosen. A specialist can help diagnose the cause and recommend appropriate settings or alternative methods.

Disagreement Between Methods

If t-SNE and UMAP produce substantially different cluster structures for the same dataset, the discrepancy should be investigated. A specialist can help determine whether the disagreement reflects genuine biological complexity or a technical artifact.

Poor Validation Metrics

If scDEED identifies a high proportion of dubious cell embeddings, or if clustering metrics such as normalized mutual information are low, the embedding may not be trustworthy. A specialist can help optimize hyperparameters or recommend alternative approaches.

Complex Batch Structure

If the data include multiple batches, samples, or experimental conditions with complex technical variation, a specialist can help design an appropriate batch correction strategy before embedding.

Large-Scale or Multi-Omic Data

If the dataset exceeds available computational resources, or if the analysis involves multi-omic data such as combined scRNA-seq and scATAC-seq, a specialist can help select appropriate methods and computational infrastructure.

Frequently Asked Questions

What is the main difference between t-SNE and UMAP for single-cell data?

The main difference lies in how each method handles global structure. t-SNE preserves local structure strongly but does not preserve global distances, so the relative positions of distant clusters are not meaningful. UMAP balances local and global structure preservation, so the arrangement of clusters in the embedding provides some information about their similarity. UMAP is also generally faster and more scalable to large datasets.

Which method should I use for my single-cell dataset?

For most exploratory analyses of large single-cell datasets, UMAP is the recommended default due to its speed, scalability, and better preservation of global structure. t-SNE remains useful for small datasets, for comparison with published results, or when the specific properties of t-SNE are desired. The choice should be validated using quantitative methods such as scDEED.

How do I choose the perplexity parameter for t-SNE?

The perplexity parameter controls the effective number of neighbors for each point. Values between 5 and 50 are commonly used, with smaller values for datasets with many rare populations and larger values for datasets with broad, continuous transitions. The scDEED method provides guidance for optimizing perplexity by minimizing the number of dubious cell embeddings.

How do I choose the n_neighbors and min_dist parameters for UMAP?

The n_neighbors parameter controls the size of the local neighborhood, with values between 5 and 50 being typical. Smaller values emphasize local structure and can reveal fine-grained subpopulations. The min_dist parameter controls how tightly points pack together, with values between 0.01 and 0.5 being typical. Smaller values produce tighter clusters. The optimal combination depends on the biological question and should be validated.

Can I interpret distances between clusters in a t-SNE plot?

No. t-SNE does not preserve global distances, so the relative positions of distant clusters are not meaningful. Distances between clusters in a t-SNE plot should never be used to infer similarity or developmental relationships. In a UMAP plot, inter-cluster distances are partially interpretable, but they should be validated using quantitative methods.

How do I know if my embedding is trustworthy?

Apply a quantitative validation method such as scDEED, which calculates a reliability score for every cell embedding based on the similarity between the cell's two-dimensional embedding neighbors and its pre-embedding neighbors [<a href="#ref-3">3</a>]. Also compare embeddings across multiple runs and across methods, and validate cluster separation using differential expression analysis.

Should I apply t-SNE or UMAP to the raw expression matrix or to a reduced representation?

Apply the embedding method to a reduced representation, such as the PCA-reduced expression matrix or a latent representation learned by an autoencoder. Direct application to the raw, sparse expression matrix often yields unstable, poorly separated clusters [<a href="#ref-2">2</a>]. The choice of input representation has a major impact on the quality of the embedding.

What should I do if my embedding shows clusters that do not correspond to known cell types?

First, validate that the clusters are supported by differential expression analysis and marker gene annotation. If the clusters lack biological support, they may be artifacts of the embedding or of technical variation such as batch effects. Consider adjusting hyperparameters, applying batch correction, or consulting a bioinformatics specialist.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

[1] [Leveraging mutual information in Variational Autoencoders for improved dimensionality reduction of single-cell RNA sequencing data: The scInfoMaxVAE approach.](https://pubmed.ncbi.nlm.nih.gov/40882444). Computational biology and chemistry, 2026. [2] [Autoencoder-Based Nonlinear Dimension Reduction for Single-Cell RNA-Seq Data: A Comparative Study of t-SNE and UMAP](https://doi.org/10.6000/1929-6029.2025.14.78). International Journal of Statistics in Medical Research, 2025. [3] [Statistical method scDEED for detecting dubious 2D single-cell embeddings and optimizing t-SNE and UMAP hyperparameters](https://doi.org/10.1038/s41467-024-45891-y). Nature Communications, 2024. [4] [Single-cell transcriptomic profiling uncovers key molecular signatures in glioma pathogenesis.](https://doi.org/10.3389/fgene.2026.1818742). 2026. [5] [Lorentz-regularized interpretable VAE for multi-scale single-cell transcriptomic and epigenomic embeddings.](https://doi.org/10.3389/fgene.2025.1713727). 2025. [6] [nf-core Documentation](https://nf-co.re/docs). nf-core. [7] [Bioconductor](https://bioconductor.org/). Bioconductor Project. [8] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project. [9] [LAIOR: a hyperbolic neural ODE variational framework for interpretable single-cell manifold learning and trajectory inference.](https://doi.org/10.3389/fgene.2026.1838613). 2026. [10] [Optimal transport analysis of high-dimensional flow cytometry data in immuno-oncology.](https://doi.org/10.3389/fimmu.2026.1856896). 2026. [11] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.