Evaluating the Quality of Dimensionality Reduction in Single-Cell Data: Metrics and Validation Techniques

By Dr. Zubair Khalid, DVM, MS, PhD ·

Evaluating the Quality of Dimensionality Reduction in Single-Cell Data: Metrics and Validation Techniques

Key Takeaways

  • Dimensionality reduction in single-cell RNA sequencing (scRNA-seq) requires objective evaluation beyond visual appeal; unvalidated embeddings can lead to biologically inaccurate interpretations by misrepresenting cell population relationships.
  • Local structure preservation is assessed by trustworthiness (embedding neighbors were original neighbors) and continuity (original neighbors remain embedding neighbors), crucial for identifying distinct cell types and continuous differentiation processes, respectively.
  • Global structure preservation, often better captured by linear methods like PCA, ensures the overall arrangement of cell populations in the embedding reflects their high-dimensional relationships, which nonlinear methods like t-SNE and UMAP may distort.
  • Quantitative metrics such as trustworthiness, continuity, and silhouette score, alongside biological validation via marker gene overlay and hyperparameter sensitivity analysis, are essential for robust decision-making in scRNA-seq data interpretation.
  • Comparing multiple dimensionality reduction methods (e.g., PCA, UMAP, t-SNE) is critical, as different algorithms capture distinct aspects of data structure, and no single method is universally superior across all datasets and biological questions.
  • Overreliance on visual cluster separation, ignoring hyperparameter sensitivity, or using the silhouette score in isolation without biological validation are common failure patterns that can lead to artifact-driven conclusions.

Dimensionality reduction is a standard step in single-cell RNA sequencing (scRNA-seq) analysis, yet most researchers select a method such as t-distributed stochastic neighbor embedding (t-SNE) or uniform manifold approximation and projection (UMAP) without an objective check of whether the resulting embedding faithfully represents the original data structure. This article provides a practical framework for evaluating embedding quality using quantitative metrics including trustworthiness, continuity, and silhouette score, along with validation strategies that can be applied in standard analysis workflows. The target reader is a biology student, researcher, or laboratory professional who has generated or downloaded single-cell data and needs to make defensible decisions about visualization and downstream interpretation.

The Problem of Unvalidated Embeddings

Single-cell transcriptomic datasets contain expression measurements for thousands of genes across tens of thousands of cells. The raw data space is so high-dimensional that direct visualization is impossible, and most downstream analyses such as clustering, trajectory inference, and cell type annotation depend on a reduced representation of the data. Principal component analysis (PCA), t-SNE, UMAP, and variational autoencoders are among the most commonly applied methods, and each makes different assumptions about what structure in the data matters most.

The central problem is that a visually appealing two-dimensional plot is not evidence of biological accuracy. A UMAP plot with well separated clusters may have been produced with hyperparameters that fragmented a continuous differentiation process, while a t-SNE plot may preserve local neighborhoods at the expense of global relationships between cell populations. Comparative studies have shown that no single dimensionality reduction method is consistently superior across datasets and that different techniques can produce discrepant representations of the same underlying biology. A 2025 evaluation of time-series scRNA-seq data found that PCA, t-SNE, UMAP, and single-cell variational inference each captured different dynamical patterns and that none of the methods reliably represented dynamics when used in isolation, underscoring the need to compare multiple perspectives (PubMed 40576031).

The practical consequence is that biological interpretations drawn from unvalidated embeddings may reflect artifacts of the algorithm instead of properties of the cells. A cluster that appears distinct in UMAP may be an artifact of excessive nearest neighbor settings, and a trajectory that appears continuous in t-SNE may be a distortion of what is actually a discrete set of cell states. Quantitative evaluation metrics provide a way to detect these problems before they propagate into downstream analysis and published conclusions.

Core Principles of Embedding Quality

Local Structure Preservation

Local structure preservation refers to whether cells that are neighbors in the high-dimensional expression space remain neighbors in the low-dimensional embedding. This property matters because most biological interpretations of single-cell data rely on the assumption that similar cells are positioned close together. If a dimensionality reduction method disrupts local neighborhoods, then clusters identified in the embedding may not correspond to genuinely similar cell populations.

Trustworthiness and continuity are the two standard metrics for assessing local structure preservation. Trustworthiness measures the extent to which cells that appear as neighbors in the low-dimensional embedding were also neighbors in the original high-dimensional space. A low trustworthiness score indicates that the embedding has placed dissimilar cells close together, which can create false clusters. Continuity measures the reverse direction, specifically whether cells that were neighbors in the high-dimensional space remain neighbors in the embedding. A low continuity score indicates that the embedding has separated cells that were originally similar, which can fragment genuine cell populations.

These two metrics are complementary because a dimensionality reduction method can sacrifice one property to preserve the other. A method that aggressively separates clusters may achieve high continuity but low trustworthiness, while a method that compresses the data may achieve high trustworthiness but low continuity. The appropriate balance depends on the biological question. For cell type identification, trustworthiness may be more important because false mixing of distinct cell types is a serious error. For trajectory analysis, continuity may be more important because fragmentation of a continuous differentiation process obscures the biological signal.

Global Structure Preservation

Global structure preservation refers to whether the overall arrangement of cell populations in the embedding reflects their relationships in the high-dimensional space. This property is distinct from local structure preservation because an embedding can correctly position individual cells relative to their immediate neighbors while distorting the relationships between distant cell populations.

Linear methods such as PCA tend to preserve global structure well because they project the data onto orthogonal axes that capture the largest sources of variance. Nonlinear methods such as t-SNE and UMAP are optimized primarily for local structure and may distort global relationships. A 2021 comparison of dimensionality reduction methods found that UMAP well preserves the original cohesion and separation of cell populations, while t-SNE yielded the best overall performance with the highest accuracy and computing cost (PubMed 33833778). The same study noted that hyperparameters must be set according to the specific situation before using dimensionality reduction methods based on nonlinear models and neural networks.

Global structure is difficult to quantify with a single metric, but several approaches can provide useful information. One approach is to compare the distances between cluster centroids in the embedding with the distances between the same centroids in the high-dimensional space. Another approach is to use a metric such as the silhouette score, which measures how similar a cell is to its own cluster compared with other clusters. A high silhouette score indicates that clusters are well separated in the embedding, but this metric does not directly assess whether the separation reflects the original data structure.

Cluster Separation and Cohesion

Cluster separation and cohesion are properties of the embedding that directly affect downstream clustering and cell type annotation. Separation refers to the degree to which distinct cell populations occupy distinct regions of the embedding, while cohesion refers to the degree to which cells from the same population remain close together.

The silhouette score is the most commonly used metric for assessing cluster quality in an embedding. For each cell, the silhouette score compares the average distance to cells in its own cluster with the average distance to cells in the nearest neighboring cluster. Scores range from negative one to positive one, with values near one indicating that the cell is well embedded within its cluster and values near zero or negative indicating that the cell is close to the boundary between clusters or possibly assigned to the wrong cluster.

The silhouette score has an important limitation in the context of dimensionality reduction evaluation. The score is computed on the embedding, so it measures whether clusters are visually distinct in the reduced representation, not whether those clusters reflect genuine biological differences. A dimensionality reduction method that artificially separates cells can produce high silhouette scores even when the underlying data do not support distinct clusters. The silhouette score should therefore be interpreted alongside local structure preservation metrics and validated against marker gene expression.

At a Glance

Metric or Validation ApproachWhat It AssessesPractical Use in Single-Cell Workflows
TrustworthinessWhether cells placed close together in the embedding were also close in the original high-dimensional spaceRun after any nonlinear embedding to detect false mixing of distinct cell populations
ContinuityWhether cells that were close in the original space remain close in the embeddingRun alongside trustworthiness to detect fragmentation of genuine cell populations
Silhouette scoreWhether clusters in the embedding are well separated and cohesiveCompute after clustering to assess whether cluster boundaries are defensible
Marker gene overlayWhether known cell type markers are expressed in the expected clustersAlways perform as a biological validation of any embedding and clustering result
Hyperparameter sensitivity analysisWhether the embedding changes substantially with different parameter settingsRun for t-SNE perplexity and UMAP n_neighbors to identify unstable configurations
Multiple method comparisonWhether different dimensionality reduction methods produce concordant representationsRun for critical conclusions to ensure findings are not method-specific artifacts

Quantitative Metrics for Embedding Evaluation

Trustworthiness and Continuity

Trustworthiness and continuity are computed by comparing the k nearest neighbors of each cell in the high-dimensional space with the k nearest neighbors in the low-dimensional embedding. The metrics are typically reported as values between zero and one, with higher values indicating better preservation of neighborhood structure.

Trustworthiness is penalized when cells that are neighbors in the embedding were not neighbors in the original space. The magnitude of the penalty depends on the rank of the cell in the high-dimensional neighborhood. A cell that was ranked far away in the original space but appears as a close neighbor in the embedding contributes a larger penalty than a cell that was only slightly outside the original neighborhood.

Continuity is penalized in the reverse situation, when cells that were neighbors in the original space are not neighbors in the embedding. The magnitude of the penalty depends on the rank of the cell in the low-dimensional neighborhood.

Both metrics depend on the choice of k, the number of neighbors considered. A small k focuses on very local structure, while a larger k captures broader neighborhood relationships. In practice, researchers should compute trustworthiness and continuity across a range of k values to understand how the embedding performs at different scales. A method that performs well for small k but poorly for larger k may be preserving only the most local structure while distorting broader relationships.

The 2019 evaluation of 18 dimensionality reduction methods on 30 publicly available scRNA-seq datasets assessed neighborhood preservation in terms of the ability to recover features of the original expression matrix, along with accuracy and robustness for cell clustering and lineage reconstruction (PubMed 31823809). The study provided guidelines for choosing dimensionality reduction methods based on these quantitative comparisons, and the analysis scripts were made publicly available for researchers to apply the same evaluation approach to their own data.

Silhouette Score

The silhouette score is computed after clustering has been performed on the embedding. For each cell, the score is calculated as the difference between the mean distance to cells in the nearest neighboring cluster and the mean distance to cells in the assigned cluster, divided by the maximum of these two values.

A mean silhouette score above 0.5 across all cells is generally interpreted as indicating reasonable cluster structure, while scores below 0.25 suggest that clusters are not well separated. However, these thresholds are heuristic and should be interpreted in the context of the biological system being studied. Some cell populations are genuinely similar to each other, and low silhouette scores may reflect real biological continuity instead of poor embedding quality.

The silhouette score is most useful as a comparative tool. Researchers can compute silhouette scores for the same clustering applied to different embeddings and compare the results. If one embedding produces substantially higher silhouette scores than another, this provides evidence that the first embedding better supports the clustering structure. The score can also be used to compare different clustering resolutions on the same embedding to identify the resolution that produces the most defensible cluster boundaries.

Neighborhood Preservation Metrics

Beyond trustworthiness and continuity, several related metrics can be used to assess neighborhood preservation. The fraction of shared nearest neighbors between the high-dimensional space and the embedding provides a measure of how well the local neighborhood structure is maintained. The Pearson correlation between pairwise distances in the original space and pairwise distances in the embedding provides a global measure of distance preservation, although this metric is sensitive to the scale differences between the original and reduced spaces.

The choice of distance metric in the original space affects all neighborhood-based evaluations. For scRNA-seq data, Euclidean distance on PCA-reduced data is a common choice because PCA removes some noise and reduces the impact of dropout events. Alternative distance metrics such as correlation-based distances may be more appropriate for certain biological questions. Researchers should document the distance metric used for evaluation and consider whether the results are robust to this choice.

Validation Strategies Beyond Single Metrics

Marker Gene Overlay

The most direct biological validation of an embedding is to overlay known marker gene expression onto the reduced representation. If the embedding accurately represents the underlying biology, then cells expressing markers for a known cell type should occupy a contiguous region of the embedding, and different cell types should occupy distinct regions.

Marker gene overlay serves as a ground truth check that does not depend on any assumptions about the dimensionality reduction algorithm. If a researcher has prior knowledge that a dataset contains T cells, B cells, and macrophages, then the embedding should separate these populations in a way that is consistent with their marker expression. Failure to observe this separation indicates either that the dimensionality reduction has distorted the data or that the cell populations are not as distinct as expected.

This validation approach is widely used in published single-cell studies. A 2023 analysis of pulmonary arterial hypertension used UMAP for dimensionality reduction and cluster identification, followed by the SingleR package for cell annotation based on marker gene expression (PubMed 38110868). A 2025 study of thyroid cancer employed PCA and UMAP for dimensionality reduction and subsequent identification of cellular clusters, with differential gene expression analysis across subclusters to validate the biological identity of each cluster (PubMed 41298873). These examples illustrate the standard practice of combining computational dimensionality reduction with biological validation.

Hyperparameter Sensitivity Analysis

Dimensionality reduction methods have hyperparameters that substantially affect the resulting embedding. For t-SNE, the perplexity parameter controls the balance between local and global aspects of the data. For UMAP, the n_neighbors parameter controls the size of the local neighborhood considered when constructing the manifold. For variational autoencoders, the latent dimensionality and network architecture affect the representation.

A hyperparameter sensitivity analysis involves running the same dimensionality reduction method with different parameter values and comparing the resulting embeddings. If the embeddings are broadly similar across a range of parameter values, the results are likely robust. If the embeddings change dramatically with small parameter changes, the results should be interpreted with caution.

The 2021 comparison of dimensionality reduction methods found that users need to set hyperparameters according to the specific situation before using methods based on nonlinear models and neural networks (PubMed 33833778). The 2019 evaluation of deep variational autoencoders for scRNA-seq data similarly emphasized that parameter tuning is a key part of dimensionality reduction with these methods (Semantic Scholar 50a9a2534e22188678d2804d8005a8919cdd387c).

A practical approach is to run UMAP with n_neighbors values of 5, 15, and 30 and t-SNE with perplexity values of 5, 15, and 30. The resulting embeddings can be compared visually and with quantitative metrics. If the same cell populations appear as distinct clusters across all parameter settings, the clustering result is likely robust. If cluster boundaries shift substantially, the researcher should investigate whether the clusters represent genuine biology or parameter-dependent artifacts.

Multiple Method Comparison

Comparing embeddings produced by different dimensionality reduction methods provides a powerful validation strategy. If PCA, t-SNE, UMAP, and a variational autoencoder all place the same cell populations in similar relative positions, the representation is likely capturing genuine structure in the data. If the methods produce conflicting representations, the researcher must determine which method is most appropriate for the specific biological question.

The 2025 study of time-series scRNA-seq data explicitly compared manifolds from PCA, t-SNE, UMAP, and single-cell variational inference and found that none of the techniques was consistently superior and that they may not reliably represent dynamics when used in isolation (PubMed 40576031). The study proposed a synthetic dynamical pattern approach based on variational autoencoders to reason about discrepancies between methods, providing a foundation for guiding future methods development.

A 2026 study of dendritic spine morphology compared PCA, ISOMAP, t-SNE, UMAP, and PCUMAP and found that the optimal dimensionality reduction strategy was dataset-dependent (PLOS ONE 10.1371/journal.pone.0349775). On the primary dataset, nonlinear approaches better preserved fine-scale structure, while PCA was more robust under increased feature-level noise in a lower-resolution secondary dataset. This finding underscores the importance of systematic, data-driven method selection instead of defaulting to a single preferred method.

Biological Transition Score

For datasets with known developmental or functional relationships between cell populations, a biological transition score can quantify how well the embedding reflects these relationships. This approach was introduced in the 2026 dendritic spine study, which used a Biological Transition Score to quantify how well low-dimensional embeddings reflect known developmental and functional relationships among spine types (PLOS ONE 10.1371/journal.pone.0349775).

The score is computed by defining expected transitions between cell populations based on prior biological knowledge and then measuring whether cells in the embedding are positioned consistently with these expected transitions. A high score indicates that the embedding preserves the known biological relationships, while a low score indicates that the embedding has distorted these relationships.

This approach is particularly valuable for trajectory and differentiation studies, where the expected relationships between cell states are often known from prior experiments or from marker gene expression. The score provides a quantitative complement to visual inspection of trajectory plots.

Practical Workflow for Embedding Evaluation

Step 1: Define the Biological Question

The appropriate evaluation metrics depend on the biological question. For cell type identification, local structure preservation and cluster separation are the primary concerns. For trajectory analysis, continuity and global structure preservation are more important. For data integration across batches or samples, the evaluation should also assess whether the embedding removes batch effects while preserving biological variation.

Researchers should document the biological question before running dimensionality reduction and select evaluation metrics that align with this question. This documentation also supports reproducibility, as it clarifies why specific methods and metrics were chosen.

Step 2: Run Multiple Dimensionality Reduction Methods

instead of committing to a single method, run at least two methods that make different assumptions about the data. A common combination is PCA for a linear baseline and UMAP or t-SNE for nonlinear structure. For datasets where trajectory structure is expected, a variational autoencoder or a method specifically designed for trajectory preservation may be appropriate.

The 2021 comparison of dimensionality reduction methods evaluated stability, accuracy, and computing cost of 10 methods using 30 simulation datasets and five real datasets (PubMed 33833778). The study found that t-SNE yielded the best overall performance with the highest accuracy and computing cost, while UMAP exhibited the highest stability with moderate accuracy and the second highest computing cost. These findings support the common practice of using UMAP for initial exploration and t-SNE for detailed local structure examination.

Step 3: Compute Quantitative Metrics

For each embedding, compute trustworthiness, continuity, and silhouette score. Use a range of k values for trustworthiness and continuity to understand performance at different scales. Compute silhouette scores for the clustering that will be used for downstream analysis.

Record the metric values in a table that also documents the dimensionality reduction method, hyperparameters, and the version of the software used. This record supports reproducibility and allows comparison across datasets or analysis updates.

Step 4: Validate with Marker Genes

Overlay known marker gene expression onto each embedding and assess whether the expected cell populations are separated and cohesive. This step is essential because quantitative metrics cannot detect all forms of biological distortion.

For datasets without well established marker genes, use differentially expressed genes between clusters as a provisional validation. The FindAllMarkers function in Seurat is commonly used for this purpose, as documented in the pulmonary arterial hypertension study (PubMed 38110868) and the thyroid cancer study (PubMed 41298873).

Step 5: Assess Hyperparameter Sensitivity

Run the primary dimensionality reduction method with different hyperparameter values and compare the resulting embeddings. Document the range of parameter values tested and the effect on the quantitative metrics and visual appearance.

If the embedding is highly sensitive to hyperparameters, consider whether the biological conclusions depend on the specific parameter choice. If the conclusions are robust across parameter values, the results are more defensible.

Step 6: Document and Report

Report the dimensionality reduction methods, hyperparameters, evaluation metrics, and validation results in the methods section of any publication or report. This documentation allows readers to assess the reliability of the embedding and to reproduce the analysis.

The Galaxy Training Network provides accessible workflow training and analysis tutorials that emphasize reproducibility in bioinformatics analyses (Galaxy Training Network). The nf-core documentation similarly emphasizes community pipeline standards and reproducible workflow context (nf-core Documentation). These resources can help researchers structure their analysis workflows for reproducibility.

Records and Measurements

What to Record

For each dimensionality reduction run, record the following information:

  • Software package and version (for example, Seurat, scanpy, or scVI)
  • Dimensionality reduction method and specific algorithm variant
  • All hyperparameter values, including perplexity for t-SNE, n_neighbors and min_dist for UMAP, and latent dimensionality for variational autoencoders
  • Input data version, including the gene filtering and normalization parameters
  • Random seed if the method uses stochastic initialization
  • Computing time and memory usage
  • Quantitative evaluation metrics including trustworthiness, continuity, and silhouette score
  • Marker genes used for biological validation and the results of the validation

This record supports reproducibility and allows the analysis to be updated when new versions of software packages are released. The Bioconductor project provides official package, workflow, installation, and reproducible genomic-analysis documentation that can help researchers structure their analysis pipelines (Bioconductor).

How to Structure Records

A spreadsheet or table with one row per dimensionality reduction run is a practical way to organize records. Columns should include the date, data version, software version, method, hyperparameters, random seed, evaluation metrics, and notes on biological validation.

For larger projects, consider using a workflow management system that automatically records the parameters and versions for each analysis step. The nf-core documentation describes community pipeline standards that support reproducible workflow configuration (nf-core Documentation). The Carpentries lessons provide foundational training in computing, data, shell, Git, and programming that can help researchers implement reproducible workflows (The Carpentries Lessons).

Interpreting Metric Values

Metric values should be interpreted in the context of the dataset and the biological question instead of against fixed thresholds. A trustworthiness value of 0.9 may be acceptable for a dataset with many closely related cell states but inadequate for a dataset with highly distinct cell types where false mixing is a serious concern.

Comparative interpretation is more informative than absolute interpretation. If one embedding has a trustworthiness of 0.85 and another has 0.95, the second embedding is likely better at preserving local structure, provided the comparison uses the same k value and the same input data.

Common Failure Patterns

Overinterpretation of Visual Separation

The most common failure pattern is treating visually distinct clusters in a UMAP or t-SNE plot as evidence of distinct cell types without quantitative or biological validation. Nonlinear dimensionality reduction methods can create the appearance of separation even when the underlying data are continuous. This is particularly problematic for UMAP, which tends to produce compact clusters with clear boundaries even for continuous data.

The 2021 comparison of dimensionality reduction methods found that UMAP well preserves the original cohesion and separation of cell populations (PubMed 33833778). However, this property can lead researchers to overinterpret cluster boundaries that are artifacts of the algorithm instead of genuine biological distinctions.

Ignoring Hyperparameter Sensitivity

A second common failure is using default hyperparameters without assessing whether the results are robust to parameter changes. Default parameters are not universally appropriate, and the 2021 comparison explicitly noted that users need to set hyperparameters according to the specific situation before using dimensionality reduction methods based on nonlinear models and neural networks (PubMed 33833778).

A related failure is not documenting hyperparameters in publications, which prevents other researchers from reproducing the analysis or assessing its reliability.

Relying on a Single Method

A third common failure is relying on a single dimensionality reduction method for all analyses. The 2025 time-series study found that different techniques may lead to discrepancies in the representation of dynamical patterns and that none of the techniques was consistently superior (PubMed 40576031). Relying on a single method risks drawing conclusions that are specific to that method instead of representative of the underlying biology.

Confusing Batch Effect Removal with Biological Variation

When integrating data from multiple batches or samples, dimensionality reduction can remove batch effects at the cost of removing genuine biological variation. The evaluation should assess whether the embedding preserves known biological differences between samples while removing technical differences.

The totalVI framework for joint analysis of CITE-seq data provides an example of an approach that probabilistically represents the data as a composite of biological and technical factors, including protein background and batch effects (PubMed 33589839). This approach explicitly models the sources of variation and can be evaluated for whether it preserves biological signal while removing technical noise.

Using Silhouette Score Without Biological Validation

A fifth common failure is using the silhouette score as the sole evidence of embedding quality. The silhouette score measures cluster separation in the embedding, but it does not assess whether the clusters correspond to genuine biological populations. A high silhouette score can be produced by an embedding that artificially separates cells that are biologically similar.

The silhouette score should always be complemented by marker gene validation and, when possible, by comparison with known biological relationships.

Limitations of Current Evaluation Approaches

No Universal Metric

No single metric captures all aspects of embedding quality. Trustworthiness and continuity assess local structure, silhouette score assesses cluster separation, and marker gene overlay assesses biological validity, but each metric has blind spots. A comprehensive evaluation requires multiple complementary metrics.

The 2026 dendritic spine study found that dimensionality reduction methods capture complementary aspects of spine morphology and that the optimal strategy is dataset-dependent (PLOS ONE 10.1371/journal.pone.0349775). This finding generalizes to other biological systems and underscores the need for systematic evaluation instead of reliance on a single preferred method.

Ground Truth Is Often Unknown

Quantitative metrics such as trustworthiness and continuity compare the embedding with the original high-dimensional space, but the original space itself contains noise and technical artifacts. The metrics assess whether the embedding preserves the structure of the input data, not whether that structure reflects biology.

Marker gene validation provides a partial ground truth, but marker genes are not available for all cell populations, and the expression of known markers does not fully define cell identity. The 2026 scKanFormer study noted that cell type annotation for scRNA-seq data needs to overcome batch effects and effectively handle large-scale datasets, and proposed a supervised framework that avoids dimensionality reduction to enable traceability from attention layers back to the original input features (PLOS Computational Biology 10.1371/journal.pcbi.1014607). This approach highlights the tradeoff between interpretability and the convenience of dimensionality reduction.

Computational Cost

Comprehensive evaluation requires running multiple dimensionality reduction methods with multiple hyperparameter settings, which can be computationally expensive for large datasets. The 2021 comparison found that t-SNE had the highest computing cost among the methods evaluated, while UMAP had the second highest (PubMed 33833778). The 2019 evaluation of 18 methods recorded computational cost as part of the comparison (PubMed 31823809).

Researchers should balance the depth of evaluation against available computing resources. For very large datasets, it may be practical to evaluate embeddings on a representative subset of cells and then apply the selected method to the full dataset.

Evaluation Metrics Are Themselves Parameter Dependent

Trustworthiness and continuity depend on the choice of k, and silhouette score depends on the clustering result. Different parameter choices can lead to different conclusions about embedding quality. Researchers should report the parameters used for evaluation and consider whether the conclusions are robust to parameter changes.

Safety and Reproducibility Context

Reproducibility Standards

Reproducibility is a core requirement for single-cell analysis, and dimensionality reduction evaluation should be documented with the same rigor as other analysis steps. The Galaxy Training Network provides accessible workflow training and analysis tutorials that emphasize reproducibility (Galaxy Training Network). The nf-core documentation describes community pipeline standards for reproducible workflow configuration (nf-core Documentation). The Carpentries lessons provide foundational training in computing, data, shell, Git, and programming that supports reproducible analysis practices (The Carpentries Lessons).

The EMBL-EBI Training program offers bioinformatics learning pathways and data-resource training that can help researchers develop the skills needed for rigorous single-cell analysis (EMBL-EBI Training). The NCBI provides official descriptions of databases, search systems, sequence resources, and analysis services that support data access and reproducibility (NCBI Data Resources).

Data Management

Dimensionality reduction evaluation requires access to the original expression data, the reduced embeddings, and the parameters used to generate them. Researchers should store these data in a structured format that supports reproducibility and reanalysis.

The Bioconductor project provides official package, workflow, installation, and reproducible genomic-analysis documentation that can help researchers structure their data and analysis pipelines (Bioconductor). Following these standards supports the long-term usability of single-cell data and analysis results.

Professional Escalation Criteria

Researchers should escalate concerns about embedding quality to a supervisor, collaborator, or statistical consultant when:

  • Quantitative metrics indicate poor local structure preservation (low trustworthiness or continuity) and the biological conclusions depend on the embedding
  • Different dimensionality reduction methods produce conflicting representations of the data and the biological conclusions depend on which method is used
  • Marker gene validation fails to confirm the expected cell populations in the embedding
  • Hyperparameter sensitivity analysis reveals that the embedding changes substantially with small parameter changes
  • The embedding is being used as the basis for clinical or translational conclusions

In these situations, the appropriate response is to conduct additional validation, consider alternative analysis approaches, or consult with a researcher who has expertise in single-cell data analysis methods.

Frequently Asked Questions

What is the difference between trustworthiness and continuity in dimensionality reduction evaluation?

Trustworthiness measures whether cells that appear as neighbors in the low-dimensional embedding were also neighbors in the original high-dimensional space. A low trustworthiness score indicates that the embedding has placed dissimilar cells close together, which can create false clusters. Continuity measures the reverse direction, specifically whether cells that were neighbors in the original high-dimensional space remain neighbors in the embedding. A low continuity score indicates that the embedding has separated cells that were originally similar, which can fragment genuine cell populations. Both metrics should be computed because a method can sacrifice one property to preserve the other.

How do I choose between UMAP and t-SNE for my single-cell data?

The choice depends on the biological question and the properties of the dataset. A 2021 comparison found that t-SNE yielded the best overall performance with the highest accuracy and computing cost, while UMAP exhibited the highest stability with moderate accuracy and the second highest computing cost (PubMed 33833778). UMAP tends to preserve the original cohesion and separation of cell populations and is generally faster for large datasets. t-SNE may provide better local structure resolution but can be more sensitive to hyperparameter choices. The recommended approach is to run both methods and compare the results using quantitative metrics and marker gene validation.

What is a good silhouette score for single-cell clustering?

Silhouette scores range from negative one to positive one, with higher values indicating better cluster separation. A mean silhouette score above 0.5 is generally interpreted as indicating reasonable cluster structure, while scores below 0.25 suggest that clusters are not well separated. However, these thresholds are heuristic and should be interpreted in the context of the biological system. Some cell populations are genuinely similar, and low silhouette scores may reflect real biological continuity instead of poor embedding quality. The silhouette score is most useful as a comparative tool across different embeddings or clustering resolutions.

How many neighbors should I use for UMAP?

The n_neighbors parameter in UMAP controls the size of the local neighborhood considered when constructing the manifold. Smaller values focus on very local structure and can produce more fragmented clusters, while larger values capture broader relationships and can produce more continuous representations. A practical approach is to run UMAP with n_neighbors values of 5, 15, and 30 and compare the resulting embeddings. If the same cell populations appear as distinct clusters across all parameter settings, the clustering result is likely robust. If cluster boundaries shift substantially, the researcher should investigate whether the clusters represent genuine biology or parameter-dependent artifacts.

Why do different dimensionality reduction methods produce different representations of the same data?

Different methods make different assumptions about the structure of the data. PCA assumes that the most important structure is captured by the directions of maximum variance. t-SNE focuses on preserving local pairwise similarities and can distort global relationships. UMAP assumes that the data lie on a low-dimensional manifold and attempts to preserve both local and global structure. Variational autoencoders learn a probabilistic mapping from the data to a low-dimensional latent space. A 2025 study of time-series scRNA-seq data found that none of these techniques was consistently superior and that they may not reliably represent dynamics when used in isolation (PubMed 40576031). The appropriate method depends on the biological question and the properties of the dataset.

How do I validate that my embedding reflects real biology instead of artifacts?

The most direct validation is to overlay known marker gene expression onto the embedding. If the embedding accurately represents the underlying biology, cells expressing markers for a known cell type should occupy a contiguous region of the embedding, and different cell types should occupy distinct regions. Additional validation approaches include comparing embeddings from multiple dimensionality reduction methods, assessing hyperparameter sensitivity, and using a biological transition score when known developmental or functional relationships exist between cell populations. The 2026 dendritic spine study introduced a Biological Transition Score to quantify how well low-dimensional embeddings reflect known developmental and functional relationships (PLOS ONE 10.1371/journal.pone.0349775).

What should I report in my methods section about dimensionality reduction?

Report the software package and version, the dimensionality reduction method and specific algorithm variant, all hyperparameter values, the input data version including gene filtering and normalization parameters, the random seed if the method uses stochastic initialization, and the quantitative evaluation metrics including trustworthiness, continuity, and silhouette score. Also report the marker genes used for biological validation and the results of the validation. This documentation allows readers to assess the reliability of the embedding and to reproduce the analysis.

When should I seek help from a bioinformatics specialist?

Seek help when quantitative metrics indicate poor local structure preservation and the biological conclusions depend on the embedding, when different dimensionality reduction methods produce conflicting representations of the data, when marker gene validation fails to confirm the expected cell populations, when hyperparameter sensitivity analysis reveals that the embedding changes substantially with small parameter changes, or when the embedding is being used as the basis for clinical or translational conclusions. A specialist can help design additional validation experiments, consider alternative analysis approaches, and interpret the results in the context of the specific biological system.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.