The Art of Using t-SNE for Single-Cell Transcriptomics: Pitfalls and Best Practices

By Dr. Zubair Khalid, DVM, MS, PhD ·

The Art of Using t-SNE for Single-Cell Transcriptomics: Pitfalls and Best Practices

Key Takeaways

  • t-SNE excels at local structure visualization but distorts global relationships: The algorithm prioritizes preserving nearest neighbors, leading to inaccurate representations of inter-cluster distances and relative positions. Claims about cell type similarity or lineage relationships require validation beyond visual inspection of t-SNE plots.
  • Perplexity is a critical hyperparameter influencing local vs. global balance: Evaluating multiple perplexity values (e.g., 5, 15, 30, 50, 100) is essential to assess cluster stability; significant structural changes across values indicate potential instability or lack of robust biological signal.
  • Stochasticity necessitates reproducible practices and robustness checks: Setting a random seed and documenting it ensures exact reproducibility of a specific run, but biological conclusions should be validated across multiple seeds to confirm robustness against initialization variability.
  • PCA initialization and a high learning rate improve embedding consistency and convergence: Utilizing PCA for initialization provides a deterministic starting point, while a sufficiently high learning rate facilitates effective optimization, preventing suboptimal embeddings or slow convergence.
  • Quality control and preprocessing choices profoundly impact t-SNE output: Rigorous filtering of low-quality cells (e.g., based on gene count, mitochondrial reads) and appropriate normalization/feature selection are critical; including artifacts can create spurious clusters, while excessive filtering can distort true biological variation.
  • t-SNE is an exploratory tool, not a substitute for formal analysis: Visual clusters must be validated with statistical clustering methods and differential gene expression analysis to identify marker genes; claims derived solely from t-SNE visualization are not scientifically robust.

t-SNE (t-distributed stochastic neighbor embedding) is a dimensionality reduction technique widely used to visualize single-cell RNA sequencing (scRNA-seq) data in two dimensions. The method excels at revealing local structure in high-dimensional data, but naive applications often suffer from severe shortcomings, including inaccurate representation of global structure, sensitivity to hyperparameter choices, and stochastic variability across runs. For biology students, researchers, and laboratory professionals analyzing single-cell transcriptomics data, this article provides a practical framework for diagnosing and fixing common t-SNE issues, with concrete management decisions, quality checks, and professional escalation criteria. The guidance applies to standard scRNA-seq and single-nucleus RNA-seq workflows, from raw count matrices through quality control, normalization, dimensionality reduction, clustering, and visualization.

Understanding t-SNE in the Single-Cell Analysis Pipeline

Single-cell transcriptomics yields ever growing data sets containing RNA expression levels for thousands of genes from up to millions of cells. Common data analysis pipelines include a dimensionality reduction step for visualizing the data in two dimensions, most frequently performed using t-SNE. The method works by converting high-dimensional similarities between cells into joint probabilities and then positioning points in a low-dimensional space to minimize the Kullback-Leibler divergence between the high-dimensional and low-dimensional probability distributions.

The primary strength of t-SNE lies in its ability to preserve local neighborhoods. Cells that are similar in high-dimensional gene expression space tend to appear close together in the t-SNE embedding. This property makes t-SNE particularly useful for identifying cell types, examining developmental trajectories, and assessing the results of clustering algorithms. However, the same properties that make t-SNE effective for local structure visualization also create systematic distortions that can mislead biological interpretation.

The distance between clusters in a t-SNE plot does not carry meaningful quantitative information. The method normalizes distances based on local density, which means that the apparent separation between clusters can be influenced by the number of cells in each cluster instead of the actual transcriptional difference between them. Similarly, the relative positions of clusters in the embedding do not reflect global relationships in the original high-dimensional space. Two clusters that appear far apart in a t-SNE plot may actually be transcriptionally similar, while clusters that appear adjacent may be quite distinct.

These limitations are well documented in the bioinformatics literature. A systematic evaluation of dimension reduction methods for transcriptomic data visualization found that t-SNE and related algorithms do not preserve important aspects of high-dimensional structure and are sensitive to arbitrary user choices. The same evaluation emphasized that dimension reduction methods should be evaluated carefully before trusting their results, considering local structure preservation, global structure preservation, sensitivity to parameter choices, sensitivity to preprocessing choices, and computational efficiency.

For researchers using t-SNE in single-cell analysis, the practical implication is clear: t-SNE plots are useful exploratory tools, but they are not reliable evidence for biological conclusions about the relationships between cell populations. Any claim about cell type relationships, developmental trajectories, or marker gene expression patterns should be supported by additional analyses, including clustering metrics, differential expression testing, and alternative visualization methods.

At a Glance: t-SNE Decision Framework

Decision PointRecommended ActionCommon MistakeDiagnostic Signal
Perplexity selectionRun multiple values (5, 15, 30, 50, 100) and compare cluster stabilityUsing a single default value without checking robustnessMajor cluster structure changes dramatically across perplexity values
InitializationUse PCA initialization for deterministic starting pointsRandom initialization without documenting seedEmbeddings differ substantially across runs with identical parameters
Learning rateSet a relatively high learning rate for effective optimizationLow learning rate causing slow convergence or poor local optimaOptimization fails to converge or embedding changes after many iterations
Large data setsApply exaggeration and downsampling-based initializationRunning naive t-SNE on millions of cells without adaptationComputational time becomes prohibitive or embedding quality degrades
Cluster interpretationValidate visual clusters with formal clustering and marker genesTreating visual separation as proof of distinct cell typesVisual clusters do not match statistically defined groups or known markers

Core Principles of Reliable t-SNE Visualization

Local Structure Preservation Versus Global Structure Distortion

The fundamental tradeoff in t-SNE is between local and global structure preservation. The algorithm prioritizes maintaining the identity of each cell's nearest neighbors in the high-dimensional space, often at the expense of accurately representing the broader relationships between groups of cells. This tradeoff is inherent to the method and cannot be eliminated by parameter tuning alone.

In practice, this means that t-SNE plots are most reliable for answering questions about local structure, such as whether a particular group of cells forms a distinct subpopulation or whether two closely related cell types can be separated. The plots are less reliable for answering questions about global structure, such as which cell types are most similar to each other overall or how different lineages relate to one another.

Researchers should document which type of question their t-SNE visualization is intended to address and choose their analytical approach accordingly. For global structure questions, alternative methods such as UMAP, principal component analysis (PCA), or force-directed layouts may provide more reliable representations, though each method has its own limitations and should be evaluated in the context of the specific data set.

The Role of Perplexity in Embedding Quality

Perplexity is the most influential hyperparameter in t-SNE and controls the balance between local and global aspects of the data. The perplexity parameter can be interpreted as a smooth measure of the effective number of neighbors considered for each cell. Low perplexity values emphasize very local structure, potentially fragmenting continuous populations into artificial clusters. High perplexity values incorporate more global information but can obscure fine-grained local distinctions.

The choice of perplexity should be guided by the number of cells in the data set and the research question. For small data sets with a few hundred cells, perplexity values in the range of 5 to 30 are commonly used. For larger data sets with thousands or tens of thousands of cells, higher perplexity values may be appropriate. However, no single perplexity value works universally, and researchers should evaluate multiple settings to understand how the embedding changes.

A practical approach is to generate t-SNE plots across a range of perplexity values, such as 5, 15, 30, 50, and 100, and compare the resulting embeddings. If the major cluster structure remains consistent across perplexity values, the visualization is likely robust. If the embedding changes dramatically with perplexity, the results should be interpreted with caution, and the source of the instability should be investigated.

Stochasticity and Reproducibility Concerns

t-SNE is a stochastic algorithm, meaning that different runs on the same data can produce different embeddings. The randomness arises from the initial positions of points in the low-dimensional space and from the optimization process itself. This variability can be substantial, particularly for data sets with many cells or complex structure.

The stochastic nature of t-SNE has practical consequences for research reproducibility. A t-SNE plot generated in one session may not match a plot generated in a later session, even with identical input data and parameters. This variability can complicate the comparison of results across studies, the verification of published figures, and the interpretation of clustering results that depend on the embedding.

To address this issue, researchers should set a random seed before running t-SNE and document the seed value in their analysis records. This practice ensures that the embedding can be reproduced exactly in future analyses. However, setting a random seed does not eliminate the underlying stochasticity, it only makes a particular run reproducible. Researchers should also examine whether their biological conclusions are robust across multiple random seeds, particularly for claims about cluster separation or cell type identity.

PCA Initialization and Learning Rate Considerations

The initialization of the low-dimensional embedding has a substantial impact on the final t-SNE result. Random initialization can lead to different solutions on each run and may converge to suboptimal embeddings that do not accurately reflect the data structure. Principal component analysis (PCA) initialization provides a deterministic starting point that captures the major axes of variation in the data, leading to more consistent and reliable embeddings.

The learning rate controls the step size during the optimization process. A learning rate that is too low can cause the optimization to converge slowly or become trapped in poor local optima. A learning rate that is too high can cause the embedding to become unstable or fail to converge. For most single-cell data sets, a relatively high learning rate is recommended to ensure that the optimization explores the embedding space effectively.

The published protocol for creating more faithful t-SNE visualizations includes PCA initialization, a high learning rate, and multi-scale similarity kernels. For very large data sets, the protocol additionally uses exaggeration and downsampling-based initialization. These choices are designed to improve the reliability and interpretability of the resulting embeddings, particularly for data sets with millions of cells.

Data Inputs and Preprocessing Decisions

Quality Control Before Dimensionality Reduction

The quality of the t-SNE embedding depends directly on the quality of the input data. Poor quality control can introduce artifacts that appear as spurious clusters or distorted relationships in the visualization. Standard quality control steps for scRNA-seq data include filtering cells based on the number of detected genes, the total number of counts, and the percentage of mitochondrial reads.

Cells with very few detected genes may represent empty droplets or damaged cells, while cells with very high counts may represent doublets or multiplets. The percentage of mitochondrial reads is often used as an indicator of cell viability, with high percentages suggesting that the cell membrane has been compromised. These quality control decisions should be made based on the specific data set and experimental context, and the thresholds should be documented in the analysis records.

The choice of quality control thresholds can have a substantial impact on the t-SNE embedding. Including low-quality cells can create artificial clusters that reflect technical artifacts instead of biological variation. Excluding too many cells can remove rare populations or distort the representation of the remaining cells. Researchers should evaluate the sensitivity of their t-SNE results to quality control choices and report the thresholds used in their analysis.

Normalization and Transformation Choices

Normalization is a critical preprocessing step that adjusts for differences in sequencing depth between cells. The goal is to make gene expression values comparable across cells while preserving biological variation. Common approaches include library size normalization, which scales each cell to a common total count, and log transformation, which stabilizes the variance and reduces the influence of extreme values.

The choice of normalization method can affect the t-SNE embedding, particularly for data sets with substantial variation in sequencing depth. Some normalization methods may overcorrect for technical variation, removing genuine biological differences between cell types. Others may undercorrect, leaving technical artifacts that appear as clusters in the embedding.

Feature selection is another important preprocessing decision. Most single-cell analysis pipelines select a subset of highly variable genes for downstream analysis, reducing the dimensionality of the data and focusing on genes that carry biological signal. The number of selected genes and the method used to identify them can influence the t-SNE embedding. Researchers should evaluate whether their results are robust to changes in feature selection parameters.

Imputation and Its Impact on Visualization

Dropout events in scRNA-seq lead to missing data and noise in the gene-cell expression matrix, adversely affecting downstream analyses. Dropout occurs when a gene that is expressed in a cell is not detected due to technical limitations, resulting in excessive zero counts in the expression matrix. These dropout events can distort the apparent similarity between cells and create artifacts in t-SNE embeddings.

Several imputation methods have been developed to recover the true gene expression levels before downstream analysis. A low-rank tensor completion-based method, termed scLRTC, exploits the similarity of single cells to build a third-order low-rank tensor and employs tensor decomposition to denoise the data. The method reconstructs cell expression by adopting a low-rank tensor completion algorithm, restoring gene-to-gene and cell-to-cell correlations. In evaluations on simulated and real scRNA-seq data sets, scLRTC outperformed other methods in imputing dropouts closest to the original expression values and achieved accurate cell classification results with different clustering methods.

Another approach, CMF-Impute, uses collaborative matrix factorization to impute dropout entries in the scRNA-seq expression matrix. The method considers the association among cells and genes and has been shown to improve cell classification accuracy and reconstruct cell-to-cell and gene-to-gene correlations. Both imputation methods have been demonstrated to be effective in cell visualization and in inferring cell lineage trajectories.

The decision to impute dropout values before t-SNE visualization should be made carefully. Imputation can improve the representation of continuous biological processes and reduce the influence of technical noise. However, imputation methods make assumptions about the underlying data structure, and these assumptions may not hold for all data sets. Researchers should compare t-SNE embeddings with and without imputation to understand how the visualization changes and whether biological conclusions are affected.

Practical Workflow for t-SNE Analysis

Step 1: Data Preparation and Quality Control

Begin with the raw count matrix from the single-cell sequencing platform. Filter cells based on quality metrics appropriate for the data set, including the number of detected genes, total counts, and mitochondrial read percentage. Document all filtering thresholds and the number of cells removed at each step.

For data sets with substantial dropout, consider whether imputation is appropriate. Evaluate imputation methods such as scLRTC or CMF-Impute on a subset of the data before applying to the full data set. Compare the distribution of expression values before and after imputation to ensure that the method is not introducing artifacts.

Normalize the data using a method appropriate for the experimental design. Log transform the normalized values and select highly variable genes for downstream analysis. Document the number of genes selected and the method used for feature selection.

Step 2: Dimensionality Reduction with PCA

Apply PCA to the normalized and scaled data to reduce the dimensionality before t-SNE. PCA provides a deterministic initialization for t-SNE and reduces the computational burden of the embedding. The number of principal components to retain should be determined based on the variance explained and the structure of the data.

PCA initialization is recommended over random initialization because it produces more consistent embeddings across runs and captures the major axes of variation in the data. The number of principal components used for initialization should be documented and reported with the t-SNE results.

Step 3: Running t-SNE with Appropriate Parameters

Select the perplexity value based on the number of cells in the data set and the research question. Generate embeddings across a range of perplexity values to assess the stability of the cluster structure. Set a high learning rate to ensure effective optimization, and use exaggeration for large data sets to improve the separation of clusters.

For very large data sets with millions of cells, consider downsampling-based initialization. This approach involves running t-SNE on a subset of the data and using the resulting embedding to initialize the full data set. This strategy can improve the quality of the embedding while reducing computational cost.

Set a random seed before running t-SNE and document the seed value in the analysis records. Run t-SNE multiple times with different random seeds to assess the variability of the embedding and ensure that biological conclusions are robust.

Step 4: Evaluating the Embedding Quality

Examine the t-SNE embedding for common artifacts, including artificial clusters, excessive fragmentation, and poor separation of known cell types. Compare the embedding with the results of clustering algorithms to assess whether the visual clusters correspond to statistically defined groups.

Evaluate the preservation of local structure by examining whether cells with similar expression profiles appear close together in the embedding. Evaluate the preservation of global structure by comparing the t-SNE embedding with PCA or other dimension reduction methods. If the global structure is not preserved, consider whether the research question requires global or local structure information.

Step 5: Integrating t-SNE with Clustering and Differential Expression

Use t-SNE as a visualization tool in conjunction with formal clustering methods, not as a substitute for them. Clustering algorithms such as SC3, Seurat, and t-SNE followed by k-means can identify cell populations based on statistical criteria, and the results can be projected onto the t-SNE embedding for visualization.

The choice of clustering method can have a substantial impact on the results. A study of ensemble clustering approaches found that different methods utilize different characteristics of data and yield varying results in terms of both the number of clusters and actual cluster assignments. The SAFE-clustering method aggregates results from multiple clustering methods to build a consensus solution, improving cluster number estimation and cluster assignment accuracy.

After clustering, perform differential expression analysis to identify marker genes for each cluster. Validate the biological identity of the clusters based on known marker genes and compare the results with published data sets. The t-SNE embedding should be considered a visualization of the clustering results, not independent evidence for the existence of cell types.

Options and Tradeoffs in t-SNE Implementation

Perplexity Selection Strategies

The choice of perplexity is the most consequential parameter decision in t-SNE analysis. Low perplexity values emphasize local structure and can reveal fine-grained distinctions between closely related cells. However, low perplexity can also fragment continuous populations into artificial clusters, particularly for data sets with many cells.

High perplexity values incorporate more global information and can produce more stable embeddings. However, high perplexity can obscure local structure and may not be appropriate for data sets with strong local heterogeneity. The optimal perplexity depends on the number of cells, the complexity of the biological system, and the research question.

A practical strategy is to run t-SNE with multiple perplexity values and compare the results. If the major cluster structure is consistent across perplexity values, the embedding is likely reliable. If the structure changes dramatically, the data may not have a clear cluster structure, or the preprocessing steps may need to be revisited.

Random Seed Management

The stochastic nature of t-SNE means that different runs can produce different embeddings. This variability can be managed by setting a random seed before each run and documenting the seed value. However, setting a seed does not eliminate the underlying stochasticity, and researchers should examine whether their conclusions are robust across multiple seeds.

For publications, the random seed and all t-SNE parameters should be reported to enable reproduction of the exact embedding. The analysis code and data should be made available to allow other researchers to verify the results. Reproducibility is a core principle of bioinformatics research, and tools such as Bioconductor provide frameworks for reproducible genomic analysis.

Computational Efficiency Considerations

t-SNE can be computationally intensive for large data sets, particularly those with millions of cells. The computational cost scales with the number of cells and the number of iterations required for convergence. For very large data sets, the analysis may require substantial computational resources and time.

Several strategies can reduce the computational burden. Downsampling-based initialization involves running t-SNE on a subset of the data and using the result to initialize the full data set. Multi-scale similarity kernels can capture structure at multiple resolutions while reducing the number of iterations required. Exaggeration can accelerate the separation of clusters during the early stages of optimization.

The choice of computational strategy should be documented and reported with the results. Researchers should also consider whether the computational cost is justified by the biological insight gained from the visualization.

Alternative Visualization Methods

t-SNE is one of several dimension reduction methods available for single-cell transcriptomics visualization. UMAP is a popular alternative that often produces embeddings with better preservation of global structure and faster computation. Other methods include PaCMAP, TriMap, and ForceAtlas2, each with its own strengths and limitations.

A systematic evaluation of dimension reduction methods for transcriptomic data visualization found that different methods have different characteristics in terms of local structure preservation, global structure preservation, sensitivity to parameter choices, sensitivity to preprocessing choices, and computational efficiency. The choice of method should be guided by the scientific goals of the user and the specific characteristics of the data set.

For research questions that require accurate representation of global structure, methods such as UMAP or PCA may be more appropriate than t-SNE. For research questions that focus on local structure, t-SNE may be the preferred choice. Researchers should consider using multiple methods and comparing the results to gain a more complete understanding of the data.

Observations and Measurements for Quality Assessment

Quantitative Metrics for Embedding Quality

Visual inspection of t-SNE plots is necessary but not sufficient for assessing embedding quality. Quantitative metrics can provide objective measures of how well the embedding preserves the structure of the high-dimensional data. These metrics can be used to compare different parameter settings, different preprocessing choices, and different dimension reduction methods.

Local structure preservation can be measured by comparing the nearest neighbors of each cell in the high-dimensional space with the nearest neighbors in the embedding. High preservation indicates that cells that are similar in gene expression space appear close together in the visualization. Global structure preservation can be measured by comparing the overall distances between clusters in the high-dimensional space with the distances in the embedding.

The evaluation framework described in the dimension reduction literature considers five components: preservation of local structure, preservation of global structure, sensitivity to parameter choices, sensitivity to preprocessing choices, and computational efficiency. Researchers should evaluate their t-SNE embeddings along each of these dimensions and document the results.

Cluster Quality Assessment

The quality of clusters identified in t-SNE embeddings can be assessed using metrics such as the adjusted rand index (ARI) and normalized mutual information (NMI). These metrics compare the cluster assignments with a reference labeling, such as known cell types or the results of an independent clustering method.

The adjusted rand index measures the similarity between two clusterings, corrected for chance. A value of 1 indicates perfect agreement, while a value of 0 indicates random agreement. The normalized mutual information measures the amount of information shared between two clusterings, normalized to a scale from 0 to 1.

These metrics are commonly used in the evaluation of clustering methods for scRNA-seq data. Studies of imputation methods and clustering algorithms have used ARI and NMI to assess the accuracy of cell classification results. Researchers should report these metrics when evaluating the quality of their t-SNE-based clustering results.

Diagnostic Plots for Embedding Assessment

Several diagnostic plots can help identify problems with t-SNE embeddings. The Kullback-Leibler divergence between the high-dimensional and low-dimensional probability distributions provides a measure of how well the embedding represents the original data. A high divergence indicates that the embedding does not accurately represent the data structure.

The number of iterations required for convergence can indicate whether the optimization is proceeding effectively. If the embedding continues to change after many iterations, the learning rate may be too high or the data may not have a clear cluster structure.

The distribution of distances between cells in the embedding can reveal artifacts. If the distances are highly variable, the embedding may be dominated by a few outliers. If the distances are too uniform, the embedding may not be capturing meaningful structure.

Records and Documentation Standards

Parameter Documentation

Every t-SNE analysis should include complete documentation of all parameters and preprocessing choices. The documentation should include the version of the software used, the random seed, the perplexity value, the learning rate, the number of iterations, the initialization method, and the number of principal components used for initialization.

The documentation should also include the quality control thresholds, normalization method, feature selection parameters, and any imputation methods applied. This information is essential for reproducing the analysis and for interpreting the results in the context of the specific data set.

Reproducibility Practices

Reproducibility is a core principle of bioinformatics research. The analysis code, data, and documentation should be organized to enable other researchers to reproduce the results. Tools such as Bioconductor provide frameworks for reproducible genomic analysis, and workflow management systems such as nf-core provide standards for community pipelines.

Version control is essential for tracking changes to analysis code and documenting the evolution of the analysis. The Carpentries offer lessons on foundational computing, data, shell, Git, and programming that provide the skills needed for reproducible research.

Reporting Standards for Publications

Publications that include t-SNE visualizations should report all parameters and preprocessing choices in the methods section. The random seed should be reported to enable reproduction of the exact embedding. The software version and analysis code should be made available to allow other researchers to verify the results.

The limitations of t-SNE should be acknowledged in the interpretation of results. Claims about the relationships between cell populations should be supported by additional analyses, including clustering metrics, differential expression testing, and alternative visualization methods.

Common Failure Patterns and Their Diagnosis

Artificial Cluster Fragmentation

One of the most common problems with t-SNE embeddings is the appearance of artificial clusters that do not correspond to genuine biological populations. This fragmentation can occur when the perplexity is too low, causing the algorithm to emphasize local structure at the expense of global relationships. It can also occur when the data contains technical artifacts, such as batch effects or dropout events.

To diagnose artificial fragmentation, compare the t-SNE embedding with the results of formal clustering methods. If the visual clusters do not correspond to statistically defined groups, the embedding may be misleading. Examine the expression of known marker genes to determine whether the apparent clusters have distinct biological identities.

Misleading Cluster Distances

The distances between clusters in a t-SNE plot do not carry quantitative meaning. Two clusters that appear far apart may be transcriptionally similar, while clusters that appear adjacent may be quite distinct. This distortion can lead to incorrect conclusions about the relationships between cell populations.

To diagnose misleading cluster distances, compare the t-SNE embedding with PCA or other dimension reduction methods that preserve global structure. Calculate the actual transcriptional distances between clusters using metrics such as Euclidean distance or correlation. If the visual distances do not match the transcriptional distances, the t-SNE embedding is not reliable for global structure interpretation.

Instability Across Runs

The stochastic nature of t-SNE can produce substantially different embeddings on different runs. This instability can be caused by random initialization, insufficient iterations, or a learning rate that is too high or too low. Instability is particularly problematic for data sets with many cells or complex structure.

To diagnose instability, run t-SNE multiple times with different random seeds and compare the resulting embeddings. If the major cluster structure is consistent across runs, the embedding is likely reliable. If the structure changes dramatically, the analysis parameters may need to be adjusted, or the data may not have a clear cluster structure.

Batch Effects and Technical Artifacts

Batch effects can create artificial clusters in t-SNE embeddings that reflect technical variation instead of biological differences. These effects can arise from differences in sequencing depth, library preparation, or other technical factors between samples or experimental batches.

To diagnose batch effects, color the t-SNE embedding by batch or sample identity. If cells from the same batch cluster together, the embedding may be dominated by technical variation. Consider batch correction methods or data integration approaches to remove technical effects before t-SNE visualization.

Limitations and Interpretation Boundaries

What t-SNE Cannot Show

t-SNE cannot accurately represent the global structure of the data. The distances between clusters in the embedding do not reflect the actual transcriptional distances in the high-dimensional space. The relative positions of clusters are influenced by local density and the specific parameters used for the analysis.

t-SNE cannot be used to infer developmental trajectories or lineage relationships. The embedding does not preserve the continuous structure of the data, and apparent trajectories in the plot may be artifacts of the algorithm. Trajectory inference should be performed using dedicated methods that model the underlying biological processes.

t-SNE cannot be used to compare the similarity of cell populations across different data sets. The embedding is specific to the input data and parameters, and the same cell type may appear in different positions in different analyses. Cross-data set comparisons should be performed using methods designed for data integration.

The Role of t-SNE in the Analysis Workflow

t-SNE is an exploratory visualization tool, not a statistical inference method. The embedding should be used to generate hypotheses about cell populations and their relationships, but these hypotheses should be tested using formal statistical methods.

The results of t-SNE analysis should be integrated with other evidence, including clustering metrics, differential expression analysis, and known biological knowledge. The t-SNE embedding should be considered one piece of evidence in a broader analytical framework, not the sole basis for biological conclusions.

When to Escalate to Professional Support

If t-SNE embeddings are consistently unstable across parameter settings, the data may have underlying quality issues that require professional attention. Consider consulting with a bioinformatics specialist or a core facility for guidance on data preprocessing and analysis.

If the biological conclusions drawn from t-SNE analysis are not supported by other evidence, the analysis may need to be revisited. Consider whether the preprocessing steps are appropriate, whether the parameters are optimal, and whether the interpretation of the embedding is correct.

If the data set is very large or complex, the computational requirements may exceed the available resources. Consider using workflow management systems such as nf-core or Galaxy Training Network for scalable analysis pipelines.

Safety and Regulatory Context

Data Management and Privacy

Single-cell transcriptomics data may contain sensitive information about research participants. Researchers should follow institutional guidelines for data management and privacy, including de-identification of samples and secure storage of data. The NCBI provides data resources and search systems for sequence data, and researchers should follow the appropriate data submission and access procedures.

Ethical Use of Computational Tools

Computational tools for single-cell analysis should be used ethically and responsibly. Researchers should not manipulate analysis parameters to produce desired results, and they should report all analysis choices transparently. The use of automated tools for clustering and visualization should be documented, and the limitations of these tools should be acknowledged.

Compliance With Publication Standards

Journals increasingly require detailed reporting of computational methods and data availability. Researchers should follow the reporting standards of their target journals and provide sufficient information for other researchers to reproduce the analysis. The EMBL-EBI Training program offers learning pathways for bioinformatics data resources and practical analysis education.

Frequently Asked Questions

What perplexity value should I use for my single-cell data set?

The optimal perplexity depends on the number of cells and the structure of the data. For small data sets with a few hundred cells, perplexity values in the range of 5 to 30 are commonly used. For larger data sets with thousands of cells, higher perplexity values may be appropriate. The best approach is to run t-SNE with multiple perplexity values and compare the resulting embeddings. If the major cluster structure remains consistent across perplexity values, the visualization is likely robust. If the embedding changes dramatically, the results should be interpreted with caution.

Why does my t-SNE plot look different every time I run it?

t-SNE is a stochastic algorithm, meaning that different runs on the same data can produce different embeddings. The randomness arises from the initial positions of points in the low-dimensional space and from the optimization process. To make the embedding reproducible, set a random seed before running t-SNE and document the seed value. However, setting a seed does not eliminate the underlying stochasticity, and you should examine whether your biological conclusions are robust across multiple random seeds.

Can I use t-SNE to determine the number of cell types in my data?

t-SNE is not a clustering method and should not be used to determine the number of cell types. The visual clusters in a t-SNE plot may not correspond to genuine biological populations, and the apparent number of clusters can be influenced by the perplexity value and other parameters. Use formal clustering methods such as SC3, Seurat, or t-SNE followed by k-means to identify cell populations, and use the t-SNE embedding for visualization of the clustering results.

How do I know if my t-SNE plot is reliable?

A reliable t-SNE plot should be stable across different parameter settings and random seeds. The major cluster structure should remain consistent when the perplexity is varied within a reasonable range. The clusters should correspond to statistically defined groups from formal clustering methods, and the expression of known marker genes should support the biological identity of the clusters. If the embedding is unstable or the clusters do not have clear biological identities, the analysis should be revisited.

Should I impute dropout values before running t-SNE?

Dropout events in scRNA-seq can create noise and missing data that adversely affect downstream analyses. Imputation methods such as scLRTC and CMF-Impute can recover the true gene expression levels and improve cell classification and visualization. However, imputation methods make assumptions about the underlying data structure, and these assumptions may not hold for all data sets. Compare t-SNE embeddings with and without imputation to understand how the visualization changes and whether biological conclusions are affected.

What is the difference between t-SNE and UMAP for single-cell visualization?

t-SNE excels at revealing local structure in high-dimensional data but does not accurately represent global structure. UMAP often produces embeddings with better preservation of global structure and is computationally faster for large data sets. The choice between the two methods depends on the research question. For questions about local structure, such as identifying closely related cell types, t-SNE may be appropriate. For questions about global structure, such as the relationships between major cell lineages, UMAP or PCA may be more reliable.

How should I report t-SNE parameters in my publication?

Report all parameters and preprocessing choices in the methods section, including the software version, random seed, perplexity value, learning rate, number of iterations, initialization method, and number of principal components used for initialization. Also report the quality control thresholds, normalization method, feature selection parameters, and any imputation methods applied. Make the analysis code and data available to allow other researchers to reproduce the results.

What should I do if my t-SNE plot shows artifacts or unexpected clusters?

First, examine whether the artifacts are consistent across different parameter settings and random seeds. If the artifacts are not stable, they may be due to the stochastic nature of the algorithm. If the artifacts are stable, investigate the source of the problem. Check the quality control metrics for the cells in the affected clusters, examine the expression of known marker genes, and compare the t-SNE embedding with PCA or other dimension reduction methods. If the problem persists, consider whether the preprocessing steps are appropriate or whether the data has underlying quality issues that require professional attention.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.