Visualizing Continuous Trajectories in Single-Cell Data: Beyond Cluster-Based UMAP and t-SNE

By Dr. Zubair Khalid, DVM, MS, PhD ·

Visualizing Continuous Trajectories in Single-Cell Data: Beyond Cluster-Based UMAP and t-SNE

Key Takeaways

  • Standard UMAP and t-SNE visualizations, optimized for local neighborhood preservation, can obscure continuous biological processes like differentiation by creating artificial cluster boundaries, leading to misinterpretations of intermediate states as distinct cell types.
  • Trajectory inference methods, such as Monocle, Slingshot, and CellRank, explicitly model cell ordering by assigning pseudotime values, which, when overlaid onto dimensionality reduction plots, reveal continuous structure and directionality of biological transitions.
  • The embedding (e.g., UMAP) and the trajectory are separate analytical outputs; trajectory information is not inherent in the dimensionality reduction coordinates and must be computed independently and then visualized in conjunction.
  • Branching biological processes require explicit modeling by trajectory inference tools that can assign cells to distinct lineages, as a single pseudotime axis cannot represent divergence points.
  • Directionality in trajectory inference is critical and can be unknown in contexts like regeneration or disease; methods like CellRank integrate RNA velocity to infer directionality, overcoming limitations of purely pseudotime-based approaches.
  • Robust trajectory interpretation necessitates rigorous quality control of input single-cell data, validation of inferred paths with known marker gene expression trends, and careful consideration of parameter sensitivity and potential batch effects.

Standard UMAP and t-SNE plots are designed to place similar cells near each other, which naturally produces visual groupings that look like discrete clusters. When a biological process such as differentiation, activation, or disease progression is continuous, those same plots can hide the ordering of cells along the process or suggest sharp boundaries where none exist. The practical problem for researchers is that cluster-based visualizations can lead to incorrect interpretations of continuous biology, such as calling intermediate states separate cell types or missing the direction of a transition. This article explains how to overlay trajectory information from tools such as Monocle, Slingshot, CellRank, and related methods onto dimensionality reduction plots, how to interpret the combined visualization correctly, and what records and quality checks are needed before those interpretations are used in a report or publication.

Why Cluster-Based Visualizations Obscure Continuous Processes

Dimensionality reduction methods like UMAP and t-SNE optimize different objectives than trajectory inference. UMAP prioritizes preserving local neighborhood structure, which makes it excellent for separating distinct cell populations but less reliable for representing global order. t-SNE emphasizes local pairwise similarities and often produces separated islands even when the underlying biology is a smooth continuum. Principal component analysis preserves global variance but does not guarantee that the first two components correspond to a biologically meaningful axis of progression.

The consequence is that a standard UMAP plot of a differentiating tissue can show several apparent clusters even when the cells actually form a continuous developmental path. A researcher who interprets those clusters as discrete cell types may then assign identities that do not reflect the true state of the cells. The problem is compounded when the same dataset is analyzed with different random seeds or different preprocessing choices, because UMAP and t-SNE can produce visually different layouts from the same input.

Trajectory inference methods address this by explicitly modeling the ordering of cells along a biological process. These methods assign each cell a pseudotime value, which is a numerical estimate of its position along the inferred trajectory. When pseudotime or branch probabilities are overlaid onto a UMAP or t-SNE plot, the continuous structure becomes visible even if the underlying layout appears clustered. The visualization then communicates both the local relationships among cells and the global ordering of the process.

The scale of this problem is substantial. The scRNA-tools database has cataloged over 1000 software tools for analyzing single-cell RNA sequencing data, and the field has shifted focus from ordering cells on continuous trajectories to integrating multiple samples and using reference datasets. This shift reflects both the maturity of trajectory methods and the recognition that continuous processes require dedicated analytical approaches instead of relying on cluster-based visualizations alone.

Core Principles for Combining Trajectories with Dimensionality Reduction

Trajectory Inference Produces an Ordering, Not a Physical Map

Pseudotime is a computational construct that orders cells by transcriptional similarity along a path. It does not measure real time, and it does not imply that cells physically move along the plotted coordinates. A cell with pseudotime 0.2 is more similar to the inferred starting state than a cell with pseudotime 0.8, but the actual elapsed time between those states is unknown. This distinction matters when interpreting developmental processes, because a trajectory inferred from snapshot data cannot reveal the rate of progression or the exact timing of gene expression changes.

The developing mouse visual cortex provides a concrete example of how trajectory reconstruction works in practice. A comprehensive atlas built from 568,654 high-quality single-cell transcriptomes and 200,061 high-quality single-nucleus Multiome nuclei, densely sampled from embryonic day 11.5 to postnatal day 56, was used to computationally reconstruct a transcriptomic developmental trajectory map of all excitatory, inhibitory, and non-neuronal cell types. Branching points that mark the emergence of new cell types at specific developmental ages were identified, and the trajectory map showed that neurogenesis, gliogenesis, and early postmitotic maturation give rise to all cell classes in a staggered parallel manner. This example demonstrates that trajectory inference can reveal continuous cell-type diversification at different stages of development, but the pseudotime values remain computational orderings instead of direct measurements of elapsed developmental time.

The Embedding and the Trajectory Are Separate Analyses

A common workflow error is to assume that the UMAP coordinates themselves contain trajectory information. They do not. The UMAP layout is a projection of the high-dimensional expression space into two dimensions, and the trajectory is a separate computational result. Overlaying the trajectory onto the embedding is a visualization step that combines two distinct analyses. The quality of the combined plot depends on both the quality of the trajectory inference and the fidelity of the embedding to the underlying data structure.

Conventional dimensionality reduction methods such as t-SNE, UMAP, PCA, and PHATE optimize different visualization objectives, resulting in tradeoffs between cluster separability, spatial organization, and temporal coherence. No single method preserves all aspects of the data structure simultaneously. This is why trajectory information must be computed separately and then overlaid, instead of expecting the embedding alone to reveal continuous processes.

Branching Requires Explicit Modeling

Many biological processes involve branching, where a common progenitor population gives rise to multiple lineages. A single pseudotime axis cannot represent a branch point, because cells on different branches share the same early trajectory but diverge later. Methods that infer branching structure, such as Slingshot or the minimum spanning tree approaches used by Monocle, assign cells to branches and compute branch-specific pseudotime. When visualizing branched trajectories, the plot must show branch identity as well as pseudotime, typically through color or line thickness.

Branching is not limited to development. In disease contexts, cellular hierarchies can underlie tumor behavior. Single-cell RNA sequencing of pediatric ependymomas revealed that these tumors are composed of a cellular hierarchy initiating from undifferentiated populations that undergo impaired differentiation toward three lineages of neuronal-glial fate specification. Prognostically favorable groups predominantly harbor differentiated cells, while aggressive groups are enriched for undifferentiated populations. This example shows that branching trajectories can have direct clinical relevance, and the visualization must communicate both the branch structure and the progression along each branch.

Direction Is Not Always Known

Trajectory inference requires that the direction of a biological process is known, which limits application to differentiating systems in normal development. In scenarios such as regeneration, reprogramming, and disease, the direction is often unknown. CellRank addresses this limitation by combining the robustness of trajectory inference with directional information from RNA velocity, taking into account the gradual and stochastic nature of cellular fate decisions and uncertainty in velocity vectors. On pancreas development data, CellRank automatically detects initial, intermediate, and terminal populations, predicts fate potentials, and visualizes continuous gene expression trends along individual lineages. Applied to lineage-traced cellular reprogramming data, predicted fate probabilities correctly recover reprogramming outcomes. CellRank also predicted a new dedifferentiation trajectory during postinjury lung regeneration, including previously unknown intermediate cell states that were confirmed experimentally.

At a Glance: Decision Table for Trajectory Visualization

ScenarioRecommended ApproachKey Visualization ElementCommon Pitfall
Linear differentiation with known start and endMonocle or Slingshot with root cell specifiedPseudotime color gradient overlaid on UMAPAssuming pseudotime equals real time
Branching development with multiple lineagesSlingshot or CellRank with branch assignmentBranch identity as color, pseudotime as intensityInterpreting branch proximity in UMAP as biological similarity
Direction unknown, such as regeneration or diseaseCellRank with RNA velocityFate probabilities and initial or terminal statesUsing a method that requires a known direction
Time-series data with multiple sampling pointsCStreet or dynamic spanning forest mixturesTime point labels connected by inferred edgesIgnoring the temporal information in the trajectory model
Multiple samples or experimental conditionsLamian for differential pseudotime analysisCondition-specific trajectory comparisonsIgnoring cross-sample variability and batch effects

Practical Workflow for Overlaying Trajectories

Step 1: Define the Biological Question and Expected Structure

Before running any trajectory method, write down the biological question in concrete terms. Are you asking whether a continuous gradient exists between two states? Do you expect a single linear path or multiple branches? Is the direction of the process known from experimental design, such as a time course, or unknown, such as in disease tissue? The answers determine which trajectory method is appropriate and how the results should be interpreted.

For time-series data, methods that explicitly use sampling time, such as CStreet, can construct connections between cells within each time point and between adjacent time points. This approach is appropriate when the experimental design includes multiple harvest times and the goal is to connect cell states across those times. CStreet estimates connection probabilities of the cell states and visualizes the trajectory, which may include multiple starting points and paths, using a force-directed graph. For snapshot data without temporal information, pseudotime methods that infer ordering from transcriptional similarity are the standard choice.

The biological question also determines whether a cluster-based approach is appropriate at all. Some datasets contain multiple independent biological processes occurring simultaneously, and a single trajectory model may force all cells onto one path. Methods that model multiple trajectories or dynamic spanning forest mixtures can represent complex relationships better than single-tree approaches. Dynamic spanning forest mixtures use decision-tree models to select genes that account for variations in multimodality, skewness, and time, then build the forest using tree agglomerative hierarchical clustering and dynamic branch cutting.

Step 2: Perform Quality Control Before Trajectory Inference

Trajectory inference is sensitive to the quality of the input data. Cells with low library complexity, high mitochondrial read fractions, or evidence of doublets can distort the inferred path. Standard single-cell quality control should be applied before any trajectory analysis. The specific thresholds depend on the tissue type and protocol, so record the criteria used and justify them in the methods.

The NCBI provides access to sequence data and analysis services that can support quality assessment and data management throughout the workflow. For training on quality control procedures and reproducible analysis practices, the EMBL-EBI training materials cover data resource usage and practical analysis education.

Quality control is particularly important for trajectory inference because low-quality cells are often transcriptionally distinct from healthy cells. They can form spurious branches or create artificial gaps in the trajectory. Doublets can appear as intermediate states between two cell types, which is exactly the kind of artifact that would distort a continuous trajectory. Record the number of cells removed at each quality control step and the criteria used for removal.

Step 3: Select the Trajectory Method Based on Data Structure

The choice of trajectory method should follow from the biological question and the structure of the data. For linear processes with a known starting population, Monocle or Slingshot with a specified root works well. For branched processes, Slingshot uses cluster-based minimum spanning trees to identify lineages, and CellRank combines trajectory inference with RNA velocity to determine direction when it is unknown.

CellRank is specifically designed for scenarios where the direction of the biological process is not known, such as regeneration, reprogramming, and disease. It combines the robustness of trajectory inference with directional information from RNA velocity and accounts for the gradual and stochastic nature of cellular fate decisions. Applied to pancreas development data, CellRank automatically detects initial, intermediate, and terminal populations, predicts fate potentials, and visualizes continuous gene expression trends along individual lineages.

For time-series data, CStreet constructs k-nearest neighbor connections between cells within each time point and between adjacent time points, then estimates connection probabilities between cell states and visualizes the trajectory using a force-directed graph. This method can represent multiple starting points and paths, which is useful when the temporal dynamics are complex.

For data where the trajectory structure is uncertain, scFocus provides an alternative approach that divides cell subpopulations based on reinforcement learning and unsupervised branching in low-dimensional latent space. The lineage component strength of scFocus coincides with the expression regions of hallmark genes, capturing differentiation processes more effectively than the original low-dimensional latent space and showing stronger subpopulation discriminative power. scFocus has been applied to ten single-cell datasets, including small-scale, common-scale, and multi-batch datasets, demonstrating applicability across different data types.

Step 4: Run the Trajectory Inference and Record Parameters

Every trajectory method has parameters that affect the result. Record all of them, including the dimensionality reduction used for the input, the number of neighbors, the minimum branch length, and any root cell specification. The Bioconductor project provides official package documentation and workflow guidance for reproducible genomic analysis, which is useful for ensuring that the analysis can be repeated.

The Galaxy Training Network offers accessible workflow training and analysis tutorials that cover reproducible analysis practices. For community pipeline standards and configuration guidance, the nf-core documentation describes how to run standardized workflows with consistent parameter tracking.

Parameter sensitivity is a genuine concern for trajectory methods. Small changes in the number of neighbors or the minimum branch length can alter the inferred topology. Run the analysis with multiple parameter settings and compare the results. If the trajectory topology changes substantially across parameter values, the inference is not robust and should be interpreted with caution. Record the results of these parameter sweeps in the analysis log.

Step 5: Overlay the Trajectory onto the Dimensionality Reduction Plot

Once the trajectory is inferred, the standard visualization is a UMAP or t-SNE plot where each cell is colored by pseudotime, branch identity, or fate probability. The color scale should be chosen so that the continuous gradient is visible. A sequential color scale from dark to light works well for pseudotime. For branched trajectories, use distinct hues for each branch and vary the intensity within a branch by pseudotime.

The overlay should be generated from the same underlying data as the trajectory inference. If the trajectory was computed on a different dimensionality reduction than the one used for the final plot, the alignment may be poor. Most trajectory methods return pseudotime values for each cell, and those values can be mapped directly onto any embedding of the same cells.

For branched trajectories, color each cell by its assigned branch and use a secondary visual channel, such as point size or opacity, for pseudotime. This approach communicates both the lineage structure and the progression along each branch. The tradeoff is that the plot becomes more complex and may be harder to read when many branches are present.

Step 6: Validate the Trajectory with Independent Evidence

A trajectory is a computational hypothesis, not a proven biological fact. Validate the inferred ordering by checking that known marker genes change expression along the expected path. For example, in a differentiation trajectory, progenitor markers should decrease and mature cell markers should increase as pseudotime advances. If the marker trends contradict the trajectory, the inference may be wrong or the root cell may be misidentified.

The DisCP-Atlas resource maps cellular processes to complex diseases using single-cell data and provides a framework for connecting continuous cell-state trajectories to disease-relevant gene expression programs. DisCP-Atlas identifies 990 cellular processes derived from 57 single-cell datasets across 35 human tissues, each annotated based on pathway enrichment and cell type-specific activities. This resource can help validate whether a trajectory corresponds to known biological processes in specific tissues.

Independent validation can also come from other data modalities. The developing mouse visual cortex atlas combined single-cell RNA sequencing with single-nucleus Multiome data to identify cell-type-specific and temporally resolved gene regulatory networks that link transcription factors and downstream targets. Throughout development, there were cooperative dynamic changes in gene expression and chromatin accessibility in specific cell types. This integration of transcriptomic and epigenomic data provides stronger evidence for trajectory structure than either data type alone.

Options and Tradeoffs in Trajectory Visualization

Pseudotime Color Overlay

The simplest approach is to color each cell by its pseudotime value on the UMAP or t-SNE plot. This works well for linear trajectories and for branched trajectories when the branches are separated in the embedding. The main limitation is that pseudotime alone does not show branch identity, so cells on different branches with similar pseudotime values will have the same color.

Branch Identity Overlay

For branched trajectories, color each cell by its assigned branch and use a secondary visual channel, such as point size or opacity, for pseudotime. This approach communicates both the lineage structure and the progression along each branch. The tradeoff is that the plot becomes more complex and may be harder to read when many branches are present.

Fate Probability Overlay

CellRank and similar methods compute fate probabilities, which estimate the likelihood that a cell will adopt each of several terminal states. These probabilities can be visualized as a ternary plot or as a color mixture on the embedding. This approach is powerful for understanding developmental potential but requires careful explanation to avoid overinterpretation.

Multiple Dimensionality Reduction Integration

Some methods integrate outputs from multiple dimensionality reduction approaches to preserve both local and global structure. GIBOOST uses a Bayesian framework and an optimized autoencoder to select and integrate the two most informative dimensionality reduction methods based on separability, spatial continuity, uniformity, cellular dynamics, and cluster sensitivity. This approach can improve the visualization of differentiation trajectories compared to any single method. GIBOOST has been demonstrated across multiple dynamic biological processes, including epithelial-mesenchymal transition, CiPSC reprogramming, spermatogenesis, and placental development, and it enhances clustering sensitivity and biological relevance compared to individual dimensionality reduction methods.

Hyperbolic Embeddings

Hyperbolic space can represent hierarchical structure more naturally than Euclidean space. LAIOR uses a hyperbolic neural ODE variational framework to learn embeddings that preserve local cell-state structure, global hierarchy, and smooth developmental trajectories simultaneously. This method combines Lorentz geometric regularization for tree-like latent hierarchy, a dual-path information bottleneck for coordinated biological programs, and neural ODE regularization for stable latent trajectories. Across 118 single-cell datasets benchmarked against 23 baseline methods on 22 complementary metrics, LAIOR improves manifold continuity, trajectory coherence, and embedding fidelity while retaining competitive clustering performance. This method is more complex to implement but may be appropriate when the biological process has a strong hierarchical component.

Morphodynamic Trajectories

For data that include time-lapse imaging, morphodynamic trajectories can link gene transcription levels to changes in cell morphology and motility observable in single-cell trajectories. The MMIST framework maps live-cell dynamics to snapshot gene transcript levels and identifies a cell state landscape bound by epithelial and mesenchymal endpoints with distinct sequences of intermediates. This approach predicts expression of thousands of RNA transcripts through extracellular signal-induced epithelial-mesenchymal transition and mesenchymal-epithelial transition with near-continuous time resolution. This option is relevant when imaging data are available alongside molecular measurements.

Observations and Measurements for Trajectory Interpretation

Pseudotime Distribution

Examine the distribution of pseudotime values across cells. A uniform distribution suggests that cells are evenly sampled along the trajectory. Peaks in the distribution indicate that many cells share similar transcriptional states, which may correspond to stable intermediate states or to sampling bias. Valleys indicate regions with few cells, which may be transitional states that are captured less frequently.

The expression of genes during normal development exhibits a high proportion of non-uniformly distributed profiles that are mostly right-skewed and multimodal, with multimodality being a characteristic of major steady states during development. This observation from dynamic spanning forest mixture analysis suggests that pseudotime distributions with multiple modes may reflect genuine biological steady states instead of technical artifacts.

Marker Gene Expression Along Pseudotime

Plot the expression of known marker genes against pseudotime. The expected pattern is a smooth transition from low to high expression or the reverse. Abrupt changes may indicate that the trajectory is discontinuous or that the pseudotime ordering is incorrect. The expression trends should be consistent with the known biology of the system.

For example, in spermatogonial stem cell populations, single-cell RNA sequencing confirmed the dynamic expression of Kit during SSC differentiation and early meiosis initiation. Co-expression analysis revealed significant interactions between Kit, Nmyc, and other pluripotency-associated genes, highlighting the role of Kit in SSC development. A trajectory visualization of this system should show Kit expression increasing along the differentiation path, with the transition occurring at the point where cells commit to differentiation instead of self-renewal.

Branch Probability Distributions

For branched trajectories, examine the distribution of branch probabilities. Cells near the branch point should have intermediate probabilities, while cells far along a branch should have probabilities near 1 for that branch. If cells far from the branch point have mixed probabilities, the branch assignment may be unreliable.

Branch probabilities can also be used to characterize changes in gene expression intensity that reflect continuous cell states without relying on prior information. The combination of branch probabilities and unsupervised clustering can effectively characterize these changes, as demonstrated by the scFocus algorithm.

Comparison with Time-Series Data

When time-series data are available, compare the inferred pseudotime with the actual sampling time. A strong correlation supports the trajectory inference. Poor correlation may indicate that the transcriptional changes are not monotonic over time or that the trajectory method is not capturing the dominant biological signal.

Time-series data can also be used directly for trajectory inference. CStreet uses time-series information to construct k-nearest neighbor connections between cells within each time point and between adjacent time points, then estimates connection probabilities of the cell states. This approach achieves high accuracy and high tolerance compared to six commonly used cell state trajectory reconstruction methods on simulated and real data.

Differential Pseudotime Analysis

When comparing trajectories across conditions, standard pseudotime analysis that ignores sample variability can produce false discoveries. Lamian provides a comprehensive and statistically rigorous computational framework for differential multi-sample pseudotime analysis. It can identify changes in a biological process associated with sample covariates, such as different biological conditions, while adjusting for batch effects, and detect changes in gene expression, cell density, and topology of a pseudotemporal trajectory. Lamian draws statistical inference after accounting for cross-sample variability and substantially reduces sample-specific false discoveries that are not generalizable to new samples.

Records and Measurements for Reproducible Trajectory Analysis

Analysis Log

Maintain a log that records the software versions, parameter settings, and input files for each trajectory analysis. The log should be detailed enough that another researcher can reproduce the analysis from the raw data. The Carpentries lessons provide foundational training in computing, data, shell, Git, and programming that supports reproducible analysis practices.

Parameter Sensitivity Records

Trajectory methods can be sensitive to parameter choices. Record the results of parameter sweeps, such as different numbers of neighbors or different minimum branch lengths. If the trajectory topology changes substantially across parameter values, the inference is not robust and should be interpreted with caution.

Quality Control Metrics

Record the quality control metrics for each cell, including library size, number of detected genes, and mitochondrial read fraction. These metrics should be reported alongside the trajectory results so that readers can assess the data quality. The NCBI provides access to sequence data and analysis services that support quality assessment.

Version Control

Use version control for analysis scripts and configuration files. The Carpentries lessons cover Git and version control practices that are essential for reproducible computational research.

Pipeline Documentation

For standardized workflows, the nf-core documentation describes community pipeline standards, usage, configuration, and reproducible workflow context. Following these standards ensures that the analysis can be run consistently across different computing environments and that parameter tracking is maintained throughout the pipeline.

Common Failure Patterns in Trajectory Visualization

Misinterpreting Cluster Proximity as Biological Similarity

In a UMAP plot, cells that appear close together are not necessarily transcriptionally similar. UMAP can bring distant cell types together in the embedding to preserve local structure. When overlaying a trajectory, cells on different branches may appear adjacent in the plot even though they are transcriptionally distinct. Interpret the trajectory based on the pseudotime and branch assignments, not on the spatial proximity in the embedding.

Using a Directional Method Without Known Direction

Some trajectory methods require the user to specify the direction of the process, either by providing a root cell or by assuming that the process proceeds from a known starting state. Applying such a method to data where the direction is unknown can produce a trajectory that is reversed or otherwise incorrect. CellRank addresses this limitation by combining trajectory inference with RNA velocity to determine direction automatically.

Ignoring Batch Effects

If the data come from multiple batches, samples, or experimental conditions, batch effects can distort the trajectory. Cells from different batches may separate in the embedding, creating apparent branches that reflect technical variation instead of biology. Methods that account for batch effects, such as Lamian for differential pseudotime analysis, should be used when multiple samples are compared.

Overinterpreting Pseudotime as Real Time

Pseudotime is an ordering, not a clock. A cell with pseudotime 0.9 is not necessarily older or more mature than a cell with pseudotime 0.8 in real time. The rate of transcriptional change along the trajectory can vary, and cells may spend different amounts of real time at different pseudotime positions. Report pseudotime as a relative ordering and avoid making claims about absolute timing.

Assuming a Single Trajectory in Heterogeneous Data

Some datasets contain multiple independent biological processes occurring simultaneously. A single trajectory model may force all cells onto one path, obscuring the true structure. Methods that model multiple trajectories or dynamic spanning forest mixtures can represent complex relationships better than single-tree approaches.

Ignoring the Cell Type Centric View Limitation

Most current resources connect disease associations only to discrete cell-type categories. This cell type centric view overlooks disease-relevant cellular processes, which constitute the primary gene expression programs underlying genetic risk. These cellular processes function along continuous cell-state trajectories and routinely extend beyond conventional cell type boundaries. When interpreting trajectory results, consider whether the relevant biology is better described by continuous processes instead of discrete cell types.

Relying on a Single Dimensionality Reduction Method

Conventional dimensionality reduction methods optimize different visualization objectives, resulting in tradeoffs between cluster separability, spatial organization, and temporal coherence. Relying on a single method may miss important aspects of the data structure. Methods that integrate outputs from multiple dimensionality reduction approaches, such as GIBOOST, can preserve both local and global structure more effectively.

Limitations of Trajectory Inference and Visualization

Snapshot Data Cannot Reveal True Dynamics

Most single-cell RNA sequencing experiments capture cells at a single time point. The resulting data are a snapshot of the transcriptional states present at that moment. Trajectory inference reconstructs a path through those states, but it cannot reveal the actual dynamics of the process, such as the rate of transition or the stability of intermediate states. Time-series data provide additional information, but even then, the sampling intervals may miss rapid transitions.

RNA Velocity Has Uncertainty

RNA velocity estimates the direction and speed of transcriptional change based on the ratio of unspliced to spliced mRNA. These estimates are noisy, and the uncertainty in velocity vectors can propagate to the trajectory inference. CellRank accounts for this uncertainty, but the resulting trajectories should still be validated with independent evidence.

Low-Density Regions Are Unreliable

Trajectory inference is least reliable in regions of the expression space where few cells are present. These regions may correspond to transitional states that are short-lived or difficult to capture. The inferred path through a low-density region is based on interpolation between distant cells and may not reflect the true biology.

Cluster-Based Methods Depend on Clustering Quality

Some trajectory methods, such as Slingshot, use clustering as a first step. The quality of the trajectory depends on the quality of the clustering. Poor clustering can create spurious branches or merge distinct lineages. The clustering parameters should be chosen carefully and the results validated.

Single-Cell Sequencing Has Throughput Limitations

Single-cell sequencing is frequently affected by omission due to limitations in sequencing throughput. Bulk RNA-seq may contain cells that are ostensibly omitted from single-cell data. The BulkTrajBlend algorithm, a component of the OmicVerse suite, leverages a Beta-Variational AutoEncoder for data deconvolution and graph neural networks for the discovery of overlapping communities to interpolate and restore the continuity of omitted cells within single-cell RNA sequencing datasets. This approach can help address the limitation that trajectory inference is only as complete as the cells captured in the single-cell data.

Deconvolution Methods Are Often Restricted to Discrete Cell Types

Current deconvolution methods are often restricted to discrete cell types and have limited power to make inferences about continuous cellular processes such as cell differentiation or immune cell activation. ConDecon provides a clustering-independent method for inferring the likelihood for each cell in a single-cell dataset to be present in a bulk tissue, representing an improvement in phenotypic resolution and functionality with respect to regression-based methods. This approach enables the deconvolution of other data modalities, such as bulk ATAC-seq data.

Safety and Regulatory Context for Trajectory Interpretation

Clinical Translation Requires Validation

Trajectory analyses of disease tissues can identify cell states and transitions that may be relevant to diagnosis or treatment. However, these findings are computational predictions that require experimental validation before they can inform clinical decisions. The DisCP-Atlas resource maps cellular processes to complex diseases and provides a framework for connecting trajectory-based findings to disease associations, but the associations are hypotheses, not established clinical facts.

The clinical relevance of trajectory analysis is illustrated by studies of T-cell exhaustion in hepatocellular carcinoma. T-cell exhaustion is a progressive decline in T-cell function due to continuous stimulation of the T-cell receptor in the presence of sustained antigen exposure. Single-cell RNA-seq analysis combined with bulk RNA sequencing was used to develop a TEX-based signature for predicting prognosis and immunotherapy response in HCC patients. The trajectory of T-cell exhaustion from functional to exhausted states is a continuous process, and understanding this trajectory has direct implications for patient prognosis and treatment decisions.

Reporting Standards for Publications

When reporting trajectory analyses in a publication, include the software versions, parameter settings, quality control metrics, and validation results. The methods section should be detailed enough that another researcher can reproduce the analysis. The Galaxy Training Network and nf-core documentation provide guidance on reproducible workflow practices that support these reporting standards.

Data Sharing and Reproducibility

Deposit the processed data and analysis code in a public repository so that other researchers can verify the results. The NCBI provides access to sequence data and analysis services that support data sharing. The EMBL-EBI training materials cover data resource usage and practical analysis education.

Educational Context

For researchers who need to build foundational skills in single-cell data analysis, the EMBL-EBI training materials provide learning pathways for bioinformatics, data-resource training, and practical analysis education. The Galaxy Training Network offers accessible workflow training and analysis tutorials. The Carpentries lessons provide foundational computing, data, shell, Git, and programming training. These resources support the development of the computational skills needed for rigorous trajectory analysis.

Professional Escalation Criteria

When to Seek Computational Expertise

If the trajectory inference produces unstable results across parameter settings, or if the inferred topology changes substantially with small changes in the input data, consult a computational biologist or bioinformatician with experience in trajectory analysis. The issue may be in the data preprocessing, the choice of method, or the interpretation of the results.

When to Question the Biological Interpretation

If the trajectory contradicts well-established biology, such as known marker gene expression patterns or experimentally validated cell state transitions, the trajectory inference may be incorrect. Re-examine the quality control, the root cell specification, and the parameter settings before accepting the result.

When to Use Specialized Methods

If the data include multiple samples or experimental conditions, standard trajectory methods that ignore sample variability can produce false discoveries. Lamian provides a statistical framework for differential pseudotime analysis that accounts for cross-sample variability and reduces sample-specific false discoveries. Use this approach when comparing trajectories across conditions.

When to Integrate Additional Data Modalities

If the trajectory inference from gene expression alone is ambiguous, consider integrating other data types. Single-nucleus Multiome data, which measures gene expression and chromatin accessibility in the same nuclei, can provide additional evidence for trajectory structure. The developing mouse visual cortex atlas demonstrates how transcriptomic and epigenomic data can be combined to reconstruct developmental trajectories.

When to Consider Alternative Visualization Frameworks

If standard dimensionality reduction methods fail to preserve both global and local structures, consider methods that integrate multiple dimensionality reduction outputs. GIBOOST systematically selects and integrates the two most informative dimensionality reduction methods by evaluating key visualization features, including separability, spatial continuity, uniformity, cellular dynamics, and cluster sensitivity. This approach can improve the visualization of differentiation trajectories and cell-cell interactions.

When to Question the Underlying Data Completeness

If the trajectory has large gaps or missing intermediate states, consider whether the single-cell data are missing cells that are present in the tissue. Bulk RNA-seq data may contain information about omitted cells, and methods like BulkTrajBlend can interpolate and restore the continuity of omitted cells within single-cell RNA sequencing datasets.

Frequently Asked Questions

What is the difference between pseudotime and real time?

Pseudotime is a computational ordering of cells along an inferred trajectory based on transcriptional similarity. It does not measure elapsed time. A cell with a higher pseudotime value is further along the inferred process, but the actual time required to progress from one pseudotime value to another is unknown and may vary across the trajectory. Pseudotime should be reported as a relative ordering, and claims about absolute timing should be avoided.

Why does my UMAP plot show clusters even though I expect a continuous process?

UMAP optimizes local neighborhood preservation, which can separate cells into apparent clusters even when the underlying biology is continuous. The clusters reflect local density differences in the expression space, not necessarily discrete cell types. Overlaying pseudotime or trajectory information onto the UMAP plot can reveal the continuous structure that the embedding obscures. Conventional dimensionality reduction methods optimize different visualization objectives, resulting in tradeoffs between cluster separability, spatial organization, and temporal coherence.

How do I choose between Monocle, Slingshot, and CellRank?

The choice depends on the biological question and the data structure. Monocle and Slingshot are appropriate when the direction of the process is known or can be specified. Slingshot uses cluster-based minimum spanning trees and can handle branching. CellRank is designed for scenarios where the direction is unknown, such as regeneration, reprogramming, and disease, and it combines trajectory inference with RNA velocity to determine direction automatically. For time-series data, CStreet explicitly uses sampling time to construct connections between cells within and between time points.

Can I use trajectory inference on single-nucleus RNA sequencing data?

Yes, trajectory inference can be applied to single-nucleus RNA sequencing data. The developing mouse visual cortex atlas used single-nucleus Multiome data to reconstruct developmental trajectories. The same quality control and validation principles apply, but the specific thresholds for quality metrics may differ from whole-cell data. Single-nucleus data can also be combined with single-cell RNA sequencing data to provide complementary views of the same biological process.

How do I validate a trajectory inference?

Validate the trajectory by checking that known marker genes change expression along the expected path. Plot marker gene expression against pseudotime and confirm that the trends match the known biology. If time-series data are available, compare the inferred pseudotime with the actual sampling time. Independent experimental validation, such as lineage tracing or functional assays, provides the strongest evidence. Resources like DisCP-Atlas can help connect trajectory-based findings to known cellular processes and disease associations.

What should I do if my trajectory changes dramatically with different parameters?

If the trajectory topology is not robust to parameter changes, the inference is unreliable. Perform a parameter sensitivity analysis and record the results. Consult a computational biologist if the instability persists. The issue may be in the data preprocessing, the choice of method, or the presence of batch effects or other technical variation. Consider whether the data contain multiple independent biological processes that a single trajectory model cannot represent.

How do I visualize branched trajectories?

For branched trajectories, color each cell by its assigned branch and use a secondary visual channel, such as point size or opacity, for pseudotime. This approach communicates both the lineage structure and the progression along each branch. Alternatively, use fate probabilities from methods like CellRank to show the likelihood that each cell will adopt each terminal state. Branch probabilities can also be combined with unsupervised clustering to characterize changes in gene expression intensity that reflect continuous cell states.

What are the limitations of trajectory inference from snapshot data?

Snapshot data capture cells at a single time point and cannot reveal the true dynamics of the process. Trajectory inference reconstructs a path through the observed states, but it cannot determine the rate of transition, the stability of intermediate states, or the direction of the process without additional information. RNA velocity and time-series data can provide additional evidence, but the results remain computational predictions that require validation. Single-cell sequencing throughput limitations can also lead to omitted cells, which may create gaps in the inferred trajectory.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.