Troubleshooting Trajectory Inference: Why Your Pseudotime Results Look Wrong and How to Fix Them
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Trajectory inference failures are predominantly upstream of the algorithm, stemming from data quality issues like excessive filtering of transitional cell states, which can disconnect trajectory components. Re-examining quality control thresholds and cell counts before filtering is crucial.
- Biologically implausible pseudotime orders often result from incorrect root cell specification; validation against marker gene expression and known developmental progression is essential to correct this.
- Dramatic trajectory changes between runs indicate algorithm stochasticity or unstable input features, necessitating the use of random seeds and thorough checks for feature stability across repeated analyses.
- Biologically implausible branching patterns can arise from over-clustering or inappropriate gene filtering, requiring a review of cluster resolution and confirmation of branch support via differential gene expression analysis.
- All cells assigned near-zero pseudotime suggests the root was placed at a terminal state or the dataset lacks true continuous variation, prompting inspection of root cell identity and assessment of data continuity.
- Trajectories ignoring known cell types may indicate missing populations after quality control or integration artifacts, necessitating verification of cell type composition pre- and post-processing steps.
Trajectory inference methods reconstruct developmental continua, cellular transitions, and dynamic gene expression programs from single-cell RNA sequencing (scRNA-seq) data. When pseudotime results appear biologically implausible, disconnected, or inconsistent across repeated runs, the cause usually lies upstream of the trajectory algorithm itself. This article provides a systematic troubleshooting framework for researchers who have completed a trajectory analysis and obtained results that do not align with known biology. The scope covers data inputs, quality control decisions, gene filtering, algorithm selection, root specification, interpretation limits, and reproducibility practices. The intended reader is a biology student, researcher, or laboratory professional who has basic familiarity with scRNA-seq analysis and needs concrete diagnostic steps instead of abstract theory.
At a Glance
The table below summarizes the most common trajectory inference failures, their typical causes, and the first diagnostic action to take. Use this table as a rapid triage tool before diving into detailed troubleshooting sections.
| Observed Problem | Most Likely Cause | First Diagnostic Step |
|---|---|---|
| Disconnected trajectory components | Excessive filtering removed transitional cell states | Re-examine QC thresholds and check cell counts per cluster before filtering |
| Pseudotime order contradicts known biology | Incorrect root cell specification | Validate root selection against marker gene expression and known developmental order |
| Trajectory changes dramatically between runs | Algorithm stochasticity or unstable input features | Set random seeds, check feature stability, and compare results across repeated runs |
| Biologically implausible branching patterns | Over-clustering or inappropriate gene filtering | Review cluster resolution and confirm branch support with differential expression |
| All cells assigned near-zero pseudotime | Root placed at terminal state or data lacks continuous variation | Inspect root cell identity and assess whether the dataset contains an actual continuum |
| Trajectory ignores known cell types | Missing cell populations after QC or integration artifacts | Verify cell type composition before and after filtering and integration |
Understanding What Trajectory Inference Actually Computes
Trajectory inference methods estimate a low-dimensional structure that represents cellular progression along a biological process such as differentiation, activation, or response to perturbation. The output is typically a pseudotime value for each cell, which orders cells along a path, and sometimes a branching structure that represents cell fate decisions. These methods do not measure real time. Pseudotime is a computational abstraction that places cells along a continuum based on transcriptional similarity. A cell with pseudotime 0 is the designated starting point, and cells with increasing values are interpreted as progressively more differentiated or activated.
The core assumption is that cells captured at different stages of a continuous process can be ordered by their gene expression profiles. This assumption fails when the underlying biology is discrete instead of continuous, when the captured cell population does not span the full transition, or when technical artifacts distort the expression measurements. Understanding this assumption is the first step in troubleshooting because many apparent algorithm failures are actually data failures.
Trajectory inference operates on the same fundamental inputs as other scRNA-seq analyses: a count matrix of genes by cells, quality-controlled cell annotations, and a set of features used for dimensionality reduction. The Bioconductor project provides extensive documentation on the packages and workflows used for these analyses, including trajectory inference tools and their dependencies. The Galaxy Training Network offers accessible tutorials that walk through the complete analysis path from raw counts to trajectory interpretation, which is useful for understanding where in the pipeline problems typically arise.
Data Input Quality: The Foundation of Reliable Trajectories
Poor quality input data produces poor trajectory results regardless of which algorithm is used. The most common source of trajectory failure is not the algorithm but the quality of the count matrix and the cell annotations fed into it. This section covers the specific quality control decisions that have outsized effects on trajectory inference.
Cell Viability and Ambient RNA Contamination
Low-quality cells, including dying cells and empty droplets, introduce noise that distorts the transcriptional continuum. Cells with high mitochondrial read fractions typically represent stressed or dying cells whose expression profiles are dominated by mitochondrial genes instead of the genes relevant to the biological transition being studied. When these cells are included in trajectory inference, they often form spurious branches or create artificial discontinuities.
The standard approach is to filter cells based on three metrics: total UMI count, number of detected genes, and mitochondrial read fraction. The thresholds for these metrics are dataset-specific and should be determined by examining the distributions instead of applying universal cutoffs. A common failure pattern is applying thresholds that are too permissive, which retains low-quality cells, or too stringent, which removes legitimate transitional states.
For trajectory inference specifically, the risk of over-filtering is particularly acute. Transitional cells often have lower total RNA content than fully differentiated cells because they are actively remodeling their transcriptomes. If the lower bound on UMI count or gene number is set too high, the very cells that define the trajectory path are removed, and the remaining cells form disconnected clusters that no algorithm can connect into a continuous path.
The NCBI Data Resources provide access to reference databases and sequence resources that support quality assessment, including annotation files and reference genomes used in alignment and quantification. Ensuring that the reference resources are current and appropriate for the organism and tissue being studied is a prerequisite for accurate count matrices.
Doublet Detection and Removal
Doublets, which are droplets containing two or more cells, create expression profiles that are mixtures of distinct cell types. In trajectory inference, doublets appear as intermediate states that do not exist biologically. They can create false branches, connect unrelated cell types, or distort pseudotime ordering by introducing cells whose expression is an average of two different states.
Doublet detection should be performed before trajectory inference, and the results should be examined in the context of the expected biology. A doublet rate of 1 to 5 percent is typical for standard microfluidic platforms, but the rate varies with loading density and cell size. The key diagnostic is whether the cells identified as doublets express marker genes from two distinct lineages that are not expected to share a common transitional state.
Ambient RNA and Background Contamination
Ambient RNA, which is free-floating RNA in the cell suspension, contaminates all droplets to some degree. This contamination adds a background expression signal that is uniform across cells, which compresses the transcriptional differences between cell states. In trajectory inference, ambient RNA can make distinct cell types appear more similar than they are, leading to false connections or incorrect branch points.
The severity of ambient RNA contamination varies with the tissue type, dissociation protocol, and platform. Computational methods exist to estimate and remove ambient RNA, but these methods require careful parameter tuning. The diagnostic sign of ambient RNA contamination is the presence of ubiquitously expressed genes, such as housekeeping genes, at similar levels across all cells, including cells where those genes should not be expressed.
Gene Filtering and Feature Selection
The genes used as input to dimensionality reduction and trajectory inference determine which biological signals are captured. Filtering decisions made early in the analysis have downstream consequences that are often not apparent until the trajectory results are examined.
Removing Genes with No Informative Signal
Genes that are not detected in any cell, or that are detected in only a handful of cells with minimal expression, contribute noise instead of signal. Standard practice is to remove genes expressed in fewer than a minimum number of cells and genes with very low total counts. The thresholds depend on the sequencing depth and the expected biology. For trajectory inference, the concern is that including too many uninformative genes dilutes the signal from the genes that actually define the transition.
The opposite failure is filtering too aggressively, removing genes that are expressed at low levels but are biologically important for the transition. Transcription factors and signaling molecules are often expressed at low levels, and their dynamic changes may define the trajectory. If these genes are removed during filtering, the trajectory algorithm cannot detect the biological process they control.
Selecting Highly Variable Genes
Most trajectory inference workflows use highly variable genes as input to principal component analysis. The selection of highly variable genes is a critical step because it determines which biological variation is retained. The standard approach identifies genes whose expression variability exceeds what would be expected from the mean-variance relationship of the data.
A common failure pattern is selecting too few highly variable genes, which captures only the most dominant sources of variation, often cell cycle or stress responses, while missing the subtler changes that define the trajectory. Another failure is selecting too many, which reintroduces noise. The optimal number depends on the dataset complexity and the biological process being studied.
For trajectory inference, the highly variable gene selection should be evaluated in the context of the expected biology. If the genes known to drive the transition are not among the selected highly variable genes, the feature selection is likely missing the relevant signal. This can happen when the transition involves coordinated but modest changes in many genes instead of large changes in a few genes.
Cell Cycle and Other Confounding Variation
Cell cycle genes are among the most highly variable genes in dividing cell populations. When these genes dominate the highly variable gene set, the resulting trajectory may order cells by cell cycle phase instead of by the biological process of interest. This produces pseudotime values that correlate with cell cycle markers and do not reflect the intended developmental or activation trajectory.
The diagnostic is to check whether the pseudotime ordering correlates with known cell cycle genes. If it does, the analysis should be repeated with cell cycle genes removed from the feature set or with cell cycle variation regressed out. The same logic applies to other sources of confounding variation, including mitochondrial gene expression, stress response genes, and batch effects.
Dimensionality Reduction and Its Impact on Trajectory Structure
Trajectory inference algorithms typically operate on a reduced-dimensional representation of the data, most commonly principal component analysis. The number of principal components retained is a key parameter that shapes the trajectory structure.
Choosing the Number of Principal Components
Too few principal components discard biological signal and force the trajectory algorithm to work with an incomplete representation of the data. Too many principal components retain noise, which can create spurious branches or distort the path. The optimal number is dataset-specific and should be determined by examining the variance explained by each component and the stability of the trajectory across different component numbers.
A practical diagnostic is to run the trajectory inference with a range of principal component numbers and compare the results. If the trajectory structure changes dramatically across a reasonable range, the data may not support a robust trajectory, or the feature selection may be capturing noise instead of signal.
Batch Effects and Integration
When data are collected across multiple batches, sequencing runs, or samples, batch effects introduce technical variation that can obscure biological trajectories. Integration methods aim to remove batch effects while preserving biological variation. The choice of integration method and its parameters has a major impact on trajectory inference.
A common failure pattern is over-integration, where the method removes genuine biological differences between samples along with the technical batch effects. This produces a trajectory that is artificially compressed, with distinct cell states pulled together into a single cloud. The diagnostic is to check whether known cell type markers are preserved after integration and whether the trajectory separates cell states that should be distinct.
Under-integration is the opposite failure, where batch effects remain and the trajectory separates cells by batch instead of by biological progression. The diagnostic is to color the trajectory by batch and check whether cells from different batches form separate branches or clusters.
The nf-core documentation describes community standards for reproducible analysis pipelines, including integration and quality control steps. Following established pipeline conventions can reduce the risk of integration artifacts, but the parameters still need to be evaluated in the context of each dataset.
Algorithm Selection and Parameter Choices
Different trajectory inference algorithms make different assumptions about the underlying structure of the data. Choosing an algorithm that does not match the biological process being studied produces results that appear wrong even when the algorithm is functioning correctly.
Linear Versus Branched Trajectories
Some algorithms assume a linear trajectory, ordering cells along a single path. Others model branching structures, allowing cells to diverge into multiple fates. If the biological process involves a branch point, such as a common progenitor giving rise to two distinct lineages, a linear algorithm will force the two lineages into a single path, producing pseudotime values that are biologically meaningless for one or both branches.
The diagnostic is to examine the dimensionality reduction plot and assess whether the cell cloud has a clear branching structure. If branches are visible, a branching algorithm should be used. If the structure is linear, a branching algorithm may introduce spurious branches that are not supported by the data.
Root Cell Specification
Most trajectory algorithms require the user to specify the root cell or root state, which defines the starting point of the trajectory. Incorrect root specification is one of the most common causes of biologically implausible pseudotime results.
The root should be the cell state that represents the earliest stage of the biological process being studied. For a differentiation trajectory, this is typically the progenitor or stem cell state. For an activation trajectory, this is the naive or resting state. If the root is placed at a terminal state, the pseudotime ordering will be reversed, and the interpretation of which genes are upregulated or downregulated along the trajectory will be inverted.
The diagnostic is to examine the expression of known marker genes along the pseudotime axis. If markers of the earliest state are highest at the end of the trajectory instead of the beginning, the root is likely specified incorrectly. Some algorithms allow the root to be inferred from the data, but these methods can fail when the data do not contain a clear root state or when the root state is rare.
Algorithm Stochasticity and Reproducibility
Many trajectory inference algorithms incorporate stochastic elements, either in the initialization of the model or in the optimization procedure. This means that repeated runs on the same data can produce slightly different results. When the differences are large, the trajectory is unstable, and the results should not be interpreted with confidence.
The standard practice is to set a random seed before running the algorithm and to document the seed in the analysis record. The trajectory should be run multiple times with different seeds to assess stability. If the trajectory structure changes substantially across runs, the data may not support a robust trajectory, or the algorithm parameters may need adjustment.
The The Carpentries lessons provide foundational training in reproducible computing practices, including the use of random seeds and version control. These practices are essential for trajectory inference, where small changes in the input or parameters can produce large changes in the output.
Practical Workflow for Diagnosing Trajectory Failures
This section provides a step-by-step workflow for diagnosing and correcting trajectory inference problems. The workflow is designed to be followed in order, with each step building on the previous one.
Step 1: Verify the Input Data
Before examining the trajectory output, verify that the input data are correct. Check the count matrix dimensions, the cell annotations, and the gene annotations. Confirm that the data have been properly normalized and that the normalization method is appropriate for the platform and protocol used.
Examine the quality control metrics for each cell, including total UMI count, gene count, and mitochondrial fraction. Plot these metrics and look for outliers. Compare the number of cells passing quality control with the expected number based on the experimental design.
Step 2: Assess the Dimensionality Reduction
Plot the cells in the reduced-dimensional space, typically using UMAP or t-SNE. Examine the structure of the cell cloud. Is there a continuous path from one cell state to another, or are the cells separated into distinct clusters with no obvious connection?
Color the cells by known cell type markers, batch, sample, and cell cycle phase. This reveals whether the dominant structure in the data corresponds to the biological process of interest or to technical artifacts.
Step 3: Evaluate Feature Selection
List the highly variable genes used as input to the dimensionality reduction. Check whether the genes known to drive the biological transition are included. If they are missing, the feature selection is likely capturing the wrong signal.
Repeat the dimensionality reduction with different numbers of highly variable genes and different numbers of principal components. Compare the resulting structures. A robust trajectory should be relatively stable across reasonable parameter choices.
Step 4: Run the Trajectory Algorithm with Multiple Settings
Run the trajectory algorithm with different parameter settings, including different root cells, different numbers of principal components, and different algorithm choices. Compare the resulting trajectories. Note which aspects of the trajectory are stable across settings and which change.
Set a random seed and run the algorithm multiple times to assess stochasticity. Document the seed and the results of each run.
Step 5: Validate the Trajectory with Independent Markers
Identify marker genes for each cell state along the expected trajectory. Plot the expression of these markers against pseudotime. The markers of the earliest state should be highest at the beginning of the trajectory, and markers of later states should increase at the appropriate positions.
Check whether the branch points in the trajectory correspond to known cell fate decisions. If a branch point separates two cell types that are not known to share a common progenitor, the branch may be an artifact.
Step 6: Document and Report
Record all parameters, seeds, and filtering decisions in the analysis documentation. Report the trajectory results with the caveats about the assumptions made and the limitations of the data. Include the diagnostic plots used to validate the trajectory.
Common Failure Patterns and Their Corrections
This section describes specific failure patterns observed in trajectory inference and the corrections that typically resolve them.
Disconnected Trajectory Components
When the trajectory algorithm produces multiple disconnected segments instead of a single continuous path, the usual cause is that intermediate cell states are missing from the data. This can happen when the cell population captured does not span the full transition, when quality control filtering removed the intermediate states, or when the sequencing depth is insufficient to detect the lowly expressed genes that define the intermediate states.
The correction depends on the cause. If filtering removed intermediate states, relax the filtering thresholds and repeat the analysis. If the cell population does not span the transition, additional samples or time points may be needed. If sequencing depth is insufficient, deeper sequencing may be required.
Pseudotime Order Contradicts Known Biology
When the pseudotime ordering places known early markers at the end of the trajectory and known late markers at the beginning, the root is likely specified incorrectly. Re-examine the root cell selection and confirm that the root represents the earliest state in the biological process.
In some cases, the contradiction arises because the trajectory algorithm has ordered cells along a different axis of variation than the biological process of interest. This can happen when a confounding source of variation, such as cell cycle or batch, dominates the feature set. Remove the confounding variation and repeat the analysis.
Trajectory Changes Dramatically Between Runs
When repeated runs of the same algorithm on the same data produce substantially different trajectories, the analysis is unstable. This instability can arise from stochastic algorithm components, from a feature set that is dominated by noise, or from a data set that does not contain a clear trajectory structure.
The correction is to stabilize the analysis by setting random seeds, reducing the feature set to the most informative genes, and increasing the number of cells if possible. If the trajectory remains unstable after these corrections, the data may not support trajectory inference, and the results should be interpreted with caution or not reported as a trajectory.
Biologically Implausible Branching Patterns
When the trajectory contains branches that do not correspond to known cell fate decisions, the branches may be artifacts of over-clustering or of including cells that do not belong to the biological process being studied. Examine the cells in each branch and check whether they express markers of distinct cell types or whether they represent a continuum that has been artificially split.
The correction is to reduce the clustering resolution, remove contaminating cell types, or use a simpler trajectory model that does not assume branching.
All Cells Assigned Near-Zero Pseudotime
When all cells receive pseudotime values close to zero, the trajectory algorithm has failed to find a meaningful ordering. This can happen when the data do not contain a continuous transition, when the root is placed at a state that is transcriptionally similar to all other cells, or when the feature set does not capture the variation that defines the trajectory.
The correction is to verify that the data contain a genuine continuum, to re-examine the root selection, and to evaluate the feature set for the presence of informative genes.
Trajectory Ignores Known Cell Types
When the trajectory does not include cell types that are known to be part of the biological process, those cell types may have been removed during quality control, may have been lost during integration, or may not be captured by the trajectory algorithm because they are transcriptionally distinct from the main path.
The correction is to verify the cell type composition before and after each filtering and integration step. If the cell types are present in the data but not in the trajectory, the algorithm parameters may need adjustment, or a different algorithm may be needed.
Records and Measurements for Trajectory Troubleshooting
Maintaining detailed records of the analysis decisions and their rationale is essential for troubleshooting trajectory inference. The following records should be kept for every trajectory analysis.
Analysis Documentation
Document the version of every software package used, including the trajectory algorithm, the normalization method, the integration method, and the visualization tools. Record the random seeds used for each run. Document the filtering thresholds, the number of highly variable genes, the number of principal components, and the root cell specification.
The EMBL-EBI Training resources provide guidance on documenting bioinformatics analyses and on using public data resources effectively. Following these practices ensures that the analysis can be reproduced and that troubleshooting can be performed systematically.
Diagnostic Plots
Save the diagnostic plots generated during the analysis, including the quality control plots, the dimensionality reduction plots, the marker gene expression plots, and the trajectory plots with cells colored by batch, sample, and cell type. These plots are essential for identifying the cause of trajectory failures and for communicating the results to collaborators.
Parameter Sensitivity Records
Record the results of running the trajectory algorithm with different parameter settings. This includes the trajectory structure, the pseudotime values, and the stability assessment. These records demonstrate that the reported trajectory is robust to reasonable parameter choices.
Version Control
Use version control for the analysis scripts and configuration files. This allows the analysis to be reproduced exactly and enables the identification of changes that may have introduced errors. The The Carpentries lessons provide training in version control with Git, which is a standard practice for reproducible bioinformatics.
Limitations of Trajectory Inference
Trajectory inference has inherent limitations that cannot be fully resolved by troubleshooting. Understanding these limitations is essential for interpreting results correctly and for avoiding overinterpretation.
Pseudotime Is Not Real Time
Pseudotime is a computational ordering based on transcriptional similarity. It does not measure the actual time a cell has spent in a particular state, and it does not provide information about the rate of transition between states. Two cells with pseudotime values of 5 and 10 are not necessarily separated by the same amount of real time as two cells with values of 50 and 55.
Trajectory Inference Cannot Establish Causality
A trajectory that orders cells along a continuum does not prove that the cells actually transition from one state to another in that order. The trajectory is a hypothesis about the relationship between cell states, and it must be validated with independent experiments, such as lineage tracing or time-course studies.
The Trajectory Depends on the Captured Cell Population
The trajectory structure depends on which cells are captured and sequenced. If the cell population does not include the full range of transitional states, the trajectory will be incomplete or disconnected. Adding more cells or more time points can improve the trajectory, but the results are always conditional on the input data.
Algorithm Assumptions Shape the Results
Every trajectory algorithm makes assumptions about the structure of the data, such as whether the trajectory is linear or branched, whether the root is known or inferred, and whether the cells are sampled uniformly along the trajectory. When these assumptions do not match the biology, the results will be misleading regardless of how carefully the analysis is performed.
Safety and Reproducibility Context
Trajectory inference is a computational analysis, and the primary risks are not physical but scientific. The main risks are overinterpretation of results, failure to document the analysis, and reporting results that cannot be reproduced.
Reproducibility Standards
Reproducibility requires that the analysis can be repeated with the same inputs and parameters to produce the same results. This requires documenting all software versions, parameters, and random seeds. The nf-core documentation describes community standards for reproducible pipelines, and the Galaxy Training Network provides tutorials that emphasize reproducible analysis practices.
Data Management
The raw sequencing data, the processed count matrices, and the analysis scripts should be stored and organized so that the analysis can be reproduced. The NCBI Data Resources provide repositories for raw and processed data, and the EMBL-EBI Training resources provide guidance on data management and sharing.
Professional Escalation Criteria
When trajectory results are biologically implausible and the troubleshooting steps described in this article do not resolve the issue, escalate the problem to a colleague with expertise in computational biology or bioinformatics. This is particularly important when the results are intended for publication or for guiding experimental decisions.
Escalate when the trajectory structure changes dramatically across repeated runs, when the pseudotime ordering contradicts well-established biology despite correct root specification, or when the trajectory cannot be validated with independent marker genes. Also escalate when the analysis involves a new algorithm or a new data type that has not been previously validated in the laboratory.
A Decision Framework for Choosing Between Trajectory Repair and Data Recollection
When troubleshooting efforts do not resolve trajectory failures, the central question becomes whether to continue repairing the existing dataset or to collect new data. Researchers often spend excessive time adjusting parameters on data that fundamentally cannot support the intended trajectory. This section provides a structured decision framework for determining when further computational repair is justified and when the experimental design itself requires revision.
The Data Sufficiency Assessment
Before investing additional time in parameter tuning, assess whether the existing dataset contains the information needed to reconstruct the intended trajectory. This assessment requires answering three questions about the data collection process.
First, does the cell population span the full biological transition? A trajectory can only be reconstructed between states that are actually present in the data. If the experiment captured only the starting and ending states without intermediate cells, no algorithm can fill the gap. The diagnostic is to examine the dimensionality reduction plot and count the number of cells in the region between known cell states. If this region contains very few cells, the trajectory will be poorly supported regardless of algorithm choice.
Second, was the sequencing depth sufficient to detect the genes that define the transition? Trajectory inference depends on detecting dynamic changes in gene expression along the path. If sequencing depth is too shallow, lowly expressed transcription factors and signaling molecules that drive the transition may be missed entirely. The diagnostic is to check whether known transition marker genes are detected in the expected cell populations. If these genes are absent from the count matrix, deeper sequencing or a different experimental approach may be required.
Third, does the experimental design capture the appropriate time points or stages? For time-course experiments, the spacing between collection points determines whether intermediate states are captured. For snapshot experiments, the composition of the sampled population determines whether the full continuum is represented. If the experimental design missed critical intermediate stages, no computational method can recover them.
The Repair Cost Matrix
Once the data sufficiency assessment is complete, use the repair cost matrix to decide between computational repair and data recollection. The matrix compares the cost of each repair option against the likelihood of success.
| Data Condition | Computational Repair Likelihood | Recommended Action |
|---|---|---|
| Missing intermediate cell states | Low, algorithms cannot create absent states | Collect additional time points or enrich for transitional populations |
| Excessive ambient RNA contamination | Moderate, computational removal may help | Attempt ambient RNA removal, then validate with known markers |
| Severe batch effects across samples | Moderate to high, integration may resolve | Test multiple integration methods and validate marker preservation |
| Incorrect root specification | High, this is a parameter error | Re-specify root and validate with marker genes |
| Inappropriate algorithm choice | High, this is a model mismatch | Switch to an algorithm matching the expected trajectory structure |
| Stochastic instability across runs | Moderate, may indicate weak signal | Assess whether the biological signal is strong enough to support inference |
| Low sequencing depth | Low, information is physically absent | Collect deeper sequencing data or use targeted enrichment |
The key distinction is between errors of analysis and errors of data collection. Analysis errors, such as incorrect root specification or algorithm mismatch, are correctable with reasonable effort. Data collection errors, such as missing cell states or insufficient sequencing depth, require new experiments.
The Escalation Decision Point
Define a clear escalation point before beginning troubleshooting. A practical approach is to limit computational repair attempts to a fixed number of parameter combinations, typically three to five, and to document the results of each attempt. If none of these attempts produces a trajectory that is stable across runs and consistent with known biology, the data likely cannot support the intended analysis.
At this escalation point, the decision shifts from computational troubleshooting to experimental redesign. The specific redesign depends on the identified data deficiency. If intermediate states are missing, consider collecting additional time points or using cell sorting to enrich for transitional populations. If sequencing depth is insufficient, increase sequencing coverage or use a targeted panel that captures the genes of interest. If the cell population is too heterogeneous, consider isolating the relevant cell types before sequencing.
The Validation Threshold for Continuing Analysis
Before accepting any trajectory result, apply a validation threshold that must be met for the analysis to proceed. This threshold has three components.
The first component is marker gene concordance. The pseudotime ordering must place known early markers at the beginning and known late markers at the end. If the ordering is reversed or inconsistent, the root specification or the feature set is likely incorrect.
The second component is branch point correspondence. Any branch points in the trajectory must correspond to known cell fate decisions. A branch that separates cell types not known to share a common progenitor is likely an artifact of over-clustering or of including contaminating cell types.
The third component is parameter stability. The trajectory structure should remain consistent across reasonable changes in the number of highly variable genes, the number of principal components, and the random seed. If the trajectory changes substantially with small parameter changes, the result is not robust enough to support biological interpretation.
The Documentation Requirement for the Decision
Document the decision process itself, beyond the final analysis. Record the data sufficiency assessment results, the repair attempts made, the parameter combinations tested, and the rationale for either continuing with the repaired analysis or collecting new data. This documentation serves two purposes.
First, it prevents repeated troubleshooting of the same data by different team members. When the decision to collect new data is documented with the specific reasons, future analysts will not waste time attempting the same repairs that already failed.
Second, it provides the evidence needed for publication or for discussions with collaborators. Reviewers and collaborators will want to know why the analysis took a particular direction and why certain decisions were made. The EMBL-EBI Training resources provide guidance on documenting bioinformatics analyses in a way that supports reproducibility and transparent reporting.
The Role of External Validation
When the decision is to continue with a computationally repaired trajectory, external validation becomes essential. This validation can take several forms. Independent experiments, such as lineage tracing or time-course studies, can confirm that the inferred trajectory reflects real biological progression. The NCBI Data Resources provide access to public datasets that can be used to check whether the inferred trajectory is consistent with data from other studies of the same biological process.
External validation is particularly important when the trajectory inference produced unexpected results. An unexpected trajectory is not necessarily wrong, but it requires stronger evidence before it can be reported with confidence. The validation experiments should be designed to test the specific predictions of the trajectory, such as the existence of an intermediate state or the order of gene expression changes along the path.
The Cost of Continuing with Unreliable Trajectories
Continuing with an unreliable trajectory carries scientific costs that extend beyond the immediate analysis. A trajectory that is not robust to parameter changes or that contradicts known biology can mislead downstream analyses, including differential expression testing, gene regulatory network inference, and experimental design decisions. The time saved by avoiding data recollection is often outweighed by the time lost to interpreting and defending results that cannot be validated.
The Galaxy Training Network and the nf-core documentation both emphasize the importance of reproducible and well-documented analyses. These resources support the broader principle that the quality of the final biological interpretation depends on the quality of the underlying data and the rigor of the analysis decisions.
The Practical Implementation of the Framework
Implement this decision framework at the start of a trajectory analysis project, not after the first failure. Define the data sufficiency criteria before collecting data, establish the repair cost matrix before beginning analysis, and set the escalation point before running the first trajectory algorithm. This proactive approach prevents the common pattern of spending weeks adjusting parameters on data that cannot support the intended analysis.
When the framework indicates that data recollection is necessary, treat this as a normal outcome instead of a failure. Many trajectory inference projects require multiple rounds of data collection before the data quality and composition support a reliable trajectory. The documentation from the initial analysis informs the design of the next experiment, making each subsequent attempt more likely to succeed.
Frequently Asked Questions
What is the most common cause of disconnected trajectories in scRNA-seq data?
The most common cause is the removal of intermediate cell states during quality control filtering. Transitional cells often have lower RNA content than fully differentiated cells, and stringent filtering thresholds can eliminate them. Relax the filtering thresholds and check whether the intermediate states reappear. Also verify that the cell population captured in the experiment actually spans the full transition.
How do I choose the correct root cell for trajectory inference?
The root cell should represent the earliest stage of the biological process being studied. For differentiation trajectories, this is typically the progenitor or stem cell state. Validate the root selection by checking that markers of the earliest state are highest at the beginning of the trajectory and that markers of later states increase along the pseudotime axis. If the ordering is reversed, the root is likely incorrect.
Why does my trajectory change every time I run the algorithm?
Many trajectory algorithms have stochastic components, and repeated runs can produce different results. Set a random seed before running the algorithm and document the seed. Run the algorithm multiple times with different seeds to assess stability. If the trajectory changes substantially across runs, the data may not support a robust trajectory, or the feature set may be dominated by noise.
How many highly variable genes should I use for trajectory inference?
There is no universal number. The optimal number depends on the dataset complexity and the biological process being studied. A practical approach is to test a range of values and compare the resulting trajectories. The trajectory should be relatively stable across reasonable choices. Check that the genes known to drive the biological transition are included in the selected feature set.
What should I do if my trajectory orders cells by cell cycle phase instead of the biological process of interest?
Cell cycle genes are often among the most highly variable genes and can dominate the feature set. Remove cell cycle genes from the feature set or regress out cell cycle variation before running the trajectory algorithm. Verify that the pseudotime ordering no longer correlates with cell cycle markers.
Can I use trajectory inference on data that has been integrated across multiple batches?
Yes, but integration must be performed carefully. Over-integration can remove genuine biological differences, while under-integration leaves batch effects that distort the trajectory. After integration, verify that known cell type markers are preserved and that cells from different batches are mixed along the trajectory instead of separated by batch.
How do I know if my trajectory results are reliable enough to report?
A reliable trajectory should be stable across repeated runs with different random seeds, robust to reasonable changes in parameters, and consistent with known biology. Validate the trajectory with independent marker genes and check that branch points correspond to known cell fate decisions. Document all parameters and seeds so that the analysis can be reproduced.
What should I do if the trajectory results contradict well-established biology?
First, verify that the root cell is specified correctly and that the feature set includes the genes known to drive the transition. Check for confounding variation from cell cycle, batch, or ambient RNA. If the contradiction persists after these checks, escalate the problem to a colleague with computational biology expertise before drawing conclusions or reporting the results.
Related Bioinformatics Guides
- How to Interpret Gene Set Enrichment Analysis Results
- Genomic Data Analysis Tools: A Comparative Guide for Researchers
- Metabolomics Data Analysis in R: A Practical Workflow
- Microbiome Data Analysis in R: A Practical Guide for Compositional Data
- Radiomics Feature Selection: Methods and Best Practices
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- Deep Learning Tools for Single Quantum Dot Tracking.. Methods in molecular biology (Clifton, N.J.), 2026.
- StrainMapperJ: An Easy-to-Use Digital Image Correlation Toolkit for Exploring and Quantifying the Mechanics of Deforming Tissues.. Results and problems in cell differentiation, 2026.
- Caring for Patients Receiving Continuous Renal Replacement Therapy in the Intensive Care Unit: A Qualitative Study.. Nursing in critical care, 2026.
- Computational protocol for quantifying body-bending amplitude and period in Caenorhabditis elegans.. 2026.
- Protocol to explore olfactory placode morphogenesis using an agent-based model.. 2026.
- Anatomical Approaches for Insertion Sites of External Ventricular Drainage Catheters.. 2026.
- LINC: a framework for maintaining high-quality passive data in digital phenotyping studies.. 2026.
- Protocol for identifying recirculating thymic regulatory T cells and characterizing the role of Eos in their function using scRNA-seq and TCR-seq.. 2026.
- "We learn while doing": informal training experiences of critical care nurses managing dialysis technologies in intensive care units.. 2026.
- Protocol for generation, time-course imaging, and automated quality control of 3D spheroid invasion using TRACEQC.. 2026.
- Ship damage submarine cable accident investigation: trajectory analysis based on computer chart operation of VTS monitoring records. Conference on Computer Graphics, Artificial Intelligence, and Data Processing, 2023.
- Failure Analysis of Open Circuit at Output Terminal of the Voltage Comparator. International Conference on Electronic Packaging Technology, 2024.
- Design and Realization of Virtual Training System for New Type Howitzer. 2012.
- Analysis of Log Files to Enable Smart-Troubleshooting in Industry 4.0: A Systematic Mapping Study. IEEE Access, 2024.
- AI-Empowered Speed Extraction via Port-Like Videos for Vehicular Trajectory Analysis. IEEE transactions on intelligent transportation systems (Print), 2023.
- Estimation of instantaneous frequency in the presence of interfering trajectories. I2mtc 2020 International Instrumentation and Measurement Technology Conference Proceedings, 2020.
- Trajectory engine: A backend for trajectory sampling. IEEE Symposium Record on Network Operations and Management Symposium, 2002.
- Anomaly Detection and Root Cause Analysis of Retainability Contextual Anomalies in Radio Access Networks. Lecture Notes in Computer Science, 2027.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.