# Gene-Level Alignment of Single-Cell Trajectories: How to Compare Differentiation Programs Across Conditions


## Key Takeaways

- Pseudotime alone is insufficient for cross-condition comparison due to arbitrary scaling, differing gene ordering, and complex branched structures, necessitating explicit trajectory alignment to establish correspondences between developmental states.
- Dynamic programming, particularly dynamic time warping, provides a foundational approach by finding optimal paths through expression matrices, though it can force spurious matches when trajectories diverge significantly.
- Bayesian information-theoretic frameworks like Genes2Genes address the limitations of definitive matching by modeling sequential matches and mismatches, enabling the identification of condition-specific gene dynamics and biological deficiencies in in vitro systems.
- Multi-omic approaches, such as ArchVelo integrating chromatin accessibility and transcriptomic data, enhance alignment accuracy by leveraging regulatory program archetypes, revealing trajectory features invisible in expression-only analyses.
- Deep learning and generative models, including variational autoencoders and conditional flow matching, offer advanced trajectory inference and forecasting capabilities by learning complex velocity fields, but require substantial computational resources and accurate splicing quantification.
- Robust trajectory alignment necessitates reproducible data processing pipelines, independent trajectory inference per condition, careful examination of gene-level alignment patterns, and rigorous validation against biological plausibility and orthogonal measurements.

---

## Direct Answer and Scope

Researchers comparing gene expression dynamics along pseudotime trajectories from different conditions face a fundamental problem: pseudotime scales are arbitrary, and the same biological process may proceed at different rates, with different gene ordering, or with condition-specific branches. Gene-level trajectory alignment addresses this by establishing correspondences between positions on a reference trajectory and positions on a query trajectory, then evaluating how individual genes behave across those matched positions. This article explains the conceptual basis of trajectory alignment, describes the main computational approaches including dynamic programming frameworks and deep learning methods, and provides a practical workflow for designing, executing, and interpreting gene-level comparisons across conditions. The target reader is a biology student, researcher, or laboratory professional who has generated single-cell RNA sequencing data and needs to determine whether a differentiation protocol, genetic perturbation, or disease state alters the sequence and timing of gene expression changes along a developmental path.

## Why Pseudotime Alone Is Insufficient for Cross-Condition Comparison

Single-cell RNA sequencing captures gene expression in individual cells, but each cell is measured once and destroyed. Researchers reconstruct dynamic processes by ordering cells along a trajectory based on transcriptional similarity, a process called pseudotime inference. The resulting pseudotime coordinate places each cell along a path from an initial state to a terminal state. However, pseudotime values are not directly comparable across datasets or conditions.

The first problem is scale. Pseudotime is typically normalized to a range such as zero to one, but the actual biological time represented by that range may differ substantially between conditions. A differentiation protocol that takes five days in vitro may produce a pseudotime trajectory that looks identical in length to one that takes three days. Without alignment, a researcher cannot determine whether a gene that peaks at pseudotime 0.5 in the reference condition is peaking at the same developmental stage in the query condition or merely at the same fraction of an arbitrarily scaled axis.

The second problem is ordering. Even when two conditions produce the same terminal cell type, the sequence of intermediate states may differ. A gene that is activated early in one condition may be activated late in another. Simple pseudotime binning will obscure these differences because both conditions place the gene somewhere along the trajectory, but the relative position carries no cross-condition meaning.

The third problem is branch structure. Many differentiation processes are not linear. Cells may choose between two or more fates, producing branched trajectories. Comparing branched trajectories across conditions requires aligning also positions along a path but also determining which branches correspond to which. This is substantially more complex than aligning linear paths.

Trajectory alignment methods address these problems by computing an explicit mapping between positions on a reference trajectory and positions on a query trajectory. The mapping is optimized according to a cost function that rewards matching similar expression patterns and penalizes mismatches. Once the mapping is established, gene-level comparison becomes possible: for each gene, the researcher can ask whether its expression dynamics along the reference trajectory match its dynamics along the query trajectory after the alignment is applied.

## Core Principles of Trajectory Alignment

### Dynamic Programming as the Foundational Approach

Dynamic programming is a computational technique that solves optimization problems by breaking them into overlapping subproblems and storing intermediate results. In the context of trajectory alignment, dynamic programming finds the optimal path through a matrix where rows represent positions on the reference trajectory and columns represent positions on the query trajectory. Each cell in the matrix contains a cost associated with matching the corresponding reference and query positions. The algorithm finds the lowest-cost path from the start of both trajectories to the end, allowing for local stretching and compression of time.

The classic formulation of this approach is dynamic time warping, originally developed for speech recognition but widely applied to biological sequence and trajectory comparison. Dynamic time warping allows a position on the reference trajectory to match multiple consecutive positions on the query trajectory, and vice versa, accommodating differences in the rate of progression through developmental states.

A limitation of standard dynamic programming approaches is the assumption that every position on the reference trajectory has a corresponding position on the query trajectory. In biological data, this assumption often fails. A gene may be expressed in one condition but entirely absent in another. A developmental state present in the reference may have no counterpart in the query. When the assumption of definitive matching is violated, standard dynamic programming can force spurious alignments that distort the biological interpretation.

### Bayesian Information-Theoretic Frameworks

The Genes2Genes method addresses the definitive match limitation by framing trajectory alignment within a Bayesian information-theoretic framework. instead of assuming every reference position must match a query position, Genes2Genes explicitly models sequential matches and mismatches between a reference and a query trajectory at single-gene resolution. This allows the alignment to capture regions where the two trajectories diverge, identifying genes that follow different dynamics across conditions.

The method produces clusters of genes with varying patterns of expression dynamics, enabling researchers to distinguish genes that are conserved between conditions from genes that are condition-specific. In a proof-of-concept application, Genes2Genes revealed that T cells differentiated in vitro matched an immature in vivo state while lacking expression of genes associated with TNF signaling. This type of finding has direct practical value: it pinpoints the specific molecular deficiencies of an in vitro system and suggests which genes or pathways must be modulated to improve culture conditions.

### Archetypal and Multi-Omic Approaches

Trajectory alignment becomes more complex when multiple data modalities are available. ArchVelo models gene regulation and infers trajectories from paired single-cell chromatin accessibility and transcriptomic data. The method represents chromatin accessibility as archetypes, which are shared regulatory programs, and models their dynamic influence on transcription. This approach improves trajectory inference accuracy and gene-level latent time alignment compared to methods that use only transcriptomic data.

The practical implication is that alignment quality depends on the information available to define the trajectory. When chromatin accessibility data are available, the regulatory programs that drive gene expression changes can be identified, and the trajectory can be aligned based on regulatory state instead of expression state alone. This can reveal trajectory features that are invisible in expression-only analysis, such as a cell population that has begun to open chromatin at lineage-specific loci before those loci are transcribed.

### Deep Learning and Generative Approaches

Recent methods apply deep learning to trajectory inference and alignment. RNA velocity analysis infers the temporal dynamics of transcriptional states from the relative abundances of spliced and unspliced mRNA. Classical approaches such as scVelo implement gene-specific kinetic modeling, while deep learning methods including DeepVelo, VeloVI, LatentVelo, SymVelo, and scTour use variational autoencoders to leverage nonlinear latent representations. Comparative evaluation of these approaches shows that variational autoencoder methods produce more directionally coherent and consistent velocity fields than classical models, although they require more computational resources and depend on accurate splicing quantification.

Generative models extend this further. Conditional flow matching approaches learn time-dependent velocity fields from unpaired snapshot populations, using optimal transport to establish couplings between adjacent time points. These methods support trajectory reconstruction at unmeasured time points and can forecast future cell states, which is valuable when experimental time points are sparse.

The choice between dynamic programming, Bayesian frameworks, archetypal models, and deep learning approaches depends on the research question, the data available, and the computational resources at hand. Each approach has distinct assumptions and outputs, and the alignment results must be interpreted within the framework that produced them.

## At a Glance: Trajectory Alignment Methods and Selection Criteria

| Method Category | Representative Tools | Data Requirements | Key Outputs | Best Use Case | Limitations |
|---|---|---|---|---|---|
| Dynamic programming alignment | Genes2Genes, dynamic time warping implementations | Single-cell expression matrices with inferred pseudotime for two conditions | Position-by-position mapping between trajectories, gene-level match and mismatch calls | Comparing differentiation programs between two conditions with known start and end states | Assumes trajectories share a common structure, may force spurious matches when conditions diverge substantially |
| Multi-omic trajectory inference | ArchVelo | Paired single-cell chromatin accessibility and transcriptomic data | Trajectories with gene-level latent time alignment, archetypal regulatory programs, transcription factor activity | Identifying regulatory drivers of trajectory differences when chromatin data are available | Requires paired multi-omic data, which is more expensive and technically demanding than expression-only data |
| Deep learning velocity and forecasting | scVelo, DeepVelo, VeloVI, LatentVelo, SymVelo, scTour, scFM | Spliced and unspliced count matrices, or time-series snapshot data | Velocity fields, predicted future cell states, trajectory reconstructions | Analyzing time-series data with sparse sampling or predicting cell fate outcomes | High computational cost, dependence on accurate splicing quantification, potential for biologically implausible predictions if training data are noisy |

## Data Inputs and Quality Requirements

### Expression Matrix Construction

Trajectory alignment begins with the same data preparation required for any single-cell RNA sequencing analysis. Raw sequencing reads must be processed into count matrices, which requires alignment to a reference genome and quantification of transcript abundance. The National Center for Biotechnology Information provides access to reference genomes, gene annotations, and sequence databases that support this step. The European Bioinformatics Institute offers training materials on data resources and practical analysis education that can guide researchers through the processing pipeline.

Quality control is essential before trajectory inference. Cells with low total read counts, high mitochondrial read fractions, or excessive doublet rates should be identified and removed. Genes detected in very few cells should be filtered. The specific thresholds depend on the tissue, the dissociation protocol, and the sequencing platform, so researchers should examine the distributions of quality metrics in their own data instead of applying fixed cutoffs from other studies.

Normalization is required to make gene expression values comparable across cells. Library size normalization divides each cell's counts by the total number of counts in that cell and multiplies by a scale factor. More sophisticated approaches account for technical variability and batch effects. The choice of normalization method affects downstream trajectory inference and alignment, so it should be documented and justified.

### Pseudotime Inference

Before trajectories can be aligned, they must be inferred. Pseudotime inference methods order cells along a path based on transcriptional similarity. Common approaches include principal component analysis followed by graph-based trajectory inference, and methods that model the transcriptional dynamics of individual genes.

The choice of trajectory inference method affects the resulting pseudotime values and therefore the alignment results. Different methods may produce different orderings of cells, particularly in regions of the trajectory where cell states are similar. Researchers should evaluate the stability of their inferred trajectories by running multiple methods or by subsampling cells and checking that the trajectory structure is preserved.

For RNA velocity analysis, the input data must include both spliced and unspliced counts. This requires that the alignment and quantification pipeline distinguish reads that map to exonic regions from reads that map to intronic regions. The accuracy of splicing quantification directly affects the quality of velocity estimates and downstream trajectory inference.

### Multi-Omic Data Integration

When chromatin accessibility data are available alongside transcriptomic data, the two modalities must be integrated before trajectory inference. ArchVelo uses paired single-cell chromatin accessibility and transcriptomic profiling, which requires experimental protocols that capture both modalities from the same cell. The computational integration must account for differences in data sparsity, dynamic range, and technical noise between the two modalities.

Researchers who do not have paired multi-omic data can still benefit from the conceptual framework of archetypal analysis. The idea that shared regulatory programs drive transcription can inform the interpretation of expression-only trajectories, even when chromatin data are not available to directly measure those programs.

## Practical Workflow for Gene-Level Trajectory Comparison

### Step 1: Define the Biological Question and Trajectory Structure

Before running any alignment software, define the comparison precisely. What are the two conditions being compared? What is the expected start state and end state of the trajectory in each condition? Is the trajectory expected to be linear or branched? What genes or pathways are of primary interest?

These decisions determine the analysis strategy. If the two conditions are expected to produce the same terminal cell type through similar intermediate states, a global alignment approach that matches the entire trajectories is appropriate. If the conditions are expected to diverge, a method that allows for mismatches, such as Genes2Genes, is preferable.

### Step 2: Process Data Through a Reproducible Pipeline

Reproducibility requires that every step of the analysis be documented and executable. Workflow management systems such as nf-core provide community standards for pipeline usage, configuration, and reproducibility. The Galaxy Training Network offers accessible workflow training and analysis tutorials that can help researchers build and validate their pipelines. The Carpentries lessons provide foundational computing, data, shell, Git, and programming training that supports reproducible analysis practices.

The Bioconductor project provides official package, workflow, installation, and reproducible genomic-analysis documentation. Many trajectory inference and alignment methods are implemented as Bioconductor packages, and the project's documentation includes worked examples that can be adapted to new datasets.

### Step 3: Infer Trajectories Independently in Each Condition

Infer pseudotime trajectories separately for each condition. Do not force the same trajectory structure onto both conditions. The goal is to discover whether the conditions differ in their developmental paths, and this requires that each condition's trajectory be inferred from its own data.

For each condition, evaluate the trajectory quality. Are the cells distributed along the trajectory as expected? Are there clear start and end states? Are there branches that correspond to known cell fates? If the trajectory is noisy or unstable, address the data quality issues before proceeding to alignment.

### Step 4: Align the Trajectories

Apply the alignment method appropriate to the data and question. For dynamic programming approaches, the input is typically the pseudotime-ordered expression matrices for the two conditions. The output is a mapping between positions on the reference and query trajectories.

For Bayesian information-theoretic approaches such as Genes2Genes, the output includes gene-level match and mismatch calls. This allows the researcher to identify which genes follow the same dynamics in both conditions and which genes diverge.

For multi-omic approaches such as ArchVelo, the alignment is integrated with trajectory inference. The method produces gene-level latent time alignment as part of its output, along with the archetypal regulatory programs that drive the trajectory.

### Step 5: Examine Gene-Level Alignment Patterns

After alignment, examine the expression dynamics of individual genes along the aligned trajectories. For each gene, plot its expression as a function of aligned position. Genes that are conserved between conditions will show similar curves. Genes that diverge will show differences in timing, amplitude, or presence.

Cluster the genes by their alignment patterns. Genes that show similar divergence patterns may be coregulated or may participate in the same biological pathway. The clustering output from Genes2Genes is designed to highlight distinct clusters of genes with varying patterns of expression dynamics.

### Step 6: Validate and Interpret the Results

Validation requires checking that the alignment is biologically plausible. Do the matched positions correspond to similar cell states? Are the genes identified as divergent consistent with known biology of the system? If the alignment produces unexpected results, examine the underlying data to determine whether the result reflects a real biological difference or an artifact of the alignment procedure.

Interpretation should connect the alignment results to the original biological question. If the goal was to determine whether an in vitro differentiation protocol reproduces the in vivo developmental program, the alignment should identify which genes are faithfully recapitulated and which are missing or mistimed. This information directly guides protocol optimization.

## Options and Tradeoffs in Alignment Methodology

### Global versus Local Alignment

Global alignment matches the entire reference trajectory to the entire query trajectory. This is appropriate when the two trajectories are expected to share the same start and end states. Local alignment finds regions of similarity without requiring the ends to match. This is appropriate when the trajectories share only a portion of their structure, such as a common early developmental phase followed by divergent fates.

The choice between global and local alignment affects the interpretation of results. Global alignment forces a correspondence between the endpoints, which may be biologically incorrect if the conditions produce different terminal states. Local alignment avoids this assumption but requires the researcher to specify the minimum length of aligned regions.

### Hard versus Soft Matching

Hard matching assigns each reference position to exactly one query position. Soft matching allows a reference position to match multiple query positions with different weights. Soft matching is more appropriate when the trajectory has continuous structure and cells at similar developmental stages may be spread across multiple pseudotime positions.

Dynamic programming approaches typically produce hard matches. Optimal transport approaches, which are used in some deep learning methods, produce soft matches. The choice affects the resolution of the alignment and the interpretation of gene-level dynamics.

### Expression-Only versus Multi-Omic Alignment

Expression-only alignment uses transcriptomic data to define the trajectory and the alignment. This is the most widely applicable approach because single-cell RNA sequencing is the most common single-cell modality. However, expression-only alignment cannot distinguish between differences in transcriptional output and differences in the regulatory programs that drive transcription.

Multi-omic alignment, as implemented in ArchVelo, uses paired chromatin accessibility and transcriptomic data. This provides a more complete picture of the regulatory state of cells along the trajectory. The tradeoff is that paired multi-omic data are more expensive and technically demanding to generate, and the computational methods are more complex.

### Static Trajectory Alignment versus Dynamic Forecasting

Static trajectory alignment compares trajectories that have already been inferred from snapshot data. Dynamic forecasting methods, such as conditional flow matching approaches, model the temporal evolution of cell states and can predict future states. Forecasting is valuable when experimental time points are sparse or when the goal is to predict the outcome of a perturbation.

The tradeoff is that forecasting methods require assumptions about the dynamics of the system and are sensitive to errors in the estimated velocity fields. Static alignment methods are more conservative but cannot predict unobserved states.

## Observations and Measurements for Alignment Assessment

### Alignment Quality Metrics

Several metrics can assess the quality of a trajectory alignment. The total alignment cost, which is the sum of the costs of the matched positions, provides a global measure of similarity. Lower cost indicates greater similarity, but the absolute value depends on the cost function and the scaling of the expression data.

The proportion of the reference trajectory that is matched to the query trajectory indicates how much of the reference structure has a counterpart in the query. A low proportion suggests substantial divergence. The proportion of genes that show conserved dynamics versus divergent dynamics provides a gene-level view of the alignment quality.

For deep learning methods, cosine similarity of velocity vectors measures directional concordance between conditions, and mean squared error of trajectory continuity measures the smoothness of the inferred trajectories. These metrics are used in comparative evaluations of velocity tools and can be applied to assess the quality of a specific analysis.

### Sensitivity Analysis

The alignment result should be robust to reasonable changes in the analysis parameters. Vary the pseudotime inference method, the number of principal components used for dimensionality reduction, the normalization method, and the alignment parameters. If the alignment changes substantially with small parameter changes, the result is not reliable.

Sensitivity analysis is particularly important for trajectory alignment because the alignment is a global optimization that can be sensitive to local features of the data. A small number of outlier cells or a noisy gene can shift the optimal alignment path.

### Biological Validation

The ultimate test of an alignment is whether it produces biologically meaningful results. Genes identified as conserved between conditions should include known markers of the developmental process. Genes identified as divergent should be plausible candidates for condition-specific regulation. If the alignment identifies unexpected genes as divergent, examine those genes in the context of the known biology of the system.

In the Genes2Genes application to T cell differentiation, the alignment revealed that in vitro differentiated T cells matched an immature in vivo state and lacked expression of genes associated with TNF signaling. This finding is biologically interpretable and directly actionable: it identifies a specific molecular deficiency of the in vitro system that could be targeted by modifying the culture conditions.

## Records and Documentation Standards

### Analysis Logs

Maintain a complete log of the analysis, including software versions, parameter settings, and input data versions. This is essential for reproducibility and for troubleshooting when results are unexpected. Workflow management systems such as nf-core provide structured logging and configuration documentation that supports this requirement.

The Bioconductor project emphasizes reproducible genomic analysis, and its documentation includes guidance on version control and session information. Recording the R or Python session information, including package versions, ensures that the analysis can be reproduced exactly.

### Data Provenance

Document the origin of every dataset used in the analysis. For public datasets, record the accession number and the date of download. The National Center for Bioinformatics Resources provides search systems and database access that support data provenance tracking. For generated data, record the experimental protocol, the sequencing platform, and the processing pipeline.

### Decision Records

Record the rationale for each analytical decision. Why was a particular trajectory inference method chosen? Why were specific quality control thresholds applied? Why was a particular alignment method selected? These decision records are valuable when the analysis is revisited or when results are questioned by reviewers or collaborators.

## Common Failure Patterns and Troubleshooting

### Failure Pattern 1: Forced Alignment of Divergent Trajectories

Standard dynamic programming approaches assume that every position on the reference trajectory has a match on the query trajectory. When the trajectories are substantially divergent, this assumption forces spurious matches that distort the alignment. The result is an alignment that appears successful by cost metrics but is biologically meaningless.

Troubleshooting: Use a method that explicitly models mismatches, such as Genes2Genes. Examine the proportion of the trajectory that is matched and the distribution of match costs. If large regions of the trajectory have high match costs, the trajectories may be more divergent than the alignment method can accommodate.

### Failure Pattern 2: Sensitivity to Pseudotime Inference

The alignment result depends on the pseudotime values assigned to cells. Different pseudotime inference methods can produce different orderings, particularly in regions of the trajectory where cell states are similar. If the alignment changes substantially when the pseudotime method is changed, the result is not robust.

Troubleshooting: Run multiple pseudotime inference methods and compare the alignments. If the alignments are consistent, the result is robust. If they differ, examine the regions of the trajectory where the pseudotime assignments differ and determine which ordering is more biologically plausible.

### Failure Pattern 3: Batch Effects Confounded with Biological Differences

If the two conditions being compared were processed in different batches, technical differences between batches can be confounded with biological differences. The alignment may identify genes as divergent when the difference is actually due to batch effects.

Troubleshooting: Apply batch correction before trajectory inference and alignment. Validate the batch correction by checking that known biological markers are preserved. If batch correction is not possible, interpret the alignment results with caution and validate key findings with an independent experiment.

### Failure Pattern 4: Overinterpretation of Gene-Level Differences

Gene-level alignment identifies genes whose dynamics differ between conditions. However, not all differences are biologically meaningful. Technical noise, dropout events, and differences in sequencing depth can produce apparent differences in gene dynamics that do not reflect true biological divergence.

Troubleshooting: Apply statistical tests to distinguish meaningful differences from noise. Examine the magnitude of the expression differences and the consistency of the differences across cells. Validate key findings with orthogonal methods such as quantitative PCR or protein-level measurements.

### Failure Pattern 5: Computational Resource Limitations

Deep learning methods for trajectory inference and alignment require substantial computational resources, including GPU acceleration for some methods. Variational autoencoder approaches are more computationally demanding than classical kinetic models. If the available computational resources are insufficient, the analysis may be impractically slow or may fail to converge.

Troubleshooting: Start with simpler methods to establish the overall trajectory structure, then apply more complex methods to refine the analysis. Use cloud computing resources if local resources are insufficient. Document the computational requirements of each method to plan resource allocation.

## Limitations and Interpretation Boundaries

### Trajectory Inference Is Not True Time

Pseudotime is an ordering of cells based on transcriptional similarity, not a measurement of real time. Cells that are placed at the same pseudotime position may have been harvested at different real times, and cells at different pseudotime positions may have been harvested at the same real time. Trajectory alignment aligns pseudotime scales, not real time scales.

This limitation affects the interpretation of gene-level dynamics. A gene that appears to be expressed earlier in one condition along the aligned trajectory may not actually be expressed earlier in real time. The alignment only establishes that the gene's expression pattern is shifted relative to the transcriptional progression of the cells.

### Alignment Cannot Establish Causality

Trajectory alignment identifies correlations between gene expression dynamics and developmental progression. It does not establish that a gene's expression causes the developmental transition. Genes that are identified as divergent between conditions may be downstream consequences of other differences instead of drivers of the divergence.

Causal inference requires perturbation experiments. If a gene is identified as a candidate driver of condition-specific differences, its role should be tested by manipulating its expression and observing the effect on the trajectory.

### Methods Have Different Assumptions

Each alignment method makes assumptions about the structure of the trajectories and the nature of the correspondence between them. Dynamic programming assumes that a match exists for every position. Bayesian information-theoretic methods relax this assumption but require specification of prior distributions. Deep learning methods assume that the dynamics can be learned from the data, which requires sufficient sample sizes and accurate splicing quantification.

The alignment results must be interpreted within the assumptions of the method that produced them. A result from a method that assumes complete matching should not be interpreted as evidence that the trajectories are completely matched.

### Multi-Omic Methods Require Multi-Omic Data

Methods such as ArchVelo that model gene regulation using chromatin accessibility data require paired single-cell chromatin accessibility and transcriptomic profiling. These data are more expensive and technically demanding to generate than expression-only data. Researchers who do not have multi-omic data cannot apply these methods and must rely on expression-only approaches.

The conceptual framework of archetypal analysis can still inform the interpretation of expression-only data. The idea that shared regulatory programs drive transcription suggests that genes with similar expression dynamics along a trajectory may be regulated by common programs, even when the regulatory programs themselves are not directly measured.

## Welfare and Safety Context for Laboratory Applications

Trajectory alignment is a computational analysis applied to data that have already been generated. The welfare and safety considerations apply to the experimental work that produces the data, not to the computational analysis itself.

For studies involving primary cells from animal tissues, the tissue collection procedures must comply with institutional animal care and use regulations. For studies involving human samples, informed consent and institutional review board approval are required. The National Center for Biotechnology Information provides access to databases that document the provenance of public datasets, which supports compliance with data use agreements.

For studies involving in vitro differentiation protocols, the culture conditions must be optimized to maintain cell viability and to produce the desired cell types. The alignment results can guide this optimization by identifying the specific molecular deficiencies of the current protocol. In the T cell differentiation example, the alignment revealed that in vitro differentiated cells lacked expression of genes associated with TNF signaling, suggesting that the culture conditions should be modified to induce this pathway.

## Professional Escalation Criteria

### When to Seek Computational Expertise

Trajectory alignment methods require familiarity with single-cell analysis workflows, statistical modeling, and programming. If the analysis team lacks these skills, consultation with a bioinformatics specialist or collaboration with a computational biology group is appropriate. The European Bioinformatics Institute offers training pathways in bioinformatics and data-resource training that can build the necessary skills.

Specific situations that warrant escalation include: the need to integrate multi-omic data when the team has only processed expression data, the need to apply deep learning methods when the team lacks GPU computing experience, and the need to develop custom alignment approaches when existing methods do not fit the data structure.

### When to Question the Data

If the alignment produces results that contradict well-established biology, the data quality should be questioned before the biology is questioned. Examine the quality control metrics, the trajectory inference results, and the alignment diagnostics. If the data quality is adequate and the alignment is robust, the unexpected result may represent a genuine biological finding that warrants further investigation.

### When to Seek Biological Validation

Computational findings should be validated experimentally before they are used to guide major decisions such as protocol redesign or clinical translation. If the alignment identifies a specific gene or pathway as divergent between conditions, validate the finding with an independent measurement. If the finding is confirmed, design experiments to test whether modulating the divergent gene or pathway improves the outcome of interest.

## Frequently Asked Questions

### What is the difference between pseudotime and aligned time?

Pseudotime is an ordering of cells along a trajectory based on transcriptional similarity. It is computed independently for each dataset and has no cross-condition meaning. Aligned time is the result of mapping pseudotime positions from one trajectory to another, establishing correspondences between developmental states across conditions. Aligned time allows direct comparison of gene expression dynamics between conditions because it places cells at the same developmental stage in correspondence.

### When should I use dynamic programming alignment versus deep learning approaches?

Dynamic programming approaches such as Genes2Genes are appropriate when you have two well-defined trajectories and want to establish a position-by-position correspondence with explicit handling of gene-level matches and mismatches. Deep learning approaches are appropriate when you have time-series data with sparse sampling, when you want to forecast future cell states, or when you have paired multi-omic data that require integrated modeling. Deep learning methods require more computational resources and more careful preprocessing.

### Can I align trajectories from different single-cell platforms?

Yes, but platform differences introduce technical variation that can confound biological differences. The data should be normalized and, if necessary, batch corrected before trajectory inference and alignment. The alignment results should be interpreted with caution because platform-specific technical effects can produce apparent gene-level differences that do not reflect true biological divergence.

### How do I know if my alignment is biologically meaningful?

Examine whether the aligned positions correspond to similar cell states, whether genes identified as conserved include known markers of the developmental process, and whether genes identified as divergent are plausible candidates for condition-specific regulation. Validate key findings with independent measurements or experiments. If the alignment produces results that contradict well-established biology, investigate the data quality and the alignment parameters before accepting the result.

### What data do I need for multi-omic trajectory alignment?

Multi-omic trajectory alignment as implemented in ArchVelo requires paired single-cell chromatin accessibility and transcriptomic profiling from the same cells. This requires experimental protocols that capture both modalities simultaneously. If you only have expression data, you can still perform trajectory alignment using expression-based methods, but you will not be able to identify the regulatory programs that drive the trajectory.

### How does RNA velocity relate to trajectory alignment?

RNA velocity estimates the temporal dynamics of transcriptional states from the relative abundances of spliced and unspliced mRNA. Velocity information can be used to infer trajectories and to align trajectories across conditions. Deep learning velocity methods produce velocity fields that can be compared across conditions using metrics such as cosine similarity of velocity vectors. Velocity-based alignment can reveal differences in the direction and rate of transcriptional change that are not visible in static pseudotime alignment.

### What should I do if my trajectories are too divergent to align?

If the trajectories share only a portion of their structure, use local alignment methods that find regions of similarity without requiring the entire trajectories to match. If the trajectories are completely divergent, the alignment may not be meaningful, and the analysis should focus on describing the differences instead of forcing a correspondence. Methods that explicitly model mismatches, such as Genes2Genes, can identify which genes are conserved and which are divergent even when the overall trajectory structures differ.

### How should I report trajectory alignment results in a publication?

Report the software versions, parameter settings, and input data versions for every step of the analysis. Describe the trajectory inference method and the alignment method, including the rationale for the choices. Report alignment quality metrics and the results of sensitivity analyses. Provide the aligned gene-level dynamics for the genes of interest, either as figures or as supplementary data. Deposit the analysis code and the processed data in a public repository to support reproducibility.

## Related Bioinformatics Guides

- [Single-Cell RNA Velocity: Inferring Cellular Dynamics](/knowledge/bioinformatics/single-cell-rna-velocity-inferring-cellular-dynamics)
- [Benchmarking Atlas-Level Data Integration in Single-Cell Genomics: Methods and Best Practices](/knowledge/bioinformatics/benchmarking-atlas-level-data-integration-in-single-cell-genomics-methods-and-best-practices)
- [Single-Cell Genomics: From Concept to Application](/knowledge/bioinformatics/single-cell-genomics-from-concept-to-application)
- [Single-Cell Annotation: A Workflow for Cell Type Identification](/knowledge/bioinformatics/single-cell-annotation-a-workflow-for-cell-type-identification)
- [Single-Cell Isolation Techniques: A Practical Comparison](/knowledge/bioinformatics/single-cell-isolation-techniques-a-practical-comparison)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)


## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [Gene-level alignment of single-cell trajectories.](https://pubmed.ncbi.nlm.nih.gov/39300283). Nature methods, 2025.
- [ArchVelo: archetypal velocity modeling for single-cell multi-omic trajectories.](https://pubmed.ncbi.nlm.nih.gov/42248920). Nature communications, 2026.
- [ArchVelo: Archetypal Velocity Modeling for Single-cell Multi-omic Trajectories.](https://pubmed.ncbi.nlm.nih.gov/41000980). bioRxiv : the preprint server for biology, 2025.
- [Serum withdrawal establishes a stress-dominant entry state during myogenic differentiation.](https://doi.org/10.3389/fcell.2026.1823860). 2026.
- [StPedf: Cell trajectory inference of spatial transcriptomics via spatial proximity embedding and spatial density-adaptive fusion.](https://doi.org/10.1371/journal.pcbi.1014346). 2026.
- [Comparison between a conventional tool and deep learning models for RNA velocity analysis of scRNA-Seq data.](https://doi.org/10.1007/s00438-026-02429-9). 2026.
- [Computational blueprints for cell fate programming.](https://doi.org/10.1016/j.stemcr.2026.102929). 2026.
- [Single-cell morphodynamical trajectories enable prediction of gene expression accompanying cell state change.](https://doi.org/10.1016/j.cels.2026.101567). 2026.
- [Gene-level alignment of single cell trajectories](https://doi.org/10.1101/2023.03.08.531713). bioRxiv, 2023.
- [scDiformer: A Difference-Aware Transformer for Time-Series Single-Cell Gene Expression Forecasting](https://doi.org/10.1109/BIBM66473.2025.11356605). IEEE International Conference on Bioinformatics and Biomedicine, 2025.
- [From Snapshots to Trajectories: Learning Single-Cell Gene Expression Dynamics via Conditional Flow Matching](https://doi.org/10.48550/arXiv.2605.22340). arXiv.org, 2026.
- [Manifold learning analysis suggests novel strategies for aligning single-cell multi-modalities and revealing functional genomics for neuronal electrophysiology](https://doi.org/10.1101/2020.12.03.410555). bioRxiv, 2020.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.