Trajectory-Based Differential Expression in Single-Cell RNA-Seq: Identifying Genes that Change Along Continuous Cell States
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Trajectory-based differential expression analysis is crucial for identifying genes that change along continuous biological processes like cell differentiation or activation, overcoming limitations of cluster-based methods that discretize these transitions.
- Pseudotime ordering, derived from dimensionality reduction techniques like PCA, is central to trajectory inference, requiring careful root cell selection and evaluation of inferred lineage structures (e.g., linear vs. branched) using algorithms like Monocle3.
- Robust data preprocessing, including stringent quality control (filtering low-quality cells and doublets), normalization, and batch effect correction (e.g., using Harmony), is paramount for accurate trajectory inference and downstream differential expression testing.
- Statistical models such as generalized additive models with
tradeSeq(handling zero inflation with observation weights) and functional principal component analysis withscFPC-DE(offering improved type I error control under zero inflation) are employed to detect genes with dynamic expression patterns along pseudotime and across lineages. - Common pitfalls include circular analysis, root selection bias, overinterpretation of spurious branches, and ignoring zero inflation; rigorous validation through independent computational methods (e.g., bootstrapping) and experimental techniques (e.g., immunohistochemistry, flow cytometry) is essential for biological conclusions.
Trajectory-based differential expression analysis addresses a specific problem in single-cell RNA sequencing: identifying genes whose expression changes along a continuous cellular process such as differentiation, activation, or disease progression, instead of comparing discrete cell clusters. Standard differential expression tests that compare cluster averages can miss genes that fluctuate smoothly across a pseudotime continuum or that change in one lineage branch but not another. This article explains the analytical workflow for trajectory-based differential expression, covering data preparation, trajectory inference, statistical modeling with methods such as tradeSeq and monocle3, interpretation of results, and common pitfalls. The target reader is a researcher or laboratory professional who has already generated single-cell or single-nucleus RNA-seq data and needs to move beyond cluster-based analysis to understand dynamic gene regulation.
The Problem with Cluster-Based Differential Expression for Continuous Processes
Single-cell RNA sequencing captures gene expression profiles from individual cells, and the most common analytical framework begins with unsupervised clustering to identify discrete cell types or states. Differential expression testing then compares the average expression of each gene between clusters. This approach works well when biological categories are truly discrete, such as comparing T cells versus B cells or healthy versus diseased tissue.
However, many biological processes are continuous. Cell differentiation proceeds through intermediate states, immune cells transition through activation stages, and disease progression involves gradual transcriptional remodeling. When cells are clustered, the continuous nature of these transitions is artificially discretized. A gene that gradually increases across a differentiation trajectory may show only modest differences between adjacent clusters, and the statistical test may fail to detect it even though the gene is clearly regulated across the process. Conversely, clustering can create arbitrary boundaries that split a continuous gradient into artificial groups, leading to false positive results when comparing cells that are actually part of the same continuum.
Trajectory inference methods address this limitation by ordering cells along a pseudotime axis that represents progress through a biological process. The term pseudotime reflects that this ordering is computationally inferred from gene expression patterns instead of measured directly through time-lapse experiments. Once cells are ordered, differential expression analysis can ask whether gene expression changes as a function of pseudotime, and whether those changes differ between lineage branches. This approach preserves the continuous resolution of the data and can identify genes that cluster-based methods miss.
The practical consequence for researchers is that choosing between cluster-based and trajectory-based differential expression depends on the biological question. If the question concerns discrete cell types, cluster-based analysis is appropriate. If the question concerns how cells transition from one state to another, or how a continuous gradient of states is established, trajectory-based analysis is the more informative approach. Many published studies now combine both strategies, using clustering for cell type annotation and trajectory analysis for understanding dynamic processes within or across cell types.
Core Principles of Trajectory Inference
Trajectory inference begins with the assumption that cells in a differentiating or transitioning population can be ordered along a low-dimensional manifold that represents the biological process. The computational challenge is to reconstruct this manifold from high-dimensional gene expression data that contains substantial technical noise.
Pseudotime Ordering and Lineage Structure
The output of trajectory inference is a pseudotime value for each cell, representing its position along the inferred developmental path. In simple cases, the trajectory is a single linear path from a starting state to an endpoint. More commonly, trajectories contain branches, representing points where cells make fate decisions and diverge toward different endpoints. The structure of a trajectory can include multiple lineages, where each lineage is a path from the root to a terminal state.
Several algorithms are available for trajectory inference, and they differ in their assumptions about the underlying structure. Monocle3, which is widely used, learns a principal graph that can represent complex trajectories with multiple branches. Other methods use different mathematical frameworks, such as local tangent space alignment for branched trajectories. The choice of algorithm can affect the resulting pseudotime ordering and lineage assignments, so researchers should evaluate multiple methods when possible and consider whether the inferred structure is biologically plausible.
The Role of Dimensionality Reduction
Trajectory inference typically operates on a reduced representation of the gene expression data. Principal component analysis is commonly used to capture the major sources of variation while reducing noise. Some methods use more specialized dimensionality reduction techniques that preserve local structure. The quality of the trajectory depends on choosing an appropriate number of dimensions and ensuring that the reduced representation captures the biological signal instead of technical artifacts.
Root Cell Selection
Most trajectory inference methods require specification of a root cell or root state, representing the beginning of the process. The choice of root can substantially affect the pseudotime ordering and therefore the results of downstream differential expression analysis. Researchers should select the root based on biological knowledge, such as a known progenitor or stem cell population, and should test the robustness of results to alternative root choices.
Data Preparation for Trajectory-Based Analysis
The quality of trajectory-based differential expression results depends heavily on the quality of the input data. Several preprocessing steps are essential before trajectory inference and differential expression testing can be performed reliably.
Quality Control and Filtering
Single-cell RNA-seq data contains technical artifacts that can distort trajectory inference. Low-quality cells, defined by low total read counts, low gene detection, or high mitochondrial read fractions, should be removed before analysis. Doublets, which are sequencing artifacts where two cells are captured together, can create spurious intermediate states that distort trajectories. Several computational tools are available for doublet detection, and their use is recommended before trajectory inference.
The quality control thresholds should be documented and reported, as they affect the final cell population and therefore the trajectory structure. Different tissue types and protocols have different expected quality metrics, so thresholds should be informed by the specific dataset instead of applied uniformly across all experiments. The Galaxy Training Network provides accessible workflow training and analysis tutorials that cover quality control procedures for single-cell data, which can serve as a reference for establishing reproducible filtering protocols.
Normalization and Batch Effect Correction
Gene expression counts must be normalized to account for differences in sequencing depth between cells. Various normalization methods are available, and the choice can affect downstream analyses. For datasets that combine multiple samples, batches, or experimental conditions, batch effect correction is often necessary before trajectory inference. Methods such as Harmony are commonly used for this purpose.
The interaction between batch correction and trajectory inference requires careful consideration. If batch effects correlate with the biological process of interest, aggressive batch correction can remove genuine biological signal. Researchers should examine whether cells from different batches are distributed across the trajectory or segregated into distinct regions, and should interpret the results accordingly.
Feature Selection
Trajectory inference and downstream differential expression analysis typically use a subset of highly variable genes instead of the full transcriptome. Feature selection reduces noise and computational burden while focusing on genes that carry biological signal. The selected genes should include those expected to be involved in the process of interest, and the feature set should be documented for reproducibility.
Statistical Models for Trajectory-Based Differential Expression
Once cells are ordered along pseudotime and assigned to lineages, the next step is to test which genes change along the trajectory. Several statistical frameworks have been developed for this purpose, each with different assumptions and capabilities.
Generalized Additive Models with tradeSeq
The tradeSeq method models gene expression as a smooth function of pseudotime using generalized additive models based on the negative binomial distribution. This approach accommodates the count nature of single-cell RNA-seq data and allows flexible modeling of nonlinear expression patterns along the trajectory. The method was introduced in a 2020 publication in Nature Communications that described its framework for trajectory-based differential expression analysis.
tradeSeq can test several types of differential expression. Within-lineage tests ask whether a gene's expression changes along a particular lineage. Between-lineage tests ask whether a gene is differentially expressed between two or more lineages, which can identify genes involved in fate decisions at branch points. The method also supports pattern testing, which groups genes with similar expression dynamics along the trajectory.
An important feature of tradeSeq is its ability to incorporate observation-level weights to account for zero inflation, a common artifact in single-cell RNA-seq data where genes are not detected in some cells due to technical dropout instead of true absence of expression. By weighting observations appropriately, the model can reduce false positives caused by excessive zeros.
Functional Principal Component Analysis with scFPC-DE
A newer approach, scFPC-DE, uses functional data analysis to model gene expression as a function of pseudotime in a function space. The method represents the covariance structure of expression functions using eigenfunctions derived from functional principal component analysis. This approach captures shared variation across genes along the trajectory and can improve statistical power while controlling false positive rates.
The developers of scFPC-DE report that it achieves superior control of type I error compared to generalized additive model approaches, particularly under zero inflation. In analyses of B cell differentiation data, the method identified temporally differentially expressed genes enriched for B cell differentiation pathways. The R package and code vignettes are publicly available.
Comparison of Methods
The choice between tradeSeq, scFPC-DE, and other trajectory-based differential expression methods depends on the specific analytical needs. tradeSeq is well established, widely used, and supports multiple types of hypothesis tests including between-lineage comparisons. scFPC-DE offers potential advantages in controlling false positives under zero inflation and in capturing shared variation across genes.
Researchers should be aware that different methods may identify different sets of differentially expressed genes, and the results should be interpreted in the context of the method's assumptions. Validation of candidate genes through independent experiments, such as immunohistochemistry or flow cytometry, is recommended before drawing strong biological conclusions.
Practical Workflow for Trajectory-Based Differential Expression
The following workflow outlines the steps for conducting trajectory-based differential expression analysis. The specific tools and parameters will depend on the data and biological question, but the overall structure provides a framework for reproducible analysis.
Step 1: Data Import and Quality Control
Import the count matrix and cell metadata into the analysis environment. Apply quality control filters to remove low-quality cells and doublets. Document the number of cells and genes retained after each filtering step. For single-nucleus RNA-seq data, verify that the quality metrics are appropriate for nuclear preparations, which may differ from whole-cell preparations.
Step 2: Normalization and Batch Correction
Normalize the count data to account for sequencing depth differences. If the dataset includes multiple batches or samples, apply batch correction methods as appropriate. Visualize the data before and after batch correction to confirm that technical variation is reduced without removing biological signal.
Step 3: Dimensionality Reduction and Clustering
Perform principal component analysis to reduce the dimensionality of the data. Use the selected principal components for clustering to identify major cell populations. Annotate cell types based on known marker genes. This step provides the biological context for interpreting the trajectory.
Step 4: Subset Cells of Interest
Trajectory analysis is typically performed on a subset of cells representing the process of interest, such as a particular cell lineage undergoing differentiation. Subsetting reduces noise from unrelated cell types and improves the resolution of the trajectory. The subset should include the starting population, intermediate states, and endpoint populations.
Step 5: Trajectory Inference
Run trajectory inference on the subset of cells. Specify the root based on biological knowledge. Examine the resulting trajectory structure, including the number of lineages and the distribution of cells along pseudotime. Evaluate whether the inferred structure is consistent with known biology.
Step 6: Trajectory-Based Differential Expression Testing
Apply the chosen differential expression method to identify genes that change along the trajectory. For tradeSeq, specify the lineages and the types of tests to perform. For scFPC-DE, prepare the data in the required format and run the analysis. Examine the distribution of p-values and apply appropriate multiple testing correction.
Step 7: Interpretation and Validation
Interpret the identified genes in the context of the biological process. Examine the expression patterns of candidate genes along pseudotime to confirm that the statistical results correspond to meaningful biological dynamics. Validate key findings through independent methods where possible.
At a Glance: Method Selection for Trajectory-Based Differential Expression
| Method | Statistical Framework | Key Strengths | Primary Use Case | Availability |
|---|---|---|---|---|
| tradeSeq | Generalized additive models with negative binomial distribution | Supports within-lineage and between-lineage tests, accommodates zero inflation with observation weights | Identifying genes associated with lineages or differentially expressed between lineages | R/Bioconductor |
| scFPC-DE | Functional principal component analysis | Controls type I error under zero inflation, captures shared variation across genes | Identifying temporally differentially expressed genes along pseudotime | R package on GitHub |
| Monocle3 | Principal graph learning with regression-based DE | Integrated trajectory inference and differential expression, handles complex branched trajectories | End-to-end trajectory analysis with visualization | R/Bioconductor |
Data Inputs and Their Impact on Results
The quality and characteristics of the input data substantially affect trajectory-based differential expression results. Understanding these effects helps researchers interpret their findings and design appropriate validation experiments.
Sequencing Platform and Protocol
Trajectory-based differential expression has been applied to data from both droplet-based and full-length sequencing protocols. Droplet-based methods capture many cells but with lower sensitivity per cell, leading to higher dropout rates. Full-length protocols provide more complete transcript coverage but for fewer cells. The choice of protocol affects the statistical power for detecting genes with low or variable expression along the trajectory.
Cell Number and Sampling Density
The number of cells sampled along the trajectory affects the resolution of pseudotime ordering and the statistical power for detecting expression changes. Sparse sampling of intermediate states can create gaps in the trajectory that make it difficult to distinguish continuous changes from discrete transitions. Researchers should aim for dense sampling across the process of interest, particularly at branch points where fate decisions occur.
Single-Cell versus Single-Nucleus Data
Single-nucleus RNA-seq is increasingly used for tissues that are difficult to dissociate into intact cells, such as brain tissue. The transcriptome of the nucleus differs from that of the whole cell, with nuclear preparations enriched for precursor and intronic RNA. Trajectory-based differential expression can be applied to single-nucleus data, but researchers should be aware that the detected expression dynamics may differ from whole-cell analyses. The analytical methods and application strategies for single-cell and single-nucleus transcriptomics have been reviewed in the context of ischemic stroke research, where both approaches have been applied to investigate cellular heterogeneity and dynamic processes.
Quantification Uncertainty
The quantification of gene expression from single-cell RNA-seq data involves inherent uncertainty due to reads that map to multiple genes. Many quantification pipelines discard multi-mapping reads, which can underestimate counts for functionally important genes. Methods that account for quantification uncertainty, such as alevin with inferential replicates, can reduce false positives in trajectory-based differential expression analyses. Compression of inferential replicates to mean and variance estimates can reduce storage and memory requirements while preserving the benefits of uncertainty modeling.
Common Failure Patterns and How to Avoid Them
Several recurring problems can compromise trajectory-based differential expression analysis. Recognizing these patterns helps researchers troubleshoot their analyses and interpret results appropriately.
Circular Analysis and Double Dipping
Using the same data to define the trajectory and then test for differential expression along that trajectory can lead to inflated significance. This problem is related to the broader issue of circular analysis in genomics, where feature selection and hypothesis testing are performed on the same dataset. Researchers should be aware that trajectory-based differential expression results are exploratory and require validation in independent datasets or through experimental confirmation.
Root Selection Bias
The choice of root cell can substantially affect pseudotime ordering and therefore the genes identified as differentially expressed. If the root is placed in the wrong cell population, the entire trajectory may be oriented incorrectly, and the interpretation of expression dynamics will be misleading. Researchers should select roots based on independent biological knowledge and test the robustness of results to alternative root choices.
Overinterpretation of Branches
Trajectory inference algorithms can produce branches that reflect technical artifacts instead of genuine biological fate decisions. Batch effects, dropout patterns, and the presence of rare cell types can all create spurious branching structure. Before interpreting branch-specific differential expression, researchers should verify that the branches are supported by multiple lines of evidence, such as known marker genes and independent trajectory inference methods.
Ignoring Zero Inflation
Single-cell RNA-seq data contains excessive zeros due to technical dropout, and this zero inflation can inflate false positive rates in trajectory-based differential expression tests. Methods that explicitly account for zero inflation, such as tradeSeq with observation-level weights and scFPC-DE, are preferable to naive approaches that treat zeros as true zero expression. Researchers should check the zero fraction for identified genes and consider whether the results are robust to dropout effects.
Failure to Validate
Trajectory-based differential expression identifies candidate genes, but computational identification alone is insufficient for biological conclusions. Validation through independent methods, such as immunohistochemistry, flow cytometry, or functional assays, is essential. Published studies that combine trajectory analysis with experimental validation provide models for this approach. For example, spatial transcriptomics analysis combined with trajectory-based differential expression and immunohistochemical staining identified SPARC as a prognostic marker in interstitial lung diseases, with validation showing high expression in young fibrotic lesions but not in old scarring lesions.
Records and Measurements for Reproducible Analysis
Reproducibility is a core requirement for trajectory-based differential expression analysis. The following records should be maintained for each analysis.
Analysis Documentation
Document the software versions, parameters, and settings used at each step of the analysis. This includes the trajectory inference algorithm and its parameters, the differential expression method and its settings, and the multiple testing correction approach. Version control for analysis scripts ensures that the exact analysis can be reproduced. The nf-core documentation provides community standards for reproducible workflow configuration and usage that can inform documentation practices for single-cell analyses.
Quality Metrics
Record the number of cells and genes at each stage of filtering, the quality control thresholds applied, and the distribution of quality metrics before and after filtering. These records allow reviewers to assess whether the filtering decisions were appropriate and whether the results are robust to alternative thresholds.
Trajectory Characteristics
Document the number of lineages, the distribution of cells along pseudotime, and the root selection rationale. If multiple trajectory inference methods are compared, record the results of each method and the degree of agreement between them.
Differential Expression Results
Maintain the full results table for trajectory-based differential expression, including gene identifiers, test statistics, p-values, adjusted p-values, and effect sizes. The complete results table, instead of only the significant genes, allows for appropriate interpretation and meta-analysis.
Interpretation of Trajectory-Based Differential Expression Results
The output of trajectory-based differential expression analysis is a list of genes with statistical evidence for expression changes along the trajectory. Interpreting these results requires biological context and careful consideration of the statistical model.
Within-Lineage Expression Dynamics
Genes identified as differentially expressed within a lineage show expression changes as a function of pseudotime. The pattern of change can be monotonic, with expression increasing or decreasing steadily, or nonmonotonic, with peaks or troughs at intermediate pseudotime values. The shape of the expression pattern provides clues about the gene's role in the process. Monotonic changes may reflect progressive commitment to a fate, while transient peaks may indicate genes involved in specific transition events.
Between-Lineage Differences
Genes differentially expressed between lineages identify molecular differences between alternative fates. These genes are particularly interesting at branch points, where they may represent the molecular basis of fate decisions. The timing of expression divergence between lineages can indicate when fate decisions are executed.
Association with Biological Processes
The identified genes should be examined for enrichment in biological pathways and processes relevant to the studied system. Gene Ontology and KEGG pathway enrichment analysis can place the differentially expressed genes in a functional context. The interpretation should connect the statistical results to the known biology of the system.
Integration with Other Data Types
Trajectory-based differential expression results can be integrated with other analyses, including gene regulatory network inference, cell-cell communication analysis, and spatial transcriptomics. This integration can reveal the regulatory mechanisms underlying expression dynamics and the intercellular signals that coordinate transitions. The analytical methods for single-cell and single-nucleus transcriptomics in disease research commonly include clustering, differential expression, trajectory inference, gene regulatory network analysis, and cell-cell communication analysis as complementary approaches.
Applications in Disease Research
Trajectory-based differential expression analysis has been applied across diverse biological systems and disease contexts. These applications illustrate the utility of the approach and provide models for study design.
Hematopoiesis and Blood Cell Development
Early applications of trajectory analysis in hematopoiesis demonstrated the power of marker-free approaches to reconstruct differentiation trajectories. Studies in zebrafish identified considerable cell-to-cell variation in the probability of transitioning to committed states within transcriptionally similar stem and progenitor populations. The analysis revealed that suppression of ribosomal gene transcription and upregulation of lineage-specific factors coordinately control lineage differentiation, and that this program is conserved between zebrafish and higher vertebrates.
Cancer Biology and Tumor Microenvironment
Trajectory-based differential expression has been applied to understand tumor heterogeneity and the dynamics of immune cells within the tumor microenvironment. Studies of colorectal cancer have used trajectory inference to characterize T cell heterogeneity and exhaustion dynamics, identifying progressive exhaustion pathways with distinct differentiation fates. The analysis revealed that exhausted T cells exhibit elevated expression of checkpoint molecules and reduced cytotoxic capacity compared to effector populations, and that exhaustion represents a terminal differentiation state with limited plasticity.
In lung adenocarcinoma, pseudo-time analysis has been used to identify pseudo-time differential genes associated with tumor development and to construct prognostic risk stratification. The identified genes were validated through multicenter datasets and immunohistochemistry, demonstrating the translational potential of trajectory-based analysis.
Fibrosis and Interstitial Lung Disease
Spatial transcriptomics combined with trajectory-based differential expression analysis identified key molecules involved in progressive fibrosis in interstitial lung diseases. The analysis identified a gene module enriched in young fibrotic lesions, with hub genes including COL1A2, COL3A1, COL1A1, and SPARC. Immunohistochemical validation showed that SPARC was highly expressed in young fibrotic lesions but not in old scarring lesions, and higher SPARC expression was associated with prognosis in a cohort of patients with unclassifiable interstitial lung diseases.
Autoimmune Disease
Single-cell RNA sequencing with trajectory analysis has been applied to characterize immune cell features in primary Sjogren's disease. The analysis identified distinct immune transcriptional reprogramming in patients with both anti-SSA and anti-SSB antibodies compared to those with only anti-SSA antibodies, with upregulation of cytotoxicity-related genes in CD8+ T cells and interferon-related transcriptional program remodeling in mucosal-associated invariant T cells.
Pancreatic Cancer
Trajectory-based analysis in pancreatic ductal adenocarcinoma revealed that PTPRR is almost exclusively expressed in pancreatic ductal epithelial cells, with higher levels in metastatic lesions than in primary tumors. The analysis identified that PTPRR-high epithelial cells showed elevated stemness and formed a specific niche with cancer-associated fibroblasts through EDN, EGF, and PDGF signaling.
Software Ecosystems and Reproducible Workflows
The practical implementation of trajectory-based differential expression analysis depends on the software environment and the reproducibility standards applied throughout the workflow.
R and Bioconductor Ecosystem
The R programming language and the Bioconductor project provide the primary software infrastructure for trajectory-based differential expression analysis. Bioconductor hosts packages for single-cell data handling, trajectory inference, and differential expression testing, including tradeSeq and related tools. The Bioconductor project provides official package documentation, workflow guides, and installation instructions that support reproducible genomic analysis. Researchers should verify that all packages are installed from the appropriate repositories and that version compatibility is maintained across the analysis environment.
Python and Scanpy Ecosystem
For researchers working in Python, the Scanpy toolkit provides scalable methods for preprocessing, visualization, clustering, pseudotime and trajectory inference, and differential expression testing. Scanpy is designed to handle datasets of more than one million cells and uses the AnnData class for managing annotated data matrices. The choice between R and Python ecosystems often depends on the specific methods required and the preferences of the research group, but both ecosystems support the core steps of trajectory-based analysis.
Workflow Management and Reproducibility
Reproducible analysis requires more than documenting commands. Workflow management systems and containerization can ensure that analyses run identically across different computing environments. The nf-core community provides standards for building reproducible bioinformatics pipelines, and the Galaxy Training Network offers accessible workflow training and analysis tutorials that emphasize reproducibility. The Carpentries lessons provide foundational training in computing, data handling, shell, Git, and programming that supports the development of reproducible analysis skills.
Limitations and Professional Escalation Criteria
Trajectory-based differential expression analysis has inherent limitations that researchers should recognize. When results are inconsistent with biological expectations or when technical issues compromise the analysis, escalation to more specialized expertise may be appropriate.
When to Seek Additional Expertise
Consult a bioinformatics specialist or computational biologist when the trajectory structure is ambiguous, when different trajectory inference methods produce substantially different results, or when the differential expression results are highly sensitive to preprocessing choices. These situations indicate that the analysis requires more sophisticated handling than standard workflows provide.
When to Question the Biological Interpretation
If the identified genes do not include expected markers of the process under study, or if the expression patterns contradict established biology, the trajectory or the differential expression model may be misspecified. Re-examine the trajectory structure, the root selection, and the cell population subset before accepting the results.
When to Consider Alternative Approaches
If the biological process does not fit a continuous trajectory model, alternative analytical approaches may be more appropriate. Some processes involve discrete state transitions that are better modeled by cluster-based comparisons. Other processes may require RNA velocity analysis to infer directionality, or spatial transcriptomics to incorporate tissue context.
Quality Control and Validation Strategies
Rigorous quality control and validation are essential for trustworthy trajectory-based differential expression results.
Computational Validation
Several computational approaches can validate trajectory-based differential expression results. Bootstrap or subsampling analyses can assess the stability of the trajectory and the identified genes. Comparison of results across different trajectory inference methods can identify genes that are robustly associated with the process. Permutation tests can assess whether the observed differential expression exceeds what would be expected by chance.
Experimental Validation
Experimental validation is the gold standard for confirming trajectory-based differential expression findings. Immunohistochemistry or immunofluorescence can confirm protein-level expression patterns along the trajectory. Flow cytometry can quantify protein expression across cell populations. Functional assays can test whether the identified genes are causally involved in the process. Published studies that combine computational trajectory analysis with experimental validation provide templates for this approach.
Reporting Standards
Reports of trajectory-based differential expression analysis should include the software versions, parameters, quality control metrics, trajectory characteristics, and complete results tables. The limitations of the analysis should be acknowledged, including the exploratory nature of trajectory inference and the need for validation.
Data Management and Resource Considerations
The scale of single-cell RNA-seq data and the computational demands of trajectory-based analysis require careful data management and resource planning.
Data Storage and Organization
Single-cell RNA-seq datasets can be large, particularly when raw sequencing data, count matrices, and intermediate analysis files are all retained. A clear directory structure and naming convention should be established at the start of a project. Raw data should be stored separately from processed data, and analysis scripts should be organized by analysis step. Public repositories such as NCBI provide infrastructure for depositing and accessing sequencing data, and researchers should follow community standards for data sharing and metadata annotation.
Computational Resource Planning
Trajectory inference and differential expression testing can be computationally intensive, particularly for datasets with hundreds of thousands of cells. Memory requirements and runtime should be estimated before starting large analyses. Cloud computing resources or institutional high-performance computing clusters may be necessary for large datasets. The scalability of tools such as Scanpy, which can handle datasets of more than one million cells, should be considered when selecting the analysis environment.
Version Control and Documentation
Version control for analysis scripts and documentation of the computing environment are essential for reproducibility. Containerization tools can capture the exact software environment used for an analysis, ensuring that the analysis can be rerun identically at a later time. The Carpentries lessons provide training in version control with Git and other foundational computing skills that support reproducible research practices.
Frequently Asked Questions
What is the difference between cluster-based and trajectory-based differential expression?
Cluster-based differential expression compares average gene expression between discrete cell clusters, which is appropriate when cell types are truly distinct. Trajectory-based differential expression models gene expression as a continuous function of pseudotime, preserving the continuous resolution of cellular transitions. This approach can identify genes that change gradually along a differentiation path or that differ between lineage branches, which cluster-based methods may miss.
When should I use trajectory-based differential expression instead of cluster-based analysis?
Use trajectory-based differential expression when your biological question concerns continuous processes such as differentiation, activation, or disease progression. If you are studying discrete cell types or comparing defined conditions, cluster-based analysis is appropriate. Many studies use both approaches, with clustering for cell type annotation and trajectory analysis for understanding dynamic processes.
What is pseudotime and how is it calculated?
Pseudotime is a computational ordering of cells along a trajectory that represents progress through a biological process. It is inferred from gene expression patterns instead of measured directly. Trajectory inference algorithms learn a low-dimensional manifold through the data and project cells onto this manifold, assigning each cell a pseudotime value based on its position.
How do I choose the root cell for trajectory inference?
The root cell should represent the beginning of the biological process, such as a known progenitor or stem cell population. Select the root based on independent biological knowledge instead of computational criteria alone. Test the robustness of your results to alternative root choices, as the pseudotime ordering and downstream differential expression can be sensitive to root selection.
What is zero inflation and why does it matter for trajectory-based differential expression?
Zero inflation refers to the excessive number of zero counts in single-cell RNA-seq data caused by technical dropout, where genes are not detected in some cells despite being expressed. Zero inflation can inflate false positive rates in differential expression tests. Methods that account for zero inflation, such as tradeSeq with observation-level weights and scFPC-DE, are recommended for trajectory-based analysis.
How many cells do I need for trajectory-based differential expression?
The required number of cells depends on the complexity of the trajectory and the statistical power needed to detect expression changes. Dense sampling across the process of interest, particularly at branch points, improves the resolution of pseudotime ordering and the power to detect differentially expressed genes. There is no universal minimum, but sparse sampling of intermediate states can compromise the analysis.
Can I apply trajectory-based differential expression to single-nucleus RNA-seq data?
Yes, trajectory-based differential expression can be applied to single-nucleus RNA-seq data. The nuclear transcriptome differs from the whole-cell transcriptome, so the detected expression dynamics may differ from whole-cell analyses. Researchers should be aware of these differences and interpret results accordingly.
How should I validate trajectory-based differential expression results?
Validation can be computational or experimental. Computational validation includes assessing the stability of results across trajectory inference methods and preprocessing choices. Experimental validation includes immunohistochemistry, flow cytometry, or functional assays to confirm the expression patterns and biological relevance of identified genes. Published studies that combine computational analysis with experimental validation provide models for this approach.
Related Bioinformatics Guides
- Single-Cell Sequencing Depth: How Much Is Enough?
- Single-Cell RNA Sequencing Quality Control: A Practical Guide to Filtering and Metrics
- Single-cell RNA-seq Trajectory Inference and Cell Lineage Tracing
- Single-Cell RNA Sequencing Depth: A Cost-Benefit Analysis for Experimental Design
- Single-Cell RNA Velocity: Inferring Cellular Dynamics
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- Trajectory-based differential expression analysis for single-cell sequencing data.. Nature communications, 2020.
- Spatial transcriptomics identifies SPARC as a prognostic marker in interstitial lung diseases.. The Journal of pathology, 2025.
- scFPC-DE: Robust Differential Expression Analysis Along Single Cell Trajectories via Functional Principal Component Analysis.. bioRxiv : the preprint server for biology, 2025.
- BLTSA: pseudotime prediction for single cells by branched local tangent space alignment.. Bioinformatics (Oxford, England), 2023.
- Single-cell RNA-sequencing uncovers transcriptional states and fate decisions in haematopoiesis.. Nature communications, 2017.
- Human surface ectoderm and amniotic ectoderm are sequentially specified according to cellular density.. Science advances, 2024.
- Compression of quantification uncertainty for scRNA-seq counts.. Bioinformatics (Oxford, England), 2021.
- Classification with Missing Data - A NIFty Pipeline for Single-Cell Proteomics.. bioRxiv : the preprint server for biology, 2026.
- Analytical Methods and Application of Single-Cell and Single-Nucleus Transcriptomics in the Study of Ischemic Stroke.. 2026.
- Characterization of Peripheral Blood Immune Cell Features in SSA/SSB Double-Positive Primary Sjögren's Disease Based on Single-Cell RNA Sequencing.. 2026.
- Single-cell transcriptomic profiling reveals CD27<,sup>,+<,/sup>, cytotoxic T Cell heterogeneity and exhaustion dynamics in colorectal cancer tumor microenvironment.. 2026.
- Single-cell and spatial transcriptomic profiling reveals the expression characteristics of PTPRR in epithelial cells and its potential implications in pancreatic cancer metastasis.. 2026.
- SCANPY: large-scale single-cell gene expression data analysis. Genome Biology, 2018.
- A Differential Evolution-based Pseudotime Estimation Method for Single-cell Data. International Journal of Advanced Computer Science and Applications, 2024.
- Tumor prognostic risk stratification based on pseudo-time analysis of single-cell sequencing for patients with lung adenocarcinoma. Journal of Thoracic Disease, 2025.
- BLASE: Bulk Linkage Analysis for Single Cell Experiments - Teasing Out the Secrets of Bulk Transcriptomics with Trajectory Analysis. bioRxiv, 2025.
- Single-cell RNA sequencing reveals intratumor heterogeneity and prognostic contributions of γδ T cells in hepatocellular carcinoma. Biomedical Signal Processing and Control, 2024.
- Single-cell transcriptome analysis reveals heterogeneity and convergence of the tumor microenvironment in colorectal cancer. Frontiers in Immunology, 2023.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.