From Trajectory to Ligand-Receptor Pairs: A Workflow for Integrating Pseudotime and Cell-Cell Communication Analysis
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Integrating pseudotime trajectory inference with cell-cell communication analysis allows for the dynamic interrogation of signaling axes during cellular transitions, moving beyond static snapshots to understand when and how cells communicate along developmental or disease paths.
- The workflow emphasizes modularity and validation at each stage, starting with rigorous quality control of scRNA-seq/snRNA-seq data, followed by robust cell type annotation, trajectory reconstruction (e.g., using Monocle, Slingshot, scVelo), and finally, communication inference (e.g., CellChat, NicheNet, CellPhoneDB).
- Validation of trajectory topology and gene expression smoothness along inferred paths is critical before communication analysis to ensure biological plausibility and prevent propagation of errors, as trajectory inference is an unsupervised reconstruction sensitive to method choice and parameters.
- Cell-cell communication inference tools provide hypothesis generation, requiring experimental validation (e.g., co-culture assays, spatial transcriptomics, immunohistochemistry) to confirm predicted ligand-receptor interactions and their functional relevance, as mRNA co-expression does not guarantee protein-level signaling or biological activity.
- Common failure patterns include insufficient cell numbers per group for reliable communication estimates, failure to address batch effects, overinterpreting computational predictions without experimental validation, and ignoring the inherent uncertainty in trajectory reconstructions.
Single-cell RNA sequencing (scRNA-seq) and single-nucleus RNA sequencing (snRNA-seq) generate high-dimensional snapshots of cellular states, but the biological questions that motivate most studies concern change: how cells transition from one state to another, and how those transitions are coordinated across cell types. Trajectory inference methods reconstruct developmental or activation paths from static snapshots, while cell-cell communication tools infer ligand-receptor interactions from expression data. Connecting these two analytical modes allows researchers to ask which signaling axes are dynamically regulated along a differentiation path and which cell populations act as senders or receivers at specific transition points. This article provides a practical workflow for integrating pseudotime analysis with ligand-receptor inference tools such as CellChat, NicheNet, and CellPhoneDB, with emphasis on data inputs, quality checks, interpretation limits, and reporting standards.
The workflow described here is intended for biology students, researchers, and laboratory professionals who have already performed basic scRNA-seq processing and clustering and now want to extend their analysis to dynamic signaling questions. The approach assumes familiarity with R or Python environments, basic quality control procedures, and standard dimensionality reduction methods. The practical outcome is a reproducible pipeline that connects trajectory position to communication output, enabling identification of ligand-receptor pairs that change along a developmental or disease progression path.
At a Glance
The table below summarizes the key decisions in the integrated workflow, the tools commonly used at each stage, and the primary outputs that feed into downstream interpretation.
| Workflow Stage | Common Tools | Primary Output | Key Quality Check |
|---|---|---|---|
| Data preprocessing and quality control | Seurat, Scanpy, Bioconductor packages | Filtered count matrix with cell metadata | Doublet detection, mitochondrial read fraction, library size distribution |
| Cell type annotation | SingleR, Garnett, manual marker inspection | Annotated cell clusters | Marker gene specificity, cluster stability across resolutions |
| Trajectory inference | Monocle, Slingshot, scVelo, PAGA | Pseudotime values and lineage assignments | Trajectory topology validation, gene expression smoothness along path |
| Cell-cell communication inference | CellChat, NicheNet, CellPhoneDB | Ligand-receptor interaction tables with significance scores | Number of cells per group, expression threshold sensitivity |
| Integration of trajectory and communication | Custom R or Python scripts | Dynamic signaling output, sender-receiver pairs per pseudotime bin | Consistency across trajectory branches, biological plausibility |
The workflow is modular. Each stage produces outputs that can be validated independently before proceeding to the next stage. This design reduces the risk of propagating errors from early analysis steps into downstream communication inference.
Understanding the Analytical Problem
Why Trajectory Inference Alone Is Insufficient
Trajectory inference methods assign each cell a position along a reconstructed path, typically representing a continuous biological process such as differentiation, activation, or disease progression. These methods are powerful for identifying genes whose expression changes monotonically or in a branch-specific manner along the path. However, trajectory analysis does not directly address intercellular signaling. A gene that changes along a trajectory may encode a ligand, a receptor, or an intracellular protein with no communication function. Knowing that a transcription factor increases in expression along a path does not reveal which neighboring cells receive the signal or which downstream pathways are activated.
Published studies illustrate this gap. In a study of macrophage and fibroblast dynamics after myocardial infarction, researchers used scVelo, PAGA, and Slingshot to reconstruct differentiation trajectories of fibroblast and macrophage subpopulations, then used CellPhoneDB and NicheNet to infer fibroblast-macrophage interactions [<a href="#ref-1">1</a>]. The trajectory analysis identified reparative cardiac fibroblasts and matrifibrocytes appearing at different times after injury, while the communication analysis revealed that SPP1-high macrophages interact with reparative cardiac fibroblasts in processes related to collagen deposition and scar formation [<a href="#ref-1">1</a>]. Neither analysis alone would have connected the temporal appearance of specific fibroblast subsets to the macrophage-derived signals that likely drive their behavior.
Why Communication Inference Alone Is Insufficient
Cell-cell communication tools infer potential ligand-receptor interactions based on co-expression of ligand genes in sender cells and receptor genes in receiver cells. These tools generate extensive interaction tables, often with hundreds or thousands of significant pairs across multiple cell types. The output is inherently static. It represents the communication potential at the time of sampling, averaged across all cells in each annotated group. If a cell population contains cells at multiple stages of differentiation, the communication output blends signals from early and late states, obscuring dynamic changes.
The integration of spatial transcriptomics with single-cell data has highlighted this limitation. In a study of lung adenocarcinoma progression from adenocarcinoma in situ to invasive adenocarcinoma, researchers combined scRNA-seq with spatial transcriptomics to characterize the invasion trajectory [<a href="#ref-2">2</a>]. They found that UBE2C-positive cancer cells increased during invasion and localized to the peripheral cancer region, while communication analysis revealed constitutive TGF-beta signaling between cancer cells and tumor microenvironment cells [<a href="#ref-2">2</a>]. The spatial component was essential for understanding which cells were physically positioned to communicate, and the trajectory component was essential for understanding when along the invasion path specific signaling became active.
The Value of Integration
Integrating trajectory and communication analyses addresses three questions that neither approach answers alone. First, which ligand-receptor pairs change significantly along a trajectory of interest? Second, which cell populations act as dominant senders or receivers at specific pseudotime positions? Third, do the dynamics of communication align with known biological transitions, such as the onset of fibrosis, immune evasion, or tissue remodeling?
A study of colorectal cancer metastasis used this integrated approach to identify the transcription factor BHLHE40 as a driver of epithelial-mesenchymal transition [<a href="#ref-3">3</a>]. The researchers analyzed single-cell data from nonmetastatic and metastatic primary tumors, used pseudotime trajectory analysis to show that malignant epithelial cells transdifferentiate into CXCL1-positive cancer-associated fibroblasts and then into SFRP2-positive fibroblasts, and used cell-cell communication analysis to identify BHLHE40 as a probable key regulator [<a href="#ref-3">3</a>]. The trajectory analysis established the differentiation path, and the communication analysis connected that path to a specific transcriptional driver with prognostic significance.
Core Principles of the Integrated Workflow
Principle 1: Define the Biological Question Before Selecting Tools
The choice of trajectory inference method and communication inference tool depends on the biological question. If the question concerns a linear differentiation path with a known start and end point, methods such as Slingshot or Monocle may be appropriate. If the question concerns branching decisions or cell fate commitment, methods such as PAGA or scVelo that can represent complex topologies may be more suitable. Similarly, if the question concerns signaling between known cell types, CellPhoneDB may provide sufficient resolution. If the question concerns ligand-receptor pairs that drive a specific downstream response, NicheNet may be more appropriate because it incorporates prior knowledge of signaling pathways and transcriptional targets.
Principle 2: Match Data Resolution to the Question
Communication inference requires sufficient cells per group to produce stable estimates of ligand and receptor expression. If a trajectory branch contains only a small number of cells, splitting that branch into pseudotime bins for dynamic communication analysis may produce unreliable results. The study of prostate cancer heterogeneity integrated single-cell RNA sequencing, spatial transcriptomics, and bulk ATAC sequencing to examine macrophage and neutrophil state transitions along cancer progression [<a href="#ref-4">4</a>]. The authors examined cell-cell communication in situ using spatial transcriptome analysis, which provided positional context that standard scRNA-seq cannot offer [<a href="#ref-4">4</a>]. When spatial data are available, they can substantially improve the interpretation of communication inference by confirming that putative interacting cells are physically proximate.
Principle 3: Validate Trajectory Results Before Communication Analysis
Trajectory inference is an unsupervised reconstruction of a continuous process from static data. The results depend heavily on the choice of starting point, the method used, and the genes included in the analysis. Before using pseudotime values as input for communication analysis, researchers should validate that the trajectory is biologically meaningful. This validation can include checking that known marker genes change expression smoothly along the path, that the trajectory topology is consistent across methods, and that the inferred direction of progression matches experimental knowledge.
Principle 4: Treat Communication Inference as Hypothesis Generation
Cell-cell communication tools generate predictions about potential interactions based on expression data. These predictions require experimental validation. The study of cervical cancer fibroblasts used functional experiments to investigate the role of SDC1, a mediator of fibroblast-tumor crosstalk identified through integrated analysis [<a href="#ref-5">5</a>]. The researchers used fibroblast-tumor cell co-culture systems and functional assays to examine the paracrine role of SDC1, and they used multiplex immunofluorescence and immunohistochemistry to characterize spatial distribution in human tissue samples [<a href="#ref-5">5</a>]. This combination of computational prediction and experimental validation represents the standard that integrated analyses should aim to meet.
Data Inputs and Preparation
Single-Cell and Single-Nucleus Data
The workflow accepts both scRNA-seq and snRNA-seq data. The choice between these approaches affects the interpretation of results. Single-nucleus data capture nuclear transcripts and may miss cytoplasmic mRNAs, which can affect the detection of certain ligand and receptor genes. A study of myocardial infarction integrated scRNA-seq and snRNA-seq datasets from 12 different studies to analyze fibroblast and macrophage dynamics [<a href="#ref-1">1</a>]. The integration of both data types increased the number of cells available for analysis and allowed cross-validation of findings across platforms.
Data should be obtained from reputable repositories. The National Center for Biotechnology Information provides access to sequence data, gene expression datasets, and associated metadata through its various databases [<a href="#ref-6">6</a>]. Researchers should document the accession numbers for all datasets used, the version of the reference genome used for alignment, and the gene annotation version. This documentation is essential for reproducibility.
Quality Control
Quality control procedures for scRNA-seq data are well established and should be completed before trajectory or communication analysis. Standard filters include removing cells with low total read counts, high mitochondrial read fractions, and evidence of doublet contamination. The specific thresholds depend on the tissue type, the dissociation protocol, and the sequencing platform. Researchers should examine the distributions of these metrics and set thresholds based on the data instead of applying arbitrary cutoffs.
The European Bioinformatics Institute provides training materials on bioinformatics data resources and practical analysis education, including guidance on quality control for single-cell data [<a href="#ref-7">7</a>]. Bioconductor hosts official package documentation and workflow resources for reproducible genomic analysis, including packages specifically designed for single-cell quality control and normalization [<a href="#ref-8">8</a>].
Data Integration Across Samples
When combining multiple samples or datasets, batch effects must be addressed before trajectory or communication analysis. Integration methods such as Harmony, Seurat's integration functions, or Scanorama can align cells across batches while preserving biological variation. The choice of integration method affects downstream results, and researchers should compare the integrated data with the original data to ensure that biological signals are preserved.
The study of metabolic dysfunction-associated steatotic liver disease progression to metabolic dysfunction-associated steatohepatitis integrated public single-cell, spatial, and bulk transcriptomic datasets [<a href="#ref-9">9</a>]. The authors mapped microenvironmental remodeling and regulatory networks during disease progression, identifying a DTNA-positive macrophage subpopulation enriched in MASH [<a href="#ref-9">9</a>]. This integration across data types required careful batch correction and validation to ensure that the identified subpopulation was not an artifact of technical variation.
Trajectory Inference Methods
Monocle and Pseudotime Reconstruction
Monocle is one of the most widely used tools for trajectory inference. It constructs a minimum spanning tree on the cell-cell distance graph and assigns each cell a pseudotime value representing its position along the reconstructed path. Monocle can identify branch points where cells make fate decisions and can order genes by their expression dynamics along the trajectory.
A study of regulatory T cells in hormone receptor-positive breast cancer used Monocle for pseudo-time analysis to visualize T cell differentiation trajectories [<a href="#ref-10">10</a>]. The researchers analyzed scRNA-seq data from breast cancer samples, extracted T cells and subsets, and used Monocle to reconstruct differentiation paths [<a href="#ref-10">10</a>]. This analysis revealed the phenotypic and functional profiles of regulatory T cells and their roles in tumor immune evasion [<a href="#ref-10">10</a>].
Slingshot and Lineage Inference
Slingshot combines cluster-based lineage inference with smooth pseudotime trajectories. It requires an initial clustering of cells and then identifies lineages that connect clusters in a biologically meaningful order. Slingshot is particularly useful when the trajectory structure is expected to be complex, with multiple branching points and lineages.
The myocardial infarction study used Slingshot alongside scVelo and PAGA to analyze differentiation trajectories of fibroblast and macrophage subpopulations [<a href="#ref-1">1</a>]. The combination of methods allowed the researchers to cross-validate trajectory topology and to identify subpopulations that appeared at different times after injury [<a href="#ref-1">1</a>].
scVelo and RNA Velocity
RNA velocity methods use the ratio of spliced to unspliced transcripts to estimate the rate and direction of transcriptional change. scVelo extends this approach with a dynamical model that can capture more complex transcriptional dynamics, including induction and repression phases. RNA velocity provides a complementary view to pseudotime methods because it is based on the actual transcriptional state of each cell instead of on similarity to other cells.
The study of gingival tissue from periodontitis patients undergoing orthodontic treatment used RNA velocity alongside pseudotime trajectory analysis to characterize macrophage polarization dynamics [<a href="#ref-11">11</a>]. The researchers identified 11 distinct cell populations with significant macrophage heterogeneity, including M1-like pro-inflammatory and M2-like tissue-remodeling subtypes with intermediate transitional states [<a href="#ref-11">11</a>]. Trajectory analysis revealed dynamic polarization pathways with multiple branching points, demonstrating phenotypic plasticity [<a href="#ref-11">11</a>].
PAGA and Topology Inference
PAGA (Partition-based Graph Abstraction) generates a coarse-grained representation of the cell-cell connectivity graph, preserving the global topology of the data while reducing noise. PAGA is useful for identifying the overall structure of the data, including disconnected components and branching points, before applying more detailed trajectory methods.
Choosing Among Methods
No single trajectory inference method is universally superior. The choice depends on the data structure, the biological question, and the computational resources available. Researchers should run multiple methods and compare the results. If different methods produce substantially different trajectories, the reasons for the discrepancy should be investigated before proceeding to communication analysis.
The study of fetal goat skeletal muscle development used trajectory analysis to trace a maturation continuum from differentiation-competent myocytes to contractile fibers [<a href="#ref-12">12</a>]. The authors identified RUNX2 mesenchymal progenitors, fibro-adipogenic progenitors, myofibroblasts, endothelial cells, macrophages, differentiating myocytes, and mature skeletal muscle fibers [<a href="#ref-12">12</a>]. Pseudotime analysis revealed sequential activation of extracellular matrix remodeling, cytoskeletal stabilization, and sarcomere assembly along the maturation path [<a href="#ref-12">12</a>]. This study demonstrates the value of trajectory analysis for understanding tissue development, but it also illustrates the need for careful annotation of cell types before trajectory reconstruction.
Cell-Cell Communication Inference Tools
CellChat
CellChat infers cell-cell communication from gene expression data using a manually curated database of ligand-receptor interactions and their associated signaling pathways. CellChat quantifies the communication probability between cell groups based on the expression of ligands in sender cells and receptors in receiver cells, and it can identify dominant senders, receivers, mediators, and influencers in the communication network.
A study of high-grade serous ovarian cancer used CellChat to analyze intercellular communication and Monocle for pseudo-time analysis [<a href="#ref-10">10</a>]. The researchers analyzed scRNA-seq data from ovarian cancer samples and identified distinct phenotypic and functional profiles of regulatory T cells [<a href="#ref-10">10</a>]. CellChat analysis revealed the communication patterns between T cells and other cell types in the tumor microenvironment [<a href="#ref-10">10</a>].
NicheNet
NicheNet uses prior knowledge of signaling pathways to predict which ligands affect the expression of target genes in receiver cells. NicheNet integrates ligand-receptor interactions, signaling pathway information, and transcriptional regulatory networks to identify the ligands most likely to drive observed gene expression changes in a cell population of interest.
The myocardial infarction study used NicheNet alongside CellPhoneDB to infer fibroblast-macrophage interactions [<a href="#ref-1">1</a>]. The researchers identified a macrophage subset expressing a gene signature conserved in both human and mouse hearts, and communication analysis indicated that SPP1-high macrophage interactions with reparative cardiac fibroblasts are mainly involved in collagen deposition and scar formation [<a href="#ref-1">1</a>]. NicheNet provided the regulatory context for these interactions by linking ligand signals to downstream transcriptional programs.
CellPhoneDB
CellPhoneDB uses a curated database of ligand-receptor complexes to identify significant interactions between cell populations. It performs a permutation test to assess whether the observed co-expression of a ligand-receptor pair is greater than expected by chance. CellPhoneDB is computationally efficient and can handle large datasets.
Choosing Among Communication Tools
The choice of communication inference tool depends on the question. CellChat provides a comprehensive network-level view of communication patterns and is useful for identifying dominant signaling pathways. NicheNet is useful for identifying the ligands most likely to drive specific gene expression changes in a receiver population. CellPhoneDB is useful for identifying statistically significant ligand-receptor pairs between defined cell populations.
The study of osteoarthritis used cell-cell communication analysis to reveal strong bidirectional interactions between M1 and M2 macrophages [<a href="#ref-13">13</a>]. Integrative analysis of macrophage subpopulations, pseudotime, weighted gene co-expression network analysis, and cell-cell communication identified SEMA4A as the only overlapping key gene [<a href="#ref-13">13</a>]. Ligand-receptor analysis showed that SEMA4A-PLXNB2 was the predominant interaction pair with the highest communication probability [<a href="#ref-13">13</a>]. This study demonstrates the value of combining multiple analytical approaches to identify key signaling axes.
The Integrated Workflow
Step 1: Define Cell Populations and Trajectories
The first step is to define the cell populations of interest and reconstruct their trajectories. This step requires careful cell type annotation. The quality of the annotation directly affects the quality of both trajectory and communication analysis. Researchers should use multiple annotation strategies, including automatic annotation tools such as SingleR and Garnett, manual inspection of marker genes, and cross-referencing with published cell type signatures.
The study of prostate cancer used SingleR and Garnett for cell type identification, with a focus on the extraction and annotation of T cells and their subsets [<a href="#ref-10">10</a>]. The researchers then used FindAllMarkers to screen for differentially expressed genes and performed gene set enrichment analysis [<a href="#ref-10">10</a>]. This systematic approach to annotation provides a solid foundation for downstream analysis.
Step 2: Reconstruct Trajectories
After annotation, trajectory inference is performed on the cell populations of interest. The choice of method depends on the expected trajectory structure. For linear differentiation paths, Monocle or Slingshot may be appropriate. For complex topologies with branching points, PAGA or scVelo may be more suitable.
The study of diabetic retinopathy reviewed the application of scRNA-seq to investigate disease pathogenesis, focusing on the reconstruction of developmental trajectories to unveil state transitions and the exploration of complex cell-cell communication [<a href="#ref-14">14</a>]. The review emphasized that trajectory reconstruction is a key step in understanding how cell states change during disease progression [<a href="#ref-14">14</a>].
Step 3: Validate Trajectories
Before proceeding to communication analysis, the reconstructed trajectories should be validated. This validation includes checking that known marker genes change expression smoothly along the trajectory, that the trajectory topology is consistent with biological knowledge, and that the inferred direction of progression is plausible.
The study of hepatocellular carcinoma integrated single-cell, bulk, and spatial transcriptome analyses to examine the diversity of cancer-associated fibroblasts [<a href="#ref-15">15</a>]. Using a training cohort of 88 scRNA-seq samples and a validation cohort of 94 samples, encompassing over 1.2 million cells, the researchers classified three fibroblast subpopulations based on highly expressed genes [<a href="#ref-15">15</a>]. Cell trajectory analysis revealed that VEGFA-positive cancer-associated fibroblasts are at the terminal stage of differentiation and are tumor-specific [<a href="#ref-15">15</a>]. The validation of this trajectory was essential for the subsequent communication analysis, which showed that VEGFA-positive fibroblasts promote intra-tumoral angiogenesis through communication with capillary endothelial cells [<a href="#ref-15">15</a>].
Step 4: Perform Communication Inference
Communication inference is performed on the annotated cell populations. The choice of tool depends on the question. CellChat provides a network-level view, NicheNet provides regulatory context, and CellPhoneDB provides statistical significance testing. Multiple tools can be used in parallel, and the results can be compared.
The study of allergic rhinitis performed single-cell RNA sequencing and single-cell ATAC sequencing on nasal mucosa samples from 39 subjects [<a href="#ref-16">16</a>]. The researchers applied differential expression analysis, differentially accessible peaks analysis, cell-cell communication, trajectory inference, and gene regulatory network reconstruction [<a href="#ref-16">16</a>]. They found that the allergic rhinitis epithelium exhibited aberrant differentiation with suppressed maturation of basal and club cells, while fibroblasts displayed inflammatory activation and matrix remodeling signatures [<a href="#ref-16">16</a>]. Epithelial-stromal crosstalk was enhanced in the allergic rhinitis group [<a href="#ref-16">16</a>].
Step 5: Integrate Trajectory and Communication Results
The integration step connects pseudotime position to communication output. This integration can be performed in several ways. One approach is to bin cells by pseudotime and perform communication inference within each bin. This approach reveals how communication patterns change along the trajectory. Another approach is to identify ligand-receptor pairs where the ligand gene changes expression along the sender cell trajectory and the receptor gene changes expression along the receiver cell trajectory. This approach identifies dynamically regulated signaling axes.
The study of sepsis used comprehensive single-cell RNA sequencing analysis, including cell clustering, differential expression analysis, cell-cell communication mapping, and pseudotime trajectory analysis, to explore the roles of identified genes within the sepsis microenvironment [<a href="#ref-17">17</a>]. The risk gene BEND7, predominantly expressed in platelets, was further analyzed using single-cell RNA sequencing, revealing strong interactions with immune cells, particularly monocytes and neutrophils, via the intercellular adhesion molecule signaling pathway [<a href="#ref-17">17</a>]. This integrated approach connected a specific gene to specific cell-cell interactions and immune modulation.
Step 6: Validate Findings
The final step is validation. Computational predictions of ligand-receptor interactions require experimental confirmation. This validation can include co-culture experiments, antibody blockade studies, spatial transcriptomics to confirm physical proximity, and immunohistochemistry to confirm protein expression.
The study of cervical cancer used functional experiments to investigate the role of SDC1, a critical mediator of fibroblast-tumor crosstalk [<a href="#ref-5">5</a>]. The researchers used fibroblast-tumor cell co-culture systems and functional assays to investigate the paracrine role of SDC1, and they used multiplex immunofluorescence and immunohistochemical analyses to characterize the spatial distribution in human cervical cancer tissue samples [<a href="#ref-5">5</a>]. This combination of computational prediction and experimental validation represents the standard for integrated analyses.
Practical Implementation Steps
Setting Up the Analysis Environment
The analysis environment should be set up with reproducibility in mind. This includes documenting the software versions, the package versions, and the analysis scripts. Bioconductor provides official package documentation and workflow resources for reproducible genomic analysis [<a href="#ref-8">8</a>]. The Galaxy Training Network provides accessible workflow training and analysis tutorials that can be adapted for single-cell analysis [<a href="#ref-18">18</a>]. The nf-core documentation describes community pipeline standards for reproducible workflow configuration [<a href="#ref-19">19</a>].
Data Management
Data management is a critical component of reproducible analysis. All raw data files should be stored in a structured directory system, with clear naming conventions. Processed data should be stored separately from raw data, and analysis scripts should be version-controlled. The Carpentries provides foundational training in computing, data, shell, Git, and programming that is directly applicable to managing bioinformatics analysis projects [<a href="#ref-20">20</a>].
Documentation
Documentation should include the source of each dataset, the accession numbers, the preprocessing steps, the quality control thresholds, the trajectory inference parameters, and the communication inference settings. This documentation enables other researchers to reproduce the analysis and to assess the validity of the conclusions.
Records and Measurements
Key Metrics to Record
The following metrics should be recorded at each stage of the workflow:
| Stage | Metric | Purpose |
|---|---|---|
| Quality control | Number of cells before and after filtering | Assess data loss |
| Quality control | Median genes per cell | Assess data quality |
| Quality control | Median UMIs per cell | Assess sequencing depth |
| Quality control | Mitochondrial read fraction distribution | Assess cell viability |
| Annotation | Number of cells per cell type | Assess group sizes for communication inference |
| Trajectory | Number of lineages identified | Assess trajectory complexity |
| Trajectory | Pseudotime range | Assess trajectory coverage |
| Communication | Number of significant ligand-receptor pairs | Assess communication complexity |
| Communication | Number of interactions per cell type | Assess sender and receiver roles |
Recording Decisions
Each analytical decision should be recorded with its rationale. For example, if a quality control threshold is set at a specific mitochondrial read fraction, the rationale for that threshold should be documented. If a trajectory method is chosen over an alternative, the reason for that choice should be recorded. This documentation is essential for reproducibility and for interpreting the results.
Common Failure Patterns
Failure Pattern 1: Insufficient Cells per Group
Communication inference requires sufficient cells per group to produce stable estimates of ligand and receptor expression. If a cell population contains fewer than approximately 50 cells, the expression estimates may be noisy, and the communication inference may produce unreliable results. This problem is exacerbated when cells are binned by pseudotime for dynamic communication analysis.
Failure Pattern 2: Ignoring Batch Effects
Batch effects can create artificial cell populations and distort trajectory inference. If samples are processed in multiple batches, batch correction should be performed before trajectory or communication analysis. The choice of batch correction method can affect the results, and the corrected data should be validated to ensure that biological signals are preserved.
Failure Pattern 3: Overinterpreting Communication Inference
Cell-cell communication tools generate predictions based on expression data. These predictions do not confirm that communication actually occurs. The physical proximity of cells, the presence of protein-level expression, and the functional relevance of the interaction must be validated experimentally. The study of high-grade serous ovarian cancer used spatial transcriptomics, single-cell sequencing, bulk deconvolution, pseudotime reconstruction, and multiplex immunofluorescence to delineate a distinct tumor-promoting macrophage phenotype [<a href="#ref-21">21</a>]. The researchers found that NDRG1-positive macrophages displayed strong VEGF- and SPP1-mediated communication with endothelial cells and occupied hypoxic, angiogenesis-enriched niches [<a href="#ref-21">21</a>]. The spatial and protein-level validation confirmed that NDRG1-positive SPP1-positive macrophages form discrete angiogenic niches [<a href="#ref-21">21</a>].
Failure Pattern 4: Ignoring Trajectory Uncertainty
Trajectory inference is an unsupervised reconstruction, and the results carry uncertainty. Different methods may produce different trajectories, and the choice of starting point can affect the pseudotime ordering. Researchers should assess the robustness of their trajectory results by running multiple methods and comparing the outcomes.
Failure Pattern 5: Using Inappropriate Pseudotime Binning
Binning cells by pseudotime for dynamic communication analysis requires a balance between resolution and statistical power. Too many bins produce noisy estimates, while too few bins obscure dynamic changes. The optimal number of bins depends on the number of cells in the trajectory and the expected dynamics of the signaling process.
Limitations and Interpretation Constraints
Computational Limitations
Trajectory inference and communication inference are computationally intensive. Large datasets may require substantial memory and processing time. The choice of methods should consider the available computational resources. Cloud-based platforms and high-performance computing clusters can address these limitations, but they require additional setup and configuration.
Biological Limitations
The integrated workflow is based on gene expression data, which provides an indirect measure of protein-level signaling. Ligand and receptor genes may be expressed at the mRNA level without corresponding protein expression, and post-translational modifications can affect signaling activity. The workflow identifies potential interactions, not confirmed signaling events.
Statistical Limitations
Communication inference tools use different statistical approaches to assess significance. CellPhoneDB uses a permutation test, while CellChat uses a probabilistic model. The results from different tools may not be directly comparable. Researchers should interpret the results in the context of the specific statistical approach used.
Interpretation Constraints
The integrated workflow identifies ligand-receptor pairs that change along a trajectory. This identification does not establish causality. A ligand-receptor pair may change along a trajectory without driving the transition, and the direction of causality may be from the trajectory to the signaling change instead of from the signaling change to the trajectory. Experimental validation is required to establish causal relationships.
Safety and Regulatory Context
Data Privacy and Ethics
Single-cell datasets derived from human samples may contain sensitive information. Researchers must comply with applicable data protection regulations and institutional review board requirements. Data should be de-identified before analysis, and access to raw data should be restricted to authorized personnel.
Reproducibility Standards
Reproducibility is a core requirement for bioinformatics analysis. The National Center for Biotechnology Information provides access to sequence data and associated metadata, enabling researchers to deposit and access datasets used in published studies [<a href="#ref-6">6</a>]. The European Bioinformatics Institute provides training on data-resource usage and practical analysis education [<a href="#ref-7">7</a>]. Bioconductor provides official package documentation and workflow resources for reproducible genomic analysis [<a href="#ref-8">8</a>]. The Galaxy Training Network provides accessible workflow training and analysis tutorials [<a href="#ref-18">18</a>]. The nf-core documentation describes community pipeline standards for reproducible workflow configuration [<a href="#ref-19">19</a>]. The Carpentries provides foundational training in computing, data, shell, Git, and programming [<a href="#ref-20">20</a>].
Professional Escalation Criteria
Researchers should seek professional guidance when they encounter the following situations:
- The trajectory inference results are inconsistent across multiple methods, and the discrepancies cannot be resolved.
- The communication inference results identify interactions that contradict established biological knowledge.
- The integrated analysis produces results that are not reproducible across computational environments.
- The analysis requires computational resources beyond the available capacity.
- The interpretation of results requires domain expertise beyond the researcher's training.
Common Failure Patterns in Integrated Analysis
Failure Pattern 6: Circular Analysis
A common error is using the same data to define cell populations and then to test hypotheses about those populations. This circularity can lead to overfitting and inflated confidence in the results. Researchers should use independent validation datasets whenever possible.
Failure Pattern 7: Ignoring Cell Type Composition Differences
Differences in cell type composition between samples can confound communication inference. If one sample contains more macrophages than another, the communication analysis may identify macrophage-related interactions that reflect composition differences instead of biological changes. Researchers should account for composition differences in their analysis.
Failure Pattern 8: Overlooking Ligand-Receptor Database Limitations
Ligand-receptor databases are incomplete and may contain errors. Interactions that are not in the database will not be detected, and interactions that are in the database may not be biologically relevant in the tissue of interest. Researchers should be aware of the limitations of the database used and should interpret the results accordingly.
Quality Controls and Validation
Internal Validation
Internal validation involves checking the consistency of the results within the dataset. This validation can include running the analysis with different parameters, comparing results across trajectory methods, and assessing the stability of communication inference across subsamples of the data.
External Validation
External validation involves confirming the results using independent data. This validation can include checking the expression of identified ligands and receptors in published datasets, performing experimental validation in cell lines or animal models, and comparing the results with findings from spatial transcriptomics.
The study of lung adenocarcinoma used spatial transcriptomics to validate the findings from scRNA-seq analysis [<a href="#ref-2">2</a>]. The researchers found that UBE2C-positive cancer cells were spatially distributed in the peripheral cancer region of invasive adenocarcinoma, representing a more malignant phenotype [<a href="#ref-2">2</a>]. This spatial validation strengthened the conclusions drawn from the single-cell analysis.
Reporting Standards
The results of the integrated analysis should be reported with sufficient detail to enable reproduction. This reporting includes the software versions, the parameters used, the quality control thresholds, and the validation procedures. The report should also describe the limitations of the analysis and the interpretation constraints.
Frequently Asked Questions
What is the difference between pseudotime analysis and RNA velocity?
Pseudotime analysis orders cells along a reconstructed trajectory based on transcriptional similarity, assigning each cell a position that represents its progress along a continuous biological process. RNA velocity uses the ratio of spliced to unspliced transcripts to estimate the rate and direction of transcriptional change for each gene, providing a more direct measure of transcriptional dynamics. Pseudotime methods such as Monocle and Slingshot reconstruct the trajectory from the overall structure of the data, while RNA velocity methods such as scVelo infer the direction of change from the balance of spliced and unspliced transcripts. Both approaches can be used to identify dynamic gene expression changes, but they rely on different biological signals and can produce complementary results.
How many cells are needed for reliable communication inference?
The number of cells needed depends on the heterogeneity of the cell populations and the sensitivity of the communication inference tool. As a general guideline, each cell population used in communication inference should contain at least 50 cells, and preferably more than 100 cells, to produce stable estimates of ligand and receptor expression. When cells are binned by pseudotime for dynamic communication analysis, each bin should also contain sufficient cells. If a trajectory branch contains only a small number of cells, splitting that branch into pseudotime bins may produce unreliable results.
Can this workflow be applied to single-nucleus RNA sequencing data?
Yes, the workflow can be applied to snRNA-seq data, but there are important considerations. Single-nucleus data capture nuclear transcripts and may miss cytoplasmic mRNAs, which can affect the detection of certain ligand and receptor genes. Researchers should verify that the genes of interest are detectable in the snRNA-seq data and should interpret the results with this limitation in mind. The myocardial infarction study integrated both scRNA-seq and snRNA-seq datasets, demonstrating that the two data types can be combined for integrated analysis [<a href="#ref-1">1</a>].
What is the best way to validate ligand-receptor interactions identified by communication inference?
The best validation approach combines multiple lines of evidence. Spatial transcriptomics can confirm that putative interacting cells are physically proximate. Immunohistochemistry or immunofluorescence can confirm protein-level expression of the ligand and receptor. Co-culture experiments can test whether the ligand from one cell type affects the behavior of the other cell type. Antibody blockade or genetic knockdown can test whether the interaction is functionally required. The cervical cancer study used co-culture systems and functional assays to investigate the paracrine role of SDC1, a mediator of fibroblast-tumor crosstalk [<a href="#ref-5">5</a>].
How do I choose between CellChat, NicheNet, and CellPhoneDB?
The choice depends on the biological question. CellChat provides a comprehensive network-level view of communication patterns and is useful for identifying dominant signaling pathways and the roles of different cell types as senders, receivers, and mediators. NicheNet is useful for identifying the ligands most likely to drive specific gene expression changes in a receiver population, because it integrates ligand-receptor interactions with signaling pathway information and transcriptional regulatory networks. CellPhoneDB is useful for identifying statistically significant ligand-receptor pairs between defined cell populations, using a permutation test to assess significance. Many studies use multiple tools in parallel and compare the results.
What should I do if different trajectory methods produce different results?
If different trajectory methods produce substantially different results, the reasons for the discrepancy should be investigated before proceeding to communication analysis. Possible explanations include differences in the underlying assumptions of the methods, sensitivity to the choice of starting point, and the presence of disconnected cell populations. Researchers should examine the trajectory topologies produced by each method, check whether known marker genes change expression smoothly along each trajectory, and consider whether the biological question is better addressed by one method than another. If the discrepancies cannot be resolved, the results should be interpreted with caution.
How do I account for batch effects when integrating multiple datasets?
Batch effects should be addressed before trajectory or communication analysis. Integration methods such as Harmony, Seurat's integration functions, or Scanorama can align cells across batches while preserving biological variation. The choice of integration method affects downstream results, and researchers should compare the integrated data with the original data to ensure that biological signals are preserved. The myocardial infarction study integrated scRNA-seq and snRNA-seq datasets from 12 different studies, requiring careful batch correction to enable cross-study comparison [<a href="#ref-1">1</a>].
What are the main limitations of the integrated workflow?
The main limitations include the indirect nature of gene expression data as a measure of protein-level signaling, the incompleteness of ligand-receptor databases, the uncertainty inherent in trajectory inference, and the need for experimental validation of computational predictions. The workflow identifies potential interactions, not confirmed signaling events. Researchers should interpret the results as hypotheses that require experimental testing.
Related Bioinformatics Guides
- Single-Cell Sequencing Workflow: From Sample Preparation to Data Analysis
- Metabolomics Data Analysis in R: A Practical Workflow
- Single-Cell Annotation: A Workflow for Cell Type Identification
- Genomic Data Analysis Tools: A Comparative Guide for Researchers
- How to Interpret Gene Set Enrichment Analysis Results
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
[1] [Macrophage and fibroblast trajectory inference and crosstalk analysis during myocardial infarction using integrated single-cell transcriptomic datasets.](https://pubmed.ncbi.nlm.nih.gov/38867219). Journal of translational medicine, 2024. [2] [Delineating the dynamic evolution from preneoplasia to invasive lung adenocarcinoma by integrating single-cell RNA sequencing and spatial transcriptomics.](https://pubmed.ncbi.nlm.nih.gov/36434043). Experimental & molecular medicine, 2022. [3] [Single-Cell and Spatial Transcriptome Profiling Identifies the Transcription Factor BHLHE40 as a Driver of EMT in Metastatic Colorectal Cancer.](https://pubmed.ncbi.nlm.nih.gov/38657117). Cancer research, 2024. [4] [Integration Analysis of Single-Cell Multi-Omics Reveals Prostate Cancer Heterogeneity.](https://pubmed.ncbi.nlm.nih.gov/38483933). Advanced science (Weinheim, Baden-Wurttemberg, Germany), 2024. [5] [Deciphering the tumor immune microenvironment: single-cell and spatial transcriptomic insights into cervical cancer fibroblasts.](https://pubmed.ncbi.nlm.nih.gov/40616092). Journal of experimental & clinical cancer research : CR, 2025. [6] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information. [7] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute. [8] [Bioconductor](https://bioconductor.org/). Bioconductor Project. [9] [Integrating multi-omics and machine learning systematically deciphers cellular heterogeneity and fibrotic regulatory networks in the progression from MASLD to MASH.](https://pubmed.ncbi.nlm.nih.gov/41545636). NPJ digital medicine, 2026. [10] [Multidimensional single-cell analysis of the molecular characteristics and functional pathways of Regulatory T cells in the microenvironment of HR+ breast cancer.](https://doi.org/10.1007/s12672-026-04947-9). 2026. [11] [Single cell transcriptomic atlas reveals macrophage polarization dynamics and intercellular communication networks in gingival tissue of periodontitis patients undergoing orthodontic treatment.](https://doi.org/10.1016/j.slast.2026.100402). 2026. [12] [Single-Cell Transcriptomic Analysis Reveals Multicellular Coordination and Signaling Rewiring During Fetal Goat Skeletal Muscle Development.](https://doi.org/10.3390/ani16091370). 2026. [13] [SEMA4A signaling in macrophage subpopulations and its implication in osteoarthritis.](https://doi.org/10.3389/fimmu.2026.1847788). 2026. [14] [Single-cell RNA sequencing in exploring the pathogenesis of diabetic retinopathy.](https://pubmed.ncbi.nlm.nih.gov/38946005). Clinical and translational medicine, 2024. [15] [Integration of single-cell and spatial transcriptomics reveals fibroblast subtypes in hepatocellular carcinoma: spatial distribution, differentiation trajectories, and therapeutic potential.](https://pubmed.ncbi.nlm.nih.gov/39966876). Journal of translational medicine, 2025. [16] [Decoding the epithelial-stromal interactome in allergic rhinitis through single-cell multi-omics integration.](https://doi.org/10.1016/j.jaci.2026.06.020). Journal of Allergy and Clinical Immunology, 2026. [17] [EXPLORING THE POTENTIAL OF BEND7 AS AN IMMUNOMODULATORY BIOMARKER IN SEPSIS THROUGH INTEGRATIVE GENOMIC AND TRANSCRIPTOMIC ANALYSIS](https://doi.org/10.1097/SHK.0000000000002529). Shock, 2024. [18] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project. [19] [nf-core Documentation](https://nf-co.re/docs). nf-core. [20] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries. [21] [Spatial and single-cell profiling identify NDRG1⁺ macrophages as key hallmark of angiogenic remodeling and platinum resistance in HGSOC.](https://doi.org/10.1016/j.tranon.2026.102905). 2026.This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.