How to Run CellChat for Cell-Cell Communication Analysis: A Step-by-Step Tutorial with Seurat Objects
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- CellChat infers cell-cell communication by modeling ligand-receptor interactions using a mass-action framework, accounting for multisubunit structures and cofactors, and quantifies communication probabilities between cell groups based on gene expression.
- The workflow necessitates a Seurat object with normalized counts and robust cell type annotations in its metadata, followed by species-specific database selection (human or mouse) for accurate ligand-receptor pair matching.
- Core analysis involves computing communication probabilities (
computeCommunProb), filtering by minimum cell count, aggregating to pathway level (computeCommunProbPathway), and summarizing networks (aggregateNet) for subsequent visualization. - Visualization options include circle plots, chord diagrams, bubble plots for specific ligand-receptor pairs, and heatmaps for signaling pathways, enabling exploration of network structure and signaling dynamics.
- Comparative analysis across conditions is achieved by merging multiple CellChat objects, allowing for the identification of differential communication patterns and signaling pathway activity using functions like
mergeCellChatandcompareInteractions. - Interpretation of CellChat output requires understanding that communication probabilities are inferences based on transcriptomic data, not direct measures of signaling activity, and necessitate experimental validation; limitations include database completeness and the absence of spatial context for dissociated single-cell data.
CellChat is a computational tool that infers and analyzes cell-cell communication networks from single-cell transcriptomic data using a simplified mass-action model that incorporates ligand-receptor interactions with multisubunit structure and modulation by cofactors. This tutorial provides a reproducible workflow for researchers who have a Seurat object and need to prepare it for CellChat, run the analysis, and interpret the communication networks that result. The workflow assumes a basic understanding of R and single-cell data analysis but does not require specialized bioinformatics training. The complete protocol typically takes about five minutes depending on dataset size.
At a Glance
The table below summarizes the key decisions and inputs required for a CellChat analysis workflow.
| Workflow Stage | Primary Input | Key Decision | Expected Output |
|---|---|---|---|
| Data preparation | Seurat object with normalized counts | Confirm data slot and cell metadata annotations | CellChat object ready for analysis |
| Database selection | Species and signaling context | Choose CellChatDB human or mouse | Ligand-receptor database for inference |
| Communication inference | CellChat object with defined cell groups | Set population size and number of permutation tests | Communication probabilities between cell groups |
| Network visualization | Inferred CellChat object | Select signaling pathway or ligand-receptor pair | Bubble plots, chord diagrams, heatmaps |
| Comparative analysis | Two or more CellChat objects | Define condition labels and comparison order | Differential communication and signaling patterns |
Understanding CellChat and Its Place in Single-Cell Analysis
CellChat was developed to address the challenge of integrating known molecular interactions with single-cell transcriptomic measurements to identify and analyze complex cell-cell communication networks. The tool quantifies the signaling communication probability between two cell groups using a mass-action-based model that accounts for the core interaction between ligands and receptors with multisubunit structure along with modulation by cofactors. CellChat performs systematic and comparative analysis of cell-cell communication using quantitative metrics and machine-learning approaches (CellChat protocol, Nature Protocols 2025).
The updated CellChat v2 includes additional comparison functionalities, an expanded database of ligand-receptor pairs with functional annotations, and an Interactive CellChat Explorer. The R implementation of CellChat v2 and its tutorials are available at the official GitHub repository. The protocol requires a basic understanding of R and single-cell data analysis but no specialized bioinformatics training (CellChat protocol, Nature Protocols 2025).
CellChat is one of several tools available for cell-cell communication inference. Other approaches include CausalCCC, a web server that reconstructs gene-gene interaction pathways across interacting cell types from single-cell or spatial transcriptomic data. CausalCCC integrates a causal network reconstruction method with internally computed ligand-receptor pairs using LIANA+ and can also accept user-defined ligand-receptor pairs from methods such as NicheNet, CellChat, and Misty (CausalCCC, Nucleic Acids Research 2025). The choice of tool depends on whether the research question requires only ligand-receptor interaction scores or a more comprehensive picture that includes upstream and downstream intracellular pathways in sender and receiver cells.
For spatial transcriptomics data, benchmarking studies have demonstrated substantial variability in tool performance across spatial resolutions, tissue contexts, and platforms (Benchmarking CCI tools, Genome Biology 2026). Tools such as CytoSignal infer the locations and dynamics of cell-cell communication at cellular resolution from spatial transcriptomic data, enabling identification of spatial gradients in signaling strength and quantification of contact-dependent versus diffusible interactions (CytoSignal, Nature Genetics 2026). These spatial tools address different questions than CellChat, which is designed primarily for dissociated single-cell data.
Preparing Your Seurat Object for CellChat
Required Seurat Object Structure
CellChat expects a Seurat object that has already undergone standard single-cell RNA sequencing analysis. The current best practices in single-cell RNA-seq analysis include pre-processing steps such as quality control, normalization, data correction, feature selection, and dimensionality reduction, followed by cell-level downstream analysis including clustering and cell type annotation (Current best practices in scRNA-seq analysis, Molecular Systems Biology 2019). Your Seurat object should have normalized data stored in the data slot and cell type annotations stored in the metadata.
The key requirement is that the Seurat object must have a metadata column that defines cell groups. These groups typically correspond to cell types identified through clustering and marker gene analysis. The quality of your cell type annotations directly affects the biological interpretability of the CellChat results. If cell types are poorly defined or contain mixed populations, the inferred communication networks will reflect those ambiguities.
Creating the CellChat Object
The first step is to load the CellChat package and create a CellChat object from your Seurat object. The createCellChat function takes the Seurat object, specifies the group-by variable that contains cell type annotations, and optionally specifies the assay to use. The function extracts the normalized data and cell group information to construct a CellChat object.
library(CellChat)
library(Seurat)
cellchat <- createCellChat(object = seurat_obj, group.by = "cell_type", assay = "RNA")
After creating the CellChat object, you must set the ligand-receptor interaction database. CellChat provides databases for human and mouse. The choice of database must match the species of your data. Using the wrong species database will produce misleading results because the ligand-receptor pairs will not correspond to the genes expressed in your cells.
CellChatDB <- CellChatDB.human
cellchat@DB <- CellChatDB
Subsetting the Expression Data
CellChat requires that the expression data be subset to the genes present in the ligand-receptor database. This step reduces the size of the data and ensures that only genes relevant to cell-cell communication are considered. The subsetData function performs this filtering.
cellchat <- subsetData(cellchat)
This step is important for computational efficiency. The ligand-receptor database contains a curated set of genes, and subsetting to those genes reduces the memory footprint and speeds up subsequent analysis steps. The subsetting does not discard information about cell identities or group structure, only genes that are not part of the communication database.
Identifying Overexpressed Genes and Communication
The next step identifies overexpressed genes and infers the communication probabilities. The identifyOverExpressedGenes function determines which ligand-receptor genes are expressed above a threshold in each cell group. The identifyOverExpressedInteractions function then identifies ligand-receptor pairs where both the ligand and receptor are overexpressed in the respective sender and receiver cell groups.
cellchat <- identifyOverExpressedGenes(cellchat)
cellchat <- identifyOverExpressedInteractions(cellchat)
These steps prepare the data for the core communication inference. The overexpressed gene identification uses a statistical approach to determine which genes are expressed at levels higher than expected by chance in each cell group. This step is analogous to differential expression analysis but is applied specifically to genes in the ligand-receptor database.
Running the Core Communication Inference
Computing Communication Probabilities
The computeCommunProb function is the core inference step in CellChat. It calculates the communication probability between each pair of cell groups for each ligand-receptor pair in the database. The function uses the mass-action model that incorporates the core interaction between ligands and receptors with multisubunit structure along with modulation by cofactors (CellChat protocol, Nature Protocols 2025).
cellchat <- computeCommunProb(cellchat, type = "triMean")
The type parameter controls how the average gene expression is calculated within each cell group. The default "triMean" uses a trimmed mean that is robust to outliers. Other options include "trimean", "median", and "mean". The choice of averaging method can affect the results, particularly for cell groups with heterogeneous gene expression.
The computeCommunProb function also performs permutation tests to assess the statistical significance of the inferred communication probabilities. The number of permutation tests can be controlled with the nboot parameter. More permutations provide more stable p-values but increase computation time.
Filtering Communication Based on Cell Number
CellChat includes a step to filter out communication events between cell groups with very few cells. The filterCommunication function removes interactions where the number of cells in either the sender or receiver group is below a minimum threshold. This filtering reduces false positives that can arise from sparse cell populations.
cellchat <- filterCommunication(cellchat, min.cells = 10)
The minimum cell number threshold is a user decision. A threshold that is too low may retain noisy communication events from rare cell populations. A threshold that is too high may discard biologically meaningful interactions involving rare cell types. The appropriate threshold depends on the total number of cells in your dataset and the representation of each cell type.
Computing Communication at the Pathway Level
After computing communication probabilities for individual ligand-receptor pairs, CellChat aggregates these results to the signaling pathway level. The computeCommunProbPathway function sums the communication probabilities for all ligand-receptor pairs that belong to the same signaling pathway.
cellchat <- computeCommunProbPathway(cellchat)
Pathway-level analysis provides a more interpretable view of cell-cell communication because signaling pathways group related ligand-receptor pairs into functional units. For example, the WNT signaling pathway includes multiple ligand-receptor pairs that collectively mediate WNT signaling between cells. Pathway-level aggregation reduces the dimensionality of the results and makes visualization and interpretation more tractable.
Aggregating Communication Networks
The final step in the core analysis is to aggregate the communication networks. The aggregateNet function creates a summarized network object that can be used for visualization and analysis.
cellchat <- aggregateNet(cellchat)
This step produces a network object that contains the aggregated communication probabilities between all pairs of cell groups. The aggregated network can be visualized as a circle plot, chord diagram, or heatmap. The network object also stores the number of interactions and the strength of interactions between each pair of cell groups.
Visualizing CellChat Results
Circle Plots and Chord Diagrams
CellChat provides several visualization functions for exploring communication networks. The netVisual_circle function creates a circle plot where cell groups are arranged around a circle and communication events are shown as edges between groups. The thickness of the edges represents the communication probability, and the color represents the sender cell group.
netVisual_circle(cellchat, weight.scale = TRUE)
The netVisual_chord function creates a chord diagram that shows the same information in a different layout. Chord diagrams are particularly useful for showing the direction of communication between cell groups. The width of the chord at each end represents the strength of communication from the sender to the receiver.
Bubble Plots for Ligand-Receptor Pairs
The netVisual_bubble function creates a bubble plot that shows the communication probability for specific ligand-receptor pairs between selected cell groups. Each bubble represents a ligand-receptor pair, with the size of the bubble representing the communication probability and the color representing the p-value from the permutation test.
netVisual_bubble(cellchat, sources.use = c(1, 2), targets.use = c(3, 4))
Bubble plots are useful for comparing the expression of specific ligand-receptor pairs across different cell group pairs. The sources.use and targets.use parameters allow you to focus on specific sender and receiver cell groups. This focused view is helpful when you have a hypothesis about which cell types are communicating through a particular signaling pathway.
Heatmaps for Signaling Pathways
The netVisual_heatmap function creates a heatmap that shows the communication probability for all pairs of cell groups for a specific signaling pathway. The rows represent sender cell groups and the columns represent receiver cell groups. The color intensity represents the communication probability.
netVisual_heatmap(cellchat, signaling = "WNT")
Heatmaps provide a comprehensive view of communication for a single pathway across all cell group pairs. This visualization is useful for identifying which cell groups are the major senders and receivers for a particular signaling pathway. The heatmap can reveal patterns such as autocrine signaling, where a cell group communicates with itself, or paracrine signaling between specific cell group pairs.
Signaling Role Analysis
CellChat includes functions to analyze the roles of different cell groups in signaling networks. The netAnalysis_signalingRole function computes the outgoing and incoming signaling strength for each cell group. This analysis identifies which cell groups are dominant senders, dominant receivers, or both for each signaling pathway.
cellchat <- netAnalysis_computeCentrality(cellchat)
netAnalysis_signalingRole_heatmap(cellchat, signaling = "WNT")
The centrality analysis computes network centrality measures for each cell group, including out-degree, in-degree, and betweenness. These measures quantify the importance of each cell group in the communication network. The signaling role heatmap visualizes the outgoing and incoming signaling strength for each cell group across all signaling pathways.
Comparative Analysis Across Conditions
Preparing Multiple CellChat Objects
CellChat supports comparative analysis of cell-cell communication across different biological conditions. This functionality is useful for identifying altered intercellular communication, signals, and cell populations between conditions (CellChat protocol, Nature Protocols 2025). The comparative analysis requires that you create a separate CellChat object for each condition and then merge them.
cellchat_control <- createCellChat(object = seurat_control, group.by = "cell_type")
cellchat_treated <- createCellChat(object = seurat_treated, group.by = "cell_type")
cellchat_control <- subsetData(cellchat_control)
cellchat_treated <- subsetData(cellchat_treated)
cellchat_control <- identifyOverExpressedGenes(cellchat_control)
cellchat_treated <- identifyOverExpressedGenes(cellchat_treated)
cellchat_control <- computeCommunProb(cellchat_control)
cellchat_treated <- computeCommunProb(cellchat_treated)
cellchat_control <- computeCommunProbPathway(cellchat_control)
cellchat_treated <- computeCommunProbPathway(cellchat_treated)
Each CellChat object must go through the same preprocessing and inference steps before merging. The cell group annotations must be consistent between conditions for the comparative analysis to be meaningful. If the cell types differ between conditions, you need to harmonize the annotations before proceeding.
Merging CellChat Objects
The mergeCellChat function combines two or more CellChat objects for comparative analysis. The function aligns the cell groups between conditions and creates a merged object that can be used for differential analysis.
object.list <- list(control = cellchat_control, treated = cellchat_treated)
cellchat_merged <- mergeCellChat(object.list, add.names = names(object.list))
The merged object stores the individual CellChat objects along with the alignment information. The add.names parameter assigns condition labels to each object. These labels are used in subsequent visualizations to distinguish between conditions.
Identifying Differential Communication
The identifyOverExpressedGenes and identifyOverExpressedInteractions functions can be applied to the merged object to identify communication events that differ between conditions. The netVisual_bubble function can then be used to compare ligand-receptor pair communication probabilities between conditions.
netVisual_bubble(cellchat_merged, sources.use = c(1, 2), targets.use = c(3, 4), comparison = c(1, 2))
The comparison parameter specifies which conditions to compare. The resulting bubble plot shows the communication probabilities for each condition side by side, making it easy to identify ligand-receptor pairs that are upregulated or downregulated between conditions.
Identifying Differential Signaling Pathways
The compareInteractions function provides a quantitative comparison of the total number and strength of interactions between conditions. The netVisual_heatmap function can be used to compare signaling pathway activity between conditions.
compareInteractions(cellchat_merged, measure = "count")
compareInteractions(cellchat_merged, measure = "weight")
The measure parameter specifies whether to compare the number of interactions or the total communication probability. The comparison results can identify signaling pathways that are globally upregulated or downregulated between conditions. These pathway-level differences can then be examined in detail using the visualization functions.
Interpreting CellChat Output
Understanding Communication Probabilities
The communication probability computed by CellChat represents the likelihood that a ligand-receptor interaction occurs between two cell groups. This probability is based on the expression levels of the ligand in the sender cell group and the receptor in the receiver cell group, adjusted for the multisubunit structure of the interaction and modulation by cofactors (CellChat protocol, Nature Protocols 2025).
The communication probability is not a direct measure of signaling activity. It is an inference based on gene expression data. High communication probability indicates that the molecular components for a signaling interaction are present in the respective cell groups, but it does not confirm that signaling actually occurs. Experimental validation is required to confirm functional signaling.
Interpreting Signaling Pathways
Pathway-level analysis aggregates ligand-receptor pairs into signaling pathways. The pathway-level communication probability represents the sum of communication probabilities for all ligand-receptor pairs in the pathway. This aggregation provides a higher-level view of communication but can obscure the contributions of individual ligand-receptor pairs.
When interpreting pathway-level results, consider which specific ligand-receptor pairs contribute most to the pathway communication probability. A pathway may show high overall communication probability due to a single highly expressed ligand-receptor pair, or it may reflect the coordinated expression of multiple pairs. The bubble plot visualization can help identify the specific pairs driving the pathway-level signal.
Identifying Dominant Signaling Populations
The signaling role analysis identifies cell groups that are dominant senders, dominant receivers, or both for each signaling pathway. Dominant sender cell groups express high levels of ligands for a pathway. Dominant receiver cell groups express high levels of receptors. Cell groups that are both dominant senders and receivers may be involved in autocrine signaling loops.
The identification of dominant signaling populations can guide experimental follow-up. For example, if a particular cell type is identified as the dominant sender for a signaling pathway of interest, you might target that cell type for perturbation experiments to test the functional role of the signaling pathway. Studies using CellChat have identified dominant signaling populations in contexts such as tumor microenvironments across multiple cancer types (Comparative oncology study, Cancers 2025) and in stem cell state regulation through BMP and NODAL signaling (BMP and NODAL signaling study, Frontiers in Cell and Developmental Biology 2025).
Common Failure Patterns and Troubleshooting
Mismatched Species Database
A common error is using the wrong species database for the data. If you have mouse data but use the human database, many ligand-receptor pairs will not be found because the gene names differ between species. Always verify that the species of your data matches the species of the CellChat database.
The gene naming conventions differ between human and mouse. Human genes use uppercase letters, while mouse genes use title case. If your Seurat object contains gene names in a different format, you may need to convert them before running CellChat. The conversion can be done using bioconductor packages that provide gene symbol mapping (Bioconductor documentation).
Inconsistent Cell Type Annotations
CellChat requires consistent cell type annotations across all samples in a comparative analysis. If the cell type labels differ between conditions, the merged analysis will not align properly. For example, if one condition labels a cell population as "T cells" and another labels the same population as "CD8 T cells", the comparative analysis will treat these as different cell groups.
Before running a comparative analysis, review the cell type annotations across all conditions. Harmonize the annotations so that equivalent cell types have the same labels. This harmonization may require re-clustering or manual annotation of some samples.
Sparse Cell Populations
Cell groups with very few cells can produce unreliable communication inferences. The permutation tests used by CellChat may not have enough power to detect significant communication from sparse cell populations. The filterCommunication function can remove interactions involving cell groups below a minimum cell number threshold.
If a cell type of interest has few cells, consider whether the cell type is truly rare or whether the clustering parameters need adjustment. Rare cell types may require higher resolution clustering or targeted re-clustering to be adequately represented.
Memory and Computation Time
CellChat can be computationally intensive for large datasets. The computeCommunProb function performs permutation tests that require repeated calculations. For datasets with many cell groups or many cells, the computation time can be substantial.
To manage computation time, consider the following approaches. Subset the data to the genes in the ligand-receptor database before running the inference. Reduce the number of permutation tests if the default number is too slow. Run the analysis on a machine with sufficient memory and processing power. The protocol typically takes about five minutes depending on dataset size (CellChat protocol, Nature Protocols 2025).
Quality Control and Reproducibility
Recording Analysis Parameters
Reproducibility requires that all analysis parameters be recorded. The key parameters for a CellChat analysis include the Seurat object version, the data slot used, the group-by variable, the species database, the averaging method, the number of permutation tests, and the minimum cell number threshold.
Record these parameters in a methods section or analysis notebook. The parameters should be sufficient for another researcher to reproduce the analysis from the same input data. The R session information, including package versions, should also be recorded. Reproducible workflow standards are emphasized in community resources such as nf-core documentation for pipeline configuration and usage (nf-core documentation) and Bioconductor package documentation (Bioconductor).
Validating Results with Independent Methods
CellChat results should be validated with independent methods before drawing biological conclusions. The communication probabilities are inferences based on gene expression and known ligand-receptor interactions. Experimental validation can include antibody staining for ligand and receptor proteins, proximity ligation assays to detect ligand-receptor binding, or functional perturbation experiments.
For spatial transcriptomics data, tools such as CytoSignal can validate the locations of ligand-receptor interactions at cellular resolution. The spatial validation can confirm that the inferred communication events occur between cells that are physically adjacent or in proximity (CytoSignal, Nature Genetics 2026).
Comparing with Other Communication Inference Tools
Multiple tools are available for cell-cell communication inference, and the results can vary between tools. Benchmarking studies have shown substantial variability in tool performance across different data types and contexts (Benchmarking CCI tools, Genome Biology 2026). Comparing CellChat results with results from other tools can identify robust communication events that are consistently inferred across methods.
Tools such as CausalCCC can extend the CellChat analysis by reconstructing intracellular causal pathways that connect ligand-receptor interactions to downstream gene expression changes (CausalCCC, Nucleic Acids Research 2025). This extended analysis provides a more comprehensive picture of cellular crosstalk that goes beyond ligand-receptor interaction scores.
Limitations of CellChat Analysis
Inference Based on Gene Expression
CellChat infers communication based on gene expression data. The presence of ligand and receptor transcripts does not guarantee that the corresponding proteins are produced, localized to the cell surface, or functionally active. Post-transcriptional regulation, protein degradation, and subcellular localization can all affect whether a ligand-receptor interaction actually occurs.
The communication probabilities should be interpreted as hypotheses about potential communication instead of confirmed signaling events. Experimental validation is essential before making strong biological claims based on CellChat results.
Database Completeness
The CellChat ligand-receptor database is curated from known molecular interactions. The database may not include all possible ligand-receptor pairs, particularly for less-studied signaling pathways or newly discovered interactions. Communication events mediated by ligand-receptor pairs not in the database will not be detected.
The database is updated in newer versions of CellChat. Using the latest version ensures that you have access to the most complete set of ligand-receptor pairs. The expanded database in CellChat v2 includes additional ligand-receptor pairs along with functional annotations (CellChat protocol, Nature Protocols 2025).
Cell Group Resolution
CellChat operates at the level of cell groups defined by the user. The resolution of the cell group annotations determines the resolution of the communication analysis. If cell groups are defined at a coarse level, the analysis may miss communication differences between subtypes within a group. If cell groups are defined at a fine level, the analysis may have insufficient cells per group for reliable inference.
The choice of cell group resolution depends on the biological question. For questions about major cell type communication, coarse annotations may be sufficient. For questions about subtype-specific communication, finer annotations are needed.
Applicability to Spatial Data
CellChat was designed for dissociated single-cell data. The tool does not use spatial information about cell locations. For spatial transcriptomics data, the communication inference may be improved by tools that incorporate spatial context. Benchmarking studies have shown that tool performance varies across spatial resolutions, tissue contexts, and platforms (Benchmarking CCI tools, Genome Biology 2026).
If you have spatial transcriptomics data, consider whether CellChat or a spatial-aware tool is more appropriate for your question. Tools such as CytoSignal can infer the locations and dynamics of ligand-receptor signaling at cellular resolution from spatial transcriptomic data (CytoSignal, Nature Genetics 2026).
Practical Workflow Summary
Step-by-Step Implementation
The following workflow summarizes the complete CellChat analysis pipeline.
Step 1. Prepare the Seurat object with normalized data and cell type annotations in the metadata.
Step 2. Create the CellChat object using createCellChat and set the species database.
Step 3. Subset the data to the genes in the ligand-receptor database using subsetData.
Step 4. Identify overexpressed genes and interactions using identifyOverExpressedGenes and identifyOverExpressedInteractions.
Step 5. Compute communication probabilities using computeCommunProb.
Step 6. Filter communication based on minimum cell number using filterCommunication.
Step 7. Compute pathway-level communication using computeCommunProbPathway.
Step 8. Aggregate the communication networks using aggregateNet.
Step 9. Visualize the results using the network visualization functions.
Step 10. For comparative analysis, create separate CellChat objects for each condition and merge them using mergeCellChat.
Records and Measurements
Maintain the following records for a reproducible CellChat analysis. The Seurat object version and the R session information. The exact parameters used for each CellChat function. The cell type annotations and the method used to derive them. The species database version. The output files from each visualization function.
These records enable another researcher to reproduce the analysis and verify the results. The records also support the interpretation of the results by documenting the decisions made during the analysis.
Professional Escalation Criteria
Consider escalating to a bioinformatics specialist or seeking additional training when the following situations arise. The Seurat object has unusual structure or missing metadata. The cell type annotations are uncertain or inconsistent across samples. The CellChat analysis produces unexpected errors or warnings. The results are biologically implausible and require expert interpretation. The dataset is very large and requires specialized computing resources.
Bioinformatics training resources are available from multiple sources. The European Bioinformatics Institute provides training for bioinformatics data-resource analysis (EMBL-EBI Training). The Galaxy Training Network offers accessible workflow training and analysis tutorials (Galaxy Training Network). The Carpentries provides foundational computing and programming lessons (The Carpentries Lessons). Bioconductor provides official package and workflow documentation for reproducible genomic analysis (Bioconductor).
Building a CellChat Analysis Decision Framework for Reproducible Research
Defining the Analysis Scope Before Running CellChat
The most common failure in CellChat projects is not a coding error but a missing decision about what question the analysis should answer. Before creating a CellChat object, define whether the goal is descriptive mapping of communication in one condition, differential comparison across conditions, or hypothesis testing of a specific ligand-receptor pair. This decision determines which functions you run, which visualizations you produce, and how you interpret the output.
For a descriptive analysis of a single dataset, the core workflow of computeCommunProb, computeCommunProbPathway, and aggregateNet is sufficient. For a comparative analysis, you need to decide whether you are comparing the number of interactions, the strength of interactions, or the identity of dominant signaling pathways between conditions. Each comparison uses different functions and produces different output tables. The CellChat protocol describes systematic and comparative analysis using quantitative metrics and machine-learning approaches (CellChat protocol, Nature Protocols 2025).
For hypothesis testing, define the specific sender cell group, receiver cell group, and signaling pathway before running the analysis. This focus prevents the common problem of multiple testing across hundreds of ligand-receptor pairs. If you test many pairs without correction, you will find significant results by chance. The bubble plot with sources.use and targets.use parameters is designed for this focused approach.
Establishing a Parameter Decision Log
CellChat has several parameters that materially affect results. The averaging method in computeCommunProb, the number of permutation tests, and the minimum cell number threshold in filterCommunication all change the output. A parameter decision log records these choices and the rationale for each one.
Create a table with columns for parameter name, value chosen, date, dataset identifier, and justification. The justification column is the most important. For example, if you choose type = "median" instead of the default "triMean", record why. If you set min.cells = 10, record whether this was based on the distribution of cells per cluster or a default from a tutorial. This log serves as the basis for the methods section of a paper and for troubleshooting when results change after a parameter adjustment.
The R session information, including package versions, should be saved with sessionInfo() and stored with the analysis output. Reproducible workflow standards from community resources emphasize recording configuration and usage details (nf-core documentation). Bioconductor package documentation also provides guidance on reproducible genomic analysis workflows (Bioconductor).
Selecting the Cell Group Resolution
The resolution of cell type annotations is the single largest determinant of CellChat output quality. Coarse annotations such as "T cells" aggregate subtypes with different communication profiles. Fine annotations such as "CD4 naive T cells" and "CD4 memory T cells" separate these profiles but reduce the cell count per group.
A practical approach is to run CellChat at two resolutions and compare the results. If the dominant signaling pathways are the same at both resolutions, the coarse annotation is adequate. If the pathways differ, the fine annotation is capturing biologically meaningful heterogeneity. This comparison takes additional computation time but prevents the common error of drawing conclusions from an annotation resolution that does not match the biological question.
The current best practices in single-cell RNA-seq analysis recommend careful attention to clustering and cell type annotation before downstream analysis (Current best practices in scRNA-seq analysis, Molecular Systems Biology 2019). CellChat inherits the quality of these upstream decisions.
Recording Cell Group Composition
For each cell group in the CellChat analysis, record the number of cells, the sample composition, and the marker genes used for annotation. This information is essential for interpreting communication probabilities. A cell group with 500 cells from one sample and 50 cells from another sample will have communication probabilities dominated by the larger sample.
The cell count per group also determines the reliability of the permutation tests in computeCommunProb. Groups with fewer than 10 cells produce unstable p-values. The filterCommunication function with min.cells = 10 removes these groups, but this threshold is arbitrary. Record the distribution of cell counts across groups before filtering and note which groups are removed.
For comparative analysis across conditions, record the cell count for each cell group in each condition. A condition with fewer total cells will have lower communication probabilities simply because the expression estimates are noisier. The compareInteractions function with measure = "count" compares the number of interactions, which is less sensitive to cell count than the measure = "weight" option that compares total communication probability.
Establishing a Validation Workflow
CellChat results require validation before biological interpretation. The validation workflow should be planned before running the analysis, not after. Decide which communication events are most important for the research question and plan the validation method for each one.
For ligand-receptor pairs, validation options include antibody staining for the ligand and receptor proteins, proximity ligation assays to detect physical binding, and functional perturbation experiments where the ligand or receptor is knocked down or blocked. The choice of validation method depends on the biological system and available reagents.
For spatial transcriptomics data, tools such as CytoSignal can validate the locations of ligand-receptor interactions at cellular resolution. CytoSignal infers the locations and dynamics of cell-cell communication and has been experimentally validated with proximity ligation assays (CytoSignal, Nature Genetics 2026). If spatial data are available for the same biological system, comparing CellChat results with spatial inference provides a cross-method validation.
Comparing CellChat with Complementary Tools
CellChat is one of several cell-cell communication inference tools, and the results can vary between tools. Benchmarking studies have shown substantial variability in tool performance across data types and contexts (Benchmarking CCI tools, Genome Biology 2026). Running a second tool on the same data and comparing the results identifies communication events that are robust across methods.
CausalCCC is a complementary tool that reconstructs intracellular causal pathways connecting ligand-receptor interactions to downstream gene expression changes. It accepts user-defined ligand-receptor pairs from methods such as CellChat and extends the analysis beyond interaction scores (CausalCCC, Nucleic Acids Research 2025). This extension is valuable when the research question concerns the downstream effects of communication instead of the communication events themselves.
The comparison with other tools should be recorded in the analysis notebook. Note which communication events are consistently inferred across tools and which are tool-specific. Tool-specific events may be artifacts of the modeling assumptions or may reflect genuine signals that only one tool can detect.
Documenting the Analysis Trail
The complete analysis trail includes the input Seurat object, the R script, the parameter decision log, the session information, and the output files. Store these together in a project directory with a clear naming convention. The directory structure should separate raw data, processed data, scripts, and results.
The R script should be written as a single file that runs from the Seurat object to the final visualizations. This script serves as the executable record of the analysis. If the analysis is run interactively in RStudio, the script should still be saved and updated as the analysis progresses. The visual and guided protocol for using RStudio for single-cell analysis emphasizes the importance of a step-by-step workflow with graphical results at each stage (R, Seurat, and CellChat workflow, Journal of Visualized Experiments 2025).
Professional Escalation Criteria
Escalate to a bioinformatics specialist when the analysis requires decisions beyond the standard workflow. Specific situations include harmonizing cell type annotations across datasets with different annotation schemes, integrating CellChat results with other omics data types, or adapting the analysis for non-model organisms where the standard ligand-receptor database may not apply.
Training resources are available for building these skills. The European Bioinformatics Institute provides training for bioinformatics data-resource analysis (EMBL-EBI Training). The Galaxy Training Network offers accessible workflow training and analysis tutorials (Galaxy Training Network). The Carpentries provides foundational computing and programming lessons (The Carpentries Lessons). Bioconductor provides official package and workflow documentation for reproducible genomic analysis (Bioconductor).
Common Failure Patterns in Decision Making
The most common decision failures in CellChat analysis are not technical but conceptual. The first is running the analysis without a defined question, which produces a large number of output tables and plots that are difficult to interpret. The second is changing parameters until the results match expectations, which invalidates the statistical inference. The third is interpreting communication probabilities as confirmed signaling events without experimental validation.
The decision framework described here prevents these failures by forcing explicit choices before the analysis runs. The parameter decision log documents the choices and the rationale. The validation workflow plans the experimental follow-up. The comparison with other tools provides cross-method confirmation. This framework converts CellChat from a black box into a documented, reproducible analysis that supports biological conclusions.
Frequently Asked Questions
What is the minimum number of cells required for a reliable CellChat analysis?
The minimum number of cells depends on the number of cell groups and the heterogeneity of gene expression within each group. CellChat includes a filterCommunication function that removes interactions involving cell groups below a minimum cell number threshold. A common threshold is 10 cells per group, but the appropriate threshold depends on your data. Cell groups with very few cells produce unreliable communication inferences because the permutation tests lack statistical power. If a cell type of interest has few cells, consider whether the clustering parameters need adjustment or whether the cell type is genuinely rare in your sample.
Can CellChat be used with single-nucleus RNA sequencing data?
CellChat can be applied to single-nucleus RNA sequencing data if the data are processed into a Seurat object with normalized counts and cell type annotations. Single-nucleus data typically have lower gene detection rates and different gene expression profiles compared to single-cell data, particularly for cytoplasmic genes. The ligand-receptor genes in the CellChat database may have lower detection rates in single-nucleus data. The analysis may still identify communication events for highly expressed ligand-receptor pairs, but the sensitivity may be reduced compared to single-cell data. Studies have applied CellChat to single-nucleus RNA sequencing data in contexts such as hepatoblastoma tumor analysis (BWS hepatoblastoma study, 2026) and neuronal-microglial communication research (Kif15 study, Journal of Biological Chemistry 2026).
How does CellChat compare with other cell-cell communication inference tools?
Multiple tools are available for cell-cell communication inference, and the results can vary between tools. Benchmarking studies have shown substantial variability in tool performance across different data types and contexts (Benchmarking CCI tools, Genome Biology 2026). CellChat uses a mass-action model that incorporates ligand-receptor multisubunit structure and cofactor modulation. Other tools use different modeling approaches and may produce different results. Comparing CellChat results with results from other tools can identify robust communication events that are consistently inferred across methods. Tools such as CausalCCC can extend the analysis by reconstructing intracellular causal pathways (CausalCCC, Nucleic Acids Research 2025).
What does the communication probability value mean?
The communication probability computed by CellChat represents the likelihood that a ligand-receptor interaction occurs between two cell groups. This probability is based on the expression levels of the ligand in the sender cell group and the receptor in the receiver cell group, adjusted for the multisubunit structure of the interaction and modulation by cofactors (CellChat protocol, Nature Protocols 2025). The probability is an inference based on gene expression data, not a direct measure of signaling activity. High communication probability indicates that the molecular components for a signaling interaction are present, but experimental validation is required to confirm functional signaling.
How should I choose the cell group annotations for CellChat?
The cell group annotations determine the resolution of the communication analysis. The annotations should reflect the biological question you are asking. For questions about major cell type communication, use coarse annotations such as T cells, B cells, macrophages, and epithelial cells. For questions about subtype-specific communication, use finer annotations such as CD4 T cells, CD8 T cells, and regulatory T cells. The annotations should be consistent across samples in a comparative analysis. The quality of the annotations directly affects the biological interpretability of the results.
What are the common errors when running CellChat?
Common errors include using the wrong species database, inconsistent cell type annotations across samples, sparse cell populations, and insufficient memory or computation time. The species database must match the species of your data. The cell type annotations must be consistent across conditions for comparative analysis. Cell groups with very few cells produce unreliable inferences. Large datasets may require substantial computation time and memory. Check the CellChat documentation and tutorials for guidance on troubleshooting specific errors.
Can CellChat be used for spatial transcriptomics data?
CellChat was designed for dissociated single-cell data and does not use spatial information about cell locations. For spatial transcriptomics data, the communication inference may be improved by tools that incorporate spatial context. Benchmarking studies have shown that tool performance varies across spatial resolutions, tissue contexts, and platforms (Benchmarking CCI tools, Genome Biology 2026). Tools such as CytoSignal can infer the locations and dynamics of ligand-receptor signaling at cellular resolution from spatial transcriptomic data (CytoSignal, Nature Genetics 2026). Consider whether CellChat or a spatial-aware tool is more appropriate for your spatial data.
How should I validate the CellChat results experimentally?
The communication probabilities inferred by CellChat should be validated with independent methods before drawing biological conclusions. Experimental validation can include antibody staining for ligand and receptor proteins, proximity ligation assays to detect ligand-receptor binding, or functional perturbation experiments. For spatial transcriptomics data, tools such as CytoSignal can validate the locations of ligand-receptor interactions at cellular resolution (CytoSignal, Nature Genetics 2026). The validation should confirm that the inferred communication events are biologically meaningful and not artifacts of the computational inference.
Related Bioinformatics Guides
- Single-Cell Sequencing Workflow: From Sample Preparation to Data Analysis
- Single-Cell Annotation: A Workflow for Cell Type Identification
- Single-Cell DNA Sequencing: Applications and Workflow Considerations
- Single-Cell Sequencing Analysis Pipeline: From Raw Data to Biological Insights
- Single-Cell RNA Sequencing Depth: A Cost-Benefit Analysis for Experimental Design
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- CellChat for systematic analysis of cell-cell communication from single-cell transcriptomics.. Nature protocols, 2025.
- CausalCCC: a web server to explore intracellular causal pathways enabling cell-cell communication.. Nucleic acids research, 2025.
- Delineation of complex gene expression patterns in single cell RNA-seq data with ICARUS v2.0.. NAR genomics and bioinformatics, 2023.
- Benchmarking tools for deciphering cellular crosstalk in spatially-resolved transcriptomics.. 2026.
- CytoSignal detects locations and dynamics of ligand-receptor signaling at cellular resolution from spatial transcriptomic data.. 2026.
- Beckwith-Wiedemann syndrome multiomic analysis of hepatoblastoma uncovers unique tumour heterogeneity and cellular landscapes, including transition cells leading to tumour formation.. 2026.
- Using R, Seurat, and CellChat to Analyze a Single-Cell Transcriptomics Dataset of Mouse Skin Wound Healing.. 2025.
- Comparative Oncology: Cross-Sectional Single-Cell Transcriptomic Profiling of the Tumor Microenvironment Across Seven Human Cancers.. 2025.
- BMP and NODAL paracrine signalling regulate the totipotent-like cell state in embryonic stem cells.. 2025.
- Kif15 orchestrates neuronal-microglial communication via CX3CL1 to impede nerve regeneration.. 2026.
- Analysis and Visualization of Single-Cell Sequencing Data with Scanpy and MetaCell: A Tutorial.. Methods in molecular biology, 2024.
- Current best practices in single-cell RNA-seq analysis: a tutorial. Molecular Systems Biology, 2019.
- Statistical Power Analysis for Designing Bulk, Single-Cell, and Spatial Transcriptomics Experiments: Review, Tutorial, and Perspectives. Biomolecules, 2023.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.