Choosing the Right Root Cell for Pseudotime Analysis: A Decision Guide for Accurate Trajectory Reconstruction

By Dr. Zubair Khalid, DVM, MS, PhD ·

Choosing the Right Root Cell for Pseudotime Analysis: A Decision Guide for Accurate Trajectory Reconstruction

Key Takeaways

  • The root cell is a critical parameter in pseudotime analysis, defining the starting point of inferred cellular trajectories; an incorrect root choice can invert biological directionality, misassign cell states, and lead to erroneous conclusions about cell fate decisions.
  • Marker-based root selection relies on prior biological knowledge of progenitor or stem cell markers, requiring validation of consistent expression in the candidate population and awareness that markers may be nonspecific or fail in aberrant biological states like disease.
  • Automated manifold geometry approaches identify root cells based on trajectory extremes or cell density, but are sensitive to outliers, loops, and disconnected components, necessitating validation against biological knowledge.
  • Robust root cell selection requires comprehensive quality control, including filtering cells by gene/transcript counts and mitochondrial read proportions, doublet detection, and addressing batch effects through data integration, as poor data quality can create spurious trajectories.
  • Sensitivity analysis, involving testing multiple alternative root cells and comparing pseudotime values and dynamically expressed genes, is essential to ensure conclusions are robust and not dependent on a single root choice.
  • Documentation of the biological question, candidate root identification methods, rationale for selection, sensitivity analysis results, and software versions is paramount for reproducibility and defensibility of trajectory inference findings.

Pseudotime analysis orders cells along a computed developmental or functional trajectory based on transcriptional similarity, and the root cell defines the starting point of that ordering. Selecting an inappropriate root cell can reverse the direction of the inferred trajectory, misassign progenitor and differentiated states, and lead to incorrect biological conclusions about cell fate decisions. This guide provides a practical framework for root cell selection in single-cell RNA sequencing (scRNA-seq) and single-nucleus RNA sequencing (snRNA-seq) analyses, covering marker-based approaches, automated tools, quality control considerations, and interpretation safeguards.

The Role of the Root Cell in Trajectory Inference

Trajectory inference methods such as Monocle 2, Monocle 3, diffusion pseudotime, and related algorithms construct a low-dimensional manifold from single-cell expression data and then order cells along a path through that manifold. The root cell anchors this path by defining which end of the trajectory represents the earliest state. Changing the root cell does not alter the underlying manifold structure, but it reverses the direction of pseudotime values assigned to every cell in the dataset.

This directional assignment carries substantial interpretive weight. When researchers ask which genes increase or decrease along a differentiation path, when they identify branch points that separate cell fates, or when they correlate pseudotime position with functional outcomes, every result depends on the root choice. A root placed in a mature cell population will make that population appear as the origin, and all other cells will appear to be progressing toward it, which inverts the biological narrative.

The practical consequence is that root cell selection must be treated as a hypothesis-driven decision, not a default parameter. Researchers should document the rationale for root choice before running trajectory inference, and they should test whether their conclusions survive alternative root placements. This sensitivity analysis is a standard expectation in peer review of trajectory-based studies.

At a Glance: Root Cell Selection Decision Framework

The following table summarizes the main approaches to root cell selection, their data requirements, and the situations where each approach is most appropriate.

ApproachData RequirementsBest Used WhenKey Limitations
Known marker genesPrior biological knowledge of progenitor or stem cell markersWell-characterized tissues with established marker panelsMarkers may be nonspecific or fail in aberrant disease states
Automated manifold geometryHigh-quality trajectory with clear extremesExploratory analyses without prior marker knowledgeSensitive to outliers, loops, and disconnected components
Gene regulatory network integrationSufficient cell numbers for network inferenceComplex differentiation programs with known regulatorsComputationally intensive and requires regulatory data
RNA velocitySpliced and unspliced transcript countsValidating root direction independentlyAssumptions about splicing kinetics may not hold

Known Marker Genes as Root Cell Anchors

The most direct approach to root cell selection uses prior biological knowledge to identify cells that express markers of the earliest state in the process under study. For a differentiation trajectory, this means finding cells that express progenitor or stem cell markers and do not yet express markers of mature progeny.

Selecting Progenitor Markers

Progenitor markers must be chosen with attention to the tissue and process being studied. A marker that identifies stem cells in one tissue may label a differentiated population in another. For example, in studies of dental-derived mesenchymal stem cells, researchers have used single-cell RNA sequencing to compare cultured stem cells from apical papilla and dental pulp, identifying subpopulations with distinct proliferative and progenitor-like characteristics based on marker coexpression patterns [<a href="#ref-1">1</a>]. The progenitor-like subset in that study coexpressed pericyte-associated markers such as NOTCH3 and PDGFRB alongside canonical mesenchymal stromal cell markers including MCAM, THY1, and DCN [<a href="#ref-1">1</a>]. This example illustrates that root cell identification often requires examining coexpression of multiple markers instead of relying on a single gene.

For epithelial differentiation trajectories, markers of basal or stem compartments are commonly used. In studies of wool fiber development in fine-wool sheep, researchers reconstructed the developmental dynamics of the epidermal lineage using pseudotime trajectory analysis, with hair follicle stem cells and outer root sheath cells representing early populations [<a href="#ref-2">2</a>]. The identification of these populations relied on known marker genes and cell type annotation before trajectory construction [<a href="#ref-2">2</a>].

Validating Marker Expression in the Root Population

After selecting candidate root cells based on marker expression, researchers should verify that the chosen population shows the expected properties. The root population should be transcriptionally distinct from the terminal populations, should express the progenitor markers at consistent levels across cells, and should occupy an appropriate position in the low-dimensional embedding.

A common validation step is to examine the expression of the progenitor markers across all clusters in the dataset. If the markers are expressed in multiple clusters, the root cell selection becomes ambiguous, and additional markers or a different approach may be needed. If the markers are expressed in a single cluster, that cluster becomes the candidate root population, and individual cells within it can be selected as root cells.

Limitations of Marker-Based Approaches

Marker-based root selection depends on the quality and completeness of prior biological knowledge. For well-studied systems such as hematopoiesis or epithelial development, reliable markers exist. For less characterized tissues, for disease states with aberrant differentiation programs, or for in vitro culture systems where cells may have drifted from their in vivo identities, marker-based approaches may fail.

In cultured stem cell systems, for instance, the use of nonspecific surface markers and technical limitations in isolation protocols can introduce considerable variability in differentiation potential [<a href="#ref-1">1</a>]. Single-cell RNA sequencing has revealed that cultured populations contain multiple subpopulations with distinct transcriptional programs, and the relationship between surface marker expression and functional state is not always straightforward [<a href="#ref-1">1</a>]. Researchers working with such systems should validate that their chosen root markers actually identify the earliest functional state instead of merely a population that expresses a convenient marker.

Automated Root Cell Selection Tools

Several computational tools attempt to identify root cells without requiring prior marker knowledge. These methods typically use properties of the trajectory manifold, such as the position of cells relative to branch points or the distribution of cells along the principal curve, to infer which cells are most likely to represent the earliest state.

Approaches Based on Manifold Geometry

One class of automated methods identifies root cells by finding cells at the extremes of the trajectory manifold. The assumption is that the earliest and latest states will be geometrically distant from the central mass of cells. These methods can be sensitive to outliers and to datasets where the trajectory forms a loop or contains multiple disconnected components.

Another approach uses the density of cells along the trajectory. Early progenitor populations often have fewer cells than the differentiated populations they give rise to, particularly in tissues where differentiation is accompanied by proliferation. Methods that identify low-density regions of the trajectory as candidate roots can be effective in such cases, but they fail when progenitor and differentiated populations have similar abundances.

Integration with Gene Regulatory Network Inference

More sophisticated approaches integrate trajectory inference with gene regulatory network analysis. The DELVE method, for example, uses a bottom-up strategy to identify molecular features that robustly recapitulate cellular trajectories, modeling cell states from dynamic gene or protein modules based on core regulatory complexes [<a href="#ref-3">3</a>]. This approach mitigates the effects of confounding sources of variation and selects features that better define cell types and cell-type transitions [<a href="#ref-3">3</a>].

In the context of root cell selection, gene regulatory network information can help identify the transcriptional programs that are active in the earliest state. If a regulatory network analysis identifies transcription factors that are predicted to drive the transition from progenitor to differentiated states, cells expressing those transcription factors at high levels and their target genes at low levels are candidate root cells.

Practical Considerations for Automated Tools

Automated root selection tools should be used with caution and always validated against biological knowledge. The tools make assumptions about trajectory geometry and cell density that may not hold for all datasets. Researchers should run automated methods, examine the proposed root cells, and compare them with marker-based expectations.

When automated and marker-based approaches disagree, the disagreement itself is informative. It may indicate that the trajectory structure is more complex than expected, that the markers are not specific to the earliest state, or that the automated method has been misled by technical artifacts. Documenting and resolving these disagreements is an important part of rigorous trajectory analysis.

Quality Control Before Root Cell Selection

Root cell selection is only meaningful if the underlying single-cell data are of sufficient quality. Poor quality data can create spurious trajectories, place cells in incorrect positions in the manifold, and lead to root cell misidentification. Quality control should be completed before trajectory inference begins.

Cell-Level Quality Metrics

Standard single-cell quality control includes filtering cells based on the number of detected genes, the total number of transcripts, and the proportion of mitochondrial reads. Cells with very low gene counts are likely to be empty droplets or damaged cells, while cells with very high mitochondrial proportions are likely to be stressed or dying. These cells should be removed before trajectory analysis.

The thresholds for these filters depend on the tissue, the protocol, and the expected biology. A rigorous approach is to examine the distributions of these metrics and set thresholds based on clear discontinuities or inflection points instead of arbitrary values. The Galaxy Training Network provides accessible tutorials on quality control and filtering for single-cell data that can serve as a reference for establishing these thresholds [<a href="#ref-4">4</a>].

Doublet Detection

Doublets, which are droplets or wells containing two or more cells, can create artificial cell states that distort trajectory inference. A doublet containing one progenitor cell and one differentiated cell will have a mixed transcriptional profile that places it in an intermediate position along the trajectory, potentially creating a false branch or altering the apparent path between states.

Doublet detection methods should be applied before trajectory inference, and predicted doublets should be removed. The proportion of doublets increases with the number of cells loaded, so loading fewer cells per reaction can reduce doublet rates at the cost of lower throughput.

Batch Effects and Integration

When single-cell data are generated across multiple samples, batches, or sequencing runs, technical variation can confound biological variation. Batch effects can create artificial separation between cells from different batches, distort the trajectory manifold, and affect root cell identification.

Data integration methods can align cells across batches while preserving biological variation. The choice of integration method and the assessment of integration quality are important steps in the workflow. The Bioconductor project provides documentation and workflows for single-cell analysis, including integration approaches, that can guide these decisions [<a href="#ref-5">5</a>].

For single-nucleus RNA sequencing data, additional considerations apply. Nuclei from different cell types may have different RNA content, and the proportion of intronic reads varies across cell types. Quality control thresholds may need to be adjusted accordingly, and the interpretation of trajectory results should account for the fact that nuclear RNA captures a different snapshot of gene expression than whole-cell RNA.

Data Integration and Its Impact on Root Cell Selection

The decision to integrate single-cell datasets before trajectory analysis has direct consequences for root cell selection. Integration can remove technical variation, but it can also remove biological variation if the integration is too aggressive.

When Integration Is Necessary

Integration is necessary when the datasets to be analyzed together were generated in different batches, by different protocols, or at different times. Without integration, cells from the same biological state but different batches may appear as distinct clusters, and the trajectory may be fragmented or distorted.

In studies of neuropathic pain, for example, researchers have integrated multiple gene expression datasets from the Gene Expression Omnibus to identify energy metabolism-related differentially expressed genes and to perform single-cell analyses [<a href="#ref-6">6</a>]. The integration of datasets from different sources requires careful attention to batch effects and to the comparability of the biological conditions being studied [<a href="#ref-6">6</a>].

Integration Methods and Their Assumptions

Different integration methods make different assumptions about the nature of the variation between datasets. Some methods assume that the biological variation is shared across datasets and that technical variation is dataset-specific. Others allow for dataset-specific biological variation, such as when one dataset contains a cell type that another does not.

The choice of integration method can affect which cells are placed near each other in the integrated space, which in turn affects trajectory inference and root cell selection. Researchers should compare the results of different integration methods and examine whether the root cell population remains consistent across methods.

Integration Validation

Integration quality should be assessed before proceeding to trajectory analysis. Useful checks include examining whether cells from the same biological state but different batches are mixed in the integrated space, whether known marker genes show consistent expression patterns across batches, and whether the overall structure of the data is preserved.

The nf-core documentation provides guidance on reproducible analysis pipelines, including quality control and integration steps, that can help ensure consistency across analyses [<a href="#ref-7">7</a>]. Reproducibility is particularly important when integration decisions affect downstream trajectory results.

Trajectory Inference Methods and Root Cell Handling

Different trajectory inference methods handle root cells differently, and the choice of method affects how root cell selection is implemented.

Monocle 2 and Monocle 3

Monocle 2 uses reversed graph embedding to learn a principal graph that describes the trajectory. The root cell is specified by the user, and pseudotime is calculated as the distance along the principal graph from the root. Monocle 2 has been used in studies of primary sclerosing cholangitis-associated cholangiocarcinogenesis to reconstruct candidate malignant-associated trajectories [<a href="#ref-8">8</a>].

Monocle 3 uses a different algorithm based on UMAP and partition-based graph abstraction. The root cell is selected interactively, and the user can specify the root node in the trajectory graph. Monocle 3 also provides functions for finding the root of a trajectory automatically, but the user should validate the automatic choice.

Diffusion Pseudotime

Diffusion pseudotime is based on diffusion maps, which embed cells in a low-dimensional space that captures the connectivity of the data. The root cell is specified by the user, and pseudotime is calculated as the diffusion distance from the root. Diffusion pseudotime has been used in integrative pipelines that combine latent-space integration, unsupervised clustering, and trajectory analysis [<a href="#ref-9">9</a>].

Diffusion pseudotime is particularly sensitive to the choice of root cell because diffusion distances depend on the global structure of the data. A root cell in a peripheral location will produce different pseudotime values than a root cell in a central location, even if both are in the same cell type.

Slingshot and Other Methods

Slingshot uses cluster-based lineage inference and then fits smooth curves to each lineage. The root cluster is specified by the user, and pseudotime is calculated along each curve. Slingshot is useful when the trajectory has multiple branches, but the root cluster choice affects the direction of pseudotime along each branch.

Other methods, such as those based on optimal transport or on RNA velocity, may not require an explicit root cell. RNA velocity uses the ratio of spliced to unspliced transcripts to infer the direction of transcriptional change, which can provide an independent estimate of the direction of the trajectory. However, RNA velocity has its own assumptions and limitations, and it should be used in conjunction with explicit root cell selection instead of as a replacement for it.

Sensitivity Analysis and Validation of Root Cell Choice

The most important step in root cell selection is testing whether the conclusions of the analysis depend on the specific root cell chosen. This sensitivity analysis should be performed for every trajectory-based study.

Alternative Root Cell Testing

Researchers should run the trajectory inference with several alternative root cells, including cells from different clusters, cells at different positions in the manifold, and cells identified by different methods. The pseudotime values and the genes identified as dynamically expressed should be compared across these runs.

If the conclusions are robust to root cell choice, the analysis is strengthened. If the conclusions change, the researcher must determine which root cell is most biologically justified and report the sensitivity of the results to this choice.

Consistency with Independent Evidence

Trajectory results should be checked against independent evidence about the direction of differentiation. This evidence can come from RNA velocity, from the expression of known markers at different stages, from proliferation assays, or from functional experiments. In studies of breast cancer, for example, researchers have integrated trajectory inference with interpretable machine learning to identify key genes and expression thresholds associated with transcriptional reprogramming during tumor evolution [<a href="#ref-10">10</a>]. The combination of trajectory inference with independent analytical approaches provides a more robust basis for interpreting cell state transitions [<a href="#ref-10">10</a>].

Reporting Sensitivity Results

The results of sensitivity analyses should be reported in the methods and results sections of papers. This reporting should include the alternative root cells tested, the criteria used to evaluate the stability of the results, and the conclusions that were robust or sensitive to root choice. Transparent reporting of root cell selection and sensitivity analysis is essential for reproducibility.

Common Failure Patterns in Root Cell Selection

Several recurring problems can undermine root cell selection and trajectory inference. Recognizing these patterns helps researchers avoid them and interpret results correctly.

Circular Trajectories

Some biological processes are cyclical, such as the cell cycle or the hair follicle growth cycle. Trajectory inference methods that assume a linear or branching structure will force a circular process into an open trajectory, and the choice of root cell will determine where the cycle is broken. This break point is arbitrary and can create artificial start and end states.

In such cases, researchers should use methods designed for cyclical trajectories or should interpret the trajectory with caution. The root cell should be chosen based on biological knowledge of the natural starting point of the cycle, and the results should be validated against independent markers of cycle position.

Disconnected Trajectory Components

If the data contain multiple disconnected components, trajectory inference methods may fail or produce misleading results. The root cell may be placed in one component, and cells in other components may be assigned pseudotime values that do not reflect their true relationship to the root.

Researchers should examine the connectivity of the data before running trajectory inference. If the data are disconnected, the analysis should be restricted to the connected component of interest, or the disconnected components should be analyzed separately.

Root Cell in a Rare Population

If the true progenitor population is rare, the root cell may be difficult to identify. The population may be represented by only a few cells, and these cells may be removed during quality control or may be placed in an ambiguous position in the manifold.

In such cases, researchers should consider whether the progenitor population is adequately captured in the data. If the population is too rare for reliable analysis, additional sequencing may be needed, or the analysis may need to be restricted to the populations that are well represented.

Technical Artifacts Mimicking Trajectories

Technical artifacts can create apparent trajectories that do not reflect biological processes. For example, cells with different levels of ambient RNA contamination may form a gradient that looks like a differentiation trajectory. Cells with different levels of mitochondrial reads may also form a gradient that can be mistaken for a biological process.

Quality control should remove the most severely affected cells, but subtle gradients may persist. Researchers should examine whether the genes driving the trajectory are biologically plausible and whether the trajectory is consistent with known biology.

Records and Documentation for Root Cell Selection

Rigorous documentation of root cell selection decisions is essential for reproducibility and for defending the analysis in peer review.

Documentation Requirements

The documentation should include the following elements:

  • The biological question and the expected direction of the trajectory
  • The method used to identify candidate root cells, including marker genes or automated tools
  • The specific root cell or cells selected and the rationale for the selection
  • The results of sensitivity analyses with alternative root cells
  • The version numbers of all software packages used
  • The parameters used for quality control, integration, and trajectory inference

Reproducibility Tools

Workflow management systems can help ensure that the analysis is reproducible. The nf-core framework provides standards for community pipelines, including configuration and usage documentation, that can be applied to single-cell analysis workflows [<a href="#ref-7">7</a>]. The Galaxy Training Network offers tutorials on reproducible analysis that cover the use of workflow tools and the documentation of analysis steps [<a href="#ref-4">4</a>].

The Carpentries lessons provide foundational training in computing, data management, and version control that is directly applicable to documenting single-cell analyses [<a href="#ref-11">11</a>]. Version control of analysis scripts and careful record keeping of parameter choices are essential for reproducibility.

Sharing Analysis Code

Sharing the analysis code, including the root cell selection steps, allows other researchers to reproduce the analysis and to test alternative root cell choices. Code should be well commented, and the key decisions should be explained in the code or in accompanying documentation.

Interpretation Limits and Reporting Standards

Trajectory inference and pseudotime analysis have inherent limitations that should be acknowledged in reporting.

Pseudotime Is Not Real Time

Pseudotime is a computational ordering based on transcriptional similarity, not a measurement of actual time. Cells that are assigned similar pseudotime values may not be at the same stage of differentiation, and the distance between cells in pseudotime does not correspond to a fixed amount of biological time.

This limitation is particularly important when interpreting the results of trajectory analysis in clinical or translational contexts. A pseudotime trajectory that places a disease-associated cell state at a particular position does not prove that the disease state arises from the preceding states in the trajectory.

Trajectory Inference Is Hypothesis Generating

Trajectory inference generates hypotheses about cell state transitions that should be validated with independent experiments. The genes identified as dynamically expressed along the trajectory are candidates for functional testing, not confirmed regulators of the transition.

In studies of prostate cancer, for example, researchers used single-cell RNA sequencing to analyze cell-type-specific expression patterns of genes identified through network toxicology and machine learning approaches [<a href="#ref-12">12</a>]. The single-cell analysis provided spatial and cell-type context for the candidate genes, but the functional roles of these genes required additional validation [<a href="#ref-12">12</a>].

Reporting Uncertainty

The uncertainty in trajectory inference should be reported. This includes the sensitivity of the results to root cell choice, to quality control thresholds, and to the parameters of the trajectory inference method. Reporting this uncertainty helps readers interpret the results appropriately and helps other researchers design validation experiments.

Professional Escalation Criteria

Some situations require escalation to more specialized expertise or to additional data generation.

When to Seek Specialized Help

Researchers should consider seeking help from bioinformatics specialists or collaborators when:

  • The trajectory structure is complex, with multiple branches, loops, or disconnected components
  • Automated root cell selection methods give inconsistent results
  • The root cell population cannot be identified with available markers
  • The sensitivity analysis shows that conclusions depend strongly on root cell choice
  • The data contain severe batch effects or technical artifacts that are difficult to resolve

Bioinformatics training resources, such as those provided by EMBL-EBI Training, can help researchers build the skills needed to address these challenges [<a href="#ref-13">13</a>]. The NCBI provides access to databases and analysis services that can support trajectory analysis and related bioinformatics tasks [<a href="#ref-14">14</a>].

When Additional Data Are Needed

Sometimes the existing data are insufficient for reliable trajectory inference. This situation arises when:

  • The progenitor population is too rare to be reliably identified
  • The trajectory spans cell states that are not captured in the data
  • The quality of the data is too low for reliable analysis
  • The biological process of interest is not represented in the sampled cells

In these cases, additional single-cell or single-nucleus sequencing may be needed. The design of the additional experiments should be informed by the gaps identified in the initial analysis.

When to Reconsider the Biological Question

If the trajectory analysis consistently fails to produce interpretable results, the biological question itself may need to be reconsidered. The process under study may not be well described by a trajectory model, or the cell states may not be related in the way assumed by the analysis.

In such cases, alternative analytical approaches should be considered. These may include clustering-based analyses that do not assume a trajectory structure, or analyses that focus on specific cell populations instead of on the relationships between populations.

Practical Workflow for Root Cell Selection

The following step-by-step workflow integrates the concepts discussed above into a practical procedure for root cell selection.

Step 1: Define the Biological Question

Before any computational analysis, write down the biological question and the expected direction of the trajectory. This statement should specify the presumed earliest state and the presumed terminal states. This documentation will guide all subsequent decisions and provide the basis for sensitivity analysis.

Step 2: Complete Quality Control

Filter cells based on gene counts, transcript counts, and mitochondrial read proportions. Apply doublet detection and remove predicted doublets. If multiple batches are present, integrate the data and validate the integration quality. Document all thresholds and parameters used.

Step 3: Annotate Cell Types

Cluster the cells and annotate the clusters using known marker genes. This annotation provides the biological context needed for root cell selection. The annotation should be based on multiple markers per cell type where possible, and the expression of markers should be examined across clusters to confirm specificity.

Step 4: Identify Candidate Root Cells

Apply at least two independent methods for root cell identification. A marker-based approach should be used when reliable markers exist. An automated approach should be used as a second method. Compare the results and document any disagreements.

Step 5: Run Trajectory Inference

Run the trajectory inference method with the candidate root cell. Record the pseudotime values and the genes identified as dynamically expressed along the trajectory.

Step 6: Perform Sensitivity Analysis

Repeat the trajectory inference with at least two alternative root cells. Compare the pseudotime values and the dynamically expressed genes across runs. Document which conclusions are robust and which depend on the root cell choice.

Step 7: Validate with Independent Evidence

Check the trajectory direction against independent evidence such as RNA velocity, known marker expression patterns, or functional data. If the independent evidence contradicts the trajectory direction, revisit the root cell selection.

Step 8: Report and Archive

Report the root cell selection rationale, the sensitivity analysis results, and all parameters in the methods section. Archive the analysis code and the documentation in a version-controlled repository.

Common Failure Patterns and Their Resolution

The following table summarizes common failure patterns in root cell selection and the recommended resolution for each pattern.

Failure PatternSymptomResolution
Circular trajectoryTrajectory forms a loop with no clear startUse cyclical trajectory methods or choose root based on biological knowledge of cycle start
Disconnected componentsCells form separate groups with no connecting pathRestrict analysis to the connected component of interest
Rare progenitor populationFew cells express progenitor markersIncrease sequencing depth or enrich for progenitor cells
Technical artifact gradientTrajectory driven by mitochondrial or ambient RNA contentTighten quality control filters and examine driver genes for biological plausibility
Marker nonspecificityProgenitor markers expressed in multiple clustersUse coexpression of multiple markers or switch to automated methods

A Practical Decision Framework for Root Cell Selection

Root cell selection is often treated as a single analytical step, but in practice it requires a structured decision process that accounts for the biological system, the available data, and the specific trajectory inference method. The following framework provides a systematic approach that researchers can apply before running trajectory inference, with clear criteria for evaluating candidate root cells and documenting the rationale for the final choice.

Tier 1: Establish the Biological Reference Frame

Before examining any computational output, define the expected direction of the biological process in explicit terms. Write a single sentence that states the presumed earliest cell state and the presumed terminal state. For example, in a study of maize root development, the expected trajectory would run from meristematic cells through the maturation zone to mature cortex cells [<a href="#ref-15">15</a>]. In a study of wool fiber development, the expected trajectory would run from hair follicle stem cells through outer root sheath cells to differentiated dermal papilla cells [<a href="#ref-2">2</a>].

This reference frame serves three purposes. First, it forces the researcher to articulate the biological hypothesis that the trajectory analysis is meant to test. Second, it provides the standard against which all candidate root cells will be evaluated. Third, it creates a written record that can be compared against the final results to check whether the analysis confirmed or contradicted the initial expectation.

The reference frame should also specify the expected relationship between cell states. Is the process linear, with a single path from start to end? Is it branching, with one progenitor population giving rise to multiple terminal states? Is it cyclical, as in the cell cycle or hair follicle growth cycle? The expected structure determines which trajectory inference methods are appropriate and how the root cell choice will be interpreted.

Tier 2: Score Candidate Root Cells Against Independent Criteria

Once the biological reference frame is established, identify candidate root cells using at least two independent methods. For each candidate, score the following criteria on a simple pass or fail basis.

Marker specificity. Does the candidate population express the progenitor markers at consistent levels across cells? Are the markers absent from the presumed terminal populations? In the dental stem cell study, the progenitor-like subset was identified by coexpression of pericyte-associated markers NOTCH3 and PDGFRB alongside canonical mesenchymal stromal cell markers MCAM, THY1, and DCN [<a href="#ref-1">1</a>]. A candidate root population that expresses only one of these markers would fail this criterion.

Transcriptional distinctness. Is the candidate population transcriptionally distinct from the terminal populations in the low-dimensional embedding? A root population that overlaps extensively with differentiated populations will produce unstable pseudotime ordering. The wool fiber study identified hair follicle stem cells and outer root sheath cells as distinct populations based on known marker genes before trajectory construction [<a href="#ref-2">2</a>], which provided a clear transcriptional boundary for root selection.

Geometric position. Does the candidate population occupy an appropriate position in the trajectory manifold? For most trajectory inference methods, the root should be at one extreme of the manifold, not in the center. A candidate root in the center of the manifold will produce pseudotime values that radiate outward in multiple directions, which is difficult to interpret.

Consistency with independent evidence. Does the candidate root align with RNA velocity estimates, known developmental timing, or functional assays? In the breast cancer study, researchers integrated trajectory inference with interpretable machine learning to identify genes and expression thresholds associated with transcriptional reprogramming [<a href="#ref-10">10</a>]. This independent analytical layer provided a check on whether the trajectory direction was biologically plausible.

A candidate root cell that passes all four criteria is a strong candidate. A candidate that fails one or more criteria should be rejected or investigated further before being used.

Tier 3: Run a Root Cell Comparison Matrix

Instead of running trajectory inference once with a single root cell, run it with multiple candidate root cells and compare the results systematically. The comparison matrix should include at least three candidates: the primary candidate identified by the scoring process, an alternative candidate from a different cell type or cluster, and a candidate from a different region of the manifold.

For each run, record the following outputs:

  • The pseudotime values assigned to all cells
  • The list of genes identified as dynamically expressed along the trajectory
  • The position of branch points, if the method identifies them
  • The correlation between pseudotime values across runs

The comparison matrix serves two purposes. First, it quantifies the sensitivity of the results to root cell choice. If the pseudotime values are highly correlated across runs and the dynamically expressed genes are largely overlapping, the conclusions are robust. Second, it identifies which aspects of the analysis are stable and which depend on the root choice. This information is essential for reporting and for interpreting the biological significance of the results.

Tier 4: Apply the Decision Rules

The following decision rules guide the interpretation of the comparison matrix.

Rule 1: Convergent results. If the primary candidate and the alternative candidates produce similar pseudotime values and similar dynamically expressed genes, use the primary candidate and report the sensitivity analysis as evidence of robustness.

Rule 2: Divergent results with clear biological justification. If the alternative candidates produce different results but the primary candidate has stronger biological justification based on the Tier 2 scoring, use the primary candidate and report the divergence. The sensitivity analysis should be described in the methods section so that readers can assess the impact of the root choice.

Rule 3: Divergent results without clear biological justification. If the candidates produce different results and none has a clear biological advantage, the trajectory analysis is not reliable for the biological question being asked. Consider whether the data are sufficient, whether the trajectory model is appropriate, or whether additional markers or data are needed.

Rule 4: Disagreement between marker-based and automated methods. If a marker-based approach identifies one root population and an automated approach identifies a different population, investigate the source of the disagreement. The automated method may be sensitive to outliers or to the density of cells along the trajectory. The marker-based approach may be using markers that are not specific to the earliest state. The resolution of this disagreement should be documented.

Records and Documentation for the Decision Framework

The decision framework produces a set of records that should be archived with the analysis code. The minimum documentation includes the biological reference frame statement, the Tier 2 scoring table for each candidate root cell, the Tier 3 comparison matrix outputs, and the decision rule applied.

The nf-core documentation provides standards for reproducible workflow configuration and usage that can be applied to trajectory analysis pipelines [<a href="#ref-7">7</a>]. The Galaxy Training Network offers tutorials on reproducible analysis that cover the documentation of analysis steps and parameter choices [<a href="#ref-4">4</a>]. The Carpentries lessons provide foundational training in version control and data management that supports the archival of analysis records [<a href="#ref-11">11</a>].

The documentation should be written so that another researcher can understand the root cell selection rationale without access to the original analyst. This includes the version numbers of all software packages, the parameters used for quality control and integration, and the specific commands used to run the trajectory inference.

Common Failure Patterns in the Decision Framework

The decision framework can fail in several recognizable ways.

Failure pattern 1: The reference frame is wrong. If the biological reference frame is incorrect, all subsequent steps will be misdirected. For example, if the researcher assumes a linear trajectory when the process is actually branching, the root cell selection will be evaluated against the wrong expected structure. The reference frame should be revisited if the trajectory results consistently contradict the initial expectation.

Failure pattern 2: The scoring criteria are too lenient. If the Tier 2 criteria are applied loosely, multiple candidate populations will pass, and the comparison matrix will show divergent results. The criteria should be applied strictly, and candidates that fail any criterion should be rejected.

Failure pattern 3: The comparison matrix is too narrow. If only two candidates are tested, the sensitivity analysis may miss important dependencies on root choice. At least three candidates should be tested, including candidates from different clusters and different regions of the manifold.

Failure pattern 4: The decision rules are ignored. If the researcher selects a root cell based on convenience or prior expectation without applying the decision rules, the sensitivity analysis is meaningless. The decision rules should be applied systematically and the results reported.

Integration with Existing Analytical Workflows

The decision framework is designed to integrate with existing single-cell analysis workflows. It does not require new software or specialized tools. The framework uses the outputs of standard quality control, clustering, and trajectory inference methods and adds a structured decision process around root cell selection.

The framework is compatible with the workflow described in the Bioconductor documentation for single-cell analysis [<a href="#ref-5">5</a>]. It can be implemented in R or Python using existing trajectory inference packages. The framework is also compatible with the pipeline standards described in the nf-core documentation [<a href="#ref-7">7</a>], which can be used to automate the comparison matrix runs.

The DELVE method for feature selection can be incorporated into the Tier 2 scoring process. DELVE identifies molecular features that robustly recapitulate cellular trajectories by modeling cell states from dynamic gene or protein modules based on core regulatory complexes [<a href="#ref-3">3</a>]. The features selected by DELVE can be used to evaluate whether the candidate root population expresses the regulatory programs expected of the earliest state [<a href="#ref-3">3</a>].

When to Escalate to Specialized Expertise

The decision framework is designed to be applied by researchers with standard bioinformatics training. However, some situations require escalation to specialized expertise.

Escalate when the comparison matrix shows divergent results across multiple candidate root cells and the source of the divergence cannot be identified. This situation may indicate that the trajectory structure is more complex than the inference method can handle, or that the data contain technical artifacts that are not removed by standard quality control.

Escalate when the biological reference frame cannot be established with confidence. This situation arises when the tissue or process under study is poorly characterized, when the markers for the earliest state are unknown, or when the disease state has disrupted the normal differentiation program.

Escalate when the trajectory results are intended to guide clinical decisions or experimental interventions. In these cases, the trajectory analysis should be reviewed by multiple experts, and the sensitivity of the results to root cell choice should be thoroughly documented.

The EMBL-EBI Training program provides learning pathways for bioinformatics analysis that can help researchers build the skills needed to apply the decision framework and to recognize when escalation is appropriate [<a href="#ref-13">13</a>]. The NCBI provides access to databases and analysis services that can support the validation of marker genes and the interpretation of trajectory results [<a href="#ref-14">14</a>].

Frequently Asked Questions

What is the root cell in pseudotime analysis?

The root cell is the cell that defines the starting point of the pseudotime trajectory. All other cells are ordered relative to the root cell, with pseudotime values increasing as cells move away from the root along the inferred trajectory. The root cell is typically chosen to represent the earliest or most progenitor-like state in the process being studied.

How do I choose the root cell using marker genes?

Identify cells that express markers of the earliest state in the process under study. These markers should be specific to the progenitor or stem cell population and should not be expressed in the differentiated populations. Validate the marker expression by examining the distribution of marker-positive cells in the low-dimensional embedding and by confirming that the marker-positive population is transcriptionally distinct from other populations.

What are the automated tools for root cell selection?

Automated tools use properties of the trajectory manifold, such as the position of cells relative to branch points or the density of cells along the trajectory, to identify candidate root cells. Some methods integrate gene regulatory network information to identify the transcriptional programs active in the earliest state. These tools should be validated against biological knowledge and marker-based approaches.

How does the choice of root cell affect the trajectory results?

The choice of root cell determines the direction of pseudotime along the trajectory. Changing the root cell reverses the direction of pseudotime, which changes which genes appear to be upregulated or downregulated along the trajectory and which cell states appear to be the origin of the differentiation process. The underlying manifold structure is not changed by the root cell choice.

What is sensitivity analysis for root cell selection?

Sensitivity analysis involves running the trajectory inference with multiple alternative root cells and comparing the results. If the conclusions are robust to root cell choice, the analysis is strengthened. If the conclusions change, the researcher must determine which root cell is most biologically justified and report the sensitivity of the results.

Can I use RNA velocity to determine the root cell?

RNA velocity uses the ratio of spliced to unspliced transcripts to infer the direction of transcriptional change. This can provide an independent estimate of the direction of the trajectory and can help validate root cell choice. However, RNA velocity has its own assumptions and limitations, and it should be used in conjunction with explicit root cell selection instead of as a replacement for it.

What should I do if my data contain multiple batches?

Integrate the data across batches before trajectory inference. The choice of integration method and the assessment of integration quality are important steps. After integration, examine whether the root cell population remains consistent across batches and whether the trajectory structure is preserved.

How should I report root cell selection in my paper?

Report the biological rationale for the root cell choice, the method used to identify candidate root cells, the specific root cell or cells selected, and the results of sensitivity analyses with alternative root cells. Include the version numbers of all software packages and the parameters used for quality control, integration, and trajectory inference.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

[1] [In Vitro Single-Cell Transcriptomic Profiling of Cultured Stem Cells From Apical Papilla and Dental Pulp Stem Cells: Unveiling Cellular Heterogeneity.](https://doi.org/10.1016/j.identj.2026.109577). 2026. [2] [Deciphering the cellular landscape and genetic underpinnings of fiber diameter determined by dermal papilla cells in fine-wool sheep.](https://doi.org/10.3389/fcell.2026.1764812). 2026. [3] [DELVE: feature selection for preserving biological trajectories in single-cell data](https://doi.org/10.1038/s41467-024-46773-z). Nature Communications, 2024. [4] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project. [5] [Bioconductor](https://bioconductor.org/). Bioconductor Project. [6] [The Integrated Transcriptome Bioinformatics Analysis of Energy Metabolism-Related Profiles for Dorsal Root Ganglion of Neuropathic Pain.](https://pubmed.ncbi.nlm.nih.gov/39406937). Molecular neurobiology, 2025. [7] [nf-core Documentation](https://nf-co.re/docs). nf-core. [8] [A Dual-Gene Signature of PMAIP1 and GADD45A for Early Detection of Intrahepatic Cholangiocarcinoma in the Context of Primary Sclerosing Cholangitis.](https://doi.org/10.3390/ijms27114826). 2026. [9] [An AI-Enabled Single-Cell Transcriptomic Analysis Pipeline for Gene Signature Discovery in Natural Killer Cells Linked to Remission Outcomes in Chronic Myeloid Leukemia.](https://doi.org/10.3390/biology15070588). 2026. [10] [Integrating trajectory inference and self-explainable predictive models to explore cell state transitions in breast cancer at single-cell resolution.](https://doi.org/10.3389/fbinf.2026.1672671). 2026. [11] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries. [12] [Integrating network toxicology, machine learning, and single-cell sequencing to reveal the FASN-mediated role of phenolic endocrine disruptors in water in promoting prostate cancer.](https://doi.org/10.1371/journal.pone.0350638). 2026. [13] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute. [14] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information. [15] [Single-cell RNA sequencing reveals cellular diversity and gene expression dynamics in maize root development](https://doi.org/10.3389/fpls.2025.1666531). Frontiers in Plant Science, 2025.

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.