Transcription Factor Activity Inference from Single-Cell Data: From SCENIC to Decoupling

By Dr. Zubair Khalid, DVM, MS, PhD ·

Transcription Factor Activity Inference from Single-Cell Data: From SCENIC to Decoupling

Key Takeaways

  • TF activity inference from single-cell RNA sequencing (scRNA-seq) is crucial because mRNA expression levels do not reliably predict functional TF activity due to post-translational modifications, nuclear localization, and chromatin accessibility.
  • SCENIC (and its Python implementation pySCENIC) infers TF activity by building gene regulatory networks (GRNs) from scRNA-seq data via co-expression and TF-motif enrichment, defining regulons and scoring their activity per cell using AUCell, thus avoiding reliance on external databases.
  • DecoupleR offers a faster alternative by scoring TF activity using curated regulator-target databases (e.g., DoRothEA) and various statistical methods (e.g., ULM, weighted mean), making it suitable for smaller datasets or when rapid, transparent assessment is needed.
  • Multi-omic approaches like FigR and Epiregulon integrate scRNA-seq with scATAC-seq data to leverage chromatin accessibility information, enabling inference of TF activity by linking distal cis-regulatory elements to genes or by requiring co-occurrence of TF expression and binding site accessibility.
  • The choice of method depends on data availability (scRNA-seq only vs. matched scRNA-seq/scATAC-seq), the maturity of regulatory knowledge for the organism, and the biological question, with SCENIC preferred for novel regulon discovery and decoupleR for known TF activity assessment.
  • Robust TF activity inference requires rigorous quality control of input data, appropriate normalization and batch correction, careful assessment of results for biological plausibility, and ultimately, experimental validation of computational predictions.

Single-cell RNA sequencing measures gene expression, but expression levels of transcription factors (TFs) do not always reflect their functional activity. Post-translational modifications, nuclear localization, chromatin accessibility, and cofactor availability all influence whether a TF actually drives target gene expression. Researchers who want to understand regulatory mechanisms must therefore infer TF activity from indirect evidence instead of read it directly from expression counts. This article compares two broad families of approaches: SCENIC, which builds gene regulatory networks (GRNs) from co-expression and motif enrichment, and decoupleR, which scores TF activity using curated regulator-target databases. It also covers related tools including FigR and Epiregulon that integrate chromatin accessibility data, and discusses inputs, outputs, quality controls, and interpretation limits for each approach.

The Core Problem: Expression Does Not Equal Activity

A TF can be highly expressed in a cell yet remain inactive because it is sequestered in the cytoplasm, bound by an inhibitor, or unable to access its target loci due to closed chromatin. Conversely, a TF with modest expression can exert strong regulatory effects if it is constitutively nuclear and its binding sites are accessible. This disconnect between mRNA abundance and functional output is a central challenge in single-cell regulatory genomics.

Researchers studying gene regulation therefore need computational methods that infer TF activity from observable features. These features include the expression of known target genes, the accessibility of TF binding motifs in chromatin, and the co-variation of TF and target expression across many cells. Each inference strategy makes different assumptions and produces different outputs, and the choice of method affects biological conclusions.

The practical question for a researcher with a new single-cell dataset is not which tool is best in the abstract, but which tool fits the data they have, the biological question they are asking, and the computational resources available. This article provides a decision framework grounded in published methods and applications.

At a Glance: Comparing TF Activity Inference Approaches

MethodPrimary Input DataCore ModelOutputKey StrengthMain Limitation
SCENIC / pySCENICscRNA-seq expression matrixCo-expression network (GRNBoost2 or GENIE3) plus motif enrichment to define regulons, then AUCell scoringPer-cell regulon activity scoresDoes not require external regulator-target databases, learns regulons from the dataRequires many cells for stable co-expression inference, motif databases must match the organism
decoupleR with DoRothEA or PROGENyscRNA-seq or bulk expression matrixCurated regulator-target signatures scored with enrichment methods such as ULM, weighted mean, or AUCellPer-cell TF activity or pathway activity scoresFast, works on small datasets, transparent model assumptionsDepends on database completeness and accuracy, cannot discover novel regulons
FigRPaired scRNA-seq and scATAC-seq from the same biological systemPairs cells across modalities, connects distal cis-regulatory elements to genes, infers GRNsTF activity scores and regulatory chromatin domains (DORCs)Uses chromatin accessibility to link enhancers to target genesRequires matched multi-omic data, computationally intensive
EpiregulonscATAC-seq and scRNA-seq dataConstructs GRNs from co-occurrence of TF expression and chromatin accessibility at TF binding sitesTF activity predictions that account for post-transcriptional modulationCan infer activity of coregulators and TFs with neomorphic mutations using ChIP-seq dataRequires ATAC-seq data, newer method with evolving documentation

SCENIC: Learning Regulons from Co-Expression and Motifs

SCENIC, implemented in Python as pySCENIC, infers TF activity through a three-step pipeline. First, it identifies potential TF targets based on co-expression across cells using either GRNBoost2 or GENIE3. Second, it performs TF-motif enrichment analysis to retain only direct targets, defining regulons as the set of genes regulated by a given TF. Third, it scores the activity of these regulons in individual cells using AUCell, which evaluates whether the genes in a regulon are enriched in the top-expressed genes of each cell [<a href="#ref-1">1</a>].

The key conceptual move in SCENIC is that it does not rely on external regulator-target databases. Instead, it learns the regulatory network from the expression data itself, then filters the co-expression links through motif enrichment to remove indirect associations. This means SCENIC can potentially discover cell type-specific regulatory relationships that curated databases might miss.

Practical Implementation of pySCENIC

The pySCENIC pipeline requires a single-cell expression matrix, typically in the form of a loom file or an AnnData object. The workflow proceeds through three commands: grn, ctx, and aucell. The grn step runs GRNBoost2 or GENIE3 to infer co-expression modules. The ctx step performs cisTarget motif enrichment to prune the network to direct binding targets. The aucell step scores regulon activity per cell.

Researchers should be aware that the co-expression step is the most computationally demanding part of the pipeline. GRNBoost2 is designed to be faster than GENIE3 and is the recommended default in pySCENIC [<a href="#ref-1">1</a>]. For datasets with hundreds of thousands of cells, the co-expression inference may require substantial memory and runtime, and users should plan their computing environment accordingly.

The motif enrichment step requires a database of TF binding motifs and their genomic targets for the organism under study. These databases are organism-specific, and using a mismatched database will produce unreliable regulons. For non-model organisms, researchers may need to build custom motif databases, which is a substantial undertaking.

SCENIC in Published Studies

SCENIC has been applied across diverse biological contexts. In a study of renal aging, researchers analyzed kidney single-cell RNA sequencing data from mice aged 8 weeks to 24 months and used SCENIC to infer Stat1 as a key age-related transcription factor promoting iron dyshomeostasis and ferroptosis in macrophages by regulating expression of the iron chaperone Pcbp1 [<a href="#ref-2">2</a>]. This example illustrates how SCENIC can nominate specific TFs for downstream functional validation.

In a study of primary Sjögren's syndrome, SCENIC analysis predicted STAT1 and EGR1 as candidate upstream transcription factors potentially associated with CCR1 expression in classical monocytes [<a href="#ref-3">3</a>]. The researchers then used molecular docking and RT-qPCR validation in clinical samples to assess the relevance of their computational predictions.

In ulcerative colitis research, investigators used AUCell scoring as part of an integrated analysis that combined differential expression, macrophage high-dimensional weighted gene co-expression network analysis, and machine learning to prioritize CCL3, CCL4, and JUNB as inflammatory regulatory signatures [<a href="#ref-4">4</a>]. This study demonstrates that SCENIC components such as AUCell can be used within broader analytical frameworks.

Limitations of SCENIC

SCENIC's reliance on co-expression means it requires sufficient cell numbers to estimate correlations reliably. Sparse datasets with few cells per cluster will produce unstable networks. The motif enrichment step also depends on the quality and completeness of motif databases, and regulons for TFs with poorly characterized binding motifs will be incomplete.

SCENIC infers activity from expression data alone and does not directly incorporate chromatin accessibility. A TF may be predicted as active based on target gene expression, but the method cannot distinguish whether the TF is acting as a pioneer factor opening chromatin or responding to already accessible loci. Studies that integrate scRNA-seq with scATAC-seq have shown that chromatin accessibility and gene expression can show opposing patterns along differentiation trajectories, with extensive epigenetic priming occurring before transcriptional commitment [<a href="#ref-5">5</a>]. This finding suggests that expression-based inference alone may miss regulatory events that occur at the chromatin level.

decoupleR: Scoring TF Activity with Curated Signatures

decoupleR takes a fundamentally different approach. Instead of learning regulatory networks from the data, it scores the activity of TFs using curated regulator-target databases such as DoRothEA, which contains signed TF-target interactions collected from literature and curated resources. The decoupleR package provides multiple statistical methods for scoring, including univariate linear model (ULM), weighted mean, and AUCell, and it can be applied to both single-cell and bulk expression data.

The primary advantage of decoupleR is speed and simplicity. Because it does not require network inference, it runs quickly even on large datasets and can be applied to datasets with relatively few cells. It also provides transparent model assumptions: the user knows exactly which target genes are being used to score each TF and how the scoring is computed.

Using decoupleR in Practice

A typical decoupleR workflow begins with a normalized expression matrix and a regulator-target database. The user selects a scoring method, runs the analysis, and obtains a matrix of TF activity scores with the same dimensions as the input expression matrix. These scores can then be used for downstream analyses such as differential activity testing between conditions, clustering, or visualization.

In a study of atherosclerotic plaque endotypes, researchers used decoupleR with DoRothEA to characterize transcriptional states in stable and unstable human carotid plaques. They projected bulk-derived inflammatory and structural signatures onto single-cell data and used transcription factor activity inference to localize inflammatory programs to specific macrophage populations [<a href="#ref-6">6</a>]. This example shows how decoupleR can bridge bulk and single-cell analyses.

Limitations of decoupleR

The main limitation of decoupleR is its dependence on curated databases. If a TF is not well represented in DoRothEA or if the database contains incorrect target assignments, the activity scores will be unreliable. Curated databases also tend to be biased toward well-studied TFs, and they may not capture cell type-specific regulatory relationships.

decoupleR cannot discover novel regulons. If a TF has no annotated targets in the database, it will receive no activity score regardless of its actual regulatory role in the data. Researchers studying non-model organisms or less-characterized TFs should verify that their organism and TFs of interest are adequately covered by the chosen database.

Integrating Chromatin Accessibility: FigR and Epiregulon

The observation that chromatin accessibility and gene expression can diverge has motivated the development of methods that integrate scATAC-seq with scRNA-seq. These methods use accessibility at TF binding sites as direct evidence of regulatory potential.

FigR: Functional Inference of Gene Regulation

FigR was developed to computationally pair scATAC-seq cells with scRNA-seq cells, connect distal cis-regulatory elements to genes, and infer gene regulatory networks to identify candidate TF regulators. In a study of resting and stimulated human blood cells, researchers generated approximately 91,000 single-cell profiles and used FigR to define domains of regulatory chromatin (DORCs) of immune stimulation. They found that cells alter chromatin accessibility and gene expression at timescales of minutes, and construction of the stimulation GRN elucidated TF activity at disease-associated DORCs [<a href="#ref-7">7</a>].

FigR addresses a limitation of expression-only methods by using chromatin accessibility to link distal regulatory elements to their target genes. This allows the method to capture regulatory interactions that are mediated by enhancers located far from gene promoters.

Epiregulon: Accounting for Post-Transcriptional Modulation

Epiregulon constructs GRNs from single-cell ATAC-seq and RNA-seq data by considering the co-occurrence of TF expression and chromatin accessibility at TF binding sites in each cell. The method can use ChIP-seq data to infer motif-agonistic activity of transcriptional coregulators or TFs harboring neomorphic mutations [<a href="#ref-8">8</a>].

The developers of Epiregulon demonstrated its utility by accurately predicting the effects of androgen receptor inhibition across different drug modalities, including an AR antagonist and an AR degrader. The method also delineated the mechanisms of a SMARCA4 degrader by identifying context-dependent interaction partners and prioritized drivers of lineage reprogramming and tumorigenesis [<a href="#ref-8">8</a>].

Epiregulon addresses a specific gap in TF activity inference: methods that rely solely on gene expression often neglect post-transcriptional modulation of TFs. By requiring both TF expression and chromatin accessibility at binding sites, Epiregulon can distinguish between a TF that is expressed but cannot access its targets and a TF that is actively engaging its binding sites.

When to Use Multi-Omic Integration

The decision to use FigR, Epiregulon, or similar integration methods depends on data availability. These methods require matched scATAC-seq and scRNA-seq data from the same biological system. Generating such data is more expensive and technically demanding than scRNA-seq alone, but it provides information that expression data cannot supply.

Studies of human developmental hematopoiesis have shown that integrative analysis of chromatin accessibility and gene expression can reveal extensive epigenetic priming of hematopoietic stem cells and multipotent progenitors prior to lineage commitment, even when transcriptional priming is not yet apparent [<a href="#ref-5">5</a>]. Similarly, research on gynecologic malignancies demonstrated that malignant cells acquire previously unannotated regulatory elements to drive hallmark cancer pathways, and that cells from within the same patient show substantial variation in chromatin accessibility linked to transcriptional output [<a href="#ref-9">9</a>]. These findings would not be accessible through expression-based inference alone.

For researchers studying Alzheimer's disease, integration of epigenomic, transcriptomic, and motif information enabled inference of upstream regulators of microglial cell states, gene regulatory networks, enhancer-gene links, and TF-driven microglial state transitions. The study of 194,000 single-nucleus microglial transcriptomes and epigenomes across 443 human subjects identified 12 microglial transcriptional states and 1,542 AD-differentially-expressed genes [<a href="#ref-10">10</a>]. The authors validated their predictions experimentally by showing that ectopic expression of predicted homeostatic-state activators induced homeostatic features in human iPSC-derived microglia-like cells.

Choosing Between Expression-Based and Multi-Omic Methods

The choice between SCENIC, decoupleR, and multi-omic integration methods should be guided by the biological question and the available data.

If the researcher has scRNA-seq data only and wants to discover cell type-specific regulatory networks, SCENIC is the appropriate choice because it learns regulons from the data. If the researcher has scRNA-seq data and wants a fast, transparent assessment of TF activity using established knowledge, decoupleR with DoRothEA is appropriate. If the researcher has matched scATAC-seq and scRNA-seq data and wants to understand how chromatin accessibility shapes regulatory output, FigR or Epiregulon should be considered.

The table below summarizes the decision criteria.

Data AvailableBiological QuestionRecommended ApproachRationale
scRNA-seq onlyWhat regulons are active in my cell types?SCENIC / pySCENICLearns regulons from co-expression and motif enrichment without external databases
scRNA-seq onlyAre known TFs active in my condition?decoupleR with DoRothEAFast scoring using curated signatures, transparent assumptions
scRNA-seq onlyWhat pathways are altered between conditions?decoupleR with PROGENyPathway activity inference from curated signatures
Matched scRNA-seq and scATAC-seqHow does chromatin accessibility regulate expression?FigRPairs cells across modalities and links distal elements to genes
Matched scRNA-seq and scATAC-seqWhich TFs are actively engaging their binding sites?EpiregulonRequires co-occurrence of TF expression and accessibility at binding sites
Matched scRNA-seq and scATAC-seqWhat are the regulatory drivers of disease states?FigR or Epiregulon with downstream validationMulti-omic integration captures regulatory mechanisms invisible to expression alone

Practical Workflow for TF Activity Inference

A robust TF activity inference workflow involves several stages beyond running the chosen tool. The following steps apply to most single-cell regulatory analyses.

Step 1: Quality Control of Input Data

TF activity inference is only as reliable as the expression or accessibility data feeding into it. Single-cell quality control should remove cells with low library complexity, high mitochondrial read fractions, or excessive doublet rates before network inference or activity scoring. The specific thresholds depend on tissue type and protocol, and researchers should examine distributions instead of applying universal cutoffs.

For scRNA-seq data, genes detected in very few cells should be filtered because they provide little information for co-expression inference. For scATAC-seq data, cells with very low fragment counts should be removed because accessibility calls at individual loci become unreliable.

Step 2: Normalization and Batch Correction

Expression data should be normalized to account for differences in sequencing depth across cells. For SCENIC, the developers recommend using raw counts and letting the pipeline handle normalization internally, but users should verify that their input format matches the expected structure. For decoupleR, normalized expression values are typically used.

If the dataset includes multiple batches, samples, or sequencing runs, batch correction should be performed before TF activity inference. Batch effects can create spurious co-expression patterns that distort network inference. However, researchers should be cautious about over-correction, which can remove genuine biological variation.

Step 3: Run the Inference Method

The specific commands depend on the chosen tool. For pySCENIC, the workflow is:

  1. Run pyscenic grn to infer co-expression modules
  2. Run pyscenic ctx to perform motif enrichment and define regulons
  3. Run pyscenic aucell to score regulon activity per cell

For decoupleR, the workflow is:

  1. Load the expression matrix and the DoRothEA database
  2. Select a scoring method such as ULM or weighted mean
  3. Run the scoring function to obtain TF activity scores

For FigR or Epiregulon, the workflow involves additional steps to pair cells across modalities and link regulatory elements to genes. These methods require careful preprocessing of both scRNA-seq and scATAC-seq data.

Step 4: Quality Assessment of Results

After running TF activity inference, researchers should assess whether the results are biologically plausible. Key checks include:

  • Do known marker TFs for major cell types show expected activity patterns?
  • Are TF activity scores stable across cells within the same cluster?
  • Do regulons contain genes with known functional relationships to the TF?

For SCENIC, the regulon specificity score and regulon quality score provide quantitative measures of regulon reliability. Regulons with very few target genes or with targets that are not supported by motif enrichment should be interpreted cautiously.

Step 5: Downstream Analysis and Validation

TF activity scores can be used for differential activity testing between conditions, trajectory analysis, clustering, or integration with other data types. However, computational predictions require experimental validation. Published studies have validated TF activity predictions through perturbation experiments, such as the Perturb-seq approach that combines single-cell RNA-seq with CRISPR-based perturbations to assay the effects of TF knockdown or knockout on transcriptional profiles [<a href="#ref-11">11</a>].

In the Alzheimer's disease study, the authors validated their predicted homeostatic-state activators by ectopic expression in human iPSC-derived microglia-like cells and showed that inhibiting activators of inflammation could block inflammatory progression [<a href="#ref-10">10</a>]. This experimental validation is the gold standard for TF activity inference.

Records and Measurements for Reproducibility

Reproducible TF activity inference requires careful record keeping. Researchers should document:

  • The exact version of the inference tool and all dependencies
  • The reference genome and annotation version
  • The motif database and its version for SCENIC
  • The regulator-target database and its version for decoupleR
  • All parameters used in each step of the pipeline
  • The computational environment, including operating system and package versions

Containerized workflows can help ensure reproducibility. The nf-core community provides standards for building reproducible bioinformatics pipelines, and their documentation describes best practices for pipeline configuration and usage [<a href="#ref-12">12</a>]. Galaxy provides accessible workflow training and analysis tutorials that emphasize reproducibility [<a href="#ref-13">13</a>]. Bioconductor offers official package and workflow documentation for reproducible genomic analysis [<a href="#ref-14">14</a>].

Version control is essential. Researchers should track changes to analysis scripts using Git, and The Carpentries provides foundational lessons on shell, Git, and programming that support reproducible research practices [<a href="#ref-15">15</a>].

Common Failure Patterns and How to Avoid Them

Several recurring problems undermine TF activity inference from single-cell data.

Failure Pattern 1: Applying Expression-Based Methods to Data with Strong Batch Effects

Co-expression inference in SCENIC assumes that expression correlations reflect regulatory relationships. If the dataset contains strong batch effects, correlations may reflect technical artifacts instead of biology. This produces regulons that are not reproducible across batches.

Prevention: Perform batch correction before network inference, or run SCENIC separately on each batch and compare regulon stability. If regulons differ substantially across batches, batch effects are likely dominating the signal.

Failure Pattern 2: Using Mismatched Motif or Regulator Databases

SCENIC requires motif databases for the organism under study. Using a human motif database for mouse data, or vice versa, will produce regulons that are largely meaningless. Similarly, decoupleR with DoRothEA requires that the database covers the TFs of interest for the organism being studied.

Prevention: Verify that the motif database matches the organism and genome version. Check DoRothEA coverage for the TFs of interest before running decoupleR.

Failure Pattern 3: Interpreting Activity Scores as Expression Levels

TF activity scores are not expression levels. A high activity score means that the TF's target genes are coordinately expressed or that its binding sites are accessible, not that the TF itself is highly expressed. Researchers who interpret activity scores as expression will draw incorrect conclusions.

Prevention: Always examine TF expression alongside activity scores. If a TF shows high activity but low expression, this may indicate post-transcriptional regulation, which is biologically interesting but requires careful interpretation.

Failure Pattern 4: Overinterpreting Regulons as Causal Networks

Regulons inferred from co-expression are correlational, not causal. A regulon indicates that a TF and its putative targets are co-expressed and that the TF's binding motif is enriched near the targets, but this does not prove that the TF regulates those targets in the studied cells.

Prevention: Treat regulons as hypotheses for experimental testing. Use perturbation data such as Perturb-seq to validate regulatory relationships [<a href="#ref-11">11</a>].

Failure Pattern 5: Ignoring the Difference Between scRNA-seq and snRNA-seq

Single-nucleus RNA sequencing (snRNA-seq) captures nuclear transcripts, which may differ from cytoplasmic mRNA captured by scRNA-seq. This distinction matters for TF activity inference because some TFs have predominantly nuclear localization, and their transcripts may be differentially captured by the two methods. The Alzheimer's disease study used single-nucleus microglial transcriptomes, and the authors noted the importance of this approach for studying disease-relevant cell states [<a href="#ref-10">10</a>].

Prevention: Understand the tissue and cell types being studied and choose the appropriate capture method. Compare results from scRNA-seq and snRNA-seq when both are available.

Limitations of TF Activity Inference

All TF activity inference methods have inherent limitations that researchers should acknowledge in their reporting.

Expression-Based Methods Cannot Detect Post-Translational Regulation

SCENIC and decoupleR infer activity from expression data. They cannot detect TFs that are regulated by phosphorylation, ubiquitination, nuclear translocation, or protein-protein interactions. Epiregulon partially addresses this by requiring chromatin accessibility at binding sites, but even this approach cannot capture all post-translational regulation [<a href="#ref-8">8</a>].

Curated Databases Are Incomplete and Biased

DoRothEA and similar databases are biased toward well-studied TFs. TFs with limited literature coverage will have sparse or missing target annotations, leading to unreliable activity scores. This bias is particularly problematic for non-model organisms and for less-characterized TFs.

Co-Expression Inference Requires Sufficient Cell Numbers

SCENIC's co-expression step requires enough cells to estimate correlations reliably. For datasets with few cells per cluster, the inferred networks will be unstable. The developers recommend using datasets with sufficient cell numbers, but the exact threshold depends on the complexity of the underlying regulatory network [<a href="#ref-1">1</a>].

Multi-Omic Methods Require Matched Data

FigR and Epiregulon require matched scATAC-seq and scRNA-seq data. Generating such data is expensive and technically challenging, and not all research questions justify the additional cost. Researchers should weigh the added information from chromatin accessibility against the cost of generating matched data.

Computational Cost Can Be Substantial

SCENIC's co-expression inference is computationally intensive, particularly for large datasets. GRNBoost2 is faster than GENIE3, but both require substantial memory and runtime [<a href="#ref-1">1</a>]. Researchers working with limited computational resources should plan accordingly or consider using decoupleR, which is much faster.

Safety and Regulatory Context

TF activity inference is a computational analysis method and does not directly involve laboratory safety or regulatory compliance. However, researchers should be aware of several contextual considerations.

Data Privacy and Consent

Single-cell datasets derived from human subjects may contain sensitive information. Researchers should ensure that their data usage complies with applicable consent agreements and data protection regulations. Public repositories such as NCBI provide data resources and search systems, but researchers are responsible for understanding the terms of use for each dataset [<a href="#ref-16">16</a>].

Reproducibility Standards

Funding agencies and journals increasingly require reproducible analysis workflows. Researchers should document their analysis steps thoroughly and consider depositing code and processed data in public repositories. EMBL-EBI provides training on bioinformatics data resources and practical analysis education that can support reproducible practices [<a href="#ref-17">17</a>].

Reporting Guidelines

When reporting TF activity inference results, researchers should describe the method version, parameters, databases, and quality control steps. This transparency allows other researchers to assess the reliability of the findings and to reproduce the analysis.

Professional Escalation Criteria

Researchers should seek expert consultation when they encounter specific challenges in TF activity inference.

When to Consult a Bioinformatics Core or Collaborator

  • The dataset has complex experimental design with multiple batches, conditions, or time points that require sophisticated correction approaches
  • The organism under study is not well covered by existing motif or regulator-target databases
  • The computational requirements exceed available local resources and require access to high-performance computing
  • The researcher needs to integrate scRNA-seq with scATAC-seq or other modalities and is unfamiliar with multi-omic analysis

When to Seek Statistical or Methodological Advice

  • TF activity scores show unexpected patterns that cannot be explained by known biology
  • Different inference methods produce conflicting results for the same dataset
  • The researcher needs to perform differential TF activity testing and is unsure which statistical approach is appropriate
  • The researcher plans to use TF activity scores as inputs to downstream modeling and needs guidance on error propagation

When to Consider Experimental Validation

  • TF activity predictions will be used to guide functional experiments
  • The researcher plans to make claims about regulatory mechanisms based on computational inference alone
  • The predicted TFs are not well characterized and require confirmation of their regulatory roles

A Practical Decision Framework for Selecting TF Activity Inference Methods

Choosing between SCENIC, decoupleR, FigR, and Epiregulon requires a structured evaluation that goes beyond the simple question of which data type is available. Researchers often discover that their initial choice based on data availability alone leads to downstream problems when the method's assumptions do not match the biological system under study. A practical decision framework should therefore incorporate three additional dimensions: the maturity of the regulatory knowledge for the organism, the expected complexity of the regulatory architecture, and the intended use of the activity scores in downstream analysis.

Dimension 1: Regulatory Knowledge Maturity

The first dimension to assess is how well the regulatory landscape of the organism and tissue is characterized. For human and mouse studies in well-studied contexts such as immune cells, cancer, or development, curated databases like DoRothEA provide substantial coverage of TF-target relationships. In these settings, decoupleR can produce reliable activity scores quickly and with transparent assumptions. The atherosclerotic plaque study demonstrated this approach by using decoupleR with DoRothEA to characterize transcriptional states in stable and unstable human carotid plaques, successfully localizing inflammatory programs to specific macrophage populations [<a href="#ref-6">6</a>].

For less-characterized organisms, tissues, or cell states, curated databases may have sparse coverage. SCENIC becomes the more appropriate choice because it learns regulons directly from the data through co-expression and motif enrichment, without requiring pre-existing target annotations [<a href="#ref-1">1</a>]. The renal aging study applied SCENIC to mouse kidney data and successfully inferred Stat1 as a key age-related transcription factor, a finding that emerged from the data instead of from database queries [<a href="#ref-2">2</a>].

The assessment question for this dimension is straightforward: does a reliable regulator-target database exist for the organism and cell types under investigation? If the answer is no, SCENIC is the preferred approach. If the answer is yes, decoupleR offers a faster and more transparent alternative.

Dimension 2: Regulatory Architecture Complexity

The second dimension concerns whether the regulatory mechanisms of interest involve distal enhancers, chromatin remodeling, or pioneer factor activity. Expression-based methods cannot distinguish between a TF that is actively opening chromatin and one that is simply maintaining expression of already accessible targets. Studies of human developmental hematopoiesis revealed extensive epigenetic priming of hematopoietic stem cells and multipotent progenitors prior to lineage commitment, with opposing patterns of chromatin accessibility and differentiation coinciding with dynamic changes in lineage-specific TF activity [<a href="#ref-5">5</a>]. These regulatory events would be invisible to expression-based inference alone.

When the biological question involves chromatin state transitions, epigenetic priming, or enhancer-mediated regulation, multi-omic integration methods become necessary. FigR was developed specifically to connect distal cis-regulatory elements to genes and infer GRNs that capture these regulatory interactions [<a href="#ref-7">7</a>]. Epiregulon goes further by requiring co-occurrence of TF expression and chromatin accessibility at TF binding sites in each cell, enabling inference of TF activity that accounts for post-transcriptional modulation [<a href="#ref-8">8</a>].

The assessment question for this dimension: is the regulatory mechanism of interest likely to involve changes in chromatin state, or is the research question focused on which TFs maintain the current transcriptional program? The former requires multi-omic integration, while the latter can be addressed with expression-based methods.

Dimension 3: Downstream Use of Activity Scores

The third dimension considers how the TF activity scores will be used after inference. Different downstream analyses impose different requirements on the activity scores.

For differential activity testing between conditions, decoupleR provides a straightforward framework because it produces activity scores with transparent model assumptions that can be directly compared across groups. The ulcerative colitis study integrated AUCell scoring with machine learning algorithms to prioritize candidate genes, demonstrating how activity scores can feed into broader analytical pipelines [<a href="#ref-4">4</a>].

For trajectory analysis or pseudotime inference, SCENIC's per-cell regulon activity scores are well suited because they capture continuous variation in regulatory activity across a differentiation continuum. The Sjögren's syndrome study combined SCENIC transcription factor network analysis with pseudotime trajectory analysis to identify STAT1 and EGR1 as candidate upstream regulators of CCR1 expression in classical monocytes [<a href="#ref-3">3</a>].

For drug response prediction or therapeutic target prioritization, Epiregulon offers unique capabilities. The developers demonstrated accurate prediction of androgen receptor inhibition effects across different drug modalities and delineated mechanisms of a SMARCA4 degrader by identifying context-dependent interaction partners [<a href="#ref-8">8</a>]. This application requires the post-transcriptional awareness that only multi-omic integration provides.

The assessment question for this dimension: will the activity scores be used for hypothesis generation, differential testing, trajectory analysis, or therapeutic prediction? Each use case has different requirements for resolution, interpretability, and biological grounding.

A Scoring Matrix for Method Selection

To operationalize this framework, researchers can score each candidate method against the three dimensions using a simple rubric. For each dimension, assign a score from 1 to 3 indicating how well the method fits the research context.

DimensionSCENICdecoupleRFigREpiregulon
Regulatory knowledge maturity3 for poorly characterized organisms, 2 for well-characterized1 for poorly characterized, 3 for well-characterized2 for both, requires matched data2 for both, requires matched data
Regulatory architecture complexity2 for expression-focused questions, 1 for chromatin questions1 for chromatin questions3 for enhancer-mediated regulation3 for post-transcriptional modulation
Downstream use flexibility3 for trajectory and regulon discovery3 for differential testing and speed2 for chromatin-focused questions3 for drug response and therapeutic targets

The method with the highest total score is the initial recommendation. However, researchers should also consider whether the data required for the top-scoring method is feasible to generate. If matched scATAC-seq and scRNA-seq data are not available, the multi-omic methods cannot be applied regardless of their fit.

Record Keeping for Method Selection

Documenting the method selection process is as important as documenting the analysis itself. Researchers should record the following for each dimension:

  • The specific databases considered and their coverage statistics for the organism and TFs of interest
  • The biological rationale for expecting or not expecting chromatin-mediated regulation in the system
  • The planned downstream analyses and any constraints they impose on activity score format or resolution
  • The data generation feasibility assessment, including cost and technical requirements for matched multi-omic data

This documentation supports reproducibility and provides a clear rationale for reviewers who may question the choice of inference method.

Troubleshooting Method Selection Failures

When the selected method produces unsatisfactory results, the failure often traces back to the decision framework instead of the method itself.

If SCENIC produces regulons that do not align with known biology, the issue may be that the regulatory knowledge for the organism is more mature than initially assessed. Switching to decoupleR with a curated database may provide more reliable activity scores for well-characterized TFs.

If decoupleR produces activity scores that are uniformly low or uninformative, the curated database may have inadequate coverage for the specific TFs driving the biological response. SCENIC may discover regulons that the database misses.

If expression-based methods produce results that conflict with known chromatin biology, the system likely involves epigenetic regulation that requires multi-omic integration. The opposing patterns of chromatin accessibility and differentiation observed in developmental hematopoiesis [<a href="#ref-5">5</a>] and the substantial variation in chromatin accessibility linked to transcriptional output in gynecologic malignancies [<a href="#ref-9">9</a>] both demonstrate situations where expression-based inference would be insufficient.

If multi-omic methods produce unstable results, the issue may be in the cell pairing across modalities or in the quality of the ATAC-seq data. Researchers should verify that the scATAC-seq data has sufficient depth and that the cell type annotation is consistent across modalities before interpreting TF activity differences.

Escalation Criteria for Method Selection

Researchers should escalate to expert consultation when the decision framework produces ambiguous recommendations. Specific situations that warrant consultation include:

  • The three dimensions point to different methods, and the trade-offs are not clear
  • The organism is a non-model system with partial regulatory knowledge and unknown chromatin architecture
  • The downstream analysis requires combining TF activity scores with other data types in ways that are not well documented
  • The research question spans multiple regulatory mechanisms that may require complementary inference approaches

In these situations, a bioinformatics core or computational collaborator can provide context-specific guidance that a general decision framework cannot capture.

Frequently Asked Questions

What is the difference between TF expression and TF activity?

TF expression refers to the measured mRNA abundance of the transcription factor gene in a cell. TF activity refers to the functional output of that TF, which depends on additional factors including protein abundance, post-translational modifications, nuclear localization, chromatin accessibility at binding sites, and availability of cofactors. A TF can be highly expressed but inactive, or weakly expressed but highly active. Computational methods infer activity from indirect evidence such as target gene expression or chromatin accessibility at binding motifs.

When should I use SCENIC instead of decoupleR?

Use SCENIC when you want to discover cell type-specific regulatory networks from your data without relying on external databases. SCENIC learns regulons from co-expression and motif enrichment, which allows it to identify regulatory relationships that curated databases might miss. Use decoupleR when you want a fast, transparent assessment of known TF activity using curated signatures. decoupleR is appropriate for datasets with few cells, for organisms with good database coverage, and when you want to avoid the computational cost of network inference.

Can I run TF activity inference on single-nucleus RNA-seq data?

Yes, TF activity inference can be applied to snRNA-seq data. The Alzheimer's disease study used single-nucleus microglial transcriptomes for regulatory analysis [<a href="#ref-10">10</a>]. However, researchers should be aware that snRNA-seq captures nuclear transcripts, which may differ from cytoplasmic mRNA. This difference can affect the detection of certain transcripts and may influence co-expression patterns. Validate that the inference results are consistent with known biology for the cell types under study.

What is the minimum number of cells needed for SCENIC?

There is no universal minimum cell number for SCENIC because the requirement depends on the complexity of the underlying regulatory network and the number of cell types in the dataset. Co-expression inference requires enough cells per cluster to estimate correlations reliably. In practice, datasets with fewer than a few hundred cells per cluster may produce unstable regulons. Researchers should assess regulon stability by running the analysis on subsampled data and checking whether regulons are consistent.

How do I validate TF activity predictions experimentally?

Experimental validation approaches include perturbation experiments such as Perturb-seq, which combines single-cell RNA-seq with CRISPR-based perturbations to assay the effects of TF knockdown or knockout on transcriptional profiles [<a href="#ref-11">11</a>]. Other approaches include ectopic expression of predicted activators, as demonstrated in the Alzheimer's disease study where expression of predicted homeostatic-state activators induced homeostatic features in iPSC-derived microglia-like cells [<a href="#ref-10">10</a>]. ChIP-seq can confirm TF binding at predicted target loci, and reporter assays can test whether TF binding sites drive expression.

What is the role of chromatin accessibility data in TF activity inference?

Chromatin accessibility data from scATAC-seq provides direct evidence of whether TF binding sites are accessible in a given cell. Methods such as FigR and Epiregulon integrate accessibility with expression to infer TF activity [<a href="#ref-7">7</a>][<a href="#ref-8">8</a>]. This integration can reveal regulatory mechanisms that are invisible to expression-based methods, such as epigenetic priming where chromatin becomes accessible before transcriptional changes occur [<a href="#ref-5">5</a>]. However, generating matched scATAC-seq and scRNA-seq data is more expensive than scRNA-seq alone.

How do batch effects affect TF activity inference?

Batch effects can create spurious co-expression patterns that distort network inference in SCENIC and bias activity scores in decoupleR. If cells from different batches have systematic differences in expression that are unrelated to biology, the inference methods may interpret these differences as regulatory signals. Perform batch correction before TF activity inference, or run the analysis separately on each batch and compare regulon stability across batches.

What should I report when publishing TF activity inference results?

Report the exact version of the inference tool and all dependencies, the reference genome and annotation version, the motif database and its version for SCENIC, the regulator-target database and its version for decoupleR, all parameters used in each step, and the computational environment. Describe quality control steps applied to the input data and any batch correction performed. Deposit code and processed data in public repositories to support reproducibility.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

[1] [Inference of Gene Regulatory Network from Single-Cell Transcriptomic Data Using pySCENIC.](https://pubmed.ncbi.nlm.nih.gov/34251625). Methods in molecular biology (Clifton, N.J.), 2021. [2] [Macrophage iron dyshomeostasis promotes aging-related renal fibrosis.](https://pubmed.ncbi.nlm.nih.gov/39016438). Aging cell, 2024. [3] [CCR1-mediated monocyte chemotaxis in the immunopathology of primary Sjögren's syndrome: multi-omics integration analysis and computational target prioritization implicating <,i>,Polygonatum odoratum<,/i>,.](https://doi.org/10.3389/fimmu.2026.1867098). 2026. [4] [Macrophage-Centered Integration of Single-Cell and Bulk Transcriptomic Data Identifies CCL3, CCL4, and JUNB as Inflammatory Regulatory Signatures in Ulcerative Colitis.](https://doi.org/10.3390/genes17070841). 2026. [5] [Integrative Single-Cell RNA-Seq and ATAC-Seq Analysis of Human Developmental Hematopoiesis.](https://pubmed.ncbi.nlm.nih.gov/33352111). Cell stem cell, 2021. [6] [Inflammatory and Structural Endotypes of Human Atherosclerotic Plaque Revealed by Integrated Transcriptomic Analysis.](https://doi.org/10.3390/genes17070779). 2026. [7] [Functional inference of gene regulation using single-cell multi-omics.](https://pubmed.ncbi.nlm.nih.gov/36204155). Cell genomics, 2022. [8] [Epiregulon: Single-cell transcription factor activity inference to predict drug response and drivers of cell states.](https://pubmed.ncbi.nlm.nih.gov/40753156). Nature communications, 2025. [9] [A multi-omic single-cell landscape of human gynecologic malignancies.](https://pubmed.ncbi.nlm.nih.gov/34739872). Molecular cell, 2021. [10] [Human microglial state dynamics in Alzheimer's disease progression.](https://pubmed.ncbi.nlm.nih.gov/37774678). Cell, 2023. [11] [Perturb-Seq: Dissecting Molecular Circuits with Scalable Single-Cell RNA Profiling of Pooled Genetic Screens.](https://pubmed.ncbi.nlm.nih.gov/27984732). Cell, 2016. [12] [nf-core Documentation](https://nf-co.re/docs). nf-core. [13] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project. [14] [Bioconductor](https://bioconductor.org/). Bioconductor Project. [15] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries. [16] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information. [17] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.