Functional Inference of Gene Regulation Using Single-Cell Multi-Omics: From Correlation to Causality
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Single-cell multi-omics (scRNA-seq and scATAC-seq) enables moving beyond correlation to causal inference of gene regulatory networks by linking chromatin accessibility to gene expression within the same cell.
- Correlation-based methods (e.g., FigR, SCENIC+) infer enhancer-gene links and TF-target edges, but require experimental validation due to potential confounding signals.
- In silico perturbation (e.g., CellOracle) simulates TF removal to predict phenotypic shifts and prioritize causal regulators, generating testable hypotheses.
- Pooled genetic perturbation with single-cell readout offers direct causal evidence by experimentally manipulating TF expression and observing cell fate changes.
- Summary-data-based Mendelian randomization (SMR) leverages genetic variants as instrumental variables to infer causal associations between gene expression and disease, reducing observational bias.
- Transcription factor footprinting identifies actively bound TFs, distinguishing them from merely expressed ones, and is crucial for prioritizing causal variants in regulatory regions.
Single-cell multi-omics technologies now allow researchers to measure chromatin accessibility and gene expression within the same individual cell, creating opportunities to move beyond correlative associations toward causal inference of gene regulatory networks. This article explains how to use matched single-cell RNA sequencing (scRNA-seq) and single-cell ATAC sequencing (scATAC-seq) data to construct regulatory networks, apply causal inference methods such as in silico perturbation and Mendelian randomization, and identify key transcription factor regulators with practical decision criteria for experimental validation.
The Problem: Correlation Does Not Equal Regulation
A common observation in single-cell genomics is that open chromatin regions correlate with expression of nearby genes. Researchers frequently interpret this correlation as evidence that a transcription factor (TF) binds the accessible region and drives transcription of the target gene. This interpretation is often incomplete. Chromatin accessibility can change without immediate transcriptional consequences, and genes can be expressed without proximal accessible regulatory elements at the time of measurement. The correlation between accessibility and expression reflects potential regulatory capacity, not necessarily active regulation.
The practical problem for researchers is distinguishing between three scenarios. First, a TF may bind an enhancer and directly activate a target gene. Second, a TF may bind an enhancer but require additional cofactors or post-translational modifications to exert effect. Third, the observed correlation may be coincidental, driven by shared upstream signals instead of direct regulatory relationships. Standard correlation-based approaches cannot distinguish these scenarios.
Causal inference methods address this gap by using additional information sources. These include the temporal ordering of regulatory events, genetic variation that perturbs regulatory elements, and experimental or in silico perturbation of transcription factors. When applied to single-cell multi-omics data, these approaches can prioritize transcription factors that are likely to be causal drivers of cell state transitions, disease phenotypes, or developmental trajectories.
At a Glance: Methods for Causal Regulatory Inference
| Method | Data Requirements | Evidence Type | Key Output | Validation Burden |
|---|---|---|---|---|
| Correlation-based network inference (FigR, SCENIC+) | Matched or computationally paired scRNA-seq and scATAC-seq | Statistical association between accessibility and expression | Enhancer-gene links, TF-target edges, regulon activity scores | High: predictions require experimental confirmation |
| In silico perturbation (CellOracle) | Inferred GRN from single-cell multi-omics | Simulated effect of TF removal on cell identity | Predicted phenotype shifts, prioritized TFs | High: predictions validated against known phenotypes or new experiments |
| Pooled genetic perturbation with single-cell readout | CRISPR perturbation plus scRNA-seq or multiome | Direct experimental manipulation of TF expression | Cell fate and state changes attributable to specific TFs | Low: direct causal evidence but requires functional experiments |
| Summary-data-based Mendelian randomization (SMR) | eQTL, pQTL, and GWAS summary statistics | Genetic variants as instrumental variables | Causal association between gene expression and disease | Medium: requires replication cohorts and functional follow-up |
What Single-Cell Multi-Omics Data Provides
Joint profiling of chromatin accessibility and gene expression in individual cells provides the foundation for enhancer-driven gene regulatory network inference. The SCENIC+ method exemplifies this approach by predicting genomic enhancers, identifying candidate upstream transcription factors, and linking enhancers to candidate target genes within the same computational framework. This method was benchmarked across diverse datasets including human peripheral blood mononuclear cells, ENCODE cell lines, melanoma cell states, and Drosophila retinal development, demonstrating applicability across species and tissue types.
The FigR framework similarly pairs scATAC-seq with scRNA-seq cells computationally, connects distal cis-regulatory elements to genes, and infers gene regulatory networks to identify candidate transcription factor regulators. When applied to resting and stimulated human blood cells, FigR enabled construction of stimulation gene regulatory networks that elucidated transcription factor activity at disease-associated domains of regulatory chromatin.
The key advantage of matched single-cell data is the ability to link variation in chromatin accessibility to gene expression quantitatively within the same cell. A study of human ovarian and endometrial tumors processed immediately following surgical resection demonstrated that malignant cells acquire previously unannotated regulatory elements to drive hallmark cancer pathways. The same study showed substantial variation in chromatin accessibility linked to transcriptional output within malignant cells from the same patient, highlighting the importance of intratumoral heterogeneity for regulatory inference.
Core Principles of Regulatory Network Inference
Cis-Regulatory Element Identification
The first step in functional inference is identifying candidate cis-regulatory elements (cCREs) from chromatin accessibility data. Accessible regions are called as peaks, then annotated based on their genomic location relative to transcription start sites. Promoter-proximal regions are distinguished from distal enhancer candidates. A study of 117,911 human lung cells from ever- and never-smokers found that candidate cis-regulatory elements are largely cell type-specific, with 37 percent detected in only one cell type. This cell type specificity means that regulatory inference must be performed within defined cell populations instead of across mixed cell types.
Enhancer-Gene Linking
Connecting distal regulatory elements to their target genes is a central challenge. Approaches include correlation between accessibility and expression across cells, proximity-based assignment, and integration of chromatin conformation data. The FigR framework connects distal cis-regulatory elements to genes using correlation-based approaches that account for the distance between elements and gene bodies. SCENIC+ uses motif information to link enhancers to candidate upstream transcription factors and target genes.
Transcription Factor Motif Analysis
Transcription factor binding site prediction requires curated motif collections. SCENIC+ curated and clustered a motif collection with more than 30,000 motifs to improve both recall and precision of transcription factor identification. The quality of motif collections directly affects the accuracy of regulatory network inference. Researchers should document which motif collection version was used and how motif matching thresholds were set.
Network Construction
Gene regulatory networks are constructed by combining enhancer-gene links with transcription factor-enhancer predictions. The resulting network contains directed edges from transcription factors to target genes, with weights reflecting confidence in the regulatory relationship. Network inference methods vary in how they handle the sparsity of single-cell data, technical noise, and the complex combinatorial logic of transcriptional regulation.
Causal Inference Methods for Regulatory Networks
In Silico Transcription Factor Perturbation
CellOracle represents a machine-learning approach that uses gene regulatory networks inferred from single-cell multi-omics data to perform in silico transcription factor perturbations. The method simulates consequent changes in cell identity using only unperturbed wild-type data. When applied to mouse and human haematopoiesis and zebrafish embryogenesis, CellOracle correctly modeled reported changes in phenotype resulting from transcription factor perturbation.
The practical value of in silico perturbation is the ability to prioritize transcription factors for experimental validation. In the developing zebrafish, systematic in silico perturbation simulated and experimentally validated a previously unreported phenotype resulting from loss of the notochord regulator noto. The approach also identified lhx1a as an axial mesoderm regulator. These findings demonstrate that in silico perturbation can generate testable hypotheses about transcription factor function.
Perturbation with Single-Cell Readout
Pooled genetic perturbation with single-cell transcriptome readout provides direct experimental evidence for transcription factor requirements. A study of human brain organoids used this approach to assess transcription factor requirements for cell fate and state regulation. The study found that certain factors regulate the abundance of cell fates, whereas other factors affect neuronal cell states after differentiation.
The same study identified the transcription factor GLI3 as required for cortical fate establishment in humans, recapitulating previous research in mammalian model systems. By measuring transcriptome and chromatin accessibility in normal or GLI3-perturbed cells, the researchers identified two distinct GLI3 regulomes central to telencephalic fate decisions. One regulome regulated dorsoventral patterning with HES4 and HES5 as direct GLI3 targets, and the other controlled ganglionic eminence diversification later in development.
Summary-Data-Based Mendelian Randomization
Mendelian randomization uses genetic variants as instrumental variables to assess causal relationships between gene expression and disease outcomes. Summary-data-based Mendelian randomization (SMR) integrates expression quantitative trait loci (eQTL) and protein quantitative trait loci (pQTL) with genome-wide association study (GWAS) data. A multi-omics and epigenome-wide study of intracranial aneurysms applied SMR to identify immune-inflammation targets, finding associations for RELT, TNFSF12, ICAM5, and ERAP2 with aneurysm risk.
The advantage of Mendelian randomization approaches is that genetic variants are randomly assigned at conception, reducing confounding that plagues observational correlation studies. When applied to single-cell data, these approaches can identify cell type-specific effects of genetic variants on gene regulation. A study of the human retina mapped eQTL, caQTL, allelic-specific expression, and allelic-specific chromatin accessibility in major retinal cell types, finding that the majority of identified single-cell eQTLs and caQTLs display cell type-specific effects.
Transcription Factor Footprinting
Transcription factor footprinting identifies regions of chromatin protected from nuclease cleavage by bound transcription factors. This approach provides evidence of active binding instead of mere accessibility. In the lung cancer susceptibility study, transcription factor footprinting was used to prioritize candidate causal variants for 68 percent of GWAS loci. Footprinting can distinguish between transcription factors that are merely expressed and those that are actively bound to regulatory elements in specific cell types.
Practical Workflow for Causal Regulatory Inference
Step 1: Define the Biological Question and Cell Types
Begin by specifying the cell types, conditions, or developmental stages of interest. Regulatory networks are cell type-specific, so analysis should be stratified by cell population. Define whether the goal is identifying regulators of a disease phenotype, developmental transition, or response to stimulation. This decision determines which causal inference methods are appropriate.
Step 2: Quality Control and Preprocessing
Apply standard quality control metrics to both scRNA-seq and scATAC-seq data. For scRNA-seq, assess library complexity, mitochondrial read fraction, and total unique molecular identifier counts. For scATAC-seq, assess transcription start site enrichment, fraction of reads in peaks, and total fragment counts. Document all filtering thresholds and justify them based on data distributions. The Bioconductor project provides official packages and workflows for reproducible genomic analysis, including single-cell quality control procedures.
Step 3: Cell Type Annotation
Annotate cell types using marker genes for scRNA-seq and marker motifs or known regulatory elements for scATAC-seq. Integration of the two modalities can improve annotation confidence. The Galaxy Training Network provides accessible workflow training for single-cell analysis, and the nf-core documentation describes community pipeline standards for reproducible analysis workflows.
Step 4: Data Integration Across Modalities
Integrate scRNA-seq and scATAC-seq data from the same biological sample. Methods differ in whether they require matched cells from the same individual or can pair cells computationally. The FigR framework computationally pairs scATAC-seq with scRNA-seq cells, which is necessary when the two assays are performed on separate cell suspensions. When matched multiome data are available from the same cell, integration is more straightforward.
Step 5: Identify Candidate Cis-Regulatory Elements
Call peaks from scATAC-seq data and annotate them relative to gene features. Filter peaks that overlap known blacklist regions. Determine which peaks are differentially accessible between cell types or conditions of interest. A study of mouse secondary palate development used single-cell multiome sequencing to profile chromatin accessibility and gene expression simultaneously within the same cells across four embryonic days, constructing five differentiation trajectories and linking open chromatin signals to gene expression changes.
Step 6: Link Enhancers to Target Genes
Apply enhancer-gene linking methods that combine distance, correlation, and motif information. Document the linking criteria and the number of enhancer-gene pairs retained. Validate links using known regulatory relationships where available. The multi-level cCRE-gene linking system used in the lung cancer study identified candidate susceptibility genes from 57 percent of GWAS loci, with most loci displaying cell-category-specific target genes.
Step 7: Infer Transcription Factor Regulatory Activity
Score transcription factor activity based on the accessibility of its predicted binding sites and the expression of its target genes. Regulon analysis can identify coordinated transcription factor programs. The hepatocellular carcinoma study used high-dimensional co-expression network analysis and regulatory network inference to delineate stemness-associated modules, identifying DUSP9 as a potential regulator enriched in a malignant subpopulation with elevated ERK activation and progenitor-like features.
Step 8: Apply Causal Inference Methods
Select causal inference methods based on available data. If only observational multi-omics data are available, use in silico perturbation methods such as CellOracle or SCENIC+ perturbation analysis. If genetic data are available, apply summary-data-based Mendelian randomization. If perturbation experiments are feasible, design pooled genetic perturbation with single-cell readout. Each method has different assumptions and limitations that should be documented.
Step 9: Prioritize Transcription Factors for Validation
Rank transcription factors by causal inference scores, network centrality, and biological plausibility. Consider whether the transcription factor is expressed in the relevant cell type and whether its binding motifs are enriched in differentially accessible regions. The in silico perturbation analysis in the palate development study identified SHOX2 and MEOX2 as important regulators of anterior and posterior palate development, respectively, providing specific candidates for experimental validation.
Step 10: Validate with Independent Experiments
Design validation experiments appropriate to the biological system. Options include genetic perturbation with single-cell readout, reporter assays for enhancer activity, and chromatin immunoprecipitation followed by sequencing. The retina study validated findings with high-throughput reporter assays. The organoid study used pooled genetic perturbation to assess transcription factor requirements. The renal cell carcinoma study used preclinical models to suggest that targeting the FABP1-PLG-PLAT axis may enhance sensitivity to tyrosine kinase inhibitor therapy.
Options and Tradeoffs in Method Selection
Matched Multiome Versus Separate Assays
Matched multiome data from the same cell provide the strongest link between chromatin accessibility and gene expression. The multiHIVE approach describes hierarchical multimodal deep generative modeling for integrating multimodal data from assays such as CITE-seq, which profiles RNA along with surface proteins, ISSAAC-seq, which captures RNA and chromatin accessibility, and TEA-seq, which simultaneously measures RNA, chromatin accessibility, and surface proteins. These assays obtain multimodal data from the same individual cell, enabling direct linking of cell states to regulatory elements.
Separate scRNA-seq and scATAC-seq assays on different cell suspensions require computational pairing of cells. This approach is less expensive per cell and allows optimization of each assay independently. However, computational pairing introduces uncertainty that can weaken regulatory inference. The choice depends on whether the biological question requires single-cell resolution of the accessibility-expression link.
Correlation-Based Versus Deep Generative Models
Correlation-based methods are transparent and computationally efficient but may miss nonlinear regulatory relationships. Deep generative models such as multiHIVE and scTFBridge can capture complex relationships and disentangle shared and modality-specific signals. scTFBridge disentangles latent spaces into shared and specific components across omics layers and integrates transcription factor motif binding knowledge to align shared embeddings with specific transcription factor regulatory activities.
The tradeoff is interpretability. Correlation-based methods produce directly interpretable regulatory scores, whereas deep generative models require explainability methods to compute regulatory scores for regulatory elements and transcription factors. Researchers should consider whether the biological question requires mechanistic interpretation or predictive accuracy.
Observational Versus Perturbation Data
Observational multi-omics data are easier to obtain and can be analyzed with in silico perturbation methods. However, in silico perturbation relies on the accuracy of the inferred network and cannot capture all biological feedback mechanisms. Experimental perturbation data provide direct evidence but are more expensive and technically challenging. The organoid study demonstrated that combining observational multi-omics with pooled genetic perturbation can identify transcription factors that regulate cell fate abundance versus those that affect cell states after differentiation.
Records and Measurements for Regulatory Inference
Essential Records
Maintain detailed records of data processing steps, including software versions, parameter settings, and filtering thresholds. The nf-core documentation describes community pipeline standards that support reproducible workflow configuration. The Carpentries lessons provide foundational training in computing, data, shell, Git, and programming that supports reproducible analysis practices.
Record the following for each analysis:
- Sequencing depth and quality metrics for each sample
- Number of cells retained after quality filtering
- Number of peaks called and retained after filtering
- Number of enhancer-gene links identified
- Number of transcription factor-target gene edges in the final network
- Causal inference scores for prioritized transcription factors
- Validation experiment results
Quality Metrics
Track quality metrics at each analysis stage. For scATAC-seq, monitor transcription start site enrichment and fraction of reads in peaks. For scRNA-seq, monitor mitochondrial read fraction and library complexity. For integrated data, assess batch effects and modality agreement. The EMBL-EBI Training provides learning pathways for data-resource training and practical analysis education that can help researchers select appropriate quality metrics.
Reproducibility Measures
Use version control for analysis code and document the computational environment. The Galaxy Training Network provides accessible workflow training that emphasizes reproducibility. Bioconductor packages support reproducible genomic-analysis workflows with versioned releases. The NCBI Data Resources provide official descriptions of databases and search systems that support data deposition and retrieval.
Common Failure Patterns in Regulatory Inference
Failure to Account for Cell Type Heterogeneity
Regulatory networks inferred across mixed cell types produce averaged signals that may not represent any actual cell population. The lung cancer study found that candidate cis-regulatory elements are largely cell type-specific, with 37 percent detected in only one cell type. Inferring networks within defined cell populations is essential for meaningful regulatory inference.
Overinterpretation of Correlation as Causation
The most common failure is treating accessibility-expression correlation as evidence of direct regulation. Correlation can arise from shared upstream signals, indirect effects, or technical artifacts. Causal inference methods should be applied before claiming regulatory relationships. The lung adenocarcinoma study explicitly noted that upstream computational inference of NF-kB and STAT signaling regulation should be regarded as hypothesis-generating exploration requiring further validation, while the primary causal evidence centered on functional validation of the CCL20-CCR6 axis.
Ignoring Transcription Factor Expression Context
Transcription factors can only regulate targets in cells where they are expressed. The retina study found that transcription factors whose binding sites are perturbed by genetic variants tend to have higher expression levels in the cell types where the variants exert their effects. Filtering transcription factors by expression in the relevant cell type reduces false positives.
Inadequate Motif Collections
Transcription factor identification depends on the quality and completeness of motif collections. SCENIC+ curated and clustered a motif collection with more than 30,000 motifs to improve recall and precision. Using outdated or incomplete motif collections can miss relevant transcription factors or produce false matches.
Failure to Validate Predictions
Computational predictions require experimental validation. The organoid study validated predictions with pooled genetic perturbation. The retina study used high-throughput reporter assays. The renal cell carcinoma study used preclinical models. Without validation, regulatory network predictions remain hypotheses.
Limitations of Current Approaches
Technical Noise and Dropout
Single-cell data contain substantial technical noise, including dropout events where genes are not detected despite expression. This noise can weaken correlations and produce spurious regulatory links. Deep generative models such as multiHIVE address this by facilitating integration, denoising, and imputation tasks.
Limited Temporal Resolution
Most single-cell multi-omics datasets represent a single time point. Regulatory dynamics occur over minutes to hours, as demonstrated by the FigR study showing that cells alter chromatin accessibility and gene expression at timescales of minutes. The cell cycle study combined single-cell multiome sequencing with biophysical modeling and deep learning to quantify rates of mRNA transcription, splicing, nuclear export, and degradation, revealing that transcriptional and post-transcriptional processes exhibit distinct oscillatory waves at specific cell cycle phases.
Genetic Variant Effects Are Context Dependent
Genetic effects on gene regulation are highly context dependent. The retina study found that cell type-dependent genetic effects are driven by precise modulation of both trans-factor expression and chromatin accessibility of cis-elements. Hierarchical collaboration among transcription factors plays a crucial role in mediating cell type-specific effects of genetic variants on gene regulation.
Incomplete Knowledge of Transcription Factor Biology
Transcription factor binding does not guarantee regulatory activity. Post-translational modifications, cofactor availability, and competitive binding all influence whether a bound transcription factor activates or represses target genes. Current computational methods cannot fully capture this complexity.
Safety and Regulatory Context
Data Deposition and Sharing
Deposit processed data and analysis code in public repositories to support reproducibility. The NCBI Data Resources provide official databases and search systems for sequence data deposition and retrieval. The EMBL-EBI Training provides guidance on data-resource training and practical analysis education.
Ethical Use of Human Data
Studies using human tissue must comply with institutional review board requirements and data protection regulations. The gynecologic malignancies study processed tumors immediately following surgical resection, requiring appropriate consent and ethical approval. The retina study used cells from 20 human donors, requiring documented consent and anonymization procedures.
Reporting Standards
Report analysis methods with sufficient detail for replication. Include software versions, parameter settings, and quality metrics. The nf-core documentation describes community pipeline standards that support reproducible workflow configuration. The Galaxy Training Network provides accessible workflow training that emphasizes reproducibility.
Professional Escalation Criteria
When to Seek Specialized Support
Escalate to bioinformatics core facilities or specialized consultants when:
- Data integration across modalities fails to produce concordant cell type annotations
- Quality metrics indicate systematic technical artifacts that standard filtering cannot resolve
- Causal inference methods produce conflicting results across approaches
- Validation experiments fail to confirm computational predictions
- The biological question requires custom analysis methods not available in standard packages
When to Reconsider the Experimental Design
Reconsider the experimental design when:
- The number of cells per cell type is too low for robust network inference
- The biological system lacks sufficient genetic variation for Mendelian randomization approaches
- Perturbation experiments are not feasible in the model system
- The temporal resolution of the data cannot capture the regulatory dynamics of interest
When to Consult Domain Experts
Consult domain experts in transcription factor biology, chromatin structure, or the specific disease or developmental system when:
- Prioritized transcription factors have no known role in the biological context
- Predicted regulatory relationships contradict established biology
- The biological interpretation of network modules requires specialized knowledge
A Decision Framework for Selecting Causal Inference Methods by Data Availability and Biological Question
Researchers frequently struggle to select the appropriate causal inference method for their specific dataset and biological question. The choice between in silico perturbation, Mendelian randomization, and experimental perturbation is not arbitrary. It depends on the data types already collected, the biological system under study, and the strength of evidence required for the intended publication or downstream application. This section provides a structured decision framework that maps data availability and biological questions to specific causal inference methods, along with a record system for tracking evidence strength and a troubleshooting approach for common method failures.
Data Availability as the Primary Decision Driver
The first decision point is what data already exist or can be feasibly collected. This constraint determines which causal inference methods are accessible. Three data availability scenarios cover most research situations.
Scenario A: Observational single-cell multi-omics only. When researchers have matched or computationally paired scRNA-seq and scATAC-seq data without genetic data or perturbation experiments, the available methods are correlation-based network inference followed by in silico perturbation. The FigR framework computationally pairs scATAC-seq with scRNA-seq cells, connects distal cis-regulatory elements to genes, and infers gene regulatory networks to identify candidate transcription factor regulators. SCENIC+ predicts genomic enhancers, identifies candidate upstream transcription factors, and links enhancers to candidate target genes within the same computational framework. CellOracle uses gene regulatory networks inferred from single-cell multi-omics data to perform in silico transcription factor perturbations, simulating consequent changes in cell identity using only unperturbed wild-type data.
The evidence strength from this scenario is hypothesis-generating. In silico perturbation relies on the accuracy of the inferred network and cannot capture all biological feedback mechanisms. The lung adenocarcinoma study explicitly noted that upstream computational inference of NF-kB and STAT signaling regulation should be regarded as hypothesis-generating exploration requiring further validation, while the primary causal evidence centered on functional validation of the CCL20-CCR6 axis. Researchers in this scenario should plan for experimental validation of prioritized transcription factors.
Scenario B: Single-cell multi-omics plus genetic data. When researchers have access to genome-wide association study summary statistics, expression quantitative trait loci, or protein quantitative trait loci data, summary-data-based Mendelian randomization becomes available. This approach uses genetic variants as instrumental variables to assess causal relationships between gene expression and disease outcomes. A multi-omics and epigenome-wide study of intracranial aneurysms applied summary-data-based Mendelian randomization using data from genome-wide association studies of gene expression from 31,684 European individuals and protein quantitative trait loci from 35,559 Icelanders, integrating these with aneurysm data from the International Stroke Genetics Consortium and the FinnGen study.
The advantage of this scenario is that genetic variants are randomly assigned at conception, reducing confounding that plagues observational correlation studies. However, Mendelian randomization requires that the genetic instruments satisfy specific assumptions, including relevance, independence, and exclusion restriction. Researchers must verify that the genetic variants are strongly associated with the exposure, are not associated with confounders, and affect the outcome only through the exposure. The retina study demonstrated how single-cell data can be integrated with genetic data by mapping eQTL, caQTL, allelic-specific expression, and allelic-specific chromatin accessibility in major retinal cell types, finding that the majority of identified single-cell eQTLs and caQTLs display cell type-specific effects.
Scenario C: Single-cell multi-omics plus perturbation capability. When researchers can perform pooled genetic perturbation with single-cell readout, they obtain the strongest causal evidence. The organoid study used this approach to assess transcription factor requirements for cell fate and state regulation in human brain organoids. The study found that certain factors regulate the abundance of cell fates, whereas other factors affect neuronal cell states after differentiation. By measuring transcriptome and chromatin accessibility in normal or GLI3-perturbed cells, the researchers identified two distinct GLI3 regulomes central to telencephalic fate decisions.
This scenario requires experimental infrastructure for CRISPR perturbation, single-cell readout, and appropriate model systems. The evidence strength is high because direct manipulation of transcription factor expression provides causal evidence. However, the cost and technical complexity are substantially higher than observational approaches.
Matching Methods to Biological Questions
Beyond data availability, the specific biological question determines which causal inference method is most appropriate. Four common question types map to distinct method preferences.
Question 1: Which transcription factors drive a cell state transition? For this question, in silico perturbation methods are well suited when only observational data are available. The palate development study used in silico perturbation analysis to identify transcription factors SHOX2 and MEOX2 as important regulators of the development of the anterior and posterior palate, respectively. The study profiled chromatin accessibility and gene expression simultaneously within the same cells from mouse secondary palate across embryonic days 12.5, 13.5, 14.0, and 14.5, constructing five trajectories representing continuous differentiation of cranial neural crest-derived multipotent cells into distinct lineages. When perturbation experiments are feasible, pooled genetic perturbation with single-cell readout provides direct evidence of which transcription factors regulate cell fate abundance versus cell states after differentiation.
Question 2: Do specific genes causally influence disease risk? For this question, summary-data-based Mendelian randomization is the preferred approach when genetic data are available. The intracranial aneurysm study identified tier 1 associations for RELT with an odds ratio of 0.14 and TNFSF12 with an odds ratio of 1.24, and tier 3 associations for ICAM5 with an odds ratio of 0.89 and ERAP2 with an odds ratio of 1.07. These effect sizes provide quantitative estimates of causal association between gene expression and disease risk. The lung cancer susceptibility study integrated single-cell data with genome-wide association study loci, using colocalization of candidate causal variants with candidate cis-regulatory elements combined with transcription factor footprinting to prioritize variants for 68 percent of the loci.
Question 3: What are the downstream targets of a specific transcription factor? For this question, correlation-based network inference combined with perturbation data provides the most complete answer. The organoid study identified two distinct GLI3 regulomes, one regulating dorsoventral patterning with HES4 and HES5 as direct GLI3 targets, and one controlling ganglionic eminence diversification later in development. The identification of direct targets required measuring transcriptome and chromatin accessibility in normal or GLI3-perturbed cells, demonstrating that perturbation data substantially improves target identification compared to observational data alone.
Question 4: Which regulatory elements are functional in specific cell types? For this question, transcription factor footprinting and allelic-specific analysis are most informative. The retina study mapped eQTL, caQTL, allelic-specific expression, and allelic-specific chromatin accessibility in major retinal cell types, identifying regulatory elements and genetic variants effective on gene regulation in individual cell types. The study found that transcription factors whose binding sites are perturbed by genetic variants tend to have higher expression levels in the cell types where the variants exert their effects. The lung cancer study used transcription factor footprinting to prioritize candidate causal variants, demonstrating that footprinting can distinguish between transcription factors that are merely expressed and those that are actively bound to regulatory elements.
A Structured Decision Matrix
The following decision matrix summarizes the mapping between data availability, biological question, and recommended methods. This matrix serves as a practical reference for planning analyses.
| Data Available | Biological Question | Recommended Method | Evidence Strength | Validation Required |
|---|---|---|---|---|
| Observational multi-omics only | Cell state transition drivers | In silico perturbation (CellOracle, SCENIC+) | Hypothesis-generating | High: experimental validation required |
| Observational multi-omics only | Transcription factor target identification | Correlation-based network inference (FigR, SCENIC+) | Hypothesis-generating | High: perturbation or reporter assays required |
| Multi-omics plus genetic data | Gene-disease causal relationships | Summary-data-based Mendelian randomization | Medium: genetic instruments reduce confounding | Medium: replication cohorts and functional follow-up |
| Multi-omics plus genetic data | Cell type-specific genetic effects | Single-cell eQTL and caQTL mapping | Medium: context-dependent effects | Medium: reporter assays for validation |
| Multi-omics plus perturbation capability | Transcription factor requirements | Pooled genetic perturbation with single-cell readout | High: direct experimental manipulation | Low: direct causal evidence |
| Multi-omics plus perturbation capability | Direct target genes of specific transcription factors | Perturbation followed by differential accessibility and expression analysis | High: direct experimental manipulation | Low: direct causal evidence |
Record System for Evidence Tracking
A structured record system helps researchers track the strength of evidence for each prioritized transcription factor and regulatory relationship. This system supports transparent reporting and prevents overinterpretation of computational predictions.
Transcription factor evidence log. For each transcription factor prioritized by any causal inference method, record the following fields:
- Transcription factor name and identifier
- Cell type and condition where prioritized
- Method that produced the prioritization
- Causal inference score or statistic
- Network centrality measures
- Expression level in the relevant cell type
- Motif enrichment in differentially accessible regions
- Known biological role from literature
- Validation status: unvalidated, in progress, or confirmed
- Validation method used
- Date of last update
Regulatory relationship evidence log. For each transcription factor-target gene edge in the inferred network, record:
- Transcription factor and target gene identifiers
- Evidence type: correlation, in silico perturbation, genetic instrument, or experimental perturbation
- Effect size and confidence interval where applicable
- Cell type specificity of the relationship
- Whether the relationship was validated experimentally
- Whether the relationship is consistent across multiple inference methods
Method comparison record. When multiple causal inference methods are applied to the same dataset, record the concordance between methods. The organoid study demonstrated that combining observational multi-omics with pooled genetic perturbation can identify transcription factors that regulate cell fate abundance versus those that affect cell states after differentiation. Disagreement between methods should trigger investigation into the assumptions and limitations of each method instead of automatic dismissal of either result.
Troubleshooting Common Method Failures
Failure pattern 1: In silico perturbation produces implausible predictions. When CellOracle or similar methods predict transcription factor effects that contradict established biology, first verify that the inferred gene regulatory network is accurate. Check whether the transcription factor is expressed in the relevant cell type. The retina study found that transcription factors whose binding sites are perturbed by genetic variants tend to have higher expression levels in the cell types where the variants exert their effects. Filtering transcription factors by expression in the relevant cell type reduces false positives. Next, verify that the motif collection used for transcription factor identification is current and complete. SCENIC+ curated and clustered a motif collection with more than 30,000 motifs to improve recall and precision. Outdated or incomplete motif collections can miss relevant transcription factors or produce false matches.
Failure pattern 2: Mendelian randomization results are inconsistent across replication cohorts. When summary-data-based Mendelian randomization produces associations that do not replicate, examine the genetic instruments used. Verify that the instruments satisfy the assumptions of relevance, independence, and exclusion restriction. Check for linkage disequilibrium between instruments and nearby variants that might violate the exclusion restriction. The intracranial aneurysm study used data from the International Stroke Genetics Consortium for discovery and the FinnGen study for replication, providing a model for replication-focused analysis. Consider whether the effect is cell type-specific and whether the relevant cell type is adequately represented in the single-cell data.
Failure pattern 3: Perturbation experiments do not confirm computational predictions. When experimental perturbation fails to reproduce predicted effects, first verify that the perturbation achieved sufficient knockdown or knockout efficiency. Check whether compensatory mechanisms might mask the effect. The organoid study found that certain factors regulate the abundance of cell fates, whereas other factors affect neuronal cell states after differentiation, suggesting that the readout measure must match the predicted effect type. If the prediction was for cell fate changes, measure cell type proportions. If the prediction was for cell state changes, measure gene expression programs within cell types. Consider whether the computational prediction captured a context-dependent effect that requires specific environmental conditions or developmental timing.
Failure pattern 4: Cell type heterogeneity obscures regulatory signals. When regulatory inference across mixed cell types produces averaged signals that do not represent any actual cell population, stratify the analysis by cell type. The lung cancer study found that candidate cis-regulatory elements are largely cell type-specific, with 37 percent detected in only one cell type. Inferring networks within defined cell populations is essential for meaningful regulatory inference. The multiHIVE approach addresses this by disentangling shared and modality-specific signals through hierarchically stacked latent variables, facilitating integration, denoising, and imputation tasks. The scTFBridge method disentangles latent spaces into shared and specific components across omics layers and integrates transcription factor motif binding knowledge to align shared embeddings with specific transcription factor regulatory activities.
Failure pattern 5: Temporal dynamics are missed in static datasets. When regulatory relationships change over time but the dataset represents a single time point, consider whether the biological question requires temporal resolution. The FigR study demonstrated that cells alter chromatin accessibility and gene expression at timescales of minutes. The cell cycle study combined single-cell multiome sequencing with biophysical modeling and deep learning to quantify rates of mRNA transcription, splicing, nuclear export, and degradation, revealing that transcriptional and post-transcriptional processes exhibit distinct oscillatory waves at specific cell cycle phases. If temporal dynamics are central to the biological question, design time course experiments instead of relying on static inference.
Integration with Existing Workflows
This decision framework integrates with the practical workflow described earlier in this article. After completing quality control, cell type annotation, and data integration, researchers should apply the decision matrix to select causal inference methods. The records and measurements section should be extended with the transcription factor evidence log and regulatory relationship evidence log described here. The common failure patterns section should be consulted when results are unexpected or inconsistent across methods.
The Bioconductor project provides official packages and workflows for reproducible genomic analysis, including single-cell quality control procedures and network inference methods. The Galaxy Training Network provides accessible workflow training for single-cell analysis that can support implementation of this decision framework. The nf-core documentation describes community pipeline standards that support reproducible workflow configuration. The Carpentries lessons provide foundational training in computing, data, shell, Git, and programming that supports reproducible analysis practices. The EMBL-EBI Training provides learning pathways for data-resource training and practical analysis education. The NCBI Data Resources provide official descriptions of databases and search systems that support data deposition and retrieval.
Escalation Criteria for Method Selection
Escalate to bioinformatics core facilities or specialized consultants when the decision framework does not produce a clear method recommendation. This situation arises when:
- The dataset combines features of multiple scenarios, such as partial genetic data or limited perturbation capability
- The biological question does not map cleanly to any single method
- Multiple methods produce conflicting results that cannot be resolved by the troubleshooting approaches described here
- The biological system has unusual features, such as extensive copy number variation or extreme cellular heterogeneity, that violate the assumptions of standard methods
Reconsider the experimental design when the available data cannot support the required evidence strength. If the biological question requires causal evidence but only observational data are available, plan for experimental validation of prioritized transcription factors. If Mendelian randomization is needed but genetic data are unavailable, consider whether public genome-wide association study data can be integrated with the single-cell data. If perturbation experiments are required but not feasible in the current model system, consider alternative model systems or validation approaches such as reporter assays.
Frequently Asked Questions
What is the difference between correlation-based and causal inference approaches in single-cell multi-omics?
Correlation-based approaches identify statistical associations between chromatin accessibility and gene expression across cells. Causal inference approaches add information that helps distinguish direct regulatory relationships from spurious correlations. In silico perturbation methods simulate the effect of removing a transcription factor from the network. Mendelian randomization uses genetic variants as instrumental variables. Experimental perturbation with single-cell readout provides direct evidence of transcription factor requirements. Each approach has different assumptions and provides different levels of evidence for causality.
How do I choose between matched multiome and separate scRNA-seq and scATAC-seq assays?
Matched multiome data from the same cell provide the strongest link between accessibility and expression and are preferred when the biological question requires single-cell resolution of the regulatory relationship. Separate assays are less expensive per cell and allow optimization of each assay independently, but require computational pairing of cells, which introduces uncertainty. Consider the cell type complexity of the system, the availability of established computational pairing methods, and the budget constraints of the study.
What quality control metrics are most important for regulatory network inference?
For scATAC-seq, monitor transcription start site enrichment and fraction of reads in peaks to assess chromatin accessibility data quality. For scRNA-seq, monitor mitochondrial read fraction and library complexity to assess cell viability and capture efficiency. After integration, assess batch effects and modality agreement. Document all filtering thresholds and justify them based on data distributions. Poor quality data at any stage propagates errors through the network inference pipeline.
How many cells do I need for reliable regulatory network inference?
The required number of cells depends on the number of cell types in the sample, the depth of sequencing, and the complexity of the regulatory programs being studied. Rare cell types require more total cells to achieve adequate representation. Network inference methods vary in their sensitivity to cell number. Assess the stability of inferred networks by subsampling cells and checking whether prioritized transcription factors remain consistent.
What is the role of transcription factor motif collections in regulatory inference?
Motif collections define the sequence patterns that transcription factors are predicted to bind. The quality and completeness of the motif collection directly affect the accuracy of transcription factor identification. SCENIC+ curated and clustered a motif collection with more than 30,000 motifs to improve recall and precision. Document which motif collection version was used and how motif matching thresholds were set.
How do I validate computational predictions of regulatory relationships?
Validation options include genetic perturbation with single-cell readout, reporter assays for enhancer activity, chromatin immunoprecipitation followed by sequencing, and preclinical models. The organoid study used pooled genetic perturbation to assess transcription factor requirements. The retina study used high-throughput reporter assays. The renal cell carcinoma study used preclinical models. Choose validation methods appropriate to the biological system and the strength of the computational prediction.
What are the limitations of in silico perturbation methods?
In silico perturbation methods rely on the accuracy of the inferred gene regulatory network. They cannot capture all biological feedback mechanisms, post-translational modifications, or context-dependent effects. CellOracle correctly modeled reported changes in phenotype for well-established paradigms, but predictions for less characterized systems require experimental validation. In silico perturbation is best used to prioritize transcription factors for experimental testing instead of as a substitute for perturbation experiments.
How do genetic variants help establish causal regulatory relationships?
Genetic variants that affect gene expression or chromatin accessibility can be used as instrumental variables in Mendelian randomization approaches. Because alleles are randomly assigned at conception, genetic associations are less susceptible to confounding than observational correlations. The intracranial aneurysm study used summary-data-based Mendelian randomization to identify immune-inflammation targets. The retina study mapped eQTL and caQTL in major retinal cell types to identify regulatory elements and genetic variants effective on gene regulation in individual cell types.
Related Bioinformatics Guides
- Single-Cell Multi-Omics Integration: Methods and Applications
- Multi-Omics Integration and Network Analysis: Uncovering Biological Relationships
- Single-Cell Genomics: From Concept to Application
- Single-Cell Annotation: A Workflow for Cell Type Identification
- Single-Cell Isolation Techniques: A Practical Comparison
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- SCENIC+: single-cell multiomic inference of enhancers and gene regulatory networks.. Nature methods, 2023.
- Functional inference of gene regulation using single-cell multi-omics.. Cell genomics, 2022.
- A multi-omic single-cell landscape of human gynecologic malignancies.. Molecular cell, 2021.
- Dissecting cell identity via network inference and in silico gene perturbation.. Nature, 2023.
- Single-cell multi-omics reveals that FABP1 + renal cell carcinoma drive tumor angiogenesis through the PLG-PLAT axis under fatty acid reprogramming.. Molecular cancer, 2025.
- Inferring and perturbing cell fate regulomes in human brain organoids.. Nature, 2023.
- Integrative multi-omics analysis reveals a novel subtype of hepatocellular carcinoma with biological and clinical relevance.. Frontiers in immunology, 2024.
- Single-cell multi-omics reveals DUSP9 as a key regulator of cancer stemness and a potential therapeutic target in hepatocellular carcinoma.. Journal of translational medicine, 2026.
- Integrating multi-omics and machine learning to uncover CCL20 as a potential regulator of the immunosuppressive microenvironment in lung adenocarcinoma.. 2026.
- multiHIVE: Hierarchical Multimodal Deep Generative Modeling for Single-cell Multiomics. 2026.
- scTFBridge: a disentangled deep generative model informed by TF-motif binding for gene regulation inference in single-cell multi-omics. bioRxiv, 2025.
- Identification of immune-inflammation targets for intracranial aneurysms: a multiomics and epigenome-wide study integrating summary-data-based Mendelian randomization, single-cell-type expression analysis, and DNA methylation regulation. International Journal of Surgery, 2024.
- Single-cell multiomics reveals the oscillatory dynamics of mRNA metabolism and chromatin accessibility during the cell cycle.. Cell Reports, 2025.
- Single-cell multiomics of the human retina reveals hierarchical transcription factor collaboration in mediating cell type-specific effects of genetic variants on gene regulation. Genome Biology, 2023.
- Context-aware single-cell multiomics approach identifies cell-type-specific lung cancer susceptibility genes. Nature Communications, 2024.
- Single-cell multiomics decodes regulatory programs for mouse secondary palate development. Nature Communications, 2024.
- A landscape of gene expression regulation for synovium in arthritis. Nature Communications, 2024.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.