# Integrating Single-Cell Multi-Omics for Gene Regulatory Inference: From ATAC + RNA to Regulatory Networks


## Key Takeaways

- Integrating single-cell ATAC sequencing (scATAC-seq) with single-cell RNA sequencing (scRNA-seq) is crucial for inferring gene regulatory networks (GRNs) by linking transcription factor (TF) binding opportunities (chromatin accessibility) to transcriptional output (gene expression).
- True paired multi-omics data, where accessibility and expression are measured from the same cell (e.g., 10x Multiome), offers superior resolution for enhancer-to-gene linkage compared to integrating independently collected datasets, which requires computational alignment of cell clusters.
- TF activity inference moves beyond motif presence in accessible regions by integrating TF expression levels and correlating motif enrichment in open chromatin with target gene upregulation, with tools like SCENIC+ and FigR employing distinct strategies for this.
- Enhancer-to-gene linkage relies on correlation between peak accessibility and gene expression, often refined by distance constraints (e.g., 100-500 kb windows) and motif presence to reduce false positives and identify potential regulatory hubs like Domains of Regulatory Chromatin (DORCs).
- Validation of inferred GRNs is essential, employing computational methods (e.g., cross-validation, comparison with ChIP-seq), in silico perturbations (e.g., CellOracle), and ultimately experimental validation (e.g., TF knockdown, CRISPR editing) to confirm causal regulatory relationships.

---

Single-cell multi-omics data, particularly the joint measurement of chromatin accessibility via single-cell ATAC sequencing (scATAC-seq) and gene expression via single-cell RNA sequencing (scRNA-seq), enables researchers to infer gene regulatory networks (GRNs) that link enhancers, transcription factors (TFs), and target genes. This article provides a practical workflow for researchers who need to move from raw multi-omics data to interpretable regulatory networks using tools such as SCENIC+, FigR, and related frameworks. The focus is on concrete decisions at each step: data input preparation, quality control thresholds, integration strategy, enhancer-to-gene linkage, TF activity inference, network construction, and validation. The guidance applies to laboratory professionals and bioinformatics practitioners working with human, mouse, or other model organism datasets where paired accessibility and expression profiles are available.

## The Core Problem: Why RNA Alone Cannot Define Regulatory Networks

Gene expression measurements from scRNA-seq reveal which genes are active in a cell, but they do not explain why those genes are active. Transcriptional regulation depends on TF binding to cis-regulatory elements such as enhancers and promoters, and this binding is conditioned on chromatin accessibility. A TF can only engage its target motif when the surrounding chromatin is open. Therefore, expression data alone cannot distinguish between a TF that is present but inactive and a TF that is actively driving transcription. Chromatin accessibility data from scATAC-seq provides the missing layer: it marks the genomic regions where regulatory proteins can bind. When accessibility and expression are measured in the same cell type or ideally the same cell, the combined signal reveals which TFs have both the opportunity (open chromatin) and the output (target gene expression) to exert regulatory control.

The practical consequence is that GRN inference from expression data alone produces high false-positive rates. Many predicted TF-target relationships are correlations that reflect shared upstream regulation instead of direct physical interaction. Integrating accessibility data filters these correlations by requiring that a candidate TF has a motif in an accessible region near the target gene. This reduces the search space and increases the likelihood that inferred edges represent genuine regulatory connections.

## Data Inputs: What You Need Before Starting

### Paired Multi-Omics Data

The ideal input is a single-cell multiome dataset where both chromatin accessibility and gene expression are measured from the same cell. Commercial platforms such as the 10x Multiome ATAC + Gene Expression kit produce this paired data directly. The advantage of true paired data is that accessibility and expression are matched at the single-cell level, eliminating the need for computational pairing of separately collected datasets. Tools such as FigR were developed with this paired data structure in mind, using the matched cells to build a joint representation before linking regulatory elements to genes.

When paired data are not available, researchers can integrate independently collected scATAC-seq and scRNA-seq datasets from the same tissue or condition. This requires computational alignment to identify corresponding cell types across modalities. The integration is less precise than true paired data because the linkage between accessibility and expression is established at the level of cell clusters instead of individual cells. However, this approach remains viable when multiome data cannot be generated due to cost, sample availability, or technical constraints.

### Reference Genome and Annotations

A reference genome assembly and gene annotation file are required for both alignment and downstream analysis. The choice of reference genome version affects motif annotation and enhancer identification. For human data, the GRCh38 assembly is standard. For mouse, GRCm39 is the current reference. The [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/) provide official access to genome assemblies, gene annotations, and sequence databases that support these analyses. Researchers should verify that their reference genome version matches the one used by their alignment and peak-calling tools to avoid coordinate mismatches.

### TF Motif Collections

GRN inference tools rely on position weight matrices (PWMs) that describe the DNA binding preferences of TFs. The quality and coverage of the motif collection directly affect TF identification accuracy. SCENIC+ curated and clustered a motif collection containing more than 30,000 motifs to improve both recall and precision of TF identification across diverse species and cell types. When selecting a motif database, consider whether it covers the species in your dataset and whether it distinguishes between closely related TF family members. Motif collections with poor coverage of a particular TF family will systematically miss those regulators in the inferred network.

## Quality Control Before Integration

### scRNA-seq Quality Control

The expression modality requires filtering for cell viability and capture quality. Standard metrics include the number of unique molecular identifiers (UMIs), the number of detected genes, and the fraction of mitochondrial reads. Cells with very low UMI counts likely represent empty droplets or damaged cells. Cells with very high mitochondrial fractions often indicate stressed or dying cells where cytoplasmic RNA has been lost. The specific thresholds depend on tissue type and dissociation protocol, so examine the distribution of these metrics across your dataset instead of applying fixed cutoffs from other studies.

Doublet detection is another essential step. Doublets are droplets that captured two cells, producing a mixed expression profile that can create spurious cell states. Computational doublet detection methods identify these profiles by comparing each cell to synthetic doublets generated from the dataset. Removing doublets before integration prevents the creation of artificial intermediate cell states that would distort regulatory network inference.

### scATAC-seq Quality Control

The accessibility modality requires filtering based on the number of unique fragments, the fraction of fragments in peaks (FRiP), and the transcription start site (TSS) enrichment score. Cells with low fragment counts have sparse accessibility profiles that may not reliably represent the regulatory landscape. FRiP measures the proportion of fragments that fall within called peaks, with higher values indicating better signal-to-noise. TSS enrichment assesses whether accessibility is concentrated at gene promoters, which is expected for healthy cells. Cells with poor TSS enrichment may represent debris or nuclei with degraded chromatin.

Peak calling is performed on the aggregated accessibility data after cell filtering. The resulting peak set defines the candidate cis-regulatory elements for the analysis. Peak calling can be performed on the full dataset or on per-cluster subsets to capture cell-type-specific regulatory elements. The latter approach is more sensitive for rare cell types whose peaks would be diluted in a global call.

### Integration Quality Checks

After filtering both modalities, verify that the cell types present in the expression data match those in the accessibility data. This is particularly important when integrating separately collected datasets. Use label transfer or clustering comparison to confirm that each cell type is represented in both modalities. Mismatches indicate batch effects, differential cell survival during processing, or annotation errors. Resolve these issues before proceeding to GRN inference because the network will inherit any systematic bias in cell type representation.

## Integration Strategies: Pairing Cells Across Modalities

### True Paired Data

For multiome data where accessibility and expression come from the same cell, integration is straightforward. The two modalities are linked by cell barcode. Downstream tools can use the paired structure directly. FigR was designed for this scenario, using the paired cells to build a joint representation that captures both regulatory potential and transcriptional output.

### Unpaired Data Integration

When integrating separate scATAC-seq and scRNA-seq datasets, the goal is to identify corresponding cells or cell states across modalities. Several strategies exist. One approach uses shared regulatory features, such as gene activity scores computed from accessibility at gene promoters, to anchor the two datasets. Another approach uses cell type labels from independent clustering of each modality and then matches clusters based on marker gene expression and marker peak accessibility.

The quality of integration determines the reliability of downstream enhancer-to-gene links. Poor integration creates mismatched cell states where accessibility from one cell type is paired with expression from a different cell type. This produces regulatory links that do not exist in any actual cell. Validate integration by checking that known marker genes show concordant expression and accessibility in the matched cells.

## Enhancer-to-Gene Linkage: Connecting Regulatory Elements to Target Genes

### The Linkage Problem

Enhancers can act over large genomic distances, sometimes hundreds of kilobases from their target promoters. Determining which gene an enhancer regulates is a central challenge in GRN inference. The linkage problem is to assign each accessible peak to one or more target genes based on evidence that the peak regulates that gene's expression.

### Correlation-Based Linkage

The most common approach uses correlation between accessibility at a peak and expression of nearby genes across cells. If a peak is accessible in the same cells where a candidate gene is expressed, the peak is linked to that gene. FigR uses this approach, connecting distal cis-regulatory elements to genes based on the coordinated variation of accessibility and expression across the cell population. The correlation is computed after accounting for the overall accessibility and expression levels to reduce bias from highly active cells.

### Distance and Motif Constraints

Correlation alone produces many spurious links because nearby genes often show correlated expression due to shared regulation. Adding distance constraints reduces false positives by requiring that a peak be within a certain genomic window of the candidate gene. The window size is a key parameter. Too small a window misses distal enhancers. Too large a window includes unrelated genes. Typical windows range from 100 to 500 kilobases, but the optimal size depends on the locus and cell type.

Motif information provides an additional filter. If a peak is linked to a gene, the TF that binds that peak should have a motif within the peak sequence. SCENIC+ uses this logic by predicting genomic enhancers, identifying candidate upstream TFs based on motif presence, and then linking enhancers to candidate target genes. The motif filter ensures that each regulatory edge has a plausible molecular mechanism.

### DORCs as Regulatory Hubs

FigR introduces the concept of domains of regulatory chromatin (DORCs), which are genomic regions where accessibility is coordinated across cells and correlates with the expression of nearby genes. DORCs represent clusters of co-accessible peaks that likely function as regulatory hubs for a target gene or gene neighborhood. In the study of immune stimulation using FigR, DORCs revealed that cells alter chromatin accessibility and gene expression at timescales of minutes after stimulation, demonstrating that regulatory changes can be rapid and coordinated. Identifying DORCs in your data can highlight the most informative regulatory regions for downstream analysis.

## Transcription Factor Activity Inference

### From Motif Presence to Regulatory Activity

Motif presence in an accessible peak indicates that a TF could bind, but it does not prove that the TF is active. TF activity depends on expression of the TF itself, nuclear localization, post-translational modifications, and the availability of cofactors. GRN inference tools estimate TF activity by combining motif information with expression data. A TF is inferred to be active when its motif is enriched in accessible peaks near genes that are upregulated in the same cells where the TF is expressed.

### SCENIC+ Approach

SCENIC+ builds enhancer-driven GRNs by integrating accessibility and expression. The method first identifies candidate enhancers from the accessibility data, then assigns TFs to those enhancers based on motif enrichment, and finally links enhancers to target genes. The curated motif collection with more than 30,000 motifs improves the resolution of TF identification, allowing the method to distinguish between TFs with similar binding preferences. SCENIC+ has been benchmarked on human peripheral blood mononuclear cells, ENCODE cell lines, melanoma cell states, and Drosophila retinal development, demonstrating applicability across species and tissue types.

### FigR Approach

FigR uses a different strategy. It first pairs scATAC-seq with scRNA-seq cells, then connects distal cis-regulatory elements to genes, and finally infers GRNs to identify candidate TF regulators. The method was applied to resting and stimulated human blood cells, generating approximately 91,000 single-cell profiles that probed the cis-regulatory landscape of the immunological response across cell types, stimuli, and time. The resulting stimulation GRN elucidated TF activity at disease-associated DORCs, showing how the method can connect regulatory inference to disease-relevant biology.

### Comparing TF Activity Across Conditions

Once TF activity is inferred, compare activity between conditions or cell types to identify regulators that distinguish states. For example, in the context of immune stimulation, TFs whose activity increases rapidly after stimulation are candidate drivers of the early response. In disease contexts, TFs with elevated activity in specific cell subpopulations may represent therapeutic targets. The comparison should be performed at the level of cell clusters or trajectories instead of individual cells to reduce noise.

## Network Construction and Visualization

### Building the Regulatory Network

The output of enhancer-to-gene linkage and TF activity inference is a set of regulatory edges: TF to enhancer, enhancer to gene, and TF to gene. These edges are assembled into a network where nodes represent TFs, enhancers, and genes, and edges represent regulatory relationships. The network can be filtered based on the strength of the underlying evidence, such as correlation coefficients, motif enrichment scores, or activity estimates.

### Cell-Type-Specific Networks

GRNs are context specific. The regulatory relationships that operate in one cell type may not operate in another, even within the same tissue. Tools such as scGATE address this by inferring context-specific TF-gene interaction networks and reconstructing Boolean logic gates that describe how combinations of TFs regulate target genes. The method uses scATAC-seq data and TF DNA binding motifs to filter out non-relevant TFs, then integrates single-cell clustering to infer networks specific to each cell state. This context specificity is essential for understanding how the same TF can have different effects in different cell types.

### Trajectory and Dynamic Networks

Cellular differentiation and disease progression involve changes in regulatory networks over time. Tools such as scMTNI leverage cellular trajectory information to infer dynamic GRNs that capture how regulatory relationships change along a developmental or disease trajectory. The method uses a multi-task learning framework to define cell-type-specific GRNs and examine their dynamics. Applying scMTNI to a cellular reprogramming dataset identified key regulators of cellular fate transitions, demonstrating the value of trajectory-aware network inference.

### Visualization Approaches

Network visualization should prioritize interpretability. For small networks with fewer than 100 nodes, standard graph layouts with node coloring by cell type or condition are effective. For larger networks, focus on subnetworks around TFs of interest or on DORCs that show condition-specific changes. Tools such as scMEGA provide integrated visualization capabilities that link the network to the underlying multi-omics data, allowing researchers to examine the accessibility and expression evidence for each regulatory edge.

## Validation Strategies: Confirming Inferred Regulatory Relationships

### Computational Validation

Several computational approaches can validate inferred GRNs. Cross-validation splits the data into training and test sets to assess whether regulatory edges generalize. Comparison with external datasets, such as ChIP-seq data for specific TFs, checks whether inferred TF binding sites overlap experimentally determined binding locations. Motif enrichment analysis confirms that the TFs identified as regulators are enriched for their known binding motifs in the relevant peaks.

### In Silico Perturbation

CellOracle uses gene regulatory networks inferred from single-cell multi-omics data to perform in silico transcription factor perturbations, simulating the consequent changes in cell identity using only unperturbed wild-type data. The method was applied to mouse and human haematopoiesis and zebrafish embryogenesis, correctly modeling reported changes in phenotype that occur as a result of TF perturbation. In the developing zebrafish, CellOracle simulated and experimentally validated a previously unreported phenotype resulting from the loss of noto, an established notochord regulator, and identified lhx1a as an axial mesoderm regulator. This approach allows researchers to prioritize TFs for experimental validation by predicting the phenotypic consequences of their perturbation.

### Experimental Validation

The gold standard for validating GRN predictions is experimental perturbation. Knockdown or overexpression of a predicted TF should produce changes in target gene expression consistent with the inferred regulatory direction. CRISPR-based approaches can validate enhancer-to-gene links by deleting or disrupting a predicted enhancer and measuring the effect on target gene expression. These experiments are resource intensive, so prioritize validation for TFs and enhancers that are central to the biological question and that show strong, consistent evidence across multiple computational metrics.

### Limitations of Validation

Computational validation cannot prove causality. Correlation-based enhancer-to-gene links may reflect indirect effects, and motif-based TF assignment does not account for competitive binding or chromatin context. Experimental validation addresses these limitations but is limited to a small number of predictions. The combination of multiple computational methods with targeted experimental validation provides the strongest evidence for regulatory relationships.

## At a Glance: Tool Selection for GRN Inference

| Tool | Input Data | Key Features | Best Use Case |
|------|------------|--------------|---------------|
| SCENIC+ | Paired scATAC + scRNA or separate datasets | Curated motif collection with 30,000+ motifs, enhancer prediction, TF identification, target gene linkage | Cross-species analysis, cell type comparison, differentiation trajectories |
| FigR | Paired scATAC + scRNA | DORC identification, distal element-to-gene linkage, stimulation response analysis | Immune response studies, disease-associated regulatory regions |
| CellOracle | Paired or integrated multi-omics | In silico TF perturbation, cell identity simulation | Prioritizing TFs for experimental validation, developmental biology |
| scGATE | scRNA-seq with scATAC-seq filtering | Boolean logic gate inference, context-specific networks | Understanding combinatorial TF regulation, reducing false positives |
| scMEGA | Paired multi-omics | End-to-end analysis, trajectory integration, visualization | Dynamic processes such as differentiation and disease remodeling |
| scMTNI | Paired multi-omics with trajectory | Multi-task learning, dynamic GRN inference | Cellular reprogramming, developmental trajectories |
| scTFBridge | Paired multi-omics | Deep generative model, TF-motif informed, disentangled latent spaces | Cell-type-specific susceptibility genes, cis and trans regulation |

## Practical Workflow: From Raw Data to Regulatory Networks

### Step 1: Data Acquisition and Preprocessing

Obtain paired multi-omics data or separately collected scATAC-seq and scRNA-seq data from the same biological system. Process each modality independently using established pipelines. For scRNA-seq, align reads to the reference genome, generate a count matrix, and perform quality control. For scATAC-seq, align reads, call peaks, and generate a fragment matrix. The [Galaxy Training Network](https://training.galaxyproject.org/) provides accessible workflow training for these preprocessing steps, and [Bioconductor](https://bioconductor.org/) offers packages for reproducible genomic analysis.

### Step 2: Quality Control and Cell Filtering

Apply quality control filters to both modalities. For scRNA-seq, filter cells based on UMI counts, gene counts, and mitochondrial fraction. For scATAC-seq, filter cells based on fragment counts, FRiP, and TSS enrichment. Remove doublets from both modalities. Document the filtering thresholds and the number of cells removed at each step.

### Step 3: Integration and Cell Type Annotation

Integrate the two modalities using the approach appropriate for your data structure. For paired data, use the cell barcodes to link modalities. For unpaired data, use computational integration to match cell states. Annotate cell types using marker genes and marker peaks. Verify that the annotation is consistent across modalities.

### Step 4: Enhancer-to-Gene Linkage

Identify candidate enhancers from the accessibility data. Link enhancers to target genes using correlation-based approaches with distance and motif constraints. Tools such as FigR and SCENIC+ implement these steps with different parameterizations. Document the linkage parameters and the number of enhancer-to-gene links generated.

### Step 5: TF Activity Inference

Infer TF activity by combining motif information with expression and accessibility data. Use the curated motif collection appropriate for your species. Compare TF activity across cell types and conditions to identify candidate regulators.

### Step 6: Network Construction and Filtering

Assemble the regulatory network from the enhancer-to-gene links and TF activity estimates. Filter edges based on evidence strength. Construct cell-type-specific networks if the data support distinct regulatory states.

### Step 7: Validation and Interpretation

Validate the network using computational approaches such as cross-validation and comparison with external data. Prioritize TFs for experimental validation using in silico perturbation where available. Interpret the network in the context of the biological question, focusing on regulators that distinguish conditions or cell types.

## Records and Measurements: What to Document

### Data Provenance

Record the source of each dataset, including the platform, library preparation kit, sequencing depth, and reference genome version. This information is essential for reproducibility and for troubleshooting downstream issues.

### Quality Control Metrics

Document the distribution of quality metrics before and after filtering. Include the number of cells passing each filter, the median UMI count, the median fragment count, and the fraction of cells removed. These records allow comparison across batches and experiments.

### Integration Parameters

Record the integration method, the anchor features used, and the number of cells matched across modalities. Document any cell types that were poorly matched or excluded from the analysis.

### Linkage and Network Parameters

Document the distance window for enhancer-to-gene linkage, the correlation threshold, the motif collection version, and the number of edges in the final network. These parameters directly affect the network structure and must be reported for interpretation.

### Analysis Code and Versions

Maintain version-controlled analysis code that reproduces the entire workflow. Record the versions of all software packages and the computational environment. The [nf-core Documentation](https://nf-co.re/docs) provides standards for reproducible workflow configuration, and [The Carpentries Lessons](https://carpentries.org/lessons) offer foundational training in version control and reproducible computing.

## Common Failure Patterns and Troubleshooting

### Poor Integration Between Modalities

When separately collected scATAC-seq and scRNA-seq datasets fail to integrate, the most common causes are batch effects, differences in cell type composition, or annotation errors. Check whether the same cell types are present in both datasets and whether marker genes show concordant patterns. Consider using more stringent integration parameters or collecting paired multi-omics data.

### Excessive or Insufficient Enhancer-to-Gene Links

Too many links usually result from a large distance window or a low correlation threshold. Too few links result from restrictive parameters or sparse accessibility data. Examine the distribution of link distances and correlation scores to select appropriate thresholds. The optimal parameters depend on the data quality and the biological question.

### TF Activity Estimates That Do Not Match Known Biology

If inferred TF activity contradicts established knowledge, check the motif collection coverage for the relevant TF family. Some TF families have similar motifs that are difficult to distinguish, leading to misassignment. Also verify that the TF is expressed in the cells where high activity is inferred. A TF cannot be active if it is not expressed.

### Networks Dominated by a Few Hub TFs

Some datasets produce networks where a small number of TFs dominate the regulatory landscape. This can reflect genuine biology, such as master regulators, or technical artifacts from library size differences or batch effects. Check whether the hub TFs are known regulators in the system and whether their dominance persists after normalization.

### Reproducibility Issues Across Runs

If the same analysis produces different networks across runs, the cause is likely stochastic elements in the analysis, such as clustering initialization or subsampling. Set random seeds and document them. Use stable clustering methods and verify that the network structure is robust to parameter perturbations.

## Limitations of Current Approaches

### Correlation Does Not Equal Causation

The fundamental limitation of GRN inference from multi-omics data is that the methods infer regulatory relationships from correlations between accessibility and expression. These correlations can arise from indirect effects, shared upstream regulation, or technical artifacts. Experimental validation is required to confirm causal relationships.

### Motif-Based TF Assignment Is Approximate

TF binding preferences are represented by position weight matrices that capture the average binding preference of a TF. Actual binding depends on chromatin context, cooperative interactions, and competitive binding with other TFs. Motif-based assignment cannot capture these complexities, leading to both false positives and false negatives.

### Enhancer-to-Gene Linkage Remains Uncertain

The rules that determine which enhancer regulates which gene are not fully understood. Correlation-based linkage assumes that regulatory interactions produce coordinated variation in accessibility and expression, but this assumption may fail for genes with complex regulatory architectures or for enhancers that act over very long distances.

### Context Specificity Limits Generalization

GRNs inferred from one cell type or condition may not apply to other contexts. The same TF can have different targets in different cell types, and the same enhancer can regulate different genes depending on the chromatin state. Networks must be interpreted within the specific context where they were inferred.

### Data Quality Constraints

The quality of GRN inference depends on the quality of the underlying data. Sparse accessibility data limits the detection of distal enhancers. Low sequencing depth limits the resolution of TF activity estimates. Batch effects can create spurious regulatory differences between conditions. These constraints must be considered when interpreting network results.

## Applications in Disease Research

### Cancer Subtyping and Prognosis

Integrative multi-omics analysis has identified novel cancer subtypes with distinct regulatory programs. In hepatocellular carcinoma, an integrative analysis combining bulk RNA sequencing, proteomics, scRNA-seq, spatial transcriptomics, and genome sequencing constructed a tumor purity and tumor microenvironment prognostic risk model. The analysis identified a novel XPO1+ epithelial subtype that expresses signatures of the high-risk subtype and is influenced primarily by fibroblasts through ligand-receptor interactions. This subtype interacts with monocytes and macrophages, T and NK cells, and endothelial cells through specific ligand-receptor pairs, demonstrating how regulatory network inference can connect molecular subtypes to the tumor microenvironment.

### Identifying Therapeutic Targets

GRN inference can identify regulators that drive disease-relevant cell states. In hepatocellular carcinoma, DUSP9 was identified as a potential regulator enriched in a malignant subpopulation characterized by elevated ERK activation, oxidative metabolism, and progenitor-like features. Co-expression and regulatory network analyses revealed a DUSP9-specific module enriched for stemness and invasion-related genes, with FOXO3 identified as a core downstream effector. Functional assays confirmed that DUSP9 enhances stemness through the ERK-FOXO3 axis, and a candidate compound inhibited HCC cell proliferation, migration, and sphere formation. This workflow demonstrates how network inference can connect computational predictions to functional validation and therapeutic discovery.

### Understanding the Tumor Microenvironment

In lung adenocarcinoma, single-cell RNA sequencing, spatial transcriptomics, and multi-omics analysis identified a metastasis-enriched malignant epithelial subpopulation characterized by a pro-inflammatory signature and high CCL20 expression. Spatial transcriptomics and cellular communication inference demonstrated that CCL20-high tumor cells co-localized with and actively recruited regulatory T cells via the CCL20-CCR6 ligand-receptor pair. The study noted that the upstream computational inference of NF-kB and STAT signaling regulation should be regarded as hypothesis-generating exploration requiring further validation, illustrating the appropriate interpretation of computational predictions.

### Neurodevelopmental Disorders

Integrative single-cell multi-omics frameworks that combine scRNA-seq, scATAC-seq, and single-cell DNA methylation data can construct cell-type-specific epigenetic regulatory networks. In neurodevelopmental disorder research, this approach identified SOX11 and CHD8 as central multi-layer master regulators whose disrupted activity contributes to aberrant lineage specification. The integrative modeling demonstrated high enhancer-gene linkage accuracy and consistent cross-species regulatory conservation, showing how multi-omics integration can reveal convergent dysregulation across molecular layers.

## Spatial Context and Emerging Approaches

### Integrating Spatial Transcriptomics

Spatial transcriptomics maps gene expression within tissue context, providing information that single-cell data lack. Methods such as ISON integrate spatial transcriptomics with single-cell multi-omics data to predict chromatin accessibility profiles for spatial spots and reconstruct spatially resolved gene regulatory networks. The method captures patterns consistent with cis and trans regulatory information and enables estimation of TF activity at the spot level, distinguishing between TFs even within the same family. Application to Alzheimer's disease data revealed disease and age-specific spatially variable gene regulatory modules, highlighting the potential to uncover spatially organized mechanisms driving complex biological processes.

### Generative Models for Multi-Omics

Deep generative models offer new approaches to multi-omics integration. scDiffusion-X uses a latent diffusion model with a dual-cross-attention module to integrate, generate, and translate multi-omics data. The model can predict one molecular modality from another with uncertainty quantification and can infer cell-type-specific heterogeneous GRNs through gradient-based interpretation. scTFBridge uses a disentangled deep generative model informed by TF-motif binding to compute regulatory scores for regulatory elements and TFs, enabling robust GRN inference that outperforms baseline methods in both cis and trans regulation inference tasks.

### Computational Blueprints for Cell Fate Programming

The integration of single-cell and spatial omics with perturbation screens and deep learning is expanding the predictive scope of regulatory inference. These approaches are being synthesized as computational blueprints embedded in an iterative design-test-learn pipeline for cell fate programming. The practical adoption of these methods for protocol design remains uneven, but the trajectory is toward increasingly predictive models of regulatory control.

## Professional Escalation Criteria

### When to Seek Expert Consultation

Consult a bioinformatics specialist or computational biologist when the analysis requires advanced statistical methods, when integration between modalities fails despite standard approaches, or when the biological interpretation of the network requires domain expertise beyond your current knowledge. The [EMBL-EBI Training](https://www.ebi.ac.uk/training) portal provides learning pathways for bioinformatics data resources and practical analysis education that can build the necessary skills.

### When to Reconsider Data Collection

If the quality control metrics indicate poor data quality, such as low fragment counts in scATAC-seq or high doublet rates in scRNA-seq, consider recollecting the data with optimized protocols. If the cell type composition is inadequate for the biological question, such as missing a rare cell type of interest, consider enrichment strategies or increased sequencing depth.

### When to Question the Network Results

If the inferred network contradicts established biology, if the network structure is highly sensitive to parameter choices, or if the network cannot be validated with external data, question the results before proceeding to experimental validation. Re-examine the quality control, integration, and linkage parameters. Consider whether the data support the level of resolution required for the biological question.

## Frequently Asked Questions

### What is the difference between SCENIC+ and FigR for GRN inference?

SCENIC+ and FigR use different strategies to integrate accessibility and expression data. SCENIC+ predicts genomic enhancers, identifies candidate upstream TFs based on a curated motif collection with more than 30,000 motifs, and links enhancers to candidate target genes. FigR computationally pairs scATAC-seq with scRNA-seq cells, connects distal cis-regulatory elements to genes, and infers GRNs to identify candidate TF regulators. FigR also identifies domains of regulatory chromatin (DORCs) that represent coordinated regulatory regions. The choice between them depends on the biological question and data structure. SCENIC+ is well suited for cross-species comparisons and datasets where enhancer prediction is a priority. FigR is effective for studies of dynamic responses, such as immune stimulation, where DORC analysis reveals rapid regulatory changes.

### Can I use scRNA-seq data alone for GRN inference?

Yes, but with important limitations. Expression data alone can identify correlations between TF expression and target gene expression, but these correlations include many false positives because they cannot distinguish direct regulation from shared upstream control. Adding chromatin accessibility data filters these correlations by requiring that a candidate TF has a motif in an accessible region near the target gene. Tools such as scGATE use scATAC-seq data and TF DNA binding motifs to filter out non-relevant TFs in gene regulation, improving the specificity of networks inferred from scRNA-seq data. If accessibility data are not available, interpret expression-based networks with caution and validate key predictions experimentally.

### What quality control metrics are most important for scATAC-seq data?

The most important metrics are the number of unique fragments per cell, the fraction of fragments in peaks (FRiP), and the transcription start site (TSS) enrichment score. Fragment count reflects the depth of accessibility information available for each cell. FRiP measures the signal-to-noise ratio by indicating what proportion of fragments fall within called peaks. TSS enrichment assesses whether accessibility is concentrated at gene promoters, which is expected for healthy nuclei. Cells with low values on these metrics should be removed before downstream analysis because they will produce unreliable enhancer-to-gene links and TF activity estimates.

### How do I choose the distance window for enhancer-to-gene linkage?

The optimal distance window depends on the genomic architecture of the system being studied. Typical windows range from 100 to 500 kilobases from the target gene promoter. Too small a window misses distal enhancers that act over long distances. Too large a window includes unrelated genes and increases false-positive links. Examine the distribution of known enhancer-to-gene distances in your system, if available, and test how the network structure changes with different window sizes. The choice should balance sensitivity for distal enhancers against specificity for true target genes.

### What is a DORC and why is it useful?

A domain of regulatory chromatin (DORC) is a genomic region where chromatin accessibility is coordinated across cells and correlates with the expression of nearby genes. DORCs represent clusters of co-accessible peaks that likely function as regulatory hubs for a target gene or gene neighborhood. FigR identifies DORCs as part of its analysis framework. In the study of immune stimulation, DORCs revealed that cells alter chromatin accessibility and gene expression at timescales of minutes, demonstrating that regulatory changes can be rapid and coordinated. Identifying DORCs in your data highlights the most informative regulatory regions for downstream analysis and can reveal disease-associated regulatory elements.

### How can I validate the regulatory relationships inferred from my data?

Validation occurs at multiple levels. Computational validation includes cross-validation, comparison with external datasets such as ChIP-seq data, and motif enrichment analysis. In silico perturbation using tools such as CellOracle simulates the consequences of TF perturbation and predicts changes in cell identity, allowing prioritization of TFs for experimental validation. Experimental validation involves perturbing the predicted TF or enhancer and measuring the effect on target gene expression. CRISPR-based approaches can validate enhancer-to-gene links by deleting or disrupting a predicted enhancer. The combination of multiple computational methods with targeted experimental validation provides the strongest evidence for regulatory relationships.

### What are the main limitations of current GRN inference methods?

The primary limitation is that correlation does not equal causation. Inferred regulatory relationships may reflect indirect effects or shared upstream regulation. Motif-based TF assignment is approximate because actual binding depends on chromatin context and cooperative interactions that position weight matrices cannot capture. Enhancer-to-gene linkage remains uncertain because the rules governing enhancer-promoter interactions are not fully understood. Networks are context specific and may not generalize across cell types or conditions. Data quality constraints, including sparse accessibility data and batch effects, further limit the resolution and reliability of inferred networks.

### When should I consider spatial transcriptomics in addition to single-cell multi-omics?

Spatial transcriptomics adds tissue context that single-cell data lack. Methods such as ISON integrate spatial transcriptomics with single-cell multi-omics data to predict chromatin accessibility profiles for spatial spots and reconstruct spatially resolved gene regulatory networks. This approach is valuable when the spatial organization of regulatory programs is relevant to the biological question, such as in tumor microenvironments where cell-cell interactions depend on spatial proximity, or in developmental systems where positional information drives cell fate decisions. If the biological question involves spatial organization, consider adding spatial transcriptomics to the analysis.

## Related Bioinformatics Guides

- [Single-Cell Multi-Omics Integration: Methods and Applications](/knowledge/bioinformatics/single-cell-multi-omics-integration-methods-and-applications)
- [Single-Cell RNA Velocity: Inferring Cellular Dynamics](/knowledge/bioinformatics/single-cell-rna-velocity-inferring-cellular-dynamics)
- [RNA-Seq vs ChIP-Seq: Complementary Approaches for Gene Regulation](/knowledge/bioinformatics/rna-seq-vs-chip-seq-complementary-approaches-for-gene-regulation)
- [Single-Cell Sequencing Depth: How Much Is Enough?](/knowledge/bioinformatics/single-cell-sequencing-depth-how-much-is-enough)
- [Single-Cell Sequencing Workflow: From Sample Preparation to Data Analysis](/knowledge/bioinformatics/single-cell-sequencing-workflow-from-sample-preparation-to-data-analysis)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)


## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [Functional inference of gene regulation using single-cell multi-omics.](https://pubmed.ncbi.nlm.nih.gov/36204155). Cell genomics, 2022.
- [SCENIC+: single-cell multiomic inference of enhancers and gene regulatory networks.](https://pubmed.ncbi.nlm.nih.gov/37443338). Nature methods, 2023.
- [Single-cell multi-omics reveals that FABP1 + renal cell carcinoma drive tumor angiogenesis through the PLG-PLAT axis under fatty acid reprogramming.](https://pubmed.ncbi.nlm.nih.gov/40518526). Molecular cancer, 2025.
- [Dissecting cell identity via network inference and in silico gene perturbation.](https://pubmed.ncbi.nlm.nih.gov/36755098). Nature, 2023.
- [Single-cell multi-omics reveals DUSP9 as a key regulator of cancer stemness and a potential therapeutic target in hepatocellular carcinoma.](https://pubmed.ncbi.nlm.nih.gov/41508037). Journal of translational medicine, 2026.
- [Integrative multi-omics analysis reveals a novel subtype of hepatocellular carcinoma with biological and clinical relevance.](https://pubmed.ncbi.nlm.nih.gov/39712016). Frontiers in immunology, 2024.
- [Single-cell multi-omics analysis identifies context-specific gene regulatory gates and mechanisms.](https://pubmed.ncbi.nlm.nih.gov/38653489). Briefings in bioinformatics, 2024.
- [scMEGA: single-cell multi-omic enhancer-based gene regulatory network inference.](https://pubmed.ncbi.nlm.nih.gov/36698768). Bioinformatics advances, 2023.
- [Inference of spatial chromatin accessibility via integration of spatial transcriptomics and single-cell multi-omics data.](https://doi.org/10.1038/s41467-026-73948-7). 2026.
- [Integrating multi-omics and machine learning to uncover CCL20 as a potential regulator of the immunosuppressive microenvironment in lung adenocarcinoma.](https://doi.org/10.3389/fimmu.2026.1870538). 2026.
- [Computational blueprints for cell fate programming.](https://doi.org/10.1016/j.stemcr.2026.102929). 2026.
- [A multi-modal diffusion model with dual-cross-attention for multi-omics data generation and translation.](https://doi.org/10.1038/s41467-026-71744-x). 2026.
- [scTFBridge: a disentangled deep generative model informed by TF-motif binding for gene regulation inference in single-cell multi-omics](https://doi.org/10.1038/s41467-025-64227-y). bioRxiv, 2025.
- [Integrative Single-Cell Multi-Omics Network Analysis to Elucidate Epigenetic Regulation in Neurodevelopmental Disorders.](https://doi.org/10.1016/j.slast.2026.100395). SLAS technology, 2026.
- [scMTNI: Leveraging cellular trajectory and context to infer dynamic GRNs from single-cell multi-omics data.](https://www.semanticscholar.org/paper/c7781fd8687ab63181efe6ab611cc8ff96108cd4). arXiv.org, 2026.
- [Single-Cell Multi-Omics Reveal Gene Regulatory Mechanisms Underlying Cardiac Embryonic Development](https://doi.org/10.3390/genes17040414). Genes, 2026.
- [Interpretable data integration for single-cell and spatial multi-omics](https://doi.org/10.1016/j.cels.2025.101479). Cell Systems, 2026.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.