# Linking Gene Expression to Regulatory Elements: How to Use Single-Cell Multi-Omics for Enhancer-Gene Linkage Inference


## Key Takeaways

- Single-cell multi-omics, specifically paired scRNA-seq and scATAC-seq, is crucial for inferring enhancer-gene linkages by resolving cell-type-specific regulatory relationships that are obscured in bulk assays.
- Co-accessibility, measured by correlating chromatin accessibility between genomic regions across single cells, and peak-to-gene correlation, linking accessible chromatin regions to expressed genes, are foundational statistical signals for linkage inference.
- Tools like Cicero leverage co-accessibility on sparse scATAC-seq data, while SCENT and Enhlink directly integrate scRNA-seq and scATAC-seq for peak-to-gene correlation and robust, condition-specific linkage inference, respectively.
- Validation strategies are critical, including motif analysis to identify relevant transcription factor binding sites within predicted enhancers and CRISPR perturbation experiments to establish causal regulatory relationships.
- Batch effects and cell type composition are common confounding factors; rigorous quality control, within-cell-type analysis, and the use of tools that account for technical covariates are essential for accurate inference.
- The ultimate output is a ranked list of enhancer-gene linkages per cell type, enabling prioritization of candidate regulatory variants, interpretation of GWAS loci, and design of functional validation experiments.

---

Single-cell multi-omics integration addresses a core question in regulatory genomics: which enhancers control which genes in which cell types. This article provides a practical workflow for researchers who have generated or plan to generate paired single-cell RNA sequencing (scRNA-seq) and single-cell ATAC sequencing (scATAC-seq) data and need to infer enhancer-gene linkages. The workflow covers data inputs, tool selection, quality controls, validation strategies, and interpretation limits, with emphasis on tools such as Cicero, SCENT, and Enhlink, and validation approaches including motif analysis and CRISPR perturbation.

## Scope and Reader Context

This article serves biology students, researchers, laboratory professionals, and life-science practitioners who need to move from raw single-cell multi-omics data to testable hypotheses about enhancer-gene regulation. The primary problem addressed is the inference of which ATAC-seq peaks function as enhancers for specific genes in specific cell types. This problem arises in diverse research contexts, including studies of developmental biology, cancer heterogeneity, complex disease genetics, and agricultural trait mapping.

The workflow described here assumes familiarity with basic single-cell analysis concepts, including quality control, normalization, dimensionality reduction, and clustering. Readers who need to build these foundations should consult training resources from the [EMBL-EBI Training portal](https://www.ebi.ac.uk/training) for structured bioinformatics learning pathways and the [Galaxy Training Network](https://training.galaxyproject.org/) for accessible workflow tutorials. Foundational computing skills, including shell navigation and version control, are covered in [The Carpentries lessons](https://carpentries.org/lessons).

The practical outcome of this workflow is a ranked list of enhancer-gene linkages for each cell type in your dataset, with supporting evidence from co-accessibility, correlation, motif enrichment, and where possible, physical interaction data or perturbation experiments. These linkages can then be used to prioritize candidate regulatory variants, interpret genome-wide association study (GWAS) loci, or design functional validation experiments.

## The Regulatory Linkage Problem in Single-Cell Data

Enhancers are regulatory DNA sequences that can activate gene transcription over genomic distances that range from a few kilobases to more than a megabase. Unlike promoters, which are typically located immediately upstream of the transcription start site, enhancers can be located in intergenic regions, introns, or even within the bodies of unrelated genes. This positional flexibility creates a fundamental assignment problem: given a set of accessible chromatin regions and a set of expressed genes, which accessible regions regulate which genes?

Bulk chromatin assays such as ChIP-seq and bulk ATAC-seq provide population-averaged signals that obscure cell-type-specific regulatory relationships. A peak that appears accessible in a bulk sample may be accessible in only a minority of cells, and the gene it regulates may be expressed in a different cellular subpopulation. Single-cell ATAC-seq resolves this ambiguity by measuring chromatin accessibility in individual cells, allowing researchers to ask whether accessibility at a putative enhancer and expression of a candidate target gene covary across cells.

The integration of scRNA-seq and scATAC-seq from the same biological system provides the statistical power needed for this covariance analysis. When both modalities are measured in the same cells, a technique known as single-cell multi-omics or single-cell multiome, the linkage inference can be made with greater confidence because the pairing of accessibility and expression is exact instead of inferred through matching cell types across separate experiments.

Recent work demonstrates the biological importance of accurate enhancer-gene linkage inference. A study of hepatocellular carcinoma integrated scRNA-seq and scATAC-seq to resolve cell-type-specific chromatin architectures, identifying enhancer-like peaks that formed long-range regulatory hubs and showing that eighteen enhancer-linked genes exhibited tumor-specific overexpression with diagnostic accuracy [11](https://pubmed.ncbi.nlm.nih.gov/40606924). In a study of neuroblastoma, single-cell multi-omics revealed extensive epigenetic priming with latent capacity for state transitions, and the resulting enhancer gene regulatory networks sustained aggressive tumor states [18](https://doi.org/10.1016/j.devcel.2025.04.013). These examples illustrate that accurate linkage inference is also a computational exercise but a prerequisite for understanding disease mechanisms.

## Core Principles of Enhancer-Gene Linkage Inference

### Co-Accessibility as a Statistical Signal

The foundational principle underlying most linkage inference methods is co-accessibility. If two genomic regions are accessible in the same cells, they are more likely to be part of the same regulatory complex. This principle extends the concept of correlation to single-cell resolution. In a population of cells, the accessibility of an enhancer and its target promoter should be positively correlated because both are open in the cell types where the gene is active and closed where it is not.

Cicero, one of the earliest tools for this purpose, operationalizes co-accessibility by constructing a k-nearest neighbor graph in single-cell ATAC-seq data and then calculating co-accessibility scores between pairs of peaks. The method accounts for the sparsity of single-cell ATAC-seq data by aggregating information across similar cells before computing correlations. The output is a set of peak-to-peak connections with associated co-accessibility scores, and peaks that connect to a gene promoter are considered candidate enhancers for that gene.

The choice of distance threshold is a critical parameter in co-accessibility analysis. Regulatory interactions can span large genomic distances, but the probability of interaction decreases with distance. Most tools use a maximum distance cutoff, typically in the range of 250 kilobases to 1 megabase, to limit the search space and reduce false positives. The appropriate threshold depends on the biological system under study and should be informed by available chromatin conformation data when possible.

### Peak-to-Gene Correlation and Activity Scores

A complementary approach uses peak-to-gene correlation, where the accessibility of each peak is correlated with the expression of nearby genes across cells. This approach requires paired or matched scRNA-seq and scATAC-seq data. In paired multi-omics data, the correlation is computed directly because each cell has both measurements. In unpaired data, the correlation is computed after matching cells across modalities, either by cell type labels or by integrative embedding.

Gene activity scores provide an intermediate step for datasets where scRNA-seq and scATAC-seq are not paired. A gene activity score aggregates accessibility signals from the promoter and nearby peaks into a single value per gene per cell, approximating expression. This score can then be used for cell type identification and for integration with transcriptome data. The predictive modeling approach implemented in MAPLE demonstrates that supervised learning can improve gene activity prediction from single-cell methylation data, and the same principle applies to accessibility-based activity scores [13](https://pubmed.ncbi.nlm.nih.gov/33219054).

### Integration of Physical Interaction Data

Co-accessibility and correlation provide statistical evidence for regulatory linkage, but they do not prove physical interaction. Chromatin conformation capture techniques such as Hi-C and promoter capture Hi-C provide direct evidence that two genomic regions are in physical proximity in the nucleus. When such data are available for the same cell types or tissues, they can validate or refine linkage predictions.

Enhlink, a computational tool for scATAC-seq analysis, was evaluated using promoter capture Hi-C data from the mouse striatum, demonstrating that its predictions align with physical enhancer-promoter interactions [14](https://pubmed.ncbi.nlm.nih.gov/37214950). The tool uses an ensemble approach that incorporates technical covariates to control for batch effects and biological covariates to infer condition-specific regulatory linkages [12](https://pubmed.ncbi.nlm.nih.gov/39223609). This integration of statistical inference with physical interaction data represents the current best practice for enhancer-gene linkage.

## At a Glance: Tool Comparison for Enhancer-Gene Linkage

| Tool | Input Data | Core Method | Output | Key Strengths | Primary Limitations |
|------|-----------|-------------|--------|---------------|-------------------|
| Cicero | scATAC-seq | k-nearest neighbor graph, co-accessibility scores | Peak-to-peak connections with co-accessibility scores | Handles sparse data, widely used, established documentation | Requires careful parameter tuning, no direct expression integration |
| SCENT | scRNA-seq and scATAC-seq | Correlation of peak accessibility with gene expression | Peak-to-gene linkages with correlation statistics | Directly integrates both modalities, cell-type specific | Requires paired or well-matched data, sensitive to batch effects |
| Enhlink | scATAC-seq, optional multi-omics | Ensemble approach with technical and biological covariates | Condition-specific linkages with p-values | Outperforms alternatives in benchmark studies, integrates multi-omic data | Newer tool, requires familiarity with ensemble methods |
| ArchR | scATAC-seq, optional scRNA-seq | Peak-to-gene linkages, co-accessibility networks | Cell-type-specific regulatory hubs | Comprehensive pipeline, handles large datasets | Complex installation, steep learning curve |

## Data Inputs and Experimental Design Considerations

### Paired Multi-Omics Data

The ideal input for enhancer-gene linkage inference is paired single-cell multi-omics data, where chromatin accessibility and gene expression are measured simultaneously in the same cells. This design eliminates the need to match cells across separate experiments and provides the strongest statistical foundation for correlation-based linkage inference. Commercial platforms offer kits for simultaneous profiling of RNA and accessible chromatin from single nuclei, and these have become standard in the field.

When designing a paired multi-omics experiment, consider the number of cells needed for adequate statistical power. Linkage inference requires sufficient cells per cell type to compute stable correlations. A common target is several thousand cells per sample, with enough samples to capture biological variability. The exact number depends on the complexity of the tissue, the abundance of rare cell types, and the expected effect sizes of regulatory relationships.

### Unpaired Data and Bridge Integration

Many research projects have existing scRNA-seq and scATAC-seq datasets generated separately, often from different samples or even different laboratories. These unpaired datasets can still support linkage inference, but the analysis requires an integration step to match cells across modalities. Methods for this integration include canonical correlation analysis, manifold alignment, and latent-variable modeling, as demonstrated in an integrative framework for neurodevelopmental disorders that combined scRNA-seq, scATAC-seq, and single-cell methylation data [15](https://doi.org/10.1016/j.slast.2026.100395).

Deep learning approaches have expanded the options for integrating unpaired data. scPairing embeds different modalities from the same cells onto a common embedding space and can generate novel multi-omics data following bridge integration, which uses an existing multi-omics bridge to link unimodal datasets [19](https://doi.org/10.1016/j.crmeth.2025.101211). This approach can increase the effective sample size for linkage inference when paired data are limited.

### Quality Control Before Linkage Inference

Quality control is the foundation of reliable linkage inference. Poor-quality cells with low library complexity or high mitochondrial content can introduce spurious correlations. The standard quality control metrics for scRNA-seq, including total counts, number of detected genes, and mitochondrial fraction, apply to the RNA modality. For the ATAC modality, additional metrics include the fraction of fragments in peaks, the transcription start site enrichment score, and the ratio of mononucleosomal to nucleosome-free fragments.

After cell-level quality control, perform doublet detection to remove cells that represent two or more nuclei captured together. Doublets can create false cell types and distort correlation structures. Several computational tools are available for doublet detection in both scRNA-seq and scATAC-seq data.

Batch effects are a major source of false linkages. If cells are processed in multiple batches, technical variation can create correlations between accessibility and expression that reflect processing artifacts instead of biology. The ensemble approach in Enhlink explicitly incorporates cell-level technical covariates to control for batch effects [14](https://pubmed.ncbi.nlm.nih.gov/37214950). For other tools, include batch information in the model or use batch correction methods before linkage inference.

## Practical Workflow for Linkage Inference

### Step 1: Data Preparation and Quality Control

Begin with count matrices for both modalities. For scRNA-seq, this is a genes-by-cells matrix of counts. For scATAC-seq, this is a peaks-by-cells matrix of counts, where peaks are called from the aggregated accessibility signal. Use standard pipelines for initial processing, and document all parameters.

Apply quality control filters to both modalities. For scRNA-seq, filter cells based on total counts, gene counts, and mitochondrial fraction. For scATAC-seq, filter cells based on total fragments, fraction of fragments in peaks, and transcription start site enrichment. Remove doublets using computational tools. Normalize each modality appropriately, using log-normalization for RNA and term frequency-inverse document frequency or similar transformations for ATAC.

Verify that the cell type composition is consistent across modalities. If using paired data, confirm that the same cell types are represented in both modalities. If using unpaired data, perform integration and check that cell types match across datasets.

### Step 2: Cell Type Identification

Identify cell types in the integrated dataset. For paired data, cluster the RNA modality and project the ATAC modality onto the same clusters. For unpaired data, perform joint embedding and clustering. Annotate clusters using canonical marker genes and transcription factor motifs.

Cell type annotation is critical because linkage inference is performed within cell types. A linkage that appears significant across all cells may reflect cell type composition instead of regulatory relationships. For example, if a peak is accessible only in T cells and a gene is expressed only in T cells, the correlation will be high even if the peak does not regulate the gene. Within-cell-type analysis avoids this confounding.

### Step 3: Peak Calling and Filtering

Call peaks from the aggregated ATAC-seq signal across all cells, then assign peaks to cell types based on the accessibility patterns. Filter peaks that are present in very few cells or that overlap known blacklist regions. Consider peak width and signal strength as additional filters.

For linkage inference, focus on peaks that are distal to promoters, typically defined as more than 2 kilobases from any transcription start site. Promoter-proximal peaks are more likely to represent promoter accessibility instead of enhancer activity. However, some enhancers are located in promoter regions, so do not exclude promoter-overlapping peaks entirely.

### Step 4: Run Linkage Inference Tools

Run at least two independent linkage inference tools to cross-validate results. A common combination is Cicero for co-accessibility and SCENT or Enhlink for peak-to-gene correlation. Document the parameters used for each tool, including distance thresholds, correlation cutoffs, and covariate adjustments.

For Cicero, construct the k-nearest neighbor graph, compute co-accessibility scores, and extract connections between distal peaks and gene promoters. For SCENT or Enhlink, compute peak-to-gene correlations within each cell type, adjusting for technical covariates as appropriate.

### Step 5: Aggregate and Rank Linkages

Combine results from multiple tools into a single ranked list. A linkage supported by both co-accessibility and peak-to-gene correlation is more credible than a linkage supported by only one method. Consider the distance between peak and promoter, the strength of the correlation, and the cell type specificity of both the peak and the gene.

Create a summary table with columns for peak coordinates, gene name, cell type, co-accessibility score, correlation coefficient, and supporting evidence. This table serves as the primary output for downstream analysis and validation.

### Step 6: Validate with Motif Analysis

Motif analysis provides biological validation for predicted linkages. If a peak is predicted to regulate a gene, the peak should contain binding motifs for transcription factors that are relevant to the cell type and the gene's function. For example, a peak predicted to regulate a hepatocyte-specific gene should contain motifs for liver-enriched transcription factors such as HNF4A, CEBPA, FOXA1, and ONECUT1, as demonstrated in the liver zonation study [22](https://doi.org/10.1038/s41556-023-01316-4).

Use motif scanning tools to identify transcription factor binding motifs within predicted enhancer peaks. Compare the motif content of predicted enhancers to the motif content of background peaks matched for accessibility and genomic location. Enriched motifs provide supporting evidence for the regulatory function of the predicted enhancers.

### Step 7: Validate with CRISPR Perturbation

CRISPR perturbation provides the strongest validation for enhancer-gene linkages. Deleting or disrupting a predicted enhancer should reduce expression of the target gene if the linkage is correct. CRISPR interference, which uses a catalytically dead Cas9 fused to a transcriptional repressor, can silence enhancer activity without deleting the underlying sequence.

Design CRISPR experiments for a prioritized subset of predicted linkages. Select linkages with strong statistical support, biological plausibility, and relevance to the research question. Include multiple guide RNAs per enhancer to control for off-target effects. Measure target gene expression by quantitative PCR or RNA sequencing after perturbation.

The hepatocellular carcinoma study provides an example of functional validation following linkage inference. After identifying DGAT1 as the hub gene most strongly regulated by chromatin accessibility, the researchers used small interfering RNA to knock down DGAT1 and demonstrated that this knockdown suppressed cell proliferation [11](https://pubmed.ncbi.nlm.nih.gov/40606924). This validation approach confirms that the regulatory relationships identified by linkage inference have functional consequences.

## Options and Tradeoffs in Tool Selection

### Cicero for Co-Accessibility

Cicero is the most established tool for co-accessibility analysis in single-cell ATAC-seq data. Its strengths include a well-documented workflow, availability through [Bioconductor](https://bioconductor.org/), and extensive use in published studies. The tool handles the sparsity of single-cell ATAC-seq data by aggregating information across similar cells, which reduces noise and improves the stability of co-accessibility estimates.

The primary limitation of Cicero is that it does not directly integrate gene expression data. The output is peak-to-peak connections, and the researcher must map connections to genes by identifying which connections involve promoter peaks. This indirect approach can miss linkages where the enhancer and promoter are not directly connected but are both connected to an intermediate element.

### SCENT for Peak-to-Gene Correlation

SCENT directly integrates scRNA-seq and scATAC-seq data to infer peak-to-gene linkages. The tool computes correlations between peak accessibility and gene expression across cells, providing a direct statistical test of the regulatory relationship. This approach is more interpretable than co-accessibility alone because it connects regulatory elements to their putative target genes.

The main limitation of SCENT is its dependence on data quality and integration. If the scRNA-seq and scATAC-seq data are not well matched, the correlations will be noisy and unreliable. The tool also requires careful handling of technical covariates to avoid false correlations driven by batch effects.

### Enhlink for Ensemble Inference

Enhlink represents a newer generation of linkage inference tools that use ensemble methods to improve robustness. The tool incorporates technical covariates to control for batch effects and biological covariates to infer condition-specific linkages [14](https://pubmed.ncbi.nlm.nih.gov/37214950). Benchmarking against simulated and real data, including promoter capture Hi-C data, demonstrated that Enhlink outperforms popular alternative strategies [12](https://pubmed.ncbi.nlm.nih.gov/39223609).

The ensemble approach in Enhlink provides several advantages. It can integrate multi-omic data when available, increasing specificity. It produces p-values for each linkage, facilitating statistical interpretation. The tool also handles condition-specific analysis, allowing researchers to identify linkages that differ between experimental conditions.

### ArchR for Comprehensive Analysis

ArchR provides a comprehensive pipeline for single-cell ATAC-seq analysis, including peak calling, cell type identification, and peak-to-gene linkage inference. The tool is designed to handle large datasets efficiently and includes visualization functions for exploring regulatory relationships. ArchR was used in the hepatocellular carcinoma study to resolve cell-type-specific chromatin architectures and identify regulatory hubs through peak-to-gene linkages and co-accessibility networks [11](https://pubmed.ncbi.nlm.nih.gov/40606924).

The tradeoff with ArchR is complexity. The tool has a steep learning curve and requires familiarity with its specific data structures and functions. However, for researchers who need a complete analysis pipeline instead of a single linkage inference step, ArchR offers an integrated solution.

## Records and Measurements for Reproducible Analysis

### Documentation Requirements

Reproducible linkage inference requires complete documentation of the analysis. Record the version of every software tool and package used, including the operating system and computing environment. Document all parameters, including quality control thresholds, normalization methods, distance cutoffs, and correlation thresholds. Save the exact commands or scripts used for each analysis step.

Version control is essential for tracking changes to analysis scripts. Use Git or a similar system to maintain a history of script versions. The [Carpentries lessons](https://carpentries.org/lessons) provide foundational training in version control and reproducible computing practices.

### Analysis Notebooks

Maintain an analysis notebook that records decisions and rationales. For each step in the workflow, note the input data, the parameters used, the output generated, and any deviations from the planned analysis. This notebook serves as the primary record for reproducing the analysis and for reporting methods in publications.

The [Galaxy Training Network](https://training.galaxyproject.org/) offers tutorials that demonstrate reproducible workflow practices, including the use of analysis histories and workflow definitions. These practices are directly applicable to single-cell multi-omics analysis.

### Data Storage and Sharing

Store raw data, processed data, and analysis outputs in organized directory structures. Use descriptive file names that include the sample identifier, data type, and processing date. Maintain a manifest file that lists all files, their checksums, and their relationships to analysis steps.

For publication, deposit raw sequencing data in public repositories such as the [NCBI databases](https://www.ncbi.nlm.nih.gov/), which provide official infrastructure for sequence data archiving and search. Processed data and analysis scripts should be shared through appropriate repositories to enable replication by other researchers.

## Common Failure Patterns and Troubleshooting

### Failure Pattern 1: Cell Type Composition Confounds Linkages

The most common failure in linkage inference is confounding by cell type composition. When linkages are computed across all cells without stratification, the results reflect differences in cell type proportions instead of regulatory relationships. A peak accessible only in one cell type will appear linked to any gene expressed in that cell type, regardless of whether the peak regulates the gene.

Troubleshooting: Always compute linkages within cell types. Verify that the cell type annotation is consistent across modalities and that each cell type has sufficient cells for stable correlation estimates. If a linkage appears significant across all cells but not within any cell type, it is likely a composition artifact.

### Failure Pattern 2: Batch Effects Create Spurious Correlations

Technical variation between batches can create correlations between accessibility and expression that do not reflect biology. If all cells from one batch have higher accessibility at a peak and higher expression of a gene, the correlation will be significant even if the peak does not regulate the gene.

Troubleshooting: Include batch as a covariate in the linkage model. Use tools that explicitly handle technical covariates, such as Enhlink [14](https://pubmed.ncbi.nlm.nih.gov/37214950). Visualize the data by batch to identify systematic differences before running linkage inference.

### Failure Pattern 3: Sparse Data Produces Unstable Estimates

Single-cell ATAC-seq data are extremely sparse, with most peaks detected in only a small fraction of cells. This sparsity can produce unstable correlation estimates, particularly for rare cell types or lowly accessible peaks.

Troubleshooting: Use tools that aggregate information across similar cells, such as Cicero. Increase the number of cells per cell type if possible. Filter peaks with very low detection rates before computing linkages. Consider using imputation methods to reduce sparsity, but validate that imputation does not introduce artifacts.

### Failure Pattern 4: Distance Thresholds Exclude Valid Linkages

Regulatory interactions can span large genomic distances, and a distance threshold that is too restrictive will miss valid linkages. Conversely, a threshold that is too permissive will include many false positives.

Troubleshooting: Use distance thresholds informed by chromatin conformation data when available. The cardiac atlas study constructed a cell-type-resolved enhancer-to-gene linkage map that refined the association of cardiomyopathy genetic risk loci to downstream target genes [17](https://doi.org/10.1186/s13059-026-04061-7), demonstrating the value of comprehensive distance modeling. Test multiple thresholds and compare the stability of top linkages across thresholds.

### Failure Pattern 5: Motif Analysis Does Not Support Predicted Linkages

If motif analysis does not reveal expected transcription factor binding sites in predicted enhancers, the linkage may be incorrect, or the relevant transcription factor may not have a known motif.

Troubleshooting: Check whether the predicted enhancer is accessible in the expected cell type. Verify that the transcription factor is expressed in the same cells where the enhancer is accessible. Consider that the relevant transcription factor may bind indirectly through cooperative interactions with other factors.

## Limitations of Enhancer-Gene Linkage Inference

### Statistical Correlation Does Not Prove Causation

The fundamental limitation of linkage inference is that correlation does not establish causation. A peak and a gene may be correlated because they are both regulated by the same upstream factor, not because the peak regulates the gene. Co-accessibility and peak-to-gene correlation provide statistical evidence that is consistent with a regulatory relationship, but they do not prove that the relationship is causal.

This limitation is particularly important when interpreting linkages for clinical or agricultural applications. A linkage that is statistically significant may not be biologically functional, and functional validation is required before using a linkage to guide intervention decisions.

### Cell Type Resolution Is Limited by Data Quality

The resolution of linkage inference is limited by the ability to distinguish cell types and states. If a tissue contains closely related cell states that are difficult to separate, linkages may be computed across a mixture of states, obscuring state-specific regulatory relationships. The neuroblastoma study demonstrated that developmental intermediate states are critical for malignant transitions, and these states may be missed if the analysis does not resolve them [18](https://doi.org/10.1016/j.devcel.2025.04.013).

### Technical Artifacts Can Mimic Regulatory Signals

Technical artifacts in single-cell data can create signals that resemble regulatory relationships. Dropout events, where a gene is not detected in a cell despite being expressed, can create spurious correlations. Amplification biases can inflate accessibility signals at certain genomic regions. These artifacts are difficult to distinguish from true regulatory signals without careful quality control and validation.

### Cross-Species Conservation Is Not Guaranteed

Enhancer-gene linkages identified in one species may not be conserved in another species. Regulatory elements evolve rapidly, and the specific enhancers that control a gene in humans may differ from those that control the orthologous gene in mice. The integrative framework for neurodevelopmental disorders demonstrated consistent cross-species regulatory conservation for some linkages [15](https://doi.org/10.1016/j.slast.2026.100395), but this conservation is not universal.

## Validation Strategies and Interpretation

### Motif Enrichment Analysis

Motif enrichment provides a first layer of validation for predicted enhancers. If a set of predicted enhancers is enriched for binding motifs of transcription factors known to be important in the relevant cell type, this supports the functional relevance of the linkages. The liver zonation study identified TCF7L1 and TBX3 as repressors alongside core hepatocyte transcription factors including HNF4A, CEBPA, FOXA1, and ONECUT1, and these factors were validated through massively parallel reporter assays [22](https://doi.org/10.1038/s41556-023-01316-4).

When performing motif enrichment, compare predicted enhancers to a matched background set. The background should be matched for accessibility level, genomic location, and sequence composition. Enrichment relative to this background provides evidence that the predicted enhancers are enriched for specific regulatory signals.

### Integration with Genetic Association Data

Linkage predictions can be validated by integration with genetic association data. If a predicted enhancer contains a genetic variant associated with a trait or disease, and the predicted target gene is biologically relevant to that trait, this provides convergent evidence for the linkage. The Enhlink study coupled linkage inference with eQTL analysis to identify a putative super-enhancer in striatal neurons [12](https://pubmed.ncbi.nlm.nih.gov/39223609).

This integration approach is particularly powerful for interpreting GWAS loci. Many disease-associated variants fall in non-coding regions, and linkage inference can connect these variants to their putative target genes. The cardiac atlas study used enhancer-to-gene linkage maps to refine the association of dilated and hypertrophic cardiomyopathy genetic risk loci to downstream target genes [17](https://doi.org/10.1186/s13059-026-04061-7).

### CRISPR Perturbation for Causal Validation

CRISPR perturbation provides the strongest validation for enhancer-gene linkages. Deleting or silencing a predicted enhancer and observing a change in target gene expression confirms the regulatory relationship. This validation is essential before using a linkage to guide therapeutic or agricultural decisions.

Design CRISPR experiments carefully to avoid off-target effects. Use multiple guide RNAs targeting different regions of the predicted enhancer. Include appropriate controls, including non-targeting guides and guides targeting known neutral regions. Measure target gene expression with sufficient replication to detect modest effect sizes.

### Interpretation of Conflicting Evidence

When different validation approaches give conflicting results, interpret the evidence carefully. A linkage supported by co-accessibility and correlation but not by motif analysis may still be valid if the relevant transcription factor has an unknown motif or binds indirectly. A linkage supported by motif analysis but not by correlation may reflect a regulatory relationship that is masked by technical noise.

Prioritize linkages with convergent evidence from multiple approaches. A linkage supported by co-accessibility, peak-to-gene correlation, motif enrichment, and genetic association data is much more credible than a linkage supported by only one line of evidence.

## Quality Controls and Reproducibility Standards

### Computational Reproducibility

Reproducible analysis requires a controlled computing environment. Use containerization tools such as Docker or Singularity to package the analysis environment, including all software dependencies and versions. Document the container image and the exact commands used to run the analysis.

Workflow management systems provide additional reproducibility guarantees. The [nf-core documentation](https://nf-co.re/docs) describes community standards for pipeline usage and configuration that ensure consistent execution across computing environments. These standards are directly applicable to single-cell multi-omics analysis.

### Biological Replication

Linkage inference should be performed on biological replicates to ensure that results are reproducible across samples. A linkage that appears in one sample but not in another may reflect sample-specific biology or technical artifacts. The cardiac atlas study integrated data from 299 donors for RNA and 106 donors for ATAC, providing substantial replication power [17](https://doi.org/10.1186/s13059-026-04061-7).

When biological replicates are not available, use statistical approaches to estimate the uncertainty of linkage predictions. Bootstrap or jackknife resampling can provide confidence intervals for correlation estimates. These intervals help distinguish robust linkages from those that depend on a small number of cells.

### Reporting Standards

Report all analysis parameters and decisions in publications. Include the versions of all software tools, the quality control thresholds, the normalization methods, and the linkage inference parameters. Provide access to analysis scripts and processed data to enable replication.

The [Bioconductor project](https://bioconductor.org/) provides official documentation for reproducible genomic analysis, including best practices for package versioning and workflow documentation. Following these standards ensures that other researchers can reproduce and build on your analysis.

## Applications in Disease Research and Trait Mapping

### Interpreting Disease-Associated Genetic Variants

Enhancer-gene linkage inference has become a standard tool for interpreting disease-associated genetic variants. Many disease-associated variants fall in non-coding regions, and linkage inference connects these variants to their putative target genes. This connection is essential for understanding disease mechanisms and identifying therapeutic targets.

The migraine study integrated GWAS data with transcriptomic, epigenomic, and spatially resolved single-cell transcriptomics profiles, using multiple analytical approaches to prioritize genes based on methodological robustness [7](https://pubmed.ncbi.nlm.nih.gov/40826382). This multi-omics framework demonstrates how linkage inference contributes to the interpretation of complex disease genetics.

### Understanding Cell-Type-Specific Disease Mechanisms

Linkage inference reveals cell-type-specific regulatory relationships that are invisible in bulk data. A disease-associated variant may affect gene regulation only in specific cell types, and this specificity has important implications for understanding disease mechanisms and designing targeted therapies.

The allergic rhinitis study performed scRNA-seq and scATAC-seq on nasal mucosa samples from 39 subjects, identifying cell subset-specific epigenetic alterations and constructing a deep learning framework that can prioritize putative cell type-specific regulatory linkages [8](https://pubmed.ncbi.nlm.nih.gov/42435887). This approach revealed disease-associated regulatory changes that would be missed in bulk analysis.

### Agricultural Trait Mapping

Linkage inference is also applicable to agricultural trait mapping. The pig backfat thickness study integrated GWAS, fine-mapping, and single-cell ATAC-seq to identify a chromatin accessibility peak that regulates PMAIP1 expression in inhibitory neurons via enhancer-mediated mechanisms [10](https://pubmed.ncbi.nlm.nih.gov/41111148). This example demonstrates how linkage inference connects genetic variants to regulatory mechanisms in agricultural species.

For agricultural applications, linkage inference can prioritize candidate causal variants and genes for breeding decisions. A variant that falls in a predicted enhancer for a gene with known effects on the trait of interest is a stronger candidate than a variant in a non-regulatory region.

## Professional Escalation Criteria

### When to Seek Specialized Bioinformatics Support

Researchers should consider seeking specialized bioinformatics support when the analysis exceeds their computational expertise or when results are inconsistent across tools. Specific situations that warrant escalation include:

- The dataset is very large and requires distributed computing resources beyond local capacity
- The integration of unpaired scRNA-seq and scATAC-seq data produces unstable cell type assignments
- Linkage predictions from different tools are highly inconsistent
- The analysis requires custom statistical modeling beyond standard tool capabilities
- The results will guide clinical or commercial decisions and require rigorous validation

### When to Consult Statistical Genetics Expertise

Statistical genetics expertise is needed when integrating linkage predictions with genetic association data. The interpretation of GWAS loci requires understanding of linkage disequilibrium, fine-mapping, and colocalization. The ferroptosis study in type 2 diabetes used genetic correlation analyses, Mendelian randomization, and colocalization analyses to complement single-cell RNA sequencing investigations [9](https://pubmed.ncbi.nlm.nih.gov/39716081), demonstrating the complexity of these integrated analyses.

### When to Engage Functional Validation Collaborators

Functional validation of enhancer-gene linkages requires specialized experimental expertise. CRISPR perturbation experiments, massively parallel reporter assays, and chromatin conformation capture require substantial experimental optimization. Engage collaborators with these capabilities before committing to validation experiments.

The liver zonation study combined single-cell multi-omics, spatial omics, massively parallel reporter assays, and deep learning to map enhancer-gene regulatory networks [22](https://doi.org/10.1038/s41556-023-01316-4). This integrated approach required expertise across multiple experimental and computational domains.

## Frequently Asked Questions

### What is the minimum number of cells needed for reliable enhancer-gene linkage inference?

The minimum number of cells depends on the complexity of the tissue and the abundance of the cell types of interest. For common cell types, several hundred cells per cell type may be sufficient for stable correlation estimates. For rare cell types, thousands of cells may be needed to achieve adequate statistical power. The sparsity of single-cell ATAC-seq data means that more cells are generally better, and tools that aggregate information across similar cells can partially compensate for limited cell numbers.

### Can enhancer-gene linkages be inferred from scATAC-seq data alone without paired scRNA-seq?

Yes, co-accessibility analysis with tools such as Cicero can infer peak-to-peak connections from scATAC-seq data alone. However, this approach does not directly connect peaks to genes. Gene activity scores can approximate expression from accessibility data, but these scores are less accurate than direct expression measurements. Paired multi-omics data provide stronger evidence for enhancer-gene linkages because the correlation between accessibility and expression is measured directly in the same cells.

### How do I choose between Cicero, SCENT, and Enhlink for my analysis?

The choice depends on your data and research question. If you have scATAC-seq data only, Cicero is the established choice for co-accessibility analysis. If you have paired scRNA-seq and scATAC-seq data, SCENT provides direct peak-to-gene correlation. If you need robust condition-specific linkages with p-values and can accommodate a newer tool, Enhlink offers an ensemble approach that outperforms alternatives in benchmark studies [12](https://pubmed.ncbi.nlm.nih.gov/39223609). Running at least two tools and cross-validating results is recommended.

### What distance threshold should I use for linking enhancers to genes?

The appropriate distance threshold depends on the biological system and the available evidence. Regulatory interactions can span more than a megabase, but most interactions occur within 250 kilobases of the target promoter. Start with a threshold of 250 to 500 kilobases and test the stability of top linkages across thresholds. If chromatin conformation data are available for your system, use these data to inform the threshold.

### How can I validate that a predicted enhancer-gene linkage is functionally relevant?

Motif analysis provides a first layer of validation by checking whether predicted enhancers contain binding motifs for transcription factors relevant to the cell type. Integration with genetic association data provides convergent evidence if the enhancer contains trait-associated variants. CRISPR perturbation provides the strongest validation by demonstrating that disrupting the enhancer changes target gene expression. Prioritize linkages with convergent evidence from multiple approaches.

### What are the most common sources of false positive linkages?

Cell type composition confounding is the most common source of false positives. When linkages are computed across all cells without stratification, the results reflect cell type proportions instead of regulatory relationships. Batch effects can also create spurious correlations. Sparse data can produce unstable estimates, particularly for rare cell types. Always compute linkages within cell types, include technical covariates, and validate results with independent approaches.

### How do I integrate enhancer-gene linkages with GWAS data to interpret disease-associated variants?

Overlap the disease-associated variants with predicted enhancer peaks, then examine the predicted target genes for biological relevance to the disease. The Enhlink study demonstrated this approach by coupling linkage inference with eQTL analysis to identify a putative super-enhancer in striatal neurons [12](https://pubmed.ncbi.nlm.nih.gov/39223609). The cardiac atlas study used enhancer-to-gene linkage maps to refine the association of cardiomyopathy risk loci to downstream target genes [17](https://doi.org/10.1186/s13059-026-04061-7). This integration requires careful handling of linkage disequilibrium and multiple testing.

### What should I do if my predicted linkages are not supported by motif analysis?

First, verify that the relevant transcription factor is expressed in the same cells where the enhancer is accessible. The transcription factor may bind indirectly through cooperative interactions with other factors, so the motif may not be directly present in the enhancer. Check whether the transcription factor has a known motif in the relevant database. Consider that the predicted linkage may be incorrect, and examine whether the correlation is driven by cell type composition or batch effects.

## Related Bioinformatics Guides

- [Single-Cell Multi-Omics Integration: Methods and Applications](/knowledge/bioinformatics/single-cell-multi-omics-integration-methods-and-applications)
- [Multi-Omics Integration: A Practical Guide to Combining Data Types](/knowledge/bioinformatics/multi-omics-integration-a-practical-guide-to-combining-data-types)
- [Spatial Transcriptomics vs. Single-Cell RNA Sequencing: Which Approach Fits Your Research?](/knowledge/bioinformatics/spatial-transcriptomics-vs-single-cell-rna-sequencing-which-approach-fits-your-research)
- [Single-Cell RNA Velocity: Inferring Cellular Dynamics](/knowledge/bioinformatics/single-cell-rna-velocity-inferring-cellular-dynamics)
- [RNA Sequencing Data Analysis: From Raw Reads to Differential Expression](/knowledge/bioinformatics/rna-sequencing-data-analysis-from-raw-reads-to-differential-expression)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)


## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [Unveiling migraine subtype heterogeneity and risk loci: integrated genome-wide association study and single-cell transcriptomics discovery.](https://pubmed.ncbi.nlm.nih.gov/40826382). The journal of headache and pain, 2025.
- [Decoding the epithelial-stromal interactome in allergic rhinitis through single-cell multiomics integration.](https://pubmed.ncbi.nlm.nih.gov/42435887). The Journal of allergy and clinical immunology, 2026.
- [Uncovering the molecular networks of ferroptosis in the pathogenesis of type 2 diabetes and its complications: a multi-omics investigation.](https://pubmed.ncbi.nlm.nih.gov/39716081). Molecular medicine (Cambridge, Mass.), 2024.
- [Multi-omics integration reveals Chr1 associated QTL mediating backfat thickness in pigs.](https://pubmed.ncbi.nlm.nih.gov/41111148). Journal of animal science and biotechnology, 2025.
- [Chromatin accessibility module identified by single-cell sequencing underlies the diagnosis and prognosis of hepatocellular carcinoma.](https://pubmed.ncbi.nlm.nih.gov/40606924). World journal of hepatology, 2025.
- [Enhlink infers distal and context-specific enhancer-promoter linkages.](https://pubmed.ncbi.nlm.nih.gov/39223609). Genome biology, 2024.
- [Predictive modeling of single-cell DNA methylome data enhances integration with transcriptome data.](https://pubmed.ncbi.nlm.nih.gov/33219054). Genome research, 2021.
- [Enhlink infers distal and context-specific enhancer-promoter linkages.](https://pubmed.ncbi.nlm.nih.gov/37214950). bioRxiv : the preprint server for biology, 2023.
- [Integrative single-cell multi-omics network analysis to elucidate epigenetic regulation in neurodevelopmental disorders.](https://doi.org/10.1016/j.slast.2026.100395). 2026.
- [Multi-omics and artificial intelligence for precision drug discovery and potential clinical applications.](https://doi.org/10.1038/s41392-026-02631-6). 2026.
- [An integrative single-nucleus multiomic atlas of the human left ventricle identifies gene regulatory network dynamics across cardiac development, aging, and disease.](https://doi.org/10.1186/s13059-026-04061-7). 2026.
- [Single-cell MultiOmics and spatial transcriptomics demonstrate neuroblastoma developmental plasticity.](https://doi.org/10.1016/j.devcel.2025.04.013). Developmental Cell, 2025.
- [Single-cell multiomics data integration and generation with scPairing](https://doi.org/10.1016/j.crmeth.2025.101211). bioRxiv, 2025.
- [Single-cell multiomics reveals the interplay of clonal evolution and cellular plasticity in hepatoblastoma](https://doi.org/10.1038/s41467-024-47280-x). Nature Communications, 2024.
- [Single-cell multiomics reveals the oscillatory dynamics of mRNA metabolism and chromatin accessibility during the cell cycle.](https://doi.org/10.1016/j.celrep.2025.116089). Cell Reports, 2025.
- [Single-cell spatial multi-omics and deep learning dissect enhancer-driven gene regulatory networks in liver zonation](https://doi.org/10.1038/s41556-023-01316-4). Nature Cell Biology, 2024.
- [Gene regulatory landscape dissected by single-cell four-omics sequencing](https://doi.org/10.1038/s41586-026-10322-z). Nature, 2026.
- [Linking chromatin acylation mark-defined proteome and genome in living cells](https://doi.org/10.1016/j.cell.2023.02.007). Cell, 2023.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.