Regulatory Potential Scoring in Single-Cell ATAC-Seq: Linking Peaks to Gene Expression
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Regulatory potential scoring in single-cell ATAC-seq (scATAC-seq) computationally infers the likelihood that an accessible chromatin region (peak) regulates a specific gene, bridging the gap between observed chromatin accessibility and unobserved gene regulation. This is achieved by integrating genomic distance, co-accessibility patterns across cells, and sometimes matched gene expression data.
- The core problem addressed is that peak accessibility alone is insufficient to determine gene regulation, especially for distal elements like enhancers, which can be located far from their target genes and regulate multiple genes. Regulatory potential scoring provides peak-gene pair resolution, distinguishing it from gene activity scores that aggregate accessibility around a gene.
- Key computational principles include genomic distance decay, where closer peaks are weighted more heavily, and co-accessibility analysis, which identifies peaks that show coordinated opening and closing with a gene's promoter across cells. Advanced methods integrate matched scRNA-seq data to directly predict gene expression from accessibility patterns.
- Practical workflows involve rigorous quality control of scATAC-seq data, cell type identification for context-specific analysis, careful selection of a regulatory potential method (e.g., ArchR for distance-weighted scores, Cicero for co-accessibility, SCARlink for expression prediction), and crucial validation of predicted links using orthogonal evidence like promoter capture Hi-C or eQTLs.
- Common failure patterns include overinterpreting co-accessibility as direct regulation, ignoring the dropout problem inherent in scATAC-seq sparsity, using inappropriate distance windows, and applying methods outside their valid data or biological domain, necessitating careful parameter selection and validation.
Single-cell ATAC-seq (scATAC-seq) measures chromatin accessibility in individual cells, but accessibility at a peak does not directly tell you which gene that peak regulates. Regulatory potential scoring addresses this problem by converting peak-level accessibility into gene-level scores that reflect the likelihood that a given cis-regulatory element influences the expression of a nearby or distal target gene. This article explains how regulatory potential scores are computed, which tools and workflows are available, how to interpret and validate the results, and what limitations you should account for when using these scores in your own single-cell epigenomics research.
The intended reader is a biology student, researcher, laboratory professional, or life-science practitioner who has generated or plans to generate scATAC-seq data and needs a practical framework for linking accessible chromatin regions to putative target genes. The focus is on the computational steps, the decisions you must make at each stage, and the quality checks that separate a defensible regulatory prediction from a spurious correlation.
What Regulatory Potential Scoring Does in scATAC-Seq Analysis
Regulatory potential scoring is a computational strategy that assigns a quantitative value to the relationship between an accessible chromatin region and a candidate target gene. The score is designed to answer a specific question: given that a peak is accessible in a particular cell or cell type, how likely is it that this peak regulates the expression of a nearby gene?
The underlying logic is that cis-regulatory elements such as enhancers and promoters exert their effects on genes through physical proximity in three-dimensional space, even when the linear genomic distance is large. In scATAC-seq data, you observe the accessibility state of each peak in each cell, but you do not directly observe which gene is being regulated. Regulatory potential scoring bridges this gap by combining genomic distance, co-accessibility patterns, and sometimes gene expression data from matched scRNA-seq assays.
The score itself is not a direct measurement of regulatory activity. It is a computational inference based on several assumptions about how chromatin accessibility relates to gene expression. Understanding these assumptions is critical because they determine the interpretation limits of the score.
The Core Problem: Peak Accessibility Does Not Equal Gene Expression
A peak in scATAC-seq data indicates that a region of chromatin was accessible to the Tn5 transposase in a particular cell. This accessibility is a necessary condition for many regulatory interactions, but it is not sufficient evidence that the region is actively regulating a gene. A peak could be accessible for reasons unrelated to the expression of a nearby gene, including the presence of structural chromatin features, the binding of proteins that maintain an open state without activating transcription, or technical artifacts from the assay itself.
The gap between accessibility and expression is especially pronounced for distal regulatory elements. Promoters are relatively straightforward to assign to genes because they overlap the transcription start site. Enhancers, by contrast, can be located hundreds of kilobases away from their target genes and can regulate multiple genes simultaneously. Without additional information, you cannot determine which gene an enhancer-like peak controls.
Regulatory potential scoring addresses this gap by integrating multiple lines of evidence. The most common approach uses a distance-weighted function where peaks closer to the transcription start site receive higher scores. More sophisticated approaches incorporate co-accessibility across cells, where peaks that open and close in concert with a gene's promoter are more likely to regulate that gene.
How Regulatory Potential Differs from Gene Activity Scores
Regulatory potential scoring is often confused with gene activity scoring, but the two approaches answer different questions. Gene activity scores, such as those implemented in tools like ArchR's addGeneScoreMatrix function, aggregate accessibility across all peaks within a defined window around a gene and produce a single value that approximates the gene's regulatory activity. These scores are useful for visualization and for comparing regulatory states across cell types, but they do not identify which specific peaks are responsible for the activity.
Regulatory potential scoring, by contrast, assigns a score to each peak-gene pair. The output is a matrix where each row is a peak, each column is a gene, and each cell contains a score representing the strength of the predicted regulatory link. This peak-level resolution is essential if your goal is to identify candidate cis-regulatory elements for downstream validation, such as CRISPR perturbation or reporter assays.
The distinction matters for practical decisions. If you need to annotate cell types based on regulatory state, gene activity scores may be sufficient. If you need to prioritize specific enhancers for experimental validation, you need regulatory potential scores that rank individual peaks.
Core Principles of Regulatory Potential Computation
The computation of regulatory potential scores rests on several core principles that are shared across most implementations. These principles determine what the score means and how you should interpret it.
Genomic Distance as a Primary Signal
The most fundamental principle is that regulatory elements tend to act on nearby genes. This is not an absolute rule, but it is a strong prior that improves the accuracy of regulatory predictions. Most regulatory potential methods use a distance decay function where the contribution of a peak to a gene's score decreases as the genomic distance between the peak and the transcription start site increases.
The specific form of the distance function varies between tools. Some use an exponential decay, others use a power-law decay, and still others use a fixed window with a hard cutoff. The choice of function affects the relative weighting of proximal versus distal peaks and can change the ranking of candidate regulatory elements.
A common implementation, used in tools like Cicero and ArchR, computes the regulatory potential as the sum of accessibility values weighted by a function of distance. The exact parameters of the distance function are often user-adjustable, which means you need to make an explicit decision about how far from a gene you are willing to look for regulatory elements.
Co-Accessibility Across Cells
A second principle is that a regulatory element and its target gene should show coordinated accessibility across cells. If a peak is accessible in the same cells where the gene's promoter is accessible, and inaccessible in cells where the promoter is closed, this co-variation is evidence of a functional connection.
Co-accessibility is computed by measuring the correlation between the accessibility profiles of two peaks across all cells in the dataset. Tools like Cicero construct a graph where peaks are nodes and co-accessibility scores are edges, then use this graph to identify cis-regulatory interactions. The assumption is that physically interacting regulatory elements will show correlated accessibility because they are opened and closed by the same regulatory programs.
This principle is particularly valuable for identifying distal enhancers that would be missed by distance-based approaches alone. An enhancer located far from its target gene may still show strong co-accessibility with the gene's promoter if the two elements are regulated together.
Integration with Gene Expression Data
When matched scRNA-seq data are available from the same cells, regulatory potential scoring can incorporate direct evidence of the relationship between accessibility and expression. This is the approach taken by methods like SCARlink, which uses regularized Poisson regression to predict gene expression from tile-level accessibility data across a gene locus.
The integration of expression data transforms the problem from an unsupervised correlation analysis to a supervised prediction task. Instead of asking which peaks are co-accessible with a promoter, the method asks which peaks best predict the observed expression level of a gene. This approach can capture nonlinear relationships and interactions between multiple regulatory elements that would be missed by pairwise correlation methods.
The tradeoff is that this approach requires matched multi-ome data where both accessibility and expression are measured in the same cells. If you only have scATAC-seq data, you cannot use expression-based methods directly, although you may be able to use a reference scRNA-seq dataset for label transfer and pseudo-multi-ome analysis.
At a Glance: Regulatory Potential Scoring Methods
The table below summarizes the main approaches to regulatory potential scoring, their data requirements, and their primary use cases. Use this table to make an initial decision about which method fits your data and your research question.
| Method | Data Requirement | Core Principle | Primary Use Case | Key Limitation |
|---|---|---|---|---|
| Distance-weighted scoring (ArchR gene scores) | scATAC-seq only | Genomic distance decay from transcription start site | Cell type annotation and regulatory state visualization | Does not identify specific peak-gene pairs |
| Co-accessibility analysis (Cicero) | scATAC-seq only | Coordinated accessibility across cells | Identifying candidate cis-regulatory interactions | Requires sufficient cell numbers for stable correlations |
| Expression prediction (SCARlink) | Matched scRNA-seq and scATAC-seq (multi-ome) | Regularized regression of expression on accessibility | Identifying enhancers that predict expression | Requires multi-ome data, computationally intensive |
| Deep learning with Shapley values (coTF-reg, scEpiLock) | scATAC-seq with optional scRNA-seq | Neural network prediction with interpretability analysis | Identifying cooperative TF pairs and refining peak boundaries | Requires training data and computational expertise |
Practical Workflow for Computing Regulatory Potential Scores
The workflow for computing regulatory potential scores follows a standard sequence of steps that you should adapt to your specific data and research question. Each step involves decisions that affect the quality and interpretability of the final scores.
Step 1: Quality Control and Preprocessing
Before you can compute regulatory potential scores, you need a high-quality scATAC-seq dataset that has passed appropriate quality control. The quality of your input data directly determines the reliability of your regulatory predictions, so this step deserves careful attention.
Quality control for scATAC-seq typically includes filtering cells based on the number of unique fragments, the fraction of fragments in peaks, the transcription start site enrichment score, and the ratio of reads in promoter regions to reads in other genomic regions. Cells with very low fragment counts are likely to produce noisy accessibility profiles that will generate spurious co-accessibility correlations.
The choice of peak calling method also matters. Most scATAC-seq pipelines use MACS2 to call peaks on pseudo-bulk data, but the sensitivity and specificity of peak calling affect downstream regulatory potential scores. Some methods, such as SCARlink, avoid the dependence on peak calling by using tile-level accessibility data directly, which can improve performance on datasets where peak boundaries are uncertain.
The sparsity of scATAC-seq data is a major challenge at this stage. Each cell contains only a small fraction of the total accessible regions in the genome, which means that the absence of a fragment at a peak does not necessarily mean the region was inaccessible. This dropout problem can inflate co-accessibility correlations and produce false regulatory predictions. Methods that explicitly model dropout, such as MOCHA's zero-inflated statistical approach, can mitigate this issue.
Step 2: Cell Type Identification and Labeling
Regulatory potential scores are most interpretable when computed within defined cell types instead of across the entire dataset. The regulatory relationships between peaks and genes can differ substantially between cell types, and averaging across a heterogeneous population will obscure these differences.
Cell type identification in scATAC-seq data is typically performed by clustering cells based on their accessibility profiles and then annotating the clusters using marker genes or label transfer from a reference scRNA-seq dataset. The label transfer approach is common because scRNA-seq cell type annotations are more mature, but the inherent differences between the two modalities create challenges.
The sparsity and high dimensionality of scATAC-seq data make label transfer difficult, and methods like scCorrect have been developed specifically to address the modal discrepancies between scRNA-seq and scATAC-seq data. These methods align the two datasets and then correct erroneous annotations produced during the initial transfer.
For regulatory potential scoring, the key decision is whether to compute scores on the full dataset or within cell type subsets. Computing within cell types is generally preferred because it captures cell type specific regulatory relationships, but it requires sufficient cells per type to produce stable estimates. A cell type with only a few hundred cells may not provide enough statistical power for co-accessibility analysis.
Step 3: Choosing a Regulatory Potential Method
The choice of method depends on your data availability and your research question. If you only have scATAC-seq data, your options are distance-weighted scoring or co-accessibility analysis. If you have matched multi-ome data, you can use expression prediction methods that directly model the relationship between accessibility and expression.
For distance-weighted scoring, tools like ArchR provide a straightforward implementation that is suitable for most applications. The method computes a gene score for each gene by aggregating accessibility across peaks within a defined window, weighted by distance. The output is a gene by cell matrix that can be used for downstream analysis.
For co-accessibility analysis, Cicero is the most widely used tool. It constructs a graph of peak co-accessibility across cells and then identifies cis-regulatory interactions where a distal peak shows coordinated accessibility with a promoter. The output includes co-accessibility scores for peak pairs, which can be used to link enhancers to target genes.
For expression prediction, SCARlink represents a more recent approach that uses regularized Poisson regression on tile-level accessibility data. This method jointly models all regulatory effects at a gene locus, avoiding the limitations of pairwise gene-peak correlations and the dependence on peak calling. SCARlink has been shown to outperform existing gene scoring methods for imputing gene expression from chromatin accessibility on high-coverage multi-ome datasets.
Step 4: Computing and Interpreting Scores
Once you have chosen a method, the computation of regulatory potential scores is typically automated by the tool. However, you need to understand the output format and the meaning of the scores to interpret them correctly.
For distance-weighted methods, the output is a gene score matrix where each value represents the aggregate regulatory potential of all peaks assigned to a gene. Higher scores indicate greater predicted regulatory activity, but the absolute value has no biological meaning. The scores are useful for comparing relative regulatory activity across genes or across cell types.
For co-accessibility methods, the output is a set of peak pairs with co-accessibility scores. A high co-accessibility score between a distal peak and a promoter suggests that the distal peak may regulate the promoter's gene. The interpretation of the score depends on the threshold you choose, and you should validate any candidate interactions using orthogonal evidence.
For expression prediction methods, the output includes both predicted expression values and feature importance scores that indicate which peaks contribute most to the prediction. Shapley value analysis can identify cell type specific enhancers that are validated by orthogonal methods such as promoter capture Hi-C.
Step 5: Validation of Predicted Regulatory Links
Regulatory potential scores are predictions, not measurements. Before you invest time and resources in experimental validation, you should assess the evidence supporting each predicted regulatory link.
Several types of orthogonal evidence can support a regulatory potential prediction. Promoter capture Hi-C data provides direct evidence of physical proximity between a peak and a promoter in three-dimensional space. Fine-mapped expression quantitative trait loci (eQTLs) provide evidence that genetic variation at a peak affects the expression of a nearby gene. Genome-wide association study (GWAS) variants that fall within a predicted enhancer and are linked to a relevant phenotype provide additional support.
The enrichment of predicted enhancers in fine-mapped eQTLs and GWAS variants is a strong validation signal. SCARlink-predicted enhancers have been shown to be 11 to 15 times enriched in fine-mapped eQTLs and 5 to 12 times enriched in fine-mapped GWAS variants, providing confidence that the method identifies functionally relevant regulatory elements.
For transcription factor analysis, you can validate predicted TF-target relationships using ChIP-seq data. If a transcription factor is predicted to regulate a gene through a specific enhancer, ChIP-seq data showing the factor bound at that enhancer provides direct support for the prediction.
Tools and Their Tradeoffs
The choice of tool for regulatory potential scoring involves tradeoffs between data requirements, computational cost, interpretability, and performance. Understanding these tradeoffs helps you select the appropriate tool for your specific situation.
ArchR for Distance-Weighted Gene Scores
ArchR is a widely used R package for scATAC-seq analysis that includes a gene score module. The method computes gene scores by aggregating accessibility across peaks within a user-defined window around each gene, with distance-based weighting. The implementation is computationally efficient and works with scATAC-seq data alone.
The main limitation of ArchR gene scores is that they do not identify specific peak-gene pairs. The score is a summary of all accessibility in the gene's vicinity, which means you cannot use it to prioritize individual enhancers for validation. For cell type annotation and regulatory state comparison, this limitation is acceptable. For enhancer discovery, you need a method that provides peak-level resolution.
Cicero for Co-Accessibility Analysis
Cicero is an R package that constructs co-accessibility networks from scATAC-seq data. The method first aggregates cells into groups to reduce sparsity, then computes co-accessibility scores between pairs of peaks, and finally identifies cis-regulatory interactions where a distal peak is co-accessible with a promoter.
The main advantage of Cicero is that it provides peak-level resolution and can identify distal enhancers that would be missed by distance-based methods. The main limitation is that co-accessibility is a correlation-based measure, and correlation does not imply causation. A peak may be co-accessible with a promoter because both are regulated by the same upstream factor, even if the peak does not directly regulate the promoter's gene.
Cicero has been used in studies of diabetic cardiomyopathy to calculate cis-coaccessibility networks and compare them to the GeneHancer database of known regulatory elements. This type of comparison provides a way to assess the biological relevance of predicted interactions.
SCARlink for Expression Prediction
SCARlink is a more recent method that uses regularized Poisson regression to predict gene expression from tile-level accessibility data in multi-ome datasets. The method jointly models all regulatory effects at a gene locus, which avoids the limitations of pairwise gene-peak correlations and the dependence on peak calling.
The main advantage of SCARlink is its superior performance in imputing gene expression from chromatin accessibility, particularly on high-coverage multi-ome datasets. The method also provides interpretable feature importance scores through Shapley value analysis, which can identify cell type specific enhancers.
The main limitation is the requirement for matched multi-ome data. If you only have scATAC-seq data, you cannot use SCARlink directly. The method is also computationally intensive, which may be a constraint for large datasets.
MOCHA for Statistical Modeling of Dropout
MOCHA is a tool that addresses the technical dropout problem in scATAC-seq data using zero-inflated statistical models. The method improves the identification of sample-specific open chromatin, mitigates false positives in single-cell analysis, and provides modules for inferring temporal gene regulatory networks from longitudinal data.
The main advantage of MOCHA is its robust handling of the sparsity that characterizes scATAC-seq data. By explicitly modeling dropout, the method reduces the false positive rate in regulatory predictions. The main limitation is that MOCHA is designed for large cohort studies and may be overkill for small datasets.
Deep Learning Approaches for Cooperative TF Analysis
Deep learning approaches such as coTF-reg and scEpiLock use neural networks to model the relationship between accessibility and expression, with interpretability methods to identify important features. coTF-reg identifies cooperative transcription factor pairs that co-regulate target genes, while scEpiLock refines peak boundaries and quantifies variant impacts.
These methods are more complex to implement and require more computational resources than simpler approaches. However, they can capture nonlinear relationships and interactions between multiple regulatory elements that would be missed by linear methods. The Shapley interaction scores from coTF-reg have identified known important TF pairs such as SOX10-TCF12 and SOX10-MYRF in oligodendrocyte differentiation.
Records and Measurements for Regulatory Potential Analysis
Keeping detailed records of your regulatory potential analysis is essential for reproducibility and for defending your conclusions. The following records should be maintained for each analysis.
Data Quality Metrics
Record the number of cells that passed quality control, the median number of unique fragments per cell, the fraction of fragments in peaks, and the transcription start site enrichment score. These metrics provide context for interpreting regulatory potential scores and help identify datasets where sparsity may compromise the analysis.
The number of cells required for reliable regulatory potential analysis depends on the method and the heterogeneity of the sample. Studies comparing bulk and single-cell ATAC-seq have determined the number of cells required to generate aggregated open chromatin profiles and to identify biologically meaningful clusters after pseudo-bulking of data. For co-accessibility analysis, you generally need more cells than for simple gene scoring because correlation estimates are unstable with small sample sizes.
Method Parameters
Record the exact parameters used for each method, including the distance window for gene scoring, the co-accessibility threshold for Cicero, the regularization parameters for SCARlink, and the number of training epochs for deep learning methods. These parameters affect the results, and different parameter choices can lead to different conclusions.
The choice of distance window is particularly important. A window that is too small will miss distal enhancers, while a window that is too large will include many peaks that are not functionally related to the gene. The optimal window depends on the genomic architecture of your system of interest and should be justified based on prior knowledge or sensitivity analysis.
Validation Results
Record the results of any validation analyses, including enrichment of predicted enhancers in eQTLs or GWAS variants, overlap with promoter capture Hi-C data, and concordance with ChIP-seq binding data. These validation results provide evidence for the biological relevance of your regulatory predictions.
The enrichment of predicted enhancers in fine-mapped eQTLs and GWAS variants is a quantitative measure that can be compared across methods and datasets. Higher enrichment indicates that the method is identifying functionally relevant regulatory elements.
Common Failure Patterns in Regulatory Potential Scoring
Several common failure patterns can compromise regulatory potential scoring. Recognizing these patterns helps you avoid them and interpret results correctly.
Overinterpretation of Co-Accessibility as Regulation
The most common failure is treating co-accessibility as evidence of direct regulation. Two peaks can be co-accessible because they are regulated by the same transcription factor, even if one does not regulate the gene associated with the other. This is a particular risk with correlation-based methods like Cicero.
To mitigate this risk, you should validate co-accessibility predictions with orthogonal evidence such as Hi-C data, eQTL analysis, or experimental perturbation. The presence of a physical interaction between the peak and the promoter provides stronger evidence for a direct regulatory relationship than co-accessibility alone.
Ignoring the Dropout Problem
The sparsity of scATAC-seq data creates a dropout problem where the absence of a fragment does not necessarily mean the region was inaccessible. Methods that do not account for dropout can produce inflated co-accessibility correlations and false regulatory predictions.
Tools like MOCHA that use zero-inflated statistical models can mitigate this issue. Alternatively, you can aggregate cells into groups before computing co-accessibility, as Cicero does, to reduce the impact of dropout.
Using Inappropriate Distance Windows
The choice of distance window for gene scoring has a major impact on the results. A window that is too small will miss distal enhancers, while a window that is too large will include many peaks that are not functionally related to the gene.
The optimal window depends on the genomic architecture of your system. Some genes have regulatory elements located hundreds of kilobases away, while others are regulated primarily by proximal elements. You should test multiple window sizes and assess the stability of your conclusions across parameter choices.
Applying Methods Outside Their Valid Domain
Each regulatory potential method has assumptions about the data and the biological system. Applying a method outside its valid domain can produce misleading results.
For example, SCARlink requires matched multi-ome data and may not perform well on low-coverage datasets. Cicero requires sufficient cells per cell type to produce stable co-accessibility estimates. Deep learning methods require training data that is representative of the system you are studying. You should verify that your data meets the requirements of the method before interpreting the results.
Limitations and Interpretation Boundaries
Regulatory potential scores are computational inferences with inherent limitations. Understanding these limitations is essential for appropriate interpretation.
Scores Are Not Measurements of Regulatory Activity
A regulatory potential score does not measure the actual regulatory activity of a peak. It is a statistical prediction based on accessibility patterns and genomic features. The score may be high for a peak that is accessible but not actively regulating a gene, and low for a peak that is regulating a gene through mechanisms not captured by the model.
The only way to confirm regulatory activity is through experimental perturbation, such as CRISPR deletion or activation of the predicted enhancer. Regulatory potential scores are useful for prioritizing candidates for such experiments, but they cannot replace them.
Cell Type Specificity Is Critical
Regulatory relationships are cell type specific. A peak that regulates a gene in one cell type may have no effect in another. Computing regulatory potential scores on a heterogeneous population will obscure these differences and may produce scores that are not representative of any cell type.
You should compute regulatory potential scores within defined cell types whenever possible. This requires sufficient cells per cell type and accurate cell type annotation.
The Reference Genome and Annotations Matter
Regulatory potential scoring depends on the reference genome and gene annotations used. Different genome builds and annotation versions can produce different results. You should use the same reference genome and annotations throughout your analysis and report these details in your methods.
The quality of gene annotations is particularly important for genes with complex structures, such as those with multiple transcription start sites. Methods that use a single transcription start site per gene may miss regulatory elements that act on alternative promoters.
Safety and Regulatory Context for Research Applications
Regulatory potential scoring is a bioinformatics analysis with no direct safety or regulatory requirements. However, the results of these analyses can inform downstream experiments that are subject to institutional oversight.
Institutional Review for Human Data
If your scATAC-seq data come from human subjects, you must ensure that the data were collected under appropriate institutional review board approval and that patient consent covers the intended use. The use of human genomic data in research is subject to privacy and confidentiality requirements.
Biosafety for Perturbation Experiments
If you use regulatory potential scores to prioritize enhancers for CRISPR perturbation experiments, you must follow institutional biosafety guidelines for the use of genome editing technologies. The specific requirements depend on the cell types and organisms involved.
Data Sharing and Reproducibility
Many journals and funding agencies require that genomic data and analysis code be shared to ensure reproducibility. You should deposit your data in appropriate public repositories such as those maintained by the National Center for Biotechnology Information and make your analysis code available through platforms like Bioconductor or Galaxy.
Professional Escalation Criteria
Regulatory potential scoring is a specialized analysis that may require consultation with experts in certain situations. The following criteria indicate when you should seek additional expertise.
When Results Are Inconsistent Across Methods
If different regulatory potential methods produce substantially different predictions for the same peaks and genes, this inconsistency suggests that the regulatory relationships are not robust. You should consult with a bioinformatics expert to understand the source of the discrepancy and determine which method is most appropriate for your data.
When Validation Fails
If predicted regulatory links show no enrichment in eQTLs or GWAS variants, and no overlap with Hi-C data, the predictions may be spurious. A bioinformatics expert can help you diagnose whether the problem is in the data quality, the method parameters, or the biological system.
When Extending to New Biological Systems
If you are applying regulatory potential scoring to a biological system where the methods have not been validated, you should consult with experts who have experience with similar systems. The assumptions underlying regulatory potential methods may not hold in all biological contexts.
When Results Will Guide Clinical Decisions
If regulatory potential scores will be used to guide clinical decisions, such as prioritizing variants for diagnostic testing or therapeutic targeting, you should involve experts in clinical genomics and regulatory affairs. The evidence standards for clinical applications are higher than for basic research.
Frequently Asked Questions
What is the difference between a gene activity score and a regulatory potential score?
A gene activity score aggregates accessibility across all peaks near a gene into a single value that approximates the gene's regulatory activity. A regulatory potential score assigns a value to each peak-gene pair, indicating the strength of the predicted regulatory link between that specific peak and that specific gene. Gene activity scores are useful for cell type annotation, while regulatory potential scores are needed to identify candidate enhancers for validation.
Can I compute regulatory potential scores with scATAC-seq data alone?
Yes, you can compute distance-weighted gene scores and co-accessibility scores with scATAC-seq data alone. Tools like ArchR and Cicero work with scATAC-seq data without requiring matched expression data. However, methods that predict expression from accessibility, such as SCARlink, require matched multi-ome data where both accessibility and expression are measured in the same cells.
How many cells do I need for reliable regulatory potential scoring?
The number of cells required depends on the method and the heterogeneity of the sample. Co-accessibility analysis requires more cells than simple gene scoring because correlation estimates are unstable with small sample sizes. Studies comparing bulk and single-cell ATAC-seq have determined the number of cells required to generate aggregated open chromatin profiles and to identify biologically meaningful clusters after pseudo-bulking of data. For heterogeneous samples, you need enough cells per cell type to produce stable estimates within each type.
How do I validate a predicted regulatory link?
You can validate predicted regulatory links using several types of orthogonal evidence. Promoter capture Hi-C data provides evidence of physical proximity between a peak and a promoter. Fine-mapped eQTLs provide evidence that genetic variation at a peak affects the expression of a nearby gene. GWAS variants that fall within a predicted enhancer and are linked to a relevant phenotype provide additional support. ChIP-seq data showing a transcription factor bound at the predicted enhancer supports TF-target predictions. The strongest validation comes from experimental perturbation, such as CRISPR deletion or activation of the enhancer.
What is the dropout problem in scATAC-seq and how does it affect regulatory potential scoring?
The dropout problem refers to the sparsity of scATAC-seq data, where each cell contains only a small fraction of the total accessible regions in the genome. The absence of a fragment at a peak does not necessarily mean the region was inaccessible. This dropout can inflate co-accessibility correlations and produce false regulatory predictions. Methods that explicitly model dropout, such as MOCHA's zero-inflated statistical approach, can mitigate this issue. Aggregating cells into groups before computing co-accessibility, as Cicero does, also reduces the impact of dropout.
What are Shapley values and how are they used in regulatory potential analysis?
Shapley values are a method from cooperative game theory that assigns credit to each feature for a prediction. In regulatory potential analysis, Shapley values are used to identify which peaks contribute most to the predicted expression of a gene. Shapley interaction scores can identify cooperative transcription factor pairs that jointly regulate target genes. These scores have been used to identify known important TF pairs such as SOX10-TCF12 and SOX10-MYRF in oligodendrocyte differentiation.
How do I choose between different regulatory potential methods?
The choice of method depends on your data availability and your research question. If you only have scATAC-seq data, use distance-weighted gene scoring for cell type annotation and co-accessibility analysis for identifying candidate regulatory interactions. If you have matched multi-ome data, use expression prediction methods like SCARlink for more accurate identification of enhancers that predict expression. Consider the computational cost, the interpretability of the output, and the validation evidence supporting each method.
What are the main limitations of regulatory potential scores?
Regulatory potential scores are computational inferences, not measurements of regulatory activity. A high score does not prove that a peak regulates a gene. The scores are cell type specific and depend on the reference genome and annotations used. Co-accessibility does not imply causation, and the dropout problem can produce spurious correlations. Experimental validation is required to confirm any predicted regulatory link.
Related Bioinformatics Guides
- Single-Cell Annotation: A Workflow for Cell Type Identification
- Single-Cell Sequencing Depth: How Much Is Enough?
- RNA-Seq vs Microarray: Choosing the Right Gene Expression Profiling Platform
- Single-Cell RNA Sequencing Quality Control: A Practical Guide to Filtering and Metrics
- Single-Cell ATAC-Seq Bioinformatics
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- Single-cell insights: pioneering an integrated atlas of chromatin accessibility and transcriptomic landscapes in diabetic cardiomyopathy.. Cardiovascular diabetology, 2024.
- Single-cell multi-ome regression models identify functional and disease-associated enhancers and enable chromatin potential analysis.. Nature genetics, 2024.
- Benchmarking Algorithms for Gene Set Scoring of Single-cell ATAC-seq Data.. Genomics, proteomics & bioinformatics, 2024.
- CellSNVReg: a multidimensional resource for SNV-mediated regulatory perturbations in single-cell and spatial omics.. Nucleic acids research, 2026.
- Single-cell multi-omics analysis reveals cooperative transcription factors for gene regulation in oligodendrocytes.. bioRxiv : the preprint server for biology, 2024.
- scEpiLock: A Weakly Supervised Learning Framework for cis-Regulatory Element Localization and Variant Impact Quantification for Single-Cell Epigenetic Data.. Biomolecules, 2022.
- CoTF-reg reveals cooperative transcription factors in oligodendrocyte gene regulation using single-cell multi-omics.. Communications biology, 2025.
- Multi-omics approach reveals the impact of prognosis model-related genes on the tumor microenvironment in medulloblastoma.. Frontiers in oncology, 2025.
- scATAC-seq generates more accurate and complete regulatory maps than bulk ATAC-seq. Scientific Reports, 2025.
- Integrative scATAC-seq and mtDNA mutation analysis reveals disease-driven regulatory aberrations in AML.. Science Bulletin, 2025.
- A scATAC-seq atlas of stasis zone in rat skin burn injury wound process. Frontiers in Cell and Developmental Biology, 2025.
- Inferring Gene Regulatory Network Based on scATAC-seq Data with Gene Perturbation. bioRxiv, 2024.
- scCorrect: Cross-modality label transfer from scRNA-seq to scATAC-seq using domain adaptation.. Analytical Biochemistry, 2025.
- MOCHA’s advanced statistical modeling of scATAC-seq data enables functional genomic inference in large human cohorts. Nature Communications, 2024.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.