Integrating scATAC-seq with scRNA-seq: A Practical Guide to Multi-Omics Analysis
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Integrating scRNA-seq and scATAC-seq data allows researchers to link chromatin accessibility to gene expression, providing a more comprehensive understanding of cellular identity and function by revealing regulatory mechanisms not discernible from either modality alone.
- Key integration methods include Seurat's weighted nearest neighbor (WNN) analysis for anchored integration, MOFA+ for identifying shared and dataset-specific factors of variation using a variational Bayesian framework, and LIGER for scalable integration via nonnegative matrix factorization.
- A critical preprocessing step for scATAC-seq data integration is the generation of a gene activity matrix, which scores gene accessibility based on promoter and regulatory element openness; the quality of this matrix directly impacts downstream integration accuracy.
- Successful integration requires rigorous quality control of individual datasets (e.g., unique molecular identifiers, gene detection for RNA-seq; unique fragments, TSS enrichment for ATAC-seq) and careful evaluation of the integrated data for preserved biological heterogeneity and removed technical batch effects.
- Common integration failures include overcorrection that removes biological variation, undercorrection that leaves batch effects, poor gene activity matrix quality, misalignment of rare cell populations, and trajectory fracture, necessitating careful parameter tuning and validation.
- Interpretation of integrated data provides correlative evidence; functional validation experiments are essential to establish causal relationships between observed chromatin accessibility patterns and gene expression changes.
Single-cell RNA sequencing (scRNA-seq) and single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq) measure different layers of cellular information. scRNA-seq captures transcriptome-wide gene expression at single-cell resolution, while scATAC-seq defines transcriptional and epigenetic changes by analyzing chromatin accessibility at the single-cell level. Integrating these two data types allows researchers to link chromatin accessibility to gene expression and infer regulatory mechanisms. This article provides a practical workflow for integration using methods such as Seurat's weighted nearest neighbor (WNN) analysis, MOFA+, and LIGER, with code examples and interpretation tips for identifying cell-type-specific regulatory programs. The intended readers are biology students, researchers, laboratory professionals, and life-science practitioners who need concrete decisions about data inputs, workflow choices, quality checks, reproducibility, and interpretation limits.
Understanding the Biological Rationale for Multi-Omics Integration
The central premise of integrating scRNA-seq and scATAC-seq is that gene expression and chromatin accessibility are complementary views of cellular state. Gene expression tells you which transcripts are present in a cell at a given moment. Chromatin accessibility tells you which regions of the genome are open and potentially available for transcription factor binding and regulatory activity. A gene can be transcribed only when its regulatory elements are accessible, but accessibility alone does not guarantee transcription. The integration of these two modalities provides a more complete picture of how cells regulate their identity and function.
Single-cell studies have transformed the ability to characterize cell states, but deep biological understanding requires more than a taxonomic listing of clusters. As new methods arise to measure distinct cellular modalities, a key analytical challenge is to integrate these datasets to better understand cellular identity and function. The strategy to anchor diverse datasets together enables integration of single-cell measurements across scRNA-seq technologies and across different modalities. For example, anchoring scRNA-seq experiments with scATAC-seq can explore chromatin differences in closely related cell subsets, and harmonizing in situ gene expression and scRNA-seq datasets allows transcriptome-wide imputation of spatial gene expression patterns.
The practical value of integration appears in multiple research contexts. In human developmental hematopoiesis, integrative analysis of chromatin accessibility and gene expression revealed extensive epigenetic but not transcriptional priming of hematopoietic stem cells and multipotent progenitors prior to their lineage commitment. This finding would not have been possible with either assay alone. In breast cancer research, integration of scRNA-seq and scATAC-seq identified a novel cell subpopulation with abnormally high expression of CXCL14 in positive lymph nodes, and cell trajectory analysis revealed that CXCL14 expression increased in late pseudo-time. The scATAC-seq data identified several transcription factors that may regulate lymph node metastasis. In clear cell renal cell carcinoma, integration of scRNA-seq and scATAC-seq identified transcriptional regulatory features of sarcomatoid differentiated tumor cells, primarily evidenced by active binding with FOS/JUND, FOSL1/JUN, and FOSL2, and identified DST, FRMD4A, and PREX2 as tumor markers.
The biological rationale extends to understanding gene regulatory networks. Cardiac embryonic development is governed by precise spatiotemporal gene regulation, and single-cell multi-omics approaches have enabled construction of high-resolution cardiac cell atlases, revealing previously unrecognized cellular heterogeneity and transitional states. Integrative analyses of transcriptomic and chromatin accessibility data have provided insights into lineage commitment, key transcription factors, enhancer-promoter interactions, and dynamic gene regulatory networks. In atherosclerosis research, integration of transcriptomic, epigenomic, proteomic, and spatial data has clarified key mechanisms of disease progression and identified cell-specific molecular pathways responsive to targeted therapy.
At a Glance: Integration Methods and Their Characteristics
| Method | Core Approach | Input Requirements | Strengths | Limitations |
|---|---|---|---|---|
| Seurat WNN | Weighted nearest neighbor analysis using anchored dataset integration | Matched or unmatched scRNA-seq and scATAC-seq, gene activity matrix for ATAC data | Established workflow, extensive documentation, handles multiple modalities | Requires careful parameter tuning, gene activity matrix quality affects results |
| MOFA+ | Multi-Omics Factor Analysis using a variational Bayesian framework | Multiple omics matrices from same or different cells | Identifies shared and dataset-specific factors, handles missing data | Requires factor number selection, interpretation can be complex |
| LIGER | Linked Inference of Genomic Experimental Relationships using integrative nonnegative matrix factorization | scRNA-seq and scATAC-seq with shared features | Fast, scalable to large datasets, good for batch correction | May oversimplify complex regulatory relationships |
| scDART | Deep learning framework that integrates data and learns cross-modality relationships simultaneously | Unmatched scRNA-seq and scATAC-seq | Preserves cell trajectories, applicable to trajectory inference | Requires deep learning expertise, computational resources |
| scBridge | Heterogeneous integration that identifies reliable cells with smaller omics differences first | scRNA-seq and scATAC-seq | Exploits cell heterogeneity to improve integration, outperforms six baselines on seven datasets | Newer method, less community validation |
| scMI | Heterogeneous graph embedding with inter-type attention mechanism | scRNA-seq and scATAC-seq | Learns cross-modality relationships without motif databases, handles unmatched data | Requires graph neural network implementation |
Core Principles of Data Integration
Data Modalities and Their Distinct Characteristics
scRNA-seq measures gene expression by quantifying transcript abundance in individual cells. The data are sparse, with many genes showing zero counts in individual cells due to dropout events. scATAC-seq measures chromatin accessibility by identifying regions of open chromatin through transposase-mediated fragmentation and sequencing. The data are also sparse, with peaks representing accessible regions across the genome.
The two data types have different feature spaces. scRNA-seq features are genes, while scATAC-seq features are genomic peaks. To integrate these data, researchers must either map peaks to genes or use a shared feature space. The most common approach is to create a gene activity matrix from scATAC-seq data, where each gene is assigned a score based on the accessibility of its promoter and nearby regulatory elements. However, pre-defined gene activity matrices are often of low quality and do not reflect the dataset-specific relationship between the two data modalities.
The Challenge of Unmatched Data
A significant practical challenge is that scRNA-seq and scATAC-seq are often performed on different cells from the same sample or from different batches. This creates unmatched data where there is no one-to-one correspondence between cells across modalities. Existing methods tend to use a pre-defined gene activity matrix to convert scATAC-seq data into scRNA-seq data, but this approach has limitations. Methods like scDART address this by integrating unmatched scRNA-seq and scATAC-seq data while learning cross-modality relationships simultaneously.
The integration of multi-omics data remains one of the most difficult tasks in bioinformatics. The difficulty arises from several sources: different technical noise profiles, different feature spaces, different cell capture efficiencies, and batch effects between experiments. A successful integration must reduce the omics difference while keeping the cell type difference. This is daunting because cell heterogeneity means that even cells of the same omics and type have various features, making the two differences less significant.
The Role of Cell Heterogeneity
Cell heterogeneity is often viewed as a nuisance in data integration, but it can be exploited to improve results. The omics difference varies in cells, and cells with smaller omics differences are easier to integrate. Methods like scBridge integrate cells in a heterogeneous manner by iterating between identifying reliable scATAC-seq cells that have smaller omics differences and integrating those reliable cells with scRNA-seq data to narrow the omics gap, thus benefiting the integration for the rest of the cells.
Practical Workflow for Data Integration
Step 1: Quality Control and Preprocessing
Before integration, each dataset must undergo quality control independently. For scRNA-seq, standard quality metrics include the number of unique molecular identifiers, the number of genes detected, and the percentage of mitochondrial reads. For scATAC-seq, quality metrics include the number of unique fragments, the fraction of fragments in peaks, and the transcription start site enrichment score.
The quality control thresholds should be determined based on the specific dataset and experimental design. There are no universal thresholds that apply to all experiments. Researchers should examine the distribution of quality metrics and set thresholds that remove clear outliers while retaining biologically relevant cells. The Bioconductor project provides official package documentation and workflow guidance for reproducible genomic analysis, including quality control procedures for single-cell data.
For scATAC-seq, the gene activity matrix is a critical preprocessing step. This matrix assigns each gene a score based on the accessibility of its promoter and nearby regulatory elements. The quality of this matrix directly affects the quality of the integration. Researchers should evaluate the gene activity matrix by checking whether known marker genes show expected accessibility patterns in known cell types.
Step 2: Feature Space Alignment
The next step is to align the feature spaces of the two datasets. For scRNA-seq, the features are genes. For scATAC-seq, the features can be peaks or genes through the gene activity matrix. When using a gene activity matrix, both datasets can be represented in gene space, allowing direct comparison.
Alternative approaches use peaks as features and map genes to nearby peaks. This approach preserves more information about regulatory elements but requires careful definition of the relationship between peaks and genes. Some methods use motif databases to establish cross-modality relationships between genes from RNA-seq data and peaks from ATAC-seq data. However, these approaches are constrained by incomplete database coverage, particularly for novel or poorly characterized relationships.
Step 3: Normalization and Scaling
Both datasets should be normalized before integration. For scRNA-seq, common normalization methods include log-normalization and SCTransform. For scATAC-seq gene activity matrices, similar normalization approaches can be applied. The choice of normalization method affects the integration results, and researchers should be consistent in their approach across datasets.
After normalization, the data should be scaled to have zero mean and unit variance. This scaling ensures that no single gene or peak dominates the integration due to its magnitude. Scaling is particularly important when integrating datasets with different dynamic ranges.
Step 4: Integration Using Seurat WNN
Seurat's weighted nearest neighbor analysis is one of the most widely used integration methods. The WNN approach uses anchored dataset integration to identify shared cell states across modalities. The workflow involves:
- Create a Seurat object for each modality
- Find anchors between the datasets using canonical correlation analysis or reciprocal PCA
- Transfer data between modalities using the anchors
- Compute weighted nearest neighbors based on the relative information content of each modality
- Perform clustering and visualization on the WNN graph
The WNN approach was developed to anchor diverse datasets together, enabling integration of single-cell measurements across scRNA-seq technologies and across different modalities. The method has been demonstrated to improve over existing methods for integrating scRNA-seq data and has been used to anchor scRNA-seq experiments with scATAC-seq to explore chromatin differences in closely related cell subsets.
Step 5: Integration Using MOFA+
MOFA+ (Multi-Omics Factor Analysis) is a variational Bayesian framework that identifies factors of variation across multiple omics datasets. The method can handle matched or unmatched data and identifies both shared and dataset-specific factors. The workflow involves:
- Prepare the data matrices for each omics layer
- Specify the number of factors to infer
- Run the MOFA+ model
- Examine factor loadings to interpret biological meaning
- Use factor values for downstream analysis such as clustering and trajectory inference
MOFA+ is particularly useful when researchers want to identify the sources of variation that are shared across modalities versus those that are specific to one modality. This can reveal regulatory relationships that are not apparent from either dataset alone.
Step 6: Integration Using LIGER
LIGER (Linked Inference of Genomic Experimental Relationships) uses integrative nonnegative matrix factorization to identify shared and dataset-specific factors. The method is designed to be fast and scalable to large datasets. The workflow involves:
- Prepare the data matrices with shared features
- Run the integrative nonnegative matrix factorization
- Identify shared factors and dataset-specific factors
- Use the shared factor space for clustering and visualization
LIGER is particularly effective for batch correction when integrating datasets from different experiments or sequencing platforms. The method has been widely used in single-cell analysis workflows.
Step 7: Evaluation of Integration Quality
After integration, researchers must evaluate whether the integration successfully removed technical differences while preserving biological differences. Common evaluation approaches include:
- Visual inspection of UMAP or t-SNE plots to check for mixing of cells from different batches within the same cell type
- Calculation of batch entropy or other mixing metrics
- Verification that known marker genes show expected expression patterns in identified clusters
- Comparison of cluster composition across batches
The Galaxy Training Network provides accessible workflow training and analysis tutorials that include evaluation approaches for single-cell data integration. These tutorials emphasize reproducibility and provide practical guidance for assessing integration quality.
Options and Tradeoffs in Integration Methods
Motif Database Dependence
A key distinction between integration methods is whether they rely on motif databases to establish cross-modality relationships. Most current methods rely on motif databases to establish relationships between genes from RNA-seq data and peaks from ATAC-seq data. However, these approaches are constrained by incomplete database coverage, particularly for novel or poorly characterized relationships.
Methods like scMI address this limitation by learning cross-modality relationships directly from the data. scMI encodes both cells and modality features from single-cell RNA-seq and ATAC-seq data into a shared latent space by learning cross-modality relationships. By modeling cells and modality features as distinct node types, scMI uses an inter-type attention mechanism to capture long-range cross-modality interactions between genes and peaks. Benchmark results demonstrate that embeddings learned by scMI preserve more biological information and achieve comparable or superior performance in downstream tasks including modality prediction, cell clustering, and gene regulatory network inference compared to methods that rely on databases.
Matched versus Unmatched Data
Some integration methods require matched data where the same cells are measured with both modalities. Other methods can handle unmatched data where different cells are measured with each modality. The choice of method depends on the experimental design.
For matched data, methods that use the correspondence between cells can achieve higher integration quality. For unmatched data, methods must infer the correspondence based on shared cell types or states. Methods like scDART are specifically designed for unmatched data and can integrate scRNA-seq and scATAC-seq data obtained from different batches.
Computational Requirements
Integration methods vary significantly in their computational requirements. Seurat WNN and LIGER are relatively efficient and can handle large datasets on standard computing infrastructure. MOFA+ requires more computational resources due to the variational inference procedure. Deep learning methods like scDART and scMI require specialized hardware such as GPUs and expertise in deep learning frameworks.
Researchers should consider their available computational resources when choosing an integration method. The nf-core documentation provides guidance on community pipeline standards, usage, and configuration for reproducible workflows, which can help researchers plan their computational infrastructure.
Observations and Measurements for Integration Quality
Metrics for Assessing Integration Success
Several quantitative metrics can be used to assess integration quality. These metrics fall into two categories: those that measure batch mixing and those that measure biological preservation.
Batch mixing metrics assess how well cells from different batches are mixed within the same cell type. Common metrics include:
- Batch entropy: measures the diversity of batches within each cluster
- kBET: tests whether the batch label distribution in a neighborhood matches the global distribution
- LISI (Local Inverse Simpson Index): measures the effective number of batches in a local neighborhood
Biological preservation metrics assess whether known biological differences are maintained after integration. Common metrics include:
- Cell type purity: measures whether cells of the same type cluster together
- Marker gene expression: verifies that known marker genes show expected patterns
- Trajectory preservation: assesses whether developmental trajectories are maintained
Recording Integration Parameters
Reproducibility requires careful recording of all integration parameters. Researchers should document:
- Software versions for all packages used
- Quality control thresholds and the number of cells removed at each step
- Normalization methods and parameters
- Integration method and all hyperparameters
- Random seeds used for stochastic procedures
- The version of the reference genome and annotation used
The EMBL-EBI Training provides bioinformatics learning pathways and practical analysis education that emphasize the importance of documentation and reproducibility in computational analysis.
Benchmarking Against Known Biology
A practical approach to evaluating integration quality is to benchmark against known biology. If the integrated data correctly identifies known cell types, marker genes, and regulatory relationships, the integration is likely successful. If known relationships are disrupted, the integration parameters may need adjustment.
For example, in human developmental hematopoiesis, integration of scRNA-seq and scATAC-seq data from over 8,000 human immunophenotypic blood cells from fetal liver and bone marrow allowed inference of differentiation trajectories and identification of three highly proliferative oligopotent progenitor populations downstream of hematopoietic stem cells and multipotent progenitors. The integration revealed opposing patterns of chromatin accessibility and differentiation that coincided with dynamic changes in the activity of distinct lineage-specific transcription factors. These findings were validated by functional experiments that refined the sorting strategy for hematopoietic stem cells and multipotent progenitors, achieving around 90% enrichment.
Common Failure Patterns in Data Integration
Failure Pattern 1: Overcorrection Removing Biological Variation
A common failure is overcorrection, where the integration removes technical differences but also biological differences between cell types. This results in clusters that are mixed across cell types and loss of meaningful biological heterogeneity.
Signs of overcorrection include:
- Known marker genes no longer distinguish cell types
- Rare cell populations disappear from the integrated data
- Clusters contain cells from multiple clearly distinct cell types
To address overcorrection, researchers can reduce the strength of the integration, use fewer dimensions in the shared space, or adjust the weighting between modalities.
Failure Pattern 2: Undercorrection Leaving Batch Effects
The opposite failure is undercorrection, where batch effects remain in the integrated data. This results in cells from the same cell type clustering by batch instead of by biological similarity.
Signs of undercorrection include:
- Cells from the same cell type separate by batch in UMAP plots
- Batch entropy values are low
- Differential expression analysis identifies many genes that are actually batch effects
To address undercorrection, researchers can increase the number of integration dimensions, use more aggressive batch correction, or add additional covariates to the integration model.
Failure Pattern 3: Poor Gene Activity Matrix Quality
The gene activity matrix for scATAC-seq data is a common source of integration failure. If the gene activity matrix does not accurately reflect the relationship between chromatin accessibility and gene expression, the integration will produce misleading results.
Signs of poor gene activity matrix quality include:
- Known marker genes do not show expected accessibility patterns
- The gene activity matrix is not correlated with gene expression for known regulatory relationships
- Integration results are highly sensitive to the gene activity matrix construction method
To address this issue, researchers should evaluate different gene activity matrix construction methods and choose the one that best reflects known biology in their system.
Failure Pattern 4: Misalignment of Rare Cell Populations
Rare cell populations are particularly vulnerable to integration failures. These populations may be lost during quality control, may not have enough cells to form anchors, or may be merged with more abundant populations during clustering.
Signs of rare cell population loss include:
- Known rare cell types are absent from the integrated data
- Rare cell markers are not detected in any cluster
- The proportion of rare cells is much lower than expected from the original data
To address this issue, researchers can use less stringent quality control thresholds, increase the number of integration anchors, or use methods specifically designed to preserve rare populations.
Failure Pattern 5: Trajectory Fracture
Integration can disrupt developmental trajectories, particularly when using methods that do not preserve continuous cell states. This results in fragmented trajectories that do not reflect the true developmental progression.
Signs of trajectory fracture include:
- Known developmental intermediates are missing from the integrated data
- Pseudotime analysis produces discontinuous trajectories
- Cells that should be in transitional states cluster with terminal states
Methods like scDART are designed to preserve cell trajectories in continuous cell populations and can be applied to trajectory inference on integrated data. The design of scDART allows it to preserve cell trajectories in continuous cell populations.
Limitations and Interpretation Boundaries
Technical Limitations of Integration
Integration methods cannot overcome fundamental limitations of the underlying data. If the scRNA-seq data has poor gene coverage or the scATAC-seq data has low sequencing depth, integration will not rescue these quality issues. Researchers should ensure that each dataset meets minimum quality standards before attempting integration.
The sparsity of single-cell data is a fundamental limitation. Both scRNA-seq and scATAC-seq data are high-dimensional, sparse, and noisy. Learning embeddings that simultaneously preserve local cell-state structure, global hierarchy, and smooth developmental trajectories remains an open problem. Existing approaches typically achieve only one of these goals: classical methods emphasize either local neighborhoods or global variance, deep generative models cluster cell types well but often fracture trajectory continuity, and hyperbolic embeddings capture hierarchy but are numerically fragile in practice.
Interpretation Boundaries
Integrated data provides correlative evidence, not causal evidence. If a transcription factor shows increased chromatin accessibility at its binding sites and increased expression of its target genes, this is consistent with regulatory activity but does not prove causation. Functional validation experiments are required to establish causal relationships.
The relationship between chromatin accessibility and gene expression is complex. Accessibility at a regulatory element does not guarantee that the element is active in regulating gene expression. Additional factors such as histone modifications, DNA methylation, and three-dimensional chromatin structure influence the relationship between accessibility and expression.
The Need for Functional Validation
Findings from integrated single-cell analysis should be validated with independent experiments. In breast cancer research, integration of scRNA-seq and scATAC-seq identified CXCL14 as a key regulator of lymph node metastasis. This finding was validated using a tissue microarray of 55 patients and the Oncomine database, confirming that CXCL14 expression was significantly higher in breast cancer patients with lymph node metastasis. In clear cell renal cell carcinoma research, the malignant role of PREX2 was validated in vitro and in vivo, demonstrating that PREX2 facilitated tumor progression by inhibiting PTEN and activating the PI3K/AKT pathway.
The NCBI Data Resources provide official descriptions of databases, search systems, sequence resources, and analysis services that can be used for validation studies. Researchers can use these resources to access public datasets for benchmarking and validation.
Safety and Regulatory Context
Data Management and Privacy
Single-cell data from human samples may contain sensitive information. Researchers must comply with applicable regulations regarding data privacy and protection. This includes obtaining appropriate informed consent, de-identifying samples, and following institutional review board requirements.
Data sharing is an important consideration. Many funding agencies and journals require data deposition in public repositories. Researchers should plan for data sharing from the beginning of their projects and ensure that data sharing complies with all applicable regulations and consent agreements.
Reproducibility Standards
Reproducibility is a core requirement for computational biology research. The The Carpentries Lessons provide foundational computing, data, shell, Git, and programming training that supports reproducible research practices. Researchers should use version control for their analysis code, document all software versions, and use containerization or workflow management tools to ensure that analyses can be reproduced.
The nf-core documentation provides community pipeline standards for reproducible workflows. These standards include requirements for documentation, testing, and containerization that support reproducibility.
Professional Escalation Criteria
Researchers should escalate to more experienced colleagues or seek professional consultation when:
- Integration results are highly unstable across different parameter settings
- Known biological relationships are consistently disrupted by integration
- The gene activity matrix quality is poor and alternative methods do not improve it
- Computational resources are insufficient for the chosen integration method
- The biological interpretation of integration results requires expertise beyond the research team
Records and Documentation Requirements
Essential Records for Integration Analysis
Researchers should maintain the following records for each integration analysis:
- Experimental design documents describing the samples, conditions, and sequencing platforms
- Quality control reports showing the number of cells and features before and after filtering
- Parameter files for all software used in the analysis
- Version information for all software packages and reference genomes
- Integration output files including cluster assignments and dimension reductions
- Evaluation metrics for integration quality
- Code and scripts used for the analysis
Data Storage and Backup
Single-cell multi-omics data are large and require substantial storage capacity. Researchers should plan for data storage and backup from the beginning of their projects. Raw sequencing data should be stored in a secure location with regular backups. Processed data and analysis outputs should be stored in a structured manner that allows easy retrieval and sharing.
Decision Framework for Selecting an Integration Method
Choosing an integration method is not a one-time decision made at the start of a project. It is a decision that should be revisited as you learn more about your data structure, quality, and biological questions. This section provides a practical decision framework that guides you through the selection process based on concrete data characteristics and analysis goals. The framework is organized as a sequence of questions that lead to specific method recommendations.
Step 1: Assess Whether Your Data Are Matched or Unmatched
The first decision point concerns the experimental design. Matched data means the same individual cells were measured with both scRNA-seq and scATAC-seq, typically through a multi-omics platform that captures both modalities simultaneously. Unmatched data means scRNA-seq and scATAC-seq were performed on different cells, either from the same sample or from different batches.
For matched data, you can use methods that leverage the known correspondence between cells. Seurat WNN is well suited for this scenario because it uses the paired measurements to compute weighted nearest neighbors that reflect the relative information content of each modality. The anchoring strategy developed for Seurat enables integration of single-cell measurements across different modalities, and the method has been demonstrated to improve over existing methods for integrating scRNA-seq data while also supporting anchoring of scRNA-seq experiments with scATAC-seq.
For unmatched data, you need methods that infer correspondence based on shared cell types or states. scDART is specifically designed for unmatched data and integrates scRNA-seq and scATAC-seq data while learning cross-modality relationships simultaneously. The design of scDART allows it to preserve cell trajectories in continuous cell populations, making it applicable to trajectory inference on integrated data. scBridge is another option for unmatched data, as it integrates cells in a heterogeneous manner by iterating between identifying reliable scATAC-seq cells that have smaller omics differences and integrating those reliable cells with scRNA-seq data to narrow the omics gap.
Step 2: Evaluate the Quality of Your Gene Activity Matrix
The gene activity matrix is a critical input for most integration methods. This matrix assigns each gene a score based on the accessibility of its promoter and nearby regulatory elements, allowing scATAC-seq data to be compared with scRNA-seq data in gene space. The quality of this matrix directly affects the quality of the integration.
Before committing to an integration method, evaluate your gene activity matrix by checking whether known marker genes show expected accessibility patterns in known cell types. If the gene activity matrix does not reflect known biology, consider methods that do not rely on a pre-defined gene activity matrix. scMI learns cross-modality relationships directly from the data without relying on motif databases or pre-defined gene activity matrices. By modeling cells and modality features as distinct node types, scMI uses an inter-type attention mechanism to capture long-range cross-modality interactions between genes and peaks. This approach is particularly valuable when the gene activity matrix is of poor quality or when you are working with novel or poorly characterized relationships.
Step 3: Determine Whether You Need Shared or Dataset-Specific Factors
Different integration methods produce different types of output. Some methods produce a shared latent space where cells from both modalities are embedded together. Other methods identify factors of variation that are shared across modalities or specific to one modality.
If your goal is to identify the sources of variation that are shared across modalities versus those that are specific to one modality, MOFA+ is a strong choice. MOFA+ uses a variational Bayesian framework to identify factors of variation across multiple omics datasets. The method can handle matched or unmatched data and identifies both shared and dataset-specific factors. This is particularly useful when you want to understand which regulatory programs are common across cell types and which are unique to specific cell types.
If your goal is primarily to create a shared embedding for clustering and visualization, Seurat WNN or LIGER may be more appropriate. LIGER uses integrative nonnegative matrix factorization to identify shared and dataset-specific factors and is designed to be fast and scalable to large datasets. The method is particularly effective for batch correction when integrating datasets from different experiments or sequencing platforms.
Step 4: Consider Your Computational Resources and Expertise
Integration methods vary significantly in their computational requirements and the expertise needed to run them effectively. Seurat WNN and LIGER are relatively efficient and can handle large datasets on standard computing infrastructure. They also have extensive documentation and community support, making them accessible to researchers with limited computational experience.
MOFA+ requires more computational resources due to the variational inference procedure. Deep learning methods like scDART and scMI require specialized hardware such as GPUs and expertise in deep learning frameworks. If your team does not have this expertise, you may need to invest in training or collaborate with researchers who have the necessary skills. The EMBL-EBI Training provides bioinformatics learning pathways and practical analysis education that can help researchers build the skills needed for advanced integration methods.
Step 5: Match the Method to Your Biological Question
The final decision point concerns your specific biological question. Different methods are better suited to different downstream analyses.
For cell type identification and clustering, Seurat WNN and LIGER are well established and produce reliable results. For trajectory inference, scDART is specifically designed to preserve cell trajectories in continuous cell populations. For gene regulatory network inference, scMI has been shown to achieve comparable or superior performance compared to methods that rely on motif databases.
Consider the examples from published studies. In human developmental hematopoiesis, integration of scRNA-seq and scATAC-seq data from over 8,000 human immunophenotypic blood cells from fetal liver and bone marrow allowed inference of differentiation trajectories and identification of three highly proliferative oligopotent progenitor populations downstream of hematopoietic stem cells and multipotent progenitors. The integration revealed opposing patterns of chromatin accessibility and differentiation that coincided with dynamic changes in the activity of distinct lineage-specific transcription factors. This type of analysis requires a method that preserves continuous cell states and supports trajectory inference.
In breast cancer research, integration of scRNA-seq and scATAC-seq identified a novel cell subpopulation with abnormally high expression of CXCL14 in positive lymph nodes. Cell trajectory analysis revealed that CXCL14 expression increased in late pseudo-time, and scATAC-seq identified several transcription factors that may regulate lymph node metastasis. This type of analysis requires a method that can identify rare cell populations and support trajectory analysis.
Practical Implementation of the Decision Framework
To implement this decision framework in practice, follow these steps:
- Document your experimental design, noting whether your data are matched or unmatched
- Generate and evaluate your gene activity matrix before choosing an integration method
- Define your primary biological question and the downstream analyses you plan to perform
- Assess your computational resources and team expertise
- Select a primary integration method and one alternative method for comparison
- Run both methods and compare the results using the evaluation metrics described in the observations and measurements section
- Document the rationale for your method choice and the results of your comparison
Comparison Table for Method Selection
| Decision Point | Seurat WNN | MOFA+ | LIGER | scDART | scBridge | scMI |
|---|---|---|---|---|---|---|
| Matched data | Strong | Strong | Moderate | Not designed | Not designed | Moderate |
| Unmatched data | Moderate | Moderate | Moderate | Strong | Strong | Strong |
| Poor gene activity matrix | Sensitive | Sensitive | Sensitive | Moderate | Moderate | Robust |
| Shared and specific factors | No | Yes | Yes | No | No | No |
| Trajectory preservation | Moderate | Moderate | Moderate | Strong | Moderate | Moderate |
| Computational requirements | Low | Moderate | Low | High | Moderate | High |
| Deep learning expertise needed | No | No | No | Yes | No | Yes |
| Gene regulatory network inference | Moderate | Moderate | Moderate | Moderate | Moderate | Strong |
Recording Your Method Selection
Document the following information when recording your method selection:
- The version of each software package used
- The specific parameters chosen for each method
- The rationale for choosing the primary method over alternatives
- The results of any comparison runs between methods
- The date and version of the reference genome and annotation used
- The random seed used for any stochastic procedures
This documentation supports reproducibility and allows you or others to revisit the decision if new information becomes available. The nf-core documentation provides community pipeline standards for reproducible workflows, including requirements for documentation and configuration that can guide your record keeping.
Common Mistakes in Method Selection
A common mistake is choosing a method based on familiarity instead of suitability. Researchers often default to the method they learned first or the method most commonly used in their field, even when the data characteristics suggest a different approach would be more appropriate.
Another common mistake is failing to evaluate the gene activity matrix before integration. If the gene activity matrix is of poor quality, the integration results will be misleading regardless of the integration method chosen. Take the time to evaluate the gene activity matrix and consider methods that do not rely on it if the quality is poor.
A third mistake is using the same integration parameters for all datasets without adjustment. Integration parameters should be tuned based on the specific characteristics of each dataset. The Galaxy Training Network provides accessible workflow training and analysis tutorials that include practical guidance for parameter tuning in single-cell data integration.
When to Escalate to Professional Support
Escalate to more experienced colleagues or seek professional consultation when:
- Your data have unusual characteristics that do not fit the standard assumptions of any integration method
- You have tried multiple integration methods and none produce biologically meaningful results
- Your gene activity matrix quality is poor and alternative construction methods do not improve it
- You need to integrate more than two modalities and are unsure how to extend your chosen method
- Your computational resources are insufficient for the integration method that best fits your data
- The biological interpretation of integration results requires expertise beyond your research team
The Bioconductor project provides official package documentation and workflow guidance for reproducible genomic analysis. If you encounter issues that are not addressed in the documentation, the community support forums can be a valuable resource for troubleshooting. The The Carpentries Lessons provide foundational computing and data training that can help you build the skills needed to troubleshoot integration issues independently.
Frequently Asked Questions
What is the difference between scRNA-seq and scATAC-seq data?
scRNA-seq measures gene expression by quantifying transcript abundance in individual cells, providing a snapshot of which genes are actively transcribed. scATAC-seq measures chromatin accessibility by identifying regions of open chromatin, providing information about which genomic regions are potentially available for regulatory activity. Gene expression tells you what transcripts are present, while chromatin accessibility tells you which regions of the genome are open and potentially available for transcription factor binding.
Why is integrating scRNA-seq and scATAC-seq data challenging?
Integration is challenging because the two data types have different feature spaces, different technical noise profiles, and often come from different cells. scRNA-seq features are genes, while scATAC-seq features are genomic peaks. To integrate these data, researchers must either map peaks to genes or use a shared feature space. The integration of multi-omics data remains one of the most difficult tasks in bioinformatics due to these differences.
What is a gene activity matrix and why is it important?
A gene activity matrix is a representation of scATAC-seq data where each gene is assigned a score based on the accessibility of its promoter and nearby regulatory elements. This matrix allows scATAC-seq data to be compared with scRNA-seq data in gene space. The quality of the gene activity matrix directly affects the quality of the integration. Pre-defined gene activity matrices are often of low quality and do not reflect the dataset-specific relationship between the two data modalities.
How do I choose between Seurat WNN, MOFA+, and LIGER?
The choice depends on your experimental design and analysis goals. Seurat WNN is a good default choice for most applications because it has extensive documentation and a well-established workflow. MOFA+ is useful when you want to identify shared and dataset-specific factors of variation across modalities. LIGER is effective for batch correction and is fast and scalable to large datasets. Consider your computational resources, the matched or unmatched nature of your data, and the specific biological questions you want to answer.
What are the signs of poor integration quality?
Poor integration quality can manifest as overcorrection, where biological variation is removed along with technical variation, or undercorrection, where batch effects remain. Signs of overcorrection include loss of known marker gene patterns and disappearance of rare cell populations. Signs of undercorrection include cells from the same cell type separating by batch and batch entropy values that are low.
Can I integrate scRNA-seq and scATAC-seq data from different samples or batches?
Yes, many integration methods are designed to handle unmatched data from different samples or batches. Methods like scDART are specifically designed for unmatched data and can integrate scRNA-seq and scATAC-seq data obtained from different batches. The key is to choose a method that can handle the specific characteristics of your data.
What downstream analyses can I perform on integrated data?
Integrated data can be used for cell type identification, trajectory inference, gene regulatory network inference, and identification of cell-type-specific regulatory programs. Integration of scRNA-seq and scATAC-seq data has been used to identify regulatory features of disease states, discover novel cell subpopulations, and infer differentiation trajectories. The specific downstream analyses depend on your biological questions.
How do I validate findings from integrated single-cell analysis?
Findings from integrated analysis should be validated with independent experiments. This can include functional validation experiments, such as knockdown or overexpression studies, and validation in independent patient cohorts or tissue microarrays. In breast cancer research, findings from integrated analysis were validated using a tissue microarray of 55 patients and the Oncomine database. In clear cell renal cell carcinoma research, the role of PREX2 was validated in vitro and in vivo.
Related Bioinformatics Guides
- Multi-Omics Integration: A Practical Workflow for Combining Proteomics, Metabolomics, and Epigenomics Data
- Multi-Omics Integration: A Practical Guide to Combining Data Types
- Proteomics Data Analysis in R: A Practical Workflow for Differential Expression and Visualization
- Gene Set Enrichment Analysis in R: A Practical Tutorial for Interpreting Omics Data
- Multi-Omics Data Integration: A Comparative Framework for Choosing the Right Method
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- Comprehensive Integration of Single-Cell Data.. Cell, 2019.
- Integrative Single-Cell RNA-Seq and ATAC-Seq Analysis of Human Developmental Hematopoiesis.. Cell stem cell, 2021.
- Integration of scATAC-Seq with scRNA-Seq Data.. Methods in molecular biology (Clifton, N.J.), 2023.
- Integrative analyses of scRNA-seq and scATAC-seq reveal CXCL14 as a key regulator of lymph node metastasis in breast cancer.. Human molecular genetics, 2021.
- Integration of scRNA-Seq and scATAC-Seq Reveals Malignant Characteristics of Sarcomatoid Clear Cell Renal Cell Carcinoma.. Cancer science, 2025.
- Integrating scRNA-seq and scATAC-seq with inter-type attention heterogeneous graph neural networks.. Briefings in bioinformatics, 2024.
- scDART: integrating unmatched scRNA-seq and scATAC-seq data and learning cross-modality relationship simultaneously.. Genome biology, 2022.
- Integrative Single-Cell Analysis Reveals Transcriptional and Epigenetic Regulatory Features of Clear Cell Renal Cell Carcinoma.. Cancer research, 2023.
- LAIOR: a hyperbolic neural ODE variational framework for interpretable single-cell manifold learning and trajectory inference.. 2026.
- Single-cell transcriptomic and chromatin accessibility atlas of peripheral blood mononuclear cells reveals immune cell heterogeneity and breed-specific characteristics in Duroc and Meishan pigs.. 2026.
- Single-cell Technologies in Atherosclerosis: Uncovering Cellular Heterogeneity, Mechanisms, and Therapeutic Opportunities.. 2026.
- Single-Cell Multi-Omics Reveal Gene Regulatory Mechanisms Underlying Cardiac Embryonic Development.. 2026.
- Multiplexed analysis of gene expression and chromatin accessibility of human umbilical cord blood using scRNA-Seq and scATAC-Seq.. Molecular Immunology, 2022.
- scBridge embraces cell heterogeneity in single-cell RNA-seq and ATAC-seq data integration. Nature Communications, 2023.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.