# Computational Integration of Single-Cell RNA and ATAC Data: A Comparison of Seurat WNN, MOFA+, and LIGER


## Key Takeaways

- Seurat WNN is optimized for paired multiome data, leveraging known cell-to-cell correspondence to learn cell-specific modality weights, making it a top performer for integrating scRNA-seq and scATAC-seq from the same nucleus.
- MOFA+ and LIGER are versatile for both paired and unpaired data; MOFA+ uses a factor-based decomposition to identify shared and modality-specific variation, while LIGER employs integrative non-negative matrix factorization with dataset-specific factors for batch effect modeling.
- For unpaired integration, LIGER requires ATAC data to be converted to gene activity scores, which can lose peak-level regulatory resolution, whereas MOFA+ can handle raw fragment counts or binary accessibility calls more directly.
- Seurat WNN is unsuitable for unpaired data due to its reliance on direct cell correspondence, necessitating alternative Seurat workflows that lack its cell-specific weighting advantage.
- Mosaic integration scenarios, where some cells are paired and others unpaired, are handled by MOFA+ and LIGER, with performance influenced by the number and diversity of paired nuclei available for anchoring.
- Peak-level regulatory analysis is best supported by Seurat WNN in paired data, as it preserves the native ATAC representation, while LIGER's gene activity score conversion limits this capability.

---

Single-cell RNA sequencing (scRNA-seq) and single-cell ATAC sequencing (scATAC-seq) measure complementary molecular layers from the same biological system. scRNA-seq captures gene expression, while scATAC-seq quantifies chromatin accessibility. When both modalities are measured from the same cell, researchers can directly link transcriptional states to regulatory landscapes. When measured from different cells, integration must account for the absence of a direct cell-to-cell correspondence. This article compares three widely used integration frameworks: Seurat's weighted nearest neighbor (WNN) analysis, MOFA+, and LIGER. The comparison focuses on practical decision criteria for bioinformaticians who need to select an appropriate method for their specific data structure, whether they are working with paired multiome data or unpaired single-modality datasets.

The choice of integration method depends on several factors that should be assessed before analysis begins. These include whether the RNA and ATAC data come from the same physical cells, the number of cells available in each modality, the expected biological complexity of the sample, and the downstream analytical goals. Seurat WNN is designed for paired data where both modalities are measured from the same cell. MOFA+ and LIGER can handle both paired and unpaired data, but they differ in their underlying statistical models and in the types of biological questions they are best suited to answer. Recent benchmarking studies have shown that integration performance varies substantially across methods and dataset characteristics, making method selection a critical step in the analytical pipeline [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>].

## Understanding the Data Structures in Single-Cell Multi-Omics

### Paired Multiome Data

Paired multiome data refers to experiments where gene expression and chromatin accessibility are measured from the same individual cell. Commercial platforms such as the 10x Multiome assay capture both RNA and ATAC libraries from a single nucleus within the same droplet. This design creates a direct molecular link between the transcriptome and the accessible chromatin landscape of each cell. The primary analytical advantage of paired data is that integration methods can use the known cell-to-cell correspondence to learn modality-specific transformations and to construct joint representations that preserve both transcriptional and epigenetic information.

Paired data also enable the direct assessment of peak-gene associations. Because the RNA and ATAC signals originate from the same nucleus, computational methods can correlate chromatin accessibility at specific genomic loci with expression of nearby genes without relying on external annotations or imputation across cells. This property makes paired multiome data particularly valuable for gene regulatory network inference and for identifying cis-regulatory elements that control cell-type-specific expression programs [<a href="#ref-3">3</a>][<a href="#ref-2">2</a>].

However, paired data introduce technical challenges that must be addressed during quality control. Ambient RNA and free-floating chromatin fragments can contaminate droplets, creating spurious signals that distort cell-type associations. Genotype-based sample multiplexing approaches have been developed to estimate ambient contamination fractions in both modalities. These methods model the probability that a read originates from ambient contamination versus from the captured nucleus, allowing researchers to classify empty droplets and to estimate droplet-specific contamination levels. In simulated datasets, such approaches have achieved high specificity and correct singlet assignment even at ambient fractions up to 60 percent, whereas models that ignore ambient contamination maintain comparable sensitivity only at fractions up to 25 percent [<a href="#ref-4">4</a>].

### Unpaired Single-Modality Data

Unpaired data arise when scRNA-seq and scATAC-seq are performed on separate cell suspensions from the same tissue or condition. This experimental design is common when samples are collected at different times, when different laboratories process different modalities, or when cost constraints prevent running both assays on the same cells. Unpaired data lack a direct cell-to-cell correspondence, so integration methods must infer the relationship between modalities based on shared cell types and states.

The absence of paired cells changes the integration problem fundamentally. Methods must align cells across modalities by identifying shared biological signals while accounting for modality-specific technical variation. Gene activity scores, which aggregate chromatin accessibility signals across gene bodies and proximal regulatory regions, provide a bridge between ATAC and RNA data. However, benchmarking studies have shown that gene activity scores correlate only weakly with measured gene expression, even though they preserve cellular neighborhoods sufficiently well for clustering purposes [<a href="#ref-5">5</a>].

The number of cells in each modality is a critical factor for unpaired integration. When a paired multiome dataset is used to guide the integration of unpaired single-modality data, the multiome dataset must contain an adequate number of nuclei to represent the full range of cell types and states present in the sample. Insufficient representation in the multiome dataset compromises the reliability of cell type annotations transferred to the single-modality data. For cell type annotation purposes, the number of cells matters more than sequencing depth when generating a multiome dataset [<a href="#ref-2">2</a>].

### Mosaic Integration Scenarios

Mosaic integration refers to experimental designs where some cells have measurements for all modalities while other cells have measurements for only one modality. This scenario arises when a researcher generates a paired multiome dataset for a subset of samples and single-modality data for additional samples or conditions. Mosaic designs are increasingly common because they balance the biological value of paired measurements against the cost and complexity of generating multiome data for every sample.

Mosaic integration methods must simultaneously learn from the paired cells, which provide ground-truth modality relationships, and from the unpaired cells, which expand the biological coverage. The paired cells serve as anchors that define how chromatin accessibility relates to gene expression, while the unpaired cells contribute additional cell states and conditions. Benchmarking studies indicate that multiome data are helpful for annotating single-modality data, but the benefit depends on the number of paired nuclei available and on the diversity of cell types represented in the paired dataset [<a href="#ref-2">2</a>].

## Core Principles of Multi-Modal Integration

### Shared and Modality-Specific Variation

Every integration method must contend with the fact that each modality captures a different aspect of cellular biology and carries its own technical noise. Gene expression measurements reflect the abundance of RNA transcripts, which are influenced by transcription rates, RNA stability, and post-transcriptional regulation. Chromatin accessibility measurements reflect the physical openness of genomic regions, which is influenced by nucleosome positioning, transcription factor binding, and chromatin remodeling. These two layers are related but not identical, and the relationship varies across genomic loci and cell types.

Integration methods model this relationship by decomposing the observed data into shared components, which represent biological signals present in both modalities, and modality-specific components, which represent signals unique to each measurement type. The shared components form the basis for joint cell embeddings, while the modality-specific components capture technical artifacts and biological processes that operate in only one layer. Methods differ in how they define and estimate these components, which affects their sensitivity to biological signals and their robustness to technical noise [<a href="#ref-1">1</a>][<a href="#ref-6">6</a>].

### Dimension Reduction as a Critical Step

Dimension reduction transforms the high-dimensional feature space of each modality into a lower-dimensional representation that preserves biologically meaningful structure while discarding noise. For scRNA-seq data, the feature space consists of thousands of genes. For scATAC-seq data, the feature space consists of hundreds of thousands of genomic peaks. Direct integration in the original feature space is computationally prohibitive and statistically unstable, so all practical methods perform dimension reduction before or during integration.

Benchmarking studies have identified dimension reduction as the most critical step in unpaired integration pipelines. Non-linear dimension reduction methods generally offer better performance in terms of preserving cellular neighborhoods and enabling accurate label transfer, while linear methods provide greater robustness and computational efficiency. The choice of dimension reduction method interacts with the choice of feature linking strategy and clustering algorithm, and the optimal combination depends on the specific dataset [<a href="#ref-5">5</a>].

### Feature Linking Between Modalities

Feature linking establishes the correspondence between genomic peaks in the ATAC data and genes in the RNA data. This correspondence is necessary because integration methods must map chromatin accessibility signals to the gene expression space to enable joint analysis. Several strategies exist for feature linking, including proximity-based approaches that assign peaks to nearby genes, correlation-based approaches that identify peaks whose accessibility correlates with gene expression across cells, and model-based approaches that learn peak-gene relationships from paired data.

The choice of feature linking strategy affects the quality of the integrated representation. Proximity-based approaches are simple and computationally efficient but may miss long-range regulatory interactions. Correlation-based approaches can capture functional relationships but require sufficient cells to estimate correlations reliably. Model-based approaches leverage paired data to learn peak-gene associations directly, but they require a paired multiome dataset and may not generalize to unpaired data from different tissues or conditions [<a href="#ref-5">5</a>][<a href="#ref-2">2</a>].

## Seurat WNN: Weighted Nearest Neighbor Analysis for Paired Data

### Method Overview

Seurat's weighted nearest neighbor (WNN) analysis is designed specifically for paired multi-modal data where both modalities are measured from the same cells. The method begins by independently processing each modality through its standard analytical pipeline, including normalization, feature selection, and dimension reduction. For RNA data, this typically involves log-normalization and principal component analysis. For ATAC data, this involves term frequency-inverse document frequency (TF-IDF) normalization and latent semantic indexing or similar dimension reduction.

After computing modality-specific embeddings, WNN analysis learns cell-specific weights that determine how much each modality contributes to the final joint representation. The weights are learned by constructing a shared nearest neighbor graph for each modality, then identifying for each cell the modality that provides the most consistent neighborhood structure. Cells where the RNA data are more informative receive higher RNA weights, while cells where the ATAC data are more informative receive higher ATAC weights. The weighted combination of modality-specific graphs forms the WNN graph, which serves as the basis for clustering and visualization [<a href="#ref-2">2</a>].

### Strengths for Same-Cell Integration

The primary strength of Seurat WNN is its ability to leverage the known cell-to-cell correspondence in paired data. Because the method knows which RNA profile belongs to which ATAC profile, it can learn modality weights that reflect the information content of each layer for each individual cell. This cell-specific weighting is particularly valuable when different cell types are better distinguished by different modalities. For example, a cell type defined by a specific transcription factor may be more clearly separated by chromatin accessibility at that factor's binding sites, while another cell type defined by a specific surface marker may be more clearly separated by gene expression.

Benchmarking studies have identified Seurat v4, which includes the WNN workflow, as the best currently available platform for integrating scRNA-seq, snATAC-seq, and multiome data. The method performs well across a range of dataset sizes and biological contexts, and it provides a complete workflow from raw data processing through clustering and visualization [<a href="#ref-2">2</a>].

### Limitations and Data Requirements

Seurat WNN requires paired data. The method cannot be applied directly to unpaired datasets because it needs the cell-to-cell correspondence to learn modality weights. For unpaired data, Seurat provides alternative workflows based on canonical correlation analysis or reciprocal principal component analysis, but these approaches do not provide the same cell-specific weighting as WNN.

The method also requires careful quality control before integration. Cells with low RNA complexity or low ATAC fragment counts can distort the modality weight learning process. Doublets, which arise when two cells are captured in the same droplet, create hybrid molecular profiles that are particularly problematic for WNN because they combine the RNA profile of one cell with the ATAC profile of another. Multiplet detection methods that integrate evidence across modalities can help identify and remove these problematic cells before WNN analysis [<a href="#ref-7">7</a>].

## MOFA+: Multi-Omics Factor Analysis for Unpaired and Paired Data

### Method Overview

MOFA+ is a statistical framework that identifies shared and modality-specific sources of variation across multiple data matrices. The method extends the original MOFA model to handle more than two modalities and to accommodate partially overlapping samples. MOFA+ decomposes each data matrix into a set of latent factors, where each factor captures a pattern of coordinated variation across features. The factors are shared across modalities, meaning that the same factor can explain variation in both RNA and ATAC data, but the factor loadings differ by modality.

The output of MOFA+ is a low-dimensional representation of each cell in terms of its factor values. These factor values can be used for clustering, visualization, and downstream analyses such as trajectory inference or differential testing. MOFA+ also provides interpretable factor loadings that identify which genes and which genomic peaks contribute most strongly to each factor, enabling biological interpretation of the identified sources of variation [<a href="#ref-1">1</a>].

### Handling Different Data Types

MOFA+ is designed to handle different data types through its likelihood framework. Gene expression counts are modeled with a Poisson or negative binomial likelihood, while chromatin accessibility data can be modeled with a Bernoulli or Poisson likelihood depending on whether the data are represented as binary accessibility calls or as fragment counts. This flexibility allows MOFA+ to integrate data that have been processed through different normalization schemes and to account for the different statistical properties of RNA and ATAC measurements.

The ability to handle different data types makes MOFA+ particularly suitable for integrating data from different experimental platforms or from different laboratories. The method does not require the features to be matched across modalities, which is important because RNA data are defined at the gene level while ATAC data are defined at the peak level. Instead, MOFA+ learns the relationship between modalities through the shared factor structure [<a href="#ref-1">1</a>].

### Strengths for Unpaired Integration

MOFA+ can integrate unpaired data because it does not require a cell-to-cell correspondence between modalities. The method aligns cells across modalities based on their factor values, which are learned from the shared sources of variation. This property makes MOFA+ suitable for mosaic integration scenarios where some cells have paired measurements and others have single-modality measurements.

The factor-based representation also provides a natural framework for identifying modality-specific biological signals. Factors that load strongly on RNA features but weakly on ATAC features capture transcriptional programs that are not reflected in chromatin accessibility, while factors with the opposite pattern capture epigenetic variation that does not manifest at the transcript level. This decomposition can reveal biological processes that would be missed by methods that force all variation into a shared representation [<a href="#ref-1">1</a>][<a href="#ref-6">6</a>].

### Limitations and Computational Considerations

MOFA+ requires careful hyperparameter selection, particularly the number of factors to retain. Too few factors fail to capture the full range of biological variation, while too many factors capture noise and complicate interpretation. Model selection criteria based on evidence lower bounds or cross-validation can guide this choice, but they require computational resources and may not always identify the biologically optimal solution.

The method also assumes that the shared factors explain the same biological processes across modalities. When the relationship between RNA and ATAC varies substantially across cell types, this assumption may be violated, and the factor structure may not accurately represent the underlying biology. In such cases, methods that allow more flexible modality relationships, such as deep generative models, may perform better [<a href="#ref-6">6</a>].

## LIGER: Linked Inference of Genomic Experimental Relationships

### Method Overview

LIGER uses integrative non-negative matrix factorization to identify shared and dataset-specific factors across multiple single-cell datasets. The method was originally developed for integrating scRNA-seq datasets from different samples, conditions, or species, but it has been extended to handle multi-modal data including ATAC-seq. LIGER factorizes each dataset into a shared factor matrix and a dataset-specific factor matrix, where the shared factors represent biological signals present across all datasets and the dataset-specific factors capture technical or biological variation unique to each dataset.

The shared factor values provide a joint representation of cells across datasets, enabling clustering, visualization, and downstream analyses. LIGER also provides a quantile normalization step that aligns the factor values across datasets, reducing batch effects and improving the alignment of cells from different experimental conditions [<a href="#ref-1">1</a>].

### Application to RNA and ATAC Integration

For RNA and ATAC integration, LIGER treats each modality as a separate dataset and identifies shared factors that capture coordinated variation across both modalities. The method requires that the features be mapped to a common space, which is typically achieved by computing gene activity scores from the ATAC data and using these scores as the ATAC modality features. This preprocessing step is critical because LIGER operates on feature-by-cell matrices and cannot directly handle the peak-level representation of ATAC data.

The choice of gene activity score calculation affects the quality of the integration. Simple approaches that sum accessibility across gene bodies and proximal promoters are computationally efficient but may miss distal regulatory elements. More sophisticated approaches that incorporate enhancer-promoter interactions or that learn peak-gene relationships from paired data can improve the biological fidelity of the gene activity scores, but they require additional data or annotations [<a href="#ref-5">5</a>].

### Strengths for Multi-Dataset Integration

LIGER is particularly well suited for integrating data from multiple samples or conditions, where batch effects and technical variation are major concerns. The dataset-specific factors explicitly model variation that is unique to each dataset, allowing the shared factors to capture biological signals that are consistent across datasets. This property makes LIGER useful for meta-analyses that combine data from multiple studies or for comparisons across disease states or developmental stages.

The non-negative matrix factorization framework also provides interpretable factor loadings that identify which genes contribute to each factor. This interpretability facilitates biological validation of the integration results and can guide downstream hypothesis generation [<a href="#ref-1">1</a>].

### Limitations and Preprocessing Requirements

LIGER requires that all modalities be represented in the same feature space, which means that ATAC data must be converted to gene activity scores before integration. This conversion loses the peak-level resolution of the ATAC data and may obscure regulatory variation that is not captured by gene-level aggregation. Researchers interested in peak-level analysis should consider methods that preserve the native ATAC representation.

The method also assumes that the shared factors represent the same biological processes across datasets. When the relationship between RNA and ATAC varies across samples or conditions, this assumption may be violated, and the integration may not accurately represent the underlying biology. In such cases, methods that allow modality-specific factor structures may be more appropriate [<a href="#ref-1">1</a>][<a href="#ref-6">6</a>].

## At a Glance: Method Comparison for Common Integration Scenarios

| Scenario | Seurat WNN | MOFA+ | LIGER |
|----------|-----------|-------|-------|
| Paired multiome data from same cells | Recommended. Uses known cell correspondence to learn cell-specific modality weights. Best-performing platform in benchmarks for paired integration [<a href="#ref-2">2</a>]. | Applicable. Factor decomposition captures shared and modality-specific variation, but does not exploit the known cell correspondence as directly [<a href="#ref-1">1</a>]. | Applicable with gene activity scores. Requires converting ATAC peaks to gene-level features, which loses peak resolution [<a href="#ref-5">5</a>]. |
| Unpaired RNA and ATAC from different cells | Not directly applicable. Requires cell-to-cell correspondence for WNN. Alternative Seurat workflows can be used but lack cell-specific weighting [<a href="#ref-2">2</a>]. | Recommended. Handles different data types and does not require cell correspondence. Factor structure aligns cells across modalities [<a href="#ref-1">1</a>]. | Recommended. Dataset-specific factors model technical variation. Requires gene activity score preprocessing [<a href="#ref-5">5</a>]. |
| Mosaic data with some paired and some unpaired cells | Applicable if paired cells are used as anchors. Performance depends on number and diversity of paired nuclei [<a href="#ref-2">2</a>]. | Applicable. Handles partially overlapping samples naturally through the factor model [<a href="#ref-1">1</a>]. | Applicable. Shared factors align paired and unpaired cells, but gene activity score conversion is required [<a href="#ref-5">5</a>]. |
| Multi-sample or multi-condition integration | Applicable with additional batch correction steps. WNN does not explicitly model dataset-specific variation [<a href="#ref-2">2</a>]. | Applicable. Factors can capture condition-specific variation, but interpretation requires careful experimental design [<a href="#ref-1">1</a>]. | Recommended. Dataset-specific factors explicitly model batch and condition effects [<a href="#ref-1">1</a>]. |
| Peak-level regulatory analysis | Applicable. WNN preserves the native ATAC representation and can identify peak-gene associations from paired data [<a href="#ref-2">2</a>]. | Applicable. Factor loadings identify peaks contributing to each factor, but peak-level interpretation requires additional analysis [<a href="#ref-1">1</a>]. | Limited. Gene activity score conversion loses peak-level resolution [<a href="#ref-5">5</a>]. |

## Practical Workflow for Method Selection

### Step 1: Assess Data Structure

Before selecting an integration method, document the structure of your data. Determine whether RNA and ATAC measurements come from the same physical cells or from different cells. If paired data are available, count the number of paired nuclei and assess whether they represent the full range of cell types and states expected in the sample. If unpaired data are available, document the number of cells in each modality and note any differences in sample processing, sequencing depth, or batch.

This assessment determines which methods are applicable. Seurat WNN requires paired data. MOFA+ and LIGER can handle both paired and unpaired data, but their performance depends on the specific characteristics of the dataset. Benchmarking studies have shown that the availability of an adequate number of nuclei in a paired multiome dataset is crucial for accurate cell type annotation when using that dataset to guide unpaired integration [<a href="#ref-2">2</a>].

### Step 2: Perform Quality Control on Each Modality

Quality control should be performed independently on each modality before integration. For RNA data, assess the number of genes detected per cell, the fraction of reads mapping to mitochondrial genes, and the total number of reads per cell. For ATAC data, assess the number of unique fragments per cell, the fraction of fragments in peaks, and the transcription start site enrichment score. Cells that fail quality thresholds in either modality should be removed before integration.

Multiplet detection is particularly important for integration because doublets create hybrid molecular profiles that distort the relationship between modalities. Methods that integrate evidence across modalities can improve multiplet detection compared with methods that analyze each modality independently. These approaches model the singlet background directly from observed signals and produce classification probabilities that enable false discovery rate control [<a href="#ref-7">7</a>].

### Step 3: Select the Integration Method Based on Data Structure and Goals

Use the decision table in the At a Glance section to select the integration method that matches your data structure and analytical goals. For paired data where the primary goal is to identify cell types and states, Seurat WNN is the recommended choice. For unpaired data where the goal is to align cells across modalities, MOFA+ or LIGER are appropriate, with the choice depending on whether you prefer the factor-based decomposition of MOFA+ or the dataset-specific modeling of LIGER.

Consider the downstream analyses you plan to perform. If you need peak-level resolution for regulatory analysis, choose a method that preserves the native ATAC representation. If you plan to integrate data from multiple samples or conditions, consider whether the method explicitly models dataset-specific variation. If you plan to perform trajectory inference or gene regulatory network analysis, verify that the integration output is compatible with the downstream tools you intend to use [<a href="#ref-3">3</a>][<a href="#ref-5">5</a>].

### Step 4: Validate Integration Results

Integration results should be validated using multiple criteria. First, assess whether known cell types are separated in the joint embedding. If you have prior knowledge of the expected cell type composition, compare the clustering results with this expectation. Second, assess whether cells from the same biological condition cluster together across modalities. For paired data, verify that the modality weights are biologically sensible, with cell types that are better defined by chromatin accessibility receiving higher ATAC weights. Third, assess the batch effect removal by checking whether cells from different batches or conditions are mixed within clusters.

For unpaired integration, additional validation is needed to confirm that cells from different modalities are correctly aligned. Gene activity scores can provide a rough check, but benchmarking studies have shown that these scores correlate only weakly with gene expression. More rigorous validation requires comparing the integrated representation with external annotations or with paired data from the same tissue [<a href="#ref-5">5</a>].

## Records and Measurements for Integration Projects

### Documentation Requirements

Maintain detailed records of the integration workflow to ensure reproducibility and to facilitate troubleshooting. Document the software versions of all tools used, including the integration method, the preprocessing packages, and the downstream analysis tools. Record all parameter settings, including normalization methods, dimension reduction parameters, and clustering resolution. Document the quality control thresholds applied to each modality and the number of cells removed at each step.

Version control is essential for reproducibility. Use a version control system to track changes to analysis scripts and document the exact commands used to generate each result. Containerization tools can help ensure that the computational environment is reproducible across different systems. Community standards for reproducible workflows provide guidance on structuring analysis pipelines and documenting computational environments [<a href="#ref-8">8</a>].

### Quality Metrics to Track

Track quality metrics at each stage of the integration workflow. Before integration, record the number of cells and features in each modality, the median number of reads or fragments per cell, and the fraction of cells passing quality control. After integration, record the number of clusters identified, the cluster sizes, and the modality weights for each cell. For methods that produce factor or component scores, record the variance explained by each factor and the number of factors retained.

For paired data, track the concordance between RNA and ATAC signals within cells. Cells where the two modalities disagree may represent technical artifacts or biologically interesting transitional states. For unpaired data, track the alignment quality by measuring the mixing of cells from different modalities within clusters. Several metrics exist for quantifying batch mixing and biological conservation, and these should be reported alongside the integration results [<a href="#ref-1">1</a>][<a href="#ref-5">5</a>].

### Reproducibility Considerations

Reproducibility requires more than documenting the analysis steps. The random seeds used in stochastic algorithms should be recorded, and the analysis should be run multiple times with different seeds to assess the stability of the results. The computational environment, including the operating system, software versions, and hardware configuration, should be documented. The input data files should be versioned, and any preprocessing steps that modify the data should be recorded.

Public data repositories provide infrastructure for sharing both raw and processed data. The National Center for Biotechnology Information maintains databases for sequence data, gene expression data, and genomic annotations, and these resources support the deposition and retrieval of single-cell datasets [<a href="#ref-9">9</a>]. Training resources from the European Bioinformatics Institute provide guidance on data management and analysis best practices for bioinformatics workflows [<a href="#ref-10">10</a>].

## Common Failure Patterns in Multi-Modal Integration

### Insufficient Paired Nuclei for Anchor-Based Integration

When paired multiome data are used to guide the integration of unpaired single-modality data, the number of paired nuclei is a critical factor. Benchmarking studies have shown that insufficient representation of nuclei in the multiome dataset compromises the reliability of cell type annotations transferred to the single-modality data. If the paired dataset lacks rare cell types or specific cell states, the integration will fail to correctly assign those cells in the unpaired data [<a href="#ref-2">2</a>].

This failure pattern can be detected by comparing the cell type composition of the paired and unpaired datasets. If the unpaired data contain cell types that are absent or underrepresented in the paired data, the integration results for those cell types should be interpreted with caution. Adding more paired nuclei, particularly from underrepresented cell types, can improve the integration performance.

### Gene Activity Score Artifacts in Unpaired Integration

Gene activity scores provide a bridge between ATAC and RNA data for unpaired integration, but they can introduce artifacts. Benchmarking studies have shown that gene activity scores correlate only weakly with gene expression, meaning that the ATAC-derived gene activity values do not accurately predict the RNA expression levels. Despite this weak correlation, gene activity scores can preserve cellular neighborhoods sufficiently well for clustering, but the biological interpretation of the integrated representation requires caution [<a href="#ref-5">5</a>].

This failure pattern manifests as clusters that are defined by ATAC-specific signals that do not correspond to transcriptional differences. Researchers should validate that clusters identified in the integrated representation are supported by both modalities and should not interpret gene activity score differences as evidence of expression differences without confirmation from RNA data.

### Modality Weight Imbalance in Paired Integration

Seurat WNN learns cell-specific modality weights that determine how much each modality contributes to the joint representation. When one modality dominates the weights across most cells, the integration effectively reduces to a single-modality analysis, and the biological value of the multi-modal data is lost. This can occur when one modality has substantially higher technical quality, when the normalization procedures create scale differences between modalities, or when the biological signal is concentrated in one modality.

This failure pattern can be detected by examining the distribution of modality weights across cells. If the weights are heavily skewed toward one modality, the integration may not be capturing the full biological signal. Adjusting the normalization or feature selection procedures for the underrepresented modality can help balance the weights.

### Batch Effects Confounded with Biological Variation

When integrating data from multiple samples or conditions, batch effects can be confounded with biological variation. This is particularly problematic when the batch structure correlates with the biological variable of interest, such as when all disease samples are processed in one batch and all control samples in another. Integration methods that remove batch effects may also remove biological variation, while methods that preserve biological variation may retain batch effects.

This failure pattern is difficult to detect without external validation. Experimental designs that randomize samples across batches can help prevent confounding, and computational methods that explicitly model batch structure can help separate technical and biological variation. LIGER's dataset-specific factors provide one approach to modeling batch effects, but the interpretation of these factors requires careful experimental design [<a href="#ref-1">1</a>].

## Limitations and Interpretation Boundaries

### Statistical Assumptions of Each Method

Each integration method makes statistical assumptions that limit its applicability. Seurat WNN assumes that the paired data provide reliable cell-to-cell correspondence and that the modality weights can be learned from the neighborhood structure. MOFA+ assumes that the shared factors explain the same biological processes across modalities and that the factor model adequately captures the data distribution. LIGER assumes that the non-negative matrix factorization captures the underlying structure and that the gene activity scores provide a faithful representation of the ATAC data.

These assumptions should be evaluated before applying the methods to new datasets. When the assumptions are violated, the integration results may be misleading, and alternative methods should be considered. Deep generative models, which make fewer assumptions about the data distribution, may provide more flexible alternatives for complex datasets [<a href="#ref-6">6</a>].

### Interpretation Limits of Integrated Representations

The integrated representation provides a joint view of the transcriptional and epigenetic state of each cell, but it does not provide a complete picture of cellular biology. The integration methods identify shared sources of variation, but they may miss modality-specific signals that are biologically important. For example, a transcription factor that regulates gene expression through chromatin remodeling may show strong ATAC signal but weak RNA signal, and this regulatory activity may be underrepresented in the integrated representation.

Researchers should interpret the integrated representation as a summary of the shared biological signal across modalities, not as a complete description of cellular state. Modality-specific analyses should be performed alongside the integrated analysis to capture signals that are unique to each layer. Gene regulatory network analysis tools that integrate RNA and ATAC data can help identify transcription factors that coordinate expression and accessibility programs [<a href="#ref-3">3</a>].

### Generalization Across Tissues and Conditions

Integration methods are typically benchmarked on specific tissues and conditions, and their performance may not generalize to other biological contexts. The relationship between chromatin accessibility and gene expression varies across cell types, developmental stages, and disease states. Methods that perform well on peripheral blood mononuclear cells may perform differently on brain tissue, tumor samples, or plant tissues.

Researchers should validate integration methods on their specific data before relying on the results. This validation can involve comparing the integration output with known biology, checking the stability of the results across parameter settings, and confirming that the identified cell types and states are consistent with external annotations. The cost and complexity of single-cell experiments make it impractical to benchmark all methods on every dataset, but targeted validation of the selected method is essential [<a href="#ref-1">1</a>][<a href="#ref-5">5</a>].

## Quality Controls and Validation Strategies

### Pre-Integration Quality Checks

Before running the integration, verify that each modality has been processed appropriately. For RNA data, confirm that normalization has been applied and that the data are on a comparable scale across cells. For ATAC data, confirm that the peak calling has been performed consistently and that the fragment counts have been normalized appropriately. Check for batch effects within each modality by visualizing the data colored by batch or sample.

For paired data, verify that the cell-to-cell correspondence is reliable. Check for cells where the RNA and ATAC signals are highly discordant, as these may represent doublets or technical artifacts. Multiplet detection methods that integrate evidence across modalities can identify problematic cells before integration [<a href="#ref-7">7</a>].

### Post-Integration Validation

After integration, validate the results using multiple independent criteria. First, assess the separation of known cell types in the joint embedding. If marker genes for expected cell types are known, check that these markers are enriched in the appropriate clusters. Second, assess the mixing of cells from different batches or conditions within clusters. Good integration should remove technical variation while preserving biological variation. Third, assess the stability of the clustering results across parameter settings and random seeds.

For unpaired integration, additional validation is needed to confirm that cells from different modalities are correctly aligned. This can involve comparing the integrated representation with external annotations, checking that cells from the same biological condition cluster together across modalities, and verifying that the gene activity scores and expression levels are consistent for known marker genes [<a href="#ref-5">5</a>].

### Escalation Criteria for Problematic Integrations

When integration results are unsatisfactory, several escalation strategies are available. First, revisit the quality control steps and consider removing additional cells or features. Second, adjust the preprocessing parameters, such as the number of dimensions retained or the normalization method. Third, consider an alternative integration method that makes different assumptions about the data structure.

If the integration continues to produce unsatisfactory results, consider whether the data are suitable for integration. Datasets with very different cell type compositions, substantial technical differences, or confounding batch effects may not be integrable with current methods. In such cases, analyzing each modality separately and comparing the results at the level of cell type annotations may be more appropriate than forcing a joint integration [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>].

## Safety and Regulatory Context for Computational Analysis

### Data Management and Privacy Considerations

Single-cell datasets may contain sensitive information, particularly when derived from human subjects. Researchers should ensure that data collection and analysis comply with applicable regulations and institutional policies. De-identification of samples, secure data storage, and controlled access to raw data are important considerations for human studies. Public data repositories provide mechanisms for controlled access to sensitive datasets while enabling sharing of de-identified data [<a href="#ref-9">9</a>].

### Reproducibility Standards in Published Research

The scientific community has developed standards for reproducible computational analysis that should be followed in single-cell integration projects. These standards include sharing analysis code, documenting the computational environment, and depositing processed data in public repositories. Training resources from the Galaxy Training Network and The Carpentries provide guidance on reproducible analysis practices and data management [<a href="#ref-11">11</a>][<a href="#ref-12">12</a>].

Community-driven pipeline frameworks provide standardized workflows for single-cell analysis that incorporate best practices for reproducibility. These frameworks include documentation on usage, configuration, and quality control, and they can help ensure that integration analyses are performed consistently across projects [<a href="#ref-8">8</a>].

### Professional Escalation for Method Selection

When the choice of integration method has substantial implications for the biological conclusions, consultation with a bioinformatics specialist or biostatistician is recommended. This is particularly important when the data have complex structure, such as mosaic designs with partial overlap, when the expected biology is poorly understood, or when the integration results will guide downstream experimental work.

Specialist consultation should also be sought when the integration results are unexpected or when different methods produce conflicting results. The interpretation of integration results requires understanding the statistical assumptions of each method and the biological context of the data. The published literature on integration methods provides guidance on method selection and interpretation, and recent benchmarking studies offer systematic comparisons across methods and datasets [<a href="#ref-1">1</a>][<a href="#ref-5">5</a>][<a href="#ref-2">2</a>].

## Frequently Asked Questions

### What is the difference between paired and unpaired single-cell multi-omics data?

Paired data come from experiments where RNA and ATAC are measured from the same physical cell, such as the 10x Multiome assay. This creates a direct molecular link between the transcriptome and the accessible chromatin landscape of each cell. Unpaired data come from experiments where RNA and ATAC are measured from different cells, either from separate cell suspensions or from different samples. Unpaired data lack a direct cell-to-cell correspondence, so integration methods must infer the relationship between modalities based on shared cell types and states [<a href="#ref-2">2</a>].

### Can Seurat WNN be used for unpaired RNA and ATAC data?

Seurat WNN requires paired data because it uses the known cell-to-cell correspondence to learn cell-specific modality weights. For unpaired data, Seurat provides alternative workflows based on canonical correlation analysis or reciprocal principal component analysis, but these approaches do not provide the same cell-specific weighting as WNN. For unpaired integration, MOFA+ or LIGER are more appropriate choices [<a href="#ref-2">2</a>].

### How do gene activity scores affect unpaired integration?

Gene activity scores convert ATAC peak-level data to gene-level features by aggregating accessibility across gene bodies and proximal regulatory regions. These scores provide a bridge between ATAC and RNA data for unpaired integration. However, benchmarking studies have shown that gene activity scores correlate only weakly with gene expression, even though they preserve cellular neighborhoods sufficiently well for clustering. The choice of gene activity score calculation affects the quality of the integration [<a href="#ref-5">5</a>].

### What is the most critical step in unpaired integration pipelines?

Benchmarking studies have identified dimension reduction as the most critical step in unpaired integration pipelines. Non-linear dimension reduction methods generally offer better performance in terms of preserving cellular neighborhoods and enabling accurate label transfer, while linear methods provide greater robustness and computational efficiency. The choice of dimension reduction method interacts with the choice of feature linking strategy and clustering algorithm [<a href="#ref-5">5</a>].

### How many paired nuclei are needed for anchor-based integration?

The number of paired nuclei needed depends on the biological complexity of the sample and the diversity of cell types present. Benchmarking studies emphasize that the availability of an adequate number of nuclei in the multiome dataset is crucial for achieving accurate cell type annotation when using that dataset to guide unpaired integration. Insufficient representation of nuclei may compromise the reliability of the annotations. For cell type annotation purposes, the number of cells matters more than sequencing depth [<a href="#ref-2">2</a>].

### What are the main differences between MOFA+ and LIGER?

MOFA+ uses factor analysis to identify shared and modality-specific sources of variation across multiple data matrices. It can handle different data types and does not require cell-to-cell correspondence. LIGER uses non-negative matrix factorization to identify shared and dataset-specific factors, and it is particularly well suited for integrating data from multiple samples or conditions. LIGER requires that all modalities be represented in the same feature space, which means that ATAC data must be converted to gene activity scores [<a href="#ref-1">1</a>][<a href="#ref-5">5</a>].

### How should integration results be validated?

Integration results should be validated using multiple criteria. First, assess whether known cell types are separated in the joint embedding. Second, assess whether cells from the same biological condition cluster together across modalities. Third, assess the batch effect removal by checking whether cells from different batches or conditions are mixed within clusters. For unpaired integration, additional validation is needed to confirm that cells from different modalities are correctly aligned [<a href="#ref-5">5</a>][<a href="#ref-2">2</a>].

### What should I do if different integration methods produce conflicting results?

Conflicting results from different integration methods indicate that the data may not support a unique integrated representation. This can occur when the relationship between RNA and ATAC varies across cell types, when batch effects are confounded with biological variation, or when the data quality is insufficient. In such cases, consult with a bioinformatics specialist, examine the assumptions of each method, and consider whether analyzing each modality separately may be more appropriate than forcing a joint integration [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>].

## Related Bioinformatics Guides

- [Benchmarking Atlas-Level Data Integration in Single-Cell Genomics: Methods and Best Practices](/knowledge/bioinformatics/benchmarking-atlas-level-data-integration-in-single-cell-genomics-methods-and-best-practices)
- [Single-Cell Multi-Omics Integration: Methods and Applications](/knowledge/bioinformatics/single-cell-multi-omics-integration-methods-and-applications)
- [Single-Cell ATAC-Seq Bioinformatics](/knowledge/bioinformatics/single-cell-atac-seq-bioinformatics)
- [Single-Cell Sequencing Depth: How Much Is Enough?](/knowledge/bioinformatics/single-cell-sequencing-depth-how-much-is-enough)
- [Single-Cell RNA Sequencing Quality Control: A Practical Guide to Filtering and Metrics](/knowledge/bioinformatics/single-cell-rna-sequencing-quality-control-a-practical-guide-to-filtering-and-metrics)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)

## References and Further Reading

<a id="ref-1"></a>[<a href="#ref-1">1</a>] [A comparison of integration methods for single-cell RNA sequencing data and ATAC sequencing data.](https://pubmed.ncbi.nlm.nih.gov/41675510). Quantitative biology (Beijing, China), 2025.

<a id="ref-2"></a>[<a href="#ref-2">2</a>] [Benchmarking algorithms for joint integration of unpaired and paired single-cell RNA-seq and ATAC-seq data](https://doi.org/10.1186/s13059-023-03073-x). Genome Biology, 2023.

<a id="ref-3"></a>[<a href="#ref-3">3</a>] [scANANSE gene regulatory network and motif analysis of single-cell clusters.](https://pubmed.ncbi.nlm.nih.gov/38116584). F1000Research, 2023.

<a id="ref-4"></a>[<a href="#ref-4">4</a>] [Integrated ambient modeling and genetic demultiplexing of single-cell RNA+ATAC multiome experiments with Ambimux.](https://pubmed.ncbi.nlm.nih.gov/40909685). bioRxiv : the preprint server for biology, 2025.

<a id="ref-5"></a>[<a href="#ref-5">5</a>] [Benchmarking component choices for unpaired single cell RNA and epigenomic integration.](https://doi.org/10.1186/s13059-026-04071-5). 2026.

<a id="ref-6"></a>[<a href="#ref-6">6</a>] [multiHIVE: Hierarchical Multimodal Deep Generative Modeling for Single-cell Multiomics](https://doi.org/10.21203/rs.3.rs-9422663/v1). 2026.

<a id="ref-7"></a>[<a href="#ref-7">7</a>] [Semi-parametric empirical bayes method for multiplet detection in snATAC-seq with probabilistic multi-omic integration.](https://doi.org/10.1371/journal.pcbi.1013653). 2026.

<a id="ref-8"></a>[<a href="#ref-8">8</a>] [nf-core Documentation](https://nf-co.re/docs). nf-core.

<a id="ref-9"></a>[<a href="#ref-9">9</a>] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.

<a id="ref-10"></a>[<a href="#ref-10">10</a>] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.

<a id="ref-11"></a>[<a href="#ref-11">11</a>] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.

<a id="ref-12"></a>[<a href="#ref-12">12</a>] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.