Benchmarking Single-Cell Multi-Omics Integration Methods: A Review of Metrics, Datasets, and Best Practices

By Dr. Zubair Khalid, DVM, MS, PhD ·

Benchmarking Single-Cell Multi-Omics Integration Methods: A Review of Metrics, Datasets, and Best Practices

Key Takeaways

  • Benchmarking single-cell multi-omics integration requires careful consideration of dataset pairing (paired, unpaired, mosaic), modality combinations (e.g., RNA+ATAC, RNA+protein), and dataset scale (<10k, 10k-100k, >100k cells), as method performance varies significantly across these factors.
  • Evaluation metrics must align with the biological question, encompassing batch mixing (e.g., silhouette width, entropy), cell type conservation (e.g., adjusted Rand index, label transfer accuracy), trajectory preservation (e.g., pseudotime correlation), and imputation/prediction accuracy (e.g., correlation, mean absolute error).
  • No single integration metric captures all aspects of quality; therefore, a multi-metric approach is essential, acknowledging trade-offs between preserving biological structure (cell type conservation, trajectory preservation) and removing technical artifacts (batch mixing).
  • Common failure patterns in benchmarks include overfitting to specific datasets, inappropriate metric selection that misaligns with the integration task, ignoring data quality variations (dropout, batch effects), and computational resource mismatches that render methods impractical for certain infrastructures.
  • Designing a robust benchmark involves defining the integration task, selecting representative datasets with diverse characteristics, choosing appropriate metrics, standardizing inputs, evaluating computational performance, and interpreting results with an understanding of method limitations and potential biases.
  • Visual inspection of embeddings (e.g., UMAP/t-SNE plots colored by batch, cell type, modality), cluster stability assessment, cross-modality consistency checks, and comparison to independent annotations are crucial quality controls for validating integration results beyond quantitative metrics.

Single-cell multi-omics integration methods combine measurements from multiple molecular layers, such as RNA expression, chromatin accessibility, protein abundance, and spatial location, to construct unified cellular representations. Researchers face a practical problem when selecting an integration tool: published methods differ in their underlying assumptions, input requirements, scalability, and performance across dataset types. This article reviews existing benchmarks, explains the metrics used to evaluate integration quality, describes the datasets commonly employed for testing, and provides a workflow for designing your own benchmark study. The guidance is intended for biology students, researchers, laboratory professionals, and life-science practitioners who need to make informed decisions about which integration method to apply to their specific data.

The Integration Problem in Single-Cell Multi-Omics

Single-cell technologies now generate measurements of multiple molecular modalities from individual cells. Some platforms measure paired modalities simultaneously from the same cell, such as gene expression and chromatin accessibility, while others produce unpaired datasets where each modality is profiled in separate cells. A third scenario involves mosaic datasets, which contain a mixture of paired and unpaired measurements across different batches or samples [<a href="#ref-1">1</a>].

The integration task is to map these diverse measurements into a shared representation that preserves biological variation while removing technical artifacts. Integration methods serve several purposes: they enable joint clustering and cell type identification, support cross-modality prediction, facilitate trajectory inference, and allow the construction of reference atlases that can be mapped to new datasets [<a href="#ref-2">2</a>][<a href="#ref-3">3</a>].

The rapid growth of integration tools has created a need for systematic evaluation. Researchers must know whether a method designed for paired data will work on unpaired inputs, whether a method that performs well on small datasets scales to millions of cells, and whether the biological structures preserved by one method are the same structures preserved by another [<a href="#ref-1">1</a>][<a href="#ref-4">4</a>]. Benchmarks address these questions by applying multiple methods to standardized datasets and comparing their outputs using quantitative metrics.

At a Glance: Benchmark Design Decisions

The table below summarizes the key decisions researchers face when designing or interpreting a benchmark study for single-cell multi-omics integration methods.

Decision PointOptionsPractical Consideration
Dataset typePaired, unpaired, mosaicPaired data allows direct validation of cross-modality correspondence, unpaired data requires methods that handle missing anchors, mosaic data tests flexibility [<a href="#ref-1">1</a>]
Integration taskVertical, horizontal, mosaicVertical integration combines different modalities from the same cells, horizontal integration combines the same modality across batches, mosaic integration handles mixed scenarios [<a href="#ref-2">2</a>]
Evaluation metricBatch mixing, cell type conservation, trajectory preservation, imputation accuracyNo single metric captures all aspects of integration quality, choose metrics that match your biological question [<a href="#ref-4">4</a>][<a href="#ref-3">3</a>]
Data scaleSmall (<10k cells), medium (10k-100k), large (>100k)Some methods fail or become computationally prohibitive at scale, test scalability before committing to a method [<a href="#ref-1">1</a>]
Modality combinationRNA + ATAC, RNA + protein, RNA + spatial, multi-modalMethod performance varies by modality combination, benchmarks often show different rankings for different modality pairs [<a href="#ref-2">2</a>][<a href="#ref-4">4</a>]

Core Principles of Integration Method Evaluation

What Integration Methods Actually Do

Integration methods differ in their mathematical frameworks and biological assumptions. Some methods embed different modalities into a common latent space using deep learning architectures. Others use matrix factorization approaches to identify shared factors across modalities. Network-based methods construct graphs of cellular relationships and propagate information across modalities [<a href="#ref-5">5</a>][<a href="#ref-6">6</a>].

The choice of framework affects what the method can and cannot do. Deep learning methods such as scPairing embed paired modalities from the same cells onto a common embedding space, which enables the generation of novel multi-omics data by linking unimodal datasets through an existing multi-omics bridge [<a href="#ref-6">6</a>]. Contrastive learning approaches enable rapid mapping to multimodal atlases at million-cell scale [<a href="#ref-7">7</a>]. Network-based methods for diagonal integration address the specific challenge of unpaired data where transcriptomes and proteomes are measured in separate cells lacking shared anchors [<a href="#ref-5">5</a>].

Why Benchmarking Is Difficult

Benchmarking integration methods presents several challenges that researchers must understand before designing their own evaluations. First, there is no ground truth for what a perfectly integrated dataset should look like. Unlike classification tasks where labels exist, integration quality is defined by multiple competing criteria that may trade off against each other [<a href="#ref-4">4</a>].

Second, methods are developed for different purposes. A method optimized for batch correction may not preserve fine-grained cell states. A method designed for cross-modality prediction may not produce embeddings suitable for trajectory inference. Benchmarks that evaluate methods across multiple tasks provide more useful guidance than those focused on a single metric [<a href="#ref-3">3</a>].

Third, dataset characteristics vary widely. The number of cells, the number of modalities, the balance between paired and unpaired measurements, and the quality of the data all influence method performance. A method that excels on clean, well-annotated reference datasets may fail on noisy, sparse real-world data [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>].

Metrics for Evaluating Integration Quality

Batch Mixing Metrics

Batch mixing metrics assess how well cells from different technical batches or experimental runs are intermingled in the integrated space. The underlying assumption is that true biological variation should be preserved while technical variation from batch effects should be removed.

Common approaches include measuring the entropy of batch labels within local neighborhoods, calculating the average silhouette width across batches, and using graph-based connectivity measures. These metrics are most relevant for horizontal integration tasks where the same modality is measured across multiple batches [<a href="#ref-4">4</a>][<a href="#ref-3">3</a>].

A limitation of batch mixing metrics is that they can be gamed. A method that aggressively merges all cells into a single cluster will achieve perfect batch mixing while destroying biological structure. For this reason, batch mixing metrics must be interpreted alongside cell type conservation metrics [<a href="#ref-4">4</a>].

Cell Type Conservation Metrics

Cell type conservation metrics evaluate whether known cell types remain distinct and correctly grouped after integration. These metrics require annotated reference data, either from the dataset itself or from external sources.

Standard approaches include calculating the adjusted Rand index between the integrated clustering and the reference labels, measuring the purity of clusters with respect to known cell types, and using label transfer accuracy when mapping to a reference atlas. The choice of metric depends on whether the benchmark uses supervised or unsupervised evaluation [<a href="#ref-2">2</a>][<a href="#ref-3">3</a>].

Cell type conservation is particularly important for vertical integration tasks where different modalities may capture different aspects of cell identity. A method that preserves RNA-based cell types but loses protein-based distinctions may be inadequate for studies focused on protein-level heterogeneity [<a href="#ref-2">2</a>].

Trajectory Preservation Metrics

Trajectory preservation metrics assess whether continuous biological processes, such as differentiation or activation, are maintained after integration. These metrics are relevant when the research question involves developmental trajectories or dynamic cellular states.

Evaluation approaches include comparing the ordering of cells along inferred trajectories before and after integration, measuring the correlation between pseudotime values across methods, and assessing whether known intermediate states remain identifiable. Trajectory preservation is often evaluated separately from cluster-based metrics because a method can preserve discrete cell types while distorting continuous transitions [<a href="#ref-4">4</a>][<a href="#ref-3">3</a>].

Imputation and Prediction Accuracy

Some integration methods perform imputation, filling in missing modality measurements based on information from other modalities. Others support cross-modality prediction, where one modality is predicted from another. These capabilities are evaluated using held-out data where known measurements are masked and the method's predictions are compared to the ground truth [<a href="#ref-2">2</a>].

Prediction accuracy metrics include correlation between predicted and observed values, mean absolute error, and classification accuracy for discrete features. The choice of metric depends on whether the predicted modality is continuous, such as protein abundance, or discrete, such as chromatin accessibility peaks [<a href="#ref-2">2</a>].

Feature Selection and Dimensionality Reduction Quality

Integration methods often produce low-dimensional embeddings that serve as inputs to downstream analyses. The quality of these embeddings can be evaluated by how well they preserve the structure present in the original high-dimensional data.

Metrics include reconstruction error when the embedding is mapped back to the original space, neighborhood preservation measures, and the ability of the embedding to separate known biological groups. Some benchmarks also evaluate whether the embedding supports downstream tasks such as clustering and classification [<a href="#ref-3">3</a>].

Datasets Used in Integration Benchmarks

Paired Multi-Omics Datasets

Paired datasets contain measurements of multiple modalities from the same individual cells. These datasets are generated by technologies such as those that simultaneously profile RNA expression and chromatin accessibility, or RNA expression and cell surface protein abundance [<a href="#ref-1">1</a>].

Paired datasets are valuable for benchmarking because they provide ground truth for cross-modality correspondence. A method that correctly integrates paired data should place the same cell in the same location regardless of which modality is used for the embedding. Benchmarks using paired data can directly evaluate whether a method preserves the known relationships between modalities [<a href="#ref-2">2</a>][<a href="#ref-6">6</a>].

The scale of paired datasets is often smaller than unpaired datasets because multi-omics technologies are more expensive than their unimodal counterparts. This limitation means that benchmarks based on paired data may not reflect the performance of methods on large-scale unimodal datasets [<a href="#ref-6">6</a>].

Unpaired and Mosaic Datasets

Unpaired datasets contain measurements of different modalities from separate cells. This scenario arises when technologies are destructive, such as mass spectrometry-based single-cell proteomics, which precludes simultaneous transcriptomic capture [<a href="#ref-5">5</a>]. Mosaic datasets contain a mixture of paired and unpaired measurements, reflecting real-world scenarios where some samples are profiled with multi-omics technologies and others with unimodal technologies [<a href="#ref-1">1</a>].

Integration methods for unpaired and mosaic data face additional challenges because they lack shared anchors between modalities. Network-based methods that propagate information across modality-specific graphs are one approach to this problem [<a href="#ref-5">5</a>]. Benchmarks that include unpaired and mosaic datasets evaluate whether methods can handle these realistic but challenging scenarios [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>].

Reference Atlases as Benchmark Resources

Large-scale reference atlases provide standardized data for benchmarking integration methods. These atlases combine multiple single-cell and single-nucleus assays with spatial imaging technologies to create comprehensive maps of cell types and states in specific tissues [<a href="#ref-8">8</a>].

For example, a human kidney atlas generated from over 400,000 nuclei or cells across healthy and diseased donors defines 51 main cell types and 28 cellular states altered in kidney injury. This atlas integrates transcriptomic profiles, regulatory factors, and spatial localizations, making it a useful benchmark for methods that must integrate diverse data types while preserving biological meaning [<a href="#ref-8">8</a>].

Reference atlases also support the evaluation of label transfer methods, where integration is used to map new datasets onto an existing annotated reference. The quality of the mapping can be assessed by comparing the transferred labels to known annotations [<a href="#ref-8">8</a>][<a href="#ref-7">7</a>].

Practical Workflow for Designing Your Own Benchmark

Step 1: Define the Integration Task

Before selecting methods or metrics, specify the integration task precisely. Determine whether the data are paired, unpaired, or mosaic. Identify the modalities to be integrated. Define the biological question that the integration should support, such as cell type identification, trajectory inference, or cross-modality prediction [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>].

The task definition determines which methods are eligible for evaluation. A method designed for vertical integration of paired RNA and protein data may not be appropriate for horizontal integration of RNA data across batches. Document the task definition before proceeding to method selection [<a href="#ref-2">2</a>][<a href="#ref-4">4</a>].

Step 2: Select Representative Datasets

Choose datasets that reflect the characteristics of the data you expect to encounter in practice. Include datasets with different numbers of cells, different modality combinations, and different levels of data quality. If possible, include both paired and unpaired datasets to test method flexibility [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>].

Consider using publicly available benchmark datasets from published studies. The datasets used in large-scale benchmarks are documented in the literature and can be reused for independent evaluations [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>][<a href="#ref-4">4</a>]. Ensure that the datasets have appropriate annotations for evaluating cell type conservation and other supervised metrics.

Step 3: Choose Evaluation Metrics

Select metrics that match the integration task and the biological question. For studies focused on cell type identification, prioritize cell type conservation metrics. For studies involving developmental processes, include trajectory preservation metrics. For studies that require cross-modality prediction, evaluate prediction accuracy [<a href="#ref-2">2</a>][<a href="#ref-4">4</a>][<a href="#ref-3">3</a>].

Use multiple metrics instead of a single score. Integration quality is multidimensional, and methods that excel on one metric may perform poorly on another. Report all metrics and discuss trade-offs between them [<a href="#ref-4">4</a>][<a href="#ref-3">3</a>].

Step 4: Run Methods with Standardized Inputs

Apply each integration method to the same input data using the method's recommended settings. Document the version of each method, the parameters used, and the computational resources required. This documentation is essential for reproducibility [<a href="#ref-9">9</a>][<a href="#ref-10">10</a>].

Be aware that some methods require specific input formats or preprocessing steps. Standardize the preprocessing where possible, but follow each method's documented requirements to avoid unfair comparisons. Record any deviations from default settings [<a href="#ref-9">9</a>][<a href="#ref-10">10</a>].

Step 5: Evaluate Computational Performance

Measure runtime, memory usage, and scalability for each method. These practical considerations are often as important as integration quality, particularly for large datasets. A method that produces excellent embeddings but cannot handle datasets with more than 100,000 cells may be impractical for many applications [<a href="#ref-1">1</a>].

Record the hardware used for benchmarking, including processor type, memory, and whether a GPU was used. This information allows others to interpret the computational results and compare them to their own infrastructure [<a href="#ref-10">10</a>].

Step 6: Interpret Results and Select Methods

Analyze the metric scores across methods and datasets. Identify methods that perform consistently well across multiple metrics and dataset types. Consider the trade-offs between integration quality, computational cost, and ease of use [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>][<a href="#ref-4">4</a>].

Document the rationale for method selection, including which metrics were prioritized and why. This documentation helps others understand the limitations of the chosen method and provides a basis for future method updates [<a href="#ref-3">3</a>].

Records and Measurements for Benchmark Studies

Documentation Requirements

A reproducible benchmark study requires detailed documentation of the data, methods, and evaluation procedures. Record the version numbers of all software packages, the exact parameters used for each method, and the preprocessing steps applied to each dataset [<a href="#ref-9">9</a>][<a href="#ref-10">10</a>].

Document the computational environment, including operating system, hardware specifications, and software dependencies. This information allows others to replicate the benchmark or adapt it to their own infrastructure [<a href="#ref-10">10</a>].

Data Management

Maintain a clear record of dataset provenance, including where each dataset was obtained, what preprocessing was applied, and how annotations were generated. Public data repositories such as those maintained by the National Center for Biotechnology Information provide standardized access to many benchmark datasets [<a href="#ref-11">11</a>].

For large-scale benchmarks, consider using workflow management systems that track data processing steps and ensure reproducibility. Community standards for workflow documentation support the sharing and reuse of benchmark protocols [<a href="#ref-10">10</a>].

Metric Calculation Records

Record the exact formulas and implementations used for each metric. Different implementations of the same metric can produce different results, particularly for metrics that involve neighborhood definitions or graph construction [<a href="#ref-4">4</a>].

Document any thresholds or parameters used in metric calculation, such as the number of neighbors for neighborhood-based metrics or the resolution parameter for clustering-based metrics. These parameters can substantially affect the results [<a href="#ref-4">4</a>][<a href="#ref-3">3</a>].

Common Failure Patterns in Integration Benchmarks

Overfitting to Benchmark Datasets

Methods that perform well on benchmark datasets may fail on new data because they have been tuned to the specific characteristics of those datasets. This risk is particularly high for methods that require extensive parameter tuning or that use supervised information during training [<a href="#ref-2">2</a>].

To mitigate this risk, evaluate methods on held-out datasets that were not used for method development. Include datasets with different biological contexts, different modalities, and different quality profiles to test generalization [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>].

Inappropriate Metric Selection

Choosing metrics that do not match the integration task can produce misleading conclusions. For example, evaluating a method designed for cross-modality prediction using only batch mixing metrics ignores the method's primary function. Conversely, evaluating a method designed for batch correction using only prediction accuracy may penalize it for characteristics that are not relevant to its intended use [<a href="#ref-4">4</a>][<a href="#ref-3">3</a>].

Select metrics based on the biological question and the intended downstream analysis. If the goal is to identify rare cell populations, evaluate whether the method preserves rare cell types. If the goal is to construct a developmental trajectory, evaluate trajectory preservation [<a href="#ref-4">4</a>][<a href="#ref-3">3</a>].

Ignoring Data Quality Variation

Benchmark datasets are often cleaner than real-world data. Real datasets contain dropout, batch effects, doublets, and other artifacts that can affect integration performance. Methods that perform well on clean benchmark data may fail on noisy data [<a href="#ref-1">1</a>].

Include datasets with varying quality levels in the benchmark. Evaluate whether methods are robust to missing data, low coverage, and technical artifacts. Document the quality characteristics of each dataset so that results can be interpreted in context [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>].

Computational Resource Mismatch

Some integration methods require substantial computational resources, including high-memory nodes or GPUs. Benchmarks that do not account for these requirements may recommend methods that are impractical for researchers with limited infrastructure [<a href="#ref-1">1</a>].

Record the computational resources used for each method and report them alongside integration quality metrics. This information allows researchers to select methods that are feasible for their computing environment [<a href="#ref-10">10</a>][<a href="#ref-1">1</a>].

Limitations of Current Benchmarks

Limited Modality Coverage

Many benchmarks focus on RNA and chromatin accessibility data, with fewer evaluations of protein, methylation, and spatial modalities. The performance of integration methods on these less-studied modality combinations is not well characterized [<a href="#ref-1">1</a>][<a href="#ref-4">4</a>].

Researchers working with non-standard modality combinations should be cautious when extrapolating from existing benchmarks. Consider running a small-scale pilot evaluation before committing to a method for a novel modality combination [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>].

Scale Limitations

Most benchmarks use datasets with fewer than 100,000 cells. The performance of integration methods on datasets with millions of cells is less well characterized, although some methods are specifically designed for large-scale data [<a href="#ref-1">1</a>][<a href="#ref-7">7</a>].

Researchers planning to integrate very large datasets should test methods at the relevant scale before committing to a workflow. Methods that perform well on small datasets may have prohibitive runtime or memory requirements at scale [<a href="#ref-1">1</a>][<a href="#ref-7">7</a>].

Rapid Method Development

The field of single-cell multi-omics integration is evolving rapidly, with new methods published frequently. Benchmarks become outdated quickly as methods are updated and new approaches are introduced [<a href="#ref-1">1</a>][<a href="#ref-3">3</a>].

When selecting an integration method, check whether the method has been updated recently and whether the benchmark results reflect the current version. Consider running a small-scale comparison of candidate methods on your own data instead of relying solely on published benchmarks [<a href="#ref-1">1</a>][<a href="#ref-3">3</a>].

Benchmark Dataset Bias

Benchmark datasets are often generated from specific tissues or biological systems, such as immune cells or cell lines. Methods that perform well on these datasets may not generalize to other biological contexts [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>].

Include datasets from diverse biological systems in your benchmark to test generalization. If your research focuses on a specific tissue or disease, consider generating or collecting data from that context for method evaluation [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>].

Quality Controls for Integration Results

Visual Inspection of Embeddings

Before relying on quantitative metrics, visually inspect the integrated embeddings. Plot the cells in two dimensions using a visualization method such as UMAP or t-SNE, colored by batch, cell type, and modality. Look for obvious artifacts such as modality-specific clusters, batch-specific separation, or the loss of known cell populations [<a href="#ref-4">4</a>].

Visual inspection can reveal problems that quantitative metrics miss. For example, a method may achieve good batch mixing scores while creating artificial structures that do not correspond to any biological grouping [<a href="#ref-4">4</a>].

Cluster Stability Assessment

Evaluate whether the clusters identified in the integrated space are stable across different parameter settings and subsamples of the data. Unstable clusters may indicate that the integration method is producing artifacts instead of capturing true biological structure [<a href="#ref-3">3</a>].

Use resampling approaches to assess cluster stability. Run the integration and clustering pipeline multiple times on subsampled data and compare the resulting cluster assignments. High variability across runs suggests that the results are not reliable [<a href="#ref-3">3</a>].

Cross-Modality Consistency Checks

For paired data, verify that the same cells are placed consistently in the integrated space regardless of which modality is used for the embedding. Inconsistencies may indicate that the integration method is not properly aligning the modalities [<a href="#ref-2">2</a>][<a href="#ref-6">6</a>].

For unpaired data, verify that known biological relationships between modalities are preserved. For example, if a transcription factor is known to regulate a target gene, cells with high expression of the transcription factor should show corresponding changes in the target gene's accessibility or expression [<a href="#ref-12">12</a>].

Comparison to Independent Annotations

Compare the integration results to independent annotations that were not used during integration. These annotations may come from external references, such as published cell type markers, or from orthogonal experimental data, such as immunohistochemistry or spatial imaging [<a href="#ref-8">8</a>].

Agreement between the integration results and independent annotations provides strong evidence that the integration has preserved biological structure. Disagreement may indicate that the integration method has introduced artifacts or that the annotations are incomplete [<a href="#ref-8">8</a>].

Safety and Reproducibility Considerations

Software Environment Management

Integration methods have complex dependencies on programming languages, libraries, and other software. Reproducible benchmarks require careful management of the software environment to ensure that results can be replicated [<a href="#ref-9">9</a>][<a href="#ref-10">10</a>].

Use containerization or virtual environments to isolate the software dependencies for each method. Document the exact versions of all software components. This documentation allows others to reproduce the benchmark in a controlled environment [<a href="#ref-9">9</a>][<a href="#ref-10">10</a>].

Data Storage and Versioning

Benchmark datasets can be large, and their storage and versioning require careful planning. Maintain a record of dataset versions and any modifications applied during preprocessing [<a href="#ref-11">11</a>][<a href="#ref-10">10</a>].

Use version control for analysis scripts and configuration files. This practice ensures that the exact analysis procedures are documented and can be reconstructed if needed [<a href="#ref-13">13</a>].

Training and Skill Development

Benchmarking integration methods requires skills in programming, data management, and statistical analysis. Researchers who lack these skills should seek training before designing their own benchmarks [<a href="#ref-14">14</a>][<a href="#ref-15">15</a>][<a href="#ref-13">13</a>].

Training resources are available from multiple sources. The European Bioinformatics Institute offers training in bioinformatics data resources and analysis [<a href="#ref-14">14</a>]. The Galaxy Training Network provides accessible workflow training and analysis tutorials [<a href="#ref-15">15</a>]. The Carpentries offers foundational computing and data skills [<a href="#ref-13">13</a>].

Professional Escalation Criteria

Some integration problems require expertise beyond what a typical research group can provide. Consider consulting with bioinformatics specialists or core facilities when:

  • The dataset involves novel modality combinations with limited published guidance
  • The integration results are inconsistent across methods and no clear best method emerges
  • The computational requirements exceed available infrastructure
  • The biological conclusions depend critically on the integration method choice
  • The data quality is poor and requires specialized preprocessing

Bioinformatics core facilities and collaborators with expertise in single-cell multi-omics analysis can provide guidance on method selection and benchmark design [<a href="#ref-14">14</a>][<a href="#ref-15">15</a>].

A Decision Framework for Matching Integration Methods to Data Characteristics

Published benchmarks provide valuable comparisons, but they do not directly tell a researcher which method to use for a specific dataset. The gap between benchmark conclusions and practical method selection arises because benchmark datasets differ from real-world data in size, quality, modality balance, and biological complexity. This section presents a structured decision framework that translates benchmark findings into concrete method choices based on observable data characteristics. The framework uses a scoring system that researchers can apply before running any integration method, reducing the risk of selecting an inappropriate tool after investing substantial computational time.

Data Characterization Checklist

Before selecting an integration method, document the following characteristics of your dataset. These characteristics determine which benchmark results are relevant to your situation.

Dataset pairing structure. Determine whether your data are fully paired, fully unpaired, or mosaic. Paired data contain multiple modalities measured from the same cells, such as RNA and chromatin accessibility from the same single cell. Unpaired data contain modalities measured from separate cells, such as transcriptomes and proteomes from different cells when mass spectrometry-based proteomics destroys the cell [<a href="#ref-5">5</a>]. Mosaic data contain a mixture of paired and unpaired measurements across samples [<a href="#ref-1">1</a>]. This distinction is the most important factor in method selection because methods are often designed specifically for one pairing scenario [<a href="#ref-2">2</a>].

Modality combination. Identify which molecular layers are measured. Common combinations include RNA with chromatin accessibility, RNA with protein abundance, and RNA with spatial location. Method performance varies substantially across modality pairs, and benchmarks consistently show different method rankings for different modality combinations [<a href="#ref-2">2</a>][<a href="#ref-4">4</a>]. A method that excels at integrating RNA and protein data may perform poorly on RNA and chromatin accessibility data.

Dataset scale. Count the total number of cells or nuclei. Scale categories are small under 10,000 cells, medium from 10,000 to 100,000 cells, and large above 100,000 cells [<a href="#ref-1">1</a>]. Some methods become computationally prohibitive at scale, while others are specifically designed for million-cell datasets [<a href="#ref-7">7</a>]. Benchmark results from small datasets may not predict performance on large datasets.

Data quality indicators. Assess dropout rates, coverage per cell, and the presence of technical artifacts. Real-world data contain dropout, batch effects, doublets, and other artifacts that can affect integration performance [<a href="#ref-1">1</a>]. Methods that perform well on clean benchmark data may fail on noisy data. Document the quality characteristics of each dataset so that results can be interpreted in context [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>].

Biological complexity. Estimate the number of expected cell types and whether the data contain continuous developmental processes. Datasets with many rare cell types require methods that preserve fine-grained structure. Datasets with developmental trajectories require methods that maintain continuous transitions instead of forcing discrete clusters [<a href="#ref-4">4</a>][<a href="#ref-3">3</a>].

Scoring Matrix for Method Selection

The scoring matrix below translates data characteristics into method suitability scores. For each data characteristic, assign a score from 1 to 5 for each candidate method based on published benchmark results and the method's documented capabilities. A score of 5 indicates strong suitability, 3 indicates moderate suitability, and 1 indicates poor suitability.

Data CharacteristicMethod A ScoreMethod B ScoreMethod C Score
Paired data support534
Unpaired data support253
Mosaic data support345
RNA + ATAC integration435
RNA + protein integration523
Large dataset scalability245
Rare cell type preservation434
Trajectory preservation352
Computational efficiency423
Total score323134

The scoring matrix should be populated using evidence from published benchmarks that match your data characteristics. For example, if your data are unpaired and you need to integrate transcriptomes with proteomes, prioritize methods that have demonstrated performance on diagonal integration tasks where transcriptomes and proteomes are measured in separate cells lacking shared anchors [<a href="#ref-5">5</a>]. If your data are paired and you need to integrate RNA with chromatin accessibility, prioritize methods that have performed well in benchmarks of vertical integration for these modalities [<a href="#ref-2">2</a>][<a href="#ref-4">4</a>].

Weighted Scoring for Research Priorities

The basic scoring matrix treats all characteristics equally, but research priorities differ. A researcher studying developmental biology should weight trajectory preservation more heavily than a researcher focused on cell type classification. Apply weights to the scoring matrix based on your research question.

Define weights that sum to 1.0 across the data characteristics. For example, a study focused on identifying rare cell populations in a large dataset might assign weights of 0.3 to rare cell type preservation, 0.3 to large dataset scalability, 0.2 to paired data support, and 0.2 to RNA and ATAC integration. Multiply each score by its weight and sum the weighted scores to obtain a final ranking.

This weighted approach prevents the common failure pattern of selecting a method based on overall benchmark performance when the benchmark did not match the specific research context. Published benchmarks consistently show that different methods have their own advantages on different aspects, and some methods outperform others only in specific scenarios [<a href="#ref-4">4</a>]. The weighted scoring framework makes these trade-offs explicit.

Thresholds for Method Elimination

Apply elimination thresholds before running any integration method to avoid wasting computational resources on clearly unsuitable tools.

Eliminate methods that do not support your pairing structure. If your data are unpaired, eliminate methods that require paired measurements for their core algorithm. If your data are mosaic, eliminate methods that cannot handle a mixture of paired and unpaired cells [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>].

Eliminate methods that cannot handle your dataset scale. If your dataset exceeds 100,000 cells, eliminate methods that have documented scalability limitations at this scale. Some methods fail or become computationally prohibitive at scale, and this information is often reported in benchmark studies [<a href="#ref-1">1</a>].

Eliminate methods that require input modalities you do not have. Some integration methods require specific preprocessing steps or additional data types. Verify that your data meet the method's input requirements before including the method in your comparison [<a href="#ref-9">9</a>][<a href="#ref-10">10</a>].

Eliminate methods with prohibitive computational requirements. If your computing environment lacks the necessary hardware, such as GPUs or high-memory nodes, eliminate methods that require these resources. Record the computational resources used for each method and report them alongside integration quality metrics [<a href="#ref-10">10</a>][<a href="#ref-1">1</a>].

Method Selection by Integration Scenario

The following scenario-based recommendations synthesize findings from published benchmarks. These recommendations are starting points, not definitive prescriptions, and should be validated on your own data.

Scenario 1: Paired RNA and protein data for cell type identification. For paired data where RNA and cell surface protein abundance are measured from the same cells, methods that embed different modalities from the same cells onto a common embedding space have shown strong performance [<a href="#ref-6">6</a>]. Benchmarks evaluating protein abundance prediction have identified methods that outperform others for this task [<a href="#ref-2">2</a>]. Prioritize methods with demonstrated accuracy in preserving protein-based distinctions while maintaining RNA-based cell type structure.

Scenario 2: Unpaired transcriptome and proteome data. When mass spectrometry-based single-cell proteomics precludes simultaneous transcriptomic capture, you face a diagonal integration challenge where transcriptomes and proteomes are measured in separate cells lacking shared anchors [<a href="#ref-5">5</a>]. Network-based methods that propagate information across modality-specific graphs are one approach to this problem [<a href="#ref-5">5</a>]. Methods designed for horizontal integration across batches may not be appropriate for this scenario because they assume shared features instead of shared cells.

Scenario 3: Mosaic data with mixed paired and unpaired measurements. Mosaic datasets reflect real-world scenarios where some samples are profiled with multi-omics technologies and others with unimodal technologies [<a href="#ref-1">1</a>]. Methods that excel in both horizontal and mosaic integration scenarios have been identified in benchmarks [<a href="#ref-2">2</a>]. Prioritize methods that can handle the flexibility required by mixed pairing structures.

Scenario 4: Large-scale reference atlas mapping. When mapping new datasets onto an existing annotated reference atlas, methods that enable rapid mapping to multimodal atlases at million-cell scale are relevant [<a href="#ref-7">7</a>]. Contrastive learning approaches have demonstrated this capability [<a href="#ref-7">7</a>]. Evaluate whether the method preserves the biological structure of the reference atlas while correctly assigning new cells to known cell types [<a href="#ref-8">8</a>].

Scenario 5: Spatial transcriptomics integration. Spatial transcriptomics adds a spatial modality that requires specialized integration approaches [<a href="#ref-1">1</a>]. Benchmarks that include spatial data evaluate whether methods can integrate spatial information with molecular measurements [<a href="#ref-1">1</a>][<a href="#ref-16">16</a>]. Consider whether the integration method preserves spatial structure and whether it can handle the unique characteristics of spatial data, such as spatial autocorrelation and limited gene coverage [<a href="#ref-1">1</a>][<a href="#ref-16">16</a>].

Validation Protocol for Selected Methods

After using the decision framework to narrow candidate methods, run a validation protocol on your own data before committing to a final method. This protocol addresses the limitation that benchmark results may not generalize to your specific dataset.

Step 1: Run candidate methods on a subsample. Select a representative subsample of your data, such as 10,000 cells, and run each candidate method on this subsample. This step tests whether the methods run successfully on your data format and whether the computational requirements are feasible [<a href="#ref-9">9</a>][<a href="#ref-10">10</a>].

Step 2: Compare integration outputs qualitatively. Visualize the integrated embeddings from each method using a two-dimensional projection. Look for obvious artifacts such as modality-specific clusters, batch-specific separation, or the loss of known cell populations [<a href="#ref-4">4</a>]. Visual inspection can reveal problems that quantitative metrics miss [<a href="#ref-4">4</a>].

Step 3: Calculate quantitative metrics on the subsample. Apply the metrics most relevant to your research question, such as cell type conservation for classification studies or trajectory preservation for developmental studies [<a href="#ref-2">2</a>][<a href="#ref-4">4</a>][<a href="#ref-3">3</a>]. Compare the metric scores across methods and identify the top performers.

Step 4: Scale up to the full dataset. Run the top-performing methods on the full dataset. Monitor runtime and memory usage to ensure the methods can handle the complete data [<a href="#ref-10">10</a>][<a href="#ref-1">1</a>]. Verify that the integration quality observed on the subsample is maintained at full scale.

Step 5: Document the validation results. Record the methods tested, parameters used, computational resources required, and metric scores obtained. This documentation supports reproducibility and provides a basis for future method updates [<a href="#ref-9">9</a>][<a href="#ref-10">10</a>][<a href="#ref-3">3</a>].

Common Decision Errors and Corrections

Error 1: Selecting a method based on overall benchmark ranking without checking dataset similarity. Benchmarks evaluate methods across diverse dataset types, and the overall ranking may not reflect performance on your specific data characteristics [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>][<a href="#ref-4">4</a>]. Correct this error by applying the scoring matrix with weights matched to your research priorities.

Error 2: Assuming that a method performing well on paired data will perform well on unpaired data. Methods are often designed for specific pairing scenarios, and performance does not transfer across scenarios [<a href="#ref-2">2</a>][<a href="#ref-5">5</a>]. Correct this error by verifying that the method supports your pairing structure before including it in the comparison.

Error 3: Ignoring computational requirements when selecting a method. A method that produces excellent embeddings but cannot handle datasets with more than 100,000 cells may be impractical for many applications [<a href="#ref-1">1</a>]. Correct this error by recording the computational resources used for each method and reporting them alongside integration quality metrics [<a href="#ref-10">10</a>][<a href="#ref-1">1</a>].

Error 4: Relying on a single metric to compare methods. Integration quality is multidimensional, and methods that excel on one metric may perform poorly on another [<a href="#ref-4">4</a>][<a href="#ref-3">3</a>]. Correct this error by using multiple metrics and discussing trade-offs between them [<a href="#ref-4">4</a>][<a href="#ref-3">3</a>].

Error 5: Failing to validate benchmark results on your own data. Benchmark datasets are often cleaner than real-world data, and methods that perform well on clean benchmark data may fail on noisy data [<a href="#ref-1">1</a>]. Correct this error by running the validation protocol described above before committing to a final method.

Records for Method Selection Decisions

Maintain a structured record of your method selection process to support reproducibility and future reference.

Data characterization record. Document the pairing structure, modality combination, dataset scale, data quality indicators, and biological complexity of your dataset. This record provides the context for interpreting method selection decisions [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>].

Scoring matrix record. Record the scores assigned to each candidate method for each data characteristic, along with the evidence from published benchmarks that supported each score. This record makes the selection process transparent and auditable [<a href="#ref-2">2</a>][<a href="#ref-4">4</a>].

Validation results record. Document the results of the validation protocol, including qualitative visualizations, quantitative metric scores, and computational performance measurements. This record provides evidence that the selected method performs appropriately on your specific data [<a href="#ref-9">9</a>][<a href="#ref-10">10</a>][<a href="#ref-3">3</a>].

Method version record. Record the version numbers of all software packages used, the exact parameters applied, and the preprocessing steps performed. This documentation is essential for reproducibility [<a href="#ref-9">9</a>][<a href="#ref-10">10</a>].

Escalation Criteria for Method Selection Uncertainty

Some method selection problems require expertise beyond what a typical research group can provide. Consider consulting with bioinformatics specialists or core facilities when:

  • The dataset involves novel modality combinations with limited published guidance
  • The scoring matrix produces similar total scores for multiple methods and no clear best method emerges
  • The validation protocol yields inconsistent results across methods
  • The computational requirements exceed available infrastructure
  • The biological conclusions depend critically on the integration method choice

Bioinformatics core facilities and collaborators with expertise in single-cell multi-omics analysis can provide guidance on method selection and benchmark design [<a href="#ref-14">14</a>][<a href="#ref-15">15</a>]. Training resources are available from multiple sources for researchers who need to develop these skills [<a href="#ref-14">14</a>][<a href="#ref-15">15</a>][<a href="#ref-13">13</a>].

Frequently Asked Questions

What is the difference between vertical, horizontal, and mosaic integration?

Vertical integration combines different modalities measured from the same cells, such as RNA and chromatin accessibility from the same single cell. Horizontal integration combines the same modality measured across different batches or samples. Mosaic integration handles datasets where some cells have paired measurements and others have only one modality [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>]. The integration task determines which methods are appropriate, as methods are often designed for specific integration scenarios [<a href="#ref-2">2</a>].

How do I choose between supervised and unsupervised evaluation metrics?

Supervised metrics require reference annotations, such as known cell types, and evaluate how well the integration preserves these known structures. Unsupervised metrics evaluate properties such as batch mixing without requiring annotations. The choice depends on whether reliable annotations are available for the benchmark datasets [<a href="#ref-4">4</a>][<a href="#ref-3">3</a>]. When annotations are available, use both types of metrics to capture different aspects of integration quality [<a href="#ref-4">4</a>].

What is the minimum number of datasets needed for a credible benchmark?

There is no fixed minimum, but benchmarks that use a single dataset provide limited evidence because method performance varies across dataset characteristics [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>]. Published benchmarks use anywhere from a handful to dozens of datasets, with larger benchmarks providing more robust conclusions [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>][<a href="#ref-4">4</a>]. Include datasets that vary in size, modality combination, and quality to test method generalization [<a href="#ref-1">1</a>].

How do I handle methods that require different input formats?

Standardize the preprocessing where possible, but follow each method's documented input requirements to avoid unfair comparisons [<a href="#ref-9">9</a>][<a href="#ref-10">10</a>]. Document any format conversions and preprocessing steps applied for each method. Some methods require specific normalization or feature selection steps that are part of the method's intended workflow [<a href="#ref-9">9</a>].

Can I trust benchmark results for methods that were developed after the benchmark was published?

Benchmark results reflect the method versions available at the time of the benchmark. Methods are frequently updated, and new versions may perform differently [<a href="#ref-1">1</a>][<a href="#ref-3">3</a>]. Check whether the method has been updated since the benchmark was published and consider running a small-scale comparison on your own data to verify performance [<a href="#ref-1">1</a>].

How do I evaluate integration methods for spatial transcriptomics data?

Spatial transcriptomics adds a spatial modality that requires specialized integration approaches [<a href="#ref-1">1</a>]. Benchmarks that include spatial data evaluate whether methods can integrate spatial information with molecular measurements [<a href="#ref-1">1</a>][<a href="#ref-16">16</a>]. Consider whether the integration method preserves spatial structure and whether it can handle the unique characteristics of spatial data, such as spatial autocorrelation and limited gene coverage [<a href="#ref-1">1</a>][<a href="#ref-16">16</a>].

What should I do when different metrics rank methods differently?

Different metrics capture different aspects of integration quality, and it is common for methods to rank differently across metrics [<a href="#ref-4">4</a>][<a href="#ref-3">3</a>]. Prioritize metrics that are most relevant to your biological question. If the question involves cell type identification, weight cell type conservation more heavily. If the question involves developmental trajectories, weight trajectory preservation more heavily [<a href="#ref-4">4</a>][<a href="#ref-3">3</a>].

How do I report benchmark results in a publication?

Report the datasets, methods, parameters, and metrics used in the benchmark. Include the version numbers of all software and the computational environment [<a href="#ref-9">9</a>][<a href="#ref-10">10</a>]. Provide the metric scores for all methods and datasets, beyond the top performers. Discuss the limitations of the benchmark and the rationale for method selection [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>][<a href="#ref-4">4</a>].

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

[1] [Benchmarking single-cell multi-modal data integrations.](https://pubmed.ncbi.nlm.nih.gov/40640531). Nature methods, 2025. [2] [Benchmarking algorithms for single-cell multi-omics prediction and integration.](https://pubmed.ncbi.nlm.nih.gov/39322753). Nature methods, 2024. [3] [Multitask benchmarking of single-cell multimodal omics integration methods](https://doi.org/10.1038/s41592-025-02856-3). bioRxiv, 2025. [4] [Benchmarking multi-omics integration algorithms across single-cell RNA and ATAC data.](https://pubmed.ncbi.nlm.nih.gov/38493343). Briefings in bioinformatics, 2024. [5] [Network methods for diagonal integration of unpaired single-cell multiomics data: a review.](https://pubmed.ncbi.nlm.nih.gov/42234838). Bioinformatics (Oxford, England), 2026. [6] [Single-cell multiomics data integration and generation with scPairing.](https://pubmed.ncbi.nlm.nih.gov/41151585). Cell reports methods, 2025. [7] [Contrastive learning enables rapid mapping to multimodal single-cell atlas of multimillion scale](https://doi.org/10.1038/s42256-022-00518-z). Nature Machine Intelligence, 2022. [8] [An atlas of healthy and injured cell states and niches in the human kidney.](https://pubmed.ncbi.nlm.nih.gov/37468583). Nature, 2023. [9] [Bioconductor](https://bioconductor.org/). Bioconductor Project. [10] [nf-core Documentation](https://nf-co.re/docs). nf-core. [11] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information. [12] [Gene regulatory network inference in the era of single-cell multi-omics.](https://pubmed.ncbi.nlm.nih.gov/37365273). Nature reviews. Genetics, 2023. [13] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries. [14] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute. [15] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project. [16] [Interpretable data integration for single-cell and spatial multi-omics](https://doi.org/10.1016/j.cels.2025.101479). Cell Systems, 2026.

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.