# Integrating Single-Cell and Spatial Transcriptomics Data: A Review of Computational Strategies and Best Practices


## Key Takeaways

- Integration of single-cell RNA sequencing (scRNA-seq) and spatial transcriptomics (ST) is crucial for a comprehensive understanding of tissue biology, as scRNA-seq provides high-resolution cell-type identification without spatial context, while ST preserves anatomical organization but often requires deconvolution to resolve cell types within mixed spots.
- Key integration strategies include label transfer (mapping known cell types from an annotated scRNA-seq reference to spatial data), deconvolution (estimating cell-type proportions within spatial spots), and joint embedding (creating a shared low-dimensional representation of both data types for exploratory analysis).
- Robust quality control for both scRNA-seq (gene detection, UMI counts, mitochondrial reads, doublet detection) and ST (spot/cell detection rates, transcript counts, tissue mapping) is a prerequisite for successful integration, alongside careful normalization and batch effect correction using methods like Harmony or mutual nearest neighbors.
- The choice of integration strategy is dictated by the research question and data characteristics: label transfer is suitable for mapping known cell types, deconvolution is necessary for mixed spots on array-based platforms, and joint embedding is ideal for exploratory discovery without predefined labels.
- Practical workflows involve sequential steps of quality control, normalization, feature selection, batch correction, integration using chosen methods, and rigorous validation of results through marker gene expression analysis and comparison with independent measurements.
- Common integration challenges include reference mismatch, batch effects dominating biological signals, platform-specific artifacts, overcorrection, and deconvolution bias, necessitating careful troubleshooting and consideration of limitations such as resolution, reference dependence, and computational demands.

---

Single-cell RNA sequencing (scRNA-seq) and spatial transcriptomics (ST) are complementary technologies that together provide a more complete picture of tissue biology than either approach alone. scRNA-seq captures the full transcriptome of individual dissociated cells, while ST preserves the anatomical context of gene expression within intact tissue sections. The integration of these two data types is a computational challenge that requires researchers to align cell populations across platforms, transfer cell-type labels from single-cell data to spatial spots, and construct joint embeddings that reveal how cellular states are organized in space. This article reviews the main computational strategies for integration, including label transfer, deconvolution, and joint embedding, and provides practical guidance for selecting an approach based on data characteristics, research questions, and available computational resources.

The intended audience includes biology students, researchers, laboratory professionals, and life-science practitioners who are planning or conducting single-cell and spatial transcriptomics experiments. The focus is on decision-making: which integration method to use, what quality controls to apply, what records to keep, and when to escalate technical problems to specialists. The article draws on peer-reviewed literature and official bioinformatics training resources to provide evidence-based recommendations.

## The Rationale for Integrating Single-Cell and Spatial Transcriptomics

### What Each Technology Provides

Single-cell RNA sequencing quantifies gene expression in individual cells after tissue dissociation. This approach reveals cellular heterogeneity at high resolution, enabling the identification of rare cell types, developmental trajectories, and cell-state transitions. However, dissociation removes cells from their native tissue context, and the process itself can introduce artifacts such as stress responses and the loss of fragile cell populations. Single-nucleus RNA sequencing (snRNA-seq) offers an alternative that improves representation of recalcitrant lineages and reduces stress signatures, making it particularly useful for frozen tissues and for cell types that are difficult to dissociate [<a href="#ref-1">1</a>].

Spatial transcriptomics measures gene expression within intact tissue sections, preserving the spatial organization of cells. Array-based platforms such as Visium and Stereo-seq capture transcripts from defined spatial locations, while imaging-based methods such as MERFISH and CosMx provide single-cell or subcellular resolution through sequential fluorescence hybridization [<a href="#ref-2">2</a>]. Spatial data reveal how gene expression is organized by position within organs, developmental niches, and disease microenvironments [<a href="#ref-3">3</a>]. The limitation is that many spatial platforms do not achieve true single-cell resolution, and the capture area of each spot may contain multiple cells [<a href="#ref-4">4</a>].

### Why Integration Is Necessary

Neither technology alone provides a complete view. scRNA-seq offers high-resolution cell-type identification but lacks spatial context. ST provides spatial organization but often requires computational deconvolution to resolve cell types within mixed spots [<a href="#ref-4">4</a>]. Integration addresses these gaps by anchoring dissociated single-cell profiles to anatomical coordinates, enabling researchers to map cell types back to their tissue locations and to study cell-cell communication within spatial niches [<a href="#ref-5">5</a>].

The value of integration is well documented across biological systems. In the tumor microenvironment, integration of scRNA-seq and ST has revealed how immune cells, stromal cells, and cancer cells are organized into spatial niches that drive disease progression and therapy resistance [<a href="#ref-4">4</a>][<a href="#ref-5">5</a>]. In neuroscience, integration has enabled the construction of brain spatial genomics atlases that define cellular heterogeneity across brain regions and characterize the spatial organization of neuronal circuits [<a href="#ref-2">2</a>][<a href="#ref-6">6</a>]. In plant biology, integration has shown that gene expression is tightly organized by position within organs, with distinct programs operating in meristems, vascular tissues, and floral organs [<a href="#ref-3">3</a>][<a href="#ref-1">1</a>].

### The Computational Problem

Integration is fundamentally a computational problem of aligning data from different platforms, each with its own technical characteristics. The challenges include differences in resolution, sensitivity, batch effects, and the presence of platform-specific artifacts [<a href="#ref-7">7</a>]. Researchers must choose among several classes of methods, each with distinct assumptions and requirements. The choice depends on whether the goal is to transfer cell-type labels, to estimate cell-type proportions within spatial spots, or to construct a joint representation of both data types in a shared embedding space [<a href="#ref-4">4</a>].

## Core Principles of Data Integration

### Data Inputs and Prerequisites

Before integration can begin, both single-cell and spatial datasets must pass quality control. For scRNA-seq data, standard quality metrics include the number of genes detected per cell, the number of unique molecular identifiers (UMIs), the proportion of mitochondrial reads, and doublet detection. For snRNA-seq data, organellar and intronic metrics require interpretation specific to nuclear preparations [<a href="#ref-1">1</a>]. Ambient RNA contamination, which arises from transcripts released during dissociation, can confound downstream analysis and should be addressed through appropriate filtering or computational correction [<a href="#ref-1">1</a>].

For spatial data, quality control includes assessment of spot or cell detection rates, total transcript counts per spot, and the proportion of reads mapping to the tissue area versus background [<a href="#ref-1">1</a>]. The choice of spatial platform affects the preprocessing steps. Array-based platforms require alignment of spots to tissue morphology, while imaging-based platforms require cell segmentation to define cell boundaries [<a href="#ref-1">1</a>].

### Batch Effects and Harmonization

Batch effects arise from differences in sample preparation, sequencing runs, and platform chemistries. These technical variations can obscure biological signals and must be addressed before integration [<a href="#ref-7">7</a>]. Harmonization methods such as canonical correlation analysis, mutual nearest neighbors, and Harmony align datasets by identifying shared cell populations across batches. The choice of harmonization method depends on the complexity of the experimental design and the degree of batch variation [<a href="#ref-1">1</a>].

For cross-platform integration, batch effects are compounded by differences in capture efficiency, gene coverage, and resolution. Single-cell data typically have higher gene detection per cell, while spatial spots may contain transcripts from multiple cells [<a href="#ref-4">4</a>]. These differences must be accounted for in the integration model.

### The Role of Reference Atlases

Reference atlases provide a shared framework for integration. A well-annotated single-cell atlas can serve as a reference for labeling spatial data through label transfer [<a href="#ref-4">4</a>]. This approach is particularly powerful when the reference covers the cell types expected in the spatial tissue. Public atlases are available for many tissues and organisms, including brain atlases for mouse, human, non-human primate, and zebrafish [<a href="#ref-2">2</a>]. Researchers can also construct their own reference from matched single-cell data generated from the same tissue [<a href="#ref-1">1</a>].

The quality of the reference is critical. A reference with incomplete cell-type coverage or inaccurate annotations will propagate errors to the spatial data [<a href="#ref-4">4</a>]. Researchers should assess reference quality by examining marker gene expression, cluster stability, and the proportion of cells assigned to each cluster [<a href="#ref-1">1</a>].

## At a Glance: Integration Strategies Compared

| Strategy | Input Requirements | Output | Best Use Case | Key Limitations |
| --- | --- | --- | --- | --- |
| Label Transfer | Annotated single-cell reference and spatial data | Cell-type labels for each spatial spot or cell | Mapping known cell types to spatial locations | Requires high-quality reference, fails for novel cell states [<a href="#ref-4">4</a>] |
| Deconvolution | Single-cell reference and spatial data | Estimated cell-type proportions per spot | Estimating cell-type composition in mixed spots | Assumes reference covers all cell types, sensitive to platform differences [<a href="#ref-4">4</a>][<a href="#ref-1">1</a>] |
| Joint Embedding | Unlabeled single-cell and spatial data | Shared low-dimensional representation | Exploring cell states across platforms without predefined labels | Computationally intensive, requires careful batch correction [<a href="#ref-7">7</a>] |

## Computational Strategies for Integration

### Label Transfer Methods

Label transfer methods assign cell-type annotations from a single-cell reference to spatial data [<a href="#ref-4">4</a>]. The process begins with a well-annotated single-cell dataset in which clusters have been assigned cell-type labels based on marker gene expression. The spatial data are then projected into the same feature space, and each spatial spot or cell is assigned the label of the most similar reference cell type [<a href="#ref-1">1</a>].

The most widely used label transfer approaches rely on shared nearest neighbor graphs or canonical correlation analysis to identify correspondences between reference and query cells [<a href="#ref-1">1</a>]. These methods are computationally efficient and work well when the reference covers the cell types present in the spatial tissue. The output is a categorical label for each spatial spot, which can be visualized on the tissue section to reveal the spatial distribution of cell types [<a href="#ref-4">4</a>].

Label transfer is most appropriate when the research question involves mapping known cell types to their locations. For example, in a tumor study, label transfer can map immune cell subtypes identified by scRNA-seq to their spatial positions within the tumor microenvironment [<a href="#ref-4">4</a>]. The main limitation is that label transfer cannot identify cell types that are absent from the reference. If the spatial tissue contains novel or disease-specific cell states, these will be misclassified or assigned to the nearest reference type [<a href="#ref-4">4</a>].

### Deconvolution Methods

Deconvolution methods estimate the proportion of each cell type within each spatial spot [<a href="#ref-4">4</a>]. This approach is necessary because array-based spatial platforms capture transcripts from multiple cells per spot [<a href="#ref-1">1</a>]. Deconvolution uses a single-cell reference to define cell-type-specific gene expression profiles, then solves an optimization problem to find the cell-type proportions that best explain the observed spot-level expression [<a href="#ref-4">4</a>].

Several deconvolution algorithms are available, ranging from regression-based approaches to more sophisticated Bayesian and deep-learning methods [<a href="#ref-7">7</a>]. The choice of algorithm depends on the number of cell types, the similarity between cell types, and the quality of the reference. Deconvolution outputs a matrix of cell-type proportions for each spot, which can be used to study the spatial organization of cell types and to identify niches where specific cell types co-occur [<a href="#ref-4">4</a>].

Deconvolution is more informative than label transfer when spots contain multiple cell types, but it is also more sensitive to reference quality and platform differences [<a href="#ref-1">1</a>]. If the reference does not include all cell types present in the tissue, the deconvolution will assign their transcripts to the nearest available cell type, leading to biased estimates [<a href="#ref-4">4</a>]. Deconvolution also assumes that the reference gene expression profiles are representative of the spatial tissue, which may not hold if dissociation or platform effects alter expression levels [<a href="#ref-1">1</a>].

### Joint Embedding Methods

Joint embedding methods construct a shared low-dimensional representation of single-cell and spatial data without requiring predefined labels [<a href="#ref-7">7</a>]. These methods learn a common feature space in which cells and spots are positioned based on their transcriptional similarity. The joint embedding can then be used for downstream analyses such as clustering, trajectory inference, and cell-cell communication [<a href="#ref-4">4</a>].

Joint embedding approaches vary in their underlying models. Some methods use variational autoencoders to learn a shared latent space, while others use graph-based approaches that align the two data types through shared cell populations [<a href="#ref-7">7</a>]. The output is a set of coordinates for each cell and spot in the joint embedding, which can be visualized using standard dimensionality reduction techniques [<a href="#ref-1">1</a>].

Joint embedding is most appropriate for exploratory analyses where the goal is to discover cell states and spatial patterns without prior annotations [<a href="#ref-7">7</a>]. This approach is computationally intensive and requires careful batch correction to avoid platform-specific clustering [<a href="#ref-7">7</a>]. The interpretation of joint embeddings can be challenging because the biological meaning of the latent dimensions is not always clear [<a href="#ref-7">7</a>].

### Choosing Among Strategies

The choice of integration strategy depends on the research question and the data characteristics [<a href="#ref-4">4</a>]. Label transfer is the simplest and most interpretable approach when a high-quality reference is available and the goal is to map known cell types [<a href="#ref-1">1</a>]. Deconvolution is necessary when spatial spots contain multiple cells and the goal is to estimate cell-type composition [<a href="#ref-4">4</a>]. Joint embedding is appropriate for exploratory analyses and for integrating data from multiple platforms or experiments [<a href="#ref-7">7</a>].

In practice, many studies use a combination of approaches [<a href="#ref-4">4</a>]. For example, a researcher might first use joint embedding to explore the overall structure of the integrated data, then apply label transfer to annotate cell types, and finally use deconvolution to estimate cell-type proportions within spatial spots [<a href="#ref-1">1</a>]. The choice of methods should be documented and justified in the analysis plan [<a href="#ref-8">8</a>].

## Practical Workflow for Integration

### Step 1: Quality Control and Preprocessing

The first step is to ensure that both single-cell and spatial datasets meet quality standards [<a href="#ref-1">1</a>]. For single-cell data, filter cells based on gene detection, UMI counts, and mitochondrial read proportion. For snRNA-seq data, adjust thresholds to account for nuclear-specific metrics [<a href="#ref-1">1</a>]. Remove doublets using computational tools or based on expected cell counts [<a href="#ref-1">1</a>].

For spatial data, assess the quality of the tissue section and the transcript capture [<a href="#ref-1">1</a>]. Remove spots with low transcript counts or those located outside the tissue area. For imaging-based platforms, verify that cell segmentation accurately defines cell boundaries [<a href="#ref-1">1</a>].

### Step 2: Normalization and Feature Selection

Normalize both datasets to account for differences in sequencing depth [<a href="#ref-1">1</a>]. Common approaches include log-normalization and scaling by total transcript counts. Select highly variable genes for integration, as these genes carry the most information about cell-type identity [<a href="#ref-1">1</a>]. The choice of genes should be consistent across datasets to ensure that the integration is based on comparable features [<a href="#ref-1">1</a>].

### Step 3: Batch Correction and Harmonization

Apply batch correction to remove technical variation between samples and platforms [<a href="#ref-7">7</a>]. The choice of method depends on the experimental design. For datasets with clear batch structure, methods such as Harmony or mutual nearest neighbors are effective [<a href="#ref-1">1</a>]. For cross-platform integration, additional steps may be needed to account for differences in gene coverage and capture efficiency [<a href="#ref-1">1</a>].

### Step 4: Integration and Label Transfer

Apply the chosen integration method [<a href="#ref-4">4</a>]. For label transfer, use the annotated single-cell reference to assign labels to spatial data [<a href="#ref-1">1</a>]. For deconvolution, estimate cell-type proportions for each spot [<a href="#ref-4">4</a>]. For joint embedding, construct the shared representation and visualize the results [<a href="#ref-7">7</a>].

### Step 5: Validation and Interpretation

Validate the integration results by examining marker gene expression in the spatial data [<a href="#ref-1">1</a>]. For label transfer, confirm that known cell-type markers are enriched in the expected spatial locations [<a href="#ref-4">4</a>]. For deconvolution, compare estimated proportions with independent measurements such as immunohistochemistry [<a href="#ref-1">1</a>]. For joint embedding, assess whether cells and spots from the same biological condition cluster together [<a href="#ref-7">7</a>].

## Options and Tradeoffs in Platform Selection

### Array-Based Spatial Platforms

Array-based platforms such as Visium and Stereo-seq capture transcripts from defined spatial locations on a slide [<a href="#ref-2">2</a>][<a href="#ref-1">1</a>]. These platforms provide whole-transcriptome coverage and are relatively straightforward to use [<a href="#ref-1">1</a>]. The main limitation is spatial resolution: each spot captures transcripts from multiple cells, requiring deconvolution to resolve cell types [<a href="#ref-4">4</a>]. Visium HD and newer iterations of Stereo-seq offer improved resolution, but the computational challenges of deconvolution remain [<a href="#ref-1">1</a>].

### Imaging-Based Spatial Platforms

Imaging-based platforms such as MERFISH, CosMx, and smFISH provide single-cell or subcellular resolution through sequential hybridization of fluorescent probes [<a href="#ref-2">2</a>][<a href="#ref-1">1</a>]. These platforms offer higher resolution than array-based methods but are typically limited to a predefined panel of genes [<a href="#ref-1">1</a>]. The choice of gene panel is critical and should be based on the cell types and pathways of interest [<a href="#ref-1">1</a>]. Imaging-based platforms are well suited for validating findings from array-based studies and for targeted analysis of specific cell populations [<a href="#ref-1">1</a>].

### Single-Cell Versus Single-Nucleus Data

The choice between scRNA-seq and snRNA-seq affects integration outcomes [<a href="#ref-1">1</a>]. scRNA-seq provides higher gene detection per cell but introduces dissociation artifacts and cell-type biases [<a href="#ref-1">1</a>]. snRNA-seq improves representation of recalcitrant lineages and reduces stress signatures, making it more suitable for frozen tissues and for cell types that are difficult to dissociate [<a href="#ref-1">1</a>]. The choice should be based on the tissue type and the research question [<a href="#ref-1">1</a>]. For integration with spatial data, snRNA-seq may be preferable when the spatial tissue is frozen and the goal is to match nuclear transcriptomes [<a href="#ref-1">1</a>].

### Microfluidics and Emerging Technologies

Microfluidics-based single-cell analysis has advanced from transcriptomics to spatiotemporal multi-omics, enabling the measurement of multiple molecular modalities from the same cell [<a href="#ref-9">9</a>]. These technologies are expanding the scope of integration by providing matched measurements of gene expression, chromatin accessibility, and protein abundance [<a href="#ref-9">9</a>]. However, these approaches are technically demanding and may not be accessible to all laboratories [<a href="#ref-9">9</a>].

## Observations and Measurements for Quality Assessment

### Metrics for Single-Cell Data Quality

The number of genes detected per cell is a primary indicator of data quality [<a href="#ref-1">1</a>]. Low gene detection may indicate poor cell viability or technical issues. The proportion of mitochondrial reads is a marker of cell stress: high proportions suggest that cells are damaged or dying [<a href="#ref-1">1</a>]. For snRNA-seq data, the proportion of reads mapping to introns is expected to be higher than for scRNA-seq, and this metric should be interpreted accordingly [<a href="#ref-1">1</a>].

Doublet detection is important because doublets can create spurious cell types that confound integration [<a href="#ref-1">1</a>]. Computational doublet detection methods estimate the likelihood that each cell is a doublet based on the number of UMIs and the co-expression of mutually exclusive marker genes [<a href="#ref-1">1</a>].

### Metrics for Spatial Data Quality

For array-based platforms, the number of transcripts per spot and the number of genes detected per spot indicate capture efficiency [<a href="#ref-1">1</a>]. Spots with very low counts may be located outside the tissue or in regions of poor tissue quality [<a href="#ref-1">1</a>]. The proportion of reads mapping to the tissue area versus background indicates the specificity of capture [<a href="#ref-1">1</a>].

For imaging-based platforms, the number of cells detected and the number of transcripts per cell indicate the sensitivity of the assay [<a href="#ref-1">1</a>]. The accuracy of cell segmentation is critical: poorly segmented cells can lead to mixed transcriptomes that confound downstream analysis [<a href="#ref-1">1</a>].

### Metrics for Integration Quality

After integration, several metrics can assess the quality of the alignment [<a href="#ref-1">1</a>]. The proportion of cells or spots that are assigned to the expected cell types indicates the accuracy of label transfer [<a href="#ref-4">4</a>]. The correlation between estimated and measured cell-type proportions indicates the accuracy of deconvolution [<a href="#ref-1">1</a>]. The mixing of cells and spots from different batches in the joint embedding indicates the effectiveness of batch correction [<a href="#ref-7">7</a>].

## Records and Documentation

### What to Record

Detailed records are essential for reproducible integration analysis [<a href="#ref-8">8</a>]. The following information should be documented for each dataset:

- Sample metadata, including tissue type, condition, and biological replicates
- Platform and chemistry used for both single-cell and spatial data
- Quality control thresholds and the number of cells or spots removed at each step
- Normalization and batch correction methods, including parameter settings
- Integration method and version, including any reference datasets used
- Software versions and computational environment

### Reproducibility Practices

Reproducibility requires that the analysis can be repeated by others [<a href="#ref-8">8</a>]. Version control for code and data is essential [<a href="#ref-10">10</a>]. Workflow management systems such as nf-core provide standardized pipelines that document each analysis step and ensure consistency across runs [<a href="#ref-8">8</a>]. Containerization ensures that software dependencies are captured and reproducible [<a href="#ref-8">8</a>].

Public data repositories such as NCBI provide a platform for sharing raw and processed data [<a href="#ref-11">11</a>]. Depositing data and code enables others to verify results and to reuse the data for new analyses [<a href="#ref-11">11</a>]. The NCBI provides search systems and sequence resources that support data sharing and reuse [<a href="#ref-11">11</a>].

### Reporting Standards

Publications should report the integration methods used, including software versions and parameter settings [<a href="#ref-8">8</a>]. The quality control metrics should be reported for both single-cell and spatial data [<a href="#ref-1">1</a>]. The validation of integration results should be described, including any independent measurements used for confirmation [<a href="#ref-1">1</a>].

## Common Failure Patterns and Troubleshooting

### Reference Mismatch

A common failure is when the single-cell reference does not match the cell types present in the spatial tissue [<a href="#ref-4">4</a>]. This can occur when the reference is generated from a different tissue, developmental stage, or disease condition [<a href="#ref-4">4</a>]. The result is misclassification of spatial spots and inaccurate deconvolution [<a href="#ref-4">4</a>]. To address this, researchers should verify that the reference covers the expected cell types and consider generating a matched reference from the same tissue [<a href="#ref-1">1</a>].

### Batch Effects Dominating Biological Signal

When batch effects are not adequately corrected, the integrated data will cluster by batch instead of by cell type [<a href="#ref-7">7</a>]. This is often visible in the joint embedding as separate clusters for each batch [<a href="#ref-7">7</a>]. To address this, researchers should apply stronger batch correction methods or increase the number of mutual nearest neighbors used for alignment [<a href="#ref-1">1</a>].

### Platform-Specific Artifacts

Differences in gene coverage and capture efficiency between platforms can create artifacts in the integrated data [<a href="#ref-1">1</a>]. For example, genes that are poorly captured by the spatial platform may appear to be downregulated in spatial data compared to single-cell data [<a href="#ref-1">1</a>]. To address this, researchers should select highly variable genes that are well captured by both platforms and should interpret platform-specific differences cautiously [<a href="#ref-1">1</a>].

### Overcorrection

Aggressive batch correction can remove biological variation along with technical variation [<a href="#ref-7">7</a>]. This is particularly problematic when the biological signal is correlated with batch, such as when different conditions are processed in different batches [<a href="#ref-7">7</a>]. To address this, researchers should use conservative batch correction settings and should validate that known biological differences are preserved after correction [<a href="#ref-7">7</a>].

### Deconvolution Bias

Deconvolution methods can produce biased estimates when the reference does not include all cell types or when cell types are highly similar [<a href="#ref-4">4</a>]. The result is that transcripts from missing cell types are assigned to the nearest available cell type [<a href="#ref-4">4</a>]. To address this, researchers should assess the completeness of the reference and should consider using methods that are robust to missing cell types [<a href="#ref-1">1</a>].

## Limitations of Current Integration Methods

### Resolution Limits

The resolution of spatial platforms limits the granularity of integration [<a href="#ref-3">3</a>][<a href="#ref-1">1</a>]. Array-based platforms capture transcripts from multiple cells per spot, requiring deconvolution to resolve cell types [<a href="#ref-4">4</a>]. Even with deconvolution, the spatial organization of individual cells within a spot cannot be resolved [<a href="#ref-1">1</a>]. Imaging-based platforms offer higher resolution but are limited to predefined gene panels [<a href="#ref-1">1</a>].

### Reference Dependence

Most integration methods depend on the quality and completeness of the single-cell reference [<a href="#ref-4">4</a>]. If the reference is incomplete or inaccurate, the integration results will be biased [<a href="#ref-4">4</a>]. This is a particular concern for disease tissues, which may contain cell states that are absent from healthy references [<a href="#ref-4">4</a>][<a href="#ref-5">5</a>].

### Computational Demands

Joint embedding methods are computationally intensive and require substantial memory and processing power [<a href="#ref-7">7</a>]. This can be a barrier for laboratories without access to high-performance computing [<a href="#ref-7">7</a>]. Cloud-based platforms and workflow management systems can help, but the computational demands remain a consideration [<a href="#ref-8">8</a>].

### Temporal Resolution

Spatial transcriptomics provides a snapshot of gene expression at a single time point [<a href="#ref-3">3</a>]. The integration of spatial data with single-cell data does not capture dynamic changes over time [<a href="#ref-3">3</a>]. Studies of development and disease progression require multiple time points, which increases the complexity and cost of experiments [<a href="#ref-3">3</a>].

### Correlation Versus Causation

Integration reveals correlations between gene expression and spatial location, but it does not establish causation [<a href="#ref-3">3</a>]. The identification of spatially defined regulators requires functional perturbation experiments, which are beyond the scope of transcriptomics alone [<a href="#ref-3">3</a>].

## Safety and Regulatory Context

### Data Sharing and Privacy

Single-cell and spatial transcriptomics data from human samples may contain sensitive information [<a href="#ref-11">11</a>]. Researchers must comply with data protection regulations and institutional review board requirements [<a href="#ref-11">11</a>]. De-identification of samples and controlled access to data are standard practices [<a href="#ref-11">11</a>]. Public repositories such as NCBI provide mechanisms for controlled access to sensitive data [<a href="#ref-11">11</a>].

### Ethical Considerations

The use of human tissues for single-cell and spatial transcriptomics requires informed consent and ethical approval [<a href="#ref-12">12</a>]. Researchers should ensure that sample collection and data sharing comply with ethical guidelines [<a href="#ref-12">12</a>]. For animal studies, institutional animal care and use committee approval is required [<a href="#ref-13">13</a>].

### Computational Reproducibility

Reproducibility is a scientific and ethical obligation [<a href="#ref-8">8</a>]. Researchers should document their analysis methods and share code and data to enable verification [<a href="#ref-8">8</a>]. Workflow management systems and containerization support reproducible analysis [<a href="#ref-8">8</a>].

## Professional Escalation Criteria

### When to Seek Specialized Support

Researchers should consider escalating to specialized bioinformatics support in the following situations:

- The integration results are inconsistent with known biology, such as cell types appearing in unexpected locations [<a href="#ref-4">4</a>]
- The quality control metrics indicate severe technical issues that cannot be resolved with standard filtering [<a href="#ref-1">1</a>]
- The computational demands exceed available resources [<a href="#ref-7">7</a>]
- The analysis requires methods that are not available in standard software packages [<a href="#ref-7">7</a>]
- The interpretation of results requires specialized expertise, such as for clinical translation [<a href="#ref-14">14</a>]

### Consulting Official Resources

Several official resources provide training and support for single-cell and spatial transcriptomics analysis [<a href="#ref-15">15</a>][<a href="#ref-16">16</a>][<a href="#ref-17">17</a>][<a href="#ref-8">8</a>][<a href="#ref-10">10</a>]. The EMBL-EBI Training program offers learning pathways for bioinformatics data resources and practical analysis education [<a href="#ref-15">15</a>]. The Galaxy Training Network provides accessible workflow training and analysis tutorials [<a href="#ref-17">17</a>]. The Carpentries offers foundational computing and data skills, including shell, Git, and programming training [<a href="#ref-10">10</a>]. The Bioconductor Project provides official package, workflow, and reproducible genomic-analysis documentation [<a href="#ref-16">16</a>]. The nf-core documentation describes community pipeline standards and usage [<a href="#ref-8">8</a>].

## Applications Across Biological Systems

### Tumor Microenvironment Research

Integration of scRNA-seq and ST has transformed the study of the tumor microenvironment [<a href="#ref-4">4</a>][<a href="#ref-5">5</a>]. These technologies jointly reveal cellular heterogeneity, stromal-immune interactions, and spatial niches that drive tumor progression and therapy resistance [<a href="#ref-4">4</a>][<a href="#ref-5">5</a>]. Deconvolution and mapping approaches have been used to characterize immune evasion, fibroblast diversity, and cell-cell communication networks [<a href="#ref-4">4</a>]. The integration of radiomics with spatial omics is an emerging direction that combines non-invasive imaging with spatially resolved molecular data for precision diagnosis and treatment [<a href="#ref-14">14</a>].

### Neuroscience and Brain Disorders

Spatial genomics technologies have enabled the construction of brain atlases that define cellular heterogeneity across brain regions [<a href="#ref-2">2</a>]. Integration with scRNA-seq has revealed the spatial organization of neuronal circuits and identified region-specific molecular signatures associated with neurological disorders [<a href="#ref-2">2</a>][<a href="#ref-6">6</a>]. These atlases provide resources for investigating brain development, function, and disease [<a href="#ref-2">2</a>].

### Plant Biology

Spatial transcriptomics has transformed plant biology by restoring the spatial context lost in dissociation-based approaches [<a href="#ref-3">3</a>][<a href="#ref-1">1</a>]. Studies of meristems, vascular tissues, and floral organs reveal spatially segregated developmental programs [<a href="#ref-3">3</a>]. Integration with single-cell data has been used to study plant-microbe interactions, photosynthesis, drought adaptation, and regeneration [<a href="#ref-3">3</a>][<a href="#ref-1">1</a>]. The plant spatial transcriptomics community faces limitations including restricted spatial resolution, reliance on computational deconvolution, and uneven taxonomic coverage [<a href="#ref-3">3</a>].

### Parasitic Nematodes

Single-cell and single-nucleus transcriptomics have transformed the study of parasitic nematodes, which are pathogens of medical and veterinary importance [<a href="#ref-13">13</a>]. These approaches have enabled the definition of cellular identity and its integration with biological function [<a href="#ref-13">13</a>]. Spatial transcriptomics provides access to tissue-specific profiling, while organoid and co-culture systems provide access to host-parasite interfaces [<a href="#ref-13">13</a>]. No single system yet provides life-cycle-wide cellular coverage, spatially resolved tissue organization, sustained culture, and reliable functional perturbation [<a href="#ref-13">13</a>].

### Islet Transplantation

Integration of scRNA-seq and ST has been applied to islet transplantation for diabetes treatment [<a href="#ref-12">12</a>]. These technologies provide a detailed view of the diversity and functionality within islet grafts, identifying cell types and states that influence graft acceptance and function [<a href="#ref-12">12</a>]. Spatial transcriptomics maps gene expression within the tissue context, revealing the microenvironment surrounding transplanted islets and their interactions with host tissues [<a href="#ref-12">12</a>].

### Atherosclerosis

Single-cell analysis technologies have clarified the cellular and molecular complexity of atherosclerotic plaques [<a href="#ref-18">18</a>]. scRNA-seq, single-cell ATAC-seq, and spatial transcriptomics have revealed diverse immune and vascular cell states [<a href="#ref-18">18</a>]. Integration of transcriptomic, epigenomic, proteomic, and spatial data has clarified mechanisms of disease progression and identified cell-specific molecular pathways responsive to targeted therapy [<a href="#ref-18">18</a>].

## Artificial Intelligence in Integration

### Current State of AI Methods

Artificial intelligence has become a common tool for bioinformatics, with hundreds of methods published in recent years [<a href="#ref-7">7</a>]. Deep-learning algorithms are particularly popular for single-cell and spatial transcriptomics because of the training data demands of these approaches [<a href="#ref-7">7</a>]. AI methods have been applied to dimensionality reduction, cross-dataset integration, data denoising, data augmentation, deconvolution, cell-cell interactions, transcriptional velocity, and the integration of transcriptomic and chromatin accessibility data [<a href="#ref-7">7</a>].

### Readiness for General Use

The readiness of AI methods for general research use varies by task [<a href="#ref-7">7</a>]. Some methods are likely to be useful for discovery researchers, while others are not yet ready for general use [<a href="#ref-7">7</a>]. Researchers should evaluate AI methods critically and should compare their performance with alternative statistical or heuristic-based approaches [<a href="#ref-7">7</a>]. The choice of method should be based on the specific analysis task and the characteristics of the data [<a href="#ref-7">7</a>].

### Challenges and Future Directions

The application of AI to integration faces several challenges, including the need for large training datasets, the interpretability of deep-learning models, and the generalization of models across platforms and tissues [<a href="#ref-7">7</a>]. Future directions include multimodal fusion, dynamic monitoring, and personalized therapeutic strategies [<a href="#ref-14">14</a>]. The integration of AI with spatial multi-omics is poised to advance precision medicine, but the full clinical potential depends on closing the gap between analytical innovation and robust clinical implementation [<a href="#ref-4">4</a>].

## A Decision Framework for Selecting Integration Methods Based on Data Characteristics

Researchers often struggle to choose among label transfer, deconvolution, and joint embedding because published comparisons emphasize algorithmic performance instead of practical fit. The decision framework below translates data characteristics and research goals into concrete method choices. It is organized around five diagnostic questions that any research team can answer before committing computational resources.

### Diagnostic Question 1: What Is the Spatial Resolution of Your Platform?

The spatial platform determines whether label transfer or deconvolution is even applicable. Array-based platforms such as Visium and Stereo-seq capture transcripts from spots that contain multiple cells, so each spot represents a mixture of cell types [<a href="#ref-1">1</a>]. Label transfer assigns a single label per spot, which is only meaningful when one cell type dominates the spot. If your platform produces multi-cell spots, deconvolution is the appropriate strategy for estimating cell-type proportions [<a href="#ref-4">4</a>].

Imaging-based platforms such as MERFISH, CosMx, and smFISH provide single-cell or subcellular resolution [<a href="#ref-2">2</a>][<a href="#ref-1">1</a>]. With these platforms, label transfer can assign a cell-type label to each segmented cell directly, and deconvolution becomes unnecessary because the spot is already a single cell [<a href="#ref-1">1</a>]. Joint embedding remains useful for both platform types when the goal is exploratory analysis of cell states across platforms [<a href="#ref-7">7</a>].

The practical decision rule is straightforward. If your spatial data are from an array-based platform with multi-cell spots, plan for deconvolution as the primary integration strategy. If your spatial data are from an imaging-based platform with single-cell resolution, label transfer is the primary strategy. Joint embedding serves as a complement in both cases for exploring cell states without predefined labels [<a href="#ref-7">7</a>].

### Diagnostic Question 2: Does Your Single-Cell Reference Cover All Expected Cell Types?

Reference completeness is the most common source of integration failure [<a href="#ref-4">4</a>]. Before choosing a method, inventory the cell types expected in your tissue based on prior literature or preliminary clustering. Compare this list against the cell types annotated in your single-cell reference.

If the reference covers all expected cell types, label transfer and deconvolution are both viable [<a href="#ref-4">4</a>]. If the reference is missing cell types that are likely present in the spatial tissue, label transfer will misclassify those cells into the nearest available type, and deconvolution will assign their transcripts to the nearest available profile [<a href="#ref-4">4</a>]. In this situation, joint embedding is the safer choice because it does not require predefined labels and can reveal novel cell states [<a href="#ref-7">7</a>].

For disease tissues, references generated from healthy tissue are often incomplete because disease-specific cell states may be absent [<a href="#ref-4">4</a>][<a href="#ref-5">5</a>]. Researchers should consider generating a matched single-cell reference from the same tissue and condition instead of relying on public atlases [<a href="#ref-1">1</a>]. The effort required to generate a matched reference is justified when the spatial tissue contains disease-specific states that are central to the research question [<a href="#ref-5">5</a>].

### Diagnostic Question 3: What Is the Primary Research Question?

The research question determines which integration output is most useful. If the goal is to map known cell types to their spatial locations, label transfer provides the most interpretable output: a categorical label for each spot or cell that can be visualized directly on the tissue section [<a href="#ref-4">4</a>].

If the goal is to quantify the composition of mixed spots, deconvolution provides the necessary output: estimated proportions of each cell type per spot [<a href="#ref-4">4</a>]. This is particularly relevant for studying cell-type co-occurrence and spatial niches in the tumor microenvironment [<a href="#ref-4">4</a>][<a href="#ref-5">5</a>].

If the goal is to discover novel cell states or to compare cell states across platforms without predefined annotations, joint embedding is the appropriate choice [<a href="#ref-7">7</a>]. Joint embedding outputs a shared low-dimensional representation that supports clustering, trajectory inference, and cell-cell communication analysis [<a href="#ref-4">4</a>].

Many studies combine strategies sequentially [<a href="#ref-4">4</a>]. A common workflow is to first apply joint embedding to explore the overall structure, then use label transfer to annotate cell types, and finally apply deconvolution to estimate proportions within multi-cell spots [<a href="#ref-1">1</a>]. This combined approach is appropriate when the research question has multiple components, such as identifying cell types and then quantifying their spatial organization.

### Diagnostic Question 4: How Large Is the Batch Effect Between Platforms?

Cross-platform integration is complicated by differences in capture efficiency, gene coverage, and resolution [<a href="#ref-4">4</a>][<a href="#ref-7">7</a>]. The magnitude of these differences affects method choice. Label transfer methods that rely on shared nearest neighbor graphs or canonical correlation analysis are relatively robust to moderate batch effects because they identify correspondences based on local structure [<a href="#ref-1">1</a>].

Deconvolution methods are more sensitive to platform differences because they assume that the reference gene expression profiles are representative of the spatial tissue [<a href="#ref-1">1</a>]. If dissociation or platform effects alter expression levels, deconvolution estimates will be biased [<a href="#ref-1">1</a>]. Researchers should assess the magnitude of platform differences by comparing the expression of known marker genes between the single-cell and spatial datasets before choosing deconvolution.

Joint embedding methods that use deep learning can model complex platform-specific effects but require careful batch correction to avoid clustering by platform instead of by cell type [<a href="#ref-7">7</a>]. The risk of overcorrection is real: aggressive batch correction can remove biological variation along with technical variation [<a href="#ref-7">7</a>]. Researchers should validate that known biological differences are preserved after correction [<a href="#ref-7">7</a>].

### Diagnostic Question 5: What Computational Resources Are Available?

Joint embedding methods are computationally intensive and require substantial memory and processing power [<a href="#ref-7">7</a>]. Label transfer is the least demanding, followed by deconvolution [<a href="#ref-1">1</a>]. Laboratories without access to high-performance computing should prioritize label transfer or deconvolution and reserve joint embedding for cases where exploratory analysis is essential [<a href="#ref-7">7</a>].

Cloud-based platforms and workflow management systems can help distribute computational load [<a href="#ref-8">8</a>]. The nf-core documentation describes community pipeline standards that support reproducible and scalable analysis [<a href="#ref-8">8</a>]. The Galaxy Training Network provides accessible workflow training that can help researchers implement integration methods without local high-performance computing [<a href="#ref-17">17</a>].

### A Scoring Matrix for Method Selection

The following scoring matrix translates the five diagnostic questions into a practical selection tool. For each question, assign a score of 1 to 3 based on your data characteristics, then sum the scores for each method.

| Diagnostic Question | Label Transfer Score 3 | Deconvolution Score 3 | Joint Embedding Score 3 |
| --- | --- | --- | --- |
| Spatial resolution | Single-cell resolution | Multi-cell spots | Either resolution |
| Reference completeness | Complete reference | Complete reference | Incomplete reference acceptable |
| Research question | Map known cell types | Quantify spot composition | Discover novel states |
| Batch effect magnitude | Moderate batch effects | Small batch effects | Large batch effects manageable |
| Computational resources | Limited resources | Moderate resources | High-performance computing |

Score each method across the five questions. The method with the highest total score is the best fit for your data. In practice, most studies will score highly on different methods for different questions, which supports the combined approach described above [<a href="#ref-4">4</a>][<a href="#ref-1">1</a>].

### Recording the Decision Process

Document the answers to all five diagnostic questions before running the integration. Record the spatial platform and its resolution, the cell-type inventory of the reference, the primary research question, the observed batch effect magnitude, and the available computational resources. This record serves two purposes. First, it justifies the method choice in publications and grant reports [<a href="#ref-8">8</a>]. Second, it provides a baseline for troubleshooting if integration results are unexpected [<a href="#ref-1">1</a>].

The decision framework should be revisited if the integration results are inconsistent with known biology [<a href="#ref-4">4</a>]. For example, if label transfer places a cell type in a location where it is not expected, the reference completeness or the spatial resolution assumptions may be incorrect [<a href="#ref-4">4</a>]. Reassess the diagnostic questions and consider switching to a different method [<a href="#ref-1">1</a>].

### Limitations of the Decision Framework

This framework assumes that the researcher has sufficient knowledge of the tissue to inventory expected cell types and to assess whether integration results are biologically plausible. For tissues with poorly characterized cell-type composition, the framework is less useful because the reference completeness question cannot be answered reliably [<a href="#ref-4">4</a>]. In these cases, joint embedding is the default choice because it does not require predefined labels [<a href="#ref-7">7</a>].

The framework also assumes that the single-cell and spatial datasets are of comparable quality. If one dataset has severe technical issues, no integration method will produce reliable results [<a href="#ref-1">1</a>]. Quality control must be completed before applying the decision framework [<a href="#ref-1">1</a>].

The framework does not address the choice of specific algorithms within each strategy. For example, deconvolution algorithms vary in their assumptions about cell-type similarity and reference completeness [<a href="#ref-7">7</a>]. Researchers should evaluate multiple algorithms within the chosen strategy and compare their outputs [<a href="#ref-7">7</a>]. The framework identifies the strategy, but algorithm selection remains an empirical decision based on the specific dataset [<a href="#ref-7">7</a>].

## Frequently Asked Questions

### What is the difference between label transfer and deconvolution?

Label transfer assigns a single cell-type label to each spatial spot or cell based on similarity to a reference single-cell dataset [<a href="#ref-4">4</a>]. Deconvolution estimates the proportion of each cell type within each spatial spot [<a href="#ref-4">4</a>]. Label transfer is appropriate when each spot contains a single dominant cell type, while deconvolution is necessary when spots contain multiple cell types [<a href="#ref-4">4</a>][<a href="#ref-1">1</a>].

### How do I choose between scRNA-seq and snRNA-seq for integration with spatial data?

The choice depends on the tissue type and the research question [<a href="#ref-1">1</a>]. scRNA-seq provides higher gene detection per cell but introduces dissociation artifacts and cell-type biases [<a href="#ref-1">1</a>]. snRNA-seq improves representation of recalcitrant lineages and reduces stress signatures, making it more suitable for frozen tissues [<a href="#ref-1">1</a>]. For integration with spatial data from frozen tissues, snRNA-seq may be preferable [<a href="#ref-1">1</a>].

### What quality controls are essential before integration?

Essential quality controls include filtering cells or spots based on gene detection, UMI counts, and mitochondrial read proportion [<a href="#ref-1">1</a>]. Doublet detection is important for single-cell data [<a href="#ref-1">1</a>]. For spatial data, assess transcript capture and remove spots outside the tissue area [<a href="#ref-1">1</a>]. Batch correction is essential when integrating data from multiple samples or platforms [<a href="#ref-7">7</a>].

### How do I validate integration results?

Validation involves examining marker gene expression in the spatial data to confirm that known cell-type markers are enriched in the expected locations [<a href="#ref-1">1</a>]. For deconvolution, compare estimated proportions with independent measurements such as immunohistochemistry [<a href="#ref-1">1</a>]. For joint embedding, assess whether cells and spots from the same biological condition cluster together [<a href="#ref-7">7</a>].

### What are the main limitations of current integration methods?

The main limitations include resolution limits of spatial platforms, dependence on reference quality, computational demands of joint embedding methods, lack of temporal resolution, and the gap between correlation and causation [<a href="#ref-3">3</a>][<a href="#ref-4">4</a>][<a href="#ref-1">1</a>][<a href="#ref-7">7</a>]. Researchers should be aware of these limitations when interpreting integration results [<a href="#ref-3">3</a>].

### When should I escalate to specialized bioinformatics support?

Escalate when integration results are inconsistent with known biology, when quality control metrics indicate severe technical issues, when computational demands exceed available resources, when the analysis requires specialized methods, or when interpretation requires specialized expertise [<a href="#ref-4">4</a>][<a href="#ref-1">1</a>][<a href="#ref-7">7</a>]. Official training resources are available from EMBL-EBI, Galaxy, The Carpentries, Bioconductor, and nf-core [<a href="#ref-15">15</a>][<a href="#ref-16">16</a>][<a href="#ref-17">17</a>][<a href="#ref-8">8</a>][<a href="#ref-10">10</a>].

### How do I ensure reproducibility of integration analysis?

Use version control for code and data, document all analysis steps and parameter settings, use workflow management systems such as nf-core, and deposit data and code in public repositories such as NCBI [<a href="#ref-11">11</a>][<a href="#ref-8">8</a>][<a href="#ref-10">10</a>]. Containerization ensures that software dependencies are captured and reproducible [<a href="#ref-8">8</a>].

### What is the role of reference atlases in integration?

Reference atlases provide a shared framework for integration [<a href="#ref-4">4</a>]. A well-annotated single-cell atlas can serve as a reference for labeling spatial data through label transfer [<a href="#ref-4">4</a>]. The quality of the reference is critical: incomplete or inaccurate references will propagate errors to the spatial data [<a href="#ref-4">4</a>]. Public atlases are available for many tissues and organisms [<a href="#ref-2">2</a>].

## Related Bioinformatics Guides

- [Benchmarking Atlas-Level Data Integration in Single-Cell Genomics: Methods and Best Practices](/knowledge/bioinformatics/benchmarking-atlas-level-data-integration-in-single-cell-genomics-methods-and-best-practices)
- [Spatial Transcriptomics vs. Single-Cell RNA Sequencing: Which Approach Fits Your Research?](/knowledge/bioinformatics/spatial-transcriptomics-vs-single-cell-rna-sequencing-which-approach-fits-your-research)
- [Spatial Transcriptomics Integration: Methods for Combining Data Across Platforms](/knowledge/bioinformatics/spatial-transcriptomics-integration-methods-for-combining-data-across-platforms)
- [Single-Cell Sequencing Methods: A Comparative Overview](/knowledge/bioinformatics/single-cell-sequencing-methods-a-comparative-overview)
- [Spatial Transcriptomics Differential Expression: Methods and Best Practices](/knowledge/bioinformatics/spatial-transcriptomics-differential-expression-methods-and-best-practices)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)

## References and Further Reading

<a id="ref-1"></a>[<a href="#ref-1">1</a>] [Why “Where” Matters as Much as “How Much”: Single-Cell and Spatial Transcriptomics in Plants](https://doi.org/10.3390/ijms262411819). International Journal of Molecular Sciences, 2025.

<a id="ref-2"></a>[<a href="#ref-2">2</a>] [Brain Spatial Genomics Atlases.](https://doi.org/10.3390/genes17070745). 2026.

<a id="ref-3"></a>[<a href="#ref-3">3</a>] [Spatial Transcriptomics in Plants: From Cellular Maps to Mechanistic Insight.](https://doi.org/10.1016/j.xplc.2026.102070). 2026.

<a id="ref-4"></a>[<a href="#ref-4">4</a>] [Single-cell and spatial transcriptomics integration: new frontiers in tumor microenvironment and cellular communication](https://doi.org/10.3389/fimmu.2025.1649468). Frontiers in Immunology, 2025.

<a id="ref-5"></a>[<a href="#ref-5">5</a>] [Refining Tumor Microenvironment Insights via Single-Cell and Spatial Transcriptomics for Precision Medicine](https://doi.org/10.54097/frv8k228). Highlights in Science Engineering and Technology, 2025.

<a id="ref-6"></a>[<a href="#ref-6">6</a>] [The Applications of Single-Cell and Spatial Transcriptomics in Neuroscience and Brain Disorders](https://doi.org/10.1016/j.neubiorev.2026.106721). Neuroscience and Biobehavioral Reviews, 2026.

<a id="ref-7"></a>[<a href="#ref-7">7</a>] [Applications of AI to single-cell and spatial transcriptomics: current state-of-the-art and challenges](https://doi.org/10.3389/fbinf.2025.1715821). Frontiers Bioinform., 2026.

<a id="ref-8"></a>[<a href="#ref-8">8</a>] [nf-core Documentation](https://nf-co.re/docs). nf-core.

<a id="ref-9"></a>[<a href="#ref-9">9</a>] [Microfluidics-based single cell analysis: from transcriptomics to spatiotemporal multi-omics](https://doi.org/10.1016/j.trac.2022.116868). Trac Trends in Analytical Chemistry, 2023.

<a id="ref-10"></a>[<a href="#ref-10">10</a>] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.

<a id="ref-11"></a>[<a href="#ref-11">11</a>] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.

<a id="ref-12"></a>[<a href="#ref-12">12</a>] [Single-cell genomics and spatial transcriptomics in islet transplantation for diabetes treatment: advancing towards personalized therapies](https://doi.org/10.3389/fimmu.2025.1554876). Frontiers in Immunology, 2025.

<a id="ref-13"></a>[<a href="#ref-13">13</a>] [From cell culture to cellular biology in animal-parasitic nematodes: Biotechnological advances.](https://doi.org/10.1016/j.biotechadv.2026.109007). 2026.

<a id="ref-14"></a>[<a href="#ref-14">14</a>] [Perspective on the integration of radiomics and spatial omics in the analysis of the tumor microenvironment of bladder cancer and prospects for precision diagnosis and treatment.](https://doi.org/10.3389/fimmu.2026.1821743). 2026.

<a id="ref-15"></a>[<a href="#ref-15">15</a>] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.

<a id="ref-16"></a>[<a href="#ref-16">16</a>] [Bioconductor](https://bioconductor.org/). Bioconductor Project.

<a id="ref-17"></a>[<a href="#ref-17">17</a>] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.

<a id="ref-18"></a>[<a href="#ref-18">18</a>] [Single-cell Technologies in Atherosclerosis: Uncovering Cellular Heterogeneity, Mechanisms, and Therapeutic Opportunities.](https://doi.org/10.1007/s11883-026-01432-0). 2026.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.