A Step-by-Step Guide to Integrating 10x Visium Spatial Transcriptomics with Single-Cell RNA-Seq Data Using Seurat

By Dr. Zubair Khalid, DVM, MS, PhD ·

A Step-by-Step Guide to Integrating 10x Visium Spatial Transcriptomics with Single-Cell RNA-Seq Data Using Seurat

Key Takeaways

  • Integration of 10x Visium spatial transcriptomics with single-cell RNA-sequencing (scRNA-seq) data via Seurat enables the spatial mapping of cell-type identities lost during tissue dissociation. This is achieved by transferring cell-type labels from a high-resolution scRNA-seq reference onto spatially barcoded spots, which capture gene expression from small cell clusters.
  • The quality of the scRNA-seq reference dataset is paramount; it must originate from the same tissue type and ideally the same biological condition, ensuring representation of all relevant cell populations and states (e.g., disease-associated states). Standard scRNA-seq quality control (UMI counts, gene detection, mitochondrial read percentage) and normalization (log-normalization or SCTransform) are critical preprocessing steps.
  • Spatial data preprocessing involves quality control of individual spots, filtering out those with low UMI counts or high mitochondrial percentages, which often indicate tissue-free areas or cellular damage. Consistent normalization with the reference dataset is essential for accurate label transfer.
  • Label transfer is facilitated by identifying "anchors" between reference and query datasets using methods like FindTransferAnchors in Seurat, followed by prediction of cell-type labels and confidence scores for each spatial spot using TransferData. Low prediction confidence across many spots signals inadequate reference coverage or significant batch effects.
  • Validation of integration results is crucial and involves assessing prediction confidence scores, comparing spatial clusters with predicted labels, and confirming the spatial expression of known cell-type marker genes. This step helps identify potential integration errors or missing cell populations in the reference.
  • Visualization of cell-type spatial organization through spatial feature plots and co-localization analysis reveals tissue architecture and cellular niches. Downstream applications include inferring cell-cell communication pathways using tools like CellChat, based on the spatially resolved gene expression profiles.

Researchers studying tissue biology face a common analytical problem: single-cell RNA sequencing (scRNA-seq) provides high-resolution molecular profiles of individual cells but loses the spatial context of where those cells reside within a tissue. Spatial transcriptomics platforms such as 10x Visium preserve tissue architecture by capturing gene expression across spatially barcoded spots, yet they typically lack the single-cell resolution needed to identify discrete cell types. Integrating these two data modalities with Seurat allows researchers to transfer cell-type labels from scRNA-seq data onto spatial spots, revealing the spatial organization of cell populations within intact tissue sections. This article provides a reproducible Seurat-based workflow covering data preprocessing, label transfer, and spatial mapping, with concrete decision criteria for quality control, troubleshooting, and interpretation.

The intended reader is a biology student, researcher, laboratory professional, or life-science practitioner who has generated or obtained 10x Visium spatial transcriptomics data and matching scRNA-seq data from the same tissue type. The workflow assumes basic familiarity with R programming and Seurat objects but does not require prior experience with spatial analysis. The practical outcome is a validated pipeline that produces spatial maps of cell-type distribution, enabling downstream analyses such as cell-cell communication inference, niche characterization, and disease mechanism studies.

Understanding the Complementary Strengths of scRNA-Seq and Spatial Transcriptomics

Single-cell RNA sequencing profiles the transcriptome of individual cells or nuclei after tissue dissociation, providing the highest resolution of cellular heterogeneity currently available. This approach has enabled researchers to identify rare cell populations, characterize developmental trajectories, and define disease-associated cellular states across diverse biological contexts. Studies in Alzheimer's disease have used scRNA-seq and single-nucleus RNA-seq (snRNA-seq) to reveal state-specific transitions in microglia and astrocytes, including shifts in transcriptional programs and pro-inflammatory polarization across disease stages. The limitation of this approach is that tissue dissociation destroys spatial architecture, so information about cellular neighborhoods and anatomical localization is lost.

Spatial transcriptomics platforms such as 10x Visium capture gene expression from spatially barcoded spots on a tissue section, preserving the anatomical context of gene activity. Each spot captures transcripts from a small cluster of cells, typically five to ten cells depending on tissue density, instead of from individual cells. This means spatial data alone cannot resolve individual cell types with confidence. Spatial transcriptomics preserves tissue architecture to map cell-type-specific gene expression within its anatomical context, which is essential for understanding how cellular organization contributes to tissue function and disease progression.

The integration of these two modalities addresses the limitations of each. scRNA-seq provides the cell-type reference, and spatial transcriptomics provides the spatial coordinates. By transferring cell-type labels from single-cell data onto spatial spots, researchers can ask questions about where specific cell populations localize, how they interact with neighboring cells, and whether spatial organization differs between healthy and diseased tissue. This integrative approach has been applied across multiple disease contexts. In dry age-related macular degeneration, researchers integrated spatial transcriptomics and single-cell RNA sequencing data to characterize immune remodeling in the retinal pigment epithelium-choroid region and identify an endothelial-macrophage signaling pathway driving disease progression. In colorectal cancer, integration of scRNA-seq and spatial transcriptomics from 36 patients revealed spatially organized epithelial-immune niches and identified signaling pathways that may serve as therapeutic targets. In rheumatoid arthritis, a three-stage analytical framework using bulk RNA-seq, single-cell transcriptome data, and spatial transcriptome data demonstrated the co-localization of CD8+ T cells and endothelial cells in the synovium.

The Seurat package provides a unified framework for this integration. Seurat supports the creation of Seurat objects from both scRNA-seq and spatial transcriptomics data, performs standard preprocessing and normalization, and implements label transfer algorithms that map cell-type annotations from a reference single-cell dataset onto query spatial data. The workflow described in this article follows the standard Seurat integration pipeline, with modifications for the specific characteristics of 10x Visium data.

At a Glance

The table below summarizes the key decisions and quality checks at each stage of the integration workflow. Use this as a quick reference before reading the detailed sections.

Workflow StagePrimary DecisionKey Quality CheckCommon Failure Mode
Reference preparationSelect scRNA-seq data from same tissue and conditionVerify all expected cell types present via marker genesMissing cell population leads to false label assignments
Spatial data loadingConfirm gene identifier format matches referenceCheck spatial image alignment with tissue sectionGene symbol mismatches cause anchor finding failure
NormalizationUse identical method for reference and queryCompare UMI distributions between datasetsInconsistent normalization reduces transfer accuracy
Label transferChoose appropriate dimensionality reductionExamine prediction score distributionLow confidence across many spots indicates poor reference coverage
ValidationCompare predicted labels with marker gene expressionVisualize spatial distribution of key markersBiologically implausible patterns signal integration errors

Preparing the Input Data

Single-Cell RNA-Seq Reference Data

The quality of the spatial mapping depends directly on the quality of the single-cell reference dataset. The reference should be generated from the same tissue type and, ideally, from the same biological condition as the spatial sample. If the spatial sample comes from diseased tissue, the reference should include disease-associated cell states. A reference that lacks relevant cell populations will produce ambiguous label assignments in the spatial data.

The reference dataset should undergo standard single-cell quality control before integration. This includes filtering cells based on the number of unique molecular identifiers (UMIs), the number of detected genes, and the percentage of mitochondrial reads. The specific thresholds depend on tissue type and dissociation protocol. For example, frozen tissues often show higher mitochondrial percentages than fresh tissues. The Galaxy Training Network provides accessible tutorials on single-cell quality control and preprocessing that can serve as a foundation for researchers new to these analyses.

After quality filtering, the reference should be normalized using log-normalization or SCTransform, depending on the downstream analysis goals. The choice of normalization method affects the label transfer results. Log-normalization is simpler and computationally faster, while SCTransform regresses out sequencing depth and technical variation more effectively. For label transfer purposes, either method can work, but consistency between the reference and query preprocessing is important.

The reference should also be checked for batch effects if it contains multiple samples or experimental conditions. Integration methods such as Harmony can be applied to the single-cell reference before label transfer. Studies have used Harmony for data preprocessing and integration in single-cell and spatial transcriptomics analyses, including the characterization of epithelial and T cell heterogeneity in colorectal cancer and the identification of high endothelial venule-related genes in bladder cancer.

10x Visium Spatial Data

The 10x Visium spatial data should be obtained from the same tissue type as the single-cell reference. The spatial data consists of two components: the gene expression matrix and the spatial coordinates. The gene expression matrix contains counts for each gene across each spatial spot, and the spatial coordinates define the position of each spot on the tissue section.

The spatial data should be loaded into Seurat using the appropriate read function for the output format. The 10x Visium output includes a filtered feature-barcode matrix, spatial image files, and tissue position files. Seurat provides functions to create a Seurat object from these files, with the spatial image information stored in the object for visualization purposes.

Quality control for spatial data differs from single-cell data. Each spot captures multiple cells, so the per-spot UMI counts are typically higher than per-cell counts. Spots with very low UMI counts may correspond to tissue-free areas or regions with poor tissue quality. Spots with very high UMI counts may represent regions with high cellular density or technical artifacts. The proportion of mitochondrial reads per spot can indicate tissue quality, with high mitochondrial proportions suggesting damaged or dying cells.

The NCBI Data Resources provide access to public single-cell and spatial transcriptomics datasets that can be used for testing workflows or for secondary analysis. Researchers who have not generated their own data can download public datasets from the Gene Expression Omnibus (GEO) or other repositories to practice the integration workflow.

Setting Up the Seurat Environment

Installing Required Packages

The integration workflow requires Seurat and several auxiliary packages. Seurat can be installed from CRAN using the standard installation command. The Seurat package includes the core functions for creating Seurat objects, normalization, dimensionality reduction, clustering, and label transfer. Additional packages may be needed for specific steps, such as Harmony for batch integration or ggplot2 for visualization.

The Bioconductor Project provides a repository of genomic analysis packages that follow reproducible analysis standards. Some packages used in single-cell and spatial analysis workflows are available through Bioconductor, and the project's documentation describes installation and usage procedures. Researchers should check whether the packages they need are available through Bioconductor or CRAN and follow the appropriate installation instructions.

The nf-core Documentation describes community standards for reproducible bioinformatics pipelines. While nf-core primarily focuses on Nextflow-based pipelines, the documentation provides useful context on pipeline configuration, parameter standardization, and reproducibility practices that can inform local analysis workflows.

Creating Seurat Objects

The first step in the workflow is to create Seurat objects for both the single-cell reference and the spatial query data. For the single-cell data, the standard CreateSeuratObject function is used with the count matrix and metadata. The metadata should include sample identifiers, condition information, and any other relevant covariates.

For the spatial data, Seurat provides a dedicated function that reads the 10x Visium output and creates a Seurat object with spatial information. This function requires the directory containing the spatial output files and the sample name. The resulting object contains the gene expression matrix, spatial coordinates, and image data needed for spatial visualization.

After creating both objects, the researcher should verify that the gene identifiers are consistent between the two datasets. Gene symbols or Ensembl IDs should match exactly. If the datasets use different gene identifier formats, conversion is necessary before integration. Mismatched gene identifiers are a common source of errors in label transfer.

Preprocessing the Single-Cell Reference

Quality Control Filtering

The single-cell reference should be filtered to remove low-quality cells and potential doublets. The standard quality metrics are the number of unique genes detected, the number of UMIs, and the percentage of mitochondrial reads. Cells with very low gene counts may be empty droplets or dying cells. Cells with very high gene counts may be doublets or multiplets. Cells with high mitochondrial percentages often indicate damaged cells where cytoplasmic RNA has been lost.

The specific thresholds depend on the tissue type and the dissociation protocol. A useful approach is to visualize the distributions of these metrics and set thresholds based on the observed distributions. For example, if the distribution of genes per cell shows a distinct population of low-gene cells, those cells should be filtered out. The The Carpentries Lessons provide foundational training on data analysis practices, including how to work with data distributions and make filtering decisions based on data exploration.

Normalization and Scaling

After filtering, the reference should be normalized. The two main options in Seurat are log-normalization and SCTransform. Log-normalization scales the counts by the total UMI count per cell, multiplies by a scale factor, and applies a log transformation. SCTransform fits a regularized negative binomial model to the counts and returns Pearson residuals that are used as normalized values.

The choice between these methods affects downstream clustering and label transfer. SCTransform generally performs better at removing technical variation and is recommended for datasets with strong batch effects or variable sequencing depth. Log-normalization is simpler and may be sufficient for well-controlled datasets. The researcher should choose one method and apply it consistently to the reference.

After normalization, the data should be scaled to regress out unwanted sources of variation. The ScaleData function in Seurat can regress out the number of UMIs, the percentage of mitochondrial reads, or other covariates. Scaling is necessary before principal component analysis (PCA) because PCA is sensitive to the scale of the features.

Dimensionality Reduction and Clustering

The reference should undergo PCA to reduce the dimensionality of the data. The number of principal components to retain can be determined using an elbow plot, which shows the variance explained by each principal component. The researcher should select the number of components where the variance explained begins to plateau.

Clustering is performed on the PCA-reduced data using a graph-based approach. The FindClusters function in Seurat uses a shared nearest neighbor graph and modularity optimization to identify clusters. The resolution parameter controls the number of clusters, with higher resolutions producing more clusters. The researcher should test multiple resolutions and select the one that produces biologically meaningful clusters.

The clusters should be annotated with cell-type labels based on marker gene expression. This annotation step is critical because the labels are what will be transferred to the spatial data. The researcher should use known marker genes for the expected cell types in the tissue and verify that the clusters express the appropriate markers. If the reference contains multiple samples or conditions, the clusters should be checked for batch effects before proceeding.

Preprocessing the Spatial Data

Quality Control for Spatial Spots

Spatial quality control focuses on identifying spots that are outside the tissue section or that have poor data quality. The 10x Visium platform includes a tissue detection algorithm that identifies spots overlapping the tissue section. Spots outside the tissue have very low UMI counts and should be excluded from analysis.

The researcher should examine the distribution of UMI counts, gene counts, and mitochondrial percentages across spots. Spots with very low UMI counts may be at the tissue edge or in regions with low cellular density. Spots with very high mitochondrial percentages may indicate damaged tissue regions. The thresholds for filtering should be based on the observed distributions and the tissue type.

Visual inspection of the spatial data is essential. Seurat provides functions to plot the spatial distribution of quality metrics, allowing the researcher to see whether low-quality spots are localized to specific regions or scattered throughout the tissue. This visualization can reveal tissue artifacts such as folds, tears, or air bubbles that affect data quality.

Normalization of Spatial Data

The spatial data should be normalized using the same method as the reference data. If the reference was normalized with log-normalization, the spatial data should also use log-normalization. If the reference used SCTransform, the spatial data should use SCTransform. Consistent normalization between reference and query is important for accurate label transfer.

The normalization should be performed on the filtered spatial object. After normalization, the data should be scaled before dimensionality reduction. The same scaling parameters should be used for the spatial data as for the reference, or the scaling should be performed independently on the spatial data.

Dimensionality Reduction for Spatial Data

The spatial data should undergo PCA to reduce dimensionality. The number of principal components may differ from the reference because the spatial data has different characteristics, such as lower resolution and higher dropout. The researcher should examine the elbow plot for the spatial data and select the appropriate number of components.

The spatial data can also be clustered independently to identify spatial domains or regions with similar gene expression. This clustering is separate from the label transfer and can provide useful context for interpreting the spatial organization. Spatial clustering can reveal tissue compartments, such as tumor regions, immune infiltrates, or stromal areas, that may not correspond exactly to cell-type boundaries.

Performing Label Transfer from Single-Cell to Spatial Data

Finding Transfer Anchors

The core of the Seurat integration workflow is the identification of transfer anchors between the reference and query datasets. The FindTransferAnchors function identifies pairs of cells or spots that are mutually similar between the reference and query. These anchors are used to transfer cell-type labels from the reference to the query.

The function requires the reference and query Seurat objects, the dimensionality reduction to use, and the normalization method. The reference should have cell-type labels stored in metadata. The function returns an anchor object that is used in the subsequent label transfer step.

The choice of dimensionality reduction for anchor finding depends on the normalization method. For log-normalized data, PCA is typically used. For SCTransform-normalized data, the SCTransform-based dimensionality reduction is used. The researcher should specify the correct reduction to match the normalization method.

Transferring Cell-Type Labels

The TransferData function uses the anchors to predict cell-type labels for each spatial spot. The function returns a matrix of prediction scores for each cell type at each spot, along with the predicted label and the prediction confidence.

The prediction scores represent the probability that a spot belongs to each cell type. The predicted label is the cell type with the highest score. The prediction confidence is the difference between the highest and second-highest scores, providing a measure of how confident the prediction is.

After label transfer, the researcher should examine the distribution of prediction scores and confidence values. Spots with low confidence may represent mixed cell populations or cell types not present in the reference. The researcher should decide whether to keep these spots in the analysis or flag them for further investigation.

Adding Predicted Labels to the Spatial Object

The predicted labels should be added to the spatial Seurat object as metadata. This allows the labels to be used for spatial visualization and downstream analysis. The prediction scores can also be stored as an assay in the spatial object, enabling the researcher to examine the spatial distribution of each cell type's prediction score.

The spatial object can then be used to visualize the cell-type distribution across the tissue section. Seurat provides functions to plot the spatial distribution of metadata, including the predicted cell-type labels. This visualization is the primary output of the integration workflow and should be examined carefully for biological plausibility.

Validating the Integration Results

Checking Prediction Confidence

The prediction confidence values provide a quantitative measure of how reliable the label transfer is for each spot. Spots with high confidence have a clear dominant cell type, while spots with low confidence may represent transitions between cell types or mixed populations.

The researcher should examine the distribution of prediction confidence across the tissue. Low-confidence spots may cluster in specific regions, suggesting that those regions contain cell types not well represented in the reference. Alternatively, low-confidence spots may be scattered, suggesting technical noise or dropout.

A useful diagnostic is to plot the prediction confidence spatially. If low-confidence spots are concentrated in specific anatomical regions, the researcher should investigate whether the reference contains the appropriate cell types for those regions. If the reference is missing a relevant cell population, the label transfer will be unreliable in those areas.

Comparing Spatial Clusters with Predicted Labels

The spatial data can be clustered independently, and the resulting spatial clusters can be compared with the predicted cell-type labels. This comparison provides a cross-validation of the integration results. If spatial clusters correspond to distinct predicted cell types, the integration is likely successful. If spatial clusters contain a mixture of predicted labels, the integration may be problematic.

The comparison can be visualized using a confusion matrix or a stacked bar plot showing the proportion of each predicted cell type within each spatial cluster. The researcher should look for spatial clusters that are dominated by a single cell type, as these represent coherent tissue regions. Spatial clusters with mixed cell types may represent transition zones or regions with high cellular heterogeneity.

Validating with Known Marker Genes

The predicted cell-type labels should be validated by examining the expression of known marker genes in the spatial data. If a spot is predicted to be a specific cell type, it should express the marker genes for that cell type. The researcher can use Seurat's feature plotting functions to visualize the spatial expression of marker genes and compare the expression patterns with the predicted labels.

This validation step is essential for identifying systematic errors in the label transfer. For example, if the reference contains a cell type that is not present in the spatial tissue, the label transfer may assign that cell type to spots with similar transcriptional profiles, producing false predictions. Marker gene validation can reveal these errors.

The Progress in the Application of Machine Learning in the Field of Single-Cell and Spatial Transcriptomics article discusses the broader context of computational methods for integrating single-cell and spatial data. While the article focuses on machine learning approaches, it provides useful background on the challenges and opportunities in this field.

Visualizing Cell-Type Spatial Organization

Spatial Feature Plots

The primary visualization for cell-type spatial organization is the spatial feature plot, which shows the expression of a gene or the prediction score of a cell type across the tissue section. Seurat provides functions to create these plots, with the spatial coordinates on the x and y axes and the feature value represented by color.

The researcher should create spatial feature plots for the predicted cell-type labels and for key marker genes. These plots reveal the spatial distribution of cell populations and can identify anatomical structures, such as tumor regions, immune infiltrates, or tissue compartments.

The spatial feature plots should be examined for biological plausibility. For example, in a tumor sample, cancer cells should be localized to the tumor region, and immune cells should be enriched at the tumor margin or in tertiary lymphoid structures. If the predicted cell-type distribution does not match the expected tissue architecture, the integration may have failed.

Co-Localization Analysis

The integration results can be used to analyze co-localization of different cell types. The researcher can calculate the frequency of co-occurrence of cell types within spatial neighborhoods or use statistical tests to identify cell types that are significantly co-localized.

Co-localization analysis can reveal cellular niches and interaction patterns. In colorectal cancer, spatial transcriptomics confirmed the in situ co-localization of CD8+ T cells with specific epithelial cells in tumor regions, revealing spatially organized epithelial-immune niches. In dry age-related macular degeneration, spatial analysis revealed that endothelial cells regulate macrophages through a specific signaling pathway, promoting the formation of a vascular-immune inflammatory niche.

The researcher should be cautious in interpreting co-localization results. Spatial spots capture multiple cells, so co-localization at the spot level does not necessarily mean direct cell-cell contact. The resolution of the spatial platform limits the conclusions that can be drawn about cellular interactions.

Cell-Cell Communication Inference

The integrated data can be used to infer cell-cell communication by combining the spatial information with ligand-receptor interaction databases. Tools such as CellChat can be applied to the integrated data to identify signaling pathways that are active between cell types in specific spatial regions.

In the colorectal cancer study, ligand-receptor signaling was assessed using CellChat, revealing that CD8+ T cells engaged in enriched communication with epithelial subsets through specific signaling pathways. The spatial transcriptomics data confirmed the in situ co-localization of these cell populations, providing evidence for spatially organized cellular interactions.

The researcher should note that cell-cell communication inference is based on gene expression and does not directly measure protein-level interactions. The results should be interpreted as hypotheses about potential interactions that require experimental validation.

Troubleshooting Common Integration Problems

Reference and Query Gene Mismatches

A common problem in label transfer is that the reference and query datasets have different gene sets. This can occur if the datasets were generated with different versions of the reference genome or if different gene annotation files were used. The FindTransferAnchors function requires that the reference and query share a sufficient number of genes.

The researcher should check the overlap between the gene sets of the reference and query. If the overlap is low, the label transfer will fail or produce unreliable results. The solution is to subset both datasets to the shared genes before integration. This step should be performed before normalization and scaling.

Batch Effects Between Reference and Query

Batch effects between the reference and query datasets can reduce the accuracy of label transfer. These effects can arise from differences in sample preparation, sequencing platform, or laboratory conditions. The FindTransferAnchors function includes options for reducing batch effects, but the researcher should also consider using batch integration methods.

Harmony is a popular method for integrating single-cell datasets across batches. The Multi-Dimensional Transcriptomics Reveals the Prominent Role of Neuroinflammation in Alzheimer's Disease review discusses the technical evolution and comparative capabilities of single-cell and spatial omics platforms, including considerations for data integration. The researcher should apply batch integration to the reference before label transfer if the reference contains multiple batches.

Low Prediction Confidence

Low prediction confidence across many spots indicates that the reference does not adequately represent the cell types in the spatial tissue. This can occur if the reference was generated from a different tissue region, a different developmental stage, or a different disease state than the spatial sample.

The researcher should examine which cell types have low prediction scores and consider whether those cell types are expected in the spatial tissue. If the reference is missing a relevant cell population, the researcher should obtain or generate additional single-cell data that includes that population. Alternatively, the researcher can use a more permissive confidence threshold, but this increases the risk of false predictions.

Spatial Artifacts

Spatial artifacts such as tissue folds, tears, or air bubbles can create regions of low data quality that produce unreliable predictions. The researcher should visually inspect the spatial data for artifacts before integration. Spots in artifact regions should be excluded from the analysis.

The Single-cell and spatial analyses reveal endothelial-macrophage inflammatory crosstalk in dry age-related macular degeneration study provides an example of how spatial data quality affects downstream analysis. The researchers characterized cellular features and regulatory pathways in the retinal pigment epithelium-choroid region, which required careful quality control to exclude artifact regions.

Recording and Reporting Integration Results

Documentation of Parameters

The integration workflow involves many parameter choices, including quality control thresholds, normalization methods, dimensionality reduction settings, and label transfer parameters. All of these choices should be documented to ensure reproducibility.

The researcher should record the version of Seurat and all other packages used in the analysis. Package versions can affect the results, so the analysis should be reproducible with the same versions. The nf-core Documentation emphasizes the importance of version control and parameter documentation for reproducible bioinformatics workflows.

The researcher should also record the specific thresholds used for quality control and the rationale for those thresholds. This documentation is important for interpreting the results and for comparing results across studies.

Output Files and Data Storage

The integration workflow produces several output files that should be stored and organized. These include the processed Seurat objects for the reference and query, the anchor object, the label transfer results, and the visualization plots.

The researcher should save the Seurat objects at each major step of the workflow. This allows the analysis to be resumed from any point without rerunning the entire workflow. The saved objects should be stored in a directory structure that reflects the analysis steps.

The The Carpentries Lessons provide training on data organization and project management that is relevant to managing bioinformatics analysis outputs. The researcher should follow these practices to ensure that the analysis is reproducible and the results are accessible.

Reporting Prediction Confidence and Limitations

The integration results should be reported with appropriate caveats about prediction confidence and limitations. The researcher should state the proportion of spots with high prediction confidence and describe the characteristics of low-confidence spots.

The limitations of the spatial platform should be acknowledged in the report. Spatial spots capture multiple cells, so the predicted cell-type labels represent the dominant cell type in each spot, not a pure cell population. The resolution of the spatial platform limits the ability to resolve fine-scale cellular interactions.

The Decoding epithelial-T cell interactions in colorectal cancer through single-cell and spatial transcriptomics study provides an example of how integration results are reported in the literature. The researchers described the cell populations identified, the spatial organization of those populations, and the limitations of their analysis.

Common Failure Patterns and How to Avoid Them

Overly Stringent Quality Control

Applying overly stringent quality control thresholds can remove legitimate cell populations from the reference, reducing the accuracy of label transfer. The researcher should examine the distributions of quality metrics and set thresholds based on the observed data instead of using arbitrary values.

A common mistake is to apply the same quality control thresholds to different tissue types. Different tissues have different cellular densities, RNA content, and dissociation sensitivities. The thresholds should be adjusted for each tissue type based on the observed distributions.

Inadequate Reference Coverage

The reference dataset must contain all cell types expected in the spatial tissue. If the reference is missing a cell type, the label transfer will assign that cell type's spots to the most transcriptionally similar cell type in the reference, producing incorrect predictions.

The researcher should verify that the reference contains the expected cell types before integration. This can be done by examining marker gene expression in the reference clusters. If a cell type is missing, the researcher should add additional single-cell data that includes that cell type.

Ignoring Batch Effects

Batch effects between the reference and query can reduce the accuracy of label transfer. The researcher should check for batch effects before integration and apply appropriate correction methods if needed.

The Integrated analysis of multiple transcriptomic approaches and machine learning integration algorithms reveals high endothelial venules as a prognostic immune-related biomarker in bladder cancer study used Harmony for data integration across multiple transcriptomic approaches. The researcher should consider similar approaches when integrating data from different sources.

Misinterpreting Prediction Scores

The prediction scores from label transfer represent probabilities, not definitive assignments. A spot with a prediction score of 0.6 for a cell type is less confidently assigned than a spot with a score of 0.9. The researcher should use the prediction scores to identify confident and uncertain predictions.

The researcher should also be aware that prediction scores are relative. A spot may have a high score for a cell type even if the absolute expression of that cell type's markers is low. The prediction scores should be interpreted in the context of the overall data quality and the marker gene expression.

Limitations of the Integration Approach

Resolution Limitations

The 10x Visium platform captures gene expression from spots that are 55 micrometers in diameter, with a center-to-center distance of 100 micrometers. Each spot captures transcripts from multiple cells, so the predicted cell-type labels represent the dominant cell type in each spot, not a pure cell population.

This resolution limitation affects the interpretation of spatial organization. Cell types that are present at low abundance may not be detected if they are mixed with more abundant cell types in the same spot. The researcher should acknowledge this limitation when interpreting the results.

Reference Dependency

The accuracy of label transfer depends on the quality and completeness of the reference dataset. If the reference does not contain the relevant cell types or if the reference has batch effects, the label transfer will be unreliable.

The researcher should validate the reference dataset before integration. This includes checking for batch effects, verifying cell-type annotations, and confirming that the reference contains the expected cell populations. The Integration of single-nuclei and spatial transcriptomics to decipher tumor phenotype predictive of relapse-free survival in Wilms tumor study demonstrates the importance of reference quality in integration analyses.

Technical Variation

Technical variation between the reference and query datasets can affect label transfer. Differences in sequencing depth, library preparation, and platform can introduce systematic biases that reduce the accuracy of label transfer.

The researcher should use normalization methods that account for technical variation and should consider using batch integration methods if the technical variation is substantial. The Deciphering immune and cellular reprogramming during the progression from inflammatory bowel disease to colorectal cancer using multi-omics single-cell and spatial transcriptomics study used multi-omics approaches to address technical variation in integration analyses.

Professional Escalation Criteria

When to Seek Expert Assistance

The integration workflow can fail in ways that are difficult to diagnose without specialized expertise. The researcher should seek expert assistance in the following situations:

If the label transfer produces results that are biologically implausible, such as cell types appearing in anatomical regions where they are not expected, the researcher should consult with a bioinformatics specialist. The specialist can help diagnose whether the problem is in the reference data, the query data, or the integration parameters.

If the prediction confidence is consistently low across the tissue, the researcher should seek advice on improving the reference dataset or adjusting the integration approach. The specialist can recommend alternative methods or additional quality control steps.

If the researcher is unsure whether the integration results are reliable, they should seek a second opinion from a colleague with experience in spatial transcriptomics analysis. The Integration of single-cell and spatial transcriptomics by SEU-TCA reveals the spatial origin of early cardiac progenitors study describes a specialized integration method that may be relevant for certain applications.

When to Consider Alternative Methods

The Seurat-based integration workflow is one of several approaches for integrating single-cell and spatial transcriptomics data. The researcher should consider alternative methods if the Seurat workflow does not produce satisfactory results.

Alternative methods include specialized integration tools such as CellTrek, which was used in the colorectal cancer study to achieve spatial mapping. Other methods use machine learning approaches to integrate single-cell and spatial data, as discussed in the Progress in the Application of Machine Learning in the Field of Single-Cell and Spatial Transcriptomics article.

The researcher should also consider whether the integration approach is appropriate for the biological question. If the goal is to identify spatial domains instead of cell-type distribution, a different analytical approach may be more suitable.

When to Validate with Experimental Methods

The integration results should be validated with experimental methods when the findings have important biological or clinical implications. Immunohistochemistry or immunofluorescence can confirm the spatial distribution of cell types predicted by the integration analysis.

The Unveiling the crucial role of CD8+ T cell and endothelial cell interaction in rheumatoid arthritis through the integrated analysis of spatial transcriptomics and bulk/single-cell RNA-seq study used spatial transcriptome analysis to validate the co-localization of cell populations identified in single-cell data. The researcher should consider similar validation approaches for their own findings.

Frequently Asked Questions

What is the minimum number of cells required in the single-cell reference for label transfer?

The minimum number of cells depends on the number of cell types expected in the tissue and the heterogeneity of those cell types. A reference with at least a few hundred cells per cell type is generally sufficient for reliable label transfer. References with fewer cells per cell type may produce unstable predictions, especially for rare cell populations. The researcher should examine the number of cells per cluster in the reference and ensure that each cluster has enough cells to support reliable label transfer.

Can I use a public single-cell dataset as the reference for my spatial data?

Yes, public single-cell datasets can be used as references if they are generated from the same tissue type and contain the relevant cell populations. The NCBI Data Resources provide access to public single-cell datasets that can be downloaded and used as references. The researcher should verify that the public dataset has adequate quality and contains the expected cell types before using it as a reference.

How do I choose between log-normalization and SCTransform for the integration?

The choice depends on the characteristics of the data. SCTransform generally performs better at removing technical variation and is recommended for datasets with strong batch effects or variable sequencing depth. Log-normalization is simpler and may be sufficient for well-controlled datasets. The researcher should use the same normalization method for both the reference and query datasets to ensure consistency in the label transfer.

What should I do if the prediction confidence is low for many spatial spots?

Low prediction confidence across many spots indicates that the reference does not adequately represent the cell types in the spatial tissue. The researcher should examine which cell types have low prediction scores and consider whether those cell types are expected in the spatial tissue. If the reference is missing a relevant cell population, the researcher should obtain or generate additional single-cell data that includes that population.

How do I know if my integration results are biologically plausible?

The integration results should be validated by comparing the predicted cell-type distribution with known tissue architecture. The researcher should examine the spatial distribution of predicted cell types and check whether the distribution matches the expected anatomy. Marker gene expression should be consistent with the predicted cell types. If the results are not biologically plausible, the integration may have failed.

Can I integrate spatial data from multiple tissue sections with a single-cell reference?

Yes, multiple spatial sections can be integrated with a single-cell reference. Each spatial section should be processed separately and then integrated with the reference. The researcher should ensure that the spatial sections are from the same tissue type and that the reference contains the relevant cell populations. The integration results for each section can be compared to identify consistent spatial patterns.

What are the main limitations of the Seurat label transfer approach?

The main limitations are the resolution of the spatial platform, the dependency on reference quality, and the effects of technical variation. Spatial spots capture multiple cells, so the predicted labels represent the dominant cell type in each spot. The accuracy of label transfer depends on the quality and completeness of the reference dataset. Technical variation between the reference and query can reduce the accuracy of label transfer.

How should I report the integration results in my publication?

The integration results should be reported with details about the reference dataset, the spatial dataset, the preprocessing steps, the integration parameters, and the validation results. The researcher should report the proportion of spots with high prediction confidence and describe the characteristics of low-confidence spots. The limitations of the spatial platform and the integration approach should be acknowledged in the discussion.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.