How to Perform Cell Type Deconvolution of Spatial Transcriptomics Data Using Single-Cell RNA-Seq References: A Step-by-Step Tutorial with R

By Dr. Zubair Khalid, DVM, MS, PhD ·

How to Perform Cell Type Deconvolution of Spatial Transcriptomics Data Using Single-Cell RNA-Seq References: A Step-by-Step Tutorial with R

Key Takeaways

  • Spatial transcriptomics data, often generated by platforms like 10x Visium, averages gene expression across spots containing multiple cells, necessitating cell type deconvolution to estimate cell type proportions within each spot.
  • Reference-based deconvolution methods, such as SPOTlight, CARD, and HarmoDecon, leverage single-cell RNA-sequencing (scRNA-seq) reference datasets to define cell-type-specific expression profiles and infer their contributions to spatial spots.
  • The quality of the scRNA-seq reference is paramount; it must accurately represent the tissue type and biological context, undergo rigorous quality control (e.g., filtering low-quality cells, doublet detection), and be appropriately normalized and annotated.
  • Gene identifier consistency between the reference and spatial transcriptomics data is critical; mismatches can lead to significant errors, and using annotation packages for conversion is a common troubleshooting step.
  • Visualization of deconvolution results, including spatial feature plots and proportion heatmaps, is essential for interpreting spatial patterns and validating findings against histology and known cell-type-specific marker gene expression.
  • Common failure patterns include reference-spatial mismatch (gene identifiers, biological context, data distribution), overbalanced proportions, loss of rare cell types, and batch effects across multiple samples, requiring careful method selection and validation.

Spatial transcriptomics platforms such as 10x Visium capture gene expression across tissue sections but measure averaged signals from spots that often contain multiple cells. Cell type deconvolution addresses this limitation by estimating the proportion of each cell type within every spatial spot, using a single-cell RNA sequencing (scRNA-seq) reference dataset to guide the estimation. This tutorial provides a reproducible R workflow for researchers who have already generated spatial transcriptomics data and a matching scRNA-seq reference, and who need to run deconvolution analysis themselves. The workflow covers reference data preparation, spatial data formatting, running deconvolution with SPOTlight or an equivalent method, visualizing results, and interpreting outputs within the constraints of the chosen algorithm.

Understanding the Deconvolution Problem in Spatial Transcriptomics

Spatial transcriptomics technologies differ in their resolution. Some platforms, such as Xenium or MERSCOPE, capture expression at near single-cell resolution, while others, including Visium and Slide-seq, measure spots that contain mixtures of cells. For platforms without single-cell resolution, the observed expression profile at each spot represents a weighted average of the cell types present in that location. Deconvolution methods reverse this averaging process by estimating the fraction of each cell type contributing to each spot's expression profile.

The need for deconvolution arises because tissue architecture and disease biology depend on knowing which cell types occupy specific spatial niches. A spot that shows high expression of a marker gene may contain many cells of that type, or it may contain few cells expressing that marker at high levels. Without deconvolution, researchers cannot distinguish between these scenarios. The NCBI Data Resources provide access to the reference datasets and sequence data that underpin these analyses, including scRNA-seq datasets deposited in the Gene Expression Omnibus and the Sequence Read Archive.

Deconvolution methods generally fall into two categories. Reference-based methods use a scRNA-seq dataset to define cell-type-specific expression profiles, then solve for the mixture proportions at each spatial spot. Reference-free methods, such as CellsFromSpace, decompose the spatial data directly using independent component analysis without requiring a single-cell reference. CellsFromSpace identifies interpretable components associated with distinct cell types or activities, enables noise reduction through component selection, and supports analysis of Visium, Slide-seq, MERSCOPE, and CosMX datasets. The method also provides a graphical interface for non-bioinformaticians to annotate components based on spatial distribution and contributor genes. Reference-free approaches are useful when a matched scRNA-seq dataset is unavailable, but reference-based methods generally provide more direct cell-type annotation because the reference defines the cell types of interest.

Choosing a Deconvolution Method and Reference Strategy

The choice of deconvolution method depends on the spatial platform, the quality of the reference dataset, and the biological question. SPOTlight is a widely used R package that performs reference-based deconvolution using non-negative matrix factorization and marker gene selection. Other methods include CARD, which uses conditional autoregressive modeling to borrow cell-type composition information across spatial locations, and HarmoDecon, a semi-supervised deep learning model that addresses biases at three scales: individual spots, entire tissue samples, and discrepancies between spatial and reference datasets.

CARD combines cell-type-specific expression information from scRNA-seq with spatial correlation in cell-type composition across tissue locations. This spatial modeling improves deconvolution accuracy even when the scRNA-seq reference is mismatched with the spatial data. CARD can also impute cell-type compositions and gene expression levels at unmeasured tissue locations, enabling construction of refined spatial maps at resolution higher than the original measurement. Applications to a pancreatic cancer dataset identified multiple cell types and molecular markers with distinct spatial localization that defined progression, heterogeneity, and compartmentalization of the disease.

HarmoDecon addresses biases that existing tools have not fully resolved. These biases produce overbalanced cell-type proportions for individual spots, mismatched cell-type fractions at the sample level, and data distribution shifts across platforms. HarmoDecon leverages pseudo-spots derived from scRNA-seq data and uses Gaussian Mixture Graph Convolutional Networks to mitigate these issues. In simulations on multi-cell spots from STARmap and osmFISH, HarmoDecon outperformed 11 state-of-the-art methods. Applied to legacy spatial platforms and 10x Visium datasets, HarmoDecon achieved the highest accuracy in spatial domain clustering and maintained strong correlations between cancer marker genes and cancer cells in human breast cancer samples. The HarmoDecon scripts and detailed tutorials are available through the study's public repository.

The reference dataset should match the tissue type and biological context of the spatial data. A reference from healthy tissue may not capture disease-specific cell states, and a reference from a different species or developmental stage will introduce systematic errors. The reference should also have sufficient sequencing depth and cell numbers to represent rare cell types. For deconvolution to estimate rare populations accurately, those populations must be present in the reference with enough cells to define their expression profiles reliably.

At a Glance: Deconvolution Workflow Overview

Workflow StepKey InputsPrimary OutputCommon Pitfall
Reference preparationscRNA-seq count matrix, cell type annotationsSeurat object with normalized data and marker genesUsing unprocessed counts or inconsistent gene identifiers
Spatial data formattingSpatial count matrix, spot coordinatesSpatialExperiment or Seurat objectMismatched gene symbols between reference and spatial data
Deconvolution runReference markers, spatial counts, method parametersCell type proportion matrix for each spotIgnoring method-specific assumptions about data distribution
Visualization and QCProportion matrix, spatial coordinatesSpatial feature plots, proportion heatmapsInterpreting proportions without checking total per spot
ValidationIndependent cell type markers, histology imagesConcordance metrics, spatial domain annotationsOverinterpreting results without orthogonal validation

Preparing the Single-Cell RNA-Seq Reference

The reference dataset must be processed with the same rigor as any scRNA-seq analysis. Quality control, normalization, and cell type annotation are prerequisites for meaningful deconvolution. The Bioconductor project provides the core R infrastructure for single-cell and spatial analysis, including packages such as SingleCellExperiment, scater, and scran that support these preprocessing steps.

Quality Control of the Reference

Filter low-quality cells before building the reference. Standard quality control metrics include the number of genes detected per cell, the total number of unique molecular identifiers (UMIs), and the percentage of mitochondrial reads. Cells with very low gene counts may be empty droplets or dying cells. Cells with very high mitochondrial percentages often represent stressed or lysed cells. The specific thresholds depend on the tissue type and the sequencing platform, so examine the distributions of these metrics before setting cutoffs.

Doublet detection is also important. Doublets are droplets that captured two cells simultaneously, and they produce expression profiles that are mixtures of two cell types. If doublets remain in the reference, they will create artificial cell type states that confuse deconvolution. Several R packages, including DoubletFinder and scDblFinder, provide doublet detection functionality.

Normalization and Feature Selection

Normalize the reference counts to account for differences in sequencing depth across cells. Log-normalization is the most common approach, but some deconvolution methods expect raw counts or specific normalization schemes. Check the method documentation before choosing a normalization strategy. The Galaxy Training Network offers accessible tutorials on single-cell quality control and normalization that translate to R workflows, and the Galaxy single-cell and spatial omics community has developed more than 120 training resources covering these analyses.

Select highly variable genes for downstream analysis. These genes capture the biological variation in the dataset and reduce the computational burden of deconvolution. The selection of marker genes for each cell type is a separate step that requires biological knowledge. Marker genes should be specific to the cell type of interest and stably expressed across the cells within that type.

Cell Type Annotation

Cell type annotation can be performed manually using known marker genes, automatically using reference-based annotation tools, or through a combination of both. Manual annotation requires examining the expression of canonical markers across clusters and assigning cell type labels based on established biology. Automatic annotation tools compare each cluster to a reference dataset and transfer labels. The quality of the annotation directly determines the quality of the deconvolution output. If the reference contains incorrect or overly broad cell type labels, the deconvolution will assign proportions to those incorrect categories.

For single-nucleus RNA-seq references, note that nuclear RNA captures a different portion of the transcriptome than whole-cell RNA. Some cell types, particularly neurons, are better represented in single-nucleus data because they are difficult to dissociate as whole cells. However, the gene expression profiles differ between nuclear and whole-cell preparations, and this difference can introduce bias when the spatial data comes from a whole-cell protocol. Document the reference preparation method and consider its implications for the deconvolution results.

Formatting Spatial Transcriptomics Data in R

Spatial transcriptomics data arrives in platform-specific formats. 10x Visium data includes a count matrix, spot coordinates, and tissue images. Slide-seq data uses bead arrays with spatial coordinates. The Bioconductor project provides the SpatialExperiment class, which stores spatial coordinates alongside expression data and supports interoperability with the broader Bioconductor ecosystem.

Creating a SpatialExperiment Object

The SpatialExperiment class extends SingleCellExperiment with spatial coordinate storage. To create a SpatialExperiment object, you need the count matrix, the spatial coordinates, and optionally the tissue image. The count matrix should have genes as rows and spots as columns, matching the orientation of the reference data. The spatial coordinates should be stored as a matrix with columns for x and y positions.

Gene identifiers must match between the reference and spatial data. If the reference uses gene symbols and the spatial data uses Ensembl IDs, convert one to the other before proceeding. The conversion can be performed using annotation packages available through Bioconductor. Mismatched gene identifiers are a common source of errors in deconvolution, and the failure often appears as poor correlation between expected and estimated proportions.

Handling Platform-Specific Differences

Different spatial platforms produce data with different characteristics. Visium spots are approximately 55 micrometers in diameter and capture expression from an estimated 1 to 10 cells per spot, depending on tissue density. Slide-seq beads are smaller and capture fewer cells per spot. The expected cell number per spot influences the interpretation of deconvolution results. A spot with an estimated total cell count of 5 will have proportion estimates that are inherently noisier than a spot with 20 cells.

The SpaceSequest pipeline provides a unified approach to processing data from five major spatial transcriptomics technologies, including Visium, Visium HD, Xenium, GeoMx, and CosMx. The pipeline performs standardized quality control, platform-specific analyses, automated cell type annotation and deconvolution, and generates publication-ready figures. SpaceSequest also integrates with cellxgene VIP and Quickomics for interactive visualization. For researchers working across multiple platforms, such a unified pipeline reduces the burden of learning platform-specific analysis procedures.

Running Deconvolution with SPOTlight

SPOTlight is an R package available through Bioconductor that performs reference-based deconvolution of spatial transcriptomics data. The method uses non-negative matrix factorization to decompose the spatial expression matrix into cell type proportion estimates. SPOTlight requires a Seurat object for the reference data and a Seurat or SpatialExperiment object for the spatial data.

Installing and Loading Required Packages

Install SPOTlight and its dependencies through Bioconductor. The installation command uses the BiocManager package. After installation, load the required libraries in your R session. The Bioconductor project provides installation instructions and package documentation for SPOTlight and related tools.

Preparing the Reference for SPOTlight

SPOTlight expects a Seurat object with normalized data and a column in the metadata that contains cell type annotations. The reference should be processed with the standard Seurat workflow: normalization, finding variable features, scaling, PCA, clustering, and annotation. The cell type column should contain labels that are consistent across all cells of the same type.

Select marker genes for each cell type before running SPOTlight. The package includes functions to identify markers using the Seurat FindAllMarkers function. The number of marker genes per cell type affects the deconvolution results. Too few markers may not capture the full expression profile of a cell type, while too many markers may include genes that are not specific. A common approach is to select the top 25 to 100 markers per cell type, but the optimal number depends on the cell types and the reference quality.

Running the Deconvolution

The core SPOTlight function takes the reference Seurat object, the spatial Seurat object, and the marker genes as inputs. The function estimates the proportion of each cell type at each spatial spot. The output includes a matrix of proportions, with spots as rows and cell types as columns. Each row should sum to approximately 1, representing the total cell composition at that spot.

The runtime depends on the number of spots, the number of cell types, and the number of marker genes. Large datasets with hundreds of thousands of spots may require substantial computational resources. The nf-core documentation describes community standards for reproducible bioinformatics pipelines, and several nf-core pipelines include spatial transcriptomics and deconvolution modules that can be run on high-performance computing clusters.

Alternative Methods and Their Parameters

CARD is available as an R package and requires a spatial count matrix, spatial coordinates, and a scRNA-seq reference. The method constructs a spatial correlation model and estimates cell type proportions using a conditional autoregressive framework. CARD can perform deconvolution without a scRNA-seq reference by using a built-in reference, but the accuracy is generally higher with a matched reference.

HarmoDecon requires Python and uses deep learning. The method constructs pseudo-spots from the scRNA-seq reference and trains a Gaussian Mixture Graph Convolutional Network to estimate cell type proportions. HarmoDecon addresses sample-level biases that other methods ignore, making it particularly useful for multi-sample spatial datasets where batch effects between samples are a concern.

The choice between methods should consider the biological question and the data characteristics. If the spatial data comes from a single sample and the reference is well matched, SPOTlight or CARD may be sufficient. If the study includes multiple samples with potential batch effects, HarmoDecon's sample-level bias correction may improve accuracy. If no reference is available, CellsFromSpace provides a reference-free alternative.

Visualizing Deconvolution Results

Visualization is essential for interpreting deconvolution results and communicating findings. The primary visualization is a spatial feature plot that shows the proportion of a cell type at each spot, overlaid on the tissue image. These plots reveal spatial patterns such as immune cell infiltration at tumor margins or neuronal layers in the cortex.

Spatial Feature Plots

Create spatial feature plots for each cell type of interest. The plots should use a color scale that highlights differences in proportion across the tissue. Spots with high proportions of a cell type will appear in one color, while spots with low proportions appear in another. The spatial distribution of these colors reveals whether the cell type is localized to specific tissue regions or distributed throughout.

Compare the spatial distribution of cell types to known tissue anatomy. If the tissue is a lymph node, B cells should localize to follicles and T cells to the paracortex. If the observed distribution contradicts known biology, investigate whether the reference annotation or the deconvolution parameters introduced errors.

Proportion Heatmaps and Bar Plots

Heatmaps showing the proportion of each cell type across all spots provide an overview of the tissue composition. Cluster the spots based on their proportion profiles to identify spatial domains with similar cell type compositions. These domains often correspond to histological structures.

Bar plots showing the average proportion of each cell type across the entire tissue provide a summary of tissue composition. Compare these averages to the expected composition based on histology or flow cytometry data. Large discrepancies may indicate problems with the reference or the deconvolution.

Spatial Domain Identification

The SpaNiche framework extends deconvolution results to spatial niche analysis. SpaNiche uses graph-regularized joint non-negative matrix factorization to integrate cell abundance and ligand-receptor expression, identifying spatial colocalization patterns among cell types and associated ligand-receptor interactions. The method uses consensus clustering to define ecotypes, which are recurring spatial microenvironments across multiple samples. Applications to colorectal cancer, prostate cancer, and cerebral cortex in early Alzheimer's disease demonstrated the utility of this approach for dissecting complex colocalization patterns within tissue microenvironments.

For hierarchical tissue organization, the HRCHY-CytoCommunity framework identifies multi-level tissue structures directly from cell-type annotated spatial maps. The graph neural network approach integrates differentiable graph pooling, adaptive edge pruning, and consistency and balance regularization to infer robust structures across multiple scales. The framework supports cross-sample hierarchy alignment and has been applied to a breast cancer cohort for hierarchical prognostic stratification of patients.

Quality Control and Validation of Deconvolution Outputs

Deconvolution results require validation before biological interpretation. The estimates are computational predictions, not direct measurements, and they carry uncertainty that varies across spots and cell types.

Checking Proportion Sums and Negative Values

Each spot's proportions should sum to approximately 1. Some methods produce small negative values due to the optimization procedure. If negative values are large or frequent, the method may not be appropriate for the data. Examine the distribution of proportion sums across all spots. Spots with sums far from 1 may represent poor-quality spots or regions with cell types not present in the reference.

Comparing to Histology and Known Markers

Compare the deconvolution results to the histology image. If the tissue has clearly defined anatomical regions, the deconvolution should assign cell types consistent with those regions. For example, a tumor section should show cancer cell proportions highest in the tumor core and immune cell proportions highest at the invasive margin.

Validate the deconvolution using independent marker genes. For each cell type, examine the spatial expression of canonical markers and compare to the estimated proportions. If the estimated proportion of T cells is high in a region but T cell marker expression is low, the deconvolution may be incorrect. This comparison is qualitative but provides a useful sanity check.

Cross-Platform Validation

If possible, validate the deconvolution results using an orthogonal method. Flow cytometry or immunohistochemistry on adjacent tissue sections can provide cell type proportions for comparison. Single-cell resolution spatial platforms, such as Xenium or MERSCOPE, can be run on adjacent sections to directly measure cell type locations. The CellsFromSpace method demonstrated the ability to identify spatially distributed cells and rare diffuse cells across Visium, Slide-seq, MERSCOPE, and CosMX datasets, providing a reference-free validation approach.

Common Failure Patterns and Troubleshooting

Deconvolution analyses frequently encounter problems that produce misleading results. Recognizing these failure patterns helps researchers diagnose issues and adjust their approach.

Reference-Spatial Mismatch

The most common failure is a mismatch between the reference and spatial data. This mismatch can occur at the level of gene identifiers, biological context, or data distribution. Gene identifier mismatches produce obvious errors, such as zero expression for many genes in the spatial data. Biological context mismatches are subtler. A reference from peripheral blood will not accurately deconvolve a tissue sample because the cell types and their expression states differ. Data distribution mismatches occur when the reference and spatial data were generated with different protocols or sequencing depths.

The HarmoDecon study specifically addressed data distribution shifts across platforms. The method uses pseudo-spots derived from scRNA-seq data to bridge the gap between reference and spatial datasets. If your data shows signs of platform-specific bias, methods that explicitly model these shifts may outperform methods that assume identical distributions.

Overbalanced Proportions

Some deconvolution methods produce overbalanced cell type proportions, where one cell type dominates the estimates across all spots. This pattern often indicates that the dominant cell type has high expression of many genes that are also present in other cell types. The reference may lack sufficient marker specificity, or the method may be assigning shared expression to the most abundant cell type.

Rare Cell Type Loss

Rare cell types are difficult to deconvolve because their contribution to each spot's expression is small. If a cell type constitutes less than 1 percent of the tissue, its signal may be below the detection threshold of the deconvolution method. The reference must contain enough cells of the rare type to define its expression profile, and the marker genes must be specific enough to distinguish it from more abundant types.

Batch Effects Across Samples

Multi-sample studies introduce batch effects that confound deconvolution. Samples processed in different batches may have systematic differences in gene expression that are unrelated to biology. These batch effects can cause the deconvolution to assign different cell type proportions to identical tissues processed in different batches. Methods that explicitly model sample-level effects, such as HarmoDecon, may be necessary for multi-sample studies.

Reproducibility and Documentation

Reproducibility is a core requirement for computational analyses. The nf-core documentation emphasizes community standards for pipeline usage, configuration, and reproducibility. For deconvolution analyses, reproducibility requires documenting the software versions, parameter settings, and reference dataset versions.

Recording Software Versions

Record the versions of R, Bioconductor, and all packages used in the analysis. The sessionInfo function in R provides a complete record of the R environment. Save this output to a file and include it with the analysis code. Package updates can change deconvolution results, so the version information is essential for reproducing the analysis.

Saving Intermediate Objects

Save the processed reference, the formatted spatial data, and the deconvolution results as R objects. These objects allow the analysis to be resumed from any point without rerunning the entire workflow. The Bioconductor project provides the RDS format for saving single R objects and the HDF5 format for large datasets.

Using Workflow Management Tools

Workflow management tools such as Snakemake or Nextflow provide structured approaches to running reproducible analyses. The nf-core documentation describes how to configure and run community pipelines, and several nf-core pipelines include spatial transcriptomics modules. The Galaxy platform offers a web-based alternative that tracks analysis history and supports reproducible sharing. The Galaxy single-cell and spatial omics community has developed more than 175 tools and 120 training resources, with over 300,000 jobs run at the time of writing.

Limitations of Deconvolution and Interpretation Boundaries

Deconvolution provides estimates, not measurements. The accuracy of these estimates depends on the quality of the reference, the appropriateness of the method, and the characteristics of the spatial data. Researchers must interpret deconvolution results within these boundaries.

Resolution Limits

Deconvolution cannot resolve cell types below the detection limit of the spatial platform. A Visium spot captures expression from multiple cells, and the deconvolution estimates the average composition of those cells. The method cannot determine the spatial arrangement of cells within a spot. Two spots with the same cell type proportions may have completely different cellular arrangements.

Reference Dependence

The deconvolution results are only as good as the reference. If the reference lacks a cell type present in the spatial data, the deconvolution will assign that cell type's expression to the most similar reference cell type. This misassignment produces biased proportions for all cell types. The reference must be comprehensive and accurately annotated.

Method-Specific Assumptions

Each deconvolution method makes assumptions about the data. Some methods assume that the reference and spatial data have similar gene expression distributions. Others assume that cell type proportions vary smoothly across the tissue. If the data violates these assumptions, the results will be biased. The benchmarking study of cellular crosstalk tools demonstrated substantial variability in tool performance across spatial resolutions, tissue contexts, and platforms, highlighting the importance of method selection.

The Nature Reviews Genetics review of cell-type deconvolution methods for spatial transcriptomics provides a comprehensive overview of the field and its limitations. The review covers the range of available methods, their assumptions, and their performance characteristics.

Professional Escalation Criteria

Some deconvolution problems require consultation with a bioinformatics specialist or the method developers. Escalate the analysis when the results are inconsistent with known biology, when the computational requirements exceed available resources, or when the analysis will inform clinical decisions.

Inconsistent Results

If the deconvolution produces proportions that contradict established biology across multiple validation approaches, the problem may require specialized expertise. A bioinformatics consultant can review the reference preparation, the method selection, and the parameter settings to identify the source of the error.

Computational Resource Limits

Large spatial datasets with millions of spots may exceed the memory or runtime limits of standard R workflows. The nf-core documentation describes how to configure pipelines for high-performance computing environments. A systems administrator or bioinformatics specialist can help optimize the computational infrastructure.

Clinical or Regulatory Decisions

Deconvolution results that inform clinical decisions or regulatory submissions require rigorous validation and documentation. The analysis must be reproducible, the methods must be appropriate for the data, and the limitations must be clearly stated. Consult with a biostatistician or regulatory specialist before using deconvolution results in these contexts.

Building a Deconvolution Decision Framework for Multi-Sample and Multi-Platform Studies

Deconvolution results become substantially harder to interpret when your study includes multiple tissue samples, multiple spatial platforms, or both. A single-sample analysis with a well-matched reference can be validated by visual inspection and marker gene comparison. Multi-sample studies introduce batch effects, platform-specific biases, and sample-level composition differences that require a structured decision framework before you run any deconvolution method. This section provides a practical framework for deciding which deconvolution approach fits your study design, how to record the decisions you make, and how to troubleshoot the specific failure patterns that emerge in multi-sample and multi-platform contexts.

Step 1: Classify Your Study Design Before Choosing a Method

The first decision is not which algorithm to run but how your study is structured. Three design categories cover most research scenarios. The first category is a single sample with a matched reference. This is the simplest case where SPOTlight, CARD, or a similar reference-based method will generally perform adequately. The second category is multiple samples from the same platform with a shared reference. This design requires attention to sample-level batch effects because the deconvolution may assign different cell type proportions to identical tissues processed in different batches. The third category is multiple samples across different spatial platforms. This design introduces platform-specific biases in addition to batch effects, and the choice of method becomes more consequential.

The HarmoDecon study explicitly addresses biases at three scales that matter for multi-sample designs: individual spots, entire tissue samples, and discrepancies between spatial and reference datasets. The sample-level bias produces mismatched cell-type fractions across samples even when the underlying tissue composition is identical. If your study includes multiple samples and you plan to compare cell type proportions between groups, you need a method that explicitly models or corrects for sample-level effects. The benchmarking study of cellular crosstalk tools demonstrated substantial variability in tool performance across spatial resolutions, tissue contexts, and platforms, reinforcing the need to match method choice to study design instead of defaulting to a single tool.

Step 2: Create a Deconvolution Decision Record

Before running any analysis, create a decision record that documents the rationale for each choice. This record serves two purposes. It forces you to articulate why a particular method and reference combination is appropriate for your data, and it provides the documentation needed for reproducibility and for troubleshooting when results are unexpected. The nf-core documentation emphasizes community standards for pipeline usage and configuration, and a decision record is a practical extension of those standards to the analysis planning phase.

The decision record should include the following fields. First, the spatial platform and the expected number of cells per spot. Second, the reference dataset source, including whether it is matched to the spatial samples or comes from a public repository. Third, the expected cell type composition based on histology or prior knowledge. Fourth, the deconvolution method selected and the specific reason for that selection. Fifth, the marker gene selection strategy and the number of markers per cell type. Sixth, the normalization approach for both reference and spatial data. Seventh, the validation approach planned before running the analysis.

Store this record as a plain text file or a structured document alongside your analysis code. The The Carpentries Lessons provide foundational training on organizing computational projects, including the use of version control and structured project directories. A decision record is most useful when it is versioned with the code and data it describes.

Step 3: Assess Reference-Spatial Compatibility Before Running Deconvolution

A common mistake is to run deconvolution immediately after basic quality control without formally assessing whether the reference and spatial data are compatible. This assessment takes less time than a full deconvolution run and can prevent hours of troubleshooting later.

Start by checking gene identifier compatibility. The reference and spatial data must use the same gene identifier system. If the reference uses gene symbols and the spatial data uses Ensembl IDs, convert one to the other before proceeding. The Bioconductor project provides annotation packages for identifier conversion. After conversion, verify that a high percentage of genes in the spatial data are present in the reference. If more than 10 percent of spatial genes are missing from the reference, investigate whether the conversion failed or whether the datasets are fundamentally incompatible.

Next, compare the distribution of expression values between reference and spatial data. The HarmoDecon study identified data distribution shifts across platforms as a major source of bias. If the reference was generated with a droplet-based scRNA-seq platform and the spatial data comes from a slide-based platform, the gene detection rates and dynamic ranges will differ. Plot the mean expression of shared genes in the reference against the mean expression in the spatial data. A strong correlation with a consistent offset indicates platform differences that can be modeled. A weak correlation suggests biological differences that may make the reference unsuitable.

Step 4: Run a Pilot Deconvolution on a Single Representative Sample

For multi-sample studies, do not run deconvolution on all samples simultaneously. Select one representative sample that has good quality metrics and clear histological structure. Run the deconvolution on this sample first and evaluate the results before scaling to the full dataset. This pilot step is analogous to testing a new protocol on a single biological replicate before processing the entire cohort.

The pilot run should use the exact parameters and reference you plan to apply to the full dataset. Evaluate the following criteria. First, do the proportion estimates sum to approximately 1 for each spot? Second, do the spatial distributions of major cell types match the histology image? Third, are the average cell type proportions consistent with published estimates for the tissue type? Fourth, are rare cell types detected at all, or are they absent from the estimates?

If the pilot run fails any of these criteria, diagnose the cause before proceeding. The CellsFromSpace method provides a reference-free alternative that can serve as an independent check on reference-based results. Running CellsFromSpace on the same pilot sample and comparing the major components to your reference-based proportions can reveal whether the reference is introducing bias or whether the spatial data itself lacks the expected cell types.

Step 5: Apply the Chosen Method to the Full Dataset and Record Per-Sample Metrics

After the pilot run passes validation, apply the method to the full dataset. For each sample, record the following metrics in a structured table. First, the number of spots analyzed. Second, the proportion of spots that passed quality control. Third, the mean and median proportion sum across spots. Fourth, the proportion of spots with negative values. Fifth, the estimated cell type composition averaged across all spots. Sixth, the runtime and peak memory usage.

These per-sample metrics serve multiple purposes. They allow you to compare sample quality across the cohort. They provide a record of computational resource requirements that will be useful for planning future analyses. They also enable you to identify outlier samples that may have failed during tissue processing or sequencing. The Galaxy Training Network provides tutorials on spatial data analysis that include guidance on quality metrics and their interpretation, and the Galaxy single-cell and spatial omics community has developed more than 120 training resources covering these analyses.

Step 6: Compare Sample-Level Composition and Investigate Outliers

After deconvolution is complete for all samples, compare the average cell type proportions across samples. Create a bar plot or heatmap showing the composition of each sample. Samples from the same tissue type and condition should have broadly similar compositions. Large differences between replicate samples warrant investigation.

The HarmoDecon study specifically addresses the problem of mismatched cell-type fractions at the sample level. If your samples show large composition differences that do not correlate with the experimental condition, the deconvolution method may be introducing sample-level bias. Consider rerunning the analysis with a method that explicitly models sample-level effects. HarmoDecon uses pseudo-spots derived from scRNA-seq data and Gaussian Mixture Graph Convolutional Networks to address these biases, and the study demonstrated improved accuracy on legacy spatial platforms and 10x Visium datasets.

For multi-platform studies, the SpaceSequest pipeline provides a unified approach to processing data from Visium, Visium HD, Xenium, GeoMx, and CosMx. The pipeline performs standardized quality control, platform-specific analyses, automated cell type annotation and deconvolution, and generates publication-ready figures. Using a unified pipeline reduces the burden of learning platform-specific procedures and ensures that the deconvolution step is applied consistently across platforms.

Step 7: Validate with Orthogonal Methods and Document Limitations

The final step in the decision framework is validation and documentation. For at least one sample per condition, validate the deconvolution results using an orthogonal method. Flow cytometry or immunohistochemistry on adjacent tissue sections can provide cell type proportions for comparison. Single-cell resolution spatial platforms, such as Xenium or MERSCOPE, can be run on adjacent sections to directly measure cell type locations. The CellsFromSpace method demonstrated the ability to identify spatially distributed cells and rare diffuse cells across Visium, Slide-seq, MERSCOPE, and CosMX datasets, providing a reference-free validation approach.

Document the limitations of your deconvolution analysis in the methods section of any manuscript or report. State the reference dataset used, the deconvolution method and version, the marker gene selection strategy, and the validation approach. Acknowledge that deconvolution provides estimates instead of direct measurements and that the accuracy depends on reference quality and method assumptions. The Nature Reviews Genetics review of cell-type deconvolution methods for spatial transcriptomics provides a comprehensive overview of method limitations that can inform this documentation.

Common Failure Patterns in Multi-Sample Deconvolution

Three failure patterns appear frequently in multi-sample deconvolution analyses. The first is sample-specific proportion inflation, where one cell type dominates the estimates in a single sample but not in others. This pattern often indicates a batch effect in the spatial data instead of a true biological difference. Check whether the affected sample was processed in a different batch or with a different reagent lot.

The second failure pattern is reference dominance, where the same cell type dominates all samples regardless of tissue composition. This pattern suggests that the reference lacks specificity for other cell types or that the dominant cell type has high expression of shared genes. Re-examine the marker gene selection and consider whether the reference includes all cell types present in the spatial data.

The third failure pattern is rare cell type loss in specific samples. If a rare cell type is detected in some samples but completely absent in others, the loss may be due to sampling or to differences in tissue quality. Compare the number of spots and the sequencing depth between samples where the rare type is detected and samples where it is absent. The HRCHY-CytoCommunity framework provides a complementary approach for identifying hierarchical tissue structures from cell-type annotated spatial maps, which can help determine whether rare cell types are organized into specific spatial niches that may be absent from some tissue sections.

When to Escalate to a Bioinformatics Specialist

Escalate the analysis to a bioinformatics specialist when the decision framework identifies problems that you cannot resolve with the available tools. Specific escalation criteria include the following. First, persistent sample-level composition differences that do not correlate with experimental conditions and persist across multiple deconvolution methods. Second, poor correlation between deconvolution estimates and orthogonal validation data across multiple samples. Third, computational resource requirements that exceed your available infrastructure for the full dataset. Fourth, the need to integrate data from more than two spatial platforms where platform-specific biases are likely to confound comparisons.

The EMBL-EBI Training provides learning pathways for bioinformatics analysis that can help researchers build the skills needed to address these challenges independently. The nf-core documentation describes how to configure pipelines for high-performance computing environments, which may be necessary for large multi-sample datasets. A bioinformatics specialist can also help implement more advanced methods such as HarmoDecon, which requires Python and deep learning infrastructure that may not be available in a standard R workflow.

Frequently Asked Questions

What is the difference between reference-based and reference-free deconvolution?

Reference-based deconvolution uses a scRNA-seq dataset to define cell-type-specific expression profiles and then estimates the proportion of each cell type at each spatial spot. Reference-free methods, such as CellsFromSpace, decompose the spatial data directly without requiring a single-cell reference. CellsFromSpace uses independent component analysis to identify interpretable components associated with distinct cell types or activities, and it supports analysis of Visium, Slide-seq, MERSCOPE, and CosMX datasets. Reference-based methods generally provide more direct cell type annotation, but reference-free methods are useful when a matched scRNA-seq reference is unavailable.

How many cells per cell type are needed in the reference for reliable deconvolution?

The required number of cells per cell type depends on the cell type's abundance in the tissue and the specificity of its markers. Rare cell types need more cells in the reference to define their expression profiles reliably. A general guideline is to include at least 50 to 100 cells per cell type, but the optimal number depends on the method and the data. Examine the stability of the deconvolution results by subsampling the reference and checking whether the proportions remain consistent.

Can I use a public scRNA-seq dataset as a reference for my spatial data?

Public scRNA-seq datasets can be used as references if they match the tissue type, species, and biological context of the spatial data. The reference should be processed with the same quality control standards as any scRNA-seq analysis. Be aware that public datasets may have been generated with different protocols or from different patient populations, which can introduce bias. The NCBI Data Resources provide access to public scRNA-seq datasets deposited in the Gene Expression Omnibus.

What should I do if my deconvolution results show one cell type dominating all spots?

Dominance by a single cell type often indicates that the reference lacks specificity for the other cell types or that the dominant cell type has high expression of shared genes. Check the marker genes for each cell type and verify that they are specific. Consider using a different deconvolution method that models spatial correlation, such as CARD, or a method that addresses sample-level biases, such as HarmoDecon.

How do I validate my deconvolution results without additional experiments?

Compare the deconvolution results to the histology image and to the spatial expression of canonical marker genes. If the estimated proportions are consistent with known tissue anatomy and marker expression, the results are likely reliable. You can also compare the average cell type proportions to published estimates for the same tissue type. The Galaxy Training Network provides tutorials on spatial data analysis and validation approaches.

What is the difference between deconvolution and cell type annotation?

Cell type annotation assigns a single cell type label to each cell or cluster based on its expression profile. Deconvolution estimates the proportion of each cell type within a spatial spot that contains multiple cells. Annotation is used for scRNA-seq data where each observation is a single cell, while deconvolution is used for spatial data where each observation is a mixture of cells.

Can deconvolution be performed on single-nucleus RNA-seq references?

Single-nucleus RNA-seq references can be used for deconvolution, but the gene expression profiles differ from whole-cell RNA-seq. Nuclear RNA captures a different portion of the transcriptome, and some genes are preferentially detected in nuclear preparations. If the spatial data comes from a whole-cell protocol, the reference should ideally also be whole-cell. Document the reference preparation method and consider its implications for the results.

How long does deconvolution take to run?

The runtime depends on the number of spots, the number of cell types, the number of marker genes, and the computational resources available. Small datasets with a few thousand spots can be processed in minutes, while large datasets with hundreds of thousands of spots may require hours or days. The nf-core documentation describes how to configure pipelines for high-performance computing environments to handle large datasets.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.