Why Are My Spatial Deconvolution Results Unreliable? Common Pitfalls and How to Fix Them
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Unreliable spatial deconvolution results stem from errors in reference data (e.g., incorrect cell types/states), spatial data, or their alignment, often due to platform-specific gene capture differences or low spot purity in capture-based technologies.
- A robust reference panel is paramount; it must accurately reflect the tissue's cell types and developmental states, with validated marker genes that are detectable in both reference and spatial datasets.
- Platform differences between reference and spatial data, such as variations in sequencing depth and gene capture efficiency, necessitate platform-aware normalization and careful assessment of gene overlap to prevent systematic over/underestimation of cell type proportions.
- Inadequate quality control of both reference (e.g., low gene counts, high mitochondrial content) and spatial data (e.g., spots outside tissue) is a critical pitfall, leading to results that change dramatically with minor input adjustments.
- Batch effects between reference and spatial datasets, arising from different preparation or sequencing runs, can skew cell type proportions; these require detection and correction using methods like domain adaptation or batch correction algorithms.
- Deconvolution results must be rigorously interpreted against tissue histology, as discrepancies (e.g., cell types in impossible locations) indicate fundamental issues with the reference, assumptions, or workflow, not the algorithm itself.
Spatial deconvolution results become unreliable when the reference data, the spatial data, or the alignment between them contains errors that the algorithm cannot correct. The most common causes are reference panels built from the wrong cell types or states, platform-specific differences in gene capture and sequencing depth, low spot purity in capture-based technologies, and inadequate quality control before deconvolution. This article explains how to identify these problems in your own data and what concrete steps you can take to correct them.
Spatial transcriptomics preserves gene expression information within the context of intact tissue architecture, which addresses a key limitation of single-cell RNA sequencing that loses the spatial position of cells. However, many spatial platforms capture RNA from multicellular regions instead of individual cells, so the measured expression at each capture spot represents a mixture of signals from multiple cells. Deconvolution is the computational process used to infer the underlying cellular composition at each spot. When deconvolution results do not match histological expectations, the problem usually lies in the inputs, the assumptions, or the workflow instead of in the algorithm itself.
This article is written for biology students, researchers, laboratory professionals, and life-science practitioners who are running spatial transcriptomics experiments and need to troubleshoot deconvolution output. The focus is on practical decisions you can make at each stage of the pipeline, from reference construction through quality control to final interpretation.
At a Glance
The table below summarizes the most common sources of unreliable spatial deconvolution results, the typical symptoms you will observe, and the first corrective action to take.
| Common Pitfall | Typical Symptom in Results | First Corrective Action |
|---|---|---|
| Reference panel built from wrong cell types or states | Cell types appear in regions where histology shows they should not exist | Rebuild the reference from tissue-matched single-cell or single-nucleus data with explicit marker gene validation |
| Platform differences between reference and spatial data | Systematic overestimation of one cell type across all spots | Apply platform-aware normalization and check gene overlap between reference and spatial datasets |
| Low spot purity in capture-based technologies | Deconvolution returns diffuse mixtures even in histologically homogeneous regions | Use higher-resolution platforms or apply super-resolution inference methods |
| Inadequate quality control before deconvolution | Results change dramatically when a small number of spots or cells are removed | Document QC thresholds, rerun deconvolution after filtering, and compare outputs |
| Batch effects between reference and spatial datasets | Cell type proportions correlate with sequencing batch instead of tissue structure | Use batch correction or domain adaptation methods during deconvolution |
| Overlapping or poorly defined cell type signatures | Adjacent cell types are confused with each other in the output | Refine the reference to include subtype markers and reduce signature overlap |
Understanding What Deconvolution Actually Computes
Deconvolution algorithms estimate the fraction of each cell type present at each spatial capture spot. The output is a matrix of cell type proportions per spot, and in some methods, an estimate of the cell type specific gene expression profile. The accuracy of these estimates depends on three things: the quality of the reference data, the quality of the spatial data, and how well the algorithm accounts for the differences between the two.
Most spatial transcriptomics platforms do not achieve single-cell resolution. Capture-based technologies such as Visium collect RNA from spots that contain multiple cells, and the measured expression at each spot is a mixture of signals from those cells. Image-based technologies can achieve higher resolution but still require computational methods to assign expression to individual cells. Deconvolution is therefore a necessary step for interpreting spatial data in terms of cellular composition.
The computational landscape of spatial deconvolution includes many algorithms that use different principles. Some methods rely on external single-cell RNA sequencing references, while others use marker genes or topic models. A comprehensive review of twenty deconvolution algorithms highlights that these methods differ in their modeling approaches, their data processing pipelines, and how they handle external references, noise, and sparsity in the data. Understanding the assumptions of your chosen method is essential because each algorithm makes different tradeoffs.
The key point is that deconvolution does not create information. It estimates cell type proportions based on the reference and the spatial data you provide. If either input is flawed, the output will be unreliable regardless of which algorithm you use.
Building a Reliable Reference Panel
The reference panel is the single most important input to deconvolution. It defines the cell types you expect to find and provides the gene expression signatures that the algorithm uses to estimate proportions. A poorly constructed reference will produce unreliable results even with a perfect spatial dataset.
Choosing Between Single-Cell and Single-Nucleus Data
The first decision is whether to use single-cell RNA sequencing or single-nucleus RNA sequencing data as the reference. This choice matters because the two methods capture different RNA populations. Single-cell RNA sequencing captures cytoplasmic mRNA from intact cells, while single-nucleus RNA sequencing captures nuclear RNA. For tissues where the cytoplasmic and nuclear transcriptomes differ substantially, the reference must match the biology of the spatial data.
For formalin-fixed paraffin-embedded tissue, the choice of reference is especially important. A protocol for spatial transcriptomics of mouse embryonic tissue describes optimized sectioning and handling of formalin-fixed paraffin-embedded tissue to preserve RNA integrity and tissue morphology for high-resolution spatial analysis. This method is compatible with both sequencing and image-based spatial transcriptomics platforms. If your spatial data comes from fixed tissue, your reference should ideally come from a similar preparation to minimize technical differences.
Matching Reference Cell Types to Tissue Biology
The reference must contain the cell types that are actually present in your tissue. If a cell type is missing from the reference, the algorithm cannot detect it. Instead, its signal will be distributed among the cell types that are present, producing inflated proportions for those types.
This problem is common when researchers use a reference from a different tissue or a different developmental stage. For example, a reference built from adult tissue will not contain the transient cell populations present during development. The spatial transcriptomics protocol for early tooth morphogenesis emphasizes that the developing tooth comprises diverse and highly specialized cell populations that work together to maintain proper form and function. Perturbations in these processes can result in congenital disorders. If your reference does not include these specialized populations, deconvolution will misassign their expression to other cell types.
Validating Marker Genes
Marker genes are the foundation of cell type identification in the reference. Before running deconvolution, you should validate that your marker genes are specific to the intended cell types and are expressed at detectable levels in both the reference and the spatial data.
Marker gene assisted deconvolution methods use known markers to guide the estimation of cell type proportions. The SMART method, for example, simultaneously infers cell type specific gene expression profiles and cellular composition at each spot using marker gene information. This approach can improve performance on cell subtypes when the markers are well chosen. However, if the markers are not specific, the method will propagate the error.
A practical validation step is to check that your marker genes are expressed in the expected cell types in the reference data and that they are detected in the spatial data. Genes that are not detected in the spatial data cannot contribute to deconvolution and should be removed from the analysis.
Addressing Platform Differences Between Reference and Spatial Data
Spatial transcriptomics platforms differ from single-cell RNA sequencing platforms in several important ways. These differences affect the gene expression measurements and must be accounted for during deconvolution.
Gene Capture and Sequencing Depth
Capture-based spatial platforms typically have lower sequencing depth per spot than single-cell platforms have per cell. This means that fewer genes are detected per spot, and the expression levels of detected genes are noisier. The mRNA capture efficiency of droplet-based single-cell RNA sequencing is estimated at 10 to 50 percent, and cell capture variability ranges from 30 to 75 percent. Spatial platforms have their own capture efficiencies, which are often lower than single-cell platforms.
When the reference data has higher sensitivity than the spatial data, deconvolution algorithms may overestimate the proportions of cell types with high expression levels and underestimate cell types with low expression levels. This is because the algorithm interprets the absence of lowly expressed genes in the spatial data as evidence that the corresponding cell types are absent.
Gene Overlap Between Datasets
The overlap between genes detected in the reference and genes detected in the spatial data is a critical quality metric. If the overlap is low, deconvolution will be based on a small set of genes and the results will be unstable.
You should calculate the number of shared genes between your reference and spatial datasets before running deconvolution. If the overlap is low, you may need to use a different reference or adjust your filtering criteria. Some deconvolution methods are more robust to low gene overlap than others, but all methods perform better with a larger shared gene set.
Platform-Aware Normalization
Normalization is the process of adjusting expression values to account for technical differences between samples. Standard single-cell normalization methods may not be appropriate for spatial data because the data has different characteristics, including spatial autocorrelation and varying spot sizes.
Some deconvolution methods incorporate platform-aware normalization into their workflow. The STDSN method uses a domain separation network with a shared-private encoder architecture to minimize feature differences between simulated spatial data generated from single-cell data and real spatial data. This approach allows the model trained on simulated data to generalize to real spatial data. If your chosen method does not account for platform differences, you may need to apply additional normalization steps before deconvolution.
Managing Spot Purity and Resolution Limitations
Spot purity refers to the degree to which a capture spot contains a single cell type instead of a mixture. Low spot purity is a fundamental limitation of capture-based spatial platforms and directly affects deconvolution accuracy.
Understanding the Resolution Problem
Capture-based technologies such as Visium collect RNA from spots that contain multiple cells. The number of cells per spot varies depending on the tissue and the platform, but it is typically in the range of several to dozens of cells. When a spot contains multiple cell types, the measured expression is a mixture that deconvolution must resolve.
The review of spatial transcriptomics deconvolution methods notes that at numerous capture spots, multiple signals from various cells are present, requiring deconvolution to deduce the underlying cellular composition. This is not a problem that can be fully solved by any algorithm. The information about the exact cellular composition at a mixed spot is not present in the data, so deconvolution can only provide an estimate.
Using Super-Resolution Methods
Super-resolution methods aim to recover single-cell gene expression profiles from low-resolution spatial data. The scResolve protocol describes a computational approach to recover single-cell gene expression profiles from spatial transcriptomics data using cluster computing. This method runs in a cluster environment and leverages parallel computing to accelerate data processing.
Super-resolution methods are useful when you need single-cell level information from capture-based data. However, they add computational complexity and require careful validation. The results should be compared with histological expectations to ensure that the inferred single-cell profiles are biologically plausible.
Choosing the Right Platform for Your Question
The choice of spatial platform should be driven by your biological question. If you need to resolve fine cellular structures, an image-based platform with higher resolution may be more appropriate than a capture-based platform. If you need whole-transcriptome coverage, a capture-based platform may be necessary.
The protocol for spatial transcriptomics of early tooth morphogenesis describes a workflow that is compatible with both sequencing and image-based platforms. This flexibility allows researchers to choose the platform that best matches their resolution requirements. For tissues with complex cellular architecture, higher resolution platforms reduce the need for deconvolution and produce more reliable results.
Performing Quality Control Before Deconvolution
Quality control is the process of identifying and removing low-quality data before analysis. In spatial transcriptomics, quality control applies to both the reference data and the spatial data. Inadequate quality control is a common cause of unreliable deconvolution results.
Quality Control for Reference Data
The reference data should be filtered to remove low-quality cells before building the reference panel. This includes removing cells with low gene counts, high mitochondrial content, or other indicators of poor quality. The specific thresholds depend on the tissue and the platform, but the goal is to retain cells that represent the true biological states.
Single-cell RNA sequencing quality control is a well-established practice. The Galaxy single-cell and spatial omics community has developed more than 175 tools and 120 training resources to help researchers perform and interpret their own analyses. These resources include guidance on quality control for single-cell data.
Quality Control for Spatial Data
Spatial data quality control includes removing spots with low gene counts, high mitochondrial content, or other indicators of poor tissue quality. Spots that fall outside the tissue section should also be removed.
The spatial transcriptomics protocol for bladder Ewing sarcoma outlines steps for sample collection, single-cell RNA sequencing, and spatial transcriptomics sequencing. The protocol emphasizes obtaining high-quality spatial transcriptomics data, which requires careful attention to tissue quality and sequencing quality.
Documenting Quality Control Decisions
Quality control decisions should be documented so that the analysis is reproducible. This includes recording the filtering thresholds, the number of cells or spots removed, and the rationale for each decision. Reproducibility is a core principle of bioinformatics analysis, and the Galaxy Training Network provides accessible workflow training that emphasizes reproducible analysis practices.
If deconvolution results change dramatically when quality control thresholds are adjusted, this indicates that the results are not robust. You should investigate the source of the instability before proceeding with interpretation.
Handling Batch Effects Between Reference and Spatial Data
Batch effects are systematic technical differences between datasets that are not related to biology. In spatial deconvolution, batch effects between the reference and spatial data can cause the algorithm to assign cell type proportions based on technical instead of biological signals.
Sources of Batch Effects
Batch effects can arise from many sources, including different sample preparation dates, different sequencing runs, different laboratories, and different platforms. The reference data and spatial data are almost always generated in different batches, so some degree of batch effect is expected.
The AddaGCN method uses graph convolutional networks to incorporate spatial information and adopts an adversarial discriminative domain adaptation approach to mitigate batch effects between spatial and single-cell reference data. This method demonstrates superior performance and robustness in cell type deconvolution compared to other methods across diverse technology platforms.
Detecting Batch Effects
You can detect batch effects by comparing the overall gene expression distributions between the reference and spatial data. If the distributions are substantially different, batch effects are likely present. Principal component analysis or other dimensionality reduction methods can help visualize these differences.
Correcting Batch Effects
Batch correction methods can be applied before deconvolution, or deconvolution methods that incorporate batch correction can be used. The choice depends on your workflow and the specific methods you are using.
The STDSN method addresses the differences in unique features between spatial transcriptomics and single-cell RNA sequencing data by using a domain separation network. This approach minimizes feature differences and enables the model trained on simulated data to generalize effectively to real spatial data. If your deconvolution results are affected by batch effects, consider using a method that explicitly handles domain differences.
Selecting the Right Deconvolution Algorithm
The choice of deconvolution algorithm affects the reliability of your results. Different algorithms make different assumptions and have different strengths and weaknesses. There is no single best algorithm for all situations.
Categories of Deconvolution Methods
Deconvolution methods can be categorized by their underlying computational principles. Some methods use regression-based approaches, some use topic models, some use neural networks, and some use marker gene information. The review of twenty deconvolution algorithms provides a comprehensive analysis of their methodological foundations, contrasting the underlying computational algorithms, modeling methods, and data processing pipelines.
Matching Algorithm to Data Characteristics
The characteristics of your data should guide your choice of algorithm. If your reference and spatial data have large platform differences, a method that handles domain adaptation may be appropriate. If you have well-validated marker genes, a marker gene assisted method may perform well. If you need single-cell resolution output, a super-resolution method may be necessary.
The Redeconve method deconvolutes spatial transcriptomics data at single-cell resolution, enabling interpretation with thousands of nuanced cell states. This method was benchmarked against state-of-the-art algorithms on diverse spatial transcriptomics platforms and datasets, demonstrating superiority in accuracy, resolution, robustness, and speed. If your biological question requires resolving subtle cell states, this type of method may be appropriate.
Testing Multiple Algorithms
A practical approach is to run deconvolution with multiple algorithms and compare the results. If different algorithms produce substantially different results, this indicates that the data may not support reliable deconvolution. If the results are consistent across algorithms, you can have more confidence in the output.
The Galaxy platform provides access to more than 175 tools for single-cell and spatial omics analysis, including multiple deconvolution methods. This allows researchers to test different algorithms within a reproducible workflow framework.
Interpreting Deconvolution Results Against Histology
Deconvolution results should always be interpreted in the context of tissue histology. If the results do not match histological expectations, there is likely a problem with the inputs or the workflow.
Using Histological Annotations
Histological annotations provide ground truth information about tissue structure. You can use these annotations to validate deconvolution results by checking that cell types appear in the expected anatomical regions.
For example, in a tissue section with clearly defined epithelial and stromal regions, deconvolution should assign epithelial cell types to the epithelial region and stromal cell types to the stromal region. If the results show epithelial cells in the stroma or vice versa, there is a problem.
Identifying Systematic Errors
Systematic errors in deconvolution results follow a pattern. For example, if one cell type is consistently overestimated across all spots, the reference may be missing a cell type or the marker genes may be too broad. If the errors are random, the problem may be technical noise or low spot purity.
The SMART method provides a covariate model that enables the identification of cell type specific differentially expressed genes across conditions. This can help elucidate biological changes at single-cell-type resolution and identify potential sources of error.
Reporting Uncertainty
Deconvolution results are estimates, not measurements. You should report the uncertainty associated with your results and avoid overinterpreting small differences in cell type proportions. The confidence intervals or other uncertainty measures provided by some algorithms can help with this.
Common Failure Patterns and Their Fixes
The following failure patterns are commonly observed in spatial deconvolution analysis. Each pattern has specific causes and corrective actions.
Cell Types Appear in Impossible Locations
If deconvolution assigns a cell type to a location where it cannot exist based on histology, the reference may contain incorrect cell type definitions. For example, if immune cells appear in regions with no blood vessels or immune infiltration, the reference may have mislabeled a cell type.
Fix: Validate the reference cell type annotations using known marker genes and compare with published data for the tissue. Consider using a reference from a more closely matched tissue or condition.
All Spots Show Similar Cell Type Proportions
If deconvolution returns nearly identical cell type proportions for all spots, the algorithm may not have enough information to distinguish cell types. This can happen when the reference signatures are too similar or when the spatial data has low gene detection.
Fix: Check the gene overlap between reference and spatial data. Increase sequencing depth if possible, or use a deconvolution method that is more sensitive to subtle expression differences.
Results Change Dramatically with Small Input Changes
If deconvolution results are highly sensitive to small changes in the input data, the analysis is not robust. This can happen when the reference contains highly correlated cell types or when the spatial data has high noise.
Fix: Test the robustness of your results by perturbing the input data. Remove a small number of cells or spots and rerun the analysis. If the results change substantially, the data may not support reliable deconvolution.
One Cell Type Dominates All Spots
If one cell type is assigned a high proportion at every spot, the reference may be missing a cell type that is actually present in the tissue. The algorithm assigns the missing cell type's signal to the most similar cell type in the reference.
Fix: Review the tissue biology and ensure that all expected cell types are represented in the reference. Consider adding cell types that are known to be present but were not included.
Subtypes Are Confused with Each Other
If closely related cell subtypes are confused in the deconvolution output, the reference signatures may be too similar. This is common for T cell subtypes, macrophage subtypes, and other closely related populations.
Fix: Use a marker gene assisted method that can enhance performance on cell subtypes. The SMART method provides a two-stage approach to improve performance on cell subtypes.
Records and Measurements for Deconvolution Analysis
Keeping detailed records of your deconvolution analysis is essential for reproducibility and troubleshooting. The following records should be maintained for each analysis.
Reference Construction Records
Document the source of the reference data, including the tissue, the platform, and the sample preparation method. Record the quality control thresholds used to filter the reference data and the number of cells retained. Document the cell type annotations and the marker genes used for each cell type.
Spatial Data Records
Document the spatial platform, the tissue, and the sample preparation method. Record the quality control thresholds used to filter the spatial data and the number of spots retained. Document the sequencing depth and the number of genes detected per spot.
Deconvolution Run Records
Record the deconvolution algorithm and version, the parameters used, and the reference and spatial data inputs. Document the runtime and the computational resources required. Save the output files, including the cell type proportion matrix and any uncertainty estimates.
Validation Records
Document the validation steps performed, including comparisons with histological annotations and consistency checks across algorithms. Record any discrepancies and the actions taken to address them.
The nf-core documentation provides standards for community pipeline usage and configuration that emphasize reproducible workflow practices. Following these standards can help ensure that your deconvolution analysis is reproducible.
Limitations of Spatial Deconvolution
Spatial deconvolution has fundamental limitations that cannot be overcome by any algorithm. Understanding these limitations is essential for interpreting results correctly.
Information Theoretic Limits
Deconvolution cannot recover information that is not present in the data. If a spot contains a mixture of cell types, the exact proportions cannot be determined with certainty. The algorithm provides an estimate based on the available information, but the estimate has uncertainty.
Reference Dependence
Deconvolution results depend entirely on the reference data. If the reference is incomplete or inaccurate, the results will be unreliable. There is no way to validate deconvolution results without independent information, such as histology or in situ hybridization.
Platform Specificity
Deconvolution methods are developed and validated on specific platforms. A method that performs well on one platform may not perform well on another. The review of deconvolution methods notes that each algorithm employs distinct computational principles and processing paradigms, and the performance varies across platforms.
Batch Effects
Batch effects between reference and spatial data can never be completely eliminated. Domain adaptation methods can reduce batch effects, but they cannot remove them entirely. The results should be interpreted with this limitation in mind.
Safety and Regulatory Context
Spatial transcriptomics involves working with biological samples that may pose safety risks. The following safety considerations should be addressed in your laboratory workflow.
Sample Handling
Tissue samples may contain infectious agents or hazardous chemicals. Formalin-fixed paraffin-embedded tissue requires careful handling to avoid exposure to formaldehyde and other fixatives. The protocol for spatial transcriptomics of early tooth morphogenesis describes optimized sectioning and handling of formalin-fixed paraffin-embedded tissue to preserve RNA integrity and tissue morphology.
RNA Integrity
RNA is easily degraded by RNases, which are present on skin and in the environment. Proper laboratory practices, including the use of RNase-free reagents and equipment, are essential for preserving RNA integrity. The csRNA-seq protocol notes that purified RNA is non-infectious and can be isolated from inactivated samples, including clinical or pathogenic specimens, allowing safe transport and analysis under standard laboratory conditions.
Data Management
Spatial transcriptomics generates large amounts of data that must be managed securely. Data management practices should comply with institutional policies and any applicable regulations. The NCBI provides data resources for storing and accessing biological data, and the EMBL-EBI Training program offers guidance on data management for bioinformatics analysis.
Professional Escalation Criteria
Some problems with spatial deconvolution cannot be solved by adjusting the analysis workflow. The following situations warrant escalation to a specialist or a change in experimental approach.
Persistent Disagreement with Histology
If deconvolution results consistently disagree with histological observations despite correcting reference composition, platform differences, and quality control, the spatial data may not be suitable for deconvolution. Consider using a higher-resolution platform or an image-based approach.
Low Gene Overlap
If the gene overlap between reference and spatial data is too low for reliable deconvolution, the experimental design may need to be changed. Consider generating reference data from the same tissue and platform as the spatial data.
Computational Resource Limitations
Some deconvolution methods, particularly super-resolution methods, require substantial computational resources. The scResolve protocol runs in a cluster environment to accelerate data processing. If your computational resources are insufficient, consider using a less computationally intensive method or accessing shared computing resources.
Regulatory Compliance
If your research involves clinical samples or regulated data, ensure that your analysis complies with applicable regulations. The csRNA-seq protocol notes that purified RNA can be isolated from inactivated samples, including clinical or pathogenic specimens, allowing safe transport and analysis under standard laboratory conditions. Consult with your institutional review board or regulatory affairs office if you have questions about compliance.
Building a Deconvolution Decision Framework for Your Specific Dataset
Choosing the right deconvolution approach requires a structured decision process that accounts for your specific data characteristics, biological question, and available resources. A systematic framework helps you avoid the common pitfall of applying a method because it is popular instead of because it is appropriate for your data. The framework below organizes the key decisions into a sequence of checkpoints that you can work through before committing to a final analysis plan.
Step 1: Define the Biological Resolution Requirement
The first decision is what level of biological resolution you actually need to answer your question. This determines whether you need single-cell resolution output, cell type level proportions, or subtype level distinctions.
If your question concerns broad tissue architecture, such as distinguishing epithelial from stromal regions, cell type level deconvolution is sufficient. If your question concerns subtle cell states, such as distinguishing T cell subtypes or macrophage activation states, you need a method that can resolve these finer distinctions. The Redeconve method deconvolutes spatial transcriptomics data at single-cell resolution, enabling interpretation with thousands of nuanced cell states. This level of resolution is necessary when your biological question requires identifying rare or closely related populations.
For questions about cell type proportions across conditions, a marker gene assisted method may be appropriate. The SMART method provides a covariate model that enables the identification of cell type specific differentially expressed genes across conditions, which is useful when you need to compare cell type composition between experimental groups.
Step 2: Assess Reference Data Compatibility
The second decision is whether your reference data is compatible with your spatial data. This assessment should be completed before running any deconvolution algorithm.
Start by calculating the gene overlap between your reference and spatial datasets. Low overlap means the deconvolution will be based on a small set of genes, which produces unstable results. If the overlap is below a level you consider acceptable for your analysis, you have three options: generate a new reference from the same tissue and platform, use a reference from a more closely matched source, or adjust your filtering criteria to increase overlap.
Next, compare the cell type composition of your reference with the expected biology of your tissue. If your reference contains cell types that are not present in your tissue, the algorithm may assign those cell types to spots where they cannot exist. If your reference is missing cell types that are present, their signal will be distributed among the cell types that are present, inflating those proportions.
Finally, assess whether the reference captures the relevant biological states. For formalin-fixed paraffin-embedded tissue, the reference should ideally come from a similar preparation. The protocol for spatial transcriptomics of early tooth morphogenesis describes optimized sectioning and handling of formalin-fixed paraffin-embedded tissue to preserve RNA integrity and tissue morphology for high-resolution spatial analysis. If your spatial data comes from fixed tissue, a reference generated from fresh tissue may introduce systematic differences.
Step 3: Evaluate Platform Differences
The third decision is how to handle platform differences between your reference and spatial data. These differences affect gene capture efficiency, sequencing depth, and the overall distribution of expression values.
If your reference and spatial data come from the same platform, platform differences are minimal and standard deconvolution methods should perform adequately. If they come from different platforms, you need to account for these differences.
For capture-based spatial platforms with lower sensitivity than single-cell platforms, the algorithm may overestimate cell types with high expression and underestimate cell types with low expression. This is because the absence of lowly expressed genes in the spatial data is interpreted as evidence that the corresponding cell types are absent.
Methods that explicitly handle domain differences between reference and spatial data can mitigate these issues. The STDSN method uses a domain separation network with a shared-private encoder architecture to minimize feature differences between simulated spatial data generated from single-cell data and real spatial data. The AddaGCN method uses graph convolutional networks and adversarial discriminative domain adaptation to mitigate batch effects between spatial and single-cell reference data.
Step 4: Determine Spot Purity and Resolution Needs
The fourth decision is whether your spatial platform provides sufficient spot purity for your biological question. Spot purity refers to the degree to which a capture spot contains a single cell type instead of a mixture.
Capture-based technologies such as Visium collect RNA from spots that contain multiple cells. The review of spatial transcriptomics deconvolution methods notes that at numerous capture spots, multiple signals from various cells are present, requiring deconvolution to deduce the underlying cellular composition. This is a fundamental limitation that no algorithm can fully overcome.
If your tissue has complex cellular architecture with many cell types in close proximity, low spot purity will produce diffuse deconvolution results even in histologically homogeneous regions. In this case, you have two options: use a higher-resolution platform or apply super-resolution methods.
Super-resolution methods aim to recover single-cell gene expression profiles from low-resolution spatial data. The scResolve protocol describes a computational approach to recover single-cell gene expression profiles from spatial transcriptomics data using cluster computing. This method runs in a cluster environment and leverages parallel computing to accelerate data processing. However, super-resolution methods add computational complexity and require careful validation against histological expectations.
Step 5: Select the Algorithm Based on Data Characteristics
The fifth decision is which deconvolution algorithm to use. This choice should be guided by the assessments from the previous steps.
If your reference and spatial data have large platform differences, choose a method that handles domain adaptation. If you have well-validated marker genes, choose a marker gene assisted method. If you need single-cell resolution output, choose a super-resolution method.
The review of twenty deconvolution algorithms provides a comprehensive analysis of their methodological foundations, contrasting the underlying computational algorithms, modeling methods, and data processing pipelines. This review is a methodological handbook that can help you understand the strengths and limitations of each approach.
A practical approach is to run deconvolution with multiple algorithms and compare the results. If different algorithms produce substantially different results, this indicates that the data may not support reliable deconvolution. If the results are consistent across algorithms, you can have more confidence in the output.
Step 6: Validate Against Independent Information
The sixth decision is how to validate your deconvolution results. Deconvolution results should always be interpreted in the context of tissue histology.
Use histological annotations to check that cell types appear in the expected anatomical regions. If the results show epithelial cells in the stroma or vice versa, there is a problem with the reference or the algorithm.
Compare your results with independent measurements if available. In situ hybridization or immunohistochemistry can provide direct evidence of cell type locations. The STARmap PLUS, RIBOmap, and TEMPOmap protocol describes imaging-based approaches for spatially resolved profiling of mRNA life cycle at transcriptome scale in intact cells and tissues. These methods can provide orthogonal validation for deconvolution results.
Step 7: Document Decisions and Record Outcomes
The seventh decision is how to document your analysis so that it is reproducible and interpretable by others. Record the rationale for each decision in the framework, including the biological resolution requirement, the reference data source, the platform differences, the spot purity assessment, the algorithm selection, and the validation results.
The nf-core documentation provides standards for community pipeline usage and configuration that emphasize reproducible workflow practices. Following these standards can help ensure that your deconvolution analysis is reproducible.
The Galaxy single-cell and spatial omics community has developed more than 175 tools and 120 training resources to help researchers perform and interpret their own analyses. These resources include guidance on reproducible analysis practices.
Implementing the Framework in Practice
To implement this framework, create a decision log that records your answers to each checkpoint. This log should be updated as you progress through the analysis and should be included in your final report or publication.
Decision Log Template
For each analysis, record the following information:
Biological resolution requirement: State whether you need cell type level, subtype level, or single-cell resolution output. Justify this choice based on your biological question.
Reference data source: Record the tissue, platform, sample preparation method, and quality control thresholds for your reference data. Note whether the reference contains all expected cell types and whether the marker genes are validated.
Gene overlap: Record the number of shared genes between reference and spatial data. Note whether this overlap is sufficient for reliable deconvolution.
Platform differences: Record the platforms used for reference and spatial data. Note any known differences in gene capture efficiency, sequencing depth, or normalization requirements.
Spot purity assessment: Record the expected number of cells per spot for your spatial platform. Note whether your tissue has complex cellular architecture that may reduce effective spot purity.
Algorithm selection: Record the deconvolution algorithm and version, the parameters used, and the rationale for this choice. Note whether multiple algorithms were tested and how the results compared.
Validation results: Record the histological annotations used for validation and whether the deconvolution results matched expectations. Note any discrepancies and the actions taken to address them.
Common Failure Patterns in Framework Implementation
The framework fails when researchers skip steps or make decisions based on convenience instead of data characteristics.
Skipping the biological resolution assessment leads to using a single-cell resolution method when cell type level output would suffice, adding unnecessary computational complexity. Conversely, using a cell type level method when subtype resolution is needed produces results that cannot answer the biological question.
Skipping the reference compatibility assessment leads to using a reference from a different tissue or developmental stage, producing cell types in impossible locations. The spatial transcriptomics protocol for early tooth morphogenesis emphasizes that the developing tooth comprises diverse and highly specialized cell populations. A reference from adult tissue will not contain these transient populations.
Skipping the platform difference assessment leads to systematic overestimation of one cell type across all spots. This pattern is often misinterpreted as a biological finding when it is actually a technical artifact.
Skipping the spot purity assessment leads to diffuse deconvolution results that do not match histological expectations. This is particularly problematic in tissues with complex cellular architecture.
Skipping the validation step leads to overinterpretation of deconvolution results that may be incorrect. Deconvolution is an estimation process, and the results depend on the quality of the inputs and the assumptions of the algorithm.
Troubleshooting Framework Failures
When deconvolution results do not match histological expectations, use the framework to identify which step failed.
If cell types appear in impossible locations, the reference compatibility assessment likely failed. Review the reference cell type annotations and ensure that all expected cell types are represented.
If one cell type dominates all spots, the reference may be missing a cell type that is present in the tissue. The algorithm assigns the missing cell type's signal to the most similar cell type in the reference.
If results change dramatically with small input changes, the data may not support reliable deconvolution. Test the robustness of your results by removing a small number of cells or spots and rerunning the analysis.
If subtypes are confused with each other, the reference signatures may be too similar. Use a marker gene assisted method that can enhance performance on cell subtypes. The SMART method provides a two-stage approach to improve performance on cell subtypes.
If results are consistent across algorithms but do not match histology, the problem may be in the reference data or the spatial data quality. Review the quality control thresholds and consider whether the data supports the analysis.
When to Escalate Beyond the Framework
Some problems cannot be solved by adjusting the analysis workflow. The following situations warrant escalation to a specialist or a change in experimental approach.
If deconvolution results consistently disagree with histological observations despite correcting reference composition, platform differences, and quality control, the spatial data may not be suitable for deconvolution. Consider using a higher-resolution platform or an image-based approach.
If the gene overlap between reference and spatial data is too low for reliable deconvolution, the experimental design may need to be changed. Consider generating reference data from the same tissue and platform as the spatial data.
If your computational resources are insufficient for the chosen method, consider using a less computationally intensive method or accessing shared computing resources. The scResolve protocol runs in a cluster environment to accelerate data processing.
If your research involves clinical samples or regulated data, ensure that your analysis complies with applicable regulations. The csRNA-seq protocol notes that purified RNA can be isolated from inactivated samples, including clinical or pathogenic specimens, allowing safe transport and analysis under standard laboratory conditions. Consult with your institutional review board or regulatory affairs office if you have questions about compliance.
Frequently Asked Questions
Why do my deconvolution results show cell types in locations where they cannot exist?
This usually indicates a problem with the reference panel. The reference may contain cell types that are not present in your tissue, or it may be missing cell types that are present. When a cell type is missing from the reference, its signal is distributed among the cell types that are present, which can cause those cell types to appear in incorrect locations. Validate your reference annotations against known marker genes and published data for your tissue.
How much gene overlap should I have between my reference and spatial data?
There is no universal threshold for gene overlap, but higher overlap generally produces more reliable deconvolution results. If the overlap is very low, the deconvolution will be based on a small set of genes and the results will be unstable. Calculate the number of shared genes before running deconvolution and consider using a different reference if the overlap is low.
Should I use single-cell or single-nucleus RNA sequencing data as my reference?
The choice depends on your tissue and your spatial platform. Single-cell RNA sequencing captures cytoplasmic mRNA, while single-nucleus RNA sequencing captures nuclear RNA. For tissues where these transcriptomes differ substantially, the reference should match the biology of the spatial data. For formalin-fixed paraffin-embedded tissue, consider using a reference generated from similar material.
Why does my deconvolution algorithm assign one cell type to almost every spot?
This pattern often indicates that the reference is missing a cell type that is present in the tissue. The algorithm assigns the missing cell type's signal to the most similar cell type in the reference. Review the tissue biology and ensure that all expected cell types are represented in the reference.
Can I trust deconvolution results without histological validation?
Deconvolution results should always be validated against independent information, such as histological annotations or in situ hybridization. Deconvolution is an estimation process, and the results depend on the quality of the inputs and the assumptions of the algorithm. Without validation, you cannot determine whether the results are reliable.
What should I do if different deconvolution algorithms give different results?
Different algorithms make different assumptions and may perform differently on your data. If the results are substantially different across algorithms, the data may not support reliable deconvolution. Investigate the source of the disagreement, which may be related to reference composition, platform differences, or data quality.
How do batch effects between reference and spatial data affect deconvolution?
Batch effects are systematic technical differences between datasets that are not related to biology. They can cause the algorithm to assign cell type proportions based on technical instead of biological signals. Methods that incorporate domain adaptation, such as AddaGCN, can mitigate batch effects, but they cannot eliminate them entirely.
When should I consider using a higher-resolution spatial platform?
If your biological question requires resolving fine cellular structures or subtle cell states, a higher-resolution platform may be necessary. Capture-based platforms have limited spot purity, and deconvolution cannot fully recover the information lost in mixed spots. Image-based platforms can achieve higher resolution but may have lower gene throughput.
Related Bioinformatics Guides
- Spatial Transcriptomics Study Design: Key Considerations for Robust Results
- Spatial Transcriptomics Methods: A Guide to Experimental Approaches
- Spatial Transcriptomics Differential Expression: Methods and Best Practices
- Spatial Transcriptomics Neighborhood Analysis: Tools and Best Practices
- Spatial Transcriptomics Data Analysis: A Guide to Preprocessing, Integration, and Interpretation
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- Spatial Transcriptomics of Early Tooth Morphogenesis in Formalin-fixed Paraffin-embedded Mouse Embryonic Tissue.. 2026.
- Galaxy single-cell & spatial omics community update: Navigating new frontiers in 2025.. 2025.
- Spatially resolved in situ profiling of mRNA life cycle at transcriptome scale in intact cells and tissues using STARmap PLUS, RIBOmap and TEMPOmap.. 2026.
- Protocol to recover single-cell gene expression profiles from spatial transcriptomics data using cluster computing.. 2025.
- Profiling active RNA polymerase II transcription start sites from total RNA by capped small RNA sequencing (csRNA-seq).. 2026.
- Protocol for single-cell RNA sequencing and spatial transcriptomics of bladder Ewing sarcoma.. 2025.
- Droplet-based single-cell RNA sequencing: decoding cellular heterogeneity for breakthroughs in cancer, reproduction, and beyond.. 2025.
- From pixels to cell types: a comprehensive review of computational methods for spatial transcriptomics deconvolution. Genomics & Informatics, 2025.
- STDSN: Domain Separation Network for Transfer Learning in Spatial Transcriptomics Deconvolution. IEEE International Conference on Bioinformatics and Biomedicine, 2025.
- SMART: spatial transcriptomics deconvolution using marker-gene-assisted topic model. Genome Biology, 2024.
- Spatial transcriptomics deconvolution at single-cell resolution using Redeconve. Nature Communications, 2023.
- AddaGCN: Spatial transcriptomics deconvolution using graph convolutional networks with adversarial discriminative domain adaptation. PLoS Computational Biology, 2026.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.