Integrating Single-Cell RNA-Seq Data with Reference Atlases for Improved Cell Type Annotation: A Step-by-Step Workflow

By Dr. Zubair Khalid, DVM, MS, PhD ·

Integrating Single-Cell RNA-Seq Data with Reference Atlases for Improved Cell Type Annotation: A Step-by-Step Workflow

Key Takeaways

  • Reference atlases, comprising large-scale integrated and annotated single-cell transcriptomic datasets (e.g., Human Lung Cell Atlas, Brain Cell Atlas), are crucial for overcoming limitations of traditional cluster-and-marker annotation, particularly for rare cell types, disease-altered states, and batch effect mitigation.
  • The selection of a reference atlas is paramount and must align with the query data's tissue type, species, developmental stage, and disease context to ensure accurate label transfer, with criteria like mapping success and agreement with manual labels being critical evaluation points.
  • Preprocessing of query data requires strict adherence to normalization methods and gene set matching with the reference atlas, alongside robust quality control filtering to remove low-quality cells that can distort integration and subsequent annotation.
  • Core integration methods like anchor-based (e.g., Seurat's FindTransferAnchors) and latent space methods (e.g., Harmony) offer distinct trade-offs in computational efficiency and ability to preserve continuous states versus discrete types, while ensemble and self-supervised approaches enhance uncertainty assessment and robustness.
  • Validation of reference-based annotations is essential, employing independent evidence such as canonical marker gene expression, spatial transcriptomics, or protein-level measurements (e.g., CITE-seq), and systematic documentation of mapping statistics, prediction score distributions, and cluster-cell type agreement is vital for reproducibility.

Cell type annotation is a decisive step in single-cell RNA sequencing (scRNA-seq) analysis. Manual annotation based on canonical marker genes becomes unreliable when your dataset contains rare populations, novel states, or cells that shift their transcriptional program due to disease or experimental perturbation. Reference atlas integration addresses this problem by projecting your query data into a harmonized coordinate system built from many annotated datasets, then transferring labels through supervised learning. This workflow is appropriate for researchers who have already performed basic quality control and normalization on their own data and now need to assign cell identities with confidence. The practical outcome is a reproducible pipeline that produces annotation scores, uncertainty estimates, and a clear record of which cells could not be confidently assigned.

The Problem with Cluster-and-Marker Annotation Alone

The conventional approach to cell type annotation begins with unsupervised clustering of your own dataset, followed by differential expression testing to find cluster markers, and then manual matching of those markers to known cell types. This strategy works well when your tissue of interest has well-characterized populations with strong, specific markers. It fails in several common situations.

First, rare cell types often do not form their own clusters. When a population represents less than one percent of your cells, standard graph-based clustering may merge it with a more abundant related type. The integrated Human Lung Cell Atlas, which combined 49 datasets spanning over 2.4 million cells from 486 individuals, demonstrated that atlas-level integration can reveal rare and previously undescribed cell types that individual studies miss. The authors showed that mapping new data to this atlas enables rapid annotation and interpretation, precisely because the reference contains the full diversity of cell states that any single experiment is unlikely to capture.

Second, disease states alter marker gene expression. A macrophage in a tumor microenvironment does not express the same marker panel as a resting macrophage from peripheral blood. The benchmarking study of five tumor immune atlases found that supervised annotation using reference atlases consistently outperformed unsupervised clustering in identifying cell states related to immunotherapy response. This finding directly supports the use of reference-based annotation when your biological question involves disease-associated cell states.

Third, batch effects and technical variation can dominate biological signal. When you cluster your data alone, the top principal components often separate sequencing batches instead of cell types. Reference integration methods are designed to remove these technical differences while preserving biological variation, allowing your cells to be placed in a shared space where annotation is meaningful.

What Reference Atlases Provide

A reference atlas is a collection of single-cell transcriptomes from many donors, studies, or tissues that have been integrated into a common coordinate system and annotated with consensus cell type labels. The value of these resources comes from their scale and their consensus annotation process.

The Human Lung Cell Atlas integrated 49 datasets to create a reference spanning over 2.4 million cells from 486 individuals. This scale captures population-level variability that a single study with a handful of donors cannot. The authors used this diversity to identify gene modules associated with age, sex, and body mass index, and to define consensus cell type annotations with matching marker genes.

The Brain Cell Atlas assembled single-cell data from 70 human and 103 mouse studies, covering over 26.3 million cells or nuclei from healthy and diseased tissues across developmental stages and brain regions. The authors used machine-learning algorithms to provide consensus cell type annotation and demonstrated the identification of putative neural progenitor cells and a specific microglia subpopulation. This atlas illustrates how integration across many studies can reveal cell types that are too rare or too variable to be detected in any single experiment.

The NBAtlas for neuroblastoma integrated seven single-cell or single-nucleus datasets into a harmonized atlas covering 362,991 cells across 61 patients. The authors explicitly showcased the utility of their atlas as a reference for data-driven cell type annotation, demonstrating that the resource can be expanded with additional data and used to annotate new datasets.

For researchers working on the endometrium, the Human Endometrial Cell Atlas combined published and new datasets from 63 women with and without endometriosis, totaling 313,527 cells. The authors assigned consensus cell type labels, identified previously unreported cell types, and validated their annotations using spatial transcriptomics and an independent single-nuclei dataset. This atlas demonstrates the standard of validation that reference resources should meet.

At a Glance

Workflow StepPrimary Tool ClassInput RequiredOutput ProducedKey Decision Point
Reference selectionAtlas repositoriesTissue type, species, disease contextChosen reference objectMatch reference to your biological question
Query preprocessingStandard scRNA-seq pipelinesRaw counts, cell metadataNormalized, scaled query objectUse same gene set as reference
IntegrationAnchor-based or latent space methodsQuery and reference objectsShared embedding or corrected queryChoose method based on dataset size
Label transferSupervised classifiersIntegrated data, reference labelsPer-cell label predictions and scoresSet confidence threshold for assignment
Uncertainty assessmentEnsemble or voting methodsMultiple label predictionsConsensus labels and uncertainty scoresIdentify cells requiring manual review
ValidationMarker expression, spatial dataAnnotated query dataConfirmed cell type assignmentsCheck against independent evidence

Selecting an Appropriate Reference Atlas

The choice of reference atlas is the most consequential decision in this workflow. A mismatched reference will produce confident but incorrect annotations. The selection criteria should include tissue type, species, developmental stage, disease context, and the granularity of annotation you need.

Tissue and Disease Context Matching

The reference should come from the same tissue and, ideally, the same disease context as your query data. If you are studying lung tissue from patients with pulmonary fibrosis, a reference built from healthy lung tissue will not contain the activated fibroblast states or profibrotic macrophages that characterize the disease. The Human Lung Cell Atlas addressed this limitation by including data from diseased tissues and identifying shared cell states across COVID-19, pulmonary fibrosis, and lung carcinoma, including SPP1-positive profibrotic monocyte-derived macrophages.

For cancer research, the choice depends on whether you need a pan-cancer or cancer-specific reference. The benchmarking study of tumor immune atlases compared two pan-cancer and three cancer-specific atlases, finding that they differ in their data sources, integration strategies, and cell type definitions. The authors recommended that users evaluate atlases based on agreement with manual labels, mapping success, accuracy, clusterability, annotatability, and stability. This benchmarking framework provides concrete criteria for selecting among available references.

Species and Developmental Stage

Most published atlases are built from human or mouse data. If you work with a non-model organism, you may need to build your own reference or use cross-species mapping approaches. The updated single cell reference atlas for the starlet anemone Nematostella vectensis demonstrates that even non-conventional model organisms can benefit from atlas construction. The authors re-mapped existing sequence data to a new chromosome-level genome assembly and incorporated additional samples, producing transcriptomic signatures for 127 distinct cell states. This example shows that reference construction is feasible for organisms with less mature genomic resources, though it requires substantial computational work.

Developmental stage matters because cell types and states change dramatically during development. The transcription factor atlas of directed differentiation mapped TF-induced expression profiles to reference cell types and validated candidate TFs for generating diverse cell types spanning all three germ layers and trophoblasts. If your query data comes from differentiating cells in vitro, you should select a reference that includes the relevant developmental stages or use an organoid-specific atlas.

Annotation Granularity

References differ in the resolution of their cell type labels. Some provide broad categories like T cell or macrophage, while others distinguish subtypes such as pulmonary venous endothelial cells versus systemic venous endothelial cells. The integrated endothelial cell atlas of the human lung identified previously indistinguishable subpopulations, including two venous populations distinguished by COL15A1 expression and two capillary populations including aerocytes characterized by EDNRB, SOSTDC1, and TBX2. If your biological question requires this level of resolution, you need a reference that provides it.

The kidney endothelial cell atlas similarly identified seven endothelial subgroups that differed in molecular characteristics and physiologic functions. The authors demonstrated that mapping new data to this atlas allows rapid data annotation and analysis, and they confirmed that endothelial cell types were highly conserved between human and mouse kidney. This cross-species conservation can be useful if you need to annotate data from multiple organisms.

Preparing Your Query Data for Integration

Before you can integrate your data with a reference atlas, your query object must meet certain preprocessing standards. The specific requirements depend on the integration method you choose, but several principles apply broadly.

Quality Control and Filtering

Your query data should have undergone standard quality control before integration. This includes filtering cells by the number of detected genes, the number of unique molecular identifiers, and the percentage of mitochondrial reads. Cells that fail these thresholds should be removed before integration because they can distort the alignment between query and reference.

The choice between single-cell and single-nucleus data affects your quality control thresholds. Single-nucleus RNA sequencing typically has lower gene detection per cell and different ambient RNA contamination patterns compared to whole-cell scRNA-seq. If your reference was built from single-nucleus data and your query is from whole cells, or vice versa, you should expect some technical differences that integration methods must correct.

Normalization and Gene Selection

Your query data should be normalized using the same method that was used to build the reference. Most modern atlases use log-normalization or SCTransform. Using a different normalization method can introduce systematic differences that complicate integration.

The gene set used for integration should match between query and reference. Most integration methods require that you restrict both datasets to a common set of highly variable genes or to the intersection of genes detected in both. The quality of integration depends on this gene set being biologically informative. If you use only the union of all genes, you will include genes with little biological signal and increase the computational burden.

Handling Batch Effects Within Your Query

If your query dataset contains multiple samples, sequencing batches, or experimental conditions, you should decide whether to integrate them before or after mapping to the reference. Some methods, such as Harmony, can correct for batch effects within your query while simultaneously aligning to the reference. Other methods require that you integrate your own batches first and then map the integrated object to the reference.

The self-supervised learning framework scDecorr addresses this challenge by learning cell embeddings independently across domains and using domain-specific batch normalization. The authors demonstrated that this approach integrates batches without losing inherent biological variance, facilitating optimal clustering, and that the representations are robust in label transfer tasks. This method illustrates the principle that batch correction and reference mapping can be accomplished in a single step.

Core Integration Methods and Their Tradeoffs

Several computational methods are available for integrating query data with reference atlases. The choice of method depends on your dataset size, the similarity between query and reference, and whether you need to preserve continuous cell states or discrete cell types.

Anchor-Based Integration with Seurat

Seurat's FindTransferAnchors and MapQuery functions implement an anchor-based approach to reference mapping. This method identifies pairs of cells between query and reference that are mutual nearest neighbors in a shared space, then uses these anchors to transform the query into the reference coordinate system.

The anchor-based approach works well when the query and reference share most cell types and the technical differences are moderate. It produces a per-cell prediction score that reflects the confidence of the label assignment. The method is computationally efficient for datasets up to several hundred thousand cells.

The main limitation of anchor-based integration is that it assumes the query can be represented as a combination of reference cell types. If your query contains a genuinely novel cell type with no counterpart in the reference, the method will assign it to the most transcriptionally similar reference type, potentially with high confidence. You must check the prediction scores and the distribution of cells in the integrated space to identify this situation.

Latent Space Methods

Harmony and similar methods learn a low-dimensional embedding that removes batch effects while preserving biological variation. These methods are particularly useful when you have multiple batches within your query that need correction in addition to reference alignment.

The product-of-experts VAE model PoE-VAE offers a flexible approach for integrating multimodal data and mapping queries onto a reference atlas. The authors demonstrated that this method can integrate CITE-seq and multiome data, accurately mapping unimodal and multimodal queries into a joint latent space. They also extended the approach to spatial data by integrating gene expression and metabolomics from paired Visium and MALDI-MSI slides. This method is appropriate when your data includes multiple modalities beyond RNA expression.

Self-Supervised and Contrastive Learning Approaches

Newer methods leverage self-supervised learning to learn representations that are robust to technical variation. scDecorr uses feature decorrelation to capture biological signal while eliminating technical noise. The authors demonstrated that this approach achieves domain-invariant representations and performs well in label transfer tasks.

For large-scale integration across many datasets, self-supervised contrastive learning has been applied to central nervous system disease data. This approach learns representations by maximizing agreement between augmented views of the same cell while decorrelating different cells. The resulting embeddings are suitable for both clustering and label transfer.

Ensemble and Voting Methods

The popV method addresses a critical limitation of single-model label transfer: the lack of uncertainty estimation. popV is an ensemble of prediction models with an ontology-based voting scheme. The authors demonstrated that popV achieves accurate cell type labeling and provides uncertainty scores, confidently annotating the majority of cells while highlighting populations that are challenging to annotate by label transfer.

The practical value of uncertainty scores is that they reduce the burden of manual inspection. Instead of reviewing every cell, you can focus on the populations that popV identifies as problematic. This approach streamlines the annotation process and is particularly valuable for large datasets where manual review of all cells is impractical.

Step-by-Step Workflow for Reference-Based Annotation

The following workflow assumes you have a processed query object with normalized counts and a reference atlas object with cell type labels. The steps are presented in the order you should execute them.

Step 1: Verify Reference and Query Compatibility

Before running any integration, confirm that your query and reference are compatible. Check the species, tissue, and gene annotation versions. If your query uses a different genome build or gene annotation than the reference, you must convert gene identifiers to a common standard.

The NCBI provides official descriptions of databases, search systems, sequence resources, and analysis services that can help you verify gene identifiers and genome builds. Cross-referencing your gene symbols against a standard annotation ensures that the integration is based on the same genes in both datasets.

Step 2: Restrict to Common Genes

Identify the intersection of genes between your query and the reference. For most integration methods, you should use the highly variable genes from the reference, restricted to those present in your query. This step reduces computational cost and focuses the integration on genes that carry biological signal.

If your query has many genes that are not in the reference, you should investigate whether this reflects a technical issue such as different gene annotation versions or a biological difference such as a different cell type composition. Large numbers of reference-exclusive genes in your query may indicate that the reference is not appropriate for your data.

Step 3: Run Integration

Execute the integration method appropriate for your data. For Seurat's anchor-based approach, this involves running FindTransferAnchors followed by MapQuery. For Harmony, you will run the integration within a dimensional reduction step. For scDecorr or PoE-VAE, you will train the model on your query and reference together.

Record the parameters you use, including the number of dimensions, the number of anchors or neighbors, and any batch correction variables. These parameters affect the quality of integration and should be reported in your methods.

Step 4: Examine the Integrated Space

After integration, visualize your query cells in the shared embedding alongside reference cells. Use UMAP or t-SNE to project the integrated data. Your query cells should overlap with reference cells of the same type. If your query cells form a separate cluster with no reference overlap, this may indicate a novel cell type or a technical failure of integration.

Check the distribution of prediction scores. Most cells should have high scores for their assigned label. A bimodal distribution of scores, with many cells at intermediate values, suggests that the integration is not cleanly separating cell types.

Step 5: Transfer Labels and Assess Confidence

Transfer the reference labels to your query cells using the integration results. Each cell will receive a predicted label and a score. Set a threshold for confident assignment based on the distribution of scores in your data.

The popV approach provides a more rigorous framework for this step. By running multiple prediction models and using an ontology-based voting scheme, popV produces consensus labels and uncertainty scores. Cells with high uncertainty should be flagged for manual review.

Step 6: Validate Annotations with Independent Evidence

Reference-based annotation is a prediction, not a ground truth. You should validate your annotations using independent evidence. This can include marker gene expression, spatial transcriptomics data, or protein-level measurements.

The bone marrow niche atlas integrated scRNA-seq data with CODEX spatial proteomic imaging to link predicted cellular signaling with spatial proximity. The authors used this integrated approach to annotate new images and uncover mesenchymal stromal cell expansion in acute myeloid leukemia patient samples. This example demonstrates the power of combining transcriptomic annotation with spatial validation.

For immune cell annotation, the review of immune cell annotation challenges emphasizes that mRNA and protein expression can diverge due to post-transcriptional regulation and technical artifacts. Technologies such as CITE-seq, which integrate transcriptomic and proteomic data, address the shortcomings of single-modality analyses. If your validation relies on protein markers, you should be aware that mRNA-based annotations may not perfectly match protein expression.

Step 7: Document and Report

Record the reference atlas version, the integration method and parameters, the confidence thresholds, and the validation results. This documentation is essential for reproducibility and for interpreting your downstream analyses.

The Galaxy Training Network provides accessible workflow training and analysis tutorials that emphasize reproducibility in bioinformatics analysis. Following these standards ensures that your annotation workflow can be reproduced by others in your field.

Records and Measurements for Annotation Quality

Systematic record keeping is essential for evaluating the quality of your reference-based annotation. The following measurements should be recorded for every integration run.

Mapping Statistics

Record the number of query cells that successfully mapped to the reference, the number that failed to map, and the number that mapped with low confidence. These statistics indicate the overall compatibility between your query and the reference.

The benchmarking study of tumor immune atlases evaluated mapping success as one of five criteria for atlas quality. A high mapping failure rate suggests that your query contains cell types or states not represented in the reference, or that technical differences between query and reference are too large for the integration method to correct.

Prediction Score Distributions

For each cell type in your query, record the distribution of prediction scores. Cell types with consistently high scores are likely well-represented in the reference. Cell types with variable or low scores may require manual review.

The popV framework provides uncertainty scores that quantify the confidence of each annotation. Cells with high uncertainty should be prioritized for manual inspection. The authors demonstrated that this approach reduces the load of manual inspection by focusing attention on the most problematic parts of the annotation.

Cluster-Cell Type Agreement

After annotation, compare your unsupervised clustering results with the transferred labels. If cells from a single cluster receive multiple different labels, this may indicate that the cluster contains multiple cell types that were not resolved by clustering alone. Conversely, if cells from multiple clusters receive the same label, this may indicate that the label is too broad or that the clusters represent different states of the same cell type.

The integrated lung endothelial cell atlas identified previously indistinguishable subpopulations among venous and capillary endothelial cells. This finding illustrates that reference-based annotation can resolve cell types that are not apparent from clustering your own data alone.

Marker Gene Validation

For each assigned cell type, check the expression of canonical marker genes in your query cells. This validation step is independent of the integration and provides confidence that the transferred labels are biologically meaningful.

The transcription factor atlas of directed differentiation validated candidate TFs by mapping TF-induced expression profiles to reference cell types. This approach demonstrates the principle that annotation should be confirmed by independent biological evidence.

Common Failure Patterns and Troubleshooting

Reference-based annotation can fail in several predictable ways. Recognizing these failure patterns allows you to diagnose problems and adjust your workflow.

Reference Mismatch

The most common failure is using a reference that does not match your biological context. If you study diseased tissue and use a healthy reference, your disease-associated cell states will be forced into the nearest healthy cell type. The Human Lung Cell Atlas addressed this by including diseased tissues and identifying shared cell states across multiple lung diseases. If your tissue of interest has a disease-specific atlas, you should use it instead of a healthy reference.

Technical Batch Effects Too Large for Correction

Integration methods can correct for moderate technical differences, but extreme differences in sequencing platform, library preparation, or sample quality may prevent successful integration. If your query cells form a separate cluster with no reference overlap, the technical differences may be too large for the integration method to handle.

The scDecorr framework addresses this by learning cell embeddings independently across domains and employing domain-specific batch normalization. If standard integration methods fail, you may need to use a method specifically designed for large technical differences.

Novel Cell Types

If your query contains a cell type that is absent from the reference, the integration method will assign it to the most transcriptionally similar reference type. This assignment may be incorrect and may have high confidence if the novel type is similar to a reference type.

The Brain Cell Atlas identified putative neural progenitor cells and a PCDH9-high microglia subpopulation that were not previously described. These discoveries were possible because the atlas integrated data from many studies, revealing cell types that individual studies missed. If you suspect your query contains novel cell types, you should examine cells with low prediction scores or high uncertainty and compare them to the reference using unsupervised clustering.

Gene Annotation Inconsistencies

If your query and reference use different gene annotation versions, many genes will not match. This can severely degrade integration quality. Always verify that gene identifiers are consistent between query and reference before running integration.

The updated Nematostella reference atlas demonstrated the importance of gene annotation quality. The authors re-mapped sequence data to a new chromosome-level genome assembly with improved gene models, which markedly improved the mapping of single-cell reads. Poor gene annotations can limit the conclusions that can be drawn from single-cell data.

Limitations of Reference-Based Annotation

Reference-based annotation is a powerful tool, but it has inherent limitations that you should understand before relying on it exclusively.

Reference Bias

The annotations you obtain are limited by the cell types and states present in the reference. If the reference does not contain a particular cell type, your query cells of that type will be misannotated. This is not a failure of the method but a fundamental limitation of supervised learning.

The benchmarking study of tumor immune atlases found that existing atlases vary in their data sources, integration and annotation strategies, and the number and definition of cell types. This variability means that different references may produce different annotations for the same query data. You should be aware of this variability and consider using multiple references or an ensemble approach.

mRNA-Protein Discrepancies

Cell type annotation based on mRNA expression may not perfectly reflect protein-level phenotypes. The review of immune cell annotation challenges detailed the mechanisms underlying mRNA-protein discrepancies, including gene expression heterogeneity, post-transcriptional regulation, and technical artifacts. These discrepancies can lead to cell misclassification, particularly in heterogeneous populations such as peripheral blood mononuclear cells.

If your downstream experiments depend on protein-level phenotypes, you should validate your mRNA-based annotations using protein measurements. Technologies such as CITE-seq integrate transcriptomic and proteomic data to address this limitation.

Continuous Cell States

Some biological processes, such as differentiation and activation, produce continuous cell states instead of discrete cell types. Reference-based annotation works best for discrete cell types and may not capture the full spectrum of continuous states.

The transcription factor atlas of directed differentiation addressed this by mapping TF-induced expression profiles to reference cell types and developing a strategy for predicting combinations of TFs that produce target expression profiles matching reference cell types. This approach acknowledges that cell states exist on a continuum and that annotation to discrete types is a simplification.

Welfare and Safety Context for Laboratory Practice

While reference-based annotation is a computational procedure, it has implications for laboratory practice and research integrity that warrant attention.

Data Management and Reproducibility

The Carpentries lessons provide foundational training in computing, data, shell, Git, and programming that is directly relevant to reproducible single-cell analysis. Version control for your analysis scripts and reference atlas versions is essential for reproducibility. The nf-core documentation describes community pipeline standards that emphasize reproducibility and configuration management.

The EMBL-EBI Training program offers bioinformatics learning pathways and data-resource training that can help you develop the computational skills needed for rigorous single-cell analysis. Investing in these skills reduces the risk of analysis errors that could compromise your biological conclusions.

Reporting Standards

When you report reference-based annotation in your publications, you should include the reference atlas version, the integration method and parameters, the confidence thresholds, and the validation results. This information allows other researchers to evaluate the reliability of your annotations and to reproduce your analysis.

The Galaxy Training Network provides accessible workflow training that emphasizes reproducibility in bioinformatics analysis. Following these standards ensures that your annotation workflow can be reproduced by others in your field.

Escalation Criteria

You should escalate to a more experienced bioinformatician or computational biologist when you encounter any of the following situations. First, if your query cells form a separate cluster with no reference overlap after integration, you may be dealing with a novel cell type or a technical failure that requires expert diagnosis. Second, if your prediction scores are uniformly low across all cell types, your reference may be inappropriate for your data. Third, if your validation markers do not match your transferred labels, you may have a systematic annotation error that requires investigation.

The Bioconductor project provides official package, workflow, installation, and reproducible genomic-analysis documentation that can help you troubleshoot integration issues. Consulting this documentation before escalating can resolve many common problems.

Practical Implementation Steps

The following steps provide a concrete implementation plan for integrating your query data with a reference atlas.

Step 1: Inventory Your Data

Document the number of cells, the number of genes, the sequencing platform, the tissue source, and the experimental conditions for your query data. This inventory informs your choice of reference atlas and integration method.

Step 2: Select a Reference

Identify candidate reference atlases for your tissue and species. Evaluate each candidate based on the number of donors, the diversity of conditions, the annotation granularity, and the availability of validation data. The benchmarking criteria from the tumor immune atlas study provide a useful framework for this evaluation.

Step 3: Preprocess Your Query

Run standard quality control, normalization, and feature selection on your query data. Ensure that your gene identifiers match the reference annotation.

Step 4: Run Integration

Execute the integration method appropriate for your data. Record all parameters and the version of the reference atlas.

Step 5: Assess Annotation Quality

Examine the integrated space, the prediction score distributions, and the cluster-cell type agreement. Validate your annotations using marker gene expression and independent evidence.

Step 6: Document and Report

Record all analysis steps, parameters, and quality metrics. Prepare a methods section that allows other researchers to reproduce your annotation workflow.

Frequently Asked Questions

What is the difference between reference-based annotation and unsupervised clustering?

Unsupervised clustering groups cells based on transcriptional similarity without using external information. Reference-based annotation uses a pre-annotated atlas to assign labels to your cells through supervised learning. The benchmarking study of tumor immune atlases found that supervised annotations consistently outperformed unsupervised clustering in identifying cell states related to immunotherapy response. Reference-based annotation is particularly valuable when your dataset contains rare cell types or disease-associated states that may not form distinct clusters in your data alone.

How do I choose between Seurat's FindTransferAnchors and Harmony for integration?

Seurat's anchor-based approach is designed for mapping a query dataset to a reference and transferring labels. Harmony is a general-purpose batch correction method that can integrate multiple datasets without requiring a designated reference. If your primary goal is label transfer from a reference atlas, the anchor-based approach is appropriate. If you need to integrate multiple batches within your query while also aligning to a reference, Harmony may be more suitable. The choice also depends on dataset size, with anchor-based methods being computationally efficient for datasets up to several hundred thousand cells.

What should I do if my query cells do not overlap with any reference cells in the integrated space?

This situation indicates that your query contains cell types or states that are not represented in the reference, or that technical differences between query and reference are too large for the integration method to correct. First, check whether your gene identifiers match the reference annotation. Second, try a different integration method that is more robust to technical differences. Third, examine the non-overlapping cells using unsupervised clustering to determine whether they represent a novel cell type. If they do, you may need to build a custom reference or use a different annotation strategy.

How do I set a confidence threshold for label transfer?

The confidence threshold should be based on the distribution of prediction scores in your data. Examine the score distribution and identify a natural cutoff that separates high-confidence from low-confidence assignments. The popV method provides uncertainty scores that can guide this decision. Cells below the threshold should be flagged for manual review or excluded from downstream analyses that require confident cell type assignments.

Can I use a reference atlas from a different species to annotate my data?

Cross-species annotation is possible when cell types are highly conserved between species. The kidney endothelial cell atlas confirmed that endothelial cell types in the human kidney were highly conserved in the mouse kidney and identified marker genes that were conserved across species. However, cross-species annotation introduces additional technical challenges, including gene name mapping and species-specific expression differences. You should validate cross-species annotations carefully using marker gene expression and other independent evidence.

What is the role of cell type ontologies in reference-based annotation?

Cell type ontologies provide standardized vocabulary for cell type names and their relationships. The Cell Type Ontologies of the Human Cell Atlas provide a framework for consistent annotation across datasets and studies. Using ontology-based labels in your reference and query ensures that annotations are comparable across analyses and can be integrated with other resources. The popV method uses an ontology-based voting scheme to reconcile labels from multiple prediction models.

How do I handle multimodal data in reference-based annotation?

If your data includes multiple modalities, such as RNA and protein expression from CITE-seq, you should use an integration method that can handle multimodal data. The PoE-VAE model integrates multimodal data and allows for the seamless mapping of both unimodal and multimodal queries onto a reference atlas. The multimodal hierarchical classification approach for CITE-seq data delineates immune cell states across lineages and tissues. These methods are appropriate when your data includes modalities beyond RNA expression.

What are the limitations of using a reference atlas for rare cell type identification?

Reference-based annotation can identify rare cell types if they are present in the reference. The Human Lung Cell Atlas identified rare and previously undescribed cell types through atlas-level integration. However, if a rare cell type is absent from the reference, your query cells of that type will be misannotated. Additionally, rare cell types may have low prediction scores because they are represented by few cells in the reference. You should examine cells with low scores or high uncertainty to identify potentially novel or rare populations.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.