Integrating Single-Cell RNA-Seq Data with Reference Atlases for Improved Cell Type Annotation: A Step-by-Step Workflow
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Reference atlases, comprising large-scale integrated and annotated single-cell transcriptomic datasets (e.g., Human Lung Cell Atlas, Brain Cell Atlas), are crucial for overcoming limitations of traditional cluster-and-marker annotation, particularly for rare cell types, disease-altered states, and batch effect mitigation.
- The selection of a reference atlas is paramount and must align with the query data's tissue type, species, developmental stage, and disease context to ensure accurate label transfer, with criteria like mapping success and agreement with manual labels being critical evaluation points.
- Preprocessing of query data requires strict adherence to normalization methods and gene set matching with the reference atlas, alongside robust quality control filtering to remove low-quality cells that can distort integration and subsequent annotation.
- Core integration methods like anchor-based (e.g., Seurat's FindTransferAnchors) and latent space methods (e.g., Harmony) offer distinct trade-offs in computational efficiency and ability to preserve continuous states versus discrete types, while ensemble and self-supervised approaches enhance uncertainty assessment and robustness.
- Validation of reference-based annotations is essential, employing independent evidence such as canonical marker gene expression, spatial transcriptomics, or protein-level measurements (e.g., CITE-seq), and systematic documentation of mapping statistics, prediction score distributions, and cluster-cell type agreement is vital for reproducibility.
Cell type annotation is a decisive step in single-cell RNA sequencing (scRNA-seq) analysis. Manual annotation based on canonical marker genes becomes unreliable when your dataset contains rare populations, novel states, or cells that shift their transcriptional program due to disease or experimental perturbation. Reference atlas integration addresses this problem by projecting your query data into a harmonized coordinate system built from many annotated datasets, then transferring labels through supervised learning. This workflow is appropriate for researchers who have already performed basic quality control and normalization on their own data and now need to assign cell identities with confidence. The practical outcome is a reproducible pipeline that produces annotation scores, uncertainty estimates, and a clear record of which cells could not be confidently assigned.
The Problem with Cluster-and-Marker Annotation Alone
The conventional approach to cell type annotation begins with unsupervised clustering of your own dataset, followed by differential expression testing to find cluster markers, and then manual matching of those markers to known cell types. This strategy works well when your tissue of interest has well-characterized populations with strong, specific markers. It fails in several common situations.
First, rare cell types often do not form their own clusters. When a population represents less than one percent of your cells, standard graph-based clustering may merge it with a more abundant related type. The integrated Human Lung Cell Atlas, which combined 49 datasets spanning over 2.4 million cells from 486 individuals, demonstrated that atlas-level integration can reveal rare and previously undescribed cell types that individual studies miss. The authors showed that mapping new data to this atlas enables rapid annotation and interpretation, precisely because the reference contains the full diversity of cell states that any single experiment is unlikely to capture.
Second, disease states alter marker gene expression. A macrophage in a tumor microenvironment does not express the same marker panel as a resting macrophage from peripheral blood. The benchmarking study of five tumor immune atlases found that supervised annotation using reference atlases consistently outperformed unsupervised clustering in identifying cell states related to immunotherapy response. This finding directly supports the use of reference-based annotation when your biological question involves disease-associated cell states.
Third, batch effects and technical variation can dominate biological signal. When you cluster your data alone, the top principal components often separate sequencing batches instead of cell types. Reference integration methods are designed to remove these technical differences while preserving biological variation, allowing your cells to be placed in a shared space where annotation is meaningful.
What Reference Atlases Provide
A reference atlas is a collection of single-cell transcriptomes from many donors, studies, or tissues that have been integrated into a common coordinate system and annotated with consensus cell type labels. The value of these resources comes from their scale and their consensus annotation process.
The Human Lung Cell Atlas integrated 49 datasets to create a reference spanning over 2.4 million cells from 486 individuals. This scale captures population-level variability that a single study with a handful of donors cannot. The authors used this diversity to identify gene modules associated with age, sex, and body mass index, and to define consensus cell type annotations with matching marker genes.
The Brain Cell Atlas assembled single-cell data from 70 human and 103 mouse studies, covering over 26.3 million cells or nuclei from healthy and diseased tissues across developmental stages and brain regions. The authors used machine-learning algorithms to provide consensus cell type annotation and demonstrated the identification of putative neural progenitor cells and a specific microglia subpopulation. This atlas illustrates how integration across many studies can reveal cell types that are too rare or too variable to be detected in any single experiment.
The NBAtlas for neuroblastoma integrated seven single-cell or single-nucleus datasets into a harmonized atlas covering 362,991 cells across 61 patients. The authors explicitly showcased the utility of their atlas as a reference for data-driven cell type annotation, demonstrating that the resource can be expanded with additional data and used to annotate new datasets.
For researchers working on the endometrium, the Human Endometrial Cell Atlas combined published and new datasets from 63 women with and without endometriosis, totaling 313,527 cells. The authors assigned consensus cell type labels, identified previously unreported cell types, and validated their annotations using spatial transcriptomics and an independent single-nuclei dataset. This atlas demonstrates the standard of validation that reference resources should meet.
At a Glance
| Workflow Step | Primary Tool Class | Input Required | Output Produced | Key Decision Point |
|---|---|---|---|---|
| Reference selection | Atlas repositories | Tissue type, species, disease context | Chosen reference object | Match reference to your biological question |
| Query preprocessing | Standard scRNA-seq pipelines | Raw counts, cell metadata | Normalized, scaled query object | Use same gene set as reference |
| Integration | Anchor-based or latent space methods | Query and reference objects | Shared embedding or corrected query | Choose method based on dataset size |
| Label transfer | Supervised classifiers | Integrated data, reference labels | Per-cell label predictions and scores | Set confidence threshold for assignment |
| Uncertainty assessment | Ensemble or voting methods | Multiple label predictions | Consensus labels and uncertainty scores | Identify cells requiring manual review |
| Validation | Marker expression, spatial data | Annotated query data | Confirmed cell type assignments | Check against independent evidence |
Selecting an Appropriate Reference Atlas
The choice of reference atlas is the most consequential decision in this workflow. A mismatched reference will produce confident but incorrect annotations. The selection criteria should include tissue type, species, developmental stage, disease context, and the granularity of annotation you need.
Tissue and Disease Context Matching
The reference should come from the same tissue and, ideally, the same disease context as your query data. If you are studying lung tissue from patients with pulmonary fibrosis, a reference built from healthy lung tissue will not contain the activated fibroblast states or profibrotic macrophages that characterize the disease. The Human Lung Cell Atlas addressed this limitation by including data from diseased tissues and identifying shared cell states across COVID-19, pulmonary fibrosis, and lung carcinoma, including SPP1-positive profibrotic monocyte-derived macrophages.
For cancer research, the choice depends on whether you need a pan-cancer or cancer-specific reference. The benchmarking study of tumor immune atlases compared two pan-cancer and three cancer-specific atlases, finding that they differ in their data sources, integration strategies, and cell type definitions. The authors recommended that users evaluate atlases based on agreement with manual labels, mapping success, accuracy, clusterability, annotatability, and stability. This benchmarking framework provides concrete criteria for selecting among available references.
Species and Developmental Stage
Most published atlases are built from human or mouse data. If you work with a non-model organism, you may need to build your own reference or use cross-species mapping approaches. The updated single cell reference atlas for the starlet anemone Nematostella vectensis demonstrates that even non-conventional model organisms can benefit from atlas construction. The authors re-mapped existing sequence data to a new chromosome-level genome assembly and incorporated additional samples, producing transcriptomic signatures for 127 distinct cell states. This example shows that reference construction is feasible for organisms with less mature genomic resources, though it requires substantial computational work.
Developmental stage matters because cell types and states change dramatically during development. The transcription factor atlas of directed differentiation mapped TF-induced expression profiles to reference cell types and validated candidate TFs for generating diverse cell types spanning all three germ layers and trophoblasts. If your query data comes from differentiating cells in vitro, you should select a reference that includes the relevant developmental stages or use an organoid-specific atlas.
Annotation Granularity
References differ in the resolution of their cell type labels. Some provide broad categories like T cell or macrophage, while others distinguish subtypes such as pulmonary venous endothelial cells versus systemic venous endothelial cells. The integrated endothelial cell atlas of the human lung identified previously indistinguishable subpopulations, including two venous populations distinguished by COL15A1 expression and two capillary populations including aerocytes characterized by EDNRB, SOSTDC1, and TBX2. If your biological question requires this level of resolution, you need a reference that provides it.
The kidney endothelial cell atlas similarly identified seven endothelial subgroups that differed in molecular characteristics and physiologic functions. The authors demonstrated that mapping new data to this atlas allows rapid data annotation and analysis, and they confirmed that endothelial cell types were highly conserved between human and mouse kidney. This cross-species conservation can be useful if you need to annotate data from multiple organisms.
Preparing Your Query Data for Integration
Before you can integrate your data with a reference atlas, your query object must meet certain preprocessing standards. The specific requirements depend on the integration method you choose, but several principles apply broadly.
Quality Control and Filtering
Your query data should have undergone standard quality control before integration. This includes filtering cells by the number of detected genes, the number of unique molecular identifiers, and the percentage of mitochondrial reads. Cells that fail these thresholds should be removed before integration because they can distort the alignment between query and reference.
The choice between single-cell and single-nucleus data affects your quality control thresholds. Single-nucleus RNA sequencing typically has lower gene detection per cell and different ambient RNA contamination patterns compared to whole-cell scRNA-seq. If your reference was built from single-nucleus data and your query is from whole cells, or vice versa, you should expect some technical differences that integration methods must correct.
Normalization and Gene Selection
Your query data should be normalized using the same method that was used to build the reference. Most modern atlases use log-normalization or SCTransform. Using a different normalization method can introduce systematic differences that complicate integration.
The gene set used for integration should match between query and reference. Most integration methods require that you restrict both datasets to a common set of highly variable genes or to the intersection of genes detected in both. The quality of integration depends on this gene set being biologically informative. If you use only the union of all genes, you will include genes with little biological signal and increase the computational burden.
Handling Batch Effects Within Your Query
If your query dataset contains multiple samples, sequencing batches, or experimental conditions, you should decide whether to integrate them before or after mapping to the reference. Some methods, such as Harmony, can correct for batch effects within your query while simultaneously aligning to the reference. Other methods require that you integrate your own batches first and then map the integrated object to the reference.
The self-supervised learning framework scDecorr addresses this challenge by learning cell embeddings independently across domains and using domain-specific batch normalization. The authors demonstrated that this approach integrates batches without losing inherent biological variance, facilitating optimal clustering, and that the representations are robust in label transfer tasks. This method illustrates the principle that batch correction and reference mapping can be accomplished in a single step.
Core Integration Methods and Their Tradeoffs
Several computational methods are available for integrating query data with reference atlases. The choice of method depends on your dataset size, the similarity between query and reference, and whether you need to preserve continuous cell states or discrete cell types.
Anchor-Based Integration with Seurat
Seurat's FindTransferAnchors and MapQuery functions implement an anchor-based approach to reference mapping. This method identifies pairs of cells between query and reference that are mutual nearest neighbors in a shared space, then uses these anchors to transform the query into the reference coordinate system.
The anchor-based approach works well when the query and reference share most cell types and the technical differences are moderate. It produces a per-cell prediction score that reflects the confidence of the label assignment. The method is computationally efficient for datasets up to several hundred thousand cells.
The main limitation of anchor-based integration is that it assumes the query can be represented as a combination of reference cell types. If your query contains a genuinely novel cell type with no counterpart in the reference, the method will assign it to the most transcriptionally similar reference type, potentially with high confidence. You must check the prediction scores and the distribution of cells in the integrated space to identify this situation.
Latent Space Methods
Harmony and similar methods learn a low-dimensional embedding that removes batch effects while preserving biological variation. These methods are particularly useful when you have multiple batches within your query that need correction in addition to reference alignment.
The product-of-experts VAE model PoE-VAE offers a flexible approach for integrating multimodal data and mapping queries onto a reference atlas. The authors demonstrated that this method can integrate CITE-seq and multiome data, accurately mapping unimodal and multimodal queries into a joint latent space. They also extended the approach to spatial data by integrating gene expression and metabolomics from paired Visium and MALDI-MSI slides. This method is appropriate when your data includes multiple modalities beyond RNA expression.
Self-Supervised and Contrastive Learning Approaches
Newer methods leverage self-supervised learning to learn representations that are robust to technical variation. scDecorr uses feature decorrelation to capture biological signal while eliminating technical noise. The authors demonstrated that this approach achieves domain-invariant representations and performs well in label transfer tasks.
For large-scale integration across many datasets, self-supervised contrastive learning has been applied to central nervous system disease data. This approach learns representations by maximizing agreement between augmented views of the same cell while decorrelating different cells. The resulting embeddings are suitable for both clustering and label transfer.
Ensemble and Voting Methods
The popV method addresses a critical limitation of single-model label transfer: the lack of uncertainty estimation. popV is an ensemble of prediction models with an ontology-based voting scheme. The authors demonstrated that popV achieves accurate cell type labeling and provides uncertainty scores, confidently annotating the majority of cells while highlighting populations that are challenging to annotate by label transfer.
The practical value of uncertainty scores is that they reduce the burden of manual inspection. Instead of reviewing every cell, you can focus on the populations that popV identifies as problematic. This approach streamlines the annotation process and is particularly valuable for large datasets where manual review of all cells is impractical.
Step-by-Step Workflow for Reference-Based Annotation
The following workflow assumes you have a processed query object with normalized counts and a reference atlas object with cell type labels. The steps are presented in the order you should execute them.
Step 1: Verify Reference and Query Compatibility
Before running any integration, confirm that your query and reference are compatible. Check the species, tissue, and gene annotation versions. If your query uses a different genome build or gene annotation than the reference, you must convert gene identifiers to a common standard.
The NCBI provides official descriptions of databases, search systems, sequence resources, and analysis services that can help you verify gene identifiers and genome builds. Cross-referencing your gene symbols against a standard annotation ensures that the integration is based on the same genes in both datasets.
Step 2: Restrict to Common Genes
Identify the intersection of genes between your query and the reference. For most integration methods, you should use the highly variable genes from the reference, restricted to those present in your query. This step reduces computational cost and focuses the integration on genes that carry biological signal.
If your query has many genes that are not in the reference, you should investigate whether this reflects a technical issue such as different gene annotation versions or a biological difference such as a different cell type composition. Large numbers of reference-exclusive genes in your query may indicate that the reference is not appropriate for your data.
Step 3: Run Integration
Execute the integration method appropriate for your data. For Seurat's anchor-based approach, this involves running FindTransferAnchors followed by MapQuery. For Harmony, you will run the integration within a dimensional reduction step. For scDecorr or PoE-VAE, you will train the model on your query and reference together.
Record the parameters you use, including the number of dimensions, the number of anchors or neighbors, and any batch correction variables. These parameters affect the quality of integration and should be reported in your methods.
Step 4: Examine the Integrated Space
After integration, visualize your query cells in the shared embedding alongside reference cells. Use UMAP or t-SNE to project the integrated data. Your query cells should overlap with reference cells of the same type. If your query cells form a separate cluster with no reference overlap, this may indicate a novel cell type or a technical failure of integration.
Check the distribution of prediction scores. Most cells should have high scores for their assigned label. A bimodal distribution of scores, with many cells at intermediate values, suggests that the integration is not cleanly separating cell types.
Step 5: Transfer Labels and Assess Confidence
Transfer the reference labels to your query cells using the integration results. Each cell will receive a predicted label and a score. Set a threshold for confident assignment based on the distribution of scores in your data.
The popV approach provides a more rigorous framework for this step. By running multiple prediction models and using an ontology-based voting scheme, popV produces consensus labels and uncertainty scores. Cells with high uncertainty should be flagged for manual review.
Step 6: Validate Annotations with Independent Evidence
Reference-based annotation is a prediction, not a ground truth. You should validate your annotations using independent evidence. This can include marker gene expression, spatial transcriptomics data, or protein-level measurements.
The bone marrow niche atlas integrated scRNA-seq data with CODEX spatial proteomic imaging to link predicted cellular signaling with spatial proximity. The authors used this integrated approach to annotate new images and uncover mesenchymal stromal cell expansion in acute myeloid leukemia patient samples. This example demonstrates the power of combining transcriptomic annotation with spatial validation.
For immune cell annotation, the review of immune cell annotation challenges emphasizes that mRNA and protein expression can diverge due to post-transcriptional regulation and technical artifacts. Technologies such as CITE-seq, which integrate transcriptomic and proteomic data, address the shortcomings of single-modality analyses. If your validation relies on protein markers, you should be aware that mRNA-based annotations may not perfectly match protein expression.
Step 7: Document and Report
Record the reference atlas version, the integration method and parameters, the confidence thresholds, and the validation results. This documentation is essential for reproducibility and for interpreting your downstream analyses.
The Galaxy Training Network provides accessible workflow training and analysis tutorials that emphasize reproducibility in bioinformatics analysis. Following these standards ensures that your annotation workflow can be reproduced by others in your field.
Records and Measurements for Annotation Quality
Systematic record keeping is essential for evaluating the quality of your reference-based annotation. The following measurements should be recorded for every integration run.
Mapping Statistics
Record the number of query cells that successfully mapped to the reference, the number that failed to map, and the number that mapped with low confidence. These statistics indicate the overall compatibility between your query and the reference.
The benchmarking study of tumor immune atlases evaluated mapping success as one of five criteria for atlas quality. A high mapping failure rate suggests that your query contains cell types or states not represented in the reference, or that technical differences between query and reference are too large for the integration method to correct.
Prediction Score Distributions
For each cell type in your query, record the distribution of prediction scores. Cell types with consistently high scores are likely well-represented in the reference. Cell types with variable or low scores may require manual review.
The popV framework provides uncertainty scores that quantify the confidence of each annotation. Cells with high uncertainty should be prioritized for manual inspection. The authors demonstrated that this approach reduces the load of manual inspection by focusing attention on the most problematic parts of the annotation.
Cluster-Cell Type Agreement
After annotation, compare your unsupervised clustering results with the transferred labels. If cells from a single cluster receive multiple different labels, this may indicate that the cluster contains multiple cell types that were not resolved by clustering alone. Conversely, if cells from multiple clusters receive the same label, this may indicate that the label is too broad or that the clusters represent different states of the same cell type.
The integrated lung endothelial cell atlas identified previously indistinguishable subpopulations among venous and capillary endothelial cells. This finding illustrates that reference-based annotation can resolve cell types that are not apparent from clustering your own data alone.
Marker Gene Validation
For each assigned cell type, check the expression of canonical marker genes in your query cells. This validation step is independent of the integration and provides confidence that the transferred labels are biologically meaningful.
The transcription factor atlas of directed differentiation validated candidate TFs by mapping TF-induced expression profiles to reference cell types. This approach demonstrates the principle that annotation should be confirmed by independent biological evidence.
Common Failure Patterns and Troubleshooting
Reference-based annotation can fail in several predictable ways. Recognizing these failure patterns allows you to diagnose problems and adjust your workflow.
Reference Mismatch
The most common failure is using a reference that does not match your biological context. If you study diseased tissue and use a healthy reference, your disease-associated cell states will be forced into the nearest healthy cell type. The Human Lung Cell Atlas addressed this by including diseased tissues and identifying shared cell states across multiple lung diseases. If your tissue of interest has a disease-specific atlas, you should use it instead of a healthy reference.
Technical Batch Effects Too Large for Correction
Integration methods can correct for moderate technical differences, but extreme differences in sequencing platform, library preparation, or sample quality may prevent successful integration. If your query cells form a separate cluster with no reference overlap, the technical differences may be too large for the integration method to handle.
The scDecorr framework addresses this by learning cell embeddings independently across domains and employing domain-specific batch normalization. If standard integration methods fail, you may need to use a method specifically designed for large technical differences.
Novel Cell Types
If your query contains a cell type that is absent from the reference, the integration method will assign it to the most transcriptionally similar reference type. This assignment may be incorrect and may have high confidence if the novel type is similar to a reference type.
The Brain Cell Atlas identified putative neural progenitor cells and a PCDH9-high microglia subpopulation that were not previously described. These discoveries were possible because the atlas integrated data from many studies, revealing cell types that individual studies missed. If you suspect your query contains novel cell types, you should examine cells with low prediction scores or high uncertainty and compare them to the reference using unsupervised clustering.
Gene Annotation Inconsistencies
If your query and reference use different gene annotation versions, many genes will not match. This can severely degrade integration quality. Always verify that gene identifiers are consistent between query and reference before running integration.
The updated Nematostella reference atlas demonstrated the importance of gene annotation quality. The authors re-mapped sequence data to a new chromosome-level genome assembly with improved gene models, which markedly improved the mapping of single-cell reads. Poor gene annotations can limit the conclusions that can be drawn from single-cell data.
Limitations of Reference-Based Annotation
Reference-based annotation is a powerful tool, but it has inherent limitations that you should understand before relying on it exclusively.
Reference Bias
The annotations you obtain are limited by the cell types and states present in the reference. If the reference does not contain a particular cell type, your query cells of that type will be misannotated. This is not a failure of the method but a fundamental limitation of supervised learning.
The benchmarking study of tumor immune atlases found that existing atlases vary in their data sources, integration and annotation strategies, and the number and definition of cell types. This variability means that different references may produce different annotations for the same query data. You should be aware of this variability and consider using multiple references or an ensemble approach.
mRNA-Protein Discrepancies
Cell type annotation based on mRNA expression may not perfectly reflect protein-level phenotypes. The review of immune cell annotation challenges detailed the mechanisms underlying mRNA-protein discrepancies, including gene expression heterogeneity, post-transcriptional regulation, and technical artifacts. These discrepancies can lead to cell misclassification, particularly in heterogeneous populations such as peripheral blood mononuclear cells.
If your downstream experiments depend on protein-level phenotypes, you should validate your mRNA-based annotations using protein measurements. Technologies such as CITE-seq integrate transcriptomic and proteomic data to address this limitation.
Continuous Cell States
Some biological processes, such as differentiation and activation, produce continuous cell states instead of discrete cell types. Reference-based annotation works best for discrete cell types and may not capture the full spectrum of continuous states.
The transcription factor atlas of directed differentiation addressed this by mapping TF-induced expression profiles to reference cell types and developing a strategy for predicting combinations of TFs that produce target expression profiles matching reference cell types. This approach acknowledges that cell states exist on a continuum and that annotation to discrete types is a simplification.
Welfare and Safety Context for Laboratory Practice
While reference-based annotation is a computational procedure, it has implications for laboratory practice and research integrity that warrant attention.
Data Management and Reproducibility
The Carpentries lessons provide foundational training in computing, data, shell, Git, and programming that is directly relevant to reproducible single-cell analysis. Version control for your analysis scripts and reference atlas versions is essential for reproducibility. The nf-core documentation describes community pipeline standards that emphasize reproducibility and configuration management.
The EMBL-EBI Training program offers bioinformatics learning pathways and data-resource training that can help you develop the computational skills needed for rigorous single-cell analysis. Investing in these skills reduces the risk of analysis errors that could compromise your biological conclusions.
Reporting Standards
When you report reference-based annotation in your publications, you should include the reference atlas version, the integration method and parameters, the confidence thresholds, and the validation results. This information allows other researchers to evaluate the reliability of your annotations and to reproduce your analysis.
The Galaxy Training Network provides accessible workflow training that emphasizes reproducibility in bioinformatics analysis. Following these standards ensures that your annotation workflow can be reproduced by others in your field.
Escalation Criteria
You should escalate to a more experienced bioinformatician or computational biologist when you encounter any of the following situations. First, if your query cells form a separate cluster with no reference overlap after integration, you may be dealing with a novel cell type or a technical failure that requires expert diagnosis. Second, if your prediction scores are uniformly low across all cell types, your reference may be inappropriate for your data. Third, if your validation markers do not match your transferred labels, you may have a systematic annotation error that requires investigation.
The Bioconductor project provides official package, workflow, installation, and reproducible genomic-analysis documentation that can help you troubleshoot integration issues. Consulting this documentation before escalating can resolve many common problems.
Practical Implementation Steps
The following steps provide a concrete implementation plan for integrating your query data with a reference atlas.
Step 1: Inventory Your Data
Document the number of cells, the number of genes, the sequencing platform, the tissue source, and the experimental conditions for your query data. This inventory informs your choice of reference atlas and integration method.
Step 2: Select a Reference
Identify candidate reference atlases for your tissue and species. Evaluate each candidate based on the number of donors, the diversity of conditions, the annotation granularity, and the availability of validation data. The benchmarking criteria from the tumor immune atlas study provide a useful framework for this evaluation.
Step 3: Preprocess Your Query
Run standard quality control, normalization, and feature selection on your query data. Ensure that your gene identifiers match the reference annotation.
Step 4: Run Integration
Execute the integration method appropriate for your data. Record all parameters and the version of the reference atlas.
Step 5: Assess Annotation Quality
Examine the integrated space, the prediction score distributions, and the cluster-cell type agreement. Validate your annotations using marker gene expression and independent evidence.
Step 6: Document and Report
Record all analysis steps, parameters, and quality metrics. Prepare a methods section that allows other researchers to reproduce your annotation workflow.
Frequently Asked Questions
What is the difference between reference-based annotation and unsupervised clustering?
Unsupervised clustering groups cells based on transcriptional similarity without using external information. Reference-based annotation uses a pre-annotated atlas to assign labels to your cells through supervised learning. The benchmarking study of tumor immune atlases found that supervised annotations consistently outperformed unsupervised clustering in identifying cell states related to immunotherapy response. Reference-based annotation is particularly valuable when your dataset contains rare cell types or disease-associated states that may not form distinct clusters in your data alone.
How do I choose between Seurat's FindTransferAnchors and Harmony for integration?
Seurat's anchor-based approach is designed for mapping a query dataset to a reference and transferring labels. Harmony is a general-purpose batch correction method that can integrate multiple datasets without requiring a designated reference. If your primary goal is label transfer from a reference atlas, the anchor-based approach is appropriate. If you need to integrate multiple batches within your query while also aligning to a reference, Harmony may be more suitable. The choice also depends on dataset size, with anchor-based methods being computationally efficient for datasets up to several hundred thousand cells.
What should I do if my query cells do not overlap with any reference cells in the integrated space?
This situation indicates that your query contains cell types or states that are not represented in the reference, or that technical differences between query and reference are too large for the integration method to correct. First, check whether your gene identifiers match the reference annotation. Second, try a different integration method that is more robust to technical differences. Third, examine the non-overlapping cells using unsupervised clustering to determine whether they represent a novel cell type. If they do, you may need to build a custom reference or use a different annotation strategy.
How do I set a confidence threshold for label transfer?
The confidence threshold should be based on the distribution of prediction scores in your data. Examine the score distribution and identify a natural cutoff that separates high-confidence from low-confidence assignments. The popV method provides uncertainty scores that can guide this decision. Cells below the threshold should be flagged for manual review or excluded from downstream analyses that require confident cell type assignments.
Can I use a reference atlas from a different species to annotate my data?
Cross-species annotation is possible when cell types are highly conserved between species. The kidney endothelial cell atlas confirmed that endothelial cell types in the human kidney were highly conserved in the mouse kidney and identified marker genes that were conserved across species. However, cross-species annotation introduces additional technical challenges, including gene name mapping and species-specific expression differences. You should validate cross-species annotations carefully using marker gene expression and other independent evidence.
What is the role of cell type ontologies in reference-based annotation?
Cell type ontologies provide standardized vocabulary for cell type names and their relationships. The Cell Type Ontologies of the Human Cell Atlas provide a framework for consistent annotation across datasets and studies. Using ontology-based labels in your reference and query ensures that annotations are comparable across analyses and can be integrated with other resources. The popV method uses an ontology-based voting scheme to reconcile labels from multiple prediction models.
How do I handle multimodal data in reference-based annotation?
If your data includes multiple modalities, such as RNA and protein expression from CITE-seq, you should use an integration method that can handle multimodal data. The PoE-VAE model integrates multimodal data and allows for the seamless mapping of both unimodal and multimodal queries onto a reference atlas. The multimodal hierarchical classification approach for CITE-seq data delineates immune cell states across lineages and tissues. These methods are appropriate when your data includes modalities beyond RNA expression.
What are the limitations of using a reference atlas for rare cell type identification?
Reference-based annotation can identify rare cell types if they are present in the reference. The Human Lung Cell Atlas identified rare and previously undescribed cell types through atlas-level integration. However, if a rare cell type is absent from the reference, your query cells of that type will be misannotated. Additionally, rare cell types may have low prediction scores because they are represented by few cells in the reference. You should examine cells with low scores or high uncertainty to identify potentially novel or rare populations.
Related Bioinformatics Guides
- Single-Cell Annotation: A Workflow for Cell Type Identification
- Single-Cell RNA-seq Clustering and Cell-Type Annotation Pipelines
- Single-Cell Sequencing Depth: How Much Is Enough?
- Single-Cell Sequencing Workflow: From Sample Preparation to Data Analysis
- RNA-Seq Data Analysis Workflow: From Raw Reads to Insights
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- Mapping the cellular biogeography of human bone marrow niches using single-cell transcriptomics and proteomic imaging.. Cell, 2024.
- A transcription factor atlas of directed differentiation.. Cell, 2023.
- NBAtlas: A harmonized single-cell transcriptomic reference atlas of human neuroblastoma tumors.. Cell reports, 2024.
- Integrated Single-Cell Atlas of Endothelial Cells of the Human Lung.. Circulation, 2021.
- An integrated cell atlas of the lung in health and disease.. Nature medicine, 2023.
- A brain cell atlas integrating single-cell transcriptomes across human brain regions.. Nature medicine, 2024.
- An integrated transcriptomic cell atlas of human neural organoids.. Nature, 2024.
- Integrated Single-Cell Transcriptomic Atlas of Human Kidney Endothelial Cells.. Journal of the American Society of Nephrology : JASN, 2024.
- Benchmarking single-cell tumor immune atlases and application for uncovering cell states related to immunotherapy response.. 2026.
- A multimodal spatial atlas of transcriptomic, morphological, and electrophysiological cell type densities in the mouse brain.. 2026.
- scDecorr: feature decorrelation based representation learning enables self-supervised alignment of multiple single-cell experiments.. 2026.
- A unified single-cell atlas of HNSCC: Toward characterizing HPV- and sex-associated TME variability.. 2026.
- Immune cell annotation in the single-cell studies: technologies, challenges, and integrative solutions.. 2026.
- scExtract: leveraging large language models for fully automated single-cell RNA-seq data annotation and prior-informed multi-dataset integration. Genome Biology, 2025.
- Integration and querying of multimodal single-cell data with PoE-VAE. bioRxiv, 2025.
- Consensus prediction of cell type labels in single-cell data with popV. Nature Genetics, 2024.
- An integrated single-cell reference atlas of the human endometrium. Nature Genetics, 2024.
- Updated single cell reference atlas for the starlet anemone Nematostella vectensis. Frontiers in Zoology, 2024.
- Singletrome enhances detection of long noncoding RNAs in single cell transcriptomes. Scientific Reports, 2025.
- Cell type ontologies of the Human Cell Atlas. Nature Cell Biology, 2021.
- Multimodal hierarchical classification of CITE-seq data delineates immune cell states across lineages and tissues. Cell Reports Methods, 2025.
- Integrating large-scale single-cell RNA sequencing in central nervous system disease using self-supervised contrastive learning. Communications Biology, 2024.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.