# Leveraging Reference Atlases for Single-Cell Cell Type Annotation: A Practical Guide to Mapping Your Data to Known Cell Types


## Key Takeaways

- Reference atlas selection is paramount and must closely align with the query data's tissue, species, developmental stage, and disease context to avoid confidently incorrect annotations. For instance, a pan-cancer T cell atlas is critical for tumor immunology, whereas a healthy adult brain atlas would be inappropriate.
- Label transfer methods offer distinct tradeoffs: integration-based methods (e.g., Symphony) excel at handling batch effects for high accuracy but are computationally intensive, while marker-based scoring (e.g., SAHA) is faster but less sensitive to subtle cell states.
- Annotation granularity must match data resolution; attempting fine-grained subtype annotation with low-sequencing depth or limited cell numbers will yield unreliable labels, necessitating a clear definition of the biological question prior to analysis.
- Mandatory validation of transferred labels is essential, employing at least one independent method such as marker gene inspection, independent clustering, or spatial transcriptomics (e.g., Human Kidney Cell Atlas integration) to confirm biological plausibility.
- Documentation of every annotation decision, including reference atlas version, label transfer parameters, confidence thresholds, and validation results, is critical for reproducibility and transparent scientific reporting.

---

Single-cell RNA sequencing generates transcriptomes for thousands to millions of individual cells, but those measurements carry no intrinsic labels. Cell type annotation is the process of assigning each cell or cluster to a biological identity, and reference atlas mapping has become a standard approach for this task. This article explains how to select appropriate reference atlases, perform label transfer, and validate annotations, with examples from popular atlases such as the Human Cell Atlas and Tabula Muris. The intended reader is a biology student, researcher, laboratory professional, or life-science practitioner who has generated or downloaded single-cell data and needs a defensible workflow for annotation.

Reference atlas mapping works by aligning your query dataset to a well-annotated reference dataset and transferring cell type labels from the reference to your data. The approach is powerful because it leverages curated knowledge from large consortia, but it also carries risks. A reference atlas built from healthy adult tissue may not represent disease states, developmental stages, or species differences. A reference built from one tissue may not generalize to another. The practical question is how to choose a reference, apply it correctly, and verify that the resulting labels are trustworthy.

## At a Glance

The table below summarizes the main decisions you will make when using reference atlases for cell type annotation.

| Decision Point | Primary Options | Key Consideration |
| --- | --- | --- |
| Reference atlas selection | Tissue-specific atlas, pan-tissue atlas, disease-specific atlas, cross-species atlas | Match the reference to your tissue, species, and biological context as closely as possible |
| Label transfer method | Integration-based mapping, marker-based scoring, machine learning classifiers | Integration-based methods handle batch effects but are computationally intensive, marker-based methods are faster but less sensitive |
| Annotation granularity | Broad cell class, fine cell subtype, cell state | Choose the level that matches your biological question and the resolution of your data |
| Validation approach | Marker gene inspection, independent clustering, spatial validation, expert review | Always validate transferred labels with at least one independent method |
| Reporting standard | Cell counts per type, marker genes used, confidence scores, method parameters | Document every annotation decision for reproducibility |

## Understanding Reference Atlases and Their Scope

A reference atlas is an integrated collection of single-cell transcriptomes that has been systematically annotated with cell type labels. The construction of such atlases typically involves aggregating data from multiple studies, applying quality control, integrating datasets to remove batch effects, clustering cells, and assigning cell type identities based on marker genes and expert knowledge. The resulting resource provides a common coordinate system against which new datasets can be mapped.

The scale and scope of reference atlases vary substantially. Some atlases cover a single tissue across many donors, while others integrate data across tissues, developmental stages, or disease states. The Human Cell Atlas initiative aims to create comprehensive reference maps of all human cells, and many tissue-specific and disease-specific atlases have been published under this umbrella. The Brain Cell Atlas, for example, assembled single-cell data from 70 human and 103 mouse studies across major developmental stages and brain regions, covering over 26.3 million cells or nuclei from both healthy and diseased tissues. This atlas provides a consensus cell type annotation and has been used to identify putative neural progenitor cells and a subpopulation of PCDH9-high microglia in the human brain [<a href="#ref-1">1</a>].

Tissue-specific atlases offer a different tradeoff. The integrated single-cell atlas of human lung endothelial cells reprocessed six datasets to identify over 15,000 vascular endothelial cells from 73 individuals. This work revealed previously indistinguishable subpopulations, including pulmonary venous endothelial cells localized to the lung parenchyma and systemic venous endothelial cells localized to the airways and visceral pleura [<a href="#ref-2">2</a>]. A researcher studying lung endothelial biology would find this atlas more useful than a pan-tissue reference because it captures the full diversity of endothelial subtypes in that tissue.

Disease-focused atlases address another gap. The pan-cancer T cell atlas assembled 308,048 transcriptomes across 16 cancer types and identified a unique stress response state characterized by heat shock gene expression. This T cell state was associated with immunotherapy resistance, and the atlas includes a web portal and automatic alignment and annotation tool [<a href="#ref-3">3</a>]. If your study involves tumor-infiltrating T cells, this atlas provides annotations that a healthy-tissue reference would miss.

Model organism atlases extend the approach beyond human biology. The Fly Cell Atlas, also called Tabula Drosophilae, includes 580,000 nuclei from 15 individually dissected sexed tissues as well as the entire head and body of adult Drosophila melanogaster, annotated to more than 250 distinct cell types [<a href="#ref-4">4</a>]. This resource supports genetic perturbation studies and disease modeling in the fruit fly.

The practical implication is that atlas selection should be driven by the biological question. A reference atlas built from the same tissue and species as your query data will almost always outperform a more general atlas. When a disease-specific atlas exists for your condition of interest, it should be considered even if a healthy-tissue atlas is more established.

## Core Principles of Reference-Based Annotation

Reference-based annotation rests on the assumption that similar transcriptional programs indicate similar cell identities. This assumption holds when the reference and query datasets share biological context, but it degrades when the datasets differ in species, tissue, developmental stage, disease state, or technical platform. Understanding these limits is essential for interpreting annotation results.

The first principle is that annotation quality depends on reference quality. A reference atlas is only as good as the annotations it carries. Atlases built through systematic manual annotation with hierarchical taxonomies tend to be more reliable than those assembled through automated pipelines alone. The Uni-TINT atlas, for example, used systematic hierarchical manual annotation to establish a unified taxonomy of T cells and innate lymphoid cells, resolving 207 cell subtypes and states across conventional and unconventional T cells, natural killer cells, helper innate lymphoid cells, thymocytes, and hematopoietic progenitors [<a href="#ref-5">5</a>]. This level of curation supports confident label transfer.

The second principle is that integration quality determines mapping accuracy. Label transfer methods align query cells to reference cells in a shared embedding space, and the quality of that alignment depends on how well batch effects are removed while biological variation is preserved. Methods that use multiple adversarial domain adaptation networks have been developed specifically for assigning cells using large-scale references [<a href="#ref-6">6</a>]. These approaches learn to remove dataset-specific variation while retaining cell type-specific signals.

The third principle is that annotation granularity must match the data resolution. A reference atlas may distinguish 200 cell types, but your query data may only support annotation at the level of broad cell classes if sequencing depth is low or cell numbers are limited. Attempting fine-grained annotation with insufficient data produces unreliable labels. The reverse problem also occurs: a reference that only distinguishes broad categories will not help you identify rare subtypes.

The fourth principle is that validation is mandatory. Transferred labels are predictions, not facts. Every annotation should be checked against marker gene expression, independent clustering results, or spatial localization data when available. The Human Kidney Cell Atlas, for example, integrated single-cell and single-nucleus RNA sequencing data with spatial transcriptomics to define the anatomical context of kidney cell populations [<a href="#ref-7">7</a>]. This kind of multimodal validation strengthens confidence in annotations.

## Selecting the Right Reference Atlas for Your Data

The selection of a reference atlas is the most consequential decision in the annotation workflow. A mismatched reference produces confidently wrong labels that are difficult to detect without careful validation. The following criteria should guide your choice.

### Tissue and Organ Match

The reference atlas should be built from the same tissue or organ as your query data whenever possible. Tissue-specific atlases capture the full repertoire of cell types present in that tissue, including rare populations that a pan-tissue atlas may miss. The human bone marrow atlas profiled 29,325 non-hematopoietic cells and discovered nine transcriptionally distinct subtypes, then integrated this data with spatial profiling of over 1.2 million cells to link cellular signaling with spatial proximity [<a href="#ref-8">8</a>]. A researcher studying bone marrow niches would need this tissue-specific resolution.

Cross-tissue atlases have their place. The Brain Cell Atlas integrates data across multiple brain regions and developmental stages, making it useful for studies that span these dimensions [<a href="#ref-1">1</a>]. The pan-cancer T cell atlas covers 16 cancer types and provides a unified framework for immune annotation across diseases [<a href="#ref-3">3</a>]. These resources are valuable when your study crosses traditional tissue boundaries.

### Species Compatibility

Reference atlases are typically built from a single species, and cross-species label transfer is challenging. While some cell types are conserved across species, the transcriptional programs that define them diverge. The Fly Cell Atlas supports Drosophila research but would not be appropriate for annotating mammalian data [<a href="#ref-4">4</a>]. Conversely, human atlases should not be used to annotate mouse data without careful consideration of orthologous gene mapping and known species differences.

Some atlases explicitly address cross-species questions. The Brain Cell Atlas includes both human and mouse data, allowing researchers to compare cell types across species [<a href="#ref-1">1</a>]. The human kidney atlas provides a framework for automated annotation of independent kidney datasets from both human and mouse [<a href="#ref-7">7</a>]. If your study requires cross-species comparison, look for atlases that were built with this purpose in mind.

### Disease Context

Healthy reference atlases may not represent disease-associated cell states. The pan-cancer T cell atlas identified a stress response state characterized by heat shock gene expression that was associated with immunotherapy resistance [<a href="#ref-3">3</a>]. This state would not be present in a healthy-tissue reference. If your study involves diseased tissue, search for atlases built from similar disease contexts.

The neuroblastoma atlas integrated seven single-cell or single-nucleus datasets into a harmonized atlas covering 362,991 cells across 61 patients [<a href="#ref-9">9</a>]. This disease-specific resource revealed associations between transcriptomic profiles and clinical outcomes within the tumor compartment. A researcher studying neuroblastoma would find this atlas more informative than a general developmental atlas.

### Developmental Stage

Cell type composition changes dramatically during development. An atlas built from adult tissue will not capture fetal or neonatal cell states. The Brain Cell Atlas spans major developmental stages, making it suitable for studies of brain development [<a href="#ref-1">1</a>]. If your data comes from a developmental time point, verify that your chosen reference includes that stage.

### Technical Platform

Single-cell and single-nucleus RNA sequencing produce different transcriptomic profiles. Single-nucleus data captures nuclear transcripts and is often used for frozen or hard-to-dissociate tissues, while single-cell data captures whole-cell transcripts. The human kidney atlas included both single-cell and single-nucleus datasets, with 816,895 cells and 215,308 nuclei, and harmonized them through uniform processing [<a href="#ref-7">7</a>]. If your data was generated with a specific platform, check whether your reference atlas includes data from that platform.

## Label Transfer Methods and Their Tradeoffs

Once you have selected a reference atlas, you must choose a method for transferring labels from the reference to your query data. The main approaches fall into three categories: integration-based mapping, marker-based scoring, and machine learning classification.

### Integration-Based Mapping

Integration-based methods embed query and reference cells in a shared space and transfer labels based on neighborhood relationships. These methods are powerful because they learn a nonlinear alignment that accounts for batch effects and technical differences. Symphony is one such method designed for efficient and precise single-cell reference atlas mapping [<a href="#ref-10">10</a>]. It builds a reference from annotated data and maps query cells onto that reference in a way that preserves biological variation while removing technical variation.

The main advantage of integration-based methods is accuracy. They can transfer fine-grained labels when the reference is well-annotated and the query data is of reasonable quality. The main disadvantage is computational cost. Integration of large datasets requires substantial memory and processing time, and the methods may be sensitive to parameter choices.

### Marker-Based Scoring

Marker-based methods score each query cell or cluster based on the expression of known marker genes for each reference cell type. These methods are fast, interpretable, and do not require the reference data itself, only the marker gene lists. The Semi-Automated Hand Annotation framework, called SAHA, uses summary statistics and user-defined hyperparameters to investigate the magnitude, directionality, and statistical significance of matches between unnamed query clusters and reference cell types [<a href="#ref-11">11</a>]. This approach works with either marker-based or marker-free methodologies and can compare across omic modalities and cluster resolutions.

The main advantage of marker-based methods is speed and simplicity. They can be applied to datasets of any size and do not require access to the raw reference data. The main disadvantage is sensitivity. Marker-based methods may miss cell types that lack well-defined markers or that are defined by combinations of genes instead of individual markers.

### Machine Learning Classification

Machine learning classifiers are trained on reference data to predict cell type labels for query cells. These methods can capture complex, nonlinear relationships between gene expression and cell identity. The scPANDA tool uses a three-layer inference approach that progressively refines cell types from broad compartments to specific clusters, leveraging a 10-million-cell blood atlas [<a href="#ref-12">12</a>]. The atlas is structured hierarchically with 16 compartments, 54 classes, and hundreds of high-level clusters.

The main advantage of machine learning methods is scalability. They can annotate millions of cells efficiently once trained. The main disadvantage is that they require a high-quality training set and may not generalize to cell types or states not represented in the reference.

### Practical Considerations for Method Choice

The choice of method depends on your data, your reference, and your computational resources. For a small dataset with a well-matched reference, integration-based mapping often provides the best accuracy. For a large dataset where computational cost is a concern, marker-based scoring or machine learning classification may be more practical. For a dataset that may contain novel cell states, marker-based methods that allow for unknown assignments are preferable to methods that force every cell into a known category.

The MAT-Cell framework addresses the problem of annotation when a target state is poorly covered by a fixed reference atlas [<a href="#ref-13">13</a>]. It separates evidence grounding from label decision, using a reverse verification query to combine tissue context, observed differentially expressed genes, and biological priors into structured candidate-specific premises. This approach is designed for batch-level annotation and provides an auditable debate trace instead of a formal proof of the annotation.

## Practical Workflow for Reference-Based Annotation

The following workflow outlines the steps for applying a reference atlas to your own single-cell data. The workflow assumes you have already performed basic quality control on your query data, including filtering low-quality cells, removing doublets, and normalizing expression values.

### Step 1: Define Your Annotation Goal

Before selecting a reference or method, define what you need from the annotation. Are you looking for broad cell classes such as T cells, B cells, and macrophages? Or do you need fine subtypes such as CD4-positive naive T cells versus CD4-positive memory T cells? The answer determines the required resolution of your reference and the granularity of your labels.

### Step 2: Select and Download a Reference Atlas

Based on the criteria in the previous section, identify one or more candidate reference atlases. Download the reference data, including the expression matrix, cell metadata, and cell type labels. Verify that the reference includes the genes you need for your analysis. Some atlases provide processed data in standard formats that can be loaded directly into analysis environments such as Bioconductor [<a href="#ref-14">14</a>].

### Step 3: Prepare Your Query Data

Ensure your query data is properly normalized and that gene symbols match the reference. If your data uses gene identifiers that differ from the reference, convert them before proceeding. Filter your query data to remove low-quality cells that could produce spurious annotations.

### Step 4: Run Label Transfer

Apply your chosen label transfer method. For integration-based methods, this step involves aligning your query data to the reference embedding. For marker-based methods, this step involves scoring each query cell or cluster against reference marker genes. For machine learning methods, this step involves applying a trained classifier to your query data.

### Step 5: Examine Confidence Scores

Most label transfer methods produce confidence scores for each assignment. Examine the distribution of these scores. Low-confidence assignments may indicate cells that do not match any reference cell type, possibly representing novel states or poor-quality cells. Decide on a threshold for accepting or rejecting annotations based on your biological question and the quality of your data.

### Step 6: Validate Annotations with Marker Genes

For each annotated cell type, examine the expression of known marker genes. This validation step is essential. If your transferred labels are correct, the marker genes should show the expected expression patterns. The Human Reference Atlas population effort, called HRApop, computed cell types and biomarker expressions for over 11 million cells from 558 single-cell transcriptomics datasets using multiple annotation tools [<a href="#ref-15">15</a>]. This kind of systematic validation builds confidence in the reference annotations themselves.

### Step 7: Validate with Independent Methods

If possible, validate your annotations with an independent method. This could involve re-clustering your data without reference information and comparing the resulting clusters to your transferred labels. It could also involve spatial transcriptomics data if available, as was done in the human bone marrow atlas where single-cell RNA sequencing data was integrated with spatial profiling to link predicted cellular signaling with spatial proximity [<a href="#ref-8">8</a>].

### Step 8: Document Your Annotation

Record every decision you made during the annotation process, including the reference atlas version, the label transfer method and parameters, the confidence threshold, and the validation results. This documentation is essential for reproducibility and for interpreting downstream analyses.

## Records and Measurements for Annotation Quality

Quantitative records of annotation quality support both scientific rigor and troubleshooting. The following measurements should be recorded for every annotation run.

### Cell Counts per Type

Record the number of cells assigned to each cell type. Unexpected distributions may indicate problems with the reference or the label transfer. For example, a complete absence of a cell type that should be present in your tissue may indicate that the reference does not capture that population or that your data preparation removed it.

### Confidence Score Distributions

Record the distribution of confidence scores for each cell type. A bimodal distribution may indicate that some cells of a given type are well-matched to the reference while others are poorly matched, possibly representing a distinct subtype or state.

### Marker Gene Expression Summary

For each annotated cell type, record the mean expression of canonical marker genes. This summary provides a quick check that annotations are biologically plausible.

### Integration Metrics

If you used an integration-based method, record metrics that describe the quality of the alignment between query and reference cells. These metrics may include the proportion of query cells that map to each reference cell type and the mixing of query and reference cells in the shared embedding.

### Computational Performance

Record the computational time and memory usage for the label transfer step. This information helps you plan for future analyses and may indicate when a different method would be more practical.

## Common Failure Patterns and Troubleshooting

Reference-based annotation can fail in predictable ways. Recognizing these failure patterns helps you diagnose problems and choose appropriate corrections.

### Reference Mismatch

The most common failure is a mismatch between the reference and query data. This can occur when the reference was built from a different tissue, species, developmental stage, or disease context than the query. Symptoms include low confidence scores across many cells, unexpected cell type distributions, and poor marker gene validation. The solution is to select a more appropriate reference or to combine multiple references.

### Batch Effects

Technical differences between the query and reference datasets can dominate biological signals, causing cells to cluster by dataset instead of by cell type. This problem is especially severe when the query and reference were generated with different protocols or sequencing platforms. Integration-based methods are designed to address this issue, but they require sufficient overlap in cell types between datasets.

### Novel Cell States

If your query data contains cell states that are not represented in the reference, the label transfer will force those cells into the nearest known category. This produces confident but incorrect labels. The MAT-Cell framework addresses this problem by allowing for structured reasoning about candidate-specific premises [<a href="#ref-13">13</a>], and the SAHA framework allows users to investigate the statistical significance of matches between query clusters and reference cell types [<a href="#ref-11">11</a>]. If you suspect novel cell states, use a method that can flag low-confidence assignments instead of forcing every cell into a known category.

### Reference Annotation Errors

Reference atlases are not infallible. Annotation errors in the reference propagate to your query data. The Uni-TINT atlas was built specifically to address inconsistent annotation across studies, highlighting that this is a recognized problem in the field [<a href="#ref-5">5</a>]. When possible, use atlases that have been built through systematic manual annotation and validated with independent methods.

### Gene Symbol Mismatches

Differences in gene nomenclature between the query and reference can cause spurious results. Always verify that gene symbols match between datasets before running label transfer. This is a simple but critical step that is often overlooked.

### Low-Quality Query Cells

Cells with low sequencing depth or high mitochondrial content may not produce reliable annotations. These cells may map to unexpected reference cell types or produce low confidence scores. Filtering these cells before annotation improves the reliability of the results.

## Limitations of Reference-Based Annotation

Reference-based annotation has inherent limitations that should be acknowledged in any study that uses this approach.

### Resolution Limits

The resolution of your annotation is limited by the resolution of the reference. If the reference distinguishes 50 cell types, you cannot annotate 200 cell types. The Human Reference Atlas v2.3 provides a quantitative 3D framework of cell types across 73 reference organs and 1,283 anatomical structures, with the HRApop effort quantifying cell types per anatomical structure using high-quality single-cell datasets [<a href="#ref-15">15</a>]. This resource provides a healthy reference for researchers, but it may not capture disease-specific states.

### Dependency on Reference Quality

The accuracy of your annotations depends entirely on the quality of the reference annotations. If the reference contains errors, your results will contain those same errors. The pan-cancer T cell atlas was built to address the problem of inconsistent annotation across studies [<a href="#ref-3">3</a>], and the Uni-TINT atlas resolved 207 cell subtypes and states through systematic hierarchical manual annotation [<a href="#ref-5">5</a>]. These efforts improve the field, but they do not eliminate the need for validation.

### Inability to Discover Novel Cell Types

Reference-based annotation is inherently conservative. It can only assign cells to categories that exist in the reference. If your data contains a genuinely novel cell type, reference-based annotation will either miss it or misclassify it. Discovery of novel cell types requires unsupervised analysis or methods specifically designed to flag cells that do not match any reference category.

### Computational Requirements

Large reference atlases and large query datasets require substantial computational resources. The scPANDA atlas contains 10 million cells [<a href="#ref-12">12</a>], and the Brain Cell Atlas covers over 26.3 million cells or nuclei [<a href="#ref-1">1</a>]. Mapping a large query dataset to these references requires significant memory and processing time. Cloud computing or high-performance computing resources may be necessary.

### Privacy and Data Sharing Concerns

Some web-based annotation tools require uploading raw data to external servers. This raises privacy concerns for sensitive datasets. The SAHA framework explicitly avoids this problem by using summary statistics instead of raw data, allowing users to perform annotation locally without sharing data [<a href="#ref-11">11</a>]. If data privacy is a concern, choose methods that can run locally.

## Quality Controls and Validation Strategies

Validation is the most important step in reference-based annotation. The following strategies provide independent evidence that your annotations are correct.

### Marker Gene Inspection

For each annotated cell type, examine the expression of canonical marker genes. This is the simplest and most direct validation. If your transferred labels are correct, the expected markers should be expressed in the expected cells. The lung endothelial cell atlas validated marker genes through fluorescent microscopy and in situ hybridization, providing strong evidence for the annotation of endothelial subpopulations [<a href="#ref-2">2</a>].

### Independent Clustering

Cluster your query data without reference information and compare the resulting clusters to your transferred labels. If the two approaches agree, your annotations are likely correct. If they disagree, investigate the source of the discrepancy. This approach is particularly useful for identifying cells that were forced into incorrect categories by the reference.

### Spatial Validation

If spatial transcriptomics data is available for your tissue, use it to validate your annotations. Cell types should localize to expected anatomical regions. The human bone marrow atlas used this approach, integrating single-cell RNA sequencing data with spatial profiling of over 1.2 million cells to link predicted cellular signaling with spatial proximity [<a href="#ref-8">8</a>]. The kidney atlas also integrated spatial transcriptomics to define the anatomical context of kidney cell populations [<a href="#ref-7">7</a>].

### Cross-Method Validation

Use multiple annotation methods and compare their results. The HRApop effort used Azimuth, CellTypist, and popV to compute cell types and biomarker expressions for over 11 million cells [<a href="#ref-15">15</a>]. Agreement between methods increases confidence in the annotations. Disagreement indicates that at least one method is producing unreliable results.

### Expert Review

If possible, have an expert in the relevant tissue or cell biology review your annotations. Expert knowledge can catch errors that automated methods miss. This is especially important for rare cell types or cell states that are difficult to distinguish based on markers alone.

## Professional Escalation Criteria

Some annotation problems require consultation with a bioinformatics specialist or the developers of the reference atlas. The following situations warrant escalation.

### Persistent Low Confidence Scores

If a substantial proportion of your cells have low confidence scores across multiple methods, the reference may not be appropriate for your data. Consult with a specialist to identify a more suitable reference or to determine whether your data contains novel cell states.

### Unexpected Cell Type Distributions

If your cell type distribution is dramatically different from what is expected for your tissue, investigate before proceeding. This could indicate a reference mismatch, a data quality problem, or a genuine biological finding. A specialist can help you distinguish between these possibilities.

### Disagreement Between Validation Methods

If your transferred labels disagree with independent clustering or marker gene validation, do not proceed with downstream analysis until the discrepancy is resolved. A specialist can help you determine which result is more reliable.

### Novel Cell State Candidates

If you suspect that your data contains cell types or states not represented in the reference, consult with a specialist before making claims. Novel cell type discovery requires rigorous validation, including spatial localization and functional studies.

### Computational Resource Limitations

If your computational resources are insufficient for the label transfer step, consult with a specialist about alternative methods or cloud computing options. Do not compromise on data quality or annotation accuracy to reduce computational cost.

## Safety and Reproducibility Context

Reproducibility is a core requirement for scientific research, and reference-based annotation introduces several reproducibility challenges. The reference atlas version must be recorded, as atlases are updated over time. The label transfer method and all parameters must be documented. The computational environment, including software versions, must be preserved.

Training resources are available for researchers who need to build these skills. The Galaxy Training Network provides accessible workflow training and analysis tutorials that cover single-cell analysis [<a href="#ref-16">16</a>]. The Carpentries offers foundational computing and data lessons that support reproducible research practices [<a href="#ref-17">17</a>]. The EMBL-EBI Training program provides bioinformatics learning pathways and data-resource training [<a href="#ref-18">18</a>]. Bioconductor provides official package, workflow, installation, and reproducible genomic-analysis documentation [<a href="#ref-14">14</a>]. The nf-core documentation describes community pipeline standards for reproducible workflows [<a href="#ref-19">19</a>]. The NCBI provides official descriptions of databases, search systems, sequence resources, and analysis services that support data access and deposition [<a href="#ref-20">20</a>].

These resources support the practical implementation of reference-based annotation in a reproducible manner. They do not replace the need for careful experimental design and biological interpretation, but they provide the technical foundation for rigorous analysis.

## Frequently Asked Questions

### What is the difference between single-cell and single-nucleus RNA sequencing for reference-based annotation?

Single-cell RNA sequencing captures whole-cell transcripts and is typically used for fresh or easily dissociable tissues. Single-nucleus RNA sequencing captures nuclear transcripts and is used for frozen tissues or tissues that are difficult to dissociate. The two platforms produce different transcriptomic profiles, and reference atlases may be built from one or both platforms. The human kidney atlas included both single-cell and single-nucleus datasets and harmonized them through uniform processing [<a href="#ref-7">7</a>]. When selecting a reference atlas, check whether it includes data from the same platform as your query data.

### How do I know if my reference atlas is appropriate for my data?

The reference atlas should match your data in tissue, species, developmental stage, and disease context. Check the atlas documentation to verify these matches. After running label transfer, examine confidence scores and validate annotations with marker genes. Low confidence scores or poor marker validation may indicate a reference mismatch.

### What should I do if my data contains cell types not present in the reference atlas?

Cells that do not match any reference cell type will produce low confidence scores or be forced into the nearest known category. Use a method that can flag low-confidence assignments, such as SAHA, which allows you to investigate the statistical significance of matches between query clusters and reference cell types [<a href="#ref-11">11</a>]. If you suspect novel cell types, perform unsupervised clustering and characterize the novel populations with marker genes and spatial validation.

### How many cells do I need for reliable reference-based annotation?

There is no universal minimum cell number. The required number depends on the complexity of your tissue, the granularity of your annotation, and the quality of your data. Rare cell types require more cells to be detected reliably. The neuroblastoma atlas integrated 362,991 cells across 61 patients to characterize the transcriptional landscape of neuroblastoma [<a href="#ref-9">9</a>], while the lung endothelial atlas identified over 15,000 vascular endothelial cells from 73 individuals [<a href="#ref-2">2</a>]. Smaller datasets can still be annotated, but confidence in rare cell type annotations will be limited.

### Can I use a reference atlas from one species to annotate data from another species?

Cross-species label transfer is challenging because transcriptional programs diverge between species. Some atlases include multiple species, such as the Brain Cell Atlas which includes both human and mouse data [<a href="#ref-1">1</a>]. The human kidney atlas provides a framework for automated annotation of independent kidney datasets from both human and mouse [<a href="#ref-7">7</a>]. If you must use a cross-species reference, map genes to orthologs and interpret results with caution.

### What is the difference between cell type and cell state annotation?

Cell type refers to a stable identity, such as a T cell or a hepatocyte. Cell state refers to a transient condition, such as activation, proliferation, or stress. Some atlases annotate both types and states. The pan-cancer T cell atlas identified a stress response state characterized by heat shock gene expression, which is a cell state instead of a distinct cell type [<a href="#ref-3">3</a>]. The Uni-TINT atlas resolved 207 cell subtypes and states, distinguishing between stable identities and transient conditions [<a href="#ref-5">5</a>].

### How should I report reference-based annotation results in my publication?

Report the reference atlas version, the label transfer method and all parameters, the confidence threshold, and the validation results. Include cell counts per type and marker gene expression summaries. This documentation allows other researchers to reproduce your analysis and assess the reliability of your annotations.

### What are the main alternatives to reference-based annotation?

The main alternatives are unsupervised clustering followed by manual annotation, and marker-based annotation without a formal reference. Unsupervised clustering identifies groups of cells with similar expression profiles, which you then annotate based on marker genes. This approach can discover novel cell types but requires substantial manual effort. Marker-based annotation scores cells based on known marker genes without requiring a full reference atlas. This approach is faster but less comprehensive than reference-based annotation.

## Related Bioinformatics Guides

- [Single-Cell Annotation: A Workflow for Cell Type Identification](/knowledge/bioinformatics/single-cell-annotation-a-workflow-for-cell-type-identification)
- [Spatial Proteomics vs. Single-Cell Proteomics: Choosing the Right Approach](/knowledge/bioinformatics/spatial-proteomics-vs-single-cell-proteomics-choosing-the-right-approach)
- [Single-Cell vs Single-Nucleus RNA Sequencing: Choosing the Right Approach](/knowledge/bioinformatics/single-cell-vs-single-nucleus-rna-sequencing-choosing-the-right-approach)
- [Benchmarking Atlas-Level Data Integration in Single-Cell Genomics: Methods and Best Practices](/knowledge/bioinformatics/benchmarking-atlas-level-data-integration-in-single-cell-genomics-methods-and-best-practices)
- [Single-Cell Isolation Techniques: A Practical Comparison](/knowledge/bioinformatics/single-cell-isolation-techniques-a-practical-comparison)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)

## References and Further Reading

<a id="ref-1"></a>[<a href="#ref-1">1</a>] [A brain cell atlas integrating single-cell transcriptomes across human brain regions.](https://pubmed.ncbi.nlm.nih.gov/39095595). Nature medicine, 2024.

<a id="ref-2"></a>[<a href="#ref-2">2</a>] [Integrated Single-Cell Atlas of Endothelial Cells of the Human Lung.](https://pubmed.ncbi.nlm.nih.gov/34030460). Circulation, 2021.

<a id="ref-3"></a>[<a href="#ref-3">3</a>] [Pan-cancer T cell atlas links a cellular stress response state to immunotherapy resistance.](https://pubmed.ncbi.nlm.nih.gov/37248301). Nature medicine, 2023.

<a id="ref-4"></a>[<a href="#ref-4">4</a>] [Fly Cell Atlas: A single-nucleus transcriptomic atlas of the adult fruit fly.](https://pubmed.ncbi.nlm.nih.gov/35239393). Science (New York, N.Y.), 2022.

<a id="ref-5"></a>[<a href="#ref-5">5</a>] [Uni-TINT: Unifying T and Innate Lymphoid Cell Taxonomy with a high resolution, pan-disease and pan-tissue single-cell transcriptomic atlas](https://doi.org/10.64898/2026.07.20.739448). 2026.

<a id="ref-6"></a>[<a href="#ref-6">6</a>] [Single-cell assignment using multiple-adversarial domain adaptation network with large-scale references](https://doi.org/10.1016/j.crmeth.2023.100577). Cell Reports Methods, 2023.

<a id="ref-7"></a>[<a href="#ref-7">7</a>] [An Integrated Atlas of the Human Kidney Spanning Health and Diseases](https://doi.org/10.64898/2026.08.12.744548). 2026.

<a id="ref-8"></a>[<a href="#ref-8">8</a>] [Mapping the cellular biogeography of human bone marrow niches using single-cell transcriptomics and proteomic imaging.](https://pubmed.ncbi.nlm.nih.gov/38714197). Cell, 2024.

<a id="ref-9"></a>[<a href="#ref-9">9</a>] [NBAtlas: A harmonized single-cell transcriptomic reference atlas of human neuroblastoma tumors.](https://pubmed.ncbi.nlm.nih.gov/39368085). Cell reports, 2024.

<a id="ref-10"></a>[<a href="#ref-10">10</a>] [Efficient and precise single-cell reference atlas mapping with Symphony](https://doi.org/10.1038/s41467-021-25957-x). Nature Communications, 2021.

<a id="ref-11"></a>[<a href="#ref-11">11</a>] [Semi-automated annotation refinement accelerates cell type identification in brain spatial and single-cell studies](https://doi.org/10.64898/2026.07.30.741795). 2026.

<a id="ref-12"></a>[<a href="#ref-12">12</a>] [scPANDA: PAN-Blood Data Annotator with a 10-Million Single-Cell Atlas.](https://doi.org/10.24920/004472). Chinese medical sciences journal = Chung-kuo i hsueh k'o hsueh tsa chih, 2025.

<a id="ref-13"></a>[<a href="#ref-13">13</a>] [MAT-Cell: A Multi-Agent Tree-Structured Reasoning Framework for Batch-Level Single-Cell Annotation](https://doi.org/10.48550/arXiv.2604.06269). arXiv.org, 2026.

<a id="ref-14"></a>[<a href="#ref-14">14</a>] [Bioconductor](https://bioconductor.org/). Bioconductor Project.

<a id="ref-15"></a>[<a href="#ref-15">15</a>] [Cell Type Populations for 3D Anatomical Structures of the Human Reference Atlas](https://doi.org/10.1038/s41597-026-06642-4). bioRxiv, 2025.

<a id="ref-16"></a>[<a href="#ref-16">16</a>] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.

<a id="ref-17"></a>[<a href="#ref-17">17</a>] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.

<a id="ref-18"></a>[<a href="#ref-18">18</a>] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.

<a id="ref-19"></a>[<a href="#ref-19">19</a>] [nf-core Documentation](https://nf-co.re/docs). nf-core.

<a id="ref-20"></a>[<a href="#ref-20">20</a>] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.