# A Beginner's Guide to Cell Type Annotation in Single-Nucleus RNA-Seq: Key Differences and Best Practices


## Key Takeaways

- Single-nucleus RNA sequencing (snRNA-seq) captures nuclear RNA, including pre-mRNA, leading to different gene detection sensitivity and marker gene performance compared to single-cell RNA sequencing (scRNA-seq), which profiles cytoplasmic and nuclear RNA. This necessitates validation of marker panels specifically for snRNA-seq data.
- Quality control for snRNA-seq requires reinterpretation of standard metrics; a high mitochondrial RNA fraction often indicates cytoplasmic contamination rather than cell stress, and lower gene detection per nucleus is expected and should not be aggressively filtered.
- Ambient RNA contamination is a distinct problem in snRNA-seq, particularly in tissues with highly expressed neuronal genes, potentially masking rare cell types or creating false signals that require computational correction (e.g., using CellBender) and careful examination of cluster marker gene co-expression.
- snRNA-seq is crucial for analyzing frozen tissues, difficult-to-dissociate samples, and large or fragile cell types, enabling studies on archived specimens and challenging biological systems where scRNA-seq is not feasible.
- Annotation workflows must incorporate multiple lines of evidence, including canonical marker expression, differential gene expression analysis, and reference-based label transfer, with rigorous validation using independent methods like spatial transcriptomics or matched snATAC-seq data.
- A structured annotation audit system, including a decision gate framework, standardized record-keeping (sample log, QC log, annotation worksheet, incident reports), and a systematic troubleshooting protocol, is essential for ensuring the trustworthiness and reproducibility of snRNA-seq cell type annotations.

---

Cell type annotation is the process of assigning biological identity to clusters of nuclei in single-nucleus RNA sequencing (snRNA-seq) data. For researchers new to this field, the central problem is that snRNA-seq data differ from single-cell RNA sequencing (scRNA-seq) data in systematic ways that affect every step of annotation, from quality control to marker selection to final cluster labeling. This article explains those differences and provides a practical workflow for accurate annotation.

Single-nucleus RNA sequencing profiles the transcriptome of individual nuclei instead of whole cells. This approach is particularly valuable for frozen tissues, difficult-to-dissociate samples, and large or fragile cell types that do not survive standard single-cell dissociation protocols. The method has been applied across diverse fields, including Alzheimer's disease research, heart failure studies, cancer biopsies, plant pathology, and model organism atlases. Each application shares the same fundamental annotation challenge: converting raw sequencing data into biologically meaningful cell type labels.

The practical outcome of this article is a decision framework that helps you choose appropriate tools, interpret quality metrics, validate your annotations, and avoid common pitfalls specific to snRNA-seq data.

## At a Glance: Key Differences Between scRNA-seq and snRNA-seq Annotation

| Feature | scRNA-seq | snRNA-seq | Annotation Impact |
|---------|-----------|-----------|-------------------|
| Input material | Fresh or minimally preserved cells | Frozen tissue, nuclei isolated from cells | Enables profiling of archived and difficult samples |
| RNA content | Cytoplasmic and nuclear RNA | Primarily nuclear RNA, including pre-mRNA | Lower gene detection per nucleus, different marker sensitivity |
| Dissociation requirement | Enzymatic or mechanical cell dissociation | Nuclear isolation, often with sucrose cushion or FANS | Reduces dissociation-induced transcriptional artifacts |
| Ambient RNA risk | Present from lysed cells | Present, with specific patterns from highly expressed neuronal genes | Can mask rare cell types and create false signals |
| Cross-dataset transfer | Established reference atlases available | Growing but fewer dedicated references | Requires careful label transfer validation |

## Understanding Why snRNA-seq Annotation Differs from scRNA-seq

### The Biological Basis of Nuclear Transcriptomes

The transcriptome of a nucleus is not identical to the transcriptome of the whole cell. Mature messenger RNA is exported to the cytoplasm, while the nucleus retains pre-mRNA, unspliced transcripts, and other nuclear RNA species. When you sequence nuclei instead of whole cells, you capture a different RNA pool. This means that genes with predominantly cytoplasmic mRNA may show reduced detection in snRNA-seq data, while genes with nuclear retention or high transcriptional activity may be overrepresented.

For annotation purposes, this difference matters because marker genes are often defined based on scRNA-seq experiments. A marker that works well for identifying a cell type in scRNA-seq data may show weaker or absent signal in snRNA-seq data if that marker's mRNA is primarily cytoplasmic. Conversely, some genes that are not useful markers in scRNA-seq data may become informative in snRNA-seq data because of their nuclear retention patterns.

The practical consequence is that you cannot assume a marker panel validated for scRNA-seq will transfer directly to snRNA-seq data. You need to validate marker expression in your own data and consult snRNA-seq-specific references when available.

### Tissue Types That Favor snRNA-seq

Single-nucleus approaches excel in specific contexts. Frozen tissue samples, including long-term archived specimens, can be processed for nuclear isolation when fresh tissue is unavailable. This capability has enabled studies of rare tumor entities using frozen pediatric glioma tissues, where a simplified nuclear isolation method allowed identification of distinct tumor cell populations and infiltrating microglia [<a href="#ref-1">1</a>]. Similarly, formalin-fixed paraffin-embedded (FFPE) tissues, which are the standard for clinical pathology archives, have become accessible to single-nucleus profiling through specialized protocols [<a href="#ref-2">2</a>][<a href="#ref-3">3</a>].

Plant tissues present another case where nuclear profiling is often preferred. Plant cells are surrounded by rigid cell walls that resist standard dissociation, and the authors of a tobacco stem snRNA-seq dataset explicitly noted that nucleus-based transcriptomic profiling provides a practical way to generate cell-resolved expression data from these samples [<a href="#ref-4">4</a>]. Heart tissue, brain tissue, and other organs with large or fragile cells similarly benefit from nuclear isolation [<a href="#ref-5">5</a>][<a href="#ref-6">6</a>].

When you choose snRNA-seq, you gain access to sample types that scRNA-seq cannot accommodate. The tradeoff is that your annotation workflow must account for the technical differences described throughout this article.

## Core Principles of snRNA-seq Cell Type Annotation

### Quality Control Starts with Nuclei, Not Cells

Quality control for snRNA-seq data begins with understanding what a good nucleus looks like in your data. The standard metrics used in scRNA-seq, including total UMI counts, number of detected genes, and mitochondrial read fraction, require reinterpretation for nuclear data.

Mitochondrial RNA is largely absent from nuclear preparations because mitochondria are cytoplasmic organelles. A high mitochondrial fraction in snRNA-seq data may indicate cytoplasmic contamination instead of cell stress, which is the usual interpretation in scRNA-seq data. This distinction changes how you set thresholds and interpret outliers.

The number of genes detected per nucleus is typically lower than per cell in scRNA-seq data from the same tissue. This reduced sensitivity means that some cell types with low transcriptional activity may be harder to resolve. Your quality control thresholds should be set based on the distribution of your own data instead of borrowed from scRNA-seq studies.

The authors of a heart failure study using snRNA-seq described a workflow that included nuclei isolation from pooled samples, genotype-based demultiplexing, and gene expression quantification before clustering and annotation [<a href="#ref-5">5</a>]. This workflow illustrates the importance of establishing sample identity and data quality before attempting cell type assignment.

### Ambient RNA Is a Distinct Problem in snRNA-seq

Ambient RNA contamination occurs when RNA from lysed or damaged cells is captured in droplets or wells alongside the intended nucleus. In snRNA-seq data from brain tissue, this contamination has been shown to cause misinterpreted and masked cell types [<a href="#ref-7">7</a>]. Highly expressed genes from abundant cell types can create spurious signals in other clusters, leading to incorrect annotations.

The problem is particularly acute in tissues where a small number of cell types dominate the transcriptome. In brain tissue, neuronal genes are highly expressed and can contaminate non-neuronal clusters. The Neuron study documenting this issue found that ambient RNA contamination caused both misinterpreted cell types and masked cell types in brain single-nuclei datasets [<a href="#ref-7">7</a>].

Your annotation workflow should include explicit steps to detect and correct ambient RNA. Computational tools such as CellBender, which was used in the heart failure snRNA-seq study, can model and remove ambient contamination [<a href="#ref-5">5</a>]. You should also examine your clusters for unexpected co-expression of markers from different cell types, which may indicate contamination instead of genuine biology.

### Marker Gene Selection Requires snRNA-seq Validation

Marker genes are the foundation of cell type annotation. In snRNA-seq data, you need markers that are detectable in nuclear transcriptomes and specific to the cell types you expect in your tissue.

For brain tissue, the Fly Cell Atlas project annotated more than 250 distinct cell types from 580,000 nuclei and provided cell type-related gene signatures and transcription factor markers [<a href="#ref-8">8</a>]. This type of resource gives you validated markers for nuclear data. Similarly, the ssREAD database for Alzheimer's disease includes cell-type-specific marker genes derived from both scRNA-seq and snRNA-seq datasets [<a href="#ref-9">9</a>].

When you select markers, consider the following criteria:

- The marker should be detected in a substantial fraction of nuclei in your target cluster
- The marker should show minimal expression in other clusters
- The marker should have published validation in snRNA-seq data when possible
- Transcription factor markers can be particularly useful because they are often nuclear-localized and well detected in snRNA-seq data

The tobacco stem dataset described marker screening as a reuse opportunity, indicating that marker validation is an ongoing process in the field instead of a solved problem [<a href="#ref-4">4</a>].

## Practical Workflow for snRNA-seq Cell Type Annotation

### Step 1: Establish Data Quality Baselines

Before any annotation attempt, you must understand the quality characteristics of your dataset. Start by examining the distribution of total UMIs, number of genes detected, and mitochondrial read fraction across all nuclei. Generate these distributions separately for each sample to identify batch effects or sample-specific issues.

Record the median and range for each metric. These values become your quality baselines. Nuclei that fall far below the median gene count may be empty droplets or damaged nuclei. Nuclei with unusually high mitochondrial fractions may have cytoplasmic contamination. Your thresholds should be informed by the shape of your data distributions and by published snRNA-seq studies from similar tissues.

The heart failure study provides a reference point: from eight pooled samples, the authors recovered 48,886 nuclei and identified 14 cell types after quality control [<a href="#ref-5">5</a>]. This recovery rate, roughly 6,000 nuclei per pooled sample, gives you a sense of expected yield from myocardial biopsies.

### Step 2: Normalize and Integrate Samples

Normalization for snRNA-seq data follows similar principles to scRNA-seq but may require adjustments for the lower gene detection rates. After normalization, you need to integrate data across samples to remove batch effects while preserving biological variation.

Integration is particularly important when you plan to annotate cell types across multiple samples or conditions. The choice of integration method affects downstream annotation accuracy. For cross-dataset annotation, methods that account for distributional differences between scRNA-seq and snRNA-seq data are needed. The ScNucAdapt method was developed specifically for cross-annotation between paired and unpaired scRNA-seq and snRNA-seq datasets, using partial domain adaptation to address distributional and cell composition differences [<a href="#ref-10">10</a>].

For snATAC-seq data, which is often generated alongside snRNA-seq in multiomic studies, a zero-controlled statistical model has been developed to account for missing data and uneven sequencing coverage [<a href="#ref-11">11</a>]. This model enables cell type classification and annotation while accounting for the high sparsity of chromatin accessibility data.

### Step 3: Cluster and Determine Resolution

Clustering groups nuclei by transcriptional similarity. The resolution parameter controls the number of clusters you obtain. For snRNA-seq data, you may need higher resolution than for scRNA-seq data to resolve cell types that show subtle transcriptional differences in nuclear profiles.

After initial clustering, examine the cluster marker genes before committing to a resolution. If clusters contain mixed markers from multiple expected cell types, increase resolution. If clusters split known cell types without clear biological justification, decrease resolution.

The Alzheimer's disease bioinformatics guidelines describe cell clustering as one of 14 major analytic directions and provide a recommended workflow for each direction [<a href="#ref-12">12</a>]. These guidelines were implemented on a large snRNA-seq dataset in AD, giving you a tested reference for clustering parameters.

### Step 4: Annotate Clusters Using Multiple Lines of Evidence

Do not rely on a single marker gene or a single annotation method. Use at least three lines of evidence for each cluster:

1. Canonical marker expression: Check whether established markers for expected cell types are enriched in the cluster
2. Differential expression: Identify the top differentially expressed genes in the cluster and compare them to published cell type signatures
3. Reference-based annotation: Use label transfer from a validated reference dataset when available

For kidney tissue, a study testing five supervised machine learning algorithms for automatic cell type annotation found high accuracies with a median F1 score of 0.94 and a median rejection rate of 1.8 percent [<a href="#ref-13">13</a>]. However, the authors noted that F1 scores were lower when models trained primarily on scRNA-seq data were tested on snRNA-seq data [<a href="#ref-13">13</a>]. This finding reinforces the importance of using snRNA-seq-specific references or validating cross-platform transfers carefully.

### Step 5: Validate Annotations with Independent Methods

Annotation validation is essential but often overlooked. Independent validation methods include:

- Spatial transcriptomics: If spatial data are available for the same tissue, you can check whether annotated cell types localize to expected anatomical regions. The Seq-Scope technology achieves resolution comparable to an optical microscope, allowing visualization of single-cell types and subtypes in tissue context [<a href="#ref-14">14</a>]
- Multiomic integration: If you have matched snATAC-seq data, you can check whether chromatin accessibility at marker gene loci supports your RNA-based annotations [<a href="#ref-15">15</a>]
- Cross-dataset comparison: Compare your annotations to published atlases from the same tissue

The metastatic breast cancer study combined single-cell or single-nucleus RNA sequencing with four spatial expression assays and H&E staining, using the coupled measurements to assess variability in cell type composition and expression [<a href="#ref-16">16</a>]. This multi-modal approach provides a model for rigorous annotation validation.

## Options and Tradeoffs in Annotation Tools

### Manual Annotation

Manual annotation involves examining cluster marker genes and assigning cell type labels based on your biological knowledge and published literature. This approach gives you maximum control and is appropriate when you have deep expertise in the tissue being studied.

Advantages include the ability to identify novel cell states and to apply nuanced biological judgment. Disadvantages include time consumption, subjectivity, and difficulty scaling to large datasets.

### Automated Annotation with Supervised Learning

Supervised machine learning methods train on labeled reference data and predict labels for new data. The kidney study tested support vector machines, random forests, multilayer perceptrons, k-nearest neighbors, and extreme gradient boosting [<a href="#ref-13">13</a>]. All five methods performed well, but performance depended on the similarity between training and test data [<a href="#ref-13">13</a>].

These methods are fast and reproducible once trained. They are appropriate when you have a high-quality reference dataset from the same tissue and platform. The main risk is that they fail silently when the test data contain cell types not present in the training data, although the kidney study found that the algorithms successfully rejected cell types not present in the training data [<a href="#ref-13">13</a>].

### Label Transfer from Reference Atlases

Label transfer uses a reference dataset with known annotations to annotate a query dataset. This approach is implemented in several tools and was used in the Pink1 knockout mouse gut study, where the authors applied anchor-based label transfer and confirmed cell-type assignments via random forest classification [<a href="#ref-17">17</a>].

Label transfer works well when the reference and query datasets are from the same tissue and species. Cross-species and cross-platform transfers require additional caution. The ScNucAdapt method addresses the specific challenge of transferring labels between scRNA-seq and snRNA-seq data [<a href="#ref-10">10</a>].

### Integration with Multiomic Data

When snRNA-seq data are generated alongside snATAC-seq data, you can use both modalities for annotation. The human adult brain study integrated nuclear transcriptomic and DNA accessibility maps for more than 60,000 single cells, revealing regulatory elements and transcription factors that underlie cell-type distinctions [<a href="#ref-15">15</a>].

Multiomic annotation is more robust than RNA-only annotation because it requires concordance between two independent molecular measurements. However, it requires additional experimental and computational resources.

## Records and Measurements for Annotation Quality

### What to Record During Annotation

Maintain a detailed annotation record for each dataset. This record should include:

- Quality control thresholds and the rationale for each threshold
- Number of nuclei before and after quality control
- Clustering parameters and resolution values tested
- Marker genes used for each cluster and their expression statistics
- Annotation confidence level for each cluster (high, medium, low)
- Any clusters that could not be confidently annotated
- Software versions and parameter settings for all tools

This record enables reproducibility and provides a basis for troubleshooting when annotations are questioned.

### Metrics for Annotation Confidence

Several metrics help you quantify annotation confidence:

- Marker gene specificity: The proportion of nuclei in a cluster expressing a marker gene compared to other clusters
- Cluster separation: How distinctly a cluster separates from its nearest neighbors in the embedding
- Label transfer concordance: The agreement between manual annotation and reference-based annotation
- Cross-validation accuracy: If using supervised methods, the accuracy on held-out data

The kidney study used F1 scores and rejection rates as performance metrics [<a href="#ref-13">13</a>]. You can apply similar metrics when evaluating automated annotation methods on your own data.

## Common Failure Patterns in snRNA-seq Annotation

### Failure Pattern 1: Overlooking Ambient RNA Contamination

The most insidious failure pattern is ambient RNA contamination that creates false cell types or masks real ones. In brain single-nuclei datasets, neuronal ambient RNA contamination has been documented to cause both misinterpreted and masked cell types [<a href="#ref-7">7</a>].

Signs of this problem include clusters that express markers from multiple unrelated cell types, unexpected expression of highly abundant genes across all clusters, and cell type proportions that do not match histological expectations.

Prevention requires explicit ambient RNA correction before clustering. Detection requires careful examination of marker gene patterns and comparison with independent data such as spatial transcriptomics.

### Failure Pattern 2: Applying scRNA-seq Quality Thresholds to snRNA-seq Data

Quality control thresholds developed for scRNA-seq data often do not transfer to snRNA-seq data. The lower gene detection rate in nuclei means that aggressive filtering based on gene counts can remove legitimate cell types with low transcriptional activity.

The mitochondrial fraction threshold is particularly problematic. In scRNA-seq data, high mitochondrial fraction indicates dying cells. In snRNA-seq data, mitochondrial RNA should be minimal, and high levels indicate cytoplasmic contamination, which is a different problem requiring different corrective action.

### Failure Pattern 3: Using Markers Without Nuclear Validation

Marker genes validated in scRNA-seq data may not perform well in snRNA-seq data. This discrepancy can lead to misannotation or failure to resolve closely related cell types.

The solution is to validate markers in your own data before relying on them. Check the detection rate and specificity of each marker across your clusters. Consult snRNA-seq-specific resources such as the Fly Cell Atlas [<a href="#ref-8">8</a>] or ssREAD database [<a href="#ref-9">9</a>] when available.

### Failure Pattern 4: Ignoring Batch Effects Between Samples

When you pool samples for snRNA-seq, batch effects can create artificial clusters that are mistaken for cell types. The heart failure study used genotype-based demultiplexing to assign nuclei to individual patients from pooled samples, which is a powerful approach for managing batch effects [<a href="#ref-5">5</a>].

If you do not demultiplex pooled samples, you need robust integration methods to remove batch effects before annotation. Examine whether clusters are composed of nuclei from multiple samples or from a single sample. Clusters dominated by one sample may represent batch effects instead of biological cell types.

### Failure Pattern 5: Over-annotating Based on Minor Transcriptional Differences

The lower gene detection in snRNA-seq data can create spurious differences between clusters that do not represent genuine cell type distinctions. Over-annotation leads to cell type labels that cannot be validated by independent methods.

Guard against over-annotation by requiring multiple lines of evidence for each cell type label. If a proposed cell type cannot be distinguished by canonical markers, spatial localization, or independent datasets, consider whether it represents a cell state instead of a distinct cell type.

## Limitations of snRNA-seq Annotation

### Gene Detection Sensitivity

The most fundamental limitation of snRNA-seq is reduced gene detection compared to scRNA-seq. This limitation affects your ability to resolve cell types with low transcriptional activity and to detect cell state changes that involve genes with primarily cytoplasmic mRNA.

The heart failure study found that cardiomyocytes and fibroblasts showed the most differentially expressed genes, while endothelial cells, pericytes, and macrophages had fewer [<a href="#ref-5">5</a>]. This pattern may reflect both biological differences and technical detection limitations.

### Reference Data Availability

While scRNA-seq has extensive reference atlases for many tissues, snRNA-seq references are less developed. The ssREAD database for Alzheimer's disease [<a href="#ref-9">9</a>] and the Fly Cell Atlas [<a href="#ref-8">8</a>] are notable exceptions, but many tissues lack dedicated snRNA-seq references.

The lack of references affects both manual annotation and automated label transfer. When you annotate snRNA-seq data from a tissue without a dedicated reference, you must rely on cross-platform transfer with appropriate caution.

### Cross-Platform Transfer Challenges

Transferring annotations between scRNA-seq and snRNA-seq data is not straightforward. The kidney study found lower F1 scores when models trained on scRNA-seq data were tested on snRNA-seq data [<a href="#ref-13">13</a>]. The ScNucAdapt method was developed specifically to address this challenge [<a href="#ref-10">10</a>].

If you must transfer labels across platforms, validate the transfer with independent methods and be prepared to manually correct misannotated clusters.

### FFPE Sample Complexity

FFPE tissues present additional challenges for snRNA-seq. The imRandom-seq platform was developed to enable single-nucleus and spatial total RNA profiling of diverse FFPE specimens, with enhanced gene detection and reduced nuclear loss compared to traditional approaches [<a href="#ref-2">2</a>]. However, FFPE-derived data still require careful quality assessment and may show tissue-specific artifacts.

The scFFPE-seq approach demonstrated comparable cell abundance, cell type annotation, and pathway characterization between FFPE and fresh tissues for colon, ileum, and skin samples [<a href="#ref-3">3</a>]. This result is encouraging but does not guarantee equivalent performance for all tissue types and preservation conditions.

## Quality and Reproducibility Context for snRNA-seq Annotation

### Reproducibility Standards

Reproducibility in snRNA-seq analysis requires documentation of every computational step. The Galaxy Training Network provides accessible workflow training and analysis tutorials that emphasize reproducibility [<a href="#ref-18">18</a>]. The nf-core documentation describes community pipeline standards for reproducible workflow configuration [<a href="#ref-19">19</a>].

For your own analyses, version control all scripts, record software versions, and archive the exact parameters used for each analysis step. The Carpentries lessons provide foundational training in version control with Git and reproducible computing practices [<a href="#ref-20">20</a>].

### Data Sharing and Database Deposition

Public data repositories are essential for validating annotations and enabling cross-study comparisons. The NCBI provides official descriptions of databases, search systems, sequence resources, and analysis services that support data deposition and retrieval [<a href="#ref-21">21</a>]. The EMBL-EBI Training program offers learning pathways for using these data resources effectively [<a href="#ref-22">22</a>].

When you publish snRNA-seq data, deposit raw data and processed count matrices in public repositories. Include detailed metadata about sample collection, nuclear isolation, library preparation, and sequencing parameters. This metadata is essential for other researchers to assess data quality and annotation validity.

### Professional Escalation Criteria

Some annotation problems require consultation with experts. Escalate to a bioinformatics specialist or a domain expert in your tissue of interest when:

- You cannot confidently annotate more than 20 percent of your clusters
- Your annotations conflict with published atlases from the same tissue
- You observe unexpected cell types that cannot be validated by independent methods
- Your quality control metrics fall far outside published ranges for similar tissues
- You need to transfer annotations across species or across scRNA-seq and snRNA-seq platforms

The Bioconductor project provides official package documentation and workflows for reproducible genomic analysis [<a href="#ref-23">23</a>]. If you are using Bioconductor tools, consult the package documentation and support forums before escalating.

## Building a Structured Annotation Audit System for snRNA-seq Data

A recurring problem for researchers new to single-nucleus RNA sequencing is knowing when an annotation is trustworthy enough to support biological conclusions. Most training materials describe how to annotate but do not provide a concrete system for auditing the annotation process itself. This section gives you a structured audit framework that treats annotation as a documented, testable procedure instead of a one-time labeling event. The framework is built around a decision gate system, a standardized record format, and a troubleshooting protocol that catches errors before they propagate into downstream analysis.

### The Decision Gate Framework for Annotation Progression

The decision gate framework divides annotation into five sequential gates. Each gate has explicit entry criteria, required outputs, and failure actions. You do not proceed to the next gate until the current gate passes its criteria. This structure prevents the common error of moving to differential expression or trajectory analysis with annotations that have not been properly validated.

**Gate 1: Data Integrity Confirmation**

Entry criteria for this gate are that you have completed quality control and understand the composition of your dataset. Before any clustering or annotation, you must confirm three things. First, the number of nuclei passing quality control matches your expectations based on the input material and isolation protocol. Second, the distribution of total UMIs and gene counts per nucleus follows a pattern consistent with published snRNA-seq data from similar tissues. Third, you have documented the ambient RNA profile of your dataset.

The output of Gate 1 is a quality control summary table that records the number of nuclei before and after filtering, the thresholds applied for each metric, and the rationale for those thresholds. If more than 30 percent of your nuclei fail quality control, stop and investigate the isolation protocol or sequencing depth before proceeding. If your mitochondrial fraction is unexpectedly high, check whether your nuclear isolation protocol adequately removed cytoplasmic contamination.

**Gate 2: Cluster Stability Assessment**

Entry criteria for Gate 2 are that you have performed clustering and have a preliminary cluster count. The purpose of this gate is to confirm that your clusters are stable and reproducible instead of artifacts of parameter choices. Run clustering at multiple resolution values and compare the results. Stable clusters persist across resolution changes while unstable clusters merge or split unpredictably.

For each cluster, calculate a stability score based on the proportion of nuclei that remain together across clustering runs at different resolutions. Clusters with stability scores below 0.8 require investigation. They may represent transitional cell states, doublets, or ambient RNA artifacts. Document the resolution values tested and the stability scores for each cluster in your annotation record.

The output of Gate 2 is a cluster stability report that lists each cluster, its stability score, and the resolution range over which it remains stable. This report becomes part of your permanent annotation record.

**Gate 3: Marker Gene Verification**

Entry criteria for Gate 3 are that you have stable clusters and have generated preliminary marker gene lists for each cluster. The purpose of this gate is to verify that your markers are genuinely specific to each cluster and detectable in nuclear transcriptomes.

For each candidate marker gene, calculate three metrics. The detection rate is the proportion of nuclei in the target cluster that express the gene. The specificity is the ratio of detection in the target cluster compared to the highest detection in any other cluster. The enrichment is the fold change in expression between the target cluster and all other clusters combined.

A marker passes verification when it meets all three criteria: detection rate above 50 percent in the target cluster, specificity ratio above 5, and enrichment above 2. Markers that fail any criterion should be flagged and replaced with alternatives. If you cannot find passing markers for a cluster, that cluster may not represent a genuine cell type.

The output of Gate 3 is a marker verification table for every cluster. This table lists each marker, its three metrics, and whether it passed or failed. This table is essential for defending your annotations in peer review.

**Gate 4: Cross-Validation with Independent Methods**

Entry criteria for Gate 4 are that you have verified marker genes for all clusters. The purpose of this gate is to confirm your annotations using methods that do not depend on the same assumptions as your primary annotation approach.

At least two independent validation methods are required. Acceptable methods include label transfer from a published reference atlas, comparison with spatial transcriptomics data if available, integration with matched snATAC-seq data, or manual review by a second researcher with domain expertise. The kidney study that tested five supervised machine learning algorithms found that all methods achieved high accuracy with a median F1 score of 0.94, but performance depended on the similarity between training and test data [<a href="#ref-13">13</a>]. This finding supports the use of automated methods for cross-validation but also warns against relying on a single method.

The output of Gate 4 is a cross-validation report that records which validation methods were used, the concordance between methods, and any clusters where methods disagreed. Disagreements must be resolved before proceeding.

**Gate 5: Final Annotation Review and Documentation**

Entry criteria for Gate 5 are that all clusters have passed Gates 1 through 4. The purpose of this gate is to produce the final annotation set and complete documentation. Assign final cell type labels to each cluster, assign a confidence level of high, medium, or low, and document any clusters that remain unannotated.

The output of Gate 5 is the complete annotation record that includes all tables and reports from previous gates. This record becomes the reference for all downstream analyses.

### The Standardized Annotation Record System

A standardized record system ensures that your annotation process is reproducible and auditable. The record system has four components that you maintain throughout the analysis.

**Component 1: Sample and Processing Log**

This log records the origin of each sample, the tissue type, the preservation method, the nuclear isolation protocol, and the sequencing platform. For frozen tissues, record the storage duration and conditions. For FFPE samples, record the fixation duration and any available information about tissue age. The long-term frozen brain tumor study demonstrated that a simplified nuclear isolation protocol can provide intact nuclei from frozen tissues for both droplet and plate-based platforms [<a href="#ref-1">1</a>]. Your log should document which protocol you used and any deviations from the published method.

**Component 2: Quality Control Decision Log**

This log records every quality control decision and the evidence supporting it. For each threshold you set, record the distribution of the metric, the rationale for the threshold, and the number of nuclei affected. This log is particularly important for snRNA-seq data because thresholds borrowed from scRNA-seq protocols often do not apply.

**Component 3: Cluster Annotation Worksheet**

This worksheet contains one row per cluster with columns for cluster identifier, number of nuclei, top marker genes, marker verification metrics, preliminary annotation, validation methods used, final annotation, and confidence level. The worksheet is the working document you update as you progress through the decision gates.

**Component 4: Troubleshooting Incident Report**

This report documents any problems encountered during annotation and how they were resolved. Each incident report includes the problem description, the evidence that identified the problem, the investigation steps, the resolution, and the impact on the final annotations. Incident reports are valuable for future projects and for training new lab members.

### Troubleshooting Method for Annotation Failures

When annotations fail validation, use this structured troubleshooting method to identify the cause. The method follows a sequence of hypotheses, each tested with specific diagnostics.

**Hypothesis 1: Ambient RNA Contamination**

Test this hypothesis when clusters show unexpected co-expression of markers from multiple cell types. Neuronal ambient RNA contamination has been documented to cause misinterpreted and masked cell types in brain single-nuclei datasets [<a href="#ref-7">7</a>]. Examine the expression of highly abundant genes across all clusters. If these genes show uniform expression regardless of cluster identity, ambient contamination is likely. Apply computational correction methods such as CellBender, which was used in the heart failure snRNA-seq study to model and remove ambient contamination [<a href="#ref-5">5</a>]. After correction, re-cluster and re-annotate.

**Hypothesis 2: Batch Effects Between Samples**

Test this hypothesis when clusters are dominated by nuclei from a single sample. Examine the sample composition of each cluster. If clusters separate primarily by sample origin instead of biological identity, batch effects are present. The heart failure study used genotype-based demultiplexing to assign nuclei to individual patients from pooled samples, which is a powerful approach for managing batch effects [<a href="#ref-5">5</a>]. If you did not demultiplex, apply integration methods to remove batch effects before re-clustering.

**Hypothesis 3: Inappropriate Resolution Parameters**

Test this hypothesis when clusters contain mixed markers from multiple expected cell types or when known cell types are split across multiple clusters. Re-run clustering at different resolution values and compare the results. The Alzheimer's disease bioinformatics guidelines describe cell clustering as one of 14 major analytic directions and provide a recommended workflow for each direction [<a href="#ref-12">12</a>]. These guidelines were implemented on a large snRNA-seq dataset in AD, giving you a tested reference for clustering parameters.

**Hypothesis 4: Marker Genes Not Detectable in Nuclear Transcriptomes**

Test this hypothesis when expected markers show low detection rates across all clusters. Some markers validated in scRNA-seq data may not perform well in snRNA-seq data because their mRNA is primarily cytoplasmic. Consult snRNA-seq-specific resources such as the Fly Cell Atlas, which provides cell type-related gene signatures and transcription factor markers for more than 250 distinct cell types [<a href="#ref-8">8</a>]. The ssREAD database for Alzheimer's disease also includes cell-type-specific marker genes derived from both scRNA-seq and snRNA-seq datasets [<a href="#ref-9">9</a>]. Replace undetectable markers with alternatives that show nuclear expression.

**Hypothesis 5: Doublets or Multiplets**

Test this hypothesis when clusters show co-expression of markers from two distinct cell types at similar levels. Doublets occur when two nuclei are captured in the same droplet or well. Examine the total UMI counts for suspect clusters. Doublets often show higher total UMIs than single nuclei. Apply doublet detection methods and remove identified doublets before re-clustering.

### Implementing the Audit System in Practice

Implementing this audit system requires planning before you begin annotation. Allocate time for each gate and record your progress systematically. The complete audit record for a typical dataset with 10 to 20 clusters requires approximately 8 to 12 hours of additional work beyond the annotation itself. This investment prevents costly errors in downstream analysis.

Start by creating the four record components before you begin any analysis. This preparation ensures that you document decisions as you make them instead of reconstructing them later. The Galaxy Training Network provides accessible workflow training and analysis tutorials that emphasize reproducibility [<a href="#ref-18">18</a>]. The nf-core documentation describes community pipeline standards for reproducible workflow configuration [<a href="#ref-19">19</a>]. These resources can help you structure your analysis pipeline to integrate with the audit system.

For version control of your annotation scripts and records, use Git as taught in The Carpentries lessons [<a href="#ref-20">20</a>]. Version control ensures that you can reconstruct exactly how each annotation was made and provides a clear audit trail for collaborators and reviewers.

### Common Failure Patterns in the Audit Process

Several failure patterns recur when researchers implement structured annotation audits. Recognizing these patterns helps you avoid them.

**Failure Pattern 1: Skipping Gates Under Time Pressure**

The most common failure is skipping validation gates when deadlines approach. This pattern leads to annotations that cannot be defended and often require complete re-analysis. The audit system only works if you complete every gate for every cluster.

**Failure Pattern 2: Confirmation Bias in Marker Selection**

Researchers often select markers that confirm their expected cell types while ignoring contradictory evidence. The marker verification metrics in Gate 3 provide objective criteria that reduce confirmation bias. If a marker fails verification, you must replace it regardless of your expectations.

**Failure Pattern 3: Treating Automated Annotation as Final**

Automated annotation methods are tools, not final answers. The kidney study found that all five tested machine learning algorithms performed well, but performance depended on the similarity between training and test data [<a href="#ref-13">13</a>]. Automated methods can mislabel clusters when the test data contain cell types not represented in the training data. Always validate automated annotations with manual review and independent methods.

**Failure Pattern 4: Incomplete Documentation**

Documentation that is incomplete or reconstructed after the fact undermines the entire audit system. The record system only provides value if you maintain it in real time. Schedule documentation time after each analysis session instead of attempting to reconstruct records at the end of the project.

### Professional Escalation Criteria for the Audit System

Some annotation problems require expertise beyond what a beginner can reasonably develop. Escalate to a bioinformatics specialist or domain expert when you encounter any of the following situations.

Escalate when you cannot confidently annotate more than 20 percent of your clusters after completing all five gates. This situation suggests either a fundamental data quality problem or a tissue type that requires specialized knowledge.

Escalate when your annotations conflict with published atlases from the same tissue and you cannot resolve the discrepancy through the troubleshooting method. The Fly Cell Atlas [<a href="#ref-8">8</a>] and ssREAD database [<a href="#ref-9">9</a>] provide reference annotations that can help identify conflicts.

Escalate when you observe unexpected cell types that cannot be validated by any independent method. Novel cell types require substantial evidence before they can be reported with confidence.

Escalate when you need to transfer annotations across species or across scRNA-seq and snRNA-seq platforms. Cross-platform transfer requires specialized methods such as ScNucAdapt, which was developed specifically for cross-annotation between paired and unpaired scRNA-seq and snRNA-seq datasets [<a href="#ref-10">10</a>].

Escalate when your quality control metrics fall far outside published ranges for similar tissues. The heart failure study recovered 48,886 nuclei from eight pooled samples and identified 14 cell types [<a href="#ref-5">5</a>]. If your recovery rates are substantially lower, a protocol issue may require expert attention.

The Bioconductor project provides official package documentation and workflows for reproducible genomic analysis [<a href="#ref-23">23</a>]. If you are using Bioconductor tools, consult the package documentation and support forums before escalating. The EMBL-EBI Training program offers learning pathways for using data resources effectively [<a href="#ref-22">22</a>], and the NCBI provides official descriptions of databases and analysis services [<a href="#ref-21">21</a>]. These resources can help you resolve many problems before seeking expert consultation.

### Records and Measurements for the Audit System

The audit system generates a specific set of records and measurements that you should maintain for every dataset. These records serve both immediate quality assurance and long-term reproducibility.

The quality control summary records the number of nuclei before and after filtering, the thresholds applied, and the rationale for each threshold. This summary is the first table in your annotation record.

The cluster stability report records the resolution values tested and the stability score for each cluster. This report is generated during Gate 2 and becomes part of the permanent record.

The marker verification table records the detection rate, specificity, and enrichment for every marker gene tested. This table is generated during Gate 3 and is essential for defending your annotations.

The cross-validation report records which validation methods were used and the concordance between methods. This report is generated during Gate 4.

The final annotation table records the cluster identifier, final cell type label, confidence level, and any notes about unresolved issues. This table is the primary output of the audit system.

The troubleshooting incident report documents any problems encountered and how they were resolved. This report is maintained throughout the analysis and provides valuable lessons for future projects.

### Integration with Reproducible Workflow Standards

The audit system integrates with established reproducibility standards for bioinformatics analysis. The nf-core documentation describes community pipeline standards for reproducible workflow configuration [<a href="#ref-19">19</a>]. These standards emphasize version control, containerization, and parameter documentation. Your audit records should reference the specific pipeline version and parameters used for each analysis step.

The Galaxy Training Network provides accessible workflow training and analysis tutorials that emphasize reproducibility [<a href="#ref-18">18</a>]. These tutorials can help you implement the audit system within a reproducible workflow framework.

The Carpentries lessons provide foundational training in version control with Git and reproducible computing practices [<a href="#ref-20">20</a>]. These skills are essential for maintaining the audit record system.

### Limitations of the Audit System

The audit system has limitations that you should acknowledge. First, it cannot detect all annotation errors. Some errors require biological validation that is beyond the scope of computational analysis. Second, the system adds time to the annotation process. This time investment is justified for research projects but may be excessive for exploratory analyses. Third, the system depends on the quality of reference data used for validation. If your reference atlases contain errors, those errors can propagate through the audit process.

The audit system is most effective when combined with domain expertise. The multi-modal approach used in the metastatic breast cancer study, which combined single-cell or single-nucleus RNA sequencing with four spatial expression assays and H&E staining, provides a model for rigorous annotation validation [<a href="#ref-16">16</a>]. While you may not have access to such extensive validation resources, the audit system helps you make the most of the resources you do have.

### Applying the Audit System to Cross-Platform Comparisons

The audit system is particularly valuable when you need to compare annotations across scRNA-seq and snRNA-seq datasets. The kidney study found lower F1 scores when models trained primarily on scRNA-seq data were tested on snRNA-seq data [<a href="#ref-13">13</a>]. This finding highlights the need for careful validation when transferring annotations across platforms.

When comparing annotations across platforms, use the audit system to document the platform-specific characteristics of each dataset. Record the gene detection rates, ambient RNA profiles, and marker performance for each platform separately. This documentation helps you interpret differences in cell type composition that may reflect technical artifacts instead of biological variation.

The ScNucAdapt method provides a practical framework for cross-domain cell type annotation between scRNA-seq and snRNA-seq data [<a href="#ref-10">10</a>]. When using this method, apply the audit system to validate the transferred labels instead of accepting them without verification.

### Summary of the Audit System

The structured annotation audit system provides a concrete framework for ensuring annotation quality in snRNA-seq data. The five decision gates create a progression from data integrity confirmation through cluster stability assessment, marker gene verification, cross-validation with independent methods, and final annotation review. The standardized record system ensures that every decision is documented and auditable. The troubleshooting method provides a systematic approach to identifying and resolving annotation failures.

This audit system addresses the specific challenges of snRNA-seq annotation, including lower gene detection, ambient RNA contamination, and the need for nuclear-validated markers. By implementing this system, you can produce annotations that are defensible, reproducible, and suitable for downstream biological analysis.

## Frequently Asked Questions

### Why do I detect fewer genes per nucleus than per cell in scRNA-seq data?

Nuclear transcriptomes contain a different RNA pool than whole cells. Mature mRNA is exported to the cytoplasm, so the nucleus retains primarily pre-mRNA and other nuclear RNA species. This biological difference, combined with the technical challenges of nuclear isolation, results in lower gene detection per nucleus. The reduced detection affects marker sensitivity and requires adjustment of quality control thresholds.

### How should I adjust quality control thresholds for snRNA-seq data?

Set thresholds based on the distributions in your own data instead of borrowing from scRNA-seq studies. Examine the distribution of total UMIs, gene counts, and mitochondrial fractions across all nuclei. Mitochondrial fraction has a different interpretation in snRNA-seq data because mitochondria are cytoplasmic. High mitochondrial RNA in nuclear preparations indicates cytoplasmic contamination instead of cell stress.

### What is ambient RNA contamination and why is it worse in snRNA-seq data?

Ambient RNA is RNA from lysed or damaged cells that is captured alongside the intended nucleus. In brain tissue, highly expressed neuronal genes can contaminate non-neuronal clusters, causing misinterpreted and masked cell types [<a href="#ref-7">7</a>]. The problem is significant in snRNA-seq data because nuclear isolation can release RNA from damaged cells, and the high expression levels of certain genes in tissues like the brain amplify the contamination signal.

### Can I use scRNA-seq marker genes to annotate snRNA-seq clusters?

You can use scRNA-seq markers as a starting point, but you must validate them in your snRNA-seq data. Some markers may show reduced detection because their mRNA is primarily cytoplasmic. Transcription factor markers are often more reliable because they are nuclear-localized. Consult snRNA-seq-specific references such as the Fly Cell Atlas [<a href="#ref-8">8</a>] or ssREAD database [<a href="#ref-9">9</a>] when available.

### What is label transfer and when should I use it?

Label transfer uses a reference dataset with known annotations to predict labels for a query dataset. It is appropriate when you have a high-quality reference from the same tissue and species. Cross-platform transfer between scRNA-seq and snRNA-seq data requires specialized methods such as ScNucAdapt because of distributional differences between the two data types [<a href="#ref-10">10</a>].

### How do I know if my annotations are correct?

Use multiple lines of evidence for each cluster. Check canonical marker expression, identify differentially expressed genes, and compare to published references. Validate with independent methods such as spatial transcriptomics [<a href="#ref-14">14</a>] or multiomic integration [<a href="#ref-15">15</a>] when available. If your annotations conflict with published atlases, investigate the discrepancy before proceeding.

### What should I do if I cannot annotate some clusters?

Document the unannotated clusters and their marker genes. Some may represent novel cell types or cell states that have not been described. Others may be artifacts of ambient RNA contamination or batch effects. Consider whether higher resolution clustering or additional quality control would resolve the issue. If unannotated clusters persist, consult a domain expert or bioinformatics specialist.

### How does snRNA-seq annotation differ for FFPE samples?

FFPE samples present additional challenges because of RNA degradation during fixation and storage. Specialized protocols such as imRandom-seq [<a href="#ref-2">2</a>] and scFFPE-seq [<a href="#ref-3">3</a>] have been developed to address these challenges. Gene detection may be lower than in fresh frozen samples, and you may need to adjust quality thresholds accordingly. Validate annotations carefully because degraded RNA can create artifacts.

## Related Bioinformatics Guides

- [Single-Cell RNA-seq Clustering and Cell-Type Annotation Pipelines](/knowledge/bioinformatics/single-cell-rna-seq-clustering-and-cell-type-annotation-pipelines)
- [Single-Cell Annotation: A Workflow for Cell Type Identification](/knowledge/bioinformatics/single-cell-annotation-a-workflow-for-cell-type-identification)
- [RNA-Seq vs DNA-Seq: Key Differences and Applications](/knowledge/bioinformatics/rna-seq-vs-dna-seq-key-differences-and-applications)
- [Single-Cell Sequencing Depth: How Much Is Enough?](/knowledge/bioinformatics/single-cell-sequencing-depth-how-much-is-enough)
- [Single-Cell vs Single-Nucleus RNA Sequencing: Choosing the Right Approach](/knowledge/bioinformatics/single-cell-vs-single-nucleus-rna-sequencing-choosing-the-right-approach)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)

## References and Further Reading

<a id="ref-1"></a>[<a href="#ref-1">1</a>] [A simplified preparation method for single-nucleus RNA-sequencing using long-term frozen brain tumor tissues](https://doi.org/10.1038/s41598-025-97053-9). Scientific Reports, 2025.

<a id="ref-2"></a>[<a href="#ref-2">2</a>] [Automated in situ microfluidic Random-seq for robust single-nucleus and spatial total RNA profiling of diverse FFPE specimens.](https://doi.org/10.1038/s41467-026-74203-9). 2026.

<a id="ref-3"></a>[<a href="#ref-3">3</a>] [Probing the Feasibility of Single-Cell Fixed RNA Sequencing from FFPE Tissue](https://doi.org/10.3390/ijms27031605). International Journal of Molecular Sciences, 2026.

<a id="ref-4"></a>[<a href="#ref-4">4</a>] [Single-nucleus transcriptomic data of susceptible and resistant tobacco stems after black shank infection.](https://doi.org/10.1016/j.dib.2026.113095). 2026.

<a id="ref-5"></a>[<a href="#ref-5">5</a>] [Single cell transcriptomic analyses of human heart failure with preserved ejection fraction.](https://pubmed.ncbi.nlm.nih.gov/40161758). bioRxiv : the preprint server for biology, 2025.

<a id="ref-6"></a>[<a href="#ref-6">6</a>] [Protocol for isolation of nuclei from murine cardiac tissue for single-nucleus multiomic sequencing.](https://doi.org/10.1016/j.xpro.2026.104615). 2026.

<a id="ref-7"></a>[<a href="#ref-7">7</a>] [Neuronal ambient RNA contamination causes misinterpreted and masked cell types in brain single-nuclei datasets](https://doi.org/10.1016/j.neuron.2022.09.010). Neuron, 2022.

<a id="ref-8"></a>[<a href="#ref-8">8</a>] [Fly Cell Atlas: A single-nucleus transcriptomic atlas of the adult fruit fly.](https://pubmed.ncbi.nlm.nih.gov/35239393). Science (New York, N.Y.), 2022.

<a id="ref-9"></a>[<a href="#ref-9">9</a>] [A single-cell and spatial RNA-seq database for Alzheimer's disease (ssREAD).](https://pubmed.ncbi.nlm.nih.gov/38844475). Nature communications, 2024.

<a id="ref-10"></a>[<a href="#ref-10">10</a>] [Partial domain adaptation enables cross domain cell type annotation between scRNA-seq and snRNA-seq.](https://doi.org/10.1371/journal.pcbi.1014223). 2026.

<a id="ref-11"></a>[<a href="#ref-11">11</a>] [Zero-Controlled Statistical Model for Single Nucleus ATAC-Seq Data Analysis and Demultiplexing](https://doi.org/10.1681/ASN.20223311S1367a). Journal of the American Society of Nephrology, 2022.

<a id="ref-12"></a>[<a href="#ref-12">12</a>] [Guidelines for bioinformatics of single-cell sequencing data analysis in Alzheimer's disease: review, recommendation, implementation and application.](https://pubmed.ncbi.nlm.nih.gov/35236372). Molecular neurodegeneration, 2022.

<a id="ref-13"></a>[<a href="#ref-13">13</a>] [Identification of kidney cell types in scRNA-seq and snRNA-seq data using machine learning algorithms](https://doi.org/10.1016/j.heliyon.2024.e38567). Heliyon, 2024.

<a id="ref-14"></a>[<a href="#ref-14">14</a>] [Microscopic examination of spatial transcriptome using Seq-Scope.](https://pubmed.ncbi.nlm.nih.gov/34115981). Cell, 2021.

<a id="ref-15"></a>[<a href="#ref-15">15</a>] [Integrative single-cell analysis of transcriptional and epigenetic states in the human adult brain.](https://pubmed.ncbi.nlm.nih.gov/29227469). Nature biotechnology, 2018.

<a id="ref-16"></a>[<a href="#ref-16">16</a>] [A multi-modal single-cell and spatial expression map of metastatic breast cancer biopsies across clinicopathological features.](https://pubmed.ncbi.nlm.nih.gov/39478111). Nature medicine, 2024.

<a id="ref-17"></a>[<a href="#ref-17">17</a>] [A single-nucleus RNA-seq dataset of the colon in Pink1-deficient and wild-type mice.](https://doi.org/10.1038/s41597-026-07193-4). 2026.

<a id="ref-18"></a>[<a href="#ref-18">18</a>] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.

<a id="ref-19"></a>[<a href="#ref-19">19</a>] [nf-core Documentation](https://nf-co.re/docs). nf-core.

<a id="ref-20"></a>[<a href="#ref-20">20</a>] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.

<a id="ref-21"></a>[<a href="#ref-21">21</a>] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.

<a id="ref-22"></a>[<a href="#ref-22">22</a>] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.

<a id="ref-23"></a>[<a href="#ref-23">23</a>] [Bioconductor](https://bioconductor.org/). Bioconductor Project.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.