How to Identify Marker Genes for Single-Cell Clusters: A Step-by-Step Guide Using Seurat and Scanpy

By Dr. Zubair Khalid, DVM, MS, PhD ·

How to Identify Marker Genes for Single-Cell Clusters: A Step-by-Step Guide Using Seurat and Scanpy

Key Takeaways

  • Marker gene identification is a critical step for annotating single-cell RNA sequencing (scRNA-seq) and single-nucleus RNA sequencing (snRNA-seq) clusters, bridging computational outputs with biological cell types.
  • Statistical tests like the Wilcoxon rank-sum (default in Seurat/Scanpy), MAST, and negative binomial are employed to identify genes with significantly differential expression between clusters, with MAST being particularly useful for accounting for dropout events common in scRNA-seq.
  • Visualization tools such as violin plots, feature plots (e.g., UMAP), and dot plots are essential for assessing the specificity and distribution of candidate marker gene expression across identified cell clusters.
  • Robust marker identification necessitates rigorous quality control to filter out low-quality cells and technical artifacts, and careful consideration of normalization and batch effect correction to ensure biological signal is not confounded by technical variation.
  • Interpretation of marker gene lists requires cross-referencing with known cell type markers from databases like NCBI Gene, and validation across independent datasets or experimental methods (e.g., flow cytometry) is crucial for confirming biological relevance.

Marker gene identification is the process of finding genes whose expression distinguishes one cell cluster from all others in a single-cell RNA sequencing dataset. This guide provides a practical workflow for researchers and laboratory professionals using Seurat in R or Scanpy in Python, covering statistical tests, visualization, and interpretation of results for downstream cell type annotation.

Scope and Reader Context

This article addresses the specific problem of moving from unsupervised clustering output to biologically meaningful cell type labels. After quality control, normalization, dimensionality reduction, and clustering, analysts face the question of what each cluster represents. Marker genes provide the bridge between computational clusters and biological identity. The workflow described here applies to both single-cell RNA sequencing (scRNA-seq) and single-nucleus RNA sequencing (snRNA-seq) data, with adjustments noted where relevant.

The intended reader is a biology student, researcher, or laboratory professional who has already generated or obtained single-cell data and needs a reproducible method for identifying cluster-defining genes. The guide assumes familiarity with basic R or Python programming but does not require advanced statistical training. Both Seurat and Scanpy are covered because many laboratories use one or the other, and the underlying statistical concepts are shared.

At a Glance

Workflow StepSeurat FunctionScanpy FunctionKey Decision Point
Identify cluster markersFindAllMarkers()sc.tl.rank_genes_groups()Choose test method and minimum expression threshold
Compare one cluster vs othersFindMarkers()sc.tl.rank_genes_groups() with groupbyDefine the comparison group and test type
Visualize marker expressionVlnPlot(), FeaturePlot(), DotPlot()sc.pl.violin(), sc.pl.umap(), sc.pl.dotplot()Select genes that show specific instead of ubiquitous expression
Score marker panelsAddModuleScore()sc.tl.score_genes()Validate that known markers align with expected cell types
Annotate clustersRenameIdents()sc.tl.map_cell_types() or manual mappingCross-check multiple markers per cluster before assigning labels

Understanding Marker Gene Identification in Single-Cell Analysis

What Constitutes a Marker Gene

A marker gene for a single-cell cluster is a gene whose expression pattern reliably distinguishes cells in that cluster from cells in other clusters. In practice, this means the gene shows elevated expression in the target cluster relative to the rest of the dataset, and ideally the expression is consistent across most cells within that cluster. The NCBI Data Resources provide access to gene annotation databases that help researchers verify whether a candidate marker has known functional relevance to a suspected cell type.

Marker genes serve multiple purposes in single-cell analysis. They enable cell type annotation, which is the assignment of biological labels such as "T cell" or "hepatocyte" to computational clusters. They also support validation of clustering results, because clusters that express coherent sets of known markers are more likely to represent genuine biological populations instead of technical artifacts. In disease studies, marker genes identified through single-cell analysis often become candidates for further investigation. For example, studies of rheumatoid arthritis have used single-cell RNA sequencing to characterize T cell heterogeneity and identify marker genes relevant to disease mechanisms [<a href="#ref-1">1</a>].

Relationship Between Clustering and Marker Identification

Clustering and marker identification are iterative processes. Initial clustering uses highly variable genes to partition cells into groups based on transcriptional similarity. Marker identification then characterizes each group by finding differentially expressed genes. The results can reveal that some clusters are too granular, representing subtypes of the same cell type, or too coarse, merging distinct populations. Analysts often adjust clustering resolution and repeat marker identification until clusters align with known biology.

The EMBL-EBI Training resources provide structured learning pathways for single-cell analysis that emphasize this iterative workflow. Their practical training materials cover the full pipeline from raw data processing through cluster annotation, which helps researchers understand where marker identification fits in the broader analysis context.

Statistical Basis of Marker Gene Tests

Marker gene identification relies on differential expression testing between groups of cells. The null hypothesis is that a gene's expression does not differ between the target cluster and the comparison group. Tests commonly used include the Wilcoxon rank-sum test, the MAST test, and the negative binomial test implemented in Seurat's FindMarkers() function.

The Wilcoxon rank-sum test is nonparametric and does not assume a specific distribution of expression values. It ranks cells by expression level and tests whether the ranks in one group are systematically higher or lower than in the other group. This test is computationally efficient and works well for large datasets. The MAST test models expression as a two-part process, one part for the probability of detection and one part for the expression level given detection. This approach can be more sensitive for genes that are expressed in only a fraction of cells. Studies integrating single-cell and bulk RNA sequencing data have used Seurat with the MAST test to identify marker genes in core cell populations [<a href="#ref-2">2</a>].

The choice of statistical test affects results, and analysts should understand the assumptions of each method. Nonparametric tests are generally safer when expression distributions are skewed, which is common in single-cell data. Parametric tests may have more power when their assumptions are met but can produce false positives when assumptions are violated.

Preparing Data for Marker Gene Identification

Quality Control Requirements

Marker gene identification is only as reliable as the underlying data. Poor quality cells with low gene counts, high mitochondrial content, or excessive ambient RNA contamination can produce spurious marker genes. Quality control should be completed before clustering and marker identification. The Galaxy Training Network offers accessible tutorials on single-cell quality control that cover filtering criteria and visualization of quality metrics.

Standard quality control steps include filtering cells by minimum and maximum gene counts, filtering by mitochondrial read fraction, and removing doublets. The specific thresholds depend on tissue type and experimental protocol. For example, frozen tissues processed for single-nucleus RNA sequencing typically show different quality metrics than fresh tissues processed for single-cell RNA sequencing. Analysts should examine distributions of quality metrics and set thresholds based on the data instead of applying universal cutoffs.

Normalization and Batch Effects

Normalization adjusts for differences in sequencing depth across cells so that expression values are comparable. Common approaches include log-normalization in Seurat and the shifted logarithm or Pearson residuals methods in Scanpy. After normalization, data integration may be necessary if samples come from multiple batches, patients, or experimental conditions.

Batch effects can create artificial clusters that reflect technical variation instead of biology. If marker identification is performed on unintegrated data, the resulting markers may distinguish batches instead of cell types. Integration methods such as Harmony, scVI, or Seurat's integration workflow align cells across batches while preserving biological variation. The nf-core Documentation describes community standards for single-cell pipelines that include integration steps, providing a reference for reproducible analysis configurations.

Feature Selection Before Clustering

Before clustering, analysts typically select highly variable genes. This step reduces dimensionality and focuses clustering on genes that carry biological signal. The number of highly variable genes selected affects clustering resolution and downstream marker identification. Selecting too few genes may miss subtle distinctions between closely related cell types. Selecting too many genes may introduce noise that obscures genuine structure.

The Bioconductor project provides R packages for single-cell analysis that include alternative feature selection methods. Researchers working in R can compare Seurat's variance-stabilizing transformation with Bioconductor approaches to determine which produces more interpretable clusters for their specific dataset.

Step-by-Step Marker Identification in Seurat

Loading and Verifying the Seurat Object

The first step is to confirm that the Seurat object contains the expected cells, genes, and metadata. The analyst should verify the number of cells, the number of genes, and the presence of cluster assignments in the metadata. This verification prevents errors that arise from analyzing an outdated or incorrectly subsetted object.

## Verify object structure
dim(seurat_obj)
table(Idents(seurat_obj))
head(seurat_obj[])

The Idents() function returns the current cluster assignments. If clusters have not been assigned, the analyst must run clustering before marker identification. The metadata columns should include sample information, batch identifiers, and any other covariates relevant to the analysis.

Running FindAllMarkers for Cluster-Specific Genes

The FindAllMarkers() function identifies markers for each cluster by comparing cells in that cluster to all other cells. This is the most common approach for initial annotation because it produces a ranked list of candidate markers for every cluster in one operation.

## Find markers for all clusters
all_markers <- FindAllMarkers(
  seurat_obj,
  only.pos = TRUE,
  min.pct = 0.25,
  logfc.threshold = 0.25
)

The only.pos = TRUE argument restricts results to genes that are upregulated in the target cluster, which is appropriate for identifying positive markers. The min.pct argument sets the minimum fraction of cells in either group that must express the gene. The logfc.threshold sets the minimum log fold change between groups. These thresholds should be adjusted based on the dataset. For datasets with many cells, more stringent thresholds reduce the number of candidate markers to a manageable set. For datasets with few cells, less stringent thresholds prevent missing genuine markers.

Interpreting the Output Table

The output of FindAllMarkers() is a data frame with one row per gene per cluster. Key columns include:

ColumnDescription
p_valRaw p-value from the statistical test
avg_log2FCAverage log2 fold change between cluster and comparison group
pct.1Fraction of cells in the target cluster expressing the gene
pct.2Fraction of cells in the comparison group expressing the gene
p_val_adjAdjusted p-value after multiple testing correction
clusterCluster for which the gene is a marker

The pct.1 and pct.2 columns are particularly informative. A good marker shows high pct.1 and low pct.2, meaning it is expressed in most cells of the target cluster and few cells elsewhere. A gene with high pct.1 and high pct.2 may be a housekeeping gene that is broadly expressed, even if the fold change is statistically significant.

Selecting Markers by Statistical and Biological Criteria

After generating the marker table, the analyst must select genes for each cluster. Statistical criteria include adjusted p-value below a threshold such as 0.05, log2 fold change above a threshold such as 0.5 or 1, and expression in a substantial fraction of cells in the target cluster. Biological criteria include prior knowledge of cell type markers and the coherence of the marker panel.

A common approach is to filter the marker table and then select the top genes for each cluster:

## Filter for significant markers
significant_markers <- all_markers %>%
  filter(p_val_adj < 0.05) %>%
  group_by(cluster) %>%
  top_n(n = 10, wt = avg_log2FC)

The number of markers selected per cluster depends on the analysis goal. For initial annotation, five to ten markers per cluster are usually sufficient. For building a marker panel for future experiments, fewer genes may be preferred. The MAGNETO framework addresses this tradeoff by optimizing marker panels for both specificity and panel size, identifying the minimum number of genes needed to discriminate a cell type of interest [<a href="#ref-3">3</a>].

Using FindMarkers for Specific Comparisons

Sometimes the analyst needs to compare one cluster against another specific cluster instead of against all other cells. This is useful for distinguishing closely related cell types or subtypes. The FindMarkers() function supports this comparison:

## Compare cluster 1 against cluster 2
cluster1_vs_cluster2 <- FindMarkers(
  seurat_obj,
  ident.1 = "1",
  ident.2 = "2",
  only.pos = TRUE
)

This approach is valuable when two clusters share many markers and the analyst needs to identify the genes that differentiate them. For example, distinguishing between monocyte and macrophage subtypes may require pairwise comparisons because both populations express common myeloid markers.

Step-by-Step Marker Identification in Scanpy

Loading and Verifying the AnnData Object

Scanpy uses the AnnData object format, which stores expression data, cell metadata, and gene metadata in a single structure. The analyst should verify the shape of the object and the presence of cluster assignments in the observation metadata.

## Verify object structure
print(adata.shape)
print(adata.obs['leiden'].value_counts())

The shape attribute gives the number of cells and genes. The cluster assignments should be stored in the observation metadata, typically in a column named leiden or louvain depending on the clustering algorithm used.

Running rank_genes_groups for Cluster Markers

The sc.tl.rank_genes_groups() function identifies marker genes for each cluster. By default, it compares each cluster against all other cells using the Wilcoxon rank-sum test.

## Find markers for all clusters
sc.tl.rank_genes_groups(
    adata,
    groupby='leiden',
    method='wilcoxon',
    use_raw=False
)

The groupby argument specifies which metadata column contains cluster assignments. The method argument selects the statistical test. Options include wilcoxon, t-test, logreg, and others. The use_raw argument determines whether the test uses the raw counts stored in the AnnData object or the normalized data in the main expression matrix.

Extracting and Filtering Results

The results of rank_genes_groups() are stored in the AnnData object's uns attribute. The analyst can extract them into a pandas DataFrame for filtering and inspection:

## Extract results into a DataFrame
result = sc.get.rank_genes_groups_df(adata, group='0')
result.head()

The resulting DataFrame contains columns for gene names, log fold change, p-values, and adjusted p-values. Filtering follows the same logic as in Seurat:

## Filter for significant markers
significant = result[
    (result['pvals_adj'] < 0.05) &
    (result['logfoldchanges'] > 0.5)
]

Comparing Specific Groups in Scanpy

Scanpy also supports comparisons between specific groups using the groupby parameter with a list of groups:

## Compare cluster 0 against cluster 1
sc.tl.rank_genes_groups(
    adata,
    groupby='leiden',
    groups=['0'],
    reference='1',
    method='wilcoxon'
)

This comparison is useful for identifying genes that distinguish two closely related clusters. The groups parameter specifies the target group, and the reference parameter specifies the comparison group.

Visualization of Marker Genes

Violin Plots for Expression Distribution

Violin plots show the distribution of expression for a gene across clusters. They are useful for verifying that a marker gene is specifically expressed in the target cluster and for assessing the consistency of expression within the cluster.

In Seurat:

VlnPlot(seurat_obj, features = c("CD3D", "CD14", "MS4A1"))

In Scanpy:

sc.pl.violin(adata, keys=['CD3D', 'CD14', 'MS4A1'], groupby='leiden')

A good marker shows a clear elevation in one cluster with relatively tight distribution. A marker that shows broad expression across many clusters is less useful for annotation, even if the statistical test reports significance.

Feature Plots on Dimensionality Reduction

Feature plots project gene expression onto a dimensionality reduction embedding such as UMAP or t-SNE. These plots show the spatial distribution of expression across the cell landscape.

In Seurat:

FeaturePlot(seurat_obj, features = c("CD3D", "CD14", "MS4A1"))

In Scanpy:

sc.pl.umap(adata, color=['CD3D', 'CD14', 'MS4A1'])

Feature plots help the analyst confirm that marker expression is localized to specific clusters instead of scattered across the embedding. A marker that lights up cells in multiple distant clusters may indicate that those clusters share a common cell type or that the marker is not specific.

Dot Plots for Multi-Gene Comparison

Dot plots combine information about expression level and expression frequency across multiple genes and clusters. Each dot's size represents the fraction of cells expressing the gene, and the color represents the average expression level.

In Seurat:

DotPlot(seurat_obj, features = c("CD3D", "CD14", "MS4A1", "NKG7", "FCGR3A"))

In Scanpy:

sc.pl.dotplot(adata, var_names=['CD3D', 'CD14', 'MS4A1', 'NKG7', 'FCGR3A'], groupby='leiden')

Dot plots are particularly useful for comparing multiple candidate markers across all clusters simultaneously. They allow the analyst to see whether a panel of known markers aligns with the clustering structure.

Heatmaps of Top Markers

Heatmaps display the expression of the top markers for each cluster across all cells or across cluster averages. They provide a global view of the marker landscape and can reveal clusters that lack distinctive markers.

In Seurat:

DoHeatmap(seurat_obj, features = top_markers$gene)

In Scanpy:

sc.pl.heatmap(adata, var_names=top_markers, groupby='leiden')

A well-separated dataset shows clear blocks of elevated expression along the diagonal, with each cluster having a distinct set of markers. Clusters that show weak or diffuse marker expression may need re-clustering or may represent transitional states.

Statistical Tests and Their Selection

Wilcoxon Rank-Sum Test

The Wilcoxon rank-sum test is the default in both Seurat and Scanpy for marker identification. It is nonparametric, meaning it does not assume a normal distribution of expression values. The test ranks all cells by expression level and compares the sum of ranks between the two groups.

Advantages of the Wilcoxon test include robustness to outliers and no distributional assumptions. Disadvantages include reduced power for detecting small but consistent differences and sensitivity to differences in cell number between groups. For datasets with highly imbalanced cluster sizes, the test may produce false positives for genes expressed in a small fraction of cells in the larger group.

MAST Test

The MAST test models expression as a hurdle model with two components. The first component models whether a gene is detected in a cell using logistic regression. The second component models the expression level given detection using a Gaussian distribution. This approach accounts for the dropout phenomenon common in single-cell data.

The MAST test is available in Seurat through the test.use = "MAST" argument. Studies have used MAST for marker gene analysis in bladder cancer single-cell data, identifying core cell populations through marker-gene analysis with Seurat and the MAST test [<a href="#ref-2">2</a>]. The test is more computationally intensive than Wilcoxon but can provide better statistical power for genes with partial detection.

Negative Binomial Test

The negative binomial test models count data directly, accounting for overdispersion relative to the Poisson distribution. This test is appropriate when working with raw counts instead of normalized or transformed data. Seurat implements this test through the test.use = "negbinom" argument.

The negative binomial test is sensitive to the choice of latent variables included in the model. Analysts should include relevant covariates such as sequencing depth or batch to avoid confounding. The test is computationally demanding for large datasets.

Logistic Regression Approach

Scanpy offers a logistic regression method for marker identification through method = "logreg". This approach fits a logistic regression model to predict cluster membership from gene expression. The coefficients from the model indicate the contribution of each gene to distinguishing the cluster.

The logistic regression approach can handle complex relationships between genes and cluster membership but requires careful regularization to avoid overfitting. It is less commonly used than Wilcoxon for routine marker identification but can be valuable for datasets with strong batch effects or complex experimental designs.

Choosing the Appropriate Test

The choice of statistical test depends on the dataset size, the distribution of expression values, and the analysis goals. For initial exploration, the Wilcoxon test provides a fast and robust default. For datasets where dropout is a major concern, MAST may provide better sensitivity. For analyses using raw counts, the negative binomial test is appropriate.

Analysts should compare results from multiple tests for critical clusters. If different tests produce substantially different marker lists, the analyst should investigate why. Discrepancies may indicate that the cluster is not well separated or that the data violate assumptions of one test.

Interpreting Marker Gene Results

Distinguishing Specific Markers from Broadly Expressed Genes

A common pitfall in marker identification is selecting genes that are statistically significant but not biologically specific. Genes with high expression in many clusters may show significant p-values simply because of large cell numbers, even when the fold change is modest. The pct.1 and pct.2 columns help identify this problem.

A gene with pct.1 = 0.9 and pct.2 = 0.8 is expressed in most cells of both groups. The difference in expression may be statistically significant but is unlikely to be useful as a marker. A gene with pct.1 = 0.8 and pct.2 = 0.1 shows clear specificity and is a better marker candidate.

Handling Clusters Without Clear Markers

Some clusters may not show distinctive markers. This situation arises when a cluster represents a transitional state, a rare population with low expression, or a technical artifact. The analyst should investigate such clusters before assigning cell type labels.

Options for clusters without clear markers include increasing the clustering resolution to split the cluster into more homogeneous subclusters, decreasing the resolution to merge the cluster with a neighboring population, or examining the expression of known markers for suspected cell types even if they do not appear in the top marker list. The The Carpentries Lessons provide foundational programming training that helps analysts develop the data manipulation skills needed for these investigations.

Using Known Markers for Annotation

Marker identification results should be interpreted in the context of known biology. Databases of cell type markers, such as those available through NCBI Data Resources, help analysts verify that candidate markers align with established cell type identities.

For example, T cells typically express CD3D, CD3E, and CD3G. B cells express MS4A1 and CD79A. Macrophages express CD14 and FCGR3A. When these known markers appear in the top marker list for a cluster, the annotation is well supported. When they do not, the analyst should question whether the cluster represents the expected cell type or whether the clustering has separated cells in an unexpected way.

Validating Markers with External Data

Marker genes identified in one dataset should be validated in independent datasets when possible. The EMBL-EBI Training resources emphasize the importance of validation for reproducible research. Validation can involve checking whether the same markers identify the same cell types in a different dataset or comparing marker expression with protein-level data from flow cytometry or immunohistochemistry.

Studies of the human placenta have systematically compared cell-type-specific genes across multiple single-cell RNA sequencing studies, identifying consensus marker genes for trophoblast subtypes [<a href="#ref-4">4</a>]. This approach demonstrates the value of cross-study validation for establishing reliable markers.

Common Failure Patterns in Marker Identification

Failure Pattern 1: Overly Stringent Thresholds

Setting thresholds too high can eliminate genuine markers, particularly for rare cell types or genes with moderate expression. The analyst may find that some clusters have very few or no significant markers. This situation often arises when min.pct is set too high or logfc.threshold is set too high.

The solution is to examine the distribution of expression for suspected markers and adjust thresholds accordingly. For rare cell types, lower thresholds may be necessary to capture markers that are expressed in only a fraction of cells.

Failure Pattern 2: Overly Permissive Thresholds

Setting thresholds too low produces long lists of markers that include many genes with marginal significance. This makes annotation difficult because the analyst cannot distinguish genuine markers from noise.

The solution is to use adjusted p-values instead of raw p-values and to require a minimum fold change. The analyst should also examine the specificity of candidate markers using the pct.1 and pct.2 columns.

Failure Pattern 3: Ignoring Batch Effects

When batch effects are present but not corrected, marker identification may produce genes that distinguish batches instead of cell types. This problem is particularly acute when different batches contain different cell type compositions.

The solution is to perform data integration before clustering and marker identification. Integration methods align cells across batches while preserving biological variation. The nf-core Documentation describes pipeline configurations that include integration steps, providing a reference for reproducible analysis.

Failure Pattern 4: Using Markers from Unfiltered Data

If quality control filtering is insufficient, clusters may form around technical artifacts such as doublets, dying cells, or ambient RNA. Markers identified for these clusters reflect technical variation instead of biology.

The solution is to revisit quality control and remove problematic cells before re-running clustering and marker identification. The Galaxy Training Network provides tutorials on quality control that help analysts identify and remove technical artifacts.

Failure Pattern 5: Overinterpreting Single Markers

Assigning cell type identity based on a single marker gene is risky. Many markers are shared across related cell types, and expression of a single gene is rarely sufficient for definitive annotation.

The solution is to use panels of markers for each cell type and to require that multiple markers align with the annotation. The MAGNETO framework addresses this need by constructing optimal marker panels that balance specificity and panel size [<a href="#ref-3">3</a>].

Limitations of Marker Gene Identification

Statistical Limitations

Marker gene identification tests for differential expression, which is not the same as biological relevance. A gene can be statistically significant without being functionally important for the cell type. Conversely, functionally important genes may not show statistically significant differential expression if they are regulated post-transcriptionally or if their expression is variable across cells.

The multiple testing problem is substantial in single-cell analysis. Testing tens of thousands of genes across dozens of clusters produces millions of comparisons. Adjusted p-values control the false discovery rate but do not eliminate false positives entirely. Analysts should treat marker lists as candidates for further investigation instead of definitive conclusions.

Biological Limitations

Marker genes are context-dependent. A gene that marks a cell type in one tissue may not mark the same cell type in another tissue. For example, markers for immune cells in blood may differ from markers for immune cells in the tumor microenvironment. Studies of the immune microenvironment in colorectal cancer have shown that resistance-related cell types and genes identified by single-cell analysis require verification in clinical samples [<a href="#ref-5">5</a>].

Cell state also affects marker expression. Cells in different activation states, differentiation stages, or cell cycle phases may express different genes. A cluster that appears homogeneous by clustering may contain cells in multiple states, and markers identified for the cluster may reflect the dominant state instead of the cell type identity.

Technical Limitations

Single-cell RNA sequencing captures only a fraction of the transcriptome per cell. Genes with low expression may be missed entirely, and genes with moderate expression may show high variability across cells. This dropout phenomenon reduces the power to detect markers for genes with low or moderate expression.

Ambient RNA contamination can create false markers by introducing transcripts from lysed cells into the droplet or well containing a different cell. This problem is particularly acute for highly expressed genes, which can appear as markers for many clusters. Computational methods for removing ambient RNA are available but add complexity to the analysis.

Interpretation Limitations

Marker gene lists require biological interpretation, which depends on prior knowledge. Analysts who are not familiar with the cell types expected in their tissue may struggle to interpret marker lists. Cross-referencing with databases and published literature is essential.

The NCBI Data Resources provide access to gene annotation and expression databases that support interpretation. The EMBL-EBI Training resources include modules on biological interpretation of single-cell data that help analysts develop these skills.

Quality Control and Reproducibility

Documenting Analysis Parameters

Reproducible marker identification requires documenting all parameters used in the analysis. This includes quality control thresholds, normalization methods, feature selection parameters, clustering resolution, and marker identification settings. The nf-core Documentation emphasizes the importance of parameter documentation for reproducible pipelines.

Analysts should record the version numbers of all software packages used. Seurat and Scanpy are updated regularly, and results can change between versions. Recording versions allows other researchers to reproduce the analysis or to understand why results differ.

Using Version Control

Version control systems such as Git help track changes to analysis scripts and document the evolution of the analysis. The The Carpentries Lessons provide training in version control that is directly applicable to single-cell analysis workflows.

Version control is particularly important when multiple analysts work on the same dataset or when the analysis is revisited after publication. It allows the team to identify which version of the analysis produced specific results and to revert to earlier versions if needed.

Validating Results Across Methods

Marker identification results should be validated using multiple approaches. This includes comparing results from different statistical tests, checking marker expression visually, and comparing with known markers from the literature. Studies that integrate single-cell and bulk RNA sequencing data often validate marker genes using external datasets and experimental methods such as quantitative PCR and Western blotting [<a href="#ref-6">6</a>].

Validation can also involve checking whether markers identified in one dataset reproduce in an independent dataset. The systematic review of placental single-cell studies identified consensus markers by comparing cell-type-specific genes across multiple studies [<a href="#ref-4">4</a>]. This approach provides stronger evidence than markers identified in a single dataset.

Recording Decisions and Rationale

Analysts should record also the parameters used but also the rationale for decisions. This includes why specific thresholds were chosen, why certain clusters were merged or split, and why particular markers were selected for annotation. This documentation supports transparency and allows other researchers to understand the analysis.

The Bioconductor project provides workflows that emphasize reproducible analysis practices. Following these workflows helps analysts produce results that can be verified and extended by others.

Records and Measurements for Marker Gene Analysis

Essential Records to Maintain

Record TypeContentsPurpose
Quality control metricsCell counts, gene counts, mitochondrial fraction before and after filteringDocument data quality and filtering decisions
Clustering parametersResolution, number of principal components, algorithmEnable reproduction of clustering results
Marker identification settingsTest method, thresholds, comparison groupsDocument statistical choices
Marker tablesFull output from marker identification with all columnsPreserve complete results for re-analysis
Annotation decisionsCell type labels, supporting markers, rationaleDocument biological interpretation

Measurements to Track

The analyst should track the number of clusters, the number of markers per cluster, and the fraction of clusters that can be confidently annotated. These measurements provide a summary of the analysis quality. Datasets where most clusters have clear markers and confident annotations are more likely to yield reliable biological conclusions.

The median gene count and UMI count per cell provide context for interpreting marker identification results. Datasets with low gene counts per cell may have reduced power for marker detection. Studies of glioma have reported median gene counts of 5,437 and UMI counts of 14,207 per cell as quality indicators [<a href="#ref-7">7</a>].

Escalation Criteria

The analyst should escalate to a supervisor or collaborator when certain situations arise. These include clusters that cannot be annotated despite extensive investigation, marker lists that conflict with established biology, and results that differ substantially between statistical tests. Escalation is appropriate when the analyst lacks the biological expertise to interpret the results or when the results have implications for downstream experiments or clinical decisions.

Safety and Regulatory Context

Data Privacy and Ethics

Single-cell RNA sequencing data from human subjects are subject to privacy and ethical regulations. Analysts must ensure that data are handled in compliance with institutional review board approvals and data use agreements. Public datasets obtained from repositories such as the Gene Expression Omnibus should be used in accordance with their data use restrictions.

The NCBI Data Resources provide access to public datasets with documented data use policies. Analysts should review these policies before downloading and using data.

Reporting Standards

Marker gene identification results should be reported with sufficient detail for others to evaluate the analysis. This includes reporting the software versions, parameters, and statistical methods used. The EMBL-EBI Training resources provide guidance on reporting standards for bioinformatics analyses.

Publications should include the marker tables as supplementary data or in a public repository. This allows other researchers to verify the results and to use the markers for their own analyses.

Professional Escalation

Analysts should seek expert advice when marker identification results will be used for clinical decisions or therapeutic targeting. The interpretation of marker genes in disease contexts requires domain expertise. For example, identifying tumor stem cell markers in lung adenocarcinoma has implications for prognosis and treatment, and the results should be interpreted by researchers with oncology expertise [<a href="#ref-8">8</a>].

Similarly, marker genes identified for diagnostic purposes, such as the immune and oxidative stress-related markers in diabetic nephropathy, should be validated by clinical researchers before any diagnostic application [<a href="#ref-9">9</a>].

Frequently Asked Questions

What is the difference between a marker gene and a differentially expressed gene?

A marker gene is a differentially expressed gene that distinguishes one cluster from others and is used for cell type annotation. All marker genes are differentially expressed, but not all differentially expressed genes are useful markers. A gene can be differentially expressed between two clusters without being specific to either cluster, for example if it shows graded expression across many clusters. Marker genes are selected for their specificity and their ability to discriminate cell types.

How many marker genes should I select for each cluster?

The number depends on the analysis goal. For initial cell type annotation, five to ten markers per cluster are usually sufficient. For building a marker panel for future experiments such as flow cytometry or spatial transcriptomics, fewer genes may be preferred. The MAGNETO framework optimizes marker panels for both specificity and panel size, identifying the minimum number of genes needed to discriminate a cell type of interest [<a href="#ref-3">3</a>].

Why do I get different markers when I use different statistical tests?

Different statistical tests make different assumptions about the data. The Wilcoxon test is nonparametric and ranks expression values. MAST models detection and expression separately. The negative binomial test models count data directly. These differences can produce different marker lists, particularly for genes with partial detection or skewed expression distributions. Comparing results across tests can help identify robust markers that are significant under multiple methods.

What should I do if a cluster has no clear markers?

Investigate the cluster before assigning a label. Check whether the cluster represents a transitional state, a rare population, or a technical artifact. Try increasing the clustering resolution to split the cluster into more homogeneous subclusters. Try decreasing the resolution to merge the cluster with a neighboring population. Examine the expression of known markers for suspected cell types even if they do not appear in the top marker list.

Can I use marker genes from published studies to annotate my clusters?

Yes, published markers are valuable for annotation. However, markers are context-dependent and may not transfer perfectly across tissues, species, or experimental conditions. Use published markers as a starting point and verify that they show the expected expression patterns in your data. Cross-study comparisons, such as the systematic review of placental single-cell studies that identified consensus markers for trophoblast subtypes, provide more reliable markers than single-study reports [<a href="#ref-4">4</a>].

How do I handle batch effects before marker identification?

Perform data integration before clustering and marker identification. Integration methods align cells across batches while preserving biological variation. After integration, clusters should reflect biology instead of batch. If integration is not performed, marker identification may produce genes that distinguish batches instead of cell types. The nf-core Documentation describes pipeline configurations that include integration steps.

What is the role of adjusted p-values in marker selection?

Adjusted p-values control the false discovery rate when testing many genes simultaneously. Single-cell datasets typically involve testing tens of thousands of genes, so raw p-values are not reliable. Use adjusted p-values for filtering markers. A common threshold is an adjusted p-value below 0.05, but the appropriate threshold depends on the dataset and the analysis goals.

How do I validate my marker genes experimentally?

Experimental validation can involve quantitative PCR, Western blotting, immunohistochemistry, or flow cytometry. Studies of lung adenocarcinoma have validated marker gene expression using quantitative PCR and Western blotting [<a href="#ref-6">6</a>]. Studies of colorectal cancer have used immunohistochemistry and immunofluorescence to verify key genes and cell marker molecules [<a href="#ref-5">5</a>]. The choice of validation method depends on the gene, the tissue, and the available antibodies.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

[1] [Integrated analysis of single-cell RNA-seq, bulk RNA-seq, Mendelian randomization, and eQTL reveals T cell-related nomogram model and subtype classification in rheumatoid arthritis.](https://pubmed.ncbi.nlm.nih.gov/38962008). Frontiers in immunology, 2024. [2] [Integrating Multi-dimensional RNA Sequencing to Construct a Prognostic Risk Model for Bladder Cancer](https://doi.org/10.21203/rs.3.rs-9821610/v1). 2026. [3] [MAGNETO: Cell type marker panel generator from single-cell transcriptomic data](https://doi.org/10.1016/j.jbi.2023.104510). Journal of Biomedical Informatics, 2023. [4] [Revealing the molecular landscape of human placenta: a systematic review and meta-analysis of single-cell RNA sequencing studies.](https://pubmed.ncbi.nlm.nih.gov/38478759). Human reproduction update, 2024. [5] [Single-cell sequencing reveals the immune microenvironment landscape related to anti-PD-1 resistance in metastatic colorectal cancer with high microsatellite instability](https://doi.org/10.1186/s12916-023-02866-y). BMC Medicine, 2023. [6] [Identification of novel gene signature for lung adenocarcinoma by machine learning to predict immunotherapy and prognosis.](https://pubmed.ncbi.nlm.nih.gov/37583701). Frontiers in immunology, 2023. [7] [Single-cell transcriptomic profiling uncovers key molecular signatures in glioma pathogenesis.](https://doi.org/10.3389/fgene.2026.1818742). 2026. [8] [Integration of single-cell and bulk RNA sequencing to identify a distinct tumor stem cells and construct a novel prognostic signature for evaluating prognosis and immunotherapy in LUAD.](https://pubmed.ncbi.nlm.nih.gov/39987127). Journal of translational medicine, 2025. [9] [Identification and validation of immune and oxidative stress-related diagnostic markers for diabetic nephropathy by WGCNA and machine learning.](https://pubmed.ncbi.nlm.nih.gov/36911691). Frontiers in immunology, 2023.

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.