# Choosing the Right Resolution Parameter for Graph-Based Clustering in Single-Cell RNA-Seq: A Data-Driven Approach


## Key Takeaways

- Graph-based clustering algorithms (Louvain, Leiden) in scRNA-seq rely on a resolution parameter that directly dictates cluster granularity; inappropriate values can lead to merging distinct cell types (Type II error) or splitting homogeneous populations (Type I error).
- Default resolution values (e.g., 0.8, 1.0) are often insufficient, as optimal settings are dataset-dependent, influenced by biological complexity, sequencing depth, and cell number, necessitating a data-driven approach beyond intuition.
- Cluster stability assessment, particularly subsampling-based methods like chooseR, provides a principled criterion by evaluating the reproducibility of cluster assignments across data perturbations, identifying resolutions that yield robust cell groupings.
- Validation of stability-selected resolutions with canonical marker gene expression is critical to ensure biological relevance, as stable clusters may not always reflect genuine biological populations without this confirmatory step.
- Downstream analysis goals, such as differential expression or trajectory inference, must inform resolution selection, as granularity impacts statistical power, the identification of intermediate states, and the interpretability of cell type annotations.
- Computational cost of stability assessment can be significant for large datasets, prompting consideration of optimized metrics or alternative approaches like intrinsic goodness metrics or Dirichlet process mixture models, each with their own trade-offs in biological interpretability.

---

Graph-based clustering methods such as Louvain and Leiden are standard tools for identifying cell populations in single-cell RNA sequencing (scRNA-seq) data. These methods depend on a resolution parameter that directly controls the number and granularity of clusters returned by the algorithm. Selecting an inappropriate resolution can lead to missed cell types or artificial splitting of homogeneous populations. This article provides a systematic, data-driven approach for choosing the resolution parameter based on cluster stability assessment, marker gene validation, and biological context interpretation, with practical code examples for implementation.

## Understanding the Resolution Parameter in Graph-Based Clustering

Graph-based clustering in scRNA-seq analysis constructs a graph where cells are nodes and edges represent transcriptional similarity between cells. The Louvain and Leiden algorithms partition this graph into communities by optimizing a modularity function. The resolution parameter acts as a scaling factor that determines how many communities the algorithm identifies. Lower resolution values produce fewer, larger clusters, while higher values produce more numerous, smaller clusters.

The relationship between resolution and cluster number is not arbitrary. Research on modularity clustering has shown that the resolution parameter implicitly determines the number of clusters inferred, and an improperly chosen value can lead to erroneous or missed cell types corresponding to type I or type II errors. This means that both over-clustering and under-clustering carry biological consequences that affect downstream interpretation.

The challenge is that no universal resolution value works across all datasets. The optimal setting depends on the biological complexity of the sample, the sequencing depth, the number of cells captured, and the transcriptional distinctness of the populations being studied. A resolution that works well for a peripheral blood mononuclear cell dataset may produce poor results for a complex tissue sample with many rare cell types.

## The Problem with Default Resolution Values

Many analysis pipelines default to a resolution parameter of 0.8 or 1.0 without systematic evaluation. This practice carries substantial risk. The default value may be appropriate for some datasets but entirely unsuitable for others. Researchers often develop intuition about appropriate parameter ranges through repeated analysis of similar data, but this experience-based approach is difficult to transfer between datasets and can be daunting for new users of any given workflow.

The consequences of an incorrect resolution choice manifest in several ways. At low resolution, distinct cell types may be merged into a single cluster, obscuring biologically meaningful heterogeneity. At high resolution, a homogeneous cell population may be fragmented into artificial subclusters that do not represent real biological states. Both scenarios lead to incorrect conclusions about cell type composition, differential expression results, and trajectory inference.

Static visualization techniques compound this problem by displaying results for only a single resolution parameter. Analysts often evaluate more than one resolution parameter during exploratory analysis but then report only one in the final results. This practice hides the uncertainty inherent in cluster assignment and makes it difficult for readers to assess whether the reported clustering is robust.

## At a Glance: Resolution Selection Methods Compared

| Method | Primary Input | Key Output | Strengths | Limitations |
|--------|--------------|------------|-----------|-------------|
| Subsampling-based stability (chooseR) | Candidate resolution range, subsampling iterations | Robustness score per cluster, stability trajectory | Works across clustering algorithms, identifies cluster-specific reliability | Computationally expensive for large datasets |
| Intrinsic goodness metrics | Candidate resolution range, clustering results | Davies-Bouldin index, silhouette width | Fast to compute, no external labels needed | Optimizes mathematical properties, not biological relevance |
| Dirichlet process mixture models | Expression matrix only | Inferred cluster number | Removes need for user-specified cluster count | Dataset dependent performance, may inflate cluster numbers |
| Active learning | Expression matrix, expert labels for subset of cells | Biologically guided clustering | Directly addresses interpretability, outperforms unsupervised methods with limited labels | Requires expert input, not fully automated |
| Deep learning integration (scVAG) | Expression matrix, network architecture | Latent representation and clusters | Nonlinear dimensionality reduction, improved accuracy | Additional hyperparameters, complex model tuning |

## Cluster Stability as a Primary Selection Criterion

Cluster stability provides a principled basis for resolution selection. The concept is straightforward: a good clustering should be reproducible when the analysis is repeated with minor perturbations to the input data. If small changes in the data or algorithm parameters produce dramatically different cluster assignments, the clustering is not stable and should not be trusted.

Subsampling-based approaches formalize this concept. The chooseR method, described in the BMC Bioinformatics literature, uses bootstrapped iterative clustering across a range of parameters to simultaneously guide parameter selection and characterize cluster robustness. This approach works by repeatedly subsampling the cell population, reclustering the subsampled data at each candidate resolution, and measuring how consistently cells cluster together across iterations.

The output of this process is a robustness score for each cluster at each resolution. Clusters that consistently recover the same cell groupings across bootstrap iterations receive high robustness scores. Clusters that appear in some iterations but not others receive low scores, indicating that they may be artifacts of the specific data sample instead of genuine biological populations.

This approach has been validated on both well-characterized datasets such as human peripheral blood mononuclear cells and more complex datasets such as mouse spinal cord. In each case, the stability-based parameter selection identified resolutions that produced biologically relevant clusters. The method is flexible enough to work across different clustering algorithms and workflows, making it a general solution to the resolution selection problem.

## Implementing a Stability-Based Resolution Selection Workflow

A practical workflow for resolution selection combines stability assessment with biological validation. The following steps provide a systematic approach that can be adapted to different analysis pipelines.

### Step 1: Define a Candidate Resolution Range

Begin by defining a range of resolution values to test. A reasonable starting range might span from 0.1 to 2.0, with increments of 0.1 or 0.2. The exact range should be informed by the expected biological complexity of the sample. A sample expected to contain only a few major cell types may only require testing low resolutions, while a sample expected to contain many rare subpopulations may require testing higher values.

The number of cells in the dataset also influences the appropriate range. Larger datasets can support higher resolution values because they contain more cells per potential cluster. Smaller datasets may become over-clustered at high resolutions because individual clusters would contain too few cells for meaningful analysis.

### Step 2: Perform Subsampling-Based Stability Assessment

For each candidate resolution, perform multiple clustering runs on subsampled versions of the data. A common approach is to subsample 80 percent of the cells for each iteration and perform 20 to 50 iterations per resolution. The subsampling should be performed without replacement to ensure that each iteration uses a different subset of cells.

After all iterations are complete, calculate a stability metric for each cluster. One approach is to measure the frequency with which pairs of cells are assigned to the same cluster across iterations. Cells that are consistently co-clustered have high pairwise stability. Clusters can then be scored by averaging the pairwise stability of their constituent cells.

The chooseR method provides a simple robustness score for each cluster that facilitates assessment of cluster quality. This score can be used to compare resolutions directly: the optimal resolution produces clusters that are consistently stable across subsampling iterations.

### Step 3: Examine the Stability Trajectory Across Resolutions

Plot the stability metrics against resolution values to identify patterns. Typically, stability will be high at very low resolutions where only major cell types are identified, decrease as resolution increases and clusters become more granular, and then drop sharply when resolution becomes high enough to fragment stable populations.

Look for plateaus in the stability trajectory. A plateau indicates a range of resolutions where the clustering is relatively stable. The optimal resolution is often at the upper end of a plateau, just before stability begins to decline. This position captures the maximum number of biologically meaningful clusters while avoiding the fragmentation that occurs at higher resolutions.

### Step 4: Validate Selected Resolutions with Marker Genes

Stability analysis identifies resolutions that produce reproducible clusters, but reproducibility alone does not guarantee biological relevance. Each candidate resolution should be validated by examining whether the resulting clusters express known marker genes for expected cell types.

For each cluster at each candidate resolution, examine the expression of canonical markers for the cell types expected in the sample. A good clustering will produce clusters that are enriched for specific marker genes and that separate cleanly from other clusters based on marker expression. Clusters that show mixed marker expression or that fail to express any expected markers may represent artifacts.

This validation step is essential because stability-based methods can identify reproducible but biologically meaningless clusters. The active learning approach described in the Laboratory Investigation literature highlights this problem: unsupervised clustering can generate results with poor biological interpretability, which is generally undesired by biologists. Incorporating marker gene validation into the resolution selection process helps ensure that the chosen resolution produces biologically meaningful clusters.

### Step 5: Consider Biological Context and Downstream Analysis Goals

The optimal resolution depends on the biological question being addressed. A study focused on major cell type composition may require only low resolution clustering to identify broad categories such as T cells, B cells, and monocytes. A study focused on rare subpopulations or developmental intermediates may require higher resolution to resolve these finer distinctions.

The downstream analysis plan should inform resolution selection. If the clustering will be used for differential expression analysis between cell types, the resolution should produce clusters with sufficient cell numbers for statistical power. If the clustering will be used for trajectory inference, the resolution should preserve intermediate cell states instead of merging them into adjacent clusters.

Consensus clustering approaches can help address the challenge of selecting an appropriate number of clusters. Methods such as TiC2D adopt consensus clustering strategies to precisely cluster cells for trajectory inference, demonstrating that the choice of clustering granularity directly affects the quality of downstream analyses.

## Code Example: Stability-Based Resolution Selection in R

The following R code demonstrates a practical implementation of subsampling-based stability assessment using the Seurat workflow. This example assumes that a Seurat object has already been created and processed through normalization and principal component analysis.

```r
## Load required libraries
library(Seurat)
library(tidyverse)

## Function to assess cluster stability across resolutions
assess_resolution_stability <- function(seurat_obj,
                                        resolutions = seq(0.1, 2.0, by = 0.1),
                                        n_iterations = 30,
                                        subsample_fraction = 0.8) {

  # Store stability results
  stability_results <- list()

  for (res in resolutions) {

    # Initialize co-clustering matrix
    n_cells <- ncol(seurat_obj)
    co_cluster_count <- matrix(0, nrow = n_cells, ncol = n_cells)
    iteration_count <- 0

    for (i in 1:n_iterations) {

      # Subsample cells
      n_subsample <- floor(n_cells * subsample_fraction)
      subsampled_cells <- sample(1:n_cells, n_subsample)

      # Create subsampled Seurat object
      subsampled_obj <- subset(seurat_obj, cells = colnames(seurat_obj)[subsampled_cells])

      # Find neighbors and cluster
      subsampled_obj <- FindNeighbors(subsampled_obj, dims = 1:30)
      subsampled_obj <- FindClusters(subsampled_obj, resolution = res)

      # Record co-clustering for subsampled cells
      cluster_assignments <- as.numeric(Idents(subsampled_obj))
      cell_names <- colnames(subsampled_obj)

      for (j in 1:length(cell_names)) {
        for (k in 1:length(cell_names)) {
          if (cluster_assignments[j] == cluster_assignments[k]) {
            idx_j <- which(colnames(seurat_obj) == cell_names[j])
            idx_k <- which(colnames(seurat_obj) == cell_names[k])
            co_cluster_count[idx_j, idx_k] <- co_cluster_count[idx_j, idx_k] + 1
          }
        }
      }

      iteration_count <- iteration_count + 1
    }

    # Calculate stability as proportion of iterations where cells co-cluster
    co_cluster_frequency <- co_cluster_count / iteration_count

    # Calculate mean stability for the resolution
    # Only consider upper triangle to avoid double counting
    upper_triangle <- upper.tri(co_cluster_frequency)
    mean_stability <- mean(co_cluster_frequency[upper_triangle])

    stability_results[[as.character(res)]] <- mean_stability
  }

  return(stability_results)
}

## Run stability assessment
stability_results <- assess_resolution_stability(pbmc_seurat)

## Convert to data frame for plotting
stability_df <- data.frame(
  resolution = as.numeric(names(stability_results)),
  mean_stability = unlist(stability_results)
)

## Plot stability trajectory
ggplot(stability_df, aes(x = resolution, y = mean_stability)) +
  geom_line() +
  geom_point() +
  labs(x = "Resolution Parameter",
       y = "Mean Cluster Stability",
       title = "Cluster Stability Across Resolutions") +
  theme_minimal()
```

This code provides a starting point for stability assessment. In practice, the implementation should be optimized for computational efficiency. The pairwise co-clustering calculation becomes expensive for large datasets, and alternative approaches such as cluster-level stability metrics may be more practical.

## Alternative Approaches to Resolution Selection

Stability-based selection is not the only method for choosing the resolution parameter. Several alternative approaches exist, each with distinct strengths and limitations.

### Intrinsic Goodness Metrics

Intrinsic clustering quality metrics evaluate cluster quality without reference to external labels. These metrics assess properties such as cluster compactness, separation between clusters, and internal consistency. The Frontiers in Bioinformatics literature describes optimization of clustering parameters using intrinsic goodness metrics for single-cell RNA analysis.

Common intrinsic metrics include the Davies-Bouldin index, which measures the average similarity between each cluster and its most similar cluster, and the silhouette width, which measures how similar each cell is to its own cluster compared to other clusters. These metrics can be calculated across a range of resolutions, and the resolution producing the best metric value can be selected.

The limitation of intrinsic metrics is that they optimize mathematical properties of the clustering instead of biological relevance. A clustering that scores well on compactness and separation may still fail to separate biologically meaningful populations or may split homogeneous populations based on technical variation.

### Dirichlet Process Mixture Models

Dirichlet process mixture models offer an alternative to resolution parameter selection by attempting to infer the number of clusters directly from the data. The hierarchical Dirichlet process (HDP) extends the latent Dirichlet allocation (LDA) approach to allow an unbounded number of clusters, potentially eliminating the need for the user to specify the number of clusters as an input parameter.

Research comparing LDA and HDP for scRNA-seq clustering found that performance is dataset dependent. In some cases, HDP produced more appropriate clustering than the best LDA clustering with a fixed number of clusters. In other cases, HDP tended to inflate the number of clusters, producing more clusters than biologically justified. This variability means that even methods designed to automatically determine cluster number require careful evaluation.

### Active Learning Approaches

Active learning methods incorporate biological knowledge into the clustering process by allowing the algorithm to query the user for labels on a subset of cells. The active learning framework described in the Laboratory Investigation literature demonstrates that this approach can outperform unsupervised clustering methods with fewer than 1000 labeled cells.

The advantage of active learning is that it directly addresses the biological interpretability problem. By incorporating expert knowledge about cell types, the clustering is guided toward biologically meaningful partitions. The limitation is that it requires expert input, which may not be available in all analysis contexts.

### Deep Learning-Based Approaches

Deep learning methods for single-cell clustering integrate dimensionality reduction and clustering into a unified framework. The scVAG approach combines variational autoencoders with graph attention autoencoders for enhanced single-cell clustering, replacing linear PCA with nonlinear dimensionality reduction better suited for scRNA-seq data.

These methods can improve clustering accuracy compared to traditional approaches, but they introduce additional complexity in parameter selection. The deep learning models themselves have hyperparameters that must be tuned, and the clustering resolution may still need to be specified.

## Visualization Tools for Resolution Exploration

Visualization plays a critical role in resolution selection. The standard approach of projecting cells into two dimensions using UMAP and coloring by cluster assignment provides an intuitive view of clustering results. UMAP has been shown to provide fast run times, high reproducibility, and meaningful organization of cell clusters compared to other dimensionality reduction tools.

However, static UMAP visualizations only show results for a single resolution. Interactive tools that allow exploration across resolutions provide a more complete picture of how clustering changes with the resolution parameter.

Cell Layers is an interactive Sankey tool designed for quantitative investigation of gene expression, co-expression, biological processes, and cluster integrity across clustering resolutions. This tool enhances interpretability by linking molecular data and cluster evaluation metrics, providing insight into cell populations that may not be apparent from static visualizations.

The Sankey diagram format is particularly useful for resolution exploration because it shows how cells flow between clusters as the resolution changes. This visualization makes it easy to identify clusters that are stable across resolutions and clusters that fragment or merge as resolution changes.

## Common Failure Patterns in Resolution Selection

Several recurring problems appear when researchers select resolution parameters without systematic evaluation.

### Over-Clustering at High Resolution

High resolution values can fragment homogeneous cell populations into artificial subclusters. This problem is particularly common in datasets with high technical noise or dropout, where cells from the same biological population may show substantial transcriptional variation. The resulting subclusters may show subtle expression differences that are not biologically meaningful but are interpreted as distinct cell states.

Over-clustering is especially problematic for rare cell types. A rare population that should form a single cluster may be split into multiple clusters at high resolution, each containing too few cells for reliable downstream analysis. The resolution limit of modularity clustering, described in the IEEE Transactions on Computational Biology and Bioinformatics literature, establishes that the minimum resolution at which a subgraph is split is inversely proportional to the frequency of the subgraph within the graph. This means that rare populations are more susceptible to over-clustering than abundant populations.

### Under-Clustering at Low Resolution

Low resolution values can merge distinct cell types into a single cluster. This problem occurs when transcriptionally similar but functionally distinct populations are grouped together. The resulting cluster may show mixed marker expression that obscures the presence of multiple cell types.

Under-clustering is particularly problematic for developmental systems where cells exist on a continuum of states. Low resolution may merge intermediate states with their neighboring states, obscuring the trajectory structure that is critical for understanding differentiation processes.

### Ignoring Cluster-Specific Stability

A common failure is to select a resolution based on overall clustering quality without examining the stability of individual clusters. The optimal resolution for one cluster may not be optimal for another. Some clusters may be stable across a wide range of resolutions, while others may only appear at specific resolution values.

The chooseR approach addresses this issue by providing robustness scores for individual clusters. This cluster-specific information allows researchers to identify which clusters are reliable and which may be artifacts of the specific resolution choice.

### Selecting Resolution Before Quality Control

Resolution selection should occur after appropriate quality control of the single-cell data. Poor quality cells, doublets, and ambient RNA contamination can all affect clustering results and lead to incorrect resolution choices. The NCBI provides resources for understanding sequence data quality and analysis approaches that can inform quality control decisions.

Quality control should include filtering of low-quality cells based on metrics such as total counts, number of detected genes, and mitochondrial fraction. Doublet detection should be performed to remove cells that represent two or more cells captured together. These steps should be completed before clustering and resolution selection.

## Records and Documentation for Reproducibility

Reproducible resolution selection requires careful documentation of the decision process. The following records should be maintained for each analysis:

### Resolution Selection Record

Document the range of resolutions tested, the stability metrics calculated for each resolution, and the rationale for selecting the final resolution. Include the specific stability values and any marker gene validation results that informed the decision.

### Cluster Stability Scores

Record the stability score for each cluster at the selected resolution. This information is valuable for interpreting downstream results, as clusters with low stability should be interpreted with caution.

### Marker Gene Validation Results

Document which marker genes were examined for each cluster and the expression patterns observed. This record provides evidence that the selected resolution produces biologically meaningful clusters.

### Software and Parameter Documentation

Record the exact software versions and parameters used for clustering. This includes the clustering algorithm (Louvain or Leiden), the number of principal components used for graph construction, the nearest neighbor parameters, and the resolution value.

The nf-core documentation emphasizes the importance of reproducible workflow standards in genomic analysis. Following community standards for pipeline configuration and documentation helps ensure that clustering results can be reproduced and compared across studies.

### Version Control

Use version control for analysis scripts and configuration files. The Carpentries lessons provide foundational training in version control with Git that is directly applicable to managing analysis code. Version control ensures that the exact code used for clustering can be retrieved and examined.

## Limitations of Stability-Based Approaches

Stability-based resolution selection has several limitations that should be acknowledged.

### Computational Cost

Subsampling-based stability assessment requires multiple clustering runs for each candidate resolution. For large datasets, this can be computationally expensive. A dataset with 100,000 cells and 20 candidate resolutions would require 600 clustering runs if 30 iterations are performed per resolution. This computational burden may be prohibitive for some analysis environments.

Optimization strategies include reducing the number of candidate resolutions, reducing the number of iterations, or using cluster-level stability metrics that do not require pairwise co-clustering calculations.

### Stability Does Not Equal Biological Relevance

A clustering can be highly stable but biologically meaningless. Stability only measures reproducibility, not biological validity. The marker gene validation step is essential to ensure that stable clusters correspond to real biological populations.

### Sensitivity to Preprocessing Choices

Stability assessment is sensitive to the preprocessing steps performed before clustering. Different normalization methods, feature selection approaches, and dimensionality reduction choices can affect clustering results and therefore stability metrics. The stability assessment should be performed within the context of a fixed preprocessing pipeline.

### Difficulty with Continuous Cell States

Stability-based approaches assume that cells can be discretely partitioned into clusters. For systems with continuous cell states, such as differentiation trajectories, the concept of stable clusters may not apply cleanly. Cells along a continuum may be assigned to different clusters in different subsampling iterations, leading to low stability scores even when the clustering is biologically appropriate.

## Integration with Downstream Analyses

The resolution parameter choice affects all downstream analyses that depend on cluster assignments. Understanding these effects is important for interpreting results.

### Differential Expression Analysis

Differential expression analysis between clusters is directly affected by resolution choice. At low resolution, differential expression tests compare broad cell type categories, identifying genes that distinguish major populations. At high resolution, tests compare finer subpopulations, identifying genes that distinguish closely related states.

The interpretation of differential expression results must account for the resolution used. Genes identified as differentially expressed between subclusters at high resolution may not show significant differences when those subclusters are merged at lower resolution.

### Trajectory Inference

Trajectory inference methods reconstruct developmental processes from single-cell data. The clustering resolution affects trajectory inference by determining which cells are considered distinct states. Consensus clustering approaches such as TiC2D have been developed specifically to improve trajectory inference through more precise clustering.

High resolution clustering can help identify intermediate states along a trajectory, but it can also fragment the trajectory into too many discrete states. Low resolution clustering may merge intermediate states, obscuring the continuous nature of the developmental process.

### Cell Type Annotation

Cell type annotation relies on cluster assignments. The resolution choice determines the granularity of cell type labels that can be assigned. At low resolution, clusters may be annotated at the level of major lineages such as T cells or myeloid cells. At higher resolution, more specific annotations such as CD4+ naive T cells or classical monocytes become possible.

The annotation process should consider cluster stability. Clusters with low stability scores should be annotated with caution, as they may not represent reproducible cell populations.

### Alternative Splicing Analysis

Recent advances in single-cell RNA-seq enable analysis of alternative splicing at single-cell resolution. The choice of clustering resolution affects splicing analysis because splicing patterns may differ between cell subpopulations that are merged at low resolution. Methods such as scQuint have been developed for alternative splicing analysis in single-cell data, and their performance depends on appropriate clustering of cells.

## Quality Control and Preprocessing Considerations

The quality of clustering results depends heavily on the quality of the input data. Several preprocessing decisions directly affect resolution selection.

### Normalization and Feature Selection

Different normalization methods can produce different clustering results at the same resolution. The choice of normalization should be made based on the data characteristics and the analysis goals. Feature selection, typically identifying highly variable genes, also affects the graph construction and therefore the clustering results.

### Dimensionality Reduction

The number of principal components used for graph construction affects clustering results. Too few components may discard meaningful biological variation, while too many may introduce noise. The choice of dimensionality should be evaluated before resolution selection.

### Batch Effect Correction

For datasets generated across multiple batches or samples, batch effect correction is often necessary before clustering. The choice of batch correction method can affect the optimal resolution. Stability assessment should be performed after batch correction to ensure that the selected resolution is appropriate for the corrected data.

## Professional Escalation Criteria

Certain situations warrant consultation with bioinformatics specialists or statisticians with expertise in single-cell analysis.

### Persistent Instability Across Resolutions

If no resolution produces stable clusters, the problem may lie in the data quality or preprocessing instead of the resolution parameter. This situation warrants escalation to a specialist who can evaluate the quality control steps and preprocessing choices.

### Discordance Between Stability and Marker Genes

If stability analysis suggests a particular resolution but marker gene validation contradicts this choice, the discrepancy should be investigated. A specialist may be needed to determine whether the marker genes are appropriate for the expected cell types or whether the stability analysis is being misled by technical factors.

### Unexpected Cluster Numbers

If the number of clusters identified at the selected resolution differs dramatically from expectations based on the biological system, this discrepancy warrants investigation. A specialist can help determine whether the unexpected cluster number reflects genuine biology or technical artifacts.

### Complex Tissue Samples

Samples from complex tissues with many cell types and continuous developmental states present particular challenges for resolution selection. Specialists with experience in similar datasets can provide valuable guidance.

## Building a Resolution Decision Log for Reproducible Cluster Selection

A recurring problem in single-cell analysis is that researchers select a resolution parameter but fail to document the reasoning behind their choice. This omission makes it difficult to revisit the decision when new biological information emerges, when reviewers question the cluster granularity, or when the same dataset is reanalyzed with updated software versions. A structured decision log addresses this gap by creating a permanent record of the evidence considered at each step of resolution selection. This section provides a practical framework for building and maintaining such a log, with specific fields, example entries, and guidance for using the log to troubleshoot problematic clustering results.

### Why a Decision Log Matters for Resolution Selection

The resolution parameter is one of the most consequential choices in a single-cell analysis pipeline, yet it is often treated as a minor configuration detail. When clustering results are later found to be inconsistent with biological expectations, the absence of documentation makes it difficult to determine whether the problem stems from the resolution value, the preprocessing steps, or the clustering algorithm itself. A decision log creates an audit trail that separates these possibilities.

The reproducibility standards emphasized by community workflow projects such as nf-core highlight the importance of documenting beyond the final parameters but also the process by which those parameters were chosen. The nf-core documentation describes configuration standards that support reproducible workflow execution, and the same principle applies to analytical decisions within a workflow. A resolution decision log extends this philosophy from pipeline configuration to data analysis judgment calls.

The log also serves a practical function during manuscript preparation. Reviewers increasingly ask authors to justify clustering parameters, and a well-maintained log provides the evidence needed to respond confidently. instead of reconstructing the decision process from memory, the researcher can present the exact stability scores, marker gene results, and biological rationale that informed the final choice.

### Core Fields for a Resolution Decision Log

A resolution decision log should capture enough information to allow another researcher to understand and potentially reproduce the decision. The following fields provide a comprehensive structure that balances completeness with usability.

#### Dataset Identifier and Version

Record the exact dataset used for resolution selection, including the version if the data has been updated. This field is essential because resolution decisions made on an early data release may not apply to later versions with additional cells or corrected annotations. Include the date the dataset was accessed and any relevant identifiers from public repositories such as the NCBI databases, which provide official descriptions of sequence data resources and search systems.

#### Preprocessing Pipeline Snapshot

Document the exact preprocessing steps applied before clustering. This includes the normalization method, the number of highly variable genes selected, the number of principal components used for graph construction, and any batch correction approach. Because stability assessment is sensitive to preprocessing choices, the log must record the pipeline state at the time of resolution evaluation. If the preprocessing changes, the resolution decision should be revisited.

#### Candidate Resolution Range and Grid

Record the range of resolutions tested and the increment between values. A typical range might span from 0.1 to 2.0 in increments of 0.1, but the specific grid should be justified based on expected biological complexity and dataset size. Note any preliminary analyses that informed the choice of range, such as a quick scan of cluster numbers across a coarse grid.

#### Stability Assessment Parameters

Document the subsampling fraction, the number of bootstrap iterations, and the stability metric used. For example, a log entry might state that 80 percent of cells were subsampled without replacement for 30 iterations per resolution, with stability measured as the mean pairwise co-clustering frequency. These parameters directly affect the stability scores and must be recorded for the results to be interpretable.

#### Stability Trajectory Summary

Record the stability values for each tested resolution, either as a table or as a summary of the trajectory shape. Note any plateaus where stability remained relatively constant and any sharp drops indicating the onset of cluster fragmentation. This summary provides the quantitative basis for the final resolution choice.

#### Marker Gene Validation Results

For each candidate resolution that passed the stability threshold, document which marker genes were examined and the expression patterns observed. Include the specific genes used for each expected cell type and whether the clusters showed the expected enrichment. This field is critical because stability alone does not guarantee biological relevance.

#### Final Resolution Selection and Rationale

State the selected resolution value and provide a concise rationale that integrates the stability results, marker gene validation, and biological context. This entry should be written so that someone reading the log without prior knowledge of the analysis can understand why this particular value was chosen over the alternatives.

#### Software Versions and Environment

Record the exact versions of the clustering software, the analysis environment, and any relevant dependencies. This includes the clustering algorithm (Louvain or Leiden), the implementation package, and the version numbers. Software updates can change clustering behavior at the same resolution value, so this information is essential for reproducibility.

### Example Decision Log Entry

The following example illustrates how these fields might be completed for a peripheral blood mononuclear cell dataset. This example is provided as a template and should be adapted to the specific details of each analysis.

| Field | Entry |
|-------|-------|
| Dataset identifier | PBMC_10x_v3, accessed 2025-06-15 from NCBI GEO accession GSE123456 |
| Preprocessing snapshot | Normalized with LogNormalize, 2000 highly variable genes, 30 principal components, no batch correction needed |
| Candidate resolution range | 0.1 to 2.0 in increments of 0.1 |
| Stability parameters | 80 percent subsampling, 30 iterations per resolution, mean pairwise co-clustering frequency |
| Stability trajectory | Stable from 0.1 to 0.6, gradual decline from 0.7 to 1.2, sharp drop after 1.3 |
| Marker gene validation | Resolution 0.6 produced clusters enriched for CD3D, CD14, MS4A1, and NKG7 with clean separation |
| Final resolution | 0.6, selected as upper end of stability plateau with clean marker separation |
| Software versions | Seurat 5.0.1, Leiden algorithm, R 4.3.1 |

This entry provides enough information for another researcher to understand the decision and to reproduce the analysis if needed. The specific values in this example are illustrative and should not be interpreted as universal recommendations.

### Using the Decision Log for Troubleshooting

The decision log becomes particularly valuable when clustering results are later questioned or when downstream analyses produce unexpected findings. The log provides a structured way to diagnose whether the resolution choice is the source of the problem.

#### When Marker Genes Do Not Match Cluster Annotations

If a cluster annotated as a specific cell type does not express the expected markers, the decision log can help determine whether the resolution was appropriate. Review the marker gene validation results recorded during resolution selection. If the cluster was not present at the selected resolution but appeared after a software update, the version information in the log will identify this as a potential cause. If the cluster was present and showed the expected markers during validation, the problem may lie in the annotation step instead of the clustering.

#### When Cluster Numbers Differ from Published Studies

If the number of clusters identified in the current analysis differs substantially from published studies of similar tissues, the decision log provides a basis for comparison. Check whether the published studies used the same preprocessing pipeline and clustering algorithm. The log records the exact pipeline state, allowing a direct comparison of methodological differences that might explain the discrepancy.

#### When Stability Scores Are Recalculated

If stability assessment is repeated with different parameters, such as a different number of bootstrap iterations, the log provides the original parameters for comparison. This is particularly useful when a reviewer requests additional validation. The log shows what was originally done and allows the researcher to determine whether the new parameters are consistent with the original approach.

#### When Software Versions Change

Software updates can alter clustering behavior even when the resolution value is unchanged. The version information in the log allows the researcher to determine whether a change in results is attributable to the software update or to other factors. If the updated software produces different clusters at the same resolution, the log provides the evidence needed to decide whether to adjust the resolution or to maintain the original choice for consistency with previous analyses.

### Integrating the Decision Log with Existing Documentation Practices

The resolution decision log should be integrated with other documentation practices instead of maintained as a separate artifact. Many research groups already maintain analysis notebooks or electronic lab notebooks that record their analytical decisions. The decision log can be incorporated into these existing systems as a structured entry that accompanies the clustering step.

The training materials from The Carpentries emphasize the importance of reproducible computing practices, including version control and documentation. The resolution decision log aligns with these principles by providing a structured record of an analytical decision that is often made informally. Version control systems such as Git can be used to track changes to the decision log itself, ensuring that the log has its own audit trail.

For groups using workflow management systems, the decision log can be linked to the specific pipeline run that produced the clustering results. The nf-core documentation describes how pipeline runs can be configured and tracked, and the decision log can be stored alongside the pipeline outputs for easy retrieval.

### Common Mistakes in Decision Log Maintenance

Several recurring problems undermine the usefulness of resolution decision logs. Awareness of these failure patterns helps researchers design logs that remain useful over time.

#### Recording Only the Final Resolution

The most common mistake is to record only the selected resolution value without documenting the alternatives that were considered. This omission makes it impossible to understand why the chosen value was preferred over nearby values. The log should always include the full range of tested resolutions and the evidence for each.

#### Failing to Update the Log After Preprocessing Changes

If the preprocessing pipeline is modified after the resolution decision is made, the log becomes outdated. The stability assessment and marker gene validation were performed on data processed with the original pipeline, and the results may not apply to the updated data. The log should be updated whenever the preprocessing changes, and the resolution decision should be revisited.

#### Omitting Negative Results

Researchers often record successful validations but omit cases where marker genes did not match expectations or where stability was poor. These negative results are valuable for understanding the limitations of the chosen resolution and for troubleshooting later problems. The log should include both positive and negative evidence.

#### Using Vague Language

Entries such as "resolution 0.8 seemed reasonable" or "clusters looked good" do not provide the specificity needed for reproducibility. The log should use quantitative language wherever possible, recording specific stability scores, marker expression values, and cluster counts.

### Practical Implementation Steps

Building a resolution decision log does not require specialized software. A structured spreadsheet or a Markdown document maintained in the analysis repository provides sufficient functionality. The following steps outline a practical implementation approach.

#### Step 1: Create a Log Template

Create a template with the core fields described above. The template can be stored in the analysis repository and copied for each new dataset or analysis. This ensures consistency across projects and makes it easy to compare decisions across datasets.

#### Step 2: Fill in the Log During Resolution Selection

Complete the log fields as the resolution selection process proceeds instead of waiting until the analysis is finished. This practice ensures that details are not forgotten and that the log reflects the actual decision process. The stability trajectory summary and marker gene validation results should be recorded as they are generated.

#### Step 3: Review the Log Before Finalizing the Analysis

Before finalizing the clustering results, review the decision log to ensure that all fields are complete and that the rationale for the selected resolution is clearly stated. This review provides an opportunity to catch any gaps in documentation while the analysis is still fresh.

#### Step 4: Store the Log with the Analysis Outputs

Store the decision log alongside the clustering results and other analysis outputs. This ensures that the log is available when the results are revisited or when the manuscript is prepared. The log should be included in the version control repository so that changes are tracked.

#### Step 5: Revisit the Log When New Data or Methods Become Available

When new data are added to the dataset or when improved clustering methods become available, revisit the decision log to determine whether the resolution decision should be updated. The log provides the baseline for evaluating whether the new data or methods produce different results.

### Relationship to Broader Reproducibility Practices

The resolution decision log is one component of a broader reproducibility strategy for single-cell analysis. The training resources from the EMBL-EBI provide learning pathways for bioinformatics data analysis that emphasize the importance of documenting analytical decisions. The Galaxy Training Network similarly offers accessible workflow training that includes guidance on reproducible analysis practices.

The decision log complements these broader practices by addressing a specific gap: the documentation of judgment calls that are not captured by standard pipeline configuration files. While software versions and parameters can be recorded automatically, the reasoning behind parameter choices requires explicit documentation. The decision log provides this documentation in a structured format that supports both immediate use and long-term reproducibility.

The log also supports the principle of transparency in scientific reporting. When clustering results are published, the decision log provides the evidence needed to justify the resolution choice to reviewers and readers. This transparency strengthens the credibility of the analysis and facilitates independent verification by other research groups.

## Frequently Asked Questions

### What is the default resolution parameter in Seurat and why is it not always appropriate?

Seurat defaults to a resolution of 0.8 for graph-based clustering. This value was chosen as a general-purpose setting that works reasonably well for many datasets, but it is not universally appropriate. The optimal resolution depends on the biological complexity of the sample, the number of cells, and the transcriptional distinctness of the populations being studied. A dataset with only a few major cell types may be over-clustered at 0.8, while a complex tissue with many rare subpopulations may be under-clustered. Systematic evaluation across a range of resolutions is necessary to identify the appropriate value for each dataset.

### How many resolution values should I test?

The number of resolution values to test depends on the expected biological complexity and the computational resources available. A reasonable starting point is to test resolutions from 0.1 to 2.0 in increments of 0.1, giving 20 candidate values. This range covers the typical spectrum from very coarse to very fine clustering. If computational resources are limited, a coarser grid with increments of 0.2 can be used initially, followed by finer sampling around promising values.

### What is the difference between Louvain and Leiden clustering for resolution selection?

Louvain and Leiden are both modularity-based clustering algorithms that use a resolution parameter. Leiden improves on Louvain by guaranteeing that communities are well-connected and by providing faster convergence. For resolution selection, the key difference is that the two algorithms may produce different clusterings at the same resolution value. Stability assessment should be performed with the specific algorithm that will be used for the final analysis. The choice between Louvain and Leiden should be made before resolution selection and kept consistent throughout the analysis.

### How do I validate that my chosen resolution produces biologically meaningful clusters?

Marker gene validation is the primary approach for confirming biological relevance. For each cluster at the candidate resolution, examine the expression of known markers for the cell types expected in the sample. A biologically meaningful clustering will produce clusters that are enriched for specific markers and that separate cleanly from other clusters based on marker expression. Additional validation can include comparison with published cell type annotations for similar tissues, examination of known developmental relationships between cell types, and assessment of whether the clustering is consistent with independent experimental data.

### Can I use the same resolution for all datasets in a comparative study?

Using the same resolution across datasets in a comparative study can be appropriate if the datasets are expected to contain the same cell types and have similar complexity. However, differences in sequencing depth, cell capture efficiency, and sample quality can affect the optimal resolution. A more rigorous approach is to perform stability assessment for each dataset independently and then compare the resulting cluster annotations. If the same resolution is used across datasets, the stability of each dataset at that resolution should be assessed to ensure that the choice is appropriate for all datasets.

### What should I do if my clusters are not stable at any resolution?

Persistent instability across all resolutions suggests that the problem lies in the data or preprocessing instead of the resolution parameter. Common causes include poor quality cells that should have been filtered during quality control, batch effects that have not been properly corrected, or excessive technical noise. Review the quality control metrics and consider additional filtering or batch correction. If the problem persists, consult with a bioinformatics specialist who can evaluate the preprocessing pipeline.

### How does the number of cells affect the appropriate resolution?

The number of cells affects the appropriate resolution in several ways. Larger datasets can support higher resolution values because they contain more cells per potential cluster, providing sufficient cells for reliable downstream analysis. Smaller datasets may become over-clustered at high resolutions because individual clusters would contain too few cells. The resolution limit of modularity clustering also depends on the frequency of subgraphs within the graph, meaning that rare populations in large datasets may require higher resolution to be resolved.

### What is the relationship between resolution and the number of clusters?

The relationship between resolution and the number of clusters is monotonic but not linear. Increasing the resolution generally increases the number of clusters, but the exact relationship depends on the structure of the data. Some resolution increases may produce no change in cluster number, while others may produce large jumps. The splitting resolution concept from the IEEE Transactions on Computational Biology and Bioinformatics literature describes the minimum resolution at which a graph or subgraph is split into multiple clusters, providing a theoretical framework for understanding this relationship.

## Related Bioinformatics Guides

- [Single-Cell vs Single-Nucleus RNA Sequencing: Choosing the Right Approach](/knowledge/bioinformatics/single-cell-vs-single-nucleus-rna-sequencing-choosing-the-right-approach)
- [Spatial Proteomics vs. Single-Cell Proteomics: Choosing the Right Approach](/knowledge/bioinformatics/spatial-proteomics-vs-single-cell-proteomics-choosing-the-right-approach)
- [Spatial Transcriptomics vs. Single-Cell RNA Sequencing: Which Approach Fits Your Research?](/knowledge/bioinformatics/spatial-transcriptomics-vs-single-cell-rna-sequencing-which-approach-fits-your-research)
- [RNA-Seq Alignment: Choosing the Right Tool and Parameters](/knowledge/bioinformatics/rna-seq-alignment-choosing-the-right-tool-and-parameters)
- [Single-Cell Sequencing Depth: How Much Is Enough?](/knowledge/bioinformatics/single-cell-sequencing-depth-how-much-is-enough)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)


## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [Selecting single cell clustering parameter values using subsampling-based robustness metrics.](https://pubmed.ncbi.nlm.nih.gov/33522897). BMC bioinformatics, 2021.
- [Dimensionality reduction for visualizing single-cell data using UMAP.](https://pubmed.ncbi.nlm.nih.gov/30531897). Nature biotechnology, 2018.
- [Resolution Tradeoffs in Modularity Clustering of Single Cell RNA-Seq Datasets.](https://pubmed.ncbi.nlm.nih.gov/41171688). IEEE transactions on computational biology and bioinformatics, 2025.
- [Cell Layers: uncovering clustering structure in unsupervised single-cell transcriptomic analysis.](https://pubmed.ncbi.nlm.nih.gov/35967929). Bioinformatics advances, 2022.
- [scVAG: Unified single-cell clustering via variational-autoencoder integration with Graph Attention Autoencoder.](https://pubmed.ncbi.nlm.nih.gov/39687165). Heliyon, 2024.
- [An active learning approach for clustering single-cell RNA-seq data.](https://pubmed.ncbi.nlm.nih.gov/34244616). Laboratory investigation, a journal of technical methods and pathology, 2022.
- [TiC2D: Trajectory Inference From Single-Cell RNA-Seq Data Using Consensus Clustering.](https://pubmed.ncbi.nlm.nih.gov/33630737). IEEE/ACM transactions on computational biology and bioinformatics, 2022.
- [Dirichlet process mixture models for single-cell RNA-seq clustering.](https://pubmed.ncbi.nlm.nih.gov/35237784). Biology open, 2022.
- [Comprehensive assessment of alternative splicing analysis methods for single-cell RNA-seq.](https://doi.org/10.1016/j.isci.2026.116090). 2026.
- [scMagnifier: Resolving fine-grained cell subtypes via GRN-informed perturbations and consensus clustering.](https://doi.org/10.1371/journal.pcbi.1014167). 2026.
- [Multiscale domain identification for spatial transcriptomics via persistent homology.](https://doi.org/10.1016/j.crmeth.2026.101376). 2026.
- [Optimization of clustering parameters for single-cell RNA analysis using intrinsic goodness metrics](https://doi.org/10.3389/fbinf.2025.1562410). Frontiers in Bioinformatics, 2025.
- [Single-molecule clustering for super-resolution optical fluorescence microscopy](https://doi.org/10.3390/photonics9010007). Photonics, 2022.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.