# Harmony for Single-Cell Data Integration: How It Works and When to Use It for Batch Effect Correction


## Key Takeaways

- Harmony is an algorithm designed to project single-cell RNA sequencing (scRNA-seq) and single-nucleus RNA sequencing (snRNA-seq) data into a shared embedding, enabling cells to cluster by cell type rather than by dataset-specific technical variations (batch effects).
- It simultaneously accounts for multiple experimental and biological factors, making it suitable for integrating datasets with complex batch structures arising from different sequencing platforms, donors, or tissue sources.
- Harmony operates iteratively by clustering cells and computing batch-specific adjustments to remove technical variation, projecting cells into a corrected space for downstream analyses like clustering and differential expression.
- The algorithm is computationally efficient, capable of integrating approximately one million cells on a personal computer, and was recommended as a first method to try in a benchmark study due to its significantly shorter runtime compared to alternatives like LIGER and Seurat 3.
- Practical implementation requires standard quality control, normalization, and dimensionality reduction (typically PCA on highly variable genes) before running Harmony, with key parameters like theta and lambda potentially requiring tuning for complex batch structures.
- Common failure patterns include overcorrection (loss of biological variation) and undercorrection (residual batch effects), which can be diagnosed through visual inspection and quantitative metrics like kBET and LISI, and require careful parameter adjustment or experimental design considerations.

---

Single-cell RNA sequencing (scRNA-seq) and single-nucleus RNA sequencing (snRNA-seq) experiments generate gene expression profiles for thousands to millions of individual cells. When researchers combine datasets from multiple samples, laboratories, sequencing platforms, or experimental conditions, systematic technical variation known as batch effects can obscure genuine biological differences. Harmony is an algorithm designed to project cells into a shared embedding where cells group by cell type instead of by dataset-specific conditions. This article explains how Harmony works, its assumptions and limitations, and provides practical guidance on when to use it for batch effect correction in single-cell data integration workflows.

Harmony addresses a distinct problem in single-cell analysis: the need to integrate large datasets with complex batch structures while preserving cell type identity. The algorithm simultaneously accounts for multiple experimental and biological factors, making it suitable for datasets where samples differ by technology, donor, tissue source, or disease state. Understanding the principles behind Harmony helps researchers decide whether it is the appropriate tool for their specific integration task and how to configure it correctly.

## At a Glance

The table below summarizes key characteristics of Harmony for batch effect correction in single-cell data analysis.

| Feature | Harmony Capability | Practical Consideration |
| --- | --- | --- |
| Input data | Normalized gene expression matrix or dimensionality reduction embedding | Typically applied after principal component analysis (PCA) on highly variable genes |
| Batch handling | Simultaneously accounts for multiple experimental and biological factors | Can integrate datasets with complex batch structures beyond simple two-group comparisons |
| Scalability | Enables integration of approximately one million cells on a personal computer | Requires fewer computational resources compared to many alternative methods |
| Output | Corrected embedding where cells group by cell type | Downstream clustering, visualization, and differential expression analysis use the corrected space |
| Performance benchmark | Recommended as first method to try among 14 compared methods | Shorter runtime than alternatives including LIGER and Seurat 3 |
| Typical workflow position | After quality control, normalization, and PCA | Before clustering, cell type annotation, and downstream analyses |
| Parameter sensitivity | Main parameters include theta, lambda, and number of Harmony iterations | Default settings work for many datasets but tuning may improve results for complex batch structures |

## Understanding Batch Effects in Single-Cell Data

Batch effects arise from systematic technical variation that is unrelated to biological differences between cells. These effects can originate from multiple sources including different sequencing platforms, library preparation protocols, sequencing depth, capture efficiency, ambient RNA contamination, and laboratory-specific conditions. When researchers analyze datasets generated across multiple studies or institutions, these technical differences become interspersed with genuine biological variation, making it difficult to distinguish true cell type signatures from technical artifacts.

The challenge of batch effect correction is particularly acute in single-cell transcriptomics because the data are sparse and high-dimensional. Each cell contains measurements for thousands of genes, but most genes are expressed at low levels or not detected at all in individual cells. This sparsity amplifies the impact of technical variation and complicates the separation of biological signal from noise. The emerging diversity of scRNA-seq datasets allows for full transcriptional characterization of cell types across a wide variety of biological and clinical conditions, but analyzing them together is challenging, particularly when datasets are assayed with different technologies, because biological and technical differences are interspersed [<a href="#ref-1">1</a>].

Batch effects manifest in several observable ways. Cells from the same biological cell type may cluster separately based on their dataset of origin instead of their shared identity. Marker gene expression may appear artificially elevated or depressed in specific batches. Cell type proportions may appear distorted when comparing across datasets. These artifacts can lead to incorrect cell type annotations, spurious differential expression results, and misinterpretation of biological variation as technical noise or vice versa.

The impact of batch effects extends beyond visualization. In studies that aim to identify rare cell populations or subtle transcriptional states, batch-related variation can mask biologically meaningful differences. For example, when integrating data from multiple cancer patients, patient-specific technical variation can be mistaken for tumor heterogeneity. Similarly, when combining data from different sequencing platforms, platform-specific biases can create artificial clusters that do not correspond to real cell types. These issues underscore the importance of selecting an appropriate batch correction method and applying it correctly within the analysis workflow.

## Core Principles of the Harmony Algorithm

Harmony operates by iteratively adjusting cell embeddings to maximize diversity within batches while preserving the overall structure of the data. The algorithm begins with an initial embedding, typically the PCA representation of the normalized gene expression matrix. It then applies a series of correction rounds that cluster cells, compute batch-specific adjustments, and update the embedding to remove batch-dependent variation.

The algorithm projects cells into a shared embedding in which cells group by cell type instead of dataset-specific conditions [<a href="#ref-1">1</a>]. This projection is achieved through a mixture model approach that assigns each cell to a cluster while simultaneously estimating batch-specific correction factors. The iterative process alternates between two main steps: assigning cells to clusters based on their current positions and computing cluster-specific batch corrections that pull cells from the same biological group together regardless of their batch of origin.

Harmony simultaneously accounts for multiple experimental and biological factors [<a href="#ref-1">1</a>]. This means the algorithm can handle datasets where cells differ by more than one technical or biological variable at the same time. For example, a study integrating peripheral blood mononuclear cells from multiple donors sequenced on different platforms can specify both donor identity and sequencing technology as batch variables. The algorithm learns correction factors for each variable and applies them jointly to produce a harmonized embedding.

The computational efficiency of Harmony is one of its distinguishing features. The algorithm enables the integration of approximately one million cells on a personal computer, making it accessible to research groups without access to high-performance computing clusters [<a href="#ref-1">1</a>]. This efficiency stems from the algorithm's design, which avoids pairwise comparisons between all cells and instead operates on cluster-level statistics.

The mathematical foundation of Harmony involves a soft clustering approach. instead of assigning each cell definitively to a single cluster, the algorithm computes probabilities of cluster membership for each cell. This soft assignment allows cells that are intermediate between clusters to contribute to multiple cluster-level corrections, which helps preserve continuous biological transitions such as differentiation trajectories. The correction factors are then estimated in a way that accounts for the uncertainty in cluster assignments.

## Comparison with Other Batch Correction Methods

Harmony is one of several algorithms developed for single-cell data integration. A benchmark study compared 14 batch correction methods in terms of computational runtime, ability to handle large datasets, and batch-effect correction efficacy while preserving cell type purity [<a href="#ref-2">2</a>]. The study designed five scenarios: identical cell types with different technologies, non-identical cell types, multiple batches, big data, and simulated data. Performance was evaluated using four benchmarking metrics including kBET, LISI, ASW, and ARI [<a href="#ref-2">2</a>].

Based on the benchmark results, Harmony, LIGER, and Seurat 3 were recommended as the most suitable methods for batch integration [<a href="#ref-2">2</a>]. Harmony was specifically recommended as the first method to try due to its significantly shorter runtime, with the other methods as viable alternatives [<a href="#ref-2">2</a>]. This recommendation reflects the practical importance of computational efficiency when working with large single-cell datasets.

The benchmark also investigated the use of batch-corrected data for differential gene expression analysis [<a href="#ref-2">2</a>]. This is an important consideration because some integration methods may overcorrect the data, removing genuine biological variation along with technical noise. The study's findings inform decisions about which downstream analyses are appropriate after batch correction with each method.

Seurat's integration workflow, which uses canonical correlation analysis and mutual nearest neighbors, is widely used in the single-cell community. LIGER employs integrative non-negative matrix factorization to identify shared and dataset-specific factors. Each method has distinct assumptions about the structure of batch effects and the relationship between biological and technical variation. The choice among these methods depends on dataset characteristics, computational resources, and downstream analysis goals.

The benchmark study's inclusion of a big data scenario is particularly relevant for researchers working with large-scale datasets. As single-cell datasets continue to grow in size, computational efficiency becomes an increasingly important consideration. Harmony's ability to integrate approximately one million cells on a personal computer positions it well for these applications [<a href="#ref-1">1</a>]. However, researchers should verify that the method performs adequately on their specific data characteristics instead of relying solely on general benchmark results.

## Practical Workflow for Harmony Integration

Implementing Harmony in a single-cell analysis workflow requires attention to data preparation, parameter selection, and quality assessment. The following steps outline a practical approach to Harmony-based batch correction.

### Data Preparation and Quality Control

Before applying Harmony, researchers must complete standard quality control steps for single-cell data. These steps include filtering low-quality cells based on mitochondrial gene content, total UMI counts, and number of detected genes. Quality control thresholds should be determined based on the specific dataset and biological context. For example, a study of knee osteoarthritis applied quality control with mitochondrial genes exceeding three median absolute deviations as a filtering criterion [<a href="#ref-3">3</a>].

After quality control, the data should be normalized to account for differences in sequencing depth across cells. Common normalization approaches include log-normalization and scaling by total UMI counts. The choice of normalization method can affect downstream integration results, and researchers should document their normalization strategy for reproducibility.

Feature selection is the next step, typically involving identification of highly variable genes. The number of highly variable genes selected influences the sensitivity of downstream clustering and integration. Some protocols recommend automated optimization of feature selection parameters, as dataset-specific tuning is often required for optimal results [<a href="#ref-4">4</a>].

### Dimensionality Reduction

Harmony operates on a low-dimensional embedding of the data instead of the full gene expression matrix. The standard approach is to perform PCA on the scaled expression values of highly variable genes. The number of principal components retained is a critical parameter that affects integration quality. Too few components may lose biological signal, while too many may include noise that complicates the integration.

The choice of dimensionality should be informed by the dataset's complexity and the expected number of cell types. Some workflows use heuristic approaches such as the elbow method on the standard deviation of principal components, while others employ more systematic optimization strategies. The scAutoTune protocol describes automated optimization of feature selection and clustering parameters, including optional Harmony batch correction, with silhouette-based performance metrics to guide parameter selection [<a href="#ref-4">4</a>].

### Running Harmony

Harmony is available as an R package and can be integrated into Seurat-based workflows. The algorithm accepts the PCA embedding and a vector or data frame specifying batch variables. The main parameters include theta, which controls the diversity penalty, lambda, which controls the regularization strength, and the maximum number of iterations.

Default parameter values work well for many datasets, but complex batch structures may require adjustment. The theta parameter influences how strongly Harmony enforces diversity within batches. Higher theta values increase the penalty for clusters that are homogeneous with respect to batch, potentially leading to more aggressive correction. Lambda controls the regularization of the correction terms and can be increased to prevent overcorrection.

After running Harmony, the corrected embedding replaces the PCA embedding for downstream analyses. Clustering is typically performed on the Harmony-corrected dimensions using graph-based methods such as Louvain or Leiden clustering. The number of clusters should be evaluated using appropriate metrics and biological knowledge.

### Cell Type Annotation

Cell type annotation after Harmony integration can be performed using marker gene expression, reference-based annotation tools, or a combination of approaches. The corrected embedding should bring cells of the same type together regardless of their batch of origin, facilitating consistent annotation across datasets.

Reference-based annotation tools such as SingleR can be applied to the integrated data to assign cell types based on known marker gene signatures. Manual annotation using canonical markers remains an important validation step, particularly for identifying rare or novel cell populations. The CellMarker database and similar resources provide curated marker gene lists for various tissues and organisms.

## Parameter Tuning and Optimization

Harmony's default parameters are suitable for many applications, but optimal performance may require dataset-specific tuning. The scAutoTune protocol demonstrates a systematic approach to parameter optimization in single-cell RNA-seq analysis [<a href="#ref-4">4</a>]. This protocol describes steps for installing the scAutoTune tool, preparing Seurat objects from public datasets, and performing grid-based parameter sweeping with optional Harmony batch correction [<a href="#ref-4">4</a>].

The protocol computes silhouette-based performance metrics to evaluate clustering quality across different parameter combinations [<a href="#ref-4">4</a>]. Silhouette scores measure how similar cells are to their own cluster compared to neighboring clusters, providing a quantitative basis for parameter selection. The results are visualized using generalized additive model-smoothed landscapes, which help identify regions of parameter space associated with high-quality clustering [<a href="#ref-4">4</a>].

After identifying optimal parameters, the protocol applies these settings to generate UMAP embeddings and interpretable clusters [<a href="#ref-4">4</a>]. This systematic approach reduces the subjectivity involved in manual parameter selection and improves reproducibility across analyses.

For Harmony specifically, the key parameters to optimize include the number of principal components used as input, the theta diversity penalty, and the number of iterations. The optimal values depend on dataset size, batch structure complexity, and the biological question being addressed. Researchers should document the parameter values used and the rationale for their choices.

The following table summarizes common parameter adjustments and their expected effects on integration outcomes.

| Parameter | Direction of Adjustment | Expected Effect | When to Consider |
| --- | --- | --- | --- |
| Theta (diversity penalty) | Increase | More aggressive batch mixing, potential overcorrection | Persistent batch structure after initial run |
| Theta (diversity penalty) | Decrease | More conservative correction, preserves more biological variation | Known biological differences between batches |
| Lambda (regularization) | Increase | Prevents overcorrection, stabilizes correction factors | Highly variable batch compositions |
| Number of principal components | Increase | Captures more biological signal, may include noise | Complex datasets with many cell types |
| Number of principal components | Decrease | Reduces noise, may lose subtle biological variation | Datasets with few expected cell types |
| Maximum iterations | Increase | Allows more correction rounds, may improve mixing | Large datasets with complex batch structures |

## Applications in Disease and Translational Research

Harmony has been applied across a wide range of biological and clinical research contexts. These applications demonstrate the algorithm's versatility and provide practical examples of its use in different experimental settings.

### Cancer Immunology Studies

In cancer research, Harmony has been used to integrate tumor and normal tissue samples to characterize the immune microenvironment. A study of colorectal cancer integrated 78 scRNA-seq datasets comprising 42 treatment-naive colorectal tumors, 13 tumor adjacent tissues, and 23 normal mucosa tissues [<a href="#ref-5">5</a>]. The researchers used standardized Seurat procedures to identify cellular components with canonical cell markers, assessed batch effects, and corrected them using the Harmony algorithm [<a href="#ref-5">5</a>].

The integration enabled identification of functional immune cell subtypes, including two states of exhausted CD8+ T cells and FOLR2+LYVE1+ macrophages associated with unfavorable prognosis [<a href="#ref-5">5</a>]. The study also revealed how KRAS/TP53 mutation status reshaped immune cell function and immune checkpoint ligand and receptor expression patterns [<a href="#ref-5">5</a>]. These findings illustrate how Harmony-based integration can reveal biological insights that would be difficult to obtain from individual datasets analyzed separately.

A study of acute myeloid leukemia used Harmony to correct batch effects across 59 samples comprising 40 AML and 19 healthy donor samples [<a href="#ref-6">6</a>]. After quality control and doublet removal, 284,687 high-quality cells were retained for analysis [<a href="#ref-6">6</a>]. The integration enabled identification of 76 cell clusters, including patient-specific malignant cells with elevated copy number variations [<a href="#ref-6">6</a>]. This analysis revealed altered immune composition in AML samples and provided insights into natural killer cell dysfunction in the tumor microenvironment [<a href="#ref-6">6</a>].

In lung adenocarcinoma research, Harmony was used to analyze single-cell RNA-seq data from the GEO database to explore cancer stem cell characteristics [<a href="#ref-7">7</a>]. The analysis identified nine cell clusters, with cancer stem cells showing higher proportion in tumor samples [<a href="#ref-7">7</a>]. The LGR5+ stem cell subpopulation was identified as a major contributor to cancer progression, with hub genes including HLA-DPB1, CD74, CTSH, and HLA-DRB5 mediating the unique transcriptional state [<a href="#ref-7">7</a>].

A melanoma study employed Harmony for batch effect correction in single-cell analysis to explore immune microenvironment and oxidative stress heterogeneity [<a href="#ref-8">8</a>]. The analysis reclassified malignant tumor cell populations into six distinct tumor subsets based on comprehensive single-cell dataset analysis [<a href="#ref-8">8</a>]. This application demonstrates Harmony's utility in integrating complex tumor datasets with heterogeneous cellular composition.

### Autoimmune and Inflammatory Disease Research

Harmony has been applied to study rheumatoid arthritis by integrating single-cell RNA sequencing datasets from patients with the disease [<a href="#ref-9">9</a>]. Researchers processed three single-cell RNA datasets of cells harvested from RA patients and integrated them using Seurat and Harmony R packages [<a href="#ref-9">9</a>]. After identifying cell types by classic marker genes, the integrated dataset was used for cell-cell communication analysis [<a href="#ref-9">9</a>].

This approach identified HBEGF+ fibroblasts that contributed to RA remission [<a href="#ref-9">9</a>]. The EGF signaling pathway was found to play a role in this process, and differential gene expression analysis revealed markers mainly expressed by HBEGF+ fibroblasts including CLIC5, PDGFD, BDH2, and ENPP1 [<a href="#ref-9">9</a>]. The study demonstrates how Harmony integration can enable discovery of cell subpopulations relevant to disease progression and remission.

In knee osteoarthritis research, Harmony was used to integrate single-cell data from 11 samples to dissect ferroptosis-driven immune remodeling [<a href="#ref-3">3</a>]. The analysis identified 12 chondrocyte clusters, including ferroptosis-active homeostasis chondrocytes that exhibited 491 differentially expressed genes linked to lipid peroxidation [<a href="#ref-3">3</a>]. Cell communication analysis revealed that these chondrocytes orchestrated synovitis via FGF signaling, amplifying extracellular matrix degradation and inflammatory cascades [<a href="#ref-3">3</a>].

### Developmental Biology

Harmony has been applied to study embryonic development, including a study of tail development in sheep [<a href="#ref-10">10</a>]. Researchers performed single-cell RNA sequencing on embryos from Hulunbuir short-tailed sheep and Hu sheep at embryonic days 16 and 19 [<a href="#ref-10">10</a>]. The analysis identified 12 cell types in E16 embryos and 15 cell types in E19 embryos, with the MDK_ITGA6+ITGB1 ligand-receptor pair consistently mediating core intercellular communication [<a href="#ref-10">10</a>].

The integration revealed that the T gene mutation in short-tailed sheep showed transcriptomic signatures of increased BMP pathway activity and reduced FGF8 expression, which may disrupt apical ectodermal ridge survival and contribute to the short-tail phenotype [<a href="#ref-10">10</a>]. This application demonstrates Harmony's utility in comparative developmental studies across genetically distinct populations.

### Pediatric Cancer Research

A study of childhood acute lymphoblastic leukemia used Seurat and Harmony R packages for quality control and scRNA-seq analysis [<a href="#ref-11">11</a>]. The researchers collected single-cell RNA-seq data from ALL and healthy samples from the Gene Expression Omnibus database [<a href="#ref-11">11</a>]. The analysis identified nine main cell clusters, with B cells showing higher infiltration proportion in ALL samples [<a href="#ref-11">11</a>].

The B cells were sub-clustered into five cell sub-groups, with B cells 1 closely associated with cell proliferation and stemness [<a href="#ref-11">11</a>]. Copy number variation analysis revealed significant amplification on chromosomes 6 and 21 that supported the stemness of this B cell subpopulation [<a href="#ref-11">11</a>]. This study illustrates how Harmony integration can facilitate identification of disease-relevant cell subpopulations in pediatric cancers.

A study of childhood medulloblastoma used Harmony alignment to reveal novel subgroup-associated subpopulations that recapitulate neurodevelopmental processes [<a href="#ref-12">12</a>]. The analysis identified photoreceptor and glutamatergic neuron-like cells in molecular subgroups GP3 and GP4, and a specific nodule-associated neuronally differentiated subpopulation in the sonic hedgehog subgroup [<a href="#ref-12">12</a>]. This application demonstrates Harmony's capacity to resolve subtle biological heterogeneity within complex tumor microenvironments.

### Sepsis and Immune Response Research

A study of preterm newborns with late-onset sepsis used Seurat and Harmony R packages for scRNA-seq analysis [<a href="#ref-13">13</a>]. The analysis identified eight cell clusters, with immune cells including neutrophils, T helper cells, and NK cells distributed more in sepsis samples [<a href="#ref-13">13</a>]. The integration revealed that T cells and NK cells exhibited greater activation of immune and detoxification pathways, such as antimicrobial humoral immune response and cellular oxidant detoxification [<a href="#ref-13">13</a>].

The study identified GZMA+ cytotoxic T cells, GRASP+ T cells, TMEM107+ NK cells, and GPR171 NK cells as having potential impact on promoting immune response in preterm infant sepsis [<a href="#ref-13">13</a>]. Key regulons including CREM, EST1, HSPA5, and STAT1 were identified as potential molecular targets for improving sepsis treatment [<a href="#ref-13">13</a>].

### Deconvolution Applications

Harmony has also been applied in the context of cell composition deconvolution, a bioinformatic task to estimate cell fractions from bulk gene expression profiles [<a href="#ref-14">14</a>]. A study developed a deep neural network model called HASCAD that was trained using bulk RNA-seq simulated from three scRNA-seq datasets normalized using a Harmony-Symphony based strategy [<a href="#ref-14">14</a>]. The study found that removal of batch effects in reference scRNA-seq datasets could benefit the task of cell composition deconvolution [<a href="#ref-14">14</a>].

## Common Failure Patterns and Troubleshooting

Despite Harmony's effectiveness, several common failure patterns can compromise integration quality. Recognizing these patterns helps researchers diagnose problems and adjust their approach.

### Overcorrection and Loss of Biological Variation

Overcorrection occurs when Harmony removes genuine biological differences along with technical batch effects. This can happen when the batch variable is confounded with a biological variable of interest. For example, if all control samples are sequenced on one platform and all treated samples on another, Harmony may remove treatment-related variation as if it were a batch effect.

Signs of overcorrection include loss of known cell type separation, merging of distinct biological populations, and failure to detect expected differentially expressed genes. Researchers should compare integration results with and without Harmony correction and validate that known biological differences are preserved.

### Undercorrection and Residual Batch Effects

Undercorrection leaves residual batch effects in the integrated data. This can occur when the number of Harmony iterations is insufficient, when the theta parameter is too low, or when the batch structure is more complex than specified. Residual batch effects manifest as cells from the same batch clustering together even after integration.

Visual inspection of UMAP or t-SNE plots colored by batch can reveal residual batch effects. Quantitative metrics such as kBET or LISI scores can provide objective assessment of batch mixing. If undercorrection is detected, increasing the number of iterations or adjusting theta may improve results.

### Confounded Batch and Biological Variables

When batch and biological variables are perfectly confounded, no computational method can fully separate technical from biological variation. This situation requires careful experimental design or acknowledgment of the limitation in interpretation. Researchers should assess the degree of confounding before applying Harmony and interpret results with appropriate caution.

### Inappropriate Input Dimensionality

The number of principal components used as input to Harmony affects integration quality. Too few components may omit important biological variation, while too many may introduce noise that complicates the integration. The optimal dimensionality depends on dataset complexity and should be evaluated systematically.

### Batch Variable Misspecification

Specifying the wrong batch variables or omitting important ones can lead to poor integration outcomes. Including biological variables of interest as batch variables can remove genuine biological signal. Omitting technical variables that contribute to batch effects leaves residual variation uncorrected. Researchers should carefully consider which variables to include as batch variables based on the experimental design and known sources of technical variation.

## Quality Assessment and Validation

Assessing the quality of Harmony integration requires multiple complementary approaches. Visual inspection of dimensionality reduction plots provides an initial qualitative assessment. Cells from different batches should be intermixed within cell type clusters, and known cell types should form coherent groups.

Quantitative metrics provide more objective evaluation. The kBET metric assesses whether the local neighborhood composition of each cell reflects the overall batch composition. The LISI metric measures local inverse Simpson index to evaluate batch mixing and cell type purity. The average silhouette width (ASW) evaluates cluster cohesion and separation. The adjusted Rand index (ARI) compares clustering results to known labels when ground truth is available.

The benchmark study comparing batch correction methods used these four metrics to evaluate performance across five scenarios [<a href="#ref-2">2</a>]. Researchers should select metrics appropriate for their specific dataset and analysis goals. For datasets with known cell type annotations, ARI provides a direct measure of how well integration preserves cell type identity. For datasets without known labels, kBET and LISI assess batch mixing without requiring ground truth.

Differential expression analysis after integration requires careful interpretation. Batch-corrected data may not be suitable for all downstream analyses, and the benchmark study specifically investigated the use of batch-corrected data for differential gene expression [<a href="#ref-2">2</a>]. Researchers should validate that differentially expressed genes identified after integration are biologically plausible and not artifacts of the correction process.

The following table summarizes common quality assessment approaches and their applications.

| Assessment Approach | What It Measures | When to Use | Interpretation Guidance |
| --- | --- | --- | --- |
| Visual inspection of UMAP or t-SNE | Qualitative assessment of batch mixing and cell type structure | Initial evaluation after integration | Cells from different batches should intermix within cell type clusters |
| kBET | Whether local neighborhood composition reflects overall batch composition | Datasets without known cell type labels | High acceptance rates indicate good batch mixing |
| LISI | Local inverse Simpson index for batch mixing and cell type purity | Comparing integration methods | Higher values indicate better mixing for batch labels |
| ASW | Cluster cohesion and separation | Evaluating clustering quality after integration | Higher values indicate more distinct clusters |
| ARI | Agreement between clustering and known labels | Datasets with ground truth annotations | Higher values indicate better preservation of cell type identity |
| Marker gene expression validation | Whether known markers are expressed in expected cell types | All integration analyses | Markers should show expected cell type specificity after integration |

## Reproducibility and Documentation

Reproducible single-cell analysis requires careful documentation of all processing steps, parameters, and software versions. The Harmony algorithm's stochastic components may produce slightly different results across runs, so setting random seeds is important for reproducibility.

Version control for analysis code and computational environments helps ensure that analyses can be reproduced or updated as new data become available. Containerization tools and workflow managers such as those provided by the nf-core community support reproducible pipeline execution [<a href="#ref-15">15</a>]. The nf-core documentation describes community pipeline standards, usage, configuration, and reproducible workflow context [<a href="#ref-15">15</a>].

Training resources from EMBL-EBI and the Galaxy Training Network provide practical guidance for researchers developing single-cell analysis skills [<a href="#ref-16">16</a>][<a href="#ref-17">17</a>]. The Carpentries lessons offer foundational computing, data, shell, Git, and programming training that supports reproducible research practices [<a href="#ref-18">18</a>]. Bioconductor provides official package, workflow, installation, and reproducible genomic-analysis documentation for R-based analyses [<a href="#ref-19">19</a>]. NCBI Data Resources provide official descriptions of databases, search systems, sequence resources, and analysis services that support data access and management [<a href="#ref-20">20</a>].

Documentation should include the version of Harmony and Seurat used, the exact parameters specified, the random seed, and the input data version. This information allows other researchers to reproduce the analysis or understand how parameter choices may have influenced the results.

## Limitations and Interpretation Caveats

Harmony is a powerful tool for batch effect correction, but it has important limitations that researchers should understand. The algorithm assumes that cell types are shared across batches and that batch effects are approximately linear in the embedding space. When these assumptions are violated, integration quality may be compromised.

Harmony may not perform optimally when integrating datasets with completely distinct cell type compositions. If one dataset contains cell types absent from another, the algorithm may struggle to correctly align shared cell types while preserving dataset-specific populations. The benchmark study included a scenario with non-identical cell types to evaluate this situation [<a href="#ref-2">2</a>].

The choice of batch variables requires careful consideration. Including too many batch variables may lead to overcorrection, while omitting important variables leaves residual batch effects. Researchers should include all known technical variables but avoid including biological variables that are of interest as batch variables.

Interpretation of results after Harmony integration requires awareness of what the correction does and does not do. Harmony produces a corrected embedding for visualization and clustering, but the corrected values do not represent direct gene expression measurements. Differential expression analysis should be performed on the original expression values with batch information included as a covariate in the statistical model.

The benchmark study's finding that Harmony, LIGER, and Seurat 3 were the recommended methods for batch integration provides useful guidance, but these results should be interpreted in the context of the specific scenarios evaluated [<a href="#ref-2">2</a>]. Different datasets may favor different methods depending on their characteristics. Researchers should evaluate integration quality on their own data instead of assuming that a method recommended in a benchmark will perform optimally in all contexts.

## Professional Escalation Criteria

Researchers should consider seeking additional expertise or alternative approaches in several situations. If Harmony integration produces results that contradict well-established biological knowledge, the analysis should be reviewed carefully before proceeding. Persistent failure to achieve adequate batch mixing despite parameter adjustment may indicate that Harmony is not suitable for the specific dataset.

When datasets have complex batch structures with many interacting variables, consultation with a bioinformatics specialist may be warranted. Similarly, when integrating data across species or very divergent biological contexts, alternative integration approaches may be more appropriate.

If downstream analyses produce results that are sensitive to small changes in Harmony parameters, the robustness of the findings should be questioned. Sensitivity analysis across parameter values can help identify stable biological conclusions versus parameter-dependent artifacts.

Researchers should also consider escalation when the computational resources required exceed what is available locally. While Harmony can integrate approximately one million cells on a personal computer [<a href="#ref-1">1</a>], larger datasets or more complex analyses may require access to high-performance computing infrastructure. In such cases, consultation with computational biology support services or collaboration with groups that have appropriate infrastructure may be necessary.

## Frequently Asked Questions

### What types of single-cell data can Harmony integrate?

Harmony can integrate scRNA-seq and single-nucleus RNA-seq data from multiple samples, donors, laboratories, and sequencing platforms. The algorithm has been applied to peripheral blood mononuclear cells, pancreatic islet cells, mouse embryogenesis datasets, and integration of scRNA-seq with spatial transcriptomics data [<a href="#ref-1">1</a>]. It can also integrate data from different technologies, as demonstrated in the original Harmony publication [<a href="#ref-1">1</a>].

### How does Harmony differ from Seurat's integration method?

Harmony and Seurat's integration approach use different mathematical frameworks. Harmony iteratively adjusts cell embeddings using a mixture model approach, while Seurat uses canonical correlation analysis and mutual nearest neighbors. The benchmark study comparing 14 methods found that Harmony, LIGER, and Seurat 3 were the recommended methods, with Harmony recommended as the first method to try due to its significantly shorter runtime [<a href="#ref-2">2</a>].

### What are the main parameters in Harmony and how should I tune them?

The main Harmony parameters include theta, which controls the diversity penalty, lambda, which controls regularization strength, and the maximum number of iterations. Default settings work for many datasets, but complex batch structures may require adjustment. Systematic parameter optimization approaches such as scAutoTune can help identify optimal settings using silhouette-based performance metrics [<a href="#ref-4">4</a>].

### Can Harmony handle datasets with more than two batches?

Yes, Harmony simultaneously accounts for multiple experimental and biological factors [<a href="#ref-1">1</a>]. The algorithm can integrate datasets with complex batch structures involving many batches and multiple batch variables. This capability was demonstrated in the original publication through analyses of peripheral blood mononuclear cells from datasets with large experimental differences and five studies of pancreatic islet cells [<a href="#ref-1">1</a>].

### How many cells can Harmony integrate?

Harmony enables the integration of approximately one million cells on a personal computer [<a href="#ref-1">1</a>]. This scalability makes it accessible to research groups without access to high-performance computing clusters. The benchmark study evaluating batch correction methods included a big data scenario to assess performance on large datasets [<a href="#ref-2">2</a>].

### Should I use Harmony for differential expression analysis?

Batch-corrected data from Harmony may not be suitable for all differential expression analyses. The benchmark study specifically investigated the use of batch-corrected data for differential gene expression and found that the choice of method affects results [<a href="#ref-2">2</a>]. Researchers should validate that differentially expressed genes identified after integration are biologically plausible and consider using original expression values with batch as a covariate in statistical models.

### How do I know if Harmony integration worked well?

Integration quality can be assessed through visual inspection of dimensionality reduction plots and quantitative metrics. Cells from different batches should be intermixed within cell type clusters. Quantitative metrics including kBET, LISI, ASW, and ARI provide objective evaluation of batch mixing and cell type preservation [<a href="#ref-2">2</a>]. The choice of metrics depends on whether ground truth cell type labels are available.

### What should I do if Harmony does not adequately correct batch effects?

If Harmony fails to achieve adequate batch mixing, several adjustments can be tried. Increasing the number of iterations, adjusting the theta parameter, or changing the number of principal components used as input may improve results. If these adjustments do not resolve the problem, alternative integration methods such as LIGER or Seurat 3 may be more suitable for the specific dataset [<a href="#ref-2">2</a>].

## Related Bioinformatics Guides

- [RNA-Seq Batch Effect Detection and Correction](/knowledge/bioinformatics/rna-seq-batch-effect-detection-and-correction)
- [Single-Cell RNA-Seq Normalization: Batch Effect Correction and Dimension Reduction (PCA, t-SNE, UMAP)](/knowledge/bioinformatics/scrna-seq-normalization-batch-correction)
- [Single-Cell Sequencing Depth: How Much Is Enough?](/knowledge/bioinformatics/single-cell-sequencing-depth-how-much-is-enough)
- [Benchmarking Atlas-Level Data Integration in Single-Cell Genomics: Methods and Best Practices](/knowledge/bioinformatics/benchmarking-atlas-level-data-integration-in-single-cell-genomics-methods-and-best-practices)
- [Single-Cell RNA Sequencing Quality Control: A Practical Guide to Filtering and Metrics](/knowledge/bioinformatics/single-cell-rna-sequencing-quality-control-a-practical-guide-to-filtering-and-metrics)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)

## References and Further Reading

<a id="ref-1"></a>[<a href="#ref-1">1</a>] [Fast, sensitive and accurate integration of single-cell data with Harmony.](https://pubmed.ncbi.nlm.nih.gov/31740819). Nature methods, 2019.

<a id="ref-2"></a>[<a href="#ref-2">2</a>] [A benchmark of batch-effect correction methods for single-cell RNA sequencing data.](https://pubmed.ncbi.nlm.nih.gov/31948481). Genome biology, 2020.

<a id="ref-3"></a>[<a href="#ref-3">3</a>] [Single-cell transcriptome and multi-omics integration reveal ferroptosis-driven immune microenvironment remodeling in knee osteoarthritis.](https://pubmed.ncbi.nlm.nih.gov/40636124). Frontiers in immunology, 2025.

<a id="ref-4"></a>[<a href="#ref-4">4</a>] [Protocol for automated optimization of feature selection and clustering parameters in single-cell RNA-seq using scAutoTune.](https://doi.org/10.1016/j.xpro.2026.104730). 2026.

<a id="ref-5"></a>[<a href="#ref-5">5</a>] [Integrated single-cell RNA-seq analysis identifies immune heterogeneity associated with KRAS/TP53 mutation status and tumor-sideness in colorectal cancers.](https://pubmed.ncbi.nlm.nih.gov/36172359). Frontiers in immunology, 2022.

<a id="ref-6"></a>[<a href="#ref-6">6</a>] [Single-cell RNA sequencing reveals immunological heterogeneity of the tumor microenvironment in acute myeloid leukemia.](https://doi.org/10.17219/acem/212574). 2026.

<a id="ref-7"></a>[<a href="#ref-7">7</a>] [The cancer stem cells characteristics analysis of LGR5 + cells that influence lung cancer risk by using single cell RNA-seq analysis](https://doi.org/10.1038/s41598-025-00585-3). Scientific Reports, 2025.

<a id="ref-8"></a>[<a href="#ref-8">8</a>] [Integrative analysis of bulk and single-cell RNA-seq reveals the molecular characterization of the immune microenvironment and oxidative stress signature in melanoma](https://doi.org/10.1016/j.heliyon.2024.e28244). Heliyon, 2024.

<a id="ref-9"></a>[<a href="#ref-9">9</a>] [Identification of HBEGF+ fibroblasts in the remission of rheumatoid arthritis by integrating single-cell RNA sequencing datasets and bulk RNA sequencing datasets.](https://pubmed.ncbi.nlm.nih.gov/36068607). Arthritis research & therapy, 2022.

<a id="ref-10"></a>[<a href="#ref-10">10</a>] [&lt,i&gt,T&lt,/i&gt, Gene Mutation Leads to Short Tail in Sheep via Premature AER Degeneration: Single-Cell Evidence from Embryos.](https://doi.org/10.3390/ani16111748). 2026.

<a id="ref-11"></a>[<a href="#ref-11">11</a>] [Single-cell RNA-seq analysis revealed the stemness of a specific cluster of B cells in acute lymphoblastic leukemia progression.](https://pubmed.ncbi.nlm.nih.gov/39465162). PeerJ, 2024.

<a id="ref-12"></a>[<a href="#ref-12">12</a>] [Neoplastic and immune single-cell transcriptomics define subgroup-specific intra-tumoral heterogeneity of childhood medulloblastoma.](https://pubmed.ncbi.nlm.nih.gov/34077540). Neuro-oncology, 2022.

<a id="ref-13"></a>[<a href="#ref-13">13</a>] [Identifying the anti-infection NK and T cell subgroups in preterm newborns with sepsis using single-cell RNA-seq analysis](https://doi.org/10.1186/s40001-025-03272-1). European Journal of Medical Research, 2025.

<a id="ref-14"></a>[<a href="#ref-14">14</a>] [HArmonized single-cell RNA-seq Cell type Assisted Deconvolution (HASCAD).](https://pubmed.ncbi.nlm.nih.gov/37907883). BMC medical genomics, 2023.

<a id="ref-15"></a>[<a href="#ref-15">15</a>] [nf-core Documentation](https://nf-co.re/docs). nf-core.

<a id="ref-16"></a>[<a href="#ref-16">16</a>] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.

<a id="ref-17"></a>[<a href="#ref-17">17</a>] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.

<a id="ref-18"></a>[<a href="#ref-18">18</a>] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.

<a id="ref-19"></a>[<a href="#ref-19">19</a>] [Bioconductor](https://bioconductor.org/). Bioconductor Project.

<a id="ref-20"></a>[<a href="#ref-20">20</a>] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.