# Doublet Detection in Single-Cell RNA-Seq: Comparing Computational Tools and Their Impact on Downstream Analysis


## Key Takeaways

- Doublets, formed by two or more cells sharing a single barcode, are a critical artifact in single-cell RNA sequencing (scRNA-seq) that distort gene expression, create false cell populations, and bias downstream analyses like differential expression and cell type proportion estimation.
- Computational doublet detection tools like Scrublet, DoubletFinder, and scDblFinder employ simulation-based classification or artificial doublet construction to identify these artifacts, with varying strengths in handling cell type heterogeneity and integration with existing bioinformatics workflows (e.g., Bioconductor).
- The choice of doublet detection tool is influenced by dataset characteristics (e.g., cell number, complexity, data modality like scRNA-seq vs. scATAC-seq), computational resources, and user expertise, with deep learning methods like Solo offering higher accuracy but requiring more specialized infrastructure.
- Robust doublet detection requires careful integration into a quality control pipeline, including initial filtering of empty droplets and low-quality cells, followed by evaluation of doublet scores, threshold setting, and assessment of the impact on downstream biological conclusions.
- Common failure patterns include overly aggressive doublet removal, under-detection of homotypic doublets, confounding batch effects, and misclassification of low-quality cells, necessitating validation through experimental methods (e.g., hashing), marker gene analysis, and comparison across multiple computational tools.

---

Doublet detection is a mandatory quality control step in single-cell RNA sequencing analysis. Doublets, also called multiplets, arise when two or more cells are captured within the same droplet or well and share one cell-identifying barcode. These hybrid molecular profiles distort gene expression measurements and can create false cell populations, alter differential expression results, and bias cell type proportion estimates. This article compares the most widely used computational doublet detection tools, including DoubletFinder, Scrublet, and scDblFinder, and explains how the choice of tool affects downstream analysis outcomes. The content is intended for biology students, researchers, laboratory professionals, and life science practitioners who need practical guidance for integrating doublet detection into single-cell RNA sequencing quality control pipelines.

## The Problem of Doublets in Single-Cell RNA Sequencing

Single-cell RNA sequencing technologies rely on the fundamental premise that each barcode corresponds to exactly one cell. When this premise fails, the resulting data violate the core assumption of the assay. Doublets can be homotypic, meaning two cells of the same type are captured together, or heterotypic, meaning two different cell types are captured together. Homotypic doublets are particularly difficult to detect because their expression profiles resemble the parent cell type. Heterotypic doublets produce mixed expression signatures that can be mistaken for transitional states, rare populations, or novel cell types.

The frequency of doublets depends on the loading concentration of cells into the microfluidic or droplet-based system. Higher cell loading increases throughput but also increases the probability of multiple cells occupying the same reaction compartment. Experimental doublet rates typically range from 1 to 10 percent depending on the platform and loading density. The 10x Genomics Chromium system and the BD Rhapsody platform both generate droplet-based libraries that are susceptible to doublet formation. A systematic comparison of these two high throughput platforms in complex tumor tissues found that both systems show similar gene sensitivity but exhibit cell type detection biases, including lower proportions of certain cell populations in one platform and reduced gene sensitivity in granulocytes for the other [16](https://doi.org/10.1016/j.heliyon.2024.e37185). These platform-associated differences underscore the need for robust computational doublet filtering regardless of the experimental platform used.

The consequences of failing to remove doublets extend beyond the creation of spurious clusters. Doublets can inflate the apparent abundance of cell types that are prone to stick together during dissociation, such as immune cells and endothelial cells. They can also generate false gene co-expression patterns that lead to incorrect inferences about cell state transitions or regulatory programs. In differential expression analysis, doublets introduce noise that reduces statistical power and can produce false positives. For these reasons, doublet detection should be considered a non-negotiable component of single-cell RNA sequencing quality control.

## Core Principles of Computational Doublet Detection

Computational doublet detection methods generally operate on one of two principles: simulation-based classification or artificial doublet construction. Simulation-based methods model the expected expression profile of a doublet by combining the profiles of two cells and then identify real cells that resemble these simulated doublets. Artificial doublet construction methods create synthetic doublets by averaging or summing the expression profiles of randomly paired cells from the dataset and then train a classifier to distinguish real cells from synthetic doublets.

Scrublet is one of the earliest widely adopted tools for doublet detection. It works by simulating doublets from the observed data and then calculating a doublet score for each cell based on its proximity to the simulated doublets in gene expression space. Scrublet also estimates the expected doublet rate and uses this to determine a threshold for classifying cells as doublets. The method is designed to work within a single sample and requires the user to specify the expected doublet rate or allow the algorithm to estimate it from the data.

DoubletFinder follows a similar simulation-based approach but incorporates additional steps to account for the heterogeneity of cell types within a sample. It first performs principal component analysis and clustering to identify cell states, then simulates doublets from the data, and finally uses a classifier to distinguish real cells from artificial doublets. DoubletFinder requires the user to specify the expected number of doublets, which can be estimated from the loading concentration or determined empirically.

scDblFinder is a more recent method that builds on the principles of Scrublet and DoubletFinder but adds several improvements. It uses a combination of clustering and iterative classification to identify doublets, and it can incorporate information about the number of cells loaded and the expected doublet rate. scDblFinder is available through the Bioconductor project, which provides official documentation for package installation and reproducible genomic analysis workflows [3](https://bioconductor.org/). The Bioconductor ecosystem offers a range of tools for single-cell analysis, and scDblFinder is designed to integrate with these workflows.

Solo represents a different approach to doublet detection. It uses a semi-supervised deep learning method that embeds cells using a variational autoencoder and then appends a feed-forward neural network layer to form a supervised classifier. The classifier is trained to distinguish simulated doublets from observed data. Solo has been shown to identify doublets with greater accuracy than existing methods and can be combined with experimental doublet detection methods to further purify single-cell RNA sequencing data [15](https://doi.org/10.1016/j.cels.2020.05.010).

The choice among these tools depends on several factors, including the number of cells in the dataset, the complexity of the cell types present, the availability of computational resources, and the user's familiarity with different programming environments. Each tool has strengths and limitations that should be considered in the context of the specific research question.

## At a Glance: Comparison of Major Doublet Detection Tools

The following table summarizes the key characteristics of the major computational doublet detection tools discussed in this article. This comparison is intended to help analysts select an appropriate tool for their specific dataset and research context.

| Tool | Underlying Principle | Input Requirements | Key Strengths | Key Limitations |
|------|---------------------|-------------------|---------------|-----------------|
| Scrublet | Simulation-based doublet scoring | Raw or normalized count matrix, expected doublet rate | Fast, works well on individual samples, estimates doublet rate from data | May struggle with highly heterogeneous samples, requires careful threshold selection |
| DoubletFinder | Simulation-based with clustering integration | Count matrix, expected doublet rate, principal component number | Accounts for cell type heterogeneity, widely used and validated | Requires specification of expected doublet number, sensitive to parameter choices |
| scDblFinder | Clustering and iterative classification | Count matrix, optional cell loading information | Integrates with Bioconductor workflows, handles multi-sample data, improved accuracy | More computationally intensive, requires familiarity with Bioconductor |
| Solo | Semi-supervised deep learning | Count matrix, simulated doublets for training | High accuracy, can be combined with experimental methods | Requires deep learning framework, less accessible to novice users |

The choice of tool should be guided by the specific characteristics of the dataset and the analytical goals. For example, a researcher working with a single sample of moderate complexity might find Scrublet sufficient, while a researcher analyzing multiple samples with diverse cell types might benefit from the more sophisticated approach of scDblFinder.

## Practical Workflow for Integrating Doublet Detection

Integrating doublet detection into a single-cell RNA sequencing quality control pipeline requires careful planning and execution. The following workflow outlines the key steps and decision points.

### Step 1: Generate the Count Matrix

The first step is to generate a count matrix from the raw sequencing data. This is typically done using platform-specific software such as Cell Ranger for 10x Genomics data or the BD Rhapsody analysis pipeline. The count matrix contains the number of unique molecular identifiers or transcripts detected for each gene in each cell barcode. The quality of the count matrix directly affects the performance of downstream doublet detection tools, so it is important to follow the manufacturer's recommendations for alignment and quantification. The National Center for Biotechnology Information provides access to sequence read archives and related databases that can support data management and retrieval for these analyses [1](https://www.ncbi.nlm.nih.gov/).

### Step 2: Perform Initial Quality Filtering

Before running doublet detection, it is advisable to perform basic quality filtering to remove empty droplets and low-quality cells. Empty droplets contain ambient RNA but no intact cell, and they can be identified by their low total transcript counts. Low-quality cells may have high mitochondrial content, indicating cell lysis or stress, or very low gene detection, indicating failed library preparation. Removing these obvious artifacts before doublet detection reduces the computational burden and improves the accuracy of the doublet classifier.

The scReady pipeline provides an automated approach to this preprocessing step. It integrates essential quality control steps, including ambient RNA removal, doublet detection, and cell and gene filtering based on mitochondrial content and customizable thresholds. scReady is designed as a containerized pipeline that can run on single machines or high performance computing clusters, making it accessible to laboratories without dedicated bioinformatics support. The pipeline outputs a fully processed Seurat object along with diagnostic plots and a comprehensive quality control report [17](https://doi.org/10.12688/wellcomeopenres.25152.1).

### Step 3: Run Doublet Detection

Once the initial quality filtering is complete, the analyst can run one or more doublet detection tools. The choice of tool should be documented, and the parameters should be recorded for reproducibility. For simulation-based tools, the expected doublet rate is a critical parameter. This can be estimated from the cell loading concentration or determined empirically by running the tool with different settings and examining the distribution of doublet scores.

For datasets with multiple samples, it is generally recommended to run doublet detection on each sample separately instead of on the pooled dataset. This is because doublet rates can vary between samples, and pooling can introduce batch effects that confound the doublet classifier. The pipeComp framework provides a general approach for evaluating pipeline performance and can be used to compare different doublet detection strategies across multiple samples [7](https://pubmed.ncbi.nlm.nih.gov/32873325).

### Step 4: Evaluate Doublet Scores and Set Thresholds

After running doublet detection, the analyst must decide where to set the threshold for classifying cells as doublets. Most tools provide a doublet score for each cell, and the threshold determines which cells are removed. Some tools provide an estimated doublet rate that can guide threshold selection. The analyst should examine the distribution of doublet scores and consider the expected doublet rate based on the experimental design.

It is important to recognize that doublet detection is not perfect. Computational methods can miss doublets, particularly homotypic doublets that resemble their parent cell type, and they can also falsely classify single cells as doublets. The tradeoff between sensitivity and specificity should be considered in the context of the downstream analysis. For example, if the goal is to identify rare cell populations, a more aggressive doublet removal strategy may be appropriate to avoid false discoveries. If the goal is to preserve as many cells as possible for differential expression analysis, a more conservative threshold may be preferred.

### Step 5: Assess the Impact on Downstream Analysis

After removing predicted doublets, the analyst should assess the impact on downstream analysis. This can be done by comparing clustering results, cell type proportions, and differential expression results with and without doublet removal. If the removal of doublets substantially changes the biological conclusions, this should be reported and discussed.

The scUmaper framework provides an automated approach to doublet removal and cell type annotation. It integrates quality control, biologically grounded doublet filtering, and marker library based cell type annotation. scUmaper codifies lineage marker incompatibility rules and applies global clustering followed by within lineage re-clustering to reveal anomalous subclusters with implausible cross lineage co-expression. Across six public human organ datasets, scUmaper removed additional high confidence heterotypic doublets that were retained by simulation based approaches and achieved annotation agreement comparable to or higher than commonly used R based baselines [8](https://doi.org/10.1016/j.isci.2026.115850).

## Options and Tradeoffs in Doublet Detection

The choice of doublet detection tool involves several tradeoffs that should be considered in the context of the specific research question and dataset characteristics.

### Simulation-Based Methods versus Deep Learning

Simulation-based methods such as Scrublet and DoubletFinder are widely used and well validated. They are relatively fast and require minimal computational resources. However, their accuracy depends on the quality of the simulated doublets, which are generated by combining the expression profiles of randomly paired cells. If the dataset contains cell types that are not well represented, the simulated doublets may not accurately reflect the true doublet population.

Deep learning methods such as Solo offer the potential for higher accuracy by learning complex patterns in the data. However, they require more computational resources and expertise to implement. Solo can be applied in combination with experimental doublet detection methods to further purify single-cell RNA sequencing data to true single cells [15](https://doi.org/10.1016/j.cels.2020.05.010).

### Single-Sample versus Multi-Sample Approaches

Some doublet detection tools are designed to work on individual samples, while others can handle multiple samples simultaneously. Scrublet and DoubletFinder are typically run on individual samples, while scDblFinder can integrate information across samples. The choice between these approaches depends on the experimental design. If samples are processed in separate batches, running doublet detection on each sample separately may be more appropriate. If samples are pooled and processed together, a multi-sample approach may be necessary.

### Parameter Sensitivity

All doublet detection tools require the user to specify certain parameters, and the results can be sensitive to these choices. For example, the expected doublet rate is a critical parameter for Scrublet and DoubletFinder. If this parameter is set too low, the tool may miss doublets. If it is set too high, the tool may falsely classify single cells as doublets. The analyst should explore the parameter space and document the choices made.

The pipeComp framework provides a systematic approach to evaluating the impact of parameter choices on pipeline performance. It can handle interactions between analysis steps and relies on multi level evaluation metrics. pipeComp can easily integrate any other step, tool, or evaluation metric, allowing extensible benchmarks and easy applications to other fields [7](https://pubmed.ncbi.nlm.nih.gov/32873325).

## Observations and Measurements in Doublet Detection

Several studies have provided empirical evidence about the performance of doublet detection tools and the impact of doublets on downstream analysis. These observations can guide the selection of tools and the interpretation of results.

### Performance on Plant Single-Cell RNA Sequencing Data

A benchmark study of plant single-cell RNA sequencing sample processing strategies compared protoplast enrichment technologies and single-cell RNA sequencing platforms using Arabidopsis roots. The study found that both the 10x Genomics Chromium and BD Rhapsody platforms captured root cell heterogeneity and yielded reproducible gene expression profiles, but showed platform associated differences in cell type composition. Notably, single nucleotide polymorphism analysis of a mixed ecotype sample revealed that among cells identified as doublets by computational algorithms, two thirds were likely to have been misclassified [9](https://doi.org/10.1038/s44318-026-00800-5). This finding highlights the limitations of computational doublet detection and the importance of experimental validation.

### Performance on Multimodal Data

The SEBULA method was developed for multiplet detection in single-nucleus ATAC sequencing data. SEBULA models the singlet background directly from observed chromatin accessibility signals using fragment level information, avoiding reliance on synthetic doublets. The method produces classification probabilities that enable direct false discovery rate control. Across simulations and seven multimodal datasets with hashing based ground truth, SEBULA demonstrated improved sensitivity and specificity compared with existing single-nucleus ATAC sequencing methods. The evidence integration framework achieved comparable or superior performance relative to state of the art multiomic approaches while maintaining computational efficiency [10](https://doi.org/10.1371/journal.pcbi.1013653).

### Impact on Cell Type Annotation

The scUmaper framework demonstrated that additional high confidence heterotypic doublets can be removed beyond those identified by simulation based approaches. By applying lineage marker incompatibility rules and global clustering followed by within lineage re-clustering, scUmaper revealed anomalous subclusters with implausible cross lineage co-expression. This approach improved cell type annotation accuracy and reduced the impact of doublets on downstream analysis [8](https://doi.org/10.1016/j.isci.2026.115850).

### Platform-Specific Considerations

The choice of sequencing platform can influence doublet detection outcomes. A comparison of the 10x Chromium and BD Rhapsody platforms in complex tissues found that the source of ambient noise differed between plate-based and droplet-based platforms, and cell type detection biases were observed between platforms [16](https://doi.org/10.1016/j.heliyon.2024.e37185). These platform-specific differences should be considered when interpreting doublet detection results and when comparing datasets generated on different platforms.

## Records and Documentation for Reproducible Doublet Detection

Reproducibility is a fundamental principle of bioinformatics analysis. Doublet detection should be documented thoroughly to ensure that results can be reproduced and interpreted correctly. The following records should be maintained for each dataset.

### Analysis Parameters

All parameters used for doublet detection should be recorded, including the tool version, the expected doublet rate, the number of principal components used, and any threshold values. This information should be included in the methods section of any publication or report. The EMBL-EBI Training program offers learning pathways that cover best practices for documenting bioinformatics analyses [2](https://www.ebi.ac.uk/training).

### Quality Control Metrics

Quality control metrics should be recorded before and after doublet removal. These include the number of cells, the number of genes detected per cell, the total number of transcripts per cell, and the mitochondrial content. Comparing these metrics before and after doublet removal can help assess the impact of the filtering step.

### Diagnostic Plots

Diagnostic plots should be generated to visualize the doublet scores and the distribution of cells before and after filtering. These plots can help identify potential issues with the doublet detection process and provide evidence for the quality of the final dataset.

### Version Control

All scripts and code used for doublet detection should be maintained under version control. This ensures that the analysis can be reproduced exactly and that any changes to the analysis can be tracked. The Carpentries provides foundational training in version control with Git, which is essential for reproducible bioinformatics analysis [6](https://carpentries.org/lessons).

### Pipeline Documentation

For laboratories using automated pipelines, documentation should include the pipeline version, configuration parameters, and any custom modifications. The nf-core documentation provides standards for community pipelines that emphasize reproducibility and configuration tracking [5](https://nf-co.re/docs). The Galaxy Training Network also offers accessible workflow training that emphasizes reproducibility in analysis [4](https://training.galaxyproject.org/).

## Common Failure Patterns in Doublet Detection

Several common failure patterns can compromise the effectiveness of doublet detection. Recognizing these patterns can help analysts troubleshoot their pipelines and interpret results correctly.

### Overly Aggressive Doublet Removal

One common failure is the removal of too many cells as predicted doublets. This can occur when the expected doublet rate is set too high or when the threshold for doublet classification is too permissive. Overly aggressive doublet removal can eliminate legitimate cell populations, particularly rare cell types that may have unusual expression profiles. This can bias downstream analysis and lead to incorrect biological conclusions.

### Under-Detection of Homotypic Doublets

Homotypic doublets, which consist of two cells of the same type, are difficult to detect because their expression profiles resemble the parent cell type. Most computational methods are better at detecting heterotypic doublets, which have mixed expression signatures. The under-detection of homotypic doublets can inflate the apparent abundance of certain cell types and distort cell type proportion estimates.

### Batch Effects Confounding Doublet Detection

When multiple samples are pooled for analysis, batch effects can confound doublet detection. Cells from different batches may have systematic differences in gene expression that are unrelated to doublet status. This can cause the doublet classifier to incorrectly identify cells from one batch as doublets or to miss doublets in another batch. Running doublet detection on each sample separately can help mitigate this issue.

### Misclassification of Low-Quality Cells

Low-quality cells, such as those with high mitochondrial content or low gene detection, can be misclassified as doublets by computational methods. This is because their expression profiles may resemble the combined profiles of two cells. Careful quality filtering before doublet detection can help reduce this problem.

### Over-Reliance on Default Parameters

Many analysts use default parameters without understanding their implications. Default parameters may not be appropriate for all datasets, particularly those with unusual cell type compositions or high ambient RNA contamination. The analyst should understand what each parameter controls and how it affects the results.

## Limitations of Computational Doublet Detection

Computational doublet detection methods have inherent limitations that should be acknowledged when interpreting results.

### Inability to Detect All Doublets

No computational method can detect all doublets. The benchmark study of plant single-cell RNA sequencing data found that two thirds of cells identified as doublets by computational algorithms were likely to have been misclassified [9](https://doi.org/10.1038/s44318-026-00800-5). This suggests that computational methods may have high false positive rates in some contexts. Conversely, computational methods may miss doublets that are not well represented in the simulated doublet population.

### Dependence on Data Quality

The accuracy of doublet detection depends on the quality of the input data. Datasets with high ambient RNA contamination, high dropout rates, or significant batch effects will produce less reliable doublet predictions. The scReady pipeline addresses this by integrating ambient RNA removal and quality filtering before doublet detection [17](https://doi.org/10.12688/wellcomeopenres.25152.1).

### Computational Resource Requirements

Some doublet detection methods, particularly deep learning approaches, require substantial computational resources. This can be a barrier for laboratories without access to high performance computing infrastructure. The nf-core documentation provides guidance on configuring and running reproducible workflows on various computing environments, which can help address this challenge [5](https://nf-co.re/docs).

### Interpretation Challenges

Doublet scores are continuous values, and the threshold for classifying cells as doublets is somewhat arbitrary. Different thresholds can produce substantially different results, and there is no universally correct threshold. Analysts should consider the tradeoff between sensitivity and specificity in the context of their specific research question.

### Context-Dependent Performance

The performance of doublet detection tools can vary depending on the tissue type, the dissociation protocol, and the sequencing platform. A tool that performs well on one dataset may perform poorly on another. The pipeComp framework provides a systematic approach to evaluating tool performance in different contexts [7](https://pubmed.ncbi.nlm.nih.gov/32873325).

## Quality Controls and Validation Strategies

Several strategies can be used to validate doublet detection results and ensure the quality of the final dataset.

### Experimental Doublet Detection

Experimental methods for doublet detection, such as hashing or multiplexing, can provide ground truth for validating computational methods. In hashing-based approaches, cells from different samples are labeled with distinct oligonucleotide tags, and doublets can be identified by the presence of multiple tags. The SEBULA study used hashing based ground truth to validate the performance of the method across seven multimodal datasets [10](https://doi.org/10.1371/journal.pcbi.1013653).

### Marker Gene Validation

After doublet removal, the analyst should verify that known marker genes for expected cell types are expressed in the appropriate cell populations. If doublet removal has eliminated a cell population that should be present based on biological knowledge, this may indicate that the doublet detection was too aggressive.

### Comparison of Results Across Tools

Running multiple doublet detection tools and comparing the results can provide confidence in the doublet calls. Cells that are consistently identified as doublets by multiple tools are more likely to be true doublets. The pipeComp framework provides a systematic approach to comparing different tools and evaluating their performance [7](https://pubmed.ncbi.nlm.nih.gov/32873325).

### Assessment of Downstream Impact

The ultimate test of doublet detection is its impact on downstream analysis. If the removal of predicted doublets changes the biological conclusions, this should be investigated carefully. The scUmaper framework provides an interpretable and extensible framework for assessing the impact of doublet removal on cell type annotation [8](https://doi.org/10.1016/j.isci.2026.115850).

### Single-Nucleus RNA Sequencing Considerations

For single-nucleus RNA sequencing data, doublet detection presents additional challenges. The SEBULA method was specifically developed for multiplet detection in single-nucleus ATAC sequencing data and demonstrated that fragment level information can improve detection accuracy [10](https://doi.org/10.1371/journal.pcbi.1013653). Analysts working with single-nucleus data should consider whether the doublet detection tool was validated for this data type.

## Safety and Regulatory Context

Doublet detection is a data analysis step and does not involve direct safety or regulatory considerations in the same way as laboratory procedures. However, there are important considerations related to data integrity and reproducibility that have implications for research quality and regulatory compliance.

### Data Integrity

Doublet detection affects the integrity of the final dataset. Removing doublets is a form of data filtering that can introduce bias if not performed carefully. Researchers should document their doublet detection approach thoroughly and be transparent about the limitations of the method.

### Reproducibility Requirements

Many funding agencies and journals now require that bioinformatics analyses be reproducible. This means that the code, parameters, and data used for doublet detection should be made available to other researchers. The Galaxy Training Network provides accessible workflow training and analysis tutorials that emphasize reproducibility [4](https://training.galaxyproject.org/), and the nf-core documentation provides standards for community pipelines [5](https://nf-co.re/docs).

### Professional Escalation Criteria

If doublet detection results are inconsistent across tools or if the removal of doublets substantially changes the biological conclusions, this should be escalated to a bioinformatics specialist or a statistician with expertise in single-cell analysis. Similarly, if the doublet rate is unexpectedly high or low, this may indicate a problem with the experimental protocol that should be investigated.

## Professional Escalation Criteria

Analysts should seek professional guidance in the following situations.

### Unexpectedly High Doublet Rates

If the estimated doublet rate is substantially higher than expected based on the cell loading concentration, this may indicate a problem with the experimental protocol. The analyst should consult with the laboratory team to investigate potential causes, such as cell aggregation during dissociation or issues with the microfluidic system.

### Inconsistent Results Across Tools

If different doublet detection tools produce substantially different results, this may indicate that the dataset has unusual characteristics that are not well handled by standard methods. A bioinformatics specialist should be consulted to investigate the cause and recommend an appropriate approach.

### Major Impact on Biological Conclusions

If the removal of predicted doublets changes the biological conclusions of the study, this should be investigated carefully. The analyst should consider whether the doublet detection approach is appropriate for the specific dataset and whether additional validation is needed.

### Lack of Reproducibility

If the doublet detection results cannot be reproduced, this indicates a problem with the analysis pipeline. The analyst should review the code and parameters to identify the source of the inconsistency.

### Integration with Downstream Analyses

When doublet detection results are used to inform downstream analyses such as cell-cell communication inference or differential detection, the impact of doublet removal should be assessed in the context of these analyses. The choice of doublet detection method can influence the results of downstream analyses, and this should be considered when interpreting findings [13](https://doi.org/10.1038/s41467-022-30755-0) [14](https://doi.org/10.1186/s12864-025-12102-x).

## A Practical Decision Framework for Selecting Doublet Detection Tools

Selecting a doublet detection tool requires a structured approach that accounts for dataset characteristics, available computational resources, and downstream analysis goals. The following decision framework provides a systematic method for choosing among Scrublet, DoubletFinder, scDblFinder, Solo, and emerging methods such as SEBULA and scUmaper. This framework is designed to be applied before running any doublet detection analysis and should be revisited if the dataset changes or if initial results suggest the chosen tool is performing poorly.

### Step 1: Characterize the Dataset

Begin by documenting the key features of the dataset that influence doublet detection tool selection. Record the number of cells per sample, the number of samples in the study, the sequencing platform used, and the expected doublet rate based on cell loading concentration. Also note whether the data are from single-cell RNA sequencing, single-nucleus RNA sequencing, or single-nucleus ATAC sequencing, as this determines which tools are appropriate. The SEBULA method was specifically developed for multiplet detection in single-nucleus ATAC sequencing data and is not designed for standard single-cell RNA sequencing analysis [10](https://doi.org/10.1371/journal.pcbi.1013653). Similarly, tools validated primarily on droplet-based platforms may perform differently on plate-based data.

For plant samples or other tissues with high cell wall content, additional considerations apply. A benchmark study of plant single-cell RNA sequencing found that computational doublet detection algorithms misclassified two thirds of identified doublets in a mixed ecotype Arabidopsis root sample [9](https://doi.org/10.1038/s44318-026-00800-5). This finding suggests that for plant datasets, analysts should consider using multiple tools and validating results with experimental approaches such as single nucleotide polymorphism analysis when possible.

### Step 2: Assess Computational Resources and User Expertise

The available computational infrastructure and the analyst's programming proficiency are practical constraints that narrow the tool selection. Scrublet and DoubletFinder require modest computational resources and can run on a standard desktop computer with R or Python installed. scDblFinder is available through the Bioconductor project and integrates with existing Bioconductor workflows, making it a natural choice for analysts already working in that ecosystem [3](https://bioconductor.org/). Solo requires a deep learning framework and is less accessible to novice users [15](https://doi.org/10.1016/j.cels.2020.05.010).

For laboratories without dedicated bioinformatics support, automated pipelines can reduce the barrier to implementing doublet detection. The scReady pipeline integrates doublet detection with ambient RNA removal and quality filtering in a containerized format that runs on single machines or high performance computing clusters [17](https://doi.org/10.12688/wellcomeopenres.25152.1). The Galaxy Training Network provides accessible workflow training that can help analysts learn to use these tools without extensive programming experience [4](https://training.galaxyproject.org/).

### Step 3: Match Tool Characteristics to Dataset Features

The following decision points guide tool selection based on dataset features.

For a single sample with fewer than 10,000 cells and moderate cell type diversity, Scrublet is often sufficient. It is fast, requires minimal parameter tuning, and estimates the doublet rate from the data. The analyst should verify that the estimated doublet rate is consistent with the expected rate based on cell loading concentration.

For a single sample with high cell type diversity or known rare populations, DoubletFinder may be more appropriate because it incorporates clustering information to account for cell type heterogeneity. The analyst must specify the expected number of doublets, which requires careful estimation from the loading concentration or empirical testing.

For multi-sample studies where samples were processed in separate batches, scDblFinder is recommended because it can handle multi-sample data and integrates with Bioconductor workflows. Running doublet detection on each sample separately is also acceptable, but the analyst should document the parameters used for each sample to ensure consistency.

For datasets with known challenges such as high ambient RNA contamination or significant batch effects, consider using scReady, which integrates ambient RNA removal before doublet detection [17](https://doi.org/10.12688/wellcomeopenres.25152.1). The source of ambient noise differs between plate-based and droplet-based platforms, and this can affect doublet detection performance [16](https://doi.org/10.1016/j.heliyon.2024.e37185).

For single-nucleus ATAC sequencing data, SEBULA is the appropriate choice because it models the singlet background directly from observed chromatin accessibility signals and produces classification probabilities that enable direct false discovery rate control [10](https://doi.org/10.1371/journal.pcbi.1013653).

For datasets where cell type annotation is a primary goal, scUmaper provides an integrated approach that combines doublet filtering with marker library based cell type annotation. It codifies lineage marker incompatibility rules and applies global clustering followed by within lineage re-clustering to reveal anomalous subclusters with implausible cross lineage co-expression [8](https://doi.org/10.1016/j.isci.2026.115850).

### Step 4: Run the Selected Tool and Document Parameters

After selecting a tool, run it with documented parameters. Record the tool version, the expected doublet rate or number of doublets, the number of principal components used, and any threshold values. The EMBL-EBI Training program offers learning pathways that cover best practices for documenting bioinformatics analyses [2](https://www.ebi.ac.uk/training). The Carpentries provides foundational training in version control with Git, which is essential for tracking changes to analysis scripts [6](https://carpentries.org/lessons).

### Step 5: Evaluate Results and Apply the Escalation Criteria

After running the selected tool, evaluate the results using the following criteria. If the estimated doublet rate is substantially higher than expected based on cell loading concentration, consult with the laboratory team to investigate potential causes such as cell aggregation during dissociation. If different tools produce substantially different results, consult a bioinformatics specialist to investigate the cause. If the removal of predicted doublets changes the biological conclusions of the study, investigate whether the doublet detection approach is appropriate for the specific dataset.

### Step 6: Validate with Multiple Tools When Uncertainty Is High

For datasets where the cost of false positives or false negatives is high, running multiple tools and comparing the results can provide confidence in the doublet calls. Cells consistently identified as doublets by multiple tools are more likely to be true doublets. The pipeComp framework provides a systematic approach to comparing different tools and evaluating their performance across multiple samples [7](https://pubmed.ncbi.nlm.nih.gov/32873325). This framework can handle interactions between analysis steps and relies on multi level evaluation metrics, making it suitable for benchmarking doublet detection strategies in the context of a complete analysis pipeline.

### Step 7: Record the Decision and Rationale

Document the tool selection decision and the rationale in the analysis notebook or methods section. This documentation should include the dataset characteristics that informed the decision, the parameters used, and the results of any validation steps. The nf-core documentation provides standards for community pipelines that emphasize reproducibility and configuration tracking [5](https://nf-co.re/docs). The National Center for Biotechnology Information provides access to sequence read archives and related databases that can support data management and retrieval for these analyses [1](https://www.ncbi.nlm.nih.gov/).

### Common Decision Errors

Several common errors can compromise the tool selection process. Selecting a tool based on familiarity instead of dataset characteristics can lead to suboptimal results. Using default parameters without understanding their implications is a frequent mistake, particularly for tools that require specification of the expected doublet rate. Applying a tool validated on one data type to a different data type without verification can produce unreliable results. Ignoring platform specific differences in doublet formation and ambient noise can bias the analysis. Failing to document the tool selection rationale makes it difficult to reproduce the analysis or explain the results to collaborators.

### Integration with Downstream Analysis Planning

The choice of doublet detection tool should be made with awareness of the downstream analysis plan. If the goal is to infer cell-cell communication, the impact of doublet removal on the inferred interactions should be considered. A comparison of cell-cell communication inference methods found that both the choice of resource and method strongly influence the predicted intercellular interactions [13](https://doi.org/10.1038/s41467-022-30755-0). If the goal is differential detection, which infers differences in the average fraction of cells in which expression is detected, doublet removal can affect the results by changing the cell population composition [14](https://doi.org/10.1186/s12864-025-12102-x). The analyst should assess the impact of doublet removal on these downstream analyses and report any substantial changes.

### Revisiting the Decision

The tool selection should be revisited if the dataset changes, if initial results suggest poor performance, or if new tools become available that are better suited to the dataset characteristics. The field of doublet detection is evolving rapidly, and new methods such as SEBULA and scUmaper have expanded the available options [8](https://doi.org/10.1016/j.isci.2026.115850) [10](https://doi.org/10.1371/journal.pcbi.1013653). Analysts should periodically review the literature and consider whether newer methods offer advantages for their specific data types and research questions.

## Frequently Asked Questions

### What is the difference between homotypic and heterotypic doublets?

Homotypic doublets consist of two cells of the same type, while heterotypic doublets consist of two different cell types. Homotypic doublets are more difficult to detect computationally because their expression profiles resemble the parent cell type. Heterotypic doublets produce mixed expression signatures that can be mistaken for transitional states or rare populations.

### How do I estimate the expected doublet rate for my dataset?

The expected doublet rate can be estimated from the cell loading concentration used in the experiment. Most droplet-based platforms provide guidelines for the relationship between loading concentration and doublet rate. Alternatively, the doublet rate can be estimated empirically by running doublet detection tools with different settings and examining the distribution of doublet scores.

### Should I run doublet detection on each sample separately or on the pooled dataset?

It is generally recommended to run doublet detection on each sample separately, particularly if samples were processed in different batches. This is because doublet rates can vary between samples, and pooling can introduce batch effects that confound the doublet classifier. Some tools, such as scDblFinder, can handle multi-sample data, but the analyst should carefully consider whether this is appropriate for the specific experimental design.

### Can computational doublet detection replace experimental methods?

Computational doublet detection cannot fully replace experimental methods. Experimental methods, such as hashing or multiplexing, provide ground truth that can be used to validate computational approaches. The benchmark study of plant single-cell RNA sequencing data found that two thirds of cells identified as doublets by computational algorithms were likely to have been misclassified [9](https://doi.org/10.1038/s44318-026-00800-5), highlighting the limitations of computational methods.

### What is the impact of doublets on differential expression analysis?

Doublets introduce noise into differential expression analysis by creating hybrid expression profiles that do not correspond to any real cell type. This can reduce statistical power and produce false positives. Removing doublets before differential expression analysis can improve the accuracy of the results.

### How do I choose between Scrublet, DoubletFinder, and scDblFinder?

The choice depends on the characteristics of the dataset and the analytical goals. Scrublet is fast and works well on individual samples. DoubletFinder accounts for cell type heterogeneity and is widely used. scDblFinder integrates with Bioconductor workflows and handles multi-sample data [3](https://bioconductor.org/). The analyst should consider the computational resources available, the complexity of the dataset, and the need for integration with other analysis tools.

### What should I do if doublet detection removes too many cells?

If doublet detection removes too many cells, the analyst should review the parameters used, particularly the expected doublet rate and the threshold for doublet classification. A more conservative threshold may be appropriate if the goal is to preserve as many cells as possible. The analyst should also verify that the removed cells are not legitimate cell populations by examining marker gene expression.

### How can I validate the results of computational doublet detection?

Computational doublet detection results can be validated using experimental methods, such as hashing or multiplexing, or by comparing results across multiple tools. The analyst should also assess the impact of doublet removal on downstream analysis, including clustering results and cell type proportions. The pipeComp framework provides a systematic approach to comparing different tools and evaluating their performance [7](https://pubmed.ncbi.nlm.nih.gov/32873325).

## Related Bioinformatics Guides

- [Single-Cell RNA Sequencing Depth: A Cost-Benefit Analysis for Experimental Design](/knowledge/bioinformatics/single-cell-rna-sequencing-depth-a-cost-benefit-analysis-for-experimental-design)
- [Single-Cell Sequencing Depth: How Much Is Enough?](/knowledge/bioinformatics/single-cell-sequencing-depth-how-much-is-enough)
- [Single-Cell RNA Sequencing Quality Control: A Practical Guide to Filtering and Metrics](/knowledge/bioinformatics/single-cell-rna-sequencing-quality-control-a-practical-guide-to-filtering-and-metrics)
- [Single-Cell RNA-Seq Analysis Pipelines for Veterinary Immunology](/knowledge/bioinformatics/single-cell-rna-seq-analysis-pipelines-for-veterinary-immunology)
- [RNA-Seq vs qPCR: Validation and Comparison](/knowledge/bioinformatics/rna-seq-vs-qpcr-validation-and-comparison)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)


## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [pipeComp, a general framework for the evaluation of computational pipelines, reveals performant single cell RNA-seq preprocessing tools.](https://pubmed.ncbi.nlm.nih.gov/32873325). Genome biology, 2020.
- [scUmaper: An automated framework for doublet removal and cell-type annotation in single-cell transcriptomics.](https://doi.org/10.1016/j.isci.2026.115850). 2026.
- [Benchmarking plant single cell RNA-sequencing sample processing strategies.](https://doi.org/10.1038/s44318-026-00800-5). 2026.
- [Semi-parametric empirical bayes method for multiplet detection in snATAC-seq with probabilistic multi-omic integration.](https://doi.org/10.1371/journal.pcbi.1013653). 2026.
- [Multi-Dimensional Transcriptomics Reveals the Prominent Role of Neuroinflammation in Alzheimer's Disease.](https://doi.org/10.3390/ijms27115020). 2026.
- [Resolving sensitivity, specificity and signal contamination in Xenium spatial transcriptomics.](https://doi.org/10.1038/s41592-026-03089-8). 2026.
- [Comparison of methods and resources for cell-cell communication inference from single-cell RNA-Seq data](https://doi.org/10.1038/s41467-022-30755-0). Nature Communications, 2022.
- [Differential detection workflows for multi-sample single-cell RNA-seq data](https://doi.org/10.1186/s12864-025-12102-x). BMC Genomics, 2025.
- [Solo: Doublet Identification in Single-Cell RNA-Seq via Semi-Supervised Deep Learning.](https://doi.org/10.1016/j.cels.2020.05.010). Cell Systems, 2020.
- [Performance comparison of high throughput single-cell RNA-Seq platforms in complex tissues](https://doi.org/10.1016/j.heliyon.2024.e37185). bioRxiv, 2024.
- [scReady - an automated and accessible pipeline for single-cell RNA-Seq preprocessing: Empowering novice bioinformaticians](https://doi.org/10.12688/wellcomeopenres.25152.1). Wellcome Open Research, 2026.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.