Doublet Identification and Mitigation in Single-Cell Sequencing: From Library Preparation to Computational Filtering

By Dr. Zubair Khalid, DVM, MS, PhD ·

Doublet Identification and Mitigation in Single-Cell Sequencing: From Library Preparation to Computational Filtering

Key Takeaways

  • Doublets, where two or more cells share a barcode, are a fundamental artifact in single-cell sequencing (scRNA-seq, scATAC-seq) that can mimic transitional states or rare cell types by generating hybrid transcriptomes.
  • Doublet formation is directly influenced by library preparation parameters, with higher cell loading densities on droplet-based platforms and multiple cells settling into wells on well-based platforms increasing doublet rates.
  • Sample multiplexing strategies, such as MULTI-seq (lipid-tagged indices) and cell hashing (antibody-tagged barcodes), are crucial for identifying cross-sample doublets and improving cost-efficiency by enabling pooled runs.
  • Computational doublet detection methods fall into categories including simulation-based (e.g., Scrublet, DoubletFinder, Solo), cluster-based (e.g., scDblFinder, DoubletDecon, scUmaper), and image-based (e.g., ImageDoubler for Fluidigm platforms), each with strengths and weaknesses depending on cell population heterogeneity.
  • Validation of computational doublet calls with orthogonal evidence, such as sample multiplexing barcodes, imaging data, or genetic variation (e.g., SNP analysis in plant studies), is critical to mitigate misclassification of genuine biological states and avoid overly aggressive doublet removal.

Single-cell RNA sequencing and single-cell ATAC sequencing rely on the assumption that each barcode corresponds to one cell. Doublets, where two or more cells share a barcode, violate this assumption and can produce hybrid transcriptomes that mimic transitional states or rare cell types. This article explains how doublets form during library preparation, how loading density and multiplexing strategies affect doublet rates, and how computational tools such as scDblFinder, DoubletFinder, Scrublet, Solo, DoubletDecon, and ImageDoubler detect them. Practical guidance covers experimental design decisions, quality control thresholds, interpretation of doublet scores, and limitations of current detection methods.

What Are Doublets and Why Do They Matter

A doublet forms when two or more cells are captured under a single barcode during library preparation. In droplet-based platforms, this occurs when a droplet contains multiple cells. In well-based platforms, it occurs when multiple cells settle into the same well. The resulting sequencing data represents a blended transcriptome that does not correspond to any real biological cell state.

Doublets are prevalent in single-cell sequencing data and can lead to artifactual findings [<a href="#ref-1">1</a>]. The most problematic doublets are heterotypic doublets, where the captured cells come from different cell types. These can appear as hybrid expression profiles that resemble transitional states, progenitor cells, or novel cell populations. Homotypic doublets, where the captured cells are the same type, are harder to detect because their expression profile resembles the parent cell type with altered gene expression levels.

The consequences of undetected doublets extend beyond individual misclassified cells. Doublets can create spurious clusters, inflate estimates of rare cell populations, distort differential expression results, and confound trajectory inference. In experiments designed to identify new cell states or developmental intermediates, doublets are a particular concern because their hybrid profiles can be mistaken for genuine biological transitions.

At a Glance: Doublet Control Across the Workflow

Workflow StagePrimary StrategyKey ConsiderationsCommon Tools or Approaches
Library PreparationControl cell loading densityLower loading reduces doublet rate but increases cost per cellManufacturer guidelines, cell counting before loading
Library PreparationSample multiplexingBarcode samples before pooling to identify cross-sample doubletsMULTI-seq lipid-tagged indices, cell hashing
Computational FilteringSimulation-based detectionCreate artificial doublets from real cells and classify by proximityScrublet, DoubletFinder, Solo
Computational FilteringCluster-based detectionIdentify clusters with implausible cross-lineage co-expressionscDblFinder, DoubletDecon, scUmaper
Computational FilteringImage-based detectionUse capture images to identify wells with multiple cellsImageDoubler for Fluidigm platforms
Quality ControlDoublet score thresholdingBalance sensitivity against loss of real cellsPer-tool score distributions, expected doublet rate

How Doublets Form During Library Preparation

Droplet-Based Platforms

Droplet-based platforms such as 10x Genomics Chromium encapsulate cells in nanoliter-scale droplets with barcoded beads. The number of cells per droplet follows a Poisson distribution, meaning that even at optimal loading densities, some droplets will contain multiple cells. The doublet rate is directly related to the number of cells loaded. Higher loading densities increase throughput but also increase the probability that a droplet contains two or more cells.

The relationship between loading density and doublet rate is well established in the field. Researchers must balance the desire for more cells against the cost of increased doublet contamination. Most droplet platforms have recommended loading densities that keep doublet rates in the range of a few percent, but these recommendations assume accurate cell counting and viability assessment before loading.

Well-Based Platforms

Well-based platforms such as Fluidigm C1 capture single cells in individual wells. Doublets occur when multiple cells settle into the same well. The doublet rate depends on the cell suspension concentration and the settling dynamics. Image-based platforms can capture images of each well before sequencing, providing an opportunity to identify wells containing multiple cells directly from the images.

ImageDoubler leverages Fluidigm single-cell sequencing image data to identify doublets and missing samples [<a href="#ref-2">2</a>]. This approach achieved doublet detection rates up to 93.87% and showed a minimum improvement of 33.1% in F1 scores compared to genomic-based methods [<a href="#ref-2">2</a>]. The image-based approach is particularly valuable for homogeneous cell populations, where simulation-based methods struggle because artificial doublets closely resemble real cells.

Single-Nucleus Sequencing

Single-nucleus RNA sequencing follows similar principles but uses isolated nuclei instead of whole cells. The same doublet risks apply, and the same computational detection tools can be used. Sample multiplexing strategies such as MULTI-seq can barcode nuclei from any species with an accessible plasma membrane [<a href="#ref-3">3</a>]. The method involves minimal sample processing, preserving cell viability and endogenous gene expression patterns [<a href="#ref-3">3</a>].

Experimental Strategies to Reduce Doublet Formation

Cell Loading Density

The most direct way to reduce doublet formation is to load fewer cells. This reduces the probability that two cells occupy the same droplet or well. However, lower loading densities increase the cost per cell and may not be feasible for experiments requiring large cell numbers.

Researchers should determine the expected doublet rate for their platform and loading density before running the experiment. This expected rate serves as a benchmark for evaluating computational doublet detection results. If the computational tools identify far more doublets than expected, this may indicate problems with cell preparation or loading.

Sample Multiplexing

Sample multiplexing involves labeling cells from different samples with unique barcodes before pooling. This approach serves two purposes: it reduces costs by allowing multiple samples in a single run, and it enables identification of cross-sample doublets. When two cells from different samples share a barcode, the barcode reads for both samples appear in the same cell, making the doublet unambiguous.

MULTI-seq uses lipid-tagged indices to barcode cells or nuclei from any species with an accessible plasma membrane [<a href="#ref-3">3</a>]. When cells are classified into sample groups using MULTI-seq barcode abundances, data quality is improved through doublet identification and recovery of cells with low RNA content that would otherwise be discarded by standard quality-control workflows [<a href="#ref-3">3</a>]. The method has been used to track T-cell activation dynamics, perform a 96-plex perturbation experiment, and multiplex cryopreserved tumors and metastatic sites [<a href="#ref-3">3</a>].

Cell hashing, a related approach, uses antibody-tagged barcodes to label cells from different samples. Both approaches enable the identification of cross-sample doublets with high confidence. However, they do not identify homotypic doublets within the same sample.

Protoplast Enrichment in Plant Studies

Plant single-cell studies face additional challenges because cell isolation requires protoplasting, which can introduce selective enrichment and sampling biases [<a href="#ref-4">4</a>]. A systematic comparison of protoplast enrichment technologies found that image-based flow cytometry offered increased precision due to customizable gating strategies, while magnetic sorting provided faster processing and enhanced representation of cell size heterogeneity [<a href="#ref-4">4</a>].

The same study compared 10x Genomics Chromium and BD Rhapsody platforms using Arabidopsis roots. Both platforms captured root cell heterogeneity and yielded reproducible gene expression profiles, but showed platform-associated differences in cell type composition [<a href="#ref-4">4</a>]. Notably, single-nucleotide polymorphism analysis of a mixed ecotype sample revealed that among cells identified as doublets by computational algorithms, two-thirds were likely to have been misclassified [<a href="#ref-4">4</a>]. This finding underscores the importance of validating computational doublet calls with orthogonal evidence when possible.

Computational Doublet Detection Methods

Simulation-Based Approaches

Simulation-based methods create artificial doublets by averaging the transcriptional profiles of randomly chosen cell pairs. They then classify real cells based on their proximity to these artificial doublets in gene expression space.

Scrublet simulates multiplets from the data and builds a nearest neighbor classifier [<a href="#ref-5">5</a>]. It avoids the need for expert knowledge or cell clustering [<a href="#ref-5">5</a>]. Scrublet was tested on several datasets with independent knowledge of cell multiplets and demonstrated utility across diverse experimental contexts [<a href="#ref-5">5</a>].

DoubletFinder identifies doublets using only gene expression data by predicting doublets according to each real cell's proximity to artificial doublets created by averaging the transcriptional profile of randomly chosen cell pairs [<a href="#ref-6">6</a>]. The method includes a procedure for estimating input parameters, allowing application across datasets with diverse distributions of cell types [<a href="#ref-6">6</a>]. DoubletFinder is insensitive to experimentally validated cell types with hybrid expression features, meaning it does not falsely classify genuine transitional cells as doublets [<a href="#ref-6">6</a>].

Solo uses a semi-supervised deep learning approach [<a href="#ref-7">7</a>]. It embeds cells unsupervised using a variational autoencoder and then appends a feed-forward neural network layer to the encoder to form a supervised classifier [<a href="#ref-7">7</a>]. The classifier is trained to distinguish simulated doublets from observed data [<a href="#ref-7">7</a>]. Solo can be applied in combination with experimental doublet detection methods to further purify data [<a href="#ref-7">7</a>].

Cluster-Based Approaches

Cluster-based methods identify doublets by examining the expression profiles of cells within clusters. Doublets between distinct cell types appear as hybrid profiles that do not match any single cell state.

scDblFinder is a fast, flexible, and accurate Bioconductor-based doublet detection method [<a href="#ref-1">1</a>]. It builds on the strengths of existing approaches and demonstrates performance on both single-cell RNA and ATAC sequencing data [<a href="#ref-1">1</a>]. Even in complex datasets, scDblFinder can accurately identify most heterotypic doublets, and an independent benchmark found it to outcompete alternatives [<a href="#ref-1">1</a>].

DoubletDecon detects doublets with a combination of deconvolution analyses and the identification of unique cell-state gene expression [<a href="#ref-8">8</a>]. It prevents the prediction of valid mixed-lineage and transitional cell states as doublets by considering their unique gene expression [<a href="#ref-8">8</a>]. DoubletDecon has a graphical user interface and is compatible with diverse species and unsupervised population detection algorithms [<a href="#ref-8">8</a>].

scUmaper integrates quality control, biologically grounded doublet filtering, and marker-library-based cell-type annotation [<a href="#ref-9">9</a>]. It codifies lineage-marker incompatibility rules and applies global clustering followed by within-lineage re-clustering to reveal anomalous subclusters with implausible cross-lineage co-expression [<a href="#ref-9">9</a>]. Across six public human organ datasets, scUmaper removed additional high-confidence heterotypic doublets that were retained by simulation-based approaches [<a href="#ref-9">9</a>].

Model-Driven Approaches

scMODD takes a different approach by using statistical models of single-cell count data to detect doublets [<a href="#ref-10">10</a>]. It tests both the negative binomial model and the zero-inflated negative binomial model as underlying statistical models [<a href="#ref-10">10</a>]. Incorporating zero inflation did not improve detection performance, suggesting that consideration of zero inflation is not necessary in the context of doublet detection [<a href="#ref-10">10</a>]. scMODD achieved similar performance compared to existing data-driven algorithms [<a href="#ref-10">10</a>].

Image-Based Approaches

ImageDoubler uses imaging data from Fluidigm single-cell sequencing to identify doublets and missing samples [<a href="#ref-2">2</a>]. This approach addresses a key limitation of genomic-based methods: their reduced effectiveness in homogeneous cell populations [<a href="#ref-2">2</a>]. The image-based model achieved notable doublet detection efficacy, with rates up to 93.87% and a minimum improvement of 33.1% in F1 scores compared to existing genomic-based methods [<a href="#ref-2">2</a>].

Doublet Detection in Single-Cell ATAC Sequencing

Doublet detection is also relevant for single-cell ATAC sequencing, which profiles chromatin accessibility. ArchR provides doublet removal as part of its comprehensive analysis suite for single-cell chromatin accessibility [<a href="#ref-11">11</a>]. ArchR enables analysis of over 1.2 million single cells within 8 hours on a standard Unix laptop [<a href="#ref-11">11</a>]. scDblFinder also demonstrates performance on single-cell accessibility sequencing data [<a href="#ref-1">1</a>].

A streamlined pipeline for single-cell chromatin accessibility analysis begins with data preprocessing using scATAC-pro or Cell Ranger ATAC, followed by peak calling with MACS2 and differential accessibility analysis [<a href="#ref-12">12</a>]. Transcription factor activity is then inferred using chromVAR, and SCENIC+ is applied to reconstruct transcriptional regulatory networks [<a href="#ref-12">12</a>]. Doublet removal should occur early in this pipeline, before clustering and cell type identification.

Practical Workflow for Doublet Mitigation

Step 1: Plan the Experiment with Doublet Rates in Mind

Before preparing libraries, determine the expected doublet rate for your platform and loading density. Use this expected rate to set expectations for downstream analysis. If you plan to use sample multiplexing, factor in the additional barcode reads and their effect on sequencing depth per cell.

Step 2: Optimize Cell Preparation

Accurate cell counting and viability assessment are critical. Dead cells and debris can clog microfluidic channels, leading to uneven cell capture and increased doublet rates. For plant samples, choose a protoplast enrichment method that balances precision, speed, and representation of cell size heterogeneity [<a href="#ref-4">4</a>].

Step 3: Load Cells at Appropriate Density

Follow manufacturer recommendations for loading density, but consider whether your experiment requires lower doublet rates. Experiments focused on rare cell types or developmental trajectories may warrant more conservative loading to minimize doublet contamination.

Step 4: Run Computational Doublet Detection

Apply at least one computational doublet detection method after initial quality control. For most datasets, scDblFinder provides a good balance of accuracy and speed [<a href="#ref-1">1</a>]. Consider running a second method to cross-validate results, particularly if you have heterogeneous cell populations.

Step 5: Validate Doublet Calls

If you used sample multiplexing, check whether cells identified as doublets show barcode reads from multiple samples. If you have imaging data, compare computational doublet calls to image-based assessments. For plant studies, consider using mixed ecotype samples to validate doublet calls through single-nucleotide polymorphism analysis [<a href="#ref-4">4</a>].

Step 6: Document and Report

Record the expected doublet rate, the number of cells identified as doublets, and the method used for detection. Report these metrics in publications to enable comparison across studies.

Records and Measurements

Expected Doublet Rate

The expected doublet rate depends on the platform and loading density. Record the loading density, the number of cells loaded, and the expected doublet rate for each experiment. This information is essential for interpreting computational doublet detection results.

Doublet Score Distributions

Most computational tools produce a doublet score for each cell. Examine the distribution of these scores to identify an appropriate threshold. A bimodal distribution, where most cells have low scores and a distinct population has high scores, suggests clear separation between singlets and doublets. A unimodal distribution indicates that doublet detection will be less certain.

Cells Removed

Record the number and percentage of cells removed as doublets. Compare this to the expected doublet rate. If the percentage removed is much higher than expected, investigate potential problems with cell preparation or loading. If it is much lower, the detection threshold may be too lenient.

Validation Metrics

If you have orthogonal evidence for doublet identity, such as sample multiplexing barcodes or imaging data, record the concordance between computational doublet calls and the orthogonal evidence. This provides a measure of detection accuracy for your specific dataset.

Common Failure Patterns

Overly Aggressive Doublet Removal

Setting doublet score thresholds too low removes many real cells, particularly those with high RNA content or complex expression profiles. This can eliminate rare cell populations and distort the remaining data. The plant benchmarking study found that two-thirds of cells identified as doublets by computational algorithms were likely misclassified in a mixed ecotype sample [<a href="#ref-4">4</a>]. This highlights the risk of over-reliance on computational doublet calls without orthogonal validation.

Under-Detection in Homogeneous Populations

Simulation-based methods struggle in homogeneous cell populations because artificial doublets closely resemble real cells. If your sample consists of a single cell type or closely related cell types, consider using image-based approaches if available [<a href="#ref-2">2</a>] or cluster-based methods that examine cross-lineage co-expression [<a href="#ref-9">9</a>].

Ignoring Expected Doublet Rates

Computational tools may identify more or fewer doublets than expected based on loading density. Large discrepancies warrant investigation. More doublets than expected may indicate problems with cell aggregation or clumping. Fewer doublets than expected may indicate that the detection threshold is too lenient or that the cell population is homogeneous.

Applying scRNA-Seq Tools to Other Modalities

Some doublet detection tools are designed specifically for scRNA-seq data. When working with single-cell ATAC data, use tools validated for that modality, such as ArchR [<a href="#ref-11">11</a>] or scDblFinder [<a href="#ref-1">1</a>]. The characteristics of doublets differ between modalities, and tools may not transfer directly.

Failure to Validate with Orthogonal Evidence

Computational doublet detection is imperfect. Whenever possible, validate doublet calls with orthogonal evidence such as sample multiplexing barcodes, imaging data, or genetic variation. The mixed ecotype analysis in plant studies provides a powerful validation approach [<a href="#ref-4">4</a>].

Limitations of Computational Doublet Detection

Homotypic Doublets

Computational methods are most effective at detecting heterotypic doublets, where the captured cells come from different cell types. Homotypic doublets, where the captured cells are the same type, produce expression profiles that resemble the parent cell type and are difficult to distinguish from real cells.

Misclassification of Genuine Cell States

Some computational methods may classify genuine transitional or mixed-lineage cells as doublets. DoubletDecon was specifically designed to prevent this by considering unique gene expression of valid cell states [<a href="#ref-8">8</a>]. DoubletFinder was shown to be insensitive to an experimentally validated kidney cell type with hybrid expression features [<a href="#ref-6">6</a>]. However, no method is perfect, and researchers should be cautious about removing cells that may represent genuine biological states.

Platform-Specific Performance

Doublet detection methods perform differently across platforms and cell types. The plant benchmarking study found platform-associated differences in cell type composition between 10x Genomics Chromium and BD Rhapsody [<a href="#ref-4">4</a>]. Methods should be validated on data similar to your own before being applied at scale.

Computational Cost

Some doublet detection methods are computationally intensive. ArchR was designed for scalability, enabling analysis of over 1.2 million single cells within 8 hours on a standard Unix laptop [<a href="#ref-11">11</a>]. For very large datasets, consider the computational requirements of different methods before choosing one.

Quality Control and Reproducibility

Reproducible Workflows

Reproducible analysis requires documented workflows. The Galaxy Training Network provides accessible workflow training and analysis tutorials [<a href="#ref-13">13</a>]. nf-core documentation describes community pipeline standards, usage, configuration, and reproducible workflow context [<a href="#ref-14">14</a>]. Bioconductor provides official package, workflow, installation, and reproducible genomic-analysis documentation [<a href="#ref-15">15</a>].

Foundational Computing Skills

Bioinformatics analysis requires foundational computing skills. The Carpentries offers lessons on foundational computing, data, shell, Git, and programming [<a href="#ref-16">16</a>]. EMBL-EBI Training provides bioinformatics learning pathways, data-resource training, and practical analysis education [<a href="#ref-17">17</a>]. NCBI provides official descriptions of databases, search systems, sequence resources, and analysis services [<a href="#ref-18">18</a>].

Documentation Standards

Document the version of each tool used, the parameters applied, and the threshold chosen for doublet classification. This enables others to reproduce your analysis and compare results across studies. Version control for analysis scripts is essential for reproducibility.

Welfare and Safety Context

Doublet identification is a data quality issue instead of a welfare or safety concern. However, the consequences of poor doublet management can affect the validity of research findings. In agricultural and veterinary contexts, single-cell studies inform understanding of mammary gland biology [<a href="#ref-19">19</a>], immune cell populations, and disease mechanisms. Misleading results from doublet contamination could lead to incorrect conclusions about cell types and states, potentially affecting downstream research directions.

For studies involving animal samples, ensure that cell isolation procedures follow institutional animal care guidelines. The quality of single-cell data depends on the quality of the input cells, which in turn depends on proper sample collection and processing.

Professional Escalation Criteria

When to Seek Expert Assistance

Consider consulting a bioinformatics specialist or core facility when:

  • Your dataset shows a doublet rate far exceeding the expected rate for your platform and loading density
  • Computational doublet detection tools produce conflicting results across methods
  • You are working with a novel cell type or tissue where doublet detection methods have not been validated
  • You plan to use doublet detection results to make conclusions about rare cell populations or developmental trajectories
  • You are integrating data across multiple batches or platforms and need consistent doublet filtering

When to Reconsider Experimental Design

If doublet rates remain high despite following recommended loading densities, investigate:

  • Cell aggregation or clumping in the suspension
  • Inaccurate cell counting
  • Debris or dead cells in the preparation
  • Platform-specific issues with your sample type

A Decision Framework for Choosing Doublet Detection Methods by Data Modality and Population Structure

Selecting the right doublet detection approach requires matching the method to the specific characteristics of your dataset. The existing guidance covers individual tools, but researchers often struggle with the practical decision of which method to apply first, when to combine methods, and how to interpret conflicting results. This section provides a structured decision framework based on data modality, cell population heterogeneity, and available orthogonal evidence.

Step 1: Classify Your Data Modality

The first decision point is whether your data comes from single-cell RNA sequencing, single-cell ATAC sequencing, or a multiomic experiment. This classification determines which detection methods are appropriate.

For scRNA-seq data, most established methods apply directly. Scrublet, DoubletFinder, Solo, DoubletDecon, and scDblFinder were all developed and validated primarily on transcriptomic data [<a href="#ref-1">1</a>, <a href="#ref-5">5</a>, <a href="#ref-6">6</a>, <a href="#ref-7">7</a>, <a href="#ref-8">8</a>]. scDblFinder demonstrates performance on both single-cell RNA and accessibility sequencing data [<a href="#ref-1">1</a>]. For scATAC-seq data, ArchR provides doublet removal as part of its comprehensive analysis suite [<a href="#ref-11">11</a>]. A streamlined pipeline for single-cell chromatin accessibility analysis begins with data preprocessing using scATAC-pro or Cell Ranger ATAC, followed by peak calling with MACS2 [<a href="#ref-12">12</a>]. Doublet removal should occur early in this pipeline, before clustering and cell type identification [<a href="#ref-12">12</a>].

For multiomic experiments that profile both RNA expression and chromatin accessibility from the same cells, apply doublet detection separately to each modality and compare results. Cells identified as doublets by both modalities have higher confidence calls. Cells identified by only one modality require closer examination, as modality-specific artifacts can produce false positives.

Step 2: Assess Cell Population Heterogeneity

The composition of your cell population strongly influences which detection method will perform best. Simulation-based methods create artificial doublets by averaging the transcriptional profiles of randomly chosen cell pairs [<a href="#ref-5">5</a>, <a href="#ref-6">6</a>]. These methods work well when the cell population contains transcriptionally distinct cell types, because artificial doublets from different cell types are clearly separable from real cells.

In homogeneous cell populations, simulation-based methods struggle because artificial doublets closely resemble real cells [<a href="#ref-2">2</a>]. ImageDoubler was developed specifically to address this limitation, using Fluidigm single-cell sequencing image data to identify doublets and missing samples [<a href="#ref-2">2</a>]. The image-based approach achieved doublet detection rates up to 93.87% and showed a minimum improvement of 33.1% in F1 scores compared to genomic-based methods [<a href="#ref-2">2</a>].

Cluster-based methods such as scDblFinder and DoubletDecon examine expression profiles within clusters to identify cells with implausible cross-lineage co-expression [<a href="#ref-1">1</a>, <a href="#ref-8">8</a>]. scUmaper codifies lineage-marker incompatibility rules and applies global clustering followed by within-lineage re-clustering to reveal anomalous subclusters [<a href="#ref-9">9</a>]. These methods are particularly useful when you have prior knowledge of expected cell types and their marker genes.

For datasets with unknown or complex cell populations, scDblFinder provides a good default choice because it is fast, flexible, and accurate [<a href="#ref-1">1</a>]. An independent benchmark found scDblFinder to outcompete alternatives [<a href="#ref-1">1</a>].

Step 3: Determine Available Orthogonal Evidence

Orthogonal evidence provides independent confirmation of doublet identity. The availability of such evidence should influence both method selection and interpretation of results.

Sample multiplexing with lipid-tagged indices or antibody-tagged barcodes enables unambiguous identification of cross-sample doublets [<a href="#ref-3">3</a>]. When two cells from different samples share a barcode, the barcode reads for both samples appear in the same cell. MULTI-seq uses lipid-tagged indices to barcode cells or nuclei from any species with an accessible plasma membrane [<a href="#ref-3">3</a>]. When cells are classified into sample groups using MULTI-seq barcode abundances, data quality is improved through doublet identification and recovery of cells with low RNA content that would otherwise be discarded by standard quality-control workflows [<a href="#ref-3">3</a>].

Imaging data provides another source of orthogonal evidence. ImageDoubler leverages Fluidigm single-cell sequencing image data to identify doublets and missing samples [<a href="#ref-2">2</a>]. If your platform captures images of wells before sequencing, you can compare computational doublet calls to image-based assessments.

Genetic variation offers a powerful validation approach. In plant studies, single-nucleotide polymorphism analysis of a mixed ecotype sample revealed that among cells identified as doublets by computational algorithms, two-thirds were likely to have been misclassified [<a href="#ref-4">4</a>]. This finding demonstrates both the value of genetic validation and the risk of over-reliance on computational doublet calls without orthogonal evidence [<a href="#ref-4">4</a>].

Step 4: Select Primary and Secondary Detection Methods

Based on the first three steps, select a primary detection method and a secondary method for cross-validation. The table below summarizes recommended combinations.

Data ModalityPopulation StructurePrimary MethodSecondary MethodRationale
scRNA-seqHeterogeneousscDblFinderDoubletFinder or ScrubletscDblFinder is fast and accurate [<a href="#ref-1">1</a>], simulation methods provide independent confirmation [<a href="#ref-5">5</a>, <a href="#ref-6">6</a>]
scRNA-seqHomogeneousImageDoubler if imaging availablescDblFinderImage-based methods excel in homogeneous populations [<a href="#ref-2">2</a>]
scRNA-seqUnknownscDblFinderDoubletDeconscDblFinder handles complex datasets [<a href="#ref-1">1</a>], DoubletDecon prevents misclassification of transitional states [<a href="#ref-8">8</a>]
scATAC-seqAnyArchRscDblFinderArchR includes doublet removal for chromatin accessibility [<a href="#ref-11">11</a>], scDblFinder works on ATAC data [<a href="#ref-1">1</a>]
scRNA-seq with multiplexingAnyBarcode-based classificationscDblFinderMultiplexing provides unambiguous cross-sample doublet identification [<a href="#ref-3">3</a>]
MultiomicAnyApply modality-specific methods separatelyCompare results across modalitiesConflicting calls require investigation

Step 5: Establish Thresholds Using Expected Doublet Rates

The expected doublet rate based on loading density serves as a benchmark for threshold selection. Most computational tools produce a doublet score for each cell. Examine the distribution of these scores to identify an appropriate threshold.

A bimodal distribution, where most cells have low scores and a distinct population has high scores, suggests clear separation between singlets and doublets. In this case, set the threshold at the valley between the two modes. A unimodal distribution indicates that doublet detection will be less certain, and you should rely more heavily on the expected doublet rate to guide threshold selection.

Compare the number of cells classified as doublets to the expected doublet rate for your platform and loading density. If the percentage removed is much higher than expected, investigate potential problems with cell preparation or loading. If it is much lower, the detection threshold may be too lenient.

Step 6: Resolve Conflicting Results Across Methods

When primary and secondary methods disagree on specific cells, use a structured approach to resolve conflicts.

Cells identified as doublets by both methods have high confidence calls and should be removed. Cells identified by only one method require closer examination. Examine the expression profiles of these cells to determine whether they show evidence of cross-lineage co-expression. If you have orthogonal evidence such as multiplexing barcodes or imaging data, use it to adjudicate.

The plant benchmarking study provides a cautionary example. Among cells identified as doublets by computational algorithms, two-thirds were likely to have been misclassified based on single-nucleotide polymorphism analysis [<a href="#ref-4">4</a>]. This finding suggests that computational doublet calls can have high false positive rates, particularly in certain experimental contexts [<a href="#ref-4">4</a>].

For cells with conflicting calls, consider whether they represent genuine transitional or mixed-lineage states. DoubletDecon was specifically designed to prevent the prediction of valid mixed-lineage and transitional cell states as doublets by considering their unique gene expression [<a href="#ref-8">8</a>]. DoubletFinder was shown to be insensitive to an experimentally validated kidney cell type with hybrid expression features [<a href="#ref-6">6</a>]. If your experiment focuses on developmental trajectories or cell state transitions, be particularly cautious about removing cells with conflicting doublet calls.

Step 7: Document Decisions and Rationale

Record the following information for each dataset:

  • Data modality and platform
  • Cell population composition and estimated heterogeneity
  • Expected doublet rate based on loading density
  • Primary and secondary detection methods with versions and parameters
  • Doublet score threshold and rationale for threshold selection
  • Number and percentage of cells removed as doublets
  • Concordance between methods and with orthogonal evidence
  • Resolution of conflicting calls

This documentation enables others to reproduce your analysis and compare results across studies. Version control for analysis scripts is essential for reproducibility.

A Record System for Doublet Management Decisions

A structured record system helps track doublet management across experiments and enables continuous improvement of protocols. The following record fields capture the essential information for each experiment.

Pre-Experiment Records

Record the expected doublet rate based on platform specifications and loading density. Record the cell counting method, viability assessment, and any observations about cell aggregation or clumping. For plant samples, record the protoplast enrichment method used and any observations about cell size heterogeneity [<a href="#ref-4">4</a>].

During-Experiment Records

Record the actual number of cells loaded, the number of droplets or wells captured, and any platform alerts or anomalies. If using sample multiplexing, record the barcode design and the number of samples pooled [<a href="#ref-3">3</a>].

Post-Sequencing Records

Record the number of cells passing initial quality control, the doublet score distribution for each detection method, and the threshold applied. Record the number and percentage of cells removed as doublets. If orthogonal evidence is available, record the concordance between computational doublet calls and the orthogonal evidence.

Cross-Experiment Comparison

Maintain a summary table across experiments showing expected doublet rate, detected doublet rate, and validation concordance. This table helps identify systematic issues with specific sample types, platforms, or protocols. If detected doublet rates consistently exceed expected rates for a particular sample type, investigate cell preparation procedures.

Troubleshooting Common Decision Framework Failures

Failure Mode 1: Method Selection Based on Habit instead of Data Characteristics

Researchers often default to the method they used previously or the method most cited in their field. This approach can lead to poor performance when data characteristics differ from the method's validation context. The plant benchmarking study found platform-associated differences in cell type composition between 10x Genomics Chromium and BD Rhapsody [<a href="#ref-4">4</a>]. Methods validated on one platform may not perform identically on another.

Troubleshooting approach: Classify your data modality and population structure before selecting methods. If your population is homogeneous, prioritize image-based approaches if available [<a href="#ref-2">2</a>] or cluster-based methods that examine cross-lineage co-expression [<a href="#ref-9">9</a>].

Failure Mode 2: Ignoring Expected Doublet Rates

Computational tools may identify more or fewer doublets than expected based on loading density. Large discrepancies warrant investigation. More doublets than expected may indicate problems with cell aggregation or clumping. Fewer doublets than expected may indicate that the detection threshold is too lenient or that the cell population is homogeneous.

Troubleshooting approach: Always record the expected doublet rate before running detection. Compare detected rates to expected rates. Investigate discrepancies before proceeding with downstream analysis.

Failure Mode 3: Over-Reliance on a Single Method

Relying on a single detection method without cross-validation increases the risk of systematic errors. The plant benchmarking study found that two-thirds of cells identified as doublets by computational algorithms were likely misclassified [<a href="#ref-4">4</a>]. This finding underscores the importance of validating computational doublet calls with orthogonal evidence when possible [<a href="#ref-4">4</a>].

Troubleshooting approach: Run at least two detection methods and compare results. Use orthogonal evidence such as multiplexing barcodes, imaging data, or genetic variation to adjudicate conflicting calls.

Failure Mode 4: Applying scRNA-Seq Tools to Other Modalities Without Validation

Some doublet detection tools are designed specifically for scRNA-seq data. When working with single-cell ATAC data, use tools validated for that modality, such as ArchR [<a href="#ref-11">11</a>] or scDblFinder [<a href="#ref-1">1</a>]. The characteristics of doublets differ between modalities, and tools may not transfer directly.

Troubleshooting approach: Check the validation context for each tool before applying it to your data. For scATAC-seq, use ArchR or scDblFinder [<a href="#ref-1">1</a>, <a href="#ref-11">11</a>]. For other modalities, seek tools validated on similar data.

Failure Mode 5: Removing Cells Without Examining Their Identity

Automated doublet removal without examining the identity of removed cells can eliminate genuine biological states. DoubletDecon was specifically designed to prevent the prediction of valid mixed-lineage and transitional cell states as doublets [<a href="#ref-8">8</a>]. DoubletFinder was shown to be insensitive to an experimentally validated kidney cell type with hybrid expression features [<a href="#ref-6">6</a>].

Troubleshooting approach: Before removing cells identified as doublets, examine their expression profiles and cluster positions. If removed cells form a distinct cluster with coherent biology, investigate whether they represent a genuine cell state instead of doublets.

Integration with Reproducible Workflow Standards

The decision framework should be implemented within reproducible workflow standards. Bioconductor provides official package, workflow, installation, and reproducible genomic-analysis documentation [<a href="#ref-15">15</a>]. The Galaxy Training Network provides accessible workflow training and analysis tutorials [<a href="#ref-13">13</a>]. nf-core documentation describes community pipeline standards, usage, configuration, and reproducible workflow context [<a href="#ref-14">14</a>].

The Carpentries offers lessons on foundational computing, data, shell, Git, and programming [<a href="#ref-16">16</a>]. EMBL-EBI Training provides bioinformatics learning pathways, data-resource training, and practical analysis education [<a href="#ref-17">17</a>]. NCBI provides official descriptions of databases, search systems, sequence resources, and analysis services [<a href="#ref-18">18</a>].

Implement the decision framework as a documented workflow with version-controlled scripts and parameters. Record the version of each tool used, the parameters applied, and the threshold chosen for doublet classification. This enables others to reproduce your analysis and compare results across studies.

Professional Escalation Criteria for the Decision Framework

Consult a bioinformatics specialist or core facility when:

  • Your dataset shows a doublet rate far exceeding the expected rate for your platform and loading density, and you cannot identify the cause through cell preparation troubleshooting
  • Computational doublet detection tools produce conflicting results across methods, and orthogonal evidence is unavailable to adjudicate
  • You are working with a novel cell type or tissue where doublet detection methods have not been validated
  • You plan to use doublet detection results to make conclusions about rare cell populations or developmental trajectories, where misclassification risks are highest
  • You are integrating data across multiple batches or platforms and need consistent doublet filtering decisions
  • Your cell population is homogeneous, and you lack imaging data or other orthogonal evidence to validate computational doublet calls

Comparison of Decision Framework Outcomes Across Scenarios

The following scenarios illustrate how the decision framework applies in practice.

Scenario 1: Heterogeneous Immune Cell Population from Peripheral Blood

A researcher profiles peripheral blood mononuclear cells using a droplet-based platform. The cell population is highly heterogeneous, containing T cells, B cells, natural killer cells, monocytes, and dendritic cells. The expected doublet rate at the chosen loading density is 4%.

The decision framework classifies this as scRNA-seq data with a heterogeneous population. The primary method is scDblFinder, with DoubletFinder as the secondary method. The researcher examines the doublet score distribution and identifies a bimodal pattern, setting the threshold at the valley between modes. The detected doublet rate is 4.5%, close to the expected rate. Cross-validation between methods shows high concordance. No orthogonal evidence is available, but the high concordance between methods and the alignment with expected rates provide confidence in the doublet calls.

Scenario 2: Homogeneous Cell Line Experiment

A researcher profiles a cultured cell line that is largely homogeneous. The expected doublet rate is 3%. Simulation-based methods produce ambiguous results because artificial doublets closely resemble real cells.

The decision framework classifies this as scRNA-seq data with a homogeneous population. If the platform captures images, ImageDoubler is the primary method [<a href="#ref-2">2</a>]. If imaging is unavailable, scDblFinder is used with careful threshold selection based on the expected doublet rate. The researcher examines the expression profiles of cells identified as doublets to check for evidence of cross-lineage co-expression. Given the homogeneous population, the researcher is cautious about removing cells and validates calls with the expected doublet rate.

Scenario 3: Plant Root Tissue with Mixed Ecotype Validation

A researcher profiles Arabidopsis roots using a droplet-based platform. The researcher includes a mixed ecotype sample to enable genetic validation of doublet calls. The expected doublet rate is 5%.

The decision framework classifies this as scRNA-seq data with a heterogeneous population. The primary method is scDblFinder. The researcher uses single-nucleotide polymorphism analysis of the mixed ecotype sample to validate doublet calls [<a href="#ref-4">4</a>]. This validation reveals that two-thirds of cells identified as doublets by computational algorithms were likely misclassified [<a href="#ref-4">4</a>]. The researcher adjusts the doublet score threshold to reduce false positives, accepting a higher false negative rate to preserve genuine cells.

Scenario 4: Single-Cell ATAC Sequencing of Brain Tissue

A researcher profiles chromatin accessibility in brain tissue using a droplet-based platform. The expected doublet rate is 6%.

The decision framework classifies this as scATAC-seq data. The primary method is ArchR, which includes doublet removal as part of its comprehensive analysis suite [<a href="#ref-11">11</a>]. scDblFinder serves as the secondary method because it demonstrates performance on accessibility sequencing data [<a href="#ref-1">1</a>]. Doublet removal occurs early in the pipeline, before clustering and cell type identification [<a href="#ref-12">12</a>]. The researcher compares doublet calls between methods and investigates any discrepancies.

Limitations of the Decision Framework

The decision framework provides structured guidance but has limitations. It cannot replace domain expertise or knowledge of specific platform characteristics. The framework assumes that expected doublet rates are known and accurate, which may not hold for novel platforms or sample types. The framework does not address all possible data modalities, and researchers working with emerging technologies should seek tools validated on similar data.

The framework also cannot resolve the fundamental limitation of computational doublet detection: homotypic doublets are difficult to distinguish from real cells. No decision framework can overcome this limitation. Researchers should be aware that even with optimal method selection and threshold setting, some doublets will remain undetected, and some real cells will be removed.

The plant benchmarking study provides a sobering reminder of these limitations. Among cells identified as doublets by computational algorithms, two-thirds were likely to have been misclassified [<a href="#ref-4">4</a>]. This finding suggests that computational doublet detection has higher false positive rates than commonly assumed, particularly in certain experimental contexts [<a href="#ref-4">4</a>]. Researchers should interpret doublet detection results with appropriate caution and validate calls with orthogonal evidence whenever possible.

Frequently Asked Questions

What is the difference between heterotypic and homotypic doublets?

Heterotypic doublets form when two cells of different types are captured under one barcode. Their hybrid expression profiles can mimic transitional states or novel cell types, making them the primary concern for data interpretation. Homotypic doublets form when two cells of the same type are captured together. Their expression profiles resemble the parent cell type, making them difficult to detect computationally and less likely to create spurious biological conclusions.

How does cell loading density affect doublet rates?

Higher loading densities increase the probability that two or more cells occupy the same droplet or well, leading to higher doublet rates. Lower loading densities reduce doublet rates but increase the cost per cell. The relationship follows a Poisson distribution for droplet-based platforms. Researchers should determine the expected doublet rate for their platform and loading density before running the experiment.

Can sample multiplexing help with doublet detection?

Sample multiplexing labels cells from different samples with unique barcodes before pooling. This enables unambiguous identification of cross-sample doublets, because a cell with barcode reads from two samples must be a doublet. MULTI-seq uses lipid-tagged indices for this purpose [<a href="#ref-3">3</a>]. However, multiplexing does not identify homotypic doublets within the same sample.

Which computational doublet detection tool should I use?

The choice depends on your data and priorities. scDblFinder is fast, flexible, and accurate for both scRNA-seq and scATAC-seq data [<a href="#ref-1">1</a>]. Scrublet and DoubletFinder are well-established simulation-based methods [<a href="#ref-5">5</a>, <a href="#ref-6">6</a>]. Solo uses deep learning and can be combined with experimental methods [<a href="#ref-7">7</a>]. DoubletDecon prevents misclassification of valid transitional states [<a href="#ref-8">8</a>]. For homogeneous populations, image-based approaches such as ImageDoubler may be more effective [<a href="#ref-2">2</a>]. Consider running two methods and comparing results.

How do I choose a doublet score threshold?

Examine the distribution of doublet scores from your chosen tool. A bimodal distribution suggests clear separation between singlets and doublets. Compare the number of cells classified as doublets to the expected doublet rate for your platform and loading density. If available, use orthogonal evidence such as sample multiplexing barcodes or imaging data to validate threshold choices.

Can computational doublet detection misclassify real cells?

Yes. Computational methods can misclassify genuine transitional or mixed-lineage cells as doublets. DoubletDecon was designed to prevent this by considering unique gene expression of valid cell states [<a href="#ref-8">8</a>]. DoubletFinder was shown to be insensitive to a validated cell type with hybrid expression features [<a href="#ref-6">6</a>]. A plant benchmarking study found that two-thirds of cells identified as doublets by computational algorithms were likely misclassified [<a href="#ref-4">4</a>]. Validate doublet calls with orthogonal evidence when possible.

Do doublet detection methods work for single-cell ATAC sequencing?

Some methods work for both scRNA-seq and scATAC-seq. scDblFinder demonstrates performance on both modalities [<a href="#ref-1">1</a>]. ArchR includes doublet removal as part of its single-cell chromatin accessibility analysis suite [<a href="#ref-11">11</a>]. However, not all scRNA-seq doublet detection tools transfer directly to other modalities. Use tools validated for your specific data type.

How should I report doublet information in publications?

Report the expected doublet rate based on loading density, the number and percentage of cells removed as doublets, the computational method and version used, and the threshold applied. If you validated doublet calls with orthogonal evidence, report the concordance. This information enables readers to assess data quality and compare results across studies.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

[1] [Doublet identification in single-cell sequencing data using scDblFinder.](https://pubmed.ncbi.nlm.nih.gov/35814628). F1000Research, 2021. [2] [ImageDoubler: image-based doublet identification in single-cell sequencing.](https://pubmed.ncbi.nlm.nih.gov/39747095). Nature communications, 2025. [3] [MULTI-seq: sample multiplexing for single-cell RNA sequencing using lipid-tagged indices.](https://pubmed.ncbi.nlm.nih.gov/31209384). Nature methods, 2019. [4] [Benchmarking plant single cell RNA-sequencing sample processing strategies.](https://doi.org/10.1038/s44318-026-00800-5). 2026. [5] [Scrublet: Computational Identification of Cell Doublets in Single-Cell Transcriptomic Data.](https://pubmed.ncbi.nlm.nih.gov/30954476). Cell systems, 2019. [6] [DoubletFinder: Doublet Detection in Single-Cell RNA Sequencing Data Using Artificial Nearest Neighbors.](https://pubmed.ncbi.nlm.nih.gov/30954475). Cell systems, 2019. [7] [Solo: Doublet Identification in Single-Cell RNA-Seq via Semi-Supervised Deep Learning.](https://pubmed.ncbi.nlm.nih.gov/32592658). Cell systems, 2020. [8] [DoubletDecon: Deconvoluting Doublets from Single-Cell RNA-Sequencing Data.](https://pubmed.ncbi.nlm.nih.gov/31693907). Cell reports, 2019. [9] [scUmaper: An automated framework for doublet removal and cell-type annotation in single-cell transcriptomics.](https://doi.org/10.1016/j.isci.2026.115850). 2026. [10] [scMODD: A model-driven algorithm for doublet identification in single-cell RNA-sequencing data](https://doi.org/10.3389/fsysb.2022.1082309). Frontiers in Systems Biology, 2023. [11] [ArchR is a scalable software package for integrative single-cell chromatin accessibility analysis.](https://pubmed.ncbi.nlm.nih.gov/33633365). Nature genetics, 2021. [12] [A pipeline for single-cell chromatin accessibility data analysis.](https://doi.org/10.1097/bs9.0000000000000259). 2026. [13] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project. [14] [nf-core Documentation](https://nf-co.re/docs). nf-core. [15] [Bioconductor](https://bioconductor.org/). Bioconductor Project. [16] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries. [17] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute. [18] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information. [19] [Deep Single-Cell Transcriptomic Profiling of Bovine Milk Somatic Cells Revealed Expression of Stem Cell Related Transcription Factors.](https://doi.org/10.3390/genes17040365). 2026.

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.