# Combinatorial Indexing for Single-Cell Sequencing: Scalable and Cost-Effective Without Specialized Instruments


## Key Takeaways

- Combinatorial indexing enables massive single-cell throughput (hundreds of thousands to millions of cells) using standard laboratory equipment (multichannel pipettes, thermocyclers, centrifuges), circumventing the need for expensive specialized microfluidic instruments characteristic of droplet-based platforms.
- This method achieves scalability through sequential rounds of pooling and splitting, where each round adds a unique barcode, exponentially increasing the total number of possible barcode combinations (e.g., 96x96x96 ≈ 884,736 with three 96-well plate rounds).
- Reagent costs are significantly reduced, often to approximately one cent per cell, making large-scale single-cell studies economically feasible for laboratories with limited budgets.
- The approach is adaptable for multimodal analyses, allowing simultaneous profiling of RNA, chromatin accessibility, and VDJ sequences, expanding its utility beyond transcriptomics.
- While offering cost-effectiveness and scalability, combinatorial indexing may exhibit lower sensitivity for detecting lowly expressed genes compared to droplet-based methods, a factor to consider based on experimental objectives.
- Data analysis involves demultiplexing sequencing reads based on the three-part barcode combination, followed by standard single-cell quality control metrics and computational analysis pipelines.

---

Single-cell RNA sequencing has transformed biological research by enabling transcriptome measurement at cellular resolution, but conventional droplet-based platforms require substantial capital investment in specialized microfluidics instruments. Combinatorial indexing offers an alternative strategy that achieves massive single-cell throughput using only standard laboratory equipment, including multichannel pipettes, PCR thermocyclers, and benchtop centrifuges. This approach assigns unique barcode combinations through sequential rounds of pooling and splitting, allowing researchers to profile hundreds of thousands of nuclei or cells in a single experiment at reagent costs that are dramatically lower than droplet-based alternatives. For biology students, researchers, and laboratory professionals who lack access to dedicated single-cell platforms, understanding combinatorial indexing principles, workflow requirements, quality control measures, and data analysis pathways enables practical implementation of large-scale single-cell studies with existing infrastructure.

## The Instrumentation Barrier in Single-Cell Sequencing

Conventional single-cell RNA sequencing methods rely on microfluidic platforms that encapsulate individual cells in nanoliter-scale droplets together with barcoded beads. These systems require specialized instruments that are expensive to purchase, maintain, and operate. The capital cost of droplet-based platforms places them beyond the reach of many individual laboratories, particularly in academic settings with limited equipment budgets or in institutions where shared core facilities have constrained capacity. Beyond the initial instrument purchase, droplet-based approaches require specialized consumables, including microfluidic chips and barcoded beads, which create ongoing per-sample costs that scale with experimental throughput.

Combinatorial indexing circumvents the instrumentation requirement entirely by replacing physical cell isolation with a logical barcoding strategy. Instead of physically separating individual cells into droplets or wells, combinatorial indexing uses sequential rounds of pooling and splitting to generate unique barcode combinations. Each round of indexing adds a new barcode to the accumulating combination, and the total number of possible barcode combinations grows exponentially with the number of rounds. This approach requires only equipment that is standard in essentially every molecular biology laboratory: multichannel pipettes for distributing cells into 96-well or 384-well plates, a thermocycler for reverse transcription and PCR amplification, and a centrifuge for cell washing and buffer exchange.

The practical consequence of this design is substantial. Laboratories that have never had access to droplet-based single-cell platforms can implement combinatorial indexing protocols using their existing equipment. The [optimized single-nucleus transcriptional profiling protocol](https://pubmed.ncbi.nlm.nih.gov/36261634) demonstrates that reagent costs can be reduced to approximately one cent per cell or less, with total hands-on time from nuclei isolation to final library preparation taking two to three days depending on sample number. This cost profile makes large-scale single-cell experiments feasible for laboratories that would otherwise be limited to bulk RNA sequencing or small-scale single-cell studies.

## Core Principles of Combinatorial Indexing

### Sequential Barcode Assembly

The fundamental principle of combinatorial indexing is the sequential assembly of barcode combinations through multiple rounds of pooling and splitting. In the first round, a population of cells or nuclei is distributed across the wells of a multiwell plate, typically a 96-well or 384-well format. Each well receives a distinct barcode that becomes covalently attached to the nucleic acid content of the cells in that well. After the first barcoding reaction, all cells are pooled together and redistributed across a fresh multiwell plate for the second round of barcoding. This second barcode is different from the first, and the combination of first-round and second-round barcodes creates a unique identifier for each cell.

The process repeats for additional rounds, with each round multiplying the number of possible barcode combinations. With three rounds of indexing using 96-well plates, the theoretical capacity reaches 96 × 96 × 96, or 884,736 possible barcode combinations. The [mouse organogenesis study](https://pubmed.ncbi.nlm.nih.gov/30787437) used single-cell combinatorial indexing to profile approximately two million cells from 61 mouse embryos staged between 9.5 and 13.5 days of gestation in a single experiment, demonstrating that the approach scales to millions of cells when sufficient barcode combinations are available.

### Split-Pool Logic and Cell Identity

The split-pool logic ensures that each cell receives a unique combination of barcodes through a probabilistic process. After the first round of barcoding, cells from different wells carry different first-round barcodes. When these cells are pooled and redistributed for the second round, cells that originated from different first-round wells are mixed together in each second-round well. The second-round barcode is added to all cells in a given well, regardless of their first-round barcode. After pooling and redistributing for the third round, the third-round barcode is added, and the complete three-barcode combination uniquely identifies each cell.

The probability that two cells receive the same barcode combination depends on the number of cells loaded relative to the number of possible combinations. Loading far fewer cells than the theoretical barcode capacity minimizes the collision rate. In practice, researchers typically load cells at a fraction of the theoretical capacity to ensure that most barcode combinations are represented by at most one cell. The [single-cell combinatorial indexing protocol](https://pubmed.ncbi.nlm.nih.gov/36261634) describes three rounds of split-pool indexing that achieve exponential scalability while maintaining robust performance across different tissue types.

### Applicability Across Modalities

Combinatorial indexing extends beyond RNA sequencing to simultaneous measurement of multiple molecular modalities from individual cells. The [UDA-seq workflow](https://pubmed.ncbi.nlm.nih.gov/39833568) integrates a post-indexing step that enhances throughput and systematically adapts existing droplet-based single-cell multimodal methods, enabling co-assay of RNA and VDJ sequences, RNA and chromatin accessibility, and RNA and CRISPR perturbation. This universal workflow generated over 100,000 high-quality single-cell datasets from three dozen frozen clinical biopsy specimens within a single-channel droplet microfluidics experiment, demonstrating that combinatorial indexing principles can be combined with droplet-based approaches for multimodal analysis.

The [easySHARE-seq method](https://doi.org/10.7554/elife.110034) represents a combinatorial indexing approach for simultaneous measurement of gene expression and chromatin accessibility. This method improved upon previous SHARE-seq protocols by enhancing the barcode design and streamlining the workflow, resulting in libraries with usable sequences of up to 300 base pairs. Applied to murine liver nuclei, easySHARE-seq recovered 19,664 nuclei with joint chromatin and expression profiles, recovering over 1.5-fold more transcripts per cell than other combinatorial indexing-based techniques while retaining high scalability and low cost.

## Practical Workflow for Combinatorial Indexing Experiments

### Nuclei Isolation and Fixation

The first step in a combinatorial indexing experiment is preparing a suspension of nuclei from the tissue or cells of interest. For fresh or frozen tissues, nuclei isolation typically involves mechanical disruption using a dounce tissue homogenizer followed by density gradient centrifugation or filtration to remove debris and intact cells. The [mouse kidney nuclear isolation protocol](https://doi.org/10.1016/j.xpro.2022.101904) describes an optimized approach using a dounce tissue homogenizer that enables nuclei extraction with high yield. Fixed nuclei are then processed for combinatorial indexing, with fixation preserving the nuclear structure and preventing RNA degradation during the multiple rounds of barcoding.

Fixation is a critical consideration because it stabilizes the nucleic acid content and allows the extended processing time required for multiple rounds of indexing. The [optimized sci-RNA-seq protocol](https://pubmed.ncbi.nlm.nih.gov/36261634) notes that improvements in the protocol allow RNA profiling from tissues rich in RNases, such as older mouse embryos or adult tissues, that were problematic for the original method. This robustness to RNase activity expands the range of tissues that can be analyzed using combinatorial indexing.

### Reverse Transcription and Barcoding Rounds

After nuclei isolation and fixation, the workflow proceeds through sequential rounds of barcoding. In the first round, nuclei are distributed across a multiwell plate, and each well receives a well-specific barcode during reverse transcription. The reverse transcription reaction converts messenger RNA into complementary DNA while simultaneously incorporating the first barcode. After reverse transcription, all nuclei are pooled, washed, and redistributed for the second round of barcoding.

The second round typically involves tagmentation, in which a transposase enzyme fragments the complementary DNA and ligates the second barcode to the fragmented molecules. The [mouse kidney library preparation protocol](https://doi.org/10.1016/j.xpro.2022.101904) describes the use of self-loaded transposome Tn5 for tagmentation in library generation. After tagmentation, nuclei are pooled again and redistributed for the third round of barcoding, which typically occurs during PCR amplification. The PCR reaction adds the third barcode along with sequencing adapters, completing the barcode combination and generating the final sequencing library.

### Library Amplification and Sequencing

The final library consists of complementary DNA fragments that carry the complete three-barcode combination identifying their cell of origin. Library amplification occurs during the final PCR step, which also adds the sequencing adapters required for cluster generation on Illumina platforms. The amplified library is then purified, quantified, and submitted for sequencing.

Sequencing requirements for combinatorial indexing libraries differ from droplet-based approaches. Because each fragment carries a three-part barcode, the sequencing read must include sufficient length to capture all barcode components. The [easySHARE-seq method](https://doi.org/10.7554/elife.110034) notes that libraries with usable sequences of up to 300 base pairs are suitable for investigation of allele-specific signals or variant discovery, and that the libraries do not require a dedicated sequencing run, saving costs. The sequencing depth per cell can be adjusted based on the experimental goals, with lower depth sufficient for cell type identification and higher depth required for detailed gene expression analysis.

## At a Glance: Combinatorial Indexing Compared to Droplet-Based Approaches

| Feature | Combinatorial Indexing | Droplet-Based Microfluidics |
|---------|------------------------|----------------------------|
| Equipment required | Standard molecular biology equipment: multichannel pipettes, thermocycler, centrifuge | Specialized microfluidic instrument and consumables |
| Reagent cost per cell | Approximately one cent or less per cell | Higher per-cell cost due to barcoded beads and microfluidic chips |
| Throughput per experiment | Hundreds of thousands to millions of cells | Typically thousands to tens of thousands of cells per run |
| Hands-on time | Two to three days from nuclei isolation to final library | Several hours for encapsulation plus library preparation |
| Sample multiplexing | Built-in through combinatorial barcode combinations | Requires additional cell hashing or multiplexing strategies |
| Multimodal capability | RNA, chromatin accessibility, CRISPR perturbations, VDJ | RNA, chromatin accessibility, protein expression, CRISPR perturbations |
| Sensitivity | Lower than droplet-based methods for some protocols | Higher sensitivity for lowly expressed genes |
| Protocol complexity | Multiple rounds of pooling and splitting require careful tracking | Single encapsulation step but requires instrument operation |

The comparison in this table reflects the practical tradeoffs that researchers must consider when selecting a single-cell approach. The [single-cell data science challenges review](https://pubmed.ncbi.nlm.nih.gov/32033589) notes that the combination of microfluidics and combinatorial indexing strategies, along with low sequencing costs, has empowered single-cell sequencing technology, with thousands or even millions of cells analyzed in a single experiment. The choice between combinatorial indexing and droplet-based approaches depends on available equipment, budget, required sensitivity, and experimental scale.

## Scaling Considerations and Throughput Limits

### Theoretical Barcode Capacity

The theoretical throughput of combinatorial indexing is determined by the number of barcode combinations that can be generated. With three rounds of indexing using 96-well plates, the theoretical capacity is approximately 884,000 combinations. Using 384-well plates increases the capacity to over 56 million combinations. In practice, the number of cells that can be profiled is lower than the theoretical capacity because loading cells at a fraction of the barcode space minimizes the probability of two cells sharing the same barcode combination.

The [mouse organogenesis cell atlas](https://pubmed.ncbi.nlm.nih.gov/30787437) demonstrated the upper end of combinatorial indexing scale by profiling approximately two million cells in a single experiment. This scale enabled the identification of hundreds of cell types and 56 developmental trajectories, many of which were detected only because of the depth of cellular coverage. The ability to profile millions of cells in a single experiment is a defining advantage of combinatorial indexing, particularly for projects that require detection of rare cell populations or comprehensive atlas construction.

### Sample Multiplexing

Combinatorial indexing provides built-in sample multiplexing because cells from different biological samples can be processed together through the indexing rounds. Each sample can be assigned a distinct set of barcode combinations, or samples can be distinguished by a separate sample barcode incorporated during the first round. The [scifi-RNA-seq method](https://doi.org/10.1038/s41592-021-01153-z) combines one-step combinatorial preindexing of entire transcriptomes inside permeabilized cells with subsequent single-cell RNA-seq using microfluidics, providing a straightforward way of multiplexing thousands of samples in a single experiment.

The [kidney fibrosis study](https://pubmed.ncbi.nlm.nih.gov/36265491) used single-cell combinatorial indexing RNA sequencing to analyze 24 mouse kidneys from two fibrosis models, profiling 309,666 cells in one experiment. This scale of sample multiplexing is difficult to achieve with droplet-based approaches without additional cell hashing strategies. The ability to process multiple biological replicates or experimental conditions in a single experiment reduces batch effects and simplifies experimental design.

### Rare Cell Population Detection

The high cellular throughput of combinatorial indexing enables detection of rare cell populations that would be missed with lower-throughput approaches. The [prokaryotic single-cell RNA sequencing study](https://pubmed.ncbi.nlm.nih.gov/32451472) applied combinatorial indexing to bacterial populations and revealed a rare subpopulation of cells undergoing prophage induction in wild-type Staphylococcus aureus. This rare subpopulation would have been difficult to detect with lower-throughput methods.

The [UDA-seq study](https://pubmed.ncbi.nlm.nih.gov/39833568) demonstrated that the robustness of combinatorial indexing approaches enables identification of rare cell subpopulations associated with clinical phenotypes and exploration of cancer cell vulnerability. For researchers studying heterogeneous tissues or seeking to identify rare cell types, the scale of combinatorial indexing provides statistical power that is difficult to achieve with alternative approaches.

## Data Analysis Workflow for Combinatorial Indexing Data

### Demultiplexing and Barcode Assignment

The first step in analyzing combinatorial indexing data is demultiplexing, in which sequencing reads are assigned to individual cells based on their barcode combinations. Each read carries the three-part barcode, and reads sharing the same barcode combination are grouped together as originating from the same cell. The demultiplexing process requires accurate barcode identification, which depends on sequencing quality and the design of the barcode sequences.

Barcode design is critical for accurate demultiplexing. Barcodes must be sufficiently distinct from each other to tolerate sequencing errors while remaining identifiable. The [easySHARE-seq method](https://doi.org/10.7554/elife.110034) describes improvements in barcode design that address limitations of previous methods, resulting in more robust barcode identification. After demultiplexing, reads are aligned to the reference genome, and gene expression counts are generated for each cell.

### Quality Control Metrics

Quality control for combinatorial indexing data follows principles similar to those used for droplet-based single-cell data, but with some differences in expected metrics. Key quality control metrics include the number of unique molecular identifiers per cell, the number of genes detected per cell, the fraction of reads mapping to mitochondrial genes, and the fraction of reads mapping to the genome. Cells with very low read counts or gene detection may represent empty barcode combinations or damaged nuclei, while cells with very high mitochondrial read fractions may represent dying or lysed cells.

The [single-cell data science challenges review](https://pubmed.ncbi.nlm.nih.gov/32033589) outlines eleven challenges central to single-cell data science, including preprocessing, quality control, normalization, batch correction, and cell type identification. These challenges apply to combinatorial indexing data as well as droplet-based data, and the analytical frameworks developed for single-cell data analysis are generally applicable across platforms.

### Computational Infrastructure and Training

Analysis of combinatorial indexing data requires computational infrastructure capable of handling large single-cell datasets. The [Bioconductor project](https://bioconductor.org/) provides official packages and workflows for single-cell genomic analysis, including packages for quality control, normalization, dimensionality reduction, clustering, and differential expression analysis. These packages are designed for reproducible genomic analysis and are maintained by an active developer community.

For researchers who need to build computational skills for single-cell data analysis, the [EMBL-EBI Training program](https://www.ebi.ac.uk/training) offers learning pathways for bioinformatics data resources and practical analysis education. The [Galaxy Training Network](https://training.galaxyproject.org/) provides accessible workflow training and analysis tutorials that can be applied to single-cell data. The [nf-core documentation](https://nf-co.re/docs) describes community pipeline standards for reproducible workflow configuration and usage, which can be applied to single-cell RNA sequencing analysis pipelines.

The [Carpentries lessons](https://carpentries.org/lessons) provide foundational training in computing, data handling, shell, Git, and programming that are useful for researchers who need to develop the computational skills required for single-cell data analysis. For laboratories that are new to single-cell sequencing, investing in computational training is as important as mastering the wet laboratory protocol.

## Options and Tradeoffs in Combinatorial Indexing Implementation

### Protocol Selection

Several combinatorial indexing protocols are available, each with distinct characteristics. The [optimized sci-RNA-seq protocol](https://pubmed.ncbi.nlm.nih.gov/36261634) is a simplified, faster, more robust, and more sensitive version of the original protocol, with reagent costs on the order of one cent per cell or less. This protocol includes a Tiny-Sci variant for experiments with very limited input material, making it suitable for precious samples.

The [scifi-RNA-seq method](https://doi.org/10.1038/s41592-021-01153-z) combines combinatorial preindexing with droplet-based single-cell RNA sequencing, allowing multiple cells to be loaded per droplet and computationally demultiplexed. This approach increases the throughput of droplet-based systems while retaining the sensitivity advantages of droplet-based methods. Compared to multiround combinatorial indexing, scifi-RNA-seq provides an easier and more efficient workflow.

The [FIPRESCI method](https://doi.org/10.1186/s13059-023-02893-1) focuses on 5-prime-end single-cell RNA sequencing, which can reveal promoter and enhancer activity and efficiently profile immune receptor repertoires. This method enables massive sample multiplexing and increases the throughput of droplet microfluidics systems by over tenfold, generating approximately 100,000 single-cell transcriptomes from E10.5 whole mouse embryos in a single-channel experiment.

### Sensitivity Considerations

A known limitation of combinatorial indexing is lower sensitivity compared to droplet-based methods. The [optimized sci-RNA-seq protocol](https://pubmed.ncbi.nlm.nih.gov/36261634) acknowledges that the original protocol exhibited lower sensitivity than alternative methods, and the optimized version improves sensitivity while maintaining the scalability advantages. Researchers who need to detect lowly expressed genes or who are studying transcripts with low abundance may need to consider whether combinatorial indexing provides sufficient sensitivity for their application.

The [easySHARE-seq method](https://doi.org/10.7554/elife.110034) demonstrates that combinatorial indexing-based techniques can recover over 1.5-fold more transcripts per cell than other combinatorial indexing-based techniques while retaining high scalability and low cost. This improvement suggests that ongoing protocol optimization is narrowing the sensitivity gap between combinatorial indexing and droplet-based methods.

### Multimodal Integration

Combinatorial indexing can be extended to multimodal measurements, allowing simultaneous profiling of multiple molecular features from individual cells. The [UDA-seq workflow](https://pubmed.ncbi.nlm.nih.gov/39833568) enables co-assay of RNA and VDJ, RNA and chromatin, and RNA and CRISPR perturbation, making it applicable to a wide range of study designs. The [easySHARE-seq method](https://doi.org/10.7554/elife.110034) simultaneously measures gene expression and chromatin accessibility, enabling investigation of cis-regulatory elements and their target genes.

The [single-cell omics technical guide](https://doi.org/10.3390/genes16121394) discusses different single-cell omic technologies and their application to postmortem human brain tissue, highlighting key findings in transcriptomics and epigenomics with emerging findings in proteomics, metabolomics, and multi-omics. For researchers studying complex tissues or diseases, multimodal combinatorial indexing approaches provide a more complete picture of cellular states than RNA sequencing alone.

## Observations and Measurements in Combinatorial Indexing Experiments

### Expected Performance Metrics

Researchers implementing combinatorial indexing protocols should track several performance metrics to assess experiment quality. The number of cells recovered, the median number of genes detected per cell, the median number of unique molecular identifiers per cell, and the fraction of reads mapping to the genome are standard metrics for evaluating single-cell experiments. The [optimized sci-RNA-seq protocol](https://pubmed.ncbi.nlm.nih.gov/36261634) reports profiling approximately 380,000 nuclei from an E16.5 mouse embryo in a single experiment, demonstrating the scale that can be achieved with optimized protocols.

The [kidney fibrosis study](https://pubmed.ncbi.nlm.nih.gov/36265491) profiled 309,666 cells from 24 mouse kidneys, representing 50 cell types and states encompassing epithelial, endothelial, immune, and stromal populations. This study demonstrates that combinatorial indexing can recover complex cellular heterogeneity from a single experiment, providing sufficient depth to identify diverse injury states of the proximal tubule and distinct early-phase populations with dysregulated lipid and amino acid metabolism.

### Troubleshooting Common Issues

Common issues in combinatorial indexing experiments include low cell recovery, high barcode collision rates, low gene detection sensitivity, and batch effects between indexing rounds. Low cell recovery can result from cell loss during the multiple pooling and washing steps, which can be mitigated by careful pipetting and minimizing transfer losses. High barcode collision rates indicate that too many cells were loaded relative to the barcode capacity, and the cell loading density should be reduced.

Low gene detection sensitivity can result from inefficient reverse transcription, suboptimal tagmentation, or RNA degradation during processing. The [optimized sci-RNA-seq protocol](https://pubmed.ncbi.nlm.nih.gov/36261634) describes improvements that increase sensitivity and yield, including changes to the reverse transcription and tagmentation steps. Batch effects between indexing rounds can be assessed by comparing the distribution of quality metrics across barcode combinations and can be addressed through careful experimental design and data analysis.

### Records and Documentation

Maintaining detailed records of combinatorial indexing experiments is essential for troubleshooting and reproducibility. Key records include the number of cells or nuclei loaded at each round, the barcode sequences used in each well, the pooling and splitting scheme, and the quality metrics for the final library. The [Galaxy Training Network](https://training.galaxyproject.org/) emphasizes the importance of reproducibility in bioinformatics analysis, and the [nf-core documentation](https://nf-co.re/docs) describes community standards for reproducible workflow configuration.

Documentation of the computational analysis is equally important. Recording the software versions, parameters, and reference genome versions used for alignment and quantification ensures that analyses can be reproduced or revisited. The [Bioconductor project](https://bioconductor.org/) provides official documentation for reproducible genomic analysis, and the [Carpentries lessons](https://carpentries.org/lessons) teach foundational practices for reproducible computational research.

## Common Failure Patterns and Mitigation Strategies

### Barcode Collisions and Doublets

Barcode collisions occur when two or more cells receive the same barcode combination, resulting in a merged expression profile that cannot be distinguished from a single cell. The probability of collisions increases with the number of cells loaded relative to the barcode capacity. Mitigation strategies include loading cells at a fraction of the theoretical barcode capacity, using more rounds of indexing to increase the barcode space, and using computational methods to detect and filter cells with unusually high read counts that may represent doublets.

The [scifi-RNA-seq method](https://doi.org/10.1038/s41592-021-01153-z) addresses the doublet problem in droplet-based approaches by resolving and retaining individual transcriptomes from overloaded droplets, instead of flagging and discarding droplets containing more than one cell. This approach increases throughput while maintaining the ability to identify individual cells.

### Low Sensitivity for Lowly Expressed Genes

Combinatorial indexing methods generally have lower sensitivity than droplet-based methods, meaning that lowly expressed genes may not be detected in all cells where they are expressed. This limitation can affect downstream analyses such as differential expression testing and cell type identification. Mitigation strategies include increasing sequencing depth, optimizing the reverse transcription and amplification steps, and using computational methods that account for dropout events.

The [optimized sci-RNA-seq protocol](https://pubmed.ncbi.nlm.nih.gov/36261634) reports improvements in sensitivity compared to the original protocol, and the [easySHARE-seq method](https://doi.org/10.7554/elife.110034) demonstrates that combinatorial indexing-based techniques can recover more transcripts per cell than previous methods. Researchers should be aware of the sensitivity characteristics of their chosen protocol and consider whether their biological questions require detection of lowly expressed genes.

### Batch Effects Across Indexing Rounds

Batch effects can arise from differences in reagent lots, processing times, or technician performance across the multiple rounds of indexing. These effects can confound biological differences if they correlate with experimental conditions. Mitigation strategies include processing all samples together through the indexing rounds, using the same reagent lots for all samples, and including technical replicates to assess batch effects.

The [single-cell data science challenges review](https://pubmed.ncbi.nlm.nih.gov/32033589) identifies batch correction as a central challenge in single-cell data science, and multiple computational methods have been developed to address batch effects in single-cell data. Researchers should assess batch effects in their data and apply appropriate correction methods when necessary.

### RNase Degradation in Difficult Tissues

Some tissues are rich in RNases, which can degrade RNA during the extended processing time required for combinatorial indexing. The [optimized sci-RNA-seq protocol](https://pubmed.ncbi.nlm.nih.gov/36261634) notes that improvements in the protocol allow RNA profiling from tissues rich in RNases, such as older mouse embryos or adult tissues, that were problematic for the original method. For difficult tissues, researchers should consider using the optimized protocol and taking additional precautions to minimize RNase activity.

## Limitations and Interpretation Boundaries

### Sensitivity Compared to Droplet-Based Methods

The lower sensitivity of combinatorial indexing compared to droplet-based methods is a well-documented limitation. Researchers should interpret gene expression measurements from combinatorial indexing data with this limitation in mind, particularly for genes with low expression levels. The [optimized sci-RNA-seq protocol](https://pubmed.ncbi.nlm.nih.gov/36261634) acknowledges that the original protocol exhibited lower sensitivity than alternative methods, and while the optimized version improves sensitivity, it may still not match the sensitivity of droplet-based approaches for all applications.

### Computational Complexity of Demultiplexing

The demultiplexing of combinatorial indexing data is computationally intensive because each read must be assigned to a cell based on its barcode combination. Errors in barcode reading can lead to incorrect cell assignment, and the accuracy of demultiplexing depends on the sequencing quality and the design of the barcode sequences. Researchers should use established demultiplexing tools and validate the accuracy of cell assignment through quality control metrics.

### Interpretation of Multimodal Data

Multimodal combinatorial indexing approaches generate complex datasets that require specialized analysis methods. The [single-cell omics technical guide](https://doi.org/10.3390/genes16121394) discusses the application of single-cell omic technologies to neuropsychiatric research, highlighting the need for careful interpretation of multimodal data. Researchers should be aware that integrating multiple molecular modalities from the same cells requires computational methods that account for the different data types and their respective noise characteristics.

## Safety and Regulatory Context

### Biosafety Considerations

Combinatorial indexing experiments involve handling biological samples, including cells, tissues, and potentially infectious agents. Researchers should follow institutional biosafety guidelines for handling biological materials, including appropriate personal protective equipment and waste disposal procedures. The [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/) provide access to sequence data and associated metadata that can inform biosafety assessments for research organisms.

### Data Sharing and Deposition

Single-cell sequencing data generated through combinatorial indexing should be deposited in public databases to enable reproducibility and data reuse. The [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/) provide official descriptions of databases, search systems, sequence resources, and analysis services for depositing and accessing genomic data. Researchers should follow data deposition guidelines from their funding agencies and journals, including providing detailed metadata about the experimental protocol and analysis methods.

### Ethical Use of Human Samples

Studies involving human samples must comply with institutional review board requirements and applicable regulations. The [single-cell omics technical guide](https://doi.org/10.3390/genes16121394) discusses the application of single-cell omic technologies to postmortem human brain tissue, highlighting the importance of ethical considerations in human tissue research. Researchers should ensure that appropriate consent and ethical approvals are in place before initiating studies involving human samples.

## Professional Escalation Criteria

Researchers should seek expert consultation when encountering specific challenges in combinatorial indexing experiments. Escalation is appropriate when cell recovery is consistently below expected levels despite protocol optimization, when barcode collision rates are unexpectedly high, when gene detection sensitivity is insufficient for the biological questions being addressed, or when computational analysis reveals patterns that cannot be explained by known technical or biological factors.

Consultation with bioinformatics core facilities or experienced single-cell researchers is recommended when planning large-scale combinatorial indexing experiments, when implementing multimodal approaches, or when integrating combinatorial indexing data with other data types. The [EMBL-EBI Training program](https://www.ebi.ac.uk/training) and the [Galaxy Training Network](https://training.galaxyproject.org/) provide educational resources that can help researchers build the skills needed to address common challenges in single-cell data analysis.

## A Practical Decision Framework for Choosing Between Combinatorial Indexing Protocols

Researchers who decide to implement combinatorial indexing still face a second decision that is often more consequential than the choice between combinatorial indexing and droplet-based methods: which combinatorial indexing protocol best fits their specific experimental constraints. The published protocols differ meaningfully in throughput, sensitivity, hands-on time, input requirements, and compatibility with downstream applications. A structured decision framework helps laboratories match protocol characteristics to their biological questions, available infrastructure, and technical expertise.

### Step 1: Define the Scale and Input Constraints

The first decision point concerns the number of cells or nuclei required and the amount of starting material available. For experiments requiring hundreds of thousands to millions of cells, the [optimized sci-RNA-seq protocol](https://pubmed.ncbi.nlm.nih.gov/36261634) is the most direct path, as it was demonstrated by profiling approximately 380,000 nuclei from an E16.5 mouse embryo in a single experiment. This protocol uses three rounds of split-pool indexing and achieves reagent costs on the order of one cent per cell or less. The same protocol includes a Tiny-Sci variant designed for experiments with very limited input material, making it suitable for precious or difficult-to-obtain samples.

For experiments that require 5-prime-end transcript information, such as promoter and enhancer activity analysis or immune receptor repertoire profiling, the [FIPRESCI method](https://doi.org/10.1186/s13059-023-02893-1) provides a 5-prime-end combinatorial indexing approach that generates approximately 100,000 single-cell transcriptomes from E10.5 whole mouse embryos in a single-channel experiment. This method also enables simultaneous identification of T cell receptor signatures from peripheral blood T cells across multiple cancer patients, demonstrating its utility for immune profiling applications.

### Step 2: Assess the Need for Droplet-Based Sensitivity

The sensitivity gap between combinatorial indexing and droplet-based methods is well documented, but the magnitude of this gap varies across protocols. The [optimized sci-RNA-seq protocol](https://pubmed.ncbi.nlm.nih.gov/36261634) acknowledges that the original protocol exhibited lower sensitivity than alternative methods, and the optimized version improves sensitivity while maintaining scalability. For researchers who need higher sensitivity for lowly expressed genes but still want the cost and throughput advantages of combinatorial indexing, hybrid approaches offer an intermediate option.

The [scifi-RNA-seq method](https://doi.org/10.1038/s41592-021-01153-z) combines one-step combinatorial preindexing of entire transcriptomes inside permeabilized cells with subsequent single-cell RNA-seq using microfluidics. This approach allows multiple cells to be loaded per droplet and computationally demultiplexed, massively increasing the throughput of droplet-based systems while retaining the sensitivity advantages of droplet-based methods. Compared to multiround combinatorial indexing, scifi-RNA-seq provides an easier and more efficient workflow, and compared to cell hashing methods that discard overloaded droplets, scifi-RNA-seq resolves and retains individual transcriptomes.

### Step 3: Determine Multimodal Requirements

If the experimental question requires simultaneous measurement of multiple molecular modalities, the protocol selection narrows considerably. The [easySHARE-seq method](https://doi.org/10.7554/elife.110034) simultaneously measures gene expression and chromatin accessibility, recovering 19,664 nuclei with joint chromatin and expression profiles from murine liver nuclei. This method produces libraries with usable sequences of up to 300 base pairs, making it suitable for allele-specific signal investigation or variant discovery, and the libraries do not require a dedicated sequencing run, saving costs.

For experiments requiring co-assay of RNA and VDJ sequences, RNA and chromatin, or RNA and CRISPR perturbation, the [UDA-seq workflow](https://pubmed.ncbi.nlm.nih.gov/39833568) integrates a post-indexing step that enhances throughput and systematically adapts existing droplet-based single-cell multimodal methods. This workflow generated over 100,000 high-quality single-cell datasets from three dozen frozen clinical biopsy specimens within a single-channel droplet microfluidics experiment, demonstrating its scalability for clinical applications.

### Step 4: Evaluate Hands-On Time and Technical Complexity

The hands-on time from nuclei isolation to final library preparation ranges from two to three days for the [optimized sci-RNA-seq protocol](https://pubmed.ncbi.nlm.nih.gov/36261634), depending on the number of samples sharing the experiment. This extended processing time reflects the multiple rounds of pooling and splitting that characterize combinatorial indexing. Laboratories with limited personnel or competing demands should factor this time commitment into their planning.

The [mouse kidney nuclear isolation and library preparation protocol](https://doi.org/10.1016/j.xpro.2022.101904) provides a step-by-step approach using a dounce tissue homogenizer for nuclei extraction with high yield, followed by sci-RNA-seq3 library preparation with self-loaded transposome Tn5 for tagmentation. This protocol is designed to allow researchers to generate scalable single-cell transcriptomic data with common laboratory supplies at low cost, but it requires careful attention to the multiple processing steps.

### Step 5: Consider Sample Multiplexing Needs

Combinatorial indexing provides built-in sample multiplexing because cells from different biological samples can be processed together through the indexing rounds. The [kidney fibrosis study](https://pubmed.ncbi.nlm.nih.gov/36265491) used single-cell combinatorial indexing RNA sequencing to analyze 24 mouse kidneys from two fibrosis models, profiling 309,666 cells in one experiment. This scale of sample multiplexing reduces batch effects and simplifies experimental design by allowing all samples to be processed together.

For experiments requiring thousands of samples in a single experiment, the [scifi-RNA-seq method](https://doi.org/10.1038/s41592-021-01153-z) provides a straightforward way of multiplexing thousands of samples through its combinatorial preindexing approach. The [FIPRESCI method](https://doi.org/10.1186/s13059-023-02893-1) also enables massive sample multiplexing, as demonstrated by the simultaneous identification of T cell receptor signatures from peripheral blood T cells of 12 cancer patients.

### Decision Matrix for Protocol Selection

| Experimental Requirement | Recommended Protocol | Key Evidence |
|--------------------------|---------------------|--------------|
| Maximum scale, minimal cost | Optimized sci-RNA-seq | Approximately 380,000 nuclei in one experiment at one cent per cell or less |
| Very limited input material | Tiny-Sci variant | Designed for experiments with very limited input material |
| 5-prime-end transcript information | FIPRESCI | Approximately 100,000 transcriptomes from E10.5 embryos, immune receptor profiling |
| Higher sensitivity with combinatorial indexing | scifi-RNA-seq | Combines preindexing with droplet microfluidics, resolves overloaded droplets |
| Simultaneous RNA and chromatin | easySHARE-seq | 19,664 nuclei with joint profiles, 300 bp usable sequences |
| Multimodal co-assay (RNA, VDJ, chromatin, CRISPR) | UDA-seq | Over 100,000 datasets from clinical biopsy specimens |
| Multiple samples in one experiment | Any protocol with built-in multiplexing | 24 mouse kidneys profiled in one experiment |

### Implementation Assessment Checklist

Before committing to a specific protocol, laboratories should complete a structured assessment of their readiness. First, verify that all required equipment is available and functional, including multichannel pipettes calibrated for the plate formats being used, a thermocycler with sufficient capacity for the number of plates per round, and a centrifuge capable of the required speeds for nuclei washing steps. Second, confirm that the laboratory has experience with the fundamental techniques required, particularly reverse transcription, PCR amplification, and careful multiwell plate handling. Third, assess whether the computational infrastructure can handle the expected data volume, including storage capacity, memory, and processing power for demultiplexing and downstream analysis.

Fourth, evaluate whether the biological question requires the sensitivity that only droplet-based methods can provide. If the genes of interest are known to be lowly expressed, or if the experiment aims to detect rare transcripts, the sensitivity limitations of combinatorial indexing may compromise the results. Fifth, consider whether the extended hands-on time of two to three days is feasible given personnel availability and competing laboratory demands. Sixth, determine whether the protocol has been validated on the tissue type of interest, as the [optimized sci-RNA-seq protocol](https://pubmed.ncbi.nlm.nih.gov/36261634) notes that the original method exhibited variable performance on different tissues, and the optimized version improves robustness for RNase-rich tissues such as older mouse embryos and adult tissues.

### Records and Measurements for Protocol Comparison

Laboratories that are comparing protocols should maintain standardized records to enable meaningful comparison. For each protocol tested, record the number of cells or nuclei loaded at each round, the number of cells recovered after demultiplexing, the median number of genes detected per cell, the median number of unique molecular identifiers per cell, the fraction of reads mapping to the genome, and the total hands-on time required. These metrics should be recorded for the same tissue type and input material to ensure that differences reflect protocol performance instead of sample variation.

The [single-cell data science challenges review](https://pubmed.ncbi.nlm.nih.gov/32033589) notes that thousands or even millions of cells analyzed in a single experiment amount to a data revolution in single-cell biology, but this scale also poses unique data science problems. Laboratories should plan for the computational demands of processing and analyzing combinatorial indexing data, including the demultiplexing step that assigns reads to cells based on barcode combinations. The [Bioconductor project](https://bioconductor.org/) provides official packages and workflows for single-cell genomic analysis, and the [Galaxy Training Network](https://training.galaxyproject.org/) offers accessible workflow training that can help laboratories build the computational skills needed for data analysis.

### Common Failure Patterns in Protocol Implementation

Several failure patterns recur when laboratories implement combinatorial indexing protocols for the first time. The most common is low cell recovery, which typically results from cumulative cell loss during the multiple pooling and washing steps. Each round of pooling and redistribution introduces opportunities for cell loss, and laboratories should monitor recovery at each step to identify where losses occur. The [mouse kidney nuclear isolation protocol](https://doi.org/10.1016/j.xpro.2022.101904) emphasizes the use of a dounce tissue homogenizer for high-yield nuclei extraction, which addresses the initial input stage of the workflow.

A second common failure pattern is high barcode collision rates, which indicate that too many cells were loaded relative to the barcode capacity. The probability of collisions increases with the number of cells loaded, and laboratories should load cells at a fraction of the theoretical barcode capacity to minimize collision rates. A third failure pattern is low gene detection sensitivity, which can result from inefficient reverse transcription, suboptimal tagmentation, or RNA degradation during processing. The [optimized sci-RNA-seq protocol](https://pubmed.ncbi.nlm.nih.gov/36261634) describes improvements that increase sensitivity and yield, and laboratories experiencing low sensitivity should review their protocol execution against the optimized version.

A fourth failure pattern is batch effects between indexing rounds, which can arise from differences in reagent lots, processing times, or technician performance. These effects can confound biological differences if they correlate with experimental conditions. Laboratories should process all samples together through the indexing rounds and use the same reagent lots for all samples to minimize batch effects. The [single-cell data science challenges review](https://pubmed.ncbi.nlm.nih.gov/32033589) identifies batch correction as a central challenge in single-cell data science, and multiple computational methods are available to address batch effects when they occur.

### Professional Escalation Criteria for Protocol Selection

Laboratories should seek expert consultation when they cannot determine which protocol best fits their experimental constraints, when they encounter persistent technical failures despite following published protocols, or when they need to implement multimodal approaches that require specialized expertise. Consultation with bioinformatics core facilities or experienced single-cell researchers is recommended when planning large-scale combinatorial indexing experiments, when implementing hybrid approaches such as scifi-RNA-seq that combine combinatorial indexing with droplet microfluidics, or when integrating combinatorial indexing data with other data types.

The [EMBL-EBI Training program](https://www.ebi.ac.uk/training) offers learning pathways for bioinformatics data resources and practical analysis education, and the [Carpentries lessons](https://carpentries.org/lessons) provide foundational training in computing, data handling, shell, Git, and programming that are useful for building the computational skills required for single-cell data analysis. For laboratories that are new to combinatorial indexing, investing in training before protocol implementation can reduce the likelihood of costly technical failures and improve the quality of the resulting data.

## Frequently Asked Questions

### What equipment is required for combinatorial indexing experiments?

Combinatorial indexing requires only standard molecular biology equipment, including multichannel pipettes, a thermocycler, a centrifuge, and standard laboratory consumables such as multiwell plates. No specialized microfluidic instruments are needed. The [optimized sci-RNA-seq protocol](https://pubmed.ncbi.nlm.nih.gov/36261634) describes a workflow that uses common laboratory supplies and achieves reagent costs on the order of one cent per cell or less.

### How many cells can be profiled in a single combinatorial indexing experiment?

The number of cells that can be profiled depends on the number of barcode combinations generated by the indexing rounds. With three rounds of indexing using 96-well plates, the theoretical capacity is approximately 884,000 cells. The [mouse organogenesis study](https://pubmed.ncbi.nlm.nih.gov/30787437) profiled approximately two million cells in a single experiment, and the [kidney fibrosis study](https://pubmed.ncbi.nlm.nih.gov/36265491) profiled 309,666 cells from 24 mouse kidneys.

### How does combinatorial indexing compare to droplet-based single-cell sequencing in cost?

Combinatorial indexing is substantially less expensive than droplet-based approaches, with reagent costs on the order of one cent per cell or less as described in the [optimized sci-RNA-seq protocol](https://pubmed.ncbi.nlm.nih.gov/36261634). Droplet-based approaches require specialized consumables including barcoded beads and microfluidic chips, which increase per-cell costs. The elimination of specialized instrument costs makes combinatorial indexing accessible to laboratories without dedicated single-cell platforms.

### What are the main limitations of combinatorial indexing?

The main limitations are lower sensitivity compared to droplet-based methods, the complexity of the multi-round protocol, and the computational demands of demultiplexing large datasets. The [optimized sci-RNA-seq protocol](https://pubmed.ncbi.nlm.nih.gov/36261634) acknowledges that the original protocol exhibited lower sensitivity than alternative methods, and while the optimized version improves sensitivity, researchers should consider whether their biological questions require detection of lowly expressed genes.

### Can combinatorial indexing be used for multimodal single-cell analysis?

Yes, combinatorial indexing can be extended to multimodal measurements. The [UDA-seq workflow](https://pubmed.ncbi.nlm.nih.gov/39833568) enables co-assay of RNA and VDJ, RNA and chromatin, and RNA and CRISPR perturbation. The [easySHARE-seq method](https://doi.org/10.7554/elife.110034) simultaneously measures gene expression and chromatin accessibility, enabling investigation of cis-regulatory elements and their target genes.

### What quality control metrics should be tracked in combinatorial indexing experiments?

Key quality control metrics include the number of cells recovered, the median number of genes detected per cell, the median number of unique molecular identifiers per cell, the fraction of reads mapping to the genome, and the fraction of reads mapping to mitochondrial genes. These metrics should be assessed across barcode combinations to identify potential batch effects or technical issues.

### How should combinatorial indexing data be analyzed?

Combinatorial indexing data should be analyzed using established single-cell RNA sequencing analysis workflows, including demultiplexing, alignment, quantification, quality control, normalization, dimensionality reduction, clustering, and differential expression analysis. The [Bioconductor project](https://bioconductor.org/) provides official packages and workflows for single-cell genomic analysis, and the [Galaxy Training Network](https://training.galaxyproject.org/) offers accessible workflow training and analysis tutorials.

### What training resources are available for researchers new to combinatorial indexing?

The [EMBL-EBI Training program](https://www.ebi.ac.uk/training) offers learning pathways for bioinformatics data resources and practical analysis education. The [Galaxy Training Network](https://training.galaxyproject.org/) provides accessible workflow training and analysis tutorials. The [Carpentries lessons](https://carpentries.org/lessons) provide foundational training in computing, data handling, shell, Git, and programming that are useful for building the computational skills required for single-cell data analysis.

## Related Bioinformatics Guides

- [Single-Cell RNA Sequencing Depth: A Cost-Benefit Analysis for Experimental Design](/knowledge/bioinformatics/single-cell-rna-sequencing-depth-a-cost-benefit-analysis-for-experimental-design)
- [Single-Cell Sequencing Depth: How Much Is Enough?](/knowledge/bioinformatics/single-cell-sequencing-depth-how-much-is-enough)
- [Single-Cell vs Single-Nucleus RNA Sequencing: Choosing the Right Approach](/knowledge/bioinformatics/single-cell-vs-single-nucleus-rna-sequencing-choosing-the-right-approach)
- [Single-Cell RNA Sequencing Quality Control: A Practical Guide to Filtering and Metrics](/knowledge/bioinformatics/single-cell-rna-sequencing-quality-control-a-practical-guide-to-filtering-and-metrics)
- [Spatial Transcriptomics vs. Single-Cell RNA Sequencing: Which Approach Fits Your Research?](/knowledge/bioinformatics/spatial-transcriptomics-vs-single-cell-rna-sequencing-which-approach-fits-your-research)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)


## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [Combinatorial single-cell CRISPR screens by direct guide RNA capture and targeted sequencing.](https://pubmed.ncbi.nlm.nih.gov/32231336). Nature biotechnology, 2020.
- [UDA-seq: universal droplet microfluidics-based combinatorial indexing for massive-scale multimodal single-cell sequencing.](https://pubmed.ncbi.nlm.nih.gov/39833568). Nature methods, 2025.
- [Optimized single-nucleus transcriptional profiling by combinatorial indexing.](https://pubmed.ncbi.nlm.nih.gov/36261634). Nature protocols, 2023.
- [The single-cell transcriptional landscape of mammalian organogenesis.](https://pubmed.ncbi.nlm.nih.gov/30787437). Nature, 2019.
- [Comprehensive single-cell transcriptional profiling defines shared and unique epithelial injury responses during kidney fibrosis.](https://pubmed.ncbi.nlm.nih.gov/36265491). Cell metabolism, 2022.
- [Transcriptomic, epigenomic, and spatial metabolomic cell profiling redefines regional human kidney anatomy.](https://pubmed.ncbi.nlm.nih.gov/38513647). Cell metabolism, 2024.
- [Prokaryotic single-cell RNA sequencing by in situ combinatorial indexing.](https://pubmed.ncbi.nlm.nih.gov/32451472). Nature microbiology, 2020.
- [Eleven grand challenges in single-cell data science.](https://pubmed.ncbi.nlm.nih.gov/32033589). Genome biology, 2020.
- [Mouse kidney nuclear isolation and library preparation for single-cell combinatorial indexing RNA sequencing.](https://doi.org/10.1016/j.xpro.2022.101904). 2022.
- [Flexible and high-throughput simultaneous profiling of gene expression and chromatin accessibility in single cells.](https://doi.org/10.7554/elife.110034). 2026.
- [A Single-Cell Omics Technical Guide for Advancing Neuropsychiatric Research.](https://doi.org/10.3390/genes16121394). 2025.
- [Imaging-Based Spatial Transcriptomics: Data Interpretation Methods and Biomedical Applications.](https://doi.org/10.3390/biology15120900). 2026.
- [Mouse kidney nuclear isolation and library preparation for single-Mouse kidney nuclear isolation and library preparation for single-cell combinatorial indexing RNA sequencing cell combinatorial indexing RNA sequencing](https://www.semanticscholar.org/paper/da4fd219148c70ea02ac3e06b26d573acea0e9d4)
- [Ultra-high-throughput single-cell RNA sequencing and perturbation screening with combinatorial fluidic indexing](https://doi.org/10.1038/s41592-021-01153-z). Nature Methods, 2021.
- [FIPRESCI: droplet microfluidics based combinatorial indexing for massive-scale 5′-end single-cell RNA sequencing](https://doi.org/10.1186/s13059-023-02893-1). Genome Biology, 2023.
- [Publisher Correction: FIPRESCI: droplet microfluidics based combinatorial indexing for massive-scale 5′-end single-cell RNA sequencing](https://doi.org/10.1186/s13059-023-02944-7). Genome Biology, 2023.
- [Overloading And unpacKing (OAK) - droplet-based combinatorial indexing for ultra-high throughput single-cell multiomic profiling](https://doi.org/10.1038/s41467-024-53227-z). Nature Communications, 2024.
- [A protocol for time-resolved transcriptomics through metabolic labeling and combinatorial indexing](https://doi.org/10.1016/j.xpro.2024.103356). STAR Protocols, 2024.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.