Multiplexing in Sequencing: How to Pool Samples Efficiently
Multiplexing in next-generation sequencing (NGS) allows multiple samples to be sequenced together in a single run by attaching unique index sequences to each sample's library. This approach reduces per-sample cost and increases laboratory throughput, but it requires careful planning around index design, pooling ratios, and coverage calculations. This article explains how multiplexing works, how to calculate coverage per sample, and how to avoid common problems such as index hopping and unbalanced read distribution. The guidance is written for laboratory students, technicians, researchers, and diagnostic professionals who need practical, evidence-based decisions for sample pooling.
At a Glance
Multiplexing is the practice of combining multiple indexed libraries into one sequencing run. The core decisions involve choosing the number of samples per run, determining the pooling strategy, and calculating the expected coverage for each sample. The table below summarizes the main pooling strategies and their appropriate applications.
| Pooling Strategy | Best Application | Key Consideration |
|---|---|---|
| Equal volume pooling | Libraries from the same species or samples with similar genome sizes | Simple to perform but can produce uneven read distribution if library concentrations vary |
| Equimolar pooling | Runs involving multiple species or samples with different genome sizes | Requires accurate quantification of each library before pooling |
| Normalized pooling | Clinical diagnostics with fixed coverage requirements | Reduces retesting risk due to insufficient coverage depth |
The choice between equal-volume and equimolar pooling depends on the biological variation in your sample set. For bacterial genome sequencing, libraries from different strains are usually multiplexed in a single run, and normalized libraries are most often pooled in equal volumes as recommended by sequencing platform manufacturers. This equal-volume strategy works well for isolates from the same species. However, for runs involving multiple microbial species, equimolar library pooling is more appropriate because of the variation in bacterial genome size. Equimolar pooling limits the retesting risk due to insufficient coverage depth, particularly when interspecies genome size difference is more than 2-fold. The use of this alternative strategy for multiplexing pathogenic bacteria should lead to more cost-effective whole-genome sequencing applications in clinical microbiology. See the comparison of library pooling strategies for multiplexing bacterial species in NGS for the supporting evidence.
Understanding the Core Principles of Multiplexing
What Multiplexing Accomplishes in NGS Workflows
Multiplexing converts a single sequencing run into many individual sample results. Each library receives a unique index, also called a barcode, during library preparation. After sequencing, bioinformatics software sorts the reads by their index sequences and assigns them to the correct sample. This process is fundamental to the economics of modern sequencing because it allows laboratories to fill a flow cell with many samples instead of running one sample at a time.
The underlying chemistry of NGS closely resembles that of Sanger sequencing, and understanding this relationship helps clarify why indexing works. Mutations that were historically discovered by analog approaches like Sanger sequencing and multiplex ligation-dependent probe amplification can now be decoded from a digital signal with NGS. The basic overview of NGS mechanisms explains how high-confidence detection of single-nucleotide polymorphisms, indels, and large deletions or duplications is possible with NGS alone.
How Indexes Work
An index is a short, known DNA sequence attached to each library fragment during preparation. During data analysis, the sequencer reads the index and assigns the associated reads to the correct sample. Indexes must be sufficiently different from one another to prevent misassignment. The design of index sets is typically provided by sequencing platform manufacturers, and these sets are validated to minimize cross-talk between indexes.
Index misassignment, also called index hopping, occurs when reads carry an index that does not match their true sample of origin. This problem has been reported on some widely used sequencing platforms at rates commonly exceeding 1%. On DNB-based platforms, single index misassignment from free indexed oligos occurs at a rate of one in 36 million reads, suggesting virtually no index hopping during DNA nanoball creation and arraying. The DNB-based NGS libraries have achieved a sample-to-sample misassignment rate of 0.0001 to 0.0004% under recommended procedures. Single indexing with DNB technology provides a simple but effective method for sensitive genetic assays with large sample numbers. See the reliable multiplex sequencing study on DNB-based platforms for these findings.
The Relationship Between Multiplexing and Coverage
Coverage, also called sequencing depth, refers to the average number of times each base in the target region is read. Higher coverage increases confidence in variant calls but costs more per sample. Multiplexing directly affects coverage because the total number of reads produced by a sequencing run is divided among all pooled samples.
The calculation is straightforward. If a sequencing run produces 100 million reads and you pool 10 samples, the average number of reads per sample is 10 million before any filtering. The actual coverage per sample depends on the library concentration of each sample in the pool, the genome size or target region size, and the efficiency of the sequencing run.
Calculating Coverage per Sample
The Basic Coverage Formula
Coverage is calculated by dividing the number of reads assigned to a sample by the size of the target region, then multiplying by the read length. The formula is:
Coverage = (Number of reads for the sample x Read length) / Target region size
For example, if a sample receives 5 million reads with 150 base pair read length and the target region is 1 megabase, the coverage is 750x. If the target is a whole human genome of approximately 3 gigabases, the same 5 million reads produce only 0.25x coverage.
Accounting for Read Losses
Not all reads produced by the sequencer pass quality filters. A portion of reads will fail due to low quality scores, adapter contamination, or index read failures. Laboratories should plan for a 10 to 20% loss during demultiplexing and quality filtering. This means the coverage calculation should include an expected loss factor.
Using a Sequencing Coverage Calculator
A sequencing coverage calculator helps plan multiplexing experiments. The inputs are total run output in reads or bases, number of samples, expected read length, target region size, and expected read loss. The output is the expected coverage per sample. Laboratories should run this calculation before preparing libraries to confirm that the planned pool size will produce sufficient coverage for the intended application.
Coverage Requirements by Application
Different applications require different coverage depths. Clinical diagnostic applications typically require higher coverage to confidently detect variants at low allele frequencies. Research applications may tolerate lower coverage depending on the question being asked. The bioanalytical method validation guidance from the FDA emphasizes that analytical methods must be validated for their intended use, and this principle applies to sequencing coverage decisions.
Pooling Strategies and Their Tradeoffs
Equal Volume Pooling
Equal volume pooling is the simplest approach. Each normalized library is added to the pool in the same volume. This strategy is recommended by sequencing platform manufacturers and works well when all libraries come from the same species or have similar genome sizes. The main advantage is simplicity, and the main risk is uneven read distribution if library concentrations are not accurately normalized.
Equimolar Pooling
Equimolar pooling adjusts the volume of each library so that the same number of molecules is added from each sample. This requires accurate quantification of each library, typically using fluorometric methods or quantitative PCR. Equimolar pooling is more appropriate for runs involving multiple microbial species because of the variation in bacterial genome size. The comparison of pooling strategies for bacterial species demonstrated that equimolar pooling limits retesting risk due to insufficient coverage depth, particularly when interspecies genome size difference is more than 2-fold.
Normalized Pooling for Clinical Diagnostics
Clinical diagnostic laboratories often use normalized pooling to meet fixed coverage requirements. Libraries are quantified and diluted to a target concentration before pooling. This approach reduces the risk of a sample failing to reach the minimum coverage needed for confident variant calling. The laboratory quality management system handbook from the World Health Organization provides guidance on quality practices that apply to library quantification and pooling in diagnostic settings.
Pooling in Specialized Applications
Pooling strategies extend beyond standard NGS workflows. In preimplantation genetic testing, NGS is performed with automatic library preparation and multiplexing up to 24 to 96 samples. The optimized NGS approach for aneuploidy detection validated protocols for detecting homogeneous and segmental aneuploidies, different degrees of mosaicism, and small deletions and duplications with high sensitivity and specificity.
In forensic genetics, NGS methods overcome some limitations of capillary electrophoresis, including low multiplexing capabilities and limited performance with challenging samples. The review of NGS and SNP microarrays in forensic practice notes that NGS enables STR sequencing and SNP typing with enhanced discriminatory power, better performance with degraded DNA, and improved mixture deconvolution. However, adoption in routine forensic practice remains limited due to high costs, technical complexity, and a lack of standardized protocols and legal frameworks.
Practical Workflow for Multiplexed Sequencing
Step 1: Define the Coverage Requirement
Determine the minimum coverage needed for your application before designing the multiplex. Clinical applications that require detection of low-frequency variants need higher coverage than applications that only need to confirm the presence of a known variant. Document the coverage requirement in your laboratory records.
Step 2: Calculate the Number of Samples per Run
Divide the total expected usable reads by the reads needed per sample. The reads needed per sample depend on the target region size and the coverage requirement. Include an expected loss factor of 10 to 20% for quality filtering and demultiplexing.
Step 3: Prepare and Quantify Libraries
Prepare each library according to your validated protocol. Quantify each library accurately before pooling. The quantification method should be appropriate for the library type and the downstream pooling strategy. For equimolar pooling, the quantification must be precise because small errors in concentration translate directly to uneven read distribution.
Step 4: Pool Libraries According to the Chosen Strategy
For equal volume pooling, combine equal volumes of each normalized library. For equimolar pooling, calculate the volume of each library needed to contribute the same number of molecules. Mix the pool thoroughly to ensure homogeneity.
Step 5: Verify the Pool Before Sequencing
Run a quality check on the pooled library. This can include a fluorometric quantification, a fragment size analysis, and optionally a low-depth test sequencing run. The verification step catches pooling errors before the full sequencing run.
Step 6: Sequence and Demultiplex
Load the pooled library onto the sequencer according to the manufacturer instructions. After sequencing, demultiplex the reads by index. Check the read distribution across samples to confirm that the pooling strategy worked as planned.
Step 7: Document and Review
Record the pooling strategy, library concentrations, expected coverage, and actual read distribution for each run. Review the results to identify any systematic issues with the pooling process.
Records and Measurements for Multiplexing
Essential Records
Laboratories should maintain records of the following for each multiplexed sequencing run:
- Library preparation date and protocol version
- Quantification method and results for each library
- Pooling strategy and volumes used
- Expected coverage calculation
- Actual read counts per sample after demultiplexing
- Quality metrics for each sample
- Any deviations from the standard protocol
These records support troubleshooting and continuous improvement. The laboratory quality management system handbook from the World Health Organization emphasizes the importance of documentation for ensuring reliable laboratory results.
Key Measurements
The most important measurement for multiplexing is the concentration of each library before pooling. Inaccurate quantification is the most common cause of uneven read distribution. Laboratories should use a quantification method that is validated for the library type being measured.
After sequencing, the read count per sample is the key outcome measurement. Compare the actual read distribution to the expected distribution. If one sample consistently receives more reads than expected, the quantification or pooling process needs adjustment.
Monitoring Index Performance
Track the rate of index misassignment for each run. While some platforms have very low misassignment rates, others may show higher rates. The DNB-based platform study demonstrated that index misassignment can be as low as 0.0001 to 0.0004% under recommended procedures. If your platform shows higher rates, investigate the cause and consider using unique dual indexes.
Common Failure Patterns in Multiplexing
Uneven Read Distribution
The most common failure in multiplexing is uneven read distribution across samples. One or a few samples consume a disproportionate share of the reads, leaving other samples with insufficient coverage. This is usually caused by inaccurate library quantification, pipetting errors during pooling, or differences in library complexity that affect amplification during sequencing.
Index Hopping
Index hopping occurs when reads are assigned to the wrong sample. This can cause false positive variant calls, particularly for low-frequency variants. The risk of index hopping varies by platform. On platforms with higher hopping rates, using unique dual indexes can help identify and filter hopped reads.
Adapter Contamination
Adapter contamination occurs when sequencing reads contain adapter sequences instead of insert DNA. This reduces the usable data from a run and can affect all samples in the pool. Adapter contamination is more common with low-input samples or when library preparation fails to remove adapters completely.
Sample Cross-Contamination
Cross-contamination between samples can occur during library preparation or pooling. This is particularly problematic for sensitive applications such as detecting low-frequency variants. The barcode-integrated reverse transcription approach demonstrates how barcoding at reverse transcription enables early pooling and authenticates RNA against DNA contamination, with the primer barcoding RNA 909-fold more often than genomic DNA.
Pooling Errors
Simple errors in pooling, such as adding the wrong volume or skipping a sample, can ruin a run. These errors are preventable with careful workflow design, including checklists and verification steps.
Quality Controls for Multiplexed Sequencing
Positive and Negative Controls
Include a positive control with a known variant or known sequence in each multiplexed run. This confirms that the entire workflow, from library preparation through sequencing and analysis, is working correctly. Include a negative control to detect contamination.
Internal Standards
Some workflows include internal standards, such as synthetic sequences with known concentrations, to monitor quantification and pooling accuracy. These standards can help identify systematic errors in the pooling process.
Replicate Samples
For clinical applications, running replicate samples can help identify variability in the workflow. The bioanalytical method validation guidance from the FDA emphasizes the importance of demonstrating precision and accuracy for analytical methods.
Quality Metrics to Monitor
Monitor the following quality metrics for each multiplexed run:
- Percentage of reads passing quality filters
- Percentage of reads assigned to each sample
- Coverage uniformity across target regions
- Rate of index misassignment
- Percentage of duplicate reads
These metrics provide an early warning of problems in the multiplexing workflow.
Limitations and Interpretation Boundaries
Coverage Is an Average
Coverage calculations provide an average depth across the target region. Actual coverage varies across the genome, with some regions covered at much higher depth and others at lower depth. Regions with extreme GC content or repetitive sequences may have lower coverage than the average.
Pooling Cannot Compensate for Poor Library Quality
Multiplexing distributes reads among samples, but it cannot improve the quality of individual libraries. A library with low complexity or high adapter contamination will produce poor data regardless of the pooling strategy.
Index Misassignment Limits Sensitivity
Index misassignment sets a floor on the detectable allele frequency. If the misassignment rate is 1%, then variants present at less than 1% frequency cannot be reliably distinguished from index hopping artifacts. Laboratories performing sensitive assays must understand the misassignment rate of their platform and design their experiments accordingly.
Sample Pooling Affects Diversity Measurements
Pooling can affect the measurement of diversity in complex samples. The mosquito virome study found that virome diversity increased with pool size, with larger pools showing greater taxonomic richness driven by the contribution of each individual. However, single-individual analyses may provide complementary information on individual-level composition and low-abundance taxa that could be less apparent in pooled samples.
Pooling Decisions Affect Downstream Analysis
The choice of pooling strategy can affect downstream analysis results. The agricultural microbiome sampling study found that sample pooling showed greater impact on fungal diversity and substantially reduced within-group variability across all treatments. Despite these effects, differential abundance analysis revealed minimal compositional changes, with only a small fraction of microbial taxa significantly affected by pooling.
Safety and Regulatory Context
Biosafety Considerations
Multiplexing involves handling multiple samples in a single workflow, which increases the risk of cross-contamination. Laboratories should follow the laboratory biosafety manual from the World Health Organization to establish appropriate practices for handling biological samples. This includes proper use of personal protective equipment, careful pipetting techniques, and decontamination of work surfaces.
Quality Management
Diagnostic laboratories should operate within a quality management system. The laboratory quality management system handbook from the World Health Organization provides guidance on organizing laboratories to produce reliable results. This includes documentation, validation, and quality control practices that apply to multiplexed sequencing.
Method Validation
Before implementing a multiplexed sequencing workflow for diagnostic purposes, the method must be validated. The bioanalytical method validation guidance from the FDA describes the expectations for demonstrating that an analytical method is suitable for its intended use. This includes assessing accuracy, precision, sensitivity, and specificity.
Regulatory Variation
Regulatory requirements for genetic testing vary by jurisdiction. The review of preimplantation genetic testing notes that the list of indications for which PGT is allowed may vary substantially from country to country, depending on PGT regulation. Laboratories must understand the regulatory context in which they operate.
Professional Escalation Criteria
When to Stop and Investigate
Stop the workflow and investigate if any of the following occur:
- The read distribution across samples deviates substantially from the expected distribution
- The rate of index misassignment exceeds the expected range for your platform
- Quality metrics fall below the thresholds established in your validation
- A positive control fails to produce the expected result
- A negative control shows contamination
When to Consult a Specialist
Consult a specialist if you encounter problems that you cannot resolve with standard troubleshooting. This includes persistent index hopping, unexplained coverage patterns, or repeated failures in library preparation. The assay guidance manual from the National Center for Advancing Translational Sciences provides general guidance on assay development and troubleshooting that may be helpful.
When to Repeat a Run
Repeat a run if the data quality is insufficient for the intended application. This decision should be based on the coverage requirements and the quality metrics of the run. Repeating a run is preferable to reporting results from inadequate data.
Frequently Asked Questions
What is the difference between single indexing and unique dual indexing?
Single indexing uses one index sequence per sample. Unique dual indexing uses two different index sequences, one on each end of the library fragment. Unique dual indexes provide better protection against index hopping because a read must match both indexes to be assigned to a sample. This is particularly important for sensitive applications where index misassignment could cause false positive variant calls.
How many samples can be multiplexed in a single sequencing run?
The number of samples depends on the sequencing platform, the run output, and the coverage requirement for each sample. Some applications multiplex up to 96 samples per run, as demonstrated in the optimized NGS approach for aneuploidy detection. The practical limit is determined by the coverage needed for the application and the total output of the sequencing run.
How do I calculate the number of samples I can pool?
Divide the total expected usable reads by the reads needed per sample. The reads needed per sample depend on the target region size and the coverage requirement. Include an expected loss factor of 10 to 20% for quality filtering and demultiplexing. For example, if a run produces 100 million usable reads and each sample needs 5 million reads, you can pool up to 20 samples.
What causes index hopping and how can I prevent it?
Index hopping is caused by index sequences being transferred between library fragments during amplification or sequencing. The rate varies by platform. On DNB-based platforms, index misassignment occurs at a rate of one in 36 million reads, as shown in the reliable multiplex sequencing study. Prevention strategies include using unique dual indexes, minimizing amplification cycles, and following the manufacturer recommended procedures.
Should I use equal volume or equimolar pooling?
Use equal volume pooling when all libraries come from the same species or have similar genome sizes. Use equimolar pooling when pooling libraries from different species or samples with different genome sizes. The comparison of pooling strategies showed that equimolar pooling limits retesting risk due to insufficient coverage depth, particularly when interspecies genome size difference is more than 2-fold.
How does multiplexing affect the detection of low-frequency variants?
Multiplexing reduces the number of reads per sample, which reduces the sensitivity for detecting low-frequency variants. The detection limit is also affected by index misassignment. If the misassignment rate is 1%, variants present at less than 1% frequency cannot be reliably distinguished from artifacts. Laboratories performing sensitive assays must understand the misassignment rate of their platform.
What is the minimum coverage needed for clinical diagnostic sequencing?
The minimum coverage depends on the application and the variant types being detected. Clinical applications that require detection of low-frequency variants need higher coverage than applications that only confirm the presence of a known variant. The bioanalytical method validation guidance from the FDA emphasizes that methods must be validated for their intended use, which includes establishing appropriate coverage thresholds.
How do I know if my pooling strategy is working correctly?
Compare the actual read distribution across samples to the expected distribution. If one sample consistently receives more reads than expected, the quantification or pooling process needs adjustment. Monitor the quality metrics for each run and review them for systematic patterns. The laboratory quality management system handbook from the World Health Organization provides guidance on using quality data to improve laboratory processes.
Related Diagnostic Guides
- How to Calculate the Number of Molecules in a DNA Sample
- NanoDrop DNA Concentration: Formula and Calculation
- How to Calculate the Specific Activity of an Enzyme: Formula and Examples
- How to Calculate Transformation Efficiency: Formula, Examples, and Common Pitfalls
- How to Calculate the Number of Bacteria in a Sample Using ATP Bioluminescence
References and Further Reading
- Laboratory Quality Management System Handbook. World Health Organization.
- Laboratory Biosafety Manual. World Health Organization.
- Assay Guidance Manual. National Center for Advancing Translational Sciences.
- Bioanalytical Method Validation Guidance. U.S. Food and Drug Administration.
- NCBI Literature Resources. National Center for Biotechnology Information.
- Preimplantation Genetic Testing for Monogenic Disorders.. Genes, 2020.
- Digital Droplet PCR in Hematologic Malignancies: A New Useful Molecular Tool.. Diagnostics (Basel, Switzerland), 2022.
- Multiplex digital spatial profiling of proteins and RNA in fixed tissue.. Nature biotechnology, 2020.
- TILLING in extremis.. Plant biotechnology journal, 2012.
- Comprehensive Analysis of PKD1 and PKD2 by Long-Read Sequencing in Autosomal Dominant Polycystic Kidney Disease.. Clinical chemistry, 2024.
- In-depth comparison of library pooling strategies for multiplexing bacterial species in NGS.. Diagnostic microbiology and infectious disease, 2019.
- First NGS-based COVID-19 diagnostic.. Nature biotechnology, 2020.
- Understanding the Basics of NGS: From Mechanism to Variant Calling.. Current genetic medicine reports, 2015.
- Protocol for total RNA sequencing analysis of extracellular RNA from biofluids.. 2026.
- Barcode-integrated reverse transcription for accurate, complete, and low-input RNA sequencing. 2026.
- Scalable Agricultural Microbiome Sampling: Operational Definitions, Pooling Strategies, and Preservation Methods. 2026.
- Evaluating the Effect of Sampling Scale on Mosquito Virome Characterization Using PacBio HiFi Long-Read Metagenomics.. 2026.
- Implementation of NGS and SNP microarrays in routine forensic practice: opportunities and barriers. BMC Genomics, 2025.
- Applications and Performance of Precision ID GlobalFiler NGS STR, Identity, and Ancestry Panels in Forensic Genetics. Genes, 2024.
- FFPE-TLC: an NGS technology for accurate DNA-based gene fusion detection in bone and soft tissue tumors.. Journal of Molecular Diagnostics, 2023.
- Partitioning for easy multiplexing: a versatile droplet PCR application for clone monitoring in tumors.. Journal of Molecular Diagnostics, 2023.
- Optimized NGS Approach for Detection of Aneuploidies and Mosaicism in PGT-A and Imbalances in PGT-SR. Genes, 2020.
- Reliable multiplex sequencing with rare index mis-assignment on DNB-based NGS platform. BMC Genomics, 2018.
- Evaluation of the ABL NGS assay for HIV-1 drug resistance testing. Heliyon, 2023.
- Inverse-multiplexing in multi-layer optical grooming networks. 2004 IEEE Sarnoff Symposium on Advances in Wired and Wireless Communication, 2004.
This article is educational and does not replace validated laboratory procedures, institutional biosafety review, manufacturer instructions, or professional interpretation.