Illumina Sequencing Workflow: From Cluster Generation to Base Calling
Illumina sequencing is a short-read next-generation sequencing technology built on a workflow of library preparation, clonal cluster generation by in situ amplification on a flow cell, and sequencing by synthesis with reversible terminator nucleotides. The platform records fluorescent signals during each nucleotide incorporation cycle and converts those signals into base calls through real-time image analysis and quality scoring. This article explains the complete workflow from cluster generation to base calling, with practical guidance on run parameter optimization, quality assessment, troubleshooting, and documentation for laboratory students, technicians, researchers, and diagnostic professionals.
The NovaSeq 6000 system exemplifies the standard Illumina workflow, enabling short-read sequencing with output up to 6 Tb per run. The platform uses the typical Illumina sequencing workflow based on library preparation, cluster generation by in situ amplification, and sequencing by synthesis. Flexibility is a major feature of the NovaSeq 6000, with several types of sequencing kits coupled with dual flow cell mode enabling high scalability of sequencing outputs to match a wide range of applications from complete genome sequencing to metagenomics analysis. Understanding each stage of this workflow is essential for optimizing run parameters and producing reliable sequencing data.
At a Glance
The Illumina sequencing workflow proceeds through distinct stages, each with specific parameters that affect data quality and run success. The table below summarizes the key stages, their primary functions, and the critical parameters to monitor.
| Workflow Stage | Primary Function | Critical Parameters | Common Quality Issues |
|---|---|---|---|
| Library Preparation | Fragment DNA or RNA, ligate adapters, amplify fragments | Input quantity, fragment size distribution, adapter concentration, PCR cycle number | Adapter dimers, fragment bias, low yield |
| Cluster Generation | Amplify individual library fragments into clonal clusters on flow cell | Loading concentration, flow cell type, amplification cycles, template distribution | Cluster density too low or too high, polyclonal clusters, failed clusters |
| Sequencing by Synthesis | Incorporate labeled nucleotides and capture fluorescent signals | Cycle number, read length, reagent quality, flow cell temperature | Signal decay, phasing, pre-phasing, intensity overlap |
| Base Calling | Convert fluorescent images into nucleotide sequences with quality scores | Image analysis algorithm, quality threshold, calibration matrix | Low Q scores, ambiguous base calls, GC bias |
| Data Output | Generate FASTQ files for downstream analysis | Read length, paired-end status, demultiplexing index quality | Index hopping, low yield, adapter contamination |
Core Principles of Illumina Sequencing
Illumina sequencing relies on three interconnected technologies: clonal amplification, reversible terminator chemistry, and fluorescence imaging. Each principle addresses a specific challenge in high-throughput DNA sequencing.
Clonal Amplification and Signal Detection
The fundamental challenge in sequencing individual DNA molecules is that the fluorescent signal from a single molecule is too weak to detect reliably. Illumina solves this problem through cluster generation, where each library fragment is amplified into a clonal cluster of approximately one thousand identical copies. This amplification step occurs on the surface of a flow cell, a glass slide with oligonucleotide-coated lanes. The in situ amplification process creates discrete clusters that each produce a detectable fluorescent signal during sequencing.
The spatial separation of clusters on the flow cell is critical. Each cluster must be physically distinct so that the imaging system can assign fluorescent signals to specific clusters. Cluster density directly affects data quality. Low density wastes sequencing capacity, while excessive density causes clusters to overlap and merge, producing polyclonal clusters that generate mixed signals and unreliable base calls.
Reversible Terminator Chemistry
Sequencing by synthesis uses modified nucleotides that carry both a fluorescent label and a blocking group. The blocking group prevents more than one nucleotide from being incorporated per cycle. After each incorporation step, the flow cell is imaged to record which nucleotide was added to each cluster. A chemical cleavage step then removes both the fluorescent label and the blocking group, allowing the next nucleotide to be incorporated.
This cyclic approach enables synchronous sequencing across all clusters on the flow cell. The reversible terminator chemistry is what distinguishes Illumina sequencing from other sequencing technologies that use different incorporation and detection strategies. The accuracy of base calling depends on the efficiency of each incorporation and cleavage step, as incomplete cleavage leads to phasing, where clusters fall out of sync.
Fluorescence Imaging and Signal Processing
After each nucleotide incorporation cycle, the flow cell is excited with lasers and imaged. The fluorescent emission from each cluster is recorded as an image, and the intensity of each color channel indicates which nucleotide was incorporated. The imaging system captures multiple images per cycle, one for each nucleotide color channel.
The raw images are processed to identify cluster positions, measure fluorescence intensities, and convert those intensities into base calls. This processing involves several computational steps, including image registration, intensity extraction, and base calling with quality score assignment. The accuracy of this processing depends on the quality of the images and the calibration of the optical system.
Library Preparation and Its Impact on Sequencing
Library preparation is the first stage of the Illumina workflow and directly influences cluster generation and base calling quality. The library consists of DNA fragments flanked by adapter sequences that enable binding to the flow cell and provide priming sites for amplification and sequencing.
Fragment Size and Distribution
The size distribution of library fragments affects cluster generation efficiency and read length capabilities. For standard short-read sequencing, fragments typically range from 200 to 500 base pairs. Fragment size influences how efficiently fragments bind to the flow cell and how well they amplify during cluster generation. Very short fragments may not bind efficiently, while very long fragments may produce clusters with uneven amplification.
The fragment size distribution should be assessed before sequencing using a bioanalyzer or similar instrument. A narrow distribution around the target size is generally preferred, as it produces more consistent cluster generation and sequencing performance. Broad distributions can lead to variable cluster densities and reduced data quality.
Adapter Ligation and Indexing
Adapters serve multiple functions in the Illumina workflow. They provide the sequences that bind to the flow cell surface, the priming sites for cluster amplification, and the sequencing primer binding sites. Adapters also contain index sequences that allow multiple libraries to be pooled and sequenced in a single run, with the index read after the main sequencing reads to assign each read to its original sample.
Index quality is critical for multiplexed sequencing. Index sequences must be sufficiently different from each other to prevent misassignment. Index hopping, where index sequences are exchanged between library molecules, can occur during cluster amplification and leads to incorrect sample assignment. Careful index design and quality assessment of index reads are necessary to minimize this problem.
Library Quantification and Normalization
Accurate library quantification is essential for optimal cluster density. Libraries are quantified using quantitative PCR or fluorometric methods, and the concentration is used to calculate the appropriate loading volume. Overestimation of library concentration leads to excessive cluster density, while underestimation leads to underutilized flow cells.
Normalization of pooled libraries ensures that each sample contributes approximately equal read counts. The NovaSeq 6000 workflow includes instructions for assembling a normalized pool of libraries for sequencing, which is important for balancing coverage across samples in multiplexed runs. Uneven normalization can result in some samples receiving insufficient sequencing depth while others receive excess coverage.
Cluster Generation on the Flow Cell
Cluster generation is the process by which individual library fragments are amplified into clonal clusters on the flow cell surface. This stage is critical because the quality and density of clusters directly determine the amount and quality of sequencing data produced.
Flow Cell Surface and Oligonucleotide Coating
The flow cell is a glass slide with channels that contain a lawn of oligonucleotides attached to the surface. These oligonucleotides are complementary to the adapter sequences on the library fragments, enabling the fragments to hybridize to the flow cell. The flow cell surface chemistry is designed to support efficient hybridization and amplification while minimizing nonspecific binding.
Different Illumina platforms use different flow cell configurations. The NovaSeq 6000 uses a patterned flow cell with discrete nanowells that physically separate clusters, while older platforms use unpatterned flow cells where clusters form randomly on the surface. Patterned flow cells provide more uniform cluster spacing and can achieve higher cluster densities without the risk of cluster overlap.
Bridge Amplification
Cluster generation uses bridge amplification, where the immobilized DNA fragment forms a bridge by hybridizing its free end to an adjacent oligonucleotide on the flow cell surface. Polymerase extends the bridge to create a double-stranded molecule, and denaturation separates the strands. Each strand then forms a new bridge, and the process repeats to create a cluster of identical copies.
The number of amplification cycles determines the cluster size. More cycles produce larger clusters with stronger signals, but also increase the risk of amplification errors and signal overlap. The optimal cycle number depends on the flow cell type and the desired cluster density. Insufficient amplification produces weak signals that are difficult to call accurately, while excessive amplification can cause signal saturation.
Cluster Density Optimization
Cluster density is the most important parameter to control during cluster generation. Optimal density varies by platform and flow cell type but generally falls within a range specified by the manufacturer. Low density wastes sequencing capacity and increases cost per base, while high density causes clusters to merge and produce polyclonal clusters that generate unreliable data.
Several factors influence cluster density, including library loading concentration, flow cell type, and amplification conditions. The loading concentration is the primary adjustable parameter. A titration experiment using a control library can establish the optimal loading concentration for a specific library type and flow cell batch. Environmental factors such as temperature and humidity can also affect cluster generation and should be controlled.
Cluster Quality Assessment
After cluster generation, the flow cell is imaged to assess cluster density and quality. The sequencing instrument performs this assessment automatically and reports metrics such as cluster density, cluster passing filter rate, and the percentage of clusters that produce usable data. Clusters that pass the quality filter are those that produce clear, distinct signals during the first few sequencing cycles.
The passing filter rate is a key quality metric. A high passing filter rate indicates that most clusters are clonal and produce clean signals, while a low rate suggests problems with cluster generation, such as polyclonal clusters or insufficient amplification. The passing filter rate can be affected by library quality, loading concentration, and flow cell condition.
Sequencing by Synthesis
Sequencing by synthesis is the core biochemical process of the Illumina workflow. Each sequencing cycle involves nucleotide incorporation, imaging, and cleavage, and the efficiency of these steps determines the accuracy and length of the reads.
Nucleotide Incorporation Cycle
Each sequencing cycle begins with the introduction of a mixture of four labeled nucleotides, each with a different fluorescent label and a reversible terminator. The polymerase incorporates one nucleotide complementary to the template strand at each cluster. The terminator prevents further incorporation, ensuring that only one nucleotide is added per cycle.
The incorporation efficiency must be very high to maintain synchrony across the cluster. If some molecules in a cluster fail to incorporate a nucleotide, the cluster becomes phased, meaning different molecules are at different positions in the sequence. Phasing increases with each cycle and eventually degrades the quality of base calls beyond a usable threshold.
Imaging and Signal Detection
After incorporation, the flow cell is imaged to detect the fluorescent signal from each cluster. The imaging system uses lasers to excite the fluorescent labels and captures images through filters that separate the emission spectra of the four nucleotides. The intensity of each color channel at each cluster position indicates which nucleotide was incorporated.
The imaging system must accurately register the cluster positions across cycles. Any drift in the flow cell position or changes in the optical system can cause misalignment and incorrect intensity measurements. The instrument performs automatic focusing and registration to maintain image quality throughout the run.
Cleavage and Removal
After imaging, a chemical solution is introduced to cleave the fluorescent label and the terminator from the incorporated nucleotide. This step prepares the cluster for the next incorporation cycle. The cleavage must be complete to prevent residual fluorescence from interfering with the next cycle's signal detection.
Incomplete cleavage leads to signal carryover, where the previous cycle's fluorescence contaminates the current cycle's signal. This problem becomes more pronounced with longer reads and can limit the maximum read length. The cleavage chemistry is optimized to achieve high efficiency while minimizing damage to the DNA template and the polymerase.
Read Length and Cycle Number
The number of sequencing cycles determines the read length. Standard Illumina runs produce reads of 50 to 300 base pairs, depending on the platform and kit configuration. Longer reads require more cycles, and each additional cycle increases the risk of phasing, signal decay, and reagent depletion.
The choice of read length depends on the application. Short reads of 50 to 75 base pairs may be sufficient for some applications such as gene expression profiling, while longer reads of 150 to 300 base pairs are preferred for genome assembly, variant detection, and metagenomics. Paired-end sequencing reads both ends of each fragment, providing additional information for alignment and assembly.
Base Calling and Quality Scoring
Base calling is the computational process of converting fluorescent images into nucleotide sequences. This process involves several steps, from image processing to quality score assignment, and the accuracy of base calling directly affects the reliability of downstream analysis.
Image Processing and Intensity Extraction
The first step in base calling is processing the raw images to identify cluster positions and extract fluorescence intensities. The instrument software identifies clusters based on their signal patterns across the first few cycles and creates a template of cluster positions. This template is then used to extract intensities from each cycle's images.
Image processing must account for several sources of noise, including background fluorescence, optical crosstalk between color channels, and intensity variations across the flow cell. The software applies corrections for these effects to produce accurate intensity measurements for each cluster at each cycle.
Base Call Assignment
The extracted intensities are used to assign a base call to each cluster at each cycle. The software compares the observed intensities to a calibration matrix that describes the expected intensity pattern for each nucleotide. The nucleotide with the highest probability is assigned as the base call.
The calibration matrix is established during the first few cycles of the run using control sequences with known composition. This calibration accounts for the specific optical properties of the instrument and the fluorescent labels. Accurate calibration is essential for reliable base calling, and problems with calibration can lead to systematic errors.
Quality Score Calculation
Each base call is assigned a quality score, typically expressed as a Phred score. The Phred score is logarithmically related to the probability that the base call is incorrect. A Phred score of 20 corresponds to an error probability of 1 in 100, while a Phred score of 30 corresponds to an error probability of 1 in 1,000.
Quality scores are calculated based on the signal intensity, the separation between the observed signal and the expected signals for other nucleotides, and the historical accuracy of similar calls. Quality scores typically decrease with read length as phasing and signal decay accumulate. The quality score distribution is a key metric for assessing run quality and determining the appropriate quality filtering thresholds for downstream analysis.
Demultiplexing and Read Assignment
For multiplexed runs, an additional sequencing read is performed to determine the index sequence for each cluster. The index read is processed similarly to the main reads, and the resulting index sequence is used to assign each read to its sample of origin. Reads with low-quality index calls or ambiguous index sequences may be discarded or assigned to an undetermined category.
Demultiplexing quality depends on the accuracy of the index read and the design of the index sequences. Index sequences that are too similar can be confused, leading to incorrect sample assignment. The index read should be assessed for quality, and samples with poor index quality should be flagged for potential reassignment or re-sequencing.
Run Parameter Optimization
Optimizing run parameters is essential for producing high-quality sequencing data efficiently. The optimal parameters depend on the application, the library type, and the specific platform being used.
Read Length and Coverage Depth
The choice of read length and coverage depth should be guided by the requirements of the downstream analysis. Genome sequencing for variant detection typically requires higher coverage than transcriptome sequencing, and metagenomics applications may require different read lengths depending on the target organisms.
Coverage depth is determined by the number of clusters that pass filter and the read length. The required coverage depends on the application and the expected error rate. Higher coverage provides more confidence in variant calls but increases cost. The optimal coverage should be determined based on the specific requirements of the analysis and the expected error rates of the sequencing platform.
Paired-End Versus Single-End Sequencing
Paired-end sequencing reads both ends of each library fragment, providing two reads per cluster. Paired-end reads improve alignment accuracy, enable detection of structural variants, and facilitate de novo assembly. The choice between paired-end and single-end sequencing depends on the application and the available budget.
Paired-end sequencing requires additional sequencing cycles and therefore increases cost. However, the additional information provided by paired-end reads often justifies the cost for applications such as genome assembly, variant detection, and metagenomics. Single-end sequencing may be sufficient for applications such as gene expression quantification where alignment to a reference genome is straightforward.
Multiplexing and Pooling Strategy
Multiplexing allows multiple samples to be sequenced in a single run, reducing cost per sample. The number of samples that can be multiplexed depends on the required coverage per sample and the total output of the run. The pooling strategy should ensure that each sample receives sufficient coverage for the intended analysis.
The NovaSeq 6000 workflow includes instructions for assembling a normalized pool of libraries for sequencing, which is important for balancing coverage across samples. Normalization ensures that each sample contributes approximately equal numbers of clusters, preventing some samples from dominating the run while others receive insufficient coverage.
Quality Control Thresholds
Quality control thresholds determine which reads are retained for downstream analysis. Common thresholds include minimum quality scores, minimum read length, and adapter contamination filters. The appropriate thresholds depend on the application and the tolerance for errors.
For diagnostic applications, stringent quality thresholds are typically required to ensure reliable results. The Laboratory Quality Management System Handbook from the World Health Organization emphasizes the importance of quality control in laboratory testing, including the use of appropriate controls and quality metrics. Sequencing runs should include control samples with known expected results to verify the accuracy of the entire workflow.
Quality Assessment and Metrics
Assessing the quality of a sequencing run is essential for determining whether the data are suitable for downstream analysis. Several metrics provide information about run quality, and these metrics should be reviewed before proceeding with analysis.
Cluster Density and Passing Filter Rate
Cluster density and passing filter rate are the first metrics to review after a run completes. Cluster density should fall within the expected range for the platform and flow cell type. The passing filter rate indicates the proportion of clusters that produce usable data, and a low passing filter rate suggests problems with cluster generation or library quality.
The relationship between cluster density and passing filter rate is important. Very high cluster densities often result in lower passing filter rates because clusters overlap and produce mixed signals. The optimal cluster density balances total output against the proportion of usable clusters.
Quality Score Distribution
The distribution of quality scores across reads provides information about the accuracy of base calling. Quality scores typically decrease with read length, and the rate of decrease indicates the efficiency of the sequencing chemistry. A rapid decline in quality scores suggests problems with phasing, signal decay, or reagent quality.
The percentage of bases above quality thresholds such as Q30 is a commonly reported metric. High-quality runs typically have a high percentage of bases above Q30, while runs with chemistry problems show lower percentages. The Q30 percentage should be reviewed for each read position to identify any systematic patterns of quality degradation.
GC Bias and Coverage Uniformity
GC bias refers to the tendency of sequencing to over- or under-represent regions with extreme GC content. This bias can arise during library preparation, cluster generation, or sequencing and can affect the accuracy of downstream analyses such as copy number variation detection and metagenomics.
Coverage uniformity describes how evenly sequencing reads are distributed across the target regions. Uneven coverage can result from GC bias, amplification bias, or problems with library preparation. Coverage uniformity should be assessed using control samples or reference standards to identify any systematic biases.
Control Samples and Reference Standards
Control samples with known expected results are essential for validating sequencing runs. These controls can include reference DNA samples with known variants, spike-in controls with known concentrations, or negative controls to detect contamination. The results from control samples should be compared to expected values to verify the accuracy of the entire workflow.
The World Health Organization Laboratory Quality Management System Handbook provides guidance on the use of quality control materials in laboratory testing. For sequencing, controls should be included in each run to monitor performance over time and to detect any changes in the workflow that could affect results.
Common Failure Patterns and Troubleshooting
Understanding common failure patterns in Illumina sequencing can help identify problems early and take corrective action. The following sections describe frequent issues and their potential causes.
Low Cluster Density
Low cluster density results in reduced sequencing output and increased cost per base. Common causes include inaccurate library quantification, insufficient library loading, degraded library quality, and problems with the flow cell or reagents.
Troubleshooting low cluster density begins with verifying library quantification and quality. The library should be re-quantified using an independent method, and the fragment size distribution should be assessed. If the library appears intact, the loading concentration can be increased, and the run repeated. If the problem persists, the flow cell and reagent lots should be checked for quality issues.
Excessive Cluster Density and Polyclonal Clusters
Excessive cluster density leads to cluster overlap and the formation of polyclonal clusters that produce mixed signals. This problem results in a low passing filter rate and reduced usable data despite high total cluster counts.
The primary cause of excessive cluster density is overloading the flow cell with too much library. The loading concentration should be reduced, and the run repeated. If the problem persists, the library may contain contaminants that interfere with cluster generation, and additional purification may be necessary.
Phasing and Pre-Phasing
Phasing occurs when some molecules in a cluster fall behind in the sequencing cycles, while pre-phasing occurs when some molecules run ahead. Both problems cause mixed signals and reduced base calling accuracy, particularly at the ends of reads.
Phasing can result from incomplete cleavage of the terminator, inefficient nucleotide incorporation, or damage to the DNA template. Pre-phasing can result from incomplete blocking of the terminator or contamination with unblocked nucleotides. These problems are often related to reagent quality and can be addressed by using fresh reagents and verifying reagent storage conditions.
Low Quality Scores and Signal Decay
Low quality scores across the entire read suggest problems with the sequencing chemistry, the optical system, or the calibration. Signal decay with read length is normal but should not be excessive. Rapid signal decay suggests problems with reagent depletion, template damage, or inefficient cleavage.
Troubleshooting low quality scores begins with reviewing the quality metrics across read positions to identify the pattern of degradation. If quality scores are uniformly low, the calibration may be incorrect, and the run should be repeated with fresh reagents. If quality scores decline rapidly with read length, the sequencing chemistry may be compromised, and the reagent lot should be checked.
Index Hopping and Sample Cross-Contamination
Index hopping occurs when index sequences are exchanged between library molecules during cluster amplification, leading to reads being assigned to the wrong sample. This problem is more common with patterned flow cells and can be detected by including negative controls or by examining the distribution of index combinations.
Sample cross-contamination can also occur during library preparation or pooling. The World Health Organization Laboratory Biosafety Manual provides guidance on preventing contamination in laboratory workflows. Sequencing facilities should implement procedures to prevent cross-contamination, including the use of separate work areas for pre- and post-amplification steps and the use of appropriate controls.
Records and Documentation
Maintaining accurate records of sequencing runs is essential for quality assurance, troubleshooting, and regulatory compliance. The following records should be maintained for each sequencing run.
Run Metadata
Run metadata includes the date of the run, the instrument used, the flow cell lot number, the reagent lot numbers, the library preparation method, and the run parameters. This information is essential for troubleshooting problems and for comparing performance across runs.
The World Health Organization Laboratory Quality Management System Handbook emphasizes the importance of documentation in laboratory quality management. Sequencing facilities should maintain a run log that records all relevant metadata and any observations made during the run.
Quality Metrics and Control Results
Quality metrics for each run should be recorded, including cluster density, passing filter rate, quality score distributions, and the results from control samples. These metrics should be reviewed regularly to identify trends and to detect any changes in performance.
Control results should be compared to expected values, and any discrepancies should be investigated. The World Health Organization Laboratory Quality Management System Handbook recommends that laboratories establish acceptance criteria for quality control results and take corrective action when results fall outside these criteria.
Troubleshooting and Corrective Actions
Any problems encountered during a sequencing run should be documented, along with the troubleshooting steps taken and the outcome. This documentation provides a valuable reference for future troubleshooting and helps identify recurring problems.
The World Health Organization Laboratory Quality Management System Handbook recommends that laboratories establish procedures for corrective action when problems are identified. These procedures should include documentation of the problem, the investigation, the corrective action taken, and the verification that the corrective action was effective.
Biosafety and Laboratory Practices
Sequencing laboratories must follow appropriate biosafety practices to protect personnel and prevent contamination. The World Health Organization Laboratory Biosafety Manual provides guidance on biosafety practices for laboratories handling biological materials.
Sample Handling and Processing
Samples for sequencing may contain infectious agents, and appropriate precautions should be taken during sample handling and processing. The World Health Organization Laboratory Biosafety Manual recommends that laboratories conduct a risk assessment and implement appropriate biosafety measures based on the nature of the samples being processed.
Standard precautions include the use of personal protective equipment, the use of biosafety cabinets for procedures that may generate aerosols, and the proper disposal of biological waste. Laboratories should also have procedures for handling spills and for decontaminating work surfaces.
Preventing Contamination
Contamination is a major concern in sequencing laboratories because even small amounts of contaminating DNA can be amplified during library preparation and cluster generation. The World Health Organization Laboratory Biosafety Manual emphasizes the importance of good laboratory practices to prevent contamination.
Sequencing facilities should separate pre-amplification and post-amplification work areas, use dedicated equipment and reagents for each area, and implement procedures for cleaning and decontaminating work surfaces. Negative controls should be included in each run to detect contamination, and any contamination should be investigated and corrected.
Waste Management
Sequencing generates biological waste, chemical waste, and electronic waste. The World Health Organization Laboratory Biosafety Manual provides guidance on the safe handling and disposal of laboratory waste. Laboratories should have procedures for the proper disposal of all waste types and should comply with applicable regulations.
Professional Escalation Criteria
Knowing when to escalate problems to supervisors, manufacturers, or other experts is important for maintaining the quality and reliability of sequencing results. The following situations warrant escalation.
Persistent Quality Problems
If quality problems persist despite troubleshooting, the issue should be escalated to the instrument manufacturer or a senior technical expert. Persistent problems may indicate a hardware issue, a reagent lot problem, or a systematic issue with the laboratory workflow.
The manufacturer's technical support should be contacted when problems cannot be resolved through standard troubleshooting. The World Health Organization Laboratory Quality Management System Handbook recommends that laboratories have procedures for escalating problems that cannot be resolved internally.
Unexpected Control Results
If control samples produce unexpected results, the issue should be escalated to the laboratory supervisor or quality manager. Unexpected control results may indicate a problem with the entire workflow, and the validity of any results generated from the run should be questioned.
The World Health Organization Laboratory Quality Management System Handbook recommends that laboratories investigate all quality control failures and take corrective action before reporting patient or research results. The investigation should determine the cause of the failure and verify that the corrective action was effective.
Regulatory or Accreditation Concerns
If a sequencing run fails to meet regulatory or accreditation requirements, the issue should be escalated to the appropriate authority. This may include the laboratory director, the quality manager, or an external regulatory body.
The World Health Organization Laboratory Quality Management System Handbook provides guidance on the requirements for laboratory accreditation and the importance of meeting regulatory standards. Laboratories should have procedures for documenting and reporting any failures to meet these standards.
Frequently Asked Questions
What is the difference between cluster generation and sequencing by synthesis?
Cluster generation is the amplification step where individual library fragments are copied into clonal clusters on the flow cell surface. This step creates enough copies of each fragment to produce a detectable fluorescent signal. Sequencing by synthesis is the subsequent step where nucleotides are incorporated one at a time and detected by fluorescence imaging. Cluster generation prepares the template, and sequencing by synthesis reads the sequence.
How does Illumina sequencing achieve single-base resolution?
Illumina sequencing uses reversible terminator nucleotides that block further incorporation after each base. Each nucleotide carries a fluorescent label and a blocking group. After one nucleotide is incorporated, the flow cell is imaged, then the label and block are removed to allow the next incorporation. This cyclic process ensures that only one base is added per cycle, providing single-base resolution.
What causes phasing in Illumina sequencing?
Phasing occurs when some molecules in a cluster fail to incorporate a nucleotide or incorporate more than one nucleotide in a cycle. This causes the molecules to fall out of sync, producing mixed signals that reduce base calling accuracy. Phasing can result from incomplete cleavage, inefficient incorporation, or damage to the template. Phasing increases with read length and limits the maximum usable read length.
How is cluster density controlled in an Illumina run?
Cluster density is controlled primarily by the library loading concentration. Higher loading concentrations produce more clusters, while lower concentrations produce fewer clusters. The optimal concentration depends on the platform, flow cell type, and library characteristics. A titration experiment using a control library can establish the optimal loading concentration for a specific workflow.
What is a passing filter rate and why is it important?
The passing filter rate is the percentage of clusters that produce usable sequencing data. Clusters that pass the filter are those that generate clear, distinct signals during the initial sequencing cycles. A high passing filter rate indicates that most clusters are clonal and produce reliable data, while a low rate suggests problems with cluster generation or library quality.
How are quality scores calculated in Illumina sequencing?
Quality scores are calculated based on the signal intensity and the separation between the observed signal and the expected signals for other nucleotides. The scores are expressed as Phred scores, which are logarithmically related to the probability of an incorrect base call. Quality scores are used to filter low-quality reads and to assess the overall quality of a sequencing run.
What is index hopping and how can it be prevented?
Index hopping occurs when index sequences are exchanged between library molecules during cluster amplification, causing reads to be assigned to the wrong sample. This problem is more common with patterned flow cells. Index hopping can be detected by including negative controls and can be minimized by using unique dual indexes and by following recommended library preparation protocols.
When should a sequencing run be repeated?
A sequencing run should be repeated when the quality metrics fall below acceptable thresholds, when control samples produce unexpected results, or when the data are insufficient for the intended analysis. The decision to repeat a run should be based on the specific requirements of the application and the tolerance for errors. The World Health Organization Laboratory Quality Management System Handbook recommends that laboratories establish acceptance criteria and repeat runs that fail to meet these criteria.
Related Diagnostic Guides
- Next-Generation Sequencing for Veterinary Pathogen Metagenomics
- Metagenomic Next-Generation Sequencing (mNGS) for Veterinary Diagnostics: Challenges and Applications
- Metagenomic Next-Generation Sequencing (mNGS) for Diagnosis of Feline Infectious Peritonitis (FIP) in Clinical Effusions
- 2D Gel Electrophoresis: Principles and Workflow for Proteomics
- Ethanol Precipitation of DNA: Protocol and Troubleshooting
References and Further Reading
- Laboratory Quality Management System Handbook. World Health Organization.
- Laboratory Biosafety Manual. World Health Organization.
- Assay Guidance Manual. National Center for Advancing Translational Sciences.
- Bioanalytical Method Validation Guidance. U.S. Food and Drug Administration.
- NCBI Literature Resources. National Center for Biotechnology Information.
- The Illumina Sequencing Protocol and the NovaSeq 6000 System.. Methods in molecular biology (Clifton, N.J.), 2021.
- Current best practices in single-cell RNA-seq analysis: a tutorial.. Molecular systems biology, 2019.
- Critical review of 16S rRNA gene sequencing workflow in microbiome studies: From primer selection to advanced data analysis.. Molecular oral microbiology, 2023.
- BD Rhapsody™ Single-Cell Analysis System Workflow: From Sample to Multimodal Single-Cell Sequencing Data.. Methods in molecular biology (Clifton, N.J.), 2023.
- Short-Read RNA-Seq.. Methods in molecular biology (Clifton, N.J.), 2024.
- Developing a Nanopore Sequencing Workflow for Protein Engineering Applications.. ACS synthetic biology, 2023.
- Evaluation of Metagenomic and Targeted Next-Generation Sequencing Workflows for Detection of Respiratory Pathogens from Bronchoalveolar Lavage Fluid Specimens.. Journal of clinical microbiology, 2022.
- High-accuracy long-read amplicon sequences using unique molecular identifiers with Nanopore or PacBio sequencing.. Nature methods, 2021.
- RiboZAP: a species-agnostic pipeline for rRNA depletion probe design in metatranscriptomics.. 2026.
- Choosing Between Short-Read 16S, Full-Length ONT 16S, and Long-Read Shotgun Metagenomics for Soil Microbiome Studies: A Critical Review of the Benchmarking Evidence.. 2026.
- Same-day tagmentation PCR-based whole genome sequencing of bacteriophage genomes from a single plaque without DNA extraction.. 2026.
- Unraveling the Taxonomic Diversity and Functional Potential of the Tunisian Salterns, Abbassia and Thyna, via Integrated 16S-18S Amplicons and Shotgun Metagenomics.. 2026.
- Probing the limits of genetic recoding using multi-omics-guided evolution.. 2026.
- Adapting the Illumina COVIDSeq for Whole Genome Sequencing of Other Respiratory Viruses in Multiple Workflows and a Single Rapid Workflow. LabMed, 2025.
- Towards a Rapid-Turnaround Low-Depth Unbiased Metagenomics Sequencing Workflow on the Illumina Platforms. medRxiv, 2023.
- Investigating fungal diversity through metabarcoding for environmental samples: assessment of ITS1 and ITS2 Illumina sequencing using multiple defined mock communities with different classification methods and reference databases. BMC Genomics, 2025.
- Robust Mutation Profiling of SARS-CoV-2 Variants from Multiple Raw Illumina Sequencing Data with Cloud Workflow. Genes, 2022.
- Operationalizing Quality Assurance for Clinical Illumina Somatic Next-Generation Sequencing Pipelines. Journal of Molecular Diagnostics, 2024.
This article is educational and does not replace validated laboratory procedures, institutional biosafety review, manufacturer instructions, or professional interpretation.