Illumina Sequencing: Principle, Chemistry, and Workflow
Illumina sequencing is a next-generation sequencing (NGS) technology that uses sequencing-by-synthesis chemistry with reversible fluorescent terminators to determine DNA sequences at massive scale. The method dominates the global sequencing market due to its high throughput, accuracy, and relatively low cost per base. This article explains the core principles of Illumina sequencing, the bridge amplification process, reversible terminator chemistry, base calling, and the complete workflow from library preparation through data analysis. It also addresses practical considerations for handling low-diversity libraries and troubleshooting common sequencing problems. The content is written for laboratory students, technicians, researchers, and diagnostic professionals who need to understand both the theoretical foundations and the operational realities of Illumina platforms.
The Place of Illumina Sequencing in Modern Genomics
DNA sequencing has transformed molecular biology and clinical diagnostics. For decades, the Sanger method based on enzymatic DNA synthesis served as the gold standard for determining nucleotide sequences. The development of next-generation sequencing technologies at the end of the twentieth century shifted the field from analyzing single genes to sequencing entire genomes. Among the competing NGS technologies, one platform has practically completely dominated the global market, and that platform is Illumina. The transition from Sanger to high-throughput sequencing enabled dramatic expansion of clinical genetic testing for inherited conditions and diseases such as cancer. Accurate variant calling in NGS data is a critical step upon which virtually all downstream analysis and interpretation processes rely. Understanding how Illumina sequencing works at the level of chemistry and workflow is therefore essential for anyone who generates, analyzes, or interprets sequencing data in research or diagnostic settings.
The practical outcome of understanding Illumina sequencing is the ability to make informed decisions about library preparation, sequencing parameters, quality control, and troubleshooting. A laboratory that understands why low-diversity libraries fail, how read length affects downstream analysis, and what quality metrics matter will produce more reliable data and avoid costly reruns. This article provides that operational knowledge with attention to the limitations and failure modes that affect real sequencing runs.
At a Glance: Illumina Sequencing Overview
| Feature | Description | Practical Implication |
|---|---|---|
| Sequencing chemistry | Sequencing-by-synthesis with reversible fluorescent terminators | Each nucleotide is incorporated one at a time and detected by its fluorescent label |
| Template preparation | Bridge amplification on a solid surface creates clonal clusters | Each cluster produces a strong signal that represents one original DNA fragment |
| Read type | Single-end or paired-end reads | Paired-end reads improve alignment accuracy and enable detection of structural variants |
| Base calling | Fluorescent signal detection followed by computational conversion | Quality scores (Phred scores) indicate the probability of incorrect base calls |
| Low-diversity limitation | Requires balanced nucleotide representation at each cycle | Low-diversity libraries produce poor cluster identification and base calling |
| Primary failure modes | Cluster density issues, phasing, index hopping, adapter contamination | Monitoring metrics and following quality control protocols prevents most failures |
Core Principles of Sequencing-by-Synthesis
The Conceptual Foundation
Illumina sequencing relies on the principle of sequencing-by-synthesis, which means the DNA sequence is determined by detecting each nucleotide as it is incorporated into a growing complementary strand. This approach differs fundamentally from the Sanger method, which uses chain-terminating dideoxynucleotides to create a series of truncated fragments that are then separated by size. In Illumina chemistry, the synthesis reaction proceeds in a controlled cycle where only one nucleotide is added at a time, and the identity of that nucleotide is read from a fluorescent signal before the next cycle begins.
The key innovation that makes this possible is the reversible terminator. Each nucleotide carries a fluorescent label and a chemical block that prevents further extension after incorporation. After the fluorescent signal is detected, the terminator is cleaved and the fluorescent label is removed, allowing the next nucleotide to be added. This cycle of incorporation, detection, and deprotection repeats for hundreds of cycles to generate read lengths that are useful for genomic analysis.
The Role of DNA Polymerase
The sequencing reaction uses a DNA polymerase that incorporates nucleotides complementary to the template strand. The polymerase must accept the modified nucleotides with their fluorescent labels and reversible terminators, which means the enzyme has been engineered for this specific purpose. The fidelity of the polymerase and the efficiency of nucleotide incorporation directly affect the quality of the sequencing data. Errors in incorporation or incomplete removal of terminators contribute to phasing, which is one of the primary quality issues in Illumina sequencing.
Signal Detection and Base Calling
Each nucleotide is labeled with a distinct fluorescent dye. After incorporation, the flow cell is imaged using lasers that excite the fluorescent labels and cameras that capture the emitted light. The four nucleotides are distinguished by their emission spectra. The images are then processed computationally to determine which nucleotide was incorporated at each cluster position. This computational step is called base calling, and it produces both the nucleotide sequence and a quality score for each base.
The quality score, typically reported as a Phred score, reflects the probability that the base call is incorrect. Higher Phred scores indicate higher confidence. For example, a Phred score of 30 corresponds to an error probability of one in one thousand, while a Phred score of 20 corresponds to an error probability of one in one hundred. These quality scores are used in downstream analysis to filter low-confidence bases and to assess the overall quality of a sequencing run.
Bridge Amplification and Cluster Generation
Preparing the Flow Cell Surface
The flow cell is a glass slide with a lawn of oligonucleotides attached to its surface. These oligonucleotides are complementary to the adapter sequences that were added to the DNA fragments during library preparation. When a library fragment is loaded onto the flow cell, its adapter sequences hybridize to the surface-bound oligonucleotides, anchoring the fragment to the flow cell.
The Bridge Amplification Process
Bridge amplification is the process by which each individual DNA fragment is amplified into a clonal cluster. The term bridge refers to the structure that forms when the free end of an anchored fragment bends over and hybridizes to a complementary oligonucleotide on the surface, creating a bridge-like structure. The polymerase extends the fragment using the surface-bound oligonucleotide as a primer, creating a double-stranded bridge. The strands are then denatured, and each strand serves as a template for further amplification. This process repeats for many cycles, with each cycle doubling the number of copies of the original fragment.
The result is a cluster of identical DNA fragments, all derived from a single original molecule, located at a discrete position on the flow cell. Each cluster contains enough copies of the template to produce a detectable fluorescent signal during sequencing. The clonal nature of the clusters is essential for accurate base calling because the signal from each cluster represents the consensus of many identical templates.
Cluster Density and Its Consequences
Cluster density is a critical parameter that affects sequencing quality. If the cluster density is too low, the flow cell is underutilized and the sequencing run produces fewer reads than expected. If the cluster density is too high, clusters may overlap or merge, making it difficult to distinguish individual clusters during image analysis. Optimal cluster density varies by platform and chemistry version, and manufacturers provide recommended ranges. Monitoring cluster density during the run and adjusting loading concentrations for future runs is a standard quality control practice.
Reversible Terminator Chemistry
The Structure of Modified Nucleotides
The reversible terminator nucleotides used in Illumina sequencing have three key components. First, the nucleotide base itself, which provides the sequence information through complementary base pairing. Second, a fluorescent dye attached to the nucleotide, which provides the detectable signal. Third, a blocking group attached to the 3' hydroxyl position, which prevents further extension after incorporation.
The blocking group is the reversible part of the terminator. After the fluorescent signal is detected, a chemical cleavage step removes both the fluorescent dye and the blocking group, regenerating a free 3' hydroxyl that can accept the next nucleotide. This cleavage step must be complete and efficient because any nucleotides that retain their blocking groups will not be extended in subsequent cycles, contributing to phasing.
The Sequencing Cycle
Each sequencing cycle consists of four steps. First, a mixture of all four labeled nucleotides and polymerase is introduced to the flow cell. Second, the polymerase incorporates one nucleotide complementary to the template at each growing strand. Third, unincorporated nucleotides are washed away, and the flow cell is imaged to detect the fluorescent signals. Fourth, the cleavage step removes the fluorescent labels and blocking groups, preparing the strands for the next cycle.
The number of cycles determines the read length. For example, a 150-cycle run produces 150-base reads. The cycle efficiency, which is the proportion of strands that successfully complete each step, determines the quality of the data at increasing read lengths. As the number of cycles increases, the cumulative effect of incomplete extensions and cleavage failures becomes more pronounced, leading to decreased data quality at the ends of reads.
Phasing and Pre-Phasing
Phasing refers to the loss of synchrony within a cluster. Some strands in a cluster may fall behind because they failed to incorporate a nucleotide or failed to cleave the terminator, while other strands may run ahead because they incorporated two nucleotides in a single cycle. Both phasing and pre-phasing cause the fluorescent signal from a cluster to become a mixture of signals from different positions in the sequence, which degrades base calling accuracy.
Phasing increases with read length and is one of the primary factors that limits the maximum read length on Illumina platforms. Manufacturers provide phasing metrics for each run, and these metrics are used to assess run quality. Excessive phasing may indicate problems with the sequencing chemistry, the library preparation, or the flow cell.
The Complete Illumina Workflow
Step 1: Nucleic Acid Extraction and Quality Assessment
The workflow begins with the extraction of DNA or RNA from the sample. The quality and quantity of the extracted nucleic acid directly affect the success of downstream steps. For RNA samples, a reverse transcription step converts RNA to complementary DNA before library preparation. Quality assessment typically includes measuring concentration, assessing purity, and evaluating fragment size distribution.
The Laboratory Quality Management System Handbook from the World Health Organization emphasizes the importance of quality control throughout the testing process, including pre-analytical steps such as sample collection and nucleic acid extraction. Laboratories performing sequencing for diagnostic purposes should have documented procedures for sample handling, nucleic acid extraction, and quality assessment to ensure reliable results.
Step 2: Library Preparation
Library preparation converts the extracted nucleic acid into a format that is compatible with the sequencing platform. The key steps in library preparation are fragmentation, end repair, adapter ligation, and amplification.
Fragmentation breaks the nucleic acid into pieces of the desired size range. The fragment size affects the sequencing output and the downstream analysis. For example, fragment size influences the performance of HLA sequencing on the Illumina MiSeq platform, as reported in studies examining the effects of fragment size, indexing, and read length. End repair creates blunt ends on the fragments, and a single adenine is added to each end to facilitate adapter ligation.
Adapter ligation attaches short oligonucleotide sequences to both ends of each fragment. These adapters serve multiple purposes. They provide the sequences that hybridize to the flow cell surface during bridge amplification. They contain the sequencing primer binding sites. They also contain index sequences, which are short unique sequences that allow multiple libraries to be pooled and sequenced in a single run. The index sequences enable sample multiplexing, which reduces the cost per sample when many samples need to be sequenced.
Amplification increases the number of copies of each fragment to ensure sufficient material for cluster generation. The number of amplification cycles must be carefully controlled to avoid over-amplification, which can introduce bias and reduce library complexity.
Step 3: Library Quantification and Quality Control
Before sequencing, the library must be quantified and its quality assessed. Quantification determines the concentration of the library, which is used to calculate the appropriate loading volume for the flow cell. Quality assessment evaluates the fragment size distribution and checks for the presence of adapter dimers or other contaminants.
The Bioanalytical Method Validation Guidance from the U.S. Food and Drug Administration provides a framework for validating analytical methods, and the principles of method validation apply to library quantification and quality assessment. Laboratories should use validated methods for these steps and document the results to ensure reproducibility.
Step 4: Cluster Generation
The quantified library is denatured to produce single-stranded fragments and loaded onto the flow cell. The fragments hybridize to the surface-bound oligonucleotides, and bridge amplification generates clonal clusters. The cluster generation process is automated on most Illumina platforms, with the flow cell and reagents contained in a single cartridge.
Step 5: Sequencing
The flow cell is loaded into the sequencing instrument, and the sequencing cycles proceed automatically. The instrument performs the incorporation, imaging, and cleavage steps for each cycle. The images are processed in real time to generate base calls and quality scores.
The sequencing run parameters, including read length and the number of reads, are determined by the instrument and the reagents used. Different Illumina platforms have different capabilities. For example, the iSeq 100 and MiniSeq platforms are compatible with low-depth metagenomics identification and can be chosen based on required turnaround times, as demonstrated in a study of a rapid unbiased metagenomics sequencing workflow. The choice of platform depends on the application, the required throughput, and the available budget.
Step 6: Data Analysis
The sequencing instrument produces raw data in the form of base calls and quality scores. The raw data are then processed through a bioinformatics pipeline that includes alignment to a reference genome, variant calling, and interpretation. The complexity of the data analysis depends on the application. A clinician's guide to bioinformatics for next-generation sequencing describes the general principles of presequencing library preparation, postsequencing alignment, and variant calling, highlighting the need for sophisticated computational methods and bioinformatics expertise.
For clinical applications, variant calling is a critical step upon which virtually all downstream analysis and interpretation processes rely. Best practices for variant calling in clinical sequencing include careful attention to the strengths and weaknesses of panel, exome, and whole-genome sequencing for variant detection, as well as guidance on variant review, validation, and benchmarking to ensure optimal performance.
Library Preparation Options and Tradeoffs
Amplicon Sequencing
Amplicon sequencing uses PCR to amplify specific target regions before sequencing. This approach is used for targeted sequencing of genes or genomic regions of interest. The targeted next-generation sequencing approach involves enrichment of target pathogens in patient samples based on multiplex PCR amplification or probe capture with excellent sensitivity. Amplicon sequencing is practical and efficient for clinical diagnostics, with high positivity rates in the detection of pathogens in respiratory samples and high sensitivity in samples with low pathogen loads, including blood and cerebrospinal fluid.
The main advantage of amplicon sequencing is its simplicity and cost-effectiveness. The main limitation is that it only examines the targeted regions, so it cannot detect unexpected variants outside those regions.
Shotgun Sequencing
Shotgun sequencing sequences all the DNA in a sample without targeting specific regions. This approach is used for whole-genome sequencing, metagenomics, and other applications where unbiased coverage is needed. Metagenomic next-generation sequencing is a transformative approach in the diagnosis of infectious diseases, utilizing unbiased high-throughput sequencing to directly detect and characterize microbial genomes from clinical samples. Shotgun sequencing enables simultaneous detection of bacteria, viruses, fungi, and parasites without prior knowledge of the infectious agent.
The main advantage of shotgun sequencing is its comprehensiveness. The main limitations are the higher cost and the greater complexity of data analysis compared to amplicon sequencing.
Metagenomic Sequencing
Metagenomic sequencing is a specific application of shotgun sequencing that analyzes the genetic material from a complex microbial community. This approach can identify rare, novel, or unculturable pathogens, providing a more comprehensive view of microbial communities compared to traditional culture-based methods. However, challenges such as data analysis complexity, high cost, and the need for optimized sample preparation protocols remain significant hurdles.
The choice between amplicon and shotgun approaches depends on the research question or clinical need. For example, 16S rRNA amplicon sequencing on Illumina platforms is a common approach for studying microbial communities, but the choice of sequencing method has a strong effect on which species are detected and how the community is described. Short-read 16S and long-read approaches generally converge on dominant taxa and on between-sample differences, but they disagree substantially on alpha diversity estimates, rare taxon detection, and the relative abundances of entire phyla.
Handling Low-Diversity Libraries
Why Low-Diversity Libraries Fail
Illumina sequencing requires a balanced representation of all four nucleotides at each sequencing cycle. This requirement exists because the base calling algorithm uses the signals from all clusters on the flow cell to calibrate the fluorescent intensity measurements. If most clusters incorporate the same nucleotide in a given cycle, the calibration is biased, and the base calls for the minority nucleotides become unreliable.
Low-diversity libraries are libraries where the nucleotide composition is not balanced. Examples include amplicon libraries where all fragments start with the same sequence, libraries from organisms with extreme GC content, and libraries prepared from samples with a dominant species. The failure mode is particularly severe in the early cycles, where the lack of diversity prevents accurate cluster identification and base calling.
Strategies for Handling Low-Diversity Libraries
Several strategies can mitigate the problems associated with low-diversity libraries. The most common approach is to add a small amount of a control library, typically PhiX, to the sequencing run. The PhiX library provides a balanced nucleotide composition that helps calibrate the base calling. The proportion of PhiX needed depends on the degree of diversity in the sample library, with higher proportions needed for more extreme cases.
Another approach is to use custom sequencing primers that shift the starting position of the read. This strategy can help when the low diversity is confined to the beginning of the read. Additionally, some library preparation methods incorporate diversity-enhancing sequences during the preparation steps.
The Assay Guidance Manual from the National Center for Advancing Translational Sciences provides general guidance on assay development and validation, and the principles of assay design apply to the development of sequencing assays. When developing a sequencing assay for a low-diversity sample type, the assay should be designed with the diversity requirement in mind.
Practical Considerations
When planning a sequencing run with low-diversity libraries, several practical considerations apply. First, the library preparation should be designed to minimize the impact of low diversity, for example by using indexing strategies that introduce sequence diversity. Second, the sequencing run should include an appropriate amount of PhiX control. Third, the sequencing parameters should be adjusted if the platform allows, for example by using custom sequencing primers.
The quality of the sequencing data should be monitored throughout the run. If the data quality is poor, the run may need to be repeated with adjusted parameters. Documentation of the run parameters and the quality metrics is essential for troubleshooting and for demonstrating the reliability of the sequencing results.
Quality Control and Quality Metrics
Key Quality Metrics
Several quality metrics are used to assess the performance of an Illumina sequencing run. The cluster density indicates how many clusters were generated on the flow cell. The percentage of clusters passing filter indicates the proportion of clusters that produced usable data. The Phred quality scores indicate the confidence in each base call. The percentage of bases above a quality threshold, such as Q30, provides an overall measure of data quality.
The coverage depth indicates how many times each position in the target region was sequenced. Adequate coverage depth is essential for reliable variant calling. The uniformity of coverage indicates how evenly the reads are distributed across the target region. Poor uniformity can result in some regions having insufficient coverage for reliable variant detection.
Quality Control Procedures
Quality control procedures should be implemented at each step of the sequencing workflow. Before sequencing, the library should be quantified and its quality assessed. During sequencing, the run metrics should be monitored to detect problems early. After sequencing, the data quality should be assessed before proceeding with downstream analysis.
The Laboratory Quality Management System Handbook from the World Health Organization provides a framework for implementing quality management in laboratories, including procedures for quality control, documentation, and continuous improvement. Laboratories performing sequencing for diagnostic purposes should have a documented quality management system that covers all aspects of the testing process.
Records and Documentation
Accurate records are essential for quality assurance and for troubleshooting. The records should include information about the sample, the library preparation, the sequencing run parameters, and the quality metrics. For diagnostic applications, the records should be sufficient to demonstrate that the testing process was performed correctly and that the results are reliable.
The Bioanalytical Method Validation Guidance from the U.S. Food and Drug Administration emphasizes the importance of documentation in method validation and sample analysis. The principles of good documentation practice apply to sequencing workflows, including the use of standard operating procedures, the recording of deviations, and the maintenance of records for a defined period.
Common Failure Patterns and Troubleshooting
Low Cluster Density
Low cluster density results in fewer reads than expected and may indicate problems with library quantification, library loading, or the flow cell. The library concentration should be verified, and the loading volume should be adjusted. If the problem persists, the library quality should be assessed, and the quantification method should be validated.
High Cluster Density
High cluster density results in overlapping clusters and poor data quality. The library loading concentration should be reduced for subsequent runs. The cluster density should be monitored during the run to detect problems early.
Poor Base Calling Quality
Poor base calling quality, indicated by low Phred scores, may result from phasing, low diversity, or problems with the sequencing chemistry. The phasing metrics should be reviewed, and the library diversity should be assessed. If low diversity is the cause, PhiX control should be added to the run.
Adapter Contamination
Adapter contamination occurs when the sequencing read extends into the adapter sequence instead of the insert sequence. This problem is common when the insert size is shorter than the read length. Adapter trimming should be performed during data analysis, and the library preparation should be optimized to produce longer inserts.
Index Hopping
Index hopping occurs when index sequences are exchanged between samples during the sequencing run, resulting in reads being assigned to the wrong sample. This problem is more common on patterned flow cells. The use of unique dual indexes can help detect and mitigate index hopping.
Addressing Challenges in Sequencing Data Production
A study addressing challenges in the production and analysis of Illumina sequencing data highlights the importance of understanding the sources of error and bias in sequencing workflows. The challenges include library preparation artifacts, sequencing errors, and analysis pipeline issues. Addressing these challenges requires a systematic approach that includes quality control at each step, careful validation of the analysis pipeline, and documentation of the entire process.
Limitations of Illumina Sequencing
Read Length Limitations
Illumina sequencing produces short reads compared to some other sequencing technologies. The maximum read length is typically a few hundred bases, which limits the ability to resolve repetitive regions and structural variants. Long-read sequencing technologies can produce reads of many kilobases, but they have different limitations, including lower accuracy and higher cost per base.
GC Bias
Illumina sequencing can exhibit bias in regions with extreme GC content. Regions with very high or very low GC content may be underrepresented in the sequencing data, which can affect the detection of variants in those regions. The GC bias can be mitigated to some extent by library preparation methods and by using appropriate bioinformatics tools.
Amplification Bias
The amplification steps in library preparation and cluster generation can introduce bias. Some sequences amplify more efficiently than others, leading to uneven coverage. This bias can affect the accuracy of variant calling and the quantification of gene expression.
Error Rates
Although Illumina sequencing has high accuracy, errors still occur. The error rate is higher in certain contexts, such as homopolymer regions and the ends of reads. The error rate should be considered when interpreting sequencing results, particularly for clinical applications where the consequences of errors can be significant.
Comparison with Other Platforms
Comparative studies have examined the performance of different sequencing platforms. For example, a comparative analysis of the MGISEQ-2000 sequencing platform versus the Illumina HiSeq 2500 for whole-genome sequencing found that the platforms have comparable performance for many applications. The choice of platform should be based on the specific requirements of the application, including throughput, read length, accuracy, and cost.
Safety and Regulatory Context
Laboratory Safety
Sequencing laboratories should follow established biosafety guidelines. The Laboratory Biosafety Manual from the World Health Organization provides guidance on the safe handling of biological materials, including the use of appropriate containment measures and personal protective equipment. The biosafety level required depends on the nature of the samples being processed.
Regulatory Considerations for Diagnostic Applications
Sequencing assays used for diagnostic purposes are subject to regulatory oversight. The Bioanalytical Method Validation Guidance from the U.S. Food and Drug Administration provides a framework for validating analytical methods used in clinical studies. Laboratories performing diagnostic sequencing should validate their assays according to applicable regulations and guidelines.
The transition of sequencing technology from research to clinical settings has been driven by technological maturity and cost reductions. The growing need for cost-effective and universally accessible sequencing assays to improve patient care and public health has led to the increasing use of targeted sequencing approaches in clinical diagnostics. However, further research to overcome technical challenges such as workflow time and cost is required.
Professional Escalation Criteria
Laboratory personnel should know when to escalate problems to supervisors or to the instrument manufacturer. Escalation is appropriate when quality metrics are consistently outside acceptable ranges, when troubleshooting steps do not resolve the problem, or when the sequencing results are inconsistent with expected results. For diagnostic applications, any indication of a potential quality failure should be escalated promptly to ensure that patient results are not compromised.
Records and Measurements for Quality Assurance
What to Record
The following information should be recorded for each sequencing run. The sample identifiers and the library preparation details, including the library preparation method, the indexing strategy, and the quantification results. The sequencing run parameters, including the platform, the chemistry version, the read length, and the number of cycles. The quality metrics, including the cluster density, the percentage of clusters passing filter, the Phred quality scores, and the percentage of bases above Q30.
How to Use the Records
The records should be reviewed regularly to identify trends and to detect problems early. For example, a gradual decrease in cluster density over multiple runs may indicate a problem with the library quantification method or the flow cell handling. A sudden decrease in data quality may indicate a problem with the sequencing chemistry or the instrument.
The records should also be used to demonstrate the reliability of the sequencing results. For diagnostic applications, the records should be sufficient to support the interpretation of the results and to defend the results if they are questioned.
Common Failure Patterns in Diagnostic Sequencing
Sample Contamination
Sample contamination can occur during sample collection, nucleic acid extraction, or library preparation. Contamination can result in the detection of variants that are not present in the sample or the failure to detect variants that are present. The use of negative controls and the monitoring of contamination levels are essential quality control measures.
Insufficient Coverage
Insufficient coverage can result in the failure to detect variants, particularly in regions with poor sequencing efficiency. The coverage depth should be assessed for each sample, and samples with insufficient coverage should be identified for repeat testing or for alternative analysis approaches.
Batch Effects
Batch effects are systematic differences between sequencing runs that are not related to the biological differences between samples. Batch effects can confound the comparison of samples from different runs. The use of appropriate experimental design and statistical methods can help mitigate batch effects.
Variant Interpretation Challenges
The interpretation of genetic variants is a complex process that requires the integration of multiple lines of evidence. A systematic approach to assessing the clinical significance of genetic variants involves searching for existing data in publications and databases, performing evidence-based assessments through statistical analyses of observations in the general population and disease cohorts, evaluating experimental data from in vivo or in vitro studies, and using computational predictions of potential impacts of each variant. The caveats and pitfalls of variant assessment should be understood by all personnel involved in the interpretation of sequencing results.
Frequently Asked Questions
What is the difference between Illumina sequencing and Sanger sequencing?
Sanger sequencing uses chain-terminating dideoxynucleotides to create truncated fragments that are separated by size, producing one sequence at a time. Illumina sequencing uses sequencing-by-synthesis with reversible fluorescent terminators to sequence millions of fragments simultaneously. Illumina sequencing provides much higher throughput at a lower cost per base, but Sanger sequencing remains useful for targeted applications and for validating specific variants.
How does bridge amplification work?
Bridge amplification starts with DNA fragments anchored to the flow cell surface through adapter sequences. The free end of each fragment bends over and hybridizes to a complementary surface-bound oligonucleotide, forming a bridge structure. A polymerase extends the fragment, creating a double-stranded bridge. The strands are denatured, and each strand serves as a template for further amplification. Repeated cycles of extension and denaturation produce a clonal cluster of identical fragments at each starting position.
What are reversible terminators and why are they important?
Reversible terminators are modified nucleotides that carry a fluorescent label and a blocking group. The blocking group prevents more than one nucleotide from being added per cycle, and the fluorescent label allows the identity of the incorporated nucleotide to be detected. After detection, the label and the blocking group are removed, allowing the next nucleotide to be added. This cycle of incorporation, detection, and deprotection is the basis of sequencing-by-synthesis.
Why do low-diversity libraries fail on Illumina platforms?
Low-diversity libraries fail because the base calling algorithm requires balanced nucleotide representation at each cycle to calibrate the fluorescent intensity measurements. When most clusters incorporate the same nucleotide, the calibration is biased, and the base calls for minority nucleotides become unreliable. Adding PhiX control library or using custom sequencing primers can mitigate the problem.
What is phasing and how does it affect sequencing quality?
Phasing is the loss of synchrony within a cluster, where some strands fall behind or run ahead of the majority. Phasing causes the fluorescent signal from a cluster to become a mixture of signals from different positions, degrading base calling accuracy. Phasing increases with read length and is a primary factor limiting maximum read length.
What quality metrics should be monitored during a sequencing run?
The key quality metrics are cluster density, the percentage of clusters passing filter, Phred quality scores, the percentage of bases above Q30, and coverage depth. These metrics should be monitored during the run and reviewed after the run to assess data quality and to identify problems.
How should adapter contamination be handled?
Adapter contamination occurs when the read extends into the adapter sequence. Adapter trimming should be performed during data analysis to remove adapter sequences from the reads. The library preparation should also be optimized to produce inserts that are longer than the read length.
What are the main limitations of Illumina sequencing?
The main limitations are short read lengths, GC bias, amplification bias, and error rates. Short reads limit the resolution of repetitive regions and structural variants. GC bias affects the representation of regions with extreme GC content. Amplification bias causes uneven coverage. Errors occur at a low rate but can affect variant calling, particularly in homopolymer regions and at the ends of reads.
Related Diagnostic Guides
- qPCR Amplification Curve Troubleshooting: Common Shape Abnormalities and Fixes
- PCR Troubleshooting: No Amplification or Weak Bands
- RT-PCR Troubleshooting: No Amplification or Multiple Bands
- Bacterial Transformation Troubleshooting: Low Efficiency and No Colonies
- BCA Assay Troubleshooting: Color Development and Compatibility Issues
References and Further Reading
- Laboratory Quality Management System Handbook. World Health Organization.
- Laboratory Biosafety Manual. World Health Organization.
- Assay Guidance Manual. National Center for Advancing Translational Sciences.
- Bioanalytical Method Validation Guidance. U.S. Food and Drug Administration.
- NCBI Literature Resources. National Center for Biotechnology Information.
- [From Sanger to genome sequencing - an overview of DNA sequencing technologies].. Postepy biochemii, 2024.
- Best practices for variant calling in clinical sequencing.. Genome medicine, 2020.
- Single-molecule regulatory architectures captured by chromatin fiber sequencing.. Science (New York, N.Y.), 2020.
- Clinical diagnostic value of targeted next-generation sequencing for infectious diseases (Review).. Molecular medicine reports, 2024.
- A Clinician's Guide to Bioinformatics for Next-Generation Sequencing.. Journal of thoracic oncology : official publication of the International Association for the Study of Lung Cancer, 2023.
- A systematic approach to assessing the clinical significance of genetic variants.. Clinical genetics, 2013.
- Application of metagenomic next-generation sequencing in the diagnosis of infectious diseases.. Frontiers in cellular and infection microbiology, 2024.
- DNA methylation methods: Global DNA methylation and methylomic analyses.. Methods (San Diego, Calif.), 2021.
- RiboZAP: a species-agnostic pipeline for rRNA depletion probe design in metatranscriptomics.. 2026.
- Choosing Between Short-Read 16S, Full-Length ONT 16S, and Long-Read Shotgun Metagenomics for Soil Microbiome Studies: A Critical Review of the Benchmarking Evidence.. 2026.
- Same-day tagmentation PCR-based whole genome sequencing of bacteriophage genomes from a single plaque without DNA extraction.. 2026.
- Unraveling the Taxonomic Diversity and Functional Potential of the Tunisian Salterns, Abbassia and Thyna, via Integrated 16S-18S Amplicons and Shotgun Metagenomics.. 2026.
- Probing the limits of genetic recoding using multi-omics-guided evolution.. 2026.
- Adapting the Illumina COVIDSeq for Whole Genome Sequencing of Other Respiratory Viruses in Multiple Workflows and a Single Rapid Workflow. LabMed, 2025.
- Towards a Rapid-Turnaround Low-Depth Unbiased Metagenomics Sequencing Workflow on the Illumina Platforms. medRxiv, 2023.
- Investigating fungal diversity through metabarcoding for environmental samples: assessment of ITS1 and ITS2 Illumina sequencing using multiple defined mock communities with different classification methods and reference databases. BMC Genomics, 2025.
- Genome sequencing of steroid-producing bacteria with Illumina technology. Methods in Molecular Biology, 2017.
- Report on the effects of fragment size, indexing, and read length on HLA sequencing on the Illumina MiSeq. Human Immunology, 2015.
- Comparative analysis of novel MGISEQ-2000 sequencing platform vs Illumina HiSeq 2500 for whole-genome sequencing. Plos One, 2020.
- Species Identification and Profiling of Complex Microbial Communities Using Shotgun Illumina Sequencing of 16S rRNA Amplicon Sequences. Plos One, 2013.
- Addressing challenges in the production and analysis of illumina sequencing data. BMC Genomics, 2011.
This article is educational and does not replace validated laboratory procedures, institutional biosafety review, manufacturer instructions, or professional interpretation.