Zubair Khalid

Virologist/Molecular Biologist | Veterinarian | Bioinformatician

Conventional & Molecular Virology • Vaccine Development • Computational Biology

Dr. Zubair Khalid is a veterinarian and virologist specializing in conventional and molecular virology, vaccine development, and computational biology. Dedicated to advancing animal health through innovative research and multi-omics approaches.

Dr. Zubair Khalid - Veterinarian, Virologist, and Vaccine Development Researcher specializing in Computational Biology, Multi-omics, Animal Health, and Infectious Disease Research

Category: Guides

Sanger Sequencing Vs Illumina Sequencing

If you need to sequence a single gene or a small set of DNA fragments with the highest per base accuracy, Sanger sequencing remains the gold standard. If you need to sequence entire genomes, transcriptomes, or many samples in parallel, Illumina sequencing provides the throughput and cost efficiency that Sanger cannot match. This guide is for laboratory researchers, bioinformaticians, and clinicians who must choose between these two core sequencing technologies and want a practical, evidence based framework for making that decision. We will cover the strengths and limits of each method, walk through the steps of a typical workflow, and highlight the common pitfalls that can undermine your results. Your choice depends on your specific research question, sample number, read length requirements, budget, and tolerance for minor error rates. For a structured foundation in sequencing principles, start with the EMBL EBI Training resources, which offer excellent modular lessons on both technologies.

Sanger sequencing uses chain terminating dideoxynucleotides to produce a nested set of labeled fragments that are separated by capillary electrophoresis. Illumina sequencing, by contrast, uses reversible terminator chemistry to sequence millions of DNA clusters in parallel on a flow cell. While both methods are based on DNA polymerase driven synthesis, they serve entirely different niches. The Galaxy Training Network provides free workflow tutorials for both Sanger trace analysis and Illumina read processing, giving you hands on practice before committing resources.

At a Glance

Feature Sanger Sequencing Illumina Sequencing
Throughput Low: 1 96 samples per run, single fragment per reaction High: millions to billions of reads per run
Read length Up to 1000 bases (typically 600 900 high quality) 50 300 bases (longer reads with dedicated kits, but standard is 150 bp paired end)
Accuracy per base 99.999% (Q40+) 99.9% (Q30) for consensus, lower for individual reads
Cost per base High: $1000+ for a full genome equivalent Very low: fractions of a cent per base when multiplexed
Best use case Single gene confirmation, small amplicons, mutation validation Whole genome, exome, transcriptome, ChIP seq, metagenomics
Turnaround time 2 4 hours for capillary run (plus PCR and cleanup) 1 3 days per run (including library prep and sequencing)

Core Concepts

Sanger sequencing relies on the incorporation of fluorescently labeled dideoxynucleotides (ddNTPs) that terminate chain elongation. The reaction produces a ladder of fragments, each ending at a specific base. Capillary electrophoresis resolves these fragments by size, and a laser excites the fluorophores to generate a chromatogram trace. The NCBI Bookshelf contains a detailed chapter on the history and biochemistry of Sanger sequencing, which remains the standard for validating variants detected by other methods. Its accuracy stems from the ability to read each base position multiple times from different fragment lengths, though the technique cannot reliably detect variants present below 15 20% frequency in a mixed sample.

Illumina sequencing uses a flow cell covered with oligonucleotides that capture adaptor ligated DNA fragments. Bridge amplification creates clonal clusters of identical templates. Sequencing by synthesis then adds fluorescently labeled reversible terminators one base at a time. After each incorporation, a camera images the entire flow cell, and the terminator is cleaved to allow the next base. This parallel process generates millions of reads simultaneously. The EMBL EBI Training module on NGS explains how Illumina’s base calling and quality scoring work, including the ubiquitous Phred quality score (Q score). Each base receives a Q score that predicts its error probability. A Q30 score corresponds to a 0.1% error rate, or one error per 1000 bases. Because individual reads have higher error rates than Sanger, confidence relies on depth of coverage typically 30x for human genome analysis.

Decision Criteria for Choosing a Method

Selecting between Sanger and Illumina depends on five primary factors: sample number, required read length, variant frequency, budget, and turnaround time.

Sample number and throughput. For fewer than 20 samples, each targeting a single amplicon of 500 900 bases, Sanger is often faster and cheaper because you avoid the cost and complexity of Illumina library preparation. For hundreds of samples or whole genome targets, Illumina’s multiplexing capability reduces per sample cost dramatically. A study comparing 16S rRNA microbiota analysis showed that Next generation sequencing (Illumina) detected far greater taxonomic diversity than Sanger due to the depth of sampling [9].

Read length requirements. Sanger excels at long reads, routinely achieving 800 bases. Illumina standard reads are 150 bases, which complicates assembly of repetitive regions or full length transcripts. For applications like full length HIV env sequencing or resolving haplotype structures, Sanger or long read alternatives may be necessary. However, for resequencing projects or variant detection in well characterized genomes, Illumina short reads are sufficient and cost effective. In HIV surveillance, Illumina based haplotype enhanced methods improved the resolution of transmission networks compared to Sanger sequencing of individual clones [6].

Variant frequency detection. Sanger sequencing cannot reliably detect variants present in less than about 15% of the DNA population in a sample. For somatic mosaicism or low frequency drug resistance mutations, Illumina’s deep coverage can detect variants down to 1% frequency or lower. A study of NLRP3 mosaicism revealed that Illumina deep sequencing identified low level variants that Sanger missed entirely [8].

Budget and turnaround. A single Sanger reaction costs $3 10 (USD) but scales linearly. An Illumina MiSeq run may cost $1000 for reagents alone but yields millions of reads. For urgent clinical confirmation of a known mutation, Sanger can return results in a few hours. For large scale discovery projects, Illumina requires 2 3 days but provides orders of magnitude more data.

Orthogonal validation. In clinical settings, many laboratories use Illumina for initial screening and then confirm positive findings with Sanger sequencing. For example, certain pMMR colorectal cancer patients should undergo additional MSI PCR testing to reduce misdiagnosis risk, illustrating that no single method is infallible [11].

Practical Workflow or Implementation Sequence

The sequence of steps differs substantially between the two methods. Below is a generalized workflow for each, starting from template preparation through data analysis.

Sanger Sequencing Workflow

  1. Template preparation. Extract DNA or RNA (reverse transcribe to cDNA). Design primers to amplify a region typically 500 900 bases. PCR amplify the target and verify product size by gel electrophoresis.
  2. Cycle sequencing. Combine the purified PCR product with a sequencing primer, DNA polymerase, dNTPs, and fluorescent ddNTPs. Run 25 30 cycles of linear amplification to generate terminated fragments of every length.
  3. Cleanup and capillary electrophoresis. Remove excess dyes and salts. Denature the samples and load onto the sequencer. Capillary electrophoresis separates fragments by size. A laser excites the terminal dye and records a four color trace.
  4. Base calling and quality assessment. Software converts the trace to a sequence using Phred base calling. Examine the chromatogram for double peaks, high background, or failed reactions. The Galaxy Training Network offers a tutorial for Sanger trace quality control. Acceptable quality requires Phred scores above 30 for most bases and clear spacing between peaks.
  5. Variant calling. Align the consensus sequence to a reference using BLAST or alignment tools. Manually inspect suspicious positions. Sanger is ideal for confirming single nucleotide variants or small indels from other methods.

Illumina Sequencing Workflow

  1. Library preparation. Fragment DNA (or use amplicons) to 200 600 bp. End repair, A tail, and ligate adaptors. Size select and PCR amplify the library to add indexing barcodes. Quantify the library using qPCR or fluorometry.
  2. Cluster generation. Load the library onto a flow cell. Bridge amplification creates clonal clusters of each fragment. This step takes several hours on the instrument.
  3. Sequencing by synthesis. The instrument cycles through addition of four fluorescently labeled reversible terminators, imaging, and cleavage. Each cycle adds one base. Dual indexing reads are performed if multiplexing.
  4. Base calling and quality filtering. The instrument’s software performs real time base calling and assigns Phred quality scores. The NCBI Sequence Read Archive accepts these raw reads (FASTQ files) with associated quality data [5].
  5. Primary analysis. Demultiplex the reads by index. Trim adaptors and low quality bases. Align to a reference genome using a short read aligner like BWA or Bowtie. Downstream analysis includes variant calling, quantification, or assembly. The Bioconductor platform provides thousands of R packages for differential expression, variant filtering, and visualization [4].

Quality Checks and Common Mistakes

Quality Checks

For Sanger sequencing, always inspect the raw chromatogram. Look for sharp, evenly spaced peaks without excessive background noise. A Phred score below 20 indicates unreliable base calls. If you see double peaks at a single position, it may indicate a heterozygous variant, but it could also be due to primer mismatch or contamination. Re run with a different primer pair if ambiguous.

For Illumina sequencing, monitor the run metrics. The most important quality metric is the percentage of bases above Q30. For a 150 bp read, expect >80% Q30. Low Q30 can result from poor library quality, overloading the flow cell, or reagent problems. Check cluster density: typical values are 800 1200 clusters per mm2 for a MiSeq, deviation can cause mixed clusters or low output. Use FastQC or the Galaxy Training Network quality control workflows to detect adapter contamination, GC bias, and duplicate reads.

Common Mistakes

Using Sanger for population level studies. Trying to Sanger sequence a metagenomic sample or a polyclonal viral population results in a consensus that masks rare variants. The PLoS Pathogens study on HIV drug resistance demonstrated that Sanger sequencing underestimated the diversity of resistant variants compared to deep Illumina sequencing [7].

Relying on Illumina for single variant validation without confirmatory method. Illumina errors can occur at low frequency, especially in homopolymer runs or GC rich regions. For clinical reporting, always confirm actionable variants with Sanger or an orthogonal method.

Ignoring read length limits in assembly projects. Illumina short reads cannot span long repetitive elements. For example, barcoding neogastropods using the COI gene with the CODEX approach required careful read overlap and assembly strategies because the target region exceeded the paired end read length [10].

Poor library normalization in Illumina. Overloading the flow cell with too many clusters degrades quality. Under loading wastes throughput. Use qPCR to accurately quantitate libraries, not just spectrophotometry.

Limits of Interpretation

Sanger sequencing delivers a consensus sequence from a bulk sample. It cannot resolve haplotypes or detect variants below 15 20% frequency. When analyzing tumor biopsies or viral quasispecies, a Sanger trace may show a mixture at several positions, but the phase of those mutations (which ones are on the same molecule) remains unknown. Illumina can partially resolve haplotypes through paired end reads and computational phasing, but in repetitive regions the short reads may map ambiguously.

Illumina’s short read length also limits interpretation in structural variant detection. Large deletions, inversions, and translocations are harder to call confidently with short reads alone. Base level errors are not random, certain motifs misincorporate more frequently. The high depth of coverage compensates for per read errors, but systematic biases (e.g., GC bias) can distort quantitative measurements like gene expression levels.

Both methods rely on a reference genome. For non model organisms without a high quality reference, Sanger sequencing of long amplicons can help assemble contigs, but complete genomes require hybrid approaches. The EMBL EBI Training notes that no sequencing technology is error free and validation strategies should always match the intended application.

Frequently Asked Questions

1. Which sequencing method is more accurate?
Sanger sequencing is more accurate per base (99.999% vs. 99.9% for a typical Illumina base call). However, Illumina achieves higher consensus accuracy by sequencing each position many times (30x or more), which can reduce the overall error rate below Sanger’s level for variant detection. The tradeoff is that Sanger is the preferred confirmatory method for single nucleotide variants due to its lower systematic error.

2. Can Illumina sequencing completely replace Sanger?
No. Illumina cannot reliably detect very low frequency variants in mixed samples without very deep coverage (1000x+), and it struggles with long repetitive regions and certain homopolymers. Sanger remains essential for validation of clinically actionable variants, for sequencing small sets of amplicons, and for full length gene sequencing when short reads cannot cover the entire coding region.

3. What is the approximate cost difference per base?
Illumina is orders of magnitude cheaper per base when run at full capacity. A single Sanger reaction for one read covering 800 bases costs about $5, or $0.006 per base. An Illumina MiSeq run costing $1000 yields 15 million reads of 150 bases each, or 2.25 billion bases, for $0.0000004 per base. For small numbers of amplicons, Sanger is cheaper because you avoid library preparation costs.

4. How long does each method take from sample to results?
Sanger sequencing: after PCR and cleanup, the capillary run takes 2 4 hours, and analysis can be done in minutes. Total time from DNA extraction to final sequence is about 4 6 hours for a small batch. Illumina: library preparation takes 4 6 hours, cluster generation and sequencing run takes 24 56 hours, and downstream bioinformatics takes hours to days depending on the analysis depth.

References and Further Reading

Related Articles

Somatic Cell Phases In Cell Cycle Endoplasmic Reticulum Cell Function Stages Of Cell Cycles Pcr Test