# Sanger Sequencing Method: Principles, Workflow, and Applications

## Introduction to Sanger Sequencing

Sanger sequencing, also known as chain-termination sequencing, is a method for determining the nucleotide order of a DNA fragment. Developed by Frederick Sanger and his colleagues in 1977, this technique was the foundational technology for [the Human Genome Project](/knowledge/bioinformatics/the-human-genome-project-computational-triumphs) and remained the dominant sequencing approach for nearly three decades. The method's enduring relevance stems from its ability to produce highly accurate, long reads from individual DNA templates, making it the gold standard for validating variants identified by high-throughput methods.

The principle underlying Sanger sequencing is elegantly simple: a DNA polymerase extends a primer along a template strand, but the reaction is stochastically terminated by the incorporation of modified nucleotides that lack the 3'-hydroxyl group required for further extension. This generates a population of fragments of varying lengths, each ending at a specific nucleotide position. When separated by size, these fragments reveal the template sequence directly.

Despite the rise of next-generation sequencing (NGS) platforms that can sequence billions of fragments in parallel, Sanger sequencing retains a critical niche. It is the method of choice for confirming single-nucleotide variants, sequencing individual clones, analyzing short genetic regions, and validating results from whole-genome or targeted NGS panels. Its accuracy per base, often exceeding 99.9%, remains unmatched by most high-throughput platforms. For a detailed comparison of throughput and cost considerations, see [Sanger Sequencing vs NGS](/knowledge/molecular-biology/sanger-sequencing-vs-ngs).

## Core Principles of the Sanger Method

### Chain Termination by Dideoxynucleotides

The central mechanism of Sanger sequencing relies on the incorporation of 2',3'-dideoxynucleoside triphosphates (ddNTPs). These molecules are structurally identical to standard deoxynucleoside triphosphates (dNTPs) except that they lack the 3'-hydroxyl (–OH) group on the deoxyribose sugar. During DNA synthesis, a polymerase forms a phosphodiester bond between the 5'-phosphate of the incoming nucleotide and the 3'-OH of the growing strand. When a ddNTP is incorporated, there is no 3'-OH available for the next phosphodiester bond formation. Chain elongation terminates at that position.

The reaction is designed so that each ddNTP is present at a low molar ratio relative to its corresponding dNTP, typically 1:100 to 1:200. This ensures that termination is a rare, stochastic event rather than a deterministic one. The result is a nested set of DNA fragments, each sharing a common 5' end (the primer) but differing in length by exactly one nucleotide at the 3' end. Every possible termination position within the amplified region is represented in the population.

In a classic four-lane format, four separate reactions are performed, each containing all four dNTPs plus a single type of ddNTP (ddATP, ddCTP, ddGTP, or ddTTP). The fragments from each reaction are then separated in adjacent lanes of a polyacrylamide gel. The sequence is read by determining the order of bands across the four lanes from the bottom (shortest fragment) to the top (longest fragment). Modern automated systems use fluorescently labeled ddNTPs in a single reaction, as described below.

### Primer Extension and Labeling

Sanger sequencing requires a primer that anneals to a known sequence adjacent to the region of interest. The primer provides the free 3'-OH group from which DNA polymerase begins extension. The primer must be unique in the template to avoid priming at multiple sites, which would produce overlapping sequence data and unreadable chromatograms.

Labeling can be achieved in two ways. In the original radioactive method, the primer was 5'-end-labeled with ³²P or ³³P, and fragments were detected by autoradiography. In modern fluorescence-based systems, each of the four ddNTPs is covalently linked to a distinct fluorescent dye with a characteristic emission spectrum. Because the dye is attached to the terminating nucleotide, every fragment carries a fluorescent label at its 3' end that identifies the terminal base. This allows all four termination reactions to be combined in a single tube and separated in a single capillary or lane, with the identity of the terminal nucleotide determined by the fluorescence emission wavelength.

The choice of DNA polymerase is critical. The enzyme must efficiently incorporate ddNTPs, which are poor substrates for many polymerases. Thermostable variants such as Taq DNA polymerase and engineered derivatives like Thermo Sequenase have been modified to accept dideoxynucleotides with higher efficiency, enabling cycle sequencing, a process in which the reaction is repeatedly heated and cooled to linearly amplify the sequencing products.

## The Sanger Sequencing Workflow

### Template Preparation and PCR Amplification

The first step in Sanger sequencing is obtaining a sufficient quantity of pure, double-stranded DNA template. For plasmid DNA, this typically involves [bacterial culture](/blog/guides/bacterial-culture) followed by alkaline lysis and column-based purification. For PCR products, the amplicon must be purified to remove residual primers, dNTPs, and polymerase that would interfere with the sequencing reaction. Enzymatic cleanup using exonuclease I (to degrade single-stranded primers) and shrimp [alkaline phosphatase](/knowledge/molecular-biology/alkaline-phosphatase) (to dephosphorylate remaining dNTPs) is a common approach.

The template concentration should be optimized: approximately 100–200 ng of plasmid DNA or 10–50 ng of purified PCR product per sequencing reaction is typical. Impurities such as salts, ethanol, and detergents must be removed, as they inhibit DNA polymerase activity and reduce signal quality. For guidance on template preparation, refer to the [Prepare Sample for Sequencing](/knowledge/molecular-biology/prepare-sample-for-sequencing) resource.

If the target region is part of a larger construct or genomic DNA, a PCR amplification step is performed first to generate a template of manageable size. The amplicon should ideally be 200–1,000 base pairs for optimal Sanger sequencing results. Larger templates can be sequenced, but read quality diminishes beyond approximately 700–900 bases from the primer.

### Cycle Sequencing Reaction

Cycle sequencing is a linear amplification method that generates single-stranded, dye-labeled fragments. The reaction mixture contains:

- Purified template (plasmid or PCR product)
- A single sequencing primer (0.5–1.0 µM final concentration)
- A mixture of dNTPs and fluorescently labeled ddNTPs
- A thermostable DNA polymerase engineered for ddNTP incorporation
- Reaction buffer containing MgCl₂ (typically 1.5–2.5 mM final concentration)

The thermal cycling protocol typically consists of 25–35 cycles of:

1. **Denaturation** at 96°C for 10–30 seconds, which separates the double-stranded template into single strands.
2. **Annealing** at 50–60°C for 5–15 seconds, allowing the primer to hybridize to its complementary sequence.
3. **Extension** at 60°C for 2–4 minutes, during which the polymerase incorporates nucleotides until a ddNTP is added, terminating the strand.

Because each cycle uses the original template as a substrate, the number of product fragments increases linearly (not exponentially) with cycle number. The ratio of ddNTPs to dNTPs determines the average fragment length; higher ddNTP concentrations produce shorter fragments, while lower concentrations yield longer reads.

After cycling, the reaction products are purified to remove unincorporated dye terminators, salts, and buffer components. This is typically done using ethanol precipitation or spin-column purification. The purified products are then resuspended in formamide or water for injection into the capillary electrophoresis system.

### Capillary Electrophoresis and Detection

The sequencing products are separated by size using capillary electrophoresis. Each sample is electrokinetically injected into a thin glass capillary (50–100 µm internal diameter) filled with a sieving polymer matrix, typically a linear polyacrylamide or polyethylene oxide solution. An electric field (typically 300 V/cm) drives the negatively charged DNA fragments through the matrix, where they separate by size, smaller fragments migrate faster.

As the fragments pass through a detection window near the anode, a laser excites the fluorescent dyes attached to the ddNTPs. The emitted light is collected by a detector and separated by wavelength using a diffraction grating or dichroic mirrors. Each of the four dyes has a distinct emission maximum (e.g., 6-FAM emits at approximately 520 nm, HEX at 555 nm, NED at 575 nm, and ROX at 605 nm in the Applied Biosystems system). The detector records fluorescence intensity at each wavelength as a function of time, generating four-color trace data.

The time at which a fragment reaches the detector is directly proportional to its size. By calibrating the system with a size standard (a mixture of known DNA fragments labeled with a fifth dye, such as LIZ), the software converts migration time to fragment length in nucleotides. The sequence is then read from the order of fluorescent peaks, starting from the shortest fragment (closest to the primer) and progressing to longer fragments.

## Key Reagents and Their Roles

### DNA Polymerase Variants

The choice of DNA polymerase is a critical determinant of Sanger sequencing success. Wild-type Taq DNA polymerase incorporates ddNTPs poorly, with a discrimination factor of approximately 1,000-fold against dideoxynucleotides relative to deoxynucleotides. This would require impractically high ddNTP concentrations to achieve adequate termination.

Engineered variants address this limitation. Thermo Sequenase, a mutant of Taq polymerase with amino acid substitutions in the nucleotide-binding pocket, incorporates ddNTPs with only a 3-fold discrimination relative to dNTPs. This allows lower ddNTP concentrations, producing longer reads and more uniform peak heights. Other variants include:

- **AmpliTaq FS**: A Taq derivative with a G46D mutation in the active site, reducing ddNTP discrimination.
- **Sequenase**: A modified T7 DNA polymerase that efficiently incorporates ddNTPs but is not thermostable, limiting its use to non-cycling reactions.
- **Phusion and other high-fidelity polymerases**: Generally not suitable for Sanger sequencing due to poor ddNTP incorporation, though they are excellent for template amplification.

The polymerase must also lack significant 3'→5' exonuclease (proofreading) activity, as this activity can remove incorporated ddNTPs, defeating chain termination. Most sequencing-grade polymerases are engineered to eliminate this activity.

### Fluorescent Dye Labeling

Modern Sanger sequencing uses four distinct fluorescent dyes, each attached to a specific ddNTP. The dyes are designed to meet several criteria:

- **Distinct emission spectra**: The emission maxima must be sufficiently separated to allow unambiguous identification by the detector.
- **High quantum yield**: Bright fluorescence enables detection of small amounts of DNA.
- **Minimal spectral overlap**: Overlapping spectra would cause cross-talk between detection channels, requiring complex matrix correction.
- **Electrophoretic mobility matching**: The dyes should not significantly alter the electrophoretic mobility of the fragments, or if they do, the effect should be uniform across fragment sizes.

The Applied Biosystems BigDye chemistry uses four dyes: dR110 (green), dR6G (yellow-green), dTAMRA (orange), and dROX (red), each linked to the corresponding ddNTP via a linker arm. The dyes are designed so that their mobility shifts are similar, allowing accurate base calling across a wide size range. The dye terminator chemistry has largely replaced the older dye primer approach, in which the primer (not the ddNTP) was labeled, because it allows all four termination reactions to be performed in a single tube.

The buffer conditions for the sequencing reaction are also important. Tris-HCl (pH 9.0–9.5) is commonly used, with MgCl₂ at 1.5–2.5 mM. The pH is critical because it affects both polymerase activity and dye fluorescence. Some formulations include betaine or dimethyl sulfoxide (DMSO) to reduce secondary structure in GC-rich templates, which can cause polymerase pausing and premature termination.

## Data Analysis and Base Calling

### Electropherogram Interpretation

The raw output of a Sanger sequencing run is an electropherogram, a plot of fluorescence intensity versus time (or fragment size) for each of the four dye channels. The base-calling software processes this data through several steps:

1. **Mobility correction**: Adjusts for differences in electrophoretic mobility between dye-labeled fragments.
2. **Spectral deconvolution**: Corrects for spectral overlap between dyes using a matrix derived from control runs.
3. **Peak detection**: Identifies fluorescence peaks above a threshold in each channel.
4. **Base assignment**: Assigns a nucleotide to each peak based on the dominant fluorescence channel.
5. **Quality scoring**: Assigns a quality value to each base call.

A well-analyzed electropherogram shows evenly spaced, single peaks with minimal background fluorescence. The first 20–40 bases after the primer are often unreliable due to unincorporated dye terminators and primer artifacts; these should be trimmed. Reliable sequence typically extends 500–900 bases from the primer, depending on template quality and reaction conditions.

The four-color trace is displayed as a chromatogram with peaks colored by base: green for A, blue for C, black for G, and red for T in the Applied Biosystems convention. Heterozygous positions in diploid samples appear as two overlapping peaks of roughly equal height at the same position. The software can call these as ambiguity codes (e.g., R for A/G, Y for C/T) when both signals exceed a threshold.

### Phred Quality Scores

Phred quality scores, developed by Phil Green and colleagues at the University of Washington, provide a numerical measure of base-calling confidence. The Phred score (Q) is defined as:

Q = −10 × log₁₀(P)

where P is the probability that the base call is incorrect. A Phred score of 20 (Q20) corresponds to an error probability of 1 in 100 (99% accuracy), Q30 to 1 in 1,000 (99.9% accuracy), and Q40 to 1 in 10,000 (99.99% accuracy).

Phred scores are calculated using a base-calling algorithm that considers peak shape, spacing, signal-to-noise ratio, and the resolution of adjacent peaks. The scores are typically reported alongside the sequence in the quality file. For most applications, a Q20 or higher is considered acceptable, while Q30 or higher is preferred for variant confirmation. The quality scores decline with increasing read length as peaks broaden and signal intensity decreases, which is why Sanger reads are typically limited to 700–900 bases.

The Phred quality score system has been widely adopted and is also used in NGS platforms, allowing direct comparison of base quality across technologies. For a discussion of how sequencing depth and coverage relate to quality in high-throughput contexts, see [Sequencing Coverage](/knowledge/molecular-biology/sequencing-coverage).

## Applications of Sanger Sequencing

### Variant Confirmation

The most common application of Sanger sequencing today is the confirmation of variants identified by NGS. While NGS platforms can identify candidate variants across large genomic regions, their error rates, particularly for indels and variants in repetitive or GC-rich regions, can produce false positives. Sanger sequencing provides an orthogonal, high-accuracy method to verify these findings.

The workflow is straightforward: primers are designed to flank the variant of interest, the region is PCR-amplified from the original sample, and the amplicon is sequenced by the Sanger method. The resulting chromatogram is examined at the variant position to confirm the presence of the expected nucleotide change. This confirmation step is particularly important in clinical diagnostics, where treatment decisions may depend on the presence of specific mutations.

For example, in oncology, Sanger sequencing is used to confirm mutations in genes such as *EGFR*, *KRAS*, and *BRAF* identified by targeted NGS panels. The high accuracy of Sanger sequencing ensures that a patient is not incorrectly classified as harboring a drug-sensitive or drug-resistant mutation based on an NGS artifact.

### Microbial Identification

Sanger sequencing of the 16S ribosomal RNA gene is a standard method for bacterial identification. The 16S rRNA gene is approximately 1,500 base pairs and contains both conserved and variable regions. Universal primers targeting conserved regions amplify the gene from any bacterial species, and the sequence of the variable regions allows taxonomic classification.

The typical workflow involves:

1. PCR amplification of the 16S rRNA gene using universal primers (e.g., 27F and 1492R).
2. Sanger sequencing of the amplicon, often using internal primers to cover the full gene.
3. Comparison of the resulting sequence against reference databases such as GenBank or the Ribosomal Database Project.

A sequence identity of ≥99% with a reference strain typically indicates species-level identification, while ≥97% identity indicates genus-level identification. This approach is widely used in [clinical microbiology](/knowledge/diagnostics/microbiology/clinical-microbiology-from-specimen-collection-to-pathogen-identification), environmental microbiology, and food safety testing. Similar approaches are used for fungal identification using the internal transcribed spacer (ITS) region and for viral genotyping using specific genomic regions.

### Forensic and Clinical Applications

In forensic science, Sanger sequencing is used for mitochondrial DNA (mtDNA) analysis, particularly in cases where nuclear DNA is degraded or present in low quantity. The hypervariable regions HV1 and HV2 of the [mitochondrial genome](/blog/guides/mitochondrial-genome) are amplified and sequenced, and the resulting haplotypes are compared against reference databases. While less discriminating than short tandem repeat (STR) analysis, mtDNA sequencing can be valuable for analyzing hair shafts, old bones, and other samples with limited nuclear DNA.

In clinical diagnostics, Sanger sequencing remains the method of choice for:

- **Single-gene disorders**: Sequencing of genes such as *CFTR* (cystic fibrosis), *HBB* (beta-thalassemia), and *HTT* (Huntington's disease) for diagnostic confirmation.
- **Pharmacogenetics**: Genotyping of variants in genes such as *CYP2C19* and *CYP2D6* that affect drug metabolism.
- **HLA typing**: High-resolution typing of human leukocyte antigen genes for organ transplantation matching.
- **Inherited cancer syndromes**: Sequencing of genes such as *BRCA1* and *BRCA2* for hereditary breast and ovarian cancer risk assessment.

The regulatory acceptance of Sanger sequencing for clinical applications is well established, and many diagnostic laboratories maintain validated Sanger-based assays for specific genes. The method's reproducibility and accuracy make it suitable for applications where a false result could have serious medical consequences.

## Advantages and Limitations

### Read Length and Accuracy

The primary advantage of Sanger sequencing is its combination of long read length and high accuracy. A single Sanger read can span 700–900 bases with per-base accuracy exceeding 99.9% (Q30 or higher). This is sufficient to cover most PCR amplicons and plasmid inserts in a single read, simplifying assembly and reducing the need for bioinformatic processing.

The long read length is particularly valuable for:

- **Resolving repetitive regions**: Sequences that are longer than the read length of short-read NGS platforms (150–300 bases) can be fully covered by a single Sanger read.
- **Phasing variants**: When two variants are present within a single read, their phase (whether they are on the same or different chromosomes) can be determined directly.
- **Sequencing GC-rich regions**: Sanger sequencing is often more successful than NGS in GC-rich templates, which are prone to amplification bias and coverage dropout in high-throughput methods.

The accuracy of Sanger sequencing is attributable to several factors: the use of a single, well-characterized polymerase; the absence of amplification bias inherent in cluster generation; and the direct detection of incorporated nucleotides rather than inference from sequencing-by-synthesis signals.

### Throughput and Cost

The principal limitation of Sanger sequencing is its low throughput. A single capillary electrophoresis instrument can process 96 or 384 samples per run, with each run taking 1–3 hours. This translates to a maximum of roughly 1,000–3,000 samples per day per instrument, with each sample yielding one sequence read. In contrast, a single NGS run can generate millions to billions of reads in the same time period.

The cost per base of Sanger sequencing is substantially higher than NGS. While the cost per reaction is relatively low (approximately $3–$10 for reagents), the cost per base is high because each reaction yields only 500–900 bases of sequence. For projects requiring more than a few hundred amplicons, NGS becomes more cost-effective despite the higher upfront instrument cost.

The choice between Sanger and NGS depends on the scale of the project:

| Parameter | Sanger Sequencing | NGS (Illumina) |
|-----------|-------------------|----------------|
| Read length | 400–900 bases | 150–300 bases (paired-end) |
| Reads per run | 96–384 | 10⁶–10⁹ |
| Throughput per run | ~0.3 Mb | 1–3,000 Gb |
| Accuracy per base | >99.9% | 99.0–99.9% |
| Cost per Mb | $50–$500 | $0.01–$0.50 |
| Time per run | 1–3 hours | 4–72 hours |
| Primary use | Variant confirmation, small-scale | Whole-genome, exome, targeted panels |

For projects requiring sequencing of more than approximately 100 amplicons, NGS is generally more cost-effective. However, the lower error rate and longer reads of Sanger sequencing make it the preferred method for validating critical variants and for applications where accuracy is paramount. The [Nanopore vs Sanger Sequencing](/knowledge/molecular-biology/nanopore-vs-sanger-sequencing) comparison provides additional context on long-read alternatives.

## Common Pitfalls and Troubleshooting

### Weak or No Signal

A failed Sanger sequencing reaction often produces no signal or a signal too weak to analyze. Common causes include:

- **Insufficient template**: The reaction requires a minimum quantity of template DNA. For plasmid DNA, 100–200 ng is typical; for PCR products, 10–50 ng. If the template concentration is too low, the number of extension products will be insufficient for detection.
- **Poor template quality**: Contaminants such as salts, ethanol, phenol, or detergents inhibit DNA polymerase. Repurify the template using a spin column or ethanol precipitation.
- **Primer problems**: The primer may not anneal due to incorrect sequence, incorrect annealing temperature, or secondary structure. Verify the primer sequence against the template and check the calculated melting temperature. If the primer has a GC content below 40% or above 60%, redesign it.
- **Degraded reagents**: ddNTPs and dNTPs are sensitive to repeated freeze-thaw cycles. Prepare fresh dilutions and store at −20°C in small aliquots.
- **Polymerase inactivation**: The polymerase may be inactive due to improper storage or exposure to high temperatures. Check the expiration date and storage conditions.

If the signal is weak but present, increasing the number of cycle sequencing reactions from 25 to 35 can help. Alternatively, increasing the template concentration within the recommended range may improve signal intensity.

### Mixed Signals and Background Noise

Mixed signals, where multiple peaks appear at a single position, can result from several causes:

- **Primer contamination**: If the template contains multiple priming sites, the reaction will produce overlapping sequences. This is common when sequencing PCR products that contain primer dimers or non-specific amplicons. Purify the PCR product by gel extraction to remove contaminating fragments.
- **Mixed template population**: If the template contains a mixture of sequences (e.g., from a mixed [bacterial culture](/blog/guides/bacterial-culture) or a plasmid with a deletion), the chromatogram will show overlapping peaks. This can be resolved by cloning the template or by using a more specific amplification strategy.
- **Secondary structure**: GC-rich templates can form secondary structures that cause the polymerase to pause or terminate prematurely, producing peaks with reduced intensity or multiple peaks. Adding DMSO (5–10%) or betaine (1–2 M) to the reaction can help denature these structures.
- **Dye blobs**: Unincorporated dye terminators can appear as large, broad peaks in the electropherogram, typically in the first 50 bases. These are removed by thorough purification of the sequencing reaction products. If dye blobs persist, repeat the purification step or use a different purification method.

Background noise, characterized by elevated fluorescence across all channels, can result from:

- **Excess template**: Too much template can saturate the reaction and produce high background. Reduce the template concentration.
- **Incomplete removal of unincorporated dye terminators**: Ensure that the purification step is performed correctly and that the wash buffer is completely removed.
- **Contaminated capillaries**: Carryover from previous runs can cause background fluorescence. Run a water blank between samples to check for contamination.

For a comprehensive troubleshooting guide, consult the [Sanger Sequencing Protocol](/knowledge/molecular-biology/sanger-sequencing-protocol) resource.

## Summary and Best Practices

Sanger sequencing remains an essential tool in [molecular biology](/blog/careers/molecular-biology), offering unmatched accuracy and read length for small-scale sequencing applications. The method's principle, chain termination by dideoxynucleotides, has remained unchanged for over four decades, though the chemistry and instrumentation have been refined substantially.

Best practices for successful Sanger sequencing include:

1. **Design primers carefully**: Primers should be 18–24 nucleotides long, with a GC content of 40–60%, and a melting temperature of 55–65°C. Avoid runs of four or more identical nucleotides and check for secondary structure.
2. **Purify templates rigorously**: Remove all contaminants that could inhibit the polymerase or interfere with electrophoresis.
3. **Optimize template concentration**: Use the recommended amount of template for the specific chemistry being used.
4. **Include appropriate controls**: Sequence a known positive control in each run to verify reagent performance.
5. **Trim low-quality sequence**: Remove the first 20–40 bases and the final 50–100 bases where quality scores decline.
6. **Verify results**: For critical applications, sequence both strands of the template to confirm the result.
7. **Maintain reagents properly**: Store enzymes at −20°C, protect fluorescent dyes from light, and avoid repeated freeze-thaw cycles.

The integration of Sanger sequencing with NGS workflows, using Sanger to confirm variants identified by high-throughput methods, represents the current best practice in genomics. This hybrid approach leverages the strengths of both technologies: the scale of NGS and the accuracy of Sanger sequencing.

## Frequently Asked Questions

### What is the Sanger sequencing method?

Sanger sequencing is a DNA sequencing technique that determines the nucleotide order of a DNA fragment by using dideoxynucleotides to terminate DNA synthesis at specific positions. Developed by Frederick Sanger in 1977, it was the first widely adopted sequencing method and remains the gold standard for validating variants and sequencing individual genes or small genomic regions.

### How does Sanger sequencing work?

Sanger sequencing works by performing a DNA polymerase reaction in the presence of a mixture of normal deoxynucleotides (dNTPs) and fluorescently labeled dideoxynucleotides (ddNTPs). When a ddNTP is incorporated into the growing DNA strand, chain elongation stops because the ddNTP lacks the 3'-hydroxyl group needed for the next phosphodiester bond. This produces a population of fragments of varying lengths, each ending at a specific nucleotide. The fragments are separated by capillary electrophoresis, and the fluorescence of each fragment identifies the terminal nucleotide, allowing the sequence to be read.

### What is the difference between Sanger sequencing and next-generation sequencing?

Sanger sequencing produces long reads (400–900 bases) with very high accuracy (>99.9%) but low throughput (96–384 reads per run). Next-generation sequencing (NGS) produces millions to billions of short reads (150–300 bases) per run with slightly lower per-base accuracy. Sanger sequencing is used for small-scale projects, variant confirmation, and sequencing individual genes, while NGS is used for whole-genome, exome, and large-scale targeted sequencing. The cost per base is much lower for NGS, but the upfront instrument cost and bioinformatics requirements are higher. See [Sanger Sequencing vs NGS](/knowledge/molecular-biology/sanger-sequencing-vs-ngs) for a detailed comparison.

### What are the applications of Sanger sequencing?

Sanger sequencing is used for confirming variants identified by NGS, sequencing individual genes for diagnostic purposes, identifying microorganisms by 16S rRNA sequencing, analyzing mitochondrial DNA in forensics, HLA typing for transplantation, and sequencing plasmid inserts and PCR products. It is also used in pharmacogenetics to genotype drug-metabolizing enzyme genes and in oncology to confirm mutations in genes such as *EGFR*, *KRAS*, and *BRAF*.

### What are the common problems in Sanger sequencing?

Common problems include weak or no signal (due to insufficient template, poor template quality, or inactive reagents), mixed signals (from primer contamination or mixed template populations), dye blobs (unincorporated dye terminators appearing as broad peaks), and background noise (from excess template or incomplete purification). Secondary structure in GC-rich templates can cause premature termination, and primer dimers can produce non-specific products. Most problems can be resolved by optimizing template concentration, improving purification, and redesigning primers.

### How do you read a Sanger sequencing chromatogram?

A Sanger sequencing chromatogram displays four colored peaks, each representing one of the four nucleotides (green for A, blue for C, black for G, red for T in the Applied Biosystems convention). The sequence is read from left to right, with each peak representing one nucleotide. The height of the peak indicates signal intensity, and the spacing between peaks should be uniform. Heterozygous positions appear as two overlapping peaks of similar height. Quality scores, typically shown as a trace above the chromatogram, indicate the confidence of each base call.

### What is the read length of Sanger sequencing?

The read length of Sanger sequencing is typically 400–900 bases per reaction, with 700–800 bases being a common target for high-quality data. The read length is limited by the resolution of capillary electrophoresis, which decreases with increasing fragment size, and by the declining signal intensity as the polymerase processivity decreases. The first 20–40 bases after the primer are often unreliable and should be trimmed, and the last 50–100 bases may have reduced quality scores.

## Key Takeaways

- Sanger sequencing uses dideoxynucleotide chain termination to generate a nested set of DNA fragments whose sizes reveal the template sequence.
- The method produces reads of 400–900 bases with per-base accuracy exceeding 99.9%, making it the gold standard for variant confirmation.
- Modern Sanger sequencing uses four fluorescently labeled ddNTPs in a single reaction, separated by capillary electrophoresis and detected by laser-induced fluorescence.
- Phred quality scores (Q20, Q30) provide a quantitative measure of base-calling confidence and are used to assess data quality.
- Sanger sequencing is primarily used for confirming NGS-identified variants, sequencing individual genes, microbial identification, and clinical diagnostics.
- The main limitations of Sanger sequencing are low throughput and high cost per base compared to NGS, making it unsuitable for large-scale projects.
- Successful Sanger sequencing requires careful primer design, rigorous template purification, optimized reaction conditions, and proper interpretation of electropherograms.

## Further Reading

- Vincent AT et al. *Next-generation sequencing (NGS) in the microbiological world: How to make the most of your money*. Journal of microbiological methods. 2017. [PubMed 26995332](https://doi.org/10.1016/j.mimet.2016.02.016)
- Furutani S et al. *Rapid DNA Sequencing Technology Based on the Sanger Method for Bacterial Identification*. Sensors (Basel, Switzerland). 2022. [PubMed 35336302](https://doi.org/10.3390/s22062130)
- Deminco F et al. *A Simplified Sanger Sequencing Method for Detection of Relevant SARS-CoV-2 Variants*. Diagnostics (Basel, Switzerland). 2022. [PubMed 36359452](https://doi.org/10.3390/diagnostics12112609)
- Guo LT, Pyle AM. *RT-based Sanger sequencing of RNAs containing complex RNA repetitive elements*. Methods in enzymology. 2023. [PubMed 37914445](https://doi.org/10.1016/bs.mie.2023.07.003)
- Sha S et al. *[Tri-primer-florescence PCR-Sanger sequencing method for screening of full and pre-mutations of FMR1 gene]*. Zhonghua yi xue yi chuan xue za zhi = Zhonghua yixue yichuanxue zazhi = Chinese journal of medical genetics. 2016. [PubMed 27984619](https://doi.org/10.3760/cma.j.issn.1003-9406.2016.06.022)
- Deng YM et al. *A simplified Sanger sequencing method for [full genome sequencing](/blog/guides/full-genome-sequencing) of multiple subtypes of human influenza A viruses*. Journal of clinical virology : the official publication of the Pan American Society for Clinical Virology. 2015. [PubMed 26071334](https://doi.org/10.1016/j.jcv.2015.04.019)



<div data-calculator="molecular-cloning"></div>

## Related Clinical & Scientific Guides

* [MAPK Pathway: Mechanism, Function, and Clinical Relevance](/knowledge/molecular-biology/mapk-pathway)
* [Mammalian Cell Culture Bioreactors: A Practical Guide](/knowledge/molecular-biology/mammalian-cell-culture-bioreactor)
* [Nucleotide Formation: Biosynthesis and Assembly of DNA/RNA Building Blocks](/knowledge/molecular-biology/nucleotide-formation)