Sanger Sequencing Validation of Variant Calls: A Practical Guide to Primer Design, PCR, and Interpretation
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Sanger sequencing validation confirms Next-Generation Sequencing (NGS) variant calls by amplifying targeted genomic regions via PCR and analyzing them with capillary electrophoresis, essential for verifying single nucleotide variants (SNVs), small insertions/deletions (indels), and some structural variants (SVs).
- Orthogonal validation is critical due to variability in NGS variant calling accuracy; for instance, Genome Analysis Toolkit (GATK) reported a 92.55% positive predictive value compared to SAMtools' 80.35% when validated against Sanger and array genotyping.
- Primer design for Sanger validation requires careful consideration of amplicon size (400-800 bp), variant position (at least 50 bp from primer binding sites), and specificity checks against the reference genome to avoid non-specific amplification from repetitive or homologous sequences.
- PCR optimization, particularly determining the optimal annealing temperature via gradient PCR and adjusting magnesium concentration, is crucial for generating a single, specific amplicon, with common failures including no amplification or multiple bands on gel electrophoresis.
- Electropherogram interpretation involves assessing peak quality, spacing, and background noise to confirm variant presence and zygosity, with heterozygous variants typically showing two overlapping peaks of roughly equal height, while discordant calls necessitate investigation into potential allele-specific amplification or NGS false positives.
- Sanger sequencing has a detection limit of approximately 10-20% for variant allele fractions, making it less suitable for validating low-frequency somatic variants; regions with high GC content or complex repeats can also pose interpretation and amplification challenges.
Direct Answer and Scope
Sanger sequencing validation of variant calls from next-generation sequencing (NGS) data is a confirmatory process in which targeted genomic regions are amplified by polymerase chain reaction (PCR) and sequenced using capillary electrophoresis to verify the presence and zygosity of candidate variants. This guide provides a practical workflow for researchers, biology students, and laboratory professionals who need to confirm NGS variant calls before downstream analysis or reporting. The protocol covers primer design using publicly available tools, PCR optimization, electropherogram interpretation, and troubleshooting of common failures. The intended outcome is a reproducible validation process that produces reliable Sanger traces for single nucleotide variants (SNVs), small insertions and deletions (indels), and selected structural variants (SVs).
The need for orthogonal validation arises from the observation that NGS variant calling pipelines vary in accuracy. A comparative study of variant calling tools reported that the Genome Analysis Toolkit (GATK) achieved a positive predictive value of 92.55% while SAMtools achieved 80.35% when validated against Sanger sequencing and array genotyping data from 130 whole exome sequencing samples [<a href="#ref-1">1</a>]. This difference demonstrates that no single bioinformatics pipeline is infallible and that confirmatory testing adds value, particularly for variants with clinical or diagnostic implications. A separate assessment of 7,601 NGS variants from 5,190 clinical samples found that 6,939 high-quality variants with at least 35x depth coverage and at least 35% heterozygous ratio were 100% confirmed by a secondary methodology [<a href="#ref-2">2</a>]. These findings support a tiered approach in which high-quality variants from validated workflows may not require routine Sanger confirmation, while lower-confidence calls or variants with clinical significance warrant orthogonal validation.
This guide assumes the reader has basic familiarity with NGS data formats, variant calling formats (VCF), and standard molecular biology techniques. The workflow described here applies to germline variant calling, somatic variant calling, and targeted panel sequencing. The principles also extend to validating structural variants, although the PCR and primer design considerations differ for large deletions, insertions, and rearrangements.
At a Glance
| Validation Step | Key Decision | Primary Consideration | Common Failure Mode |
|---|---|---|---|
| Variant selection | Which variants require Sanger confirmation | Clinical significance, variant quality metrics, allele frequency | Confirming variants that do not need validation, wasting time and cost |
| Primer design | Choose primers flanking the variant site | Amplicon size 400-800 bp, variant position not near primer binding site, GC content 40-60% | Primers binding to repetitive regions or homologous sequences producing non-specific products |
| PCR optimization | Annealing temperature and magnesium concentration | Gradient PCR to determine optimal annealing temperature | No amplification or multiple bands on gel electrophoresis |
| PCR cleanup | Remove unincorporated primers and dNTPs | Enzymatic cleanup with exonuclease I and shrimp alkaline phosphatase | Residual primers interfering with sequencing reaction |
| Cycle sequencing | Prepare single-stranded products with labeled terminators | Use forward or reverse primer separately, never both in one reaction | Weak signal or noisy baseline from excess template or primer |
| Electropherogram interpretation | Assess peak quality and zygosity | Peak height, spacing, background noise, heterozygous peak ratio | Overlapping peaks, pull-up artifacts, or dye blobs misread as variants |
| Result reporting | Document confirmation or discordance | Compare Sanger genotype to NGS genotype call | Discordant calls not investigated or resolved |
Variant Calling Context and the Role of Sanger Validation
Germline Variant Calling
Germline variant calling identifies variants present in constitutional DNA, typically from blood, saliva, or tissue samples. The expectation in germline calling is that variants are present in either a heterozygous or homozygous state, with allele fractions near 50% or 100% respectively. Sanger sequencing validation for germline variants is most valuable when the variant has a low quality score, is located in a region with poor coverage, or has clinical significance for a hereditary condition.
A study validating a targeted gene panel for hereditary chronic liver diseases compared targeted sequencing (TS) with whole exome sequencing (WES) and confirmed all variants by Sanger sequencing except two uniquely detected by TS [<a href="#ref-3">3</a>]. The study reported that discrepancies in variant calling were mainly due to variability in read depth and insufficient coverage in the corresponding target regions [<a href="#ref-3">3</a>]. This finding underscores the importance of examining coverage metrics before deciding which variants to validate. Variants in regions with low depth or uneven coverage are prime candidates for Sanger confirmation.
Somatic Variant Calling
Somatic variant calling identifies variants present in tumor or diseased tissue that are not present in the germline. Somatic variants often have lower allele fractions due to tumor heterogeneity, stromal contamination, or subclonal populations. The validation of somatic variants by Sanger sequencing is complicated by the possibility that the variant allele fraction falls below the detection limit of Sanger sequencing, which is generally considered to be around 10-20% for reliable detection.
When validating somatic variants, the laboratory must consider whether the Sanger assay has sufficient sensitivity to detect the variant at the expected allele fraction. If the NGS data indicate an allele fraction below 15%, Sanger validation may produce a false negative result. In such cases, the validation strategy should include a discussion of the limitations and may require alternative approaches such as droplet digital PCR or targeted deep resequencing.
Variant Filtering and Selection for Validation
The decision to validate a variant by Sanger sequencing should be guided by a filtering strategy that prioritizes variants based on quality metrics, biological relevance, and clinical significance. The following criteria are commonly applied:
- Variant quality score below the pipeline-specific threshold
- Read depth below 30x for germline calls or below 100x for somatic calls
- Allele balance inconsistent with expected zygosity
- Variants located in homopolymer regions, low complexity regions, or segmental duplications
- Variants with clinical significance or actionability
- Novel variants not present in public databases
- Variants that fail quality control metrics in the VCF file
A study comparing variant calling pipelines found that mapping quality, read depth, and allele balance correlate with SNV call accuracy [<a href="#ref-1">1</a>]. However, the same study noted that if best practices are used in data processing, additional filtering based on these metrics provides little gain and accuracies above 99% are achievable [<a href="#ref-1">1</a>]. This observation supports the position that high-quality variants from well-validated pipelines may not require routine Sanger confirmation, while lower-confidence calls warrant validation.
Core Principles of Sanger Sequencing Validation
The Chemistry of Sanger Sequencing
Sanger sequencing relies on the incorporation of dideoxynucleotide triphosphates (ddNTPs) that terminate DNA chain elongation. The sequencing reaction contains a mixture of standard deoxynucleotides and fluorescently labeled ddNTPs. When the DNA polymerase incorporates a ddNTP, chain elongation stops, producing fragments of varying lengths. Capillary electrophoresis separates these fragments by size, and a laser excites the fluorescent labels to produce a chromatogram with peaks corresponding to each nucleotide position.
The quality of a Sanger trace depends on several factors: the purity of the PCR product, the ratio of template to primer, the annealing temperature of the sequencing primer, and the performance of the DNA polymerase. Poor quality traces can result from residual PCR primers, excess template, or suboptimal reaction conditions.
Why Orthogonal Validation Matters
Orthogonal validation using an independent technology provides confidence that the variant call is accurate and not an artifact of the NGS workflow. Sanger sequencing is considered the gold standard for variant confirmation because it uses a different chemistry and detection method than NGS. The validation process can identify false positive calls arising from alignment errors, PCR duplicates, sequencing errors, or bioinformatics artifacts.
A comprehensive assessment of NGS variant validation using a secondary technology reported that high-quality variants with sufficient depth and heterozygous ratio were 100% confirmed by a secondary methodology [<a href="#ref-2">2</a>]. The same study found that 1.5% of validated variants required mass spectrometry genotyping because Sanger results were unavailable or inconsistent with NGS calls [<a href="#ref-2">2</a>]. This finding indicates that Sanger validation is not always straightforward and that alternative methods may be needed for certain variants.
Limitations of Sanger Validation
Sanger sequencing has inherent limitations that must be acknowledged when designing a validation strategy. The method is not suitable for validating variants in regions with complex repeat structure, high GC content, or multiple homologous sequences. Large structural variants, such as copy number changes or inversions, may require specialized PCR strategies or alternative validation methods.
A study establishing benchmark structural variant calls used PCR amplification and Sanger sequencing to validate 544 randomly selected SVs, demonstrating the robustness of the SV calls [<a href="#ref-4">4</a>]. This work shows that Sanger validation can be applied to structural variants when the breakpoints are known and PCR primers can be designed to flank the junction. However, the study also noted that benchmarking procedures are needed to confidently assess the performance of SV detection approaches [<a href="#ref-4">4</a>].
Primer Design for Sanger Validation
Selecting the Target Region
The first step in primer design is to extract the genomic sequence flanking the variant of interest. The reference sequence should be obtained from a reliable source such as the NCBI databases, which provide access to reference genomes, gene sequences, and variation data [<a href="#ref-5">5</a>]. The sequence should include at least 200 base pairs upstream and downstream of the variant position to allow for primer placement and adequate read length.
For SNV validation, the amplicon should be designed so that the variant site is at least 50 base pairs from either primer binding site. This spacing ensures that the variant is read after the initial low-quality region of the sequencing run and before the signal degrades at the end of the read. For indel validation, the variant should be positioned centrally in the amplicon to allow for clear visualization of the insertion or deletion in the electropherogram.
Primer Design Tools and Parameters
Several web-based and standalone tools are available for primer design. The NCBI Primer-BLAST tool, accessible through the NCBI website, combines primer design with a specificity check against the reference genome [<a href="#ref-5">5</a>]. The tool allows the user to specify the target region, amplicon size range, and primer melting temperature. The specificity check identifies primers that may bind to unintended locations in the genome, which is critical for avoiding non-specific amplification.
Recommended primer design parameters for Sanger validation:
- Amplicon size: 400-800 base pairs
- Primer length: 18-24 nucleotides
- Melting temperature: 55-65 degrees Celsius
- GC content: 40-60%
- Maximum self-complementarity: low
- Maximum pair complementarity: low
- Primer specificity: check against the reference genome
The EMBL-EBI training resources provide educational materials on sequence analysis and primer design that can help researchers understand the underlying principles [<a href="#ref-6">6</a>]. These resources are useful for students and early-career researchers who need to build foundational skills in bioinformatics.
Avoiding Common Primer Design Pitfalls
Primers that bind to repetitive elements, pseudogenes, or homologous sequences will produce non-specific amplification products that interfere with Sanger sequencing. The specificity check in Primer-BLAST is essential for identifying these problematic primers. If the target region contains repetitive elements, the primer should be placed in unique flanking sequence, and the amplicon should be designed to span the repeat without including it in the primer binding sites.
For variants located in GC-rich regions, the primers should be designed with a higher melting temperature and the PCR should include additives such as dimethyl sulfoxide (DMSO) or betaine to improve amplification. For variants in AT-rich regions, shorter primers with lower melting temperatures may be more effective.
Primer Design for Structural Variants
Structural variant validation requires a different primer design strategy. For deletions, primers should flank the predicted breakpoints so that the deletion allele produces a shorter amplicon than the reference allele. For insertions, primers should flank the insertion site so that the insertion allele produces a longer amplicon. For inversions and translocations, primers should be designed to amplify across the junction.
The study that established benchmark structural variant calls used PCR amplification and Sanger sequencing to validate selected SVs [<a href="#ref-4">4</a>]. The validation strategy required careful primer design based on the predicted breakpoints from the NGS data. This approach is feasible when the breakpoints are known with sufficient precision, but it may fail when the breakpoints are uncertain or when the rearrangement is complex.
PCR Amplification for Sanger Validation
PCR Reaction Setup
The PCR reaction for Sanger validation typically uses a high-fidelity DNA polymerase to minimize the introduction of errors during amplification. The reaction contains template DNA, forward and reverse primers, deoxynucleotide triphosphates, buffer, and magnesium chloride. The template concentration should be optimized to produce sufficient product without excessive background.
Standard PCR cycling conditions for Sanger validation:
- Initial denaturation: 95 degrees Celsius for 2-5 minutes
- Denaturation: 95 degrees Celsius for 30 seconds
- Annealing: 55-65 degrees Celsius for 30 seconds
- Extension: 72 degrees Celsius for 30-60 seconds
- Final extension: 72 degrees Celsius for 5-10 minutes
- Number of cycles: 30-35
The annealing temperature should be determined empirically using a gradient PCR. The optimal annealing temperature is typically 3-5 degrees Celsius below the lowest primer melting temperature. If the PCR produces non-specific bands, the annealing temperature should be increased or the magnesium concentration should be adjusted.
PCR Optimization
PCR optimization is an iterative process that adjusts reaction components and cycling conditions to produce a single, specific amplicon of the expected size. The following variables can be adjusted:
- Magnesium chloride concentration: 1.5-3.0 mM
- Primer concentration: 0.1-0.5 micromolar
- Template concentration: 1-50 nanograms per reaction
- Annealing temperature: gradient around the calculated melting temperature
- Extension time: adjusted based on amplicon size and polymerase processivity
The PCR product should be evaluated by agarose gel electrophoresis before proceeding to the sequencing reaction. A single band of the expected size indicates successful amplification. Multiple bands or smears indicate non-specific amplification that must be resolved before sequencing.
PCR Cleanup
The PCR product must be purified before the sequencing reaction to remove unincorporated primers and deoxynucleotide triphosphates. Enzymatic cleanup using exonuclease I and shrimp alkaline phosphatase is the most common method. Exonuclease I degrades single-stranded DNA, including residual primers, while shrimp alkaline phosphatase dephosphorylates unincorporated dNTPs. The cleanup reaction is incubated at 37 degrees Celsius for 15-30 minutes, followed by enzyme inactivation at 80 degrees Celsius for 15 minutes.
Alternative cleanup methods include column-based purification or bead-based purification. These methods are more expensive but may be necessary for PCR products with high background or for samples that require higher purity.
Cycle Sequencing and Capillary Electrophoresis
Cycle Sequencing Reaction
The cycle sequencing reaction uses the purified PCR product as the template and a single sequencing primer. The reaction contains fluorescently labeled ddNTPs, a DNA polymerase optimized for sequencing, buffer, and the sequencing primer. The cycling conditions are similar to PCR but with fewer cycles, typically 25-30.
The sequencing primer can be the same as the forward or reverse PCR primer, or it can be an internal primer designed to anneal within the amplicon. Using the PCR primers as sequencing primers is convenient and cost-effective, but internal primers may produce cleaner traces if the PCR primers have secondary structure or if the variant is located close to the primer binding site.
Capillary Electrophoresis
The cycle sequencing products are purified to remove unincorporated ddNTPs and then loaded onto a capillary electrophoresis instrument. The instrument separates the fluorescently labeled fragments by size and detects the fluorescence at each position. The output is an electropherogram with peaks corresponding to each nucleotide.
The quality of the electropherogram depends on the amount of template loaded, the performance of the capillary array, and the quality of the sequencing reaction. The instrument software assigns a quality score to each base call, and the researcher should inspect the trace manually to confirm the calls.
Data Analysis and Base Calling
The electropherogram is analyzed by base calling software that converts the fluorescence signal into a nucleotide sequence. The software assigns quality scores to each base call based on peak height, peak spacing, and signal-to-noise ratio. The researcher should review the trace to confirm that the variant site is clearly readable and that the zygosity call is consistent with the peak pattern.
For a heterozygous variant, the electropherogram shows two overlapping peaks at the variant position, each approximately half the height of the surrounding homozygous peaks. For a homozygous variant, the electropherogram shows a single peak at the variant position. For a hemizygous variant, such as a variant on the X chromosome in a male, the electropherogram shows a single peak, but the interpretation depends on the sex of the sample.
Interpreting Sanger Electropherograms
Assessing Trace Quality
The first step in interpretation is to assess the overall quality of the trace. A high-quality trace has evenly spaced peaks, consistent peak heights, low background noise, and clear separation between peaks. The trace should be readable from the beginning to the end of the sequencing read, with no ambiguous base calls.
Common quality issues include:
- Weak signal: peaks are too low to distinguish from background noise
- Strong signal: peaks are saturated and may show pull-up artifacts
- Noisy baseline: background fluorescence interferes with base calling
- Compressions: peaks are not evenly spaced, often in GC-rich regions
- Dye blobs: large broad peaks caused by unincorporated dye terminators
If the trace quality is poor, the sequencing reaction should be repeated with adjusted template concentration or a different sequencing primer.
Confirming the Variant Call
The variant site should be examined in both the forward and reverse sequencing reactions when both are available. Confirmation in both directions provides stronger evidence than a single-direction read. The genotype call from the Sanger trace should be compared to the NGS genotype call.
For a heterozygous SNV, the Sanger trace should show two peaks at the variant position with approximately equal heights. The ratio of the two peaks may vary depending on the quality of the sequencing reaction, but a ratio between 0.5 and 2.0 is generally considered acceptable. If the ratio is outside this range, the variant may be a mosaic or the sequencing reaction may have failed to amplify one allele preferentially.
For a homozygous SNV, the Sanger trace should show a single peak at the variant position. The peak should be clean, with no evidence of a second peak at the same position.
For an indel, the Sanger trace will show overlapping peaks downstream of the variant site because the insertion or deletion shifts the reading frame. The trace should be examined carefully to determine the exact nature of the indel and to confirm that the NGS call is correct.
Discordant Calls and Resolution
If the Sanger result does not match the NGS call, the discrepancy must be investigated before reporting. Possible explanations include:
- Primer binding site polymorphism that causes allele-specific amplification
- The NGS variant is a false positive call
- The Sanger result is a false negative due to low allele fraction
- Sample mix-up or contamination
- The variant is located in a region with complex structure
The investigation should include a review of the NGS alignment, an examination of the Sanger trace quality, and possibly a repeat of the Sanger sequencing with different primers. If the discrepancy persists, the variant should be reported as discordant and the limitations should be documented.
Records and Measurements for Validation Documentation
Required Records
Laboratory records for Sanger validation should include the following information for each variant:
- Variant identifier and genomic coordinates
- Gene name and transcript identifier
- NGS genotype call and quality metrics
- Primer sequences and amplicon size
- PCR conditions and cleanup method
- Sequencing reaction conditions
- Electropherogram trace file identifier
- Sanger genotype call and interpretation
- Discordance resolution or confirmation status
These records support reproducibility and provide the documentation needed for clinical reporting or publication. The records should be stored in a laboratory information management system or a structured electronic format.
Quality Metrics to Track
The following quality metrics should be tracked for each validation run:
- PCR amplification success rate
- Sequencing reaction success rate
- Trace quality score
- Variant confirmation rate
- Discordance rate
- Time from variant selection to final report
Tracking these metrics over time allows the laboratory to identify systematic issues in the validation workflow and to implement corrective actions. For example, a high discordance rate may indicate a problem with the NGS variant calling pipeline, while a low PCR success rate may indicate a problem with primer design.
Reproducibility Considerations
Reproducibility in Sanger validation requires standardized protocols, validated reagents, and documented procedures. The laboratory should use the same primer design parameters, PCR conditions, and sequencing protocols for all validation runs. Any changes to the protocol should be documented and validated before implementation.
The Galaxy Training Network provides accessible workflow training and analysis tutorials that emphasize reproducibility in bioinformatics [<a href="#ref-7">7</a>]. Similarly, the nf-core documentation describes community pipeline standards and usage that support reproducible analysis workflows [<a href="#ref-8">8</a>]. These resources are relevant for researchers who want to ensure that their NGS variant calling and validation workflows are reproducible.
Common Failure Patterns and Troubleshooting
No PCR Amplification
If the PCR produces no product, the following troubleshooting steps should be considered:
- Verify that the template DNA is of adequate quality and quantity
- Check the primer sequences for errors or misannotation
- Confirm that the annealing temperature is appropriate for the primer melting temperatures
- Increase the number of PCR cycles
- Add PCR additives such as DMSO or betaine for GC-rich templates
- Verify that the polymerase is active and the buffer is correct
Multiple PCR Products
If the PCR produces multiple bands, the following troubleshooting steps should be considered:
- Increase the annealing temperature to improve specificity
- Decrease the primer concentration
- Decrease the magnesium concentration
- Redesign the primers to avoid non-specific binding
- Use a hot-start polymerase to prevent primer dimer formation
Poor Sequencing Traces
If the sequencing trace is of poor quality, the following troubleshooting steps should be considered:
- Increase or decrease the amount of template in the sequencing reaction
- Purify the PCR product more thoroughly
- Use a different sequencing primer
- Adjust the cycling conditions for the sequencing reaction
- Verify that the capillary electrophoresis instrument is performing correctly
Heterozygous Peak Imbalance
If the heterozygous peaks are not balanced, the following troubleshooting steps should be considered:
- Verify that the variant is not located near the primer binding site
- Check for allele-specific amplification due to a polymorphism in the primer binding site
- Repeat the sequencing with the opposite primer to confirm the genotype
- Consider the possibility of a mosaic variant or copy number alteration
Limitations of Sanger Validation
Detection Limits
Sanger sequencing has a detection limit of approximately 10-20% for variant allele fractions. Variants with allele fractions below this threshold may not be detected, producing false negative results. This limitation is particularly relevant for somatic variant validation, where tumor heterogeneity and stromal contamination can reduce allele fractions.
The study that assessed NGS variant validation using a secondary technology reported that high-quality variants with at least 35% heterozygous ratio were 100% confirmed [<a href="#ref-2">2</a>]. This finding suggests that variants with allele fractions below 35% may require additional consideration before Sanger validation is attempted.
Regions Not Amenable to Sanger Sequencing
Some genomic regions are not amenable to Sanger sequencing due to high GC content, repetitive sequence, or complex structure. These regions may produce poor quality traces or may not amplify by PCR. In such cases, alternative validation methods should be considered, including targeted deep resequencing, digital PCR, or mass spectrometry genotyping.
Interpretation Challenges
The interpretation of Sanger traces requires expertise and experience. Automated base calling software can misassign bases in regions with compressions, pull-up artifacts, or high background noise. The researcher must review the trace manually and make a judgment about the genotype call. This manual review introduces a subjective element that can lead to inter-observer variability.
Safety and Regulatory Context
Laboratory Safety
Standard molecular biology laboratory safety practices apply to Sanger validation workflows. These practices include the use of personal protective equipment, proper handling of biological samples, and safe disposal of chemical waste. The PCR and sequencing reactions involve the use of hazardous chemicals, including ethidium bromide for gel electrophoresis and formamide for capillary electrophoresis. Material safety data sheets should be reviewed before handling these chemicals.
Regulatory Considerations for Clinical Validation
In clinical laboratories, Sanger validation of NGS variants must comply with regulatory requirements for laboratory developed tests. These requirements include validation of the assay performance characteristics, documentation of the validation process, and participation in proficiency testing programs. The laboratory must establish the sensitivity, specificity, and accuracy of the Sanger assay for the intended use.
The study that validated a targeted gene panel for hereditary chronic liver diseases described the validation process for a clinical diagnostic assay [<a href="#ref-3">3</a>]. The study compared targeted sequencing with whole exome sequencing and confirmed variants by Sanger sequencing, demonstrating the importance of orthogonal validation in clinical diagnostics [<a href="#ref-3">3</a>].
Professional Escalation Criteria
The following situations warrant escalation to a senior researcher, laboratory director, or clinical geneticist:
- Discordant results between NGS and Sanger that cannot be resolved
- Variants with clinical significance that fail validation
- Evidence of sample mix-up or contamination
- Systematic failures in the validation workflow
- Unexpected results that suggest a problem with the NGS pipeline
The escalation should include a summary of the findings, the steps taken to resolve the discrepancy, and a recommendation for further action.
Practical Implementation Steps
Step 1: Select Variants for Validation
Review the VCF file and select variants for validation based on the filtering criteria described earlier. Prioritize variants with clinical significance, low quality scores, or unusual allele fractions. Document the selection criteria and the rationale for each variant.
Step 2: Extract Flanking Sequence
Extract the genomic sequence flanking each variant from the reference genome. The sequence should include at least 200 base pairs on each side of the variant. Use the NCBI databases to obtain the reference sequence and to verify the genomic coordinates [<a href="#ref-5">5</a>].
Step 3: Design Primers
Use Primer-BLAST or a similar tool to design primers for each variant. Verify the specificity of the primers against the reference genome. Document the primer sequences, amplicon size, and melting temperatures.
Step 4: Optimize PCR
Perform a gradient PCR to determine the optimal annealing temperature for each primer pair. Evaluate the PCR products by agarose gel electrophoresis. Adjust the reaction conditions as needed to produce a single, specific amplicon.
Step 5: Purify PCR Products
Purify the PCR products using enzymatic cleanup or column-based purification. Verify the quality and quantity of the purified products before proceeding to the sequencing reaction.
Step 6: Perform Cycle Sequencing
Set up the cycle sequencing reaction with the purified PCR product and the sequencing primer. Use the recommended cycling conditions for the sequencing chemistry. Purify the sequencing products to remove unincorporated ddNTPs.
Step 7: Run Capillary Electrophoresis
Load the purified sequencing products onto the capillary electrophoresis instrument. Run the instrument according to the manufacturer's instructions. Collect the electropherogram data.
Step 8: Analyze and Interpret
Analyze the electropherograms using the base calling software. Review the traces manually to confirm the genotype calls. Compare the Sanger results to the NGS calls and document the confirmation status.
Step 9: Report Results
Document the validation results in the laboratory records. Report confirmed variants, discordant variants, and failed validations. Escalate unresolved discrepancies to the appropriate personnel.
Bioinformatics Resources for Validation Workflows
NCBI Resources
The NCBI provides access to reference genomes, variation databases, and sequence analysis tools [<a href="#ref-5">5</a>]. The Primer-BLAST tool is essential for primer design, and the Variation Viewer can be used to examine variant context. The NCBI databases also provide access to reference sequences for primer design and variant annotation.
EMBL-EBI Training
The EMBL-EBI training portal offers learning pathways and practical analysis education for bioinformatics [<a href="#ref-6">6</a>]. These resources are valuable for researchers who need to build skills in sequence analysis, variant interpretation, and data management.
Bioconductor
Bioconductor provides packages and workflows for reproducible genomic analysis [<a href="#ref-9">9</a>]. The Bioconductor project offers official documentation for package installation and usage, which is relevant for researchers who want to integrate Sanger validation data with their NGS analysis pipelines.
Galaxy Training Network
The Galaxy Training Network provides accessible workflow training and analysis tutorials [<a href="#ref-7">7</a>]. The tutorials cover variant calling, quality control, and data visualization, which are relevant for understanding the NGS data that informs the validation process.
nf-core Documentation
The nf-core documentation describes community pipeline standards and usage [<a href="#ref-8">8</a>]. The nf-core pipelines provide reproducible analysis workflows that can be used for NGS variant calling, and the documentation supports the configuration and execution of these pipelines.
The Carpentries Lessons
The Carpentries lessons provide foundational training in computing, data, shell, Git, and programming [<a href="#ref-10">10</a>]. These skills are essential for managing the large data files generated by NGS and for automating the validation workflow.
Frequently Asked Questions
What is the difference between germline and somatic variant validation?
Germline variant validation confirms variants present in constitutional DNA, where the expected allele fraction is approximately 50% for heterozygous variants and 100% for homozygous variants. Somatic variant validation confirms variants present in tumor or diseased tissue, where the allele fraction may be lower due to tumor heterogeneity and stromal contamination. The validation strategy must account for these differences, particularly the detection limit of Sanger sequencing for low allele fraction variants.
How many variants should be selected for Sanger validation?
The number of variants selected for validation depends on the purpose of the study and the quality of the NGS data. For clinical validation, all clinically significant variants should be confirmed. For research studies, a representative sample of variants across quality score ranges and allele fractions may be sufficient. The study that assessed 7,601 NGS variants from 5,190 clinical samples provides a reference point for the scale of validation efforts in clinical settings [<a href="#ref-2">2</a>].
What is the minimum read depth for a variant to be considered high quality?
The study that assessed NGS variant validation reported that variants with at least 35x depth coverage and at least 35% heterozygous ratio were 100% confirmed by a secondary methodology [<a href="#ref-2">2</a>]. This threshold can serve as a guide for determining which variants may not require Sanger validation. However, the appropriate depth threshold depends on the sequencing platform, the variant calling pipeline, and the intended use of the results.
Can Sanger sequencing validate all types of variants?
Sanger sequencing can validate single nucleotide variants, small insertions and deletions, and selected structural variants when the breakpoints are known. The method is not suitable for validating large copy number changes, complex rearrangements, or variants in regions with high GC content or repetitive sequence. The study that established benchmark structural variant calls used PCR and Sanger sequencing to validate selected SVs, demonstrating the feasibility of this approach for certain structural variants [<a href="#ref-4">4</a>].
What should I do if the Sanger result does not match the NGS call?
If the Sanger result does not match the NGS call, the discrepancy should be investigated before reporting. Review the NGS alignment to confirm the variant call, examine the Sanger trace quality, and consider possible explanations such as primer binding site polymorphisms, allele-specific amplification, or sample mix-up. If the discrepancy persists, repeat the Sanger sequencing with different primers or use an alternative validation method.
How long does Sanger validation take?
The time required for Sanger validation depends on the number of variants, the efficiency of the PCR optimization, and the availability of sequencing capacity. A single variant can be validated in one to two days if the primers are already designed and the PCR conditions are established. A panel of variants may require one to two weeks, including primer design, PCR optimization, and sequencing.
What are the costs of Sanger validation?
The costs of Sanger validation include primer synthesis, PCR reagents, sequencing reagents, and instrument time. The cost per variant decreases as the number of variants increases because the PCR optimization and primer design costs are amortized. The cost of Sanger validation should be weighed against the cost of reporting an unvalidated variant that may be a false positive.
Is Sanger validation always necessary for NGS variants?
Sanger validation is not always necessary for high-quality variants from well-validated NGS workflows. The study that assessed 7,601 NGS variants found that high-quality variants with sufficient depth and heterozygous ratio were 100% confirmed by a secondary methodology [<a href="#ref-2">2</a>]. The study concluded that a new variant with high quality from a well-validated capture-based NGS workflow can be reported directly without validation [<a href="#ref-2">2</a>]. However, variants with clinical significance, low quality scores, or unusual characteristics should be validated by Sanger sequencing or an alternative method.
Related Bioinformatics Guides
- Oxford Nanopore Sequencing: From Sample to Base Calls
- From Raw Reads to Variants: A Diagnostic Blueprint for Next-Generation Sequencing (NGS) Workflows
- The Advent of Next-Generation Sequencing (NGS)
- Olink Proteomics: A Practical Guide to Panel Selection and Data Interpretation
- Digital Pathology Validation: A Practical Guide to CAP and RCPath Compliance
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
[1] [Validation and assessment of variant calling pipelines for next-generation sequencing.](https://pubmed.ncbi.nlm.nih.gov/25078893). Human genomics, 2014. [2] [A comprehensive assessment of Next-Generation Sequencing variants validation using a secondary technology.](https://pubmed.ncbi.nlm.nih.gov/31165590). Molecular genetics & genomic medicine, 2019. [3] [Validation of a targeted gene panel sequencing for the diagnosis of hereditary chronic liver diseases.](https://pubmed.ncbi.nlm.nih.gov/37388930). Frontiers in genetics, 2023. [4] [Robust Benchmark Structural Variant Calls of An Asian Using State-of-the-art Long-read Sequencing Technologies.](https://pubmed.ncbi.nlm.nih.gov/33662625). Genomics, proteomics & bioinformatics, 2022. [5] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information. [6] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute. [7] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project. [8] [nf-core Documentation](https://nf-co.re/docs). nf-core. [9] [Bioconductor](https://bioconductor.org/). Bioconductor Project. [10] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.