Troubleshooting Discordant Variant Calls: Why Sanger and NGS Results Differ and How to Resolve Them
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Discordant variant calls between Sanger and Next-Generation Sequencing (NGS) arise from fundamental differences in signal detection and sensitivity; Sanger relies on peak height ratios in a consensus trace, making low-frequency variants (e.g., <15-20% VAF) indistinguishable from noise, while NGS quantifies individual reads, enabling detection of variants present in a small fraction of molecules (e.g., 5% VAF).
- Biological causes of discordance include allelic dropout, where variants in Sanger primer binding sites prevent amplification of one allele leading to a false homozygous call, and mosaic variants or low-level heteroplasmy (e.g., mitochondrial) present in a subset of cells, which are readily detectable by NGS's higher sensitivity but missed by Sanger.
- Technical factors contributing to discordance include coverage gaps and low depth in NGS, library preparation artifacts (e.g., polymerase slippage in repetitive regions like RPGR ORF15), PCR errors, and distinct error profiles of each platform (e.g., dye blobs in Sanger, mapping errors in NGS).
- Bioinformatics pipeline differences, including choice of alignment software, variant callers, and filtering thresholds, can lead to variant calls being present in raw data but filtered out by one pipeline, necessitating review of raw alignments and adjustment of filters.
- Resolving discordance involves a tiered approach: Tier 1 data review (quality assessment, raw alignment inspection, primer site checks), Tier 2 targeted wet lab confirmation (primer redesign, deep resequencing), and Tier 3 orthogonal methods (e.g., droplet digital PCR) when initial steps fail.
- Documentation of discordant calls, including sample details, assay conditions, coverage, VAF, and resolution outcome, is crucial for measuring discordance rates, establishing quality thresholds, and identifying recurring issues in variant validation workflows.
Researchers frequently encounter a situation where a variant confidently called by next-generation sequencing (NGS) is not confirmed by Sanger sequencing, or the reverse occurs. These discordant results can delay projects, undermine confidence in data, and complicate clinical reporting. This article explains the biological, technical, and analytical reasons why Sanger and NGS results differ and provides a systematic approach to determining the true genotype.
The scope covers both germline and somatic variant calling workflows, with emphasis on practical troubleshooting steps that laboratory professionals and researchers can apply immediately. The guidance applies to targeted gene panels, exome sequencing, and amplicon-based assays where Sanger validation is commonly used as a confirmatory method.
Understanding the Core Differences Between Sanger and NGS Platforms
Sanger sequencing and NGS operate on fundamentally different principles, and these differences explain many discordant results. Sanger sequencing produces a single consensus trace representing the average signal from many template molecules. NGS generates millions of individual reads that can be analyzed independently, allowing detection of variants present in a minority of molecules.
Signal Detection and Sensitivity
Sanger sequencing detects variants by measuring fluorescent signal peaks at each position along the sequenced fragment. The relative height of peaks indicates the proportion of each base in the template population. Heterozygous variants typically appear as two peaks at roughly equal height. Low-level variants, however, may be indistinguishable from background noise.
NGS detects variants by counting individual reads that carry the alternate allele. This digital counting approach provides quantitative information about variant allele frequency (VAF). A variant present in 5% of reads is clearly detectable with sufficient coverage, whereas the same variant in a Sanger trace might be dismissed as noise.
The sensitivity difference between platforms is well documented. In a comparison of exome sequencing and Sanger sequencing across 258 genes, Sanger sequencing achieved 99% sensitivity for true variant detection while NGS experiments achieved 97 to 100% sensitivity depending on the enrichment kit used. Both methods produced false-positive and false-negative calls, demonstrating that neither platform is infallible. The study authors emphasized that high clinical suspicion for a specific diagnosis should override negative results from either method alone.
Read Length and Phasing Information
Sanger sequencing produces reads of 500 to 1000 bases with information about the physical linkage of variants along a single molecule. NGS read lengths vary by platform, with short-read technologies typically producing 150 base pair reads. This difference matters when variants are located in repetitive regions or when phase information is needed to determine whether two variants are on the same chromosome.
For example, in the RPGR ORF15 region, which contains highly repetitive sequence, standard NGS library preparation methods produced discordant results in 24 cases due to false negatives, incorrectly called variants, and a false positive. The discordance was resolved by changing the library preparation method to one that produced more complete coverage of the repetitive region. This case illustrates that sequence context, beyond platform choice, determines whether Sanger and NGS will agree.
Error Profiles and Artifacts
Each platform has characteristic error patterns. Sanger sequencing errors often arise from dye blobs, secondary structures, and homopolymer regions that cause peak broadening. NGS errors include base substitution errors from polymerase incorporation, mapping errors in repetitive regions, and alignment artifacts.
In formalin-fixed, paraffin-embedded (FFPE) tumor specimens, the background noise of variant detection was approximately twofold higher compared with cell line DNA. This elevated noise required optimization of bioinformatic algorithms to accommodate the increased error rate. Researchers working with FFPE samples should expect higher discordance rates and should adjust their validation strategies accordingly.
At a Glance: Common Causes of Discordance
| Cause Category | Specific Cause | Typical Direction of Discordance | Primary Resolution Strategy |
|---|---|---|---|
| Allelic dropout | Primer binding site variant prevents PCR amplification of one allele | Variant present in NGS but absent in Sanger | Redesign primers in conserved regions, verify primer binding sites |
| Low variant allele frequency | Somatic variant present in small fraction of cells | Variant present in NGS but absent or ambiguous in Sanger | Use NGS with high depth, confirm with sensitive orthogonal method |
| Repetitive or GC-rich regions | Polymerase slippage, sequencing artifacts, poor mapping | Variant calls differ in either direction | Use long-range PCR, optimized library preparation, manual review |
| Sample quality issues | FFPE damage, DNA degradation, contamination | False positives or false negatives in either platform | Assess quality metrics, use FFPE-optimized workflows, repeat extraction |
| Bioinformatics filtering | Overly stringent or lenient variant filters | Variant present in raw data but filtered in one pipeline | Review raw alignments, adjust filters, use multiple callers |
Biological Causes of Discordance
Allelic Dropout and Primer Binding Site Variants
Allelic dropout occurs when a variant in a primer binding site prevents PCR amplification of that allele. If the primer fails to anneal due to sequence mismatch, only the other allele is amplified and sequenced. The result is a false homozygous call in Sanger sequencing while NGS, which may use different primers or capture probes, correctly identifies the heterozygous variant.
This mechanism was described in the context of STR marker characterization, where variants in the flanking region of PCR amplicons produced null alleles. The researchers noted that sequencing could identify genomic variations within PCR amplicons, particularly variants that result in null alleles and alleles that do not migrate within allele sizing bins provided by kit manufacturers. When validating NGS variants by Sanger sequencing, always check whether the Sanger primers overlap any known variants in the region.
Mosaic Variants and Low-Level Heteroplasmy
Mosaic variants arise from post-zygotic mutations and are present in only a fraction of cells. The variant allele frequency depends on when the mutation occurred during development and which tissues are examined. A variant present in 10% of blood cells may be undetectable by Sanger sequencing but clearly visible in NGS data with adequate depth.
Mitochondrial heteroplasmy provides a well-documented example. In a study comparing NGS and Sanger-type sequencing of mitochondrial genomes, 99.9996% concordance was observed across the full mitochondrial genome. The only discordant calls involved low-level point heteroplasmies, with differences resulting from stochastic variation and the increased sensitivity of NGS. The higher sensitivity of NGS also allowed detection of a mixed sample that was not detected by Sanger sequencing.
Copy Number Variation and Large Deletions
Standard variant calling pipelines focus on single nucleotide variants and small insertions or deletions. Larger structural variants, including exon-level deletions and duplications, are often missed by both Sanger sequencing and standard NGS analysis. When a deletion removes a primer binding site or an entire exon, Sanger sequencing may amplify only the normal allele while NGS read depth analysis reveals the copy number change.
In the RPGR ORF15 study, duplication analysis using an in-silico array was needed to supplement the NGS variant detection. This additional analysis resolved discordance between Sanger and NGS data in all cases. Researchers investigating inherited retinal dystrophy should consider copy number analysis as part of their variant confirmation workflow.
Technical Causes of Discordance
Coverage Gaps and Low Depth Regions
NGS coverage is not uniform across the genome or targeted regions. Some regions consistently show low coverage due to GC content, repetitive sequence, or capture probe design. If a variant falls in a low coverage region, the NGS call may be unreliable even if the variant passes quality filters.
The exome sequencing comparison study found that NGS mean coverage exceeded 20x for more than 98% of regions targeted by one enrichment kit and more than 91% for another kit. This means that 2 to 9% of targeted regions had suboptimal coverage. Variants in these regions are more likely to produce discordant results when compared with Sanger sequencing.
Library Preparation Artifacts
The method used to fragment DNA and prepare sequencing libraries can introduce artifacts that affect variant calling. In the RPGR ORF15 study, the Nextera library preparation method produced 24 discordant cases. The discordance was resolved by switching to an enzymatic fragmentation method that produced more complete coverage of the highly repetitive region. The improvement in variant detection accuracy was largely attributed to improvement in random fragmentation offered by the enzymatic method.
Researchers should document the library preparation method used for each sample and be aware that changing methods can affect variant detection in difficult regions. When troubleshooting discordant calls, consider whether the library preparation method is appropriate for the sequence context of the variant.
PCR Errors and Amplification Bias
Both Sanger and NGS workflows typically include PCR amplification steps. PCR errors can introduce base substitutions that are subsequently detected as variants. Polymerase errors are more likely in GC-rich regions and homopolymer stretches. Amplification bias can also skew the representation of alleles, particularly when the two alleles have different GC content.
For somatic variant calling, the error rate is compounded by the low input DNA amounts often available from tumor specimens. The FFPE study noted that background noise was elevated approximately twofold in FFPE DNA compared with cell line DNA. This elevated noise can lead to false-positive variant calls that are not confirmed by Sanger sequencing.
Bioinformatics Pipeline Differences
The choice of alignment software, variant caller, and filtering thresholds can produce different results from the same raw sequencing data. Each variant caller uses different statistical models and quality scoring systems. A variant that passes filters in one pipeline may be rejected by another.
The mitochondrial genome study demonstrated that variant calls were reproducible between sequencing sets and between software analysis versions, with variant frequency differing by only 0.23% between sequencing sets and 0.01% between software versions. This high reproducibility was achieved with a standardized analysis protocol. Laboratories that use different pipelines for research and clinical samples may observe higher discordance rates.
Practical Workflow for Resolving Discordant Calls
Step 1: Verify the Quality of Both Datasets
Before investigating biological causes, confirm that both the NGS and Sanger data meet quality standards. For NGS data, check the coverage at the variant position, the mapping quality of reads supporting the variant, and the variant quality score. For Sanger data, examine the trace quality at the variant position, looking for overlapping peaks, high background, or other artifacts.
The Galaxy Training Network provides accessible tutorials on quality assessment for sequencing data. These resources can help researchers systematically evaluate whether poor data quality explains the discordance.
Step 2: Review the Raw Alignments
Open the NGS alignment file in a genome browser and visually inspect the region around the variant. Look for the following patterns:
- Reads with mismatches clustered at read ends, suggesting alignment artifacts
- Variants supported only by reads with low mapping quality
- Strand bias, where the variant appears only on one strand
- Variants in homopolymer regions or near indels
For Sanger data, review the electropherogram to confirm that the variant call was made correctly. Check whether the variant is present in both forward and reverse traces if both were generated.
Step 3: Check Primer Binding Sites
If the Sanger assay failed to detect a variant that is clearly present in NGS data, examine the Sanger primer sequences. Use a genome browser or primer analysis tool to check whether any known variants fall within the primer binding sites. The NCBI Data Resources provides access to dbSNP and other variant databases that can be used to check for common polymorphisms in primer regions.
If a variant is found in a primer binding site, redesign the primers in a conserved region and repeat the Sanger sequencing. This approach resolves many cases of allelic dropout.
Step 4: Assess Variant Allele Frequency
For somatic samples, determine whether the variant allele frequency is consistent with the expected biology. A variant present at 5% VAF in a tumor sample with 50% tumor content may be real, but it will be difficult to confirm by Sanger sequencing. The FFPE study found that three discordant calls between NGS platforms were each represented at less than 10% of reads.
If the VAF is low, consider using a more sensitive orthogonal method for confirmation. Options include droplet digital PCR, targeted deep resequencing, or amplicon-based NGS with high depth. The Bioconductor project provides packages for analyzing low-frequency variants and assessing the statistical confidence of variant calls.
Step 5: Consider the Sequence Context
Variants in repetitive regions, GC-rich regions, or regions with high sequence homology to other genomic locations are more likely to produce discordant results. The RPGR ORF15 study demonstrated that highly repetitive regions require specialized library preparation and analysis approaches.
For variants in difficult regions, consider the following approaches:
- Use long-range PCR followed by Sanger sequencing to span the entire repetitive region
- Use an NGS library preparation method with enzymatic fragmentation instead of transposase-based methods
- Perform manual review of the alignment to confirm that reads are uniquely mapped
Step 6: Repeat the Assay
Sometimes discordance results from a one-time technical failure. Repeat both assays if possible. For NGS, this may mean resequencing the sample or running a targeted amplicon assay. For Sanger, repeat the PCR and sequencing reaction with fresh reagents.
The nf-core documentation describes standards for reproducible bioinformatics workflows. Following these standards can reduce run-to-run variability and make it easier to distinguish technical failures from true biological differences.
Step 7: Use an Orthogonal Method
When Sanger and NGS disagree and the cause is not immediately apparent, use a third method to determine the true genotype. Options include:
- Droplet digital PCR for targeted variant confirmation
- Pyrosequencing for quantitative allele detection
- Restriction fragment length polymorphism analysis if the variant creates or destroys a restriction site
- Cloning and sequencing individual molecules to determine phase and confirm low-level variants
The choice of orthogonal method depends on the variant type, the expected allele frequency, and the resources available in the laboratory.
Records and Measurements for Discordance Tracking
Documenting Discordant Calls
Maintain a systematic record of all discordant calls encountered during variant validation. For each discordant case, record the following information:
- Sample identifier and tissue type
- Gene and variant coordinates
- NGS platform and pipeline version
- Sanger assay conditions and primer sequences
- Coverage and VAF from NGS data
- Trace quality metrics from Sanger data
- Resolution outcome and final genotype determination
This documentation supports quality improvement efforts and helps identify recurring causes of discordance. The EMBL-EBI Training resources provide guidance on data management practices that support reproducible analysis.
Measuring Discordance Rates
Calculate the discordance rate for each assay or pipeline as the number of discordant calls divided by the total number of variants validated. Track this metric over time to identify changes that may result from reagent lot changes, equipment maintenance, or pipeline updates.
The exome sequencing comparison study provides a useful benchmark. Of 449 variants identified in at least one experiment, 407 (90.6%) were detected by all methods. The 42 discordant variants included 23 that were determined to be true calls. This means that approximately 9% of variants showed some discordance across methods, with roughly half of the discordant variants being true calls that were missed by one method.
Establishing Quality Thresholds
Use the discordance data to establish quality thresholds for both NGS and Sanger assays. For NGS, determine the minimum coverage and VAF required for confident variant calls in your sample types. For Sanger, establish criteria for trace quality that indicate reliable base calling.
The mitochondrial genome study provides an example of rigorous quality assessment. Both sequencing sets resulted in 99.998% of positions with greater than 10x coverage when 96 samples were multiplexed. This level of completeness required careful optimization of the sequencing workflow.
Common Failure Patterns and Their Resolutions
Pattern 1: Variant Present in NGS but Absent in Sanger
This pattern most commonly results from one of three causes:
- Low variant allele frequency: The variant is present in a small fraction of cells and falls below the detection limit of Sanger sequencing. Confirm by checking the VAF in NGS data and considering whether the expected biology supports a low VAF.
- Allelic dropout: A variant in the Sanger primer binding site prevents amplification of the variant allele. Check primer binding sites for known variants and redesign primers if needed.
- NGS false positive: The variant is a sequencing or alignment artifact. Review the raw alignments for evidence of mapping errors, strand bias, or other artifacts.
Pattern 2: Variant Present in Sanger but Absent in NGS
This pattern can result from:
- Coverage gap in NGS: The variant falls in a region with insufficient NGS coverage. Check the coverage at the variant position and consider whether the region is difficult to sequence.
- Overly stringent NGS filters: The variant is present in the raw data but was filtered out by the variant calling pipeline. Review the raw alignments and adjust filters if appropriate.
- Sanger false positive: The variant call in Sanger is an artifact of poor trace quality. Review the electropherogram for evidence of background noise or mixed signals.
Pattern 3: Discordant Genotype Calls
When both methods detect the variant but assign different genotypes, the cause is often related to allele-specific amplification or sequencing bias. For example, if one allele amplifies more efficiently than the other, the Sanger trace may show a skewed peak ratio that is interpreted as a homozygous call.
For NGS, allele-specific bias can result from capture probe design or PCR amplification during library preparation. The Bioconductor project provides tools for detecting and correcting allele-specific bias in sequencing data.
Pattern 4: Discordance in Repetitive Regions
Variants in repetitive regions present unique challenges for both platforms. Sanger sequencing may produce unreadable traces due to polymerase slippage. NGS may produce alignment errors or coverage gaps.
The RPGR ORF15 study provides a model for addressing this challenge. The researchers tested multiple library preparation methods and selected one that produced complete coverage of the repetitive region. They also supplemented the analysis with duplication analysis to detect copy number changes.
Somatic Variant Calling Considerations
Tumor Content and Heterogeneity
Somatic variant calling from tumor specimens is complicated by the mixture of tumor and normal cells in the sample. The tumor content determines the maximum VAF for somatic variants. A variant present in all tumor cells will have a VAF approximately equal to half the tumor content for heterozygous variants.
When validating somatic variants by Sanger sequencing, consider whether the expected VAF is within the detection range of Sanger. The FFPE study found that NGS could detect mutations at 1% allele frequency with optimized algorithms. Sanger sequencing typically requires VAF above 15 to 20% for reliable detection.
FFPE Sample Artifacts
Formalin fixation introduces crosslinks and base damage that can create sequencing artifacts. The FFPE study found that background noise was elevated approximately twofold in FFPE DNA compared with cell line DNA. This elevated noise can produce false-positive variant calls, particularly C to T transitions that are characteristic of formalin damage.
When working with FFPE samples, use the following practices:
- Assess DNA quality before library preparation
- Use FFPE-optimized library preparation kits
- Apply FFPE-specific variant filtering criteria
- Confirm clinically significant variants with an orthogonal method
Clonal Hematopoiesis
In blood-derived DNA samples, clonal hematopoiesis can produce somatic variants that are present in a fraction of blood cells. These variants may be detected by NGS but not by Sanger sequencing. When validating variants from blood samples, consider whether the variant may represent clonal hematopoiesis instead of a germline variant.
Germline Variant Calling Considerations
Confirming Germline Status
Germline variants are expected to be present at approximately 50% VAF for heterozygous calls and 100% VAF for homozygous calls. Variants with VAF significantly different from these expected values should be investigated further.
The exome sequencing comparison study provides context for germline variant validation. The study compared Sanger sequencing results of 258 genes to NGS results from two exome enrichment kits using DNA from a single individual. The high overall concordance between Sanger and NGS performances supports the use of either method for germline variant detection.
De Novo Variants
De novo variants are present in the affected individual but not in either parent. Confirming de novo status requires sequencing both parents. When a de novo variant is suspected, confirm the variant in the proband by an orthogonal method and then test both parents.
Variant Filtering in Germline Pipelines
Germline variant calling pipelines typically apply filters for quality score, depth, and allele balance. Overly stringent filters can remove true variants, while overly lenient filters can produce false positives. The Galaxy Training Network provides tutorials on germline variant calling and filtering that can help researchers optimize their pipelines.
Quality Control and Reproducibility
Controls and Reference Materials
Include appropriate controls in every sequencing run. Positive controls with known variants confirm that the assay can detect variants in the expected regions. Negative controls confirm that the assay does not produce false positives.
The NCBI Data Resources provides access to reference materials and control samples that can be used for assay validation. These resources support the development of robust quality control programs.
Replicate Analysis
Run replicate samples to assess reproducibility. The mitochondrial genome study demonstrated high reproducibility between replicate sequencing sets, with variant frequency differing by only 0.23%. This level of reproducibility requires standardized protocols and careful quality control.
For clinical samples, consider running replicates when discordant results are obtained. A replicate run can distinguish between stochastic variation and systematic errors.
Pipeline Version Control
Document the exact versions of all software used in the analysis pipeline. Pipeline updates can change variant calls, even when the underlying sequencing data are identical. The nf-core documentation describes standards for pipeline versioning and reproducibility that can be applied to any bioinformatics workflow.
The Carpentries lessons provide training on version control and reproducible analysis practices. These skills are essential for maintaining consistent variant calling across projects and over time.
Limitations of Discordance Resolution
When Discordance Cannot Be Resolved
Some discordant calls cannot be resolved definitively. This situation arises when:
- The variant falls in a region that cannot be reliably sequenced by any available method
- The sample quantity is insufficient for additional testing
- The variant is present at a level near the detection limit of all available methods
In these cases, document the discordance and the efforts made to resolve it. For clinical samples, report the uncertainty to the requesting clinician and recommend appropriate follow-up testing.
Interpretation of Low-Level Variants
Low-level variants present a particular challenge for interpretation. A variant present at 2% VAF may be a true somatic mutation, a clonal hematopoiesis event, or a sequencing artifact. The biological significance of the variant depends on the clinical context.
The EMBL-EBI Training resources provide guidance on variant interpretation and the use of population databases for assessing variant significance. These resources support evidence-based interpretation of low-level variants.
Reporting Uncertain Results
When discordant results cannot be resolved, report the uncertainty clearly. For research studies, document the discordance in the methods section and consider sensitivity analyses that exclude discordant variants. For clinical samples, follow laboratory policies for reporting uncertain results and recommend appropriate follow-up testing.
Professional Escalation Criteria
When to Consult a Bioinformatics Specialist
Seek assistance from a bioinformatics specialist when:
- The cause of discordance is not apparent after completing the troubleshooting steps
- The discordance involves multiple variants in the same sample or region
- The discordance pattern suggests a systematic problem with the analysis pipeline
- The variant is in a difficult-to-sequence region that requires specialized analysis
When to Consult a Clinical Geneticist
For clinical samples, consult a clinical geneticist when:
- The discordant variant has potential clinical significance
- The discordance affects the interpretation of a genetic test result
- Additional testing is needed to resolve the discordance
- The result may affect patient management decisions
When to Consider Alternative Technologies
Consider alternative technologies when:
- The variant is in a region that cannot be reliably sequenced by current methods
- The variant is present at a level below the reliable detection limit of available methods
- The discordance persists despite multiple troubleshooting attempts
- The clinical significance of the variant warrants investment in specialized testing
The Bioconductor project and Galaxy Training Network provide resources for learning about alternative analysis approaches and emerging technologies.
Building a Discordance Decision Framework for Your Laboratory
A structured decision framework transforms discordance troubleshooting from an ad hoc exercise into a repeatable process that produces defensible genotype determinations. The framework below organizes the technical and biological considerations from the preceding sections into a sequence of decisions that any laboratory technician or researcher can follow without requiring deep bioinformatics expertise. The framework prioritizes actions by cost, time, and information value, so the least expensive and most informative steps are completed first.
Framework Overview and Guiding Principles
The framework operates on three principles derived from published concordance data. First, both platforms produce false positives and false negatives, so neither result should be treated as automatically correct. The exome sequencing comparison study found that Sanger sequencing had a mean false-positive rate of 3.7E-6 while the two NGS enrichment methods had rates of 2.5E-6 and 5.2E-6, demonstrating that errors occur on both platforms at comparable frequencies. Second, the true genotype is determined by weighing all available evidence, not by defaulting to one platform. Third, the resolution pathway depends on the variant class, the sample type, and the clinical or research context, so the framework branches accordingly.
The framework uses a tiered approach. Tier 1 actions are data review steps that require no additional wet lab work. Tier 2 actions involve targeted wet lab confirmation. Tier 3 actions require orthogonal methods or specialized analysis. Most discordant calls are resolved at Tier 1 or Tier 2, which keeps costs low and turnaround times short.
Tier 1: Data Review and Quality Assessment
Begin every discordance investigation by completing the following steps in order. Document each step and its outcome in the discordance log described later in this section.
Step 1A: Confirm the variant coordinates and annotation. Verify that both the NGS and Sanger results refer to the same genomic position and the same transcript or genome build. Discrepancies in genome build versions or transcript annotations are a common but easily overlooked cause of apparent discordance. Check the chromosome, position, reference allele, and alternate allele against the current genome build using NCBI Data Resources to confirm the coordinates match.
Step 1B: Assess NGS data quality at the variant position. Record the depth, variant allele frequency, mapping quality, and variant quality score from the NGS pipeline. Apply the following thresholds as initial screening criteria. For germline variants, require a minimum depth of 20 reads and a VAF between 25% and 75% for heterozygous calls. For somatic variants, require a minimum depth of 100 reads and a VAF above the laboratory validated detection limit. If the variant fails these thresholds, the NGS call is suspect and the discordance may be explained by insufficient evidence in the NGS data.
Step 1C: Assess Sanger trace quality at the variant position. Examine the electropherogram for the region surrounding the variant. Look for overlapping peaks, high background fluorescence, dye blobs, or compressed peaks that indicate secondary structure. A clean Sanger trace with well separated peaks and low background supports the Sanger call. A noisy trace with ambiguous peaks weakens the Sanger call.
Step 1D: Review the raw NGS alignments. Open the alignment file in a genome browser and visually inspect the reads supporting the variant. Look for the following red flags. Variants supported only by reads with low mapping quality suggest mapping artifacts. Variants showing strong strand bias, where the alternate allele appears predominantly on one strand, suggest sequencing artifacts. Variants clustered at read ends suggest alignment errors. Variants in homopolymer regions or adjacent to indels suggest polymerase errors. The Galaxy Training Network provides tutorials on visual alignment review that are useful for researchers who are not yet familiar with this step.
Step 1E: Check primer binding sites for the Sanger assay. Obtain the primer sequences used for the Sanger reaction and check whether any known variants fall within the primer binding regions. Use dbSNP through NCBI Data Resources to identify common polymorphisms in the primer regions. If a variant is present in a primer binding site, allelic dropout is the likely cause of discordance.
Tier 2: Targeted Wet Lab Confirmation
If Tier 1 steps do not resolve the discordance, proceed to targeted wet lab experiments. These experiments are designed to address the most likely causes identified during data review.
Step 2A: Redesign Sanger primers in conserved regions. If the original Sanger primers overlap known variants or fall in regions with poor conservation, design new primers in flanking conserved sequences. Verify the new primer binding sites do not contain known variants before ordering. Repeat Sanger sequencing with the redesigned primers. This step resolves most cases of allelic dropout.
Step 2B: Repeat NGS analysis with relaxed filters. If the NGS variant was flagged as low quality or filtered by the pipeline, rerun the variant caller with relaxed thresholds to determine whether the variant is present in the raw data. Review the raw alignments to assess whether the variant is supported by genuine reads or represents an artifact. The Bioconductor project provides packages for exploring variant calls and adjusting filtering parameters in a reproducible manner.
Step 2C: Perform targeted deep resequencing. For variants where the NGS depth is borderline or the VAF is near the detection limit, design a targeted amplicon assay that sequences the variant region at high depth. This approach provides additional evidence about whether the variant is real and what its true allele frequency is. The RPGR ORF15 study demonstrated that changing the library preparation method to improve coverage in difficult regions resolved all discordant cases, highlighting the value of resequencing with optimized methods.
Step 2D: Test both DNA strands with Sanger sequencing. If the original Sanger reaction used only one primer, repeat the reaction with the reverse primer to obtain sequence from both strands. This step can resolve artifacts that appear on only one strand and provides additional evidence for the genotype call.
Tier 3: Orthogonal Methods and Specialized Analysis
When Tier 1 and Tier 2 steps do not resolve the discordance, or when the variant has clinical significance that warrants additional confirmation, use an orthogonal method.
Step 3A: Droplet digital PCR for quantitative variant confirmation. Droplet digital PCR partitions the DNA sample into thousands of individual reactions and counts the number of droplets containing the variant allele. This method provides absolute quantification of the variant allele frequency without relying on amplification efficiency or sequencing chemistry. It is particularly useful for confirming low-level somatic variants and for determining whether a variant is present at a level consistent with mosaicism or clonal hematopoiesis.
Step 3B: Long-range PCR followed by Sanger sequencing. For variants in repetitive regions or regions with complex structure, design long-range PCR primers that span the entire difficult region. Sequence the long-range PCR product with internal primers to obtain coverage across the variant. This approach was used successfully in the RPGR ORF15 study, where long-range PCR products were fragmented and sequenced by NGS to resolve discordant calls.
Step 3C: Cloning and sequencing individual molecules. For variants where phase information is needed or where the variant is present at very low levels, clone individual PCR products into a plasmid vector and sequence multiple clones. This approach provides definitive evidence about the presence and phase of variants but is labor intensive and time consuming.
Step 3D: Copy number analysis. If the discordance involves a region where copy number variation is suspected, perform copy number analysis using the NGS depth data or a dedicated assay. The RPGR ORF15 study supplemented NGS variant detection with duplication analysis using an in-silico array, which resolved discordance in all cases. Copy number changes can cause apparent discordance when one allele is deleted and the remaining allele is sequenced as homozygous.
Decision Points and Branching Logic
The framework uses explicit decision points to guide the investigator through the resolution process.
Decision Point 1: Are the variant coordinates and annotations consistent between platforms? If no, correct the annotation discrepancy and re-evaluate. If yes, proceed to Decision Point 2.
Decision Point 2: Does the NGS data meet quality thresholds at the variant position? If no, the NGS call is suspect. Proceed to Tier 2 Step 2B to determine whether the variant is present in raw data. If yes, proceed to Decision Point 3.
Decision Point 3: Does the Sanger trace meet quality standards at the variant position? If no, the Sanger call is suspect. Repeat Sanger sequencing with fresh reagents or redesigned primers. If yes, proceed to Decision Point 4.
Decision Point 4: Are there known variants in the Sanger primer binding sites? If yes, allelic dropout is likely. Redesign primers and repeat Sanger sequencing. If no, proceed to Decision Point 5.
Decision Point 5: Is the variant allele frequency consistent with the expected biology? For germline variants, the VAF should be approximately 50% for heterozygous calls. For somatic variants, the VAF should be consistent with the tumor content and the expected clonality. If the VAF is unexpectedly low, consider mosaicism, clonal hematopoiesis, or tumor heterogeneity. If the VAF is consistent, proceed to Decision Point 6.
Decision Point 6: Is the variant in a difficult-to-sequence region? If yes, consider specialized library preparation, long-range PCR, or orthogonal methods. If no, proceed to Tier 3 orthogonal confirmation.
Record System for Discordance Tracking
A systematic record system is essential for identifying recurring causes of discordance and for demonstrating the reliability of variant confirmation procedures. The following record structure captures the information needed for quality improvement and for responding to inquiries about specific variant calls.
Discordance Log Fields
For each discordant call, record the following fields in a laboratory notebook or electronic database:
- Unique case identifier
- Sample identifier and tissue type
- Gene name and variant coordinates using the current genome build
- Reference allele and alternate allele
- NGS platform and pipeline version
- NGS depth and variant allele frequency at the variant position
- Sanger assay identifier and primer sequences
- Sanger trace quality assessment
- Tier 1 findings including alignment review results and primer binding site check
- Tier 2 actions taken and results
- Tier 3 actions taken and results
- Final genotype determination and the evidence supporting it
- Date of resolution and investigator name
Quarterly Discordance Summary
Aggregate the discordance log quarterly to calculate the following metrics:
- Total variants validated by Sanger sequencing
- Number and percentage of discordant calls
- Discordance rate by variant class, gene, and sample type
- Most common causes of discordance
- Median time to resolution
- Resolution rate by tier
These metrics identify systematic problems such as recurring primer design issues, problematic genomic regions, or pipeline errors. The EMBL-EBI Training resources provide guidance on data management practices that support this type of quality improvement analysis.
Common Failure Patterns in Framework Application
Several failure patterns recur when laboratories implement a discordance resolution framework.
Pattern 1: Skipping Tier 1 steps. Laboratories that proceed directly to wet lab confirmation without completing data review waste time and resources. Most discordant calls are resolved by careful data review, particularly by checking primer binding sites and reviewing raw alignments.
Pattern 2: Treating one platform as the gold standard. The exome sequencing comparison study demonstrated that both Sanger and NGS produce false positives and false negatives. Laboratories that automatically trust one platform over the other will misclassify a portion of discordant calls.
Pattern 3: Failing to document the resolution process. Without documentation, the laboratory cannot learn from discordant calls or demonstrate the reliability of its variant confirmation procedures. The discordance log is essential for quality improvement.
Pattern 4: Applying germline thresholds to somatic samples. Somatic variant calling requires different quality thresholds than germline calling due to the lower allele frequencies and elevated background noise in tumor samples. The FFPE study found that background noise was approximately twofold higher in FFPE DNA compared with cell line DNA, requiring optimized bioinformatic algorithms.
Professional Escalation Criteria Within the Framework
The framework includes explicit criteria for escalating a discordance investigation to a specialist.
Escalate to a bioinformatics specialist when:
- Tier 1 data review reveals complex alignment patterns that require specialized analysis
- The discordance involves multiple variants in the same sample or genomic region
- The discordance pattern suggests a systematic pipeline error
- The variant is in a region requiring specialized analysis tools
Escalate to a clinical geneticist when:
- The discordant variant has potential clinical significance
- The discordance affects the interpretation of a genetic test result
- Additional testing is needed to resolve the discordance
- The result may affect patient management decisions
Consider alternative technologies when:
- The variant is in a region that cannot be reliably sequenced by current methods
- The variant is present at a level below the reliable detection limit of available methods
- The discordance persists despite completing all three tiers of the framework
- The clinical significance of the variant warrants investment in specialized testing
The nf-core documentation describes standards for reproducible bioinformatics workflows that support consistent application of the framework across different samples and time points. The Carpentries lessons provide training on the computing skills needed to implement and maintain the record system described above.
Frequently Asked Questions
Why does Sanger sequencing fail to confirm a variant that is clearly present in NGS data?
The most common causes are low variant allele frequency, allelic dropout due to primer binding site variants, and NGS false positives from sequencing or alignment artifacts. Check the VAF in the NGS data, review the Sanger primer binding sites for known variants, and inspect the raw NGS alignments for evidence of artifacts.
Can Sanger sequencing detect variants at low allele frequency?
Sanger sequencing has limited sensitivity for low-level variants. Variants present at less than 15 to 20% allele frequency may be difficult to distinguish from background noise. The exome sequencing comparison study found that Sanger sequencing had 99% sensitivity for true variant detection, but this was in the context of germline variants where allele frequencies are typically 50% or 100%.
What is allelic dropout and how does it cause discordant results?
Allelic dropout occurs when a variant in a PCR primer binding site prevents the primer from annealing, so only the other allele is amplified. This produces a false homozygous call. NGS may use different primers or capture probes that do not overlap the variant, allowing correct detection of the heterozygous genotype.
How should I handle discordant results in FFPE tumor samples?
FFPE samples have elevated background noise compared with fresh samples. Use FFPE-optimized library preparation and variant calling workflows. Confirm clinically significant variants with an orthogonal method. The FFPE study found that background noise was approximately twofold higher in FFPE DNA compared with cell line DNA.
What is the best way to confirm a somatic variant present at low allele frequency?
Use a sensitive orthogonal method such as droplet digital PCR or targeted deep resequencing. Sanger sequencing is generally not suitable for confirming variants below 15 to 20% allele frequency. The FFPE study demonstrated that NGS could detect mutations at 1% allele frequency with optimized algorithms.
Why do NGS and Sanger results disagree in repetitive regions?
Repetitive regions are difficult to sequence with both platforms. Sanger sequencing may produce unreadable traces due to polymerase slippage. NGS may produce alignment errors or coverage gaps. The RPGR ORF15 study demonstrated that specialized library preparation methods can improve coverage in repetitive regions.
Should I always validate NGS variants by Sanger sequencing?
The need for Sanger validation depends on the context. For research studies, validation may be limited to variants used for downstream experiments. For clinical samples, validation policies vary by laboratory and regulatory requirements. The exome sequencing comparison study found high overall concordance between Sanger and NGS, suggesting that validation can be targeted to specific variant classes or regions.
How can I reduce the frequency of discordant calls in my workflow?
Use standardized protocols for library preparation and sequencing, document pipeline versions, include appropriate controls, and establish quality thresholds for both NGS and Sanger data. Track discordance rates over time to identify systematic problems. The nf-core documentation and Galaxy Training Network provide resources for establishing reproducible workflows.
Related Bioinformatics Guides
- RNA-Seq vs qPCR: Validation and Comparison
- How to Interpret Gene Set Enrichment Analysis Results
- Genomic Data Analysis Tools: A Comparative Guide for Researchers
- Oxford Nanopore Sequencing: From Sample to Base Calls
- Genomic Data Visualization Tools: Choosing and Using Them Effectively
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- Performance comparison: exome sequencing as a single test replacing Sanger sequencing.. Molecular genetics and genomics : MGG, 2021.
- Variant Allele Characterization in STR Markers Using Next-Generation Sequencing.. Genes, 2026.
- Development of High-Throughput Clinical Testing of RPGR ORF15 Using a Large Inherited Retinal Dystrophy Cohort.. Investigative ophthalmology & visual science, 2018.
- Concordance and reproducibility of a next generation mtGenome sequencing method for high-quality samples using the Illumina MiSeq.. Forensic science international. Genetics, 2016.
- Targeted, high-depth, next-generation sequencing of cancer genes in formalin-fixed, paraffin-embedded and fine-needle aspiration tumor specimens.. The Journal of molecular diagnostics : JMD, 2013.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.