Haplotype-Based Germline Variant Calling: How GATK HaplotypeCaller and DeepVariant Reconstruct Diploid Genotypes

By Dr. Zubair Khalid, DVM, MS, PhD ·

Haplotype-Based Germline Variant Calling: How GATK HaplotypeCaller and DeepVariant Reconstruct Diploid Genotypes

Key Takeaways

  • GATK HaplotypeCaller reconstructs diploid genotypes by performing local de novo assembly of reads within "active regions" and then employing a pairwise Hidden Markov Model (HMM) to evaluate candidate haplotypes against observed read data, accounting for base quality scores and indel probabilities.
  • DeepVariant utilizes a convolutional neural network (CNN) that processes aligned reads as multi-channel image tensors, learning to predict genotype probabilities directly from read pileup patterns without explicit haplotype assembly, requiring platform-specific trained models.
  • HaplotypeCaller's explicit assembly provides interpretable output for complex variants and allows for local haplotype phasing, whereas DeepVariant's learned approach may offer higher accuracy for simple variants by recognizing subtle error patterns but lacks direct phasing information.
  • Both tools require aligned BAM/CRAM files with accurate base quality scores and matching reference genomes; HaplotypeCaller benefits from GATK's standard preprocessing like Base Quality Score Recalibration, while DeepVariant is more tolerant of minimal preprocessing if the training model is appropriate.
  • Computational requirements differ significantly: HaplotypeCaller is CPU-intensive and parallelizable by interval, whereas DeepVariant is optimized for GPU acceleration, offering substantial speedups for large cohorts with GPU infrastructure.
  • For clinical applications, rigorous validation of the chosen variant calling pipeline (either HaplotypeCaller or DeepVariant) against known variant sets is mandatory, alongside robust quality management systems to monitor metrics like transition/transversion ratios and genotype concordance.

Germline variant calling from short-read whole-genome or exome sequencing data requires reconstructing the two parental haplotypes that make up a diploid genome. GATK HaplotypeCaller and DeepVariant are two widely used tools that approach this reconstruction problem through fundamentally different mechanisms. GATK HaplotypeCaller performs local de novo assembly of reads spanning each active genomic region, then evaluates candidate haplotypes against a pairwise hidden Markov model to compute genotype likelihoods. DeepVariant instead converts aligned reads into multi-channel image tensors and uses a convolutional neural network trained on labeled examples to predict genotype probabilities directly from the read pileup. This article explains how each caller reconstructs diploid genotypes, compares their practical strengths and limitations, and provides concrete guidance for choosing between them in germline analysis workflows.

The intended reader is a bioinformatics practitioner who needs to understand algorithmic differences to make informed tool selection decisions. The scope covers data inputs, workflow choices, quality controls, reproducibility considerations, interpretation limits, and practical decision criteria. Somatic variant calling and structural variant detection are discussed only where they clarify boundaries of germline-focused tools.

The Diploid Genotype Reconstruction Problem

A diploid genome contains two copies of each autosome, one inherited from each parent. When sequencing reads are aligned to a reference genome, the aligner places each read at its best matching position, but the alignment alone does not reveal which parental chromosome a read originated from. Variant calling is the process of determining, at each genomic position, whether the sequenced sample differs from the reference and what the two alleles are.

For a diploid sample, the possible genotypes at a single nucleotide position are homozygous reference, heterozygous, and homozygous alternate. Indels add complexity because the alternate allele involves a length change that shifts the reading frame of alignment. Small complex variants such as microinversions involve multiple nucleotide changes in a localized region. The core challenge is that sequencing errors, alignment artifacts, and repetitive sequence create ambiguity about which alleles are truly present.

Haplotype-based callers address this ambiguity by reconstructing the actual sequence of each chromosome copy in a local region. Instead of evaluating each position independently, they assemble reads into candidate haplotypes, then determine which combination of two haplotypes best explains the observed read data. This approach is particularly valuable for indels and complex variants because it resolves the alignment ambiguity that plagues position-by-position methods.

The practical consequence is that haplotype-based callers generally produce more accurate indel calls than simple pileup-based methods. They also provide phased genotype information within local regions, meaning they can determine which alleles co-occur on the same chromosome copy. This phasing information is useful for downstream analyses such as compound heterozygote detection and haplotype-based association studies.

GATK HaplotypeCaller: Local Assembly and Pairwise Hidden Markov Model

GATK HaplotypeCaller implements a multi-stage algorithm that reconstructs diploid genotypes through local de novo assembly. The Broad Institute developed this tool as part of the Genome Analysis Toolkit, and it has become a standard for germline variant discovery in research and clinical settings.

Active Region Determination

The first stage identifies genomic intervals where variation is likely present. HaplotypeCaller scans the aligned reads and flags regions where the read data disagree with the reference sequence beyond what sequencing error alone would explain. These active regions are the only intervals subjected to assembly, which conserves computational resources.

The active region determination uses a statistical model that accounts for base quality scores and expected error rates. Regions with sufficient evidence of alternate alleles become candidates for assembly. The size of each active region is bounded to keep the assembly problem tractable, and overlapping regions are merged when they are close together.

Local De Novo Assembly

For each active region, HaplotypeCaller extracts all reads that overlap the interval and performs local de novo assembly using a De Bruijn graph. The assembly process builds a graph where k-mers from the reads are nodes and connections represent adjacent k-mers observed in the read data. The graph is then simplified by removing low-coverage branches and resolving cycles where possible.

The assembled contigs represent candidate haplotypes for the region. Each contig is a sequence that could plausibly be present in the sample, given the observed reads. The number of candidate haplotypes depends on the complexity of the region and the diversity of the read data. In simple regions, the assembly may produce only the reference haplotype and one alternate. In complex regions such as those containing repeats or multiple nearby variants, the assembly can produce many candidates.

The key advantage of assembly is that it naturally handles indels. When reads contain an insertion or deletion, the assembly process incorporates the length change into the contig sequence. This avoids the alignment ambiguity that occurs when a read with an indel is aligned to a linear reference. The assembled haplotypes represent the actual sequence content of the region, including any length differences.

Pairwise Hidden Markov Model Genotyping

Once candidate haplotypes are assembled, HaplotypeCaller must determine which two haplotypes are present in the diploid sample. This is accomplished with a pairwise hidden Markov model (HMM) that computes the likelihood of each read given each possible haplotype pair.

The HMM models the process by which a read could have been generated from a haplotype, accounting for base quality scores, insertion and deletion probabilities, and the possibility of sequencing errors. For each read and each haplotype, the HMM computes a likelihood score. The read likelihood given a diploid genotype is then the average of the read likelihoods given each of the two haplotypes, under the assumption that the read is equally likely to have come from either chromosome copy.

The genotype likelihood for each possible diploid combination is the product of read likelihoods across all reads in the region. HaplotypeCaller then applies Bayes' rule with a prior on genotype frequencies to compute posterior genotype probabilities. The genotype with the highest posterior probability is emitted as the call, along with a quality score derived from the posterior probability.

This pairwise HMM approach is computationally intensive but provides accurate genotype likelihoods that account for the full read evidence. The model naturally handles reads that span multiple variant sites, which improves phasing accuracy within the assembled region.

Variant Representation and Output

HaplotypeCaller represents variants in the Variant Call Format (VCF) using a normalized representation. Each variant record includes the reference allele, alternate allele, genotype, genotype quality, and per-sample depth and allele balance information. The output also includes the assembled haplotypes in the assembly format field, which can be useful for validation and interpretation.

The tool can operate in two modes. The single-sample mode processes one sample at a time and produces a VCF with genotype calls. The cohort mode processes multiple samples jointly, first generating per-sample genomic VCF (gVCF) files that contain reference confidence blocks, then combining them in a joint genotyping step. The gVCF approach allows new samples to be added to a cohort without recalling all previous samples.

DeepVariant: Image-Based Neural Network Genotyping

DeepVariant takes a fundamentally different approach to diploid genotype reconstruction. Instead of assembling haplotypes and computing likelihoods with a probabilistic model, DeepVariant converts the aligned read data into image-like tensors and uses a deep convolutional neural network to predict genotypes directly.

Read Pileup to Image Conversion

The first stage of DeepVariant constructs a multi-channel image for each candidate variant site. The image is centered on the position of interest and spans a window of approximately 100 base pairs. Each row of the image represents a sequencing read that overlaps the position, and each column represents a genomic position within the window.

The image has multiple channels that encode different aspects of the read data. One channel encodes the base call at each position, using a one-hot encoding of the four nucleotides. Another channel encodes the base quality score. Additional channels encode read strand, read mapping quality, whether the base matches the reference, and whether the position is an insertion or deletion. The result is a three-dimensional tensor that captures the full read pileup information in a format suitable for convolutional neural network processing.

The image construction is deterministic and does not involve assembly. Reads are taken directly from the aligned BAM file and placed in the pileup according to their alignment positions. This means DeepVariant does not explicitly reconstruct haplotypes. Instead, the neural network learns to recognize patterns in the pileup that distinguish true variants from sequencing errors and alignment artifacts.

Convolutional Neural Network Architecture

The neural network used by DeepVariant is a convolutional architecture inspired by image recognition models. The network applies a series of convolutional filters that detect local patterns in the pileup, followed by pooling layers that reduce dimensionality and fully connected layers that produce the final classification.

The network is trained on labeled examples derived from real sequencing data. Training examples are generated by aligning reads to a reference, calling variants with existing tools, and using the resulting calls as labels. The training process adjusts the network weights to minimize the difference between predicted and actual genotypes.

The output layer of the network produces probabilities for each possible genotype: homozygous reference, heterozygous, and homozygous alternate. For sites with multiple alternate alleles, the network can produce probabilities for each allele combination. The genotype with the highest probability is emitted as the call.

Training Data and Model Transfer

DeepVariant's accuracy depends heavily on the training data used to fit the network weights. The original model was trained on whole-genome sequencing data from the Genome in a Bottle reference samples, which have high-confidence variant calls established through multiple sequencing technologies and orthogonal methods.

A critical practical consideration is that the model may not transfer perfectly to data generated with different sequencing platforms, library preparation methods, or coverage levels. The network learns platform-specific error patterns, and applying the model to data with different error characteristics can reduce accuracy. DeepVariant provides separate models for different sequencing platforms, including Illumina whole-genome, Illumina exome, and Pacific Biosciences data.

The image-based approach has a notable advantage for certain variant types. Because the network learns directly from labeled examples, it can recognize complex patterns that are difficult to model explicitly. This includes variants in repetitive regions and variants with unusual read support patterns. The tradeoff is that the network is a black box, and the reasons for specific calls are not directly interpretable from the model parameters.

Side-by-Side Mechanistic Comparison

The two callers differ in their fundamental approach to the genotype reconstruction problem. GATK HaplotypeCaller builds explicit haplotypes through assembly and evaluates them with a probabilistic model. DeepVariant learns to recognize variant patterns from labeled training data without explicit haplotype reconstruction.

Assembly and Haplotype Representation

GATK HaplotypeCaller produces explicit haplotype sequences for each active region. These haplotypes are available in the output and can be examined to understand the variant structure. The assembly process naturally handles complex variants such as microinversions and multi-nucleotide polymorphisms because the assembled contigs represent the actual sequence content.

DeepVariant does not produce explicit haplotypes. The network operates on the read pileup and predicts genotypes without reconstructing the underlying chromosome sequences. This means the caller cannot directly phase variants across a region, and complex variant structures must be inferred from the pileup pattern.

The practical consequence is that GATK HaplotypeCaller provides more interpretable output for complex variants, while DeepVariant may achieve higher accuracy for simple variants because the network can learn subtle error patterns that are difficult to model explicitly.

Genotype Likelihood Computation

GATK HaplotypeCaller computes genotype likelihoods using a pairwise HMM that models the read generation process. The model parameters include base error rates, indel rates, and the probability of reads originating from each chromosome copy. These parameters are estimated from the data and can be adjusted by the user.

DeepVariant replaces the explicit likelihood model with a learned mapping from image tensors to genotype probabilities. The network learns the relationship between read patterns and true genotypes from training data. This approach can capture complex interactions between error sources that are difficult to specify in a hand-built model.

The tradeoff is that the HMM approach has well-understood statistical properties and can be adapted to new data types by adjusting parameters. The neural network approach requires retraining or fine-tuning when applied to data with different error characteristics.

Computational Requirements

GATK HaplotypeCaller is computationally intensive because assembly and HMM evaluation are expensive operations. The tool scales with the number of active regions and the complexity of the assembly graphs. Runtime can be reduced by parallelizing across genomic intervals, but memory usage can be high for complex regions.

DeepVariant is also computationally intensive because the neural network inference requires substantial matrix operations. The tool is optimized for graphics processing unit (GPU) acceleration, and GPU usage can dramatically reduce runtime. On CPU-only systems, DeepVariant can be slower than HaplotypeCaller for large datasets.

The choice between the two tools may depend on available computational infrastructure. Laboratories with GPU access may find DeepVariant more practical for large cohorts, while those with CPU-only clusters may prefer HaplotypeCaller.

At a Glance

FeatureGATK HaplotypeCallerDeepVariant
Core mechanismLocal de novo assembly plus pairwise hidden Markov modelConvolutional neural network on read pileup images
Haplotype reconstructionExplicit assembly of candidate haplotypesNo explicit haplotype assembly
Genotype likelihoodProbabilistic model with explicit error parametersLearned mapping from training data
Indel handlingAssembly resolves alignment ambiguityNetwork learns indel patterns from training examples
Phasing informationLocal phasing within assembled regionsNo direct phasing output
Training data requirementNone, uses explicit model parametersRequires platform-matched training model
Computational profileCPU intensive, parallelizable by intervalGPU accelerated, CPU fallback available
Output formatVCF with assembly detailsVCF with genotype probabilities
Best suited forComplex variants, interpretable output, custom data typesHigh accuracy on standard platforms, large cohorts with GPU

Practical Workflow for Germline Variant Calling

A complete germline variant calling workflow involves multiple stages before and after the variant caller itself. Understanding where HaplotypeCaller and DeepVariant fit in the pipeline is essential for producing reliable results.

Input Data Requirements

Both callers require aligned sequencing data in BAM or CRAM format. The alignment must include a reference genome index and read group information. Read groups are essential because they identify the sample and library, and both callers use this information for quality control and error modeling.

Base quality scores must be present in the BAM file. These scores are used by HaplotypeCaller in the HMM likelihood computation and are encoded in the image tensors used by DeepVariant. Low-quality base calls reduce the confidence of variant calls, and both tools account for this in their respective models.

The reference genome must match the one used for alignment. Mismatches between the alignment reference and the variant calling reference can cause spurious calls or missed variants. Both tools require the reference FASTA file and its associated index files.

Preprocessing Steps

Before variant calling, the aligned reads should undergo preprocessing to improve accuracy. The standard preprocessing steps include marking duplicate reads, which are PCR or optical duplicates that should not be counted as independent evidence. Both callers can operate on BAM files with marked duplicates, and the duplicate flag is used to exclude these reads from analysis.

Base quality score recalibration is a step that adjusts base quality scores based on empirical error patterns. This step is part of the GATK best practices workflow and is particularly important for HaplotypeCaller because the HMM relies on accurate quality scores. DeepVariant can also benefit from recalibrated quality scores, although the network may learn to compensate for systematic errors.

The preprocessing requirements differ between the two tools. GATK HaplotypeCaller is designed to work with the full GATK preprocessing pipeline, including base quality score recalibration. DeepVariant is more tolerant of minimal preprocessing and can produce accurate calls from raw aligned reads, provided the training model matches the sequencing platform.

Running HaplotypeCaller

HaplotypeCaller is typically run in one of two modes. The single-sample mode is appropriate for analyzing one sample at a time and produces a VCF with genotype calls. The cohort mode uses gVCF files and joint genotyping, which is recommended for multi-sample projects because it improves genotype accuracy at low-coverage sites.

The command line for HaplotypeCaller requires specification of the input BAM, reference genome, and output file. Optional parameters control the confidence thresholds for emitting variants, the minimum base quality score, and the size of active regions. Default parameters are appropriate for most germline applications, but adjustments may be needed for unusual data types.

The runtime of HaplotypeCaller depends on the number of active regions and the complexity of the assembly. Whole-genome samples can take many hours to process on a single CPU core. Parallelization across genomic intervals is supported and is the primary strategy for reducing runtime.

Running DeepVariant

DeepVariant is run through a command line interface that requires specification of the input BAM, reference genome, output directory, and model type. The model type must match the sequencing platform and library preparation method used to generate the data.

The tool produces a VCF file with genotype calls and an accompanying report file that summarizes the number of variants called and the distribution of genotype qualities. The report can be used for quality control and to identify potential issues with the input data.

DeepVariant benefits significantly from GPU acceleration. The neural network inference operations are highly parallel and map well to GPU hardware. On systems without GPUs, the tool can run on CPU but may be substantially slower.

Post-Calling Filtering and Quality Control

Both callers produce VCF files that require filtering before downstream analysis. The raw calls include low-confidence variants that should be removed based on quality scores, depth, and allele balance.

The standard filtering approach uses the variant quality score, which is derived from the posterior probability of the genotype call. Variants with low quality scores are more likely to be false positives and should be filtered. The specific threshold depends on the application, with clinical applications requiring higher confidence than research applications.

Additional filters include depth filters that remove variants with insufficient read support, allele balance filters that remove variants with skewed allele fractions, and strand bias filters that remove variants supported primarily by one read strand. These filters are applied after variant calling and are not specific to either caller.

Options and Tradeoffs in Tool Selection

The choice between HaplotypeCaller and DeepVariant depends on multiple factors including data type, computational infrastructure, accuracy requirements, and the need for interpretable output.

Accuracy Considerations

Both tools achieve high accuracy for standard germline variant calling on Illumina whole-genome data. The relative performance depends on the specific variant type and genomic context. HaplotypeCaller has an advantage for complex variants that benefit from explicit assembly, while DeepVariant may have an advantage for simple variants in repetitive regions where the network has learned to recognize error patterns.

The training data used for DeepVariant models is an important consideration. Models trained on data from one sequencing platform may not perform optimally on data from another platform. HaplotypeCaller does not require training data and uses explicit model parameters that can be adjusted for different data types.

For clinical applications where accuracy is critical, benchmarking on a validation set with known variants is recommended. The validation set should be representative of the sequencing platform, coverage, and variant types expected in the clinical samples.

Computational Infrastructure

HaplotypeCaller runs on CPU and can be parallelized across genomic intervals. The tool is memory intensive for complex regions, and memory usage should be monitored when processing whole-genome data. The runtime is predictable and scales with the number of active regions.

DeepVariant is optimized for GPU acceleration and can achieve substantial speedups on GPU hardware. The tool can run on CPU but may be slower than HaplotypeCaller for large datasets. The memory requirements are moderate and depend on the batch size used for neural network inference.

Laboratories with GPU infrastructure may prefer DeepVariant for large cohorts, while those with CPU-only clusters may find HaplotypeCaller more practical. The total cost of ownership should consider both hardware and runtime.

Interpretability and Validation

HaplotypeCaller provides explicit haplotype sequences in its output, which can be examined to understand the variant structure. This interpretability is valuable for validating complex variants and for understanding why a particular genotype was called.

DeepVariant does not provide explicit haplotypes, and the neural network is a black box. The reasons for specific calls are not directly interpretable from the model. However, the tool provides genotype probabilities that can be used to assess confidence.

For applications where variant validation is important, the interpretability of HaplotypeCaller may be an advantage. For applications where throughput and accuracy on standard variants are the primary concerns, DeepVariant may be preferred.

Observations and Measurements for Quality Assessment

Monitoring the performance of variant calling requires systematic measurement of quality metrics and comparison against expected values.

Variant Count and Transition Transversion Ratio

The number of variants called and the ratio of transitions to transversions are useful quality indicators. Transition mutations (A to G, C to T) are more common than transversions in the human genome, and the transition transversion ratio for whole-genome data is typically around 2.0. A significantly lower ratio may indicate an excess of false positive calls.

The variant count should be consistent with expectations for the sample type and sequencing platform. Whole-genome samples typically have approximately 3 to 4 million variants relative to the reference genome. Exome samples have far fewer variants because they cover only the protein-coding regions.

These measurements should be compared across samples in a cohort to identify outliers. A sample with an unusually high variant count may have contamination or alignment issues, while a sample with a low count may have low coverage or poor data quality.

Genotype Concordance

When multiple samples from the same individual are available, genotype concordance can be measured. This includes samples sequenced on different platforms, at different coverage levels, or processed with different pipelines. High concordance indicates reliable genotyping, while discordance may indicate systematic errors.

Genotype concordance is also used to compare different variant callers. Running both HaplotypeCaller and DeepVariant on the same data and comparing the calls can identify regions where the tools disagree. These discordant regions may warrant manual inspection or orthogonal validation.

Depth and Allele Balance Distributions

The depth distribution across called variants should be consistent with the overall sequencing depth. Variants with extremely low depth may be unreliable, while variants with extremely high depth may be in repetitive regions with alignment artifacts.

Allele balance is the fraction of reads supporting the alternate allele. For heterozygous variants, the expected allele balance is approximately 0.5, although variation is expected due to sampling. Homozygous variants should have allele balance near 1.0. Variants with allele balance values that deviate significantly from expectations may be false positives or may indicate mosaicism.

Records and Documentation for Reproducible Analysis

Reproducible variant calling requires careful documentation of the analysis parameters, software versions, and reference genome used.

Software Version Control

Both HaplotypeCaller and DeepVariant are actively developed, and new versions may change the algorithm or default parameters. Recording the exact software version is essential for reproducibility and for comparing results across analyses.

The version information is typically available from the command line interface or the software documentation. This information should be recorded in the analysis log or the project metadata.

Parameter Documentation

The parameters used for variant calling should be documented, including any deviations from default values. This includes the confidence thresholds, minimum base quality, and any platform-specific settings.

For DeepVariant, the model type is a critical parameter that must be documented. The model determines the expected error patterns and should match the sequencing platform.

Reference Genome Version

The reference genome version must be documented because variant coordinates are relative to the reference. Different reference versions have different coordinate systems, and variants called against one version cannot be directly compared to variants called against another.

The reference genome version should be recorded in the VCF header and in the project metadata. This information is essential for downstream analysis and for data sharing.

Common Failure Patterns and Troubleshooting

Several common problems can arise in germline variant calling, and recognizing these patterns is essential for producing reliable results.

Low Variant Count or Missing Variants

A low variant count may indicate that the variant caller is missing true variants. This can occur when the sequencing coverage is too low, when the base quality scores are poor, or when the alignment is incorrect.

For HaplotypeCaller, low variant counts may result from overly aggressive active region determination or from HMM parameters that are too conservative. For DeepVariant, low variant counts may result from using a model that does not match the sequencing platform.

The first troubleshooting step is to examine the coverage and quality metrics for the sample. Low coverage regions should be identified, and the base quality score distribution should be checked.

Excess Variant Calls

An excess of variant calls may indicate false positives from sequencing errors, alignment artifacts, or contamination. The transition transversion ratio is a useful diagnostic, as a low ratio suggests an excess of false positives.

Contamination from another sample can cause an excess of heterozygous calls with skewed allele balance. This can be detected by examining the allele balance distribution and by comparing the sample to expected population frequencies.

Discordance Between Callers

When HaplotypeCaller and DeepVariant produce discordant calls, the discordant regions should be examined to determine which caller is correct. This may require manual inspection of the read alignment or orthogonal validation with a different sequencing technology.

Discordance is more common in complex regions, repetitive regions, and regions with low coverage. The cause of the discordance should be documented, and the results should be interpreted with appropriate caution.

Runtime and Memory Issues

HaplotypeCaller can encounter runtime and memory issues in complex genomic regions. The assembly process can generate large graphs that consume substantial memory. If the tool crashes or runs out of memory, the problematic region should be identified and the analysis should be rerun with adjusted parameters.

DeepVariant can encounter runtime issues on CPU-only systems, particularly for large datasets. GPU acceleration is recommended for production use, and the batch size should be adjusted to fit the available memory.

Limitations and Boundaries of Haplotype-Based Germline Callers

Both HaplotypeCaller and DeepVariant have limitations that should be understood before applying them to research or clinical questions.

Somatic Variant Calling

Haplotype-based callers were designed primarily for germline variant detection in diploid populations. Somatic variant calling requires detecting variants present in a subset of cells, which presents different statistical challenges. The genotype model in HaplotypeCaller assumes diploid germline variation and is not optimal for somatic mutation detection.

Specialized somatic callers use tumor-normal comparison and allele fraction modeling to detect low-frequency mutations. These tools are distinct from germline callers and should be used for cancer sequencing applications. The distinction between germline and somatic calling is critical for clinical interpretation.

Structural Variants

Both HaplotypeCaller and DeepVariant are designed for small variants including single nucleotide variants and indels. Structural variants, which include large deletions, duplications, inversions, and translocations, require different detection methods.

Structural variant calling often uses read depth, split read, and assembly-based approaches. Haplotype-based assembly can contribute to structural variant detection, but the standard germline callers do not provide comprehensive structural variant calling. Specialized tools such as Aquila_stLFR use haplotype-based assembly of linked-read data to reconstruct structural variants across the genome.

Complex Genomic Regions

Both callers have difficulty in complex genomic regions such as segmental duplications, centromeres, and telomeres. These regions have high sequence similarity between different genomic locations, which causes alignment ambiguity and assembly errors.

The accuracy of variant calls in these regions should be interpreted with caution. Orthogonal validation with long-read sequencing or other methods may be necessary for variants in complex regions.

Haplotype Phasing Limitations

The local phasing provided by HaplotypeCaller is limited to the assembled regions and does not provide chromosome-scale phasing. Full haplotype phasing requires specialized methods that use population reference panels, linked-read sequencing, or long-read sequencing.

The distinction between local phasing and chromosome-scale phasing is important for applications such as compound heterozygote detection. Local phasing can determine whether two variants are on the same chromosome copy within a small region, but cannot phase variants across large genomic distances.

Safety and Regulatory Context for Clinical Applications

Variant calling for clinical applications requires additional considerations beyond research use.

Validation Requirements

Clinical laboratories must validate variant calling pipelines before use in patient care. Validation typically involves testing the pipeline on samples with known variants and demonstrating acceptable sensitivity and specificity. The validation should be performed on data representative of the clinical samples that will be analyzed.

The validation results should be documented and reviewed by qualified personnel. Any changes to the pipeline, including software version updates or parameter changes, should trigger revalidation.

Quality Management

Clinical laboratories should have a quality management system that includes regular monitoring of variant calling performance. This includes tracking quality metrics, investigating discordant results, and documenting corrective actions.

The quality management system should also include procedures for handling samples that fail quality thresholds. Samples with insufficient coverage or poor data quality should be identified and either rerun or reported with appropriate limitations.

Professional Escalation Criteria

Variant calls that are unexpected, discordant between methods, or in clinically significant genes should be escalated for review by qualified personnel. This includes variants that are inconsistent with the clinical presentation, variants in genes with established disease associations, and variants with unusual quality metrics.

The escalation process should be documented and should include procedures for orthogonal validation, literature review, and consultation with clinical experts. The goal is to ensure that clinical decisions are based on accurate and well-understood variant calls.

Frequently Asked Questions

What is the main algorithmic difference between GATK HaplotypeCaller and DeepVariant?

GATK HaplotypeCaller performs local de novo assembly of reads in active regions to construct candidate haplotypes, then uses a pairwise hidden Markov model to compute genotype likelihoods for each possible diploid combination. DeepVariant converts the aligned read pileup into multi-channel image tensors and uses a convolutional neural network trained on labeled examples to predict genotype probabilities directly. The fundamental difference is that HaplotypeCaller builds explicit haplotype sequences and evaluates them with a probabilistic model, while DeepVariant learns to recognize variant patterns from training data without explicit haplotype reconstruction.

Which tool is more accurate for germline variant calling?

Both tools achieve high accuracy for standard germline variant calling on Illumina whole-genome data. The relative performance depends on the variant type and genomic context. HaplotypeCaller has an advantage for complex variants that benefit from explicit assembly, while DeepVariant may have an advantage for simple variants in repetitive regions where the network has learned to recognize error patterns. Benchmarking on a validation set representative of the intended data is recommended to determine which tool performs better for a specific application.

Does DeepVariant require GPU hardware?

DeepVariant is optimized for GPU acceleration and achieves substantial speedups on GPU hardware. The tool can run on CPU-only systems but may be substantially slower for large datasets. Laboratories with GPU infrastructure may prefer DeepVariant for large cohorts, while those with CPU-only clusters may find HaplotypeCaller more practical.

Can HaplotypeCaller and DeepVariant be used for somatic variant calling?

Both tools were designed primarily for germline variant detection in diploid populations. Somatic variant calling requires detecting variants present in a subset of cells, which presents different statistical challenges. Specialized somatic callers use tumor-normal comparison and allele fraction modeling to detect low-frequency mutations. The distinction between germline and somatic calling is critical for cancer sequencing applications.

What preprocessing steps are required before variant calling?

The standard preprocessing steps include marking duplicate reads and base quality score recalibration. HaplotypeCaller is designed to work with the full GATK preprocessing pipeline, including base quality score recalibration. DeepVariant is more tolerant of minimal preprocessing and can produce accurate calls from raw aligned reads, provided the training model matches the sequencing platform.

How should discordant calls between the two tools be resolved?

Discordant calls should be examined to determine which tool is correct. This may require manual inspection of the read alignment or orthogonal validation with a different sequencing technology. Discordance is more common in complex regions, repetitive regions, and regions with low coverage. The cause of the discordance should be documented, and the results should be interpreted with appropriate caution.

What is the transition transversion ratio and why is it useful?

The transition transversion ratio is the ratio of transition mutations to transversion mutations in the called variants. Transition mutations are more common than transversions in the human genome, and the ratio for whole-genome data is typically around 2.0. A significantly lower ratio may indicate an excess of false positive calls. This measurement is useful for quality assessment and for comparing samples across a cohort.

How does local phasing from HaplotypeCaller differ from chromosome-scale phasing?

Local phasing from HaplotypeCaller is limited to the assembled regions and determines which alleles co-occur on the same chromosome copy within a small region. Chromosome-scale phasing determines the parental origin of alleles across entire chromosomes and requires specialized methods that use population reference panels, linked-read sequencing, or long-read sequencing. The distinction is important for applications such as compound heterozygote detection.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.