Variant Quality Score Recalibration (VQSR) for Germline Variants: A Practical Guide to Training Sets and Parameters
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Variant Quality Score Recalibration (VQSR) is a Gaussian mixture model-based filtering approach in GATK for distinguishing true germline variants from sequencing artifacts, relying on annotation profiles of known true positives (e.g., HapMap, Omni, 1000 Genomes) and false positives (e.g., dbSNP).
- VQSR is not suitable for small variant call sets, such as single-sample exomes or targeted panels, which typically lack the minimum number of sites (historically ~30,000) required for reliable model training; hard filtering is recommended in such cases.
- The selection of training resources and annotations (e.g., QD, MQ, FS, SOR, ReadPosRankSum) is critical, as biases or errors in training data directly propagate into filtering decisions, and specific annotations may be excluded for exome data due to constrained read placement.
- The VQSR workflow involves two main steps:
VariantRecalibratorto build the model andApplyVQSRto filter variants based on user-defined tranche sensitivity levels (commonly 99.0% or 99.9%), with separate recalibration for SNPs and indels due to differing annotation distributions. - Common failure modes include insufficient variant sites, mismatched reference genomes between call sets and training resources, missing annotations, and poor annotation distributions, necessitating careful input file verification and data quality assessment.
- Reproducibility is paramount, requiring detailed documentation of GATK versions, reference genomes, training resources, priors, annotations, and parameters used, alongside validation of filtered call sets via Ti/Tv ratio assessment and concordance with known variant sites.
Variant Quality Score Recalibration (VQSR) is a Gaussian mixture model-based filtering approach used in the Genome Analysis Toolkit (GATK) to separate true germline variants from sequencing and alignment artifacts. This guide explains how to apply VQSR correctly to germline variant calling workflows, with specific attention to training set selection, annotation configuration, parameter tuning, and common failure modes. The intended reader is a bioinformatics practitioner who has generated a raw variant call set and needs to produce a high-confidence germline VCF for downstream analysis.
VQSR differs from hard filtering because it learns from the distribution of variant annotations in your own call set and compares that distribution against known true-positive and false-positive training resources. The method requires a minimum number of variant sites to build a reliable model, which means it is not appropriate for single-sample exomes or small targeted panels. Understanding when VQSR applies and when it does not is the first practical decision in this workflow.
Scope and Context for Germline VQSR
Germline variant calling aims to identify inherited sequence differences from a diploid genome or exome. The raw output of a variant caller such as GATK HaplotypeCaller contains true variants mixed with artifacts arising from read misalignment, sequencing errors, PCR duplicates, and reference biases. VQSR is one layer of the filtering strategy that sits between raw variant calling and final variant annotation.
The Genome Analysis Toolkit has been a foundational tool in short-read variant calling since its release in 2010, and it remains widely used across the genomics community for its integrated approach to variant discovery [<a href="#ref-1">1</a>]. The toolkit provides the VariantRecalibrator and ApplyVQSR tools that implement the VQSR workflow. These tools are part of a larger collection of analysis utilities that support the full variant calling process from raw reads to filtered variant sets [<a href="#ref-1">1</a>].
VQSR is designed for germline data because it relies on known polymorphic sites from population databases. Somatic variant calling has different requirements because tumor samples contain subclonal mutations and copy number alterations that distort the annotation distributions VQSR depends on. If you are working with somatic data, you should use somatic-specific filtering approaches instead of VQSR. The VariFAST method, for example, provides an alternative automated scoring approach that handles both germline and somatic variant filtering through weighted metrics and machine learning [<a href="#ref-2">2</a>].
The practical scope of this guide covers germline whole-genome sequencing (WGS) and whole-exome sequencing (WES) data processed through GATK best practices. The principles apply to both single-sample and cohort-level joint calling, though the minimum site requirements differ. A cohort approach with joint genotyping produces a multi-sample VCF that provides more sites for VQSR modeling than any single sample alone [<a href="#ref-3">3</a>].
Core Principles of VQSR
VQSR operates on the principle that true variants and false positives have distinguishable annotation profiles. The VariantRecalibrator tool builds a Gaussian mixture model using variant annotations such as quality by depth (QD), mapping quality (MQ), Fisher strand bias (FS), strand odds ratio (SOR), and read position rank sum test (ReadPosRankSum). These annotations capture different aspects of variant support quality.
The model is trained using three categories of input. The first category is the variant call set itself, which provides the full distribution of annotations across all called sites. The second category consists of known true-positive sites from resources such as HapMap, Omni, and the 1000 Genomes project. The third category consists of known false-positive sites, typically from the dbSNP database, which contains both true polymorphisms and artifacts that have been observed across many studies.
The Gaussian mixture model assigns each variant a quality score based on its position in the annotation space relative to the training distributions. Variants that cluster with known true positives receive high scores. Variants that cluster with known false positives or that fall in ambiguous regions receive lower scores. The ApplyVQSR tool then applies a tranche threshold to filter variants below a specified sensitivity level.
The tranche system is a distinctive feature of VQSR. Instead of applying a single quality cutoff, VQSR divides the variant set into tranches of increasing stringency. The first tranche contains the highest-confidence variants, and subsequent tranches include progressively more variants at the cost of including more false positives. The user selects a tranche sensitivity level, commonly 99.0 percent or 99.9 percent, to determine the final filtered call set.
Training Sets and Resources
The choice of training resources is the most consequential decision in VQSR. The model learns what a true variant looks like from these resources, so errors or biases in the training data propagate directly into the filtering decision.
HapMap and Omni Sites
HapMap sites are polymorphic loci that have been genotyped across multiple populations with high confidence. These sites are valuable training data because they represent validated true positives with well-characterized allele frequencies. Omni sites come from the Illumina Omni genotyping array and provide another set of high-confidence polymorphic loci. Both resources are commonly used in germline VQSR training.
The National Center for Biotechnology Information (NCBI) maintains databases and search systems that support the discovery and validation of genetic variation [<a href="#ref-4">4</a>]. These resources provide the underlying data infrastructure for population variation studies, and they are relevant when you need to understand the provenance of training resources or access variation data for validation purposes.
1000 Genomes Sites
The 1000 Genomes Project provides a large catalog of variants identified through low-coverage and exome sequencing of diverse populations. These sites are useful for training because they represent variants observed across many individuals and sequencing platforms. However, the 1000 Genomes call set contains some false positives, particularly at low allele frequencies, so it is often used in combination with HapMap and Omni instead of as the sole training resource.
dbSNP as False-Positive Resource
The dbSNP database is used in VQSR as a source of known false positives. This may seem counterintuitive because dbSNP contains many true variants. The rationale is that dbSNP contains a large number of sites that have been observed in sequencing studies but that have not been validated as true polymorphisms. These sites include sequencing artifacts that have been deposited in the database over many years. The VQSR model uses the presence of a variant in dbSNP as evidence that the site may be a false positive, while the absence from dbSNP combined with strong annotation values supports a true-positive call.
The NCBI provides official documentation and access to the dbSNP database through its data resources portal [<a href="#ref-4">4</a>]. When you download dbSNP for use in VQSR, you should verify the build version and ensure it matches the reference genome version used in your variant calling pipeline.
Resource Selection by Data Type
Whole-genome and whole-exome data have different annotation distributions, and the training resource strategy should reflect this difference. Exome data have fewer variant sites overall, which makes VQSR modeling more challenging. Some practitioners use a reduced training resource set for exomes, focusing on HapMap and Omni sites that are well represented in exonic regions.
The minimum number of variant sites required for VQSR is a practical constraint. GATK documentation historically recommended at least 30,000 variant sites for building a reliable model. This requirement is easily met by whole-genome data, which typically produces millions of variant sites. Exome data produce fewer sites, often in the range of 20,000 to 50,000 variants, which may be marginal for VQSR. A cohort of exomes that has been jointly genotyped produces a multi-sample VCF with more sites than any single exome, which improves VQSR performance [<a href="#ref-3">3</a>].
Annotations for VQSR Modeling
The annotations used in VQSR define the feature space for the Gaussian mixture model. Selecting appropriate annotations and ensuring they are present in the input VCF is essential for successful recalibration.
Core Annotation Set
The standard VQSR annotation set includes quality by depth (QD), mapping quality (MQ), Fisher strand bias (FS), strand odds ratio (SOR), mapping quality rank sum test (MQRankSum), read position rank sum test (ReadPosRankSum), and variant quality (QUAL). These annotations capture different aspects of variant support:
QD measures the variant quality divided by the depth of coverage at the site. True variants tend to have high QD values, while artifacts often have low QD because they are supported by few reads with weak evidence.
MQ measures the root mean square mapping quality of reads supporting the variant allele. Low MQ values indicate that supporting reads map poorly to the reference, which is a hallmark of alignment artifacts.
FS and SOR measure strand bias, which is the tendency for variant-supporting reads to come predominantly from one strand. True variants should have support from both strands, so high strand bias values suggest artifacts.
MQRankSum and ReadPosRankSum compare the mapping quality and read position distributions of reference-supporting and variant-supporting reads. Significant differences between these distributions indicate potential artifacts.
Annotation Availability
The annotations must be present in the input VCF before running VariantRecalibrator. If you used GATK HaplotypeCaller to generate the raw variants, these annotations are included by default. If you used a different variant caller, you may need to add annotations using the VariantAnnotator tool before running VQSR.
The European Bioinformatics Institute provides training materials on data resources and practical analysis education that cover variant annotation and quality control concepts [<a href="#ref-5">5</a>]. These training resources can help you understand how annotations are calculated and interpreted in the context of variant filtering.
Annotation Selection for Specific Data Types
Whole-genome data benefit from the full annotation set because the model has enough sites to learn from all annotation dimensions. Exome data may benefit from a reduced annotation set because the smaller number of sites makes the model more sensitive to overfitting. Some practitioners exclude MQRankSum and ReadPosRankSum for exome data because these annotations behave differently in exonic regions where read placement is constrained by capture probes.
The decision to include or exclude specific annotations should be based on your data type and the distribution of annotation values in your call set. You can examine annotation distributions using plotting tools before running VQSR to identify annotations with unusual distributions that may indicate data quality problems.
Practical VQSR Workflow
The VQSR workflow consists of two main steps: VariantRecalibrator and ApplyVQSR. The workflow requires careful preparation of input files and parameter settings.
Step 1: Prepare Input Files
Before running VariantRecalibrator, you need the following inputs:
A raw variant call set in VCF format. This should be a single VCF containing all variant sites from your sample or cohort. For cohort analysis, joint genotyping produces a multi-sample VCF that serves as the input [<a href="#ref-3">3</a>].
Training resource files in VCF format. These include HapMap, Omni, 1000 Genomes, and dbSNP files that have been aligned to the same reference genome version as your call set.
A reference genome file in FASTA format with associated index files.
A known-sites resource for the variant caller. This is used during the calling step, not during VQSR, but it should be consistent with the training resources used in VQSR.
Step 2: Configure VariantRecalibrator
The VariantRecalibrator tool requires a resource argument for each training resource, specifying the file path, the resource type, and the prior probability. The prior probability reflects the expected proportion of true variants in your call set that are represented by the resource.
A typical configuration for germline WGS data includes:
HapMap with a prior of 15.0, indicating high confidence in these sites as true positives.
Omni with a prior of 12.0, indicating slightly lower confidence than HapMap.
1000 Genomes with a prior of 10.0, reflecting the higher error rate in this resource.
dbSNP with a prior of 2.0, reflecting its use as a false-positive resource.
The prior values are adjustable, and the optimal values depend on your data and the version of the training resources. The GATK documentation provides recommended values, but you should treat these as starting points instead of fixed parameters.
Step 3: Run VariantRecalibrator
The VariantRecalibrator tool builds the Gaussian mixture model and outputs a recalibration file that contains the model parameters and the quality scores assigned to each variant. The tool also produces a tranche file that defines the sensitivity levels for filtering.
The tool requires a mode argument that specifies whether you are recalibrating SNPs, indels, or both. In germline workflows, SNP and indel recalibration are typically run separately because the annotation distributions differ between these variant types.
Step 4: Run ApplyVQSR
The ApplyVQSR tool uses the recalibration file and tranche file to filter the input VCF. The tool applies the tranche threshold and outputs a filtered VCF that contains only variants passing the selected sensitivity level.
The tranche sensitivity level is specified with the tranche argument. A common choice is 99.0 percent, which retains 99 percent of true positives while removing most false positives. More stringent thresholds, such as 99.9 percent, retain more true positives but also retain more false positives.
Step 5: Evaluate the Filtered Call Set
After applying VQSR, you should evaluate the filtered call set to confirm that the filtering performed as expected. Evaluation metrics include the transition-to-transversion ratio (Ti/Tv), the number of variants passing each tranche, and the concordance with known variant sites.
The Ti/Tv ratio is a useful quality metric for germline data. Whole-genome data typically have a Ti/Tv ratio around 2.0 to 2.1, while exome data have a higher ratio around 2.8 to 3.0. A Ti/Tv ratio that is much lower than expected suggests that the call set contains many false positives, which may indicate problems with the VQSR configuration.
At a Glance: VQSR Decision Table
| Data Type | Minimum Sites for VQSR | Recommended Training Resources | Typical Tranche Threshold |
|---|---|---|---|
| Whole-genome single sample | 30,000 or more | HapMap, Omni, 1000 Genomes, dbSNP | 99.0 percent |
| Whole-exome single sample | Marginal, often insufficient | HapMap, Omni, reduced 1000 Genomes | 99.0 percent or hard filtering |
| Whole-exome cohort with joint calling | 30,000 or more from multi-sample VCF | HapMap, Omni, 1000 Genomes, dbSNP | 99.0 percent |
The decision table summarizes the practical considerations for applying VQSR across common data types. Single-sample exomes often lack sufficient sites for reliable VQSR modeling, and hard filtering may be more appropriate in this scenario. Cohort-based exome analysis with joint genotyping produces a multi-sample VCF with enough sites for VQSR [<a href="#ref-3">3</a>].
Parameter Configuration and Tuning
The parameters of VQSR have a substantial impact on the filtering outcome. Understanding each parameter and how to adjust it for your data is essential for producing a reliable variant call set.
Prior Probabilities
The prior probability for each training resource controls how much influence that resource has on the model. Higher priors give the resource more weight in the model, meaning that variants matching the resource are more likely to be classified as true positives.
The default priors in GATK best practices are based on the expected accuracy of each resource. HapMap sites are highly accurate, so they receive a high prior. dbSNP contains many artifacts, so it receives a low prior. If you observe that your filtered call set retains too many false positives, you may need to increase the dbSNP prior to give it more weight as a false-positive indicator.
Tranche Threshold
The tranche threshold determines the sensitivity of the final filtered call set. A threshold of 99.0 percent means that the filtered set is expected to contain 99 percent of the true variants in the unfiltered call set. The remaining 1 percent of true variants are removed along with the false positives.
The choice of tranche threshold involves a tradeoff between sensitivity and precision. A lower threshold, such as 99.5 percent, retains more true variants but also retains more false positives. A higher threshold, such as 98.0 percent, removes more false positives but also removes more true variants. The appropriate threshold depends on your downstream analysis. Clinical applications may require higher sensitivity, while population genetics studies may tolerate lower sensitivity in exchange for higher precision.
SNP and Indel Recalibration
SNPs and indels have different annotation distributions and should be recalibrated separately. The VariantRecalibrator tool can be run once for SNPs and once for indels, with the indel run using the SNP-recalibrated call set as input. This two-step approach ensures that the model for each variant type is trained on appropriate annotations.
Indel recalibration is more challenging than SNP recalibration because indels are less abundant and have more variable annotation distributions. The VariFAST study found that their machine learning approach outperformed VQSR specifically in indel filtering, suggesting that VQSR may have limitations for indel variant refinement [<a href="#ref-2">2</a>]. If you observe poor indel filtering performance, you may consider alternative approaches.
Max Gaussians Parameter
The max-gaussians parameter controls the number of Gaussian distributions that the mixture model can use to describe the annotation space. The default value is 8, which is appropriate for most datasets. Increasing this value allows the model to capture more complex annotation distributions but also increases the risk of overfitting. Decreasing the value simplifies the model and may improve performance with small datasets.
If you receive a warning that the model is using fewer Gaussians than the maximum, this is not necessarily a problem. The model automatically selects the number of Gaussians that best fits the data. However, if the model uses only one or two Gaussians, this may indicate that the annotation distributions are not well separated, which could be a sign of data quality problems.
Common Failure Patterns in VQSR
VQSR can fail in several ways, and recognizing these failure patterns is important for troubleshooting.
Insufficient Variant Sites
The most common cause of VQSR failure is an insufficient number of variant sites in the input call set. When the model is trained on too few sites, it cannot accurately estimate the annotation distributions, leading to poor separation between true variants and false positives.
The symptom of this failure is a filtered call set that retains too many false positives or that removes too many true variants. You may also see warnings from VariantRecalibrator about the small number of sites used for training.
The solution is to use hard filtering instead of VQSR for small call sets, or to combine samples into a cohort for joint calling to increase the number of sites [<a href="#ref-3">3</a>]. The GermVarX workflow demonstrates a joint germline variant discovery approach that produces a multi-sample VCF suitable for downstream filtering [<a href="#ref-3">3</a>].
Mismatched Reference Genomes
Training resources must be aligned to the same reference genome version as your variant call set. If the training resources use a different reference version, the genomic coordinates will not match, and the model will not correctly identify known variant sites.
The symptom of this failure is a low concordance between your call set and the training resources, which results in poor model performance. You may also see errors when VariantRecalibrator attempts to intersect your call set with the training resources.
The solution is to verify the reference genome version of all input files before running VQSR. The NCBI provides reference genome assemblies and associated annotation resources that can help you confirm the correct version [<a href="#ref-4">4</a>].
Missing Annotations
VQSR requires specific annotations to be present in the input VCF. If these annotations are missing, VariantRecalibrator will fail or produce a model that does not capture the relevant annotation distributions.
The symptom of this failure is an error message from VariantRecalibrator indicating that a required annotation is missing. The solution is to add the missing annotations using the VariantAnnotator tool before running VQSR.
Overlapping Training Resources
Training resources may contain overlapping sites, which can cause the model to double-count certain variants. This is not usually a problem in practice, but it can affect the prior probabilities if the same site appears in multiple resources with different priors.
The solution is to be aware of the overlap between resources and to adjust priors accordingly. The GATK documentation provides guidance on handling overlapping resources.
Poor Annotation Distributions
If the annotation distributions in your call set are unusual, the VQSR model may not perform well. This can happen with data from unusual sequencing platforms, samples with extreme GC content, or samples with high contamination levels.
The symptom of this failure is a filtered call set with unexpected characteristics, such as a low Ti/Tv ratio or a high proportion of variants in repetitive regions. The solution is to examine the annotation distributions before running VQSR and to address any underlying data quality issues.
Records and Measurements for VQSR
Keeping detailed records of your VQSR configuration and results is essential for reproducibility and troubleshooting.
Configuration Records
Record the following information for each VQSR run:
The version of GATK used for recalibration.
The reference genome version and build.
The training resource files and their versions.
The prior probability assigned to each training resource.
The annotation set used for modeling.
The max-gaussians parameter value.
The tranche threshold applied.
The input VCF file and the number of variant sites it contains.
Output Records
Record the following information from the VQSR output:
The number of variants passing each tranche.
The Ti/Tv ratio of the filtered call set.
The number of SNPs and indels in the filtered call set.
The concordance with known variant sites.
The number of variants removed by filtering.
These records allow you to compare VQSR performance across different configurations and to identify changes that improve or degrade filtering quality.
Reproducibility Considerations
Reproducibility is a major concern in bioinformatics workflows. The nf-core community provides documentation on pipeline standards and reproducible workflow practices that are relevant to variant calling pipelines [<a href="#ref-6">6</a>]. Following these standards helps ensure that your VQSR workflow produces consistent results across runs and across different computing environments.
The Galaxy Training Network provides accessible workflow training and analysis tutorials that cover reproducible analysis practices [<a href="#ref-7">7</a>]. These resources are useful for learning how to structure your VQSR workflow for reproducibility.
Quality Controls and Validation
VQSR is not a substitute for other quality control measures. You should validate your filtered variant call set using multiple approaches.
Concordance with Known Variants
Comparing your filtered call set with known variant sites provides an estimate of sensitivity. If your call set is missing a large proportion of known variants, this may indicate that VQSR is too stringent or that the underlying variant calling had problems.
The NCBI maintains databases of known genetic variation that can be used for concordance checks [<a href="#ref-4">4</a>]. These resources provide independent validation of your variant calls.
Ti/Tv Ratio Assessment
The Ti/Tv ratio is a useful summary statistic for germline variant quality. A ratio that is too low suggests false-positive contamination, while a ratio that is too high may indicate that true variants are being removed.
The expected Ti/Tv ratio depends on the genomic region and the population. Whole-genome data typically have a ratio around 2.0 to 2.1, while exome data have a higher ratio around 2.8 to 3.0. Deviations from these expected values warrant investigation.
Visual Inspection of Variants
Manual review of a subset of variants in a genome browser can reveal systematic problems that are not apparent from summary statistics. The VariFAST study noted that manual review is a common approach for false-positive variant filtering, though it is labor-intensive and subject to inter- and intra-lab variability [<a href="#ref-2">2</a>]. Automated approaches such as VariFAST aim to reduce the need for manual review while maintaining consistency with manual review decisions [<a href="#ref-2">2</a>].
Comparison with Alternative Callers
Comparing your VQSR-filtered call set with variants called by an alternative caller can identify systematic biases. The GermVarX workflow integrates GATK HaplotypeCaller and DeepVariant, with consensus generation between callers to increase reliability [<a href="#ref-3">3</a>]. This approach provides a cross-validation of variant calls that is not available from a single caller.
Limitations of VQSR
VQSR has several limitations that you should understand before applying it to your data.
Not Suitable for Small Call Sets
VQSR requires a minimum number of variant sites to build a reliable model. For single-sample exomes or targeted panels, the number of sites is often insufficient, and hard filtering is more appropriate.
Reliance on Training Resources
VQSR depends on the quality and completeness of training resources. If the training resources are outdated, misaligned, or biased toward certain populations, the model will reflect these biases. The NCBI provides current versions of variation databases, and you should use the most recent versions that are compatible with your reference genome [<a href="#ref-4">4</a>].
Population Bias
Training resources are derived from populations that may not represent your sample. If your sample comes from a population that is underrepresented in the training resources, VQSR may misclassify true variants as false positives. This is a particular concern for non-European populations, which are underrepresented in many variation databases.
Limited to Germline Data
VQSR is designed for germline variant calling and is not appropriate for somatic variant calling. Somatic data have different annotation distributions due to subclonal mutations and copy number alterations. The VariFAST approach provides an alternative that handles both germline and somatic variant filtering [<a href="#ref-2">2</a>].
Model Interpretability
The Gaussian mixture model used by VQSR is a black box in the sense that it does not provide a simple rule for why a variant passed or failed filtering. This lack of interpretability can be a problem in clinical applications where you need to explain filtering decisions.
Alternatives to VQSR
Several alternatives to VQSR have been developed, and you should be aware of these options when deciding on your filtering strategy.
Hard Filtering
Hard filtering applies fixed thresholds to individual annotations. For example, you might filter out variants with QD below 2.0 or FS above 60. Hard filtering is simpler than VQSR and does not require training resources, but it is less flexible and may not capture complex annotation interactions.
Hard filtering is appropriate for small call sets where VQSR cannot be applied. It is also useful as a complement to VQSR for removing obvious artifacts.
Machine Learning Approaches
Several machine learning approaches have been developed for variant filtering. The VariFAST method uses a predictive model trained with the XGBOOST algorithm for germline variant refinement, and it demonstrated better Matthews correlation coefficient and area under the curve than VQSR, particularly for indel filtering [<a href="#ref-2">2</a>].
The VariantTransformer framework uses a Transformer architecture to model dependencies among variant features and processes VCF files directly, enabling integration with standard pipelines such as BCFTools and GATK4 [<a href="#ref-8">8</a>]. This approach demonstrated consistent improvements over baseline filtering accuracy across tested samples [<a href="#ref-8">8</a>].
These machine learning approaches offer potential advantages over VQSR, but they require training data and may not generalize to all data types. You should evaluate these approaches on your own data before adopting them.
Consensus Approaches
Consensus approaches combine variant calls from multiple callers and retain only variants that are called by more than one caller. The GermVarX workflow supports consensus generation between GATK HaplotypeCaller and DeepVariant, coupled with sample- and cohort-level quality control [<a href="#ref-3">3</a>]. This approach increases reliability by requiring agreement between independent callers.
Consensus approaches are computationally more expensive than single-caller approaches, but they can improve variant calling accuracy, particularly for challenging genomic regions.
Safety and Regulatory Context
Variant filtering decisions have implications for downstream analysis, and in clinical contexts, these decisions can affect patient care. You should be aware of the regulatory and ethical considerations that apply to your work.
Clinical Variant Interpretation
If your variant calls are used for clinical interpretation, the filtering strategy must be documented and validated. The American College of Medical Genetics and Genomics provides guidelines for variant interpretation that specify the evidence levels required for different classifications. VQSR filtering is one component of the variant calling process, and the filtering parameters should be reported in any clinical variant report.
Data Privacy and Security
Germline variant data are sensitive personal information. You should ensure that your data storage and analysis procedures comply with applicable privacy regulations. The NCBI provides information on data submission and access policies for genetic data [<a href="#ref-4">4</a>].
Reproducibility Requirements
Regulatory and funding agencies increasingly require reproducible analysis workflows. The nf-core documentation provides standards for pipeline development and usage that support reproducibility [<a href="#ref-6">6</a>]. The Carpentries lessons provide foundational training in computing and data skills that support reproducible research practices [<a href="#ref-9">9</a>].
Professional Escalation Criteria
Knowing when to escalate a VQSR problem to a more experienced colleague or to seek external support is important for avoiding wasted effort and incorrect results.
Escalate When VQSR Fails Repeatedly
If VQSR fails with the same error multiple times despite troubleshooting, you should escalate the problem. Repeated failures may indicate a fundamental issue with the input data or the configuration that requires expert attention.
Escalate When Filtered Results Are Unexpected
If the filtered call set has unexpected characteristics, such as a very low Ti/Tv ratio or a high proportion of variants in known artifact regions, you should escalate the problem. Unexpected results may indicate problems with the variant calling step, the training resources, or the VQSR configuration.
Escalate When Clinical Decisions Are Affected
If your variant calls are used for clinical decisions, any uncertainty about the filtering strategy should be escalated to a clinical genomics expert. The consequences of incorrect variant filtering in a clinical context can be severe, and expert review is warranted.
Escalate When You Cannot Explain Filtering Decisions
If you cannot explain why a specific variant passed or failed VQSR filtering, you should escalate the question. The ability to explain filtering decisions is important for quality assurance and for regulatory compliance.
Practical Implementation Steps
The following steps provide a practical implementation guide for VQSR in a germline variant calling workflow.
Step 1: Assess Your Data
Before running VQSR, assess whether your data are suitable for this approach. Count the number of variant sites in your raw call set. If you have fewer than 30,000 sites, consider hard filtering instead of VQSR.
Step 2: Verify Input Files
Verify that all input files use the same reference genome version. Check that the training resources are compatible with your reference genome and that the annotation fields are present in your input VCF.
Step 3: Configure the Recalibration
Configure the VariantRecalibrator tool with the appropriate training resources, priors, and annotations for your data type. Use the recommended values from GATK documentation as starting points, and adjust based on your data characteristics.
Step 4: Run SNP Recalibration
Run VariantRecalibrator for SNPs first. Examine the output to confirm that the model converged and that the tranche file contains the expected sensitivity levels.
Step 5: Run Indel Recalibration
Run VariantRecalibrator for indels using the SNP-recalibrated call set as input. Indel recalibration is more challenging, and you should examine the output carefully for signs of poor model performance.
Step 6: Apply the Filter
Run ApplyVQSR with the selected tranche threshold. Verify that the filtered call set contains the expected number of variants and that the Ti/Tv ratio is within the expected range.
Step 7: Validate the Results
Validate the filtered call set using concordance checks, Ti/Tv ratio assessment, and comparison with alternative callers if available. Document the results for reproducibility.
Step 8: Document the Configuration
Record the full VQSR configuration, including tool versions, resource versions, parameters, and validation results. This documentation is essential for reproducibility and for troubleshooting future problems.
Common Failure Patterns and Troubleshooting
The following table summarizes common VQSR failure patterns and their solutions.
| Failure Pattern | Symptom | Likely Cause | Troubleshooting Step |
|---|---|---|---|
| Model does not converge | Warning from VariantRecalibrator | Insufficient sites or poor annotation distributions | Check site count, examine annotation distributions, consider hard filtering |
| Too many false positives in filtered set | Low Ti/Tv ratio, high variant count | Insufficient dbSNP prior, poor training resources | Increase dbSNP prior, verify training resource versions |
| Too many true variants removed | Low sensitivity, missing known variants | Tranche threshold too stringent, poor model fit | Lower tranche threshold, check annotation set |
| Error about missing annotations | VariantRecalibrator fails | Annotations not present in input VCF | Add annotations with VariantAnnotator |
| Error about mismatched contigs | VariantRecalibrator fails | Reference genome mismatch between files | Verify reference genome versions of all inputs |
The troubleshooting table provides a starting point for diagnosing VQSR problems. If the suggested steps do not resolve the issue, escalate to a more experienced colleague or seek support from the GATK community.
Welfare and Ethical Considerations
While VQSR is a technical bioinformatics procedure, it operates within a broader context of responsible genomics research. The following considerations are relevant to the ethical conduct of variant calling and filtering.
Transparency in Methods
Transparent reporting of variant filtering methods is essential for scientific reproducibility. The nf-core community emphasizes the importance of transparent and reproducible pipeline standards [<a href="#ref-6">6</a>]. When you publish results based on VQSR-filtered variants, you should report the exact configuration used, including tool versions, resource versions, and parameters.
Data Sharing
Sharing variant data and filtering configurations contributes to the collective knowledge of the genomics community. The NCBI provides data submission systems that support the sharing of genetic variation data [<a href="#ref-4">4</a>]. The EMBL-EBI provides training on data resource usage that covers data sharing practices [<a href="#ref-5">5</a>].
Avoiding Bias
VQSR training resources may contain population biases that affect filtering decisions. You should be aware of these biases and consider their impact on your results, particularly if your samples come from populations that are underrepresented in training resources.
Continuous Learning
The field of variant calling is evolving rapidly, and new methods and resources are continually being developed. The EMBL-EBI training program provides learning pathways for bioinformatics data resources and practical analysis education [<a href="#ref-5">5</a>]. The Galaxy Training Network provides accessible workflow training that covers current best practices [<a href="#ref-7">7</a>]. Engaging with these resources helps you stay current with developments in variant filtering.
Frequently Asked Questions
What is the minimum number of variant sites required for VQSR?
VQSR requires a sufficient number of variant sites to build a reliable Gaussian mixture model. In practice, at least 30,000 variant sites are recommended for stable model estimation. Whole-genome data typically produce millions of variant sites and easily meet this requirement. Single-sample exomes often produce fewer than 30,000 sites and may not be suitable for VQSR. A cohort of exomes that has been jointly genotyped produces a multi-sample VCF with more sites than any single exome, which improves VQSR performance [<a href="#ref-3">3</a>]. If your call set has too few sites, use hard filtering instead of VQSR.
Why does VQSR use dbSNP as a false-positive resource?
The dbSNP database contains a large number of sites that have been observed in sequencing studies but that have not been validated as true polymorphisms. These sites include sequencing artifacts that have accumulated in the database over many years. The VQSR model uses the presence of a variant in dbSNP as evidence that the site may be a false positive, while the absence from dbSNP combined with strong annotation values supports a true-positive call. The NCBI provides official access to the dbSNP database and documentation on its contents [<a href="#ref-4">4</a>].
Can VQSR be used for somatic variant calling?
VQSR is designed for germline variant calling and is not appropriate for somatic variant calling. Somatic data have different annotation distributions due to subclonal mutations and copy number alterations that distort the annotation distributions VQSR depends on. The VariFAST method provides an alternative automated scoring approach that handles both germline and somatic variant filtering through weighted metrics and machine learning [<a href="#ref-2">2</a>]. If you are working with somatic data, use somatic-specific filtering approaches instead of VQSR.
What is the difference between VQSR and hard filtering?
VQSR uses a Gaussian mixture model to learn the annotation distributions of true variants and false positives from training resources, then applies a tranche threshold to filter variants. Hard filtering applies fixed thresholds to individual annotations, such as filtering out variants with QD below 2.0 or FS above 60. VQSR is more flexible because it captures complex interactions between annotations, but it requires a sufficient number of variant sites and high-quality training resources. Hard filtering is simpler and does not require training resources, making it appropriate for small call sets where VQSR cannot be applied.
How do I choose the tranche threshold for VQSR?
The tranche threshold determines the sensitivity of the final filtered call set. A threshold of 99.0 percent means that the filtered set is expected to contain 99 percent of the true variants in the unfiltered call set. The choice of threshold involves a tradeoff between sensitivity and precision. Clinical applications may require higher sensitivity, while population genetics studies may tolerate lower sensitivity in exchange for higher precision. A common choice is 99.0 percent, with more stringent thresholds such as 99.9 percent used when false positives are particularly problematic.
What annotations are required for VQSR?
The standard VQSR annotation set includes quality by depth (QD), mapping quality (MQ), Fisher strand bias (FS), strand odds ratio (SOR), mapping quality rank sum test (MQRankSum), read position rank sum test (ReadPosRankSum), and variant quality (QUAL). These annotations must be present in the input VCF before running VariantRecalibrator. If you used GATK HaplotypeCaller to generate the raw variants, these annotations are included by default. If you used a different variant caller, you may need to add annotations using the VariantAnnotator tool before running VQSR.
Why does VQSR perform poorly for indel filtering?
Indel recalibration is more challenging than SNP recalibration because indels are less abundant and have more variable annotation distributions. The VariFAST study found that their machine learning approach outperformed VQSR specifically in indel filtering, suggesting that VQSR may have limitations for indel variant refinement [<a href="#ref-2">2</a>]. If you observe poor indel filtering performance, consider alternative approaches such as VariFAST or the VariantTransformer framework [<a href="#ref-8">8</a>].
What should I do if VQSR fails with an error?
If VQSR fails with an error, first check the error message for specific information about the cause. Common causes include missing annotations, mismatched reference genomes, and insufficient variant sites. Verify that all input files use the same reference genome version, that the required annotations are present in the input VCF, and that the call set contains enough variant sites for model building. If the error persists after troubleshooting, escalate the problem to a more experienced colleague or seek support from the GATK community.
Related Bioinformatics Guides
- Single-Cell RNA Sequencing Quality Control: A Practical Guide to Filtering and Metrics
- Detecting Structural Variants with Long-Read Sequencing: Methods and Considerations
- Metabolomics Data Analysis in R: A Practical Workflow
- Metagenomics Tools: A Practical Guide to Software and Pipelines
- Microbiome Data Analysis in R: A Practical Guide for Compositional Data
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
[1] [Fifteen Years of the Genome Analysis Toolkit as the De Facto Standard in Short-Read Variant Calling.](https://doi.org/10.3390/ijms27093754). 2026. [2] [VariFAST: a variant filter by automated scoring based on tagged-signatures.](https://pubmed.ncbi.nlm.nih.gov/31888441). BMC bioinformatics, 2019. [3] [GermVarX: A Robust Workflow for Joint Germline Variant Exploration in whole-exome sequencing cohorts.](https://doi.org/10.1371/journal.pone.0345561). 2026. [4] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information. [5] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute. [6] [nf-core Documentation](https://nf-co.re/docs). nf-core. [7] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project. [8] [A Transformers-based framework for refinement of genetic variants.](https://doi.org/10.3389/fbinf.2025.1694924). 2025. [9] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.