Variant Filtering Strategies: Hard Filters vs. Machine Learning Approaches (VQSR) in GATK

By Dr. Zubair Khalid, DVM, MS, PhD ·

Variant Filtering Strategies: Hard Filters vs. Machine Learning Approaches (VQSR) in GATK

Key Takeaways

  • Hard filtering applies fixed thresholds to individual variant annotations (e.g., QD < 2.0, FS > 60.0) and is suitable for small cohorts or single samples due to its transparency and minimal computational requirements.
  • Variant Quality Score Recalibration (VQSR) employs machine learning to model annotation distributions, requiring large cohorts (millions of variant sites, typically 50+ exomes or whole genomes) and public variant databases for training to effectively distinguish true variants from artifacts.
  • The primary decision criterion between hard filters and VQSR is dataset size; VQSR's ability to model joint annotation distributions offers higher sensitivity for rare variants in large datasets, while hard filters provide a transparent, reproducible approach for smaller cohorts.
  • Implementing hard filters involves using GATK's VariantFiltration tool with user-defined thresholds, while VQSR requires GATK's VariantRecalibrator for model training and ApplyVQSR for applying recalibration scores based on a chosen sensitivity threshold (e.g., 99.0%).
  • Both methods have limitations: hard filters lack adaptability to cohort-specific errors and treat annotations independently, whereas VQSR is computationally intensive, requires substantial training data, and its model-based nature is less transparent.
  • Quality control metrics like the Transition to Transversion (Ti/Tv) ratio (expected ~2.0-2.1 for WGS, ~2.8-3.0 for exomes) and genotype concordance are crucial for evaluating the performance of either filtering strategy.

Germline variant filtering is the step in a variant calling workflow where candidate variants are separated from sequencing and alignment artifacts before downstream interpretation. The decision between hard filters and Variant Quality Score Recalibration (VQSR) depends on dataset size, computational resources, and the tolerance for false positives in your specific analysis. For small cohorts or single samples, hard filters are practical and transparent. For large cohorts, VQSR leverages machine learning to model error patterns across many samples and typically preserves more true variants while removing systematic artifacts. This article provides a direct comparison of both approaches, concrete implementation steps, and decision criteria for choosing between them.

Scope and Reader Context

This article addresses bioinformatics students, researchers, laboratory professionals, and life-science practitioners who need to make informed decisions about germline variant filtering in GATK-based pipelines. The practical outcome is a clear framework for choosing between hard filtering and VQSR based on your data and infrastructure. The content assumes familiarity with basic variant calling concepts but does not require advanced statistical training. We focus on germline variant calling workflows, where the distinction between hard filters and VQSR is most consequential. Somatic variant calling follows different principles and is mentioned only for contrast.

The decision between hard filters and VQSR is not a one-time choice. It affects downstream analyses including population genetics, disease association studies, and clinical genomics. Accurate identification of germline variants from whole-exome sequencing data is foundational to these applications, yet variant calling across cohorts poses challenges in scalability, consistency, and reproducibility. The filtering strategy you select directly influences whether your results can be compared across studies and whether your conclusions are robust to technical variation.

Understanding Variant Filtering in Germline Calling Workflows

Variant filtering occurs after variant calling and before functional annotation or statistical analysis. The goal is to distinguish true biological variants from artifacts introduced during library preparation, sequencing, alignment, or the variant calling algorithm itself. In germline workflows, the expectation is that most variants are inherited and present in a predictable fraction of reads. This expectation enables both hard filters and machine learning approaches to identify outliers that likely represent errors.

The raw output of a variant caller such as GATK HaplotypeCaller is a VCF file containing every candidate variant that passed the caller's internal thresholds. These candidates include true variants, sequencing errors that happen to align cleanly, alignment artifacts near indels, and systematic errors related to sequence context. Without filtering, downstream analyses will be contaminated by false positives that can obscure real biological signals.

The filtering step is distinct from variant calling itself. Variant calling determines genotypes and assigns quality scores. Filtering applies thresholds or models to decide which variants are reliable enough for downstream use. In GATK, each variant record contains annotation fields such as quality by depth (QD), mapping quality (MQ), Fisher strand bias (FS), and strand odds ratio (SOR). These annotations describe different aspects of evidence supporting each variant. Hard filters apply fixed thresholds to these annotations. VQSR combines many annotations into a single probability score using a trained model.

The choice of filtering strategy affects both sensitivity and precision. Sensitivity refers to the proportion of true variants retained. Precision refers to the proportion of retained variants that are true. Hard filters with stringent thresholds reduce false positives but may discard true variants with unusual annotation values. VQSR can model the joint distribution of annotations and identify variants that are likely true even when individual annotations fall outside typical ranges.

Core Principles of Hard Filtering

Hard filtering applies fixed thresholds to individual variant annotations. Each variant is evaluated independently, and variants failing any threshold are flagged or removed. The approach is deterministic, transparent, and easy to implement in any pipeline. It requires no training data and no special computational resources beyond the variant calling step itself.

The standard hard filter recommendations for germline data target specific failure modes. Quality by depth (QD) measures the variant quality score divided by the depth of coverage at the variant site. Low QD values indicate variants supported by few reads or by reads with low quality. Mapping quality (MQ) measures the average mapping quality of reads supporting the variant. Low MQ suggests the variant is in a region where reads align ambiguously. Fisher strand bias (FS) and strand odds ratio (SOR) measure whether variant-supporting reads are concentrated on one strand. Strand bias often indicates artifacts from PCR duplication or alignment errors. Read position rank sum (ReadPosRankSum) compares the position of variant-supporting reads within their reads to the position of reference-supporting reads. Variants near read ends are more likely to be artifacts.

Hard filtering is appropriate when you have a small number of samples, when you need a transparent and reproducible filter, or when you lack the computational resources for VQSR. The thresholds are applied uniformly, which means the same criteria are used for every variant regardless of the overall data quality. This uniformity is both a strength and a limitation. It ensures consistency across analyses but cannot adapt to cohort-specific error patterns.

The main limitation of hard filtering is that it treats each annotation independently. A variant might fail one annotation threshold while having strong evidence across all other annotations. Hard filtering would discard this variant even though it is likely true. Conversely, a variant might pass all individual thresholds while having a combination of annotation values that is characteristic of systematic artifacts. Hard filtering would retain this variant even though it is likely false.

Core Principles of VQSR

Variant Quality Score Recalibration (VQSR) uses machine learning to model the distribution of annotations for true variants and for probable artifacts. The method requires a set of known true variants, typically from databases such as those maintained by the National Center for Biotechnology Information (NCBI). These known variants serve as training data. The model learns which annotation combinations are characteristic of true variants in your specific dataset.

VQSR works by building a Gaussian mixture model that separates variants into clusters based on their annotation values. The model is trained on variants that match known true sites and variants that do not. After training, every variant in the dataset receives a probability of being true. You then apply a sensitivity threshold that determines the tradeoff between retaining true variants and accepting false positives.

The key requirement for VQSR is a sufficient number of variants for training. The GATK documentation recommends at least 30 million sites for the model to perform reliably. This requirement means VQSR is appropriate for whole genome sequencing data or for large whole-exome cohorts where the combined variant count across all samples is high. For a single exome or a small cohort, the number of variant sites is typically too low for the model to converge on a stable error distribution.

VQSR is computationally intensive. The model training step requires substantial memory and processing time. For large cohorts, this cost is justified by the improved accuracy. For small datasets, the cost may exceed the benefit because hard filters can achieve comparable accuracy with less complexity.

The advantage of VQSR is its ability to model the joint distribution of annotations. A variant with an unusual value for one annotation but strong evidence across all others can be retained if the model determines it is likely true. This capability is especially valuable for detecting rare variants with modest effects, which are increasingly recognized as important contributors to complex disease risk. Rare variants often have annotation values that differ from common variants because they are supported by fewer reads and may occur in difficult-to-sequence regions.

At a Glance: Hard Filters vs. VQSR

Decision CriterionHard FiltersVQSR
Recommended dataset sizeSingle samples, small cohorts, exomes with fewer than 30 samplesLarge cohorts, whole genomes, datasets with millions of variant sites
Training data requiredNoneKnown true variant sites from public databases
Computational resourcesMinimal, runs in seconds to minutesSubstantial, requires significant memory and processing time
TransparencyFully transparent, fixed thresholds are easy to documentModel-based, thresholds are derived from data and harder to explain
Adaptability to cohort-specific errorsLow, same thresholds apply to all dataHigh, model learns from your specific data
Sensitivity for rare variantsMay discard true variants with unusual annotationsCan retain true variants with unusual annotation combinations
Reproducibility across studiesHigh, same thresholds produce same resultsModerate, model depends on training data and cohort composition
Best use caseClinical single-sample analysis, small research studies, limited infrastructurePopulation studies, large cohort analyses, whole genome sequencing

Dataset Size and the Decision Boundary

The most important factor in choosing between hard filters and VQSR is the number of variant sites in your dataset. This number depends on the number of samples, the sequencing depth, and whether you are analyzing whole genomes or whole exomes. A whole genome from a single individual contains roughly 3 to 4 million variant sites. A whole exome from a single individual contains roughly 20,000 to 50,000 variant sites. A cohort of 100 exomes might contain 200,000 to 500,000 variant sites after joint genotyping.

VQSR requires enough variant sites for the Gaussian mixture model to distinguish true variants from artifacts. With too few sites, the model cannot estimate the parameters of the error distribution reliably. The result is a model that either overfits the training data or fails to separate true variants from artifacts. In practice, VQSR becomes reliable when the dataset contains at least several hundred thousand variant sites. This threshold is typically reached with whole genome data from a single sample or with whole-exome cohorts of 50 or more samples.

Hard filters do not have this minimum data requirement. They can be applied to a single variant or to millions of variants with identical logic. The thresholds are fixed and do not depend on the data. This property makes hard filters the default choice for small datasets and for clinical applications where transparency and reproducibility are paramount.

The decision boundary is not absolute. Some researchers use hard filters for small cohorts and VQSR for large cohorts within the same study. Others use hard filters exclusively because their computational infrastructure cannot support VQSR. The choice should be documented in the methods section of any report or publication, along with the specific thresholds or sensitivity settings used.

Practical Workflow for Hard Filtering

Implementing hard filtering in GATK follows a straightforward sequence. The steps assume you have already completed variant calling and joint genotyping if you have multiple samples. The filtering step operates on the VCF file produced by GenotypeGVCFs.

Step 1: Select Variant Annotations

Before applying filters, confirm that your VCF contains the annotations needed for filtering. The standard annotations are QD, MQ, FS, SOR, and ReadPosRankSum. These annotations are calculated by HaplotypeCaller and included in the output by default. If you are using a different variant caller, verify that the annotations are present or calculate them separately.

Step 2: Apply Hard Filters

Use GATK VariantFiltration to apply thresholds to each annotation. The tool evaluates each variant against the specified thresholds and adds a filter status to variants that fail. The filter status is recorded in the FILTER column of the VCF. Variants that pass all thresholds are marked as PASS.

The specific thresholds depend on your data and your tolerance for false positives. Common starting points for germline data are QD less than 2.0, FS greater than 60.0, MQ less than 40.0, SOR greater than 3.0, and ReadPosRankSum less than negative 8.0. These values are starting points, not universal standards. You should examine the distribution of annotations in your own data and adjust thresholds based on the observed patterns.

Step 3: Evaluate Filtering Results

After applying filters, examine the number of variants that passed and failed. Compare the transition to transversion ratio (Ti/Tv) for passed and failed variants. A Ti/Tv ratio around 2.0 to 2.1 for whole genome data and around 2.8 to 3.0 for exome data is typical for true variants. A lower Ti/Tv ratio among failed variants suggests the filters are removing artifacts. A similar Ti/Tv ratio between passed and failed variants suggests the filters may be too lenient or too stringent.

Step 4: Document Thresholds

Record the exact thresholds used and the version of GATK in your methods. This documentation is essential for reproducibility. Other researchers should be able to apply the same thresholds to their data and obtain comparable results.

Practical Workflow for VQSR

Implementing VQSR requires more preparation than hard filtering. The workflow assumes you have a large dataset with sufficient variant sites and access to known true variant databases.

Step 1: Prepare Training Resources

VQSR requires a set of known true variants for training. These are typically obtained from public databases such as those maintained by the NCBI. The training set should contain variants that are well validated and unlikely to be artifacts. The GATK best practices recommend using a specific set of known variants, but the exact resource depends on the reference genome version and the population being studied.

Step 2: Run Variant Recalibration

Use GATK VariantRecalibrator to build the model. The tool takes your variant calls and the training resources as input. It calculates a score for each variant based on the model. The output is a recalibration file that assigns a probability of being true to each variant.

The VariantRecalibrator step requires substantial memory. For whole genome data, the tool may need 16 to 32 gigabytes of RAM. For large cohorts, the memory requirement can be higher. Ensure your computing environment can accommodate this requirement before starting.

Step 3: Apply the Recalibration

Use GATK ApplyVQSR to assign filter status to each variant based on the recalibration scores. You specify a sensitivity threshold, typically 99.0 or 99.5 percent. This threshold determines the proportion of true variants that are retained. A higher sensitivity threshold retains more true variants but also accepts more false positives.

Step 4: Evaluate the Model

Examine the output of VariantRecalibrator to assess model quality. The tool produces plots showing the distribution of variants by annotation value and the separation between true variants and artifacts. Poor separation suggests the model is not distinguishing well, which may indicate insufficient training data or problematic annotations.

Step 5: Document Settings

Record the training resources, sensitivity threshold, and GATK version in your methods. VQSR results depend on these settings, and other researchers need this information to reproduce your analysis.

Options and Tradeoffs in Filtering Strategies

The choice between hard filters and VQSR is not the only decision in variant filtering. Several related options affect the outcome and should be considered together.

Joint Genotyping and Filtering

Joint genotyping of multiple samples produces a single multi-sample VCF that can be filtered as a unit. This approach enables consistent filtering across all samples and improves the accuracy of genotype calls in low-coverage regions. A workflow such as GermVarX implements joint variant calling with GATK HaplotypeCaller and DeepVariant, followed by joint genotyping with GATK or GLnexus. The workflow supports consensus generation between callers and sample-level and cohort-level quality control. This approach is appropriate for large whole-exome sequencing cohorts where consistency and reproducibility are priorities.

Joint genotyping is particularly important for VQSR because the model benefits from the larger number of variant sites in a multi-sample VCF. A single sample exome may have too few sites for VQSR, but a cohort of 100 exomes may have enough. Joint genotyping also enables the detection of variants that are present in only one sample but are supported by the overall data quality across the cohort.

Consensus Filtering Between Callers

Some workflows combine variant calls from multiple callers and retain only variants that are detected by more than one caller. This consensus approach reduces false positives but may also reduce sensitivity for true variants that one caller detects and another misses. The GermVarX workflow supports consensus generation between GATK HaplotypeCaller and DeepVariant, providing an additional layer of filtering beyond the per-caller quality scores.

Consensus filtering is a complement to hard filters or VQSR, not a replacement. You can apply hard filters to each caller's output and then take the intersection of the filtered calls. This approach is more conservative than using a single caller and is appropriate when false positives are especially costly.

Functional Annotation and Filtering

After variant filtering, functional annotation adds biological context to each variant. Tools such as the Variant Effect Predictor (VEP) classify variants by their predicted effect on genes and proteins. This annotation is essential for interpreting the clinical or biological significance of retained variants. The GermVarX workflow integrates VEP annotation and produces unified reports through MultiQC, facilitating downstream interpretation.

Functional annotation can also inform filtering decisions. For example, you might apply more stringent filters to variants in coding regions and more lenient filters to variants in non-coding regions. This approach prioritizes variants most likely to have biological impact.

Observations and Measurements for Filtering Decisions

The choice of filtering strategy should be informed by measurements from your own data. Several metrics provide insight into whether your filtering approach is working as intended.

Transition to Transversion Ratio

The Ti/Tv ratio compares the number of transitions (A to G, C to T) to transversions (all other substitutions). Transitions are more common than transversions in the human genome, so the Ti/Tv ratio is typically above 1.0. For whole genome data, a Ti/Tv ratio around 2.0 to 2.1 is expected. For exome data, the ratio is higher, around 2.8 to 3.0, because coding regions are enriched for transitions.

A low Ti/Tv ratio among passed variants suggests that false positives remain in the dataset. A high Ti/Tv ratio among failed variants suggests that true variants are being removed. Comparing the Ti/Tv ratio between passed and failed variants provides a quick check on filter performance.

Variant Count and Density

The number of variants per sample and the density of variants across the genome provide context for filtering decisions. An unusually high number of variants in one sample may indicate contamination or sample mix-up. An unusually low number may indicate poor sequencing quality or excessive filtering.

For population studies, the number of variants should be consistent with the expected diversity of the population being studied. A founder population with a recent bottleneck may have fewer variants and longer runs of homozygosity than an outbred population. These population characteristics affect the distribution of variant annotations and should be considered when evaluating filter performance.

Genotype Concordance

If you have genotype data from an independent platform, such as array-based genotyping, you can compare your variant calls to this reference. Concordance between the sequencing-based calls and the array-based calls provides a direct measure of accuracy. Low concordance suggests problems with variant calling or filtering.

Genotype concordance is especially valuable for validating filtering decisions in clinical applications. A filter that removes true variants will reduce concordance. A filter that retains false positives will also reduce concordance. The concordance rate provides an objective measure of filter performance.

Records and Documentation for Filtering

Reproducibility in variant filtering requires careful documentation of every decision and setting. The following records should be maintained for each analysis.

Pipeline Configuration

Record the exact version of GATK or other tools used, the reference genome version, and the parameters for each step. This information should be stored in a configuration file that is version controlled. Workflow management systems such as Nextflow, as documented by nf-core, provide a framework for reproducible pipeline configuration. The nf-core documentation describes community standards for pipeline usage and configuration that support reproducibility across computing environments.

Filter Thresholds and Settings

Record the specific thresholds applied for hard filters or the sensitivity threshold and training resources used for VQSR. This information should be included in the methods section of any report or publication. Without this documentation, other researchers cannot reproduce your filtering decisions.

Variant Counts at Each Stage

Record the number of variants before filtering, after filtering, and after any additional steps such as consensus calling or functional annotation. These counts provide a summary of the filtering effect and help identify unexpected losses or gains.

Quality Metrics

Record the Ti/Tv ratio, genotype concordance, and other quality metrics for the filtered variant set. These metrics provide evidence that the filtering approach is working and allow comparison across analyses.

Common Failure Patterns in Variant Filtering

Several recurring problems affect variant filtering in germline workflows. Recognizing these patterns helps you diagnose and correct issues in your own analysis.

Overly Stringent Filters

Applying thresholds that are too strict removes true variants along with artifacts. This problem is common when researchers use thresholds from a different dataset without examining their own data. The result is a reduced variant count, a low Ti/Tv ratio among passed variants, and reduced power for downstream analyses.

The solution is to examine the distribution of annotations in your own data before setting thresholds. Plot the distribution of each annotation and identify the values that separate the main cluster of variants from the tail of outliers. Set thresholds at the boundaries of the main cluster instead of at arbitrary values.

Overly Lenient Filters

Applying thresholds that are too permissive retains artifacts in the variant set. This problem is common when researchers are concerned about losing true variants and set thresholds to retain as many variants as possible. The result is an inflated variant count, a low Ti/Tv ratio among passed variants, and false associations in downstream analyses.

The solution is to compare the Ti/Tv ratio of passed and failed variants. If the ratios are similar, the filters are not distinguishing true variants from artifacts. Tighten the thresholds until the failed variants show a clear enrichment for transversions, which are more common in artifacts.

Applying VQSR to Small Datasets

Using VQSR on a dataset with too few variant sites produces an unreliable model. The model may fail to converge, or it may produce scores that do not separate true variants from artifacts. The result is a filter that removes true variants or retains artifacts in an unpredictable manner.

The solution is to check the number of variant sites before running VQSR. If the dataset has fewer than several hundred thousand sites, use hard filters instead. If you must use VQSR, consider combining your data with other samples to increase the variant count.

Ignoring Strand Bias

Strand bias is a common indicator of artifacts, especially in targeted sequencing data. Variants supported predominantly by reads on one strand are often the result of PCR duplication or alignment errors. Ignoring strand bias annotations in hard filters or excluding them from VQSR allows these artifacts to pass.

The solution is to include strand bias annotations in your filtering strategy. For hard filters, apply thresholds to FS and SOR. For VQSR, ensure that strand bias annotations are included in the model.

Inconsistent Filtering Across Batches

Applying different filters to different batches of samples within the same study introduces batch effects that can be mistaken for biological variation. This problem is common when samples are processed at different times or by different facilities.

The solution is to apply the same filtering strategy to all samples in a study. If you must use different filters for different batches, document the differences and account for them in downstream analyses.

Limitations of Hard Filtering and VQSR

Both filtering approaches have limitations that should be acknowledged in any analysis.

Limitations of Hard Filtering

Hard filtering cannot adapt to cohort-specific error patterns. The same thresholds are applied to all variants regardless of the overall data quality. This limitation means that hard filters may be too stringent for high-quality data and too lenient for low-quality data.

Hard filtering also treats each annotation independently. A variant that fails one annotation threshold is removed even if all other annotations indicate high quality. This limitation reduces sensitivity for true variants with unusual annotation values, including some rare variants with modest effects that are increasingly recognized as important in complex disease.

Limitations of VQSR

VQSR requires a large number of variant sites and substantial computational resources. These requirements make VQSR impractical for small datasets and for researchers with limited computing infrastructure.

VQSR also depends on the training data. If the known true variant database is incomplete or biased toward certain populations, the model will be biased accordingly. This limitation is especially relevant for non-European populations, which are underrepresented in many public databases.

The model-based nature of VQSR makes it less transparent than hard filtering. The thresholds are derived from the data and are harder to explain to reviewers or clinicians. This limitation is a concern in clinical applications where transparency is important.

General Limitations

Both approaches assume that the annotations in the VCF accurately reflect variant quality. If the variant caller produces inaccurate annotations, neither hard filters nor VQSR can compensate. This limitation underscores the importance of using a well-validated variant caller and checking annotation quality before filtering.

Both approaches also assume that the reference genome and the variant calling pipeline are appropriate for the data being analyzed. Errors in alignment or reference genome choice can produce systematic artifacts that are not captured by the standard annotations.

Quality Controls and Validation

Quality controls should be applied before, during, and after variant filtering to ensure the reliability of your results.

Pre-Filtering Controls

Before filtering, verify that the variant calling step completed successfully. Check the depth of coverage across samples and genomic regions. Low coverage in some regions may produce unreliable variant calls that no filtering strategy can fully correct.

Check for sample contamination and sample mix-ups. Contamination produces an excess of heterozygous variants and unusual allele frequencies. Sample mix-ups produce genotype patterns that are inconsistent with the expected relationships between samples.

During-Filtering Controls

Monitor the number of variants passing and failing each filter. A sudden drop in the number of passing variants at a specific threshold may indicate that the threshold is too stringent. A lack of change in the number of passing variants may indicate that the threshold is not discriminating.

For VQSR, examine the model output plots. The plots should show clear separation between the clusters of true variants and artifacts. Poor separation indicates that the model is not working well.

Post-Filtering Controls

After filtering, validate the retained variants using independent methods. Genotype concordance with array-based data provides a direct measure of accuracy. Sanger sequencing of a subset of variants provides experimental validation. The GermVarX workflow and similar pipelines support these validation approaches by producing outputs that are compatible with downstream analysis tools.

For clinical applications, validation is especially important. Variants that will be used for diagnosis or treatment decisions should be confirmed by an independent method before reporting.

Safety and Regulatory Context

Variant filtering is a computational step with no direct physical safety implications. However, the results of variant filtering can influence clinical decisions, and errors in filtering can have serious consequences.

In clinical genomics, false positive variants can lead to incorrect diagnoses and inappropriate treatments. False negative variants can lead to missed diagnoses and missed opportunities for intervention. The filtering strategy should be chosen to minimize these risks, and the limitations of the chosen approach should be communicated to clinicians.

Regulatory frameworks for clinical genomics vary by jurisdiction. In general, clinical laboratories are expected to validate their variant calling and filtering pipelines and to document the performance characteristics of these pipelines. The transparency of hard filters is an advantage in this context, as the thresholds can be clearly documented and validated.

For research applications, the regulatory context is less stringent, but the same principles apply. The filtering strategy should be documented in sufficient detail that other researchers can reproduce the analysis. The limitations of the approach should be acknowledged in the interpretation of results.

Professional Escalation Criteria

Variant filtering problems can sometimes indicate issues that require escalation to a specialist or a change in approach. The following criteria suggest that escalation is appropriate.

Escalate When VQSR Fails to Converge

If VariantRecalibrator fails to converge or produces a model with poor separation between true variants and artifacts, escalate the issue. The problem may indicate insufficient variant sites, problematic training data, or issues with the variant calling step. A bioinformatics specialist should review the data and recommend an alternative approach.

Escalate When Filtering Produces Unexpected Results

If the number of variants passing filters is dramatically different from expectations, escalate the issue. The problem may indicate sample contamination, sample mix-up, or a systematic error in the variant calling pipeline. These issues require investigation before the filtering strategy can be adjusted.

Escalate When Genotype Concordance Is Low

If genotype concordance with independent data is below acceptable levels, escalate the issue. The problem may indicate a fundamental issue with the variant calling or filtering approach that cannot be corrected by adjusting thresholds.

Escalate When Clinical Results Are Affected

If variant filtering results will be used for clinical decisions, involve a clinical genomics specialist in the review of the filtering strategy. The specialist can assess whether the approach meets the standards for clinical use and can recommend additional validation steps.

Frequently Asked Questions

What is the minimum dataset size for VQSR?

VQSR requires enough variant sites for the Gaussian mixture model to estimate the error distribution reliably. In practice, this requirement means at least several hundred thousand variant sites. A single whole genome provides enough sites, but a single whole exome does not. For whole-exome cohorts, you typically need 50 or more samples to reach the required variant count. If your dataset has fewer sites, use hard filters instead.

Can I use hard filters on large cohorts?

Yes, hard filters can be applied to datasets of any size. The thresholds are fixed and do not depend on the number of samples. However, for large cohorts, VQSR typically provides better accuracy because it models the joint distribution of annotations and can adapt to cohort-specific error patterns. If you have the computational resources, VQSR is usually the better choice for large datasets.

What annotations should I include in hard filters?

The standard annotations for germline hard filtering are quality by depth (QD), mapping quality (MQ), Fisher strand bias (FS), strand odds ratio (SOR), and read position rank sum (ReadPosRankSum). These annotations capture different aspects of variant evidence and are calculated by GATK HaplotypeCaller by default. You should examine the distribution of each annotation in your own data before setting thresholds.

How do I choose sensitivity thresholds for VQSR?

The sensitivity threshold for VQSR determines the tradeoff between retaining true variants and accepting false positives. A higher sensitivity threshold, such as 99.5 percent, retains more true variants but also accepts more false positives. A lower threshold, such as 99.0 percent, removes more false positives but may discard some true variants. The choice depends on your tolerance for false positives and the consequences of missing true variants.

Does VQSR work for non-human genomes?

VQSR can be applied to non-human genomes if you have a suitable set of known true variants for training. The training data must be appropriate for the species and the reference genome being used. For many non-human species, such training data are not available, and hard filters are the only practical option.

How do I validate my filtering results?

Validation methods include comparing the Ti/Tv ratio of passed and failed variants, checking genotype concordance with independent data, and confirming a subset of variants by Sanger sequencing. These methods provide evidence that the filtering approach is working and help identify problems that need correction.

What is the difference between germline and somatic variant filtering?

Germline variant filtering assumes that variants are inherited and present in a predictable fraction of reads. Somatic variant filtering must account for the possibility that variants are present in only a subset of cells and at low allele fractions. The filtering strategies differ because the error models are different. This article focuses on germline variant filtering, which is the context where hard filters and VQSR are most directly comparable.

How should I document my filtering strategy?

Document the exact thresholds or sensitivity settings, the version of GATK or other tools used, the reference genome version, and the training resources for VQSR. Include this information in the methods section of any report or publication. Store the configuration in a version-controlled file so that the analysis can be reproduced exactly.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.