Hard Filters vs. VQSR for Germline Variant Calling: Which Approach Works Best for Your Data?

By Dr. Zubair Khalid, DVM, MS, PhD ·

Hard Filters vs. VQSR for Germline Variant Calling: Which Approach Works Best for Your Data?

Key Takeaways

  • VQSR leverages machine learning for complex quality metric modeling, offering potential for higher sensitivity and specificity in large human cohorts (hundreds of samples) by learning intricate relationships between annotations and known truth sets. This approach requires substantial variant density for stable model training and is sensitive to platform differences and batch effects.
  • Hard filtering employs fixed, transparent thresholds on variant annotations, making it the practical and defensible choice for small cohorts (under 30 samples), single-sample calls, non-human organisms, or when explicit, auditable documentation is paramount. Its simplicity and reproducibility are advantageous where VQSR's assumptions are not met or truth sets are unavailable.
  • The choice hinges on data characteristics: VQSR demands sufficient variant count and high-quality, organism-specific truth sets (primarily human), whereas hard filtering is universally applicable across species and data types with a reference genome. Exome data, with its inherently lower variant density, often favors hard filtering due to VQSR model instability.
  • VQSR's accuracy is contingent on well-matched truth sets and data homogeneity; deviations can lead to model non-convergence or biased results, necessitating careful validation. Hard filtering's accuracy is directly tied to the judicious selection of thresholds, which may require iterative refinement based on assessment metrics like transition-transversion ratio and concordance with orthogonal data.
  • Computational cost and reproducibility differ significantly: VQSR training is resource-intensive, while hard filtering is computationally inexpensive and inherently transparent, facilitating easier documentation and auditing for clinical reporting. Workflow integration systems like Nextflow can automate both approaches, but careful parameter management and version control are critical for reproducibility.

Germline variant calling produces a raw set of candidate variants that requires quality stratification before downstream analysis. The two dominant strategies for this stratification are hard filtering, which applies fixed thresholds to variant annotations, and Variant Quality Score Recalibration (VQSR), which uses machine learning to model the distribution of variant quality metrics against known truth sets. The choice between these approaches depends on cohort size, available computational resources, the variant caller used, and the downstream application of the variant calls. For small cohorts, single-sample calls, or non-human organisms, hard filtering is often the more practical and defensible choice. For large cohorts with hundreds or thousands of samples, VQSR can provide better sensitivity and specificity when its assumptions are met. This article provides a decision framework for selecting between these approaches, with concrete criteria for implementation, quality assessment, and troubleshooting.

Scope and Reader Context

This guidance addresses bioinformaticians, laboratory professionals, and life-science researchers who need to decide how to filter germline variants in their calling pipelines. The focus is on short-read sequencing data, primarily whole-exome sequencing (WES) and whole-genome sequencing (WGS), processed through standard germline calling workflows. The decision between hard filters and VQSR is one of the most consequential choices in variant calling because it directly determines which variants enter downstream analyses such as association studies, population genetics, and clinical interpretation.

The practical outcome of this article is a clear decision process. You will learn the data requirements for each approach, the specific quality metrics that matter, how to evaluate whether your filtering strategy is working, and when to escalate to alternative strategies. The guidance applies to both research and clinical genomics contexts, though clinical applications require additional considerations beyond the scope of this article.

Understanding the Two Filtering Paradigms

Germline variant callers produce a VCF file containing candidate variants with quality annotations. These annotations include depth, genotype quality, mapping quality, strand bias, and other metrics that distinguish true variants from sequencing and alignment artifacts. The filtering step decides which candidates are reliable enough for downstream use.

Hard Filtering: Fixed Thresholds and Transparent Rules

Hard filtering applies predetermined thresholds to variant annotations. A variant passes if its metrics meet all thresholds and fails if any metric falls outside the acceptable range. For example, a typical hard filter might require a minimum depth of 10 reads, a minimum genotype quality of 20, and a maximum strand bias of 0.01. These thresholds are applied uniformly to all variants.

The primary advantage of hard filtering is transparency. Every variant is evaluated against the same explicit criteria, and the filtering decisions are fully reproducible and explainable. This is particularly valuable in clinical settings where the rationale for variant inclusion or exclusion must be documented. Hard filtering also requires no training data, no special computational resources, and works equally well for any organism with a reference genome.

The primary limitation of hard filtering is that fixed thresholds do not account for the varying difficulty of calling different genomic regions. A depth threshold that works well in uniquely mappable regions may be too stringent in repetitive regions or too permissive in regions with high coverage. Hard filters also discard variants that fail one metric even if all other metrics are excellent, which can reduce sensitivity for true variants with unusual characteristics.

VQSR: Machine Learning with Truth Sets

VQSR uses a Gaussian mixture model to learn the relationship between variant annotations and variant quality. The model is trained on known true variants from truth sets such as the Genome in a Bottle Consortium reference materials, along with known false positives from sites that are homozygous reference in the truth set. The trained model then assigns a quality score to every variant in the call set, and variants below a chosen threshold are filtered out.

The advantage of VQSR is its ability to capture complex interactions between quality metrics. Instead of applying independent thresholds, VQSR learns how combinations of metrics distinguish true variants from artifacts. This can improve sensitivity for true variants that would fail a single hard threshold while improving specificity by identifying artifacts that pass all individual thresholds.

The limitations of VQSR are substantial. It requires a large number of variants to build a stable model, typically tens of thousands or more. It requires access to high-quality truth sets that match the sequencing technology and variant types being called. It is sensitive to batch effects and sequencing platform differences. And it is designed for human data, with truth sets that may not transfer to other organisms. For small cohorts, exome capture panels, or non-human species, VQSR may produce unreliable results or fail entirely.

At a Glance: Filtering Approach Decision Table

ScenarioRecommended ApproachPrimary Rationale
Single sample or small cohort (fewer than 30 samples)Hard filteringVQSR requires large variant counts for stable model training, hard filters are transparent and reproducible
Large WGS cohort (hundreds of samples) with human dataVQSRSufficient variant density and truth set availability enable model-based quality stratification
Non-human organism or non-standard referenceHard filteringTruth sets for VQSR are primarily human-focused and may not transfer across species
Clinical reporting with documentation requirementsHard filteringFixed thresholds provide explicit, auditable inclusion and exclusion criteria
Exome data with targeted capture panelsHard filteringLower variant counts and capture bias reduce VQSR model stability
Mixed sequencing platforms or batch effectsHard filteringVQSR is sensitive to platform-specific error profiles and batch variation

Core Principles of Variant Quality Stratification

Before choosing between hard filters and VQSR, it is essential to understand the quality metrics that both approaches use. These metrics describe different aspects of sequencing and alignment evidence for each variant.

Depth and Allele Balance

Depth is the number of reads covering a genomic position. For germline heterozygous variants, the alternate allele should be present in approximately half of the reads. Allele balance, the fraction of reads supporting the alternate allele, is a key indicator of genotype accuracy. Low depth reduces confidence in both variant detection and genotype assignment. Extreme allele balance, either very low or very high, may indicate mapping artifacts or copy number variation instead of true heterozygosity.

Mapping Quality and Strand Bias

Mapping quality reflects the confidence that reads are correctly placed in the reference genome. Low mapping quality suggests reads may originate from multiple genomic locations, which is common in repetitive regions. Strand bias measures whether the alternate allele appears predominantly on one sequencing strand. True variants should appear on both strands, while artifacts often show strong strand bias.

Genotype Quality and Phred-Scaled Quality

Genotype quality reflects confidence in the assigned genotype, while the variant quality score reflects confidence that a variant exists at the position. Both are Phred-scaled, meaning higher values indicate higher confidence. These scores incorporate depth, allele balance, and mapping quality into a single metric.

Context-Specific Artifacts

Certain genomic contexts produce characteristic artifacts. Homopolymer runs and short tandem repeats cause polymerase slippage during library preparation, producing false indels. Regions with high sequence similarity to other genomic locations produce mapping artifacts. Methylated CpG sites are prone to deamination during library preparation, producing C-to-T transitions that may be artifacts instead of true variants. Both hard filters and VQSR must account for these context-specific error patterns.

Data Requirements for Each Approach

The choice between hard filters and VQSR begins with an assessment of your data. The critical factors are cohort size, variant density, sequencing technology, and reference genome.

Cohort Size and Variant Count

VQSR requires a large number of variants to train its model. The model needs enough true variants and enough false positives to learn the distinction between them. With too few variants, the model overfits or fails to converge. The exact minimum depends on the variant caller, the truth set, and the diversity of the data, but cohorts with fewer than 30 samples typically produce too few variants for stable VQSR. Exome data produces fewer variants than whole-genome data because only the coding regions are sequenced, so exome cohorts need more samples to reach the same variant count.

Hard filtering has no minimum variant count. It can be applied to a single sample with a few thousand variants or to a large cohort with millions of variants. The thresholds are set a priori and applied uniformly.

Sequencing Technology and Platform

VQSR models are trained on data from specific sequencing platforms. The error profiles of different platforms differ, and a model trained on one platform may not transfer to another. If your data comes from a platform that differs from the one used to generate the truth set, VQSR may produce biased results. Hard filtering is platform-agnostic because the thresholds are based on variant annotations that are computed by the caller, though the interpretation of those annotations may differ across platforms.

Reference Genome and Organism

VQSR truth sets are primarily developed for the human genome. The Genome in a Bottle Consortium provides high-confidence variant calls for several human reference samples, and these are used to train VQSR models. For non-human organisms, appropriate truth sets may not exist. Hard filtering can be applied to any organism with a reference genome, and thresholds can be adjusted based on the expected variant characteristics of the species.

Practical Workflow for Hard Filtering

Implementing hard filtering requires selecting thresholds, applying them to the call set, and evaluating the results. The following workflow provides a structured approach.

Step 1: Define Filtering Thresholds

Threshold selection should be based on the variant caller's documentation, the sequencing depth of your data, and the downstream application. For GATK HaplotypeCaller, common hard filters include quality by depth, Fisher strand bias, mapping quality, and read position rank sum. For DeepVariant, the quality score itself is often used as the primary filter. The thresholds should be conservative enough to remove obvious artifacts while retaining true variants.

Step 2: Apply Filters to the Call Set

Filters are applied using tools such as GATK SelectVariants or bcftools. The filtering process should be scripted and version-controlled so that the exact filtering criteria are documented and reproducible. The output is a filtered VCF file with variants that pass all thresholds.

Step 3: Evaluate Filtering Performance

After filtering, assess the transition and transversion ratio, the number of variants retained, and the distribution of remaining variants across functional categories. A sudden drop in the transition-transversion ratio may indicate that true variants are being removed. Comparing the filtered call set to known variants in your sample, if available, provides a direct measure of sensitivity and specificity.

Step 4: Document and Iterate

Record the filtering thresholds, the number of variants before and after filtering, and any observations about which variants were removed. If downstream analyses reveal problems, such as an excess of apparent false positives or a loss of known true variants, adjust the thresholds and repeat the process.

Practical Workflow for VQSR

Implementing VQSR requires more preparation than hard filtering. The workflow includes building the model, applying it to the call set, and evaluating the results.

Step 1: Verify Data Suitability

Confirm that your cohort has enough variants for VQSR. Count the number of variants in your raw call set and compare to the recommended minimum for your variant caller. Confirm that your data matches the sequencing platform and variant types used to generate the truth set. For human WGS data with hundreds of samples, VQSR is likely suitable. For exome data or small cohorts, consider hard filtering instead.

Step 2: Prepare Truth Sets

VQSR requires two sets of known variants: true positives and false positives. The true positive set contains variants that are known to exist in the sample, typically from the Genome in a Bottle Consortium. The false positive set contains sites that are known to be homozygous reference, where any variant call is an artifact. These truth sets must be in VCF format and must be compatible with your reference genome build.

Step 3: Train the Model

The VQSR tool builds a Gaussian mixture model using the variant annotations from your call set and the labels from the truth sets. The training process produces a recalibration file that assigns a quality score to every variant. The model should be inspected for convergence and for the number of variants in each training cluster.

Step 4: Apply the Sensitivity Threshold

After training, choose a sensitivity threshold that determines which variants are retained. The threshold is expressed as a sensitivity level, such as 99 percent sensitivity for true variants. Higher sensitivity retains more variants but also retains more false positives. The choice of threshold depends on the downstream application. Research applications may tolerate more false positives, while clinical applications require higher specificity.

Step 5: Evaluate and Document

Assess the number of variants retained, the transition-transversion ratio, and the concordance with known variants. Document the model parameters, the truth sets used, and the sensitivity threshold. If the model fails to converge or produces unexpected results, investigate the cause before proceeding.

Options and Tradeoffs: Hard Filters versus VQSR

The choice between hard filters and VQSR involves tradeoffs across several dimensions. Understanding these tradeoffs helps match the approach to your specific data and analysis goals.

Accuracy and Sensitivity

VQSR has the potential to improve accuracy by modeling complex interactions between quality metrics. In large cohorts with sufficient data, VQSR can identify true variants that fail individual hard thresholds and remove artifacts that pass all thresholds. However, this accuracy depends on the model being well-trained and the truth sets being appropriate. A poorly trained VQSR model can be less accurate than well-chosen hard filters.

Hard filtering is simpler and more predictable. The accuracy depends entirely on the quality of the thresholds. Well-chosen thresholds based on the variant caller's documentation and the characteristics of your data can achieve high accuracy, particularly for high-quality sequencing data.

Scalability and Computational Cost

VQSR requires substantial computational resources for model training, particularly for large cohorts. The training step is computationally intensive and may require significant memory and processing time. The model must be retrained for each new cohort or when the sequencing platform changes.

Hard filtering is computationally inexpensive. The filtering step processes each variant independently and can be completed quickly even for large call sets. The thresholds can be applied to new data without retraining.

Reproducibility and Transparency

Hard filtering is fully transparent. Every variant is evaluated against explicit thresholds that can be documented and audited. This is a significant advantage in clinical settings and in any analysis where the filtering decisions must be explained.

VQSR is a machine learning model, and the filtering decisions are based on the model's learned patterns instead of explicit thresholds. The model parameters and truth sets can be documented, but the rationale for individual variant inclusion or exclusion is less transparent.

Applicability Across Organisms and Data Types

Hard filtering applies to any organism and any sequencing platform. The thresholds may need adjustment, but the approach is universally applicable.

VQSR is primarily designed for human data with appropriate truth sets. For non-human organisms, the lack of truth sets makes VQSR impractical. For exome data, the reduced variant count may make VQSR unstable.

Observations and Measurements for Filtering Assessment

Evaluating whether your filtering approach is working requires systematic measurement of the call set before and after filtering. The following metrics provide a basis for assessment.

Transition-Transversion Ratio

The transition-transversion ratio compares the number of transition mutations to transversion mutations. In human germline data, transitions are more common than transversions, with a typical ratio around 2.0 to 2.1 for whole-genome data. A significantly lower ratio may indicate an excess of artifacts, while a significantly higher ratio may indicate over-filtering of true variants.

Variant Count by Functional Category

Counting variants by functional category, such as synonymous, missense, and loss-of-function, provides a check on the biological plausibility of the call set. An excess of loss-of-function variants may indicate artifacts, while a deficit may indicate over-filtering.

Concordance with Known Variants

If your samples have known variants from orthogonal methods, such as array genotyping or Sanger sequencing, compare the filtered call set to these known variants. The concordance rate provides a direct measure of sensitivity. Discordant calls should be investigated to determine whether they are false positives or false negatives.

Genotype Concordance Across Samples

In cohort studies, check for unexpected patterns in genotype frequencies. An excess of homozygous alternate genotypes at rare variant sites may indicate genotyping errors. Hardy-Weinberg equilibrium tests can identify systematic genotyping problems.

Records and Documentation Requirements

Maintaining detailed records of the filtering process is essential for reproducibility and for troubleshooting downstream analysis problems.

Record the Filtering Parameters

Document the exact thresholds or model parameters used for filtering. For hard filters, record each threshold and the annotation it applies to. For VQSR, record the truth sets, the sensitivity threshold, and the model training parameters.

Record the Input and Output Variant Counts

Record the number of variants in the raw call set, the number after filtering, and the number removed by each filter. This provides a basis for comparing filtering performance across cohorts and for identifying unexpected losses.

Record the Software Versions

Variant callers and filtering tools change between versions, and these changes can affect results. Record the exact software versions used for variant calling and filtering, along with the reference genome build.

Record the Assessment Metrics

Record the transition-transversion ratio, variant counts by functional category, and concordance with known variants for each filtered call set. These metrics provide a baseline for evaluating future filtering decisions.

Common Failure Patterns and Troubleshooting

Both hard filtering and VQSR can fail in predictable ways. Recognizing these failure patterns helps diagnose problems and choose corrective actions.

Hard Filtering Failure Patterns

Over-filtering occurs when thresholds are too stringent, removing true variants along with artifacts. Signs include a low variant count, a low transition-transversion ratio, and poor concordance with known variants. Corrective action involves relaxing thresholds and re-evaluating.

Under-filtering occurs when thresholds are too permissive, retaining artifacts. Signs include an excess of variants in repetitive regions, a high proportion of singletons, and poor validation rates. Corrective action involves tightening thresholds and adding filters for specific artifact types.

Threshold mismatch occurs when thresholds are appropriate for one data type but not another. For example, thresholds tuned for high-depth WGS data may be too stringent for low-depth exome data. Corrective action involves adjusting thresholds based on the depth and characteristics of each dataset.

VQSR Failure Patterns

Model non-convergence occurs when the Gaussian mixture model fails to reach a stable solution. This often results from too few variants or from truth sets that do not match the data. Signs include warnings from the VQSR tool and unusual quality score distributions. Corrective action involves increasing the variant count, using different truth sets, or switching to hard filtering.

Truth set mismatch occurs when the truth sets do not represent the variants in your data. This can happen when the truth sets are from a different population, sequencing platform, or reference build. Signs include poor concordance between the model's predictions and known variants. Corrective action involves obtaining appropriate truth sets or switching to hard filtering.

Batch effects occur when the model is trained on data from one sequencing batch and applied to data from another batch. The model may not generalize across batches with different error profiles. Signs include systematic differences in variant quality between batches. Corrective action involves training separate models for each batch or using hard filtering.

Limitations and Contextual Considerations

Both filtering approaches have limitations that should be acknowledged when interpreting results.

Hard Filtering Limitations

Hard filtering cannot capture complex interactions between quality metrics. A variant with slightly low depth but excellent mapping quality and allele balance may be filtered out even though it is likely true. Conversely, a variant that passes all individual thresholds may still be an artifact if the combination of metrics is unusual.

Hard filtering thresholds are somewhat arbitrary. Different practitioners may choose different thresholds for the same data, leading to different call sets. This reduces comparability across studies and complicates meta-analyses.

VQSR Limitations

VQSR is sensitive to the quality and appropriateness of the truth sets. If the truth sets contain errors or do not represent the variant spectrum in your data, the model will learn incorrect patterns.

VQSR is designed for germline variant calling in human data. Its application to somatic variant calling, non-human organisms, or non-standard sequencing technologies requires careful validation.

VQSR does not eliminate the need for manual review of clinically significant variants. Even well-calibrated models can misclassify individual variants, and clinical interpretation requires additional evidence beyond the quality score.

Data Type Considerations

Whole-exome sequencing produces fewer variants than whole-genome sequencing because only coding regions are captured. This reduced variant count can make VQSR unstable for exome data, even with large cohorts. The capture process also introduces biases in coverage and allele balance that differ from whole-genome sequencing.

Third-generation sequencing technologies, such as those from Pacific Biosciences and Oxford Nanopore, have different error profiles than short-read sequencing. The quality metrics and filtering approaches developed for short-read data may not transfer directly to long-read data. Best practices for these technologies are still evolving, and careful validation is required.

Safety and Regulatory Context

In clinical genomics, the choice of filtering approach has direct implications for patient care. Variant calls that enter clinical reports must meet rigorous standards for accuracy and reproducibility.

Clinical Reporting Requirements

Clinical laboratories must document the rationale for variant inclusion and exclusion. Hard filtering provides explicit, auditable criteria that can be reviewed by laboratory directors and external auditors. VQSR, as a machine learning approach, requires additional documentation of the model training process and validation results.

Validation Requirements

Clinical laboratories must validate their variant calling and filtering pipelines against known reference materials. The validation should demonstrate that the filtering approach achieves acceptable sensitivity and specificity for the intended clinical use. This validation must be repeated when the filtering approach changes or when the pipeline is updated.

Professional Escalation Criteria

If the filtering approach produces unexpected results, such as an excess of apparent false positives, poor concordance with known variants, or model non-convergence, escalate the issue to a bioinformatics specialist or laboratory director. Do not proceed with downstream analysis or clinical reporting until the filtering problem is resolved.

Workflow Integration and Automation

Both hard filtering and VQSR can be integrated into automated analysis pipelines. The choice of workflow management system affects how filtering is implemented and documented.

Reproducible Workflow Systems

Workflow management systems such as Nextflow and Common Workflow Language provide structured frameworks for implementing variant calling and filtering pipelines. These systems support containerization, which ensures that the same software versions are used across analyses. The GermVarX workflow, implemented in Nextflow DSL2, demonstrates how joint germline variant calling and filtering can be automated for cohort studies, with support for both GATK HaplotypeCaller and DeepVariant, and joint genotyping through GATK or GLnexus. This workflow also supports consensus generation between callers, sample- and cohort-level quality control, functional annotation using the Variant Effect Predictor, and unified reporting through MultiQC, with PLINK-compatible outputs for downstream statistical analyses.

Common Workflow Language pipelines have been shown to achieve high accuracy in reproducing published results and detecting germline SNP and small INDEL variants when validated against Genome in a Bottle reference materials. These pipelines combine containerization with explicit workflow definitions to overcome software incompatibility and configuration issues.

Training and Skill Development

Implementing and troubleshooting variant filtering requires foundational bioinformatics skills. Training resources from the Galaxy Training Network provide accessible tutorials for variant calling workflows and quality assessment. The European Bioinformatics Institute offers training pathways for data-resource usage and practical analysis education. The Carpentries provides foundational computing and data skills that support reproducible analysis practices.

Documentation and Version Control

All filtering parameters, software versions, and assessment metrics should be recorded in version-controlled files. This documentation supports reproducibility and enables troubleshooting when downstream analyses reveal problems. The nf-core community standards provide guidance for pipeline documentation and configuration that can be adapted to variant filtering workflows.

Decision Framework for Your Data

The following decision framework integrates the considerations discussed above into a practical process for choosing between hard filters and VQSR.

Step 1: Assess Your Data

Count the number of samples in your cohort and the number of variants in your raw call set. Determine whether your data is whole-genome or whole-exome. Identify the sequencing platform and reference genome build.

Step 2: Assess Your Requirements

Determine the downstream application of the variant calls. Research applications may tolerate more false positives, while clinical applications require higher specificity. Determine whether you need to document the filtering rationale for regulatory or publication purposes.

Step 3: Evaluate VQSR Suitability

If your cohort has sufficient variants, your data is human, and appropriate truth sets are available, VQSR may be suitable. If your cohort is small, your data is exome, or your organism is non-human, hard filtering is likely the better choice.

Step 4: Implement and Evaluate

Implement the chosen approach and evaluate the filtered call set using the assessment metrics described above. Compare the results to known variants if available. If the results are unsatisfactory, adjust the approach or switch to the alternative.

Step 5: Document and Archive

Record all filtering parameters, assessment metrics, and software versions. Archive the filtered call sets and the documentation for future reference.

A Practical Decision Framework for Filtering Strategy Selection

Choosing between hard filters and VQSR requires a structured evaluation that goes beyond simple cohort size considerations. The following framework provides a step-by-step process for matching your filtering strategy to your specific data characteristics, computational environment, and downstream analysis goals. This framework is designed to be applied before you commit to either approach, saving time and computational resources by identifying potential problems early.

Step 1: Inventory Your Data Characteristics

Begin by documenting the fundamental properties of your dataset. Record the number of samples, the sequencing platform and instrument model, the mean depth of coverage, and the total number of raw variant calls. For whole-exome data, note the capture kit version and target region size, as these directly influence variant density. For whole-genome data, record the read length and insert size distribution, since these affect mapping quality and variant calling accuracy.

Count the number of variants in your raw call set before any filtering. This count is the single most important determinant of whether VQSR is viable. The GermVarX workflow, which supports joint variant calling across cohorts, provides a practical example of how to generate and assess multi-sample VCF files before committing to a filtering strategy. Joint calling produces a single high-confidence multi-sample VCF that can be evaluated for variant density and quality metric distributions prior to filtering decisions.

Step 2: Evaluate VQSR Eligibility Criteria

Apply the following eligibility checklist to determine whether VQSR is appropriate for your data. Each criterion must be satisfied for VQSR to produce reliable results.

The first criterion is variant count. Your raw call set must contain enough variants for the Gaussian mixture model to converge. While the exact minimum depends on the variant caller and data type, whole-genome data from a cohort of at least 30 samples typically provides sufficient variant density. Whole-exome data requires substantially larger cohorts because the captured regions produce far fewer variants per sample.

The second criterion is truth set availability. You must have access to high-confidence variant calls from reference materials that match your sequencing platform and reference genome build. The Genome in a Bottle Consortium provides such truth sets for human samples, but these are primarily developed for specific sequencing technologies and may not transfer across platforms.

The third criterion is data homogeneity. VQSR models are sensitive to batch effects and platform differences. If your cohort was sequenced across multiple instruments, flow cells, or library preparation protocols, the model may learn platform-specific error patterns instead of generalizable quality distinctions. The best practices for germline variant analysis emphasize meticulous data handling and inspection of problematic variants, particularly in clinical settings where accuracy directly affects patient care.

The fourth criterion is organism and variant type compatibility. VQSR truth sets are primarily human-focused, and the model assumptions may not hold for other organisms or for variant types that are underrepresented in the training data.

Step 3: Assess Computational and Documentation Constraints

Evaluate your computational environment and documentation requirements before selecting an approach. VQSR model training is computationally intensive and requires substantial memory and processing time, particularly for large cohorts. If your computing environment has limited resources or if you need results quickly, hard filtering may be more practical.

Documentation requirements also influence the choice. Hard filtering provides explicit, auditable criteria that can be reviewed by laboratory directors, collaborators, and external auditors. VQSR requires documentation of the model training process, truth sets, sensitivity threshold, and validation results. In clinical settings where the rationale for variant inclusion must be explained, hard filtering is often preferred because each filtering decision can be traced to a specific threshold.

Step 4: Run a Pilot Comparison

If your data satisfies the VQSR eligibility criteria and you have the computational resources, run a pilot comparison before committing to a full analysis. Apply both hard filtering and VQSR to a subset of your data, such as a single chromosome or a random sample of 10,000 variants. Compare the filtered call sets using the assessment metrics described in the observations and measurements section.

The pilot comparison should evaluate the transition-transversion ratio, the number of variants retained, and the concordance with known variants if available. A well-calibrated VQSR model should retain a similar number of variants as well-chosen hard filters while potentially improving the balance of sensitivity and specificity. If VQSR produces a substantially different call set than hard filtering, investigate the cause before proceeding.

Step 5: Document the Decision Rationale

Record the reasoning behind your filtering strategy selection, including the data characteristics, eligibility criteria, pilot comparison results, and any limitations identified. This documentation supports reproducibility and provides a basis for revisiting the decision if your data or analysis requirements change.

The documentation should include the variant counts before and after filtering, the specific thresholds or model parameters used, the software versions, and the assessment metrics. The nf-core community standards provide guidance for pipeline documentation and configuration that can be adapted to variant filtering workflows. The Common Workflow Language pipelines developed for germline variant calling demonstrate how containerization and explicit workflow definitions support reproducibility and cross-platform compatibility.

A Record System for Filtering Decisions

Maintaining a structured record of filtering decisions across cohorts and projects enables continuous improvement and facilitates troubleshooting when downstream analyses reveal problems. The following record system provides a template for tracking the key parameters and outcomes of each filtering run.

Cohort and Data Metadata

Record the cohort identifier, the number of samples, the sequencing platform, the mean depth, the capture kit or sequencing strategy, and the reference genome build. This metadata provides context for interpreting filtering results and for comparing across cohorts.

Raw Call Set Summary

Record the total number of variants in the raw call set, the number of SNPs and indels separately, and the transition-transversion ratio. This baseline summary allows you to detect unexpected changes in variant composition after filtering.

Filtering Parameters

For hard filtering, record each threshold and the annotation it applies to. For VQSR, record the truth sets used, the sensitivity threshold, the model training parameters, and any warnings or errors generated during training.

Filtered Call Set Summary

Record the number of variants retained after filtering, the number removed by each filter or by the VQSR threshold, and the transition-transversion ratio of the filtered set. Compare these values to the raw call set summary to quantify the filtering effect.

Assessment Metrics

Record the concordance with known variants if available, the variant counts by functional category, and any Hardy-Weinberg equilibrium test results. These metrics provide evidence of filtering quality and support downstream interpretation.

Troubleshooting Notes

Record any anomalies observed during filtering, such as unexpected variant loss, model non-convergence, or batch effects. Include the corrective actions taken and the outcome. This troubleshooting log becomes a valuable reference for future filtering decisions.

Common Failure Patterns in Filtering Strategy Selection

Recognizing the common failure patterns in filtering strategy selection helps avoid costly mistakes and guides corrective action.

VQSR Applied to Insufficient Data

The most common failure pattern is applying VQSR to a cohort with too few variants. The model fails to converge or produces a quality score distribution that does not separate true variants from artifacts. Signs include warnings from the VQSR tool, an unusual number of variants at extreme quality scores, and poor concordance with known variants. The corrective action is to switch to hard filtering or to increase the cohort size.

Hard Filters Applied Without Data-Specific Tuning

Applying default hard filtering thresholds without adjusting for your data characteristics can produce poor results. Thresholds tuned for high-depth whole-genome data may be too stringent for low-depth exome data, removing true variants along with artifacts. Signs include a low variant count, a low transition-transversion ratio, and poor concordance with known variants. The corrective action is to adjust thresholds based on the depth and characteristics of your data.

Truth Set Mismatch in VQSR

Using truth sets that do not match your sequencing platform, reference build, or population can produce biased VQSR results. The model learns incorrect patterns and misclassifies variants. Signs include systematic differences between the model's predictions and known variants, and unexpected variant composition in the filtered set. The corrective action is to obtain appropriate truth sets or switch to hard filtering.

Ignoring Batch Effects

Applying a single filtering strategy to data with significant batch effects can produce inconsistent results across batches. VQSR models trained on one batch may not generalize to another, and hard filtering thresholds may not account for batch-specific error profiles. Signs include systematic differences in variant quality between batches and unexpected genotype frequency patterns. The corrective action is to evaluate batch effects before filtering and consider batch-specific filtering or separate models.

Escalation Criteria for Filtering Problems

Certain filtering problems require escalation to a bioinformatics specialist, laboratory director, or other qualified professional. The following criteria indicate when to seek additional expertise.

Model Non-Convergence or Instability

If VQSR fails to converge after multiple attempts with different parameters, or if the model produces unstable results across repeated runs, escalate the issue. This may indicate fundamental problems with the data or truth sets that require expert investigation.

Unexplained Variant Loss or Gain

If filtering removes an unexpectedly large proportion of variants, or if the filtered set contains an excess of apparent artifacts, escalate the issue. This may indicate problems with the variant caller, the reference genome, or the filtering parameters that require expert diagnosis.

Discordance with Orthogonal Validation

If the filtered call set shows poor concordance with known variants from orthogonal methods, such as array genotyping or Sanger sequencing, escalate the issue. This indicates that the filtering approach is not achieving acceptable accuracy and may require a different strategy.

Clinical Reporting Implications

If the filtering problem affects variants that are being considered for clinical reporting, escalate the issue immediately. Do not proceed with clinical interpretation until the filtering problem is resolved and the pipeline is validated.

Integration with Reproducible Workflow Systems

Both hard filtering and VQSR can be integrated into automated analysis pipelines using workflow management systems. The choice of workflow system affects how filtering is implemented, documented, and reproduced.

Nextflow and nf-core Standards

Nextflow workflows, such as the GermVarX pipeline, support joint germline variant calling and filtering for cohort studies. These workflows integrate GATK HaplotypeCaller and DeepVariant with joint genotyping through GATK or GLnexus, and support consensus generation between callers. The nf-core community standards provide guidance for pipeline configuration, documentation, and reproducibility that can be adapted to filtering workflows.

Common Workflow Language Pipelines

Common Workflow Language pipelines provide a standardized format for defining analysis workflows. These pipelines have been shown to achieve high accuracy in reproducing published results and detecting germline SNP and small INDEL variants when validated against Genome in a Bottle reference materials. Containerization combined with explicit workflow definitions overcomes software incompatibility and configuration issues.

Training for Filtering Implementation

Implementing and troubleshooting variant filtering requires foundational bioinformatics skills. The Galaxy Training Network provides accessible tutorials for variant calling workflows and quality assessment. The European Bioinformatics Institute offers training pathways for data-resource usage and practical analysis education. The Carpentries provides foundational computing and data skills that support reproducible analysis practices.

Limitations of the Decision Framework

This decision framework provides practical guidance but has limitations that should be acknowledged. The framework assumes that your variant caller produces accurate quality annotations, which may not hold for all callers or data types. The framework does not address somatic variant calling, which has different quality considerations than germline calling. The framework also does not account for the specific requirements of every downstream analysis, and you may need to adjust the approach based on your analysis goals.

The framework is based on current best practices for short-read sequencing data. Third-generation sequencing technologies have different error profiles and quality considerations, and the filtering approaches described here may not transfer directly. Best practices for these technologies are still evolving, and careful validation is required.

Frequently Asked Questions

What is the minimum cohort size for VQSR to work reliably?

VQSR requires enough variants to train a stable model. The exact minimum depends on the variant caller, the data type, and the diversity of the cohort. Whole-genome data produces more variants than exome data, so WGS cohorts can be smaller than WES cohorts. In practice, cohorts with fewer than 30 samples often produce too few variants for stable VQSR, and hard filtering is a safer choice. The GermVarX workflow supports joint variant calling across cohorts and can be used to assess variant counts before deciding on a filtering approach.

Can VQSR be used for non-human organisms?

VQSR relies on truth sets of known true and false variants. These truth sets are primarily developed for the human genome, and appropriate truth sets for non-human organisms may not exist. Without truth sets, VQSR cannot be trained. For non-human organisms, hard filtering with thresholds adjusted to the species is the practical approach.

How do I choose hard filtering thresholds for my data?

Threshold selection should be based on the variant caller's documentation, the sequencing depth of your data, and the downstream application. Start with the thresholds recommended by the variant caller and adjust based on the assessment metrics. The transition-transversion ratio, variant count, and concordance with known variants provide feedback on whether the thresholds are appropriate.

What is the difference between hard filtering and VQSR in terms of reproducibility?

Hard filtering is fully reproducible because the thresholds are explicit and applied uniformly. VQSR is reproducible if the model training process is documented and the same truth sets and parameters are used. However, VQSR models are sensitive to the training data, and models trained on different cohorts may produce different results.

Does VQSR work for exome sequencing data?

VQSR can work for exome data if the cohort is large enough to produce sufficient variants. Exome data produces fewer variants than whole-genome data, so larger cohorts are needed. The capture process also introduces biases that may affect the model. For many exome studies, hard filtering is more practical and reliable.

How do I know if my filtering approach is removing too many true variants?

Compare the filtered call set to known variants from orthogonal methods, such as array genotyping or Sanger sequencing. A low concordance rate indicates that true variants are being removed. The transition-transversion ratio can also indicate over-filtering if it drops below the expected range for your organism and data type.

What should I do if VQSR fails to converge?

VQSR non-convergence typically results from too few variants or from truth sets that do not match the data. Check the variant count in your raw call set and verify that the truth sets are appropriate for your sequencing platform and reference build. If the problem persists, switch to hard filtering.

Is hard filtering or VQSR better for clinical variant reporting?

Hard filtering is often preferred for clinical reporting because the filtering criteria are explicit and auditable. VQSR can be used in clinical settings if the model training and validation are thoroughly documented, but the reduced transparency of machine learning approaches requires additional validation effort. Clinical laboratories must validate their entire pipeline, including the filtering approach, against known reference materials.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.