VQSR vs. Hard Filters: Which Variant Filtering Approach Should You Use for Your Germline Data?

By Dr. Zubair Khalid, DVM, MS, PhD ·

VQSR vs. Hard Filters: Which Variant Filtering Approach Should You Use for Your Germline Data?

Key Takeaways

  • VQSR is a machine learning approach that models error profiles using known variant sites, adapting to dataset characteristics for improved precision, but requires a minimum of ~30 samples for reliable model training. It excels with large Whole Genome Sequencing (WGS) cohorts where variant density is high, but is less suitable for small cohorts or targeted sequencing due to insufficient training data.
  • Hard filters apply fixed thresholds to quality metrics (e.g., QD, FS, MQ, SOR) and are simpler, reproducible, and require no training data, making them suitable for small cohorts, single samples, or clinical reporting where auditability is paramount. Their primary limitation is their inability to adapt to dataset-specific error profiles or capture complex interactions between metrics.
  • The choice between VQSR and hard filters is critically dependent on cohort size and data type (WES vs. WGS), with VQSR generally recommended for WGS cohorts of 30+ samples and hard filters for smaller cohorts or Whole Exome Sequencing (WES) due to lower variant density. Joint calling in large WES cohorts can increase variant density, making VQSR a potential option.
  • Reproducibility is achieved by meticulously documenting and version-controlling filtering parameters, including specific thresholds for hard filters and reference database versions/sensitivity thresholds for VQSR. This ensures that analyses can be precisely replicated, which is crucial for both research and clinical applications.
  • Quality assessment metrics such as the transition-to-transversion (Ti/Tv) ratio (expected ~2.0 for WGS, ~2.8-3.0 for WES) and concordance with known variant databases are essential for evaluating filtering efficacy. Deviations from expected ratios can indicate over- or under-filtering, necessitating adjustments to the chosen approach.
  • Upstream errors in alignment, base quality recalibration, or variant calling cannot be corrected by filtering; therefore, robust quality control at every stage of the bioinformatics pipeline is fundamental. Systematic artifacts or sample contamination can confound filtering results, highlighting the importance of cohort-level quality control.

Variant filtering separates true biological variants from sequencing and alignment artifacts in a germline variant calling workflow. The two dominant approaches are Variant Quality Score Recalibration (VQSR), a machine learning method that uses known variant sites to model error patterns, and hard filters, which apply fixed thresholds to metrics like quality score, depth, and allele balance. For most researchers, the decision depends on cohort size, available computational resources, and whether you are working with whole-exome sequencing (WES) or whole-genome sequencing (WGS) data. This article provides a decision framework grounded in published workflows and official bioinformatics resources, with practical guidance for implementation, quality control, and troubleshooting.

Understanding the Two Filtering Approaches

Variant filtering operates on the raw variant call set produced by callers such as GATK HaplotypeCaller or DeepVariant. The goal is to reduce false positives while retaining true positives, and the choice of filtering strategy directly affects downstream analyses including association studies, population genetics, and clinical interpretation.

Hard Filters: Fixed Thresholds and Direct Rules

Hard filters apply predetermined thresholds to variant annotations. A typical hard filter set for germline data might include thresholds for quality by depth (QD), Fisher strand bias (FS), mapping quality (MQ), and strand odds ratio (SOR). These thresholds are applied uniformly to every variant in the call set, regardless of the genomic context or the overall distribution of variant qualities in your specific dataset.

The primary advantage of hard filters is simplicity and reproducibility. The same thresholds apply to every sample and every run, which makes the filtering step easy to document and audit. Hard filters also require no training data, no reference panel of known variants, and no additional computational overhead beyond the filtering step itself.

The primary limitation is that fixed thresholds do not adapt to the characteristics of your specific dataset. A threshold that works well for high-coverage WGS data may be too permissive or too stringent for low-coverage WES data. Hard filters also treat each variant independently, so they do not capture the joint distribution of multiple quality metrics that can distinguish true variants from artifacts.

VQSR: Machine Learning with Known Variant Sites

VQSR uses a machine learning approach to model the distribution of variant quality metrics. The method requires a set of known variant sites, typically derived from databases such as dbSNP, HapMap, and the 1000 Genomes Project, which are available through official NCBI Data Resources. The algorithm trains a Gaussian mixture model on these known sites, then applies the model to score every variant in your call set. Variants that fall below a chosen sensitivity threshold are filtered out.

The key advantage of VQSR is that it adapts to the specific error profile of your dataset. It can capture complex interactions between quality metrics that hard filters miss, and it can be tuned to different sensitivity levels depending on your downstream analysis goals. For example, a study prioritizing recall might choose a more permissive sensitivity threshold, while a study prioritizing precision might choose a more stringent one.

The primary limitation is the requirement for a sufficient number of known variant sites in your data. VQSR performs poorly on small cohorts or targeted sequencing datasets where the number of variants available for training is too low to build a reliable model. The GATK documentation and community workflows generally recommend a minimum of 30 samples for VQSR, though the exact number depends on the data type and the density of known variants in the targeted regions.

At a Glance: Decision Table for Filtering Approach

ScenarioRecommended ApproachRationale
WGS cohort with 30 or more samplesVQSRSufficient variant density for model training, adaptive error modeling improves precision
WES cohort with fewer than 30 samplesHard filtersInsufficient known variant sites for reliable VQSR training, fixed thresholds are more predictable
Single sample or small family studyHard filtersVQSR cannot be trained reliably, hard filters with manual review are more appropriate
Clinical or diagnostic reportingHard filters with manual reviewFixed thresholds are easier to document and audit, VQSR model parameters are harder to explain in a clinical report
Large WES cohort with joint callingVQSR or hard filters with cohort-level QCJoint calling increases variant density, GermVarX workflows demonstrate consensus approaches with both callers
Reproducible pipeline deploymentHard filtersFixed thresholds are deterministic and easier to version-control across pipeline releases

Core Principles of Germline Variant Filtering

The Variant Calling Workflow Context

Variant filtering does not operate in isolation. It sits within a larger workflow that includes read alignment, variant calling, joint genotyping, and functional annotation. The choice of filtering approach should be made in the context of the entire pipeline, because upstream decisions affect the distribution of variant quality metrics.

For example, the choice of variant caller influences the quality metric distributions. GATK HaplotypeCaller and DeepVariant produce different error profiles, and a filtering strategy calibrated for one caller may not transfer directly to the other. The GermVarX workflow, which integrates both callers with joint genotyping via GATK or GLnexus, demonstrates that consensus generation between callers can increase reliability, but it also means that filtering decisions must account for the strengths and weaknesses of each caller.

Cohort Size and Variant Density

The most important factor in choosing between VQSR and hard filters is the number of variants available for model training. VQSR requires a large number of known variant sites to build a reliable model. In WGS data, where variants are densely distributed across the genome, a cohort of 30 samples typically provides enough training sites. In WES data, where only the exonic regions are sequenced, the variant density is much lower, and even a larger cohort may not provide enough known sites for reliable VQSR training.

The Faroese whole genome sequencing study, which analyzed 40 participants, demonstrates the kind of cohort where VQSR can be applied. The study identified numerous putatively functional private alleles, including stop gain variants and high impact missense variants, which required careful filtering to distinguish true variants from artifacts. The density of variants in WGS data across 40 samples provided sufficient training material for a reliable filtering approach.

Data Type: WES versus WGS

WES and WGS present different challenges for variant filtering. WES data has higher coverage in targeted regions but covers only a small fraction of the genome. This means that the total number of variants is lower, and the distribution of quality metrics may be influenced by the capture kit and the specific enrichment protocol used.

WGS data has more uniform coverage across the genome and a higher total number of variants, which makes VQSR more feasible. However, WGS data also includes variants in repetitive and low-complexity regions that are more prone to alignment artifacts, and these regions may require additional filtering beyond what VQSR or hard filters provide.

The whole exome sequencing study of a Saudi family with familial hypercholesterolemia and sitosterolemia illustrates the WES context. The study used a platform with more than 98 percent of targeted bases covered at 20x or higher, and the bioinformatics analysis used a standardized pipeline. In this kind of targeted WES setting, hard filters with manual review are often more practical than VQSR, because the variant density is too low for reliable model training.

Practical Workflow for Filtering Germline Variants

Step 1: Assess Your Cohort and Data Type

Before choosing a filtering approach, document the following characteristics of your dataset:

  • Number of samples in the cohort
  • Sequencing platform and coverage depth
  • WES or WGS data type
  • Variant caller used for initial calling
  • Whether joint genotyping was performed
  • Availability of known variant databases for training

This assessment determines whether VQSR is feasible and whether hard filters are more appropriate. For cohorts with fewer than 30 samples, or for WES data with low variant density, hard filters are the safer choice.

Step 2: Generate the Raw Variant Call Set

The filtering step operates on the raw variant call set produced by the variant caller. For GATK HaplotypeCaller, this means generating per-sample GVCF files and then performing joint genotyping to produce a multi-sample VCF. For DeepVariant, the output is a VCF that can be used directly or combined with other callers through tools like GLnexus.

The GermVarX workflow demonstrates a reproducible approach to this step. It uses Nextflow DSL2 for pipeline orchestration, which ensures portability across workstations, HPC clusters, and cloud platforms. The workflow supports both GATK HaplotypeCaller and DeepVariant, with joint genotyping performed via GATK or GLnexus, and it generates a single high-confidence multi-sample VCF optimized for downstream analyses.

Step 3: Apply the Chosen Filtering Approach

For hard filters, apply fixed thresholds to the variant annotations in the VCF. Common annotations used for germline filtering include:

  • Quality by depth (QD): the variant quality score divided by the depth of coverage at the variant site
  • Fisher strand bias (FS): a measure of strand bias in the supporting reads
  • Mapping quality (MQ): the root mean square of the mapping quality of the reads supporting the variant
  • Strand odds ratio (SOR): an alternative measure of strand bias
  • Depth of coverage (DP): the total depth at the variant site
  • Genotype quality (GQ): the confidence in the genotype call

For VQSR, use the GATK VariantRecalibrator tool with the appropriate known variant databases. The tool builds a model using the known sites, then applies the model to score all variants. The output includes a VQSLOD score for each variant, and you choose a sensitivity threshold to determine which variants pass the filter.

Step 4: Evaluate the Filtering Results

After filtering, evaluate the results using both quantitative and qualitative measures. Quantitative measures include the transition-to-transversion ratio (Ti/Tv), the number of variants retained, and the concordance with known variant databases. Qualitative measures include visual inspection of variants in the Integrative Genomics Viewer or similar tools.

The Ti/Tv ratio is a useful sanity check for germline data. For WGS data, a Ti/Tv ratio around 2.0 is expected, while for WES data, a ratio around 2.8 to 3.0 is typical. A significantly lower Ti/Tv ratio suggests that many false positives remain in the call set, while a significantly higher ratio may indicate that true variants are being filtered out.

Step 5: Document and Version-Control the Filtering Parameters

Reproducibility requires that the filtering parameters be documented and version-controlled. This is particularly important for hard filters, where the specific thresholds used should be recorded in the pipeline configuration. For VQSR, the version of the known variant databases and the sensitivity threshold should be recorded.

The nf-core documentation provides guidance on reproducible pipeline standards, including configuration management and version control. Following these standards ensures that the filtering step can be reproduced exactly, which is essential for both research and clinical applications.

Options and Tradeoffs in Filtering Strategies

Consensus Filtering with Multiple Callers

An emerging approach to variant filtering is consensus generation between multiple callers. The GermVarX workflow supports this approach by integrating GATK HaplotypeCaller and DeepVariant, with consensus generation between callers to increase reliability. The idea is that variants called by both callers are more likely to be true positives, while variants called by only one caller require additional scrutiny.

The tradeoff is that consensus filtering can reduce sensitivity. A variant that is a true positive but is missed by one caller due to its specific error profile will be filtered out by the consensus approach. This is a particular concern for variants in difficult genomic regions, such as homopolymers or GC-rich regions, where different callers may have different strengths and weaknesses.

The transcript SNV classifier study, which integrated short-read and long-read RNA-seq data, demonstrates the value of complementary sequencing approaches for variant detection. The study found that the integrated approach increased the total detected transcript SNVs by 31.83 percent on average, with genomic and RNA editing variants exceeding conventional methods by more than one-fold and more than two-fold, respectively. While this study focused on transcript SNVs instead of germline DNA variants, it illustrates the principle that integrating multiple data sources can improve variant detection.

Cohort-Level Quality Control

Joint variant calling enables cohort-level quality control that is not possible with per-sample calling. When multiple samples are genotyped together, you can assess the distribution of quality metrics across the entire cohort and identify samples that are outliers. This is particularly useful for detecting sample contamination, mislabeled samples, or systematic sequencing artifacts.

The GermVarX workflow includes sample- and cohort-level quality control as part of its standard pipeline. This includes generating MultiQC reports that aggregate quality metrics across all samples, making it easier to identify problematic samples before they affect downstream analyses.

The tradeoff is that joint calling requires more computational resources and more complex pipeline orchestration. The GermVarX workflow addresses this by using Nextflow DSL2, which enables efficient parallelization across diverse computing environments, including workstations, HPC clusters, and cloud platforms.

Functional Annotation and Filtering

Functional annotation can be used as a post-filtering step to prioritize variants for downstream analysis. Tools like the Variant Effect Predictor (VEP), which is integrated into the GermVarX workflow, annotate variants with their predicted functional effects, including missense, stop gain, and splice site variants.

The Faroese genome study identified numerous putatively functional private alleles, including stop gain variants and high impact missense variants. These variants are of particular interest for population genetics and disease association studies, and functional annotation helps prioritize them for further analysis.

The tradeoff is that functional annotation is not a substitute for quality filtering. A variant with a predicted high impact effect is still unreliable if it is a sequencing artifact, and filtering should always be performed before functional annotation.

Observations and Measurements for Filtering Assessment

Transition-to-Transversion Ratio

The transition-to-transversion ratio is a standard quality metric for germline variant call sets. Transitions (A to G, C to T) are more common than transversions (A to T, A to C, G to T, G to C) in the human genome, and the expected ratio depends on the genomic context.

For WGS data, a Ti/Tv ratio around 2.0 is expected. For WES data, the ratio is higher, typically around 2.8 to 3.0, because exonic regions are enriched for transitions. A significantly lower ratio suggests that false positives remain in the call set, while a significantly higher ratio may indicate over-filtering.

Variant Density and Known Variant Concordance

The density of variants in your call set and the concordance with known variant databases provide additional quality checks. For WGS data, a typical human genome has around 3 to 4 million variants, with about 85 to 90 percent of these found in dbSNP. For WES data, the number of variants is much lower, typically around 20,000 to 30,000.

Concordance with known variant databases is a useful check, but it should be interpreted with caution. Novel variants that are not in known databases can be true positives, particularly in populations that are underrepresented in reference databases. The Faroese genome study identified numerous private alleles that were not found in other European populations, highlighting the importance of considering population-specific variation when assessing concordance.

Genotype Concordance with Array Data

If you have genotype array data for the same samples, you can assess the concordance between the sequencing-based genotypes and the array-based genotypes. High concordance, typically above 99 percent, indicates that the variant calling and filtering pipeline is performing well. Low concordance at specific sites may indicate genotyping errors or filtering issues.

Sample-Level Quality Metrics

Sample-level quality metrics, such as the number of variants called per sample, the Ti/Tv ratio per sample, and the heterozygous-to-homozygous ratio, can identify problematic samples. Samples that are outliers on these metrics may have contamination, low coverage, or other issues that affect variant calling quality.

Records and Documentation for Filtering Decisions

Pipeline Configuration Files

The filtering parameters should be recorded in the pipeline configuration files. For hard filters, this includes the specific thresholds for each annotation. For VQSR, this includes the version of the known variant databases, the sensitivity threshold, and the specific annotations used for model training.

The nf-core documentation provides guidance on pipeline configuration and version control. Following these standards ensures that the filtering step can be reproduced exactly, which is essential for both research and clinical applications.

Filtering Reports

Generate a filtering report that documents the number of variants before and after filtering, the number of variants removed by each filter, and the quality metrics of the retained variants. This report should be stored with the pipeline outputs and made available to collaborators or reviewers.

Version Control for Reference Data

The reference genome and known variant databases used for filtering should be version-controlled. Changes to these resources can affect the filtering results, and documenting the versions used ensures that the analysis can be reproduced.

Audit Trail for Clinical Applications

For clinical applications, the filtering decisions should be documented in an audit trail that includes the specific parameters used, the rationale for the filtering approach, and the results of any manual review. This audit trail is essential for regulatory compliance and for explaining the analysis to clinicians and patients.

Common Failure Patterns in Variant Filtering

Over-Filtering with Stringent Hard Filters

A common failure pattern is applying hard filters that are too stringent, resulting in the loss of true variants. This is particularly problematic for rare variants, which may have lower quality scores due to lower allele frequencies. The Faroese genome study identified numerous private alleles, including stop gain variants and high impact missense variants, which are exactly the kind of variants that can be lost with overly stringent filtering.

To avoid over-filtering, evaluate the filtering results using the Ti/Tv ratio and known variant concordance. If the Ti/Tv ratio is higher than expected, or if the concordance with known variants is lower than expected, the filters may be too stringent.

Under-Filtering with Permissive Hard Filters

The opposite failure pattern is applying hard filters that are too permissive, resulting in a high number of false positives. This is particularly problematic for downstream analyses that are sensitive to false positives, such as association studies and clinical variant interpretation.

To avoid under-filtering, evaluate the filtering results using the Ti/Tv ratio and the number of variants retained. If the Ti/Tv ratio is lower than expected, or if the number of variants is higher than expected, the filters may be too permissive.

VQSR Failure on Small Cohorts

VQSR can fail when the cohort is too small to provide sufficient training sites. The model may be poorly calibrated, resulting in either over-filtering or under-filtering. This is a particular risk for WES data, where the variant density is lower than in WGS data.

To avoid this failure, assess the number of known variant sites in your dataset before running VQSR. If the number is too low, use hard filters instead.

Batch Effects and Systematic Artifacts

Systematic artifacts that affect multiple samples in a cohort can confound both VQSR and hard filters. For example, a sequencing batch with elevated error rates may produce variants that pass quality filters but are still artifacts. Cohort-level quality control, as implemented in the GermVarX workflow, can help identify these systematic issues.

Sample Contamination and Mix-Ups

Sample contamination and mix-ups can produce spurious variants that pass quality filters. Contamination from another sample can produce variants with intermediate allele frequencies that are difficult to distinguish from true heterozygous variants. Sample mix-ups can produce genotype discordance that is only detectable when comparing multiple samples or data types.

Limitations of Both Filtering Approaches

Hard Filters Do Not Capture Joint Metric Distributions

Hard filters treat each quality metric independently, so they do not capture the joint distribution of metrics that can distinguish true variants from artifacts. A variant with a borderline quality score but excellent strand bias and mapping quality may be a true positive, while a variant with the same quality score but poor strand bias may be an artifact. Hard filters cannot make this distinction.

VQSR Requires Large Training Sets

VQSR requires a large number of known variant sites for reliable model training. This limits its applicability to small cohorts and targeted sequencing datasets. The exact minimum number of training sites depends on the data type and the density of known variants, but a general guideline is that VQSR is not recommended for cohorts with fewer than 30 samples.

Both Approaches Are Sensitive to Reference Database Quality

Both VQSR and hard filters depend on the quality of reference databases. VQSR uses known variant databases for training, and hard filters are often calibrated using known variant sets. If the reference databases contain errors or are incomplete for the population being studied, the filtering results will be affected.

The Faroese genome study highlights this limitation. The study identified numerous private alleles that were not found in other European populations, suggesting that reference databases may be incomplete for founder populations. This incompleteness can affect both VQSR training and the assessment of filtering quality.

Filtering Cannot Correct for Upstream Errors

Variant filtering cannot correct for errors introduced upstream in the pipeline. Poor read alignment, incorrect base quality scores, or systematic sequencing artifacts will produce variants that pass quality filters regardless of the filtering approach. This is why quality control at every step of the pipeline is essential.

Safety and Regulatory Context for Variant Filtering

Clinical Reporting Requirements

For clinical applications, variant filtering decisions must be documented and reproducible. The audit trail should include the specific filtering parameters, the rationale for the filtering approach, and the results of any manual review. This documentation is essential for regulatory compliance and for explaining the analysis to clinicians and patients.

The whole exome sequencing study of the Saudi family with familial hypercholesterolemia and sitosterolemia illustrates the clinical context. The study used a standardized pipeline with Sanger sequencing validation of identified variants, and variant pathogenicity was evaluated using multiple in silico tools. This kind of rigorous validation is essential for clinical reporting.

Data Privacy and Security

Germline variant data is sensitive personal information, and its handling is subject to data privacy regulations. Researchers should ensure that their variant filtering workflows comply with applicable data protection requirements, including secure storage, access controls, and data sharing agreements.

Reproducibility Standards

Reproducibility is a core requirement for both research and clinical applications. The nf-core documentation provides standards for reproducible pipeline development, including version control, containerization, and configuration management. Following these standards ensures that the filtering step can be reproduced exactly.

Professional Escalation Criteria

When to Consult a Bioinformatics Specialist

If you encounter any of the following situations, consult a bioinformatics specialist:

  • VQSR fails to converge or produces unexpected results
  • The Ti/Tv ratio is significantly outside the expected range for your data type
  • The concordance with known variant databases is much lower than expected
  • You observe systematic artifacts that affect multiple samples
  • You are uncertain whether your cohort size is sufficient for VQSR

When to Consider Alternative Filtering Approaches

If both VQSR and hard filters produce unsatisfactory results, consider alternative approaches:

  • Consensus filtering with multiple callers, as implemented in the GermVarX workflow
  • Machine learning classifiers that integrate multiple data types, as demonstrated in the transcript SNV classifier study
  • Manual review of variants in genomic regions that are prone to artifacts

When to Revisit the Entire Pipeline

If filtering problems persist despite adjustments to the filtering parameters, the issue may be upstream in the pipeline. Consider revisiting the read alignment, base quality score recalibration, or variant calling steps. The Galaxy Training Network provides accessible workflow training that can help identify and address upstream issues.

Building a Filtering Decision Log and Audit Trail for Your Cohort

A filtering decision log is a structured record that captures why you chose VQSR or hard filters, which parameters you applied, and how the filtering performed against measurable quality targets. This log serves as the bridge between the filtering step and the downstream analyses that depend on it, including association studies, population genetics, and clinical interpretation. Without a decision log, you cannot reliably compare results across pipeline versions, justify filtering choices to collaborators or reviewers, or diagnose why a downstream analysis produced unexpected findings.

What to Record Before Filtering

The decision log begins before you run any filtering tool. Record the following items in a plain text file or spreadsheet that is stored with your pipeline outputs:

  • Cohort size and composition, including the number of samples and any known population structure
  • Sequencing platform, coverage depth, and whether the data is WES or WGS
  • Variant caller version and the specific command used for initial calling
  • Whether joint genotyping was performed and which tool was used
  • Reference genome build and version
  • Known variant database versions, including dbSNP build and any population-specific resources
  • The filtering approach selected and the rationale for that choice

This pre-filtering record is essential because the choice between VQSR and hard filters depends on cohort characteristics that may change as you add samples or update your pipeline. The nf-core documentation provides guidance on pipeline configuration and version control that supports this kind of record keeping, and the Galaxy Training Network offers accessible tutorials on reproducible analysis practices that include documentation standards.

Recording Filtering Parameters and Thresholds

For hard filters, record every threshold applied to every annotation. A typical record entry might look like this:

  • QD less than 2.0: filter applied
  • FS greater than 60: filter applied
  • MQ less than 40: filter applied
  • SOR greater than 3.0: filter applied
  • DP less than 10: filter applied
  • GQ less than 20: filter applied

For VQSR, record the version of the VariantRecalibrator tool, the known variant databases used for training, the annotations included in the model, and the sensitivity threshold selected. Also record the number of variants used for training and the number of known sites that were available, because this information explains why VQSR may have performed well or poorly.

The Bioconductor project provides official documentation for R packages that can help you generate and visualize these records. Many Bioconductor packages include functions for reading VCF files and extracting the annotation values that you need to document, and the project's reproducible research guidance supports the creation of analysis reports that combine code, results, and documentation.

Tracking Variant Counts Through the Filtering Process

The most basic record is the number of variants at each stage of the filtering process. Record the following counts:

  • Total variants in the raw call set
  • Variants passing the initial quality filters
  • Variants removed by each individual filter
  • Variants retained after all filters
  • Variants passing VQSR at the selected sensitivity threshold
  • Variants failing VQSR and the distribution of their VQSLOD scores

These counts provide a quantitative basis for evaluating whether the filtering is too stringent or too permissive. For example, if a single hard filter removes an unexpectedly large fraction of variants, that filter may be poorly calibrated for your dataset. The GermVarX workflow demonstrates how automated pipelines can generate these counts as part of the standard output, and its integration with MultiQC provides aggregated reports that make it easier to track variant counts across multiple samples and pipeline runs.

Recording Quality Metrics Before and After Filtering

In addition to variant counts, record the key quality metrics that you use to evaluate filtering performance. These include:

  • Transition-to-transversion ratio before and after filtering
  • Number of variants per sample before and after filtering
  • Heterozygous-to-homozygous ratio per sample
  • Concordance with known variant databases
  • Genotype concordance with array data if available

The transition-to-transversion ratio is a particularly useful metric to track over time. For WGS data, a ratio around 2.0 is expected, while for WES data, a ratio around 2.8 to 3.0 is typical. If the ratio changes significantly after a pipeline update, the filtering parameters may need adjustment. The Faroese whole genome sequencing study provides an example of how population-specific variation can affect these metrics, since the study identified numerous private alleles that were not found in other European populations.

Using the Decision Log to Diagnose Filtering Problems

The decision log becomes most valuable when filtering produces unexpected results. If the transition-to-transversion ratio is lower than expected, the log tells you which filters were applied and allows you to test whether a specific filter is responsible. If VQSR fails to converge, the log tells you how many training sites were available and which databases were used, which helps you determine whether the failure is due to insufficient training data.

The whole exome sequencing study of a Saudi family with familial hypercholesterolemia and sitosterolemia illustrates the importance of documentation in clinical contexts. The study used a standardized pipeline with Sanger sequencing validation of identified variants, and the bioinformatics analysis was documented in sufficient detail to support the clinical interpretation. A decision log provides the same level of documentation for the filtering step.

Creating a Filtering Comparison Table

For cohorts where you are uncertain whether VQSR or hard filters is the better approach, create a comparison table that records the performance of both methods on the same dataset. This table should include:

  • Number of variants retained by each method
  • Transition-to-transversion ratio for each method
  • Concordance with known variant databases for each method
  • Number of variants passing both methods
  • Number of variants passing only one method
  • Computational time and resources required for each method

This comparison provides empirical evidence for your filtering decision instead of relying on general guidelines. The GermVarX workflow supports this kind of comparison by integrating both GATK HaplotypeCaller and DeepVariant with joint genotyping via GATK or GLnexus, and its consensus generation between callers provides an additional point of comparison.

Recording Manual Review Decisions

For clinical applications or small cohorts where manual review is part of the filtering process, record every manual review decision. This includes:

  • The variant identifier and genomic position
  • The reason the variant was flagged for review
  • The evidence reviewed, including read alignments and quality metrics
  • The decision to retain or filter the variant
  • The name of the reviewer and the date of the review

This record is essential for clinical reporting and for explaining filtering decisions to clinicians and patients. The audit trail should be detailed enough that another reviewer can understand the rationale for each decision.

Version Control for the Decision Log

The decision log itself should be version-controlled along with the pipeline configuration files. When you update the pipeline, create a new version of the log that documents the changes and the rationale for those changes. This allows you to compare filtering performance across pipeline versions and to identify when a change in filtering parameters affected downstream results.

The nf-core documentation provides standards for pipeline version control that can be applied to the decision log. The Carpentries lessons offer foundational training in Git and version control that is useful for managing these records, and the EMBL-EBI training resources provide additional guidance on reproducible bioinformatics practices.

Integrating the Decision Log with Pipeline Outputs

Store the decision log with the pipeline outputs so that it is available when you or others need to interpret the results. A common approach is to include the log in the same directory as the filtered VCF files, with a naming convention that links the log to the specific pipeline run. The Galaxy Training Network provides examples of how to structure analysis outputs for reproducibility, and the nf-core documentation describes how to organize pipeline outputs in a consistent way.

Reviewing the Decision Log After Pipeline Updates

When you update any part of the variant calling pipeline, review the decision log to determine whether the filtering parameters need adjustment. Changes to the variant caller, reference genome, or known variant databases can all affect the distribution of quality metrics and the performance of the filtering approach. The decision log provides the baseline against which you can evaluate the impact of these changes.

Using the Decision Log for Collaboration and Publication

A well-maintained decision log supports collaboration by making it easy to share filtering decisions with other researchers. When you publish results, the decision log can be included as supplementary material or referenced in the methods section. This transparency is increasingly expected in genomic research, and it supports the reproducibility standards promoted by the nf-core documentation and the Bioconductor project.

Common Mistakes in Decision Log Maintenance

The most common mistake is failing to record the rationale for the filtering choice. A log that records only the parameters without explaining why those parameters were chosen is of limited value when you need to revisit the decision months later. Record the reasoning at the time of the decision, including any comparisons you performed between VQSR and hard filters.

Another common mistake is failing to update the log when the pipeline changes. If you update the variant caller or change the reference genome, the old log no longer reflects the current pipeline. Create a new version of the log for each pipeline version and clearly indicate which log corresponds to which pipeline run.

A third mistake is recording only the final filtering results without recording the intermediate steps. If you need to diagnose a filtering problem, the intermediate counts and quality metrics are essential. Record the variant counts and quality metrics at each stage of the filtering process, beyond the final result.

Professional Escalation Criteria for Decision Log Issues

If you discover that your decision log is incomplete or that you cannot reconstruct the filtering parameters used for a previous analysis, consult a bioinformatics specialist. This situation can arise when pipelines are run by different people over time or when the original documentation has been lost. A specialist can help you reconstruct the filtering parameters from the output files or recommend a re-analysis if the documentation cannot be recovered.

If you find that the filtering parameters in your decision log produce different results when re-run on the same data, this indicates a reproducibility problem that requires immediate attention. The issue may be in the pipeline code, the reference data, or the computing environment. The nf-core documentation provides guidance on containerization and environment management that can help resolve these issues, and the Carpentries lessons offer training in reproducible computing practices.

Frequently Asked Questions

What is the minimum cohort size for VQSR?

The minimum cohort size for VQSR depends on the data type and the density of known variants in the targeted regions. For WGS data, a cohort of 30 samples is generally sufficient because the variant density is high. For WES data, the variant density is much lower, and even a larger cohort may not provide enough known sites for reliable VQSR training. If you are working with WES data or a small cohort, hard filters are the safer choice.

Can I use VQSR for a single sample or small family study?

VQSR is not recommended for single samples or small family studies because the number of known variant sites is too low for reliable model training. Hard filters with manual review are more appropriate for these scenarios. The filtering parameters should be documented carefully, and the results should be evaluated using the Ti/Tv ratio and known variant concordance.

How do I choose the sensitivity threshold for VQSR?

The sensitivity threshold for VQSR depends on your downstream analysis goals. A more permissive threshold prioritizes recall, which is useful for discovery studies where you want to minimize false negatives. A more stringent threshold prioritizes precision, which is useful for validation studies or clinical applications where false positives are more problematic. The threshold should be chosen based on the specific requirements of your analysis.

What are the best hard filter thresholds for germline WES data?

The best hard filter thresholds depend on your specific dataset, including the sequencing platform, coverage depth, and variant caller. Common thresholds used in published workflows include a quality by depth of 2.0, a Fisher strand bias of 60, and a mapping quality of 40. However, these thresholds should be evaluated and adjusted based on the quality metrics of your specific dataset.

How do I evaluate whether my filtering approach is working?

Evaluate your filtering approach using multiple metrics, including the Ti/Tv ratio, the number of variants retained, the concordance with known variant databases, and the genotype concordance with array data if available. A Ti/Tv ratio around 2.0 for WGS data or 2.8 to 3.0 for WES data suggests that the filtering is working well. Significant deviations from these expected values indicate that the filtering parameters may need adjustment.

What is the difference between joint calling and per-sample calling for filtering?

Joint calling genotypes multiple samples together, which enables cohort-level quality control and produces a single multi-sample VCF. This approach, as implemented in the GermVarX workflow, allows you to assess the distribution of quality metrics across the entire cohort and identify problematic samples. Per-sample calling genotypes each sample independently, which is simpler but does not provide the same cohort-level context for filtering decisions.

How do I handle variants in repetitive or low-complexity regions?

Variants in repetitive or low-complexity regions are prone to alignment artifacts and may require additional filtering beyond what VQSR or hard filters provide. Consider applying region-specific filters or flagging variants in these regions for manual review. The specific approach depends on your downstream analysis goals and the genomic regions of interest.

Should I use functional annotation before or after filtering?

Functional annotation should be performed after filtering. Filtering removes sequencing artifacts and low-quality variants, and functional annotation then prioritizes the retained variants based on their predicted functional effects. Annotating before filtering can waste computational resources on variants that will be removed, and it can complicate the interpretation of functional annotations if artifacts are included.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.