How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Hard filtering applies fixed thresholds to variant quality annotations (QD, FS, MQ, MQRankSum, ReadPosRankSum) to remove low-confidence germline variant calls, offering transparency and reproducibility for smaller cohorts (<30 samples) or when Variant Quality Score Recalibration (VQSR) is not feasible due to insufficient data size.
- Specific hard filter thresholds differ for Single Nucleotide Polymorphisms (SNPs) and Insertions/Deletions (indels), with indels requiring more lenient FS (>200.0) and ReadPosRankSum (< -20.0) thresholds due to distinct error profiles and alignment artifact patterns.
- Depth (DP) filters are critical and must be adjusted based on data type; whole-genome sequencing (WGS) typically uses a minimum DP of 10, while whole-exome sequencing (WES) requires higher minimums (e.g., 20) and stricter maximums (e.g., 300-500) to account for capture bias and uneven coverage.
- For multi-sample cohorts, joint calling followed by filtering is recommended, but filtering on cohort-level DP is less informative than per-sample genotype quality (GQ) filters (e.g., GQ < 20) to ensure consistent genotype call confidence across samples.
- Reproducibility necessitates meticulous documentation of pipeline versions, exact filter expressions and thresholds, variant counts before and after filtering, and coverage statistics, enabling validation through metrics like the transition/transversion ratio (Ti/Tv) and known variant concordance.
- Common filtering failures include applying SNP thresholds to indels, ignoring depth filters, using default thresholds without data-specific adjustment, filtering before joint calling, and overlooking specific challenges like strand bias in FFPE samples.
Direct Answer and Scope
Hard filtering is the process of applying fixed threshold values to variant quality annotations in a VCF file to remove low-confidence calls before downstream analysis. For germline variant calling with GATK HaplotypeCaller, the practical question is which thresholds to use and how to apply them correctly with command-line tools. This article provides concrete parameter values, exact commands, and adjustment strategies for whole-genome and exome data. The guidance applies to researchers, laboratory professionals, and bioinformatics students who need reproducible filtering decisions instead of vague recommendations.
The GATK toolkit has remained a foundational standard in short-read variant calling since its release in 2010, with continuous maintenance and adaptation to evolving sequencing technologies [<a href="#ref-1">1</a>]. However, the official documentation has shifted emphasis toward Variant Quality Score Recalibration (VQSR) for large datasets, leaving many researchers uncertain about when and how to apply hard filters. This article addresses that gap with actionable thresholds, command examples, and decision criteria grounded in published evidence.
At a Glance: Hard Filtering Decision Table
| Data Type | Recommended Approach | Key Parameters | Primary Consideration |
|---|---|---|---|
| Whole-genome sequencing (WGS), single sample | Hard filtering with GATK VariantFiltration | QD < 2.0, FS > 60.0, MQ < 40.0, MQRankSum < -12.5, ReadPosRankSum < -8.0 | Lower depth per site requires relaxed depth filters but strict quality filters |
| Whole-exome sequencing (WES), single sample | Hard filtering with adjusted thresholds | QD < 2.0, FS > 60.0, MQ < 40.0, MQRankSum < -12.5, ReadPosRankSum < -8.0, DP < 10 | Capture bias and uneven coverage require explicit depth minimums |
| Multi-sample cohort (WES or WGS) | Joint calling followed by hard filtering or VQSR | Same quality thresholds plus cohort-level DP filters | VQSR requires sufficient variant sites, hard filtering is more transparent for small cohorts |
The choice between hard filtering and VQSR depends on dataset size and the need for interpretability. Hard filtering applies fixed thresholds that are easy to document and reproduce. VQSR uses machine learning to model variant quality but requires large-scale genomic data to train effectively [<a href="#ref-2">2</a>]. For cohorts with fewer than 30 samples, hard filtering is often the more practical choice.
Understanding Germline Variant Calling Workflows
The Germline Calling Pipeline
Germline variant calling identifies genetic variants present in an individual's constitutional DNA. The standard GATK workflow involves several stages: alignment of raw sequencing reads to a reference genome, marking duplicates, base quality score recalibration, variant calling with HaplotypeCaller, and finally filtering. Each stage affects the quality of the final variant set, and filtering is the last opportunity to remove artifacts before downstream interpretation.
The HaplotypeCaller tool performs local reassembly of haplotypes in regions showing evidence of variation, which improves sensitivity in complex regions compared to older caller algorithms. The output is a VCF file containing variant records with multiple annotation fields. These annotations describe sequencing depth, mapping quality, strand bias, and other metrics that distinguish true variants from sequencing or alignment artifacts.
Germline Versus Somatic Calling
Germline variant calling assumes the variant is present in all cells of the individual and typically uses a diploid model. Somatic variant calling, used in cancer research, compares tumor and normal samples to identify variants acquired during the individual's lifetime. The filtering strategies differ substantially. Somatic callers like Mutect2 have their own filtering logic and require paired sample analysis. This article focuses exclusively on germline variants, where the expected allele frequency is approximately 0.5 for heterozygous calls and 1.0 for homozygous alternate calls.
Joint Calling for Cohorts
For multi-sample projects, joint calling improves consistency across samples. The GermVarX workflow demonstrates a modern implementation of joint germline variant discovery, using GATK HaplotypeCaller in GVCF mode followed by joint genotyping with GATK GenotypeGVCFs [<a href="#ref-3">3</a>]. This approach produces a single multi-sample VCF where all samples are genotyped at the same variant sites, enabling consistent filtering across the cohort. Joint calling also improves sensitivity for variants present at low frequency in the population.
Core Principles of Hard Filtering
What Hard Filters Evaluate
Hard filters assess the evidence supporting each variant call. The key annotations used in GATK best practices include:
QD (Quality by Depth): The variant quality score divided by the depth of alternate allele observations. This normalizes quality by coverage, preventing high-depth sites from receiving inflated quality scores. Low QD values indicate variants supported by few reads relative to total depth.
FS (Fisher Strand): A measure of strand bias using Fisher's exact test. High FS values indicate that alternate alleles appear predominantly on one sequencing strand, which is a common artifact pattern.
MQ (Mapping Quality): The root mean square mapping quality of reads supporting the variant. Low MQ values suggest reads map ambiguously to multiple genomic locations.
MQRankSum (Mapping Quality Rank Sum Test): A statistical comparison of mapping quality between reads supporting the reference versus alternate alleles. Extreme negative values indicate the alternate-supporting reads have notably lower mapping quality.
ReadPosRankSum (Read Position Rank Sum Test): A comparison of the position of alternate alleles within reads. Variants near read ends are more likely to be artifacts.
DP (Depth): The total read depth at the variant site. Very low depth provides insufficient evidence, while extremely high depth in exome data may indicate mapping issues.
Why Thresholds Matter
The choice of thresholds directly affects the precision and sensitivity of the final variant set. Stringent thresholds remove more false positives but also discard true variants. Relaxed thresholds retain more variants but increase the false positive burden for downstream analysis. Published evidence shows that different variant calling pipelines produce discordant results, and filtering choices significantly impact the final call set [<a href="#ref-2">2</a>]. The challenge is selecting thresholds that balance these competing objectives for your specific data.
The Limitations of Hard Filtering
Hard filtering applies the same thresholds to all variants regardless of context. This is both its strength and weakness. The approach is transparent and reproducible, but it cannot adapt to regional differences in coverage or sequence complexity. Variant quality score recalibration offers an alternative that models the joint distribution of multiple annotations, but it requires large datasets to train effectively [<a href="#ref-2">2</a>]. For small projects, hard filtering remains the standard approach.
Practical Workflow: Applying Hard Filters with GATK
Step 1: Prepare the Input VCF
Before filtering, ensure your VCF file contains the necessary annotation fields. HaplotypeCaller and GenotypeGVCFs produce these annotations by default. If you are using a VCF from another source, verify that QD, FS, MQ, MQRankSum, ReadPosRankSum, and DP are present. The command grep -v "^#" yourfile.vcf | head -1 shows the FORMAT and INFO columns for the first variant.
Step 2: Apply Filters with VariantFiltration
The GATK VariantFiltration tool applies hard filters to a VCF file. The standard command for SNP filtering is:
gatk VariantFiltration \
-R reference.fasta \
-V input.vcf.gz \
-O output_snps_filtered.vcf.gz \
--filter-expression "QD < 2.0" --filter-name "LowQD" \
--filter-expression "FS > 60.0" --filter-name "HighFS" \
--filter-expression "MQ < 40.0" --filter-name "LowMQ" \
--filter-expression "MQRankSum < -12.5" --filter-name "LowMQRankSum" \
--filter-expression "ReadPosRankSum < -8.0" --filter-name "LowReadPosRankSum"
This command adds filter flags to the VCF but does not remove variants. Variants that pass all filters receive the PASS flag in the FILTER column. Variants failing any filter receive the corresponding filter name. Downstream tools can then include or exclude filtered variants as needed.
Step 3: Apply Depth Filters
Depth filtering requires separate consideration. For whole-genome data, the standard GATK recommendation includes a depth filter, but the appropriate threshold depends on your sequencing depth. A common approach is:
gatk VariantFiltration \
-R reference.fasta \
-V input.vcf.gz \
-O output_filtered.vcf.gz \
--filter-expression "DP < 10" --filter-name "LowDepth" \
--filter-expression "DP > 100" --filter-name "HighDepth"
The high depth filter is particularly relevant for exome data, where extreme depth can indicate reads from paralogous regions or other mapping artifacts. Adjust the thresholds based on your mean coverage. For 30x whole-genome data, a minimum depth of 10 provides reasonable sensitivity. For exome data with higher mean coverage, you may increase the minimum to 15 or 20.
Step 4: Filter Indels Separately
Indels require different thresholds than SNPs because their error profiles differ. The standard GATK recommendations for indels are:
gatk VariantFiltration \
-R reference.fasta \
-V input_indels.vcf.gz \
-O output_indels_filtered.vcf.gz \
--filter-expression "QD < 2.0" --filter-name "LowQD" \
--filter-expression "FS > 200.0" --filter-name "HighFS" \
--filter-expression "ReadPosRankSum < -20.0" --filter-name "LowReadPosRankSum"
The FS threshold is more lenient for indels because strand bias patterns differ from SNPs. The ReadPosRankSum threshold is also relaxed because indels are more likely to occur near read ends due to alignment artifacts.
Step 5: Separate SNPs and Indels Before Filtering
Best practice is to split the VCF into SNP and indel subsets before applying filters. The GATK SelectVariants tool accomplishes this:
gatk SelectVariants \
-R reference.fasta \
-V input.vcf.gz \
--select-type-to-include SNP \
-O snps.vcf.gz
gatk SelectVariants \
-R reference.fasta \
-V input.vcf.gz \
--select-type-to-include INDEL \
-O indels.vcf.gz
Filtering SNPs and indels separately prevents the different quality distributions from confounding threshold selection.
Step 6: Extract Passing Variants
After filtering, extract variants that passed all filters for downstream analysis:
gatk SelectVariants \
-R reference.fasta \
-V output_filtered.vcf.gz \
--exclude-filtered \
-O final_pass.vcf.gz
This command removes all variants with any filter flag, leaving only PASS variants.
Adjusting Filters for Whole-Genome Versus Exome Data
Whole-Genome Sequencing Considerations
Whole-genome sequencing provides relatively uniform coverage across the genome, with the exception of repetitive and low-complexity regions. The standard GATK hard filter thresholds work well for WGS data because the annotation distributions are relatively consistent across the genome. The primary adjustment involves depth filters, which should reflect the mean coverage of your experiment.
For 30x WGS data, a minimum depth of 10 retains most true variants while removing low-confidence calls. The maximum depth filter is less critical for WGS because extreme depth is less common outside of repetitive regions. However, applying a high depth filter around 100 can remove variants in duplicated regions that produce spurious calls.
Whole-Exome Sequencing Considerations
Exome sequencing introduces capture bias, where some regions are sequenced at much higher depth than others. This uneven coverage affects the distribution of quality annotations. The QD annotation partially accounts for this by normalizing quality by depth, but depth filters require careful adjustment.
For exome data with mean coverage of 100x, a minimum depth of 20 is appropriate. This threshold removes variants in regions with insufficient capture efficiency while retaining the majority of true variants. The maximum depth filter is more important for exome data because capture artifacts can produce extreme depth in specific regions. A maximum depth of 300 to 500 is reasonable, depending on your capture kit and sequencing platform.
Multi-Sample Cohort Adjustments
Joint calling across multiple samples changes the depth distribution at each variant site. The DP annotation in a multi-sample VCF represents the total depth across all samples, not the per-sample depth. Filtering on this combined DP is less informative than filtering on per-sample depth. The GermVarX workflow addresses this by incorporating sample-level quality control alongside cohort-level filtering [<a href="#ref-3">3</a>].
For multi-sample VCFs, consider using the per-sample genotype quality (GQ) field for additional filtering. The GQ field indicates the confidence in each genotype call. A common threshold is GQ < 20, which removes low-confidence genotype calls while retaining the variant site for other samples.
Options and Tradeoffs in Filtering Strategy
Hard Filtering Versus VQSR
Variant quality score recalibration uses machine learning to model the distribution of annotations in high-quality variant sites, then applies a sensitivity threshold to classify variants. This approach can achieve better precision than hard filtering because it accounts for the joint distribution of annotations [<a href="#ref-2">2</a>]. However, VQSR requires a large number of variant sites for training, typically at least 30 samples of whole-genome data or 100 samples of exome data.
For smaller datasets, hard filtering is more appropriate. The thresholds are transparent, easy to document, and reproducible across analyses. The tradeoff is that hard filtering may discard some true variants that fall outside the fixed thresholds, particularly in regions with unusual coverage or sequence context.
Consensus Approaches
Multiple variant callers produce discordant results, and combining callers can improve accuracy. The VariantMetaCaller approach uses support vector machines to combine information from multiple calling pipelines, achieving higher sensitivity and precision than individual callers [<a href="#ref-2">2</a>]. This approach is more complex than hard filtering but may be appropriate for projects where accuracy is critical.
For most research applications, a single caller with well-chosen hard filters provides acceptable accuracy. The additional complexity of multi-caller consensus is justified when the cost of false positives or false negatives is high, such as in clinical applications.
Annotation-Based Filtering
Beyond the standard GATK annotations, additional annotations can improve filtering. Population allele frequency from databases like gnomAD can identify common variants that are unlikely to be pathogenic. The NCBI provides access to multiple variation databases that can inform filtering decisions [<a href="#ref-4">4</a>]. However, population frequency filtering is a separate step from quality filtering and should be applied based on the research question.
Records and Measurements for Filtering Decisions
Essential Records to Maintain
Documenting filtering decisions is essential for reproducibility. Maintain the following records for each analysis:
Pipeline version information: Record the GATK version, reference genome build, and all tool versions used in the workflow. The GATK toolkit undergoes continuous updates, and version differences can affect results [<a href="#ref-1">1</a>].
Filter thresholds: Document the exact filter expressions and thresholds applied, including any adjustments made for data type or coverage.
Variant counts: Record the number of variants before filtering, after each filter, and in the final PASS set. This information helps evaluate whether thresholds are appropriate.
Coverage statistics: Document mean depth, percentage of target regions covered at specified depths, and other sequencing quality metrics.
Measuring Filter Performance
Several metrics help evaluate whether filtering thresholds are appropriate:
Transition/transversion ratio (Ti/Tv): The ratio of transitions to transversions in the filtered variant set. For human whole-genome data, this ratio is typically around 2.0 to 2.1. A lower ratio suggests false positive variants remain in the call set.
Known variant concordance: The proportion of filtered variants that match known variants in databases. The NCBI maintains databases of known human variants that can be used for this comparison [<a href="#ref-4">4</a>].
Genotype concordance: For samples with array-based genotype data, compare the sequencing-derived genotypes to the array genotypes. Discordance indicates genotyping errors.
Using Records to Adjust Thresholds
If the Ti/Tv ratio is lower than expected, consider tightening quality filters. If the known variant concordance is high but the total variant count is lower than expected, consider relaxing filters to improve sensitivity. These adjustments should be made systematically, changing one threshold at a time and documenting the effect on variant counts and quality metrics.
Common Failure Patterns in Hard Filtering
Applying SNP Thresholds to Indels
One of the most common errors is applying SNP-specific thresholds to indel calls. Indels have different quality distributions, and the standard SNP thresholds for FS and ReadPosRankSum are too stringent for indels. This results in excessive false negative calls in indel regions. Always split the VCF and apply separate filters to each variant type.
Ignoring Depth Filters
Some researchers omit depth filters entirely, assuming that quality scores adequately account for coverage. This is problematic for exome data, where low-depth regions are common. Variants in these regions may have acceptable quality scores but insufficient evidence for reliable genotyping. Always include a minimum depth filter appropriate for your data type.
Using Default Thresholds Without Adjustment
The standard GATK thresholds are starting points, not universal constants. Data from different sequencing platforms, capture kits, and sample types may require adjustment. The published evidence on variant calling accuracy emphasizes that filtering choices significantly impact results [<a href="#ref-2">2</a>]. Evaluate your data's annotation distributions before finalizing thresholds.
Filtering Before Joint Calling
Applying hard filters to individual sample VCFs before joint calling removes low-quality variants from consideration in the cohort. This can reduce sensitivity for variants that are poorly supported in one sample but well-supported in others. Filter after joint calling to maintain maximum sensitivity.
Overlooking Strand Bias in FFPE Data
Formalin-fixed paraffin-embedded (FFPE) samples present special challenges for variant calling due to formaldehyde-induced DNA damage. A study comparing fresh frozen and FFPE samples found that fresh frozen samples showed superior variant call accuracy, although FFPE data was sufficient to identify causal variants in some cases [<a href="#ref-5">5</a>]. If working with FFPE samples, expect higher artifact rates and consider more stringent strand bias filters.
Quality Controls and Validation
Pre-Filtering Quality Checks
Before applying filters, verify that the input VCF is well-formed and contains the expected annotations. Check the number of variant records, the distribution of quality scores, and the presence of required annotation fields. The Galaxy Training Network provides accessible tutorials for variant calling workflows that include quality assessment steps [<a href="#ref-6">6</a>].
Post-Filtering Validation
After filtering, validate the results using multiple approaches:
Visual inspection: Use a genome browser to examine a sample of filtered and passing variants. This manual check can identify systematic issues that automated metrics miss.
Comparison to known variants: Compare the filtered call set to known variant databases. The NCBI provides access to dbSNP and other variation resources for this purpose [<a href="#ref-4">4</a>].
Replicate concordance: If multiple samples from the same individual are available, compare the variant calls. Discordance indicates technical artifacts.
Reproducibility Controls
Reproducibility requires more than documenting commands. Use workflow management systems to ensure consistent execution. The nf-core documentation describes community standards for reproducible bioinformatics pipelines [<a href="#ref-7">7</a>]. The GermVarX workflow demonstrates a production-ready implementation using Nextflow, including containerization and parallelization across computing environments [<a href="#ref-3">3</a>].
Limitations and Interpretation Boundaries
What Hard Filtering Cannot Do
Hard filtering removes low-quality variant calls, but it cannot correct for systematic errors in alignment or variant calling. If the reference genome is incorrect in a region, or if reads are misaligned due to structural variation, hard filters will not fix these problems. Copy-number variations present particular challenges because they affect read depth in ways that quality annotations may not capture [<a href="#ref-8">8</a>].
The Limits of Annotation-Based Filtering
Quality annotations provide indirect evidence about variant accuracy. A variant with excellent quality scores can still be a false positive if the underlying alignment is incorrect. Conversely, true variants can fail quality filters in difficult regions. Hard filtering is a heuristic approach, not a definitive classification.
Population-Specific Considerations
Filtering thresholds developed for human data may not transfer directly to other species. The canine FFPE study demonstrates that variant calling workflows require adaptation for non-human samples [<a href="#ref-5">5</a>]. If working with non-human data, evaluate annotation distributions empirically instead of relying on human-derived thresholds.
When to Escalate to Professional Support
Seek additional expertise when:
Variant calls are used for clinical decisions: Clinical applications require validated pipelines and may require CLIA certification or equivalent regulatory compliance.
Filtering decisions affect diagnosis: If the variant call set is used to identify disease-causing variants, involve a molecular geneticist or clinical geneticist in the filtering strategy.
Unexpected quality metrics persist: If Ti/Tv ratios, known variant concordance, or other metrics remain abnormal after filtering adjustments, consult a bioinformatics specialist.
Structural variant analysis is needed: Copy-number variants and other structural variants require specialized tools beyond the standard GATK workflow [<a href="#ref-8">8</a>].
Safety and Regulatory Context
Research Versus Clinical Use
The filtering strategy described in this article is appropriate for research applications. Clinical use of variant calls requires additional validation, documentation, and regulatory compliance. The distinction between research and clinical genomics is important for determining appropriate quality thresholds and documentation requirements.
Data Management and Privacy
Germline variant data is sensitive personal information. Ensure that your data management practices comply with applicable regulations and institutional policies. The EMBL-EBI training resources provide guidance on responsible data management in bioinformatics [<a href="#ref-6">6</a>].
Reproducibility Standards
Funding agencies and journals increasingly require reproducible analysis workflows. The Carpentries lessons provide foundational training in reproducible computing practices, including version control and documentation [<a href="#ref-9">9</a>]. The Bioconductor project emphasizes reproducible genomic analysis through versioned packages and workflow documentation [<a href="#ref-10">10</a>].
Building a Decision Framework for Filter Threshold Selection
Why Fixed Thresholds Fail Without a Decision Framework
The standard GATK hard filter thresholds represent starting points derived from whole-genome sequencing data with specific coverage characteristics. Applying these values without understanding your data's annotation distributions leads to two failure modes: excessive false positives when thresholds are too lenient, or excessive false negatives when thresholds are too strict. Published evidence shows that different variant calling pipelines produce discordant results, and the choice of filtering thresholds significantly impacts the final call set [<a href="#ref-2">2</a>]. A systematic decision framework replaces guesswork with evidence-based threshold selection.
The core problem is that annotation distributions vary by sequencing platform, capture kit, coverage depth, sample quality, and even the reference genome build. A threshold that works for 30x Illumina whole-genome data may perform poorly on 100x exome data from a different capture kit. The decision framework described here uses your own data to derive thresholds, then validates those thresholds against known quality metrics before finalizing the filter set.
Step 1: Characterize Your Annotation Distributions
Before selecting any threshold, generate distribution summaries for each annotation you plan to filter on. This requires extracting the relevant fields from your VCF file and computing summary statistics. The GATK VariantsToTable tool converts VCF annotations to a tabular format suitable for analysis:
gatk VariantsToTable \
-R reference.fasta \
-V input.vcf.gz \
-F CHROM -F POS -F QD -F FS -F MQ -F MQRankSum -F ReadPosRankSum -F DP \
-O annotation_table.txt
This command produces a tab-delimited file with one row per variant and columns for each requested annotation. For multi-sample VCFs, add the -GF flag to extract genotype-level fields such as GQ and per-sample depth.
Once the table is generated, compute summary statistics for each annotation. The R programming environment, available through Bioconductor for reproducible genomic analysis [<a href="#ref-10">10</a>], provides straightforward functions for this purpose:
data <- read.table("annotation_table.txt", header = TRUE)
summary(data$QD)
quantile(data$QD, probs = c(0.01, 0.05, 0.10, 0.50, 0.90, 0.95, 0.99))
The quantile information is particularly valuable. The 1st and 5th percentiles of QD indicate where the lowest-quality variants cluster. The 95th and 99th percentiles of FS show where the most strand-biased variants fall. These empirical distributions guide threshold selection far better than generic recommendations.
For researchers who prefer graphical assessment, histograms and density plots reveal the shape of each annotation distribution. The Galaxy Training Network provides accessible tutorials for variant calling workflows that include visualization steps [<a href="#ref-6">6</a>]. A bimodal distribution for QD, for example, suggests two distinct populations of variants that may warrant different filtering approaches.
Step 2: Separate Variants by Type Before Distribution Analysis
SNPs and indels have fundamentally different annotation distributions. Combining them in a single distribution analysis produces misleading summary statistics. The GATK SelectVariants tool separates variant types before analysis:
gatk SelectVariants \
-R reference.fasta \
-V input.vcf.gz \
--select-type-to-include SNP \
-O snps_only.vcf.gz
gatk SelectVariants \
-R reference.fasta \
-V input.vcf.gz \
--select-type-to-include INDEL \
-O indels_only.vcf.gz
Generate separate annotation tables for each variant type and analyze them independently. The FS distribution for indels, for example, typically shows higher values than SNPs because alignment algorithms handle insertions and deletions differently. Applying SNP-derived FS thresholds to indels removes a substantial proportion of true indel calls.
The separation also matters for depth distributions. In exome data, indels often occur in regions with different capture efficiency than SNPs. Analyzing depth distributions separately reveals whether a single depth threshold is appropriate or whether separate thresholds are needed.
Step 3: Apply the Percentile-Based Threshold Selection Method
The percentile-based method uses your data's empirical distributions to derive thresholds that balance sensitivity and precision. The procedure follows a systematic sequence:
Quality by Depth (QD): Examine the 5th percentile of QD for SNPs. Variants below this value represent the lowest-quality calls in your dataset. The standard threshold of QD < 2.0 corresponds approximately to the 1st to 5th percentile in typical 30x whole-genome data. For exome data with higher coverage, the QD distribution shifts upward, and a threshold of QD < 2.0 may be too lenient. Consider setting the SNP QD threshold at the 5th percentile of your empirical distribution, then compare the resulting variant count and Ti/Tv ratio to the standard threshold.
Fisher Strand (FS): Examine the 95th percentile of FS for SNPs. The standard threshold of FS > 60.0 removes variants with significant strand bias. If your data shows a higher baseline FS due to library preparation artifacts, the 95th percentile may exceed 60.0, indicating that the standard threshold is appropriate. If the 95th percentile falls below 60.0, your data has less strand bias than typical datasets, and you may consider a more stringent threshold.
Mapping Quality (MQ): The standard threshold of MQ < 40.0 removes variants supported by reads with ambiguous mapping. Examine the 5th percentile of MQ. Values below 40.0 at the 5th percentile suggest alignment issues that warrant investigation before filtering. Low MQ across many variants may indicate problems with the reference genome or alignment parameters instead of variant calling artifacts.
MQRankSum and ReadPosRankSum: These rank sum tests compare the mapping quality and read position distributions between reference-supporting and alternate-supporting reads. The standard thresholds of MQRankSum < -12.5 and ReadPosRankSum < -8.0 for SNPs remove variants where alternate-supporting reads have notably lower quality or occur preferentially at read ends. Examine the 5th percentile of each annotation. If your data shows more extreme values, consider whether the underlying alignment or library preparation introduced systematic biases.
Depth (DP): The depth threshold requires separate consideration because it depends on sequencing coverage. For whole-genome data at 30x mean coverage, the 1st percentile of DP typically falls around 8 to 10. For exome data at 100x mean coverage, the 1st percentile may fall around 20 to 30. Set the minimum depth threshold at the 1st percentile of your empirical DP distribution, then adjust based on downstream analysis requirements.
The percentile-based method does not replace the standard thresholds. It provides a data-driven check on whether the standard thresholds are appropriate for your specific dataset. When the empirical percentiles align closely with the standard thresholds, use the standard values. When they diverge substantially, investigate the cause before adjusting thresholds.
Step 4: Validate Thresholds with Known Variant Concordance
After applying candidate thresholds, validate the resulting variant set against known variant databases. The NCBI provides access to dbSNP and other variation resources that document known human variants [<a href="#ref-4">4</a>]. The validation process compares your filtered call set to these databases to estimate the proportion of true positives.
The transition/transversion ratio (Ti/Tv) provides a quick validation metric. For human whole-genome data, the Ti/Tv ratio typically falls between 2.0 and 2.1. A ratio below 2.0 suggests false positive variants remain in the call set. A ratio above 2.2 may indicate over-filtering that removed true variants. The Ti/Tv ratio for exome data is typically higher, around 2.5 to 3.0, because exome capture targets coding regions with different base composition.
Calculate the Ti/Tv ratio using the filtered VCF:
gatk CollectVariantCallingMetrics \
-I final_pass.vcf.gz \
-R reference.fasta \
-O variant_metrics
The output includes the Ti/Tv ratio for the filtered variant set. Compare this value to the expected range for your data type. If the ratio falls outside the expected range, adjust thresholds systematically and recalculate.
Known variant concordance provides a more direct validation. The NCBI maintains databases of known variants that can be used for comparison [<a href="#ref-4">4</a>]. The proportion of your filtered variants that match known variants estimates the precision of your call set. A concordance rate above 95 percent for common variants suggests the filtering thresholds are appropriate. Lower concordance indicates false positives remain.
Step 5: Document Threshold Decisions and Rationale
The decision framework requires documentation of beyond the final thresholds but the reasoning behind them. Maintain a filtering decision record for each analysis that includes:
Data characteristics: Sequencing platform, coverage depth, capture kit version, sample type, and any known quality issues.
Annotation distributions: The percentile values for each annotation, separated by variant type. This information allows others to understand why specific thresholds were chosen.
Threshold selection rationale: Whether the standard thresholds were used, adjusted based on empirical distributions, or replaced with data-derived values.
Validation metrics: The Ti/Tv ratio, known variant concordance, and any other metrics used to assess filter performance.
Version information: GATK version, reference genome build, and all tool versions used in the filtering process. The GATK toolkit undergoes continuous updates, and version differences can affect results [<a href="#ref-1">1</a>].
This documentation serves multiple purposes. It enables reproducibility when the analysis is repeated or shared with collaborators. It provides a basis for comparing filtering decisions across different datasets. It also supports troubleshooting when downstream analysis reveals unexpected results.
Step 6: Establish a Threshold Adjustment Protocol
When validation metrics indicate problems, adjust thresholds systematically instead of arbitrarily. The adjustment protocol follows a defined sequence:
Start with the most impactful annotation. QD typically has the largest effect on variant counts because it combines quality and depth information. Adjust QD first and evaluate the effect on variant counts and Ti/Tv ratio.
Change one threshold at a time. Adjusting multiple thresholds simultaneously makes it impossible to determine which change produced the observed effect. Document the variant count and validation metrics after each single-threshold change.
Evaluate the direction of change. If the Ti/Tv ratio is too low, tighten thresholds that remove false positives, particularly QD and FS. If the Ti/Tv ratio is too high, relax thresholds that may be removing true variants, particularly MQRankSum and ReadPosRankSum.
Revalidate after each adjustment. Recalculate the Ti/Tv ratio and known variant concordance after each threshold change. Continue adjusting until the validation metrics fall within the expected range.
Set limits on adjustments. The standard thresholds represent reasonable starting points based on years of accumulated experience with GATK. Adjustments beyond approximately 50 percent of the standard value warrant investigation into underlying data quality issues instead of further threshold changes.
Common Failure Patterns in Threshold Selection
Using thresholds from a different data type. The standard GATK thresholds were developed primarily for whole-genome sequencing data. Applying them directly to exome data without examining annotation distributions leads to suboptimal filtering. Exome data typically has higher depth, different FS distributions, and different ReadPosRankSum patterns due to capture bias.
Ignoring the effect of coverage on QD. QD normalizes quality by depth, but the relationship is not perfectly linear. At very high depth, QD can be artificially inflated or deflated depending on the variant caller's quality calculation. Examine the relationship between QD and DP in your data before finalizing thresholds.
Applying the same thresholds to all samples in a cohort. Sample quality varies within a cohort. A sample with lower mean coverage or higher duplication rates may require different thresholds than a high-quality sample. The GermVarX workflow addresses this by incorporating sample-level quality control alongside cohort-level filtering [<a href="#ref-3">3</a>].
Overlooking the effect of reference genome build. Different reference genome builds have different base composition and repeat content, which affects annotation distributions. Thresholds validated on GRCh37 may not transfer directly to GRCh38.
Failing to revalidate after pipeline changes. Any change to alignment parameters, duplicate marking, base quality recalibration, or variant calling parameters affects annotation distributions. Revalidate thresholds after any pipeline modification.
Records and Measurements for Ongoing Filter Management
Maintain a filtering log that records the following information for each analysis run:
Input VCF statistics: Total variant count, SNP count, indel count, and the number of variants at each quality score bin.
Annotation distribution summaries: The percentile values for each annotation, separated by variant type. Store these as machine-readable files for automated comparison across runs.
Threshold values applied: The exact filter expressions and thresholds used, including any adjustments from the standard values.
Output VCF statistics: The number of variants passing filters, the number failing each filter, and the proportion of variants in each filter category.
Validation metrics: Ti/Tv ratio, known variant concordance, and any other metrics used to assess filter performance.
Pipeline version information: GATK version, reference genome build, and all tool versions used in the analysis.
These records enable systematic comparison across datasets and over time. If a new sequencing run produces annotation distributions that differ substantially from previous runs, the filtering log helps identify whether the difference stems from sample quality, library preparation, or pipeline changes.
When to Escalate Beyond Hard Filtering
The decision framework identifies situations where hard filtering is insufficient and alternative approaches are needed:
Persistent low Ti/Tv ratio despite threshold adjustment. If the Ti/Tv ratio remains below 2.0 after systematic threshold adjustment, the problem likely lies in alignment or variant calling instead of filtering. Investigate alignment quality, reference genome choice, and variant caller parameters before continuing to adjust filters.
High discordance between variant callers. Published evidence shows that different variant calling pipelines produce discordant results, and the choice of filtering thresholds significantly impacts the final call set [<a href="#ref-2">2</a>]. If multiple callers produce substantially different variant sets, consider consensus approaches that combine information from multiple pipelines.
Need for quantitative precision-based filtering. The VariantMetaCaller approach uses support vector machines to combine information from multiple calling pipelines, achieving higher sensitivity and precision than individual callers [<a href="#ref-2">2</a>]. This approach is more complex than hard filtering but may be appropriate when the cost of false positives or false negatives is high.
Structural variant analysis required. Copy-number variants and other structural variants require specialized tools beyond the standard GATK workflow [<a href="#ref-8">8</a>]. Hard filtering of SNVs and indels does not address the challenges of structural variant detection.
Clinical applications. Clinical use of variant calls requires additional validation, documentation, and regulatory compliance beyond what hard filtering provides. Involve a molecular geneticist or clinical geneticist in the filtering strategy when variant calls are used for clinical decisions.
FFPE sample data. Formalin-fixed paraffin-embedded samples present special challenges due to formaldehyde-induced DNA damage. A study comparing fresh frozen and FFPE samples found that fresh frozen samples showed superior variant call accuracy, although FFPE data was sufficient to identify causal variants in some cases [<a href="#ref-5">5</a>]. If working with FFPE samples, expect higher artifact rates and consider more stringent strand bias filters, or escalate to specialized FFPE-aware workflows.
Integrating the Decision Framework with Reproducible Workflows
The decision framework described here works best when integrated into a reproducible workflow management system. The nf-core documentation describes community standards for reproducible bioinformatics pipelines [<a href="#ref-7">7</a>]. The GermVarX workflow demonstrates a production-ready implementation using Nextflow, including containerization and parallelization across computing environments [<a href="#ref-3">3</a>].
A reproducible implementation of the decision framework includes:
Automated annotation distribution reporting. Generate annotation distribution summaries automatically after variant calling, before filtering decisions are made.
Versioned threshold configuration. Store threshold values in version-controlled configuration files that document when and why thresholds changed.
Automated validation reporting. Calculate Ti/Tv ratio and known variant concordance automatically after filtering, with alerts when metrics fall outside expected ranges.
Documentation generation. Produce filtering decision records automatically, including input statistics, threshold values, and validation metrics.
The Carpentries lessons provide foundational training in reproducible computing practices, including version control and documentation [<a href="#ref-9">9</a>]. The Bioconductor project emphasizes reproducible genomic analysis through versioned packages and workflow documentation [<a href="#ref-10">10</a>]. These resources support the implementation of reproducible filtering workflows.
The decision framework transforms hard filtering from a fixed set of commands into an evidence-based process that adapts to your data while maintaining reproducibility and documentation standards. The framework does not replace the standard GATK thresholds but provides a systematic method for evaluating whether those thresholds are appropriate for your specific dataset and for adjusting them when they are not.
Frequently Asked Questions
What is the difference between hard filtering and VQSR?
Hard filtering applies fixed thresholds to individual variant annotations, such as requiring QD greater than 2.0. VQSR uses machine learning to model the joint distribution of multiple annotations and classify variants based on their similarity to known true and false variant sites. VQSR can achieve better precision but requires large datasets for training [<a href="#ref-2">2</a>]. For small cohorts, hard filtering is more practical and transparent.
Why do SNP and indel filters use different thresholds?
SNPs and indels have different error profiles. Indel calls are more prone to strand bias and read position artifacts because alignment algorithms handle insertions and deletions differently than single nucleotide changes. The standard GATK recommendations use more lenient FS and ReadPosRankSum thresholds for indels to account for these differences.
How do I choose the minimum depth threshold?
The minimum depth threshold depends on your sequencing coverage and data type. For 30x whole-genome data, a minimum depth of 10 is common. For exome data with higher mean coverage, use a threshold that removes variants in poorly captured regions while retaining the majority of true variants. Examine the depth distribution in your data and choose a threshold that balances sensitivity and specificity.
Can I apply hard filters to a multi-sample VCF?
Yes, but be aware that the DP annotation in a multi-sample VCF represents total depth across all samples. Filtering on this combined depth is less informative than per-sample depth. Consider using genotype quality (GQ) for per-sample filtering and reserve DP filtering for extreme outliers.
What should I do if my filtered variant set has a low Ti/Tv ratio?
A low transition/transversion ratio suggests false positive variants remain in the call set. Tighten quality filters, particularly QD and FS thresholds. Also verify that you are using the correct reference genome and that the variant caller was run with appropriate parameters.
How do I handle FFPE sample data?
FFPE samples produce lower quality data due to formaldehyde-induced DNA damage. A study comparing fresh frozen and FFPE samples found that fresh frozen samples had superior variant call accuracy, but FFPE data was sufficient to identify causal variants in some cases [<a href="#ref-5">5</a>]. Expect higher artifact rates and consider more stringent strand bias filters. If possible, validate FFPE results with fresh frozen samples.
Should I filter before or after joint calling?
Filter after joint calling. Applying hard filters to individual sample VCFs before joint calling removes low-quality variants from consideration in the cohort, which can reduce sensitivity for variants that are poorly supported in one sample but well-supported in others.
What annotations are essential for hard filtering?
The essential annotations are QD, FS, MQ, MQRankSum, ReadPosRankSum, and DP. These annotations capture different aspects of variant quality: overall quality normalized by depth, strand bias, mapping quality, and read position effects. Verify that your VCF contains these annotations before applying filters.
Related Bioinformatics Guides
- Variant Calling Pipelines: GATK Best Practices, FreeBayes, and DeepVariant Comparison
- Gene Set Enrichment Analysis in R: A Practical Tutorial for Interpreting Omics Data
- Detecting Structural Variants with Long-Read Sequencing: Methods and Considerations
- Variant Calling GATK: Structural Analysis and Computational Methodologies in Bioinformatics
- Metabolomics Data Analysis in R: A Practical Workflow
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- Quality Control in Single-Cell RNA-Seq: Common Pitfalls and How to Avoid Them
References and Further Reading
[1] [Fifteen Years of the Genome Analysis Toolkit as the De Facto Standard in Short-Read Variant Calling.](https://doi.org/10.3390/ijms27093754). 2026. [2] [VariantMetaCaller: automated fusion of variant calling pipelines for quantitative, precision-based filtering.](https://pubmed.ncbi.nlm.nih.gov/26510841). BMC genomics, 2015. [3] [GermVarX: A Robust Workflow for Joint Germline Variant Exploration in whole-exome sequencing cohorts.](https://doi.org/10.1371/journal.pone.0345561). 2026. [4] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information. [5] [Germline Variant Call Accuracy in Whole Genome Sequence Data from Canine Formalin-Fixed Paraffin-Embedded Tissue Samples.](https://doi.org/10.3390/genes16111371). 2025. [6] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute. [7] [nf-core Documentation](https://nf-co.re/docs). nf-core. [8] [A Comparison of Tools for Copy-Number Variation Detection in Germline Whole Exome and Whole Genome Sequencing Data.](https://pubmed.ncbi.nlm.nih.gov/34944901). Cancers, 2021. [9] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries. [10] [Bioconductor](https://bioconductor.org/). Bioconductor Project.This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.