How to Filter Germline Variants Using Population Frequency Databases: A Step-by-Step Guide

By Dr. Zubair Khalid, DVM, MS, PhD ·

How to Filter Germline Variants Using Population Frequency Databases: A Step-by-Step Guide

Key Takeaways

  • Population frequency filtering leverages large reference cohorts (e.g., gnomAD, 1000 Genomes) to remove germline variants with allele frequencies exceeding a predefined threshold, based on the principle that rare Mendelian disease-causing mutations are unlikely to be common in healthy populations.
  • The optimal frequency threshold is assay-specific and influenced by factors such as inheritance model (stricter for dominant, adjusted for recessive), gene function, and the capture space of targeted sequencing panels, necessitating empirical validation rather than adoption of generic cutoffs.
  • Exome-derived and genome-derived allele frequencies can differ due to read-depth variations across gene exons; therefore, assessing read-depth independently for each data source is crucial for accurate frequency estimation.
  • Clonal hematopoiesis in blood-derived population data can inflate allele frequency estimates with somatic variants, necessitating manual review of variants in CH-associated genes to avoid filtering bona fide germline mutations.
  • Documentation of database version, reference genome build, filtering thresholds, and tool versions is paramount for ensuring reproducibility and enabling accurate interpretation of variant call sets across different analyses.
  • Variants in genes with reduced penetrance or late-onset disorders, and those with high carrier frequencies in specific subpopulations, may require adjusted filtering strategies or manual review, as they can be falsely filtered out by standard frequency thresholds.

Population frequency filtering removes variants from a germline variant call set when their allele frequency in large reference cohorts exceeds a predefined threshold, based on the assumption that common variants are unlikely to be rare disease-causing mutations. This workflow applies to germline variant call sets derived from whole-genome, whole-exome, or targeted sequencing, and uses gnomAD and 1000 Genomes data to reduce false positives before downstream interpretation. The intended reader is a bioinformatics student, researcher, laboratory professional, or life-science practitioner who has a VCF file and needs a practical, reproducible filtering protocol.

Scope and Reader Context

Germline variant calling identifies genetic variants inherited from parents and present in essentially all cells of an individual. A typical human genome contains millions of variants relative to the reference genome, and the vast majority are common polymorphisms with no clinical significance. Population frequency databases such as the Genome Aggregation Database (gnomAD) and 1000 Genomes provide allele frequency estimates from large cohorts of presumably healthy individuals. Filtering against these databases is a standard first step in reducing a variant call set to a manageable number of candidate variants for clinical or research interpretation.

This article covers the data inputs you need, the core principles of frequency-based filtering, a practical command-line workflow using common bioinformatics tools, options and tradeoffs for different inheritance models, quality controls, common failure patterns, limitations, and professional escalation criteria. The focus is on germline variant call sets, though the principles apply to tumour-only sequencing with appropriate caveats. The EMBL-EBI Training portal offers learning pathways for bioinformatics data-resource training that cover VCF handling and variant interpretation fundamentals, and the Galaxy Training Network provides accessible workflow training that includes tutorials on variant annotation and filtering using population databases.

At a Glance

Workflow ComponentRecommended PracticeKey Consideration
Population databaseUse gnomAD exome and genome datasets plus 1000 GenomesExome-derived and genome-derived frequencies can differ for specific gene exons due to read-depth variation
Frequency thresholdStart with 1% to 2% for autosomal dominant filteringThe optimal cutoff is assay-specific and influenced by cancer type and capture space
Inheritance model adjustmentApply stricter thresholds for recessive disordersCarrier frequency considerations require different logic than dominant disease filtering
Quality metric checkAssess read-depth separately for exome and genome datasetsA 30X read-depth achieves acceptable precision for substitutions but poor recall for small indels
Clonal hematopoiesis awarenessReview variants in CH-associated genes manuallySomatic variants in blood-derived population data can distort frequency estimates
DocumentationRecord database version, build, and filtering thresholdReproducibility requires exact version tracking

Data Inputs and Preparation

Variant Call Format Files

The starting point for germline variant filtering is a VCF file produced by a validated germline variant calling pipeline. The VCF format stores variant position, reference and alternate alleles, genotype information, and quality metrics. Before applying population frequency filters, confirm that your VCF has been produced with a validated germline calling workflow and that the reference genome build is recorded. The reference build matters because population frequency databases are aligned to specific builds, and liftover between builds can introduce artifacts.

The NCBI Data Resources provide official descriptions of sequence resources and analysis services that can help you confirm the reference build and understand the data formats used in variant analysis. The The Carpentries Lessons provide foundational training in shell and data manipulation that is useful for constructing and debugging command-line workflows.

Population Frequency Databases

The two primary population frequency resources used in germline variant filtering are gnomAD and 1000 Genomes. gnomAD aggregates exome and genome sequencing data from hundreds of thousands of individuals and provides allele frequencies for multiple ancestral populations. 1000 Genomes provides whole-genome sequencing data from a smaller but well-characterized set of global populations.

The Bioconductor project hosts packages for genomic analysis that can retrieve and annotate variants with population frequency data in a reproducible manner. The nf-core Documentation describes community pipeline standards that incorporate population frequency filtering into reproducible analysis workflows.

Database Version and Build Consistency

Record the exact version of the population database you use, including the build (for example GRCh37 or GRCh38) and the release number. Different versions of gnomAD have different cohort sizes and allele frequency estimates. A variant that is absent from an older release may appear at low frequency in a newer release as cohort size increases. Version tracking is essential for reproducibility and for comparing results across studies.

Core Principles of Population Frequency Filtering

The Rationale for Frequency Thresholds

The underlying logic is that a truly pathogenic variant causing a rare Mendelian disease should be rare in the general population. If a variant is observed at high frequency in a large cohort of presumably healthy individuals, it is unlikely to be a highly penetrant disease-causing mutation. This logic is sound for highly penetrant dominant disorders but requires careful adjustment for recessive disorders, reduced penetrance, and late-onset conditions.

The ESMO Precision Medicine Working Group recommendations describe how population variant frequency filters are applied to tumour-detected variants to identify those of true germline origin. In their analysis of 49,264 paired tumour-normal samples, they applied filters based on variant allele frequency, predicted pathogenicity, and population variant frequency to 58 cancer-susceptibility genes. The germline conversion rate, which is the proportion of filtered variants of true germline origin, varied substantially by gene. For BRCA1, BRCA2, and PALB2 the germline conversion rate exceeded 80%, while for TP53, APC, and STK11 it was below 2%. This demonstrates that the utility of population frequency filtering depends heavily on the gene being examined.

Frequency Threshold Selection

The choice of frequency threshold is a balance between sensitivity and specificity. A high threshold, such as 5%, will remove more common variants and reduce the number of false positives, but it may also remove true pathogenic variants that occur at appreciable frequency in certain populations. A low threshold, such as 0.1%, will retain more rare variants but will also retain more common benign variants that need manual review.

The optimization study on population frequency cutoffs demonstrated that the 1% to 2% cutoffs widely used in bioinformatic pipelines resulted in high sensitivity for classification of somatic variants but unnecessarily reduced sensitivity for germline variants. Using optimized cutoffs, the source of variants in The Cancer Genome Atlas data could be predicted with greater than 95% accuracy. Importantly, the optimal cutoff was influenced by both cancer type and the assay region of interest, and cutoffs were not transferable between assays even when the gene set was held constant. This means that the frequency threshold should be optimized for your specific assay and clinical context instead of adopted from a published pipeline without validation.

Filtering Allele Frequency versus Filtering Allele Count

Population databases report both allele frequency and allele count. Allele frequency is the proportion of chromosomes carrying the variant in the cohort. Allele count is the raw number of chromosomes carrying the variant. For very rare variants, allele count can be more informative than allele frequency because a variant observed once in a small subpopulation may have a misleadingly high frequency in that subpopulation. Some filtering approaches use a maximum allele count threshold in addition to or instead of a frequency threshold.

Practical Workflow for Filtering Germline Variants

Step 1: Annotate the VCF with Population Frequency Data

The first step is to annotate your VCF with allele frequencies from gnomAD and 1000 Genomes. Several tools can perform this annotation, including bcftools, VEP, and SnpEff. The annotation adds fields to the VCF that contain the allele frequency from each database and each population subcategory.

A typical command using bcftools with a gnomAD VCF as the annotation source would look like this:

bcftools annotate -a gnomad.exomes.vcf.gz -c INFO,gnomad_AF input.vcf.gz -o annotated.vcf.gz

The exact syntax depends on the version of bcftools and the structure of the annotation file. The The Carpentries Lessons provide foundational training in shell and data manipulation that is useful for constructing and debugging these command-line workflows.

Step 2: Apply the Frequency Filter

Once the VCF is annotated, apply the frequency filter using bcftools or a similar tool. The filter removes variants where the allele frequency exceeds the chosen threshold.

bcftools filter -i 'INFO/gnomad_AF < 0.01' annotated.vcf.gz -o filtered.vcf.gz

This command retains variants with a gnomAD allele frequency below 1%. For a recessive disorder analysis, you may need to adjust the threshold based on carrier frequency considerations.

Step 3: Apply Inheritance Model Specific Filters

Different inheritance models require different filtering logic. For autosomal dominant disorders, a variant observed at any appreciable frequency in the population is unlikely to be pathogenic, so a threshold of 1% or lower is appropriate. For autosomal recessive disorders, the carrier frequency can be higher, and the relevant threshold depends on the disease prevalence. For X-linked disorders, consider the allele frequency in males separately from the overall frequency.

Step 4: Review Variants That Fail the Filter

Variants that fail the frequency filter are not automatically benign. The cautionary study on TP53 variants in gnomAD demonstrated that a significant number of certified oncogenic TP53 variants are included in gnomAD and ExAC, most corresponding to TP53 hotspot variants occurring as somatic and germline events in human cancer. Disease-associated variants for BRCA1, BRCA2, APC, PTEN, and MLH1 were also identified in population databases. This means that population databases must be used with caution and should be annotated for the presence of oncogenic variants to improve their clinical utility.

Step 5: Document the Filtering Process

Record the database versions, the filtering threshold, the reference build, and the tool versions used. This documentation is essential for reproducibility and for interpreting results in a clinical context.

Options and Tradeoffs in Filtering Approaches

Exome-Derived versus Genome-Derived Frequencies

gnomAD provides separate allele frequency estimates from exome sequencing and genome sequencing. These estimates can differ for specific variants and genes. The study on population frequency data considerations found that exome-derived and genome-derived datasets exhibited low read-depth for different gene exons. Individual variants were mostly assigned to non-divergent frequency bins in more than 95% of cases, but two major bin divergences were resolved by applying a minimal acceptable read-depth threshold. This means that you should assess read-depth separately for population datasets sourced from different short-read sequencing technologies before assigning a frequency-based classification code.

Filtering on Maximum Subpopulation Frequency

Some variants have a high allele frequency in a specific subpopulation but are absent or rare in the overall cohort. Filtering on the overall allele frequency may retain these variants, while filtering on the maximum subpopulation frequency would remove them. The choice depends on whether you are concerned about variants that are common in any population or only variants that are common overall. For clinical applications where the ancestry of the sample is known, the relevant subpopulation frequency may be more appropriate than the overall frequency.

Hard Filters versus Probabilistic Approaches

The workflow described above uses hard filters, where variants above a threshold are removed and variants below the threshold are retained. An alternative approach is to use probabilistic methods that incorporate population frequency as one piece of evidence in a variant prioritization score. These methods can be more nuanced but are also more complex to implement and interpret.

Quality Controls and Measurements

Read-Depth Assessment

Population frequency estimates are only reliable if the underlying sequencing data has adequate read-depth at the variant position. The population frequency data considerations study found that a 30X read-depth achieved acceptable precision and recall for detection of substitutions but poor recall for small insertions and deletions. Before relying on a population frequency estimate for a specific variant, check the read-depth in the population database at that position.

Genotype Quality Checks

The genotype quality in your own variant call set is equally important. A variant with low genotype quality may have an inaccurate allele frequency estimate in your sample, and filtering decisions based on that estimate may be unreliable. Review the genotype quality metrics in your VCF and consider filtering out variants with low quality before applying population frequency filters.

Strand Bias and Mapping Quality

Variants with strand bias or low mapping quality may be artifacts instead of true variants. Population frequency filtering should be applied after basic quality filtering to remove sequencing artifacts. The order of filtering steps matters, and applying frequency filters before quality filters can retain artifacts that happen to be rare in the population.

Common Failure Patterns

Filtering Out True Pathogenic Variants

The most serious failure pattern is removing a true pathogenic variant because it appears at appreciable frequency in a population database. This can occur for several reasons. The variant may be a founder mutation that is common in a specific population. The variant may have reduced penetrance, meaning that carriers do not always develop the disease. The variant may be in a gene where clonal hematopoiesis in the population cohort produces somatic variants that inflate the frequency estimate.

The clonal hematopoiesis study explains that clonal hematopoiesis obstructs variant interpretation because somatic variants that provide a proliferative advantage will affect variant frequencies, depletion scores, and downstream filtering. Default filtering of variants or genes associated with clonal hematopoiesis risks filtering bona fide germline variants, as variants associated with clonal hematopoiesis can also cause Mendelian conditions. The study provides recommendations for careful review of 36 established clonal hematopoiesis genes associated with neurodevelopmental conditions.

Retaining Too Many Common Variants

The opposite failure pattern is retaining too many common variants because the frequency threshold is set too low. This increases the manual review burden and may obscure the true pathogenic variant among many benign variants. The optimal threshold depends on the assay and the clinical context, as demonstrated by the population frequency cutoff optimization study.

Ignoring Population Substructure

A variant that is rare in the overall gnomAD cohort may be common in a specific ancestral population. If your sample is from that population, the variant may be a common polymorphism instead of a pathogenic variant. Filtering on the overall allele frequency will retain this variant, while filtering on the maximum subpopulation frequency would remove it. The choice depends on whether you are concerned about variants that are common in any population or only variants that are common overall.

Using Inconsistent Database Versions

If different members of a team use different versions of gnomAD, they may reach different conclusions about the same variant. Version tracking is essential for consistency and reproducibility.

Records and Documentation

What to Record

For each filtering run, record the following information:

  • The version and build of the population database used
  • The filtering threshold applied
  • The tool versions used for annotation and filtering
  • The reference genome build
  • The date of the analysis
  • The number of variants before and after filtering
  • The number of variants removed by the filter

Why Documentation Matters

Documentation supports reproducibility and enables comparison across samples and studies. If a variant is reported as pathogenic in one analysis but filtered out in another, the documentation allows you to determine whether the difference is due to a different database version, a different threshold, or a different annotation tool.

Limitations of Population Frequency Filtering

Reduced Penetrance and Late-Onset Disorders

Population frequency filtering assumes that a pathogenic variant will be rare in a cohort of presumably healthy individuals. This assumption fails for variants with reduced penetrance, where many carriers do not develop the disease, and for late-onset disorders, where carriers may be healthy at the time of cohort recruitment. The ESMO Precision Medicine Working Group recommendations note that for genes such as TP53, APC, and STK11, the germline conversion rate in filtered tumour-detected variants was below 2%, meaning that most filtered variants were not of true germline origin. This does not mean that population frequency filtering is useless for these genes, but it does mean that the filter will retain many variants that are not true germline pathogenic variants.

Clonal Hematopoiesis Contamination

Population databases are built from blood-derived DNA, and clonal hematopoiesis can introduce somatic variants into the population frequency estimates. The clonal hematopoiesis study recommends careful review of variants in genes affected by clonal hematopoiesis and provides a list of 36 established clonal hematopoiesis genes associated with neurodevelopmental conditions. If your variant of interest is in one of these genes, the population frequency estimate may be inflated by somatic variants from clonal hematopoiesis.

Assay-Specific Cutoffs

The population frequency cutoff optimization study demonstrated that frequency cutoffs are not transferable between assays, even when the gene set is held constant. A cutoff that works well for a whole-exome assay may not work well for a targeted panel with a different capture space. The optimal cutoff is influenced by both cancer type and the assay region of interest. This means that you should validate the frequency cutoff for your specific assay instead of adopting a cutoff from a published pipeline.

Database Artifacts

Population databases contain artifacts, including sequencing errors, mapping errors, and annotation errors. A variant that appears to be rare in the population may be a sequencing artifact in the population database instead of a true rare variant. Conversely, a variant that appears to be common may be an artifact that inflates the frequency estimate. The NCBI Data Resources provide official descriptions of sequence resources that can help you understand the limitations of the data underlying population frequency estimates.

Safety and Regulatory Context

Clinical Reporting Considerations

If the filtered variant call set is used for clinical reporting, the filtering process must be validated and documented. The ESMO Precision Medicine Working Group recommendations provide updated recommendations around germline follow-up of tumour-only sequencing, including a revision to 5% for the minimum per-gene germline conversion rate, inclusion of actionable intermediate penetrance genes ATM and CHEK2, and definition of a set of seven most actionable cancer-susceptibility genes in which germline follow-up is recommended regardless of the filtering outcome.

Professional Escalation Criteria

Escalate to a clinical geneticist or molecular pathologist when:

  • A variant in a high-penetrance gene is retained after filtering but has conflicting evidence
  • A variant in a clonal hematopoiesis gene is being interpreted as germline
  • The frequency threshold needs to be adjusted for a specific clinical context
  • A variant is identified that may be a founder mutation in a specific population

Building a Population Frequency Filtering Decision Framework for Your Assay

A population frequency filter is not a single universal setting that can be copied from one pipeline to another. The evidence from the optimization study on population frequency cutoffs shows that the optimal cutoff is influenced by both cancer type and the assay region of interest, and that cutoffs are not transferable between assays even when the gene set is held constant. This means that every laboratory or research group needs a structured way to decide which frequency threshold to use, how to apply it across different inheritance models, and when to override the filter for specific variants or genes. This section provides a practical decision framework that you can implement with your own data, a record system for tracking filtering decisions, and a troubleshooting method for common problems that arise after filtering.

The Core Decision Points in Frequency Filtering

Before you run any filtering command, you need to make four decisions that will determine the behavior of your filter. These decisions are interconnected, and changing one often requires revisiting the others.

Decision 1: What is the purpose of your filter?

The purpose determines the stringency of the threshold. If you are filtering to reduce the variant call set for research prioritization, you can afford a more permissive threshold because downstream manual review will catch false positives. If you are filtering for clinical reporting, the threshold must be validated for your specific assay and the consequences of removing a true pathogenic variant are more serious. The ESMO Precision Medicine Working Group recommendations demonstrate that the germline conversion rate, which is the proportion of filtered variants of true germline origin, varies substantially by gene. For BRCA1, BRCA2, and PALB2 the germline conversion rate exceeded 80%, while for TP53, APC, and STK11 it was below 2%. This means that the same frequency filter can have very different performance depending on which genes are in your assay and what you are trying to achieve.

Decision 2: Which population frequency metric will you filter on?

The choice between overall allele frequency, maximum subpopulation frequency, and filter allele frequency affects which variants are retained. Overall allele frequency is the frequency across all individuals in the database. Maximum subpopulation frequency is the highest frequency observed in any single ancestral population. Filter allele frequency is a gnomAD-specific metric that excludes variants that failed quality filters in the contributing studies. For clinical applications where the ancestry of the sample is known, the relevant subpopulation frequency may be more appropriate than the overall frequency. For research applications where ancestry is unknown or admixed, the maximum subpopulation frequency is a more conservative choice because it removes variants that are common in any population.

Decision 3: What is your primary frequency threshold?

The evidence from the optimization study on population frequency cutoffs shows that the 1% to 2% cutoffs widely used in bioinformatic pipelines resulted in high sensitivity for classification of somatic variants but unnecessarily reduced sensitivity for germline variants. Using optimized cutoffs, the source of variants in The Cancer Genome Atlas data could be predicted with greater than 95% accuracy. This does not mean that 1% is always wrong or that a higher threshold is always better. It means that you need to determine the threshold empirically for your assay instead of assuming that a published cutoff will work.

Decision 4: How will you handle variants that fall near the threshold?

Variants with allele frequencies just below your threshold are retained, and variants just above are removed. This binary decision can be arbitrary when the frequency estimate has uncertainty. A variant observed once in a small subpopulation may have a frequency estimate with wide confidence intervals. Some filtering approaches use a buffer zone around the threshold where variants are flagged for manual review instead of being automatically retained or removed. This is particularly important for variants in genes where the germline conversion rate is low, as demonstrated by the ESMO Precision Medicine Working Group recommendations.

A Step-by-Step Decision Framework

The following framework walks you through the process of establishing a frequency filtering protocol for your specific assay. The framework is designed to be implemented once when you set up your pipeline and then revisited when you change assays, update databases, or encounter unexpected filtering results.

Step 1: Define your gene set and inheritance model

List the genes in your assay and classify each gene by the inheritance model that is relevant for your clinical or research question. Autosomal dominant genes require a different filtering logic than autosomal recessive genes. For autosomal dominant disorders, a variant observed at any appreciable frequency in the population is unlikely to be highly penetrant and pathogenic, so a threshold of 1% or lower is appropriate. For autosomal recessive disorders, the carrier frequency can be higher, and the relevant threshold depends on the disease prevalence. For X-linked disorders, consider the allele frequency in males separately from the overall frequency.

The ESMO Precision Medicine Working Group recommendations provide a useful model for this classification. They defined a set of seven most actionable cancer-susceptibility genes, including BRCA1, BRCA2, PALB2, MLH1, MSH2, MSH6, and RET, in which germline follow-up is recommended regardless of the filtering outcome. They also included actionable intermediate penetrance genes ATM and CHEK2. This gene-level classification allows the filter to be applied with different stringency for different genes in the same assay.

Step 2: Determine the read-depth profile of your assay

The population frequency data considerations study found that a 30X read-depth achieved acceptable precision and recall for detection of substitutions but poor recall for small insertions and deletions. The study also found that exome-derived and genome-derived datasets exhibited low read-depth for different gene exons. Before you can trust a population frequency estimate for a specific variant, you need to know whether the population database has adequate read-depth at that position.

For your own assay, calculate the median read-depth across all target regions and identify any genes or exons that fall below your quality threshold. These regions will have less reliable variant calls, and frequency filtering decisions based on variants in these regions should be treated with caution. The Bioconductor project hosts packages for genomic analysis that can calculate coverage statistics and identify low-coverage regions in a reproducible manner.

Step 3: Establish a baseline frequency threshold

Start with a baseline threshold based on the inheritance model and disease prevalence. For autosomal dominant disorders, a threshold of 1% is a reasonable starting point. For autosomal recessive disorders, calculate the maximum allowable carrier frequency based on the disease prevalence using the Hardy-Weinberg equation. For example, if the disease prevalence is 1 in 10,000, the carrier frequency is approximately 2%, and a variant with a population frequency above 2% is unlikely to be a pathogenic recessive allele.

The optimization study on population frequency cutoffs demonstrated that the optimal cutoff is influenced by both cancer type and the assay region of interest. This means that your baseline threshold is a starting point, not a final answer. You will need to validate and adjust the threshold using your own data.

Step 4: Validate the threshold with known positive and negative controls

The most reliable way to determine whether your frequency threshold is appropriate is to test it against a set of variants with known classification. Use a set of confirmed pathogenic variants from ClinVar or a similar curated database and a set of confirmed benign variants. Apply your filtering pipeline to both sets and measure the sensitivity and specificity of your threshold.

The population frequency data considerations study provides a model for this validation. The study compared exome-derived and genome-derived allele frequencies for six key cancer genes selected for variant curation by ClinGen expert panels. Individual variants were mostly assigned to non-divergent frequency bins in more than 95% of cases, but two major bin divergences were resolved by applying a minimal acceptable read-depth threshold. This validation approach can be adapted to your own assay by testing whether your threshold correctly classifies known variants in your gene set.

Step 5: Adjust the threshold for gene-specific considerations

The ESMO Precision Medicine Working Group recommendations demonstrate that the germline conversion rate varies substantially by gene. For genes with high germline conversion rates, such as BRCA1, BRCA2, and PALB2, a standard frequency filter will perform well because most filtered variants are of true germline origin. For genes with low germline conversion rates, such as TP53, APC, and STK11, the filter will retain many variants that are not of true germline origin, and you may need to apply additional evidence or a different threshold.

The cautionary study on TP53 variants in gnomAD demonstrated that a significant number of certified oncogenic TP53 variants are included in gnomAD and ExAC, most corresponding to TP53 hotspot variants occurring as somatic and germline events in human cancer. Disease-associated variants for BRCA1, BRCA2, APC, PTEN, and MLH1 were also identified in population databases. This means that for these genes, you should not assume that a variant present in gnomAD is benign. You may need to override the frequency filter for specific variants that are known oncogenic hotspots.

Step 6: Document the decision process

Record the rationale for each decision in the framework. This documentation is essential for reproducibility and for defending filtering decisions in a clinical context. The documentation should include the gene set, the inheritance model for each gene, the baseline threshold, the validation results, and any gene-specific adjustments.

A Record System for Filtering Decisions

A frequency filtering run produces a set of decisions that need to be tracked. The following record system captures the information needed to reproduce a filtering run and to troubleshoot unexpected results.

Run-level records

For each filtering run, record the following information:

  • The version and build of the population database used, including the release number
  • The filtering threshold applied and whether it was applied to overall allele frequency, maximum subpopulation frequency, or filter allele frequency
  • The tool versions used for annotation and filtering
  • The reference genome build
  • The date of the analysis
  • The number of variants before and after filtering
  • The number of variants removed by the filter

Variant-level records

For each variant that is retained after filtering, record the following information:

  • The allele frequency from each population database used
  • The subpopulation frequencies, particularly if the maximum subpopulation frequency was used for filtering
  • The read-depth in the population database at the variant position
  • The read-depth in your own sample at the variant position
  • Any gene-specific considerations that affected the filtering decision

Decision override records

When a variant is retained or removed despite the frequency filter, record the reason for the override. Common reasons include:

  • The variant is a known oncogenic hotspot that appears in population databases due to somatic events or reduced penetrance
  • The variant is in a gene affected by clonal hematopoiesis, and the population frequency estimate may be inflated by somatic variants
  • The variant is in a gene with a low germline conversion rate, and additional evidence is needed before making a filtering decision

The clonal hematopoiesis study explains that clonal hematopoiesis obstructs variant interpretation because somatic variants that provide a proliferative advantage will affect variant frequencies, depletion scores, and downstream filtering. The study provides recommendations for careful review of 36 established clonal hematopoiesis genes associated with neurodevelopmental conditions. If your variant of interest is in one of these genes, the population frequency estimate may be inflated by somatic variants from clonal hematopoiesis, and you should record this consideration in your decision override records.

Troubleshooting Common Filtering Problems

When a filtering run produces unexpected results, the following troubleshooting method can help you identify the cause and correct the problem.

Problem 1: Too many variants are retained after filtering

If your filtered variant call set is still too large for practical review, check the following:

  • Is the frequency threshold set correctly? A threshold that is too low will retain many common variants. The optimization study on population frequency cutoffs found that the 1% to 2% cutoffs widely used in bioinformatic pipelines resulted in high sensitivity for classification of somatic variants but unnecessarily reduced sensitivity for germline variants. If you are using a threshold below 1%, you may be retaining too many common variants.
  • Are you filtering on the correct frequency metric? If you are filtering on overall allele frequency, variants that are common in a specific subpopulation but rare overall will be retained. Filtering on maximum subpopulation frequency will remove these variants.
  • Are you applying the filter after quality filtering? If you apply the frequency filter before removing low-quality variants, you may be retaining artifacts that happen to be rare in the population. Apply quality filters first, then frequency filters.

Problem 2: True pathogenic variants are being removed by the filter

If you suspect that the filter is removing true pathogenic variants, check the following:

  • Is the variant in a gene affected by clonal hematopoiesis? The clonal hematopoiesis study explains that somatic variants in blood-derived population data can distort frequency estimates. If the variant is in one of the 36 established clonal hematopoiesis genes, the population frequency may be inflated by somatic variants.
  • Is the variant a known oncogenic hotspot? The cautionary study on TP53 variants in gnomAD demonstrated that a significant number of certified oncogenic TP53 variants are included in gnomAD and ExAC. If the variant is a known hotspot, you may need to override the frequency filter.
  • Is the variant in a gene with reduced penetrance or late-onset disease? Population frequency filtering assumes that a pathogenic variant will be rare in a cohort of presumably healthy individuals. This assumption fails for variants with reduced penetrance, where many carriers do not develop the disease, and for late-onset disorders, where carriers may be healthy at the time of cohort recruitment.

Problem 3: The same variant is filtered differently in different runs

If the same variant is retained in one run and removed in another, check the following:

  • Are you using the same database version? Different versions of gnomAD have different cohort sizes and allele frequency estimates. A variant that is absent from an older release may appear at low frequency in a newer release as cohort size increases.
  • Are you using the same frequency metric? Filtering on overall allele frequency versus maximum subpopulation frequency can produce different results for the same variant.
  • Are you using the same threshold? A threshold of 1% versus 2% will produce different results for variants with allele frequencies between 1% and 2%.

Problem 4: The filter performs differently for different genes in the same assay

This is expected behavior, not a problem. The ESMO Precision Medicine Working Group recommendations demonstrated that the germline conversion rate varies substantially by gene. For genes such as BRCA1, BRCA2, and PALB2, the germline conversion rate in filtered tumour-detected variants exceeded 80%, while for TP53, APC, and STK11 it was below 2%. If the filter is performing poorly for a specific gene, consider whether that gene needs a different threshold or additional evidence beyond population frequency.

Implementing the Framework in Practice

The decision framework described above can be implemented using standard bioinformatics tools. The Galaxy Training Network provides accessible workflow training that includes tutorials on variant annotation and filtering using population databases. The nf-core Documentation describes community pipeline standards that incorporate population frequency filtering into reproducible analysis workflows. The Bioconductor project hosts packages for genomic analysis that can retrieve and annotate variants with population frequency data in a reproducible manner.

The EMBL-EBI Training portal offers learning pathways for bioinformatics data-resource training that cover VCF handling and variant interpretation fundamentals. The The Carpentries Lessons provide foundational training in shell and data manipulation that is useful for constructing and debugging command-line workflows. The NCBI Data Resources provide official descriptions of sequence resources and analysis services that can help you confirm the reference build and understand the data formats used in variant analysis.

Professional Escalation Criteria for Filtering Decisions

Escalate to a clinical geneticist or molecular pathologist when:

  • A variant in a high-penetrance gene is retained after filtering but has conflicting evidence from functional studies or segregation analysis
  • A variant in a clonal hematopoiesis gene is being interpreted as germline without additional confirmation
  • The frequency threshold needs to be adjusted for a specific clinical context, such as a founder population or a family with a strong history of disease
  • A variant is identified that may be a founder mutation in a specific population, and the population frequency estimate may not reflect the true carrier frequency in that population
  • A variant in a gene with a low germline conversion rate, such as TP53, APC, or STK11, is being considered for clinical reporting

The ESMO Precision Medicine Working Group recommendations provide updated recommendations around germline follow-up of tumour-only sequencing, including a revision to 5% for the minimum per-gene germline conversion rate, inclusion of actionable intermediate penetrance genes ATM and CHEK2, and definition of a set of seven most actionable cancer-susceptibility genes in which germline follow-up is recommended regardless of the filtering outcome. These recommendations provide a useful framework for deciding when a variant warrants escalation regardless of the frequency filter outcome.

Frequently Asked Questions

What is the difference between filtering on allele frequency and filtering on allele count?

Allele frequency is the proportion of chromosomes carrying the variant in the population cohort. Allele count is the raw number of chromosomes carrying the variant. For very rare variants, allele count can be more informative because a variant observed once in a small subpopulation may have a misleadingly high frequency in that subpopulation. Some filtering approaches use a maximum allele count threshold in addition to or instead of a frequency threshold.

How do I choose the frequency threshold for my analysis?

The optimal threshold depends on your assay, the disease context, and the inheritance model. The population frequency cutoff optimization study found that the 1% to 2% cutoffs widely used in bioinformatic pipelines resulted in high sensitivity for classification of somatic variants but unnecessarily reduced sensitivity for germline variants. The optimal cutoff is influenced by both cancer type and the assay region of interest, and cutoffs are not transferable between assays. Validate the threshold for your specific assay instead of adopting a cutoff from a published pipeline.

Should I use exome-derived or genome-derived allele frequencies from gnomAD?

The population frequency data considerations study found that exome-derived and genome-derived datasets exhibited low read-depth for different gene exons. Individual variants were mostly assigned to non-divergent frequency bins in more than 95% of cases, but two major bin divergences were resolved by applying a minimal acceptable read-depth threshold. Assess read-depth separately for exome-derived and genome-derived datasets before assigning a frequency-based classification code.

Why are some pathogenic variants present in population databases?

The cautionary study on TP53 variants in gnomAD demonstrated that a significant number of certified oncogenic TP53 variants are included in gnomAD and ExAC, most corresponding to TP53 hotspot variants occurring as somatic and germline events in human cancer. Disease-associated variants for BRCA1, BRCA2, APC, PTEN, and MLH1 were also identified. Germline variants in the human population are more frequent than previously thought, and population databases must be used with caution.

What is clonal hematopoiesis and how does it affect population frequency filtering?

Clonal hematopoiesis is the presence of somatic variants in blood cells that provide a proliferative advantage. The clonal hematopoiesis study explains that these somatic variants will affect variant frequencies, depletion scores, and downstream filtering in population databases built from blood-derived DNA. Default filtering of variants or genes associated with clonal hematopoiesis risks filtering bona fide germline variants, as variants associated with clonal hematopoiesis can also cause Mendelian conditions.

How do I handle variants in genes with low germline conversion rates?

The ESMO Precision Medicine Working Group recommendations found that for genes such as TP53, APC, and STK11, the germline conversion rate in filtered tumour-detected variants was below 2%. For these genes, population frequency filtering will retain many variants that are not of true germline origin. Consider whether the gene is appropriate for population frequency filtering or whether additional evidence is needed.

What quality metrics should I check before relying on a population frequency estimate?

The population frequency data considerations study found that a 30X read-depth achieved acceptable precision and recall for detection of substitutions but poor recall for small insertions and deletions. Check the read-depth in the population database at the variant position, and also check the genotype quality in your own variant call set before making filtering decisions.

How do I document the filtering process for reproducibility?

Record the version and build of the population database, the filtering threshold, the tool versions used for annotation and filtering, the reference genome build, the date of the analysis, and the number of variants before and after filtering. This documentation supports reproducibility and enables comparison across samples and studies.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.