How to Use gnomAD Allele Frequencies for Variant Filtering: Thresholds, Subpopulations, and Best Practices
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- gnomAD allele frequencies are critical for germline and somatic variant filtering, distinguishing common polymorphisms from rare disease-causing variants by comparing observed frequencies to population baselines.
- Threshold selection is context-dependent, requiring consideration of disease prevalence, inheritance mode (e.g., 0.6% for autosomal recessive, 0.1% for autosomal dominant as starting points), gene contribution, and penetrance, rather than applying universal fixed values.
- Filtering Allele Frequency (Grpmax), representing the maximum frequency across all subpopulations, is the conservative and preferred metric for filtering to mitigate misclassification of ancestry-specific common variants.
- Subpopulation frequencies are essential as variant prevalence can differ significantly across genetic ancestries due to founder effects or genetic drift, necessitating the use of the filtering allele frequency for accurate assessment.
- Disease prevalence directly informs allele frequency thresholds; for instance, a variant in a dominant disorder with 1:5000 prevalence cannot exceed a frequency of ~1:10000, while recessive disorders relate carrier frequency to the square root of disease prevalence.
- Documentation of filtering decisions is paramount, including the gnomAD version, specific thresholds used, rationale based on disease parameters, and observed allele frequencies, to ensure reproducibility and defensibility.
Variant filtering using gnomAD allele frequencies is a core decision step in germline and somatic variant calling workflows. The practical problem is straightforward: common polymorphisms must be removed from consideration while rare disease-causing variants must be retained. This article provides concrete guidance on accessing gnomAD data, interpreting subpopulation frequencies, and selecting thresholds based on disease prevalence and inheritance mode. The guidance applies to biology students, researchers, laboratory professionals, and life-science practitioners who work with variant calling pipelines and need defensible filtering decisions.
The Role of gnomAD in Variant Filtering Workflows
Population reference databases serve as the baseline against which observed variants are compared. The Genome Aggregation Database (gnomAD) aggregates exome and genome sequencing data from tens of thousands of individuals across multiple genetic ancestries. When a variant appears at high frequency in the general population, it is unlikely to be the sole cause of a rare monogenic disorder. When a variant is absent or extremely rare in the population, it remains a candidate for disease causation.
The filtering decision is not a single universal threshold. It depends on the disease mechanism, the inheritance pattern, the prevalence of the condition, and the gene contribution to the phenotype. A variant in a dominant disease gene with 1% population frequency may be benign, while the same frequency in a recessive disease gene may still be relevant. The threshold must be derived from the biology of the condition, not applied as a fixed number across all analyses.
The National Center for Biotechnology Information provides access to sequence databases, search systems, and analysis services that support variant interpretation workflows. Researchers can use these resources to cross-reference gnomAD findings with other population datasets and clinical databases. The European Bioinformatics Institute offers training pathways for bioinformatics data resources, including practical education on using population frequency data in genomic analysis. These official training materials help laboratory professionals build reproducible variant filtering workflows.
At a Glance: Threshold Selection by Inheritance Mode
The following table summarizes threshold considerations for common inheritance patterns. These values are starting points derived from published evaluations and must be adjusted for the specific disease context.
| Inheritance Mode | Suggested Starting Threshold | Rationale | Key Consideration |
|---|---|---|---|
| Autosomal recessive | 0.6% allele frequency | Based on Hardy-Weinberg equilibrium calculations for nonsyndromic hearing loss genes | Carrier frequency in the general population directly affects the threshold |
| Autosomal dominant | 0.1% allele frequency | Based on high-resolution variant frequency framework for dominant disease genes | Penetrance and prevalence of the specific disorder must be considered |
| Secondary findings genes | Disease-specific DAFT with BA1 at 10 times DAFT | Calculated from prevalence, gene contribution, and penetrance | No pathogenic variant filtering allele frequency exceeded the BA1 threshold in published evaluations |
A systematic evaluation of variants linked to hearing loss applied allele frequency thresholds of 0.6% for recessive genes and 0.1% for dominant genes. The study identified 48 variants in 23 genes with filtering allele frequencies above these thresholds and determined that most were benign or likely benign. This work demonstrates that threshold selection based on inheritance mode and Hardy-Weinberg calculations can systematically identify common polymorphisms that should be filtered from diagnostic consideration.
Accessing gnomAD Data for Variant Filtering
Navigating the gnomAD Browser
The gnomAD browser provides variant-level allele frequencies across multiple subpopulations. Each variant entry displays the overall allele frequency, the filtering allele frequency, and population-specific frequencies. The filtering allele frequency represents the maximum allele frequency observed across any genetic ancestry group, which is a conservative measure for variant filtering.
When accessing gnomAD data, record the following fields for each variant:
- Variant identifier with chromosome position, reference allele, and alternate allele
- Overall allele frequency across all populations
- Filtering allele frequency (Grpmax in gnomAD v4)
- Population-specific frequencies for each genetic ancestry group
- Number of alleles observed and total alleles genotyped
- Coverage metrics that indicate the quality of the genotype call
The Galaxy Training Network provides accessible workflow training for genomic analysis, including tutorials on variant annotation and filtering. These training materials help researchers build reproducible analysis pipelines that incorporate gnomAD data. The nf-core documentation describes community pipeline standards for variant calling and annotation, which include population frequency filtering steps.
Subpopulation Frequencies and Genetic Ancestry
Subpopulation frequencies matter because disease prevalence and variant frequencies differ across genetic ancestries. A variant that is rare globally may be common in one subpopulation due to founder effects or genetic drift. Using the filtering allele frequency, which is the maximum across all subpopulations, provides a conservative approach that reduces false positive variant classifications.
The filtering allele frequency is particularly important for variants that are common in one ancestry group but absent in others. If a variant has 0.01% frequency in European populations but 2% frequency in South Asian populations, the filtering allele frequency of 2% should be used for filtering decisions. This approach prevents the misclassification of ancestry-specific common variants as pathogenic.
The NCBI data resources provide access to multiple population databases and search systems that can be used to cross-reference gnomAD findings. Researchers should compare gnomAD frequencies with other population datasets when making filtering decisions for variants near threshold boundaries.
Core Principles of Allele Frequency Threshold Selection
Disease Prevalence and Allele Frequency Relationship
The relationship between disease prevalence and allele frequency follows population genetics principles. For a rare autosomal dominant disorder with prevalence of 1 in 5000, the disease-causing allele frequency is approximately 1 in 10000, assuming full penetrance and no selection against heterozygotes. A variant with population frequency substantially above this level cannot be the sole cause of the disorder.
For autosomal recessive disorders, the carrier frequency is related to the square root of the disease prevalence. A disorder with prevalence of 1 in 10000 has a carrier frequency of approximately 1 in 50, corresponding to an allele frequency of 1 in 100. Variants with allele frequencies above this level are unlikely to be fully penetrant recessive disease causes.
A study of hereditary hemorrhagic telangiectasia calculated prevalence estimates from gnomAD allele frequencies of predicted pathogenic variants in ENG and ACVRL1. The calculated prevalence ranged from 2.1 in 5000 to 11.9 in 5000, which is 2 to 12 times higher than current clinical estimates. This work demonstrates that summing allele frequencies of predicted pathogenic variants can estimate disease prevalence and reveal potential underdiagnosis.
Inheritance Mode and Threshold Derivation
The inheritance mode determines the maximum tolerable allele frequency for a disease-causing variant. For autosomal dominant disorders, the allele frequency of a fully penetrant disease-causing variant cannot exceed the disease prevalence divided by two. For autosomal recessive disorders, the allele frequency cannot exceed the square root of the disease prevalence.
The hearing loss study applied this logic using Hardy-Weinberg equilibrium calculations. The researchers determined thresholds of 0.6% for recessive genes and 0.1% for dominant genes based on the prevalence of nonsyndromic hearing loss and the number of genes contributing to the condition. These thresholds were then used to systematically evaluate all reported pathogenic variants in 97 hearing loss genes.
Gene Contribution and Penetrance Adjustments
The gene contribution to a phenotype affects the threshold calculation. When multiple genes can cause the same disorder, the allele frequency threshold for each gene must account for the proportion of cases attributable to that gene. A gene that causes 10% of cases of a disorder with prevalence 1 in 5000 has a lower maximum allele frequency than a gene that causes 50% of cases.
Penetrance also affects threshold selection. Reduced penetrance allows disease-causing variants to exist at higher population frequencies because not all carriers manifest the disorder. For variants with incomplete penetrance, the allele frequency threshold should be adjusted upward to account for the proportion of carriers who do not develop the condition.
A study of secondary findings genes calculated disease allele frequency thresholds considering prevalence, gene contribution, and penetrance. The American College of Medical Genetics and Genomics criterion BS1 was set at the calculated DAFT value, with BA1 set at 10 times the calculated DAFT. No pathogenic or likely pathogenic variant filtering allele frequency exceeded the relevant BA1 threshold in the evaluation.
Practical Workflow for Variant Filtering with gnomAD
Step 1: Define the Disease Context
Before selecting a threshold, document the following parameters:
- Inheritance mode for the gene of interest
- Disease prevalence in the population being studied
- Proportion of cases attributable to the specific gene
- Penetrance of the disorder
- Mode of inheritance for the specific variant being evaluated
These parameters determine the maximum allele frequency that a disease-causing variant can have in the general population. Without this context, threshold selection is arbitrary and may lead to incorrect variant classification.
Step 2: Select the Appropriate Threshold
Use the disease context to calculate the maximum allele frequency for a disease-causing variant. For autosomal dominant disorders, divide the disease prevalence by two and adjust for gene contribution and penetrance. For autosomal recessive disorders, take the square root of the disease prevalence and adjust for gene contribution.
The hearing loss study provides a practical example. For recessive genes, the threshold was set at 0.6%. For dominant genes, the threshold was set at 0.1%. These values were derived from the prevalence of nonsyndromic hearing loss and the number of genes contributing to the condition.
Step 3: Query gnomAD for Each Variant
For each variant under evaluation, record the following data from gnomAD:
- Overall allele frequency
- Filtering allele frequency (Grpmax)
- Population-specific frequencies
- Allele count and total allele number
- Coverage and quality metrics
The filtering allele frequency is the primary value for filtering decisions because it represents the maximum frequency across all genetic ancestry groups. This conservative approach reduces the risk of misclassifying ancestry-specific common variants as pathogenic.
Step 4: Apply the Threshold
Compare the filtering allele frequency to the calculated threshold. If the filtering allele frequency exceeds the threshold, the variant is unlikely to be a fully penetrant disease-causing variant for the specified condition. This variant should be filtered from further consideration unless there is strong evidence for reduced penetrance or a modifying effect.
If the filtering allele frequency is below the threshold, the variant remains a candidate for disease causation. Additional evidence from functional studies, segregation analysis, and clinical databases should be considered before making a final classification.
Step 5: Document the Filtering Decision
Record the threshold used, the gnomAD version accessed, the filtering allele frequency observed, and the rationale for the filtering decision. This documentation supports reproducibility and allows other researchers to understand the basis for variant classification.
The Bioconductor project provides official package documentation and workflow guidance for reproducible genomic analysis. Researchers can use Bioconductor packages to automate variant filtering steps and document the parameters used in each analysis.
Options and Tradeoffs in Threshold Selection
Fixed Thresholds versus Disease-Specific Thresholds
Fixed thresholds such as 1% or 5% are simple to apply but may be inappropriate for rare diseases. A variant with 1% population frequency cannot cause a disorder with prevalence 1 in 10000 under a fully penetrant dominant model. Fixed thresholds also fail to account for differences in gene contribution and penetrance across disorders.
Disease-specific thresholds derived from prevalence, gene contribution, and penetrance provide more accurate filtering decisions. The secondary findings study calculated DAFTs for 58 monogenic disease entities across 47 genes, demonstrating that thresholds vary substantially across conditions. Using disease-specific thresholds improves variant classification consistency and reduces misclassification.
Filtering Allele Frequency versus Overall Allele Frequency
The filtering allele frequency is the maximum allele frequency observed across any genetic ancestry group. This value is more conservative than the overall allele frequency because it captures ancestry-specific variation. For variant filtering, the filtering allele frequency is preferred because it prevents the misclassification of variants that are common in specific subpopulations.
The overall allele frequency can be misleading for variants that are common in one ancestry group but rare in others. A variant with 0.5% overall frequency but 3% frequency in one subpopulation may appear rare when only the overall frequency is considered. Using the filtering allele frequency of 3% correctly identifies this variant as common in the relevant population.
Stringent versus Relaxed Thresholds
Stringent thresholds filter more variants but may remove true disease-causing variants with higher population frequencies. Relaxed thresholds retain more variants but increase the number of common polymorphisms that must be evaluated through other evidence.
The choice between stringent and relaxed thresholds depends on the analysis context. Diagnostic testing for rare disorders benefits from stringent thresholds that reduce false positive classifications. Research studies exploring variant pathogenicity may use relaxed thresholds to retain more candidates for functional evaluation.
Observations and Measurements in Variant Filtering
Variant Frequency Distributions in Disease Genes
Published evaluations of disease gene variants in gnomAD reveal that pathogenic and likely pathogenic variants are present in population databases more frequently than expected. A study of autosomal dominant disorder genes detected 2653 pathogenic or likely pathogenic variants in 253 genes within gnomAD. This finding demonstrates that population databases contain variants with clinical annotations that require careful evaluation.
The presence of clinically annotated variants in gnomAD does not automatically invalidate their pathogenicity. Reduced penetrance, variable expressivity, and late-onset disorders can allow disease-causing variants to persist in the population. The filtering decision must account for these factors when evaluating variants with frequencies near the calculated threshold.
Filtering Allele Frequency Comparisons
The secondary findings study compared gnomAD Grpmax filtering allele frequencies for pathogenic and likely pathogenic variants against calculated DAFTs. No pathogenic or likely pathogenic variant had a filtering allele frequency greater than the relevant BA1 threshold. This observation supports the use of BA1 at 10 times the calculated DAFT as a robust exclusion criterion.
For variants with filtering allele frequencies between the DAFT and BA1 values, the BS1 criterion applies. These variants are unlikely to be disease-causing but may require additional evidence for definitive classification. The study recommends using these frequency specifications for secondary findings genes until full ClinGen Variant Curation Expert Panel specifications are available.
Copy Number Variant Frequency Considerations
Population frequency filtering also applies to copy number variants. A study of gnomAD-reported copy number deletions identified common deletions associated with quantitative traits including uric acid levels, HDL cholesterol, red blood cell traits, and childhood obesity. These findings demonstrate that common structural variants can have phenotypic effects and should not be automatically filtered from consideration.
The study filtered copy number deletions with minor allele frequency below 0.05 and Hardy-Weinberg equilibrium P values below 1.0 x 10^-6 before association analysis. This approach retained common deletions while removing low-quality or rare variants that would not have sufficient statistical power for association testing.
Records and Documentation for Variant Filtering
Essential Records for Each Filtering Decision
Maintain the following records for each variant filtering decision:
- Gene name and variant identifier
- gnomAD version and access date
- Overall allele frequency and filtering allele frequency
- Population-specific frequencies for relevant ancestry groups
- Threshold used and the basis for threshold selection
- Disease prevalence, gene contribution, and penetrance values used in calculations
- Final filtering decision and rationale
These records support reproducibility and allow re-evaluation when updated gnomAD versions become available. The nf-core documentation emphasizes the importance of reproducible workflow configuration and parameter documentation for genomic analysis pipelines.
Version Control for gnomAD Data
gnomAD releases update allele frequencies as more samples are added and quality filters are refined. A variant that is rare in gnomAD v2 may have different frequencies in gnomAD v4 due to expanded population representation and improved variant calling. Document the gnomAD version used for each analysis and re-evaluate filtering decisions when new versions are released.
The Galaxy Training Network provides tutorials on reproducible genomic analysis that include version control for reference data and analysis parameters. These practices ensure that filtering decisions can be traced to specific data versions and analysis configurations.
Quality Metrics and Coverage Assessment
Variant quality metrics affect the reliability of allele frequency estimates. Low coverage at a variant site can produce inaccurate allele frequency calculations. Record the coverage metrics for each variant and consider excluding variants with insufficient coverage from filtering decisions.
The NCBI data resources provide access to sequence quality information and analysis services that support variant quality assessment. Researchers should verify that allele frequency estimates are based on adequate allele counts and coverage before making filtering decisions.
Common Failure Patterns in Variant Filtering
Applying a Universal Threshold Across All Diseases
The most common failure pattern is applying a single allele frequency threshold to all variants regardless of disease context. A threshold appropriate for a common disorder may be too lenient for a rare disorder, allowing common polymorphisms to remain in the candidate list. A threshold appropriate for a rare disorder may be too stringent for a common disorder, removing true disease-causing variants.
The hearing loss study demonstrated that thresholds of 0.6% for recessive genes and 0.1% for dominant genes were appropriate for that condition. These values would not be appropriate for disorders with different prevalence, gene contribution, or penetrance. Threshold selection must be based on the specific disease context.
Ignoring Subpopulation Frequencies
Using only the overall allele frequency without considering subpopulation frequencies can lead to incorrect filtering decisions. Variants that are common in specific genetic ancestry groups may appear rare when only the overall frequency is considered. The filtering allele frequency captures this ancestry-specific variation and should be used for filtering decisions.
The hereditary hemorrhagic telangiectasia study found similar disease prevalence across genetic ancestries when using machine learning-based classification of missense variants. This finding underscores the importance of considering subpopulation frequencies when evaluating variant pathogenicity.
Confusing Filtering Allele Frequency with Pathogenicity
A filtering allele frequency below the threshold does not prove pathogenicity. It only indicates that the variant is rare enough in the population to be a candidate for disease causation. Additional evidence from functional studies, segregation analysis, clinical databases, and in silico predictions is required for definitive classification.
The fibrillinopathy study identified pathogenic and likely pathogenic variants in gnomAD that were challenging to classify according to ACMG/AMP guidelines. The presence of these variants in the population database did not resolve their pathogenicity, and the study emphasized the need for careful variant classification beyond frequency filtering.
Neglecting Digenic and Multigenic Inheritance
Variant filtering that considers only single variant effects may miss cases where multiple variants contribute to the phenotype. The fibrillinopathy study discovered two families with co-occurring FBN1 and FBN2 variants causing phenotypes with mixed or modified clinical features. These digenic cases would be missed by filtering strategies that evaluate each variant independently.
When filtering variants for disorders with potential digenic inheritance, consider whether combinations of variants in related genes could contribute to the phenotype. The study concluded that neglecting digenic variants may lead to incomplete or missed diagnoses.
Limitations of Allele Frequency Filtering
Reduced Penetrance and Late-Onset Disorders
Allele frequency filtering assumes that disease-causing variants are under negative selection and therefore rare in the population. This assumption fails for variants with reduced penetrance or late-onset disorders. A variant that causes disease in only 10% of carriers can exist at higher population frequencies than a fully penetrant variant.
For disorders with reduced penetrance, the allele frequency threshold should be adjusted upward to account for the proportion of carriers who do not manifest the disorder. The threshold calculation should incorporate penetrance estimates when available.
Population Structure and Founder Effects
Population databases may not fully represent all genetic ancestry groups. Variants that are common in underrepresented populations may have inaccurate frequency estimates. The filtering allele frequency provides some protection against this limitation by using the maximum frequency across available subpopulations.
The hereditary hemorrhagic telangiectasia study found similar disease prevalence across genetic ancestries, suggesting that pathogenic variant frequencies are comparable across populations for this disorder. However, this finding may not generalize to all disorders and genes.
Database Artifacts and Sequencing Errors
Population databases contain sequencing errors and artifacts that can produce inaccurate allele frequency estimates. Variants in difficult-to-sequence regions may have unreliable frequencies. Quality metrics and coverage information should be reviewed before making filtering decisions based on gnomAD data.
The NCBI data resources provide access to sequence quality information and analysis services that support variant quality assessment. Researchers should verify that allele frequency estimates are based on adequate allele counts and coverage before making filtering decisions.
Somatic Variant Calling Considerations
Somatic variant calling requires different filtering approaches than germline variant calling. Somatic variants are present in tumor tissue and may not be represented in population databases. The absence of a variant from gnomAD does not provide the same evidence for somatic variant calling as it does for germline variant calling.
For somatic variant calling, population frequency filtering is used primarily to remove common germline polymorphisms that contaminate tumor sequencing data. The threshold for this purpose is typically higher than for germline variant calling because the goal is to remove common variants instead of to identify rare disease-causing variants.
Quality Controls and Reproducibility in Variant Filtering
Pipeline Configuration and Parameter Documentation
Variant filtering pipelines should document all parameters used in the analysis, including gnomAD version, threshold values, and quality filters. The nf-core documentation provides standards for pipeline configuration and usage that support reproducible genomic analysis. Following these standards ensures that filtering decisions can be reproduced by other researchers.
The Galaxy Training Network offers accessible workflow training that emphasizes reproducibility in genomic analysis. Researchers can use these training materials to build variant filtering workflows that document all analysis parameters and data versions.
Validation with Known Variants
Validate variant filtering workflows using known pathogenic variants and known benign polymorphisms. Confirm that the workflow retains known pathogenic variants and filters known benign variants. This validation provides confidence that the filtering parameters are appropriate for the intended analysis.
The hearing loss study used ClinVar and HGMD annotations to evaluate the performance of their filtering approach. The study confirmed that 47 of 48 variants with filtering allele frequencies above the suggested thresholds were benign or likely benign, validating the threshold selection.
Cross-Reference with Multiple Databases
Cross-reference gnomAD frequencies with other population databases to confirm variant frequency estimates. The NCBI data resources provide access to multiple sequence databases and search systems that can be used for this purpose. Discrepancies between databases may indicate quality issues or population representation differences that should be investigated.
The hearing loss study used gnomAD, ExAC, EVS, and 1000 Genomes data to evaluate variant frequencies. The use of multiple databases provided a more complete picture of variant frequencies across populations and supported the filtering decisions.
Professional Escalation Criteria
When to Seek Additional Expertise
Escalate variant filtering decisions to a clinical geneticist or molecular pathologist when:
- The variant frequency is near the calculated threshold and the classification would change based on small frequency differences
- The disorder has reduced penetrance or variable expressivity that complicates threshold calculation
- Multiple variants in the same or related genes are identified and digenic inheritance is possible
- The variant is in a gene with conflicting evidence for disease association
- The filtering decision would affect clinical reporting or patient management
The fibrillinopathy study demonstrated that variant classification according to ACMG/AMP guidelines can be challenging even for experienced researchers. Professional expertise may be required for variants that do not have clear frequency-based classifications.
When to Re-evaluate Filtering Decisions
Re-evaluate filtering decisions when:
- A new gnomAD version is released with updated allele frequencies
- New evidence about disease prevalence, gene contribution, or penetrance becomes available
- Additional family members are tested and segregation information is obtained
- Functional studies provide new evidence about variant effects
- Clinical presentation suggests a diagnosis that was not initially considered
The secondary findings study noted that frequency specifications can be used until full ClinGen Variant Curation Expert Panel specifications are available. As new specifications are published, filtering decisions should be updated to reflect the latest guidance.
A Decision Framework for Threshold Selection When Disease Prevalence Is Unknown
The threshold calculations described earlier depend on knowing disease prevalence, gene contribution, and penetrance. In practice, these parameters are often uncertain or unavailable for the specific disorder under investigation. This section provides a structured decision framework for selecting allele frequency thresholds when prevalence data are incomplete, when multiple genes contribute to the same phenotype, and when published thresholds from related disorders must be adapted to a new gene.
Tiered Threshold Selection Based on Evidence Availability
When disease prevalence is unknown, use a tiered approach that starts with conservative assumptions and relaxes them only when supporting evidence justifies the adjustment. This framework prevents arbitrary threshold selection while allowing refinement as more information becomes available.
Tier 1: No prevalence data available. Use the published threshold from the closest analogous disorder with the same inheritance mode. For autosomal recessive disorders, start with 0.6% allele frequency. For autosomal dominant disorders, start with 0.1% allele frequency. These values come from the systematic evaluation of nonsyndromic hearing loss variants, which derived thresholds using Hardy-Weinberg equilibrium calculations and a high-resolution variant frequency framework. Document that the threshold is borrowed from a related condition and note the assumptions about prevalence and gene contribution that underlie this choice.
Tier 2: Prevalence estimate available but gene contribution unknown. Calculate the maximum allele frequency using the disease prevalence and inheritance mode, then divide by the estimated number of genes that can cause the disorder. For a disorder with prevalence 1 in 10000 and 10 known causative genes with roughly equal contribution, the gene-specific prevalence is approximately 1 in 100000. For an autosomal dominant disorder, the maximum allele frequency would be approximately 0.0005% before penetrance adjustment. This calculation assumes equal gene contribution, which should be stated explicitly in the analysis documentation.
Tier 3: Prevalence and gene contribution both available. Use the full calculation described in the core principles section. For autosomal dominant disorders, divide the disease prevalence by two, multiply by the gene contribution proportion, and divide by the penetrance. For autosomal recessive disorders, take the square root of the disease prevalence, multiply by the gene contribution proportion, and divide by the penetrance. This tier produces the most defensible threshold but requires the most complete disease knowledge.
The tiered approach aligns with the secondary findings study that calculated disease allele frequency thresholds considering prevalence, gene contribution, and penetrance for 58 monogenic disease entities across 47 genes. That study demonstrated that thresholds vary substantially across conditions and that disease-specific calculations improve variant classification consistency.
Handling Genes with Multiple Associated Phenotypes
Some genes are associated with multiple distinct disorders that have different prevalences, inheritance modes, or penetrance values. The secondary findings study addressed this situation by combining disease allele frequency threshold values for genes associated with multiple monogenic disease entities without clear genotype-phenotype correlation. The combined threshold was set at a value that would be appropriate for the most common or most penetrant phenotype.
Apply the same logic when a gene under evaluation causes multiple phenotypes. Calculate the threshold for each phenotype separately, then use the most lenient threshold that remains defensible. This approach retains variants that could cause any of the associated phenotypes while still filtering variants that are too common to cause the most prevalent or most penetrant condition. Document the phenotype-specific thresholds and the rationale for selecting the final value.
For example, a gene that causes both a severe early-onset dominant disorder with prevalence 1 in 50000 and a milder late-onset dominant disorder with prevalence 1 in 5000 would have different maximum allele frequencies under each model. The threshold for the milder disorder would be approximately 10 times higher than the threshold for the severe disorder. Using the more lenient threshold retains variants that could cause either phenotype, while using the stricter threshold would incorrectly filter variants that cause only the milder condition.
Adapting Published Thresholds to New Genes
When published thresholds exist for a related gene or disorder, adapt them systematically instead of applying them without adjustment. The hearing loss study provides a model for this adaptation. The researchers determined thresholds of 0.6% for recessive genes and 0.1% for dominant genes based on the prevalence of nonsyndromic hearing loss and the number of genes contributing to the condition. These thresholds were then applied across 97 hearing loss genes.
To adapt a published threshold to a new gene, document the following adjustments:
- Difference in disease prevalence between the published disorder and the target disorder
- Difference in the number of genes contributing to each phenotype
- Difference in penetrance between the two conditions
- Difference in the proportion of cases attributable to the specific gene
For each parameter that differs, calculate the proportional adjustment to the threshold. If the target disorder has half the prevalence of the published disorder, divide the threshold by two. If the target gene contributes to 20% of cases while the published gene contributes to 10%, multiply the threshold by two. Apply all adjustments sequentially and document each step.
The fibrillinopathy study illustrates why this adaptation is necessary. The researchers detected 2653 pathogenic or likely pathogenic variants in 253 genes associated with autosomal dominant disorders within gnomAD. These variants had filtering allele frequencies that required careful evaluation against disease-specific thresholds. A universal threshold applied across all 253 genes would have misclassified a substantial proportion of these variants.
Decision Matrix for Threshold Selection
Use the following decision matrix when selecting a threshold for a new gene or disorder. This matrix structures the decision process and ensures that all relevant parameters are considered.
| Parameter | Data Available | Action |
|---|---|---|
| Disease prevalence | Yes | Calculate maximum allele frequency from prevalence and inheritance mode |
| Disease prevalence | No | Use published threshold from closest analogous disorder |
| Gene contribution | Yes | Multiply threshold by the proportion of cases attributable to the gene |
| Gene contribution | No | Divide threshold by the estimated number of causative genes |
| Penetrance | Yes | Divide threshold by the penetrance estimate |
| Penetrance | No | Assume full penetrance and document this assumption |
| Multiple phenotypes | Yes | Calculate threshold for each phenotype and use the most lenient defensible value |
| Multiple phenotypes | No | Use the threshold for the most prevalent or most penetrant phenotype |
Record the data availability for each parameter in the analysis documentation. This record allows other researchers to understand the basis for the threshold and to update the threshold when new prevalence, gene contribution, or penetrance data become available.
Validation of Adapted Thresholds
After selecting a threshold through this decision framework, validate it using known variants in the gene of interest. The hearing loss study provides a validation model. The researchers identified 48 variants in 23 genes with filtering allele frequencies above the suggested thresholds and confirmed that 47 of these variants were benign or likely benign using NSHL-optimized ACMG guidelines. Only one variant, the high-frequency GJB2 mutation c.109G greater than A, p.Val37Ile, remained classified as pathogenic despite exceeding the threshold.
This validation approach has two components. First, confirm that known benign polymorphisms in the gene have filtering allele frequencies above the selected threshold. Second, confirm that known pathogenic variants with strong clinical evidence have filtering allele frequencies below the threshold. Variants that violate either expectation require investigation. The variant may have reduced penetrance, the threshold may be incorrectly calibrated, or the clinical classification may be wrong.
The hereditary hemorrhagic telangiectasia study used a different validation approach. The researchers summed allele frequencies of predicted pathogenic variants in ENG and ACVRL1 using three methods and calculated prevalence estimates between 2.1 in 5000 and 11.9 in 5000. The calculated prevalence was 2 to 12 times higher than current clinical estimates, supporting the hypothesis that the disorder is underdiagnosed. This approach validates thresholds by comparing calculated disease prevalence with clinical estimates, revealing discrepancies that warrant investigation.
Documentation Requirements for Adapted Thresholds
Maintain the following records when using this decision framework:
- The tier used for threshold selection and the reason for that tier
- The published threshold used as a starting point and the source of that threshold
- Each adjustment applied and the parameter value that justified the adjustment
- The data availability for prevalence, gene contribution, and penetrance
- The validation results showing how known variants compare with the selected threshold
- The gnomAD version used for validation and the access date
The nf-core documentation emphasizes the importance of reproducible workflow configuration and parameter documentation. Apply the same standard to threshold selection documentation. The Galaxy Training Network provides tutorials on reproducible genomic analysis that include version control for reference data and analysis parameters. These practices ensure that threshold selection decisions can be traced to specific data sources and analysis configurations.
Common Errors in Threshold Adaptation
The most frequent error in adapting published thresholds is applying them without adjustment for differences in disease prevalence. A threshold of 0.6% for recessive hearing loss genes cannot be applied to a recessive disorder with 10 times lower prevalence without dividing the threshold by 10. This error leads to retention of common polymorphisms that should be filtered.
A second common error is ignoring gene contribution differences. A gene that causes 50% of cases of a disorder can tolerate a higher allele frequency threshold than a gene that causes 5% of cases. Applying the same threshold to both genes will filter variants in the minor gene that could be disease-causing.
A third error is failing to document the assumptions underlying the adapted threshold. Without documentation, other researchers cannot evaluate whether the threshold is appropriate for the disease context. The Bioconductor project provides official package documentation and workflow guidance that emphasizes reproducible analysis practices. Apply these practices to threshold selection documentation.
Escalation Criteria for Threshold Selection
Escalate threshold selection to a clinical geneticist or molecular pathologist when the decision framework produces a threshold that conflicts with published variant classifications, when the disorder has highly variable penetrance that cannot be estimated reliably, when multiple genes with different inheritance modes contribute to the same phenotype, or when the threshold selection would affect clinical reporting. The fibrillinopathy study demonstrated that variant classification according to ACMG/AMP guidelines can be challenging even for experienced researchers. Professional expertise may be required when the decision framework does not produce a clear threshold or when validation reveals unexpected variant frequency patterns.
Frequently Asked Questions
What is the difference between allele frequency and filtering allele frequency in gnomAD?
The allele frequency is the overall frequency of a variant across all individuals in the database. The filtering allele frequency is the maximum allele frequency observed across any genetic ancestry group. The filtering allele frequency is more conservative for variant filtering because it captures ancestry-specific variation that may be hidden in the overall frequency. For example, a variant with 0.1% overall frequency but 2% frequency in one subpopulation would have a filtering allele frequency of 2%. Using the filtering allele frequency prevents the misclassification of ancestry-specific common variants as rare disease-causing variants.
How do I calculate an allele frequency threshold for a specific disease?
Calculate the threshold based on disease prevalence, inheritance mode, gene contribution, and penetrance. For autosomal dominant disorders, divide the disease prevalence by two and adjust for the proportion of cases attributable to the gene and the penetrance of the disorder. For autosomal recessive disorders, take the square root of the disease prevalence and apply the same adjustments. The hearing loss study provides a practical example with thresholds of 0.6% for recessive genes and 0.1% for dominant genes based on Hardy-Weinberg equilibrium calculations.
Why should I use the filtering allele frequency instead of the overall allele frequency?
The filtering allele frequency is the maximum frequency across all genetic ancestry groups in gnomAD. This value is more conservative than the overall allele frequency because it captures variants that are common in specific subpopulations. A variant that is rare globally but common in one ancestry group may appear rare when only the overall frequency is considered. Using the filtering allele frequency prevents the misclassification of ancestry-specific common variants as pathogenic.
What threshold should I use for somatic variant calling?
Somatic variant calling requires different filtering approaches than germline variant calling. Population frequency filtering for somatic variants is used primarily to remove common germline polymorphisms that contaminate tumor sequencing data. The threshold for this purpose is typically higher than for germline variant calling. The specific threshold depends on the tumor type, sequencing depth, and the goal of the analysis. Population databases may not contain somatic variants, so the absence of a variant from gnomAD does not provide the same evidence for somatic variant calling.
Can a pathogenic variant have a high allele frequency in gnomAD?
Yes, pathogenic variants can have high allele frequencies in gnomAD under certain conditions. Reduced penetrance allows disease-causing variants to persist in the population because not all carriers manifest the disorder. Late-onset disorders can also allow disease-causing variants to reach higher frequencies because affected individuals may reproduce before symptoms appear. The fibrillinopathy study detected 2653 pathogenic or likely pathogenic variants in 253 autosomal dominant disorder genes within gnomAD, demonstrating that clinically annotated variants are present in population databases.
How do I account for reduced penetrance when selecting a threshold?
Reduced penetrance allows disease-causing variants to exist at higher population frequencies because only a proportion of carriers manifest the disorder. When calculating the allele frequency threshold, divide the expected disease-causing allele frequency by the penetrance. For example, a variant with 50% penetrance can exist at twice the frequency of a fully penetrant variant. The threshold calculation should incorporate penetrance estimates when available from the literature or clinical databases.
What should I do if a variant has conflicting frequencies across gnomAD versions?
Compare the variant frequencies across gnomAD versions and investigate the reason for the discrepancy. Differences may result from expanded population representation, improved variant calling, or changes in quality filters. Document the gnomAD version used for the filtering decision and re-evaluate the decision when new versions are released. Cross-reference with other population databases through the NCBI data resources to confirm the variant frequency estimate.
When should I escalate a variant filtering decision to a clinical geneticist?
Escalate when the variant frequency is near the calculated threshold and small frequency differences would change the classification, when the disorder has reduced penetrance or variable expressivity, when multiple variants in related genes are identified, when the variant is in a gene with conflicting disease association evidence, or when the filtering decision would affect clinical reporting or patient management. The fibrillinopathy study demonstrated that variant classification can be challenging even for experienced researchers, and professional expertise may be required for complex cases.
Related Bioinformatics Guides
- Detecting Structural Variants with Long-Read Sequencing: Methods and Considerations
- Radiomics Feature Selection: Methods and Best Practices
- Longitudinal Microbiome Data Analysis: Methods and Best Practices
- Metagenomic Contamination Control: Best Practices for Clean Data
- Spatial Transcriptomics Differential Expression: Methods and Best Practices
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- Systematic evaluation of gene variants linked to hearing loss based on allele frequency threshold and filtering allele frequency.. Scientific reports, 2019.
- Hereditary Hemorrhagic Telangiectasia Prevalence Estimates Calculated From GnomAD Allele Frequencies of Predicted Pathogenic Variants in ENG and ACVRL1.. Circulation. Genomic and precision medicine, 2025.
- Variant filtering, digenic variants, and other challenges in clinical sequencing: a lesson from fibrillinopathies.. Clinical genetics, 2020.
- Specification of frequency criteria for secondary findings genes to improve variant classification concordance.. Genetics in medicine : official journal of the American College of Medical Genetics, 2026.
- Exploring quantitative traits-associated copy number deletions through reanalysis of UK10K consortium whole genome sequencing cohorts.. BMC genomics, 2023.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.