ExAC vs. gnomAD: What Changed and Why It Matters for Variant Interpretation

By Dr. Zubair Khalid, DVM, MS, PhD ·

ExAC vs. gnomAD: What Changed and Why It Matters for Variant Interpretation

Key Takeaways

  • gnomAD significantly expands upon ExAC in both sample size (over 700,000 exomes and 76,000 genomes) and population diversity, offering more statistically robust allele frequency estimates, particularly crucial for distinguishing rare pathogenic variants from common benign polymorphisms in germline variant interpretation.
  • The transition to gnomAD's joint calling approach enhances genotype consistency across samples, leading to more reliable allele frequency data for low-frequency variants compared to ExAC's per-sample calling.
  • gnomAD incorporates enhanced site quality metrics (e.g., inbreeding coefficient, allele balance) and includes genome sequencing data, providing a more comprehensive assessment of variant reliability and enabling interpretation of non-coding variants absent from ExAC.
  • For somatic variant calling, gnomAD's larger dataset improves the filtering of germline polymorphisms from tumor sequencing data, reducing false positives, while also offering better representation of rare population variants to distinguish them from true somatic mutations.
  • When implementing variant interpretation workflows, it is critical to use the latest gnomAD version, consider population-specific allele frequencies, and critically evaluate site quality metrics to avoid misclassification of variants.

The Exome Aggregation Consortium (ExAC) and its successor, the Genome Aggregation Database (gnomAD), are population reference datasets that provide allele frequency information for genetic variants observed in large numbers of sequenced human exomes and genomes. For researchers and laboratory professionals interpreting sequence variants, the transition from ExAC to gnomAD changed the scale, diversity, and processing of reference population data. This article explains the concrete differences between the two resources, how those differences affect allele frequency filtering in variant calling workflows, and when each resource remains useful for germline and somatic variant interpretation.

The Core Difference: Scale and Population Representation

ExAC was released in 2014 and contained exome sequencing data from 60,706 unrelated individuals. gnomAD, first released in 2017 and updated in subsequent versions, expanded this to over 125,000 exomes and 15,000 genomes in its initial release, with later versions exceeding 700,000 exomes and 76,000 genomes. This increase in sample size directly affects the statistical confidence in allele frequency estimates, particularly for rare variants.

The practical consequence for variant interpretation is straightforward. A variant observed once in ExAC had an allele frequency of approximately 1 in 121,000 alleles. The same single observation in a larger gnomAD release represents a lower frequency with tighter confidence intervals. For rare disease research, where distinguishing a truly rare pathogenic variant from a common benign polymorphism is central, the larger reference population reduces the chance of misclassifying a common variant as rare.

Population diversity also expanded between ExAC and gnomAD. ExAC drew primarily from European and African American populations, with smaller contributions from East Asian, South Asian, and Latino groups. gnomAD increased the representation of Finnish Europeans, Ashkenazi Jews, and other populations, and later versions added more individuals from African, East Asian, and Latino backgrounds. This matters because allele frequencies vary substantially across ancestral groups. A variant that is rare in one population may be common in another, and using a reference dataset that lacks adequate representation of the patient's ancestry can lead to incorrect frequency-based filtering decisions.

The NCBI maintains related population frequency resources and search systems that researchers can use to cross-check allele frequency data across databases [<a href="#ref-1">1</a>]. The EMBL-EBI training materials provide structured learning pathways for understanding how population reference data are generated and applied in variant interpretation [<a href="#ref-2">2</a>].

Data Processing and Quality Control Changes

The transition from ExAC to gnomAD involved substantial changes to the processing pipeline that affect how allele frequencies should be interpreted.

Joint Calling Versus Per-Sample Calling

ExAC used a per-sample variant calling approach followed by aggregation. gnomAD moved to joint calling, where all samples are processed together in a single variant calling run. Joint calling improves genotype consistency across samples and reduces batch effects that can arise when samples are processed separately. For users, this means gnomAD allele frequencies are generally more reliable for low-frequency variants because the genotype calls are more uniform across the dataset.

Site Quality Metrics

gnomAD introduced a more sophisticated set of quality metrics for each variant site. These include inbreeding coefficient, allele balance, and read depth metrics that are used to flag sites where the variant call may be unreliable. The gnomAD browser displays these metrics alongside allele frequencies, allowing users to assess whether a particular frequency estimate is trustworthy.

For example, a variant with a high allele frequency but poor site quality metrics may be an artifact of sequencing or mapping error instead of a true population variant. Users who rely solely on the allele frequency number without checking site quality can be misled. The gnomAD interface provides a quality flag system that marks sites failing specific filters, and these flags should be reviewed before making interpretation decisions.

Coverage Differences Between Exomes and Genomes

gnomAD includes both exome and genome sequencing data. The genome data provide coverage across intronic and intergenic regions that are not captured by exome sequencing. This is relevant for variant interpretation because some pathogenic variants fall outside coding regions, and the genome component of gnomAD provides frequency information for these variants that ExAC could not offer.

However, coverage depth differs between the exome and genome components. Exome sequencing typically provides higher depth in coding regions, while genome sequencing provides more uniform but lower depth across the genome. Users should check whether a variant's frequency is derived from the exome, genome, or combined dataset, as the confidence in the frequency estimate depends on the underlying coverage.

Allele Frequency Filtering in Variant Calling Workflows

Allele frequency filtering is a standard step in germline variant calling workflows. The principle is that a variant observed at high frequency in the general population is unlikely to be a rare disease-causing variant. The threshold for filtering depends on the disease model and the expected allele frequency of pathogenic variants in the relevant gene.

Germline Variant Calling

In germline variant calling, researchers typically filter out variants with allele frequencies above a threshold such as 1% or 0.1%, depending on whether the disease is dominant or recessive and how rare the condition is. The choice of threshold should account for the possibility that some pathogenic variants are present at low frequencies in the population, particularly for recessive conditions or for variants with reduced penetrance.

The larger gnomAD dataset provides more precise frequency estimates for variants at these filtering thresholds. A variant with a frequency of 0.5% in ExAC might have a frequency of 0.3% or 0.7% in gnomAD, and the larger sample size gives more confidence in which estimate is closer to the true population frequency. For variants near the filtering threshold, this precision can change whether a variant is retained or filtered out.

A study of hereditary cancer susceptibility genes in a Chinese cohort identified pathogenic or likely pathogenic germline variants in 1.57% of 10,456 cancer-free adults, with eight novel variants that were unreported in gnomAD and ClinVar [<a href="#ref-3">3</a>]. This finding illustrates that novel pathogenic variants can be absent from population databases, and allele frequency filtering should not be the sole criterion for variant classification. The study also found substantial ethnic disparities in carrier rates, with non-Han ethnic groups showing a 4.56-fold higher carrier rate than Han Chinese participants [<a href="#ref-3">3</a>]. This underscores the importance of using population reference data that includes the relevant ancestral group.

Somatic Variant Calling

Somatic variant calling, used in cancer genomics, has different filtering considerations. Somatic variants are acquired mutations present in tumor tissue but not in the germline. The allele frequency of a somatic variant reflects the proportion of tumor cells carrying the mutation, which is influenced by tumor purity and copy number changes.

Population databases such as gnomAD are used in somatic calling to filter out germline polymorphisms that are captured in the tumor sequencing data. A variant present at high frequency in gnomAD is likely a germline variant instead of a somatic mutation, even if it appears in the tumor sample. However, somatic callers must be careful not to filter out true somatic variants that happen to match a low-frequency population variant.

The key difference for somatic workflows is that gnomAD provides a more comprehensive list of common polymorphisms to filter against, reducing the number of germline variants that pass through somatic filters. The larger dataset also provides better representation of rare population variants, which can help distinguish between a rare germline variant and a true somatic mutation.

Variant Calling Workflow Integration

Integrating gnomAD into a variant calling workflow requires downloading the appropriate files and configuring the filtering steps. The gnomAD data are available as VCF files that can be used with standard bioinformatics tools. The nf-core documentation provides standards for community pipelines that incorporate population frequency filtering as part of reproducible variant calling workflows [<a href="#ref-4">4</a>].

The Galaxy Training Network offers accessible tutorials for variant calling workflows that include population frequency filtering steps [<a href="#ref-5">5</a>]. These tutorials demonstrate how to annotate variants with gnomAD allele frequencies and apply filtering thresholds in a reproducible manner.

For researchers using R for downstream analysis, the Bioconductor project provides packages for working with variant annotation data, including tools for filtering based on population frequencies [<a href="#ref-6">6</a>]. These packages allow programmatic access to gnomAD data and integration with statistical analysis workflows.

At a Glance: ExAC Versus gnomAD

FeatureExACgnomAD
Initial sample size60,706 exomes125,748 exomes and 15,708 genomes in v2, expanding to over 700,000 exomes in later versions
Population representationPrimarily European and African American, with smaller East Asian, South Asian, and Latino groupsExpanded Finnish, Ashkenazi Jewish, and additional African, East Asian, and Latino representation
Variant calling approachPer-sample calling followed by aggregationJoint calling across all samples for improved genotype consistency
Genome dataNot includedIncluded, providing coverage of non-coding regions
Site quality metricsBasic quality filtersEnhanced quality metrics including inbreeding coefficient, allele balance, and quality flags
Typical use in filteringHistorical reference for older studies and legacy pipelinesCurrent standard for allele frequency filtering in new analyses

When to Use ExAC Versus gnomAD

The practical question for researchers is which resource to use in a given analysis. In most cases, gnomAD is the preferred choice because it offers a larger sample size, better population diversity, and improved quality metrics. However, there are situations where ExAC remains relevant.

Use gnomAD for New Analyses

For any new variant interpretation project, gnomAD should be the primary population frequency reference. The larger sample size provides more precise frequency estimates, and the improved quality metrics allow users to assess the reliability of each variant site. The inclusion of genome data extends the utility of gnomAD to non-coding variants, which is important for interpreting variants in regulatory regions or deep intronic regions that may affect splicing.

The gnomAD browser provides a user-friendly interface for looking up individual variants, while the VCF files support large-scale analysis. The browser displays allele frequencies across populations, quality metrics, and links to related resources. For researchers who need to query many variants, the VCF files can be used with standard bioinformatics tools for annotation and filtering.

Use ExAC for Legacy Comparisons

ExAC remains useful for comparing results with older studies that used ExAC as their reference. If a published study reported allele frequencies from ExAC, replicating the analysis with ExAC data can help verify results before transitioning to gnomAD. This is particularly relevant for meta-analyses or replication studies that need to maintain consistency with the original methodology.

ExAC data are still available for download, and the ExAC browser remains accessible. However, ExAC is no longer updated, and new variants cannot be added to the dataset. For this reason, ExAC should not be used as the primary reference for new analyses.

Consider the Ancestry of Your Samples

The choice of reference dataset should account for the ancestry of the samples being analyzed. If a study focuses on a population that is underrepresented in gnomAD, the allele frequency estimates may be less reliable for that population. In such cases, researchers should check the population-specific frequencies in gnomAD and consider whether additional population-specific reference data are needed.

A study of LRRK2 variants in Han Chinese patients with Parkinson's disease used whole exome sequencing and identified variants including p.A419V and p.G2385R as potential risk factors for increased PD susceptibility [<a href="#ref-7">7</a>]. The study found that 14.74% of PD cases and 7.24% of controls carried at least one LRRK2 variant, with an odds ratio of 2.2144 [<a href="#ref-7">7</a>]. This type of population-specific analysis depends on accurate allele frequency data for the relevant ancestral group, and researchers should verify that the reference dataset includes adequate representation of that population.

Practical Implementation Steps

Implementing gnomAD in a variant interpretation workflow involves several concrete steps. The following approach can be adapted to different analysis pipelines.

Step 1: Download the Appropriate gnomAD Files

The gnomAD data are available from the gnomAD website in VCF format. The files are organized by genome build and by exome or genome data. For most analyses, the exome VCF is sufficient, but the genome VCF should be used when analyzing non-coding variants.

The files are large, and downloading them requires sufficient storage space and bandwidth. The VCF files are compressed and indexed, allowing tools to access specific regions without loading the entire file. The gnomAD website provides checksums for verifying file integrity after download.

Step 2: Annotate Variants with gnomAD Allele Frequencies

Variant annotation tools can add gnomAD allele frequencies to a VCF file. These tools match each variant to the corresponding entry in the gnomAD VCF and add the allele frequency information to the variant record. The annotation should include the overall allele frequency and the population-specific frequencies for relevant ancestral groups.

The annotation process should also capture the quality metrics from gnomAD, including the quality flags and coverage information. This allows downstream filtering steps to consider data quality in addition to allele frequency.

Step 3: Apply Filtering Thresholds

The filtering threshold should be set based on the disease model and the expected allele frequency of pathogenic variants in the relevant genes. For rare monogenic diseases, a threshold of 1% is commonly used for dominant conditions, while a threshold of 0.1% may be used for recessive conditions where pathogenic variants are expected to be very rare.

The filtering step should also consider the quality flags from gnomAD. Variants with poor site quality should be reviewed manually instead of automatically filtered, as the allele frequency estimate may be unreliable.

Step 4: Document the Filtering Decisions

Reproducibility requires documenting the filtering decisions, including the gnomAD version used, the filtering threshold, and the rationale for the threshold choice. This documentation should be included in the analysis report so that other researchers can understand how the filtering was performed.

The Carpentries lessons provide foundational training in reproducible data analysis practices, including version control and documentation [<a href="#ref-8">8</a>]. These practices are essential for ensuring that variant interpretation workflows can be reproduced and verified.

Records and Measurements for Variant Interpretation

Maintaining accurate records is essential for variant interpretation. The following measurements and records should be captured for each variant analyzed.

Allele Frequency Records

The allele frequency from gnomAD should be recorded for each variant, including the overall frequency and the population-specific frequencies. The gnomAD version and the specific dataset (exome or genome) should also be recorded, as these affect the frequency estimate.

Quality Metric Records

The quality metrics from gnomAD should be recorded for each variant, including the read depth, allele balance, and quality flags. These metrics provide context for assessing the reliability of the allele frequency estimate.

Filtering Decision Records

The filtering decisions should be recorded, including the threshold used and whether the variant passed or failed the filter. The rationale for the threshold choice should be documented, particularly if the threshold differs from standard practice.

Interpretation Records

The final interpretation of each variant should be recorded, including the classification and the evidence supporting the classification. This record should include the allele frequency data, the quality metrics, and any other evidence used in the interpretation.

A case report of hypertrophic cardiomyopathy caused by compound heterozygous ALPK3 mutations described novel variants identified through whole exome sequencing [<a href="#ref-9">9</a>]. The report documented the specific variants, c.4234C>T and c.3491G>A, and their clinical significance [<a href="#ref-9">9</a>]. This type of detailed variant documentation is essential for building the evidence base for variant interpretation.

Common Failure Patterns in Allele Frequency Filtering

Several common errors can undermine the effectiveness of allele frequency filtering. Recognizing these patterns can help researchers avoid them.

Using an Outdated Reference Dataset

Continuing to use ExAC when gnomAD is available can lead to incorrect filtering decisions. The smaller ExAC dataset provides less precise frequency estimates, and the lack of genome data limits the analysis of non-coding variants. Researchers should transition to gnomAD for new analyses and use ExAC only for legacy comparisons.

Ignoring Population-Specific Frequencies

Using only the overall allele frequency from gnomAD can be misleading when the variant frequency differs substantially across populations. A variant that is rare overall but common in a specific population may be filtered incorrectly if the population-specific frequency is not considered. Researchers should check the population-specific frequencies for the relevant ancestral group.

Overlooking Quality Flags

The quality flags in gnomAD indicate sites where the variant call may be unreliable. Ignoring these flags can lead to incorrect frequency estimates being used in filtering decisions. Researchers should review the quality flags for each variant and consider whether the frequency estimate is trustworthy.

Applying a Single Threshold to All Genes

The appropriate allele frequency threshold depends on the gene and the disease. Applying a single threshold to all genes can lead to filtering out pathogenic variants in genes where pathogenic variants are relatively common in the population, or retaining benign variants in genes where pathogenic variants are very rare. The threshold should be adjusted based on the gene-specific context.

Failing to Account for Reduced Penetrance

Some pathogenic variants have reduced penetrance, meaning that not all individuals carrying the variant develop the disease. These variants may be present at higher frequencies in the population than expected for fully penetrant pathogenic variants. Using a strict allele frequency threshold can filter out these variants incorrectly. Researchers should consider the penetrance of the variant when setting the filtering threshold.

Limitations of Population Frequency Data

Population frequency data from gnomAD have several limitations that should be considered in variant interpretation.

Underrepresentation of Certain Populations

Despite the expanded diversity in gnomAD, some populations remain underrepresented. This is particularly true for populations from Africa, the Middle East, and parts of Asia. For variants in these populations, the allele frequency estimates may be less reliable, and the absence of a variant from gnomAD does not necessarily mean the variant is rare in the relevant population.

Incomplete Capture of Rare Variants

Even with hundreds of thousands of samples, gnomAD cannot capture all rare variants in the human population. Many rare variants are unique to specific families or individuals and will not appear in any population database. The absence of a variant from gnomAD should not be interpreted as evidence that the variant is not present in the population.

Sequencing and Alignment Artifacts

The gnomAD data are generated from sequencing data, and sequencing and alignment artifacts can create false variant calls. The quality metrics in gnomAD help identify these artifacts, but some artifacts may pass the quality filters. Researchers should be aware that some variants in gnomAD may be artifacts instead of true population variants.

Differences in Variant Representation

The representation of variants in gnomAD depends on the reference genome build and the variant calling pipeline. Variants may be represented differently in different versions of gnomAD, and researchers should ensure that the variant representation matches the reference genome build used in their analysis.

A study of familial hypercholesterolemia used targeted next-generation sequencing with a 9-gene panel and found that rule-based diagnostic approaches showed inconsistent concordance with molecular confirmation [<a href="#ref-10">10</a>]. The study found that machine learning models showed higher discrimination for reportable pathogenic or likely pathogenic variant positivity than dichotomized rule-based criteria [<a href="#ref-10">10</a>]. This finding illustrates that variant interpretation requires integrating multiple lines of evidence, and population frequency data are only one component of the interpretation process.

Quality Controls for Variant Interpretation

Implementing quality controls in the variant interpretation workflow can reduce errors and improve the reliability of results.

Verify Variant Representation

Before filtering, verify that the variant representation in the analysis matches the representation in gnomAD. This includes checking the reference genome build, the variant type, and the genomic coordinates. Mismatches in representation can lead to variants being missed or incorrectly filtered.

Check Coverage at Variant Sites

The coverage at a variant site affects the confidence in the variant call. Low coverage can lead to false negative calls, while high coverage with imbalanced allele reads can indicate artifacts. The gnomAD quality metrics include coverage information that should be reviewed for each variant.

Use Multiple Lines of Evidence

Population frequency data should not be the sole basis for variant interpretation. Other evidence, including functional studies, segregation data, and computational predictions, should be considered. The interpretation should integrate all available evidence to reach a conclusion.

Document Version Control

The gnomAD version used in the analysis should be documented, as allele frequencies can change between versions. This documentation allows the analysis to be reproduced and updated when new versions of gnomAD are released.

Review Filtering Decisions

Filtering decisions should be reviewed by a second person or through a structured review process. This review can catch errors in threshold selection or variant representation that might otherwise go unnoticed.

Professional Escalation Criteria

There are situations where the standard allele frequency filtering approach is insufficient, and professional guidance should be sought.

Variants in Underrepresented Populations

If a variant is identified in a population that is underrepresented in gnomAD, the allele frequency data may be unreliable. In this case, consultation with a genetic counselor or a specialist in the relevant population genetics should be considered.

Variants Near Filtering Thresholds

If a variant has an allele frequency close to the filtering threshold, the interpretation may be uncertain. The confidence intervals for the frequency estimate should be considered, and the variant may need to be classified as a variant of uncertain significance instead of filtered or retained based on the frequency alone.

Novel Variants in Known Disease Genes

If a novel variant is identified in a known disease gene, the absence of the variant from gnomAD does not establish pathogenicity. The variant should be evaluated using all available evidence, and consultation with a clinical geneticist or molecular pathologist should be considered.

Discrepant Results Across Databases

If the allele frequency for a variant differs substantially between gnomAD and other population databases, the discrepancy should be investigated. The variant representation and quality metrics should be compared across databases to identify the source of the discrepancy.

Variants with Conflicting Interpretations

If different lines of evidence lead to conflicting interpretations of a variant, the case should be escalated for review. This may involve consultation with a variant curation expert or submission of the variant to a clinical variant database.

Safety and Regulatory Context

Variant interpretation has direct implications for patient care, and the use of population frequency data must be conducted within an appropriate regulatory and ethical framework.

Clinical Validation Requirements

Laboratories performing clinical variant interpretation must validate their workflows, including the use of population frequency data. The validation should demonstrate that the filtering thresholds and interpretation criteria produce accurate and reproducible results.

Data Privacy and Security

Population frequency data are derived from large numbers of individuals, and the use of these data must comply with data privacy and security requirements. Researchers should ensure that their use of gnomAD data complies with applicable regulations and ethical guidelines.

Reporting Requirements

Clinical variant interpretation reports should include the evidence supporting the interpretation, including the population frequency data. The report should document the gnomAD version used and the filtering decisions applied.

Professional Oversight

Variant interpretation for clinical purposes should be conducted under the oversight of qualified professionals, including clinical geneticists, molecular pathologists, and genetic counselors. The interpretation should be reviewed by these professionals before being reported to patients or healthcare providers.

A systematic review of next-generation sequencing for the molecular diagnosis of inborn errors of immunity in Brazil found considerable variability among studies in reporting methodological details of NGS workflows, including sequencing platforms, bioinformatics pipelines, and quality control metrics [<a href="#ref-11">11</a>]. This finding highlights the importance of standardized reporting and quality control in clinical sequencing applications [<a href="#ref-11">11</a>].

A Decision Framework for Choosing Filtering Thresholds Across Disease Models

Selecting the correct allele frequency threshold is the most consequential decision in a variant interpretation workflow, yet many laboratories apply a single default value across all genes and disease models. The transition from ExAC to gnomAD changed the statistical basis for these thresholds because the larger dataset provides tighter confidence intervals around frequency estimates. This section presents a practical decision framework that connects the disease model, the gene-specific constraint metrics, and the gnomAD population frequency data to a defensible filtering threshold.

Threshold Selection by Inheritance Pattern

The inheritance pattern of the disease should drive the initial threshold selection. For autosomal dominant disorders, pathogenic variants are typically rare in the general population because they cause disease in heterozygotes and are subject to negative selection. A common starting threshold is 0.1% or 1 in 1,000 alleles, but this must be adjusted based on the penetrance of the condition. For highly penetrant dominant disorders such as familial adenomatous polyposis, a threshold of 0.01% may be more appropriate because truly pathogenic variants are almost never observed in population databases. For dominant disorders with reduced penetrance, such as hereditary breast and ovarian cancer associated with certain BRCA1 and BRCA2 variants, slightly higher thresholds may be acceptable because some pathogenic variants are present at low frequencies in the population.

For autosomal recessive disorders, the threshold logic differs. Pathogenic variants in recessive disease genes can be present at higher frequencies in the population because carriers are unaffected. A threshold of 1% or even 2% may be appropriate for genes where carrier frequencies are known to be elevated in specific populations. The gnomAD dataset provides population-specific frequencies that are essential for this assessment. A variant with a frequency of 1.5% in a specific ancestral group may be a common benign polymorphism in that group, while the same frequency in a different population may warrant further investigation.

Gene-Specific Constraint and Dosage Sensitivity

The gnomAD resource includes gene-level constraint metrics that quantify how depleted a gene is of certain types of variants relative to expectation. The observed over expected ratio for missense variants and the probability of loss-of-function intolerance are particularly useful for threshold selection. Genes with high constraint scores are under strong purifying selection, meaning that damaging variants in these genes are rarely observed in the population. For these genes, a stricter allele frequency threshold is appropriate because even low-frequency variants may be pathogenic.

Conversely, genes with low constraint scores tolerate variation, and higher allele frequency thresholds are appropriate to avoid filtering out benign variants that are common in the population. The constraint metrics should be reviewed for each gene in the analysis panel before setting the filtering threshold. This gene-specific approach reduces the risk of applying a single threshold that is too strict for some genes and too lenient for others.

Population-Specific Frequency Assessment

The gnomAD browser provides allele frequencies for multiple ancestral populations, and the decision framework should incorporate these population-specific values instead of relying solely on the overall frequency. The relevant population for threshold assessment is the population that matches the ancestry of the patient or sample being analyzed. A variant that is absent from the overall gnomAD dataset but present at 0.5% in a specific population may be a common polymorphism in that population and should not be filtered based on the overall frequency alone.

The confidence intervals around population-specific frequencies are wider for smaller population groups within gnomAD. For populations with fewer than 10,000 alleles represented, the frequency estimate should be interpreted with caution. The gnomAD interface displays the allele count and allele number for each population, and these values should be reviewed to assess the reliability of the frequency estimate. A frequency based on 2 observations out of 200 alleles is less reliable than a frequency based on 200 observations out of 20,000 alleles, even if the calculated frequency is similar.

A Stepwise Decision Procedure

The following procedure provides a structured approach to threshold selection that can be documented and reproduced.

First, determine the inheritance pattern of the disease and the mode of action of the gene. This information should be recorded for each gene in the analysis. Second, review the gnomAD constraint metrics for the gene, including the missense observed over expected ratio and the probability of loss-of-function intolerance. Third, identify the ancestral population that matches the sample being analyzed and record the population-specific allele counts and allele numbers from gnomAD. Fourth, select an initial threshold based on the inheritance pattern, using 0.1% for dominant disorders and 1% for recessive disorders as starting points. Fifth, adjust the threshold based on the constraint metrics, using stricter thresholds for highly constrained genes and more permissive thresholds for genes with low constraint. Sixth, review the population-specific frequencies for variants near the threshold and consider whether the confidence intervals overlap the threshold value. Seventh, document the threshold selection rationale for each gene in the analysis protocol.

Recording Threshold Decisions

The threshold selection process should be recorded in a structured format that allows review and audit. For each gene, the record should include the inheritance pattern, the constraint metrics used, the initial threshold, the adjusted threshold, and the rationale for any adjustments. This record should be maintained as part of the laboratory quality management system and reviewed periodically as new versions of gnomAD are released.

The record should also include the gnomAD version used for the threshold assessment. Allele frequencies and constraint metrics can change between gnomAD versions as more samples are added and the processing pipeline is updated. A threshold that was appropriate for gnomAD version 2.1 may need revision when using version 4.0, and the record should reflect the version-specific basis for the threshold decision.

Troubleshooting Threshold Failures

When a variant with an unexpectedly high allele frequency is retained as a candidate pathogenic variant, or when a variant with a very low allele frequency is filtered out, the threshold decision should be reviewed. The following troubleshooting steps can help identify the source of the problem.

First, verify that the variant representation matches between the analysis file and the gnomAD file. Differences in reference genome build, variant normalization, or representation of insertions and deletions can cause a variant to be missed or incorrectly matched in the gnomAD data. Second, check the quality metrics for the variant site in gnomAD. A variant with poor site quality may have an unreliable frequency estimate that should not be used for filtering decisions. Third, review the population-specific frequencies to determine whether the variant is common in a specific ancestral group that is relevant to the sample. Fourth, confirm that the constraint metrics used for the threshold adjustment are from the same gnomAD version as the allele frequency data. Fifth, consider whether the disease model or gene classification has changed since the threshold was set.

Common Threshold Selection Errors

Several recurring errors undermine threshold selection in practice. The most common is applying a single threshold across all genes without considering constraint or inheritance pattern. This approach is simple but leads to both false positives and false negatives. A second common error is using the overall gnomAD frequency without reviewing population-specific values, which can cause variants that are common in a specific population to be retained or filtered incorrectly. A third error is failing to update thresholds when transitioning from ExAC to gnomAD, using thresholds that were calibrated for the smaller ExAC dataset. The larger gnomAD dataset provides more precise frequency estimates, and thresholds should be recalibrated accordingly.

A fourth error is ignoring the confidence intervals around frequency estimates for rare variants. A variant observed once in gnomAD has a wide confidence interval, and the true population frequency could be substantially higher or lower than the point estimate. Threshold decisions should account for this uncertainty, particularly for variants near the threshold boundary. A fifth error is using constraint metrics without understanding their limitations. The constraint metrics are derived from the gnomAD dataset itself and are influenced by the same sequencing and processing artifacts that affect allele frequency estimates.

Integration with Variant Classification

The allele frequency threshold is one component of the variant classification process, and the decision framework should be integrated with the broader classification criteria. A variant that passes the allele frequency filter should still be evaluated using other evidence, including functional studies, segregation data, and computational predictions. Conversely, a variant that fails the allele frequency filter should not be automatically classified as benign without considering other evidence, particularly for genes with reduced penetrance or for variants in underrepresented populations.

The gnomAD dataset also provides information about the number of homozygotes observed for each variant. The presence of homozygotes in gnomAD is relevant for recessive disease assessment because it indicates that the variant is compatible with survival in the homozygous state. For autosomal dominant disorders, the presence of homozygotes may suggest reduced penetrance or a milder phenotype than expected. This information should be incorporated into the classification decision alongside the allele frequency data.

Documentation and Audit Requirements

Laboratories should maintain documentation of the threshold selection process for each gene in their analysis panels. The documentation should include the version of gnomAD used, the constraint metrics reviewed, the population-specific frequencies considered, and the rationale for the final threshold. This documentation supports reproducibility and allows the threshold decisions to be audited by external reviewers or accreditation bodies.

The documentation should also include the date of the threshold review and the name of the individual who performed the review. Threshold decisions should be reviewed whenever a new version of gnomAD is released or when new disease associations are established for a gene. The review should assess whether the existing threshold remains appropriate given the updated data and should document any changes made.

Professional Escalation for Threshold Uncertainty

When the threshold selection process produces an uncertain result, professional guidance should be sought. This includes situations where the constraint metrics are ambiguous, where the population-specific frequencies are based on very small allele numbers, or where the disease model is not well established. Consultation with a clinical geneticist, molecular pathologist, or genetic counselor should be considered for these cases.

The escalation should be documented in the analysis record, including the reason for the escalation and the outcome of the consultation. This documentation supports continuous improvement of the threshold selection process and provides a basis for future decisions involving similar genes or variants.

Frequently Asked Questions

What is the main difference between ExAC and gnomAD?

The main difference is the scale of the reference population. ExAC contained 60,706 exomes, while gnomAD expanded to over 125,000 exomes and 15,000 genomes in its initial release, with later versions exceeding 700,000 exomes. gnomAD also uses joint calling across all samples, which improves genotype consistency, and includes genome data that cover non-coding regions.

Should I use ExAC or gnomAD for my variant analysis?

For new analyses, gnomAD should be the primary population frequency reference. The larger sample size provides more precise frequency estimates, and the improved quality metrics allow users to assess the reliability of each variant site. ExAC should be used only for legacy comparisons with older studies that used ExAC as their reference.

How does the larger sample size in gnomAD affect allele frequency filtering?

The larger sample size provides more precise frequency estimates, particularly for rare variants. A variant observed once in ExAC had a frequency of approximately 1 in 121,000 alleles, while the same observation in a larger gnomAD release represents a lower frequency with tighter confidence intervals. This precision reduces the chance of misclassifying a common variant as rare.

Why does population diversity matter for variant interpretation?

Allele frequencies vary substantially across ancestral groups. A variant that is rare in one population may be common in another. Using a reference dataset that lacks adequate representation of the patient's ancestry can lead to incorrect frequency-based filtering decisions. gnomAD expanded the population representation compared to ExAC, but some populations remain underrepresented.

How do I integrate gnomAD into my variant calling workflow?

Integration involves downloading the appropriate gnomAD VCF files, annotating variants with gnomAD allele frequencies, applying filtering thresholds based on the disease model, and documenting the filtering decisions. The nf-core documentation provides standards for community pipelines that incorporate population frequency filtering [<a href="#ref-4">4</a>], and the Galaxy Training Network offers accessible tutorials for variant calling workflows [<a href="#ref-5">5</a>].

What are the quality metrics in gnomAD and why are they important?

gnomAD provides quality metrics for each variant site, including inbreeding coefficient, allele balance, and read depth. These metrics are used to flag sites where the variant call may be unreliable. Reviewing these metrics is important because a variant with a high allele frequency but poor site quality may be an artifact of sequencing or mapping error instead of a true population variant.

Can I use gnomAD for somatic variant calling?

Yes, gnomAD is used in somatic variant calling to filter out germline polymorphisms that are captured in tumor sequencing data. A variant present at high frequency in gnomAD is likely a germline variant instead of a somatic mutation. However, somatic callers must be careful not to filter out true somatic variants that happen to match a low-frequency population variant.

What should I do if a variant is absent from gnomAD?

The absence of a variant from gnomAD does not establish pathogenicity. Many rare variants are unique to specific families or individuals and will not appear in any population database. The variant should be evaluated using all available evidence, including functional studies, segregation data, and computational predictions, and consultation with a clinical geneticist or molecular pathologist should be considered for novel variants in known disease genes.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

[1] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information. [2] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute. [3] [Germline variant spectrum of hereditary cancer susceptibility genes in a Chinese cohort.](https://doi.org/10.1186/s40246-025-00873-z). 2025. [4] [nf-core Documentation](https://nf-co.re/docs). nf-core. [5] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project. [6] [Bioconductor](https://bioconductor.org/). Bioconductor Project. [7] [Genetic analysis of LRRK2 variants in Han Chinese patients with Parkinson's disease.](https://doi.org/10.1371/journal.pone.0340448). 2026. [8] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries. [9] [Novel compound heterozygous <,i>,ALPK3<,/i>, mutations (c.4234C>,T and c.3491G>,A), causing hypertrophic cardiomyopathy treated with the liwen procedure: case report.](https://doi.org/10.3389/fcvm.2025.1671882). 2025. [10] [Integrated Clinical, Molecular, and Machine Learning Assessment of Familial Hypercholesterolemia.](https://doi.org/10.3390/life16040633). 2026. [11] [Overview of next-generation sequencing to the molecular diagnosis of inborn errors of immunity in Brazil: a systematic review.](https://doi.org/10.3389/fimmu.2026.1794921). 2026.

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.