A Practical Guide to Using gnomAD for Variant Interpretation in Clinical Research: From Allele Frequency to ACMG Criteria

By Dr. Zubair Khalid, DVM, MS, PhD ·

A Practical Guide to Using gnomAD for Variant Interpretation in Clinical Research: From Allele Frequency to ACMG Criteria

Key Takeaways

  • gnomAD's population allele frequency data is critical for variant interpretation, primarily informing ACMG/AMP evidence codes PM2 (moderate pathogenic), BA1 (stand-alone benign), and BS1 (moderate benign).
  • Selecting the correct gnomAD population subset, typically matching the patient's genetic ancestry, is paramount, as a variant's frequency can differ significantly across groups, impacting pathogenicity assessment.
  • PM2 requires a variant to be absent or at extremely low frequency in gnomAD, with the threshold being disorder-specific and influenced by disease prevalence and penetrance, while BA1 applies when allele frequency substantially exceeds the maximum credible frequency for a rare Mendelian disorder.
  • BS1 is applied when a variant's allele frequency is higher than expected for a specific disorder but does not meet the BA1 threshold, necessitating combination with other benign evidence.
  • Reproducible interpretation mandates meticulous record-keeping, including the gnomAD version, specific population subset used, allele counts, and the rationale for disorder-specific frequency thresholds.
  • Limitations of gnomAD, such as potential selection bias and underrepresentation of certain populations, must be considered, and interpretation should always integrate allele frequency data with computational predictions, functional assays, and segregation data.

Clinical researchers face a recurring problem when interpreting sequence variants: a variant is observed in a patient, but its clinical significance is unclear. The Genome Aggregation Database (gnomAD) provides population allele frequency data that can help resolve this uncertainty, but applying that data correctly requires understanding which population subset to use, how to interpret frequency thresholds, and how those frequencies map to specific ACMG/AMP criteria. This article provides a practical walkthrough of using gnomAD for the PM2, BA1, and BS1 criteria, including how to select appropriate population subsets and interpret allele frequencies in a clinical research context.

The intended reader is a biology student, researcher, laboratory professional, or life-science practitioner who has access to variant call data and needs to apply gnomAD correctly in variant interpretation workflows. The focus is on germline variant interpretation in clinical research settings, with attention to the practical decisions that determine whether an allele frequency observation supports pathogenicity or benignity.

The Role of Population Frequency Data in Variant Interpretation

Population databases serve as reference points for how common or rare a variant is in the general population. The underlying logic is straightforward: a variant that causes a severe, early-onset, or highly penetrant Mendelian disease should be rare in the general population because affected individuals are less likely to reproduce or because the variant is under purifying selection. Conversely, a variant that is common in the general population is unlikely to be a high-penetrance cause of a rare disease.

gnomAD aggregates exome and genome sequencing data from hundreds of thousands of individuals across diverse global populations. The resource provides allele frequencies for single-nucleotide variants, small insertions and deletions, and structural variants. For clinical interpretation, the key columns are the overall allele frequency, the allele frequency within specific genetic ancestry groups, and the number of alleles observed at each site. The structural variant reference from gnomAD, built from nearly 15,000 high-coverage genomes, demonstrates that structural variants are responsible for a substantial fraction of rare protein-truncating events per genome, which means population frequency data for structural variants is also relevant to clinical interpretation workflows.

The practical implication for a researcher is that gnomAD is not a single number. It is a set of frequency estimates stratified by population, and the choice of which estimate to use changes the interpretation. A variant that is absent from the overall database but present at low frequency in one ancestry group may warrant different evidence weighting than a variant that is truly absent from all groups.

The ACMG/AMP Framework and Where Allele Frequency Fits

The American College of Medical Genetics and Genomics and the Association for Molecular Pathology published a framework for classifying sequence variants into five categories: pathogenic, likely pathogenic, uncertain significance, likely benign, and benign. The framework uses evidence types that combine to support a classification. Allele frequency data contributes primarily to three evidence codes: PM2, BA1, and BS1.

PM2 is a moderate evidence code for pathogenicity. It applies when a variant is absent in a large population database or is present only at an extremely low frequency. The rationale is that a truly pathogenic variant for a rare Mendelian disorder should not be observed at appreciable frequency in the general population. The original ACMG/AMP guidance did not specify an exact frequency threshold for PM2, which has led to inconsistent application across laboratories.

BA1 is a stand-alone evidence code for benignity. It applies when the allele frequency of a variant is too high to be consistent with a pathogenic role in a rare Mendelian disease. The original guidance suggested a threshold of greater than 5 percent allele frequency, but this threshold was intended for autosomal dominant disorders with high penetrance and may not be appropriate for all disease mechanisms.

BS1 is a moderate evidence code for benignity. It applies when the allele frequency of a variant is higher than expected for the disorder. The expected frequency depends on the disease prevalence, penetrance, and inheritance pattern. BS1 requires a disorder-specific calculation instead of a universal threshold.

The challenge for clinical researchers is that these criteria require judgment about what constitutes a large database, what counts as an extremely low frequency, and what frequency is inconsistent with pathogenicity for a specific disorder. The sections below provide a practical framework for making those judgments with gnomAD data.

Selecting the Appropriate gnomAD Population Subset

The choice of population subset is one of the most consequential decisions in applying gnomAD data to variant interpretation. Using the overall allele frequency can mask population-specific variation. A variant that is rare globally but common in one ancestry group may be a benign polymorphism in that group, even if it is absent from other groups. Conversely, a variant that is absent from the overall database but present in a small number of individuals from one ancestry group may still warrant PM2 if the disease is very rare and the variant is predicted to be damaging.

The practical approach is to examine the allele frequency in the population subset that most closely matches the patient's genetic ancestry, while also considering the overall frequency. For a variant observed in a patient of European ancestry, the non-Finnish European subset is the most relevant comparison. For a patient of East Asian ancestry, the East Asian subset is the primary reference. The gnomAD browser displays these subsets as separate columns, and the researcher should record both the subset frequency and the overall frequency.

A second consideration is the number of alleles observed at the site. A variant observed in one allele out of 10,000 alleles in a population subset is a different observation than a variant observed in 50 alleles out of 10,000. The allele count provides context for whether the frequency estimate is stable or based on a single observation. For PM2, a variant that is truly absent from all populations is stronger evidence than a variant that is present in a single allele, even if the frequency is very low.

A third consideration is the gnomAD version. The database is updated periodically, and allele frequencies change as more individuals are added. A variant that was absent in an earlier version may be present at low frequency in a later version. The researcher should record the gnomAD version used for interpretation, because this affects reproducibility and re-evaluation. The PTPN1 case series provides an example of this practice: the authors specified that variants had a frequency on gnomAD version 4.1.0 of less than 1.25 times 10 to the negative sixth, which allowed other researchers to reproduce their filtering criteria.

At a Glance: gnomAD Evidence Codes and Application Decisions

Evidence CodeEvidence StrengthPrimary QuestionTypical gnomAD ObservationDecision Context
PM2Moderate pathogenicIs the variant absent or extremely rare in the general population?Absent from all populations or present below a disorder-specific thresholdMust be combined with other pathogenic evidence, not sufficient alone
BA1Stand-alone benignIs the variant far too common to cause a rare Mendelian disorder?Allele frequency substantially exceeds the maximum credible allele frequencyCan classify a variant as benign without additional evidence
BS1Moderate benignIs the variant more common than expected for the specific disorder?Allele frequency exceeds the disorder-specific expected frequencyMust be combined with other benign evidence for final classification

Applying PM2: Absence or Extreme Rarity

PM2 is applied when a variant is absent from large population databases or present at an extremely low frequency. The practical question is what counts as a large database and what counts as extremely low.

For a large database, gnomAD currently contains hundreds of thousands of alleles, which provides substantial power to detect variants that are present at even very low frequencies. A variant that is truly absent from gnomAD has a very low upper bound on its population frequency. For example, if a variant is absent from 200,000 alleles, the upper bound of the 95 percent confidence interval for the population frequency is approximately 1.5 per 100,000 alleles. This calculation assumes the variant is not present in the database at all.

For an extremely low frequency, the threshold depends on the disease. For a rare autosomal recessive disorder with a prevalence of 1 in 100,000, the carrier frequency is approximately 1 in 158, which corresponds to an allele frequency of approximately 0.3 percent. A variant present at 0.3 percent in the general population could be a pathogenic founder variant for that disorder, so PM2 would not apply. For a very rare autosomal dominant disorder with a prevalence of 1 in 1,000,000, the allele frequency of pathogenic variants should be well below 0.01 percent, and a variant present at 0.1 percent would be inconsistent with pathogenicity.

The practical workflow for PM2 is as follows. First, confirm that the variant is absent from gnomAD or present at a frequency below the disorder-specific threshold. Second, record the total number of alleles examined in the relevant population subset. Third, record the gnomAD version. Fourth, document the disorder-specific threshold used to define extremely low frequency. Fifth, apply PM2 only if the variant is also supported by other evidence for pathogenicity, such as a predicted loss-of-function mechanism or a previously reported pathogenic variant in the same gene.

A common error is applying PM2 to every variant that is absent from gnomAD, regardless of the disease mechanism. For a gene with a high rate of benign loss-of-function variation, absence from gnomAD may not be informative. The gnomAD constraint metrics, such as the probability of loss-of-function intolerance, provide context for whether a gene is sensitive to loss-of-function variation. A variant in a gene that is highly intolerant to loss-of-function variation is more likely to be pathogenic when absent from gnomAD than a variant in a gene that tolerates loss-of-function variation.

Applying BA1: Stand-Alone Benign Evidence

BA1 is the strongest benign evidence code because it can stand alone to classify a variant as benign. The original ACMG/AMP guidance suggested a threshold of greater than 5 percent allele frequency, but this threshold was designed for rare Mendelian disorders with high penetrance. For disorders with lower penetrance or later onset, a variant present at 5 percent could still be pathogenic.

The practical approach is to calculate the maximum credible allele frequency for the disorder and compare it to the gnomAD allele frequency. If the gnomAD allele frequency exceeds the maximum credible allele frequency by a substantial margin, BA1 applies. The maximum credible allele frequency is calculated from the disease prevalence, the inheritance pattern, the penetrance, and the genetic heterogeneity.

For an autosomal dominant disorder with a prevalence of 1 in 10,000, full penetrance, and no genetic heterogeneity, the maximum credible allele frequency is approximately 0.005 percent. A variant present at 5 percent in gnomAD is 1,000 times more common than expected and would clearly meet BA1. For an autosomal recessive disorder with a prevalence of 1 in 10,000, the maximum credible allele frequency is approximately 1 percent, and a variant present at 5 percent would still exceed the threshold but by a smaller margin.

The practical workflow for BA1 is as follows. First, determine the disease prevalence, inheritance pattern, penetrance, and genetic heterogeneity for the disorder. Second, calculate the maximum credible allele frequency. Third, compare the gnomAD allele frequency in the relevant population subset to the maximum credible allele frequency. Fourth, apply BA1 only if the gnomAD allele frequency substantially exceeds the maximum credible allele frequency. Fifth, document the calculation so that the interpretation can be reviewed.

A common error is applying BA1 without considering the disease mechanism. For a disorder with reduced penetrance, a variant present at 5 percent could still be pathogenic. For a disorder with late onset, a variant present at 5 percent could be pathogenic because affected individuals may reproduce before disease onset. The BA1 threshold should be adjusted for these factors.

Applying BS1: Moderate Benign Evidence

BS1 is applied when the allele frequency of a variant is higher than expected for the disorder, but not high enough to meet the BA1 threshold. BS1 is a moderate evidence code, meaning it must be combined with other evidence to classify a variant as likely benign or benign.

The calculation for BS1 is similar to BA1, but the threshold is less stringent. The variant allele frequency in gnomAD should exceed the maximum credible allele frequency for the disorder, but the margin does not need to be as large. The original ACMG/AMP guidance suggested that BS1 applies when the allele frequency is greater than expected for the disorder, without specifying a precise margin.

The practical workflow for BS1 is as follows. First, calculate the maximum credible allele frequency for the disorder using the same approach as for BA1. Second, compare the gnomAD allele frequency to the maximum credible allele frequency. Third, apply BS1 if the gnomAD allele frequency exceeds the maximum credible allele frequency. Fourth, combine BS1 with other benign evidence, such as a lack of segregation with disease in affected family members or a benign functional assay result.

A common error is applying BS1 without considering the population subset. A variant that is common in one ancestry group but rare in another may be a benign polymorphism in the first group but a candidate pathogenic variant in the second. The researcher should examine the allele frequency in the population subset that matches the patient's ancestry and consider whether the variant is present across multiple populations or restricted to one.

Quantitative Customization of ACMG/AMP Criteria

The standard ACMG/AMP framework was designed for general use, but disease-specific customization can improve diagnostic yield. A study of inherited arrhythmias demonstrated this approach by comparing rare variant frequencies from large case cohorts to population-specific gnomAD data and developing disease-specific criteria for the PM2 and BS1 rules. The study found that stringent application of the standard guidelines led to high rates of variants of uncertain significance for genetically heterogeneous diseases such as long QT syndrome and Brugada syndrome.

The customization involved two steps. First, the researchers calculated the expected frequency of pathogenic variants in the general population based on disease prevalence and penetrance. Second, they compared the observed frequency of rare variants in case cohorts to the expected frequency and adjusted the PM2 and BS1 thresholds accordingly. This quantitative approach reduced the rate of variants of uncertain significance and increased the diagnostic yield.

The practical implication for a clinical researcher is that the standard ACMG/AMP thresholds may not be appropriate for all disorders. For a disorder with high genetic heterogeneity, the expected frequency of any single pathogenic variant is lower, which means the PM2 threshold should be more stringent. For a disorder with reduced penetrance, the expected frequency of pathogenic variants is higher, which means the BS1 threshold should be less stringent.

The quantitative approach requires access to case cohorts with known pathogenic variants. For a researcher without access to such cohorts, the alternative is to use published disease-specific criteria or to apply the standard thresholds with careful documentation of the limitations.

Using gnomAD for Non-Coding Variants

The interpretation of non-coding variants presents a distinct challenge because the standard bioinformatic prediction algorithms that assess effects on protein function are not applicable. A study of RNU4ATAC, a non-coding gene transcribed into a spliceosomal RNA component, illustrated this challenge. The gene has a single non-coding exon, so prediction algorithms that assess splicing or protein function are irrelevant, making variant interpretation challenging for molecular diagnostic laboratories.

The study used gnomAD to analyze genetic variation affecting RNU4ATAC and compared the pathogenicity prediction performances of an RNA structure prediction tool and the Combined Annotation Dependent Depletion tool. The study also developed a cellular assay to measure the effect of variants on splicing efficiency of a minor intron.

The practical implication for a clinical researcher is that gnomAD allele frequency data is useful for non-coding variants, but the interpretation framework differs. For a non-coding variant that is absent from gnomAD, PM2 may apply, but the evidence for pathogenicity must come from functional assays or RNA analysis instead of protein prediction algorithms. For a non-coding variant that is common in gnomAD, BA1 or BS1 may apply, but the researcher should consider whether the variant affects a regulatory element that is under selective constraint.

The gnomAD structural variant reference also provides context for non-coding variants. The reference demonstrated modest selection against non-coding structural variants in cis-regulatory elements, which suggests that some non-coding variants are under selective pressure. A non-coding variant that disrupts a cis-regulatory element and is rare in gnomAD may warrant further investigation.

Integrating gnomAD with Other Evidence Types

Allele frequency data is one component of variant interpretation, and it must be integrated with other evidence types to reach a classification. The ACMG/AMP framework includes evidence from population data, computational predictions, functional assays, segregation data, and case-control studies. The final classification depends on the combination of evidence, not on any single observation.

The practical workflow for integrating gnomAD data with other evidence is as follows. First, collect all available evidence for the variant, including allele frequency, computational predictions, functional data, segregation data, and case reports. Second, assign evidence codes based on the strength of each observation. Third, combine the evidence codes according to the ACMG/AMP rules to reach a classification. Fourth, document the evidence and the classification in the variant report.

A common error is over-weighting allele frequency data. A variant that is absent from gnomAD and predicted to be damaging by multiple computational tools may still be a benign variant in a gene that is not relevant to the patient's phenotype. Conversely, a variant that is present at low frequency in gnomAD may be pathogenic if it is a founder variant in a specific population or if it has reduced penetrance.

The quantitative customization study provides an example of how to integrate allele frequency data with case enrichment data. The researchers used the excess of rare variants in case cohorts compared to gnomAD to assign the PS4 evidence code, which applies when the prevalence of a variant in affected individuals is significantly increased compared to controls. This approach requires access to case cohorts, but it demonstrates the value of combining population data with disease-specific data.

Practical Workflow for gnomAD-Based Variant Interpretation

The following workflow provides a step-by-step approach to using gnomAD for variant interpretation in a clinical research setting. The workflow assumes the researcher has a variant call file or a list of candidate variants and needs to apply ACMG/AMP criteria.

Step 1: Confirm the variant coordinates and reference genome build. gnomAD uses specific genome builds, and the researcher must confirm that the variant coordinates match the build used by gnomAD. A mismatch in genome build will produce incorrect allele frequency data.

Step 2: Query gnomAD for the variant. The gnomAD browser accepts variant coordinates, rsIDs, or gene names. The researcher should record the gnomAD version, the overall allele frequency, the allele frequency in each population subset, and the number of alleles observed at the site.

Step 3: Select the relevant population subset. The choice depends on the patient's genetic ancestry and the disease mechanism. For a variant in a gene associated with a disorder that has different prevalence across populations, the researcher should examine the allele frequency in the population that matches the patient's ancestry.

Step 4: Calculate the maximum credible allele frequency for the disorder. This calculation requires the disease prevalence, inheritance pattern, penetrance, and genetic heterogeneity. The researcher should document the calculation and the sources for each parameter.

Step 5: Apply the relevant evidence codes. If the variant is absent from gnomAD or present at a frequency below the disorder-specific threshold, apply PM2. If the variant is present at a frequency that substantially exceeds the maximum credible allele frequency, apply BA1. If the variant is present at a frequency that exceeds the maximum credible allele frequency but does not meet the BA1 threshold, apply BS1.

Step 6: Integrate the allele frequency evidence with other evidence types. The researcher should collect computational predictions, functional data, segregation data, and case reports for the variant and assign evidence codes based on the strength of each observation.

Step 7: Combine the evidence codes according to the ACMG/AMP rules to reach a classification. The researcher should document the evidence and the classification in the variant report.

Step 8: Record the interpretation in a variant database or spreadsheet. The record should include the variant coordinates, the gnomAD version, the allele frequencies, the evidence codes applied, and the final classification.

Records and Measurements for Reproducible Interpretation

Reproducibility requires that another researcher can repeat the interpretation and reach the same conclusion. The following records should be maintained for each variant interpretation.

The variant record should include the genomic coordinates, the reference and alternate alleles, the gene name, and the transcript identifier. The gnomAD record should include the version number, the overall allele frequency, the allele frequency in each population subset, and the number of alleles observed at the site. The evidence record should include the evidence codes applied, the rationale for each code, and the sources for any disorder-specific parameters used in the calculation.

The interpretation record should include the final classification, the date of the interpretation, and the name of the researcher who performed the interpretation. If the interpretation is re-evaluated after a gnomAD update, the record should include the previous interpretation and the reason for the change.

A practical approach is to use a spreadsheet or a variant interpretation database with columns for each data element. The spreadsheet should include a column for the gnomAD version, because allele frequencies change between versions and the version is essential for reproducibility. The spreadsheet should also include a column for the population subset used for the interpretation, because the choice of subset affects the conclusion.

Common Failure Patterns in gnomAD-Based Interpretation

Several recurring errors lead to incorrect variant interpretations when using gnomAD data. Recognizing these patterns can help the researcher avoid them.

The first failure pattern is using the overall allele frequency without examining population subsets. This error can lead to a variant being classified as benign when it is actually rare in the patient's ancestry group, or classified as pathogenic when it is common in the patient's ancestry group. The researcher should always examine the allele frequency in the population subset that matches the patient's ancestry.

The second failure pattern is applying PM2 to every variant that is absent from gnomAD. PM2 is moderate evidence for pathogenicity, and it should be applied only when the variant is also supported by other evidence. A variant that is absent from gnomAD but has no other evidence for pathogenicity should not be classified as pathogenic based on PM2 alone.

The third failure pattern is applying BA1 or BS1 without calculating the maximum credible allele frequency for the disorder. The original ACMG/AMP threshold of 5 percent for BA1 was designed for rare Mendelian disorders with high penetrance, and it may not be appropriate for all disorders. The researcher should calculate the maximum credible allele frequency and compare it to the gnomAD allele frequency.

The fourth failure pattern is using an outdated gnomAD version without noting the version in the record. Allele frequencies change between versions, and an interpretation based on an outdated version may be incorrect. The researcher should record the gnomAD version and re-evaluate the interpretation when a new version is released.

The fifth failure pattern is ignoring the number of alleles observed at the site. A variant observed in one allele out of 10,000 is a different observation than a variant observed in 50 alleles out of 10,000. The allele count provides context for whether the frequency estimate is stable, and the researcher should record the allele count along with the frequency.

The sixth failure pattern is applying the same thresholds to all disease mechanisms. A variant in a gene associated with an autosomal recessive disorder has a different expected frequency than a variant in a gene associated with an autosomal dominant disorder. The researcher should adjust the thresholds based on the inheritance pattern and the disease prevalence.

Limitations of gnomAD Data

gnomAD is a powerful resource, but it has limitations that affect variant interpretation. The researcher should be aware of these limitations and document them in the variant report.

The first limitation is that gnomAD is not a random sample of the general population. The individuals in gnomAD were recruited for various studies, and some may have been selected for specific phenotypes. This selection bias can affect allele frequencies, particularly for variants associated with common diseases.

The second limitation is that gnomAD includes individuals with pediatric diseases. The gnomAD cohort includes some individuals who were recruited for pediatric disease studies, which means that variants causing severe pediatric diseases may be underrepresented. This limitation is relevant for PM2, because a variant that is absent from gnomAD may still be present in the general population at a frequency that is too high for pathogenicity.

The third limitation is that gnomAD does not include all human populations. The database has substantial representation from European, African, East Asian, South Asian, and Latino populations, but other populations are underrepresented. A variant that is common in an underrepresented population may appear to be rare in gnomAD, leading to an incorrect PM2 application.

The fourth limitation is that gnomAD allele frequencies are based on the number of alleles observed at each site, and some sites have low coverage. A variant at a site with low coverage may be missed, leading to an incorrect absence call. The researcher should check the coverage at the site before concluding that a variant is absent.

The fifth limitation is that gnomAD includes structural variants, but the structural variant reference is based on a smaller number of genomes than the single-nucleotide variant reference. The structural variant reference was constructed from approximately 15,000 genomes, which is smaller than the exome and genome cohorts used for single-nucleotide variants. The researcher should be cautious when using gnomAD structural variant frequencies for interpretation.

The sixth limitation is that gnomAD does not provide phenotype data for the individuals in the database. The researcher cannot determine whether a variant observed in gnomAD is present in an affected or unaffected individual. This limitation is relevant for BA1 and BS1, because a variant that is common in gnomAD may still be pathogenic if it has reduced penetrance.

Professional Escalation Criteria

The researcher should escalate a variant interpretation to a more experienced colleague or a molecular diagnostic laboratory when the interpretation has significant clinical consequences and the evidence is ambiguous. The following criteria indicate that escalation is appropriate.

Escalate when the variant is in a gene associated with a disorder for which the researcher does not have disease-specific expertise. The calculation of the maximum credible allele frequency requires knowledge of disease prevalence, penetrance, and genetic heterogeneity, and an incorrect calculation can lead to an incorrect classification.

Escalate when the variant is in a gene with high genetic heterogeneity and the researcher cannot determine whether the variant is a known cause of the disorder. The quantitative customization study demonstrated that stringent application of standard guidelines can lead to high rates of variants of uncertain significance for genetically heterogeneous diseases, and disease-specific criteria may be required.

Escalate when the variant is in a non-coding region and the researcher cannot determine the functional effect. The RNU4ATAC study demonstrated that non-coding variants require functional assays for interpretation, and the researcher should escalate to a laboratory with the capacity to perform such assays.

Escalate when the variant is a structural variant and the researcher is uncertain about the interpretation. The gnomAD structural variant reference provides population frequency data, but the interpretation of structural variants requires specialized expertise.

Escalate when the variant is in a gene associated with a disorder that has medicolegal implications, such as a disorder with predictive testing or prenatal testing implications. The researcher should ensure that the interpretation meets the standards of the relevant regulatory body.

Escalate when the researcher cannot reproduce the interpretation using the recorded data. If the records are incomplete or the interpretation cannot be reproduced, the researcher should escalate to a colleague for review.

Relevant Welfare and Safety Context

Variant interpretation in clinical research has direct implications for patient care, and the researcher has a responsibility to ensure that the interpretation is accurate and appropriately communicated. The following considerations are relevant to the welfare and safety of patients and research participants.

The researcher should ensure that the variant interpretation is performed in accordance with the standards of the relevant regulatory body. In many jurisdictions, clinical variant interpretation is regulated, and the researcher should confirm that the interpretation is performed by a qualified professional.

The researcher should ensure that the variant interpretation is communicated to the patient or research participant in a way that is understandable and appropriate. The interpretation should include the evidence supporting the classification and the limitations of the evidence.

The researcher should ensure that the variant interpretation is re-evaluated when new evidence becomes available. gnomAD is updated periodically, and new disease-specific data may change the interpretation. The researcher should have a process for re-evaluating interpretations when new data is released.

The researcher should ensure that the variant interpretation is stored in a secure and accessible format. The interpretation should be recorded in a way that allows other researchers to review the evidence and reproduce the classification.

The researcher should ensure that the variant interpretation is not overinterpreted. A variant classified as pathogenic based on limited evidence may be reclassified as benign when more data becomes available. The researcher should communicate the uncertainty in the interpretation to the patient or research participant.

Frequently Asked Questions

What is the difference between PM2, BA1, and BS1 in the ACMG/AMP framework?

PM2 is a moderate evidence code for pathogenicity that applies when a variant is absent from large population databases or present at an extremely low frequency. BA1 is a stand-alone evidence code for benignity that applies when the allele frequency of a variant is too high to be consistent with a pathogenic role in a rare Mendelian disease. BS1 is a moderate evidence code for benignity that applies when the allele frequency of a variant is higher than expected for the disorder. The key difference is that BA1 can stand alone to classify a variant as benign, while PM2 and BS1 must be combined with other evidence.

How do I choose the appropriate gnomAD population subset for variant interpretation?

The choice of population subset depends on the patient's genetic ancestry and the disease mechanism. The researcher should examine the allele frequency in the population subset that most closely matches the patient's ancestry, while also considering the overall frequency. A variant that is rare globally but common in one ancestry group may be a benign polymorphism in that group, even if it is absent from other groups. The researcher should record both the subset frequency and the overall frequency.

What is the maximum credible allele frequency and how do I calculate it?

The maximum credible allele frequency is the highest frequency at which a pathogenic variant could exist in the general population for a specific disorder. It is calculated from the disease prevalence, the inheritance pattern, the penetrance, and the genetic heterogeneity. For an autosomal dominant disorder with a prevalence of 1 in 10,000, full penetrance, and no genetic heterogeneity, the maximum credible allele frequency is approximately 0.005 percent. The researcher should document the calculation and the sources for each parameter.

Can I apply PM2 to a variant that is absent from gnomAD but has no other evidence for pathogenicity?

PM2 is moderate evidence for pathogenicity, and it should be applied only when the variant is also supported by other evidence. A variant that is absent from gnomAD but has no other evidence for pathogenicity should not be classified as pathogenic based on PM2 alone. The researcher should collect computational predictions, functional data, segregation data, and case reports for the variant before applying PM2.

How do I interpret allele frequency data for non-coding variants?

The interpretation of non-coding variants is challenging because the standard bioinformatic prediction algorithms that assess effects on protein function are not applicable. gnomAD allele frequency data is useful for non-coding variants, but the evidence for pathogenicity must come from functional assays or RNA analysis instead of protein prediction algorithms. The researcher should consider whether the variant affects a regulatory element that is under selective constraint.

What should I do if the gnomAD allele frequency changes between versions?

The researcher should record the gnomAD version used for the interpretation and re-evaluate the interpretation when a new version is released. Allele frequencies change between versions as more individuals are added to the database, and a variant that was absent in an earlier version may be present at low frequency in a later version. The researcher should have a process for re-evaluating interpretations when new data is released.

How do I calculate the maximum credible allele frequency for a disorder with reduced penetrance?

For a disorder with reduced penetrance, the maximum credible allele frequency is higher than for a disorder with full penetrance, because affected individuals may not be under strong selective pressure. The calculation should include the penetrance as a parameter, and the researcher should document the source for the penetrance estimate. The quantitative customization study demonstrated that disease-specific criteria can improve diagnostic yield for disorders with reduced penetrance.

When should I escalate a variant interpretation to a more experienced colleague?

The researcher should escalate a variant interpretation when the variant is in a gene associated with a disorder for which the researcher does not have disease-specific expertise, when the variant is in a gene with high genetic heterogeneity, when the variant is in a non-coding region and the functional effect is unknown, when the variant is a structural variant, or when the variant is in a gene associated with a disorder that has medicolegal implications. The researcher should also escalate when the interpretation cannot be reproduced using the recorded data.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.