Applying the ACMG/AMP Guidelines for Variant Classification: A Step-by-Step Framework for Clinical Laboratories
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- The ACMG/AMP guidelines provide a five-tier classification system (pathogenic, likely pathogenic, VUS, likely benign, benign) for germline sequence variants in Mendelian disease genes, requiring systematic application of evidence codes with defined strength levels (very strong, strong, moderate, supporting).
- Variant classification necessitates a structured workflow including confirming variant identity and annotation, determining variant type and predicted consequence, querying population and disease databases (e.g., ClinVar, locus-specific databases), applying evidence codes, and determining the final classification based on code combinations.
- Common interpretation errors include overcounting related evidence (e.g., null variant and functional loss of function from the same observation), applying criteria without gene-specific context (e.g., null variant in a gain-of-function gene), ignoring database version changes, treating computational predictions as definitive, and failing to periodically reclassify Variants of Uncertain Significance (VUS).
- Gene-specific modifications and expert panel criteria (e.g., ClinGen VCEPs) are crucial for genes with unique challenges, as the standard ACMG/AMP framework may not adequately address all variant types or disease mechanisms.
- Meticulous record-keeping is paramount, requiring documentation of variant identification, evidence collection (databases queried, versions, dates), classification decisions (codes applied, rationale), and reclassification reviews to ensure reproducibility and enable external audit.
- Quality control measures such as dual review, interpretation policy compliance audits, defined escalation criteria for complex cases, and external consultation are essential for professional oversight and consistent application of the ACMG/AMP framework.
Clinical laboratories face a recurring problem: how to convert raw variant observations into classifications that clinicians can use for diagnosis and management. The American College of Medical Genetics and Genomics and the Association for Molecular Pathology published a joint framework in 2015 that defines five classification tiers: pathogenic, likely pathogenic, variant of uncertain significance (VUS), likely benign, and benign. This article provides a practical walkthrough for laboratory scientists who need to apply those criteria systematically, combine evidence codes correctly, and avoid common interpretation errors. The framework applies to germline sequence variants identified through diagnostic testing, and the workflow described here assumes you have already completed variant calling and filtering steps that produce a high-confidence variant list.
Scope and Laboratory Context
The ACMG/AMP framework was designed for germline sequence variants in Mendelian disease genes. It does not directly address somatic variants, copy number variants, or variants of unknown clinical relevance in pharmacogenomic contexts. Laboratories that classify somatic variants or structural variants need separate adaptation protocols. For example, researchers have developed CNV-specific modifications of the ACMG/AMP criteria because the original guidelines address copy number interpretation only sparingly, and three diagnostic laboratories independently tested these modified criteria to improve concordance for single-gene CNVs 10.
Your laboratory should establish a written variant interpretation policy before you begin classifying individual variants. That policy should specify which genes you will interpret, which evidence sources you will consult, how you will handle conflicting evidence, and who has authority to resolve disagreements. The policy also needs a version number and a review schedule, because the evidence base for variant interpretation changes continuously.
At a Glance: The Five-Tier Classification System
The table below summarizes the five classification tiers, the typical evidence weight required, and the clinical action that usually follows each classification.
| Classification Tier | Evidence Standard | Typical Clinical Consequence |
|---|---|---|
| Pathogenic | Very strong evidence for pathogenicity, often with multiple supporting lines of evidence | Used for diagnostic confirmation, carrier testing, and predictive testing of at-risk relatives |
| Likely pathogenic | Strong evidence for pathogenicity but not conclusive | May guide clinical management while additional family or functional studies are pursued |
| Variant of uncertain significance | Insufficient or conflicting evidence | No change in clinical management based on this variant alone, periodic reclassification review required |
| Likely benign | Strong evidence against pathogenicity but not conclusive | May be reported but should not drive clinical decisions |
| Benign | Very strong evidence against pathogenicity | May be excluded from clinical consideration and often not reported |
The classification of a variant can change over time. The ALPL gene variant classification project was established specifically to reclassify VUS and continuously assess genetic, phenotypic, and functional variant information in an open-access database, using a multi-step process that adheres to the ACMG/AMP guidelines 8. Your laboratory needs a reclassification workflow that checks previously reported VUS against new evidence on a defined schedule.
Core Principles of the ACMG/AMP Framework
The ACMG/AMP framework uses evidence codes that combine to produce a final classification. Each code has a strength level: very strong, strong, moderate, or supporting. The framework assigns specific combinations of codes to each classification tier. Understanding the logic behind these combinations matters more than memorizing the code list, because the same code can appear in different combinations with different outcomes.
Evidence Categories
The framework organizes evidence into two broad categories: pathogenicity evidence and benign evidence. Pathogenicity evidence includes null variants, de novo occurrence, functional studies, segregation data, allele frequency data, and computational predictions. Benign evidence includes high population allele frequency, observations in healthy individuals, and functional studies showing no damaging effect.
The original framework established 33 rules. Subsequent refinement efforts have shown that a simple implementation of the guidelines in their current form is insufficient for consistent and comprehensive variant classification, which led to the development of Sherloc, a refinement that separates ambiguous criteria into discrete related rules, groups certain criteria to prevent overcounting of conceptually related evidence, and replaces clinical criteria with additive semiquantitative criteria 7. Your laboratory should be aware that the 2015 framework is a starting point, not a final protocol.
Combining Evidence Codes
The framework specifies which combinations of codes reach each classification tier. For example, a pathogenic classification typically requires one very strong code plus several supporting codes, or multiple strong codes in combination. A likely pathogenic classification requires fewer or weaker combinations. The exact combinations are defined in the original ACMG/AMP publication, and your laboratory should have that publication available at every interpretation workstation.
The critical error to avoid is double counting. If two evidence codes derive from the same underlying observation, they should not both contribute to the classification. For example, if a functional study shows loss of function and the variant is a null variant in a gene where loss of function is the disease mechanism, you should not count both the functional study evidence and the null variant evidence as independent lines of support. The Sherloc refinement specifically addresses this problem by grouping related criteria to protect against overcounting 7.
Practical Workflow for Variant Classification
The workflow below describes the sequence of steps your laboratory should follow for each variant. Each step produces a record that becomes part of the variant interpretation file.
Step 1: Confirm Variant Identity and Annotation
Before you apply any classification criteria, confirm that the variant call is accurate. Review the sequencing quality metrics for the variant position, including read depth, base quality, and mapping quality. If the variant was identified through next-generation sequencing, consider whether Sanger confirmation is required by your laboratory policy.
Annotate the variant using a consistent transcript reference. Different transcripts can produce different protein-level consequences for the same genomic variant. Your laboratory policy should specify which transcript you use for each gene and how you handle variants that affect multiple transcripts. The NCBI provides sequence resources and database search systems that support consistent variant annotation 1.
Step 2: Determine Variant Type and Predicted Consequence
Classify the variant by type: single nucleotide variant, insertion, deletion, or indel. Determine the predicted consequence on the protein: missense, nonsense, frameshift, splice site, synonymous, or inframe. This determination drives which evidence codes are applicable.
For splice variants, you need to consider the position relative to the canonical splice sites. Variants in the invariant splice donor and acceptor positions typically qualify for null variant evidence, while variants deeper in the intron require additional analysis. Computational splice prediction tools can provide supporting evidence, but your laboratory should have a defined policy for how much weight these predictions receive.
Step 3: Search Population Databases
Query population frequency databases to determine whether the variant is observed in healthy populations. High allele frequency in a general population database is strong evidence for a benign classification, provided the database is appropriate for the patient population and the disease has a known prevalence.
The NCBI provides access to population variant databases and sequence resources that support this step 1. Your laboratory should record the database version and the allele frequency observed, because database versions change and a variant that appears rare in one version may be common in a later version.
Step 4: Search Disease and Literature Databases
Search ClinVar, locus-specific databases, and the published literature for previous reports of the variant. Record whether the variant has been reported before, what classifications were assigned, and what evidence supported those classifications.
The ALPL gene variant database demonstrates the value of locus-specific resources. It contains coding and non-coding variants, including single nucleotide variants, insertions/deletions, and structural variants, with each variant displayed alongside details explaining the corresponding pathogenicity and all reported genotypes and phenotypes 8. Your laboratory should identify the locus-specific database for each gene you interpret and consult it routinely.
Step 5: Apply Evidence Codes
Apply each evidence code that is supported by your collected data. Document the rationale for each code in the variant interpretation file. If a code could apply but you decide not to apply it, document that decision and the reason.
The evidence codes are applied independently, and the final classification is determined by the combination of codes. Your laboratory should use a standardized worksheet or software tool that enforces the combination rules and prevents arithmetic errors.
Step 6: Determine Final Classification
Combine the applied evidence codes according to the ACMG/AMP combination rules. If the combination reaches a classification tier, assign that tier. If the combination does not reach any tier, the variant is classified as VUS.
The VUS classification is not a failure. It is an honest statement that current evidence is insufficient. The ALPL variant classification project was established specifically because VUS cause diagnostic delay and uncertainty among patients and health care providers, and the project uses a multi-step process including clinical phenotype assessment, deep literature research, molecular genetic assessment, and in-vitro functional testing to reclassify submitted VUS 8. Your laboratory should have a similar process for periodic VUS review.
Step 7: Document and Report
Document the full evidence trail for each variant, including the databases queried, the versions used, the evidence codes applied, and the rationale for each code. This documentation must be sufficient for another laboratory scientist to reproduce your classification decision.
The report to the clinician should include the classification, the gene and variant description, and a brief summary of the evidence. The report should also indicate when the classification was last reviewed and whether a reclassification review is scheduled.
Evidence Code Application in Detail
The following sections describe how to apply specific evidence codes in practice. These descriptions focus on the decisions your laboratory scientist will make at the workbench.
Null Variant Evidence
Null variants include nonsense variants, frameshift variants, canonical splice site variants, and initiation codon variants. The very strong pathogenicity code applies when the null variant is in a gene where loss of function is a known disease mechanism.
The critical decision is whether loss of function is truly the disease mechanism for the gene in question. Some genes cause disease through dominant negative or gain of function mechanisms, and null variants in those genes may not be pathogenic. Your laboratory should have a curated list of genes with established loss of function disease mechanisms, and that list should be reviewed and updated regularly.
De Novo Variant Evidence
A de novo variant is one that is present in the proband but absent in both biological parents. The strength of this evidence depends on the confirmation status. Confirmed de novo occurrence provides strong evidence for pathogenicity, while unconfirmed de novo occurrence provides moderate evidence.
Parental relationships must be verified before de novo evidence is applied. If relationship testing has not been performed, the de novo status is not confirmed. Your laboratory policy should specify how you handle unconfirmed de novo variants and whether you require parental sample testing to confirm the variant call.
Population Frequency Evidence
Population frequency data provide both pathogenic and benign evidence. A variant that is absent from or extremely rare in large population databases supports pathogenicity, while a variant that is common in the general population supports a benign classification.
The benign strength of population frequency evidence depends on the allele frequency threshold and the disease prevalence. A variant observed at high frequency in a general population database is unlikely to be pathogenic for a rare Mendelian disease. Your laboratory should define the allele frequency thresholds for each strength level in your interpretation policy.
Functional Study Evidence
Functional studies can provide strong evidence for either pathogenicity or benign effect, depending on the study design and the assay used. Well-established functional assays that show a damaging effect support pathogenicity, while assays showing no damaging effect support a benign classification.
The quality of the functional study matters. Studies that use patient-derived cells or validated assays carry more weight than studies using overexpression systems or artificial constructs. Your laboratory should have criteria for evaluating functional study quality and should document how each study was assessed.
Computational Prediction Evidence
Computational prediction tools provide supporting evidence only. These tools include missense prediction algorithms, splice prediction algorithms, and conservation scores. The ACMG/AMP framework assigns supporting weight to computational evidence, and multiple tools that agree can be combined.
The limitation of computational tools is that they predict effect, not pathogenicity. A variant predicted to be damaging by multiple tools may still be benign if population frequency or functional studies show otherwise. Your laboratory should not allow computational predictions to override stronger evidence categories.
Gene-Specific and Disease-Specific Modifications
The standard ACMG/AMP framework does not fit every gene equally well. Some genes have unique challenges that require gene-specific modifications. Expert panels have developed customized criteria for specific gene groups, and your laboratory should use these modifications when they exist.
The ClinGen Rett/Angelman-like expert panel developed gene-specific variant interpretation methods for MECP2, CDKL5, FOXG1, UBE3A, SLC9A6, and TCF4 because these genes present unique challenges for the standard ACMG/AMP guidelines 9. In a pilot study, multiple curators obtained the same interpretation for approximately 90 percent of variants using the customized criteria, and the classification of 13 variants changed compared to original curations, with only one clinically relevant change from likely pathogenic to VUS 9.
Your laboratory should check whether a ClinGen expert panel has published specifications for each gene you interpret. If specifications exist, your interpretation policy should incorporate them. If no specifications exist, your laboratory may need to develop internal gene-specific rules, and those rules should be documented and reviewed.
Records and Measurements
Variant interpretation requires meticulous record keeping. Each variant interpretation file should contain the following elements:
Variant Identification Records
Record the genomic coordinates, the reference and alternate alleles, the gene name, the transcript identifier, and the protein-level consequence. Record the date of the interpretation and the version of the reference genome used.
Evidence Collection Records
Record every database queried, the version of each database, the date of the query, and the result. Record every publication reviewed, including the PubMed identifier and the relevant findings. Record every functional study evaluated and the assessment of study quality.
Classification Decision Records
Record each evidence code applied, the strength assigned, and the rationale. Record the final classification and the combination of codes that produced it. Record the names of the laboratory scientists who performed the interpretation and the review.
Reclassification Records
Record the date of each reclassification review, the new evidence considered, and the outcome. Record whether the classification changed and what triggered the change. The ALPL gene variant database provides a model for this process, with a submission system for clinicians, geneticists, genetic counselors, and researchers to submit VUS for classification by an international multidisciplinary consortium 8.
Common Failure Patterns in Variant Classification
Laboratories that apply the ACMG/AMP framework encounter recurring problems. Recognizing these patterns helps you avoid them.
Overcounting Related Evidence
The most common error is counting the same underlying observation multiple times through different evidence codes. For example, a variant that is both a null variant and predicted to cause nonsense-mediated decay should not receive separate evidence for both observations, because the null variant call already captures the loss of function consequence. The Sherloc refinement specifically addresses this problem by grouping related criteria 7.
Applying Criteria Without Gene-Specific Context
The standard ACMG/AMP framework assumes certain gene characteristics that do not hold for every gene. Applying null variant evidence to a gene where gain of function is the disease mechanism produces incorrect classifications. Your laboratory must verify the disease mechanism for each gene before applying mechanism-dependent criteria.
Ignoring Database Version Changes
Population databases and disease databases update regularly. A variant that was absent from a database in one version may be present in a later version. Your laboratory must record database versions and recheck variants when databases update.
Treating Computational Predictions as Definitive
Computational prediction tools provide supporting evidence, not definitive evidence. A variant predicted to be damaging by multiple tools still requires additional evidence for a pathogenic classification. Machine learning approaches that combine ACMG/AMP guidelines with variant annotation features can provide a probabilistic score of pathogenicity and may help prioritize variants that would otherwise be interpreted as uncertain 11, but these approaches support instead of replace the guideline framework.
Failing to Review VUS on Schedule
VUS classifications are provisional. New evidence accumulates continuously, and a variant that is a VUS today may be classifiable tomorrow. Your laboratory needs a defined VUS review schedule and a process for checking new evidence against previously classified variants.
Quality Controls and Professional Oversight
Variant interpretation is a professional judgment activity, and quality controls should focus on the interpretation process instead of only the final classification.
Dual Review
Each variant interpretation should be performed by one laboratory scientist and reviewed by a second qualified individual. The reviewer should have access to all evidence collected and should independently verify the evidence code assignments and the final classification. The Rett/AS VCEP pilot study demonstrated that multiple curators can obtain the same interpretation for approximately 90 percent of variants when using customized criteria 9, which supports the value of structured criteria in reducing inter-curator variability.
Interpretation Policy Compliance
Your laboratory should audit a sample of variant interpretations against the written interpretation policy on a defined schedule. The audit should check that all required evidence sources were queried, that evidence codes were applied correctly, and that the final classification follows from the applied codes.
Escalation Criteria
Define the situations that require escalation to a senior laboratory scientist or a variant interpretation committee. Typical escalation triggers include conflicting evidence that cannot be resolved, a classification that would change a previously reported result, a variant in a gene with no established disease mechanism, and any case where the laboratory scientist is uncertain about the appropriate evidence code application.
External Consultation
For difficult variants, consider consulting external experts. The ALPL gene variant classification project demonstrates the value of an international multidisciplinary consortium that includes clinicians, geneticists, genetic counselors, and researchers who reclassify VUS using a multi-step process 8. Your laboratory should identify external consultation pathways for genes where internal expertise is limited.
Training and Competency Requirements
Laboratory scientists who perform variant interpretation need specific training in the ACMG/AMP framework and in the genes they interpret. Training should cover the evidence code definitions, the combination rules, the gene-specific modifications, and the common failure patterns.
The EMBL-EBI provides training resources for bioinformatics and data-resource analysis that can support foundational skills in sequence analysis and variant interpretation 2. The Galaxy Training Network offers accessible workflow training and analysis tutorials that can help laboratory scientists build reproducible analysis skills 4. The Carpentries provides foundational computing and data lessons that support the computational skills needed for variant analysis 6.
Competency assessment should include both knowledge testing and practical case review. A laboratory scientist should demonstrate the ability to apply evidence codes correctly, to recognize when gene-specific modifications apply, and to document the interpretation process completely.
Reproducibility and Workflow Management
Variant interpretation should be reproducible. Another laboratory scientist should be able to follow your documentation and reach the same classification. Reproducibility requires standardized workflows and complete documentation.
Standardized Analysis Workflows
Your laboratory should use standardized analysis workflows for variant calling and annotation. Community workflow standards can support this goal. The nf-core documentation describes community pipeline standards, usage, configuration, and reproducible workflow context 5. Bioconductor provides official package, workflow, installation, and reproducible genomic-analysis documentation 3.
Version Control
Record the version of every tool, database, and reference file used in the interpretation process. Version control is essential for reproducibility because tool updates can change variant calls and annotations.
Documentation Standards
Your laboratory should define the minimum documentation required for each variant interpretation. The documentation should be sufficient for an external auditor to reconstruct the interpretation process and verify the classification.
Limitations of the ACMG/AMP Framework
The ACMG/AMP framework has known limitations that your laboratory should acknowledge in its interpretation policy.
Variant Types Not Covered
The framework was designed for sequence variants. Copy number variants require separate adaptation, as demonstrated by the development of CNV-specific modifications that showed improved concordance and usability for single-gene CNVs compared with using the sequence variant guidelines 10. Somatic variants also require separate protocols.
Gene-Specific Challenges
Some genes present challenges that the standard framework does not address. The Rett/AS VCEP was formed specifically because MECP2, CDKL5, FOXG1, UBE3A, SLC9A6, and TCF4 present unique challenges for current ACMG/AMP variant interpretation guidelines 9. Your laboratory should identify genes in your test menu that may require gene-specific modifications.
Interpretation Variability
Even with standardized criteria, interpretation can vary between curators. The Rett/AS VCEP pilot study found that multiple curators obtained the same interpretation for approximately 90 percent of variants, which means approximately 10 percent of variants had discordant interpretations 9. Your laboratory should have a process for resolving discordant interpretations.
Evolving Evidence
The evidence base for variant interpretation changes continuously. New population data, new functional studies, and new clinical reports can change the classification of a previously interpreted variant. Your laboratory needs a systematic process for tracking new evidence and reclassifying variants when appropriate.
Safety and Regulatory Context
Variant classification results affect clinical decisions. A pathogenic classification can lead to diagnostic confirmation, predictive testing of relatives, and management changes. A benign classification can lead to exclusion of the variant from clinical consideration. An incorrect classification can cause harm.
Your laboratory should understand the regulatory requirements that apply to variant interpretation in your jurisdiction. These requirements typically include quality system standards, proficiency testing, and reporting requirements. Your laboratory should also understand the professional standards for variant interpretation, including the expectation that classifications are based on current evidence and are reviewed periodically.
The NCBI provides access to databases and resources that support variant interpretation, including population frequency data, disease databases, and sequence resources 1. Your laboratory should use these resources as part of its standard interpretation workflow.
Professional Escalation Criteria
Your laboratory should define clear escalation criteria for variant interpretation cases that exceed routine decision-making. The following situations should trigger escalation to a senior laboratory scientist or a variant interpretation committee:
Conflicting Strong Evidence
When strong evidence supports both pathogenicity and benign effect for the same variant, the interpretation requires committee review. The committee should evaluate the quality of the conflicting evidence and determine which evidence carries more weight.
Classification Changes for Previously Reported Variants
When new evidence changes the classification of a variant that has already been reported to a clinician, the change requires review and a plan for notifying the clinician and the patient. The Rett/AS VCEP pilot study found that classification changes occurred when criteria specifications were applied, with changes in strength of criteria being one cause 9.
Variants in Genes Without Established Disease Mechanisms
When a variant is identified in a gene where the disease mechanism is unknown or controversial, the interpretation requires expert consultation. Applying mechanism-dependent evidence codes without established disease mechanisms risks incorrect classifications.
Variants Requiring Gene-Specific Modifications
When a variant falls into a gene category that requires gene-specific modifications, and no expert panel specifications exist, the interpretation requires committee review. Your laboratory may need to develop internal specifications and document the rationale.
Reclassification of VUS
When a VUS is reclassified to pathogenic or likely pathogenic, the change requires review and a plan for clinical notification. The ALPL gene variant classification project demonstrates the value of a structured reclassification process that includes clinical phenotype assessment, deep literature research, molecular genetic assessment, and in-vitro functional testing 8.
Building a Structured Evidence Matrix for Consistent ACMG/AMP Code Assignment
The most frequent source of discordant variant classifications is not a misunderstanding of the final combination rules but inconsistent application of individual evidence codes during the data collection phase. Two curators can review the same variant and assign different strengths to the same evidence because they have not agreed in advance on what constitutes a qualifying observation. A structured evidence matrix solves this problem by forcing the curator to document each piece of evidence against predefined criteria before any code is assigned. This section describes how to build, validate, and maintain such a matrix for your laboratory.
Defining the Evidence Matrix Structure
An evidence matrix is a tabular instrument that lists every ACMG/AMP evidence code in rows and the qualifying criteria for each code in columns. The columns should capture the evidence source, the specific observation, the strength level assigned, and the justification for that strength. The matrix serves as the intermediate artifact between raw data collection and final code assignment, and it becomes the permanent record of why each code was applied.
The matrix should be organized by evidence category. Pathogenicity codes occupy one section, and benign codes occupy another. Within each section, codes are ordered by their default strength level, from very strong to supporting. This ordering helps the curator recognize when evidence might qualify for a higher or lower strength than the default, which is a common source of error. For example, population frequency evidence can support either pathogenicity or benign effect depending on the observed frequency and the disease prevalence, and the matrix should make both possibilities visible to the curator.
Each row of the matrix should contain a field for the evidence source identifier. For population databases, this is the database name and version. For literature, this is the PubMed identifier. For functional studies, this is the study identifier and the assay type. For clinical data, this is the internal case identifier or the external report reference. The source identifier ensures that the evidence can be retrieved and re-examined during review or audit.
Building the Matrix for Your Laboratory
Start by listing every evidence code your laboratory will use. The standard ACMG/AMP framework defines 33 rules, and the Sherloc refinement expands these into 108 detailed refinements that separate ambiguous criteria into discrete related rules and group certain criteria to prevent overcounting of conceptually related evidence 7. Your matrix should reflect the framework version your laboratory has adopted, and it should be updated whenever your interpretation policy changes.
For each evidence code, define the qualifying criteria in operational terms. A qualifying criterion is a measurable or verifiable condition that must be met before the code can be applied. For example, the null variant code requires that the variant is a nonsense, frameshift, canonical splice site, or initiation codon variant, and that loss of function is a known disease mechanism for the gene. The matrix should list both conditions separately so the curator can verify each one independently.
The matrix should also define the exclusion criteria for each code. An exclusion criterion is a condition that prevents the code from being applied even when the qualifying criteria are met. For example, a null variant in the last exon of a gene may not trigger nonsense-mediated decay, and some laboratories exclude the null variant code in this context unless functional data support a loss of function effect. The Sherloc refinement specifically addresses such edge cases by capturing exceptions and refining the classification criteria 7. Your matrix should document these exclusions so that curators apply them consistently.
Populating the Matrix During Variant Review
When you begin a new variant interpretation, create a fresh copy of the matrix template and populate it as you collect evidence. The matrix should be populated in the same order as your data collection workflow, which typically proceeds from variant confirmation through population databases, disease databases, literature review, and functional study evaluation.
For each piece of evidence, record the observation in the matrix before deciding whether it qualifies for a code. This separation of observation from interpretation is critical. A curator who records the observation first and then evaluates the code is less likely to force ambiguous evidence into a code than a curator who starts with a hypothesis about the classification and searches for supporting evidence.
The matrix should include a field for evidence strength assessment. The default strength for each code is defined by the ACMG/AMP framework, but the framework allows strength adjustment in specific circumstances. For example, the de novo evidence code has different strengths depending on whether the de novo occurrence is confirmed by parental relationship testing and whether the phenotype in the proband matches the expected disease spectrum. The matrix should prompt the curator to document the factors that support the assigned strength.
Using the Matrix to Prevent Double Counting
The most common error in variant classification is applying multiple evidence codes that derive from the same underlying observation. The evidence matrix helps prevent this error by making the evidence source visible for every code. When two codes reference the same source identifier, the curator and the reviewer can immediately recognize the potential for double counting.
The Sherloc refinement specifically addresses this problem by grouping certain criteria to protect against the overcounting of conceptually related evidence 7. Your matrix should implement these groupings by adding a grouping field to each row. Codes that belong to the same evidence group should be flagged, and the matrix should include a summary view that shows which groups have been used for the current variant.
For example, a functional study showing loss of function and a null variant call in a gene where loss of function is the disease mechanism derive from related evidence. The functional study confirms the mechanism, but the null variant code already captures the predicted consequence. Applying both codes at full strength overcounts the evidence. The matrix should prompt the curator to identify such relationships and to document whether the codes are applied independently or whether one code subsumes the other.
Validating the Matrix With Known Variants
Before your laboratory adopts an evidence matrix for routine use, validate it with a set of known variants. Select 10 to 20 variants with established classifications from ClinVar or from your laboratory's internal records. Have two curators independently populate the matrix for each variant and compare their code assignments. The goal is not to achieve perfect agreement on the final classification but to identify where the matrix allows divergent code assignments.
The Rett/AS VCEP pilot study provides a model for this validation approach. Multiple curators obtained the same interpretation for 78 out of 87 variants, approximately 90 percent, when using customized variant interpretation criteria 9. The remaining discordant cases were analyzed to identify which criteria were ambiguous, and the criteria were refined to reduce future discordance. Your laboratory should follow the same iterative process, using discordant cases to improve the matrix instead of to assign blame.
The validation process should also test the matrix against edge cases. Include variants in genes with unusual disease mechanisms, variants in the last exon, variants in non-canonical splice sites, and variants with conflicting functional data. These edge cases are where the standard framework provides the least guidance and where the matrix provides the most value by forcing explicit documentation of the curator's reasoning.
Maintaining the Matrix Over Time
The evidence matrix is not a static document. It must be updated whenever the ACMG/AMP framework is revised, when new gene-specific specifications are published, and when your laboratory identifies recurring interpretation problems. The ClinGen Rett/Angelman-like expert panel developed gene-specific modifications for MECP2, CDKL5, FOXG1, UBE3A, SLC9A6, and TCF4 because these genes present unique challenges for the standard guidelines 9. When such specifications exist, your matrix should incorporate them for the affected genes.
Assign a matrix owner who is responsible for tracking updates and communicating changes to the interpretation team. The matrix should have a version number and a change log that records what changed, when it changed, and why. The change log is essential for audit purposes because it allows reviewers to determine which matrix version was in effect when a particular variant was classified.
Schedule a periodic matrix review, typically annually or when a significant framework update is published. The review should examine whether any evidence codes have caused recurring discordance, whether new evidence types should be added, and whether the qualifying criteria need refinement. The ALPL gene variant classification project demonstrates the value of continuous assessment, as the project was established to continuously assess and update genetic, phenotypic, and functional variant information in the database 8. Your matrix should follow the same principle of continuous improvement.
Integrating the Matrix With Laboratory Information Systems
The evidence matrix can be implemented as a paper worksheet, a spreadsheet, or a module within your laboratory information system. The implementation choice depends on your laboratory's volume and existing infrastructure. Laboratories with high variant volumes should integrate the matrix into their information system so that the matrix data can be searched, audited, and exported for quality reviews.
Regardless of the implementation, the matrix must be accessible to all curators and reviewers at the time of interpretation. A matrix that exists only as a policy document on a shared drive will not be used consistently. The matrix should be embedded in the interpretation workflow so that the curator cannot complete a variant interpretation without populating the matrix.
The matrix data should also feed into your laboratory's quality metrics. Track the frequency of code application across variants, the rate of discordant interpretations, and the types of evidence that most often lead to classification changes. These metrics help identify where the matrix needs refinement and where additional curator training is needed.
Common Matrix Implementation Failures
Laboratories that adopt an evidence matrix sometimes encounter implementation failures that undermine its value. Recognizing these patterns helps you avoid them.
The first failure is creating a matrix that is too complex to use efficiently. A matrix with dozens of columns and extensive free-text fields becomes a burden, and curators will complete it superficially or after the fact. The matrix should capture the essential information, source, observation, strength, and justification, without requiring unnecessary detail.
The second failure is treating the matrix as a documentation exercise instead of a decision tool. The matrix adds value only when it is populated during the evidence collection process and used to guide the code assignment. A matrix that is completed after the classification has already been determined provides no protection against biased interpretation.
The third failure is failing to update the matrix when the evidence base changes. A matrix that reflects outdated framework versions or outdated gene-specific specifications will produce classifications that do not reflect current knowledge. The matrix owner must monitor the literature and the ClinGen expert panel publications for updates that affect the matrix content.
The fourth failure is not training curators on matrix use. The matrix is only as good as the people who use it. Curators need training on the matrix structure, the qualifying criteria, and the documentation expectations. The EMBL-EBI provides training resources for bioinformatics and data-resource analysis that can support foundational skills in sequence analysis and variant interpretation 2, and the Galaxy Training Network offers accessible workflow training and analysis tutorials that can help laboratory scientists build reproducible analysis skills 4. Your laboratory should incorporate matrix training into its competency program.
Records and Measurements for Matrix Use
Your laboratory should track specific measurements to evaluate whether the evidence matrix is achieving its purpose. The primary measurement is the inter-curator agreement rate, which is the proportion of variants for which two independent curators assign the same evidence codes and reach the same final classification. The Rett/AS VCEP pilot study achieved approximately 90 percent agreement using customized criteria 9, and your laboratory should track this rate over time to identify trends.
A second measurement is the code application frequency, which is the proportion of variants for which each evidence code is applied. This measurement helps identify codes that are overused or underused. A code that is applied to nearly every variant may be applied too liberally, while a code that is rarely applied may be too restrictive.
A third measurement is the matrix completion rate, which is the proportion of variant interpretations that have a fully populated evidence matrix. This measurement identifies compliance problems before they affect classification quality. A low completion rate indicates that curators are not using the matrix as intended and that additional training or workflow changes are needed.
A fourth measurement is the classification change rate, which is the proportion of variants whose classification changes upon re-review. This measurement helps evaluate whether the matrix is capturing the evidence needed for stable classifications. A high change rate may indicate that the matrix is missing important evidence categories or that the qualifying criteria are not specific enough.
Professional Escalation Criteria for Matrix Disagreements
When two curators disagree on the evidence codes for a variant, the disagreement should be resolved through a defined process. The first step is to compare the populated matrices and identify where the code assignments diverge. The divergence may result from different evidence sources, different assessments of study quality, or different interpretations of the qualifying criteria.
If the disagreement results from different evidence sources, the curators should review the sources together and determine which sources are reliable. If the disagreement results from different assessments of study quality, the curators should apply the laboratory's study quality criteria to the specific study. If the disagreement results from different interpretations of the qualifying criteria, the case should be escalated to the matrix owner or the variant interpretation committee.
The matrix owner should track recurring disagreements and use them to refine the matrix. A criterion that consistently produces divergent interpretations needs to be clarified. The Sherloc refinement process, which evolved through eight major and minor revisions based on the classification of more than 40,000 clinically observed variants, demonstrates the value of iterative refinement 7. Your laboratory should adopt the same iterative approach to matrix improvement.
Frequently Asked Questions
What is the difference between pathogenic and likely pathogenic classifications?
A pathogenic classification means the evidence strongly supports a disease-causing role for the variant, while a likely pathogenic classification means the evidence supports a disease-causing role but with less certainty. The distinction matters for clinical action. Pathogenic variants can be used for diagnostic confirmation and predictive testing, while likely pathogenic variants may guide management while additional evidence is gathered. The ACMG/AMP framework defines specific evidence combinations for each tier, and your laboratory should apply those combinations consistently.
How should a laboratory handle a variant that does not reach any classification tier?
A variant that does not reach the evidence threshold for any classification tier is classified as variant of uncertain significance. This classification is appropriate when evidence is insufficient or conflicting. The VUS classification should trigger a reclassification review schedule, and the laboratory should document the evidence gaps that prevented a definitive classification. The ALPL gene variant classification project demonstrates how structured reclassification processes can convert VUS to definitive classifications over time 8.
Can computational prediction tools replace functional studies in variant classification?
Computational prediction tools provide supporting evidence only and cannot replace functional studies. The ACMG/AMP framework assigns supporting weight to computational evidence, while well-established functional studies can provide strong evidence. Machine learning approaches that combine ACMG/AMP guidelines with variant annotation features can provide a probabilistic score of pathogenicity and may help prioritize variants 11, but these approaches support instead of replace the guideline framework.
How often should a laboratory review previously classified variants?
Your laboratory should define a VUS review schedule in its interpretation policy. The schedule should account for the rate at which new evidence accumulates for the genes in your test menu. Some laboratories review VUS annually, while others review on a rolling basis as new evidence becomes available. The ALPL gene variant classification project was established to continuously assess and update genetic, phenotypic, and functional variant information 8, which reflects the expectation that variant classifications are not static.
What should a laboratory do when different curators reach different classifications for the same variant?
Discordant interpretations should be resolved through a defined process. The laboratory should document the evidence each curator applied and identify the source of disagreement. The disagreement may result from different evidence code applications, different assessments of study quality, or different interpretations of gene-specific context. The Rett/AS VCEP pilot study found that multiple curators obtained the same interpretation for approximately 90 percent of variants using customized criteria 9, which suggests that structured criteria reduce but do not eliminate discordance.
How do gene-specific expert panel specifications affect variant classification?
Gene-specific expert panel specifications modify the standard ACMG/AMP criteria to address gene-specific challenges. When specifications exist for a gene, your laboratory should use them instead of the standard criteria. The Rett/AS VCEP developed gene-specific modifications for MECP2, CDKL5, FOXG1, UBE3A, SLC9A6, and TCF4 because these genes present unique challenges for the standard guidelines 9. Your laboratory should check for expert panel specifications for every gene you interpret.
What evidence is needed to classify a variant as benign?
A benign classification requires strong evidence against pathogenicity. This evidence typically includes high population allele frequency, observations in healthy individuals, or functional studies showing no damaging effect. The ACMG/AMP framework defines specific evidence codes for benign classification, and your laboratory should apply those codes consistently. A likely benign classification requires less evidence than a benign classification but still requires evidence against pathogenicity.
How should a laboratory document variant interpretation for regulatory compliance?
Variant interpretation documentation should include the variant identification, the evidence sources queried with versions and dates, the evidence codes applied with rationale, the final classification, and the names of the individuals who performed and reviewed the interpretation. The documentation should be sufficient for an external auditor to reconstruct the interpretation process. Your laboratory should also document the interpretation policy version and any gene-specific specifications that were applied.
Related Bioinformatics Guides
- FAIR Data Maturity Model: A Practical Assessment Framework for Bioinformatics Workflows
- Digital Pathology Validation: A Practical Guide to CAP and RCPath Compliance
- Detecting Structural Variants with Long-Read Sequencing: Methods and Considerations
- Digital Pathology Guidelines: A Reference for Implementation
- Metabolomics Data Analysis in R: A Practical Workflow
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- Sherloc: a comprehensive refinement of the ACMG-AMP variant classification criteria.. Genetics in medicine : official journal of the American College of Medical Genetics, 2017.
- The Global ALPL gene variant classification project: Dedicated to deciphering variants.. Bone, 2024.
- Recommendations by the ClinGen Rett/Angelman-like expert panel for gene-specific variant interpretation methods.. Human mutation, 2022.
- Adapting ACMG/AMP sequence variant classification guidelines for single-gene copy number variants.. Genetics in medicine : official journal of the American College of Medical Genetics, 2020.
- A machine learning approach based on ACMG/AMP guidelines for genomic variant classification and prioritization.. Scientific reports, 2022.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.