A Beginner's Guide to Variant Annotation: From VCF to Functional Insights
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Variant annotation transforms raw VCF data into biologically interpretable information by linking variants to genes, transcripts, protein changes, population frequencies (e.g., gnomAD), and clinical significance (e.g., ClinVar). Tools like Ensembl VEP and SnpEff are essential for this process, predicting molecular consequences such as missense, frameshift, or splice region variants.
- The accuracy of annotation is critically dependent on the quality and version of reference databases (e.g., Ensembl, GENCODE, RefSeq) and genome assemblies (e.g., GRCh37, GRCh38); using mismatched assemblies is a primary source of annotation errors. Maintaining meticulous records of tool and database versions is crucial for reproducibility.
- Interpreting annotation outputs requires considering multiple lines of evidence: consequence terms describe molecular impact, population frequencies (e.g., gnomAD allele frequency) provide context for commonality, and clinical annotations (e.g., ClinVar classification) directly link variants to known diseases.
- Deleteriousness prediction scores (e.g., SIFT, CADD) offer supplementary evidence for functional impact but should not be used as standalone decision criteria, as different tools can yield conflicting predictions.
- A systematic triage system, often employing a decision matrix, is vital for prioritizing variants, distinguishing between Tier 1 (strong evidence), Tier 2 (partial evidence), and Tier 3 (weak evidence) findings, and determining when to escalate for further analysis or clinical review.
- For pharmacogenomic applications, specialized tools like PharmCAT integrate patient genotypes with clinical guidelines (e.g., CPIC, DPWG) to inform drug therapy decisions, requiring interpretation by qualified healthcare professionals.
Variant annotation is the process of adding biological context to the variants listed in a VCF file, transforming raw genomic coordinates and allele calls into interpretable information about genes, transcripts, protein changes, and potential clinical relevance. For a researcher who has just completed variant calling and holds a VCF file with thousands of candidate variants, annotation answers the immediate question of which variants matter and why. This guide walks through the practical steps of annotating a VCF using SnpEff and Ensembl VEP, explains the key output fields, and provides concrete criteria for interpreting results and deciding when to escalate findings for further analysis.
The intended reader is a biology student, researcher, or laboratory professional who understands the basics of sequencing but has not yet performed variant annotation. The workflow assumes you have a VCF file from a germline or somatic variant calling pipeline and need to add functional information before filtering and prioritization. By the end of this guide, you will know how to run annotation tools, read their outputs, interpret consequence terms, and avoid common mistakes that lead to misinterpreted variants.
What Variant Annotation Actually Does
Variant annotation connects each variant in your VCF to reference databases and prediction algorithms. The raw VCF contains chromosome position, reference allele, alternate allele, quality scores, and genotype information. Annotation adds layers of meaning: which gene the variant falls within, which transcript is affected, whether the change alters the protein sequence, whether population databases show this variant in healthy individuals, and whether clinical databases have classified it as pathogenic or benign.
The Ensembl Variant Effect Predictor (VEP) is a freely available open-source tool that performs this annotation and filtering work. VEP predicts the molecular consequences of variants using the Ensembl, GENCODE, or RefSeq gene sets. It also reports phenotype associations from databases such as ClinVar, allele frequencies from studies including gnomAD, and deleteriousness predictions from tools such as Sorting Intolerant From Tolerant (SIFT) and Combined Annotation Dependent Depletion (CADD). VEP is updated roughly quarterly to incorporate the latest gene, variant, and phenotype association information, which matters for keeping annotations current [<a href="#ref-1">1</a>].
The National Center for Biotechnology Information (NCBI) maintains a suite of databases and analysis services that support variant interpretation, including reference sequences, gene records, and clinical variation resources. These databases provide the underlying reference data that annotation tools draw upon, and understanding their structure helps you interpret annotation outputs correctly [<a href="#ref-2">2</a>].
For researchers who prefer graphical interfaces or structured training, the Galaxy Training Network offers accessible workflow training and analysis tutorials that cover variant annotation in a reproducible environment. The European Bioinformatics Institute provides bioinformatics learning pathways and data-resource training that can supplement your understanding of the underlying databases [<a href="#ref-3">3</a>][<a href="#ref-4">4</a>].
The Variant Calling Context
Variant annotation does not happen in isolation. It sits at the end of a variant calling workflow that begins with sequencing reads and proceeds through alignment, variant calling, and filtering. The quality of your annotation output depends directly on the quality of your input VCF.
Germline variant calling identifies variants inherited from parents and present in all cells of an individual. These variants are typically heterozygous or homozygous and are called against a reference genome using tools that account for expected allele frequencies. Somatic variant calling identifies mutations acquired during an individual's lifetime, often in cancer samples, and requires comparing tumor tissue to matched normal tissue to distinguish true somatic mutations from inherited variants and sequencing artifacts.
The complexity of variant analysis pipelines, terminology, and tool selection remains a major barrier for researchers new to the field or working in translational settings. A practical framework for variant interrogation in tumor samples breaks the process into four phases: planning the experimental design and understanding sequencing outputs, gathering the tools and reference data required for analysis, filtering and validating variants systematically, and disseminating findings through transparent reporting and data sharing [<a href="#ref-5">5</a>].
Your annotation strategy should account for whether you are working with germline or somatic variants. Germline analysis often focuses on rare variants with high predicted impact that segregate with disease in families. Somatic analysis must contend with tumor heterogeneity, varying variant allele fractions, and the need to distinguish driver mutations from passenger mutations. The annotation fields you prioritize will differ accordingly.
Core Principles of Variant Annotation
Annotation Is Only as Good as the Reference Data
Every annotation tool depends on reference databases that map genomic coordinates to genes, transcripts, and regulatory elements. These databases are updated regularly as genome assemblies improve and gene models are refined. VEP is updated roughly quarterly to incorporate the latest gene, variant, and phenotype association information, which means running the same VCF through different versions of VEP can produce different annotations [<a href="#ref-1">1</a>].
You should record the version of the annotation tool and the reference database version in your analysis notes. This practice ensures that your results can be reproduced and compared across time points. If you revisit a VCF file months later, you need to know which annotation version produced your original results.
Consequence Terms Describe Molecular Impact
The core output of variant annotation is a consequence term that describes the predicted molecular effect of a variant. Common terms include missense variant, which changes one amino acid in the protein, synonymous variant, which does not change the amino acid, frameshift variant, which alters the reading frame, stop gained, which introduces a premature stop codon, and splice region variant, which affects splicing.
These terms are assigned based on the transcript model. A variant may have different consequences for different transcripts of the same gene. VEP reports the most severe consequence across all transcripts in its canonical output, but the full output lists consequences for each affected transcript. You should examine transcript-specific annotations when a variant falls in a gene with multiple isoforms.
Population Frequency Provides Context
Population databases such as gnomAD report how often a variant appears in large cohorts of sequenced individuals. A variant that is common in the general population is less likely to be a rare disease-causing variant, although there are exceptions for variants with incomplete penetrance or late onset. VEP reports allele frequencies from studies including gnomAD, allowing you to filter out common variants during prioritization [<a href="#ref-1">1</a>].
Population frequency filtering is particularly useful for germline analysis. A variant that appears in more than one percent of the population is unlikely to be a rare Mendelian disease variant, though this threshold varies by disease and inheritance model. For somatic analysis, population frequency helps distinguish rare germline variants from true somatic mutations.
Clinical Annotations Link to Known Disease
Clinical databases such as ClinVar contain classifications of variants based on evidence from clinical testing and literature. VEP reports phenotype associations from databases such as ClinVar, giving you direct access to established clinical interpretations [<a href="#ref-1">1</a>].
A ClinVar classification of pathogenic or likely pathogenic provides strong evidence for variant significance. A classification of benign or likely benign suggests the variant is not disease-causing. Many variants have conflicting classifications or no clinical annotation at all, and these require additional scrutiny.
Deleteriousness Predictions Estimate Functional Impact
Computational prediction tools estimate whether a variant is likely to damage protein function. SIFT and CADD are among the tools whose predictions VEP can report. These predictions are based on sequence conservation, protein structure, and other features, and they provide supporting evidence when clinical and population data are insufficient [<a href="#ref-1">1</a>].
Prediction scores are not definitive. Different tools can disagree on the same variant, and predictions are only as good as the training data and features used to build them. Use prediction scores as one line of evidence among many, not as a standalone decision criterion.
At a Glance
| Annotation Task | Recommended Tool | Key Output Fields | Primary Use Case |
|---|---|---|---|
| General variant annotation with clinical and population data | Ensembl VEP | Consequence, gene symbol, ClinVar classification, gnomAD frequency, CADD score | Germline and somatic variant interpretation with broad database coverage |
| Rapid annotation with a lightweight tool | SnpEff | Consequence, gene symbol, transcript ID, impact category | Quick screening of large VCF files before detailed analysis |
| Pharmacogenomic interpretation | PharmCAT | Allele, genotype, phenotype, CPIC or DPWG guideline recommendation | Clinical decision support for drug response variants |
| Custom prioritization of annotated variants | VPOT | Custom pathogenicity ranking score from user-defined annotations | Cohort analysis with tailored variant ranking |
Practical Workflow for Annotating a VCF
Step 1: Prepare Your Input VCF
The annotation process begins with a VCF file that has passed basic quality filtering. Your VCF should contain high-confidence variant calls with genotype information. Remove low-quality calls and obvious artifacts before annotation to reduce noise in your output.
For somatic variant calling, ensure that your VCF contains the somatic calls you want to annotate and that you have documented the filtering criteria used to produce them. The planning phase of variant interrogation includes understanding sequencing outputs and designing the analysis around the specific question you are asking [<a href="#ref-5">5</a>].
Check that your VCF uses the same reference genome assembly as your annotation tool. A VCF built against GRCh37 cannot be annotated correctly with tools configured for GRCh38. This mismatch is a common source of annotation errors that are difficult to detect after the fact.
Step 2: Choose Your Annotation Tool
SnpEff and VEP are the two most commonly used annotation tools for beginners. Both are freely available and well documented.
SnpEff is a fast, lightweight tool that annotates variants with gene information and consequence terms. It is suitable for quick screening of large VCF files and produces a straightforward output that is easy to parse. SnpEff requires downloading a prebuilt database for your reference genome, and the database version must match your VCF assembly.
VEP offers more comprehensive annotation with access to clinical databases, population frequencies, and prediction scores. It can be run through a web interface, a command-line tool, or an application programming interface, making it accessible to users with different levels of bioinformatics experience. The web interface is suitable for small numbers of variants, while the command-line tool handles larger datasets efficiently [<a href="#ref-1">1</a>].
For pharmacogenomic applications, the Pharmacogenomics Clinical Annotation Tool (PharmCAT) incorporates patient genotypes, annotates pharmacogenomic information including allele, genotype, and phenotype, and generates a report with guideline recommendations from the Clinical Pharmacogenetics Implementation Consortium (CPIC) and the Dutch Pharmacogenetics Working Group (DPWG). PharmCAT includes a VCF Preprocessor that can prepare biobank-scale data for analysis [<a href="#ref-6">6</a>].
Step 3: Run the Annotation
For VEP through the web interface, upload your VCF file, select the appropriate species and genome assembly, and choose the annotation sources you want to include. The web interface is designed to enable sophisticated analysis through a simple interface, making it a good starting point for beginners [<a href="#ref-1">1</a>].
For VEP through the command line, the basic command structure is:
vep -i input.vcf -o output.txt --cache --offline
This command uses the local cache for faster processing. Additional flags add specific annotation sources, such as --clinvar for clinical classifications, --gnomad for population frequencies, and --cadc for CADD scores. Consult the VEP documentation for the full list of options.
For SnpEff, the basic command structure is:
snpEff ann GRCh38.105 input.vcf > output.ann.vcf
This command annotates the input VCF using the specified database and writes the annotated VCF to standard output. The annotated VCF contains the original information plus an ANN field in the INFO column with the annotation details.
Step 4: Examine the Output
VEP produces a tab-delimited text file by default, with one line per variant and transcript. Key columns include:
- Uploaded variation: the variant identifier from your input
- Location: chromosome and position
- Allele: the alternate allele
- Gene: the Ensembl gene identifier
- Feature: the transcript identifier
- Consequence: the predicted molecular consequence
- IMPACT: a high-level category of predicted impact
- SYMBOL: the gene symbol
- ClinVar: clinical classification if available
- gnomAD_AF: allele frequency in gnomAD
SnpEff adds an ANN field to the INFO column of the VCF. This field contains a comma-separated list of annotations, each with the format:
Allele | Consequence | Impact | Gene | Gene ID | Feature type | Feature ID | Transcript biotype | Rank | HGVS notation
The ANN field can be parsed with tools like BCFtools or custom scripts to extract specific annotation fields for filtering.
Step 5: Filter and Prioritize
After annotation, you need to filter variants based on your research question. The filtering and validation phase of variant interrogation involves executing a systematic approach to prioritize meaningful variants [<a href="#ref-5">5</a>].
For germline analysis, a typical filtering strategy includes:
- Remove variants with population frequency above a threshold relevant to your disease
- Retain variants with high or moderate predicted impact
- Prioritize variants in genes with known relevance to your phenotype
- Consider inheritance patterns and family segregation data
For somatic analysis, additional considerations include:
- Focus on variants with adequate sequencing depth and variant allele fraction
- Prioritize variants in known cancer driver genes
- Consider the functional impact of the variant in the context of the tumor type
- Validate candidate variants with an orthogonal method
The Variant Prioritization Ordering Tool (VPOT) allows researchers to create a single customizable pathogenicity ranking score from any number of annotation values, each with a user-defined weighting. VPOT can be informative when analyzing entire cohorts, as variants in a cohort can be prioritized. It also provides functions for filtering based on a candidate gene list or affected status in a family pedigree [<a href="#ref-7">7</a>].
Options and Tradeoffs in Annotation Tools
SnpEff versus VEP
SnpEff is faster and simpler, making it ideal for initial screening and for researchers who need a quick annotation of a large VCF. Its output is compact and easy to parse programmatically. However, SnpEff does not include clinical annotations or population frequencies by default, so you will need to supplement its output with additional data sources.
VEP provides a more complete annotation with clinical, population, and prediction data in a single run. It is more configurable and extensible, with options to add custom annotations and plugins. The tradeoff is speed and complexity. VEP runs are slower than SnpEff, and the command-line configuration requires more learning.
The choice between tools depends on your specific needs. If you need clinical annotations and population frequencies, VEP is the better choice. If you need a quick screen of a large dataset and plan to add annotations later, SnpEff is sufficient.
Web Interface versus Command Line
VEP offers three access methods: a web interface, a command-line tool, and an application programming interface. These methods are designed to suit different levels of bioinformatics experience and meet different needs in terms of data size, visualization, and flexibility [<a href="#ref-1">1</a>].
The web interface is the easiest way to start. You can upload a small VCF file, select your options through a form, and view results in a browser. The web interface is suitable for datasets with a few thousand variants. For larger datasets, the command-line tool is necessary.
The command-line tool requires installation and configuration but offers full control over annotation sources and output formats. It can process large VCF files efficiently and can be integrated into automated pipelines.
The application programming interface allows programmatic access to VEP for custom applications. This option is for advanced users who need to integrate annotation into their own software.
Reproducibility Considerations
Reproducibility is a core concern in bioinformatics analysis. The Galaxy Training Network emphasizes accessible workflow training and reproducibility context, and the nf-core documentation describes community pipeline standards for reproducible workflow execution [<a href="#ref-4">4</a>][<a href="#ref-8">8</a>].
For reproducible variant annotation, you should:
- Record the exact version of your annotation tool
- Record the version of the reference database
- Record all command-line options used
- Store the annotated output files with your analysis
- Document the filtering criteria applied after annotation
The Carpentries lessons provide foundational training in computing, data, shell, Git, and programming that supports reproducible analysis practices. These skills are directly applicable to managing annotation workflows and version control [<a href="#ref-9">9</a>].
Records and Measurements for Annotation Quality
What to Record
Maintain a laboratory notebook or electronic record that documents:
- The VCF file used as input, including its version and how it was generated
- The annotation tool and version number
- The reference genome assembly and database version
- All command-line options and parameters
- The date of annotation
- The output file names and locations
- Any filtering steps applied after annotation
This record ensures that your annotation results can be reproduced and audited. If a collaborator or reviewer asks how a particular annotation was generated, you can provide the exact commands and versions.
Quality Checks After Annotation
After running annotation, perform these quality checks:
- Verify that the number of annotated variants matches the number of variants in your input VCF
- Check that gene symbols and transcript identifiers are present for a sample of variants
- Confirm that the reference genome assembly in your annotation output matches your input VCF
- Examine a few known variants to confirm that their annotations match expected results
- Check for unexpected patterns, such as a high proportion of variants with no annotation
These checks catch common errors such as assembly mismatches, database version problems, and input formatting issues.
Measuring Annotation Completeness
Annotation completeness refers to the proportion of variants that receive a meaningful annotation. A variant may lack annotation if it falls in an intergenic region, if the gene model does not include that position, or if the annotation database does not cover the variant.
For most analyses, a small proportion of unannotated variants is expected. Intergenic variants and variants in poorly characterized regions will not have gene annotations. However, a high proportion of unannotated variants may indicate a problem with your reference assembly or database version.
Common Failure Patterns in Variant Annotation
Reference Assembly Mismatch
The most common and damaging error is using a VCF built against one reference assembly with an annotation tool configured for another. This error produces annotations that map to incorrect genomic positions, leading to wrong gene assignments and consequence predictions.
Prevention: Always verify the reference assembly of your VCF before annotation. Check the header of your VCF for assembly information, and confirm that your annotation tool is configured for the same assembly.
Outdated Database Versions
Annotation databases are updated regularly. Using an outdated database means missing newly discovered variants, updated clinical classifications, and refined gene models. VEP is updated roughly quarterly, and using an older version means your annotations do not reflect the latest information [<a href="#ref-1">1</a>].
Prevention: Update your annotation databases regularly and record the version used for each analysis. If you compare results across time points, ensure that you use the same database version or account for differences.
Misinterpreting Consequence Terms
Consequence terms describe molecular impact, but they do not directly indicate disease relevance. A missense variant in a gene with no known disease association is less informative than a missense variant in a well-characterized disease gene. Beginners often overinterpret consequence terms without considering gene context.
Prevention: Always interpret consequence terms in the context of gene function, population frequency, and clinical annotations. A high-impact consequence in a gene with no known phenotype is not necessarily disease-causing.
Ignoring Transcript Diversity
Many genes have multiple transcripts, and a variant may have different consequences for different transcripts. Focusing only on the canonical transcript can miss important effects on alternative isoforms.
Prevention: Examine transcript-specific annotations when a variant falls in a gene with multiple isoforms. VEP reports consequences for all affected transcripts, and you should review these when the canonical transcript annotation is ambiguous.
Overreliance on Prediction Scores
Prediction scores from tools like SIFT and CADD are useful supporting evidence, but they are not definitive. Different tools can disagree, and predictions can be wrong. Using prediction scores as the sole basis for variant prioritization leads to false positives and false negatives.
Prevention: Use prediction scores as one line of evidence among many. Combine them with population frequency, clinical annotations, and gene context to make informed decisions.
Failure to Document the Analysis
Without proper documentation, annotation results cannot be reproduced or audited. This failure undermines the credibility of your analysis and makes it difficult to troubleshoot problems.
Prevention: Record all versions, parameters, and commands as described in the records section. Store annotated outputs with your analysis files.
Interpreting Annotation Outputs
Reading a VEP Output Line
A typical VEP output line for a missense variant might look like this:
1 123456 rs123456 A G 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100
The columns are defined in the VEP output header, which is included in the output file. Key columns to examine include the consequence, gene symbol, ClinVar classification, and population frequency.
Reading a SnpEff ANN Field
The SnpEff ANN field in the INFO column of the annotated VCF contains the annotation details. For a missense variant, the ANN field might look like this:
ANN=A|missense_variant|MODERATE|BRCA1|ENSG00000012048|transcript|NM_007294.3|protein_coding|11/22|c.123A>G|p.Leu41Phe|1234/5678|...
This field tells you the alternate allele, the consequence, the impact category, the gene symbol, the gene identifier, the feature type, the transcript identifier, the transcript biotype, the exon and total exon count, the coding change in HGVS notation, the protein change in HGVS notation, and the cDNA position.
Using Clinical Annotations
When a variant has a ClinVar classification, this information provides direct clinical context. A classification of pathogenic or likely pathogenic indicates that the variant has been associated with disease in clinical testing. A classification of benign or likely benign indicates that the variant has been observed in healthy individuals or has been shown not to cause disease.
Variants without ClinVar classifications require additional evidence to interpret. Population frequency, prediction scores, and gene context become more important for these variants.
Using Population Frequencies
Population frequency data from gnomAD and similar databases tells you how common a variant is in the general population. A variant that is common in the population is unlikely to be a rare disease-causing variant, though there are exceptions.
For germline analysis, a common filtering approach is to remove variants with a population frequency above a threshold. The threshold depends on the disease prevalence and inheritance model. For rare diseases, a threshold of one percent is commonly used, but you should adjust based on your specific question.
For somatic analysis, population frequency helps distinguish rare germline variants from true somatic mutations. A variant that is common in the population is unlikely to be a somatic mutation, though it could be a germline variant in the tumor sample.
Safety and Regulatory Context
Variant annotation is a research tool, not a clinical diagnostic. The interpretation of variants for clinical decision-making requires additional validation, regulatory oversight, and clinical expertise.
For pharmacogenomic applications, PharmCAT generates reports with guideline recommendations from CPIC and DPWG. These recommendations are intended to support clinical decision-making, but they must be interpreted by qualified healthcare professionals in the context of the individual patient [<a href="#ref-6">6</a>].
The dissemination and storage phase of variant interrogation emphasizes transparent reporting and data sharing. When you report variant findings, you should clearly state the limitations of your analysis, including the annotation tools and databases used, the filtering criteria applied, and the evidence supporting each interpretation [<a href="#ref-5">5</a>].
If your analysis produces a variant with potential clinical significance, you should escalate the finding to a qualified professional. This escalation is appropriate when:
- A variant has a ClinVar classification of pathogenic or likely pathogenic
- A variant has a high predicted impact in a gene with known disease association
- A pharmacogenomic variant has a guideline recommendation that could affect drug therapy
- A variant is novel but has strong functional evidence suggesting disease relevance
The decision to escalate should be documented, and the evidence supporting the escalation should be recorded.
Professional Escalation Criteria
You should seek additional expertise or escalate findings when:
- You identify a variant with a ClinVar classification of pathogenic or likely pathogenic in a gene relevant to the phenotype under study
- You identify a pharmacogenomic variant with a CPIC or DPWG guideline recommendation that could affect patient care
- Your analysis produces conflicting evidence, such as a variant with a high predicted impact but a high population frequency
- You are uncertain about the interpretation of a variant with potential clinical significance
- Your findings need to be reported to a clinical team or research oversight committee
For research settings, escalation may involve consulting with a bioinformatics specialist, a clinical geneticist, or a domain expert. For clinical settings, escalation involves reporting findings through the appropriate clinical channels.
A Practical Decision Framework for Variant Annotation
The Annotation Triage System
After you have run SnpEff or VEP and generated annotated output, the next challenge is deciding which variants deserve deeper investigation. A structured triage system prevents two common failure modes: spending hours on variants that will never pass biological scrutiny and missing a high-priority variant buried in a large output file. The framework below adapts the four-phase approach described in the variant interrogation literature, which emphasizes planning, gathering resources, filtering and validation, and dissemination [<a href="#ref-5">5</a>].
The triage system uses three tiers. Tier 1 variants have strong evidence from clinical databases, clear functional impact, and supportive population frequency data. These variants warrant immediate documentation and possible escalation. Tier 2 variants have partial evidence, such as a high-impact consequence in a gene with no clinical annotation or a moderate-impact variant in a known disease gene. These variants require additional evidence gathering before a decision. Tier 3 variants have weak or conflicting evidence and should be deprioritized unless new information emerges.
Building Your Annotation Decision Matrix
Create a decision matrix before you begin filtering. This matrix converts your research question into explicit inclusion and exclusion criteria. For a germline rare disease study, the matrix might include:
- Population frequency below one percent in gnomAD
- Consequence term of missense, frameshift, stop gained, splice donor, or splice acceptor
- Gene with known association to the phenotype of interest
- Segregation with affected status in family members if pedigree data is available
For a somatic cancer study, the matrix shifts to include:
- Variant allele fraction above the threshold established during variant calling
- Presence in a curated cancer gene list
- Evidence of recurrence across samples in a cohort
- Predicted functional impact from multiple prediction tools
The Variant Prioritization Ordering Tool (VPOT) formalizes this process by allowing you to create a single customizable pathogenicity ranking score from any number of annotation values, each with a user-defined weighting. VPOT also provides filtering based on a candidate gene list or affected status in a family pedigree, which is useful when analyzing entire cohorts [<a href="#ref-7">7</a>].
Step-by-Step Triage Workflow
Step 1: Separate Variants by Clinical Annotation Status
Start by partitioning your annotated VCF into three groups based on clinical database entries. The first group contains variants with a ClinVar classification of pathogenic or likely pathogenic. The second group contains variants with conflicting or uncertain classifications. The third group contains variants with no clinical annotation at all.
VEP reports phenotype associations from databases such as ClinVar, so this partition is straightforward when you have included clinical sources in your annotation run [<a href="#ref-1">1</a>]. For the first group, document the clinical assertion and move these variants to your high-priority list. For the second group, note the conflicting classifications and investigate the underlying evidence. For the third group, proceed to the next step.
Step 2: Apply Consequence and Impact Filters
Filter the unannotated and uncertain variants by consequence term and impact category. Retain variants with high or moderate predicted impact for further analysis. These include missense variants, frameshift variants, stop gained, start lost, splice donor and acceptor variants, and inframe insertions or deletions.
Variants with low or modifier impact, such as synonymous variants and intronic variants, generally do not warrant immediate investigation unless they fall in a known regulatory region or have other supporting evidence. The consequence terms assigned by VEP and SnpEff describe the predicted molecular effect, and the impact category provides a high-level summary of severity [<a href="#ref-1">1</a>].
Step 3: Apply Population Frequency Thresholds
Apply your population frequency threshold to the remaining variants. For rare germline disease analysis, remove variants with a population frequency above your chosen threshold. The threshold depends on disease prevalence and inheritance model. A one percent threshold is common for rare Mendelian disorders, but you should adjust based on your specific question.
For somatic analysis, population frequency helps distinguish rare germline variants from true somatic mutations. A variant that is common in the general population is unlikely to be a somatic mutation, though it could represent a germline variant in the tumor sample.
Step 4: Evaluate Gene Context
For the variants that survive the first three steps, evaluate the gene context. A high-impact consequence in a gene with no known disease association is less informative than a moderate-impact consequence in a well-characterized disease gene. Check whether the gene has established links to your phenotype of interest using resources such as NCBI gene records and clinical databases [<a href="#ref-2">2</a>].
Consider the biological plausibility of the gene in the context of your study. Does the gene have a function that could explain the observed phenotype? Is the gene expressed in the relevant tissue? Are there animal models or functional studies that support a role for this gene?
Step 5: Integrate Prediction Scores
For missense variants that pass the first four steps, integrate prediction scores from tools such as SIFT and CADD. VEP can report these predictions as part of its annotation output [<a href="#ref-1">1</a>]. Use these scores as supporting evidence, not as standalone decision criteria.
When prediction tools disagree, examine the underlying features that drive each prediction. A variant that is predicted deleterious by conservation-based tools but tolerated by structure-based tools may have different biological implications than a variant with concordant predictions across all tools.
Step 6: Document and Escalate
For variants that reach the top of your priority list, document the evidence supporting their significance. Record the consequence term, population frequency, clinical annotation, gene context, and prediction scores. This documentation supports the dissemination and storage phase of variant interrogation, which emphasizes transparent reporting and data sharing [<a href="#ref-5">5</a>].
Escalate findings to a qualified professional when you identify a variant with a ClinVar classification of pathogenic or likely pathogenic, a pharmacogenomic variant with a guideline recommendation that could affect drug therapy, or a variant with strong evidence suggesting clinical significance.
Record System for Annotation Decisions
Maintain a structured record of your triage decisions. A spreadsheet or table with the following columns supports reproducibility and audit:
| Variant ID | Consequence | Gene | gnomAD AF | ClinVar | CADD | SIFT | Triage Tier | Decision | Rationale |
|---|---|---|---|---|---|---|---|---|---|
| chr1:123456A>G | missense | BRCA1 | 0.0001 | Pathogenic | 25.3 | Deleterious | Tier 1 | Escalate | ClinVar pathogenic, low frequency, high impact |
| chr2:789012C>T | synonymous | TP53 | 0.45 | Benign | 3.2 | Tolerated | Tier 3 | Deprioritize | Common, no functional impact |
| chr3:345678G>A | missense | NOVEL | 0.0002 | None | 18.7 | Deleterious | Tier 2 | Investigate | Low frequency, high impact, unknown gene |
Record the annotation tool version, database version, and all filtering parameters in your analysis notes. This record ensures that your triage decisions can be reproduced and audited. The Carpentries lessons provide foundational training in computing, data, shell, Git, and programming that supports reproducible analysis practices [<a href="#ref-9">9</a>].
Troubleshooting the Triage Process
Problem: Too Many Tier 1 Variants
If your triage produces an unmanageable number of high-priority variants, your inclusion criteria are too permissive. Tighten the population frequency threshold, require clinical annotations from multiple sources, or add a minimum CADD score. For somatic analysis, require a minimum variant allele fraction or evidence of recurrence across samples.
Problem: No Variants Survive Filtering
If your triage eliminates all variants, your criteria are too strict. Loosen the population frequency threshold, expand the consequence term list, or include genes with weaker but plausible disease associations. Re-examine your input VCF for quality issues that may have removed true variants during variant calling.
Problem: Conflicting Evidence Across Sources
When clinical annotations, prediction scores, and population frequencies disagree, document the conflict and investigate the underlying evidence. A variant with a high population frequency but a pathogenic ClinVar classification may represent a common variant with incomplete penetrance. A variant with a high CADD score but a benign clinical classification may be a false positive prediction.
Problem: Annotation Output Does Not Match Expectations
If your annotation output contains unexpected patterns, such as a high proportion of unannotated variants or incorrect gene assignments, check your reference assembly and database versions. A VCF built against one assembly cannot be annotated correctly with tools configured for another. Verify that your annotation tool and database versions match your input VCF.
Integrating the Triage System with Existing Workflows
The triage system integrates with the broader variant interrogation framework described in the literature. The planning phase establishes your research question and experimental design. The gathering resources phase assembles the tools, reference data, and annotation sets required for analysis. The filtering and validation phase executes the triage system described here. The dissemination and storage phase ensures findings are reproducible and accessible through transparent reporting and data sharing [<a href="#ref-5">5</a>].
For pharmacogenomic applications, the triage system adapts to the specific requirements of PharmCAT. PharmCAT incorporates patient genotypes, annotates pharmacogenomic information including allele, genotype, and phenotype, and generates a report with guideline recommendations from CPIC and DPWG. The VCF Preprocessor can prepare biobank-scale data for analysis, and the triage system helps prioritize which pharmacogenomic findings warrant clinical attention [<a href="#ref-6">6</a>].
For researchers using community pipelines, the nf-core documentation describes standards for reproducible workflow execution that complement the triage system. These standards support consistent annotation and filtering across samples and studies [<a href="#ref-8">8</a>]. The Galaxy Training Network offers accessible workflow training that can help you implement the triage system in a reproducible environment [<a href="#ref-4">4</a>].
Common Mistakes in the Triage Process
Mistake: Skipping the Clinical Annotation Partition
Jumping directly to consequence and population frequency filters can cause you to miss a variant with a pathogenic ClinVar classification but a moderate consequence term. Always partition by clinical annotation status first.
Mistake: Applying the Same Thresholds to Germline and Somatic Variants
Germline and somatic analyses have different biological contexts and require different filtering strategies. A population frequency threshold appropriate for rare germline disease may be too strict for somatic analysis, where the goal is to distinguish true somatic mutations from germline variants and artifacts.
Mistake: Treating Prediction Scores as Definitive
Prediction scores from tools like SIFT and CADD are supporting evidence, not definitive answers. Different tools can disagree on the same variant, and predictions are only as good as the training data and features used to build them. Use prediction scores as one line of evidence among many.
Mistake: Failing to Document Triage Decisions
Without documentation, your triage decisions cannot be reproduced or audited. Record the evidence supporting each decision, including the annotation tool version, database version, and filtering parameters. This documentation supports the dissemination and storage phase of variant interrogation [<a href="#ref-5">5</a>].
When to Seek Additional Expertise
You should seek additional expertise or escalate findings when:
- You identify a variant with a ClinVar classification of pathogenic or likely pathogenic in a gene relevant to the phenotype under study
- You identify a pharmacogenomic variant with a CPIC or DPWG guideline recommendation that could affect patient care
- Your triage produces conflicting evidence, such as a variant with a high predicted impact but a high population frequency
- You are uncertain about the interpretation of a variant with potential clinical significance
- Your findings need to be reported to a clinical team or research oversight committee
For research settings, escalation may involve consulting with a bioinformatics specialist, a clinical geneticist, or a domain expert. For clinical settings, escalation involves reporting findings through the appropriate clinical channels. The decision to escalate should be documented, and the evidence supporting the escalation should be recorded.
Frequently Asked Questions
What is the difference between variant calling and variant annotation?
Variant calling identifies the genetic variants present in a sample by comparing sequencing reads to a reference genome. The output is a VCF file with genomic coordinates, alleles, and quality scores. Variant annotation adds biological context to those variants by mapping them to genes, transcripts, and protein changes, and by adding information from clinical and population databases. Variant calling answers the question of what variants are present. Variant annotation answers the question of what those variants mean.
Do I need to annotate both germline and somatic variants?
Yes, both germline and somatic variants benefit from annotation, but the annotation priorities differ. Germline analysis focuses on rare variants with high predicted impact that may cause inherited disease. Somatic analysis focuses on variants that drive cancer, which requires distinguishing true somatic mutations from germline variants and sequencing artifacts. The same annotation tools work for both, but the filtering and interpretation strategies differ.
Which annotation tool should I start with?
Start with the Ensembl VEP web interface if you have a small number of variants and want a comprehensive annotation with clinical and population data. The web interface is user-friendly and requires no installation. If you have a large VCF file or need to annotate many samples, move to the VEP command-line tool. SnpEff is a good alternative if you need a fast, lightweight annotation and plan to add clinical and population data separately.
How do I know which reference genome assembly my VCF uses?
Check the header of your VCF file. The header contains information about the reference genome, usually in lines starting with ##reference or ##contig. You can also check the documentation from your variant calling pipeline. If you are unsure, compare a few known variant positions to the expected positions in each assembly. Using the wrong assembly is a common source of annotation errors.
What does a consequence term like missense variant actually mean?
A missense variant is a single nucleotide change that results in a different amino acid being incorporated into the protein. The consequence term describes the molecular effect of the variant on the protein product. Other common terms include synonymous variant, which does not change the amino acid, and frameshift variant, which alters the reading frame and usually produces a nonfunctional protein. The consequence term is a starting point for interpretation, not a final answer about disease relevance.
How should I use population frequency data in variant filtering?
Population frequency data tells you how common a variant is in the general population. For germline analysis of rare diseases, you typically remove variants that are common in the population, because a common variant is unlikely to cause a rare disease. The specific frequency threshold depends on your disease and inheritance model. For somatic analysis, population frequency helps distinguish rare germline variants from true somatic mutations.
What is the role of prediction scores like SIFT and CADD?
Prediction scores estimate whether a variant is likely to damage protein function. SIFT uses sequence conservation to predict whether an amino acid change is tolerated. CADD combines multiple annotations into a single deleteriousness score. These scores provide supporting evidence when clinical and population data are insufficient, but they are not definitive. Different tools can disagree, and predictions can be wrong. Use prediction scores as one line of evidence among many.
When should I escalate a variant finding to a qualified professional?
Escalate when you identify a variant with a ClinVar classification of pathogenic or likely pathogenic, a pharmacogenomic variant with a guideline recommendation that could affect drug therapy, or a variant with strong evidence suggesting clinical significance. Also escalate when you are uncertain about the interpretation of a potentially significant variant. Document the evidence supporting the escalation and report through the appropriate channels for your research or clinical setting.
Related Bioinformatics Guides
- Functional Annotation of Metagenomes: A Guide to Databases and Pipelines
- Deep Learning for Annotating Structural Variants in Viral Genomes
- Lipidomic Analysis: A Beginner's Guide to Workflows and Data Interpretation
- Single-Cell Annotation: A Workflow for Cell Type Identification
- Proteomics Analysis Tools: A Comparative Guide for Functional Interpretation
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
[1] [Annotating and prioritizing genomic variants using the Ensembl Variant Effect Predictor-A tutorial.](https://pubmed.ncbi.nlm.nih.gov/34816521). Human mutation, 2022. [2] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information. [3] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute. [4] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project. [5] [Tutorial for variant interrogation in tumor samples.](https://pubmed.ncbi.nlm.nih.gov/41701771). PLoS computational biology, 2026. [6] [How to Run the Pharmacogenomics Clinical Annotation Tool (PharmCAT).](https://pubmed.ncbi.nlm.nih.gov/36350094). Clinical pharmacology and therapeutics, 2023. [7] [VPOT: A Customizable Variant Prioritization Ordering Tool for Annotated Variants.](https://pubmed.ncbi.nlm.nih.gov/31765830). Genomics, proteomics & bioinformatics, 2019. [8] [nf-core Documentation](https://nf-co.re/docs). nf-core. [9] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.