A Practical Guide to Using VEP for Variant Annotation: Command Line and Web Interface
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- VEP annotates VCF files by predicting molecular consequences (e.g., missense_variant, frameshift_variant), impact levels (HIGH, MODERATE, LOW, MODIFIER), and integrating data from databases like ClinVar (clinical significance) and gnomAD (allele frequencies).
- The command-line interface offers granular control over annotation sources (e.g.,
--af_gnomad,--clin_sig,--sift,--cadd) and filtering (e.g.,--filter_common,--max_af), essential for reproducible pipelines and large datasets. - The web interface provides an accessible, installation-free option for beginners and small datasets, offering visual output and interactive filtering, but is unsuitable for whole-genome or large exome analyses.
- Successful VEP implementation requires meticulous attention to matching the
--assemblyoption with downloaded cache files (e.g., GRCh38, GRCh37) to prevent coordinate mismatch errors and ensure accurate annotation. - Performance optimization for large datasets relies heavily on using local cache files (
--cache,--offline) and parallel processing (--fork), while restricting annotation sources to only those relevant to the specific biological question. - Reproducibility is paramount; document the exact VEP version, Ensembl release, cache version, and all command-line options used, and leverage
--stats_filefor summary statistics and--verbosefor detailed logging.
Researchers analyzing genomic variants face a common problem: a variant call format (VCF) file lists genomic positions and alleles, but it does not explain what those changes mean biologically. The Ensembl Variant Effect Predictor (VEP) addresses this gap by annotating variants with molecular consequences, allele frequencies, phenotype associations, and deleteriousness predictions. This guide explains how to run VEP through both the command line and web interface, with concrete options, output formats, and troubleshooting steps for variant calling workflows.
VEP is a freely available, open-source tool for the annotation and filtering of genomic variants. It predicts variant molecular consequences using the Ensembl/GENCODE or RefSeq gene sets, reports phenotype associations from databases such as ClinVar, allele frequencies from studies including gnomAD, and predictions of deleteriousness from tools such as Sorting Intolerant From Tolerant (SIFT) and Combined Annotation Dependent Depletion (CADD). VEP includes filtering options to customize variant prioritization and is updated roughly quarterly to incorporate the latest gene, variant, and phenotype association information. Analysis can be performed using a highly configurable command-line tool, a Representational State Transfer (REST) application programming interface, and a web interface, each designed to suit different levels of bioinformatics experience and different needs in terms of data size, visualization, and flexibility.
At a Glance
The table below summarizes the three primary ways to access VEP and the situations where each approach is most appropriate.
| Access Method | Best For | Input Size | Key Advantage | Main Limitation |
|---|---|---|---|---|
| Web Interface | Beginners, small datasets, exploratory analysis | Up to a few thousand variants per submission | No installation required, visual output, interactive filtering | Not suitable for whole-genome or large exome datasets |
| Command Line Tool | Batch processing, large datasets, reproducible pipelines | Hundreds of thousands to millions of variants | Full configuration options, scriptable, cache-based speed | Requires Linux environment and installation |
| REST API | Programmatic access, integration into custom scripts | Moderate volumes with rate limits | Language-agnostic, no local installation | Dependent on remote server availability and API limits |
Understanding VEP Annotation Output
VEP produces a tab-delimited output file by default, with one line per variant transcript. Each line contains core columns that describe the variant and its predicted effect. Understanding these columns is essential before you design your annotation strategy.
Core Consequence Fields
The most important output column is "Consequence," which uses Sequence Ontology terms to describe the molecular effect of a variant. Common terms include missense_variant, synonymous_variant, frameshift_variant, stop_gained, splice_acceptor_variant, and regulatory_region_variant. The "IMPACT" column provides a high-level classification: HIGH, MODERATE, LOW, or MODIFIER. This classification helps you prioritize variants for downstream analysis.
The "SYMBOL" column gives the gene name, while "Gene" provides the Ensembl gene identifier. "Feature" identifies the specific transcript, and "BIOTYPE" indicates whether the transcript is protein-coding, non-coding, or another type. The "cDNA_position" and "CDS_position" columns show where the variant falls within the transcript sequence, and "Protein_position" indicates the affected amino acid position.
Additional Annotation Columns
When you enable additional plugins and databases, VEP appends more columns. The "Existing_variation" column lists known variant identifiers such as rs numbers. "CLIN_SIG" reports clinical significance from ClinVar. "gnomAD_AF" and related columns provide allele frequencies from the Genome Aggregation Database. "SIFT" and "PolyPhen" columns give deleteriousness predictions, and "CADD_PHRED" provides a scaled CADD score.
Installing VEP for Command Line Use
The command-line version of VEP requires a Linux operating system with Perl installed. The installation process involves downloading the VEP code, installing dependencies, and optionally downloading cache files for faster annotation.
System Requirements and Dependencies
VEP requires Perl version 5.10 or higher, along with several Perl modules. The Ensembl VEP installation script, called install.pl, checks for these dependencies and installs missing modules automatically. You also need the Ensembl API, which the installer fetches from the Ensembl Git repository. For full functionality, including plugin support, you need additional Perl modules such as BioPerl, DBI, and JSON.
The installation script supports several options. The --AUTO flag with the letter c installs the core API and VEP script. Adding the letter f installs the cache files for a specified species and release. The --CACHE_VERSION option lets you specify which Ensembl release to use. The --SPECIES option restricts cache downloads to particular species, which saves disk space and installation time.
Downloading Cache Files
VEP can query the Ensembl databases directly over the internet, but this approach is slow for large datasets. The recommended approach is to download cache files, which store gene, transcript, and variant data locally. Cache files are organized by Ensembl release and species. For human data, you typically download the Homo_sapiens cache, which includes the GRCh37 and GRCh38 assemblies.
The cache download command uses the install.pl script with the --AUTO f option. You can specify the assembly with --ASSEMBLY GRCh38 and the species with --SPECIES homo_sapiens. The cache files are compressed and require approximately 20 to 30 gigabytes of disk space for human data, depending on the release and whether you include regulatory build data.
Verifying the Installation
After installation, you can verify that VEP works by running a simple test. The command ./vep --help displays the full list of options. A more practical test involves running VEP on a small test VCF file. The Ensembl team provides example files in the examples directory of the VEP installation. Running ./vep -i examples/example.vcf -o test_output.txt --cache --offline should produce an annotated output file without errors.
Running VEP from the Command Line
The command-line interface gives you full control over annotation options. The basic command structure is vep -i input.vcf -o output.txt --cache --offline. The --cache flag tells VEP to use local cache files, and --offline prevents VEP from attempting to connect to the Ensembl databases for features not available in the cache.
Essential Input and Output Options
The -i option specifies the input file, which can be in VCF, BED, GFF, or FASTA format. The -o option specifies the output file. If you omit -o, VEP writes to standard output. The --format option lets you explicitly specify the input format, though VEP usually detects it automatically.
For VCF input, VEP preserves the original VCF information and adds annotation data. The --vcf option produces output in VCF format, which is useful when you want to keep the annotated variants in a format compatible with downstream tools. The default output is tab-delimited text.
Species and Assembly Selection
The --species option specifies the species, using the scientific name with an underscore, such as homo_sapiens or mus_musculus. The --assembly option specifies the genome assembly, such as GRCh38 or GRCm39. These options must match the cache files you downloaded. Using mismatched species or assembly options produces errors or incorrect annotations.
Transcript and Gene Set Selection
VEP uses the Ensembl/GENCODE gene set by default. The --refseq option switches to the RefSeq gene set, which some researchers prefer for consistency with NCBI-based analyses. The --gencode_basic option restricts annotation to the GENCODE basic gene set, which includes only the primary transcripts for each gene. This option reduces output complexity when you are interested in protein-coding effects only.
Frequency and Clinical Database Options
The --af, --af_gnomad, --af_esp, and --af_1kg options add allele frequency data from gnomAD, the Exome Sequencing Project, and 1000 Genomes respectively. These options require the corresponding cache files or internet access. The --clin_sig option adds clinical significance annotations from ClinVar. The --pubmed option adds PubMed identifiers for publications associated with the variant.
Deleteriousness Prediction Options
VEP can add predictions from multiple deleteriousness tools. The --sift and --polyphen options add predictions from SIFT and PolyPhen, which are based on protein sequence conservation and structural properties. The --cadd option adds CADD scores, which integrate multiple annotations into a single deleteriousness metric. These options require the corresponding data files, which you download separately from the respective tool websites.
Plugin Usage
VEP plugins extend the core functionality with additional annotations. The --plugin option loads a plugin, and you can pass parameters to the plugin in parentheses. For example, --plugin LoFtool adds a loss-of-function intolerance score. The --plugin dbNSFP adds scores from the dbNSFP database, which includes multiple prediction tools. Plugins require the plugin Perl module and often require additional data files.
Using the VEP Web Interface
The VEP web interface provides the same core annotation functionality without requiring installation. You access it through the Ensembl website, where you can paste variants, upload files, or provide a URL to a VCF file.
Input Methods and Format Requirements
The web interface accepts several input formats. You can paste variants directly into a text box, with one variant per line in a format such as chr 1 123456 A G. You can upload a VCF file, a BED file, or a tab-delimited file. The interface also accepts a URL pointing to a remotely hosted file. For large files, the web interface recommends using the command-line tool or REST API.
Configuring Web Interface Options
The web interface presents a form with options organized into sections. You select the species and assembly, choose the gene set, and specify which additional annotations to include. The interface includes checkboxes for allele frequencies, clinical significance, and deleteriousness predictions. You can also enable plugins, though the selection is more limited than the command-line version.
Submitting Jobs and Retrieving Results
When you submit a job, the web interface assigns a job ID and displays a results page. Small jobs complete within seconds, while larger jobs may take several minutes. The results page displays a summary table with the annotated variants, and you can download the full results as a text file, VCF file, or Excel-compatible format. The interface also provides a link to view results in the Ensembl browser for individual variants.
Web Interface Limitations
The web interface has practical limits on input size. Submitting more than a few thousand variants at once may cause the job to fail or time out. The interface is not designed for whole-genome or large exome datasets. For those applications, the command-line tool or REST API is more appropriate. The web interface also provides less control over output formatting and plugin configuration.
Integrating VEP into Variant Calling Workflows
VEP is typically one step in a larger variant analysis pipeline. Understanding where VEP fits in the workflow helps you design efficient and reproducible analyses.
Position in the Variant Calling Pipeline
A standard germline variant calling workflow proceeds from raw sequencing reads through alignment, variant calling, and variant filtering. VEP annotation occurs after variant calling and before final variant prioritization. The input to VEP is a VCF file produced by a variant caller such as GATK HaplotypeCaller, FreeBayes, or Samtools mpileup. The annotated output feeds into downstream filtering and interpretation steps.
For somatic variant calling in tumor samples, the workflow includes additional steps for comparing tumor and normal samples. The variant caller produces a VCF file with somatic mutations, which VEP annotates with the same consequence and frequency information. The annotated output supports the filtering and validation steps that prioritize meaningful variants for clinical or research follow-up.
Preparing Input Files for VEP
Before running VEP, you should ensure your VCF file is properly formatted. The VCF header must include the ##fileformat=VCFv4.2 line and appropriate contig definitions. Variant records must include chromosome, position, reference allele, and alternate allele fields. VEP can handle multi-allelic variants, but it may split them into separate records depending on the options you use.
The --check_existing option checks whether your variants have been previously identified and reported in dbSNP or other databases. This option adds the "Existing_variation" column to the output. The --check_alleles option performs a more stringent check that requires the alleles to match exactly.
Filtering Variants Before and After Annotation
VEP includes filtering options that help you prioritize variants. The --filter option applies a filter expression to the annotated output, keeping only variants that match the criteria. For example, --filter "IMPACT is HIGH" keeps only high-impact variants. The --filter_common option removes variants with allele frequencies above a specified threshold, which is useful for removing common polymorphisms from rare variant analyses.
The --max_af option sets a maximum allele frequency for variants to be included in the output. This option is particularly useful for rare variant studies, where you want to exclude common variants. The --af_gnomad option must be enabled for the --max_af filter to work with gnomAD frequencies.
Reproducibility Considerations
Reproducible variant annotation requires documenting the VEP version, Ensembl release, cache version, and all options used. The --verbose option prints detailed information about the VEP version and configuration to the log file. The --stats_file option writes a statistics file that summarizes the annotation results, including counts of variants by consequence type.
For pipeline reproducibility, you should record the exact VEP command in your analysis documentation. The Ensembl release number and cache version must match between the VEP script and the cache files. Using different versions produces different annotations, which can affect downstream interpretation.
Common VEP Options and Their Effects
The table below lists frequently used VEP options and their practical effects on annotation output.
| Option | Effect on Output | Use Case |
|---|---|---|
--vcf | Output in VCF format instead of tab-delimited text | Downstream tools that require VCF input |
--refseq | Use RefSeq gene set instead of Ensembl/GENCODE | Consistency with NCBI-based analyses |
--gencode_basic | Restrict to primary transcripts only | Reducing output complexity |
--af_gnomad | Add gnomAD allele frequencies | Rare variant filtering |
--clin_sig | Add ClinVar clinical significance | Clinical variant interpretation |
--sift and --polyphen | Add deleteriousness predictions | Variant prioritization |
--filter_common | Remove common variants | Rare disease studies |
--plugin | Add plugin-specific annotations | Custom annotation needs |
Troubleshooting Common VEP Errors
VEP installation and execution can produce errors that are confusing for new users. Understanding the most common failure patterns helps you diagnose and fix problems quickly.
Cache and Assembly Mismatch Errors
A frequent error occurs when the cache files do not match the assembly or species specified in the command. VEP reports an error such as "No cache found for species homo_sapiens and assembly GRCh37." This error means you downloaded the GRCh38 cache but requested GRCh37 assembly, or you did not download the cache at all. The solution is to download the correct cache files or change the --assembly option to match your cache.
Missing Perl Module Errors
VEP requires several Perl modules that may not be installed on your system. The installation script usually installs these modules automatically, but system-level restrictions can prevent installation. The error message lists the missing module name. You can install the module manually using CPAN or your system package manager. The --NO_UPDATE option prevents the installer from attempting to update modules, which can help in restricted environments.
Memory and Resource Errors
Large input files can exhaust system memory, especially when you enable multiple annotation sources. VEP reports an "Out of memory" error or the process is killed by the operating system. The --fork option enables parallel processing, which distributes the workload across multiple CPU cores. The --buffer_size option controls the number of variants processed at once, and reducing this value can lower memory usage.
Plugin Data File Errors
Plugins that require external data files produce errors when the data files are missing or incorrectly formatted. The error message typically indicates which file is missing. You must download the data file from the plugin documentation and specify its path in the plugin parameters. For example, the dbNSFP plugin requires the dbNSFP database file, which you download from the dbNSFP website.
Performance Optimization for Large Datasets
Whole-genome and large exome datasets contain millions of variants, and VEP annotation can take hours without optimization. Several strategies reduce runtime and resource usage.
Using Cache Files and Offline Mode
The --cache and --offline options are the most important performance optimizations. Querying the Ensembl databases over the internet for each variant is extremely slow. Local cache files provide near-instant access to gene and transcript data. The --offline option prevents VEP from making any internet connections, which also improves reliability in environments with restricted network access.
Parallel Processing with Fork
The --fork option enables parallel processing by splitting the input into chunks and processing them simultaneously. The number after --fork specifies the number of processes to run. For example, --fork 4 uses four processes. The optimal number depends on your CPU core count and available memory. Using more processes than available cores can slow down the analysis due to context switching.
Restricting Annotation Sources
Each additional annotation source adds runtime. If you do not need a particular annotation, omit the corresponding option. For example, if you are working with a non-model organism, you can omit gnomAD and ClinVar options, which are human-specific. The --regulatory option adds regulatory region annotations, which are useful for non-coding variants but increase runtime.
Using the REST API for Moderate Workloads
The REST API provides a middle ground between the web interface and the command-line tool. You can send HTTP requests to the Ensembl REST server with your variants, and the server returns annotations in JSON or VCF format. The REST API is suitable for moderate workloads and integrates well with custom scripts. However, the API has rate limits, and large submissions may be rejected.
Quality Control and Validation of Annotations
VEP annotation results should be validated before you use them for downstream analysis. Several quality checks help you identify problems in the annotation process.
Reviewing the Statistics File
The --stats_file option writes a summary of the annotation run, including the total number of variants processed, the number of variants with no annotation, and the distribution of consequence types. Reviewing this file helps you identify unexpected patterns. For example, a high proportion of variants with no annotation may indicate a mismatch between your input coordinates and the assembly used by VEP.
Checking Known Variants
The --check_existing option adds known variant identifiers to your output. You can validate your annotation by checking that well-known variants receive the expected annotations. For example, a known pathogenic variant in BRCA1 should receive a HIGH impact consequence and a ClinVar pathogenic classification. If known variants do not receive expected annotations, your VEP version or cache may be outdated.
Verifying Strand and Allele Orientation
VEP reports variants on the forward strand of the reference genome. If your input file uses a different strand convention, the annotations will be incorrect. The --strand option controls how VEP reports strand information. You should verify that your input VCF file uses the same reference genome and allele orientation as the VEP cache.
Comparing with Independent Annotations
For critical variants, you can compare VEP annotations with those from independent sources. The NCBI databases provide variant information that can be cross-checked with VEP output. Discrepancies may indicate differences in gene sets, transcript versions, or annotation algorithms. Understanding these differences is important for interpreting conflicting results.
Limitations and Interpretation Boundaries
VEP provides powerful annotations, but it has limitations that affect interpretation. Understanding these boundaries prevents overinterpretation of results.
Transcript and Gene Set Dependencies
VEP annotations depend on the gene set and transcript versions in the Ensembl or RefSeq databases. Different gene sets may produce different consequences for the same variant. The Ensembl/GENCODE gene set includes more transcripts than the GENCODE basic set, which can produce multiple annotations for a single variant. RefSeq annotations may differ from Ensembl annotations for genes with complex structures.
Prediction Tool Limitations
Deleteriousness predictions from SIFT, PolyPhen, and CADD are computational predictions, not experimental validations. These tools have known false positive and false negative rates. A variant predicted to be deleterious by SIFT may have no functional effect, and a variant predicted to be benign may be pathogenic. These predictions should be used for prioritization, not as definitive evidence of pathogenicity.
Frequency Database Limitations
Allele frequencies from gnomAD and other databases reflect the populations included in those studies. Frequencies may not be representative of all populations, and rare variants in one population may be common in another. The absence of a variant from frequency databases does not prove that it is rare, as databases have incomplete coverage of global genetic diversity.
Clinical Significance Interpretation
ClinVar annotations reflect the clinical interpretations submitted by laboratories and researchers. These interpretations can change over time as new evidence emerges. A variant classified as pathogenic in one submission may be reclassified as benign in a later submission. You should check the ClinVar submission details and review dates when interpreting clinical significance annotations.
Professional Escalation Criteria
VEP annotation results often inform important decisions in research and clinical contexts. Knowing when to escalate concerns to a supervisor, bioinformatics specialist, or clinical geneticist is essential.
When to Consult a Bioinformatics Specialist
You should consult a bioinformatics specialist when VEP produces unexpected results that you cannot explain. Examples include a high proportion of variants with no annotation, systematic errors in consequence assignment, or performance problems that persist after optimization. A specialist can help diagnose cache mismatches, plugin configuration issues, and pipeline integration problems.
When to Consult a Clinical Geneticist
For clinical applications, you should consult a clinical geneticist when VEP annotations suggest a potentially pathogenic variant that could affect patient care. The geneticist can review the variant in the context of the patient's clinical presentation, family history, and other genetic findings. The geneticist can also order confirmatory testing and provide genetic counseling.
When to Update Your VEP Installation
You should update VEP and cache files when a new Ensembl release becomes available, approximately every three months. Updates include new gene models, variant annotations, and phenotype associations. If you are working on a project that requires the latest annotations, you should plan for regular updates. However, updating mid-project can introduce inconsistencies, so you should document the VEP version used for each analysis.
Training and Learning Resources
VEP is a complex tool, and ongoing learning is important for effective use. Several official resources provide training and documentation.
Ensembl and EMBL-EBI Training
The EMBL-EBI training portal offers courses and tutorials on bioinformatics topics, including variant annotation. These resources provide structured learning pathways for researchers at different skill levels. The Ensembl website also provides detailed documentation for VEP, including the full option list and example commands.
Galaxy Training Network
The Galaxy Training Network provides interactive tutorials for bioinformatics workflows, including variant calling and annotation. These tutorials use the Galaxy platform, which provides a graphical interface for running VEP and other tools. The tutorials include step-by-step instructions and example datasets, making them useful for beginners.
Bioconductor and Community Resources
Bioconductor provides R packages for genomic analysis, including packages that work with VEP output. The Bioconductor documentation includes workflows for variant annotation and filtering. The nf-core community provides standardized pipelines that include VEP as a component, with documentation on configuration and usage.
Foundational Computing Skills
Effective use of the VEP command-line tool requires basic Unix shell skills. The Carpentries lessons provide foundational training in shell, Git, and programming that supports bioinformatics work. These skills help you navigate the command line, manage files, and write scripts for reproducible analysis.
Building a Structured VEP Annotation Decision Framework
Running VEP successfully requires more than knowing individual options. Researchers often struggle because they apply the same annotation strategy to every project without considering how the biological question, data source, and downstream analysis goals should shape the VEP configuration. A structured decision framework helps you select the right options deliberately instead of by habit, document your choices for reproducibility, and avoid common annotation pitfalls that compromise downstream interpretation.
Defining the Annotation Objective
Before you construct any VEP command, write down the specific biological question your analysis must answer. This single step prevents wasted compute time and reduces the risk of misinterpreting results. The question determines which annotation sources matter, which filters apply, and which output format you need.
For a rare disease study searching for causal variants in a small family cohort, the annotation objective centers on identifying high-impact variants with very low population frequency. This objective requires gnomAD allele frequencies, ClinVar clinical significance, and deleteriousness predictions from SIFT or CADD. The --filter_common option becomes essential to remove variants present at appreciable frequency in the general population.
For a population genetics project examining selective pressure across many individuals, the objective shifts toward comprehensive allele frequency reporting across multiple ancestral groups. You need gnomAD subpopulation frequencies, and you may not need clinical significance annotations at all. The filtering strategy differs because common variants are often the variants of interest.
For a somatic cancer study, the objective involves distinguishing driver mutations from passenger mutations. This requires different annotation priorities, including COSMIC identifiers when available, and careful attention to variant allele frequency within the tumor sample. The filtering approach for somatic data differs fundamentally from germline analysis because you cannot simply remove common variants, as some common variants are recurrent somatic mutations in cancer.
Write your objective as a single sentence before configuring VEP. For example, "Identify rare protein-altering variants in the exome that segregate with affected status in this family." This sentence directly maps to specific VEP options and filters.
Selecting the Annotation Tier
Once you have defined the objective, assign your analysis to one of three annotation tiers. Each tier represents a different balance between annotation depth and computational cost.
Tier 1 is the minimal annotation set for exploratory analysis or initial quality checks. It includes the core consequence fields, gene symbols, and transcript identifiers. You run VEP with the default options plus --symbol and --check_existing. This tier is appropriate when you are verifying that your VCF file is properly formatted or when you need a quick overview of variant types in a dataset.
Tier 2 is the standard annotation set for most research applications. It adds allele frequencies from gnomAD, clinical significance from ClinVar, and at least one deleteriousness prediction. The command includes --af_gnomad, --clin_sig, and --sift or --cadd. This tier suits rare disease studies, candidate gene projects, and most publication-oriented analyses.
Tier 3 is the comprehensive annotation set for clinical interpretation or final variant prioritization. It includes all Tier 2 annotations plus multiple deleteriousness tools, regulatory region annotations, and custom plugins relevant to your disease area. This tier uses options such as --regulatory, --plugin LoFtool, and --plugin dbNSFP. Tier 3 runs take substantially longer and require more disk space for plugin data files, so reserve it for the final shortlist of candidate variants instead of the full dataset.
The tier system creates a natural workflow. Run Tier 1 on your full variant set to verify data quality, apply initial filters to reduce the variant count, then run Tier 2 or Tier 3 on the reduced set. This staged approach saves hours of compute time compared to running comprehensive annotation on every variant from the start.
Matching Input Data Source to VEP Configuration
The origin of your VCF file determines several VEP configuration decisions. A VCF from whole-genome sequencing contains variants in coding and non-coding regions, so you need regulatory region annotations and must consider non-coding consequence types. A VCF from whole-exome sequencing is enriched for coding variants, so you can focus on protein-coding consequences and may not need the --regulatory option.
The variant caller used to produce your VCF also matters. GATK HaplotypeCaller output includes genotype quality fields that VEP preserves in VCF output format. FreeBayes output may contain more multi-allelic sites that require careful handling. Samtools mpileup output has different quality score conventions. Review your VCF header before running VEP to understand what information is already present, so you do not request redundant annotations.
The reference genome build used during alignment and variant calling must match the VEP assembly option exactly. If your reads were aligned to GRCh37 but you run VEP with --assembly GRCh38, every variant coordinate will be wrong and annotations will be meaningless. This mismatch is one of the most common and damaging errors in variant annotation. Confirm the reference build by checking your VCF header for the ##reference line or by examining a known variant position.
Building the Annotation Command Systematically
Construct your VEP command in layers instead of typing a long string of options from memory. This layered approach reduces errors and makes your command easier to document and reproduce.
Start with the input and output specification layer. This includes -i input.vcf, -o output.txt, and the format options. Decide whether you need VCF output for downstream tool compatibility or tab-delimited text for readability.
Add the cache and performance layer. Include --cache, --offline, and --fork with a number appropriate to your system. The --fork setting should match your available CPU cores but leave one core free for system processes. For a machine with eight cores, --fork 7 is a reasonable starting point.
Add the species and assembly layer. Specify --species homo_sapiens and --assembly GRCh38 explicitly even if you believe VEP will detect them automatically. Explicit specification prevents subtle errors when your cache contains multiple species or assemblies.
Add the annotation source layer. Select the frequency databases, clinical databases, and deleteriousness tools that match your annotation tier. Each option in this layer adds runtime, so include only what your objective requires.
Add the filtering layer. Apply --filter_common or --max_af if your objective requires rare variant enrichment. Apply --filter expressions to restrict output to specific consequence types. Document the exact filter expressions you use, as these directly shape your final variant list.
Add the output enhancement layer. Include --stats_file to generate the summary statistics file, --verbose for detailed logging, and --check_existing to add known variant identifiers. These options support quality control and reproducibility.
Recording the Annotation Run
A complete record of your VEP run is essential for reproducibility and for troubleshooting unexpected results. Create a standard annotation log for each VEP run that captures the following information.
Record the VEP version number and the Ensembl release. The command vep --help displays the version at the top of the output. The cache version must match the VEP version, and mismatches produce errors or incorrect annotations.
Record the exact command line, including every option and parameter. Do not rely on memory or on shell history files that may be cleared. Paste the command into your analysis notebook or write it to a text file at the time you run it.
Record the input file checksum and the output file checksum. The md5sum command on Linux computes these values. Checksums let you verify that output files correspond to the exact input files you intended to annotate, which matters when you process many samples in a batch.
Record the runtime and resource usage. The time command reports how long VEP took to run. The --stats_file output includes variant counts by consequence type. These records help you estimate resources for future runs and detect performance regressions after VEP updates.
Record the cache path and the plugin data file paths. If you move your analysis to a different machine, these paths must be recreated. Documenting them now prevents configuration errors later.
Implementing the Decision Framework in Practice
The following workflow applies the decision framework to a typical rare disease exome study. This example shows how the framework translates into concrete commands and decisions.
Define the objective as identifying rare, protein-altering variants that segregate with disease in a family of four. The input is a VCF file with 80,000 variants from whole-exome sequencing aligned to GRCh38.
Run Tier 1 annotation on the full dataset to verify data quality. The command is vep -i family.vcf -o family_tier1.txt --cache --offline --fork 7 --species homo_sapiens --assembly GRCh38 --symbol --check_existing --stats_file family_tier1_stats.html. Review the stats file to confirm the variant count matches expectations and that the consequence distribution looks reasonable.
Apply initial filters to reduce the variant set. Remove variants with gnomAD allele frequency above 0.01 and variants with MODIFIER impact. This step requires a Tier 2 run with --af_gnomad and --filter_common. The command becomes vep -i family.vcf -o family_tier2.txt --cache --offline --fork 7 --species homo_sapiens --assembly GRCh38 --symbol --check_existing --af_gnomad --filter_common --clin_sig --sift --stats_file family_tier2_stats.html.
Review the Tier 2 output and select candidate variants for detailed evaluation. For the final shortlist of 20 to 50 variants, run Tier 3 annotation with additional plugins. The command adds --plugin LoFtool and --plugin dbNSFP with the path to the dbNSFP data file.
Document each run in the annotation log with the version, command, checksums, and runtime. This record supports the methods section of your manuscript and enables another researcher to reproduce your exact analysis.
Common Failure Patterns in the Decision Framework
Several recurring mistakes undermine the decision framework and produce unreliable annotations. Recognizing these patterns helps you avoid them.
The first pattern is applying the same VEP command to every project regardless of the biological question. This habit produces output that either lacks necessary annotations or includes irrelevant ones that complicate interpretation. The decision framework prevents this by forcing you to define the objective before selecting options.
The second pattern is skipping the Tier 1 quality check and running comprehensive annotation immediately. This approach wastes hours of compute time when the input VCF has formatting problems or coordinate mismatches that would be obvious in the Tier 1 stats file. The staged approach catches these problems early.
The third pattern is failing to record the VEP version and cache version. When you return to the analysis months later or when a collaborator asks about your methods, you cannot reproduce the exact annotation environment. The annotation log solves this problem.
The fourth pattern is ignoring the stats file. The stats file contains valuable diagnostic information, including the number of variants with no annotation and the distribution of consequence types. A sudden increase in variants with no annotation may indicate a cache problem or a coordinate mismatch. Review the stats file after every run.
Validation Checks Within the Framework
Build validation steps into your workflow at each tier. These checks confirm that VEP is producing sensible annotations before you invest time in downstream analysis.
At Tier 1, verify that known variants in your dataset receive expected consequences. If your VCF contains a known pathogenic variant in a disease gene, confirm that VEP reports the expected consequence and gene symbol. This check validates that your cache and assembly are correct.
At Tier 2, verify that allele frequencies match expectations for your population. If you are studying a European population, common variants should have gnomAD allele frequencies consistent with known population genetics. Wildly discrepant frequencies may indicate a strand issue or a reference build mismatch.
At Tier 3, verify that multiple prediction tools agree on the most damaging variants. When SIFT, PolyPhen, and CADD all predict a variant is deleterious, you have higher confidence than when the tools disagree. Disagreement between tools is common and should be documented instead of ignored.
The NCBI databases provide an independent source for cross-checking variant annotations. You can look up specific variants in dbSNP or ClinVar through the NCBI website and compare the reported consequences with your VEP output. Discrepancies may reflect differences in transcript versions or gene sets between Ensembl and NCBI resources, which is expected and should be noted in your interpretation.
Adapting the Framework for Somatic Variant Analysis
Somatic variant annotation in tumor samples requires adjustments to the decision framework. The objective differs because you are identifying mutations acquired during cancer development instead of inherited variants.
The filtering strategy changes fundamentally. You cannot apply --filter_common to remove common variants, because some common germline variants are present in the tumor sample and must be distinguished from somatic mutations. The variant caller should have already performed this distinction by comparing tumor and normal samples. Your VEP annotation focuses on the somatic mutations that remain.
The annotation priorities shift toward cancer-specific resources. COSMIC annotations, when available, provide information about recurrent mutations in cancer. The clinical significance annotations from ClinVar remain useful, but you must interpret them in the somatic context instead of the germline context.
The decision framework still applies, but the objective statement changes to something like, "Identify somatic mutations in protein-coding regions that are predicted to be deleterious and have known cancer associations." This objective maps to a different set of VEP options and filters than the germline rare disease example.
The four-phase framework for variant interrogation in tumor samples provides additional structure for somatic analyses. The planning phase establishes the experimental design and sequencing outputs, the gathering resources phase assembles the tools and reference data, the filtering and validation phase prioritizes meaningful variants, and the dissemination and storage phase ensures reproducible reporting. VEP annotation fits within the gathering resources and filtering phases of this larger framework.
Maintaining the Framework Across VEP Updates
VEP releases quarterly updates that include new gene models, variant annotations, and phenotype associations. These updates improve annotation quality but introduce a reproducibility challenge. The decision framework must account for version changes.
When you update VEP and cache files, rerun a small validation set of known variants through the new version. Compare the annotations with the previous version. Most variants will produce identical annotations, but some will change due to updated transcript models or revised clinical classifications. Document these changes in your annotation log.
For active projects, avoid updating VEP mid-analysis. Complete the current analysis with the existing version, then update for the next project. This practice ensures that all variants in a single project receive consistent annotations from the same VEP version.
For long-term projects spanning multiple VEP releases, record the VEP version used for each analysis batch. When you compare results across batches, account for version differences in your interpretation. The annotation log makes this comparison possible.
Frequently Asked Questions
What is the difference between the VEP web interface and the command-line tool?
The web interface requires no installation and provides a visual form for configuring options, making it suitable for small datasets and exploratory analysis. The command-line tool requires installation but supports large datasets, full configuration options, and scriptable reproducible workflows. The command-line tool is the recommended choice for batch processing and pipeline integration.
How do I choose between Ensembl and RefSeq gene sets in VEP?
The Ensembl/GENCODE gene set is the default and includes comprehensive transcript annotations. The RefSeq gene set is curated by NCBI and may be preferred for consistency with NCBI-based analyses. The choice depends on your downstream analysis tools and the gene set used in your reference genome annotation. You should document which gene set you use, as consequences can differ between gene sets.
Why does VEP report multiple consequences for a single variant?
A single variant can affect multiple transcripts of the same gene or transcripts of different genes. VEP reports one line per variant transcript pair, so a variant in a gene with multiple isoforms produces multiple output lines. The --gencode_basic option restricts annotation to primary transcripts, which reduces the number of consequences reported.
How do I add gnomAD allele frequencies to my VEP output?
Use the --af_gnomad option in the command-line tool or check the corresponding box in the web interface. This option requires the gnomAD data to be available in your cache files or accessible over the internet. The allele frequency appears in the "gnomAD_AF" column of the output.
What does the IMPACT column mean in VEP output?
The IMPACT column provides a high-level classification of the variant's predicted effect. HIGH impact includes variants that likely cause loss of function, such as stop_gained and frameshift_variant. MODERATE impact includes missense variants and in-frame indels. LOW impact includes synonymous variants. MODIFIER impact includes variants in non-coding regions or with uncertain effects.
Can I use VEP for non-human species?
Yes, VEP supports many species. You need to download the cache files for your species of interest using the --SPECIES option during installation. The available annotations depend on the species, and some human-specific databases such as gnomAD and ClinVar are not available for other species.
How do I filter VEP output to keep only high-impact variants?
Use the --filter option with an expression such as --filter "IMPACT is HIGH". This option keeps only variants with HIGH impact consequences. You can combine multiple filter criteria using logical operators. The --filter_common option removes common variants based on allele frequency thresholds.
Why does my VEP job fail with an out-of-memory error?
Large input files and multiple annotation sources can exhaust system memory. Reduce memory usage by using the --buffer_size option to process fewer variants at once, using the --fork option to distribute work across cores, and omitting unnecessary annotation options. If the problem persists, consider splitting your input file into smaller chunks.
Related Bioinformatics Guides
- Gene Set Enrichment Analysis in R: A Practical Tutorial for Interpreting Omics Data
- Medical Image Annotation Tools: A Practical Guide for Building Segmentation Datasets
- Metabolomics Data Analysis in R: A Practical Workflow
- Metagenomics Tools: A Practical Guide to Software and Pipelines
- Microbiome Data Analysis in R: A Practical Guide for Compositional Data
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- Annotating and prioritizing genomic variants using the Ensembl Variant Effect Predictor-A tutorial.. Human mutation, 2022.
- Tutorial for variant interrogation in tumor samples.. 2026.
- Protocol for detecting rare and common genetic associations in whole-exome sequencing studies using MAGICpipeline.. 2024.
- Client Applications and Server-Side Docker for Management of RNASeq and/or VariantSeq Workflows and Pipelines of the GPRO Suite.. 2023.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.