How to Annotate a Eukaryotic Genome from Scratch: A Step-by-Step Pipeline with MAKER and InterProScan

By Dr. Zubair Khalid, DVM, MS, PhD ·

How to Annotate a Eukaryotic Genome from Scratch: A Step-by-Step Pipeline with MAKER and InterProScan

Key Takeaways

  • MAKER integrates multiple evidence types for robust eukaryotic genome annotation, synthesizing repeat masking, ab initio predictions, and extrinsic data (protein and RNA-seq alignments) into evidence-based gene models with quantitative Annotation Edit Distance (AED) scores.
  • Protein evidence from closely related species is paramount for accurate gene model prediction, significantly increasing gene count and accuracy, and can partially compensate for limited transcriptome data.
  • Iterative predictor training is crucial for MAKER's effectiveness, utilizing preliminary gene models to refine ab initio predictors like Augustus and SNAP, thereby improving gene structure prediction accuracy in subsequent annotation rounds.
  • Assembly quality assessment (e.g., BUSCO completeness, contiguity metrics) and repeat masking are prerequisite steps, as fragmented assemblies and unmasked repetitive elements directly lead to fragmented or spurious gene models.
  • Functional annotation with InterProScan enriches gene models with domain and GO assignments, providing biological context, though a high proportion of unassigned proteins may indicate poor gene model quality or rapidly evolving sequences.

A researcher with a new eukaryotic genome assembly needs a practical path from raw sequence to annotated gene models with functional information. This article walks through the complete MAKER-based annotation process, covering repeat masking, training ab initio predictors, integrating RNA-seq evidence, and functional annotation with InterProScan. The target reader has a genome assembly in hand, has basic command-line familiarity, and needs reproducible steps that produce defensible gene annotations for downstream biological interpretation.

Genome annotation is the process of identifying the functional elements along a genome sequence, primarily protein-coding genes, and assigning biological meaning to those elements. For a newly assembled eukaryotic genome, annotation transforms raw nucleotide sequence into structured gene models with evidence-based support. The MAKER pipeline was designed for this purpose, allowing investigators to independently annotate eukaryotic genomes and create genome databases without requiring extensive bioinformatics infrastructure [<a href="#ref-1">1</a>]. MAKER identifies repeats, aligns ESTs and proteins to a genome, produces ab initio gene predictions, and automatically synthesizes these data into gene annotations having evidence-based quality indices [<a href="#ref-1">1</a>].

This article provides a practical workflow that moves from assembly quality assessment through repeat masking, evidence preparation, ab initio training, MAKER annotation rounds, and functional annotation with InterProScan. Each section includes concrete decisions, command-level considerations, quality checks, and common failure patterns. The goal is to produce gene models that are supported by multiple lines of evidence and that can withstand scrutiny during manuscript review or database submission.

At a Glance

The table below summarizes the major pipeline stages, their primary inputs, expected outputs, and the key quality check at each step.

Pipeline StagePrimary InputsExpected OutputsKey Quality Check
Assembly quality assessmentGenome FASTA, raw readsAssembly statistics, completeness estimatesBUSCO completeness score and contiguity metrics
Repeat maskingGenome FASTA, repeat databaseSoft-masked genome, repeat annotation GFFProportion of genome masked and repeat family distribution
Evidence preparationRNA-seq reads, protein sets from related speciesAligned transcript alignments, protein alignmentsAlignment coverage and mapping rates
Ab initio gene prediction trainingMasked genome, high-confidence transcript alignmentsTrained predictor parametersGene structure statistics compared to known genes
MAKER annotation roundsMasked genome, evidence files, trained predictorsGene models with annotation edit distance (AED) scoresAED distribution and number of gene models
Functional annotationPredicted protein sequencesInterProScan domain and GO annotationsProportion of proteins with functional assignments

The MAKER pipeline scales to match available computational resources and can rapidly annotate genomes of any size [<a href="#ref-2">2</a>]. The workflow described here follows the iterative approach that has proven effective across diverse eukaryotic taxa, from insects to plants [<a href="#ref-3">3</a>][<a href="#ref-4">4</a>][<a href="#ref-5">5</a>].

Understanding the Annotation Problem

Why Eukaryotic Genomes Require a Specialized Pipeline

Eukaryotic genomes present annotation challenges that prokaryotic pipelines do not address. Introns interrupt coding sequences, alternative splicing produces multiple transcripts from a single locus, and repetitive elements can comprise a substantial fraction of the genome. A pipeline designed for emerging model organism genomes must handle these features while remaining portable and easily configurable [<a href="#ref-1">1</a>].

The MAKER pipeline was developed to address exactly this problem. It was designed to allow investigators to independently annotate eukaryotic genomes and create genome databases [<a href="#ref-1">1</a>]. The pipeline identifies repeats, aligns ESTs and proteins to a genome, produces ab initio gene predictions, and automatically synthesizes these data into gene annotations having evidence-based quality indices [<a href="#ref-1">1</a>]. This synthesis of multiple evidence types is what distinguishes MAKER from simpler gene-finding approaches.

The Role of Evidence in Annotation Quality

Gene models produced by ab initio predictors alone are often inaccurate because they lack experimental support. The quality of an annotation depends on the quantity and type of extrinsic data provided to the annotation pipeline. A study validating the International Weed Genomics Consortium annotation pipeline through reannotation of Arabidopsis thaliana found that the pipeline produced more accurate annotated genes with more input proteins, especially from closely related species, in the gene model prediction step [<a href="#ref-5">5</a>]. The combination of proteins from several closely related species increased the number of annotated genes [<a href="#ref-5">5</a>].

This finding has direct practical implications. When annotating a new genome, the researcher should invest time in collecting protein sets from the closest available relatives. The study also found that the number or source of Iso-seq reads did not have a significant effect if many proteins from closely related species were utilized [<a href="#ref-5">5</a>]. This suggests that protein evidence can partially compensate for limited transcriptome data, though transcript evidence remains valuable for refining gene structures.

What MAKER Actually Does

MAKER integrates multiple data types into a unified annotation. The pipeline takes as input a genome sequence, repeat annotations, transcript alignments, protein alignments, and ab initio gene predictions. It then synthesizes these data into gene annotations having evidence-based quality indices [<a href="#ref-1">1</a>]. The annotation edit distance (AED) score provides a quantitative measure of how well a gene model agrees with the available evidence. An AED of zero indicates perfect agreement with evidence, while an AED of one indicates no support.

MAKER is also easily trainable. Outputs of preliminary runs are used to automatically retrain its gene-prediction algorithm, producing higher-quality gene models on subsequent runs [<a href="#ref-1">1</a>]. This iterative training approach means that the first annotation round is not the final product. The researcher should plan for multiple rounds, using each round to improve the training data for the next.

Preparing the Genome Assembly

Assessing Assembly Quality Before Annotation

Annotation quality depends on assembly quality. A fragmented assembly with many gaps will produce fragmented gene models regardless of the annotation pipeline used. Before beginning annotation, the researcher should assess the assembly using standard metrics.

The primary assembly quality metrics include contiguity measures such as N50 and L50, which describe the length distribution of contigs or scaffolds. Completeness can be assessed using BUSCO (Benchmarking Universal Single-Copy Orthologs), which searches for a set of conserved genes expected to be present in the target lineage. A high BUSCO completeness score indicates that most expected genes are present in the assembly.

The NCBI provides official descriptions of sequence resources and analysis services that can support assembly quality assessment [<a href="#ref-6">6</a>]. The researcher should document assembly statistics before annotation so that any gene model deficiencies can be interpreted in the context of assembly quality.

Handling Assembly Polishing

Long-read genome assemblies often contain base-level errors that can disrupt gene models. Assembly polishing uses raw reads to correct these errors. The decision to polish before annotation depends on the assembly method and error rate.

If the assembly was produced with high-accuracy long reads and has been polished during assembly, additional polishing may provide minimal benefit. If the assembly contains regions with low coverage or high error rates, polishing can improve annotation accuracy. The researcher should assess per-base quality and consider polishing if gene models show premature stop codons or frameshifts that cannot be explained by biological variation.

Creating the Working Directory Structure

A reproducible annotation project requires organized file management. The following directory structure provides a logical organization for the pipeline:

project/
├── genome/
│   └── assembly.fasta
├── evidence/
│   ├── proteins/
│   ├── rnaseq/
│   └── repeats/
├── repeat_mask/
├── train/
├── maker_round1/
├── maker_round2/
├── functional/
└── logs/

This structure keeps inputs separate from outputs and allows the researcher to track which version of each file was used in each pipeline round. The Carpentries provides foundational computing and data lessons that cover file organization and command-line practices [<a href="#ref-7">7</a>].

Repeat Masking

Why Repeat Masking Matters

Repetitive elements, including transposable elements and simple sequence repeats, can confuse gene predictors. If repeats are not masked, ab initio predictors may generate spurious gene models within repeat regions. MAKER identifies repeats as part of its standard workflow [<a href="#ref-1">1</a>], but providing a pre-computed repeat annotation gives the pipeline better information.

Repeat masking serves two purposes. First, it prevents gene predictors from modeling repetitive elements as protein-coding genes. Second, it allows the annotation pipeline to distinguish between genuine coding sequence and repeat-derived sequence.

Building a Species-Specific Repeat Library

The most accurate repeat masking uses a species-specific repeat library. RepeatModeler can build such a library from the genome assembly itself. The resulting library contains consensus sequences for the repeat families present in the genome.

For species with well-characterized repeat content, existing databases such as RepBase or Dfam can supplement the de novo library. The researcher should combine the de novo library with known repeat sequences from related species to maximize sensitivity.

Running RepeatMasker

RepeatMasker uses the repeat library to identify and mask repetitive elements. The output includes a masked genome sequence and a GFF file describing repeat locations.

The choice between hard masking and soft masking affects downstream analysis. Hard masking replaces repeat bases with N characters, which prevents aligners and predictors from using those positions. Soft masking converts repeat bases to lowercase, which allows aligners to use the sequence but marks it as repetitive. Soft masking is generally preferred for MAKER because it allows the pipeline to use repeat regions for training while still marking them as low confidence.

The masked genome should be checked for the proportion of sequence masked. Eukaryotic genomes vary widely in repeat content, from a few percent in compact genomes to more than half in some plant and animal genomes. The repeat annotation GFF should be retained for downstream analysis and for interpreting gene model locations relative to repeat content.

Preparing Evidence Files

Protein Evidence from Related Species

Protein evidence from related species provides the most important extrinsic data for gene model prediction. The International Weed Genomics Consortium pipeline study found that more input proteins, especially from closely related species, improved annotation accuracy [<a href="#ref-5">5</a>]. The combination of proteins from several closely related species increased the number of annotated genes [<a href="#ref-5">5</a>].

The researcher should collect protein sets from the closest available relatives. For a newly sequenced species, this might include proteins from the same genus, family, or order. Public databases such as NCBI provide access to protein sequences from annotated genomes [<a href="#ref-6">6</a>]. The EMBL-EBI provides training and data-resource documentation that can help researchers navigate protein sequence databases [<a href="#ref-8">8</a>].

Protein sets should be filtered to remove redundant sequences and sequences with low quality. Redundancy can be reduced using tools such as CD-HIT or similar clustering approaches. The final protein file should be formatted as a FASTA file with unique identifiers.

Transcript Evidence from RNA-seq

RNA-seq data provides direct evidence of transcribed regions and exon-intron boundaries. The OMIGA pipeline for insect genome annotation mapped RNA-Seq reads to genomic scaffolds to determine transcribed regions using Bowtie, and the putative transcripts were assembled using Cufflinks [<a href="#ref-3">3</a>]. This approach identifies transcribed regions that can be used to train gene predictors and to validate gene models.

For a new genome, the researcher should generate or obtain RNA-seq data from multiple tissues or conditions to maximize transcript coverage. The woodland strawberry re-annotation used RNA-Seq transcriptome sequences from 25 diverse tissue types to improve annotation accuracy [<a href="#ref-4">4</a>]. This extensive transcriptome data uncovered new genes, added exons to current genes, and extended existing gene models [<a href="#ref-4">4</a>].

RNA-seq reads should be aligned to the masked genome using a splice-aware aligner such as HISAT2 or STAR. The resulting alignments should be converted to BAM format and sorted. MAKER can use BAM files directly as transcript evidence.

Iso-seq Data as High-Quality Transcript Evidence

Isoform sequencing (Iso-seq) provides full-length transcript sequences that can substantially improve gene model accuracy. The International Weed Genomics Consortium pipeline is based on Iso-seq data and was shown to annotate almost all genes without manual curation when informed with an Iso-seq dataset and proteins of related species [<a href="#ref-5">5</a>].

If Iso-seq data are available, they should be processed to generate high-quality transcript models. This typically involves clustering full-length reads, polishing consensus sequences, and mapping the resulting transcripts to the genome. The resulting transcript models provide strong evidence for exon-intron structure.

Organizing Evidence for MAKER

MAKER accepts evidence in specific formats. Protein alignments can be provided as a FASTA file of protein sequences, which MAKER will align to the genome using Exonerate. Transcript alignments can be provided as BAM files or GFF files. The OMIGA pipeline used Exonerate to refine gene structure and to determine near exact exon/intron boundaries in the genome [<a href="#ref-3">3</a>].

The researcher should organize evidence files with clear names that indicate the evidence type and source. This documentation is essential for reproducing the annotation and for interpreting gene model quality scores.

Training Ab Initio Gene Predictors

Why Train Predictors

Ab initio gene predictors use statistical models of gene structure to identify genes in genomic sequence. Default parameters are often poorly suited to a new species because gene structure varies across taxa. Training the predictors on species-specific data improves their accuracy.

MAKER is easily trainable because outputs of preliminary runs are used to automatically retrain its gene-prediction algorithm, producing higher-quality gene models on subsequent runs [<a href="#ref-1">1</a>]. This means that the researcher does not need to create training data from scratch. The first MAKER run can be used to generate training data for the second run.

Selecting Training Genes

High-quality training genes are essential for accurate predictor training. The OMIGA pipeline selected highly reliable transcripts with intact coding sequences to train de novo gene prediction software, including Augustus [<a href="#ref-3">3</a>]. This selection step ensures that the training set contains complete, correctly structured genes.

For the first training round, the researcher can use transcript evidence to identify high-confidence gene models. Transcripts with full-length coding sequences that align cleanly to the genome provide the best training data. The researcher should filter for transcripts that have complete open reading frames and that do not overlap repetitive elements.

Training Augustus

Augustus is a widely used ab initio gene predictor that can be trained for new species. The training process involves generating a set of gene structures from the training data, optimizing the parameters, and validating the trained model.

The training procedure typically follows these steps:

  1. Generate training gene structures from high-confidence transcript alignments
  2. Create a training set with proper format for Augustus
  3. Run the Augustus training scripts to optimize parameters
  4. Validate the trained model on a held-out set of genes

The trained Augustus parameters should be saved and used in subsequent MAKER rounds. The researcher should document which training data were used and the resulting gene structure statistics.

Training SNAP

SNAP is another ab initio gene predictor that MAKER supports. SNAP training follows a similar procedure to Augustus training but uses a different underlying model. The MAKER documentation provides specific instructions for training SNAP using MAKER output.

The iterative training approach means that the researcher should plan for at least two MAKER rounds. The first round uses default or preliminary predictor parameters. The second round uses predictors trained on the first round output. Additional rounds can further improve accuracy if the gene models continue to improve.

Running MAKER

Configuring MAKER

MAKER uses control files to specify parameters for each run. The three control files are maker_opts.ctl, maker_bopts.ctl, and maker_exe.ctl. The maker_opts.ctl file specifies the genome, evidence files, and predictor settings. The maker_bopts.ctl file specifies BLAST and Exonerate parameters. The maker_exe.ctl file specifies the paths to external programs.

The researcher should create a new set of control files for each MAKER round. This documentation allows the researcher to track which parameters were used for each annotation version.

Running the First MAKER Round

The first MAKER round uses the masked genome, evidence files, and default or preliminary predictor parameters. The purpose of this round is to generate initial gene models that can be used for predictor training.

The first round should be run with the following settings:

  • Genome: soft-masked assembly
  • Protein evidence: filtered protein set from related species
  • Transcript evidence: RNA-seq alignments and/or Iso-seq transcript models
  • Repeat GFF: repeat annotation from RepeatMasker
  • Ab initio predictors: default parameters or preliminary trained parameters

The MAKER run can take substantial computational time depending on genome size and the amount of evidence. The pipeline scales to match available computational resources [<a href="#ref-2">2</a>], so the researcher should allocate sufficient compute time and disk space.

Evaluating First Round Output

The first round output should be evaluated before proceeding to training. Key metrics include:

  • Number of gene models
  • Distribution of AED scores
  • Proportion of gene models with protein support
  • Gene model completeness

Gene models with low AED scores (below 0.25) have strong evidence support. Gene models with high AED scores (above 0.5) have weak support and may be artifacts. The researcher should examine the AED distribution to assess overall annotation quality.

Training Predictors from First Round Output

The first round gene models provide training data for the second round. The researcher should select gene models with strong evidence support for training. The MAKER documentation provides scripts for extracting training data from MAKER output.

The training set should include gene models with:

  • AED scores below 0.25
  • Complete open reading frames
  • No overlap with repetitive elements
  • Support from multiple evidence types

The selected gene models should be converted to the format required by each predictor. Augustus and SNAP have different training procedures, and the researcher should follow the documentation for each.

Running the Second MAKER Round

The second MAKER round uses the trained predictors. This round should produce higher-quality gene models than the first round because the predictors are now species-specific.

The second round should be run with the same evidence files as the first round, but with the trained predictor parameters. The researcher should compare the second round output to the first round output to assess improvement.

The re-annotation of the woodland strawberry genome using MAKER described many more predicted protein coding genes compared to the GeneMark generated annotation [<a href="#ref-4">4</a>]. This demonstrates that the choice of annotation pipeline and parameters can substantially affect the final gene count.

Additional MAKER Rounds

Additional MAKER rounds can further improve annotation quality. Each round uses the previous round output to retrain predictors. The researcher should continue iterating until the gene model quality metrics stabilize.

The decision to stop iterating should be based on the rate of improvement. If the AED distribution and gene model statistics do not change substantially between rounds, additional rounds are unlikely to provide significant benefit.

Functional Annotation with InterProScan

Preparing Protein Sequences for InterProScan

The MAKER output includes predicted protein sequences for each gene model. These protein sequences are the input for functional annotation with InterProScan.

The researcher should extract the longest protein isoform for each gene model to create a non-redundant protein set. This reduces the computational burden of InterProScan and avoids redundant functional annotations.

Running InterProScan

InterProScan searches protein sequences against multiple protein domain and family databases. The output includes domain annotations, Gene Ontology (GO) terms, and pathway information.

InterProScan can be run in parallel across multiple processors to reduce runtime. The output should be saved in a structured format that can be parsed for downstream analysis.

Interpreting InterProScan Output

The InterProScan output provides functional information for each protein. The researcher should examine the proportion of proteins with functional assignments. A high proportion of proteins with no functional assignment may indicate:

  • Poor gene model quality
  • Rapidly evolving proteins with no detectable homologs
  • Species-specific proteins

The functional annotations should be integrated with the gene models to create a complete annotation file. This file can be used for downstream analysis such as gene ontology enrichment or pathway analysis.

Integrating Functional Annotations with Gene Models

The final annotation should include both structural information (gene models) and functional information (InterProScan results). The researcher should create a combined annotation file that includes:

  • Gene model coordinates
  • Transcript and protein sequences
  • Functional annotations
  • Evidence support scores

This combined file serves as the primary annotation product for the genome project.

Quality Assessment and Validation

Using BUSCO to Assess Annotation Completeness

BUSCO completeness should be assessed on the predicted protein set. The BUSCO analysis searches for conserved single-copy orthologs expected to be present in the target lineage. A high BUSCO completeness score on the predicted proteins indicates that most expected genes are present in the annotation.

The researcher should compare the BUSCO score on the genome assembly to the BUSCO score on the predicted proteins. A substantial drop in completeness between the genome and the proteins indicates that some genes were missed during annotation.

Examining Gene Model Statistics

Gene model statistics provide insight into annotation quality. Key statistics include:

  • Gene density (genes per megabase)
  • Mean gene length
  • Mean exon count per gene
  • Mean exon length
  • Mean intron length

These statistics should be compared to related species with well-annotated genomes. Substantial deviations may indicate annotation problems.

Checking for Common Artifacts

Common annotation artifacts include:

  • Gene models with premature stop codons
  • Gene models with frameshifts
  • Gene models that overlap repeats
  • Gene models with no evidence support

The researcher should examine a sample of gene models manually to check for these artifacts. The Apollo genome browser provides a means to annotate, view, and edit individual contigs and BACs without the overhead of a database [<a href="#ref-1">1</a>].

Comparing to Related Annotations

If a related species has a well-annotated genome, the researcher can compare gene content between the two species. Orthologous gene pairs should have similar structures. Substantial differences may indicate annotation errors in either genome.

The International Weed Genomics Consortium pipeline was validated by reannotating Arabidopsis thaliana and comparing the reannotation to the published annotation [<a href="#ref-5">5</a>]. This validation approach can be applied to any new annotation project.

Common Failure Patterns and Troubleshooting

Low Gene Model Count

A low gene model count relative to related species may indicate:

  • Overly stringent filtering
  • Poor assembly quality
  • Insufficient evidence
  • Incorrect repeat masking

The researcher should examine the AED distribution and the proportion of gene models with protein support. If many gene models were filtered due to high AED scores, the evidence may be insufficient or the predictors may be poorly trained.

High Proportion of Gene Models with No Functional Assignment

A high proportion of proteins with no InterProScan hits may indicate:

  • Poor gene model quality
  • Truncated gene models
  • Rapidly evolving proteins

The researcher should examine the length distribution of proteins with no functional assignment. Truncated proteins may indicate incomplete gene models.

Gene Models with Premature Stop Codons

Premature stop codons in gene models may indicate:

  • Base errors in the assembly
  • Incorrect exon-intron boundaries
  • Pseudogenes

The researcher should examine the genomic context of gene models with premature stop codons. If they are located in repeat-rich regions, they may be artifacts of repeat masking.

Poor Predictor Training

Poor predictor training can result from:

  • Too few training genes
  • Training genes with incorrect structures
  • Training genes that are not representative of the genome

The researcher should examine the training gene statistics and compare them to known gene structure statistics for the lineage.

Computational Resource Limitations

MAKER can be computationally intensive for large genomes. The pipeline scales to match available computational resources [<a href="#ref-2">2</a>], but the researcher should plan for substantial compute time and disk space.

If computational resources are limited, the researcher can:

  • Reduce the amount of evidence used
  • Use a smaller training set
  • Run MAKER on individual chromosomes or scaffolds

The Galaxy Training Network provides accessible workflow training and analysis tutorials that can help researchers implement annotation pipelines in a shared computing environment [<a href="#ref-9">9</a>]. The nf-core documentation describes community pipeline standards for reproducible workflows [<a href="#ref-10">10</a>].

Reproducibility and Documentation

Recording Pipeline Parameters

Reproducible annotation requires complete documentation of pipeline parameters. The researcher should record:

  • Genome assembly version and source
  • Repeat library version and parameters
  • Evidence file versions and sources
  • Predictor training parameters
  • MAKER control file settings
  • InterProScan version and parameters

This documentation allows the annotation to be reproduced or updated as new evidence becomes available.

Version Control for Annotation Files

Version control is essential for tracking annotation changes. The researcher should use Git or a similar system to track changes to control files, scripts, and annotation outputs. The Carpentries provides foundational Git training that covers version control practices [<a href="#ref-7">7</a>].

Containerization and Workflow Management

Containerization can improve reproducibility by capturing the software environment. The nf-core documentation describes community pipeline standards that emphasize reproducibility and portability [<a href="#ref-10">10</a>]. Bioconductor provides official package, workflow, and installation documentation for reproducible genomic analysis [<a href="#ref-11">11</a>].

The researcher should consider using containers or workflow managers to ensure that the annotation pipeline can be rerun with the same software versions.

Limitations and Interpretation

Annotation Is a Model, Not Ground Truth

Genome annotations are computational models of gene content. They contain errors and should be interpreted with appropriate caution. Gene models with strong evidence support are more reliable than those with weak support.

The AED score provides a quantitative measure of evidence support. Gene models with AED scores below 0.25 have strong support and are likely to be accurate. Gene models with AED scores above 0.5 should be treated as provisional.

Annotation Quality Depends on Evidence Quality

The quality of an annotation depends on the quality and quantity of evidence provided. The International Weed Genomics Consortium study found that more input proteins, especially from closely related species, improved annotation accuracy [<a href="#ref-5">5</a>]. The researcher should invest in collecting high-quality evidence before running the annotation pipeline.

Manual Curation May Be Required

Automated annotation pipelines produce draft annotations that may require manual curation. The OMIGA pipeline used MAKER to integrate data from RNA-Seq, de novo gene prediction, and protein alignment to produce an official gene set [<a href="#ref-3">3</a>]. Even with this integration, manual curation may be needed for genes with complex structures or weak evidence support.

The Apollo genome browser provides a means to annotate, view, and edit individual contigs and BACs [<a href="#ref-1">1</a>]. Manual curation should focus on gene models with high AED scores or those with biological importance.

Annotation Is an Ongoing Process

Genome annotations are not static products. As new evidence becomes available, annotations should be updated. The woodland strawberry genome was re-annotated using MAKER with extensive transcriptome data, providing improvements over the first generation annotation [<a href="#ref-4">4</a>]. The researcher should plan for annotation updates as new data are generated.

Professional Escalation Criteria

When to Seek Additional Expertise

The researcher should consider seeking additional expertise when:

  • BUSCO completeness on predicted proteins is substantially lower than on the genome assembly
  • Gene model statistics deviate substantially from related species
  • The annotation pipeline produces inconsistent results across runs
  • Manual curation reveals systematic errors in gene models

Bioinformatics support teams or collaborators with annotation experience can provide guidance on troubleshooting and pipeline optimization.

When to Consider Alternative Pipelines

MAKER is one of several annotation pipelines available. The researcher should consider alternative approaches when:

  • MAKER produces poor results despite adequate evidence
  • The target genome has unusual features that MAKER does not handle well
  • The research question requires specialized annotation features

The OMIGA pipeline was developed specifically for insect genomes with high levels of heterozygosity [<a href="#ref-3">3</a>]. Other pipelines may be better suited to specific taxonomic groups or genome features.

When to Involve the Community

For community resource genomes, the researcher should consider involving the broader community in annotation. Community annotation efforts can improve annotation quality through distributed manual curation. The i5k initiative for insect genomes demonstrates the value of community-based annotation efforts [<a href="#ref-3">3</a>].

Building a Decision Framework for Evidence Selection and Annotation Rounds

Defining the Evidence Hierarchy for Your Project

The success of a MAKER annotation depends on the strategic selection and prioritization of evidence files. instead of simply collecting all available data, the researcher should establish a clear evidence hierarchy before the first MAKER run. This hierarchy determines which evidence types receive priority when MAKER synthesizes conflicting signals into gene models.

The International Weed Genomics Consortium validation study provides a practical basis for this hierarchy. The study found that the pipeline produced more accurate annotated genes with more input proteins, especially from closely related species, in the gene model prediction step [<a href="#ref-5">5</a>]. The combination of proteins from several closely related species increased the number of annotated genes [<a href="#ref-5">5</a>]. This finding establishes protein evidence from closely related species as the primary evidence tier for gene model prediction.

The same study found that the number or source of Iso-seq reads did not have a significant effect if many proteins from closely related species were utilized [<a href="#ref-5">5</a>]. This result supports a tiered evidence strategy where protein evidence from close relatives takes precedence over transcript evidence when both are available. Transcript evidence remains essential for refining exon-intron boundaries and identifying alternative splicing, but it should not be the sole basis for gene model construction.

The woodland strawberry re-annotation demonstrates the value of extensive transcriptome data for improving existing annotations. Using RNA-Seq transcriptome sequences from 25 diverse tissue types, the re-annotation pipeline improved existing annotations by increasing the annotation accuracy based on extensive transcriptome data [<a href="#ref-4">4</a>]. It uncovered new genes, added exons to current genes, and extended existing gene models [<a href="#ref-4">4</a>]. This finding supports the inclusion of diverse tissue transcriptomes as a secondary evidence tier that adds value beyond protein evidence alone.

Constructing the Evidence Decision Matrix

The researcher should construct a decision matrix that maps evidence availability to annotation strategy. This matrix guides the choice of MAKER parameters and the expected quality of the final annotation.

Evidence ScenarioPrimary Evidence TierSecondary Evidence TierExpected Annotation QualityMAKER Configuration
Close relative proteins plus Iso-seqProteins from same genus or familyFull-length Iso-seq transcriptsHighest quality, near-complete gene modelsUse both evidence types, train predictors on high-confidence models
Close relative proteins plus RNA-seqProteins from same genus or familyRNA-seq from multiple tissuesHigh quality with accurate exon boundariesUse both evidence types, train predictors on transcript-supported models
Distant relative proteins plus RNA-seqProteins from same order or classRNA-seq from multiple tissuesModerate quality, may miss lineage-specific genesPrioritize transcript evidence for gene structure, use proteins for homology support
Proteins only, no transcript dataProteins from same genus or familyNoneDraft quality, gene models depend on homologyUse protein evidence exclusively, expect higher AED scores
Transcript data only, no proteinsRNA-seq or Iso-seq from target speciesNoneGene models limited to expressed regionsUse transcript evidence exclusively, expect incomplete gene space

This matrix serves as a planning tool before the first MAKER run. The researcher should document which scenario applies to their project and adjust expectations accordingly. The matrix also guides the interpretation of quality metrics after each annotation round.

Setting Quality Thresholds by Evidence Scenario

Quality thresholds should be calibrated to the evidence scenario. A project with close relative proteins and Iso-seq data should achieve a higher proportion of gene models with low AED scores than a project with only distant relative proteins.

For projects with close relative proteins and transcript evidence, the researcher should expect at least 70 percent of gene models to have AED scores below 0.25 after the second MAKER round. For projects with only distant relative proteins, a lower threshold of 50 percent may be realistic. The researcher should record these thresholds before running MAKER to avoid post hoc adjustment of quality criteria.

The BUSCO completeness score on predicted proteins should also be interpreted in the context of the evidence scenario. A project with close relative proteins should achieve BUSCO completeness close to the assembly completeness. A project with only transcript evidence will likely show lower BUSCO completeness because genes not expressed in the sampled tissues will be missing from the annotation.

Implementing the Iterative Decision Loop

The annotation process should follow a structured decision loop that determines when to stop iterating. This loop prevents both premature termination and unnecessary computational expenditure.

The decision loop operates as follows:

  1. Run MAKER with the current evidence and predictor configuration
  2. Calculate quality metrics including AED distribution, gene model count, and BUSCO completeness on predicted proteins
  3. Compare metrics to the thresholds established for the evidence scenario
  4. If metrics meet thresholds, proceed to functional annotation
  5. If metrics fall below thresholds, identify the limiting factor and adjust

The limiting factor analysis should examine three potential causes of poor quality. First, insufficient evidence may require collecting additional protein sets from closer relatives or generating more transcript data. Second, poor predictor training may require selecting better training genes or running additional training rounds. Third, assembly errors may require polishing or targeted correction of problematic regions.

The OMIGA pipeline provides an example of this iterative approach. The pipeline first mapped RNA-Seq reads to genomic scaffolds to determine transcribed regions, then selected highly reliable transcripts with intact coding sequences to train Augustus [<a href="#ref-3">3</a>]. This selection step ensured that the training data contained complete gene models. The trained predictor was then used in MAKER to integrate RNA-Seq, de novo gene prediction, and protein alignment [<a href="#ref-3">3</a>].

Recording the Decision Trail

A complete decision trail is essential for reproducibility and for defending the annotation during manuscript review. The researcher should maintain a decision log that records each major choice and the rationale behind it.

The decision log should include the following entries for each annotation round:

  • Evidence files used and their version identifiers
  • Predictor training parameters and training gene selection criteria
  • Quality metrics before and after the round
  • Decisions made based on quality metric comparisons
  • Changes to parameters or evidence for the next round

This log serves multiple purposes. It allows the researcher to reconstruct the annotation process months later when preparing the manuscript. It provides a basis for updating the annotation when new evidence becomes available. It also provides documentation for reviewers who may question specific annotation decisions.

The Carpentries provides foundational lessons on data organization and documentation practices that support this record-keeping approach [<a href="#ref-7">7</a>]. The nf-core documentation describes community pipeline standards that emphasize reproducibility and portability [<a href="#ref-10">10</a>]. These resources can help the researcher establish documentation practices before beginning the annotation project.

Handling Conflicting Evidence

Conflicting evidence arises when different evidence types support different gene structures at the same locus. The researcher needs a systematic approach to resolving these conflicts instead of making ad hoc decisions.

The first step is to examine the evidence support for each conflicting model. MAKER assigns AED scores that reflect agreement with all available evidence. A gene model supported by protein homology and transcript evidence should receive a lower AED score than a model supported by only one evidence type.

The second step is to consider the biological plausibility of each model. A model with a complete open reading frame and conserved protein domains is more plausible than a model with a truncated reading frame. The researcher should examine the protein alignment to determine whether the predicted protein matches known homologs across its full length.

The third step is to document the conflict and the resolution. This documentation is particularly important for gene models that will be highlighted in the manuscript or used for downstream functional analysis. The researcher should record which evidence types supported each model and why one model was selected over the other.

When to Stop Iterating

The decision to stop iterating should be based on the rate of improvement between rounds instead of an absolute quality threshold. If the quality metrics improve substantially between round one and round two, an additional round may provide further improvement. If the metrics are stable between round two and round three, additional rounds are unlikely to provide significant benefit.

The researcher should also consider the purpose of the annotation. A draft annotation for a genome report may require fewer rounds than an annotation intended for detailed comparative genomics or functional studies. The decision to stop should balance quality goals against computational cost and project timeline.

The MAKER documentation describes the pipeline as easily trainable because outputs of preliminary runs are used to automatically retrain its gene-prediction algorithm [<a href="#ref-1">1</a>]. This design supports multiple rounds of iteration. The researcher should plan for at least two rounds and evaluate the improvement after each round before deciding whether to continue.

Common Decision Errors

Several common errors undermine the decision framework. The first error is changing quality thresholds after seeing results. This practice invalidates the quality assessment and makes it impossible to determine whether the annotation actually improved between rounds.

The second error is adding new evidence files after the first round without documenting the change. This practice makes it impossible to attribute quality improvements to specific changes. The researcher should freeze the evidence set before the first round and only add new evidence with clear documentation.

The third error is training predictors on gene models that include artifacts. The OMIGA pipeline selected highly reliable transcripts with intact coding sequences for training [<a href="#ref-3">3</a>]. Including gene models with premature stop codons or frameshifts in the training set will propagate these errors to the trained predictors.

The fourth error is ignoring the evidence scenario when interpreting quality metrics. A project with only distant relative proteins will not achieve the same quality as a project with close relative proteins and Iso-seq data. The researcher should interpret quality metrics in the context of the evidence available instead of comparing to published annotations from well-resourced projects.

Professional Escalation for Persistent Quality Problems

When quality metrics do not improve despite multiple annotation rounds, the researcher should escalate the problem instead of continuing to iterate with the same approach. Persistent quality problems may indicate issues that cannot be resolved through parameter adjustment alone.

The researcher should seek additional expertise when the BUSCO completeness on predicted proteins remains substantially lower than on the genome assembly after three MAKER rounds. This gap suggests systematic gene model errors that may require manual curation or assembly improvement.

The researcher should also seek expertise when gene model statistics deviate substantially from related species despite adequate evidence. This deviation may indicate unusual genome features that require specialized annotation approaches. The OMIGA pipeline was developed specifically for insect genomes with high levels of heterozygosity [<a href="#ref-3">3</a>], demonstrating that some genome features require tailored solutions.

Bioinformatics support teams or collaborators with annotation experience can provide guidance on troubleshooting and pipeline optimization. The EMBL-EBI provides training and data-resource documentation that can help researchers build the skills needed to diagnose annotation problems [<a href="#ref-8">8</a>]. The Galaxy Training Network provides accessible workflow training and analysis tutorials that can help researchers implement alternative approaches in a shared computing environment [<a href="#ref-9">9</a>].

Frequently Asked Questions

What is the minimum evidence needed to annotate a eukaryotic genome?

Protein evidence from related species is the most important evidence type. The International Weed Genomics Consortium study found that the pipeline produced more accurate annotated genes with more input proteins, especially from closely related species [<a href="#ref-5">5</a>]. A single protein set from a closely related species can provide sufficient evidence for a draft annotation. Transcript evidence improves gene model accuracy but is not strictly required for an initial annotation.

How many MAKER rounds should I run?

The number of MAKER rounds depends on the rate of improvement between rounds. MAKER is easily trainable because outputs of preliminary runs are used to automatically retrain its gene-prediction algorithm [<a href="#ref-1">1</a>]. Most projects benefit from at least two rounds. The researcher should continue iterating until the gene model quality metrics stabilize between rounds.

What is a good AED score cutoff for filtering gene models?

AED scores range from zero to one, with zero indicating perfect agreement with evidence. Gene models with AED scores below 0.25 have strong evidence support. Gene models with AED scores above 0.5 have weak support and should be treated as provisional. The researcher should examine the AED distribution and select a cutoff appropriate for the project goals.

How do I handle gene models that overlap repetitive elements?

Gene models that overlap repetitive elements may be artifacts of repeat masking. The researcher should examine these gene models manually to determine whether they represent genuine genes or repeat-derived artifacts. If the gene model has strong protein support, it may be a genuine gene that has been misclassified as repetitive.

What should I do if my BUSCO completeness on proteins is much lower than on the genome?

A substantial drop in BUSCO completeness between the genome and the predicted proteins indicates that some genes were missed during annotation. The researcher should examine the missing BUSCO genes to determine why they were not annotated. Possible causes include insufficient evidence, poor predictor training, or assembly errors in the relevant genomic regions.

Can I use MAKER for non-model organisms?

MAKER was designed for emerging model organism genomes for which extensive bioinformatics resources may not be readily available [<a href="#ref-1">1</a>]. The pipeline is portable and easily configurable, making it suitable for non-model organisms. The OMIGA pipeline was used to annotate the draft genome of an important insect pest, Chilo suppressalis, yielding 12,548 genes [<a href="#ref-3">3</a>].

How long does a MAKER annotation run take?

The runtime depends on genome size, the amount of evidence, and available computational resources. MAKER can rapidly annotate genomes of any size and scales to match available computational resources [<a href="#ref-2">2</a>]. A small eukaryotic genome with modest evidence can be annotated in hours. A large genome with extensive evidence may require days or weeks.

What is the difference between MAKER and MAKER-P?

MAKER-P is an extension of MAKER designed for plant genome annotation. The MAKER-P protocol describes how to use MAKER and MAKER-P to annotate protein-coding and noncoding RNA genes in newly assembled genomes [<a href="#ref-2">2</a>]. The researcher should use MAKER-P for plant genomes and MAKER for other eukaryotic genomes.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

[1] [MAKER: an easy-to-use annotation pipeline designed for emerging model organism genomes.](https://pubmed.ncbi.nlm.nih.gov/18025269). Genome research, 2008. [2] [Genome Annotation and Curation Using MAKER and MAKER-P.](https://pubmed.ncbi.nlm.nih.gov/25501943). Current protocols in bioinformatics, 2014. [3] [OMIGA: Optimized Maker-Based Insect Genome Annotation.](https://pubmed.ncbi.nlm.nih.gov/24609470). Molecular genetics and genomics : MGG, 2014. [4] [Re-annotation of the woodland strawberry (Fragaria vesca) genome.](https://pubmed.ncbi.nlm.nih.gov/25623424). BMC genomics, 2015. [5] [Validation of the International Weed Genomics Consortium genome annotation pipeline through reannotation of the model species Arabidopsis thaliana.](https://pubmed.ncbi.nlm.nih.gov/42374877). The plant genome, 2026. [6] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information. [7] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries. [8] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute. [9] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project. [10] [nf-core Documentation](https://nf-co.re/docs). nf-core. [11] [Bioconductor](https://bioconductor.org/). Bioconductor Project.

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.