Integrating RNA-Seq Data into Genome Annotation: A Guide to Using Transcript Evidence for Gene Model Refinement

By Dr. Zubair Khalid, DVM, MS, PhD ·

Integrating RNA-Seq Data into Genome Annotation: A Guide to Using Transcript Evidence for Gene Model Refinement

Key Takeaways

  • RNA-Seq data provides direct evidence of transcribed regions, crucial for correcting errors in ab initio gene predictions such as misidentified exon-intron boundaries or missed alternative isoforms.
  • The core workflow involves three stages: splice-aware alignment of RNA-Seq reads to the reference genome (using tools like STAR or HISAT2), assembly of aligned reads into transcript models (e.g., with StringTie), and subsequent refinement of gene models using these transcript structures.
  • High-quality RNA-Seq data, characterized by paired-end reads, sufficient coverage depth across multiple tissues/conditions, and rigorous adapter trimming/quality control, is paramount for accurate transcript assembly and subsequent gene model refinement.
  • Tools like MAKER and PASA integrate transcript evidence (GTF files from StringTie) with protein alignments to update and refine gene models, while fully automated pipelines such as BRAKER3 offer a reproducible, hands-off approach for comprehensive annotation.
  • Critical quality checks throughout the process include assessing alignment rates, proper pair rates, multi-mapping proportions, transcript assembly metrics (e.g., multi-exon transcript proportion), and evidence support scores for refined gene models.

Gene prediction algorithms produce preliminary models that often misidentify exon-intron boundaries, merge distinct transcripts, or miss alternative isoforms entirely. RNA-seq data provides direct evidence of transcribed regions that can correct these errors. This guide explains how to align RNA-seq reads to a reference genome, assemble transcripts, and use evidence-based annotation tools to refine gene models. The target reader is a researcher or laboratory professional who has RNA-seq data and a genome assembly and wants to produce more accurate gene annotations.

The Role of Transcript Evidence in Gene Model Refinement

Ab initio gene predictors use statistical models of coding sequence to identify putative genes. These predictions are useful starting points but carry inherent uncertainty. Transcript evidence grounds predictions in observed biological reality. When RNA-seq reads align to the genome, the alignment pattern reveals which genomic regions are actually transcribed, where splicing occurs, and how exons are joined. This information directly confirms or rejects predicted exon-intron boundaries.

The core workflow involves three stages. First, RNA-seq reads are aligned to the reference genome using a splice-aware aligner. Second, aligned reads are assembled into transcript models. Third, these transcript models are used to update or replace existing gene predictions. Each stage has distinct tool choices, parameter decisions, and quality considerations.

The value of this approach is well established in the bioinformatics community. Automated annotation pipelines that integrate RNA-seq evidence with protein evidence have shown substantial improvements over methods using either data type alone. One benchmark study across eleven eukaryotic species found that a pipeline combining short-read RNA-seq with a large protein database improved transcript-level accuracy by approximately twenty percentage points on average compared to earlier methods, with the largest gains observed in species with large and complex genomes [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>]. This magnitude of improvement justifies the computational investment required for transcript-based refinement.

Preparing Input Data for Transcript-Based Annotation

RNA-seq Data Requirements

The quality of transcript assembly depends directly on the quality and quantity of input RNA-seq data. Paired-end reads provide substantially better assembly results than single-end reads because the insert size information helps resolve splice junctions and exon boundaries. Read length matters as well. Longer reads span more exonic sequence and provide stronger evidence for exon connectivity.

Coverage depth is a critical consideration. Low coverage leads to fragmented assemblies with missing exons. High coverage of a single condition captures only the transcripts expressed in that condition. For genome annotation purposes, RNA-seq data from multiple tissues, developmental stages, or treatment conditions is preferable because it captures a broader representation of the transcriptome. The NCBI Sequence Read Archive provides access to publicly deposited RNA-seq datasets that can supplement your own data when broader transcript coverage is needed.

Genome Assembly Quality Assessment

Before investing computational resources in transcript alignment, assess the quality of the reference genome assembly. Fragmented assemblies with many short contigs produce ambiguous read alignments and unreliable transcript models. Key metrics include contig N50, completeness of the assembly, and the proportion of reads that map uniquely. If the genome assembly is highly fragmented, consider whether assembly polishing or scaffolding is needed before proceeding with annotation. The Galaxy Training Network offers accessible tutorials on genome assembly quality assessment that can guide this evaluation.

Read Trimming and Quality Control

Raw RNA-seq reads contain adapter sequences and low-quality bases that interfere with alignment. Trimming removes these artifacts. Standard practice involves removing adapter sequences, trimming low-quality bases from read ends, and filtering reads that become too short after trimming. Quality control metrics should be recorded before and after trimming, including total read count, retained read proportion, and per-base quality scores.

The choice of trimming parameters affects downstream results. Aggressive trimming removes more low-quality data but may discard useful sequence. Conservative trimming preserves more data but may leave adapter contamination that causes spurious alignments. A reasonable approach is to trim adapters and remove bases with quality scores below a moderate threshold, then assess whether the retained data meets alignment quality expectations.

Aligning RNA-seq Reads to the Reference Genome

Splice-Aware Alignment Principles

RNA-seq reads span exon-exon junctions. A read that crosses a splice junction contains sequence from two non-contiguous genomic regions. Standard aligners that require contiguous matches will fail to align such reads. Splice-aware aligners handle this by allowing large gaps in alignments that correspond to introns. The aligner identifies reads that split across two genomic locations and infers the splice junction from the alignment pattern.

Two widely used splice-aware aligners are HISAT2 and STAR. Both accept trimmed FASTQ files and a reference genome index as input. The choice between them involves tradeoffs in speed, memory usage, and sensitivity. STAR is known for high alignment accuracy but requires substantial memory. HISAT2 uses a graph-based index that is more memory efficient. Both produce SAM or BAM files that serve as input to transcript assemblers.

Alignment Parameter Decisions

Several parameters materially affect alignment quality. Maximum intron length should be set based on knowledge of the organism. Genomes with large introns require higher maximum intron settings. The default values in most aligners are reasonable for mammalian genomes but may need adjustment for species with unusually large or small introns.

Read mapping stringency controls how many mismatches are tolerated between the read and the genome. Higher stringency reduces spurious alignments but may miss alignments for reads with sequencing errors or those from divergent alleles. Lower stringency captures more alignments but increases false positives. For annotation purposes, moderate stringency is appropriate because the goal is to identify genuine transcript structure instead of to quantify expression.

Multi-mapping reads, which align equally well to multiple genomic locations, present a particular challenge. These reads provide ambiguous evidence and are often excluded from transcript assembly. The proportion of multi-mapping reads should be recorded as a quality metric. High multi-mapping rates may indicate repetitive genome content or assembly problems.

Alignment Quality Assessment

After alignment, examine several metrics to determine whether the data is suitable for transcript assembly. The overall alignment rate indicates what proportion of reads mapped to the genome. Low alignment rates suggest contamination, adapter problems, or genome assembly issues. The proportion of reads that align in proper pairs provides information about library quality. The distribution of reads across genomic features indicates whether the library is enriched for coding regions.

Visual inspection of alignments in a genome browser is a valuable quality check. Examine several known genes to confirm that reads align in expected patterns, with coverage across exons and gaps at introns. Unexpected patterns may indicate sample mix-ups, genome assembly errors, or annotation problems. The EMBL-EBI Training resources provide practical guidance on interpreting alignment data and diagnosing common problems.

Assembling Transcripts from Aligned Reads

Reference-Based Assembly with StringTie

StringTie assembles transcripts from aligned RNA-seq reads. It uses the alignment information in a BAM file to construct transcript models, estimating transcript abundance as part of the assembly process. The output is a GTF file containing transcript models with exon coordinates and splice junctions.

StringTie operates in two modes. In reference-guided mode, it uses an existing annotation to guide assembly, which helps assemble lowly expressed transcripts and reduces artifacts. In de novo mode, it assembles transcripts without reference guidance, which can identify novel transcripts but may produce more fragmented or spurious models. For gene model refinement, running StringTie in both modes and comparing results can be informative.

Merging Assemblies Across Samples

Individual samples produce transcript assemblies that reflect the transcripts expressed in that particular condition. Merging assemblies across multiple samples produces a more complete transcriptome representation. StringTie includes a merge function that combines multiple GTF files into a single unified transcript set. This merge step is essential when working with multiple tissues or conditions.

The merge process resolves conflicts between transcript models. When two samples produce slightly different models for the same gene, the merge produces a consensus model. Parameters control how strictly transcripts must match to be merged. StringTie's merge function is the standard approach for this step.

Assembly Quality Metrics

Transcript assemblies should be evaluated before use in annotation refinement. Key metrics include the number of transcripts assembled, the proportion of multi-exon transcripts, the distribution of transcript lengths, and the proportion of transcripts that match known genes. Assemblies with excessive numbers of single-exon transcripts may indicate genomic DNA contamination. Assemblies with very short transcripts may indicate fragmentation due to low coverage.

Comparing assembled transcripts to known annotations from related species can provide a sanity check. If the assembly produces transcripts that are dramatically different in structure from orthologous genes in related species, investigate whether the differences reflect genuine biological variation or assembly artifacts. The Bioconductor project provides R packages for comparing and visualizing transcript assemblies.

Using Transcript Evidence to Refine Gene Models

Evidence-Based Annotation with MAKER

MAKER is an annotation pipeline that integrates multiple sources of evidence, including transcript alignments, protein alignments, and ab initio predictions. It uses evidence to refine gene models, producing annotations that are supported by observed data. MAKER can be run in an iterative fashion, where the output of one round informs the next round.

The MAKER workflow involves configuring a control file that specifies the evidence sources and parameters. Transcript evidence from StringTie assemblies is provided as GTF files. Protein evidence from related species can be provided as FASTA files. MAKER aligns both types of evidence to the genome and uses the alignments to construct gene models.

MAKER produces several output files, including GFF3 annotation files, quality metrics for each predicted gene, and evidence alignments. The quality metrics indicate how well each gene model is supported by evidence. Genes with strong support can be accepted with confidence. Genes with weak support may require additional investigation.

Transcript Alignment with PASA

PASA is specifically designed to align transcript assemblies to the genome and update gene models based on transcript evidence. It takes transcript assemblies as input and produces updated gene models that incorporate the transcript structures. PASA is particularly useful for identifying alternative splicing isoforms and correcting exon-intron boundaries.

The PASA workflow involves aligning transcripts to the genome, validating the alignments, and using the validated alignments to update gene models. PASA can add new isoforms to existing genes, extend or shorten existing models, and create new gene models for transcripts that do not overlap existing annotations.

PASA is often used in conjunction with MAKER. The MAKER pipeline can incorporate PASA-updated annotations as evidence for subsequent rounds of annotation. This iterative approach progressively refines gene models, with each round incorporating additional evidence.

Automated Pipelines for Integrated Annotation

Fully automated pipelines that integrate RNA-seq and protein evidence have become increasingly accessible. BRAKER3 is one such pipeline that combines short-read RNA-seq alignments with protein database searches to annotate protein-coding genes in eukaryotic genomes [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>]. It uses GeneMark-ETP for initial gene prediction, AUGUSTUS for refinement, and TSEBRA to combine predictions from multiple sources.

The advantage of automated pipelines is reproducibility and ease of use. BRAKER3 is available as a Docker container, which simplifies installation and ensures consistent execution across systems [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>]. The pipeline learns statistical models specific to the target genome, which improves accuracy compared to using generic models.

Benchmark results indicate that BRAKER3 outperforms earlier versions of BRAKER and other annotation tools including MAKER2, Funannotate, and FINDER [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>]. The performance advantage is most pronounced for species with large and complex genomes, where transcript evidence provides critical information for resolving gene structure.

Practical Workflow for Gene Model Refinement

Step-by-Step Implementation

The following workflow represents a practical approach to using RNA-seq data for gene model refinement. Each step includes specific decisions and quality checks.

Step 1: Prepare the reference genome. Index the genome assembly for the chosen aligner. Record assembly statistics including total length, contig N50, and GC content. Verify that the assembly is in the expected format and that sequence identifiers are consistent.

Step 2: Process RNA-seq reads. Trim adapters and low-quality bases from raw reads. Record the number of reads before and after trimming. Assess read quality metrics to confirm that the data meets minimum standards for alignment.

Step 3: Align reads to the genome. Run the splice-aware aligner with parameters appropriate for the organism. Record alignment statistics including overall alignment rate, unique alignment rate, and multi-mapping rate. Inspect alignments at known genes to verify expected patterns.

Step 4: Assemble transcripts. Run StringTie on each sample individually. Record assembly statistics for each sample. Merge assemblies across samples to produce a unified transcript set.

Step 5: Assess transcript assembly quality. Compare assembled transcripts to known annotations from related species. Examine the distribution of transcript lengths and exon counts. Identify potential artifacts such as excessive single-exon transcripts.

Step 6: Refine gene models. Run MAKER or PASA with the transcript assemblies as evidence. Record the number of gene models before and after refinement. Examine changes to understand how transcript evidence affected the annotation.

Step 7: Validate refined models. Compare refined models to the original predictions. Assess the proportion of models that changed and the nature of the changes. Verify that refined models are consistent with the transcript evidence.

Computational Resource Considerations

Transcript-based annotation is computationally intensive. The alignment step is typically the most resource-demanding, requiring substantial memory and CPU time. STAR alignment of a typical RNA-seq sample can require tens of gigabytes of memory. StringTie assembly is less demanding but still requires significant CPU time for large datasets.

The nf-core community provides standardized pipelines for RNA-seq analysis that can be run on high-performance computing clusters. These pipelines implement best practices for alignment, quantification, and quality control. Using a standardized pipeline ensures reproducibility and reduces the burden of pipeline maintenance.

For researchers without access to high-performance computing, cloud-based options are available. The Galaxy Training Network provides accessible workflows that can be run on public Galaxy servers. These workflows implement the core steps of RNA-seq analysis and transcript assembly without requiring command-line expertise.

At a Glance

Workflow StagePrimary ToolsKey InputsMain OutputsCritical Quality Checks
Read processingTrimmomatic, fastpRaw FASTQ filesTrimmed FASTQ filesRetained read proportion, per-base quality scores
Splice-aware alignmentHISAT2, STARTrimmed FASTQ, genome indexBAM alignment filesOverall alignment rate, proper pair rate, multi-mapping rate
Transcript assemblyStringTieBAM filesGTF transcript modelsTranscript count, multi-exon proportion, length distribution
Assembly mergingStringTie mergeMultiple GTF filesUnified GTF fileTranscript redundancy, model consistency
Gene model refinementMAKER, PASA, BRAKER3GTF files, genome, protein databaseUpdated GFF3 annotationEvidence support scores, model structure changes

Options and Tradeoffs in Tool Selection

HISAT2 versus STAR

HISAT2 and STAR represent two different approaches to splice-aware alignment. HISAT2 uses a hierarchical graph-based index that is memory efficient and fast. STAR uses a sequential maximum mappable seed search that is highly accurate but memory intensive. For genomes with high repetitive content, STAR may provide better alignment accuracy. For large genomes where memory is constrained, HISAT2 is often the practical choice.

Both aligners produce BAM files that are compatible with StringTie and downstream tools. The choice between them is unlikely to dramatically affect the final annotation quality if parameters are set appropriately. The more important decision is ensuring that alignment parameters match the characteristics of the target genome.

StringTie Assembly Modes

StringTie's reference-guided mode uses existing annotations to improve assembly. This mode is appropriate when a preliminary annotation exists and the goal is refinement. The reference guides the assembler toward known gene structures, reducing fragmentation and artifacts. However, reference guidance can bias the assembly toward existing models, potentially missing novel transcripts that do not match the reference.

De novo mode assembles without reference guidance. This mode is appropriate when no annotation exists or when the goal is to identify novel transcripts. De novo assembly produces more complete transcript models for novel genes but may also produce more spurious models. Running both modes and comparing results provides the most complete picture.

MAKER versus PASA versus BRAKER3

MAKER is a flexible annotation pipeline that integrates multiple evidence types. It is appropriate when you want to control the annotation process and incorporate diverse evidence sources. MAKER requires more manual configuration than automated pipelines but provides greater control.

PASA specializes in using transcript alignments to update gene models. It is appropriate when the primary goal is refining existing models based on transcript evidence. PASA is particularly effective at identifying alternative isoforms and correcting exon boundaries.

BRAKER3 is a fully automated pipeline that integrates RNA-seq and protein evidence. It is appropriate when you want a reproducible, hands-off approach. BRAKER3 requires less manual intervention than MAKER or PASA but provides less control over the annotation process. The benchmark results showing substantial accuracy improvements make BRAKER3 an attractive option for many applications [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>].

Records and Measurements for Annotation Projects

Documentation Requirements

Systematic documentation is essential for reproducible annotation projects. Record the version of every tool used, the parameters for each run, and the input data versions. This information allows the analysis to be reproduced and provides context for interpreting results.

The The Carpentries lessons on reproducible research provide practical guidance on project organization and documentation. Adopting these practices from the start of an annotation project prevents problems later when results need to be verified or extended.

Quality Metrics to Track

Track the following metrics throughout the annotation process:

Read-level metrics. Total read count, trimmed read count, retained proportion, average read length before and after trimming.

Alignment metrics. Overall alignment rate, unique alignment rate, multi-mapping rate, proper pair rate, median insert size.

Assembly metrics. Number of transcripts, number of multi-exon transcripts, transcript length distribution, exon count distribution, proportion of transcripts overlapping known genes.

Annotation metrics. Number of gene models before and after refinement, proportion of models changed, proportion of models with strong evidence support, number of novel transcripts identified.

These metrics provide a quantitative basis for evaluating whether the annotation refinement was successful. Comparing metrics across samples and runs helps identify problems early in the process.

Version Control and Data Storage

Annotation projects generate large intermediate files. BAM files for multiple samples can consume hundreds of gigabytes. Plan storage capacity accordingly. Compressed formats such as CRAM can reduce storage requirements for alignment files.

Version control for scripts and configuration files is essential. The The Carpentries lessons on Git provide foundational training for version control. Tracking changes to analysis scripts ensures that the analysis can be reproduced and that parameter changes are documented.

Common Failure Patterns and Troubleshooting

Low Alignment Rates

Low alignment rates indicate that reads are not matching the reference genome. Possible causes include adapter contamination that was not fully removed, contamination from other organisms, or genome assembly errors. Check the proportion of reads that fail to align and investigate whether the unaligned reads share common characteristics.

If contamination is suspected, run a taxonomic classification on unaligned reads to identify the contaminating organism. If the genome assembly is the problem, consider whether assembly polishing or improvement is needed before proceeding with annotation.

Fragmented Transcript Assemblies

Fragmented assemblies with many short transcripts typically indicate insufficient coverage or poor read quality. Check the coverage distribution across the genome. Low coverage regions will produce fragmented assemblies. Consider whether additional RNA-seq data is needed to improve coverage.

Fragmentation can also result from incorrect intron length settings during alignment. If introns are longer than the maximum setting, reads spanning the intron will not align correctly, leading to fragmented assemblies. Check whether the maximum intron setting is appropriate for the target organism.

Excessive Single-Exon Transcripts

While single-exon transcripts can be genuine, an excessive proportion often indicates genomic DNA contamination in the RNA-seq library. Genomic DNA produces reads that align continuously across intronic regions, creating false single-exon transcript models. Check whether the single-exon transcripts overlap known intronic regions or intergenic space. If contamination is suspected, review the library preparation protocol and consider whether DNase treatment was adequate.

Discrepancies Between Transcript Evidence and Predictions

When transcript evidence conflicts with ab initio predictions, investigate the source of the discrepancy. The prediction may be wrong, the transcript assembly may be wrong, or both may be partially correct. Examine the raw read alignments at the conflicting locus to determine which model is supported by the evidence.

In some cases, the discrepancy reflects genuine biological complexity such as alternative splicing or overlapping genes. In other cases, it reflects errors in one of the models. Careful examination of the evidence is required to resolve these conflicts.

Limitations and Interpretation Boundaries

Condition-Specific Transcript Representation

RNA-seq data reflects the transcripts expressed in the specific conditions sampled. Genes that are not expressed in the sampled conditions will have no transcript evidence. The absence of transcript evidence for a gene does not mean the gene is not real. It may simply mean the gene is not expressed in the sampled conditions.

This limitation is particularly important for genes with narrow expression patterns. Sampling multiple tissues and conditions reduces but does not eliminate this problem. The NCBI databases provide access to RNA-seq data from diverse conditions that can supplement your own data.

Short-Read Assembly Limitations

Short-read RNA-seq data has inherent limitations for transcript assembly. Short reads cannot span very long exons or resolve complex alternative splicing patterns. Transcripts with very low expression may not be assembled. These limitations are inherent to the technology and cannot be fully overcome by computational methods.

Long-read RNA-seq technologies such as Iso-Seq can overcome some of these limitations by sequencing full-length transcripts. If the goal is comprehensive transcript discovery, consider whether long-read data would provide valuable supplementary evidence.

Annotation Quality Depends on Evidence Quality

The quality of the final annotation depends on the quality of the input evidence. Poor quality RNA-seq data produces poor quality transcript assemblies, which in turn produce poor quality gene models. The adage garbage in, garbage out applies directly to genome annotation.

Investing in high-quality RNA-seq data is the most effective way to improve annotation quality. This includes using appropriate library preparation methods, sufficient sequencing depth, and appropriate biological sampling.

Single-Cell and Spatial Transcriptomics Considerations

Single-cell RNA-seq (scRNA-seq) and spatial transcriptomics technologies generate data with distinct characteristics compared to bulk RNA-seq. These approaches enable characterization of cellular heterogeneity and identification of rare cell types [<a href="#ref-3">3</a>]. However, scRNA-seq data is characterized by sparsity and low expression levels, which can complicate transcript assembly for genome annotation purposes [<a href="#ref-3">3</a>]. The drop-out events and technical noise in single-cell data require specialized analysis approaches that differ from bulk RNA-seq workflows [<a href="#ref-3">3</a>].

For genome annotation, bulk RNA-seq from multiple tissues remains the primary recommended evidence source. Single-cell data can provide complementary information about cell-type-specific transcript usage but should be analyzed with appropriate tools designed for the sparsity and noise characteristics of the technology [<a href="#ref-3">3</a>]. Databases such as ssREAD demonstrate how single-cell and spatial transcriptomics data can be organized and made accessible for downstream analysis, including cell clustering and identification of cell-type-specific marker genes [<a href="#ref-4">4</a>].

Safety and Reproducibility Context

Computational Reproducibility

Genome annotation results should be reproducible. This requires documenting the exact versions of all tools and parameters used. Containerization tools such as Docker and Singularity provide a mechanism for ensuring that the same software environment is used across runs. BRAKER3 is available as a Docker container, which simplifies reproducibility [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>].

The nf-core community emphasizes reproducibility as a core principle. Their pipelines are designed to be run consistently across different computing environments. Adopting these practices for annotation projects ensures that results can be verified and extended by other researchers.

Data Management and Sharing

Annotation results should be shared in standard formats that other researchers can use. GFF3 is the standard format for genome annotations. The NCBI provides repositories for depositing genome annotations and supporting evidence data.

When sharing annotation results, include the evidence data that supports the annotations. This allows other researchers to evaluate the quality of the annotations and to update them as new evidence becomes available.

Professional Escalation Criteria

Some annotation problems require specialized expertise. Consider escalating to a bioinformatics core facility or collaborator with genome annotation experience when:

The genome assembly is highly fragmented. Annotation of fragmented assemblies requires specialized approaches. A bioinformatics professional can assess whether assembly improvement is needed before annotation.

RNA-seq data quality is poor. If multiple samples show low alignment rates or other quality problems, the issue may be in the library preparation or sequencing. A specialist can diagnose the problem and recommend corrective action.

Annotation results are inconsistent with biological expectations. If the annotation produces gene models that are dramatically different from related species, a specialist can investigate whether the differences reflect genuine biology or technical artifacts.

The project requires regulatory-grade annotation. Some applications require annotation that meets specific quality standards. A specialist can ensure that the annotation process meets these standards.

A Decision Framework for Choosing Between Manual Refinement and Automated Annotation Pipelines

The choice between manual refinement with MAKER or PASA and fully automated pipelines such as BRAKER3 is not always straightforward. Each approach has distinct strengths and limitations that become apparent only when applied to specific annotation problems. This section provides a practical decision framework based on project goals, available evidence, and computational constraints.

Core Decision Criteria

The first decision point concerns the nature of the existing annotation. If you have a well-curated annotation that has been manually reviewed and you want to preserve its structure while adding transcript evidence, manual refinement with PASA is the appropriate choice. PASA aligns transcript assemblies to the genome and updates gene models while preserving existing models that are supported by evidence. This approach is conservative and maintains continuity with previous annotation versions.

If the existing annotation is purely computational with no manual curation, automated pipelines offer a more thorough reassessment. BRAKER3 builds gene models from scratch using RNA-seq and protein evidence, which can correct systematic errors in the original prediction [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>]. The benchmark data showing approximately twenty percentage point improvements in transcript-level F1 scores compared to earlier methods suggests that automated pipelines can produce substantially better models than incremental refinement of poor initial predictions [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>].

The second decision criterion is the diversity of transcript evidence available. Automated pipelines such as BRAKER3 are designed to handle heterogeneous evidence and learn statistical models specific to the target genome [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>]. This makes them well suited for projects with RNA-seq data from multiple tissues, developmental stages, or treatment conditions. Manual refinement with PASA is more appropriate when you have a focused set of transcript assemblies that you want to integrate into an existing annotation without extensive model rebuilding.

The third criterion is the need for manual curation. Some annotation projects require human review of gene models, particularly for genes of special interest or for genomes that will serve as community reference resources. Manual refinement workflows allow curators to examine each gene model and make informed decisions about accepting, rejecting, or modifying transcript evidence. Automated pipelines produce annotations that are accurate on average but do not provide the same level of per-gene scrutiny.

Evidence Quality Assessment Before Pipeline Selection

Before selecting a refinement approach, assess the quality and completeness of your transcript evidence. The alignment statistics described in the previous section provide the first indication of evidence quality. High alignment rates with low multi-mapping proportions suggest that the RNA-seq data is suitable for annotation refinement. Low alignment rates or high multi-mapping rates indicate problems that will affect any downstream analysis.

The depth and breadth of transcript coverage is equally important. A single RNA-seq library from one condition captures only the transcripts expressed in that condition. For genome annotation purposes, evidence from multiple tissues and conditions is strongly preferred because it captures a broader representation of the transcriptome. The NCBI Sequence Read Archive provides access to publicly deposited RNA-seq datasets that can supplement your own data when broader transcript coverage is needed.

Assess the completeness of transcript assemblies before deciding on a refinement strategy. Fragmented assemblies with many short transcripts will produce poor gene models regardless of the refinement approach. If assembly quality is poor, address the underlying causes such as insufficient coverage or incorrect alignment parameters before proceeding with annotation refinement.

Comparative Workflow for Pipeline Selection

The following workflow provides a structured approach to selecting between manual refinement and automated annotation.

Step 1: Inventory available evidence. List all RNA-seq datasets, protein databases, and existing annotations available for the target genome. Record the number of samples, tissues represented, and sequencing depth for each dataset. Note the quality of the existing annotation, including whether it has been manually curated.

Step 2: Assess evidence quality. Run the alignment and assembly quality checks described in the previous sections. Record alignment rates, multi-mapping proportions, transcript counts, and assembly completeness metrics. If evidence quality is poor, address the underlying problems before proceeding.

Step 3: Define annotation goals. Determine whether the primary goal is correcting errors in an existing annotation, identifying novel transcripts, or producing a comprehensive annotation from scratch. Manual refinement is appropriate for the first goal. Automated pipelines are well suited for the second and third goals.

Step 4: Evaluate computational resources. Automated pipelines such as BRAKER3 require substantial computational resources, including memory for alignment and CPU time for model training [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>]. Manual refinement with PASA is less computationally demanding but requires more analyst time. Consider which constraint is more limiting for your project.

Step 5: Run a pilot comparison. If resources permit, run both approaches on a subset of the data. Compare the resulting gene models for a set of well-characterized genes. This pilot comparison provides direct evidence about which approach works better for your specific genome and data.

Step 6: Select and document the approach. Based on the pilot comparison and the criteria above, select the refinement approach. Document the rationale for the selection, including the evidence quality metrics and the results of the pilot comparison.

Record Keeping for Pipeline Decisions

Systematic documentation of the pipeline selection process is essential for reproducibility. Record the version of every tool used, the parameters for each run, and the input data versions. This information allows the analysis to be reproduced and provides context for interpreting results.

The The Carpentries lessons on reproducible research provide practical guidance on project organization and documentation. Adopting these practices from the start of an annotation project prevents problems later when results need to be verified or extended.

Track the following metrics when comparing refinement approaches:

Pipeline performance metrics. Wall clock time, peak memory usage, and CPU hours for each approach. These metrics inform future resource planning.

Annotation quality metrics. Number of gene models produced, proportion of models with strong evidence support, number of novel transcripts identified, and proportion of models that match known genes from related species.

Manual curation effort. Number of gene models requiring manual review and the time required for review. This metric is often overlooked but is critical for project planning.

Common Failure Patterns in Pipeline Selection

Several recurring problems emerge when researchers select annotation refinement approaches without adequate assessment.

Selecting an automated pipeline with poor quality input data. Automated pipelines such as BRAKER3 are powerful but cannot compensate for poor quality RNA-seq data. Low alignment rates, high multi-mapping proportions, or fragmented assemblies will produce poor annotations regardless of the pipeline used. Assess evidence quality before committing to a pipeline.

Using manual refinement when the existing annotation is fundamentally flawed. If the existing annotation has systematic errors such as incorrect splice sites or merged gene models, incremental refinement with PASA will preserve these errors. A more thorough reassessment with an automated pipeline may be necessary.

Assuming that more evidence always produces better annotations. While additional evidence generally improves annotation quality, the relationship is not linear. Adding low-quality evidence can degrade annotation quality by introducing spurious transcript models. Focus on evidence quality instead of quantity.

Neglecting to validate the final annotation. Regardless of the refinement approach, the final annotation should be validated against independent evidence. Compare refined models to annotations from related species, check for consistency with protein evidence, and examine gene models in a genome browser to confirm that they are supported by the underlying read alignments.

Integration with Existing Annotation Workflows

The decision framework described here is designed to complement existing annotation workflows instead of replace them. Many annotation projects benefit from a hybrid approach that uses automated pipelines for initial annotation and manual refinement for specific gene families or regions of interest.

The Galaxy Training Network provides accessible workflows that implement the core steps of RNA-seq analysis and transcript assembly. These workflows can be used to generate transcript evidence that feeds into either manual refinement or automated annotation pipelines. The nf-core community provides standardized pipelines for RNA-seq analysis that can be run on high-performance computing clusters, ensuring reproducibility and reducing the burden of pipeline maintenance.

The Bioconductor project provides R packages for comparing and visualizing transcript assemblies and gene models. These tools are useful for evaluating the output of different refinement approaches and for documenting the changes made to gene models during the refinement process.

Professional Escalation Criteria for Pipeline Selection

Some annotation problems require specialized expertise beyond what this decision framework can provide. Consider escalating to a bioinformatics core facility or collaborator with genome annotation experience when:

The genome is highly repetitive or polyploid. These genomes present unique challenges for transcript assembly and gene model refinement. Specialized approaches may be needed to handle paralogous genes and repetitive regions.

The RNA-seq data comes from multiple closely related species. Cross-species transcript evidence requires careful handling to avoid mapping artifacts. A specialist can design an appropriate analysis strategy.

The annotation will serve as a community reference resource. Community reference annotations require particularly rigorous validation and documentation. A specialist can ensure that the annotation meets community standards.

The project requires regulatory-grade annotation. Some applications require annotation that meets specific quality standards. A specialist can ensure that the annotation process meets these standards.

The EMBL-EBI Training resources provide practical guidance on bioinformatics analysis and can help identify when specialized expertise is needed. The NCBI provides access to reference annotations and supporting evidence data that can inform annotation refinement decisions.

Frequently Asked Questions

What is the difference between reference-guided and de novo transcript assembly?

Reference-guided assembly uses an existing genome annotation to guide the assembly process. This approach produces more complete assemblies for known genes but may miss novel transcripts. De novo assembly builds transcript models without reference guidance, which can identify novel transcripts but may produce more fragmented or spurious models. For gene model refinement, running both modes and comparing results provides the most complete picture.

How much RNA-seq data is needed for genome annotation?

The amount of RNA-seq data needed depends on the genome complexity and the diversity of the transcriptome. More data from diverse conditions provides better transcript coverage. A practical approach is to start with data from multiple tissues or conditions and assess assembly quality. If assemblies are fragmented or missing expected transcripts, additional data may be needed.

Can I use publicly available RNA-seq data for annotation?

Yes. Public repositories such as the NCBI Sequence Read Archive provide access to RNA-seq data from diverse organisms and conditions. Using public data can supplement your own data and improve transcript coverage. Ensure that the data is from the same species and that the conditions are relevant to your annotation goals.

What is the difference between MAKER and BRAKER3?

MAKER is a flexible annotation pipeline that integrates multiple evidence types and allows manual control. BRAKER3 is a fully automated pipeline that integrates RNA-seq and protein evidence using a specific combination of tools. BRAKER3 requires less manual configuration and has shown strong benchmark performance, while MAKER provides greater flexibility for custom annotation strategies.

How do I know if my refined gene models are accurate?

Assess the evidence support for each gene model. Models supported by multiple independent lines of evidence, such as transcript alignments and protein alignments, are more reliable than models supported by a single evidence type. Compare refined models to annotations from related species as a sanity check. The proportion of models with strong evidence support is a useful overall quality metric.

What should I do if transcript evidence conflicts with existing predictions?

Examine the raw read alignments at the conflicting locus to determine which model is supported by the evidence. The transcript evidence may reveal that the existing prediction has incorrect exon boundaries or missing isoforms. In some cases, the conflict reflects genuine biological complexity such as alternative splicing. Careful examination of the evidence is required to resolve these conflicts.

Do I need protein evidence in addition to RNA-seq data?

Protein evidence provides complementary information to RNA-seq data. Protein alignments can confirm that predicted coding sequences are translated and can help identify coding regions in transcripts. Automated pipelines such as BRAKER3 integrate both RNA-seq and protein evidence, which has been shown to improve annotation accuracy [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>]. Including protein evidence is recommended when suitable protein databases are available.

How long does a typical annotation refinement project take?

The time required depends on the genome size, the amount of RNA-seq data, and the computational resources available. Alignment and assembly of a typical RNA-seq dataset can take hours to days. Gene model refinement with MAKER or BRAKER3 adds additional time. The total project can range from days to weeks depending on the complexity of the analysis and the need for troubleshooting.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

[1] [BRAKER3: Fully automated genome annotation using RNA-seq and protein evidence with GeneMark-ETP, AUGUSTUS and TSEBRA.](https://pubmed.ncbi.nlm.nih.gov/37398387). bioRxiv : the preprint server for biology, 2024. [2] [BRAKER3: Fully automated genome annotation using RNA-seq and protein evidence with GeneMark-ETP, AUGUSTUS, and TSEBRA.](https://pubmed.ncbi.nlm.nih.gov/38866550). Genome research, 2024. [3] [A Review of Single-Cell RNA-Seq Annotation, Integration, and Cell-Cell Communication.](https://pubmed.ncbi.nlm.nih.gov/37566049). Cells, 2023. [4] [A single-cell and spatial RNA-seq database for Alzheimer's disease (ssREAD).](https://pubmed.ncbi.nlm.nih.gov/38844475). Nature communications, 2024.

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.