How to Annotate Variants from RNA-Seq Data: Challenges and Best Practices
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- RNA-seq variant annotation necessitates accounting for transcript isoform diversity and splicing, distinguishing it from DNA-seq by interpreting variants within the context of transcription, not just the genome. Splice-aware aligners are critical for accurately mapping reads that span exon junctions, and junction support must be verified for variant calls.
- Allele-specific expression (ASE) is a common phenomenon in RNA-seq where heterozygous variants exhibit skewed read counts due to biological or technical factors, complicating the interpretation of variant support and requiring assessment of allele balance.
- RNA editing, particularly A-to-I editing, can introduce non-genomic nucleotide changes (e.g., A-to-G) that mimic genomic variants, necessitating cross-checking RNA-seq variant calls against DNA-seq data or known editing sites.
- Mapping artifacts from pseudogenes and repetitive regions can lead to false positive variant calls; filtering by mapping quality and ensuring unique alignment are essential controls.
- Strand-specificity and library preparation methods impact variant interpretation, particularly for antisense transcripts, requiring confirmation of the library strand protocol and annotation with strand awareness.
- Comprehensive transcript models are crucial for accurate annotation, and when these are incomplete or novel isoforms are present, RNA-seq assembly can supplement reference annotations to provide a more complete context for variant interpretation.
RNA sequencing produces reads that reflect the transcribed portion of the genome, and variant calling from these data requires annotation strategies that account for splicing, transcript isoform diversity, and allele-specific expression. This article provides a practical workflow for researchers who need to annotate variants identified from RNA-seq data, with emphasis on the biological and technical challenges that distinguish RNA-seq variant annotation from DNA-based approaches.
At a Glance
| Challenge | Impact on Variant Annotation | Recommended Control |
|---|---|---|
| Spliced read alignment | Reads span exon junctions and may misalign to paralogous regions | Use splice-aware aligners and verify junction support for variant calls |
| Transcript isoform selection | A variant may fall in a constitutive exon or an alternative exon depending on isoform | Annotate against transcript models and report isoform context |
| Allele-specific expression | Heterozygous variants may show skewed read counts due to biological or technical factors | Assess allele balance and consider ASE when interpreting variant support |
| RNA editing and post-transcriptional modification | A-to-I editing can create apparent variants that are not genomic | Cross-check RNA-seq variants against DNA-seq or known editing sites |
| Mapping artifacts from pseudogenes and repeats | Multi-mapping reads produce false positive calls | Filter by mapping quality and unique alignment |
| Strand-specificity and library preparation | Read orientation affects variant interpretation in antisense transcripts | Confirm library strand protocol and annotate with strand awareness |
Context for RNA-Seq Variant Annotation
RNA-seq has become a standard method for measuring gene expression across species and conditions, and the same data can support variant discovery when the analysis is designed appropriately. The HISAT, StringTie and Ballgown protocol describes a complete path from raw reads to transcript assembly and differential expression, and this same family of tools illustrates how spliced alignment underpins transcript-level analysis [<a href="#ref-1">1</a>]. Variant annotation from RNA-seq builds on these alignment and assembly steps but adds the requirement to interpret each variant in the context of transcription.
The central distinction between DNA-seq and RNA-seq variant calling is that RNA-seq reads represent the transcriptome, not the genome. A variant observed in RNA-seq data may reflect a genuine genomic variant, an RNA editing event, a mapping artifact, or a technical error introduced during library preparation or sequencing. Annotation must therefore consider the transcript context of each variant, including the gene, the transcript isoform, the exon-intron structure, and the strand of transcription.
For researchers working with model organisms, livestock, or human samples, the same principles apply. A cattle multi-tissue atlas of regulatory variants demonstrates how RNA-seq data can be used to map genetic associations with gene expression and alternative splicing across more than 100 tissues, and this work shows the scale at which transcriptome-informed variant interpretation becomes feasible [<a href="#ref-2">2</a>]. The practical lesson is that RNA-seq variant annotation is most powerful when it integrates transcript models, expression context, and genetic variation.
Core Principles of RNA-Seq Variant Annotation
Transcript Context Determines Variant Interpretation
A genomic variant can have different functional consequences depending on the transcript in which it is expressed. A variant in a constitutive exon affects all isoforms of a gene, while a variant in an alternative exon affects only a subset. Annotation tools must therefore report the transcript or transcripts in which a variant falls, and researchers must interpret consequences relative to the relevant isoform.
The availability of comprehensive transcript models is essential. NCBI provides sequence databases, gene models, and variation resources that support this annotation step, and researchers should confirm that their annotation pipeline uses current transcript builds [<a href="#ref-3">3</a>]. When transcript models are incomplete or when novel isoforms are expressed, RNA-seq assembly can identify transcripts that are not present in reference annotations, and these novel transcripts require separate consideration.
Splicing Creates Alignment and Annotation Challenges
RNA-seq reads that span exon junctions cannot be aligned to the genome as contiguous sequences. Splice-aware aligners handle this by allowing reads to be split across exon boundaries, but the alignment process introduces ambiguity. A read that spans a junction may align to multiple locations in the genome, particularly when the exons are short or when the junction sequence is similar to another genomic region.
The HISAT, StringTie and Ballgown protocol emphasizes the importance of spliced alignment for transcript analysis, and the same alignment principles apply to variant calling [<a href="#ref-1">1</a>]. Variants called from spliced reads require junction-level quality assessment. A variant that appears only in reads spanning a particular junction may reflect an alignment artifact instead of a genuine variant, especially if the junction is not well supported by other reads.
Allele-Specific Expression Confounds Variant Support
In a diploid organism, a heterozygous variant should ideally show approximately equal read support from each allele. RNA-seq data frequently deviate from this expectation because the two alleles of a gene may be expressed at different levels. Allele-specific expression can arise from cis-regulatory variation, genomic imprinting, or allelic imbalance in specific cell types or conditions.
The single-cell RNA-seq study of lupus identified cell type-specific expression quantitative trait loci and linked disease-associated variants to cell type-specific expression, demonstrating that allele-specific effects can be cell type dependent [<a href="#ref-4">4</a>]. For variant annotation, this means that a heterozygous variant with skewed allele balance in RNA-seq data is not necessarily a false positive. The skew may be biologically meaningful, but it complicates the distinction between a true heterozygous variant and a technical artifact.
RNA Editing Produces Non-Genomic Variants
RNA editing is a post-transcriptional process that changes the nucleotide sequence of RNA molecules. The most common form in mammals is A-to-I editing, which appears as A-to-G changes in RNA-seq data when compared to the genome reference. These edited sites are not present in the DNA sequence, and they can be mistaken for genomic variants if the annotation pipeline does not account for them.
Researchers annotating RNA-seq variants should be aware that certain positions in the transcriptome are prone to editing, and they should cross-check candidate variants against known editing sites or DNA-seq data when available. The long-read RNA-seq study of human tissues identified novel transcripts and characterized allele-specific expression and transcript structure events, and this work highlights the resolution that long reads provide for distinguishing true variants from post-transcriptional modifications [<a href="#ref-5">5</a>].
Practical Workflow for Annotating RNA-Seq Variants
Step 1: Define the Biological Question and Sample Design
Before starting the analysis, define whether the goal is germline variant calling, somatic variant calling, or transcriptome-wide variant discovery. Germline variant calling from RNA-seq is appropriate when DNA is not available, but it requires careful filtering because RNA-seq coverage is limited to expressed regions. Somatic variant calling from RNA-seq is more challenging because tumor or disease samples may have variable expression and contamination from normal cells.
The sample design should include biological replicates when possible, and the number of samples should be sufficient to distinguish genuine variants from technical noise. The cattle multi-tissue atlas used thousands of RNA-seq samples to map regulatory variants across tissues, and this scale enabled tissue-specific variant interpretation [<a href="#ref-2">2</a>]. Smaller studies can still produce useful variant annotations, but the limitations of sample size should be acknowledged.
Step 2: Select Reference Genome and Transcript Annotation
The reference genome and transcript annotation determine the coordinate system for variant reporting and the transcript context for annotation. Use the most recent reference build available for the species of interest, and confirm that the transcript annotation matches the reference build. NCBI provides reference sequences and annotation resources that can be downloaded for this purpose [<a href="#ref-3">3</a>].
For species with well-annotated genomes, the reference transcript models provide the primary context for variant annotation. For species with incomplete annotations, or for studies that expect novel transcripts, RNA-seq assembly can supplement the reference annotation. The HISAT, StringTie and Ballgown protocol describes how to assemble transcripts from RNA-seq data, including novel splice variants, and these assembled transcripts can be used as the annotation basis for variant interpretation [<a href="#ref-1">1</a>].
Step 3: Align Reads with a Splice-Aware Aligner
Splice-aware alignment is a prerequisite for RNA-seq variant calling. The aligner must be able to split reads across exon junctions, and it should report mapping quality scores that reflect alignment confidence. The choice of aligner affects downstream variant calling, and researchers should use an aligner that is compatible with their variant caller and annotation tools.
The Galaxy Training Network provides accessible workflow training for RNA-seq analysis, including alignment and downstream steps, and these tutorials can help researchers implement a reproducible alignment pipeline [<a href="#ref-6">6</a>]. The nf-core documentation describes community standards for pipeline configuration and reproducibility, and these standards apply to RNA-seq variant calling workflows [<a href="#ref-7">7</a>].
Step 4: Call Variants with RNA-Seq Appropriate Settings
Variant calling from RNA-seq data requires settings that account for the non-uniform coverage of the transcriptome. Coverage is limited to expressed regions, and coverage depth varies across genes and transcripts. Variant callers should be configured to require minimum coverage and allele balance thresholds that are appropriate for RNA-seq data.
The distinction between germline and somatic variant calling is important at this stage. Germline variant calling assumes that variants are present in all cells, while somatic variant calling allows for variants that are present in a subset of cells. The filtering thresholds differ between these two modes, and researchers should select the mode that matches their biological question.
Step 5: Annotate Variants with Transcript Context
Variant annotation tools assign functional consequences to variants based on their position relative to transcripts. The annotation should report the gene, the transcript, the exon or intron location, and the predicted consequence such as synonymous, missense, or splice site variant. The annotation should also report whether the variant falls in a constitutive or alternative exon.
The EMBL-EBI Training program provides learning pathways for bioinformatics data resources, including variation annotation, and these materials can help researchers understand the annotation process and the interpretation of variant consequences [<a href="#ref-8">8</a>]. Bioconductor provides packages for genomic analysis and annotation, and these packages can be integrated into reproducible workflows [<a href="#ref-9">9</a>].
Step 6: Filter Artifacts and Validate Candidate Variants
The final step is filtering and validation. Filtering criteria should include mapping quality, read depth, allele balance, and proximity to splice junctions. Variants that fail these filters should be flagged for manual review or excluded from downstream analysis.
Validation can be performed by comparing RNA-seq variants to DNA-seq data when available, by checking known variant databases, or by experimental validation. The long-read RNA-seq study of human tissues demonstrated how long reads can enhance variant interpretation and identify rare variants leading to aberrant splicing patterns, and this approach can be used to validate variants that affect splicing [<a href="#ref-5">5</a>].
Options and Tradeoffs in RNA-Seq Variant Annotation
Short-Read versus Long-Read Sequencing
Short-read RNA-seq is the most common approach, and it provides high base-level accuracy for well-expressed transcripts. The limitations of short reads include difficulty in resolving complex splice variants and ambiguity in assigning reads to highly similar transcripts. The HISAT, StringTie and Ballgown protocol is designed for short-read data and provides a complete analysis path [<a href="#ref-1">1</a>].
Long-read RNA-seq provides full-length transcript information, which resolves isoform structure and enables allele-specific analysis across entire transcripts. The long-read study of human tissues identified over 70,000 novel transcripts and characterized allele-specific expression and transcript structure events, demonstrating the resolution gained from long-read data [<a href="#ref-5">5</a>]. The tradeoff is higher cost and lower throughput compared to short-read sequencing.
Reference-Based versus Assembly-Based Annotation
Reference-based annotation uses existing transcript models to interpret variants. This approach is fast and works well for well-annotated species, but it misses variants in novel transcripts. Assembly-based annotation uses RNA-seq data to reconstruct transcripts, which can identify novel isoforms, but assembly is computationally intensive and may produce incomplete or incorrect transcripts.
The HISAT, StringTie and Ballgown protocol combines reference-based alignment with transcript assembly, allowing the identification of novel splice variants while maintaining the reference coordinate system [<a href="#ref-1">1</a>]. This hybrid approach is recommended for studies that expect novel transcript diversity.
Single-Cell versus Bulk RNA-Seq
Bulk RNA-seq provides average expression across many cells, and variant calling from bulk data reflects the dominant alleles in the sample. Single-cell RNA-seq provides cell type-specific information, and the lupus study demonstrated how single-cell data can map cell type-specific expression quantitative trait loci and link variants to cell type-specific expression [<a href="#ref-4">4</a>].
The tradeoff is that single-cell RNA-seq has lower coverage per cell and higher technical noise, which complicates variant calling. The probabilistic harmonization and annotation method described in the scVI and scANVI study provides an approach for integrating single-cell datasets and assigning cell type labels, which can support cell type-specific variant interpretation [<a href="#ref-10">10</a>].
Observations and Measurements for RNA-Seq Variant Annotation
Coverage and Depth Metrics
The depth of coverage at a variant position determines the confidence in the variant call. RNA-seq coverage is highly variable across the transcriptome, with highly expressed genes having deep coverage and lowly expressed genes having shallow or no coverage. Researchers should record the depth at each variant position and use this information to filter low-confidence calls.
The minimum depth threshold depends on the variant calling mode and the expected allele frequency. Germline heterozygous variants should have sufficient depth to observe both alleles, while somatic variants may be present at low allele frequency and require deeper coverage for detection.
Allele Balance and Expression Skew
Allele balance is the proportion of reads supporting the alternative allele at a heterozygous position. In DNA-seq data, allele balance is expected to be near 50 percent for germline heterozygotes. In RNA-seq data, allele balance can deviate from 50 percent due to allele-specific expression, and this deviation can be biologically meaningful.
The cattle multi-tissue atlas identified hundreds of thousands of genetic associations with gene expression and alternative splicing, and these associations reflect allele-specific regulatory effects [<a href="#ref-2">2</a>]. When annotating RNA-seq variants, researchers should record allele balance and consider whether skew is consistent with known regulatory variation.
Junction Support and Splice Site Context
Variants near splice junctions require special attention. A variant in a splice donor or acceptor site can affect splicing, and the annotation should report the distance to the nearest splice site. Variants that are called only in reads spanning a particular junction may reflect alignment artifacts.
The long-read RNA-seq study demonstrated how allele-specific analysis of long reads can characterize transcript structure events and identify variants that lead to aberrant splicing patterns [<a href="#ref-5">5</a>]. This approach provides a model for validating splice-related variants.
Records and Documentation for Reproducibility
Pipeline Configuration and Version Control
RNA-seq variant annotation involves multiple software tools, and the versions of these tools affect the results. Record the exact versions of the aligner, variant caller, and annotation tools, and document the parameters used for each step. The nf-core documentation describes community standards for pipeline configuration and reproducibility, and these standards can be applied to RNA-seq variant calling workflows [<a href="#ref-7">7</a>].
Version control for analysis scripts is essential. The Carpentries lessons provide foundational training in shell, Git, and programming, and these skills support reproducible analysis [<a href="#ref-11">11</a>]. Researchers should commit their analysis scripts to a version control system and document the data inputs and outputs for each analysis run.
Data Management and Storage
RNA-seq data files are large, and the intermediate files from alignment and variant calling add to the storage burden. Develop a data management plan that specifies file naming conventions, directory structure, and backup procedures. NCBI provides data resources for sequence data deposition and retrieval, and these resources can be used for data sharing and archival [<a href="#ref-3">3</a>].
The Galaxy Training Network provides tutorials on data management and analysis workflows, and these tutorials emphasize the importance of documenting data provenance [<a href="#ref-6">6</a>]. Reproducible analysis requires that every result can be traced back to the specific data and parameters that produced it.
Analysis Logs and Quality Reports
Maintain analysis logs that record the commands run, the parameters used, and the output files generated. Quality reports should summarize the alignment rate, the number of variants called, the transition-transversion ratio, and other metrics that indicate data quality. These reports provide a record of the analysis and support troubleshooting when results are unexpected.
Quality Controls for RNA-Seq Variant Annotation
Alignment Quality Control
The alignment rate indicates the proportion of reads that map to the reference genome. Low alignment rates may indicate contamination, adapter contamination, or a mismatch between the sample and the reference. The alignment rate should be recorded for each sample, and samples with unusually low alignment rates should be investigated.
Mapping quality scores reflect the confidence in each read alignment. Reads with low mapping quality are more likely to be misaligned, and variants called from these reads are more likely to be false positives. Filtering by mapping quality is a standard quality control step in RNA-seq variant calling.
Variant Call Quality Control
Variant callers produce quality scores for each variant, and these scores reflect the probability that the variant is genuine. The quality score distribution should be examined, and variants with low quality scores should be filtered. The transition-transversion ratio is a useful metric for assessing variant call quality, as transitions are generally more common than transversions in genomic data.
The allele balance distribution should also be examined. In germline variant calling, heterozygous variants should cluster around 50 percent allele balance, while homozygous variants should cluster near 0 or 100 percent. Deviations from these expectations may indicate contamination or technical artifacts.
Cross-Sample Consistency
When multiple samples are analyzed, variants that are called in multiple samples are more likely to be genuine than variants called in a single sample. The cattle multi-tissue atlas used thousands of samples to map regulatory variants, and the consistency of variant calls across tissues and samples provided confidence in the results [<a href="#ref-2">2</a>]. Researchers should examine the overlap of variant calls across samples and investigate variants that appear in only one sample.
Common Failure Patterns in RNA-Seq Variant Annotation
Failure to Account for Splicing
The most common failure in RNA-seq variant annotation is the use of aligners or annotation tools that do not account for splicing. Reads that span exon junctions will not align correctly with unspliced aligners, and variants in these reads will be missed or misannotated. Splice-aware alignment is a prerequisite for RNA-seq variant calling.
Misinterpretation of Allele-Specific Expression
Researchers who are accustomed to DNA-seq data may interpret skewed allele balance in RNA-seq data as evidence of a false positive variant. This interpretation can be incorrect, as allele-specific expression is a biological phenomenon that affects the read counts at heterozygous positions. The lupus study demonstrated that cell type-specific expression quantitative trait loci can link variants to cell type-specific expression, and this context is important for interpreting allele balance [<a href="#ref-4">4</a>].
Confusion of RNA Editing with Genomic Variation
RNA editing events appear as nucleotide changes in RNA-seq data that are not present in the genome. A-to-I editing is the most common form in mammals, and it can produce apparent A-to-G variants. Researchers who do not account for RNA editing may report false positive genomic variants. Cross-checking against DNA-seq data or known editing sites is the most reliable way to distinguish editing from genomic variation.
Overreliance on Reference Annotations
Reference transcript annotations are incomplete for many species, and novel transcripts are common. The long-read RNA-seq study identified over 70,000 novel transcripts for annotated genes, demonstrating the extent of transcript diversity that is missing from reference annotations [<a href="#ref-5">5</a>]. Variant annotation that relies exclusively on reference annotations will miss variants in novel transcripts.
Inadequate Filtering of Multi-Mapping Reads
Reads that map to multiple locations in the genome are a source of false positive variants. These multi-mapping reads are common in repetitive regions and in gene families with high sequence similarity. Filtering by mapping quality and requiring unique alignment can reduce this source of error.
Limitations of RNA-Seq Variant Annotation
Coverage Bias Toward Expressed Genes
RNA-seq only captures transcribed regions, and coverage is proportional to expression level. Genes with low expression may have insufficient coverage for variant calling, and genes that are not expressed in the sampled tissue will have no coverage. This coverage bias means that RNA-seq variant annotation is incomplete by design.
Difficulty in Distinguishing Biological and Technical Variation
The distinction between a genuine genomic variant and a technical artifact is not always clear in RNA-seq data. RNA editing, allele-specific expression, and mapping artifacts can all produce apparent variants, and the annotation process must consider these possibilities. The probabilistic harmonization and annotation approach described in the scVI and scANVI study accounts for uncertainty caused by biological and measurement noise, and this principle applies to variant annotation as well [<a href="#ref-10">10</a>].
Tissue and Cell Type Specificity
Gene expression varies across tissues and cell types, and variants that are expressed in one tissue may not be expressed in another. The cattle multi-tissue atlas mapped regulatory variants across more than 100 tissues, and this work showed that regulatory effects can be tissue-specific [<a href="#ref-2">2</a>]. Variant annotation from a single tissue provides an incomplete picture of the variant landscape.
Reference Genome Limitations
The reference genome is a composite representation of the species, and it may not represent the specific individual being studied. Structural variants and regions of high diversity may be poorly represented in the reference, leading to alignment artifacts and missed variants. NCBI provides reference sequences and variation resources that can help researchers understand the limitations of the reference [<a href="#ref-3">3</a>].
Safety and Regulatory Context for RNA-Seq Variant Annotation
Data Privacy and Confidentiality
Human RNA-seq data contain genetic information that is subject to privacy regulations. Researchers must ensure that data handling complies with applicable regulations and institutional policies. NCBI provides data resources that support secure data deposition and controlled access for sensitive data [<a href="#ref-3">3</a>].
Reporting Standards for Clinical or Diagnostic Use
Variant annotation from RNA-seq data may be used in clinical or diagnostic contexts, and the reporting standards for these contexts are more stringent than for research. Variants that are reported for clinical use should be validated by an orthogonal method, and the limitations of RNA-seq-based variant calling should be clearly stated.
Ethical Considerations for Genetic Data
The interpretation of genetic variants has ethical implications, particularly when variants are associated with disease risk or other sensitive traits. Researchers should consider the ethical context of their work and ensure that their analysis and reporting are consistent with ethical standards.
Professional Escalation Criteria
When to Seek Expert Consultation
Researchers should seek expert consultation when they encounter variants that are difficult to interpret, when the results of the analysis are unexpected, or when the variant annotation will be used for clinical or regulatory decisions. Bioinformatics support teams and statistical genetics experts can provide guidance on complex cases.
When to Validate with Additional Data
Variants that will be used for important decisions should be validated with additional data. DNA-seq validation is the most reliable approach for confirming that an RNA-seq variant is genomic. The long-read RNA-seq study demonstrated how long reads can enhance variant interpretation, and this approach can be used to validate variants that affect splicing [<a href="#ref-5">5</a>].
When to Revisit the Analysis Pipeline
If the results of the analysis are inconsistent with known biology, or if quality control metrics indicate problems, the analysis pipeline should be revisited. The Galaxy Training Network provides tutorials that can help researchers troubleshoot their analysis workflows [<a href="#ref-6">6</a>], and the nf-core documentation describes community standards for pipeline configuration [<a href="#ref-7">7</a>].
Decision Framework for Variant Annotation Triage in RNA-Seq Projects
Variant annotation from RNA-seq data produces a large number of candidate variants, and not all of them require the same depth of analysis. A structured triage system helps researchers allocate effort toward variants that matter for their biological question while avoiding wasted time on artifacts or low-information calls. This section presents a practical decision framework that separates variants into priority tiers based on evidence strength, biological relevance, and downstream use.
Tier Assignment Based on Evidence Strength
The first triage decision separates variants by the strength of the evidence supporting them. Evidence strength combines read depth, mapping quality, allele balance, and junction context into a single assessment. Variants with deep coverage, high mapping quality, and consistent allele balance across multiple samples receive the highest evidence scores. Variants with shallow coverage, low mapping quality, or allele balance that conflicts with the expected genotype receive lower scores.
For germline variant calling from RNA-seq, a practical threshold separates variants that can be interpreted with confidence from those that require additional validation. The exact depth threshold depends on the sequencing platform, the library preparation method, and the expression level of the gene. A variant in a highly expressed gene with hundreds of supporting reads is more reliable than a variant in a lowly expressed gene with only a handful of reads. The HISAT, StringTie and Ballgown protocol demonstrates that RNA-seq analysis produces variable coverage across the transcriptome, and this variability must be accounted for when setting evidence thresholds [<a href="#ref-1">1</a>].
The tier assignment should also consider whether the variant is supported by reads from multiple independent alignment contexts. A variant that appears in reads spanning multiple different splice junctions is more reliable than a variant that appears only in reads spanning a single junction. The long-read RNA-seq study of human tissues demonstrated how allele-specific analysis of long reads can characterize transcript structure events and identify variants that lead to aberrant splicing patterns, and this work shows the value of examining variant support across transcript contexts [<a href="#ref-5">5</a>].
Tier Assignment Based on Biological Relevance
The second triage decision separates variants by their biological relevance to the research question. A variant in a gene that is central to the study hypothesis deserves more attention than a variant in a gene that is expressed but not relevant to the question. Biological relevance should be assessed using the gene annotation, the predicted functional consequence of the variant, and the expression context in the sampled tissue.
The cattle multi-tissue atlas of regulatory variants demonstrated how RNA-seq data can link gene expression in different tissues to economically important traits using transcriptome-wide association and colocalization analyses [<a href="#ref-2">2</a>]. This work shows that the biological relevance of a variant depends on the tissue context and the trait of interest. A variant that affects gene expression in one tissue may have no effect in another tissue, and the annotation should reflect this tissue specificity.
For studies focused on disease or trait associations, variants that fall in genes with known relevance to the phenotype should be prioritized. The single-cell RNA-seq study of lupus identified cell type-specific expression quantitative trait loci and linked disease-associated variants to cell type-specific expression, demonstrating that the biological relevance of a variant can be cell type specific [<a href="#ref-4">4</a>]. Researchers should consider whether the sampled tissue or cell type is relevant to the phenotype being studied.
Tier Assignment Based on Downstream Use
The third triage decision separates variants by their downstream use. Variants that will be used for clinical or diagnostic decisions require the highest level of validation and documentation. Variants that will be used for hypothesis generation in a research context require less stringent validation. Variants that will be used for population-level analyses require consistent calling across samples and careful quality control.
The nf-core documentation describes community standards for pipeline configuration and reproducibility, and these standards apply to the downstream use of variant calls [<a href="#ref-7">7</a>]. When variants will be shared with collaborators or deposited in public databases, the annotation must include complete provenance information, including the reference build, the transcript annotation version, and the software versions used.
Practical Triage Workflow
A practical triage workflow assigns each variant to one of three tiers. Tier 1 variants have strong evidence support, clear biological relevance, and a defined downstream use. These variants receive full annotation, including transcript context, predicted functional consequence, and allele-specific expression assessment. Tier 2 variants have moderate evidence support or uncertain biological relevance. These variants receive standard annotation but are flagged for additional review if they become relevant to the research question. Tier 3 variants have weak evidence support or no clear biological relevance. These variants are recorded but not prioritized for further analysis.
The triage workflow should be documented and applied consistently across all samples in a study. The Galaxy Training Network provides accessible workflow training for RNA-seq analysis, and these tutorials emphasize the importance of reproducible analysis workflows [<a href="#ref-6">6</a>]. A documented triage process ensures that the same criteria are applied to all variants and that the results can be compared across samples and studies.
Record System for Triage Decisions
A record system for triage decisions should capture the evidence that supports each tier assignment. For each variant, the record should include the read depth, the mapping quality, the allele balance, the number of supporting samples, and the junction context. The record should also include the biological relevance assessment, including the gene, the predicted functional consequence, and the relevance to the research question.
The record system should be implemented in a structured format that can be queried and analyzed. A spreadsheet or database with one row per variant and columns for each evidence metric provides a simple and effective record system. The Carpentries lessons provide foundational training in data organization and management, and these skills support the implementation of a structured record system [<a href="#ref-11">11</a>].
The record system should also capture the decisions made during triage, including the tier assignment and the rationale for the assignment. This documentation supports transparency and reproducibility, and it allows other researchers to understand why certain variants were prioritized for analysis.
Common Failure Patterns in Triage
A common failure in triage is applying the same evidence thresholds to all variants without considering the expression context. A variant in a lowly expressed gene may have low read depth simply because the gene is not highly expressed, not because the variant is unreliable. The triage process should account for the expected coverage based on gene expression level.
Another common failure is prioritizing variants based on biological relevance without considering evidence strength. A variant in a highly relevant gene with weak evidence support may be a false positive, and prioritizing it for analysis can waste resources. The triage process should require both evidence strength and biological relevance for Tier 1 assignment.
A third common failure is failing to document triage decisions. Without documentation, the triage process cannot be reproduced, and the rationale for prioritizing certain variants is lost. The record system should be maintained throughout the analysis and updated as new information becomes available.
Integration with Existing Annotation Workflows
The triage framework integrates with existing annotation workflows by adding a decision layer between variant calling and final annotation. After variants are called and before full annotation is performed, the triage process assigns each variant to a tier. Tier 1 variants receive full annotation, while Tier 2 and Tier 3 variants receive reduced annotation or are deferred.
The EMBL-EBI Training program provides learning pathways for bioinformatics data resources, including variation annotation, and these materials can help researchers understand how to integrate triage into their existing workflows [<a href="#ref-8">8</a>]. Bioconductor provides packages for genomic analysis and annotation, and these packages can be used to implement the triage process programmatically [<a href="#ref-9">9</a>].
The triage framework should be flexible enough to accommodate different research questions and data types. A study focused on a specific gene may prioritize variants in that gene regardless of evidence strength, while a genome-wide study may require consistent evidence thresholds across all variants. The framework should be adapted to the specific needs of each study.
Escalation Criteria for Triage Decisions
The triage framework should include escalation criteria for variants that are difficult to classify. A variant with conflicting evidence, such as high read depth but low mapping quality, should be escalated for manual review. A variant with strong evidence support but unclear biological relevance should be escalated for consultation with a domain expert.
The professional escalation criteria should be documented and applied consistently. When a variant is escalated, the record system should capture the reason for escalation and the outcome of the review. This documentation supports continuous improvement of the triage process and provides a record of the decisions made during the analysis.
Validation of Triage Decisions
The triage decisions should be validated periodically to ensure that the criteria are producing reliable results. Validation can be performed by comparing triage assignments to known variants, by examining the concordance of triage assignments across samples, or by reviewing a random sample of Tier 2 and Tier 3 variants to confirm that they were correctly classified.
The long-read RNA-seq study of human tissues demonstrated how long reads can enhance variant interpretation and identify rare variants leading to aberrant splicing patterns [<a href="#ref-5">5</a>]. This approach can be used to validate triage decisions by providing additional evidence for variants that were classified as Tier 2 or Tier 3 based on short-read data.
Practical Implementation Steps
To implement the triage framework, researchers should first define the evidence thresholds and biological relevance criteria for their study. These criteria should be documented and shared with the research team. Next, the record system should be implemented, either as a spreadsheet or a database, and the triage process should be applied to a test set of variants to ensure that it works as intended.
The triage process should be integrated into the analysis pipeline so that it runs automatically after variant calling. The nf-core documentation describes community standards for pipeline configuration, and these standards can be applied to the triage step [<a href="#ref-7">7</a>]. The pipeline should produce a triage report that summarizes the number of variants in each tier and the evidence supporting the tier assignments.
Finally, the triage process should be reviewed and updated as the research project evolves. New information about the biological relevance of certain genes or variants may change the triage criteria, and the process should be flexible enough to accommodate these changes. The record system should capture the version of the triage criteria used for each analysis run, so that results can be compared across runs.
Frequently Asked Questions
What is the difference between germline and somatic variant calling from RNA-seq?
Germline variant calling assumes that variants are present in all cells of an individual, and it is appropriate when the goal is to identify inherited variants. Somatic variant calling allows for variants that are present in a subset of cells, such as tumor cells, and it requires different filtering thresholds. The choice between germline and somatic calling depends on the biological question and the sample type.
How does splicing affect variant annotation from RNA-seq?
Splicing creates reads that span exon junctions, and these reads require splice-aware alignment. Variants in spliced reads may be misaligned if the aligner does not account for junctions, and variants near splice sites may affect splicing. Annotation must consider the exon-intron structure of the transcript and the distance of the variant to the nearest splice site.
What is allele-specific expression and why does it matter for variant annotation?
Allele-specific expression is the unequal expression of the two alleles of a gene. This phenomenon can cause skewed allele balance at heterozygous variant positions in RNA-seq data, and it can be biologically meaningful. The lupus study demonstrated that cell type-specific expression quantitative trait loci can link variants to cell type-specific expression [<a href="#ref-4">4</a>], and this context is important for interpreting allele balance.
How can I distinguish RNA editing from genomic variants?
RNA editing is a post-transcriptional process that changes RNA sequences, and the most common form in mammals is A-to-I editing, which appears as A-to-G changes. Cross-checking RNA-seq variants against DNA-seq data is the most reliable way to distinguish editing from genomic variation. Known editing sites can also be used to flag potential editing events.
What coverage is needed for reliable variant calling from RNA-seq?
The coverage needed depends on the variant calling mode and the expected allele frequency. Germline heterozygous variants require sufficient depth to observe both alleles, while somatic variants may be present at low allele frequency and require deeper coverage. RNA-seq coverage is variable across the transcriptome, and lowly expressed genes may not have sufficient coverage for variant calling.
Should I use long-read or short-read RNA-seq for variant annotation?
Short-read RNA-seq is more common and provides high base-level accuracy for well-expressed transcripts. Long-read RNA-seq provides full-length transcript information and can resolve isoform structure and allele-specific expression across entire transcripts. The long-read study of human tissues identified over 70,000 novel transcripts and demonstrated the resolution gained from long-read data [<a href="#ref-5">5</a>]. The choice depends on the research question and available resources.
How do I handle variants in novel transcripts that are not in the reference annotation?
RNA-seq assembly can identify transcripts that are not present in reference annotations. The HISAT, StringTie and Ballgown protocol describes how to assemble transcripts including novel splice variants [<a href="#ref-1">1</a>]. Variants in novel transcripts should be annotated relative to the assembled transcript models, and the novel transcript should be reported in the analysis.
What quality controls are essential for RNA-seq variant annotation?
Essential quality controls include alignment rate assessment, mapping quality filtering, variant quality score examination, allele balance assessment, and cross-sample consistency checks. The transition-transversion ratio is a useful metric for assessing variant call quality. Researchers should document all quality control metrics and investigate samples or variants that deviate from expectations.
Related Bioinformatics Guides
- Data Annotation for AI in Life Sciences: Roles, Challenges, and Best Practices
- Alternative Splicing Analysis from RNA-Seq Data
- RNA-Seq Databases: Accessing and Using Public RNA-Seq Data
- RNA-Seq Data Analysis in Galaxy: A User-Friendly Platform
- RNA-Seq Data Analysis Workflow: From Raw Reads to Insights
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
[1] [Transcript-level expression analysis of RNA-seq experiments with HISAT, StringTie and Ballgown.](https://pubmed.ncbi.nlm.nih.gov/27560171). Nature protocols, 2016. [2] [A multi-tissue atlas of regulatory variants in cattle.](https://pubmed.ncbi.nlm.nih.gov/35953587). Nature genetics, 2022. [3] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information. [4] [Single-cell RNA-seq reveals cell type-specific molecular and genetic associations to lupus.](https://pubmed.ncbi.nlm.nih.gov/35389781). Science (New York, N.Y.), 2022. [5] [Transcriptome variation in human tissues revealed by long-read sequencing.](https://pubmed.ncbi.nlm.nih.gov/35922509). Nature, 2022. [6] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project. [7] [nf-core Documentation](https://nf-co.re/docs). nf-core. [8] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute. [9] [Bioconductor](https://bioconductor.org/). Bioconductor Project. [10] [Probabilistic harmonization and annotation of single-cell transcriptomics data with deep generative models.](https://pubmed.ncbi.nlm.nih.gov/33491336). Molecular systems biology, 2021. [11] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.