# Somatic Variant Calling from RNA-Seq Data: Challenges, Tools, and Best Practices


## Key Takeaways

- RNA-seq variant calling identifies expressed mutations in transcribed regions only, fundamentally differing from DNA-seq by being a filtered view of the genome and thus unable to detect variants in untranscribed regions or silenced genes.
- Key artifacts in RNA-seq variant calling include splicing misalignments, RNA editing events (e.g., A-to-G conversions), and PCR artifacts from library preparation, each requiring distinct filtering strategies.
- Allele-specific expression significantly distorts variant allele frequencies in RNA-seq, meaning observed frequencies do not directly reflect cellular prevalence and can differ substantially from matched DNA exome data.
- Splice-aware alignment, particularly using a two-pass strategy as recommended by GATK, is critical for accurate read placement across exon-exon junctions, mitigating alignment artifacts that lead to false positive variant calls.
- The absence of matched normal RNA samples in many RNA-seq studies elevates the risk of germline variant misclassification and artifactual calls, necessitating more rigorous filtering and orthogonal validation methods like RT-PCR or targeted DNA sequencing.
- Specialized tools like RNAIndel are crucial for accurate somatic indel detection in tumor RNA-seq data, as general-purpose callers exhibit high false positive rates for this variant class due to library preparation artifacts.

---

Somatic variant calling from RNA sequencing data allows researchers to identify mutations expressed in tumor transcriptomes, but the approach differs fundamentally from DNA-based calling. RNA-seq variant calling requires accounting for splicing, allele-specific expression, and uneven coverage across transcripts. This article explains the core differences from DNA-based workflows, describes tools including GATK and VarScan2, and provides a practical filtering workflow for reducing artifacts. The guidance applies to biology students, researchers, laboratory professionals, and life-science practitioners who need to make informed decisions about when RNA-seq variant calling is appropriate and how to interpret its results.

## At a Glance

RNA-seq variant calling detects expressed variants in transcribed regions only. The table below summarizes the key decisions researchers face when planning a somatic variant calling experiment using tumor RNA-seq data.

| Decision Point | DNA-Seq Somatic Calling | RNA-Seq Somatic Calling | Practical Consideration |
|---|---|---|---|
| Input material | Tumor and matched normal DNA | Tumor RNA, often without matched normal RNA | Absence of matched normal RNA increases artifact risk and complicates germline filtering |
| Variant types detectable | SNVs, indels, structural variants across genome | SNVs and indels in expressed exons and splice junctions | Variants in untranscribed regions or silenced genes are invisible to RNA-seq |
| Primary artifacts | Sequencing errors, alignment errors | Splicing misalignment, RNA editing, allele-specific expression, PCR artifacts from library preparation | Each artifact class requires distinct filtering strategies |
| Coverage characteristics | Relatively uniform across genome | Highly uneven, dependent on gene expression level | Lowly expressed genes may have insufficient depth for confident calls |
| Recommended tools | GATK Mutect2, VarScan2 with matched normal | GATK HaplotypeCaller per-sample workflow, VarScan2, RNAIndel for indels | Tool choice depends on whether matched normal data exist and whether indels are the focus |
| Validation approach | Sanger sequencing or targeted deep sequencing | RT-PCR, targeted RNA sequencing, or orthogonal DNA sequencing | RNA-level validation confirms expression but not genomic status |

## Why RNA-Seq Variant Calling Differs from DNA-Based Calling

RNA sequencing has become a central assay in cancer research because it provides information beyond what DNA sequencing alone can offer. The transcriptome reflects which variants are actually expressed in a tumor, and this expression information can influence understanding of disease mechanisms and therapeutic strategies. However, the features that make RNA-seq valuable also create analytical complications.

### The Transcriptome Is a Filtered View of the Genome

DNA sequencing surveys the full genome regardless of whether a region is transcribed. RNA sequencing only captures sequences from genes that are actively expressed in the sampled tissue at the time of collection. A somatic variant in a tumor suppressor gene that is silenced through epigenetic mechanisms will not appear in RNA-seq data. Similarly, variants in regulatory regions, introns, or intergenic sequences are generally absent from standard RNA-seq libraries prepared with poly-A selection or ribosomal RNA depletion.

This filtered view means that absence of a variant in RNA-seq data does not indicate absence of that variant in the genome. Researchers must interpret negative results with caution. The practical consequence is that RNA-seq variant calling answers a specific question: which variants are expressed in this tumor at detectable levels. It does not answer the broader question of which variants exist in the tumor genome.

### Splicing Creates Alignment Challenges

RNA-seq reads span exon-exon junctions. When a read crosses a splice junction, it does not align contiguously to the reference genome. Alignment tools must split such reads across exons, and this process introduces opportunities for misalignment. Reads that span junctions can be incorrectly placed, particularly in regions with homologous sequences or pseudogenes. These misalignments generate false variant calls that resemble genuine somatic mutations.

The GATK best practices for RNA-seq variant calling address this problem through a two-pass alignment strategy. The first pass identifies splice junctions, and the second pass uses those junctions to improve alignment accuracy. This approach reduces spurious calls caused by reads mapping across exon boundaries. Researchers who skip the two-pass strategy risk inflating their variant counts with alignment artifacts.

### Allele-Specific Expression Distorts Variant Allele Frequencies

In DNA sequencing, the variant allele frequency reflects the proportion of cells carrying the mutation. In RNA sequencing, the observed allele frequency reflects both the cellular proportion and the relative expression of the two alleles. A somatic variant may be present in all tumor cells but expressed at low levels compared to the reference allele. Conversely, a variant present in a small subclone may show high allele frequency if the mutant allele is strongly transcribed.

Recent work has demonstrated that variant allele frequencies in RNA-seq data can differ substantially from those observed in matched DNA exome data. This phenomenon appears particularly pronounced in cancer-driving genes, where RNA analysis can reveal variant allele expression much higher than expected based on DNA sequencing. Researchers who interpret RNA-seq allele frequencies as direct proxies for cellular prevalence will draw incorrect conclusions about clonal architecture.

### RNA Editing Produces Variants That Are Not Genomic Mutations

RNA editing is a post-transcriptional process that changes nucleotide sequences in RNA molecules. The most common form in humans converts adenosine to inosine, which is read as guanosine during reverse transcription and sequencing. These editing events appear as A-to-G variants in RNA-seq data but do not correspond to mutations in the genomic DNA.

Distinguishing RNA editing from genuine somatic mutations requires either matched DNA sequencing or careful filtering against known editing sites. Databases of common RNA editing events can help flag suspicious calls. Without such filtering, RNA editing events will be misclassified as somatic variants, inflating the true mutation burden and potentially leading to incorrect biological interpretations.

## Core Principles of RNA-Seq Variant Calling

Successful RNA-seq variant calling depends on understanding the data generation process, the bioinformatics pipeline, and the biological context. The following principles guide practical decision-making.

### Read Alignment Must Account for Splicing

The choice of alignment tool and strategy is the first major decision. Spliced aligners such as STAR or HISAT2 are designed to handle exon-exon junctions. The GATK workflow for RNA-seq variant calling specifies a particular approach: align reads with a splice-aware aligner, mark duplicates, split reads at junctions with known or novel splice sites, and then apply base quality score recalibration before variant calling.

The two-pass alignment approach used in the GATK RNA-seq workflow improves sensitivity for variants near splice junctions. In the first pass, the aligner identifies candidate junctions. In the second pass, the aligner uses these junctions as annotations, allowing more accurate placement of reads that span exon boundaries. This approach is particularly important for detecting variants in genes with complex splicing patterns.

### Base Quality Recalibration Reduces Systematic Errors

RNA-seq data contain systematic base quality errors that vary by sequencing platform, library preparation method, and sequence context. Base quality score recalibration uses known variant sites to model these errors and adjust quality scores accordingly. The GATK workflow includes this step for RNA-seq data, and it improves the accuracy of downstream variant calls.

Researchers working outside the GATK ecosystem should verify that their chosen pipeline includes an equivalent quality recalibration step. Skipping recalibration does not necessarily invalidate results, but it increases the risk of calling systematic sequencing errors as genuine variants.

### Variant Calling Tools Differ in Their Assumptions

GATK HaplotypeCaller is the most widely used tool for RNA-seq variant calling. The GATK toolkit provides state-of-the-art pipelines for germline and somatic variant discovery and genotyping. For RNA-seq data, the fully validated GATK pipeline is a per-sample workflow that does not include joint genotyping across multiple samples. Recent work has described how modern GATK commands can be combined to perform joint genotyping on RNA-seq samples, extending the per-sample approach to multi-sample analyses.

VarScan2 is another commonly used tool that can call variants from RNA-seq data. VarScan2 uses a heuristic approach based on read depth, allele frequency, and statistical tests to distinguish genuine variants from sequencing errors. It can operate in somatic mode when matched normal data are available, or in germline mode for single-sample analysis.

For indel calling specifically, RNAIndel was developed to address the high false positive rate of standard RNA-seq variant calling pipelines for insertions and deletions. RNAIndel uses machine learning to classify indels as somatic, germline, or artifact based on sequence context and predicted biological effect. This tool outperforms general-purpose pipelines for indel detection in tumor RNA-seq data.

### Matched Normal Data Transform the Analysis

The availability of matched normal samples changes the analytical approach. When matched normal DNA or RNA is available, somatic variant callers can subtract germline variants and focus on tumor-specific mutations. When no matched normal is available, researchers must rely on population databases and filtering heuristics to exclude common germline polymorphisms.

In practice, many tumor RNA-seq studies lack matched normal RNA. The absence of a matched normal increases the importance of careful filtering. Researchers should expect higher false positive rates and should validate candidate somatic variants through orthogonal methods.

## Practical Workflow for Somatic Variant Calling from RNA-Seq

The following workflow describes the steps from raw RNA-seq reads to filtered somatic variant calls. This workflow follows the GATK best practices adapted for RNA-seq data and incorporates additional filtering for somatic analysis.

### Step 1: Quality Control of Raw Reads

Start with quality assessment of the raw sequencing data. Check per-base quality scores, GC content, adapter contamination, and duplication rates. Low-quality samples should be flagged before proceeding. The [Galaxy Training Network](https://training.galaxyproject.org/) provides accessible tutorials for quality control and downstream analysis steps, which are useful for researchers learning the workflow.

Record the following metrics for each sample: total read count, percentage of reads passing quality filters, mean read length, and duplication rate. These metrics inform later decisions about coverage thresholds and filtering stringency.

### Step 2: Splice-Aware Alignment

Align reads to the reference genome using a splice-aware aligner. Use the two-pass strategy recommended by the GATK workflow to improve junction detection. The reference genome and annotation files should match the version used for downstream analysis.

After alignment, assess mapping statistics. The percentage of uniquely mapped reads, the percentage of reads mapping to exonic regions, and the distribution of reads across genes provide important quality indicators. Samples with low unique mapping rates or excessive multi-mapping reads may produce unreliable variant calls.

### Step 3: Post-Alignment Processing

Mark duplicate reads, which arise from PCR amplification during library preparation. Duplicate reads do not represent independent observations and should not contribute to variant allele frequency calculations.

Split reads at exon-exon junctions using the GATK SplitNCigarReads tool. This step ensures that reads spanning splice junctions are processed correctly during variant calling. Reads that remain unsplit can cause false variant calls at junction boundaries.

Apply base quality score recalibration using known variant sites appropriate for the organism under study. This step corrects systematic quality score errors and improves the accuracy of variant calling.

### Step 4: Variant Calling

Call variants using GATK HaplotypeCaller in per-sample mode. The GATK workflow for RNA-seq variant calling is designed for per-sample analysis, and this approach is the fully validated pipeline. For multi-sample joint genotyping, combine modern GATK commands as described in recent methodological work.

For indel-focused analyses, consider running RNAIndel in addition to or instead of general-purpose callers. RNAIndel was specifically designed for somatic indel detection in tumor RNA-seq data and addresses the high false positive rate of standard pipelines for this variant class.

If matched normal data are available, use a somatic caller such as VarScan2 in somatic mode. This approach subtracts germline variants and produces a tumor-specific variant list.

### Step 5: Filtering and Annotation

Apply hard filters to remove low-quality calls. Common filters include minimum read depth, minimum variant allele frequency, and minimum base quality. The specific thresholds depend on the research question and the quality of the data.

Annotate variants with their predicted functional effects. Tools such as SnpEff or VEP provide gene names, transcript information, and predicted protein changes. This annotation helps prioritize variants for downstream analysis.

For somatic analysis without matched normal, filter against population databases to remove common germline polymorphisms. The [NCBI](https://www.ncbi.nlm.nih.gov/) provides access to dbSNP and other variation resources that support this filtering step.

### Step 6: Artifact Filtering

Apply additional filters specific to RNA-seq artifacts. These include filters for strand bias, which can indicate alignment artifacts, and filters for known RNA editing sites. The [EMBL-EBI Training](https://www.ebi.ac.uk/training) resources provide guidance on variant annotation and filtering strategies.

For indels, use the RNAIndel classification to distinguish somatic indels from germline and artifact calls. RNAIndel leverages features derived from indel sequence context and biological effect in a machine-learning framework, providing a confidence score for each call.

### Step 7: Validation and Reporting

Validate a subset of candidate variants using an orthogonal method. RT-PCR followed by Sanger sequencing can confirm that the variant is present in the RNA. Targeted deep DNA sequencing can confirm whether the variant is present in the genome. The choice of validation method depends on the research question.

Report the validation rate, the number of variants passing each filtering stage, and the final variant list. This information allows readers to assess the reliability of the calls.

## Tools and Their Tradeoffs

The choice of variant calling tool depends on the research question, the availability of matched normal data, and the variant classes of interest. The following sections describe the main options.

### GATK HaplotypeCaller for RNA-Seq

The Genome Analysis Toolkit developed at the Broad Institute provides state-of-the-art pipelines for germline and somatic variant discovery and genotyping. For RNA-seq data, the fully validated GATK pipeline is a per-sample workflow. This workflow includes splice-aware alignment, duplicate marking, junction splitting, base quality recalibration, and variant calling with HaplotypeCaller.

The per-sample nature of the validated RNA-seq workflow means that joint genotyping across multiple samples requires combining commands from distinct workflows. Recent methodological work has described how to perform joint genotyping on RNA-seq samples, extending the standard approach. This extension is useful for studies with multiple samples where consistent genotyping across samples is important.

GATK HaplotypeCaller performs local realignment around candidate variants and uses a Bayesian model to assign genotype likelihoods. This approach is robust to alignment errors and provides accurate variant calls in most contexts. The main limitation is that the per-sample workflow does not leverage information from other samples, which can reduce sensitivity for low-frequency variants.

### VarScan2 for Somatic Calling

VarScan2 uses a different algorithmic approach based on read depth and allele frequency thresholds. It can operate in multiple modes: germline calling for single samples, somatic calling for tumor-normal pairs, and copy number analysis. For RNA-seq data, VarScan2 can be applied to tumor samples with or without matched normals.

In somatic mode, VarScan2 compares the tumor and normal samples and classifies variants as somatic, germline, or loss of heterozygosity. This classification requires matched normal data. Without a matched normal, VarScan2 can still call variants but cannot distinguish somatic from germline events.

VarScan2 is computationally efficient and works well with RNA-seq data. Its main limitation is that it does not perform local realignment, so it may miss variants in regions with complex alignment patterns.

### RNAIndel for Indel Detection

RNAIndel was developed specifically for somatic indel detection in tumor RNA-seq data. The motivation for this tool was the unmet need for reliable identification of expressed somatic indels. PCR-based RNA-seq library preparation generates artifacts, and the lack of normal RNA-seq data presents analytical challenges for somatic indel discovery.

RNAIndel uses a machine-learning framework that incorporates features derived from indel sequence context and biological effect. The tool classifies indels as somatic, germline, or artifact. In evaluations across diverse test datasets of pediatric and adult cancers, RNAIndel robustly predicted a high percentage of somatic indels, including subclonal driver indels missed by targeted deep sequencing. The tool outperformed the current best practice for RNA-seq variant calling, which had lower sensitivity and substantially more false positives.

Researchers studying indels in tumor RNA-seq data should use RNAIndel instead of relying solely on general-purpose callers. The tool is freely available and can be integrated into existing pipelines.

### Emerging Machine Learning Approaches

Newer methods are applying machine learning to the RNA-seq variant calling problem. One recent approach, VarRNA, uses two XGBoost machine learning models to classify transcriptome variants as germline, somatic, or artifact. This method was trained and validated using pediatric cancer samples with paired tumor and normal DNA exome sequencing as ground truth.

VarRNA identifies a substantial fraction of the variants detected by exome sequencing and also detects unique RNA variants absent in paired tumor and normal DNA exome data. Some variants classified by VarRNA exhibit variant allele frequencies distinct from the corresponding DNA exome data. This phenomenon is prevalent in cancer-driving genes, where RNA analysis reveals variant allele expression much higher than expected based on exome sequencing.

These machine learning approaches represent an active area of development. Researchers should monitor the literature for new tools, but established methods such as GATK and VarScan2 remain appropriate for most applications.

## Observations and Measurements

Systematic measurement of pipeline performance is essential for reliable somatic variant calling from RNA-seq data. The following metrics provide a basis for evaluating and comparing different approaches.

### Sensitivity and Precision

Sensitivity measures the fraction of true somatic variants that the pipeline detects. Precision measures the fraction of called variants that are true somatic events. These two metrics trade off against each other: increasing sensitivity typically decreases precision and vice versa.

For RNA-seq variant calling, sensitivity is inherently limited by expression level. Variants in lowly expressed genes may fall below the detection threshold even with deep sequencing. The recent VarRNA study found that RNA-based calling identifies approximately half of the variants detected by exome sequencing, reflecting this expression-based limitation.

Precision is affected by artifacts from splicing, RNA editing, and library preparation. Standard RNA-seq variant calling pipelines have been shown to produce high false positive rates for indels, motivating the development of specialized tools like RNAIndel.

### Variant Allele Frequency Distributions

The distribution of variant allele frequencies in RNA-seq data differs from DNA-seq data due to allele-specific expression. Researchers should examine the allele frequency distribution of their calls to identify potential artifacts. A cluster of variants at very low allele frequencies may indicate sequencing errors, while a cluster at intermediate frequencies may indicate germline contamination or RNA editing.

Comparing RNA-seq allele frequencies to matched DNA-seq data, when available, reveals the extent of allele-specific expression. Variants with substantially different allele frequencies between RNA and DNA may be subject to regulatory mechanisms that affect expression.

### Validation Rates

The validation rate is the fraction of called variants confirmed by an orthogonal method. This metric provides the most direct assessment of pipeline accuracy. Researchers should validate a representative sample of calls, including variants across the allele frequency spectrum and in different sequence contexts.

Validation rates below 50 percent indicate that the pipeline is producing excessive false positives and requires more stringent filtering. Validation rates above 90 percent suggest that the pipeline is conservative and may be missing true variants.

## Records and Documentation

Maintaining detailed records of the variant calling pipeline is essential for reproducibility and interpretation. The following records should be preserved for each analysis.

### Pipeline Configuration

Record the exact versions of all software tools, the reference genome version, and all parameter settings. The [nf-core Documentation](https://nf-co.re/docs) describes community standards for pipeline usage and configuration that support reproducible analysis. Version control of analysis scripts using Git, as taught in [The Carpentries Lessons](https://carpentries.org/lessons), ensures that the analysis can be reproduced exactly.

### Quality Metrics

Record quality metrics at each pipeline stage. These include raw read quality, alignment rates, duplication rates, coverage statistics, and variant call counts. These metrics allow researchers to identify samples that deviate from expected patterns and may require additional processing.

### Filtering Decisions

Document every filtering step and the rationale for each threshold. Filtering decisions substantially affect the final variant list, and transparent documentation allows others to assess the impact of these decisions. The [Bioconductor](https://bioconductor.org/) project provides reproducible genomic-analysis documentation standards that support this level of record keeping.

### Variant Call Files

Preserve the unfiltered variant call files in addition to the filtered versions. This practice allows re-analysis with different filtering criteria without rerunning the computationally expensive alignment and variant calling steps.

## Common Failure Patterns

Understanding common failure patterns helps researchers diagnose problems in their variant calling pipelines.

### Excessive False Positives from Splicing Artifacts

The most common failure pattern is a high false positive rate caused by splicing artifacts. Reads spanning exon-exon junctions can be misaligned, producing apparent variants that do not exist in the genome. This problem is most severe in genes with complex splicing patterns or in regions with homologous sequences.

Diagnosis: Examine the genomic distribution of called variants. A cluster of variants near splice junctions or in repetitive regions suggests alignment artifacts. Strand bias, where variants appear predominantly on one strand, also indicates alignment problems.

Remediation: Ensure that the two-pass alignment strategy is used, that reads are properly split at junctions, and that strand bias filters are applied.

### Low Sensitivity Due to Expression Level

RNA-seq variant calling misses variants in lowly expressed genes. This limitation is inherent to the approach and cannot be fully overcome by deeper sequencing. Researchers who need comprehensive variant detection should use DNA sequencing.

Diagnosis: Compare the expression levels of genes with detected variants to the overall expression distribution. If variants are only detected in highly expressed genes, the pipeline is working as expected but the data cannot support detection in lowly expressed genes.

Remediation: If specific genes are of interest, consider targeted RNA sequencing with higher coverage for those genes, or use DNA sequencing for comprehensive variant detection.

### Indel False Positives from Library Preparation

PCR-based RNA-seq library preparation generates artifacts that appear as indels. These artifacts are particularly problematic because they can occur at homopolymer runs and other repetitive sequences. Standard variant calling pipelines have high false positive rates for indels in RNA-seq data.

Diagnosis: Examine the sequence context of called indels. A high proportion of indels in homopolymer regions suggests PCR artifacts. Compare indel calls to known germline indel databases to assess the fraction of calls that are likely artifacts.

Remediation: Use RNAIndel for indel calling, which was specifically designed to distinguish somatic indels from artifacts in tumor RNA-seq data.

### Germline Contamination Misclassified as Somatic

Without matched normal data, germline variants can be misclassified as somatic. This problem is particularly acute for common polymorphisms that are present in the population at high frequency.

Diagnosis: Compare called variants to population databases. A high fraction of calls matching common polymorphisms suggests that germline filtering is insufficient.

Remediation: Apply more stringent filtering against population databases, or obtain matched normal samples for definitive somatic classification.

## Limitations and Interpretation Boundaries

RNA-seq variant calling has inherent limitations that researchers must acknowledge when interpreting results.

### Expression-Dependent Detection

RNA-seq can only detect variants in expressed genes. The detection limit depends on the expression level of the gene, the sequencing depth, and the variant allele frequency. Variants in genes with low or absent expression are invisible to this approach. This limitation means that RNA-seq variant calling cannot provide a complete picture of the tumor genome.

### Allele-Specific Expression Confounds Frequency Interpretation

The variant allele frequency in RNA-seq data reflects both the proportion of cells carrying the variant and the relative expression of the variant allele. Researchers cannot infer cellular prevalence from RNA-seq allele frequencies without additional information. The recent finding that cancer-driving genes can show variant allele expression much higher than expected based on DNA sequencing highlights this interpretational challenge.

### RNA Editing Mimics Somatic Mutation

RNA editing events produce sequence changes that are not present in the genome. Without filtering against known editing sites or validating with DNA sequencing, these events will be misclassified as somatic mutations. The biological significance of RNA editing differs fundamentally from genomic mutation, and conflating the two leads to incorrect interpretations.

### Absence of Evidence Is Not Evidence of Absence

A variant that is not detected in RNA-seq data may still be present in the genome. The variant may be in a gene that is not expressed in the sampled tissue, or the expression level may be too low for detection. Negative results from RNA-seq variant calling should be interpreted with this limitation in mind.

## Quality Controls and Reproducibility

Reproducibility requires careful attention to computational environments, documentation, and quality control at each pipeline stage.

### Computational Environment Management

Use container technologies or package managers to ensure that the computational environment is reproducible. The [nf-core Documentation](https://nf-co.re/docs) describes community standards for pipeline usage and configuration that support reproducible analysis. Container images that include all software dependencies can be archived and reused.

### Version Control for Analysis Scripts

Maintain analysis scripts in a version control system such as Git. The [Carpentries Lessons](https://carpentries.org/lessons) provide foundational training in Git and other computing skills that support reproducible research. Each analysis should be associated with a specific commit that can be retrieved and re-executed.

### Quality Control Checkpoints

Establish quality control checkpoints at each pipeline stage. Define pass and fail criteria before running the analysis, and document the outcomes at each checkpoint. This approach prevents low-quality data from propagating through the pipeline and producing unreliable results.

### Training and Skill Development

Researchers new to RNA-seq variant calling should complete structured training before analyzing their own data. The [EMBL-EBI Training](https://www.ebi.ac.uk/training) program offers learning pathways for bioinformatics data analysis, and the [Galaxy Training Network](https://training.galaxyproject.org/) provides accessible workflow tutorials. The [Bioconductor](https://bioconductor.org/) project offers documentation for reproducible genomic analysis in R.

## Professional Escalation Criteria

Researchers should seek expert assistance when they encounter situations beyond their current expertise. The following criteria indicate when escalation is appropriate.

### Unusual Variant Patterns

If the variant calling results show unexpected patterns, such as an extremely high variant count, a skewed allele frequency distribution, or variants concentrated in specific genomic regions, consult with a bioinformatics specialist before proceeding with interpretation.

### Validation Failures

If a substantial fraction of called variants fail validation, the pipeline requires revision. A bioinformatics specialist can help diagnose whether the problem lies in alignment, variant calling, or filtering.

### Clinical or Diagnostic Applications

RNA-seq variant calling for clinical or diagnostic purposes requires specialized expertise and validation. Researchers should consult with molecular pathology or clinical genomics specialists before using RNA-seq variant calls for patient-related decisions.

### Novel or Unusual Sample Types

Samples with unusual characteristics, such as very low RNA quality, high degradation, or contamination, may require specialized processing. Consult with a laboratory specialist to determine whether the data are suitable for variant calling.

## A Practical Decision Framework for RNA-Seq Somatic Variant Calling

Choosing whether to call somatic variants from RNA-seq data and which pipeline to run requires a structured decision process that accounts for the biological question, available samples, and data characteristics. The following framework translates the technical considerations described above into concrete decisions a researcher can make before spending compute time and analytical effort.

### Step 1: Define the Biological Question

The first decision determines whether RNA-seq variant calling is appropriate at all. Ask what the variant information will be used for. If the goal is to identify which expressed mutations drive tumor behavior, RNA-seq variant calling provides direct evidence of expression that DNA sequencing cannot offer. If the goal is comprehensive genomic characterization, DNA sequencing is the appropriate method and RNA-seq should not be used as a substitute.

Record the specific question in one sentence before starting. This sentence guides every subsequent decision. A question such as "Which expressed somatic mutations distinguish this tumor subtype from others" justifies RNA-seq variant calling. A question such as "What is the complete mutational landscape of this tumor" requires DNA sequencing.

### Step 2: Assess Sample Availability

Inventory the samples available for the study. The presence or absence of matched normal material changes the analytical approach substantially. Three scenarios exist.

The first scenario has matched normal DNA exome data available for the same tumor. This scenario is the strongest design because the DNA data provide ground truth for validating RNA-seq calls and for distinguishing somatic from germline variants. The VarRNA method was trained and validated using paired tumor and normal DNA exome sequencing as ground truth, demonstrating the value of this design.

The second scenario has matched normal RNA available. This design allows direct subtraction of germline variants expressed in normal tissue, but it does not resolve the question of whether a variant is truly somatic or simply expressed in both tissues.

The third scenario has no matched normal material. This is the most common situation in retrospective RNA-seq studies. Without matched normal data, the analysis must rely on population databases and artifact filtering to exclude germline polymorphisms. Expect higher false positive rates and plan for orthogonal validation of candidate variants.

### Step 3: Evaluate Data Characteristics Before Calling

Before running any variant caller, examine the alignment files and assess whether the data can support variant calling. Three metrics matter most.

First, check the percentage of reads mapping to exonic regions. RNA-seq libraries with substantial intronic or intergenic content may indicate genomic DNA contamination during library preparation. High contamination levels produce variants that appear somatic but actually reflect the tumor genome instead of the transcriptome, confounding the interpretation of expressed variants.

Second, examine the distribution of coverage across genes of interest. If the genes most relevant to the biological question have mean depths below 20 reads, variant calling in those genes will be unreliable regardless of the tool used. The expression-dependent detection limit means that lowly expressed genes may fall below the threshold for confident calls.

Third, assess the duplication rate. High duplication rates reduce the effective number of independent observations supporting each variant call. PCR-based RNA-seq library preparation generates artifacts that appear as indels, and high duplication rates amplify this problem.

### Step 4: Select the Calling Strategy Based on Variant Class

The variant class of primary interest determines the tool selection. For single nucleotide variants, GATK HaplotypeCaller following the per-sample RNA-seq workflow is the established approach. The fully validated GATK pipeline for RNA-seq data is a per-sample workflow, and recent methodological work has described how to extend this approach to joint genotyping across multiple samples.

For insertions and deletions, general-purpose callers produce unacceptably high false positive rates in RNA-seq data. RNAIndel was developed specifically for somatic indel detection in tumor RNA-seq data and uses machine learning to classify indels as somatic, germline, or artifact. In evaluations across diverse test datasets, RNAIndel robustly predicted a high percentage of somatic indels while the current best practice for RNA-seq variant calling had lower sensitivity and substantially more false positives.

For studies requiring both SNVs and indels, run both tools and integrate the results. Use GATK HaplotypeCaller for SNVs and RNAIndel for indels, then combine the high-confidence calls from each tool.

### Step 5: Apply Tiered Filtering Based on Confidence

Filtering should proceed in tiers, with each tier removing a distinct class of artifacts. The first tier removes low-quality calls based on read depth, base quality, and mapping quality. The second tier removes alignment artifacts by filtering variants near splice junctions and applying strand bias filters. The third tier removes RNA editing events by filtering against known editing sites. The fourth tier removes germline polymorphisms by filtering against population databases.

Document the number of variants removed at each tier. This record provides transparency about the filtering stringency and allows readers to assess the reliability of the final variant list. A pipeline that removes 90 percent of initial calls at the germline filtering tier is behaving differently from one that removes 10 percent, and this difference matters for interpretation.

### Step 6: Plan Validation Before Running the Pipeline

Decide on the validation strategy before generating variant calls. This decision prevents the common failure of completing the analysis and then discovering that no appropriate validation method is available. For confirming that a variant is expressed, RT-PCR followed by Sanger sequencing is appropriate. For confirming that a variant exists in the genome, targeted deep DNA sequencing is required.

Select a representative sample of variants for validation across the allele frequency spectrum. Validating only high-confidence calls overestimates pipeline accuracy. Include variants at low allele frequencies and in challenging sequence contexts to obtain a realistic assessment of performance.

### Step 7: Record Decisions and Outcomes

Maintain a decision log that records the biological question, sample inventory, data quality metrics, tool selections, filtering thresholds, and validation results. This log serves two purposes. It supports reproducibility by documenting the exact analysis decisions, and it provides a basis for comparing results across studies.

The [nf-core Documentation](https://nf-co.re/docs) describes community standards for pipeline usage and configuration that support reproducible analysis. Version control of analysis scripts using Git, as taught in [The Carpentries Lessons](https://carpentries.org/lessons), ensures that the analysis can be reproduced exactly.

### Decision Matrix for Common Scenarios

The following matrix summarizes the recommended approach for common study designs.

| Scenario | Recommended Approach | Primary Limitation |
|---|---|---|
| Tumor RNA only, SNV focus | GATK HaplotypeCaller per-sample workflow with population database filtering | Cannot definitively distinguish somatic from germline variants |
| Tumor RNA only, indel focus | RNAIndel for indel classification | Requires careful interpretation of artifact calls |
| Tumor RNA with matched normal DNA | GATK HaplotypeCaller plus DNA-based validation of candidate variants | DNA data may not cover all expressed variants |
| Tumor RNA with matched normal RNA | GATK HaplotypeCaller with normal subtraction | Normal RNA may not express all germline variants |
| Multiple tumor RNA samples | GATK HaplotypeCaller with joint genotyping extension | Requires careful batch effect assessment |

### Escalation Criteria Within the Decision Framework

Escalate to a bioinformatics specialist when the decision framework reveals conditions beyond routine analysis. These conditions include samples with very low exonic mapping rates, studies requiring clinical or diagnostic interpretation, and situations where validation rates fall below 50 percent. The [EMBL-EBI Training](https://www.ebi.ac.uk/training) program offers learning pathways for bioinformatics data analysis, and the [Galaxy Training Network](https://training.galaxyproject.org/) provides accessible workflow tutorials for researchers developing these skills.

## Frequently Asked Questions

### What is the main difference between somatic variant calling from RNA-seq and DNA-seq?

DNA-seq surveys the full genome regardless of expression, while RNA-seq only detects variants in transcribed regions. RNA-seq data also contain artifacts from splicing, RNA editing, and allele-specific expression that are absent from DNA-seq data. These differences require specialized alignment strategies and filtering approaches for RNA-seq variant calling.

### Can I call somatic variants from tumor RNA-seq data without a matched normal sample?

Yes, but the analysis is more challenging. Without a matched normal, you cannot definitively distinguish somatic variants from germline polymorphisms. Filtering against population databases can remove common germline variants, but rare germline variants will remain in the candidate list. Validation with orthogonal methods is particularly important when no matched normal is available.

### Why do variant allele frequencies differ between RNA-seq and DNA-seq data?

Variant allele frequency in RNA-seq reflects both the proportion of cells carrying the variant and the relative expression of the variant and reference alleles. Allele-specific expression can push the RNA-seq allele frequency higher or lower than the DNA-seq allele frequency. This effect can be pronounced in cancer-driving genes, where the mutant allele may be preferentially expressed.

### How should I handle RNA editing events in my variant calls?

RNA editing produces A-to-G changes that are not present in the genome. Filter against known RNA editing sites using appropriate databases, and consider validating candidate variants with DNA sequencing. If DNA validation is not possible, flag RNA editing as a possible explanation for A-to-G variants.

### What is the best tool for calling indels from tumor RNA-seq data?

RNAIndel was specifically designed for somatic indel detection in tumor RNA-seq data and outperforms general-purpose pipelines for this variant class. The tool uses machine learning to classify indels as somatic, germline, or artifact. For comprehensive variant calling, run RNAIndel in addition to a general-purpose caller such as GATK HaplotypeCaller.

### How much sequencing depth do I need for RNA-seq variant calling?

The required depth depends on the expression level of the genes of interest and the variant allele frequency you need to detect. Highly expressed genes can support variant calling at modest sequencing depths, while lowly expressed genes require much deeper sequencing. There is no universal depth threshold, and researchers should assess coverage for their genes of interest before drawing conclusions.

### Can RNA-seq variant calling detect subclonal mutations?

RNA-seq can detect subclonal mutations if the mutant allele is expressed at sufficient levels. RNAIndel has been shown to recover subclonal driver indels with variant allele frequencies in the range of 0.01 to 0.15 that were missed by targeted deep DNA sequencing. However, detection depends on expression level, and subclonal mutations in lowly expressed genes will be missed.

### How should I validate somatic variants called from RNA-seq data?

Validation methods include RT-PCR followed by Sanger sequencing to confirm the variant is present in RNA, and targeted deep DNA sequencing to confirm the variant is present in the genome. The choice depends on whether you need to confirm expression or genomic status. Validate a representative sample of calls across the allele frequency spectrum to assess overall pipeline accuracy.

## Related Bioinformatics Guides

- [Genomic Data Analysis Tools: A Comparative Guide for Researchers](/knowledge/bioinformatics/genomic-data-analysis-tools-a-comparative-guide-for-researchers)
- [RNA-Seq Alignment Tools: STAR, HISAT2, and Beyond](/knowledge/bioinformatics/rna-seq-alignment-tools-star-hisat2-and-beyond)
- [RNA-Seq Databases: Accessing and Using Public RNA-Seq Data](/knowledge/bioinformatics/rna-seq-databases-accessing-and-using-public-rna-seq-data)
- [Alternative Splicing Analysis from RNA-Seq Data](/knowledge/bioinformatics/alternative-splicing-analysis-from-rna-seq-data)
- [RNA-Seq Data Analysis in Galaxy: A User-Friendly Platform](/knowledge/bioinformatics/rna-seq-data-analysis-in-galaxy-a-user-friendly-platform)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)


## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [Variant Calling from RNA-seq Data Using the GATK Joint Genotyping Workflow.](https://pubmed.ncbi.nlm.nih.gov/35751817). Methods in molecular biology (Clifton, N.J.), 2022.
- [Variant calling from RNA-Seq data reveals allele-specific differential expression of pathogenic cancer variants.](https://pubmed.ncbi.nlm.nih.gov/40437029). Communications medicine, 2025.
- [Somatic variant calling from single-cell DNA sequencing data.](https://pubmed.ncbi.nlm.nih.gov/35782734). Computational and structural biotechnology journal, 2022.
- [RNAIndel: discovering somatic coding indels from tumor RNA-Seq data.](https://pubmed.ncbi.nlm.nih.gov/31593214). Bioinformatics (Oxford, England), 2020.
- [Identifying cancer cells from calling single-nucleotide variants in scRNA-seq data.](https://pubmed.ncbi.nlm.nih.gov/39163479). Bioinformatics (Oxford, England), 2024.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.


<div data-calculator="genetics"></div>