# Filtering Variants from RNA-Seq Data: Best Practices for Handling Mapping Artifacts and Splice Junctions


## Key Takeaways

- **Splice-aware alignment is foundational:** Unlike DNA-seq, RNA-seq reads span exon-exon junctions, necessitating aligners (e.g., STAR, HISAT2) that split reads and report gapped alignments to avoid misinterpretation of sequence continuity.
- **Coverage reflects expression, not genome representation:** Depth-based filters must be adjusted for gene expression levels, as low coverage in lowly expressed genes can lead to the exclusion of true variants, while high coverage in repetitive regions can generate false positives.
- **Base quality recalibration requires caution:** Standard recalibration using DNA-derived variant sites may not accurately reflect RNA-seq error profiles due to reverse transcription and PCR biases; hard filtering or RNA-specific recalibration is often preferred.
- **Strand information interpretation is protocol-dependent:** Stranded library preparation preserves RNA orientation, providing biological context for strand bias annotations, whereas unstranded libraries necessitate careful interpretation to distinguish biological patterns from mapping artifacts.
- **GATK's RNA-seq workflow addresses specific challenges:** The pipeline includes critical steps like splitting reads at intron boundaries (using `SplitNCigarReads`) and recommends hard filtering over variant quality score recalibration due to differing annotation distributions in RNA-seq data.
- **RNA editing must be distinguished from genomic variants:** Post-transcriptional A-to-I editing events appear as A-to-G variants in RNA-seq data and require comparison with DNA sequencing or filtering against known editing sites to avoid confounding genomic variant calls.

---

RNA sequencing generates expression data, but researchers increasingly want to extract genetic variant information from the same datasets. Calling variants from RNA-seq data requires a different analytical approach than DNA sequencing because transcriptome reads span splice junctions, coverage tracks expression levels instead of genomic representation, and mapping errors concentrate in repetitive and paralogous regions. This article provides practical filtering strategies for RNA-seq variant calling, covering junction-aware alignment, depth and quality filters, and the use of established pipelines such as the Genome Analysis Toolkit RNA-seq short variant calling workflow.

The intended reader is a biology student, researcher, or laboratory professional who has generated RNA-seq data and wants to call variants reliably. The focus is on germline variant calling, with attention to somatic variant calling where relevant. The guidance assumes familiarity with basic bioinformatics concepts but does not require prior experience with variant calling. The practical outcome is a defensible filtering workflow that separates true biological variants from alignment artifacts and sequencing errors.

## Why RNA-Seq Variant Calling Differs from DNA-Based Approaches

DNA-based variant calling pipelines assume that reads align continuously to the reference genome. RNA-seq data violates this assumption because introns are removed during splicing. A read that spans an exon-exon junction will not align as a continuous sequence to the reference genome. Instead, the alignment software must split the read across the junction, creating a gapped alignment with distinct segments.

Standard aligners that do not account for splicing will either discard junction-spanning reads or map them incorrectly. This produces two problems. First, coverage drops in regions where many reads span junctions, reducing confidence in variant calls. Second, misaligned reads can create false variant calls when a read from one exon is forced to align to a similar sequence elsewhere in the genome.

The Genome Analysis Toolkit, commonly called GATK, has been a standard tool for short-read variant calling since its release in 2010. It remains widely used because it provides a coordinated set of analysis tools that handle the full variant calling process. The toolkit includes specific procedures for RNA-seq data that address the splice junction problem. The [GATK review](https://doi.org/10.3390/ijms27093754) describes the toolkit as a collection of more than 400 analysis tools that have unified variant calling practices across the field through continuous adaptation to new sequencing technologies and analysis challenges.

RNA-seq variant calling also differs from DNA approaches because the data generation process introduces distinct error modes. Library preparation for RNA-seq involves reverse transcription, which can introduce errors and biases. PCR amplification during library preparation creates duplicate reads that inflate coverage estimates. The transcriptome itself contains post-transcriptional modifications such as RNA editing that appear as variants in the sequencing data but are not present in the genomic DNA.

## Core Principles of RNA-Seq Variant Filtering

### Junction-Aware Alignment Is the First Filter

The most important decision in RNA-seq variant calling happens before variant calling begins. The aligner must be splice-aware. Tools such as STAR and HISAT2 split reads across exon-exon junctions and report the genomic coordinates of each segment. This alignment information determines everything that follows.

When a read is split across a junction, the alignment record contains supplementary alignments. Variant callers must be configured to handle these records correctly. If the caller treats a split read as two independent fragments, it may double-count evidence or fail to recognize that both segments come from the same original read.

The [Galaxy Training Network](https://training.galaxyproject.org/) provides accessible tutorials on RNA-seq analysis workflows, including alignment and variant calling steps. These tutorials demonstrate how to configure splice-aware aligners and how to inspect alignment quality before proceeding to variant calling.

### Base Quality Scores Drive Variant Confidence

Base quality scores indicate the probability that a sequenced base is correct. These scores are assigned by the sequencing instrument and can be recalibrated using known variant sites. In RNA-seq data, base quality recalibration requires care because the variant sites used for recalibration are typically derived from DNA sequencing projects.

The [EMBL-EBI Training](https://www.ebi.ac.uk/training) program offers learning pathways that cover quality assessment and data processing for sequencing projects. Understanding how quality scores are generated and how to assess them is essential for interpreting variant calls.

### Read Depth Requirements Reflect Expression, Not Genomic Representation

DNA sequencing aims for uniform coverage across the genome. RNA-seq coverage reflects gene expression levels. Highly expressed genes may have hundreds of reads at a position, while lowly expressed genes may have only a few. This variation means that depth-based filters must be applied with awareness of expression levels.

A variant in a lowly expressed gene may have only five or ten supporting reads, while a variant in a highly expressed gene may have hundreds. Setting a single minimum depth threshold will discard legitimate variants in lowly expressed genes. Conversely, high depth in repetitive regions can produce false variant calls that pass depth filters.

### Strand Information Has Different Meaning in RNA-Seq

RNA-seq library preparation can preserve strand information, depending on the protocol used. Stranded libraries retain the orientation of the original RNA molecule, while unstranded libraries lose this information. Variant callers use strand bias annotations to detect artifacts where variant-supporting reads come predominantly from one strand.

In RNA-seq data, strand bias must be interpreted with caution. Genes with strand-specific expression patterns may show apparent strand bias that reflects biology instead of mapping error. The interpretation of strand bias annotations depends on whether the library was prepared with a stranded or unstranded protocol.

## The GATK RNA-Seq Short Variant Calling Pipeline

The GATK best practices for RNA-seq variant calling provide a structured workflow that addresses the unique features of transcriptome data. The [GATK review](https://doi.org/10.3390/ijms27093754) describes how the toolkit has maintained its position as a standard approach through continuous adaptation to new sequencing technologies and analysis challenges.

### Pipeline Overview

The RNA-seq short variant calling workflow includes these major steps:

1. Splice-aware alignment of reads to the reference genome
2. Marking duplicate reads that arise from PCR amplification
3. Splitting reads that span introns into exon segments
4. Recalibrating base quality scores
5. Calling variants with HaplotypeCaller in a mode appropriate for RNA-seq
6. Filtering variants based on quality metrics

The split-read step is critical. GATK provides a tool that identifies reads spanning introns and splits them into separate records at exon boundaries. This prevents the variant caller from being confused by the large gaps in alignment that introns create.

### Variant Filtering with Variant Quality Score Recalibration

GATK offers two approaches to filtering variants after calling. The first is variant quality score recalibration, which uses machine learning to model the relationship between variant annotations and the likelihood that a variant is real. This approach requires a set of known variant sites for training.

The second approach is hard filtering, which applies fixed thresholds to variant annotations. Hard filtering is simpler and does not require training data. For RNA-seq data, hard filtering is often more practical because the variant annotations behave differently than in DNA data.

A study of [Parkinson's disease exome data](https://pubmed.ncbi.nlm.nih.gov/32348865) used the GATK best practices pipeline for variant discovery and then applied filter-based annotation to categorize variants. This study identified thousands of SNPs and INDELs and prioritized a subset of missense variants based on predicted functional effects. The workflow demonstrates how variant filtering connects to biological interpretation.

## At a Glance: RNA-Seq Variant Filtering Decisions

| Filtering Decision | DNA-Seq Standard Practice | RNA-Seq Recommended Practice | Rationale |
| --- | --- | --- | --- |
| Alignment | Continuous alignment to reference | Splice-aware alignment with junction support | Reads spanning exon-exon junctions require gapped alignment |
| Read handling | Keep full reads | Split reads at intron boundaries before variant calling | Prevents confusion from large alignment gaps |
| Depth threshold | Uniform genome-wide threshold | Expression-aware thresholds per gene or transcript | Coverage reflects expression level, not genomic representation |
| Base quality recalibration | Standard recalibration with known variant sites | Recalibrate with caution or use hard filters | Known variant sites may not represent transcriptome diversity |
| Variant filtering | Variant quality score recalibration | Hard filtering or recalibration with RNA-specific training | RNA-seq variant annotations have different distributions |
| Strand bias interpretation | Apply standard strand bias filters | Interpret with knowledge of library strand protocol | Stranded libraries produce biologically meaningful strand patterns |

## Practical Workflow for RNA-Seq Variant Calling

### Step 1: Assess Input Data Quality

Before alignment, examine the raw sequencing data. Check per-base quality scores, GC content, adapter contamination, and duplication rates. The [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/) provide access to sequence databases and analysis services that can help with quality assessment and data management.

Poor quality bases at read ends will create false variant calls. Trimming low-quality bases and adapter sequences before alignment reduces this problem. However, trimming must be balanced against the need to retain junction-spanning reads, which may be short after trimming.

Document the sequencing platform, read length, and library preparation protocol. These details affect downstream filtering decisions. For example, longer reads provide more reliable junction-spanning alignments, while shorter reads require more careful handling of multi-mapping reads.

### Step 2: Align with a Splice-Aware Aligner

Choose an aligner that explicitly models splicing. STAR and HISAT2 are common choices. Both produce alignment files in SAM or BAM format that include junction information.

After alignment, inspect the alignment statistics. The percentage of reads that map uniquely, the number of reads that span junctions, and the distribution of insert sizes all indicate alignment quality. The [nf-core documentation](https://nf-co.re/docs) describes community standards for pipeline configuration and quality control that apply to alignment steps.

Record the aligner version, reference genome version, and all alignment parameters. These records are essential for reproducing the analysis and for troubleshooting unexpected results.

### Step 3: Mark Duplicates

PCR amplification during library preparation creates duplicate reads that represent the same original fragment. These duplicates inflate coverage and can bias variant allele frequency estimates. Marking duplicates allows the variant caller to account for them.

For RNA-seq data, duplicate marking is complicated by the fact that reads from highly expressed genes may share the same start position by chance. This is especially true for short reads from the same transcript. Some pipelines skip duplicate marking for RNA-seq for this reason, while others mark duplicates but interpret the results with caution.

The decision to mark duplicates should be documented and justified. If duplicates are marked, compare the variant calls with and without duplicate marking to understand the impact on the final results.

### Step 4: Split Reads at Junctions

Use the GATK SplitNCigarReads tool to split reads at intron boundaries. This tool examines the CIGAR strings in the alignment file and identifies reads that span introns. It then splits these reads into separate records, each corresponding to one exon segment.

The split step also filters out reads that do not meet quality thresholds. Reads with low mapping quality or excessive mismatches are removed at this stage.

After splitting, verify that the number of reads spanning junctions is consistent with expectations for the organism and tissue type. An unusually low number of junction-spanning reads may indicate that the aligner was not configured correctly for splice-aware alignment.

### Step 5: Recalibrate Base Qualities

Base quality score recalibration adjusts quality scores based on empirical error rates. The standard GATK recalibration workflow uses known variant sites from databases such as dbSNP. For RNA-seq data, the known variant sites may not be representative because they are derived from DNA sequencing of diverse populations.

An alternative is to skip recalibration and rely on hard filtering of variant calls. This approach is simpler and avoids the risk of overfitting the recalibration model to DNA-derived variant sets.

If recalibration is performed, document the training set used and the number of sites available for building the recalibration model. A small training set may produce an unreliable model.

### Step 6: Call Variants

Run HaplotypeCaller with the mode appropriate for RNA-seq data. The caller will use the split reads to assemble haplotypes and identify variant sites. The output is a VCF file containing raw variant calls with annotations such as read depth, allele frequency, and quality scores.

The [Bioconductor project](https://bioconductor.org/) provides R packages for variant annotation and filtering that can be integrated into RNA-seq analysis workflows. These packages allow researchers to apply reproducible filtering criteria and generate reports.

### Step 7: Filter Variant Calls

Apply filters to remove false positives. Common filters for RNA-seq variants include:

- Read depth at the variant site
- Allele balance, which measures the proportion of reads supporting each allele
- Mapping quality of reads supporting the variant
- Strand bias, which detects variants supported only by reads on one strand
- Quality by depth, which normalizes variant quality by read depth

Document the threshold chosen for each filter and the number of variants removed by each filter. This documentation allows reviewers to assess whether the filtering strategy was appropriate for the data.

### Step 8: Annotate and Interpret

After filtering, annotate the remaining variants with gene information, predicted functional effects, and population frequencies. The [Parkinson's disease study](https://pubmed.ncbi.nlm.nih.gov/32348865) used ANNOVAR for filter-based annotation and then applied SIFT and PolyPhen2 to predict the damaging effects of missense variants. This approach connects variant filtering to biological interpretation.

The [sarcoma multi-omics study](https://doi.org/10.1007/s00262-026-04395-y) used RNA-seq as an orthogonal validation check for variants identified from whole-exome sequencing. The study used RNA-seq to confirm that variants were transcribed and to distinguish functional amplifications from technical depth artifacts. This approach demonstrates how RNA-seq can complement DNA-based variant calling.

## Options and Tradeoffs in Variant Filtering

### Hard Filtering versus Machine Learning Filtering

Hard filtering applies fixed thresholds to variant annotations. For example, a researcher might require a minimum read depth of 10 and a minimum quality score of 30. Hard filtering is transparent and reproducible, but it requires the researcher to choose thresholds that may not be optimal for all regions of the transcriptome.

Machine learning filtering, such as variant quality score recalibration, learns the relationship between annotations and variant quality from training data. This approach can capture complex patterns that hard thresholds miss. However, it requires a reliable training set and can produce unpredictable results when applied to data that differs from the training set.

For RNA-seq data, hard filtering is often the safer choice because the annotation distributions differ from DNA data. The [GATK review](https://doi.org/10.3390/ijms27093754) notes that the toolkit has adapted its recommendations over time as the community has learned which approaches work best for different data types.

### Germline versus Somatic Variant Calling

Germline variant calling assumes that variants are present in all cells and follows Mendelian inheritance patterns. Somatic variant calling looks for variants that are present in a subset of cells, such as tumor cells within a normal tissue sample.

RNA-seq data can support both types of calling, but the filtering strategies differ. Somatic calling requires paired normal and tumor samples to distinguish somatic mutations from germline variants and technical artifacts. The allele frequency threshold for somatic variants is often lower because the variant may be present in only a fraction of cells.

The [sarcoma study](https://doi.org/10.1007/s00262-026-04395-y) demonstrated that integrating transcriptomic data can recover actionable targets in tumors with low mutation burden. The study used RNA-seq for expression filtering and as an orthogonal validation check for variant transcription, distinguishing functional amplifications from technical depth artifacts.

### Reference Genome Choice

The choice of reference genome affects variant calling. A reference genome that matches the sample's ancestry or strain will produce fewer false variants than a mismatched reference. For non-model organisms, the reference may be incomplete or contain errors.

The [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/) provide access to reference genomes and annotation databases. Researchers should document the reference genome version and build used in their analysis, as this information is essential for reproducibility.

### Stranded versus Unstranded Library Preparation

The choice of library preparation protocol affects how strand information is interpreted. Stranded libraries preserve the orientation of the original RNA molecule, allowing the researcher to distinguish sense from antisense transcription. Unstranded libraries lose this information.

For variant calling, stranded libraries provide additional information for interpreting strand bias annotations. However, the strand protocol must be documented and accounted for in the analysis. Misinterpreting strand bias in unstranded data can lead to incorrect filtering decisions.

## Records and Measurements for Variant Calling

### Essential Records

Maintain detailed records of the variant calling process. These records should include:

- Sequencing platform and chemistry
- Read length and insert size
- Library preparation protocol, including whether the library is stranded
- Alignment software and version
- Reference genome version and build
- Variant caller and version
- Filtering thresholds and rationale
- Number of variants before and after each filtering step

The [The Carpentries lessons](https://carpentries.org/lessons) teach foundational data management practices that apply to bioinformatics projects. Organizing analysis files, documenting software versions, and maintaining a clear directory structure are essential for reproducible variant calling.

### Quality Metrics to Track

Track these metrics at each stage of the pipeline:

- Total read count before and after trimming
- Alignment rate and percentage of uniquely mapped reads
- Number of reads spanning splice junctions
- Duplicate rate
- Mean coverage across expressed genes
- Number of variant calls before filtering
- Number of variant calls after each filter
- Transition to transversion ratio, which indicates variant quality

A transition to transversion ratio that deviates significantly from the expected value for the organism suggests systematic errors in variant calling. For human data, a ratio around 2.0 is typical for high-quality variant calls.

### Validation Approaches

Validation of RNA-seq variant calls can use several approaches:

- Sanger sequencing of selected variants
- Comparison with DNA sequencing data from the same sample
- Genotyping arrays for known variant sites
- RNA-seq data from replicate samples

The [sarcoma study](https://doi.org/10.1007/s00262-026-04395-y) used RNA-seq to validate variants identified from whole-exome sequencing. This orthogonal validation approach increases confidence in variant calls by confirming that the variant is transcribed.

The [poison frog study](https://doi.org/10.1186/s12864-026-12957-8) demonstrated that RNA-seq can provide thousands of exonic SNPs for population genetics analysis. The study validated the applicability of mRNA sequences to elucidate population structure and demography, showing that transcriptomic genotyping can be a cost-efficient alternative for non-model species.

## Common Failure Patterns in RNA-Seq Variant Calling

### Failure Pattern 1: Excessive False Variants in Repetitive Regions

Repetitive regions of the genome produce ambiguous alignments. Reads from different copies of a repeat may align to the same location, creating apparent variants that are actually alignment artifacts.

**Detection:** Examine the genomic distribution of variant calls. Clusters of variants in known repeat regions suggest alignment problems.

**Response:** Apply mapping quality filters to remove reads with low mapping quality. Consider masking repeat regions or excluding them from analysis.

### Failure Pattern 2: Variant Calls in Lowly Expressed Genes

Lowly expressed genes have few reads, so variant calls in these genes have low depth. These calls may be real but lack statistical support.

**Detection:** Review the depth distribution of variant calls. Variants with very low depth are generally unreliable.

**Response:** Apply a minimum depth filter appropriate for the analysis goals. For germline calling, a minimum depth of 10 is common. For somatic calling, lower depth may be acceptable if the variant has other supporting evidence.

### Failure Pattern 3: Strand Bias in Variant Calls

Strand bias occurs when variant-supporting reads come predominantly from one strand. This pattern often indicates mapping artifacts instead of true variants.

**Detection:** Examine the strand distribution of reads supporting each variant. The Fisher strand bias annotation in VCF files provides a statistical test for strand bias.

**Response:** Apply a strand bias filter to remove variants with significant strand bias. Be cautious with this filter for variants in genes with strand-specific expression and for unstranded library preparations.

### Failure Pattern 4: Variant Calls at Splice Junctions

Variants called at or near splice junctions may be artifacts of misalignment. Reads that span junctions can be misaligned if the junction is not annotated or if the splice site is unusual.

**Detection:** Examine the position of variant calls relative to annotated splice junctions. Variants within a few bases of a junction warrant scrutiny.

**Response:** Use the split-read information to verify that variant-supporting reads align correctly across the junction. Consider excluding variants within a defined distance of splice junctions.

### Failure Pattern 5: Overfiltering of True Variants

Aggressive filtering can remove true variants along with false ones. This is especially problematic for variants in lowly expressed genes or variants with unusual allele frequencies.

**Detection:** Compare the number of variants called with the expected number for the organism and sample type. A dramatic reduction after filtering may indicate overfiltering.

**Response:** Review the distribution of filtered variants. If many variants are removed by a single filter, examine whether the filter threshold is appropriate for RNA-seq data.

### Failure Pattern 6: RNA Editing Confounded with Genomic Variants

RNA editing is a post-transcriptional process that changes the sequence of RNA molecules. The most common form in humans is A-to-I editing, which appears as A-to-G variants in RNA-seq data. These variants are not present in the DNA and will not be detected by DNA sequencing.

**Detection:** Compare RNA-seq variant calls with DNA sequencing data from the same sample. Variants present only in RNA-seq data may represent RNA editing events.

**Response:** If the analysis goal is genomic variant discovery, filter out known RNA editing sites or compare with DNA data to identify editing events.

## Limitations of RNA-Seq Variant Calling

### Coverage Is Expression-Dependent

RNA-seq coverage reflects gene expression, not genomic representation. Genes with low expression may have insufficient coverage for reliable variant calling. Genes with very high expression may have coverage that exceeds the dynamic range of the sequencing platform, leading to duplicate reads and biased allele frequencies.

This limitation means that RNA-seq variant calling cannot provide complete genome-wide variant information. Variants in non-expressed genes will be missed entirely. The [poison frog study](https://doi.org/10.1186/s12864-026-12957-8) demonstrated that RNA-seq can provide thousands of exonic SNPs for population genetics, but the authors acknowledged that this approach samples only the transcribed portion of the genome.

### Allele-Specific Expression Can Bias Variant Calls

Allele-specific expression occurs when one allele of a gene is expressed at a higher level than the other. This can happen due to genomic imprinting, regulatory variation, or nonsense-mediated decay. Allele-specific expression will bias the allele frequency of variants in the affected gene, making it difficult to distinguish true heterozygotes from homozygotes or from artifacts.

### RNA Editing Creates Apparent Variants

RNA editing is a post-transcriptional process that changes the sequence of RNA molecules. The most common form in humans is A-to-I editing, which appears as A-to-G variants in RNA-seq data. These variants are not present in the DNA and will not be detected by DNA sequencing.

RNA editing can create false variant calls if the researcher is looking for genomic variants. Comparing RNA-seq variant calls with DNA sequencing data from the same sample can identify RNA editing events.

### Mapping Bias in Paralogous Genes

Paralogous genes share sequence similarity because they arose from gene duplication events. Reads from one paralog may align to another paralog, creating false variant calls. This problem is especially severe for recently duplicated genes with high sequence similarity.

The [sarcoma study](https://doi.org/10.1007/s00262-026-04395-y) addressed this issue by defining a quality-controlled callable territory and normalizing tumor mutational burden metrics. This approach acknowledges that some regions of the genome cannot be reliably assessed with short-read sequencing.

### Variant Calling Is Limited to Expressed Regions

RNA-seq variant calling only provides information about regions of the genome that are transcribed in the sampled tissue. Variants in regulatory regions, introns, and genes that are not expressed in the sampled tissue will not be detected. This limitation must be considered when interpreting the absence of variants in specific genomic regions.

## Quality Controls and Reproducibility

### Containerized Workflows

Containerized workflows package software and dependencies into a single image that can be run on any system. This approach ensures that the analysis environment is identical across runs and across researchers.

The [nf-core documentation](https://nf-co.re/docs) describes community standards for building and running reproducible bioinformatics pipelines. These standards include version control, containerization, and automated testing.

### Version Control for Analysis Code

Version control systems such as Git track changes to analysis code and configuration files. This allows researchers to document exactly what analysis was performed and to reproduce previous results.

The [The Carpentries lessons](https://carpentries.org/lessons) provide training in version control with Git, including how to manage repositories, track changes, and collaborate with other researchers.

### Documentation of Parameters

Every filtering threshold and analysis parameter should be documented. This documentation should include the rationale for each choice and the expected impact on results. The [EMBL-EBI Training](https://www.ebi.ac.uk/training) program emphasizes the importance of documentation for reproducible research.

### Benchmarking Against Known Variants

If the sample has known variants from previous DNA sequencing, compare the RNA-seq variant calls with the DNA-derived variants. This comparison provides a direct measure of sensitivity and specificity for the RNA-seq variant calling pipeline.

The [Parkinson's disease study](https://pubmed.ncbi.nlm.nih.gov/32348865) used whole-exome sequencing data and the GATK best practices pipeline to identify variants. The study then prioritized variants based on predicted functional effects, demonstrating how variant calling connects to biological interpretation.

### Reproducibility in Practice

Reproducibility requires more than documenting parameters. The analysis environment must be preserved, including software versions, reference genome files, and annotation databases. Containerized workflows address this requirement by packaging the entire analysis environment.

The [Galaxy Training Network](https://training.galaxyproject.org/) provides accessible tutorials that emphasize reproducibility in bioinformatics workflows. These tutorials demonstrate how to document analysis steps and share workflows with other researchers.

## Professional Escalation Criteria

### When to Seek Expert Assistance

Consult a bioinformatics specialist or core facility when:

- The variant calling results are inconsistent with biological expectations
- The number of variants is dramatically higher or lower than expected
- The transition to transversion ratio deviates significantly from expected values
- Variant calls cannot be validated by orthogonal methods
- The analysis requires specialized tools or pipelines beyond standard approaches

### When to Reconsider the Analysis Approach

Reconsider the analysis approach when:

- The alignment rate is below 70 percent
- More than 20 percent of reads are duplicates
- The variant calls are dominated by a single type of artifact
- The filtering thresholds require extreme values to produce reasonable results

### When to Regenerate Data

Consider regenerating data when:

- The sequencing library has clear quality problems
- The coverage is too low for the intended analysis
- The sample was contaminated or mislabeled
- The sequencing platform produced systematic errors

## Safety and Regulatory Context

### Data Privacy and Consent

RNA-seq data may contain identifiable genetic information. Researchers must ensure that data collection and analysis comply with applicable privacy regulations and that sample donors have provided appropriate consent. The [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/) provide guidance on data submission and access policies.

### Data Sharing Requirements

Many funding agencies and journals require data sharing. RNA-seq variant calls should be deposited in appropriate databases with sufficient metadata to allow reuse. The [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/) provide repositories for sequence data and variant calls.

### Clinical Applications

RNA-seq variant calling for clinical applications requires additional validation and quality control. Variants that may guide clinical decisions should be confirmed by an orthogonal method, such as Sanger sequencing or DNA-based genotyping. The [editorial on MiniResource papers](https://doi.org/10.1016/j.mocell.2026.100344) emphasizes the importance of quantitative validation and analytical rigor in molecular and cellular biology research.

### Research Integrity

Variant calling results should be reported with sufficient detail to allow independent verification. This includes documenting the reference genome version, alignment software, variant caller, and filtering thresholds. The [EMBL-EBI Training](https://www.ebi.ac.uk/training) program provides guidance on reporting standards for bioinformatics analyses.

## A Decision Framework for Choosing RNA-Seq Variant Filtering Strategies

Selecting the right filtering strategy for RNA-seq variant calls requires a structured approach that accounts for the biological question, data characteristics, and available validation resources. A practical decision framework helps researchers avoid the common error of applying DNA-seq filtering defaults to transcriptome data without adjustment.

### Step 1: Define the Primary Analysis Goal

The filtering strategy begins with the analysis goal, not the data. Three distinct goals require different filtering priorities:

**Goal A: Germline variant discovery in expressed genes.** The priority is sensitivity with controlled false positives. Filters should emphasize mapping quality, base quality, and allele balance. Depth thresholds can be modest because the goal is to identify variants that exist in the genome and are transcribed.

**Goal B: Somatic variant detection in disease samples.** The priority is specificity because somatic variants are often present at low allele fractions. Filters must address sequencing artifacts aggressively, and orthogonal validation becomes essential. The [sarcoma multi-omics study](https://doi.org/10.1007/s00262-026-04395-y) demonstrated that RNA-seq can serve as an orthogonal validation check for variants identified from whole-exome sequencing, distinguishing functional amplifications from technical depth artifacts.

**Goal C: Population genetics and genotyping.** The priority is consistency across samples. Filters should be applied uniformly, and the analysis should focus on a callable territory that is reliably covered in all samples. The [poison frog study](https://doi.org/10.1186/s12864-026-12957-8) genotyped thousands of exonic SNPs using RNA-seq data to elucidate population structure, validating the applicability of mRNA sequences for population-level analysis.

### Step 2: Assess Data Characteristics That Constrain Filtering

Three data characteristics determine which filters are feasible and how thresholds should be set:

**Coverage distribution.** Examine the distribution of read depth across expressed genes. If the median depth is below 20 reads, aggressive depth filters will remove most variants. If the distribution is highly skewed, with some genes at hundreds of reads and others at fewer than five, expression-aware thresholds are required.

**Library strand protocol.** Stranded libraries preserve the orientation of the original RNA molecule, while unstranded libraries lose this information. Strand bias filters must be interpreted differently for each protocol. For unstranded libraries, strand bias annotations are less informative and should not be used as a primary filter.

**Read length and insert size.** Longer reads provide more reliable junction-spanning alignments. Shorter reads, typically below 75 bases, produce more multi-mapping reads and require stricter mapping quality filters.

### Step 3: Select the Filtering Approach Based on Validation Resources

The availability of validation data determines whether machine learning filtering is appropriate:

**No validation data available.** Use hard filtering with conservative thresholds. Document every threshold and its rationale. The [GATK review](https://doi.org/10.3390/ijms27093754) describes how the toolkit has adapted its recommendations over time, with hard filtering often being more practical for RNA-seq data because variant annotations behave differently than in DNA data.

**DNA sequencing data from the same sample available.** Use the DNA-derived variants as a truth set. Compare RNA-seq variant calls against this set to measure sensitivity and specificity. This comparison can guide threshold selection and identify systematic biases in the RNA-seq pipeline.

**Genotyping array data available.** Use array genotypes for known variant sites to calibrate filters. This approach provides a large set of known heterozygous and homozygous sites for evaluating allele balance and depth filters.

**Replicate RNA-seq samples available.** Use concordance between replicates as a filter. Variants called in only one replicate warrant scrutiny, while variants called consistently across replicates are more likely to be real.

### Step 4: Apply Filters in a Defined Order

Filter order matters because each filter changes the variant set that subsequent filters evaluate. A recommended order is:

1. **Mapping quality filter.** Remove variants supported primarily by reads with low mapping quality. This is the first filter because misaligned reads create artifacts that distort all other annotations.

2. **Depth filter.** Apply an expression-aware depth threshold. For germline calling, require a minimum depth that reflects the lower tail of the coverage distribution for expressed genes.

3. **Allele balance filter.** Remove variants where the alternate allele fraction deviates significantly from expected values. For heterozygous germline variants, the expected fraction is approximately 0.5, though allele-specific expression can shift this.

4. **Strand bias filter.** Apply with caution, considering the library strand protocol and the possibility of strand-specific expression.

5. **Quality by depth filter.** Remove variants where the quality score is high but the supporting depth is low, which often indicates a single high-quality read driving the call.

6. **Context-specific filters.** Apply filters for known artifact regions, such as repetitive elements, paralogous genes, and RNA editing sites.

### Step 5: Evaluate Filter Performance with Metrics

After applying filters, evaluate whether the filtering strategy performed appropriately:

**Transition to transversion ratio.** For human data, a ratio around 2.0 is typical for high-quality variant calls. A ratio below 1.5 suggests systematic errors, while a ratio above 3.0 may indicate overfiltering of true variants.

**Variant density by gene expression level.** Plot the number of variants per gene against expression level. True variants should be distributed across expression levels, while artifacts often concentrate in highly expressed genes due to PCR duplicates and mapping errors.

**Heterozygous to homozygous ratio.** For germline data, an excess of homozygous variants may indicate allele-specific expression or reference bias. An excess of heterozygous variants may indicate mapping artifacts.

**Concordance with known variants.** If a truth set is available, calculate sensitivity and specificity. The [Parkinson's disease study](https://pubmed.ncbi.nlm.nih.gov/32348865) used the GATK best practices pipeline for variant discovery and then applied filter-based annotation to categorize variants, demonstrating how filtering connects to biological interpretation.

### Step 6: Document Filtering Decisions for Reproducibility

Record the following for each filtering decision:

- The filter name and the tool used to apply it
- The threshold value and the units
- The rationale for choosing that threshold
- The number of variants removed by the filter
- The number of variants remaining after the filter

The [EMBL-EBI Training](https://www.ebi.ac.uk/training) program emphasizes the importance of documentation for reproducible research. The [nf-core documentation](https://nf-co.re/docs) describes community standards for pipeline configuration and quality control that apply to filtering steps.

### Common Decision Scenarios

**Scenario 1: Low coverage RNA-seq data.** When median depth is below 15 reads, avoid aggressive depth filters. Instead, rely on mapping quality, base quality, and allele balance filters. Consider whether the analysis goal can be met with the available depth or whether additional sequencing is needed.

**Scenario 2: High duplication rate.** When the duplicate rate exceeds 20 percent, evaluate whether duplicate marking is appropriate. For RNA-seq, reads from highly expressed genes may share start positions by chance. Compare variant calls with and without duplicate marking to understand the impact.

**Scenario 3: Non-model organism without a high-quality reference.** Use a reference genome that matches the sample as closely as possible. Apply stricter mapping quality filters because the reference may contain errors or gaps. The [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/) provide access to reference genomes and annotation databases for many organisms.

**Scenario 4: Clinical or diagnostic application.** Require orthogonal validation for any variant that may guide clinical decisions. The [editorial on MiniResource papers](https://doi.org/10.1016/j.mocell.2026.100344) emphasizes the importance of quantitative validation and analytical rigor in molecular and cellular biology research.

### When to Escalate to Expert Assistance

Seek bioinformatics consultation when:

- The filtering strategy requires thresholds that conflict with the analysis goals
- The variant calls cannot be validated by any available orthogonal method
- The transition to transversion ratio remains abnormal after filtering
- The analysis requires specialized tools beyond standard pipelines
- The results will inform clinical or regulatory decisions

The [Galaxy Training Network](https://training.galaxyproject.org/) provides accessible tutorials on RNA-seq analysis workflows that can help researchers understand filtering options before seeking expert assistance. The [Bioconductor project](https://bioconductor.org/) offers R packages for variant annotation and filtering that can be integrated into reproducible workflows.

## Frequently Asked Questions

### Why does RNA-seq variant calling require a different pipeline than DNA-seq variant calling?

RNA-seq reads span splice junctions, so they cannot be aligned as continuous sequences to the reference genome. A splice-aware aligner must split reads across exon-exon junctions. The variant caller must then handle these split reads correctly. Additionally, RNA-seq coverage reflects gene expression levels, so depth-based filters must account for expression variation across genes.

### What is the minimum read depth for calling a variant from RNA-seq data?

There is no universal minimum depth. The appropriate threshold depends on the analysis goals, the expression level of the gene, and the quality of the sequencing data. A common minimum depth for germline variant calling is 10 reads, but variants in lowly expressed genes may require a lower threshold. The key is to document the threshold and understand its impact on sensitivity and specificity.

### How do I handle reads that span splice junctions during variant calling?

Use a splice-aware aligner such as STAR or HISAT2 for alignment. After alignment, split reads at intron boundaries using a tool such as GATK SplitNCigarReads. This step creates separate alignment records for each exon segment, preventing the variant caller from being confused by large gaps in alignment.

### Can I use RNA-seq data for somatic variant calling?

Yes, but somatic variant calling requires paired normal and tumor samples to distinguish somatic mutations from germline variants. The allele frequency threshold for somatic variants is often lower because the variant may be present in only a fraction of cells. RNA-seq can also serve as an orthogonal validation check for variants identified from DNA sequencing.

### What causes false variant calls in RNA-seq data?

Common causes include misalignment of reads in repetitive regions, reads spanning splice junctions that are aligned incorrectly, RNA editing events that create apparent variants, and allele-specific expression that biases allele frequencies. Base quality errors and PCR duplicates can also create false variant calls.

### How do I validate RNA-seq variant calls?

Validation approaches include Sanger sequencing of selected variants, comparison with DNA sequencing data from the same sample, genotyping arrays for known variant sites, and RNA-seq data from replicate samples. The choice of validation approach depends on the analysis goals and available resources.

### Should I use hard filtering or machine learning filtering for RNA-seq variants?

Hard filtering is often the safer choice for RNA-seq data because the variant annotation distributions differ from DNA data. Machine learning filtering, such as variant quality score recalibration, requires a reliable training set and can produce unpredictable results when applied to data that differs from the training set.

### How do I account for allele-specific expression in variant calling?

Allele-specific expression biases the allele frequency of variants in affected genes. Comparing RNA-seq variant calls with DNA sequencing data from the same sample can identify genes with allele-specific expression. Alternatively, tools that model allele-specific expression can be used to adjust variant calling.

## Related Bioinformatics Guides

- [RNA-Seq Data Analysis Workflow: From Raw Reads to Insights](/knowledge/bioinformatics/rna-seq-data-analysis-workflow-from-raw-reads-to-insights)
- [RNA-Seq Databases: Accessing and Using Public RNA-Seq Data](/knowledge/bioinformatics/rna-seq-databases-accessing-and-using-public-rna-seq-data)
- [RNA-Seq Data Analysis in Galaxy: A User-Friendly Platform](/knowledge/bioinformatics/rna-seq-data-analysis-in-galaxy-a-user-friendly-platform)
- [RNA Sequencing Data Analysis: From Raw Reads to Differential Expression](/knowledge/bioinformatics/rna-sequencing-data-analysis-from-raw-reads-to-differential-expression)
- [Data Annotation for AI in Life Sciences: Roles, Challenges, and Best Practices](/knowledge/bioinformatics/data-annotation-for-ai-in-life-sciences-roles-challenges-and-best-practices)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)


## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [Next generation sequencing exome data analysis aids in the discovery of SNP and INDEL patterns in Parkinson's disease.](https://pubmed.ncbi.nlm.nih.gov/32348865). Genomics, 2020.
- [Fifteen Years of the Genome Analysis Toolkit as the De Facto Standard in Short-Read Variant Calling.](https://doi.org/10.3390/ijms27093754). 2026.
- [Editorial: MiniResource papers from Molecules and Cells.](https://doi.org/10.1016/j.mocell.2026.100344). 2026.
- [Repurposing public sarcoma multi-omics for neoantigen discovery.](https://doi.org/10.1007/s00262-026-04395-y). 2026.
- [Transcriptomic genotyping elucidates the population structure and demographic history of the endangered poison frog Oophaga vicentei (Anura: Dendrobatidae).](https://doi.org/10.1186/s12864-026-12957-8). 2026.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.