Troubleshooting Strand Bias in Variant Calling: How to Identify and Filter Strand-Specific Artifacts
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Strand bias in variant calling signifies an uneven distribution of alternative allele reads across forward and reverse DNA strands, often indicating a systematic sequencing artifact rather than a true biological variant.
- Common sources of strand-specific artifacts include chemical modifications in Formalin-Fixed Paraffin-Embedded (FFPE) tissues (e.g., C to T deamination) and oxidative damage during DNA fragmentation (e.g., C to A substitutions from 8-oxoguanine lesions).
- Key metrics for assessing strand bias include Fisher's exact test p-value (FS), with GATK recommending filtering SNPs with FS > 60, and strand bias odds ratio (SB), where values near 1.0 indicate balance.
- Practical filtering guidance involves applying coverage-dependent thresholds, such as the 1:3 forward/reverse read ratio for mitochondrial DNA analysis, and considering substitution type context, as FFPE artifacts exhibit characteristic C>T/G>A patterns while oxidative damage favors C>A/G>T.
- Visual inspection of read pileups and utilizing tools like SOBDetector for FFPE samples are crucial for distinguishing true variants from artifacts, especially in borderline cases or low-coverage regions where statistical metrics are less reliable.
- Batch effects can significantly influence strand bias; monitoring average strand bias across sequencing runs and comparing distributions between batches is essential for identifying and mitigating systematic issues.
Strand bias in variant calling occurs when the alternative allele is observed predominantly on one DNA strand instead of being distributed evenly across both forward and reverse strands. This asymmetry often signals a systematic sequencing artifact instead of a true biological variant. When your variant call format (VCF) file flags a variant for strand bias, you must decide whether to filter it out or adjust parameters. Filtering too aggressively removes true variants, particularly in regions with low coverage or extreme GC content. Filtering too leniently retains artifacts that waste validation resources and contaminate downstream analyses. This article explains the statistical basis of strand bias metrics, provides practical thresholds for filtering, and describes visualization techniques that help you distinguish true variants from artifacts in germline and somatic variant calling workflows.
Understanding Strand Bias in Sequencing Data
Strand bias arises from the chemistry and physics of library preparation and sequencing. During Illumina sequencing, DNA fragments are amplified and sequenced from both ends. A true heterozygous variant should appear on both forward and reverse strands at roughly equal proportions, because the original DNA molecule contains both alleles and both strands are sequenced independently. When the alternative allele appears almost exclusively on one strand, the observation is suspicious.
Sources of Strand-Specific Artifacts
Several mechanisms produce strand-specific errors. Formalin-fixed paraffin-embedded (FFPE) tissue samples are a well-documented source of artifacts in clinical sequencing. The formalin fixation process introduces crosslinks and deamination events that create characteristic C to T and G to A substitutions. These artifacts are not randomly distributed across strands. The Strand Orientation Bias Detector (SOBDetector) tool was developed specifically to assess whether a detected mutation in FFPE samples is an artifact, using a Bayesian logistic regression model trained on The Cancer Genome Atlas whole exomes. This tool is compatible with common somatic single nucleotide variant calling pipelines and is implemented in Java 1.8 [<a href="#ref-1">1</a>].
Oxidative damage during DNA fragmentation is another major source of strand-specific artifacts. Ultrasonic fragmentation, a standard step in library preparation, can introduce 8-oxoguanine lesions that lead to C to A and G to T substitutions. Research on mitochondrial DNA sequencing demonstrated that these low-frequency substitutions occur frequently and show strong batch effects and correlations with overall mutation burden and local GC content. The same study found that enzyme-based fragmentation markedly reduced oxidative damage and improved mutation detection fidelity compared to ultrasonic fragmentation. Temperature, intensity, and cycle number during sonication influenced the extent of damage [<a href="#ref-2">2</a>].
Why Strand Bias Matters for Variant Confidence
Systematic sequencing errors that pass initial quality filters can be mistaken for true polymorphisms. A study of SNP discovery in the deep-sea fish orange roughy found that filtering out systematic sequencing errors substantially improved the efficiency of SNP discovery. The researchers examined five predictors for their effect on obtaining an assayable SNP: depth of coverage, number of reads supporting a variant, polymorphism type, strand bias, and Illumina SNP probe design score. Their results showed that accounting for strand bias reduced the number of false polymorphisms that would otherwise waste genotyping resources [<a href="#ref-3">3</a>].
For mitochondrial DNA analysis, strand bias assessment is a standard quality metric. A forensic validation study using the ForenSeq mtDNA Whole Genome Kit reported forward and reverse strand bias values near 1.0 across nine sequencing runs, indicating balanced strand representation. The same study identified five amplicons with low read depth and high strand bias as the most vulnerable positions in the mitochondrial genome. These positions required additional scrutiny during variant interpretation [<a href="#ref-4">4</a>].
At a Glance: Strand Bias Metrics and Filtering Decisions
The table below summarizes the common strand bias metrics you will encounter in variant calling outputs, what they measure, and practical guidance for interpretation.
| Metric | What It Measures | Typical Filtering Guidance | Caveats |
|---|---|---|---|
| Fisher's exact test p-value (FS) | Probability that the observed strand distribution of alleles occurs by chance | Filter variants with FS greater than 60 for SNPs in GATK recommendations | Low coverage positions can produce extreme p-values even for true variants |
| Strand bias odds ratio (SB) | Ratio of alternative allele support on forward versus reverse strand | Values near 1.0 indicate balanced representation, extreme values warrant review | Some true variants in repetitive or GC-rich regions show inherent bias |
| Strand orientation bias score (SOB) | Probability that a variant is an FFPE artifact based on substitution pattern | Higher scores indicate higher artifact probability in FFPE samples | Only applicable to FFPE-derived data and specific substitution types |
| Forward/reverse read ratio | Raw proportion of alternative allele reads on each strand | Ratios exceeding 1:3 suggest strand-specific error | Coverage-dependent thresholds are recommended for heteroplasmic variants |
Core Principles of Strand Bias Assessment
The Statistical Foundation
Strand bias metrics compare the number of reference and alternative allele reads on the forward strand against the number on the reverse strand. The Fisher's exact test evaluates whether the allele distribution differs significantly between strands. A low p-value indicates that the observed asymmetry is unlikely to occur by chance, which suggests a systematic artifact. However, the test is sensitive to read depth. At low coverage, even a true variant can show apparent strand bias because only a few reads support the alternative allele and they happen to fall on one strand.
The strand bias odds ratio provides a complementary measure. An odds ratio near 1.0 means the alternative allele is equally likely to appear on either strand. Values far from 1.0 indicate asymmetry. The forensic mitochondrial genome study reported average strand bias values near 1.0 for a well-performing sequencing run, providing a benchmark for expected performance [<a href="#ref-4">4</a>].
Coverage-Dependent Interpretation
Coverage interacts with strand bias in ways that affect filtering decisions. A variant with 100 total reads and 50 alternative allele reads distributed as 25 forward and 25 reverse is clearly balanced. A variant with 10 total reads and 5 alternative allele reads distributed as 5 forward and 0 reverse shows complete strand bias, but the small read count makes the observation unreliable. The single-cell RNA sequencing pipeline for mitochondrial DNA variants addressed this by using coverage-dependent thresholds for detecting heteroplasmic variants. The pipeline removed strand bias errors exceeding a 1:3 ratio, meaning the alternative allele had to appear on both strands with no more than three times more support on one strand than the other [<a href="#ref-5">5</a>].
Substitution Type Context
Different artifact mechanisms produce characteristic substitution patterns. Oxidative damage primarily generates C to A and G to T substitutions. FFPE deamination generates C to T and G to A substitutions. When you observe a variant with a substitution type that matches a known artifact mechanism and the variant also shows strand bias, the evidence for an artifact strengthens. The mtDeOxoGer model incorporated substitution type, variant allele frequency, strand orientation bias score, local GC content, and sequence context to distinguish authentic mitochondrial mutations from oxidative artifacts. This multi-feature approach outperformed strand bias alone [<a href="#ref-2">2</a>].
Practical Workflow for Identifying Strand Bias
Step 1: Examine the Variant Call Format Fields
Your VCF file contains strand bias information in the INFO column. The specific fields depend on the variant caller you used. GATK HaplotypeCaller and Mutect2 output the FisherStrand (FS) and StrandOddsRatio (SOR) annotations. SAMtools mpileup outputs the DP4 field, which contains counts of reference and alternative alleles on forward and reverse strands. FreeBayes outputs the SAF, SAR, SRF, and SRR fields for strand-specific allele counts.
Open your VCF file and identify which strand bias fields are present. If the fields are missing, you need to run the variant caller with the appropriate annotation options or use a separate tool to calculate strand bias from the alignment files. The Galaxy Training Network provides accessible workflow training for variant calling that includes guidance on interpreting VCF fields and applying quality filters [<a href="#ref-6">6</a>].
Step 2: Calculate Strand-Specific Allele Counts
For variants where you need a direct assessment, calculate the strand distribution manually from the alignment data. Use a genome browser or a command-line tool to visualize the reads supporting the variant. Count the number of forward-strand reads and reverse-strand reads that carry the alternative allele. Compare these counts to the total read depth at the position.
A practical approach is to use the Integrated Genomics Viewer or similar visualization software. Load the alignment file and the variant call file, navigate to the variant position, and inspect the read pileup. The reads are color-coded by strand orientation. Look for the alternative allele and note whether it appears on both strands.
Step 3: Apply Thresholds Based on Your Data Type
For germline variant calling in whole-genome or whole-exome data, the GATK recommended threshold for FisherStrand is 60 for SNPs. Variants with FS values above this threshold are more likely to be artifacts. For somatic variant calling, particularly in FFPE samples, use the SOBDetector tool to calculate the probability that a detected mutation is an artifact. The tool outputs a probability score that you can use to prioritize variants for validation [<a href="#ref-1">1</a>].
For mitochondrial DNA analysis, the 1:3 strand bias ratio provides a practical cutoff. Variants where the alternative allele appears on only one strand or with a ratio exceeding 1:3 should be removed. The single-cell RNA sequencing pipeline for mitochondrial variants applied this threshold and successfully filtered sequencing errors while retaining true homoplasmic and heteroplasmic variants [<a href="#ref-5">5</a>].
Step 4: Visualize and Confirm
Visual inspection is essential for borderline cases. Generate a read pileup image for the variant position and examine the strand distribution. Look for additional evidence of artifacts, such as the presence of multiple alternative alleles at the same position, which can indicate RNA modification-induced errors or sequencing errors. The single-cell mitochondrial pipeline specifically removed variants with multiple alternative alleles at the same position as a separate filtering step [<a href="#ref-5">5</a>].
For FFPE samples, check whether the variant is a C to T or G to A transition, which are the characteristic FFPE artifacts. If the variant shows strand bias and matches the FFPE substitution pattern, treat it as a likely artifact. The SOBDetector tool provides a systematic way to make this assessment based on the posterior distribution of a Bayesian logistic regression model [<a href="#ref-1">1</a>].
Options and Tradeoffs in Strand Bias Filtering
Hard Filtering Versus Soft Flagging
You have two main approaches to handling strand bias. Hard filtering removes variants that exceed a threshold from the final variant set. Soft flagging retains the variants but marks them with a quality score or annotation that indicates the strand bias concern. Hard filtering is simpler and produces a cleaner dataset for downstream analysis. Soft flagging preserves information and allows you to revisit borderline variants if additional data becomes available.
The choice depends on your downstream application. If you are building a SNP panel for genotyping, hard filtering is appropriate because you need high-confidence variants. The orange roughy SNP discovery study demonstrated that filtering out systematic sequencing errors improved the efficiency of SNP discovery and reduced the number of polymorphisms that failed to convert to assayable assays [<a href="#ref-3">3</a>]. If you are doing exploratory analysis or variant prioritization, soft flagging allows you to retain potentially interesting variants while acknowledging the uncertainty.
Adjusting Parameters Versus Changing the Pipeline
When you observe excessive strand bias across many variants, the problem may be systematic instead of variant-specific. Adjusting filtering thresholds will not fix the underlying issue. Instead, examine your library preparation and sequencing protocols. The oxidative damage study found that ultrasonic fragmentation parameters, including temperature, intensity, and cycle number, influenced the extent of artifacts. Switching to enzyme-based fragmentation reduced oxidative damage and improved mutation detection fidelity [<a href="#ref-2">2</a>].
For FFPE samples, the artifact burden is inherent to the sample type. SOBDetector was developed specifically for this context because FFPE artifacts are pervasive and cannot be eliminated by protocol changes alone. The tool provides a probability score that you can use to filter variants, but you should also consider whether the sample type is appropriate for your research question [<a href="#ref-1">1</a>].
Coverage Thresholds and Strand Bias Interaction
Low coverage regions present a particular challenge for strand bias assessment. At low depth, the strand distribution of reads is subject to high variance. A true variant with 8 supporting reads might show 7 reads on the forward strand and 1 on the reverse strand, producing a strand bias flag even though the variant is real. The coverage-dependent thresholds used in the single-cell mitochondrial pipeline address this by requiring different levels of strand balance depending on the total read depth [<a href="#ref-5">5</a>].
For your own data, consider stratifying strand bias filtering by coverage. Apply stricter strand bias thresholds in high coverage regions and more lenient thresholds in low coverage regions. Alternatively, require a minimum read depth before applying strand bias filtering at all.
Observations and Measurements for Strand Bias Monitoring
Tracking Strand Bias Across Sequencing Runs
Strand bias is a useful quality metric for monitoring sequencing performance over time. The forensic mitochondrial genome study reported average strand bias values near 1.0 across nine consecutive sequencing runs, demonstrating consistent performance. If you observe a run with average strand bias values significantly different from 1.0, investigate the library preparation and sequencing conditions for that run [<a href="#ref-4">4</a>].
Record the following measurements for each sequencing run:
- Average strand bias across all variants
- Number of variants with extreme strand bias values
- Distribution of strand bias by chromosome or amplicon
- Correlation between strand bias and GC content
- Correlation between strand bias and read depth
These measurements help you identify systematic issues. For example, if strand bias is elevated in specific amplicons across multiple runs, the problem is likely in the primer design or amplicon structure instead of a random sequencing error.
Identifying Vulnerable Regions
Certain genomic regions are more prone to strand bias than others. The forensic mitochondrial study identified five amplicons with low read depth and high strand bias as the most vulnerable positions in the mitochondrial genome. These regions required additional scrutiny during variant interpretation [<a href="#ref-4">4</a>]. For your own data, identify regions with consistently high strand bias and document them for future reference.
Common vulnerable regions include:
- Regions with extreme GC content
- Repetitive sequences
- Regions with secondary structure
- Amplicon boundaries
- Regions with low mappability
When you encounter a variant in a known vulnerable region, apply additional scrutiny regardless of the strand bias metric value.
Recording Filtering Decisions
Maintain a record of your strand bias filtering decisions. For each variant that you filter based on strand bias, record:
- The variant position and substitution type
- The strand bias metric value
- The read depth and allele frequency
- The reason for the filtering decision
- Whether the variant was subsequently validated
This record helps you refine your thresholds over time. If you find that many filtered variants are later validated as true, your thresholds are too strict. If you find that many retained variants fail validation, your thresholds are too lenient.
Common Failure Patterns in Strand Bias Filtering
Overfiltering True Variants
The most common failure pattern is filtering true variants because they show strand bias in low coverage regions. This is particularly problematic for somatic variant calling, where true variants may be present at low allele frequencies. A true somatic variant with 5% allele frequency will have few supporting reads, and those reads may by chance fall predominantly on one strand.
To avoid this failure, examine the read depth before applying strand bias filters. If the total read depth is below 20, interpret strand bias metrics with caution. Consider whether the variant has other supporting evidence, such as a known biological role or presence in multiple samples.
Underfiltering Artifacts
The opposite failure pattern is retaining artifacts because the strand bias metric does not exceed the threshold. This occurs when the artifact mechanism produces balanced strand representation. Some FFPE artifacts and oxidative damage artifacts can appear on both strands, particularly when the damage occurs before library amplification.
To avoid this failure, combine strand bias with other artifact indicators. The mtDeOxoGer model used substitution type, variant allele frequency, strand orientation bias score, local GC content, and sequence context to identify oxidative artifacts [<a href="#ref-2">2</a>]. A variant that matches the artifact substitution pattern and occurs in a high GC region should be treated with suspicion even if strand bias is not extreme.
Ignoring Batch Effects
Strand bias can vary systematically between sequencing batches. The oxidative damage study found strong batch effects in the occurrence of low-frequency C to A and G to T substitutions [<a href="#ref-2">2</a>]. If you process samples in multiple batches, compare strand bias distributions between batches. A batch with elevated strand bias may have experienced different library preparation conditions.
Record the batch information for each sample and include it in your quality control analysis. If one batch shows consistently higher strand bias, investigate the library preparation conditions for that batch and consider whether the data should be re-sequenced or processed with different filtering thresholds.
Applying Germline Thresholds to Somatic Data
Germline and somatic variant calling have different expectations for strand bias. Germline variants are typically present at 50% or 100% allele frequency, so strand bias is easier to assess with confidence. Somatic variants may be present at low allele frequencies, making strand bias assessment less reliable. The GATK recommended thresholds for germline data may not be appropriate for somatic data.
For somatic variant calling, use the SOBDetector tool for FFPE samples and consider lower confidence in strand bias metrics for low allele frequency variants [<a href="#ref-1">1</a>]. Validate somatic variants with an orthogonal method before accepting them as true.
Limitations of Strand Bias Metrics
Statistical Power at Low Coverage
Strand bias metrics have limited statistical power at low read depths. The Fisher's exact test requires sufficient counts in each category to detect a significant difference. At low coverage, the test may fail to detect true strand bias or may produce false positives due to small sample sizes. This limitation is inherent to the statistical approach and cannot be fully overcome by adjusting thresholds.
Artifact Mechanisms Without Strand Bias
Not all sequencing artifacts produce strand bias. Some artifacts, such as those introduced during PCR amplification, can affect both strands equally. The absence of strand bias does not guarantee that a variant is true. You must use additional quality metrics and validation approaches to confirm variant calls.
Sample Type Specificity
Strand bias metrics are calibrated for specific sample types and sequencing platforms. The SOBDetector model was trained on The Cancer Genome Atlas whole exomes and is designed for FFPE samples [<a href="#ref-1">1</a>]. Applying this model to fresh frozen tissue or to RNA sequencing data may produce unreliable results. Similarly, the 1:3 strand bias ratio used in the single-cell mitochondrial pipeline was developed for that specific data type and may not transfer directly to other applications [<a href="#ref-5">5</a>].
Reference Genome Dependence
Strand bias assessment depends on the alignment of reads to the reference genome. Regions with low mappability or structural variation can produce spurious strand bias because reads align incorrectly. If your reference genome has errors or your alignment parameters are suboptimal, strand bias metrics will be unreliable.
Quality Controls for Strand Bias Assessment
Alignment Quality Checks
Before assessing strand bias, verify that your alignments are of high quality. Check the mapping quality distribution, the percentage of reads that map uniquely, and the insert size distribution. Poor alignments produce spurious variant calls and unreliable strand bias metrics. The EMBL-EBI training resources provide structured learning pathways for alignment quality assessment and variant calling best practices [<a href="#ref-7">7</a>].
Duplicate Read Handling
PCR duplicates can inflate strand bias because duplicate reads originate from the same original molecule and therefore share the same strand orientation. Mark and remove duplicates before variant calling to avoid this artifact. The single-cell mitochondrial pipeline noted that duplicate reads could be retained because the majority were valid biological duplicates in that context, but this is specific to that application [<a href="#ref-5">5</a>]. For standard DNA sequencing, duplicate removal is recommended.
Base Quality Score Recalibration
Base quality scores affect variant calling and strand bias assessment. Recalibrate base quality scores using known variant sites before variant calling. This step corrects for systematic errors in base quality estimation and improves the accuracy of variant calls and associated metrics.
Strand Bias as a Run-Level Quality Metric
Monitor strand bias at the run level as part of your sequencing quality control. The forensic mitochondrial study reported average strand bias near 1.0 as a performance metric for the sequencing system [<a href="#ref-4">4</a>]. If your run shows average strand bias significantly different from 1.0, investigate the cause before proceeding with variant interpretation.
Safety and Regulatory Context
Clinical and Diagnostic Applications
If your variant calling results are used for clinical or diagnostic purposes, strand bias filtering decisions have direct patient care implications. Filtering a true variant can lead to a missed diagnosis. Retaining an artifact can lead to an incorrect diagnosis. In this context, follow established guidelines and consult with clinical genomics professionals before making filtering decisions.
The SOBDetector tool was developed for clinical FFPE samples, which are the most common tissue specimens stored in clinical practice [<a href="#ref-1">1</a>]. Using this tool for FFPE samples provides a systematic approach to artifact assessment, but the output should be interpreted in the context of the specific clinical question.
Forensic Applications
Forensic mitochondrial DNA analysis requires rigorous quality control because the results may be used as legal evidence. The forensic validation study reported specific performance metrics, including strand bias values near 1.0, as part of the internal validation for routine application [<a href="#ref-4">4</a>]. If you are conducting forensic analysis, follow the validation requirements for your jurisdiction and document all quality control measures.
Research Data Repositories
If you deposit your variant calls in public databases such as those maintained by NCBI, the quality of your variant calls affects the utility of the data for other researchers [<a href="#ref-8">8</a>]. Variant calls contaminated with strand bias artifacts reduce the reliability of the database. Apply appropriate filtering and document your filtering decisions so that data users can assess the quality of your calls.
Professional Escalation Criteria
When to Seek Expert Assistance
You should escalate strand bias issues to a bioinformatics specialist or a clinical genomics professional in the following situations:
- You observe strand bias in a high proportion of variants across multiple samples
- You are working with FFPE samples and need to implement SOBDetector or similar artifact filtering
- You are developing a variant calling pipeline for clinical or diagnostic use
- You are uncertain whether a specific variant is a true variant or an artifact
- You need to validate variants using an orthogonal method
When to Re-sequence
Consider re-sequencing a sample when:
- The sample shows excessive strand bias across many variants
- The library preparation conditions were suboptimal
- The sample is critical for your research question and the variant calls are unreliable
- The strand bias pattern suggests a systematic issue that cannot be resolved by filtering
When to Change Protocols
Change your library preparation or sequencing protocols when:
- You observe consistent strand bias patterns across multiple runs
- The strand bias correlates with specific fragmentation conditions
- You are using ultrasonic fragmentation and observe oxidative damage artifacts [<a href="#ref-2">2</a>]
- You are working with FFPE samples and need to implement artifact-specific filtering
A Practical Decision Framework for Strand Bias Filtering
When your VCF file flags variants for strand bias, you need a structured approach that prevents both overfiltering and underfiltering. A tiered decision framework based on evidence strength, read depth, and biological context provides a reproducible method for handling flagged variants consistently across samples and projects.
Tier 1: Automatic Filtering for High-Confidence Artifacts
The first tier applies to variants that meet criteria strongly associated with known artifact mechanisms. These variants can be filtered automatically without manual review, provided your data matches the conditions where the evidence is robust.
For FFPE samples, variants that show strand bias and match the characteristic deamination substitution pattern of C to T or G to A transitions warrant automatic filtering. The SOBDetector tool provides a systematic probability score for this assessment, trained on The Cancer Genome Atlas whole exomes using a Bayesian logistic regression model [<a href="#ref-1">1</a>]. When the tool assigns a high artifact probability and the substitution type matches the FFPE signature, the evidence for an artifact is strong enough to justify removal without further review.
For oxidative damage artifacts, the characteristic pattern is C to A and G to T substitutions accompanied by strand bias. Research on mitochondrial DNA sequencing demonstrated that these low-frequency substitutions occur frequently and correlate with overall mutation burden and local GC content [<a href="#ref-2">2</a>]. If your library preparation used ultrasonic fragmentation and you observe this substitution pattern with strand bias, automatic filtering is appropriate. The same study found that enzyme-based fragmentation markedly reduced oxidative damage, so if you have already switched protocols, the prior evidence base for automatic filtering changes.
For mitochondrial DNA analysis, the 1:3 strand bias ratio provides a practical automatic filtering threshold. The single-cell RNA sequencing pipeline for mitochondrial variants removed strand bias errors exceeding this ratio, meaning the alternative allele had to appear on both strands with no more than three times more support on one strand than the other [<a href="#ref-5">5</a>]. This threshold worked effectively for that data type and provides a starting point for your own mitochondrial analyses.
Tier 2: Manual Review for Borderline Cases
The second tier applies to variants that show strand bias but do not meet the automatic filtering criteria. These variants require manual review using visualization and contextual evidence before making a filtering decision.
Start by visualizing the variant position in a genome browser or read pileup viewer. Examine the strand distribution of reads supporting the alternative allele. Look for additional evidence of artifacts, such as the presence of multiple alternative alleles at the same position, which can indicate RNA modification-induced errors or sequencing errors. The single-cell mitochondrial pipeline specifically removed variants with multiple alternative alleles at the same position as a separate filtering step [<a href="#ref-5">5</a>].
Check the read depth at the variant position. Low coverage positions produce unreliable strand bias metrics because the small number of reads creates high variance in strand distribution. A true variant with 8 supporting reads might show 7 reads on the forward strand and 1 on the reverse strand, producing a strand bias flag even though the variant is real. For low coverage positions, consider whether the variant has other supporting evidence, such as presence in multiple samples or a known biological role.
Examine the local sequence context. Regions with extreme GC content, repetitive sequences, or secondary structure are prone to alignment artifacts that produce spurious strand bias. The forensic mitochondrial genome study identified five amplicons with low read depth and high strand bias as the most vulnerable positions in the mitochondrial genome, demonstrating that certain regions consistently require additional scrutiny [<a href="#ref-4">4</a>]. If your variant falls in a known vulnerable region, apply more conservative filtering decisions.
Tier 3: Retain with Flagging for Uncertain Cases
The third tier applies to variants where the evidence is genuinely ambiguous. These variants should be retained in the dataset but flagged for downstream consideration. This approach preserves information while acknowledging the uncertainty.
Flagged variants should be annotated with the strand bias metric values, the read depth, the allele frequency, and the reason for the uncertainty. This annotation allows you to revisit these variants if additional data becomes available or if validation results provide new information.
For somatic variant calling, low allele frequency variants present a particular challenge. A true somatic variant with 5% allele frequency will have few supporting reads, and those reads may by chance fall predominantly on one strand. The strand bias metric is unreliable at this allele frequency, so flagging instead of filtering is appropriate. The mtDeOxoGer model incorporated variant allele frequency alongside strand orientation bias score, substitution type, local GC content, and sequence context to distinguish authentic mutations from artifacts, demonstrating that multiple features together provide better discrimination than strand bias alone [<a href="#ref-2">2</a>].
Implementing the Decision Framework in Practice
To implement this framework, create a decision matrix that you apply consistently to every flagged variant. The matrix should include the following columns:
- Variant position and substitution type
- Sample type (FFPE, fresh frozen, mitochondrial, other)
- Read depth at the variant position
- Allele frequency
- Strand bias metric values
- Substitution pattern match to known artifact mechanisms
- Presence in a known vulnerable region
- Decision tier and rationale
Record the decision for each variant in a tracking spreadsheet or database. This record becomes your audit trail and allows you to refine your thresholds over time. If you find that many Tier 3 flagged variants are later validated as true, your Tier 1 criteria may be too aggressive. If you find that many Tier 1 filtered variants would have been validated as true, your criteria need adjustment.
Calibrating Thresholds for Your Specific Data
The thresholds provided in this framework are starting points, not universal constants. Your specific data type, sequencing platform, library preparation method, and sample source all influence the appropriate thresholds. The forensic mitochondrial study reported average strand bias values near 1.0 across nine consecutive sequencing runs, providing a benchmark for expected performance on that platform [<a href="#ref-4">4</a>]. Your own data may show different baseline values.
To calibrate thresholds for your data, start with a training set of variants that you have validated using an orthogonal method. Sanger sequencing is the standard validation approach for single nucleotide variants. Alternatively, use a different sequencing platform or a targeted genotyping assay. Compare the strand bias metric distributions between validated true variants and validated artifacts. Identify the threshold that best separates the two groups.
The orange roughy SNP discovery study provides a model for this calibration approach. The researchers examined five predictors for their effect on obtaining an assayable SNP: depth of coverage, number of reads supporting a variant, polymorphism type, strand bias, and Illumina SNP probe design score. Their results showed that filtering out systematic sequencing errors substantially improved the efficiency of SNP discovery [<a href="#ref-3">3</a>]. By tracking which filtered variants converted to successful genotyping assays, they could assess whether their filtering decisions were appropriate.
Batch-Level Assessment and Adjustment
Strand bias can vary systematically between sequencing batches. The oxidative damage study found strong batch effects in the occurrence of low-frequency C to A and G to T substitutions [<a href="#ref-2">2</a>]. If you process samples in multiple batches, compare strand bias distributions between batches before applying filtering thresholds.
For each batch, calculate the distribution of strand bias metric values across all variants. If one batch shows consistently higher strand bias, investigate the library preparation conditions for that batch. Temperature, intensity, and cycle number during ultrasonic fragmentation influenced the extent of oxidative damage in the mitochondrial study [<a href="#ref-2">2</a>]. A batch processed under different conditions may require different filtering thresholds.
Record the batch information for each sample and include it in your quality control analysis. If you identify a batch with elevated strand bias, consider whether the data should be re-sequenced or processed with different filtering thresholds. The decision depends on the criticality of the samples and the extent of the bias.
Documenting the Decision Process
Documentation is essential for reproducible variant calling. For each project, record the following information:
- The version of the variant caller and the specific strand bias annotations used
- The decision framework and thresholds applied
- The number of variants filtered at each tier
- The number of variants flagged for downstream consideration
- The validation results for a sample of filtered and retained variants
This documentation allows other researchers to assess the quality of your variant calls and to apply different filtering decisions if they disagree with your thresholds. If you deposit your variant calls in public databases such as those maintained by NCBI, the quality of your variant calls affects the utility of the data for other researchers [<a href="#ref-8">8</a>]. Transparent documentation of filtering decisions improves the reliability of the database.
The Galaxy Training Network provides accessible workflow training for variant calling that includes guidance on interpreting VCF fields and applying quality filters [<a href="#ref-6">6</a>]. The EMBL-EBI training resources provide structured learning pathways for alignment quality assessment and variant calling best practices [<a href="#ref-7">7</a>]. These resources can help you implement a reproducible decision framework in your own workflow.
Common Mistakes in Decision Framework Implementation
One common mistake is applying the same thresholds to all variant types without considering the biological context. Germline variants are typically present at 50% or 100% allele frequency, so strand bias is easier to assess with confidence. Somatic variants may be present at low allele frequencies, making strand bias assessment less reliable. The GATK recommended thresholds for germline data may not be appropriate for somatic data.
Another common mistake is ignoring the interaction between coverage and strand bias. At low depth, the strand distribution of reads is subject to high variance. A true variant with 8 supporting reads might show 7 reads on the forward strand and 1 on the reverse strand, producing a strand bias flag even though the variant is real. The coverage-dependent thresholds used in the single-cell mitochondrial pipeline address this by requiring different levels of strand balance depending on the total read depth [<a href="#ref-5">5</a>]. For your own data, consider stratifying strand bias filtering by coverage.
A third mistake is failing to revisit filtering decisions as new information becomes available. If you validate a sample of filtered variants and find that many are true, your thresholds are too strict. If you validate a sample of retained variants and find that many are artifacts, your thresholds are too lenient. The decision framework should be treated as a living process that you refine based on validation results.
When to Escalate Beyond the Framework
The decision framework handles routine strand bias filtering decisions. Some situations require escalation to a bioinformatics specialist or a clinical genomics professional. Escalate when you observe strand bias in a high proportion of variants across multiple samples, when you are working with FFPE samples and need to implement SOBDetector or similar artifact filtering, when you are developing a variant calling pipeline for clinical or diagnostic use, or when you are uncertain whether a specific variant is a true variant or an artifact.
For clinical and diagnostic applications, strand bias filtering decisions have direct patient care implications. Filtering a true variant can lead to a missed diagnosis. Retaining an artifact can lead to an incorrect diagnosis. In this context, follow established guidelines and consult with clinical genomics professionals before making filtering decisions. The SOBDetector tool was developed for clinical FFPE samples, which are the most common tissue specimens stored in clinical practice [<a href="#ref-1">1</a>]. Using this tool for FFPE samples provides a systematic approach to artifact assessment, but the output should be interpreted in the context of the specific clinical question.
For forensic applications, rigorous quality control is required because the results may be used as legal evidence. The forensic validation study reported specific performance metrics, including strand bias values near 1.0, as part of the internal validation for routine application [<a href="#ref-4">4</a>]. If you are conducting forensic analysis, follow the validation requirements for your jurisdiction and document all quality control measures.
Frequently Asked Questions
What does a strand bias flag in my VCF file actually mean?
A strand bias flag indicates that the alternative allele is distributed unevenly between the forward and reverse strands. The variant caller has calculated a statistical metric, such as Fisher's exact test p-value or strand odds ratio, that suggests the observed distribution is unlikely to occur by chance. This flag is a warning that the variant may be a sequencing artifact instead of a true biological variant. You should examine the read depth and visualize the reads before deciding whether to filter the variant.
Should I always filter variants with high strand bias?
No. High strand bias alone is not sufficient to filter a variant. Consider the read depth, the substitution type, the sample type, and the biological context. A variant in a low coverage region may show apparent strand bias due to chance. A variant in an FFPE sample may show strand bias because of formalin-induced artifacts [<a href="#ref-1">1</a>]. A variant in a repetitive region may show strand bias because of alignment issues. Evaluate each variant in context before making a filtering decision.
What is the difference between FisherStrand and StrandOddsRatio?
FisherStrand is the p-value from Fisher's exact test comparing the allele distribution between strands. A low p-value indicates that the observed strand distribution is unlikely to occur by chance. StrandOddsRatio is the odds ratio for the alternative allele appearing on one strand versus the other. An odds ratio near 1.0 indicates balanced representation. Both metrics measure strand bias but use different statistical formulations. GATK provides recommended thresholds for both metrics, but you should adjust these thresholds based on your data type and coverage.
How do I calculate strand bias from a BAM file?
You can calculate strand bias from a BAM file by counting the number of forward-strand and reverse-strand reads supporting the reference and alternative alleles at a variant position. Use a genome browser to visualize the reads or use a command-line tool to extract the strand-specific counts. The DP4 field in SAMtools mpileup output contains these counts. You can then calculate the strand distribution and apply your filtering thresholds.
Why do FFPE samples show more strand bias than fresh frozen samples?
FFPE samples undergo formalin fixation, which introduces crosslinks and deamination events that create characteristic C to T and G to A substitutions. These artifacts are not randomly distributed across strands, producing strand bias. The SOBDetector tool was developed specifically to assess the probability that a detected mutation in an FFPE sample is an artifact [<a href="#ref-1">1</a>]. If you are working with FFPE samples, use this tool to filter artifacts instead of relying on generic strand bias thresholds.
Can strand bias affect mitochondrial DNA variant calling?
Yes. Mitochondrial DNA sequencing can show strand bias, particularly in specific amplicons or regions. The forensic mitochondrial genome study identified five amplicons with low read depth and high strand bias as the most vulnerable positions in the mitochondrial genome [<a href="#ref-4">4</a>]. The single-cell mitochondrial pipeline used a 1:3 strand bias ratio threshold to filter errors [<a href="#ref-5">5</a>]. If you are calling mitochondrial variants, monitor strand bias by amplicon and apply appropriate filtering.
What is the relationship between strand bias and oxidative damage?
Oxidative damage during DNA fragmentation introduces 8-oxoguanine lesions that lead to C to A and G to T substitutions. These artifacts show strand bias and are influenced by the fragmentation method. Ultrasonic fragmentation produces more oxidative damage than enzyme-based fragmentation. The mtDeOxoGer model was developed to filter these artifacts using strand orientation bias score, substitution type, variant allele frequency, local GC content, and sequence context [<a href="#ref-2">2</a>]. If you observe C to A and G to T substitutions with strand bias, consider oxidative damage as the cause.
How do I validate a variant that shows strand bias?
Use an orthogonal method to validate the variant. Sanger sequencing is the standard validation approach for single nucleotide variants. Alternatively, use a different sequencing platform or a targeted genotyping assay. If the variant is confirmed by the orthogonal method, it is likely a true variant despite the strand bias. If the orthogonal method does not confirm the variant, it was likely an artifact. Record the validation results for future reference.
Related Bioinformatics Guides
- Detecting Structural Variants with Long-Read Sequencing: Methods and Considerations
- From Raw Reads to Variants: A Diagnostic Blueprint for Next-Generation Sequencing (NGS) Workflows
- Genomic Data Analysis Tools: A Comparative Guide for Researchers
- Genomic Data Visualization Tools: Choosing and Using Them Effectively
- Volcano Plot Proteomics: How to Create and Interpret Them Effectively
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
[1] [Strand Orientation Bias Detector to determine the probability of FFPE sequencing artifacts.](https://pubmed.ncbi.nlm.nih.gov/34015811). Briefings in bioinformatics, 2021. [2] [mtDeOxoGer: A machine learning-based strategy to effectively filter artifactual mutations induced by oxidative damage in high-throughput mtDNA sequencing data.](https://pubmed.ncbi.nlm.nih.gov/41183721). Free radical biology & medicine, 2026. [3] [SNP discovery in nonmodel organisms: strand bias and base-substitution errors reduce conversion rates.](https://pubmed.ncbi.nlm.nih.gov/25388640). Molecular ecology resources, 2015. [4] [Mitochondrial genome sequencing with ForenSeq™ mtDNA Whole Genome Kit.](https://pubmed.ncbi.nlm.nih.gov/40117916). Forensic science international. Genetics, 2025. [5] [A bioinformatics pipeline for identifying homoplasmic and heteroplasmic mitochondrial DNA SNVs in single-cell RNA-Seq datasets.](https://pubmed.ncbi.nlm.nih.gov/41061772). Genomics, 2025. [6] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project. [7] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute. [8] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.