Artifact Filtering in Somatic Variant Calling: Distinguishing True Mutations from Sequencing and Alignment Errors
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Formalin-fixed paraffin-embedded (FFPE) tissue introduces C to T transition artifacts due to cytosine deamination, which can be distinguished from true mutations by strand orientation bias analysis, as FFPE artifacts preferentially affect one DNA strand.
- Oxidative damage, particularly 8-oxoguanine formation, leads to G to T transversions that also exhibit strand bias, necessitating strand orientation analysis and optimized DNA extraction/library preparation protocols.
- PCR amplification errors, prevalent in targeted sequencing due to extensive amplification, are random and harder to filter but can be partially addressed by analyzing PCR duplicate families to identify variants present in multiple independent amplification events.
- Alignment artifacts, arising from incorrect read placement in repetitive regions, are often clustered and can be effectively filtered using a panel of normals, which identifies recurrent variants across multiple control samples processed through the same pipeline.
- Single-cell whole-genome sequencing presents unique amplification artifact challenges, requiring specialized callers like SCcaller to differentiate true single-nucleotide variations from amplification-induced errors.
- Strand orientation bias analysis, using probabilistic models, and panel of normals filtering are core computational strategies to differentiate true somatic mutations from technical noise across various sequencing workflows.
Somatic variant calling aims to identify mutations that arise in tissues during life, yet sequencing and alignment errors routinely produce false positive calls that resemble genuine biological variation. This article defines the major artifact classes that confound somatic variant detection, explains their molecular and computational origins, and presents concrete filtering strategies including orientation bias filters and panel of normals approaches. The practical outcome for researchers and laboratory professionals is a decision framework for distinguishing true somatic mutations from technical noise across whole genome, whole exome, targeted, and single-cell sequencing workflows.
The Artifact Problem in Somatic Variant Calling
Somatic variant calling differs fundamentally from germline variant calling because the mutations of interest are often present at low allele fractions within a mixed cell population. A true somatic mutation may exist in only a fraction of the sequenced cells, while sequencing artifacts can mimic low-frequency variants across many genomic positions. The challenge is compounded by the fact that somatic mutations are implicated in cancer and have been associated with noncancerous diseases and aging, making accurate discrimination between real mutations and technical errors a prerequisite for meaningful biological interpretation [<a href="#ref-1">1</a>].
The scale of this problem becomes apparent when considering that most somatic mutations do not expand clonally and are unique to individual cells [<a href="#ref-1">1</a>]. In bulk sequencing, clonally expanded mutations can be detected through deep sequencing, but the vast majority of mutations exist at frequencies that overlap substantially with the error profiles of common sequencing platforms. Without rigorous artifact filtering, a typical somatic variant call set will contain a mixture of true mutations, sequencing errors, alignment artifacts, and sample preparation damage.
The consequences of inadequate filtering extend beyond wasted validation effort. In clonal hematopoiesis research, for example, differentiating true CHIP mutations from sequencing artifacts and germline variants is a considerable bioinformatic challenge [<a href="#ref-2">2</a>]. Small changes in filtering parameters can considerably increase misclassification and reduce the effect size of epidemiological associations [<a href="#ref-2">2</a>]. This observation underscores that artifact filtering is not a minor quality control step but a central determinant of downstream biological and clinical conclusions.
Major Artifact Classes and Their Origins
Formalin-Fixed Paraffin-Embedded Induced Deamination
Formalin-fixed paraffin-embedded tissue is the most common tissue specimen stored in clinical practice, yet formalin fixation introduces characteristic artifacts that complicate analysis [<a href="#ref-3">3</a>]. The fixation process causes deamination of cytosine bases, converting them to uracil. During sequencing library preparation and amplification, these uracils are read as thymine, producing C to T transitions that can be mistaken for genuine somatic mutations.
The artifact pattern from FFPE damage is distinctive because it preferentially affects one DNA strand. When the damage occurs on the original template strand, the resulting C to T change appears only in reads originating from that strand. This strand bias provides a computational handle for distinguishing FFPE artifacts from true mutations, which typically show representation on both strands.
The practical impact of FFPE artifacts is substantial because most clinical tissue specimens are stored in this format [<a href="#ref-3">3</a>]. Researchers working with archival material must therefore expect a higher background of C to T artifacts and apply dedicated filtering approaches. The Strand Orientation Bias Detector was developed specifically to assess the probability that a detected mutation is an artifact based on strand orientation patterns, using a Bayesian logistic regression model trained on The Cancer Genome Atlas whole exomes [<a href="#ref-3">3</a>].
Oxidative Damage During DNA Extraction and Library Preparation
Oxidative damage introduces another class of base modifications that generate sequencing artifacts. Guanine oxidation produces 8-oxoguanine, which mispairs with adenine during PCR amplification, leading to G to T transversions in the final sequencing data. Unlike FFPE deamination, oxidative damage can occur in fresh frozen samples if DNA extraction and library preparation are not optimized.
The key distinction between oxidative damage and true mutations is again strand bias. Oxidative lesions on one strand produce apparent variants only in reads derived from that strand. This pattern is analogous to FFPE deamination but produces a different base substitution spectrum. Laboratories processing fresh frozen tissue should still monitor for oxidative artifacts, particularly when samples have undergone multiple freeze-thaw cycles or prolonged storage.
PCR Amplification Errors
Polymerase chain reaction amplification introduces errors during library construction and target enrichment. DNA polymerases have characteristic error rates and mutational spectra, with some enzymes prone to specific substitution types. These errors become problematic when they occur early in the amplification process, because they are then replicated in many daughter molecules and can reach allele fractions that pass typical variant calling thresholds.
PCR artifacts are particularly challenging in targeted sequencing workflows where limited input DNA requires extensive amplification. The errors are random in position and strand, making them harder to distinguish from true mutations than FFPE or oxidative damage. However, PCR duplicates provide a partial solution: if the same position shows the same variant in multiple independent amplification events, the variant is more likely to be genuine. Conversely, variants appearing in only one PCR duplicate family are suspect.
Alignment Errors and Mapping Artifacts
Sequence alignment errors generate false variant calls when reads are placed incorrectly in the reference genome. Repetitive regions, segmental duplications, and regions with high homology to other genomic locations are particularly prone to misalignment. A read originating from one genomic location may align to a similar but nonidentical location, creating apparent mismatches that are actually sequence differences between paralogous regions.
Mapping artifacts are distinct from base-level sequencing errors because they arise from the computational placement of reads instead of from the sequencing chemistry itself. These artifacts often cluster in specific genomic regions and can be consistent across samples, making them detectable through panel of normals approaches. Regions with low mappability should be treated with caution regardless of the variant caller used.
Amplification Artifacts in Single-Cell Whole-Genome Sequencing
Single-cell whole-genome sequencing introduces unique artifact challenges because the entire genome must be amplified from a single copy of DNA. Multiple displacement amplification, a common method for single-cell genome amplification, has a characteristic error profile that differs from bulk sequencing approaches. The SCcaller software tool was developed specifically to filter out amplification artifacts when calling single-nucleotide variations and small insertions and deletions from single-cell sequencing data [<a href="#ref-1">1</a>].
The single-cell context amplifies the artifact problem because there is no population of cells to average out errors. A single amplification error early in the process can appear as a high-confidence variant in the final data. The protocol for single-cell whole-genome sequencing using SCMDA emphasizes both the efficiency and high fidelity of the amplification method and the computational filtering provided by SCcaller [<a href="#ref-1">1</a>]. Researchers entering single-cell somatic mutation analysis must recognize that the artifact landscape differs substantially from bulk sequencing.
Core Principles of Artifact Filtering
Strand Orientation Bias Analysis
Strand orientation bias is one of the most powerful discriminators between true somatic mutations and sequencing artifacts. True mutations exist in the original DNA template and should be detectable in reads from both the forward and reverse strands. Artifacts introduced during library preparation, such as FFPE deamination and oxidative damage, preferentially affect one strand and therefore show an imbalanced distribution across strands.
The Strand Orientation Bias Detector operationalizes this principle through a Bayesian logistic regression model trained on The Cancer Genome Atlas whole exomes [<a href="#ref-3">3</a>]. The tool is compatible with all common somatic SNV-calling pipelines and provides a probability that a given detected mutation is an artifact [<a href="#ref-3">3</a>]. This probabilistic output is more informative than a simple binary filter because it allows researchers to set thresholds appropriate for their specific application.
Orientation bias filtering is most effective for C to T and G to T artifacts, which are the signatures of deamination and oxidation respectively. True mutations can occasionally show strand bias due to sampling effects at low allele fractions, so the filter should be applied with awareness of sequencing depth. At very low coverage, even genuine mutations may appear strand-biased by chance.
Panel of Normals Filtering
A panel of normals is a collection of sequencing data from normal samples processed through the same pipeline as the study samples. The rationale is that systematic artifacts, including alignment errors and platform-specific noise, will appear recurrently across normal samples. Variants detected in the study samples that also appear in the panel of normals are likely artifacts instead of true somatic mutations.
The panel of normals approach is particularly effective for filtering alignment artifacts and platform-specific errors that are consistent across samples. These artifacts often cluster in specific genomic regions with low mappability or high sequence complexity. A well-constructed panel of normals captures these systematic errors and allows their removal from the study call set.
The utility of panel of normals filtering depends on the panel's size and relevance. A panel should ideally include samples processed with the same library preparation kit, sequencing platform, and bioinformatics pipeline as the study samples. Small panels may miss rare artifacts, while panels from different tissue types or preparation methods may introduce inappropriate filtering. The clonal hematopoiesis field has highlighted that recurrent artifactual variants can be unmasked when calling at scale, emphasizing the importance of specialized filtering approaches for several genes including TET2 and ASXL1 [<a href="#ref-2">2</a>].
Variant Annotation and Population Frequency Filtering
Variant annotation provides contextual information that helps distinguish true somatic mutations from artifacts and germline variants. Population frequency databases can identify variants that are common in the general population, which are unlikely to be somatic mutations in a single individual. However, population frequency filtering must be applied carefully because some true somatic mutations occur in genes that also harbor germline polymorphisms.
The challenge of distinguishing CHIP mutations from germline variants is well documented [<a href="#ref-2">2</a>]. CHIP mutations can be identified in peripheral blood samples sequenced using approaches that cover the whole genome, the whole exome, or targeted genetic regions [<a href="#ref-2">2</a>]. A stepwise method that combines filtering based on sequencing metrics, variant annotation, and population-based associations has been shown to increase the accuracy of CHIP calls [<a href="#ref-2">2</a>]. This combined approach recognizes that no single filter is sufficient and that the integration of multiple evidence types improves accuracy.
Sequencing Metrics and Quality Filters
Basic sequencing metrics provide the first line of defense against artifacts. Read depth, base quality scores, mapping quality, and variant allele fraction all contribute to variant confidence. Low-depth positions are more susceptible to sequencing errors, while low base quality scores indicate uncertainty in the underlying base calls.
The challenge in somatic variant calling is that true somatic mutations can exist at low allele fractions, particularly in heterogeneous tissue samples or small clonal populations. Overly aggressive quality filtering can remove genuine low-frequency mutations, while insufficient filtering retains artifacts. The ArCH pipeline addresses this challenge by combining the output of four variant calling tools and filtering based on variant characteristics and sequencing error rate estimation [<a href="#ref-4">4</a>]. This approach improves sensitivity and positive predictive value at low allele frequencies compared to standard application of commonly used variant calling approaches [<a href="#ref-4">4</a>].
Practical Workflow for Artifact Filtering
Step 1: Define the Artifact Landscape for Your Data Type
Before applying filters, characterize the expected artifact profile for your specific data type. Formalin-fixed paraffin-embedded tissue requires attention to deamination artifacts [<a href="#ref-3">3</a>]. Fresh frozen tissue should be monitored for oxidative damage. Single-cell whole-genome sequencing demands filtering of amplification artifacts [<a href="#ref-1">1</a>]. Targeted sequencing panels may have region-specific artifacts related to primer design and capture efficiency.
Document the sample type, DNA extraction method, library preparation kit, sequencing platform, and bioinformatics pipeline for each batch of samples. This metadata is essential for interpreting artifact patterns and for constructing appropriate panels of normals.
Step 2: Apply Strand Orientation Bias Filtering
Run strand orientation bias analysis on all candidate somatic variants. The Strand Orientation Bias Detector provides a probability that a detected mutation is an artifact based on strand orientation patterns [<a href="#ref-3">3</a>]. Set a threshold appropriate for your application, recognizing that more stringent thresholds remove more artifacts but may also remove true mutations with genuine strand bias due to sampling.
For FFPE samples, pay particular attention to C to T transitions, which are the hallmark of formalin-induced deamination. For samples with suspected oxidative damage, monitor G to T transversions. Record the strand bias metric for each variant to allow retrospective adjustment of thresholds.
Step 3: Construct and Apply a Panel of Normals
Assemble a panel of normal samples processed through the same pipeline as your study samples. The panel should include samples from the same tissue type, prepared with the same methods, and sequenced on the same platform. For targeted sequencing, include samples captured with the same panel design.
Filter out any study variant that appears in the panel of normals, particularly if it appears recurrently. The clonal hematopoiesis field has demonstrated that recurrent artifactual variants can be unmasked when calling at scale [<a href="#ref-2">2</a>]. A variant appearing in multiple normal samples is almost certainly an artifact or a germline polymorphism instead of a true somatic mutation.
Step 4: Apply Variant Annotation and Population Frequency Filters
Annotate all remaining variants with gene information, variant type, predicted functional impact, and population frequency. Filter out variants that are common in population databases, as these are unlikely to be somatic mutations. However, be cautious with genes known to harbor both germline polymorphisms and somatic mutations.
The stepwise method combining sequencing metrics, variant annotation, and population-based associations has been shown to increase the accuracy of CHIP calls [<a href="#ref-2">2</a>]. This approach is applicable beyond clonal hematopoiesis and provides a template for somatic variant filtering in other contexts.
Step 5: Integrate Multiple Caller Outputs
Consider using multiple variant callers and integrating their outputs. The ArCH pipeline combines the output of four variant calling tools and filters based on variant characteristics and sequencing error rate estimation [<a href="#ref-4">4</a>]. This approach improves sensitivity and positive predictive value compared to standard application of commonly used variant calling approaches [<a href="#ref-4">4</a>].
Variants called by multiple independent tools are more likely to be genuine, while variants called by only one tool warrant additional scrutiny. The integration of multiple callers also provides a form of internal validation that is particularly valuable when orthogonal validation data are not available.
Step 6: Validate with Orthogonal Methods
When possible, validate candidate somatic variants with orthogonal methods. Technical replicates provide a straightforward validation approach: if the same variant is detected in independent library preparations from the same sample, it is more likely to be genuine. The ArCH pipeline was validated using deep targeted sequencing data from tumor-normal dilutions, blood samples with orthogonal validation, and blood samples with technical replicates [<a href="#ref-4">4</a>].
For single-cell whole-genome sequencing, the SCMDA protocol with SCcaller provides high genomic coverage and high accuracy in single-nucleotide variation and small insertion and deletion calling from the same single-cell genome [<a href="#ref-1">1</a>]. The protocol takes 12 to 15 hours from single-cell isolation to library preparation and 3 to 7 days of data processing [<a href="#ref-1">1</a>].
At a Glance: Artifact Types and Filtering Strategies
| Artifact Type | Molecular Origin | Characteristic Signature | Primary Filtering Strategy |
|---|---|---|---|
| FFPE deamination | Formalin-induced cytosine deamination to uracil | C to T transitions with strand bias | Strand orientation bias analysis [<a href="#ref-3">3</a>] |
| Oxidative damage | 8-oxoguanine mispairing with adenine | G to T transversions with strand bias | Strand orientation bias analysis, antioxidant protocols |
| PCR amplification errors | Polymerase misincorporation during amplification | Random substitutions, often low allele fraction | Duplicate family analysis, multiple caller integration [<a href="#ref-4">4</a>] |
| Alignment artifacts | Incorrect read placement in repetitive or homologous regions | Variants clustered in low-mappability regions | Panel of normals filtering [<a href="#ref-2">2</a>] |
| Single-cell amplification artifacts | Errors during whole-genome amplification | Amplification-specific error profile | Dedicated single-cell callers such as SCcaller [<a href="#ref-1">1</a>] |
| Germline contamination | Constitutional variants mistaken for somatic | Variants present in population databases | Population frequency filtering, matched germline comparison [<a href="#ref-5">5</a>] |
Options and Tradeoffs in Artifact Filtering
Matched Germline Comparison Versus Tumor-Only Calling
The gold standard for somatic variant calling is comparison against a matched germline sample, typically from blood or adjacent normal tissue. This approach directly excludes germline variants and focuses the analysis on somatic changes. However, matched germline samples are not always available, particularly in clinical settings where only tumor tissue is sequenced.
PipeIT2 addresses this clinical need by providing a tumor-only somatic variant calling workflow that achieves greater than 95 percent recall for variants with variant allele fraction above 10 percent, reliably detects driver and actionable mutations, and filters out most germline mutations and sequencing artifacts [<a href="#ref-5">5</a>]. The workflow is specific for Ion Torrent sequencing data and is enclosed in a Singularity container for reproducibility [<a href="#ref-5">5</a>].
The tradeoff between matched and tumor-only approaches is sensitivity versus practicality. Matched germline comparison provides the most definitive exclusion of germline variants but requires additional sequencing. Tumor-only calling is more practical but requires more aggressive filtering to exclude germline variants and artifacts. The choice depends on the research question, sample availability, and clinical context.
Stringency Thresholds and Their Consequences
The choice of filtering thresholds has profound consequences for downstream analysis. The clonal hematopoiesis field has demonstrated that small changes in filtering parameters can considerably increase misclassification and reduce the effect size of epidemiological associations [<a href="#ref-2">2</a>]. More stringent filtering reduces false positives but may remove true low-frequency mutations. Less stringent filtering retains more candidate variants but increases the validation burden.
The appropriate threshold depends on the application. Clinical diagnostic workflows require high specificity to avoid reporting false mutations, while discovery research may tolerate more false positives to maximize sensitivity. The ArCH pipeline provides customizable parameters adaptable to multiple sequencing technologies, research questions, and datasets [<a href="#ref-4">4</a>], allowing researchers to tune filtering stringency to their specific needs.
Computational Resources and Reproducibility
Artifact filtering adds computational overhead to variant calling workflows. Multiple caller integration, panel of normals construction, and strand orientation bias analysis all require additional processing time and storage. Reproducibility is a separate concern: the same data processed through different pipelines can produce different results, complicating cross-study comparisons.
Containerized workflows address reproducibility concerns by encapsulating the analysis environment. PipeIT2 is enclosed in a Singularity container [<a href="#ref-5">5</a>], and the ArCH pipeline is available as an end-to-end cloud-based pipeline [<a href="#ref-4">4</a>]. These approaches ensure that the same inputs produce the same outputs regardless of the computing environment. The nf-core documentation provides standards for community pipeline usage and configuration that support reproducible workflow execution [<a href="#ref-6">6</a>].
Records and Measurements for Artifact Filtering
Essential Records for Each Variant Call Set
Maintain comprehensive records for each somatic variant call set to enable retrospective analysis and threshold adjustment. For each variant, record the chromosome, position, reference allele, alternate allele, variant allele fraction, read depth, base quality scores, mapping quality, strand bias metrics, and the presence in panel of normals. Also record the variant caller used, the filtering thresholds applied, and the version of all software tools.
For clonal hematopoiesis studies, record the sequencing approach used, whether whole genome, whole exome, or targeted genetic regions [<a href="#ref-2">2</a>]. The prevalence and clone size of CHIP mutations vary with sequencing depth and coverage, and these parameters must be documented for meaningful interpretation.
Quality Metrics for Filtering Performance
Track the number of variants removed at each filtering step to assess the performance of the filtering pipeline. A well-functioning pipeline should remove a predictable fraction of variants at each stage, with the largest removals occurring at the artifact-specific filters. Sudden changes in the number of variants removed may indicate a batch effect, a change in sample quality, or a problem with the filtering pipeline itself.
For single-cell whole-genome sequencing, record the amplification method used and the variant calling software applied. The SCMDA protocol ensures efficiency and high fidelity in amplification, and the SCcaller software filters out amplification artifacts [<a href="#ref-1">1</a>]. These methodological details are essential for interpreting mutation burden estimates and comparing results across studies.
Documentation of Filtering Parameters
Document all filtering parameters and their rationale. The clonal hematopoiesis field has shown that small changes in filtering parameters can considerably increase misclassification [<a href="#ref-2">2</a>], so parameter choices must be recorded and justified. Include the version numbers of all software tools, the reference genome build, and the annotation databases used.
For the Strand Orientation Bias Detector, record the threshold used for artifact probability and the training data underlying the model [<a href="#ref-3">3</a>]. For panel of normals filtering, document the panel composition, including the number of samples, tissue types, and processing methods.
Common Failure Patterns in Artifact Filtering
Overfiltering True Low-Frequency Mutations
The most common failure pattern is applying filters so aggressively that true low-frequency mutations are removed. This is particularly problematic in clonal hematopoiesis research, where CHIP mutations are generally present at low allelic fractions [<a href="#ref-4">4</a>]. Overly stringent filters can eliminate the very mutations of interest, biasing downstream analyses toward high-frequency mutations and distorting biological conclusions.
The solution is to validate filtering thresholds against known positive controls. Dilution series, where tumor DNA is mixed with normal DNA at known ratios, provide a quantitative assessment of sensitivity at different allele fractions. The ArCH pipeline was validated using deep targeted sequencing data generated from six acute myeloid leukemia patient tumor-normal dilutions [<a href="#ref-4">4</a>], demonstrating the value of this approach.
Underfiltering Recurrent Artifacts
The opposite failure pattern is retaining recurrent artifacts that masquerade as genuine mutations. The clonal hematopoiesis field has highlighted that recurrent artifactual variants can be unmasked when calling at scale [<a href="#ref-2">2</a>]. These artifacts may appear in multiple samples and can be mistaken for true recurrent mutations, leading to incorrect biological conclusions.
Panel of normals filtering is the primary defense against recurrent artifacts. A well-constructed panel captures systematic errors and allows their removal from study call sets. The challenge is that panel of normals construction requires access to a sufficient number of normal samples processed through the same pipeline, which may not be available in all research settings.
Ignoring Batch Effects
Batch effects in sample processing can introduce artifacts that are consistent within a batch but absent across batches. Changes in reagent lots, instrument calibration, or personnel can all affect the artifact landscape. Failure to account for batch effects can lead to spurious associations between variants and biological or clinical variables.
The solution is to process samples in a randomized or balanced design, to include batch information in the analysis, and to monitor quality metrics across batches. The panel of normals should include samples from all batches to capture batch-specific artifacts.
Applying Germline Filters to Somatic Variants
Germline variant filters, such as population frequency thresholds, can inadvertently remove true somatic mutations in genes that also harbor common germline polymorphisms. This is particularly problematic in genes like TET2 and ASXL1, where specialized filtering approaches are required [<a href="#ref-2">2</a>]. The distinction between germline and somatic variants requires careful consideration of the biological context and the variant characteristics.
Matched germline comparison provides the most definitive resolution of this issue [<a href="#ref-5">5</a>]. When matched germline data are unavailable, tumor-only workflows like PipeIT2 apply additional filtering to exclude germline mutations [<a href="#ref-5">5</a>], but these filters must be calibrated to avoid removing true somatic mutations.
Limitations of Artifact Filtering Approaches
Incomplete Knowledge of Artifact Signatures
Artifact filtering relies on known signatures of technical errors, but the full spectrum of artifacts is not completely characterized. New library preparation methods, sequencing platforms, and analysis pipelines may introduce artifacts with unfamiliar signatures. The Bayesian logistic regression model underlying the Strand Orientation Bias Detector was trained on The Cancer Genome Atlas whole exomes [<a href="#ref-3">3</a>], and its performance on other data types may differ.
Researchers should remain alert for unexpected variant patterns and investigate instead of dismiss them. The clonal hematopoiesis field has demonstrated that recurrent artifactual variants can be unmasked when calling at scale [<a href="#ref-2">2</a>], suggesting that some artifacts are only recognized when large datasets are analyzed.
Difficulty Distinguishing Low-Frequency Mutations from Artifacts
The fundamental limitation of artifact filtering is that true low-frequency mutations and sequencing artifacts can be indistinguishable at the level of individual variants. Both may appear at similar allele fractions, with similar quality metrics, and in similar genomic contexts. The ArCH pipeline improves sensitivity and positive predictive value at low allele frequencies compared to standard approaches [<a href="#ref-4">4</a>], but it does not eliminate the ambiguity.
Orthogonal validation remains the definitive approach for resolving this ambiguity. Technical replicates, independent sequencing methods, and functional assays can confirm whether a candidate variant is genuine. However, orthogonal validation is expensive and time-consuming, and it is not feasible for every candidate variant.
Population Frequency Databases and Diversity
Population frequency databases are essential for distinguishing somatic mutations from germline variants, but their utility depends on the diversity of the underlying populations. The clonal hematopoiesis field has found a significantly lower prevalence of CHIP in individuals of self-reported Latino or Hispanic ethnicity in the All of Us Research Program, highlighting the importance of including diverse populations [<a href="#ref-2">2</a>]. Variants that are rare in one population may be common in another, and databases that lack diversity may lead to incorrect filtering decisions.
Researchers should use population frequency databases with awareness of their limitations and should consider the ancestry of their study samples when applying frequency filters.
Quality and Reproducibility Controls
Containerized Workflows
Containerization provides a practical solution to reproducibility challenges in artifact filtering. PipeIT2 is enclosed in a Singularity container [<a href="#ref-5">5</a>], ensuring that the analysis environment is consistent across different computing platforms. The nf-core documentation provides standards for community pipeline usage and configuration [<a href="#ref-6">6</a>], supporting reproducible workflow execution.
Containers also facilitate the sharing of analysis pipelines among collaborators and the publication of reproducible analyses. A containerized workflow can be versioned, archived, and re-executed to verify results.
Training and Skill Development
Artifact filtering requires a combination of molecular biology knowledge and computational skills. The EMBL-EBI Training program provides bioinformatics learning pathways and practical analysis education [<a href="#ref-7">7</a>], while The Carpentries Lessons offer foundational computing, data, shell, Git, and programming training [<a href="#ref-8">8</a>]. The Galaxy Training Network provides accessible workflow training and analysis tutorials [<a href="#ref-9">9</a>].
Laboratory professionals entering somatic variant calling should invest in training that covers both the biological basis of artifacts and the computational tools for filtering them. The Bioconductor Project provides official package, workflow, installation, and reproducible genomic-analysis documentation [<a href="#ref-10">10</a>], and the NCBI Data Resources provide official descriptions of databases, search systems, sequence resources, and analysis services [<a href="#ref-11">11</a>].
Version Control and Documentation
Version control is essential for reproducible artifact filtering. Record the versions of all software tools, reference genomes, and annotation databases used in the analysis. The nf-core documentation emphasizes community pipeline standards and reproducible workflow context [<a href="#ref-6">6</a>], and these principles apply to individual analyses as well.
Documentation should include also the final filtering parameters but also the rationale for those parameters. The clonal hematopoiesis field has shown that small changes in filtering parameters can considerably increase misclassification [<a href="#ref-2">2</a>], so parameter choices must be justified and recorded.
Professional Escalation Criteria
When to Seek Specialized Consultation
Certain situations warrant consultation with specialized bioinformaticians or molecular pathologists. If a variant call set contains an unexpectedly high number of candidate variants, if recurrent artifacts are suspected in specific genes, or if filtering decisions have clinical implications, seek expert input. The clonal hematopoiesis field has highlighted the importance of specialized filtering approaches for several genes, including TET2 and ASXL1 [<a href="#ref-2">2</a>], and similar gene-specific considerations may apply in other contexts.
When to Revisit the Filtering Pipeline
The filtering pipeline should be revisited when new data types are introduced, when the artifact landscape changes, or when validation results indicate problems. If orthogonal validation reveals a high false positive rate, the filtering thresholds should be adjusted. If new artifact signatures are identified in the literature, the pipeline should be updated to address them.
When to Consider Alternative Approaches
If artifact filtering cannot adequately distinguish true mutations from technical noise, consider alternative experimental approaches. Single-cell whole-genome sequencing with SCMDA and SCcaller provides high genomic coverage and high accuracy in single-nucleotide variation and small insertion and deletion calling [<a href="#ref-1">1</a>], and may be appropriate when bulk sequencing cannot resolve the artifact problem. Alternatively, orthogonal validation with an independent sequencing method may provide the necessary discrimination.
Frequently Asked Questions
What is the difference between somatic and germline variant calling?
Somatic variant calling identifies mutations that arise in tissues during life and are present in only a subset of cells, while germline variant calling identifies variants inherited from parents and present in all cells. Somatic mutations are implicated in cancer and have been associated with noncancerous diseases and aging [<a href="#ref-1">1</a>]. The distinction matters for filtering because germline variants can be excluded by population frequency filters or matched germline comparison, while somatic mutations require different filtering approaches focused on distinguishing true mutations from sequencing artifacts.
Why do formalin-fixed paraffin-embedded samples produce so many artifacts?
Formalin fixation causes deamination of cytosine bases, converting them to uracil, which is then read as thymine during sequencing. This produces C to T transitions that can be mistaken for genuine somatic mutations [<a href="#ref-3">3</a>]. The artifacts show strand bias because the damage preferentially affects one DNA strand. The Strand Orientation Bias Detector was developed specifically to assess the probability that a detected mutation is an artifact based on strand orientation patterns [<a href="#ref-3">3</a>].
What is a panel of normals and why is it important?
A panel of normals is a collection of sequencing data from normal samples processed through the same pipeline as the study samples. It is used to identify and filter out systematic artifacts, including alignment errors and platform-specific noise, that appear recurrently across normal samples. The clonal hematopoiesis field has demonstrated that recurrent artifactual variants can be unmasked when calling at scale [<a href="#ref-2">2</a>], highlighting the importance of panel of normals filtering.
How does strand orientation bias distinguish true mutations from artifacts?
True mutations exist in the original DNA template and should be detectable in reads from both the forward and reverse strands. Artifacts introduced during library preparation, such as FFPE deamination and oxidative damage, preferentially affect one strand and therefore show an imbalanced distribution across strands. The Strand Orientation Bias Detector uses a Bayesian logistic regression model trained on The Cancer Genome Atlas whole exomes to assess the probability that a detected mutation is an artifact [<a href="#ref-3">3</a>].
Can tumor-only somatic variant calling be reliable?
Tumor-only somatic variant calling can be reliable when appropriate filtering is applied. PipeIT2 is a tumor-only somatic variant calling workflow that achieves greater than 95 percent recall for variants with variant allele fraction above 10 percent, reliably detects driver and actionable mutations, and filters out most germline mutations and sequencing artifacts [<a href="#ref-5">5</a>]. The workflow is specific for Ion Torrent sequencing data and is enclosed in a Singularity container for reproducibility [<a href="#ref-5">5</a>].
What are the special challenges of single-cell whole-genome sequencing for somatic mutations?
Single-cell whole-genome sequencing requires amplification of the entire genome from a single copy of DNA, which introduces amplification artifacts that differ from bulk sequencing errors. The SCcaller software tool was developed to filter out amplification artifacts when calling single-nucleotide variations and small insertions and deletions from single-cell sequencing data [<a href="#ref-1">1</a>]. The SCMDA protocol ensures efficiency and high fidelity in amplification [<a href="#ref-1">1</a>].
How do filtering parameters affect downstream analyses?
Small changes in filtering parameters can considerably increase misclassification and reduce the effect size of epidemiological associations in clonal hematopoiesis research [<a href="#ref-2">2</a>]. More stringent filtering reduces false positives but may remove true low-frequency mutations, while less stringent filtering retains more candidate variants but increases the validation burden. The appropriate threshold depends on the application, and parameter choices must be documented and justified.
What is the role of multiple variant callers in artifact filtering?
Using multiple variant callers and integrating their outputs can improve the accuracy of somatic variant calling. The ArCH pipeline combines the output of four variant calling tools and filters based on variant characteristics and sequencing error rate estimation [<a href="#ref-4">4</a>]. This approach improves sensitivity and positive predictive value at low allele frequencies compared to standard application of commonly used variant calling approaches [<a href="#ref-4">4</a>].
Related Bioinformatics Guides
- Detecting Structural Variants with Long-Read Sequencing: Methods and Considerations
- Oxford Nanopore Sequencing: From Sample to Base Calls
- Single-Cell RNA Sequencing Quality Control: A Practical Guide to Filtering and Metrics
- Variant Calling in Whole Exome Sequencing (WES): Principles, Algorithms, and Veterinary Applications
- Metagenomics Sequencing: Technologies and Considerations
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
[1] [Analyzing somatic mutations by single-cell whole-genome sequencing.](https://pubmed.ncbi.nlm.nih.gov/37996541). Nature protocols, 2024. [2] [A practical approach to curate clonal hematopoiesis of indeterminate potential in human genetic data sets.](https://pubmed.ncbi.nlm.nih.gov/36652671). Blood, 2023. [3] [Strand Orientation Bias Detector to determine the probability of FFPE sequencing artifacts.](https://pubmed.ncbi.nlm.nih.gov/34015811). Briefings in bioinformatics, 2021. [4] [ArCH: improving the performance of clonal hematopoiesis variant calling and interpretation.](https://pubmed.ncbi.nlm.nih.gov/38485690). Bioinformatics (Oxford, England), 2024. [5] [PipeIT2: A tumor-only somatic variant calling workflow for molecular diagnostic Ion Torrent sequencing data.](https://pubmed.ncbi.nlm.nih.gov/36796655). Genomics, 2023. [6] [nf-core Documentation](https://nf-co.re/docs). nf-core. [7] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute. [8] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries. [9] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project. [10] [Bioconductor](https://bioconductor.org/). Bioconductor Project. [11] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.