How to Filter Somatic Variants Using Population Databases: Distinguishing Germline Contamination from True Somatic Mutations
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Population frequency databases (e.g., gnomAD, 1000 Genomes) are crucial for filtering germline variants from somatic mutation call sets, but naive application can erroneously remove true somatic mutations or retain pathogenic germline variants.
- A tiered filtering approach is recommended, starting with removal of common polymorphisms (e.g., >1% allele frequency), followed by rare polymorphisms not annotated as clinically relevant, and retaining variants with known cancer relevance or very low population frequencies (<0.1%) for further scrutiny.
- Clonal hematopoiesis (CH) complicates filtering as somatic CH variants can appear in population databases; variants in established CH genes require careful review to distinguish somatic origin from germline inheritance, as CH variants can also cause Mendelian conditions.
- Integrating clinical annotation databases (e.g., ClinVar) and cancer-specific databases (e.g., COSMIC) is essential to retain known oncogenic somatic variants (e.g., TP53 hotspot mutations) that may exceed population frequency thresholds.
- Overestimation of Tumor Mutational Burden (TMB) is a significant risk with tumor-only filtering, necessitating validation against matched normal data or using filtering strategies calibrated to germline subtraction results to ensure accurate biomarker assessment for immunotherapy response.
- Sampling bias in population databases and the presence of oncogenic variants within them are inherent limitations, requiring careful consideration of population-specific frequencies and annotation for known germline cancer predisposition variants.
Somatic variant calling from tumor sequencing data frequently produces a mixed call set containing both genuine somatic mutations and germline variants that were not removed during alignment or initial variant detection. Population frequency databases such as gnomAD and 1000 Genomes provide the reference data needed to filter out common germline polymorphisms, but naive filtering approaches can remove true somatic mutations or retain pathogenic germline variants that masquerade as somatic events. This article describes a practical workflow for using population databases to separate germline contamination from true somatic mutations, with specific attention to tumor purity, clonality, and the known limitations of population frequency data.
The workflow presented here is intended for biology students, researchers, laboratory professionals, and life-science practitioners who work with targeted sequencing panels, whole-exome sequencing, or whole-genome sequencing data from tumor samples. The core problem is straightforward: somatic variant callers identify variants that differ from a reference genome, but they cannot distinguish between a variant that arose in the tumor and a variant that was inherited from a parent or acquired early in development. Population databases provide allele frequencies across large cohorts, and variants that appear at high frequency in the general population are unlikely to be somatic mutations driving cancer. However, the relationship between population frequency and somatic origin is not simple, and several documented failure modes can compromise the filtering process.
The Problem of Germline Contamination in Somatic Call Sets
Germline contamination in somatic variant call sets occurs when variants present in the normal genome of the patient are reported alongside true somatic mutations. This happens for several reasons. When tumor-only sequencing is performed without a matched normal sample, there is no direct way to subtract germline variants from the somatic call set. Even when matched tumor-normal pairs are sequenced, incomplete subtraction can occur due to differences in coverage, allelic fraction, or bioinformatic processing between the tumor and normal samples.
The clinical and research consequences of germline contamination are substantial. A germline variant that is mistakenly classified as somatic can lead to incorrect interpretation of the mutational landscape of a tumor, false assignment of driver status to a benign polymorphism, and inaccurate tumor mutational burden estimates. Conversely, a true somatic mutation that is filtered out because it appears in a population database can cause a researcher to miss a clinically actionable alteration. The challenge is to design a filtering strategy that maximizes sensitivity for true somatic variants while minimizing the retention of germline contamination.
Population databases were developed specifically to address this problem. The 1000 Genomes Project, the Exome Aggregation Consortium (ExAC), and the Genome Aggregation Database (gnomAD) provide large-scale reference data of genetic variation across diverse human populations. These resources allow researchers to filter out common benign variants and identify rare variants of clinical importance based on their frequency in the human population. The underlying assumption is that a variant observed at high frequency in thousands of healthy individuals is unlikely to be a somatic mutation driving cancer in a single patient.
However, this assumption has documented exceptions. A study examining TP53 variants in gnomAD and ExAC found that a significant number of known oncogenic TP53 variants are present in these population databases, including hotspot variants that occur as both somatic and germline events in human cancer. The study demonstrated that germline TP53 variants in the human population are more frequent than previously thought and concluded that population databases must be used with caution and annotated for the presence of oncogenic variants to improve their clinical utility. This finding has direct implications for somatic variant filtering: a filtering strategy that removes all variants above a population frequency threshold will also remove bona fide pathogenic germline variants that are relevant to cancer predisposition, and it may remove rare somatic mutations that happen to be present at low frequency in population cohorts due to clonal hematopoiesis or other biological processes.
At a Glance: Population Database Filtering Decision Table
The following table summarizes the key decisions involved in filtering somatic variants using population databases. Each row represents a common scenario encountered in somatic variant analysis and the recommended approach based on current evidence.
| Scenario | Recommended Approach | Key Consideration |
|---|---|---|
| Tumor-only sequencing with no matched normal | Use population frequency filtering as a primary germline subtraction method, but retain variants with known cancer relevance regardless of frequency | Population databases contain oncogenic variants, so frequency alone cannot distinguish somatic from germline |
| Paired tumor-normal sequencing | Use population databases as a secondary filter after germline subtraction to remove residual contamination | Germline subtraction is the criterion standard, but incomplete subtraction can occur due to coverage differences |
| Variants in clonal hematopoiesis genes | Review population frequency data carefully and do not automatically filter variants in CH-associated genes | Somatic variants from clonal hematopoiesis can affect population frequencies and filtering decisions |
| High tumor mutational burden estimates | Apply tiered filtering that removes common variants first, then rare variants, then internal pipeline filters | Naive removal of all variants observed in population databases can overestimate TMB |
| Known cancer hotspot variants | Retain variants in cancer hotspot regions even if they appear in population databases | Oncogenic TP53 and other tumor suppressor variants are present in gnomAD and ExAC |
Core Principles of Population Database Filtering
Understanding Allele Frequency Thresholds
The most common approach to filtering somatic variants with population databases is to set an allele frequency threshold. Variants with a population allele frequency above the threshold are considered germline and removed from the somatic call set. The choice of threshold depends on the sequencing context, the population being studied, and the goals of the analysis.
For targeted cancer panels, a common threshold is 1% population allele frequency. Variants observed at or above this frequency in any population database are typically classified as germline polymorphisms and removed from the somatic call set. More stringent thresholds of 0.1% or 0.01% may be used when the goal is to identify very rare somatic mutations, but these thresholds increase the risk of removing true somatic variants that happen to be present at low frequency in population cohorts.
The 1000 Genomes Project, ExAC, and gnomAD each provide allele frequency data across multiple populations. The choice of which database to use and which population subset to query depends on the ancestry of the patient or sample being analyzed. Using a global allele frequency instead of a population-specific frequency can mask variants that are common in one population but rare in another. For example, a variant that is common in African populations but rare in European populations would be filtered out using a global threshold but retained using a European-specific threshold.
The Role of Variant Annotation in Filtering Decisions
Population frequency alone is insufficient for distinguishing germline contamination from true somatic mutations. Variant annotation with clinical databases such as ClinVar and cancer-specific databases such as COSMIC provides additional context that can improve filtering decisions. A variant that appears at low frequency in population databases but is a known recurrent somatic mutation in cancer should be retained even if it exceeds a strict frequency threshold.
A study of tumor-only variant filtration strategies found that the optimal algorithm required an ordinal filtering approach using information from variant population databases (1000 Genomes Phase 3, ESP6500, ExAC), clinical mutation databases (ClinVar), and information on recurring clinically relevant somatic variants. The study demonstrated that this approach could define clinically relevant somatic variants from tumor-only analysis with sensitivity of 97% to 99% and specificity of 87% to 94%. The ordinal approach is important because it applies filters in a specific sequence: first remove common population variants, then remove variants with benign clinical annotations, then retain variants with known somatic relevance.
Clonal Hematopoiesis and Its Effect on Population Frequencies
Clonal hematopoiesis (CH) is a biological process that complicates the use of population databases for somatic variant filtering. CH occurs when somatic mutations in hematopoietic stem cells provide a proliferative advantage, leading to the expansion of a clone of blood cells carrying the mutation. These somatic variants can appear in population databases because the databases are built from blood-derived DNA, and the CH-associated variants will affect variant frequencies, depletion scores, and downstream filtering.
The presence of CH-associated variants in population databases creates a specific problem for somatic variant filtering. A variant that is truly somatic in the tumor but also present in the blood due to CH will appear in population databases at a frequency that reflects the prevalence of CH in the population. Default filtering of variants or genes associated with CH risks filtering bona fide germline variants, because variants associated with CH can also cause Mendelian conditions. The recommendation is to carefully review variants in established CH genes and to consider whether a variant might be somatic in origin despite appearing in population databases.
Practical Workflow for Filtering Somatic Variants
Step 1: Prepare the Input Variant Call Set
The filtering workflow begins with a variant call set produced by a somatic variant caller. Common callers include Mutect2, Strelka2, VarScan2, and others. The input file should be in VCF format and should contain all variants detected in the tumor sample, including those that may be germline contamination. Before applying population frequency filters, the call set should be normalized to ensure consistent representation of variants. Normalization includes left-aligning indels, trimming variant alleles to a minimal representation, and ensuring that all variants are represented in the same genomic coordinate system.
Quality filters should be applied before population frequency filtering. Variants with low read depth, low mapping quality, or high strand bias should be removed or flagged, as these variants are more likely to be sequencing artifacts than true biological variants. The specific quality thresholds depend on the sequencing platform and the variant caller used, but a minimum read depth of 20 to 30 reads and a minimum variant allele frequency of 5% are common starting points for targeted panels.
Step 2: Annotate Variants with Population Frequency Data
The next step is to annotate the variant call set with population frequency data from gnomAD, 1000 Genomes, and other databases. Several annotation tools are available for this purpose. The Ensembl Variant Effect Predictor (VEP) can annotate variants with allele frequencies from multiple population databases in a single run. ANNOVAR provides similar functionality with a different interface. The choice of annotation tool depends on the user's familiarity with the software and the specific databases that need to be queried.
The annotation should include the following fields for each variant: the allele frequency in gnomAD (overall and by population), the allele frequency in 1000 Genomes (overall and by population), and the allele frequency in any other relevant databases such as ExAC or ESP6500. The annotation should also include the number of alleles observed in each database, as a variant observed in a single allele out of thousands is less informative than a variant observed in hundreds of alleles.
Step 3: Apply Tiered Frequency Filters
The filtering process should be applied in tiers instead of as a single threshold. The tiered approach reduces the risk of removing true somatic variants while still eliminating the majority of germline contamination.
Tier 1 removes variants with a population allele frequency above 1% in any population database. This tier removes common polymorphisms that are almost certainly germline in origin. Variants removed at this tier are unlikely to be true somatic mutations, although rare exceptions exist for variants that are common in specific populations and also recurrent in cancer.
Tier 2 removes variants with a population allele frequency between 0.1% and 1% that are not annotated as clinically relevant in cancer databases. This tier removes rare polymorphisms that are still more likely to be germline than somatic. Variants in this frequency range that are known cancer hotspots or that have evidence of somatic recurrence should be retained.
Tier 3 retains variants with a population allele frequency below 0.1% for further analysis. These variants are rare enough that they could be true somatic mutations, but they could also be rare germline variants. The decision to retain or remove variants in this tier depends on additional evidence, including the variant's presence in cancer databases, the predicted functional impact, and the variant allele frequency in the tumor sample.
Step 4: Integrate Clinical and Cancer-Specific Annotations
After applying population frequency filters, the remaining variants should be annotated with clinical and cancer-specific information. ClinVar provides information about the clinical significance of variants, including whether a variant is pathogenic, likely pathogenic, benign, or of uncertain significance. COSMIC provides information about the frequency of variants in cancer samples and whether a variant is known to be recurrent in specific cancer types.
The integration of these annotations serves two purposes. First, it identifies variants that should be retained despite appearing in population databases. A variant that is a known cancer hotspot and is present in COSMIC should be retained even if it has a population allele frequency above the filtering threshold. Second, it identifies variants that should be removed despite being rare in the population. A variant that is classified as benign in ClinVar and is not present in COSMIC is more likely to be a rare germline polymorphism than a true somatic mutation.
Step 5: Review Variants in Clonal Hematopoiesis Genes
Variants in genes associated with clonal hematopoiesis require special attention during the filtering process. The 36 established CH genes associated with neurodevelopmental conditions should be reviewed carefully, as should other genes known to be recurrently mutated in CH. The presence of a variant in a CH gene does not automatically mean the variant is germline, but it does mean that the population frequency data should be interpreted with caution.
For variants in CH genes, the following considerations apply. If the variant is a known CH driver mutation and appears at a population allele frequency consistent with CH prevalence, it may be somatic in origin and should be retained. If the variant is in a CH gene but is not a known driver mutation and appears at a population allele frequency consistent with germline variation, it may be germline and should be removed. The distinction requires knowledge of the specific gene and variant, and consultation with the literature or a clinical genomics expert may be necessary.
Step 6: Validate the Filtered Call Set
The final step in the workflow is to validate the filtered call set. Validation can take several forms. If matched normal data are available for a subset of samples, the filtered somatic variants can be compared against the germline subtraction results to assess concordance. If validation data are not available, the filtered call set can be assessed for internal consistency, such as the expected distribution of variant allele frequencies and the expected mutation spectrum for the tumor type.
The validation step is particularly important for tumor mutational burden estimation. A study comparing tumor-only TMB with matched germline-subtracted TMB found that tumor-only filtering approaches can overestimate TMB. The study evaluated three levels of filtering: removing variants with an allelic fraction of at least 1% in ExAC, removing all variants observed in population databases, and using an internal tumor-only pipeline. The results showed significantly higher estimates of TMB with the first level of filtering compared with germline subtraction. This finding underscores the importance of validating tumor-only filtering approaches against matched normal data when possible.
Options and Tradeoffs in Population Database Filtering
Tumor-Only Versus Paired Tumor-Normal Sequencing
The choice between tumor-only and paired tumor-normal sequencing has a direct impact on the filtering strategy. Paired tumor-normal sequencing is the criterion standard for somatic variant identification because it allows direct subtraction of germline variants. However, paired testing has challenges, including increased cost of dual sample testing and the identification of germline cancer predisposing variants that may have implications for the patient and their family.
Tumor-only sequencing avoids the cost and complexity of paired testing but requires in silico filtration to identify somatic variants. The barrier to tumor-only variant filtration is defining a reliable approach with high sensitivity and specificity. The ordinal filtering approach described earlier, which uses population databases, clinical mutation databases, and information on recurring somatic variants, has been shown to achieve sensitivity of 97% to 99% and specificity of 87% to 94% for identifying clinically relevant somatic variants. These performance metrics are encouraging, but they also indicate that tumor-only filtering is not perfect and that some true somatic variants will be missed while some germline variants will be retained.
Choice of Population Database
The choice of population database affects the filtering results. gnomAD is the largest and most comprehensive population database, with data from over 100,000 exomes and 15,000 genomes. ExAC is the predecessor to gnomAD and contains data from over 60,000 exomes. The 1000 Genomes Project provides whole-genome sequencing data from approximately 2,500 individuals across 26 populations.
Each database has strengths and limitations. gnomAD provides the most detailed population frequency data, including allele frequencies for multiple ancestral populations and filtering recommendations based on sequencing quality. ExAC provides similar data but with a smaller sample size. The 1000 Genomes Project provides whole-genome data, which is useful for variants in non-coding regions that may not be well represented in exome-based databases.
The choice of database should be guided by the specific analysis. For targeted cancer panels that focus on coding regions, gnomAD or ExAC are appropriate. For whole-genome sequencing, the 1000 Genomes Project provides complementary data. In practice, most filtering workflows annotate variants with data from multiple databases and use the maximum allele frequency across databases as the filtering criterion.
Stringency of Frequency Thresholds
The stringency of the population frequency threshold is a tradeoff between sensitivity and specificity. A high threshold (such as 5% or 10%) removes more variants and reduces the risk of retaining germline contamination, but it also increases the risk of removing true somatic mutations that happen to be common in the population. A low threshold (such as 0.01% or 0.1%) retains more variants and increases sensitivity for rare somatic mutations, but it also increases the risk of retaining germline contamination.
The optimal threshold depends on the context. For clinical diagnostic testing, where the consequences of missing a true somatic mutation are severe, a lower threshold may be appropriate. For research applications where the goal is to identify the complete mutational landscape of a tumor, a higher threshold may be acceptable. The tiered filtering approach described in this article provides a compromise by applying different thresholds at different stages of the filtering process.
Observations and Measurements for Filtering Quality
Variant Allele Frequency Distributions
The distribution of variant allele frequencies in the filtered call set provides information about the quality of the filtering process. True somatic mutations in a tumor sample typically have variant allele frequencies that reflect the purity of the tumor sample and the clonality of the mutation. Clonal mutations that are present in all tumor cells will have variant allele frequencies close to the tumor purity, while subclonal mutations will have lower variant allele frequencies.
Germline variants, in contrast, typically have variant allele frequencies close to 50% (for heterozygous variants) or 100% (for homozygous variants), regardless of tumor purity. After filtering, the remaining variants should show a distribution of variant allele frequencies that is consistent with somatic mutations. If the filtered call set still contains a substantial number of variants with allele frequencies near 50%, this suggests that germline contamination has not been fully removed.
Transition to Transversion Ratio
The transition to transversion ratio (Ti/Tv ratio) is a quality metric that can be used to assess the filtered call set. Germline variants typically have a Ti/Tv ratio of approximately 2.0 to 2.1, reflecting the mutational processes that generate germline variation. Somatic mutations in cancer can have different Ti/Tv ratios depending on the mutational signatures present in the tumor.
If the filtered call set has a Ti/Tv ratio close to 2.0, this may indicate that germline contamination remains. If the Ti/Tv ratio is lower or higher, this may indicate that the filtering has successfully removed germline variants and retained somatic mutations. The expected Ti/Tv ratio for somatic mutations depends on the tumor type and the mutational signatures present, so this metric should be interpreted in the context of the specific analysis.
Comparison with Matched Normal Data
The most direct measurement of filtering quality is comparison with matched normal data. If matched normal sequencing data are available for a subset of samples, the filtered somatic variants can be compared with the variants identified by germline subtraction. The concordance between the two approaches provides a quantitative measure of filtering performance.
The comparison should assess both sensitivity and specificity. Sensitivity is the proportion of true somatic variants (identified by germline subtraction) that are retained by the filtering approach. Specificity is the proportion of variants retained by the filtering approach that are true somatic variants. The study of tumor-only filtration strategies reported sensitivity of 97% to 99% and specificity of 87% to 94%, providing a benchmark for evaluating filtering performance.
Records and Documentation for Reproducible Filtering
Documenting Filtering Parameters
Reproducible filtering requires documentation of all parameters used in the analysis. The documentation should include the version of the population database used, the specific allele frequency thresholds applied, the annotation tool and version, and the reference genome build. Changes in any of these parameters can affect the filtering results, so the documentation should be sufficiently detailed to allow the analysis to be repeated exactly.
The documentation should also include the rationale for the chosen parameters. For example, if a 1% population frequency threshold was chosen, the documentation should explain why this threshold was selected and what the expected impact on sensitivity and specificity would be. This information is important for interpreting the results and for making adjustments in future analyses.
Version Control for Filtering Workflows
Version control is essential for reproducible filtering workflows. The workflow should be stored in a version control system such as Git, and all changes to the workflow should be tracked. This allows researchers to identify which version of the workflow was used for a particular analysis and to reproduce the analysis if needed.
The Carpentries provides lessons on version control with Git that are relevant to bioinformatics workflows. These lessons cover the basics of Git, including creating repositories, committing changes, and branching. For more complex workflows, the nf-core documentation provides guidance on developing and using reproducible pipelines with Nextflow. The nf-core community has established standards for pipeline development that include version control, containerization, and automated testing.
Containerization and Environment Reproducibility
Containerization ensures that the filtering workflow runs in a consistent environment regardless of the computing infrastructure. Docker and Singularity are the most commonly used containerization tools in bioinformatics. Containers package the workflow code, dependencies, and reference data into a single image that can be run on any system with the appropriate container runtime.
The Bioconductor project provides documentation on reproducible genomic analysis, including guidance on managing R package versions and environments. The Bioconductor approach uses the BiocManager package to install and manage package versions, and the sessionInfo function to document the exact versions of all packages used in an analysis. This level of documentation is important for ensuring that filtering results can be reproduced by other researchers.
Common Failure Patterns in Population Database Filtering
Overfiltering of True Somatic Mutations
The most common failure pattern in population database filtering is the removal of true somatic mutations. This occurs when a somatic mutation is present in a population database at a frequency above the filtering threshold. The TP53 study provides a clear example of this failure mode: a significant number of oncogenic TP53 variants are included in gnomAD and ExAC, and most of them correspond to TP53 hotspot variants occurring as somatic and germline events in human cancer.
Overfiltering can be reduced by integrating cancer-specific annotations into the filtering process. Variants that are known cancer hotspots or that are recurrent in COSMIC should be retained regardless of population frequency. The ordinal filtering approach described earlier addresses this by applying population frequency filters first and then using clinical and cancer-specific annotations to retain variants of known relevance.
Underfiltering of Germline Contamination
The opposite failure pattern is underfiltering, where germline contamination remains in the somatic call set. This occurs when the population frequency threshold is too low or when the population database does not contain the germline variant. Rare germline variants that are not present in population databases will not be removed by population frequency filtering, and these variants will remain in the somatic call set as contamination.
Underfiltering can be reduced by using multiple population databases and by applying the maximum allele frequency across databases as the filtering criterion. However, even with comprehensive population databases, some germline variants will not be captured. For clinical applications, the risk of underfiltering can be mitigated by using paired tumor-normal sequencing when possible.
Misinterpretation of Clonal Hematopoiesis Variants
Variants in clonal hematopoiesis genes present a specific failure pattern. These variants can be somatic in origin but appear in population databases because they are present in the blood of the individuals included in the database. Default filtering of variants or genes associated with CH risks filtering bona fide germline variants, as variants associated with CH can also cause Mendelian conditions.
The recommendation for handling CH variants is to review them carefully instead of applying automatic filters. The 36 established CH genes associated with neurodevelopmental conditions should be reviewed individually, and the decision to retain or remove a variant should be based on the specific gene, the specific variant, and the clinical context.
Overestimation of Tumor Mutational Burden
Tumor mutational burden estimation is particularly sensitive to filtering decisions. The study comparing tumor-only TMB with germline-subtracted TMB found that tumor-only filtering approaches can overestimate TMB. The study evaluated three levels of filtering and found significantly higher estimates of TMB with the first level of filtering, which removed variants with an allelic fraction of at least 1% in ExAC.
The overestimation of TMB has clinical implications because TMB is used as a biomarker for immunotherapy response. An overestimated TMB could lead to inappropriate treatment decisions. The study recommends validating tumor-only TMB estimates against matched normal data when possible and using a filtering approach that has been calibrated against germline subtraction results.
Limitations of Population Database Filtering
Population Database Composition and Sampling Bias
Population databases are not representative of all human populations. gnomAD and ExAC are built primarily from individuals of European ancestry, with smaller contributions from African, East Asian, South Asian, and Latino populations. The 1000 Genomes Project includes more diverse populations but has a smaller total sample size. This sampling bias means that variants common in underrepresented populations may be absent from population databases, leading to underfiltering of germline contamination in samples from these populations.
The sampling bias also affects the interpretation of allele frequencies. A variant that is rare in the European population but common in an African population will have a low overall allele frequency in gnomAD, even though it is a common polymorphism in the African population. Using population-specific allele frequencies instead of overall frequencies can reduce this problem, but the population-specific frequencies are only available for populations that are well represented in the database.
Presence of Oncogenic Variants in Population Databases
The presence of oncogenic variants in population databases is a documented limitation that affects filtering decisions. The TP53 study demonstrated that germline TP53 variants in the human population are more frequent than previously thought and that population databases must be used with caution and annotated for the presence of oncogenic variants to improve their clinical utility.
This limitation means that population frequency alone cannot distinguish between a germline variant that predisposes to cancer and a somatic mutation that drives cancer. A filtering strategy that removes all variants above a population frequency threshold will also remove pathogenic germline variants that are relevant to cancer predisposition. The integration of clinical annotations and cancer-specific databases is essential for addressing this limitation.
Clonal Hematopoiesis and Its Effect on Frequency Estimates
Clonal hematopoiesis affects the allele frequencies of somatic variants in population databases. Somatic variants that provide a proliferative advantage in hematopoietic stem cells will be present in the blood of a subset of individuals in the database, and their allele frequencies will reflect the prevalence of CH in the population. This can lead to the misclassification of somatic variants as germline variants based on their population frequency.
The effect of CH on population frequencies is particularly problematic for genes that are recurrently mutated in both CH and cancer. A variant in a CH gene that is also a cancer driver will appear in population databases at a frequency that reflects CH prevalence, and a filtering strategy that removes variants above a frequency threshold will remove this variant from the somatic call set. The recommendation is to review variants in CH genes carefully and to consider the possibility that a variant is somatic in origin despite appearing in population databases.
Safety and Regulatory Context for Clinical Applications
Clinical Validation Requirements
The use of population database filtering in clinical diagnostic testing is subject to regulatory requirements that vary by jurisdiction. In the United States, laboratory-developed tests are regulated by the Centers for Medicare and Medicaid Services under the Clinical Laboratory Improvement Amendments (CLIA). Laboratories performing clinical tumor sequencing must validate their bioinformatics pipelines, including the filtering approach, to ensure that the results are accurate and reliable.
The validation process should include an assessment of the filtering approach's sensitivity and specificity for identifying clinically relevant somatic variants. The study of tumor-only filtration strategies provides a framework for this validation, with reported sensitivity of 97% to 99% and specificity of 87% to 94%. Laboratories should document their validation results and make them available for inspection by regulatory authorities.
Reporting of Germline Findings
The use of population database filtering in tumor-only sequencing has implications for the reporting of germline findings. When a variant that is known to be pathogenic in the germline is identified in a tumor sample, the laboratory must decide whether to report this finding to the clinician. The identification of germline cancer predisposing variants is a challenge of paired tumor-normal testing, and tumor-only testing can inadvertently identify germline variants that have implications for the patient and their family.
The American College of Medical Genetics and Genomics (ACMG) has published recommendations for the reporting of secondary findings from clinical sequencing. These recommendations identify a list of genes in which pathogenic variants should be reported regardless of the indication for testing. Laboratories performing tumor sequencing should be aware of these recommendations and should have policies in place for handling germline findings that are identified incidentally.
Professional Escalation Criteria
The filtering workflow described in this article is intended for use by researchers and laboratory professionals with appropriate training in bioinformatics and genomics. When the filtering results are ambiguous or when the interpretation has clinical implications, escalation to a professional with appropriate expertise is recommended.
Escalation is appropriate in the following situations: when a variant in a clonal hematopoiesis gene has uncertain significance and the decision to retain or remove the variant affects clinical interpretation, when a variant is present in a population database at a frequency that is inconsistent with its predicted functional impact, when the tumor mutational burden estimate is borderline and the filtering approach could affect treatment decisions, and when a germline pathogenic variant is identified incidentally in a tumor sample. In these situations, consultation with a molecular pathologist, clinical geneticist, or other qualified professional is recommended.
Frequently Asked Questions
What is the difference between germline and somatic variants in sequencing data?
Germline variants are present in the DNA of every cell in the body and are inherited from one or both parents. They are typically present at an allele frequency of 50% (heterozygous) or 100% (homozygous) in all tissues. Somatic variants arise during the lifetime of an individual and are present only in the cells that descend from the cell in which the mutation occurred. In a tumor sample, somatic variants are present at allele frequencies that depend on tumor purity and clonality. Somatic variant callers identify variants that differ from the reference genome, but they cannot distinguish between germline and somatic variants without additional information.
How do population databases like gnomAD and 1000 Genomes help filter somatic variants?
Population databases provide allele frequency data for genetic variants observed in large cohorts of individuals. The 1000 Genomes Project, ExAC, and gnomAD were developed to provide large-scale reference data of genetic variations for various populations. The underlying assumption is that a variant observed at high frequency in thousands of healthy individuals is unlikely to be a somatic mutation driving cancer in a single patient. By removing variants that are common in the general population, researchers can reduce germline contamination in somatic call sets. However, this approach has documented limitations, including the presence of oncogenic variants in population databases and the effect of clonal hematopoiesis on allele frequencies.
What allele frequency threshold should I use for filtering somatic variants?
The choice of allele frequency threshold depends on the sequencing context and the goals of the analysis. A common threshold is 1% population allele frequency, which removes common polymorphisms that are almost certainly germline in origin. More stringent thresholds of 0.1% or 0.01% can be used to identify very rare somatic mutations, but these thresholds increase the risk of removing true somatic variants. A tiered filtering approach is recommended, where common variants are removed first, then rare variants, and then variants with clinical annotations are reviewed individually. The optimal threshold should be validated against matched normal data when possible.
Why do oncogenic variants appear in population databases?
Oncogenic variants appear in population databases for several reasons. Some variants that are pathogenic in the germline are present in the population at low frequencies, and these variants are included in population databases because the databases are designed to capture all genetic variation. The TP53 study demonstrated that a significant number of oncogenic TP53 variants are included in gnomAD and ExAC, including hotspot variants that occur as somatic and germline events in human cancer. Additionally, clonal hematopoiesis can introduce somatic variants into the blood of individuals included in population databases, and these variants will appear in the database even though they are somatic in origin.
How does clonal hematopoiesis affect population database filtering?
Clonal hematopoiesis affects population database filtering because somatic variants that provide a proliferative advantage in hematopoietic stem cells will be present in the blood of a subset of individuals in the database. These variants will affect variant frequencies, depletion scores, and downstream filtering. Default filtering of variants or genes associated with CH risks filtering bona fide germline variants, as variants associated with CH can also cause Mendelian conditions. The recommendation is to review variants in established CH genes carefully and to consider whether a variant might be somatic in origin despite appearing in population databases.
What is the difference between tumor-only and paired tumor-normal filtering?
Paired tumor-normal sequencing is the criterion standard for somatic variant identification because it allows direct subtraction of germline variants. However, paired testing has challenges, including increased cost of dual sample testing and identification of germline cancer predisposing variants. Tumor-only sequencing avoids these challenges but requires in silico filtration to identify somatic variants. The barrier to tumor-only variant filtration is defining a reliable approach with high sensitivity and specificity. Studies have shown that an ordinal filtering approach using population databases, clinical mutation databases, and information on recurring somatic variants can achieve sensitivity of 97% to 99% and specificity of 87% to 94%.
How can I validate my somatic variant filtering approach?
Validation of a somatic variant filtering approach requires comparison with a criterion standard. If matched normal data are available for a subset of samples, the filtered somatic variants can be compared with the variants identified by germline subtraction. The concordance between the two approaches provides a quantitative measure of filtering performance. If matched normal data are not available, the filtered call set can be assessed for internal consistency, such as the expected distribution of variant allele frequencies and the expected mutation spectrum for the tumor type. The tumor mutational burden estimate can also be compared with published values for the tumor type.
What should I do when a variant appears in both population databases and cancer databases?
When a variant appears in both population databases and cancer databases, the interpretation depends on the specific variant and the context. A variant that is a known cancer hotspot and is present in COSMIC should be retained even if it has a population allele frequency above the filtering threshold. A variant that is classified as benign in ClinVar and is not present in COSMIC is more likely to be a rare germline polymorphism than a true somatic mutation. The ordinal filtering approach addresses this by applying population frequency filters first and then using clinical and cancer-specific annotations to retain variants of known relevance. When the interpretation is ambiguous, consultation with a molecular pathologist or clinical geneticist is recommended.
Related Bioinformatics Guides
- Detecting Structural Variants with Long-Read Sequencing: Methods and Considerations
- Functional Annotation of Metagenomes: A Guide to Databases and Pipelines
- Metagenomic Contamination Control: Best Practices for Clean Data
- Genomic Data Analysis Tools: A Comparative Guide for Researchers
- Metagenomics Functional Profiling: Tools and Databases for Pathway Analysis
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- High prevalence of cancer-associated TP53 variants in the gnomAD database: A word of caution concerning the use of variant filtering.. Human mutation, 2019.
- Interpreting variants in genes affected by clonal hematopoiesis in population data.. Human genetics, 2024.
- Somatic Tumor Variant Filtration Strategies to Optimize Tumor-Only Molecular Profiling Using Targeted Next-Generation Sequencing Panels.. The Journal of molecular diagnostics : JMD, 2019.
- Integrative Neoepitope Discovery in Glioblastoma via HLA Class I Profiling and AlphaFold2-Multimer.. Biomedicines, 2025.
- Tumor Mutational Burden From Tumor-Only Sequencing Compared With Germline Subtraction From Paired Tumor and Normal Specimens.. JAMA network open, 2020.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.