How to Build a Panel of Normals for Somatic Variant Filtering: A Practical Guide

By Dr. Zubair Khalid, DVM, MS, PhD ·

How to Build a Panel of Normals for Somatic Variant Filtering: A Practical Guide

Key Takeaways

  • A Panel of Normals (PoN) is critical for somatic variant calling by cataloging recurrent technical artifacts and germline polymorphisms specific to a given sequencing workflow, thereby distinguishing true mutations from systematic errors.
  • Artifacts are workflow-specific, necessitating PoN construction from normal samples processed through the identical library preparation, sequencing, and bioinformatics pipeline as the somatic samples to accurately capture platform-specific biases.
  • Adequate sample size is crucial, with a minimum of 20-30 samples recommended for targeted panels and 40+ for whole exome sequencing to ensure robust capture of rare artifacts and platform-specific errors.
  • Joint genotyping across all normal samples is strongly recommended over per-sample calling to improve variant calling accuracy and provide a consistent framework for identifying recurrent artifacts.
  • Validation using positive control samples with known somatic variants is essential to confirm the PoN effectively removes artifacts without compromising sensitivity for genuine somatic mutations.
  • PoN maintenance requires regular updates, version control, and monitoring of quality metrics to account for changes in sequencing platforms, reagents, or bioinformatics software that can alter the artifact profile.

Scope and Reader Context

A panel of normals (PoN) is a collection of sequencing data from non-diseased samples that captures the recurrent technical artifacts and germline polymorphisms present in your specific sequencing workflow. Building a PoN is a prerequisite for reliable somatic variant calling because sequencing platforms introduce systematic errors that mimic true mutations. This guide provides a practical protocol for constructing a PoN, determining adequate sample size, integrating the PoN into variant calling workflows, and validating that the panel effectively removes artifacts without discarding genuine somatic variants. The target reader is a laboratory professional or bioinformatics practitioner who has access to raw sequencing data, a compute environment, and a working knowledge of command-line tools. The guidance applies to whole exome sequencing, targeted gene panels, and whole genome sequencing workflows used in cancer diagnostics, rare disease research, and emerging applications of somatic variant calling beyond oncology.

Why Somatic Variant Calling Requires a Panel of Normals

Somatic variant calling identifies mutations that are present in a diseased tissue sample but absent from a matched normal sample from the same individual. The fundamental challenge is that sequencing errors, alignment artifacts, and library preparation biases produce variant calls that are indistinguishable from true low-frequency mutations without additional context. A PoN provides that context by cataloging the positions where your specific workflow routinely produces false positive calls.

The scale of the artifact problem is substantial. In a benchmarking study of whole exome sequencing across 21 centers in the German network for personalized medicine, somatic variant calling achieved a mean positive percentage agreement of 76 percent against a consensus variant list, with variant filtering identified as the main cause of divergent calls. Adjusting filter criteria and re-analysis increased agreement to 88 percent for all variants and 97 percent for clinically relevant variants. This finding demonstrates that filtering decisions, including the use of a PoN, are the primary determinant of somatic variant calling accuracy.

The need for a PoN extends beyond cancer diagnostics. Somatic variant callers are increasingly repurposed for detecting low-allele-fraction variants in developmental biology, reproductive medicine, virology, and mitochondrial genetics. In these applications, allele fractions frequently fall below 5 percent, where sequencing artifacts are most prevalent. A review of somatic callers used outside cancer contexts found that panel-of-normal filtering markedly improves specificity when detecting post-zygotic variants at 1 to 3 percent allele fraction, mitochondrial heteroplasmy as low as 0.4 to 0.5 percent, and minority viral variants around 5 percent. The statistical assumptions of somatic callers, including diploid genomes and moderate allele fractions, require careful adaptation for these applications, and a PoN is a central component of that adaptation.

At a Glance: Panel of Normals Construction Decisions

Decision PointRecommended ApproachEvidence BasisCommon Error
Sample sourceMatched normal tissue or blood from healthy individuals processed with the same workflowArtifacts are workflow-specific, not sample-specificUsing public datasets from different sequencing platforms
Sample sizeMinimum 20 to 30 samples for targeted panels, 40 or more for whole exomeLarger panels capture rare artifacts and platform-specific errorsBuilding a PoN from fewer than 10 samples
Variant calling strategyGermline calling with joint genotyping across all normal samplesCaptures population polymorphisms and recurrent artifactsCalling variants per sample and merging without joint genotyping
Artifact filteringRemove variants present in 2 or more PoN samples or exceeding a population frequency thresholdRecurrent artifacts appear across multiple normalsRemoving all variants found in any single normal sample
ValidationCompare somatic calls with and without PoN filtering on a positive control sampleConfirms the PoN removes artifacts without losing true variantsDeploying a PoN without validation on known variants

Core Principles of Panel of Normals Construction

Artifacts Are Workflow-Specific

Sequencing artifacts arise from multiple sources that are unique to your laboratory's protocols. Library preparation introduces errors during PCR amplification, particularly in GC-rich regions and homopolymer tracts. Sequencing by synthesis generates base-specific errors that vary by instrument model and chemistry version. Alignment algorithms mis-map reads in repetitive regions, creating false variant calls at specific genomic positions. Each combination of these factors produces a characteristic artifact profile that cannot be transferred between laboratories or even between different runs on the same instrument.

The practical implication is that a PoN must be built from data generated by the same workflow that produces your somatic samples. Using a public PoN from a different sequencing platform or library preparation method will not capture your specific artifacts and may introduce false assumptions about which positions are reliable. The NCBI provides access to extensive sequencing datasets, but these are useful for understanding general artifact patterns, not for substituting your own PoN.

Germline Polymorphisms Confound Somatic Calls

A second category of false somatic calls comes from germline variants that are present in the tumor sample but absent from the matched normal. This situation arises when the normal sample has low coverage at a polymorphic position, when the variant is a rare polymorphism not captured in population databases, or when the normal sample is contaminated. A study of 1,417 tumors with matched tumor-normal exomes found that 33.8 percent of germline single nucleotide variants within a clinically oriented gene panel were rare and at risk of being falsely reported as somatic. This finding underscores that a PoN serves a dual purpose: filtering technical artifacts and flagging positions where germline variation is common enough to warrant caution.

Statistical Foundations of Artifact Detection

Somatic variant callers use statistical models to distinguish true variants from noise. Most callers assume diploid genomes and moderate allele fractions, using Beta-binomial or Poisson-based models to compare tumor and normal samples. The PoN extends these models by providing an empirical distribution of allele fractions at each position across many normal samples. A variant call in the tumor sample that matches the allele fraction distribution observed in the PoN is likely an artifact or a germline polymorphism, while a variant that is absent from the PoN and exceeds the expected error rate is more likely to be a true somatic mutation.

The statistical power of a PoN depends on both the number of samples and the sequencing depth. More samples provide better estimates of the artifact allele fraction distribution at each position. Higher depth in each normal sample increases the sensitivity for detecting low-level artifacts. A PoN built from 20 samples at 100x coverage will detect artifacts present at 1 percent allele fraction in multiple samples, but will miss artifacts that occur in only one sample or at very low allele fractions.

Sample Selection for Panel of Normals Construction

Source Tissue Considerations

The ideal PoN samples are normal tissues from healthy individuals or matched normal samples from patients whose tumors are being analyzed. Blood-derived DNA is the most common source because it is easy to obtain and represents the germline genome. For cancer applications, matched normal samples from the same patients provide the additional benefit of capturing patient-specific germline variants, although these are typically handled by the tumor-normal comparison instead of the PoN.

Tissue-specific considerations apply when the PoN is used for non-cancer applications. For detecting post-zygotic variants in developmental biology, the PoN should include normal samples from the same tissue type as the diseased samples, because different tissues accumulate different patterns of somatic variation. For mitochondrial heteroplasmy detection, the PoN must account for the high copy number and maternal inheritance patterns of mitochondrial DNA.

Sample Size Determination

The minimum sample size for a PoN depends on the target sequencing breadth and the desired sensitivity for artifact detection. For targeted gene panels, where the number of interrogated positions is limited, 20 to 30 normal samples provide adequate coverage of recurrent artifacts. For whole exome sequencing, where the target space is larger and artifacts are more diverse, 40 or more samples are recommended.

The statistical rationale for these numbers relates to the frequency of artifacts. A recurrent artifact that appears in 5 percent of normal samples has a high probability of being detected in a PoN of 20 samples. The probability of detecting an artifact present at frequency f in a PoN of n samples is 1 minus (1 minus f) to the power of n. For an artifact present at 5 percent frequency, a PoN of 20 samples has a 64 percent chance of containing at least one instance, while a PoN of 40 samples has an 87 percent chance. For rare artifacts present at 1 percent frequency, a PoN of 40 samples has only a 33 percent chance of detection, which is why additional filtering strategies beyond the PoN are necessary.

Sample Quality Requirements

Each normal sample in the PoN must meet the same quality standards as the somatic samples in your workflow. Minimum coverage requirements should match or exceed the coverage expected for somatic samples. Samples with excessive duplication rates, contamination, or poor alignment metrics should be excluded. The Galaxy Training Network provides accessible tutorials on quality assessment and quality control for sequencing data that can be applied to PoN sample selection.

For clinical applications, the samples used to build the PoN should be de-identified and collected under appropriate ethical approvals. The EMBL-EBI Training resources include modules on responsible data handling and the ethical considerations of using human genomic data in research and diagnostics.

Variant Calling Workflow for Panel of Normals Construction

Alignment and Preprocessing

The first step in building a PoN is aligning the raw sequencing reads from each normal sample to the reference genome. The alignment parameters must be identical to those used for somatic samples in your workflow. Differences in alignment software, reference genome version, or alignment parameters will introduce artifacts that are not representative of your somatic variant calling pipeline.

After alignment, each sample should undergo the same preprocessing steps used for somatic samples, including marking duplicates, base quality score recalibration, and indel realignment. These steps correct systematic errors introduced during library preparation and sequencing. The nf-core documentation provides standardized pipeline configurations that ensure consistent preprocessing across samples, which is valuable for PoN construction because it eliminates a source of variability.

Germline Variant Calling

The PoN is built by calling germline variants in each normal sample using a germline variant caller. The choice of caller depends on your workflow and the type of variants you need to detect. Haplotype-based callers are generally preferred because they handle indels more accurately than allele-specific callers. The caller should be configured to output all candidate variants, including those with low quality scores, because the PoN needs to capture artifacts that might otherwise be filtered.

Joint genotyping across all PoN samples is strongly recommended over per-sample calling followed by merging. Joint genotyping considers the genotype likelihoods across all samples simultaneously, which improves the accuracy of genotype calls at positions with low coverage in individual samples. This approach also provides a consistent framework for identifying variants that are present in multiple samples, which is the primary signal for artifact detection.

Variant Representation and Normalization

Before building the PoN, all variant calls must be normalized to a consistent representation. Variants at the same position can be represented differently depending on the caller and the surrounding sequence context. Left-alignment and trimming of variants to a minimal representation ensures that the same variant is represented identically across all samples. Without normalization, the same artifact might appear as a different variant in different samples and escape detection by the PoN.

Normalization is particularly important for indels in repetitive regions, where the same biological variant can be represented at multiple positions. The Bioconductor project provides packages for variant normalization and annotation that can be integrated into PoN construction workflows.

Building the Panel of Normals

Combining Variant Calls Across Samples

The core of PoN construction is combining the variant calls from all normal samples into a single panel. This process involves collecting all variants found in any sample, determining which samples carry each variant, and calculating summary statistics for each variant position. The summary statistics include the number of samples carrying the variant, the allele fraction distribution across samples, and the total depth at the position.

The output of this process is a file that lists each variant position with its frequency across the PoN samples. This file is used by somatic variant callers to filter candidate variants. Variants that appear in the PoN at a frequency above a threshold are flagged as likely artifacts or germline polymorphisms and are removed from the somatic call set.

Artifact Filtering Thresholds

The threshold for removing variants based on the PoN depends on the application and the tolerance for false positives versus false negatives. A common approach is to remove any variant that appears in two or more PoN samples, regardless of allele fraction. This threshold captures recurrent artifacts while preserving variants that appear in only one normal sample, which are more likely to be rare germline variants or true somatic mutations.

For applications requiring high sensitivity, such as detecting minimal residual disease or low-allele-fraction variants, a more stringent threshold may be appropriate. The review of somatic callers used outside cancer contexts found that panel-of-normal filtering markedly improves specificity for low-allele-fraction variants, but the optimal threshold depends on the specific application and the acceptable trade-off between sensitivity and specificity.

Population Frequency Filtering

In addition to the PoN, population frequency databases provide a complementary filtering strategy. Variants that are common in the general population are unlikely to be somatic mutations and can be filtered based on their frequency in databases such as those available through NCBI. The study of germline variants falsely reported as somatic found that 33.8 percent of germline single nucleotide variants within a gene panel were rare and not found in variant information domains, meaning they would escape population frequency filtering. This finding highlights the limitation of population databases and the need for a PoN to capture rare germline variants that are specific to your patient population.

Integrating the Panel of Normals into Somatic Variant Calling

Caller-Specific Integration

The method for integrating the PoN into somatic variant calling depends on the variant caller being used. Most somatic callers accept a PoN file as input and use it to filter candidate variants during the calling process. The caller compares each candidate variant in the tumor sample against the PoN and applies the configured filtering threshold.

Some callers use the PoN to estimate the local error rate at each position, which improves the statistical model for distinguishing true variants from artifacts. This approach is more sophisticated than simple filtering because it accounts for the depth and allele fraction of the candidate variant relative to the expected error distribution from the PoN.

Post-Calling Filtering

When the variant caller does not support direct PoN integration, or when additional filtering is needed, the PoN can be applied as a post-calling filter. In this approach, the somatic variant caller produces an unfiltered list of candidate variants, and a separate filtering step removes variants that match the PoN. This approach provides more control over the filtering thresholds and allows for manual review of filtered variants.

Post-calling filtering is also useful for validating the PoN. By comparing the variant calls with and without PoN filtering, you can assess how many variants are removed and whether any known true variants are lost. This validation step is essential before deploying the PoN in a clinical workflow.

Machine Learning Post-Processing

Emerging approaches use machine learning classifiers to distinguish true somatic variants from artifacts. These classifiers integrate multiple features, including the PoN frequency, allele fraction, read position, base quality, and sequence context. The review of somatic callers found that machine-learning post-processing markedly improves specificity when combined with panel-of-normal filtering. The Bioconductor project provides packages for building and evaluating such classifiers, although these approaches require careful validation to avoid overfitting.

Validation of the Panel of Normals

Positive Control Samples

The most important validation step is testing the PoN on samples with known somatic variants. Commercial reference standards with characterized variants are available for this purpose. The benchmarking study of whole exome sequencing used four commercial reference standards and found that variant filtering was the main cause of divergent calls across centers. Testing your PoN on these standards will reveal whether the filtering thresholds are too aggressive or too permissive.

A positive control sample should contain a mix of true somatic variants at various allele fractions, including low-allele-fraction variants near the detection limit. The PoN should remove artifacts while preserving these true variants. The validation should assess both sensitivity, the proportion of true variants retained, and specificity, the proportion of artifacts removed.

Negative Control Samples

Negative controls, such as normal samples processed through the somatic variant calling workflow, should produce few or no somatic variant calls after PoN filtering. A high number of calls in negative controls indicates that the PoN is not capturing the artifacts in your workflow. This situation can arise when the PoN samples are not representative of the somatic samples, such as when the PoN is built from blood DNA but the somatic samples are from formalin-fixed paraffin-embedded tissue.

Formalin fixation introduces characteristic artifacts, including C to T transitions and deamination-induced errors. The benchmarking study of whole exome sequencing used formalin-fixed paraffin-embedded tissue specimens and found that wet-lab differences contributed to variability in copy-number alterations and complex biomarkers. If your somatic samples are from formalin-fixed tissue, the PoN should include normal samples processed through the same formalin fixation and DNA extraction protocols.

Cross-Validation

Cross-validation involves splitting the PoN samples into training and test sets. The PoN is built from the training set and evaluated on the test set. This approach provides an unbiased estimate of the PoN's performance because the test samples were not used to build the panel. Cross-validation is particularly useful for determining the optimal sample size and filtering thresholds.

A practical cross-validation approach is to build PoNs from increasing numbers of samples, such as 10, 20, 30, and 40 samples, and evaluate the performance of each on a held-out set of normal samples. The performance should improve with increasing sample size until a plateau is reached, indicating the minimum sample size needed for your workflow.

Records and Measurements for Panel of Normals Maintenance

Documentation Requirements

A PoN is not a static resource. Sequencing platforms change, reagent lots vary, and new artifact patterns emerge over time. Maintaining a PoN requires documentation of the samples, protocols, and software versions used to build the panel. This documentation should include the alignment software and version, the reference genome version, the variant caller and version, the preprocessing steps, and the filtering thresholds.

For clinical applications, the documentation should also include the ethical approvals for the normal samples, the de-identification process, and the quality metrics for each sample. The EMBL-EBI Training resources provide guidance on documenting bioinformatics workflows for reproducibility and regulatory compliance.

Quality Metrics to Track

Each normal sample in the PoN should have quality metrics recorded, including mean coverage, percentage of target bases covered at minimum depth, duplication rate, contamination estimate, and transition to transversion ratio. These metrics should be monitored over time to detect changes in the sequencing workflow that might affect the PoN's performance.

The PoN itself should have summary metrics, including the number of variant positions, the distribution of variant frequencies across samples, and the number of variants removed from a standard somatic sample. Changes in these metrics over time can indicate that the PoN is becoming outdated or that the sequencing workflow has changed.

Version Control and Updates

The PoN should be versioned, and the version used for each somatic variant calling run should be recorded. When the sequencing workflow changes, such as a change in library preparation kit or sequencing instrument, the PoN should be updated to reflect the new artifact profile. The nf-core documentation emphasizes the importance of version control for reproducible workflows, and this principle applies to PoN maintenance.

A practical approach is to rebuild the PoN quarterly or whenever a significant workflow change occurs. The new PoN should be validated against the previous PoN using the same positive and negative control samples. If the new PoN performs comparably or better, it can replace the previous version.

Common Failure Patterns in Panel of Normals Construction

Insufficient Sample Size

The most common failure is building a PoN from too few samples. A PoN built from fewer than 10 samples will miss many recurrent artifacts, particularly those that occur at low frequency. The result is an excess of false positive somatic calls that are not filtered by the PoN. This failure is often detected during validation when negative control samples produce many somatic calls.

The solution is to increase the number of normal samples in the PoN. The required sample size depends on the artifact frequency in your workflow, which can be estimated from the validation results. If the PoN is not adequately filtering artifacts, adding more samples is the first corrective action.

Non-Representative Normal Samples

A second common failure is using normal samples that are not representative of the somatic samples. This situation arises when the PoN is built from samples processed with a different library preparation kit, sequencing instrument, or bioinformatics pipeline than the somatic samples. The PoN will not capture the artifacts present in the somatic samples, leading to false positive calls.

The solution is to ensure that the PoN samples are processed through the identical workflow as the somatic samples. This requirement can be challenging when the workflow changes over time, which is why the PoN must be updated whenever the workflow changes.

Overly Aggressive Filtering

A third failure pattern is filtering too aggressively, removing true somatic variants along with artifacts. This situation arises when the filtering threshold is set too low, such as removing any variant that appears in a single PoN sample. Rare germline variants that happen to be present in one PoN sample will be removed, potentially including clinically relevant variants.

The solution is to set the filtering threshold based on the validation results. The threshold should be adjusted to maximize the retention of true variants in positive control samples while minimizing the retention of artifacts in negative control samples. The optimal threshold balances sensitivity and specificity for your specific application.

Failure to Update the Panel

A fourth failure pattern is continuing to use an outdated PoN after the sequencing workflow has changed. Changes in reagent lots, instrument maintenance, or bioinformatics software can alter the artifact profile. An outdated PoN will not capture new artifacts, leading to an increase in false positive calls over time.

The solution is to monitor the performance of the PoN on an ongoing basis and to rebuild the panel when the workflow changes. Tracking quality metrics over time can help detect changes in the artifact profile before they become problematic.

Limitations of Panel of Normals Filtering

Rare Artifacts and Sample-Specific Errors

A PoN cannot capture artifacts that occur in only one sample or that are specific to a particular sample. Sample-specific artifacts can arise from contamination, degradation, or unusual library preparation conditions. These artifacts will not be filtered by the PoN and require other filtering strategies, such as read-level filters or manual review.

The review of somatic callers used outside cancer contexts noted that unique molecular identifiers markedly improve specificity when combined with panel-of-normal filtering. Unique molecular identifiers tag individual DNA molecules before amplification, allowing the removal of PCR duplicates and the identification of true variants that appear in multiple independent molecules. This approach is complementary to PoN filtering and is particularly valuable for low-allele-fraction variants.

Population-Specific Germline Variants

A PoN built from one population may not be representative of another population. Germline variants that are rare in the PoN population may be common in the patient population, leading to false somatic calls. This limitation is particularly relevant for clinical applications serving diverse patient populations.

The study of germline variants falsely reported as somatic found that 33.8 percent of germline single nucleotide variants within a gene panel were rare and not found in variant information domains. This finding suggests that population databases are incomplete for rare variants and that a PoN built from a different population may not capture the germline variants relevant to your patients.

Copy Number Alterations and Structural Variants

A PoN is primarily designed for filtering single nucleotide variants and small indels. Copy number alterations and structural variants require different filtering approaches. The benchmarking study of whole exome sequencing found that copy-number alteration calls were concordant for 82 percent of genomic regions, and this variability was attributed mainly to wet-lab differences instead of bioinformatics processing.

For applications that require copy number alteration detection, the PoN should be supplemented with a set of normal samples processed through the copy number calling workflow. The normal samples provide a baseline for read depth and allelic balance that is used to identify copy number changes in the somatic samples.

Safety and Regulatory Context

Clinical Diagnostic Applications

When the PoN is used in a clinical diagnostic workflow, the construction and validation of the PoN must meet the requirements of the relevant regulatory framework. The benchmarking study of whole exome sequencing in the German network for personalized medicine recommended continuous optimization of bioinformatic workflows and participation in round robin tests. These recommendations apply to PoN construction and maintenance.

The validation of the PoN should be documented, including the sample selection criteria, the variant calling workflow, the filtering thresholds, and the performance on positive and negative controls. This documentation should be available for audit and should be updated when the PoN is revised.

Research Applications

For research applications, the PoN construction should follow the same rigorous standards as clinical applications, even when regulatory requirements do not apply. The Galaxy Training Network and The Carpentries provide training on reproducible research practices that are relevant to PoN construction.

The Carpentries lessons on data management and reproducible analysis are particularly relevant for documenting the PoN construction process and ensuring that the results can be reproduced by other researchers.

Ethical Considerations for Normal Samples

The normal samples used to build the PoN are human biological specimens and must be collected under appropriate ethical approvals. The samples should be de-identified, and the consent process should cover the use of the samples for assay development and quality control. The EMBL-EBI Training resources include modules on the ethical and legal considerations of using human genomic data.

For clinical applications, the use of normal samples for PoN construction should be reviewed by the institutional review board or ethics committee. The review should consider the privacy risks of including the samples in a shared panel and the measures taken to protect participant confidentiality.

Professional Escalation Criteria

When to Seek Additional Expertise

Building a PoN requires expertise in bioinformatics, sequencing technology, and statistical analysis. If your laboratory does not have this expertise in-house, consider consulting with a bioinformatics core facility or a clinical genomics specialist. The EMBL-EBI Training resources can help identify the skills needed and the training available.

Specific situations that warrant escalation include persistent false positive calls that are not resolved by PoN filtering, poor concordance with reference standards, and the need to detect variants at very low allele fractions. These situations may require advanced approaches such as unique molecular identifiers, machine learning classifiers, or long-read sequencing.

When to Rebuild the Panel

The PoN should be rebuilt when the sequencing workflow changes, when the performance on validation samples degrades, or when the panel has been in use for an extended period without updates. The specific triggers for rebuilding include changes in library preparation kits, sequencing instruments, alignment software, variant callers, or reference genome versions.

The decision to rebuild should be based on data instead of a fixed schedule. Monitoring the PoN's performance on positive and negative controls provides the evidence needed to determine when a rebuild is necessary. If the performance metrics remain stable, the PoN can continue to be used. If the metrics degrade, a rebuild is indicated.

When to Consider Alternative Approaches

A PoN is one component of a somatic variant calling workflow, and it has limitations that cannot be overcome by adding more samples. When the limitations of PoN filtering become the bottleneck for your application, consider alternative or complementary approaches. Unique molecular identifiers provide a fundamentally different approach to artifact removal that is complementary to PoN filtering. Long-read sequencing can resolve variants in repetitive regions that are difficult to align with short reads. Machine learning classifiers can integrate multiple features to distinguish true variants from artifacts.

The choice of approach depends on your specific application, the allele fractions you need to detect, and the resources available. The review of somatic callers used outside cancer contexts provides a synthesis of the available approaches and their performance characteristics for various applications.

Frequently Asked Questions

What is the minimum number of normal samples needed for a panel of normals?

The minimum number depends on your sequencing workflow and the allele fractions you need to detect. For targeted gene panels, 20 to 30 normal samples provide adequate coverage of recurrent artifacts. For whole exome sequencing, 40 or more samples are recommended. The statistical rationale is that an artifact present at 5 percent frequency has a 64 percent chance of being detected in a PoN of 20 samples and an 87 percent chance in a PoN of 40 samples. You can determine the optimal sample size for your workflow by building PoNs from increasing numbers of samples and evaluating the performance on held-out validation samples.

Can I use public sequencing data to build a panel of normals?

Public sequencing data from NCBI can be useful for understanding general artifact patterns, but it should not substitute for a PoN built from your own workflow. Sequencing artifacts are specific to the library preparation kit, sequencing instrument, and bioinformatics pipeline. A PoN built from public data generated with different protocols will not capture the artifacts in your somatic samples and may introduce incorrect assumptions about which positions are reliable.

How often should I update my panel of normals?

The PoN should be updated whenever the sequencing workflow changes, including changes in library preparation kits, sequencing instruments, reagent lots, alignment software, variant callers, or reference genome versions. You should also monitor the PoN's performance on positive and negative controls and rebuild the panel if the performance degrades. A practical approach is to rebuild the PoN quarterly or whenever a significant workflow change occurs.

What is the difference between a panel of normals and population frequency filtering?

A PoN captures the artifacts and germline variants specific to your sequencing workflow and patient population. Population frequency filtering uses databases of variants found in the general population to remove common polymorphisms. The two approaches are complementary. A study of germline variants falsely reported as somatic found that 33.8 percent of germline single nucleotide variants within a gene panel were rare and not found in variant information domains, meaning they would escape population frequency filtering but could be captured by a PoN built from a representative population.

How do I validate that my panel of normals is working correctly?

Validation involves testing the PoN on positive and negative control samples. Positive controls, such as commercial reference standards with characterized variants, should retain their true somatic variants after PoN filtering. Negative controls, such as normal samples processed through the somatic variant calling workflow, should produce few or no somatic variant calls after PoN filtering. Cross-validation, where the PoN is built from a training set and evaluated on a held-out test set, provides an unbiased estimate of performance.

Can a panel of normals be used for copy number alteration detection?

A PoN is primarily designed for filtering single nucleotide variants and small indels. Copy number alterations require a different approach, typically using a set of normal samples processed through a copy number calling workflow to establish a baseline for read depth and allelic balance. The benchmarking study of whole exome sequencing found that copy-number alteration calls were concordant for 82 percent of genomic regions, with variability attributed mainly to wet-lab differences.

What should I do if my panel of normals is not filtering enough artifacts?

If negative control samples produce many somatic calls after PoN filtering, the PoN is not capturing the artifacts in your workflow. The first corrective action is to increase the number of normal samples in the PoN. You should also verify that the PoN samples are processed through the identical workflow as the somatic samples, including the same library preparation kit, sequencing instrument, and bioinformatics pipeline. If the problem persists, consider complementary approaches such as unique molecular identifiers or machine learning classifiers.

How does a panel of normals perform for low-allele-fraction variant detection?

A review of somatic callers used outside cancer contexts found that panel-of-normal filtering markedly improves specificity for detecting low-allele-fraction variants, including post-zygotic variants at 1 to 3 percent allele fraction, mitochondrial heteroplasmy as low as 0.4 to 0.5 percent, and minority viral variants around 5 percent. The performance depends on the sample size and depth of the PoN, as well as the filtering thresholds. Unique molecular identifiers and machine learning post-processing can further improve specificity when combined with PoN filtering.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.