Evaluating Variant Filtering Performance: Metrics and Methods to Assess Precision and Recall
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Variant filtering performance is critically assessed using precision (fraction of filtered variants that are true positives) and recall (fraction of true variants that are identified). A stringent filter increases precision but decreases recall, necessitating a balance tailored to the research question, such as prioritizing precision in clinical diagnostics to avoid false positive diagnoses.
- Gold standard datasets, like those from the Genome in a Bottle (GIAB) consortium for human samples, are essential for benchmarking filtering strategies by providing a high-confidence truth set against which filtered variant calls are compared. For non-human organisms, custom truth sets generated via long-read sequencing or consensus from multiple callers are necessary.
- The choice of variant caller significantly impacts filtering performance, as different callers have distinct error profiles and assumptions; therefore, filtering thresholds must be evaluated in the context of the specific caller used, rather than assuming transferability from published studies.
- Machine learning filters, which consider interactions between multiple metrics, often outperform traditional hard-cutoff filters by capturing more nuanced variant characteristics, but require a well-defined truth set for training.
- A systematic workflow for evaluation involves defining the specific question, obtaining or generating a suitable truth set, running the pipeline and filters, comparing results to calculate metrics (precision, recall, F1 score), and meticulously documenting all steps and limitations to ensure reproducibility and detect pipeline drift.
- Overfitting to the truth set, using an inappropriate truth set, confusing precision with accuracy, or failing to separate variant types (e.g., SNVs vs. indels) are common pitfalls that can lead to misleading performance evaluations and unreliable downstream analyses.
Variant filtering is the step in a bioinformatics pipeline where candidate genetic variants are separated into those that are likely real and those that are likely artifacts. Researchers who skip a formal evaluation of this step risk publishing false associations, missing genuine variants, or wasting resources on validation experiments. This article provides concrete metrics and methods for measuring how well a variant filtering strategy performs, with a focus on precision and recall. The target reader is a biology student, researcher, or laboratory professional who has generated variant calls and needs to know whether the filtering decisions made are defensible. The practical outcome is a workflow for benchmarking a filtering strategy against known truth sets, interpreting the resulting metrics, and documenting the performance for reporting purposes.
The Core Problem: Filtering Without Measurement
Most variant calling pipelines produce a raw set of candidate variants that includes both true biological variants and technical artifacts. The artifacts arise from sequencing errors, misalignment, paralogous sequences, and other sources. Filtering is the process of applying thresholds to remove suspected artifacts. Common filters include depth of coverage, mapping quality, allele balance, strand bias, and proximity to indels. The challenge is that every filter threshold involves a tradeoff. A stringent filter removes more artifacts but also removes true variants. A lenient filter retains more true variants but also retains more artifacts.
The central problem addressed here is that many researchers apply filters without measuring the consequences. They use default thresholds from a pipeline or thresholds from a published study without checking whether those thresholds are appropriate for their data. This is a measurable problem. The performance of a filtering strategy can be quantified using precision and recall, and those quantities can be estimated using gold standard datasets or internal consistency checks. Without such measurement, the researcher cannot know whether the final variant set is dominated by false positives or missing a substantial fraction of true variants.
The consequences of unmeasured filtering are concrete. In a clinical context, a false positive variant could lead to an incorrect diagnosis. In a population genetics study, false positives inflate diversity estimates. In a plant breeding program, a missed true variant could mean losing a trait-associated marker. The methods described in this article allow a researcher to quantify these risks and make filtering decisions based on evidence instead of habit.
Metrics for Filtering Performance
Precision and Recall Defined in the Variant Calling Context
Precision and recall are the two primary metrics for evaluating variant filtering performance. Precision answers the question: of the variants that passed the filter, what fraction is truly present in the sample? Recall answers the question: of the variants truly present in the sample, what fraction passed the filter? Both are needed because a filter can achieve high precision by being extremely stringent, but that stringency may come at the cost of low recall.
Precision is calculated as the number of true positive calls divided by the total number of calls that passed the filter. The total includes both true positives and false positives. Recall is calculated as the number of true positive calls divided by the total number of true variants in the sample. The total includes both true positives and false negatives. A perfect filter would have both precision and recall equal to 1.0, meaning every call is real and every real variant is called.
In practice, precision and recall are inversely related for most filtering strategies. Increasing the stringency of a filter typically increases precision while decreasing recall. This tradeoff is well documented in the variant calling literature. A benchmarking study of short-read variant calling in Mycobacterium tuberculosis found that tuning the mapping quality threshold produced a maximum recall of 89.0 percent with a minimum precision of 98.5 percent across the parameters evaluated. The approach that maximized recall while maintaining precision below 99 percent was a mapping quality threshold of at least 40, which gave a recall of 85.8 percent and a precision of 99.1 percent. A more conservative approach that masked repetitive sequence content increased precision to 99.6 percent but reduced recall to 70.2 percent. These numbers illustrate the direct tradeoff between the two metrics.
The F1 Score as a Combined Measure
The F1 score is the harmonic mean of precision and recall. It provides a single number that summarizes filtering performance when both false positives and false negatives are considered undesirable. The F1 score is calculated as two times the product of precision and recall divided by their sum. A high F1 score indicates that the filter achieves a good balance between the two metrics.
The F1 score is useful for comparing different filtering strategies when the researcher does not have a strong preference for either precision or recall. However, the appropriate balance depends on the research question. A clinical diagnostic test may prioritize precision because a false positive result has serious consequences. A discovery study may prioritize recall because missing a rare variant could mean missing the variant of interest. The F1 score should be interpreted in the context of the research goals.
Additional Metrics for Specific Filtering Decisions
Beyond precision and recall, several other metrics are useful for evaluating specific aspects of filtering performance. The false discovery rate is the complement of precision and represents the proportion of passing calls that are false positives. The false negative rate is the complement of recall and represents the proportion of true variants that were filtered out. The Matthews correlation coefficient is a more balanced measure that accounts for all four categories of calls: true positives, false positives, true negatives, and false negatives. It is particularly useful when the number of true negatives is large, which is common in variant calling because most genomic positions do not contain variants.
The choice of metric should match the decision being made. If the question is whether a filter threshold should be tightened, the researcher should examine how precision and recall change across a range of thresholds. If the question is whether one filtering strategy is better than another, the F1 score or Matthews correlation coefficient provides a summary comparison. If the question is whether the final variant set is suitable for a specific downstream analysis, the researcher should consider the metric that is most relevant to that analysis.
Gold Standard Datasets for Benchmarking
The Genome in a Bottle Consortium
The most widely used gold standard datasets for benchmarking human variant calling come from the Genome in a Bottle (GIAB) consortium. GIAB provides highly curated variant calls for several human reference samples. These calls are generated by integrating data from multiple sequencing technologies, multiple read lengths, and multiple variant calling algorithms. The resulting variant sets are considered the best available approximation of the true variants in those samples.
The GIAB datasets are hosted by the National Center for Biotechnology Information (NCBI), which provides official descriptions of its databases and search systems. Researchers can download the GIAB variant calls and use them as a truth set for evaluating their own variant calling and filtering pipelines. The standard approach is to run the pipeline on the GIAB sample's raw sequencing data, apply the filtering strategy, and then compare the filtered calls against the GIAB truth set. This comparison yields the counts needed to calculate precision and recall.
The limitation of GIAB is that it covers only a small number of human samples. The truth set is not available for every population or every genomic region. Researchers working on non-human organisms cannot use GIAB directly. They must either generate their own truth set or use alternative benchmarking approaches.
Organism-Specific Benchmarking Studies
For non-human organisms, benchmarking studies published in the literature provide guidance on expected performance. A study of variant calling in inbred mouse strains used simulated genomes to evaluate multiple variant calling tools. The study found a tradeoff between recall and precision across tools and showed that an optimal call set could be obtained by taking the intersection of variants reported by multiple callers. The study also identified empirical filters that improved performance for both rare variant discovery and strain polymorphism identification. These findings are relevant to any researcher working with inbred organisms, where the assumption of heterozygosity built into many tools does not apply.
A benchmarking study of plant variant calling evaluated the entire pipeline from alignment through variant calling, filtering, and imputation. The study found that the choice of variant caller affected precision and recall differently depending on the levels of diversity, sequence coverage, and genome complexity. The study also found that a machine-learning-based filtering strategy outperformed a traditional hard-cutoff strategy, producing a higher number of true positive variants and fewer false positive variants. These results demonstrate that filtering performance is not universal and must be evaluated in the context of the organism and data type.
A study of Mycobacterium tuberculosis variant calling used 36 clinical isolates that were sequenced with both Illumina short reads and PacBio long reads. The long-read data served as the truth set for evaluating the short-read variant calls. This study design is a practical approach for organisms where a gold standard dataset does not exist. The researcher generates high-confidence calls using a more accurate technology and then evaluates the performance of the routine pipeline against those calls.
Generating a Custom Truth Set
When no gold standard dataset exists for the organism or sample type, the researcher can generate a custom truth set. The most reliable approach is to sequence the same sample with a different technology, preferably one with longer reads or higher accuracy. The long-read data can be used to generate a high-confidence variant set that serves as the truth set for evaluating the short-read pipeline. This approach was used in the Mycobacterium tuberculosis study described above.
An alternative approach is to use a combination of multiple variant callers on the same data. Variants that are called by multiple independent tools are more likely to be true than variants called by only one tool. The intersection of calls from multiple tools can serve as a provisional truth set. However, this approach has a limitation: it is biased toward variants that are easy to call and may miss variants that all tools fail to detect. The inbred mouse study found that the intersection approach produced an optimal call set, but this finding was based on simulated data where the true variants were known.
A third approach is to validate a subset of variants experimentally. PCR amplification and Sanger sequencing of a random sample of passing variants provides an estimate of precision. The fraction of validated variants among those tested is an estimate of precision. This approach does not provide an estimate of recall because it does not reveal how many true variants were missed. However, it is a practical option when no truth set is available and the researcher needs at least a precision estimate.
The Transition from Raw Calls to Filtered Variants
Understanding the Raw Call Set
Before evaluating filtering performance, the researcher must understand the nature of the raw call set. The raw call set is the output of the variant caller before any filtering is applied. It typically contains a large number of candidate variants, many of which are artifacts. The number of raw calls depends on the variant caller, the sequencing depth, the genome complexity, and the diversity of the sample relative to the reference genome.
The raw call set is not a random collection of variants. It is biased toward variants in regions that are easy to align and away from variants in repetitive or low-complexity regions. The Mycobacterium tuberculosis study found that genetic divergence from the reference genome, repetitive sequences, and sequencing bias reduce the performance of variant calling using short-read alignment. The study also found that 68 percent of genomic positions typically excluded for Mycobacterium tuberculosis were accurately called using Illumina whole-genome sequencing, including 52 of 168 PE/PPE genes. This finding demonstrates that some excluded regions can be recovered with appropriate filtering.
The researcher should examine the distribution of quality metrics in the raw call set before applying filters. This examination provides a sense of where the filtering thresholds should be set. For example, if most true variants have a depth of coverage above 30 but many artifacts have a depth below 10, a depth filter at 10 or 20 may be appropriate. The examination should include the metrics that are available in the variant call format file, such as quality score, depth, mapping quality, and allele balance.
The Role of the Variant Caller in Filtering Performance
The choice of variant caller has a substantial effect on filtering performance. Different callers have different error profiles, and the filters that work well for one caller may not work well for another. The plant benchmarking study found that the choice of variant caller affected precision and recall differently depending on the levels of diversity, sequence coverage, and genome complexity. The study compared GATK HaplotypeCaller and SAMtools mpileup and found that neither was universally better.
The inbred mouse study found that many variant calling tools assume outbred genomes and implicit heterozygosity, conditions that do not apply to inbred laboratory models. This finding is a reminder that the assumptions built into a tool may not match the biology of the sample. The researcher should understand the assumptions of the chosen caller and consider whether they are appropriate for the data.
The practical implication is that filtering thresholds should be evaluated in the context of the specific caller being used. A threshold that works well for GATK HaplotypeCaller output may not work well for SAMtools mpileup output. The researcher should not assume that a filtering strategy from a published study will transfer directly to their pipeline.
Hard Filters versus Machine Learning Filters
Traditional filtering uses hard thresholds on individual metrics. A variant passes the filter if its depth is above a threshold, its mapping quality is above a threshold, and its allele balance is within a range. Hard filters are simple to implement and interpret. However, they do not capture interactions between metrics. A variant with slightly low depth but very high mapping quality might be real, but a hard filter would remove it.
Machine learning filters use a model that considers multiple metrics simultaneously. The model is trained on a dataset where the true status of each variant is known. The plant benchmarking study found that a machine-learning-based filtering strategy outperformed the traditional hard-cutoff strategy, resulting in a higher number of true positive variants and fewer false positive variants. This result suggests that the interactions between metrics contain information that hard filters do not capture.
The limitation of machine learning filters is that they require a training dataset with known true variants. This requirement is the same as the requirement for benchmarking. If a gold standard dataset is available, the researcher can train a machine learning filter and evaluate its performance. If no gold standard dataset is available, the researcher must rely on hard filters or generate a truth set first.
At a Glance: Filtering Performance Metrics and Their Interpretation
| Metric | What It Measures | How to Calculate | Interpretation Guidance |
|---|---|---|---|
| Precision | Fraction of passing calls that are true variants | True positives divided by all passing calls | High precision means few false positives in the final variant set |
| Recall | Fraction of true variants that pass the filter | True positives divided by all true variants | High recall means few true variants were filtered out |
| F1 Score | Combined measure of precision and recall | Two times precision times recall divided by their sum | Useful for comparing strategies when both errors matter |
| False Discovery Rate | Fraction of passing calls that are artifacts | One minus precision | Directly estimates the proportion of bad calls in the final set |
| False Negative Rate | Fraction of true variants that were filtered out | One minus recall | Directly estimates the proportion of missed true variants |
| Matthews Correlation Coefficient | Balanced measure accounting for all call categories | Calculated from true positives, false positives, true negatives, and false negatives | Useful when true negatives are numerous, as in most genomic comparisons |
Practical Workflow for Evaluating Filtering Performance
Step 1: Define the Evaluation Question
The first step is to define what the evaluation is trying to determine. The evaluation question should be specific. For example, the researcher might ask whether a depth filter at 10 or 20 produces a better final variant set. Or the researcher might ask whether the filtering strategy achieves a precision of at least 99 percent while maintaining a recall of at least 90 percent. Defining the question in advance prevents the researcher from choosing the metric that makes the filtering strategy look best after the fact.
The evaluation question should also specify the variant types of interest. Single nucleotide variants and small indels have different error profiles and may require different filters. The researcher should decide whether the evaluation covers both types or focuses on one. The evaluation should also specify whether the focus is on all variants or on a subset, such as rare variants or variants in coding regions.
Step 2: Obtain or Generate a Truth Set
The second step is to obtain or generate a truth set. For human samples, the GIAB datasets are the standard choice. For other organisms, the researcher may need to generate a custom truth set using long-read sequencing or multi-caller consensus. The truth set should be as accurate as possible because errors in the truth set propagate to errors in the precision and recall estimates.
The truth set should also be matched to the sample being evaluated. If the pipeline is being run on a sample that is not the same as the truth set sample, the evaluation is indirect. The researcher is assuming that the filtering performance observed on the truth set sample will transfer to the actual sample. This assumption is reasonable when the samples are similar in diversity and sequencing quality but becomes less reliable as the samples diverge.
Step 3: Run the Pipeline and Apply the Filter
The third step is to run the variant calling pipeline on the raw sequencing data and apply the filtering strategy being evaluated. The pipeline should be run exactly as it will be run for the actual samples. Any differences between the evaluation run and the production run will reduce the validity of the evaluation.
The filtering strategy should be applied in a way that allows the researcher to see the effect of each filter. One approach is to apply filters incrementally and calculate precision and recall after each addition. This approach reveals which filters are improving performance and which are removing true variants without removing many artifacts. Another approach is to apply the complete filtering strategy and calculate the final precision and recall.
Step 4: Compare Filtered Calls to the Truth Set
The fourth step is to compare the filtered calls to the truth set. The comparison requires matching variants between the two sets. Variants are typically matched by genomic position and allele. The comparison should account for small differences in position representation between callers, particularly for indels. Tools for variant comparison are available in the Bioconductor ecosystem, which provides official documentation for genomic analysis packages.
The comparison produces four counts: true positives, false positives, false negatives, and true negatives. True positives are variants in both the filtered set and the truth set. False positives are variants in the filtered set but not in the truth set. False negatives are variants in the truth set but not in the filtered set. True negatives are positions where neither set has a variant. True negatives are usually not counted explicitly because the number of non-variant positions is enormous.
Step 5: Calculate and Interpret the Metrics
The fifth step is to calculate precision, recall, and any additional metrics from the four counts. The calculations are straightforward. Precision is true positives divided by the sum of true positives and false positives. Recall is true positives divided by the sum of true positives and false negatives.
The interpretation of the metrics should be tied to the evaluation question defined in step 1. If the question was whether a depth filter at 10 or 20 is better, the researcher compares the precision and recall for the two thresholds. If the question was whether the filtering strategy meets a performance target, the researcher checks whether the calculated metrics meet the target.
Step 6: Document the Results
The sixth step is to document the evaluation results. The documentation should include the truth set used, the pipeline version, the filter thresholds, the four counts, and the calculated metrics. This documentation is essential for reproducibility. A researcher who cannot reproduce the evaluation cannot verify the filtering performance.
The documentation should also include the limitations of the evaluation. If the truth set was generated from a different sample, that limitation should be noted. If the truth set covers only a subset of the genome, that limitation should be noted. The documentation should be detailed enough that another researcher could repeat the evaluation and obtain the same results.
Records and Measurements for Filtering Evaluation
What to Record During the Evaluation
The evaluation produces a set of records that should be preserved. The most important records are the raw sequencing data identifiers, the pipeline version and parameters, the truth set version and source, and the comparison results. These records allow the evaluation to be repeated and verified.
The pipeline version is critical because variant calling tools change rapidly. A filtering strategy that performs well with one version of a tool may perform differently with the next version. The researcher should record the exact version of every tool in the pipeline, including the aligner, the variant caller, and any filtering tools.
The truth set version is equally critical. Gold standard datasets are updated as new data become available. A comparison against version 3 of a truth set may give different results than a comparison against version 4. The researcher should record the exact version and the date it was downloaded.
Measurements to Track Across Samples
When the same filtering strategy is applied to multiple samples, the researcher should track the precision and recall estimates across samples. Variation in these estimates across samples indicates that the filtering strategy is not performing consistently. This variation could be due to differences in sequencing depth, sample diversity, or library quality.
The researcher should also track the number of variants that pass the filter for each sample. A sudden increase or decrease in the number of passing variants may indicate a problem with the sequencing or the pipeline. The number of passing variants is not a direct measure of precision or recall, but it is a useful monitoring metric.
Using Records to Detect Pipeline Drift
Pipeline drift occurs when the performance of a pipeline changes over time without an intentional change to the pipeline. Drift can be caused by updates to tools, changes in sequencing chemistry, or changes in the reference genome. Regular evaluation against a truth set is the most reliable way to detect drift.
The researcher should establish a schedule for re-evaluating the filtering strategy. The schedule depends on how often the pipeline or the sequencing technology changes. A reasonable approach is to re-evaluate whenever a tool is updated, whenever the sequencing platform or chemistry changes, and at regular intervals such as annually. The records from previous evaluations provide the baseline for detecting drift.
Common Failure Patterns in Filtering Evaluation
Overfitting to the Truth Set
A common failure pattern is overfitting the filtering strategy to the truth set. The researcher tunes the filters until the precision and recall on the truth set sample are excellent. However, the tuned filters may not perform well on other samples. The filters may be exploiting idiosyncrasies of the truth set sample instead of general properties of true variants.
The risk of overfitting is highest when the truth set is small or when the researcher evaluates many filter combinations. Each additional filter combination that is tested increases the chance of finding a combination that performs well by chance. The researcher should evaluate the final filtering strategy on a sample that was not used for tuning. This evaluation provides a more realistic estimate of performance.
Using an Inappropriate Truth Set
Another common failure pattern is using a truth set that does not match the sample being evaluated. The truth set may be from a different population, a different tissue type, or a different sequencing platform. The precision and recall estimates from such an evaluation may not reflect the performance on the actual samples.
The researcher should consider whether the truth set is representative of the samples in the study. If the study includes samples from multiple populations, the filtering strategy should be evaluated on truth sets from each population. If the study uses a specific sequencing platform, the truth set should be generated from data on that platform.
Ignoring the Confidence of Truth Set Calls
Gold standard datasets often include confidence scores for each variant. The confidence reflects the evidence supporting the variant. High-confidence variants are supported by multiple technologies and multiple callers. Low-confidence variants have weaker support.
The researcher should decide whether to include low-confidence truth set variants in the evaluation. Including them provides a more complete truth set but introduces noise because some low-confidence variants may be false. Excluding them provides a cleaner truth set but may bias the evaluation toward easy-to-call variants. The decision should be documented and justified.
Confusing Precision with Accuracy
A subtle failure pattern is confusing precision with accuracy. Accuracy is the fraction of all calls that are correct, including both positive and negative calls. In variant calling, the number of true negatives is enormous, so accuracy is almost always very high even when precision is poor. A filtering strategy that removes all variants would have high accuracy because almost all positions are non-variant. The researcher should focus on precision and recall instead of accuracy.
Failing to Separate Variant Types
Another failure pattern is evaluating filtering performance on all variants combined. Single nucleotide variants and indels have different error profiles. A filtering strategy may have high precision for single nucleotide variants but low precision for indels. Combining the two types obscures this difference.
The researcher should evaluate filtering performance separately for each variant type. This separation reveals which variant types are causing problems and allows the filters to be adjusted accordingly. The separation is particularly important for indels, which are generally harder to call than single nucleotide variants.
Limitations of Filtering Performance Evaluation
The Truth Set Is Not Perfect
The most fundamental limitation of filtering performance evaluation is that the truth set is not perfect. Even the best gold standard datasets contain errors. The Genome in a Bottle datasets are highly curated, but they are not guaranteed to be complete or correct. Errors in the truth set cause errors in the precision and recall estimates.
The researcher should consider the direction of the bias caused by truth set errors. If the truth set contains false variants, the evaluation will underestimate precision because some passing calls will be counted as false positives when they are actually true. If the truth set is missing true variants, the evaluation will underestimate recall because some filtered calls will be counted as false negatives when they are actually true.
Performance Does Not Transfer Perfectly Across Samples
Filtering performance measured on one sample does not transfer perfectly to other samples. The performance depends on the sequencing depth, the library quality, the sample diversity, and the genome complexity. A filtering strategy that performs well on a high-depth sample may perform poorly on a low-depth sample.
The researcher should evaluate the filtering strategy on samples that are representative of the range of data quality in the study. If the study includes samples with a wide range of depths, the evaluation should include samples at both the high and low ends of the range. This approach provides a more complete picture of the filtering performance.
The Evaluation Is Only as Good as the Comparison Method
The comparison between the filtered calls and the truth set requires a method for matching variants. Different comparison methods can give different results. Some methods require exact position and allele matches. Others allow for small differences in position. The choice of comparison method affects the precision and recall estimates.
The researcher should use a comparison method that is appropriate for the variant types being evaluated. For single nucleotide variants, exact position and allele matching is usually appropriate. For indels, the comparison is more complex because the same variant can be represented in multiple ways. The researcher should understand how the comparison method handles these complexities.
Relevant Welfare and Safety Context
Clinical and Diagnostic Applications
In clinical and diagnostic applications, the consequences of filtering errors are direct and serious. A false positive variant could lead to an incorrect diagnosis, unnecessary treatment, or inappropriate genetic counseling. A false negative variant could lead to a missed diagnosis and a missed opportunity for intervention. The precision and recall of the filtering strategy are therefore beyond technical metrics but patient safety metrics.
The researcher working on clinical samples should set performance targets that reflect the clinical context. The targets should be based on the consequences of each type of error. If a false positive has serious consequences, the precision target should be high. If a false negative has serious consequences, the recall target should be high. The targets should be documented and the filtering strategy should be evaluated against them.
Public Health Surveillance
In public health surveillance, variant calling is used to track the spread of pathogens and to identify drug resistance mutations. The Mycobacterium tuberculosis study noted that benchmarking results have broad implications for the inference of transmission in public health surveillance systems. A false positive variant could lead to an incorrect inference of transmission between cases. A false negative variant could lead to a missed transmission link.
The researcher working on pathogen surveillance should evaluate the filtering strategy in the context of the surveillance question. The evaluation should consider whether the precision and recall are adequate for the specific inferences being made. The evaluation should also consider the diversity of the pathogen population, which can affect variant calling performance.
Research Reproducibility
The evaluation of filtering performance is part of the broader requirement for research reproducibility. A study that reports variant calls without reporting the filtering performance cannot be fully reproduced or verified. The reader of the study cannot know whether the variant calls are reliable.
The researcher should report the filtering performance metrics in publications and in data releases. The report should include the truth set used, the pipeline version, the filter thresholds, and the precision and recall estimates. This reporting allows other researchers to assess the reliability of the variant calls and to compare the filtering strategy with their own.
Professional Escalation Criteria
When to Seek Expert Assistance
The evaluation of filtering performance can be technically challenging, particularly for researchers who are new to bioinformatics. The researcher should consider seeking expert assistance when the evaluation reveals problems that cannot be resolved with standard approaches. Examples include persistent low precision despite extensive filtering, persistent low recall despite lenient filtering, or inconsistent performance across samples.
Expert assistance is available from multiple sources. The Galaxy Training Network provides accessible workflow training and analysis tutorials that cover variant calling and filtering. The EMBL-EBI Training program offers bioinformatics learning pathways and data-resource training. The nf-core documentation describes community pipeline standards and usage, which can help the researcher understand how established pipelines handle filtering.
When to Question the Pipeline
The researcher should question the pipeline when the evaluation results are surprising. A filtering strategy that produces a precision below 90 percent or a recall below 80 percent may indicate a problem with the pipeline instead of with the filters. The problem could be in the alignment, the variant calling, or the reference genome.
The researcher should also question the pipeline when the evaluation results change substantially after a minor change to the pipeline. This instability suggests that the pipeline is sensitive to parameters that should not matter. The researcher should investigate the source of the instability before proceeding with the analysis.
When to Reconsider the Study Design
The evaluation results may indicate that the study design is not adequate for the research question. For example, if the sequencing depth is too low to achieve the required recall, the researcher may need to sequence the samples at higher depth. If the sample diversity is too high relative to the reference genome, the researcher may need to use a different reference or a different variant calling approach.
The researcher should treat the evaluation results as information about the study design, beyond about the filtering strategy. The evaluation may reveal that the study needs more sequencing, a different reference genome, or a different variant caller. These changes are better made before the full analysis than after.
Frequently Asked Questions
What is the difference between precision and recall in variant filtering?
Precision measures the fraction of variants that pass the filter and are actually true. Recall measures the fraction of true variants that pass the filter. A filter with high precision removes most artifacts but may also remove true variants. A filter with high recall retains most true variants but may also retain artifacts. Both metrics are needed to understand filtering performance because a filter can achieve high precision by being extremely stringent or high recall by being extremely lenient.
How do I obtain a gold standard dataset for my organism?
For human samples, the Genome in a Bottle datasets hosted by the National Center for Biotechnology Information are the standard choice. For non-human organisms, you may need to generate a custom truth set. One approach is to sequence the same sample with long-read technology and use the long-read calls as the truth set. Another approach is to use the intersection of calls from multiple variant callers as a provisional truth set. A third approach is to validate a subset of variants experimentally using PCR and Sanger sequencing.
What is a reasonable precision and recall target for variant filtering?
The appropriate targets depend on the research question and the consequences of each type of error. A clinical diagnostic application may require precision above 99 percent because false positives have serious consequences. A discovery study may accept lower precision to achieve higher recall because missing a rare variant is the greater risk. The Mycobacterium tuberculosis benchmarking study reported a maximum recall of 89.0 percent with a minimum precision of 98.5 percent across the parameters evaluated, which provides a reference point for short-read variant calling.
Why does my filtering strategy perform differently on different samples?
Filtering performance depends on sequencing depth, library quality, sample diversity, and genome complexity. Samples with lower depth will have more false negatives because true variants are not covered by enough reads. Samples with higher diversity relative to the reference genome will have more false positives because reads from divergent regions align poorly. The filtering strategy should be evaluated on samples that represent the range of data quality in the study.
Should I use hard filters or machine learning filters?
Hard filters are simple to implement and interpret but do not capture interactions between metrics. Machine learning filters consider multiple metrics simultaneously and can capture these interactions. The plant benchmarking study found that a machine-learning-based filtering strategy outperformed the traditional hard-cutoff strategy. However, machine learning filters require a training dataset with known true variants. If you have a gold standard dataset, machine learning filters are worth considering. If not, hard filters are the practical choice.
How often should I re-evaluate my filtering strategy?
You should re-evaluate the filtering strategy whenever a tool in the pipeline is updated, whenever the sequencing platform or chemistry changes, and at regular intervals such as annually. Pipeline drift can occur without an intentional change to the pipeline. Regular evaluation against a truth set is the most reliable way to detect drift. The records from previous evaluations provide the baseline for comparison.
What should I report about filtering performance in my publication?
You should report the truth set used, the pipeline version, the filter thresholds, the counts of true positives, false positives, and false negatives, and the calculated precision and recall. This reporting allows other researchers to assess the reliability of the variant calls and to compare the filtering strategy with their own. The report should also include the limitations of the evaluation, such as whether the truth set was generated from the same sample or a different sample.
What should I do if my precision or recall is much lower than expected?
First, check whether the truth set is appropriate for your sample. A truth set from a different population or a different sequencing platform may not be representative. Second, check whether the comparison method is appropriate for your variant types. Indel comparisons are particularly sensitive to the matching method. Third, consider whether the pipeline has a problem in the alignment or variant calling steps. If the problem persists, seek expert assistance from resources such as the Galaxy Training Network or the EMBL-EBI Training program.
Related Bioinformatics Guides
- Evaluating Genome Assembly Quality: Metrics and Tools
- Single-Cell RNA Sequencing Quality Control: A Practical Guide to Filtering and Metrics
- Detecting Structural Variants with Long-Read Sequencing: Methods and Considerations
- Radiomics Feature Selection: Methods and Best Practices
- Spatial Transcriptomics Methods: A Guide to Experimental Approaches
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- Benchmarking the empirical accuracy of short-read sequencing across the M. tuberculosis genome.. Bioinformatics (Oxford, England), 2022.
- Benchmarking variant identification tools for plant diversity discovery.. BMC genomics, 2019.
- Memory Visualization-Based Malware Detection Technique.. Sensors (Basel, Switzerland), 2022.
- Benchmarking genomic variant calling tools in inbred mouse strains: recommendations and considerations.. Genetics, 2026.
- Benchmarking Genomic Variant Calling Tools in Inbred Mouse Strains: Recommendations and Considerations.. bioRxiv : the preprint server for biology, 2025.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.