Essential Quality Metrics for Evaluating Variant Calling Results: Mapping Quality, Depth, and Variant Quality Scores
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- QUAL score represents the Phred-scaled probability of a variant call being incorrect; a score of 20 implies a 1% false positive rate, while 30 indicates a 0.1% rate, with higher scores signifying greater confidence.
- Depth of Coverage (DP) directly impacts variant detection power; insufficient depth (e.g., <10-20x for germline) increases false negatives, while excessive depth offers diminishing returns and can indicate artifacts.
- Genotype Quality (GQ) quantifies confidence in the called genotype, crucial for distinguishing heterozygotes from homozygotes, especially at low depths where low GQ signals uncertainty.
- Mapping Quality (MQ) reflects read alignment accuracy; low MQ (<20-40) suggests potential misalignment artifacts in repetitive or paralogous genomic regions, necessitating careful evaluation.
- Germline variant calling relies on allele frequencies near 50% (heterozygote) or 100% (homozygote) for validation, with sufficient depth (e.g., 20-30x) critical to differentiate true variants from sequencing errors.
- Somatic variant calling for low-frequency variants demands significantly higher depth (e.g., 50-100x) to distinguish true signals from sequencing errors, often requiring more stringent quality filters and consideration of tumor purity.
Variant calling results in VCF files contain multiple quality metrics that determine whether a reported genetic variant is reliable enough for downstream analysis. This article explains the core metrics, how to interpret them, and how to apply them for filtering in germline and somatic variant calling workflows.
Scope and Reader Context
Researchers analyzing next-generation sequencing data must evaluate variant calling output before using it for biological interpretation. The VCF format includes per-variant and per-sample fields that describe mapping quality, sequencing depth, genotype quality, and overall variant quality scores. Understanding these metrics allows researchers to filter false positives, assess pipeline performance, and make informed decisions about which variants to validate or report. This article covers the definitions, practical thresholds, common failure patterns, and decision criteria for using quality metrics in variant filtering.
The Role of Quality Metrics in Variant Calling Pipelines
Variant calling pipelines process raw sequencing data through several stages: quality control of raw reads, alignment to a reference genome, variant calling, and annotation. Each stage generates metrics that influence the final variant list. The quality metrics stored in VCF files reflect the confidence that a variant call is accurate instead of an artifact of sequencing error, misalignment, or library preparation issues.
The importance of quality metrics extends beyond individual variant filtering. Quality assessment at the alignment and variant calling stages is essential for meaningful and successful studies, yet monitoring of QC metrics in specific steps including alignment and variant calling is neglected in certain pipelines such as the Best Practices Workflows in GATK. Researchers should therefore implement their own quality checks instead of assuming that a pipeline produces reliable output by default.
Quality metrics serve two distinct purposes. First, they allow filtering of individual variants based on confidence thresholds. Second, they provide an overall assessment of whether the sequencing run and analysis pipeline performed adequately. Both purposes require understanding what each metric measures and what values indicate problems.
Core Quality Metrics in VCF Files
The VCF format includes fixed fields and sample-specific fields. The fixed fields include POS, ID, REF, ALT, QUAL, and FILTER. Sample-specific fields appear in the FORMAT column and include DP, GQ, and genotype information. The INFO column contains additional annotations such as mapping quality (MQ), depth of coverage, and allele frequency.
QUAL Score
The QUAL field represents the Phred-scaled probability that the variant call is incorrect. A QUAL score of 20 indicates a 1 in 100 chance that the variant is a false positive. A score of 30 indicates a 1 in 1000 chance. Higher QUAL scores correspond to higher confidence in the variant call.
QUAL scores are calculated differently by different variant callers. Some callers incorporate read depth, mapping quality, base quality, and strand bias into the QUAL calculation. Others use simpler models. Researchers should understand how their chosen caller calculates QUAL before applying a universal threshold.
Depth of Coverage (DP)
The DP field represents the number of reads that cover the variant position. In the INFO column, DP typically refers to the total depth across all samples. In the FORMAT column, DP refers to the depth for that specific sample.
Depth of coverage directly affects the statistical power to detect variants. Low depth increases the chance that a true variant is missed or that a sequencing error is called as a variant. High depth provides more evidence for or against a variant call. The relationship between depth and variant confidence is not linear, and beyond a certain depth, additional reads provide diminishing returns.
Genotype Quality (GQ)
The GQ field represents the Phred-scaled confidence that the called genotype is correct. A GQ score of 20 means a 1 in 100 chance that the genotype is wrong. GQ is distinct from QUAL because a variant can be confidently called as present while the specific genotype assignment is uncertain.
GQ is particularly important for heterozygous calls. At low depth, distinguishing between a true heterozygote and a homozygote with sequencing errors becomes difficult. Low GQ values indicate that the genotype assignment should be treated with caution.
Mapping Quality (MQ)
The MQ field represents the root mean square mapping quality of reads across the variant position. Mapping quality reflects the probability that a read is correctly placed in the genome. Reads that map to multiple locations or map with many mismatches receive lower mapping quality scores.
Low mapping quality can indicate that the variant lies in a repetitive or paralogous region of the genome. Variants in such regions are more likely to be artifacts of misalignment instead of true biological variants. Tailored solutions for paralogous or low-complexity areas of the genome are needed to improve variant calling accuracy in these regions.
Additional Quality-Related Fields
Several other fields commonly appear in VCF files and provide useful filtering information. The allele frequency (AF) field reports the proportion of reads supporting the alternate allele. The read depth of the alternate allele (AD) reports the count of reads supporting each allele. Strand bias fields indicate whether the alternate allele appears predominantly on one strand, which can suggest an artifact.
The FILTER field in the VCF header defines which variants passed or failed quality thresholds. Variants that fail filters are marked with the filter name instead of PASS. Researchers can customize filter definitions based on their specific quality requirements.
At a Glance: Quality Metric Decision Table
| Metric | What It Measures | Typical Use | Common Threshold | Interpretation Guidance |
|---|---|---|---|---|
| QUAL | Phred-scaled probability that the variant call is incorrect | Overall variant confidence filtering | 20 to 30 for germline, higher for somatic | Lower thresholds increase sensitivity but allow more false positives |
| DP (depth) | Number of reads covering the variant position | Minimum coverage filtering | 10 to 20 for germline, 50 to 100 for somatic low-frequency detection | Depth below threshold increases false negative rate |
| GQ | Phred-scaled confidence in the genotype assignment | Genotype reliability filtering | 20 to 30 | Low GQ indicates uncertain heterozygote versus homozygote distinction |
| MQ | Root mean square mapping quality of reads at the position | Misalignment artifact filtering | 20 to 40 depending on aligner and region | Low MQ suggests repetitive or paralogous region artifacts |
Germline Variant Calling Quality Considerations
Germline variant calling aims to identify variants present in the germline genome, typically from blood or normal tissue samples. Germline variants are expected to be present at allele frequencies near 50 percent for heterozygotes or 100 percent for homozygotes. This expectation simplifies quality assessment because the allele frequency provides a check on variant validity.
Depth Requirements for Germline Calling
Germline variant calling requires sufficient depth to distinguish true heterozygotes from sequencing errors. At a depth of 10 reads, a heterozygous variant should have approximately 5 reads supporting the alternate allele. The probability that sequencing errors produce this pattern is low but not negligible. Increasing depth to 20 or 30 reads substantially reduces the false positive rate.
Targeted sequencing panels for clinical applications typically achieve higher depths than whole-genome sequencing. The choice of depth depends on the intended use of the variant calls. Research applications may accept lower depth with higher false positive rates, while clinical reporting requires higher confidence.
Genotype Quality in Germline Calls
GQ is particularly relevant for germline variant calling because genotype assignment affects downstream analysis. A variant that is truly heterozygous but called as homozygous will produce incorrect results in association studies, linkage analysis, and clinical interpretation. Low GQ values should prompt closer examination of the read support for each allele.
Mapping Quality in Germline Calls
Germline variant calling across the whole genome encounters repetitive regions, segmental duplications, and paralogous genes. Variants in these regions often have low mapping quality because reads cannot be uniquely assigned. The optimization of alignment algorithms and attention to quality-coverage metrics are required to adapt genomics strategies for clinical needs. Researchers should treat variants with low mapping quality as candidates for validation or exclusion.
Somatic Variant Calling Quality Considerations
Somatic variant calling identifies variants present in tumor or diseased tissue that are absent from the normal germline genome. Somatic variants can be present at low allele frequencies because tumor samples often contain a mixture of tumor and normal cells. This mixture complicates quality assessment because low allele frequency variants are difficult to distinguish from sequencing errors.
Low-Frequency Variant Detection
The detection of low-frequency somatic variants requires higher depth than germline variant calling. At an allele frequency of 5 percent, a depth of 100 reads provides approximately 5 reads supporting the alternate allele. Distinguishing this signal from sequencing error requires careful quality filtering and often specialized variant callers.
Benchmarking studies of low-frequency variant calling demonstrate the challenges involved. In a study using nanopore sequencing for low-frequency mitochondrial DNA variant detection, the choice of aligner affected both the F1 score and the allele frequencies of false-positive calls. One aligner showed higher F1 scores but also higher allele frequencies of false-positive calls across the mixtures compared to another aligner. This finding illustrates that aligner choice influences the quality metrics and error patterns in variant calling results.
Somatic Variant Quality Filters
Somatic variant calling pipelines typically apply more stringent quality filters than germline pipelines. The QUAL threshold may be set higher to reduce false positives. Additional filters based on strand bias, read position, and allele frequency may be applied. The specific thresholds depend on the variant caller and the intended use of the results.
Tumor Purity and Variant Detection
The proportion of tumor cells in a sample affects the allele frequency of somatic variants. A sample with 50 percent tumor purity and a heterozygous somatic variant will show an allele frequency of approximately 25 percent. Lower tumor purity reduces the allele frequency and makes variant detection more difficult. Quality metrics alone cannot compensate for low tumor purity, and researchers should consider sample composition when interpreting variant calls.
Variant Calling Workflow Stages and Quality Assessment
Quality assessment should occur at multiple stages of the variant calling workflow, also at the final VCF output. Each stage generates metrics that can identify problems before they propagate through the pipeline.
Raw Sequence Data Quality Control
The first stage involves assessing the quality of raw sequencing reads. Metrics include per-base quality scores, GC content, adapter contamination, and duplication rates. Poor quality raw data produces poor variant calls regardless of downstream filtering. The Galaxy Training Network provides accessible workflow training for quality control and analysis steps that researchers can adapt to their specific data types.
Alignment Quality Assessment
After alignment, researchers should assess mapping rates, insert size distributions, and coverage uniformity. Low mapping rates may indicate contamination, adapter problems, or reference genome issues. Coverage uniformity affects the reliability of variant calls across the genome. Regions with very low or very high coverage require special attention.
Post-Alignment Processing
Reads may undergo additional processing after alignment, including duplicate removal, base quality score recalibration, and local realignment. These steps improve the accuracy of variant calling by reducing systematic errors. The choice of processing steps affects the quality metrics in the final VCF file.
Variant Calling and Post-Calling Filters
The variant calling stage produces the initial VCF file with raw quality metrics. Post-calling filters refine the variant list based on quality thresholds and annotation information. The nf-core documentation describes community pipeline standards that include quality control and filtering steps for reproducible variant calling workflows.
Practical Implementation Steps for Quality Metric Assessment
Researchers should implement a systematic approach to quality metric assessment instead of relying on default pipeline settings. The following steps provide a framework for evaluating variant calling results.
Step 1: Examine Overall VCF Statistics
Before filtering individual variants, examine the overall statistics of the VCF file. Count the number of variants, the transition to transversion ratio, the distribution of QUAL scores, and the number of variants failing each filter. Unusual distributions may indicate pipeline problems.
Step 2: Assess Depth and Coverage Distribution
Calculate the distribution of depth across all variant positions. Identify regions with very low or very high depth. Low depth regions may have false negative variants. High depth regions may indicate PCR duplicates or alignment artifacts.
Step 3: Evaluate Mapping Quality Distribution
Examine the distribution of mapping quality scores across variants. A large proportion of variants with low mapping quality suggests problems with the reference genome, the aligner settings, or the presence of repetitive regions in the target areas.
Step 4: Apply Quality Filters
Apply quality filters based on the intended use of the variant calls. Document the filtering thresholds and the number of variants removed at each step. Reproducible analysis requires that filtering decisions are recorded and justified.
Step 5: Validate a Subset of Variants
Select a subset of variants for validation using an orthogonal method such as Sanger sequencing or an independent sequencing platform. The validation rate provides an estimate of the false positive rate in the filtered variant set.
Step 6: Document Quality Metrics in Reports
Include quality metrics in analysis reports so that downstream users can assess the reliability of the variant calls. The report should state the depth, mapping quality, and variant quality thresholds used for filtering.
Records and Measurements for Quality Assessment
Maintaining records of quality metrics allows researchers to track pipeline performance over time and across samples. The following measurements should be recorded for each variant calling run.
Per-Sample Metrics
For each sample, record the total number of variants called, the number passing filters, the mean depth across variant positions, the transition to transversion ratio, and the number of variants with low mapping quality. These metrics provide a baseline for detecting sample-specific problems.
Per-Run Metrics
For each sequencing run or batch, record the overall mapping rate, the duplication rate, the mean coverage, and the percentage of the target region covered at specified depth thresholds. Batch effects can be detected by comparing these metrics across runs.
Pipeline Version Tracking
Record the versions of all software used in the variant calling pipeline, including the aligner, the variant caller, and any post-processing tools. Software updates can change quality metric calculations and variant calling behavior. The Bioconductor project provides official package and workflow documentation that supports reproducible genomic analysis with version tracking.
Quality Metric Trends
Track quality metrics over time to identify gradual changes in sequencing or analysis performance. A gradual decline in mapping quality or depth may indicate instrument problems, reagent degradation, or reference genome issues.
Common Failure Patterns in Variant Calling Quality
Several recurring problems appear in variant calling results. Recognizing these patterns allows researchers to diagnose and correct pipeline issues.
Low Depth in Specific Genomic Regions
Certain genomic regions consistently show low depth, including GC-rich regions, repetitive elements, and regions with high sequence similarity to other parts of the genome. Variants in these regions may be missed entirely or called with low confidence. Researchers should be aware of the limitations of their sequencing platform and aligner in these regions.
Strand Bias Artifacts
Strand bias occurs when the alternate allele appears predominantly on one strand. This pattern often indicates an artifact instead of a true variant. Many variant callers report strand bias statistics that can be used for filtering. The double-masking approach for bisulfite sequencing data demonstrates how per-strand analysis can distinguish true polymorphisms from artificial mutations induced by chemical treatment.
Mapping Quality Collapse in Paralogous Regions
Paralogous genes share high sequence similarity and can cause reads to map incorrectly. Variants in these regions often have low mapping quality and may represent differences between paralogs instead of true variants in the target gene. The optimization of alignment algorithms and tailored solutions for paralogous or low-complexity areas of the genome are needed to address this problem.
Contamination Effects
Sample contamination with DNA from another individual or species can produce mixed allele frequencies and unusual quality metric patterns. Contamination is particularly problematic for somatic variant calling because it mimics low tumor purity. Quality metrics alone may not detect contamination, and researchers should use dedicated contamination detection tools.
Batch Effects Across Sequencing Runs
Variants called from samples sequenced in different runs may show systematic differences in quality metrics. These batch effects can produce false associations in downstream analysis. Comparing quality metric distributions across batches helps identify and correct these effects.
Limitations of Quality Metrics
Quality metrics provide useful information but have inherent limitations that researchers must understand.
Metrics Cannot Detect All Errors
Quality metrics reflect the probability of error based on the sequencing and analysis model used by the variant caller. Systematic errors that violate the model assumptions may not be reflected in the quality scores. For example, alignment errors in repetitive regions may produce high QUAL scores because the reads appear to support the variant consistently.
Thresholds Are Context Dependent
Universal quality thresholds do not exist. The appropriate threshold depends on the sequencing platform, the variant caller, the genomic region, and the intended use of the variant calls. Researchers should establish thresholds based on their specific data and validate them with known variants.
Quality Metrics Do Not Capture Biological Validity
A variant can have high quality scores and still be biologically incorrect. For example, a variant may be a sequencing artifact that consistently appears across many samples. Quality metrics cannot distinguish between a true variant and a systematic artifact. Variant annotation and comparison with known variant databases provide additional evidence for biological validity.
Low-Frequency Variant Detection Limits
The detection of low-frequency variants is fundamentally limited by sequencing error rates. Even with high depth, distinguishing a true variant at 1 percent allele frequency from sequencing error is challenging. Benchmarking studies of low-frequency variant calling with long-read data on mitochondrial DNA identified current limitations in detecting variants at 1 percent and 2 percent mixture levels, with performance varying by variant caller and aligner choice.
Quality Controls and Professional Escalation Criteria
Researchers should establish quality control procedures and know when to escalate problems to more experienced colleagues or seek additional validation.
Quality Control Triggers
The following situations should trigger additional quality investigation: a sudden change in the number of variants called, a large proportion of variants failing quality filters, unusual transition to transversion ratios, or inconsistent results between replicate samples. These patterns may indicate problems with the sequencing run, the analysis pipeline, or the sample itself.
Escalation Criteria
Researchers should escalate quality problems when they cannot identify the cause through standard troubleshooting. Situations that warrant escalation include persistent low mapping quality across multiple samples, contamination that cannot be removed by filtering, and discrepancies between variant calls and orthogonal validation results. In clinical settings, variants of uncertain significance or those near detection thresholds often require validation before reporting.
Clinical Reporting Considerations
For clinical applications, quality metrics are critical for accurate variant calling, with validation often required for variants of uncertain significance or those near detection thresholds. The integration of next-generation sequencing into diagnostic frameworks for conditions such as acute myeloid leukemia and myelodysplastic neoplasms has made quality assessment a standard part of clinical reporting. Bioinformatic pipelines and quality metrics including read length, sequencing depth, and coverage are critical for accurate variant calling in these settings.
Training and Reproducibility Resources
Researchers seeking to improve their variant calling quality assessment skills can use several training resources. The EMBL-EBI Training program provides bioinformatics learning pathways and data-resource training for practical analysis education. The Galaxy Training Network offers accessible workflow training and analysis tutorials that cover variant calling and quality assessment. The Carpentries Lessons provide foundational computing and data skills that support reproducible analysis practices. The nf-core documentation describes community pipeline standards for reproducible variant calling workflows.
These resources support the development of reproducible analysis practices, which are essential for reliable variant calling quality assessment. Reproducible workflows ensure that quality metrics can be compared across samples and runs and that filtering decisions can be justified and reviewed.
Building a Structured Quality Metric Review System for Variant Calling Runs
Quality metrics in VCF files only become useful when researchers apply them through a consistent, documented review process. Many variant calling pipelines produce VCF files with quality scores, but researchers often lack a structured method for examining those metrics systematically across samples and runs. This section provides a practical decision framework for reviewing variant calling quality, a record system for tracking metrics over time, and troubleshooting methods for common quality failures. The framework is designed for researchers who need to make defensible filtering decisions and detect pipeline problems before they compromise downstream analysis.
The Three Tier Quality Review Framework
A structured quality review should operate at three distinct levels: the individual variant level, the sample level, and the batch or run level. Each tier answers a different question and requires different thresholds and actions.
Tier 1: Individual Variant Assessment
The first tier examines each variant in the VCF file against the core quality metrics described in the main article. This tier answers the question of whether a specific variant call is reliable enough to include in downstream analysis. The assessment uses the QUAL, DP, GQ, and MQ fields as primary filters, with additional fields such as allele balance and strand bias as secondary checks.
For each variant, the researcher should record the values of the core metrics and compare them against pre-established thresholds. The thresholds should be defined before the review begins, not adjusted during the review to accommodate borderline variants. Adjusting thresholds during review introduces bias and makes the filtering process difficult to reproduce.
The individual variant assessment should also include a check for allele balance consistency. For germline heterozygous variants, the allele balance should approximate 50 percent. For germline homozygous variants, the allele balance should approach 100 percent. Significant deviation from these expected values suggests either a genotyping error or sample contamination. For somatic variants, the expected allele balance depends on tumor purity and copy number status, making this check more complex.
Tier 2: Sample Level Assessment
The second tier examines the distribution of quality metrics across all variants within a single sample. This tier answers the question of whether the sample itself was sequenced and processed adequately. A sample with poor overall quality metrics may produce unreliable variant calls even for individual variants that pass Tier 1 filters.
Key sample level metrics include the total number of variants called, the transition to transversion ratio, the distribution of depth across variant positions, and the proportion of variants failing each quality filter. The transition to transversion ratio provides a useful sanity check because transitions are generally more common than transversions in human genomes. A ratio well outside the expected range for the species and sequencing platform suggests systematic calling errors.
The sample level assessment should also examine the distribution of QUAL scores across all variants. A healthy sample typically shows a distribution with most variants at moderate to high QUAL scores and a tail of lower quality variants. A sample with an unusually large proportion of low QUAL variants may indicate problems with sequencing depth, library quality, or sample degradation.
Tier 3: Batch and Run Level Assessment
The third tier compares quality metrics across multiple samples processed in the same batch or sequencing run. This tier answers the question of whether the batch or run introduced systematic artifacts that affect all samples equally. Batch effects are particularly problematic because they can produce false associations in downstream analysis.
The batch level assessment should compare the distributions of depth, mapping quality, and variant counts across all samples in the batch. Samples that fall well outside the batch distribution warrant investigation. The assessment should also compare metrics across batches processed at different times or on different instruments to detect instrument specific patterns.
The nf-core documentation describes community pipeline standards that include quality control and filtering steps for reproducible variant calling workflows. These standards provide a useful reference for establishing batch level quality expectations and for ensuring that the review process itself is reproducible across runs.
Implementing the Review Process
The three tier review process should be implemented as a documented procedure with clear decision points and escalation criteria. The following steps describe a practical implementation that can be adapted to different research settings.
Step 1: Define Thresholds Before Review
Before examining any VCF file, define the quality thresholds for each tier of the review. The thresholds should be based on the sequencing platform, the variant caller, the expected allele frequencies, and the intended use of the variant calls. Document the thresholds and the rationale for each threshold in the analysis protocol.
For Tier 1, define minimum thresholds for QUAL, DP, GQ, and MQ. For Tier 2, define acceptable ranges for the transition to transversion ratio, the proportion of variants passing filters, and the mean depth across variant positions. For Tier 3, define acceptable variation across samples within a batch.
Step 2: Run the Tier 1 Variant Filter
Apply the Tier 1 thresholds to the VCF file and record the number of variants passing and failing each filter. The filtering can be performed using standard bioinformatics tools or custom scripts. The Bioconductor project provides official package and workflow documentation that supports reproducible genomic analysis with version tracking, which is useful for implementing and documenting the filtering process.
Record the number of variants removed by each filter separately, beyond the total number removed. This information helps identify which quality metric is causing the most variant loss and whether the threshold is appropriate for the data.
Step 3: Generate Sample Level Summary Statistics
For each sample, generate summary statistics for the variants that passed Tier 1 filters. Include the total variant count, the transition to transversion ratio, the mean and median depth, the distribution of QUAL scores, and the proportion of variants at each genotype class.
Compare these statistics against the thresholds defined in Step 1. Flag any sample that falls outside the acceptable range for any metric.
Step 4: Compare Across Samples and Batches
Group the sample level statistics by batch or sequencing run and compare the distributions across groups. Look for systematic differences in depth, variant count, or quality score distributions. Flag any batch that shows a distinct pattern compared to other batches.
The Galaxy Training Network provides accessible workflow training and analysis tutorials that cover variant calling and quality assessment. These tutorials can help researchers implement the batch comparison step using reproducible workflows.
Step 5: Document the Review Results
Record the results of each tier of the review in a structured format that can be compared across runs. The record should include the date of the review, the pipeline version, the thresholds used, the number of variants passing and failing each filter, and any samples or batches that were flagged for investigation.
The EMBL-EBI Training program provides bioinformatics learning pathways and data-resource training for practical analysis education. These resources can help researchers develop the documentation skills needed for maintaining consistent quality review records.
Record System for Tracking Quality Metrics
A structured record system allows researchers to track quality metrics over time and detect gradual changes in sequencing or analysis performance. The record system should capture both per sample and per run metrics in a format that supports trend analysis.
Per Sample Records
For each sample, record the following information in a structured table or database: sample identifier, sequencing run identifier, pipeline version, total variants called, variants passing Tier 1 filters, transition to transversion ratio, mean depth across variant positions, median depth, proportion of variants with MQ below threshold, and any flags raised during the review.
The per sample record should also include the date of sequencing and the date of analysis. This information allows researchers to correlate quality changes with instrument maintenance, reagent lot changes, or software updates.
Per Run Records
For each sequencing run or batch, record the overall mapping rate, the duplication rate, the mean coverage across the target region, the percentage of the target region covered at specified depth thresholds, and the distribution of quality metrics across all samples in the run.
The per run record should also include information about the sequencing instrument, the flow cell lot, the reagent kit lot, and any instrument or protocol deviations that occurred during the run. This information helps identify the cause of quality problems when they appear.
Trend Analysis
Review the per run records regularly to identify trends in quality metrics. A gradual decline in mapping quality or depth across multiple runs may indicate instrument degradation or reagent problems. A sudden change in the transition to transversion ratio may indicate a software update or a change in the analysis pipeline.
The Carpentries Lessons provide foundational computing and data skills that support reproducible analysis practices. These skills are essential for maintaining the record system and performing the trend analysis effectively.
Troubleshooting Common Quality Failures
When the quality review identifies problems, researchers need a systematic method for diagnosing the cause. The following troubleshooting approach addresses the most common quality failure patterns.
Pattern 1: Low Depth Across the Entire Sample
If the sample level assessment shows low depth across most variant positions, the cause is likely at the sequencing or library preparation stage. Check the raw sequence yield, the mapping rate, and the duplication rate. Low raw yield indicates a sequencing problem. Low mapping rate indicates contamination or adapter problems. High duplication rate indicates library amplification issues.
The comprehensive review of somatic variant calling and quality management strategies for human cancer genomes emphasizes that monitoring of QC metrics in specific steps including alignment and variant calling is neglected in certain pipelines. Researchers should therefore check the alignment metrics directly instead of relying only on the VCF quality fields.
Pattern 2: Low Depth in Specific Genomic Regions
If low depth appears only in specific regions, the cause is likely related to sequence composition or alignment. GC rich regions, repetitive elements, and regions with high sequence similarity to other parts of the genome commonly show low depth. Check whether the low depth regions share common sequence features.
The optimization of alignment algorithms and attention to quality coverage metrics are required to adapt genomics strategies for clinical needs, particularly for paralogous or low complexity areas of the genome. Consider whether the aligner settings need adjustment for these regions.
Pattern 3: High Proportion of Variants Failing Mapping Quality Filters
If many variants have low MQ scores, the cause may be misalignment in repetitive or paralogous regions. Check whether the low MQ variants cluster in specific genomic regions. If they do, the problem is likely alignment related. If they are distributed across the genome, the problem may be with the reference genome or the aligner settings.
Pattern 4: Unexpected Allele Balance Patterns
If germline variants show allele balance values that deviate significantly from the expected 50 percent for heterozygotes or 100 percent for homozygotes, check for sample contamination or copy number variation. Contamination from another individual produces mixed allele frequencies. Copy number variation produces allele balance values that deviate from the expected ratios.
Pattern 5: Batch Specific Quality Patterns
If one batch of samples shows quality metrics that differ systematically from other batches, investigate the sequencing run conditions, the reagent lots, and the analysis pipeline version used for that batch. The benchmarking study of low frequency variant calling with long read data on mitochondrial DNA demonstrated that the choice of aligner affected both the F1 score and the allele frequencies of false positive calls. Software version differences between batches can produce similar effects.
Professional Escalation Criteria
The quality review process should include clear criteria for escalating problems to more experienced colleagues or seeking additional validation. The following situations warrant escalation.
Persistent Unexplained Quality Problems
If the troubleshooting process cannot identify the cause of a quality problem, escalate to a bioinformatics specialist or the sequencing facility. Persistent problems that affect multiple samples or runs may indicate instrument issues that require professional maintenance.
Discrepancies Between Variant Calls and Validation Results
If orthogonal validation of a subset of variants shows a high false positive rate despite good quality metrics, escalate the problem. This situation suggests a systematic error that the quality metrics do not capture. The discrepancy may indicate problems with the variant caller, the reference genome, or the validation method itself.
Clinical Reporting Situations
For clinical applications, quality problems that affect variant interpretation should be escalated immediately. The review of next generation sequencing reports for acute myeloid leukemia and myelodysplastic neoplasms notes that bioinformatic pipelines and quality metrics including read length, sequencing depth, and coverage are critical for accurate variant calling, with validation often required for variants of uncertain significance or those near detection thresholds. In clinical settings, the escalation criteria should be defined in the laboratory standard operating procedures.
Contamination Suspected
If the quality review suggests sample contamination, escalate the problem before proceeding with downstream analysis. Contamination can produce false variant calls and incorrect genotype assignments. The double masking approach for bisulfite sequencing data demonstrates how per strand analysis can distinguish true polymorphisms from artificial mutations, but contamination requires dedicated detection tools and may require sample resequencing.
Integrating the Review Process into Existing Pipelines
The quality review process should be integrated into the variant calling pipeline instead of applied as an afterthought. The review can be implemented as a series of automated steps that generate reports at each tier, with manual review reserved for samples or batches that trigger flags.
The nf-core documentation describes community pipeline standards that include quality control and filtering steps for reproducible variant calling workflows. These standards provide a useful model for integrating quality review into automated pipelines.
The review process should also be documented in the analysis protocol so that the filtering decisions can be justified and reproduced. The documentation should include the thresholds used, the rationale for each threshold, and the results of the review for each sample and batch.
Limitations of the Structured Review Approach
The structured review approach has limitations that researchers should understand. The thresholds defined in the review are context dependent and may not transfer across sequencing platforms, variant callers, or sample types. The review process cannot detect all systematic errors, particularly those that produce high quality scores for false variants. The review process also requires ongoing maintenance as pipelines and sequencing technologies evolve.
The review process should therefore be treated as a component of a broader quality management strategy that includes orthogonal validation, comparison with known variant databases, and ongoing monitoring of pipeline performance. The comprehensive review of somatic variant calling and quality management strategies emphasizes that multidimensional QC of sequencing data is essential for a meaningful and successful study. The structured review process described here provides a practical framework for implementing that multidimensional QC in the context of variant calling quality assessment.
Frequently Asked Questions
What is the difference between QUAL and GQ in a VCF file?
QUAL represents the Phred-scaled probability that the variant call itself is incorrect, meaning the variant is not truly present at that position. GQ represents the Phred-scaled probability that the specific genotype assignment is incorrect. A variant can have high QUAL, indicating confidence that a variant exists, but low GQ, indicating uncertainty about whether the sample is heterozygous or homozygous for the alternate allele.
What depth of coverage is needed for reliable germline variant calling?
The required depth depends on the application and the expected allele frequency. For germline variants at 50 percent allele frequency, depths of 10 to 30 reads provide reasonable confidence for most applications. Clinical applications typically require higher depth, often 50 to 100 reads, to ensure accurate genotype calls. Lower depth increases the false negative rate and the chance of incorrect genotype assignment.
How does mapping quality affect variant calling accuracy?
Mapping quality reflects the probability that a read is correctly placed in the genome. Reads with low mapping quality may be placed incorrectly, producing false variant calls. Variants in repetitive or paralogous regions often have low mapping quality because reads cannot be uniquely assigned. Low mapping quality variants should be treated with caution and may require validation.
Why do somatic variant calling pipelines use different quality thresholds than germline pipelines?
Somatic variants can be present at low allele frequencies because tumor samples contain a mixture of tumor and normal cells. Distinguishing low-frequency variants from sequencing errors requires higher depth and often more stringent quality filters. Somatic pipelines may also apply additional filters based on strand bias and read position to reduce false positives.
Can quality metrics identify all false positive variants?
No. Quality metrics reflect the probability of error based on the variant caller model, but systematic errors that violate model assumptions may not be detected. Alignment errors in repetitive regions, sample contamination, and batch effects can produce false variants with high quality scores. Orthogonal validation and comparison with known variant databases provide additional evidence for variant validity.
What should I do if a large proportion of my variants fail quality filters?
A large proportion of variants failing quality filters suggests a systematic problem instead of random variation. Check the raw sequence quality, the alignment metrics, and the variant caller settings. Compare the quality metric distributions with previous runs using the same pipeline. If the problem persists, escalate to a bioinformatics specialist or consider resequencing the samples.
How do I choose quality thresholds for my specific dataset?
Quality thresholds should be based on the sequencing platform, the variant caller, the genomic region, and the intended use of the variant calls. Start with thresholds recommended by the variant caller documentation or published studies using similar data. Validate the filtered variant set using known variants or orthogonal methods and adjust thresholds based on the observed false positive and false negative rates.
What quality metrics should I report with my variant calling results?
Report the total number of variants called and passing filters, the mean depth across variant positions, the distribution of QUAL scores, the transition to transversion ratio, and the number of variants with low mapping quality. For somatic calling, also report the tumor purity estimate and the minimum allele frequency threshold used for variant detection. These metrics allow downstream users to assess the reliability of the variant calls.
Related Bioinformatics Guides
- Evaluating Genome Assembly Quality: Metrics and Tools
- Detecting Structural Variants with Long-Read Sequencing: Methods and Considerations
- How to Interpret Gene Set Enrichment Analysis Results
- RNA-Seq Quality Control: Essential Checks and Tools
- Single-Cell RNA Sequencing Quality Control: A Practical Guide to Filtering and Metrics
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- Comprehensive fundamental somatic variant calling and quality management strategies for human cancer genomes.. Briefings in bioinformatics, 2021.
- Towards precision medicine.. Nature reviews. Genetics, 2016.
- Manipulating base quality scores enables variant calling from bisulfite sequencing alignments using conventional bayesian approaches.. BMC genomics, 2022.
- How to Read a Next-Generation Sequencing Report for AML and MDS? What Hematologists Need to Know.. Journal of clinical medicine, 2025.
- Benchmarking Low-Frequency Variant Calling With Long-Read Data on Mitochondrial DNA.. Frontiers in genetics, 2022.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.