# Using Conservation Scores to Prioritize Variants: PhyloP, GERP++, and CADD Explained


## Key Takeaways

- Conservation scores (PhyloP, GERP++) quantify evolutionary constraint at genomic positions by comparing observed substitutions against neutral evolution models, with higher positive scores indicating stronger purifying selection and functional importance.
- CADD integrates conservation with numerous other annotations (e.g., regulatory elements, protein structure) using a machine learning approach to predict deleteriousness, offering a broader assessment than conservation alone.
- PhyloP provides base-by-base scores that can detect both conservation (positive scores) and accelerated evolution (negative scores), while GERP++ focuses on rejected substitutions (RS scores) within constrained elements, offering comparable constraint metrics.
- Practical variant prioritization workflows involve annotating with multiple scores, applying study-specific thresholds (e.g., PhyloP > 2, GERP++ RS > 2, scaled CADD > 15-20), and cross-validating results by requiring agreement across different scoring systems.
- Integrating conservation scores with population allele frequencies (e.g., from gnomAD) and functional annotations (e.g., predicted protein changes via SnpEff) is crucial for robust variant filtering, as common variants at conserved sites may be tolerated.
- Documentation of score versions, genome builds, and threshold selection is paramount for reproducibility, and discordant scores across methods necessitate a tiered classification approach based on agreement to guide further investigation.

---

## Scope and Reader Context

Researchers analyzing whole-genome or whole-exome sequencing data face a common problem: a single human genome contains millions of genetic variants, but only a small fraction are likely to have functional consequences. Conservation scores provide one layer of evidence for prioritizing which variants merit further investigation. This article explains three widely used scoring systems, PhyloP, GERP++, and CADD, describes how each is calculated, and outlines practical strategies for combining them in variant filtering workflows. The content is directed at biology students, researchers, laboratory professionals, and life-science practitioners who need to make informed decisions about which variants to carry forward into validation studies or clinical interpretation.

Variant annotation is the process of enriching raw variant calls with biological information. This includes predicting amino acid changes, identifying potential splice disruptions, and combining variant data with genomic databases, conservation scores, and population allele frequencies. The goal is to filter and prioritize variants into a reduced set that is highly relevant for the study or clinical assay at hand. Conservation scores serve as one component of this annotation pipeline, providing evolutionary evidence that a genomic position has been maintained across species and therefore may be functionally important.

## What Conservation Scores Measure

Conservation scores quantify how similar a genomic position is across multiple species. The underlying logic is that functionally important regions of the genome, such as protein-coding exons, regulatory elements, and RNA genes, tend to change slowly over evolutionary time because mutations in these regions are more likely to be deleterious. Conversely, positions that are free to vary across species are less likely to be under strong functional constraint.

Conservation is not a direct measurement of function. A conserved position may be important for reasons that are not yet understood, and a non-conserved position may still be functional in a species-specific manner. Conservation scores are therefore best interpreted as probabilistic evidence that a position is under purifying selection, which increases the prior probability that a variant at that position has functional consequences.

The three scoring systems covered in this article approach conservation measurement differently. PhyloP and GERP++ are both phylogenetic conservation scores that compare sequence alignments across species, but they use different statistical frameworks. CADD is a broader annotation method that integrates conservation with many other types of information into a single score.

## PhyloP: Phylogenetic Conservation with Base-by-Base Resolution

PhyloP is a conservation scoring method that evaluates each nucleotide position in a multiple sequence alignment. It compares the observed rate of substitution at a position with the rate expected under neutral evolution. The score is derived from a phylogenetic model that accounts for the evolutionary relationships among the species in the alignment.

PhyloP scores can be positive or negative. A positive score indicates that a position is more conserved than expected under neutral evolution, meaning fewer substitutions have occurred than would be predicted by chance. A negative score indicates that a position is less conserved than expected, meaning more substitutions have occurred than would be predicted. The magnitude of the score reflects the strength of the conservation signal.

One important feature of PhyloP is that it can detect accelerated evolution as well as conservation. Positions with significant negative scores may be undergoing positive selection or relaxed constraint, which can be biologically interesting in its own right. For variant prioritization purposes, researchers typically focus on positions with positive PhyloP scores, with higher scores indicating stronger conservation.

PhyloP scores are available for multiple genome builds and can be downloaded from the UCSC Genome Browser or accessed through annotation tools. The scores are precomputed for each position in the genome, which makes them straightforward to incorporate into variant annotation pipelines. The [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/) provide access to the underlying sequence alignments and genome annotations needed to interpret PhyloP scores in context.

## GERP++: Rejected Substitutions as a Measure of Constraint

GERP++ (Genomic Evolutionary Rate Profiling) takes a different approach to measuring conservation. Instead of scoring each position independently, GERP++ identifies constrained elements within a multiple sequence alignment and then estimates the level of constraint at each position within those elements.

The key statistic produced by GERP++ is the rejected substitution (RS) score. This score represents the difference between the number of substitutions actually observed at a position and the number that would be expected under neutral evolution. A positive RS score indicates that fewer substitutions have occurred than expected, which suggests the position is under constraint. A higher RS score indicates stronger constraint.

GERP++ also provides a measure of the number of constrained elements and their boundaries. This element-level information can be useful for understanding whether a variant falls within a broader region of evolutionary constraint instead of being an isolated conserved position.

One practical difference between GERP++ and PhyloP is that GERP++ RS scores are designed to be comparable across positions, with higher scores always indicating stronger constraint. This makes GERP++ scores easier to use in threshold-based filtering than PhyloP scores, which require careful interpretation of the sign and magnitude.

The [EMBL-EBI Training](https://www.ebi.ac.uk/training) resources provide guidance on accessing and interpreting conservation data from Ensembl and related databases, which can help researchers understand how GERP++ scores are generated and where to retrieve them.

## CADD: Integrating Multiple Annotations into a Single Score

CADD (Combined Annotation-Dependent Depletion) differs fundamentally from PhyloP and GERP++ because it does not rely solely on conservation. Instead, CADD integrates many diverse annotations into a single measure for each variant. The method was developed to address a limitation of earlier annotation approaches, which tended to exploit a single information type or were restricted in scope to specific variant classes such as missense changes.

CADD is implemented as a support vector machine trained to differentiate high-frequency human-derived alleles from simulated variants. The training data includes 14.7 million high-frequency human alleles and 14.7 million simulated variants. The model learns to distinguish variants that are depleted in the human population, which are more likely to be deleterious, from variants that are common and therefore more likely to be neutral or tolerated.

The output of CADD is a C score for each variant. C scores are precomputed for all 8.6 billion possible human single-nucleotide variants, and the method also enables scoring of short insertions and deletions. The scores correlate with allelic diversity, annotations of functionality, pathogenicity, disease severity, experimentally measured regulatory effects, and complex trait associations. CADD has been shown to highly rank known pathogenic variants within individual genomes.

The ability of CADD to prioritize functional, deleterious, and pathogenic variants across many functional categories, effect sizes, and genetic architectures is unmatched by any single-annotation method. This is because CADD captures information from conservation scores, regulatory annotations, protein structure predictions, and other sources simultaneously, allowing it to identify variants that might be missed by any one annotation type alone.

The original CADD paper, published in Nature Genetics in 2014, describes the method in detail and provides the framework for interpreting C scores. The [PubMed record for the CADD paper](https://pubmed.ncbi.nlm.nih.gov/24487276) is the primary reference for understanding the method's development and validation.

## At a Glance: Comparison of PhyloP, GERP++, and CADD

| Feature | PhyloP | GERP++ | CADD |
|---------|--------|--------|------|
| Primary output | Per-position score, positive or negative | Rejected substitution (RS) score, positive for constraint | C score, higher indicates more deleterious |
| Statistical basis | Phylogenetic model comparing observed vs. expected substitution rates | Difference between observed and expected substitutions within constrained elements | Support vector machine trained on high-frequency vs. simulated variants |
| Information sources | Multiple sequence alignment across species | Multiple sequence alignment across species | Conservation, regulatory annotations, protein structure, population frequency, and other annotations |
| Variant types scored | Single-nucleotide positions | Single-nucleotide positions | Single-nucleotide variants and short insertions/deletions |
| Interpretation | Positive scores indicate conservation, negative scores indicate acceleration | Higher RS scores indicate stronger constraint | Higher C scores indicate greater likelihood of deleteriousness |
| Best use in filtering | Identifying conserved positions for further analysis | Ranking variants within constrained elements | Broad prioritization across all variant types and functional categories |

## How Conservation Scores Are Calculated

### Multiple Sequence Alignments as the Foundation

Both PhyloP and GERP++ depend on multiple sequence alignments that show the corresponding genomic positions across a set of species. The quality of the alignment directly affects the reliability of the conservation scores. Alignments that include closely related species provide fine-grained information about recent evolutionary changes, while alignments that include distantly related species provide information about deep conservation.

The choice of species included in the alignment is a critical decision. Including too many distantly related species can make alignments noisy and difficult to interpret, while including too few species reduces the statistical power to detect conservation. Most standard conservation tracks are precomputed using a carefully selected set of species and are available for download from genome browsers and databases.

The [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/) provide access to the reference genomes and comparative genomics data needed to understand the species composition of standard alignments. Researchers who need to generate custom conservation scores for non-model organisms must construct their own alignments, which requires careful attention to orthology assignment and alignment quality.

### Phylogenetic Models and Neutral Evolution

Conservation scores are calculated relative to a model of neutral evolution. The neutral model predicts how many substitutions would be expected at a position given the evolutionary distances among the species in the alignment and the mutation rate. Positions that deviate from this expectation are either conserved (fewer substitutions than expected) or accelerated (more substitutions than expected).

The phylogenetic tree describing the relationships among the species is an essential input to this calculation. The tree determines the expected number of substitutions along each branch, and errors in the tree topology can lead to incorrect conservation estimates. Standard conservation tracks use well-established phylogenies, but custom analyses require careful attention to tree construction.

PhyloP uses a statistical framework that evaluates each position independently, which allows it to detect conservation at single-nucleotide resolution. GERP++ first identifies constrained elements and then estimates the level of constraint within those elements, which provides a different perspective on the same underlying data.

### Simulated Variants and Machine Learning in CADD

CADD takes a fundamentally different approach to calculation. Instead of relying on a phylogenetic model, CADD uses machine learning to integrate many annotations. The training process involves generating simulated variants that represent the expected distribution of mutations under neutral evolution. These simulated variants are compared with high-frequency human alleles, which are presumed to be largely neutral because they have reached high frequency in the population.

The support vector machine learns to distinguish the two classes based on the annotation values at each variant. The resulting model can then be applied to any variant, including those not seen in the training data. The C score reflects the position of a variant relative to the decision boundary learned by the model.

One advantage of this approach is that CADD can incorporate annotations that are not directly related to conservation, such as regulatory element annotations, protein structure predictions, and measures of chromatin state. This allows CADD to identify functional variants that are not conserved, which would be missed by conservation-only approaches.

## Practical Workflow for Using Conservation Scores

### Step 1: Annotate Variants with Multiple Scores

The first step in any variant prioritization workflow is to annotate the variant call set with the relevant scores. This can be done using annotation tools that add conservation scores and CADD scores to variant call format (VCF) files. The [Bioconductor](https://bioconductor.org/) project provides R packages for genomic annotation and analysis that can retrieve and append conservation scores to variant sets.

The annotation process should include all three scores when possible. PhyloP and GERP++ provide complementary views of evolutionary constraint, while CADD integrates these with other functional annotations. Having all three scores available allows for flexible filtering strategies and cross-validation of results.

### Step 2: Apply Initial Filters Based on Study Design

The thresholds used for conservation scores depend on the study design and the tolerance for false positives and false negatives. A study looking for rare disease-causing variants might use stringent thresholds to minimize the number of candidates, while a study of regulatory evolution might use more permissive thresholds to capture a broader range of potentially functional variants.

For PhyloP, a common approach is to select positions with positive scores, with higher scores indicating stronger conservation. Some studies use a threshold of PhyloP greater than 2 or 3, which corresponds to approximately 99th percentile conservation in some alignments. However, the appropriate threshold depends on the specific alignment and the distribution of scores in the dataset.

For GERP++, RS scores greater than 2 are often used as a threshold for constraint, with scores greater than 4 indicating strong constraint. These thresholds are based on the observation that most neutral positions have RS scores near zero, while constrained positions have positive scores that can range up to 6 or higher.

For CADD, the C score is often converted to a scaled score that ranges from 1 to 99, representing the rank of the variant relative to all possible substitutions in the genome. A scaled CADD score of 20 means the variant is in the top 10% of deleteriousness, while a score of 30 means it is in the top 1%. Many studies use a scaled CADD score of 15 or 20 as a threshold for further analysis.

### Step 3: Combine Scores for Cross-Validation

Using multiple conservation scores together provides more robust evidence than relying on any single score. A variant that is conserved according to both PhyloP and GERP++ and has a high CADD score is more likely to be functional than a variant that is conserved according to only one method.

One practical strategy is to require that a variant meet thresholds for at least two of the three scores. This reduces the impact of method-specific artifacts while retaining sensitivity for variants that are captured by different approaches. For example, a variant might have a high GERP++ score but a moderate PhyloP score due to differences in how the two methods handle alignment gaps or species composition. Requiring agreement between two methods helps identify variants with robust conservation signals.

The [Galaxy Training Network](https://training.galaxyproject.org/) provides accessible tutorials on variant annotation and filtering that demonstrate how to combine multiple annotation sources in reproducible workflows. These tutorials are useful for researchers who are new to variant prioritization or who want to implement best practices in their analysis pipelines.

### Step 4: Integrate with Population Frequency and Functional Annotations

Conservation scores should not be used in isolation. Population frequency data from large sequencing projects provides complementary information about whether a variant is likely to be tolerated. Variants that are common in the population are less likely to be highly deleterious, even if they occur at conserved positions. Conversely, rare variants at conserved positions are strong candidates for functional significance.

Functional annotations such as gene location, variant type, and predicted protein changes should also be considered. A synonymous variant at a conserved position may be less concerning than a missense variant at the same position, although synonymous variants can affect splicing or regulatory elements. The [SnpEff annotation method](https://pubmed.ncbi.nlm.nih.gov/35751823) provides a framework for combining these functional predictions with conservation scores and population frequency data to prioritize variants.

### Step 5: Document Filtering Decisions and Thresholds

Reproducibility requires careful documentation of all filtering decisions. The specific versions of the conservation scores, the genome build, and the thresholds applied should be recorded for each analysis. This documentation allows other researchers to understand how the final variant list was generated and to compare results across studies.

The [nf-core Documentation](https://nf-co.re/docs) provides standards for reproducible bioinformatics workflows that include version tracking and parameter documentation. Adopting these standards for variant prioritization analyses ensures that the filtering process can be repeated and audited.

## Options and Tradeoffs in Score Selection

### Choosing Between PhyloP and GERP++

PhyloP and GERP++ both measure evolutionary constraint, but they have different strengths and limitations. PhyloP provides per-position scores that can detect both conservation and acceleration, which makes it useful for identifying positions under positive selection. GERP++ provides RS scores that are more directly interpretable as measures of constraint, and the element-level information can help identify broader regions of functional importance.

In practice, many researchers use both scores and require agreement between them. This approach reduces false positives from method-specific artifacts. However, it can also reduce sensitivity, because some genuinely constrained positions may be missed by one method due to alignment or model differences.

The choice between PhyloP and GERP++ may also depend on the species being studied. Both methods were originally developed for human genome analysis, but they can be applied to other species if appropriate alignments and phylogenetic models are available. The [Dog10K consortium analysis](https://pubmed.ncbi.nlm.nih.gov/37582787) used Zoonomia phyloP constraint scores to prioritize functional variants in canine genomes, demonstrating that these methods can be adapted to non-human species.

### When CADD Is More Appropriate

CADD is the preferred choice when the goal is broad prioritization across all variant types and functional categories. Because CADD integrates many annotations, it can identify variants that are functional but not conserved, such as variants in recently evolved regulatory elements or variants that affect protein structure without being evolutionarily constrained.

CADD is also useful for scoring insertions and deletions, which cannot be scored by PhyloP or GERP++ in their standard implementations. Short indels can have significant functional consequences, and CADD provides a way to prioritize these variants alongside single-nucleotide variants.

The tradeoff is that CADD is a more complex model that is harder to interpret. The C score does not directly indicate why a variant was classified as deleterious, and the contribution of individual annotations to the score is not transparent. Researchers who need to understand the biological basis for prioritization may prefer to examine the individual annotations separately.

### Combining Conservation with Other Evidence Types

Conservation scores are most powerful when combined with other types of evidence. Population frequency, segregation in affected families, predicted protein effects, and experimental functional data all provide complementary information. The [critical assessment of variant prioritization methods](https://pubmed.ncbi.nlm.nih.gov/38685113) conducted within the Rare Genomes Project demonstrated that top-performing methods often incorporate multiple evidence types and sometimes include manual review.

The challenge of variant prioritization is that no single method is perfect. The diagnostic odyssey for rare disease families often lasts over five years, and causal variants are identified in under 50% of cases even with genome-wide sequencing. Computational methods can help narrow the search space, but they cannot replace careful clinical and biological interpretation.

## Observations and Measurements in Practice

### Score Distributions in Real Datasets

The distribution of conservation scores in a typical variant call set depends on the sequencing strategy and the population being studied. Whole-genome sequencing produces millions of variants, most of which are rare and have low conservation scores. Whole-exome sequencing produces fewer variants, but a higher proportion fall in conserved coding regions.

Researchers should examine the distribution of scores in their own data before applying thresholds. A variant set that is enriched for rare variants from a population with recent population growth may have a different score distribution than a variant set from a more stable population. The [EMBL-EBI Training](https://www.ebi.ac.uk/training) resources provide guidance on interpreting score distributions and selecting appropriate thresholds.

### Correlation Between Scores

PhyloP, GERP++, and CADD scores are correlated with each other, but the correlation is not perfect. Variants that are highly conserved according to all three methods are relatively rare, and the degree of agreement varies by genomic region and variant type. Coding variants tend to show higher agreement than non-coding variants, because coding regions are more likely to be under strong purifying selection.

The correlation between conservation scores and CADD is stronger for variants that are clearly deleterious, such as nonsense variants in essential genes, and weaker for variants with subtle effects. This is expected, because CADD integrates many annotations and may classify a variant as deleterious even if it is not strongly conserved.

### Performance in Benchmark Studies

Benchmark studies provide evidence about the performance of variant prioritization methods in real clinical settings. The Rare Genomes Project challenge evaluated 52 models from 16 teams, with top performers recalling causal variants in up to 13 of 14 solved families within the top 5 ranked variants. This demonstrates that computational methods can be highly effective when used appropriately.

However, the performance of any method depends on the specific dataset and the types of variants being sought. Methods that perform well for rare Mendelian diseases may perform differently for complex traits or for variants with moderate effect sizes. Researchers should evaluate methods in the context of their specific study design and variant types of interest.

## Records and Documentation for Reproducibility

### Tracking Score Versions and Genome Builds

Conservation scores are updated as new genome assemblies and alignments become available. The version of the score track and the genome build used for annotation must be recorded for each analysis. Using different versions of scores across studies can lead to inconsistent results and make comparisons difficult.

The [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/) provide versioned access to genome assemblies and annotation tracks, allowing researchers to identify the exact data used in their analyses. Recording these version identifiers in analysis documentation ensures that the analysis can be reproduced or updated when new data become available.

### Documenting Threshold Selection

The rationale for threshold selection should be documented for each analysis. Thresholds may be based on published recommendations, empirical examination of score distributions, or a balance between sensitivity and specificity that is appropriate for the study question. The documentation should include the specific threshold values, the proportion of variants passing each threshold, and the rationale for the chosen values.

### Maintaining Analysis Pipelines

Reproducible analysis pipelines should be version-controlled and documented. The [nf-core Documentation](https://nf-co.re/docs) provides standards for pipeline development that include version tracking, parameter documentation, and containerization. Adopting these standards for variant prioritization workflows ensures that analyses can be rerun and audited.

The [Bioconductor](https://bioconductor.org/) project provides R packages that support reproducible genomic analysis, including packages for retrieving and annotating conservation scores. Using these packages within a version-controlled workflow ensures that the analysis is reproducible and that the specific versions of the scores are recorded.

## Common Failure Patterns in Variant Prioritization

### Overreliance on a Single Score

A common mistake is to rely on a single conservation score for variant prioritization. Each method has specific limitations, and variants that are important may be missed by any one method. Using multiple scores and requiring agreement between them reduces this risk.

### Applying Inappropriate Thresholds

Thresholds that are appropriate for one study design may not be appropriate for another. A threshold that is too stringent will miss true functional variants, while a threshold that is too permissive will produce too many candidates for manual review. Thresholds should be selected based on the study question and the tolerance for false positives and false negatives.

### Ignoring Population Frequency

Conservation scores do not account for population frequency. A variant that is common in the population is unlikely to be highly deleterious, even if it occurs at a conserved position. Integrating population frequency data with conservation scores provides a more complete picture of variant impact.

### Failing to Validate with Functional Data

Conservation scores provide computational evidence, but they do not prove that a variant is functional. Variants that pass conservation-based filters should be validated with functional assays when possible. The [critical assessment of variant prioritization methods](https://pubmed.ncbi.nlm.nih.gov/38685113) demonstrated that manual review and functional interpretation can improve the performance of computational methods.

### Misinterpreting Negative Scores

Negative PhyloP scores indicate acceleration, not absence of function. A position with a negative PhyloP score may be under positive selection or may be functional in a species-specific manner. Researchers should not automatically exclude variants at positions with negative PhyloP scores, particularly when studying traits that may be under recent selection.

## Limitations of Conservation Scores

### Conservation Does Not Equal Function

A conserved position is more likely to be functional than a non-conserved position, but conservation is not proof of function. Some conserved positions may be conserved for reasons unrelated to function, such as biased gene conversion or low mutation rates. Conversely, some functional positions may not be conserved because they have evolved new functions in specific lineages.

### Species-Specific Functional Elements

Functional elements that are unique to a species will not be conserved across species. Human-specific regulatory elements, for example, may not be detected by conservation-based methods. CADD partially addresses this limitation by incorporating annotations that do not depend on conservation, but no computational method can fully capture species-specific function.

### Alignment and Model Dependencies

Conservation scores depend on the quality of the multiple sequence alignment and the accuracy of the phylogenetic model. Errors in alignment or tree construction can produce incorrect scores. Standard precomputed tracks use carefully curated alignments, but custom analyses require careful quality control.

### Limited Resolution for Non-Coding Variants

Conservation scores are less informative for non-coding variants than for coding variants. Non-coding regions are generally less conserved than coding regions, and the functional significance of non-coding conservation is harder to interpret. CADD incorporates regulatory annotations that can help prioritize non-coding variants, but the interpretation remains challenging.

### Score Saturation and Ceiling Effects

Some conservation scores have ceiling effects, where the maximum score is reached for positions that are perfectly conserved across all species in the alignment. This limits the ability to distinguish between strongly conserved and extremely strongly conserved positions. CADD scores are less affected by this limitation because they integrate multiple annotations.

## Safety and Regulatory Context

### Clinical Interpretation Requires Caution

Conservation scores are computational predictions, not clinical diagnoses. Variants that are prioritized by conservation scores must be interpreted in the context of clinical findings, family history, and other evidence before they are used for clinical decision-making. The [critical assessment of variant prioritization methods](https://pubmed.ncbi.nlm.nih.gov/38685113) highlights the challenges of achieving genetic diagnoses and the importance of careful interpretation.

### Reporting Standards for Research Studies

Research studies that use conservation scores for variant prioritization should report the specific scores, versions, and thresholds used. This transparency allows other researchers to evaluate the evidence and to compare results across studies. The [SnpEff annotation framework](https://pubmed.ncbi.nlm.nih.gov/35751823) provides a structured approach to variant annotation that supports transparent reporting.

### Data Sharing and Reproducibility

Conservation score data are publicly available from multiple sources, and the methods for calculating scores are published. Researchers should share their analysis pipelines and parameters to support reproducibility. The [Galaxy Training Network](https://training.galaxyproject.org/) and [nf-core Documentation](https://nf-co.re/docs) provide resources for building and sharing reproducible analysis workflows.

## Professional Escalation Criteria

### When to Seek Expert Consultation

Researchers should consider consulting with a bioinformatics specialist or clinical geneticist when conservation scores produce conflicting results, when the variant of interest falls in a poorly annotated region, or when the study involves clinical interpretation. Expert consultation can help interpret ambiguous results and avoid common pitfalls.

### When to Re-evaluate the Analysis

The analysis should be re-evaluated when new versions of conservation scores become available, when the genome assembly is updated, or when new population frequency data are released. Re-evaluation ensures that the prioritization reflects the current state of knowledge.

### When to Use Additional Methods

Additional methods should be considered when conservation scores do not provide sufficient discrimination, when the variant set includes many non-coding variants, or when the study involves complex traits with many contributing variants. Methods that integrate additional annotation types or that use machine learning approaches may provide better performance in these contexts.

## Decision Framework for Triaging Variants with Discordant Conservation Scores

### The Problem of Score Discordance

Researchers frequently encounter variants where PhyloP, GERP++, and CADD disagree. A variant may show strong conservation according to GERP++ but a modest PhyloP score, or a high CADD score despite weak evolutionary constraint. These discordant results create practical uncertainty about whether to advance the variant for validation or discard it from consideration. instead of treating discordance as a failure of the scoring methods, researchers can use a structured decision framework that assigns different evidentiary weight depending on the variant class, genomic context, and study objective.

### Tiered Classification Based on Score Agreement

A practical approach is to classify variants into three tiers based on the level of agreement among the three scoring systems. Tier 1 variants meet thresholds for all three scores, indicating robust evidence of functional potential. Tier 2 variants meet thresholds for at least two scores, suggesting moderate evidence that warrants further examination. Tier 3 variants meet thresholds for only one score or none, indicating weak or conflicting evidence that should be deprioritized unless other factors justify investigation.

The tier assignment should be recorded for each variant in the analysis output. This classification provides a transparent basis for deciding which variants to carry forward into validation studies, family segregation analysis, or functional assays. The [critical assessment of variant prioritization methods](https://pubmed.ncbi.nlm.nih.gov/38685113) conducted within the Rare Genomes Project demonstrated that top-performing methods often combine multiple evidence types and that manual review can improve performance, which supports the use of tiered classification as a structured approach to manual review.

### Context-Dependent Weighting of Scores

The relative importance of each score should vary by genomic context. For protein-coding variants, CADD often provides the most informative signal because it integrates protein structure predictions and regulatory annotations alongside conservation. For non-coding variants, PhyloP and GERP++ may be more informative because they directly measure evolutionary constraint in regulatory regions, while CADD's integration of diverse annotations can sometimes obscure the conservation signal.

For synonymous variants, conservation scores become particularly important because CADD may not strongly prioritize variants that do not alter the amino acid sequence. A synonymous variant at a highly conserved position may affect splicing regulatory elements or codon usage, and the conservation evidence provides the primary justification for further investigation. The [SnpEff annotation framework](https://pubmed.ncbi.nlm.nih.gov/35751823) provides a structured approach to combining functional predictions with conservation scores, which is useful for evaluating synonymous and other non-coding variants.

### Decision Rules for Discordant Scores

When scores disagree, specific decision rules can guide the interpretation. If PhyloP and GERP++ agree on conservation but CADD is low, the variant may be conserved for reasons unrelated to variant deleteriousness, such as structural constraints on the DNA sequence itself. This scenario warrants caution before discarding the variant, particularly for non-coding variants where CADD's training data may be less comprehensive.

If CADD is high but both conservation scores are low, the variant may be functional through mechanisms that are not evolutionarily constrained, such as recently evolved regulatory elements or variants that affect protein structure without being subject to purifying selection across species. This scenario is more common in human-specific regulatory regions and in genes that have undergone recent adaptive evolution. The [evolutionary analyses of head-to-body ratio variants](https://pubmed.ncbi.nlm.nih.gov/41444482) demonstrated that trait-associated variants can be enriched in conserved regions and human accelerated regions, highlighting the importance of considering both conservation and acceleration signals.

If PhyloP and GERP++ disagree with each other, the discrepancy often reflects methodological differences in how the two methods handle alignment gaps, species composition, or model parameters. In this case, the CADD score can serve as a tiebreaker, because it integrates many annotations and may capture information that resolves the ambiguity.

### Implementing the Framework in Practice

The decision framework should be implemented as a structured annotation step in the variant prioritization pipeline. After annotating variants with PhyloP, GERP++, and CADD scores, researchers can assign tier classifications using a simple set of rules encoded in the analysis script. The [Bioconductor](https://bioconductor.org/) project provides R packages that support this type of structured annotation and classification, allowing researchers to implement the framework reproducibly.

The framework should also include a step for recording the rationale for any manual overrides of the tier assignment. If a Tier 3 variant is advanced for validation because of strong family segregation evidence or a compelling biological hypothesis, this decision should be documented. The [nf-core Documentation](https://nf-co.re/docs) provides standards for reproducible bioinformatics workflows that include parameter documentation and version tracking, which supports this type of audit trail.

### Validation of the Framework with Known Variants

Before applying the decision framework to a novel variant set, researchers should validate it using variants with known functional consequences. This validation can be performed using a set of confirmed pathogenic variants and a set of common benign variants. The framework should correctly classify the pathogenic variants into Tier 1 or Tier 2 and the benign variants into Tier 3.

The validation step provides an empirical check on the threshold values and the tier assignment rules. If the framework fails to classify known pathogenic variants appropriately, the thresholds or the weighting scheme should be adjusted. This iterative refinement ensures that the framework is calibrated for the specific study context and variant types of interest.

### Limitations of the Decision Framework

The decision framework does not eliminate the need for biological interpretation. A Tier 1 classification indicates strong computational evidence of functional potential, but it does not prove that a variant causes disease or has a measurable functional effect. Conversely, a Tier 3 classification does not prove that a variant is benign, particularly for variants in poorly annotated regions or for traits that are not well captured by existing annotations.

The framework also assumes that the three scoring systems are available for the species and genome build being analyzed. For non-human species, the availability of precomputed scores may be limited. The [Dog10K consortium analysis](https://pubmed.ncbi.nlm.nih.gov/37582787) used Zoonomia phyloP constraint scores for canine genomes, demonstrating that conservation scores can be adapted to non-human species, but researchers working with less well-characterized species may need to generate custom scores or rely on a reduced set of annotations.

### Integration with Population Frequency and Segregation Data

The decision framework should be applied alongside population frequency filtering and family segregation analysis. A Tier 1 variant that is common in the population is less likely to be highly deleterious than a Tier 1 variant that is rare or absent from population databases. Similarly, a Tier 2 variant that segregates with disease in an affected family may be more compelling than a Tier 1 variant that does not segregate.

The [critical assessment of variant prioritization methods](https://pubmed.ncbi.nlm.nih.gov/38685113) demonstrated that top-performing methods in the Rare Genomes Project challenge often incorporated multiple evidence types and sometimes included manual review. The decision framework provides a structured approach to combining conservation scores with these other evidence types, supporting the type of integrated analysis that performs best in benchmark studies.

### Documentation and Reporting of Tier Assignments

The tier assignment for each variant should be included in the analysis output and reported in publications. This documentation allows other researchers to understand how variants were prioritized and to compare results across studies. The report should include the specific thresholds used for each score, the tier assignment rules, and the number of variants in each tier.

The [Galaxy Training Network](https://training.galaxyproject.org/) provides accessible tutorials on variant annotation and filtering that demonstrate how to document filtering decisions in reproducible workflows. Adopting these practices ensures that the decision framework can be audited and refined as new data and methods become available.

## Frequently Asked Questions

### What is the difference between PhyloP and GERP++?

PhyloP evaluates each nucleotide position independently using a phylogenetic model, producing scores that can be positive (conserved) or negative (accelerated). GERP++ first identifies constrained elements and then estimates the rejected substitution score at each position, which represents the difference between observed and expected substitutions. Both measure evolutionary constraint, but they use different statistical frameworks and provide different types of information.

### How should I choose thresholds for conservation scores?

Thresholds should be based on the study design and the tolerance for false positives and false negatives. Common thresholds include PhyloP greater than 2 or 3, GERP++ RS scores greater than 2 or 4, and scaled CADD scores greater than 15 or 20. Researchers should examine the distribution of scores in their own data and select thresholds that balance sensitivity and specificity for their specific question.

### Can conservation scores be used for non-human species?

Yes, conservation scores can be calculated for any species with appropriate multiple sequence alignments and phylogenetic models. The [Dog10K consortium](https://pubmed.ncbi.nlm.nih.gov/37582787) used Zoonomia phyloP constraint scores to prioritize functional variants in canine genomes. However, the quality of the scores depends on the availability of suitable alignments and the accuracy of the phylogenetic model.

### Why do PhyloP and GERP++ sometimes give different results?

PhyloP and GERP++ use different statistical frameworks and may handle alignment gaps, species composition, and model parameters differently. A position may be classified as conserved by one method but not the other due to these methodological differences. Using both scores and requiring agreement can reduce method-specific artifacts.

### What does a negative PhyloP score mean?

A negative PhyloP score indicates that a position has more substitutions than expected under neutral evolution, which suggests accelerated evolution or relaxed constraint. This does not necessarily mean the position is non-functional. Accelerated regions can be functionally important, particularly in traits that have undergone recent adaptive evolution.

### How does CADD differ from conservation scores?

CADD integrates many annotations, including conservation, regulatory elements, protein structure, and population frequency, into a single score using a support vector machine. Conservation scores use only evolutionary sequence data. CADD can identify functional variants that are not conserved, and it can score insertions and deletions, which conservation scores cannot.

### Should I use conservation scores for non-coding variants?

Conservation scores can be used for non-coding variants, but they are less informative than for coding variants. Non-coding regions are generally less conserved, and the functional significance of conservation in non-coding regions is harder to interpret. CADD incorporates regulatory annotations that can help prioritize non-coding variants.

### How do I document conservation score usage for publication?

Record the specific score versions, genome build, alignment used, and thresholds applied. Describe the rationale for threshold selection and report the number of variants passing each filter. This documentation allows other researchers to reproduce the analysis and to compare results across studies.

## Related Bioinformatics Guides

- [Detecting Structural Variants with Long-Read Sequencing: Methods and Considerations](/knowledge/bioinformatics/detecting-structural-variants-with-long-read-sequencing-methods-and-considerations)
- [Genomic Data Analysis Tools: A Comparative Guide for Researchers](/knowledge/bioinformatics/genomic-data-analysis-tools-a-comparative-guide-for-researchers)
- [Functional Annotation of Metagenomes: A Guide to Databases and Pipelines](/knowledge/bioinformatics/functional-annotation-of-metagenomes-a-guide-to-databases-and-pipelines)
- [Proteomics Analysis Tools: A Comparative Guide for Functional Interpretation](/knowledge/bioinformatics/proteomics-analysis-tools-a-comparative-guide-for-functional-interpretation)
- [Digital Pathology Scanners: A Buyer's Guide for Clinical and Research Use](/knowledge/bioinformatics/digital-pathology-scanners-a-buyer-s-guide-for-clinical-and-research-use)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)


## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [A general framework for estimating the relative pathogenicity of human genetic variants.](https://pubmed.ncbi.nlm.nih.gov/24487276). Nature genetics, 2014.
- [Variant Annotation and Functional Prediction: SnpEff.](https://pubmed.ncbi.nlm.nih.gov/35751823). Methods in molecular biology (Clifton, N.J.), 2022.
- [Critical assessment of variant prioritization methods for rare disease diagnosis within the rare genomes project.](https://pubmed.ncbi.nlm.nih.gov/38685113). Human genomics, 2024.
- [Genome sequencing of 2000 canids by the Dog10K consortium advances the understanding of demography, genome function and architecture.](https://pubmed.ncbi.nlm.nih.gov/37582787). Genome biology, 2023.
- [Genetic Insights into Head-to-Body Ratios Via Deep Learning-Based Image Segmentation and Implications for Common Diseases.](https://pubmed.ncbi.nlm.nih.gov/41444482). Nature communications, 2025.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.