Purifying vs Positive Selection: Mechanisms and Detection

By Dr. Zubair Khalid, DVM, MS, PhD ·

Purifying vs Positive Selection: Mechanisms and Detection

Introduction to Purifying and Positive Selection

Definitions and Core Concepts

Natural selection operates on heritable phenotypic variation through two primary directional modes. Purifying selection (also called negative selection) removes deleterious alleles from a population because they reduce organismal fitness. Positive selection (also called Darwinian selection or directional selection) increases the frequency of beneficial alleles because they confer a fitness advantage. Both processes are fundamental to Evolution by Natural Selection, yet they leave profoundly different genomic footprints and require distinct analytical approaches to detect.

The distinction matters at multiple levels. At the level of an individual nucleotide site, purifying selection acts when a new mutation is harmful—perhaps it destabilizes a protein fold, disrupts a splice site, or alters a regulatory element. Positive selection acts when a mutation is beneficial—perhaps it confers resistance to a pathogen, enables exploitation of a new food source, or improves enzymatic efficiency under novel environmental conditions. Between these extremes lies a vast middle ground of nearly neutral mutations whose fate is governed primarily by genetic drift, as formalized by the Neutral Theory of Molecular Evolution.

A critical conceptual point is that selection acts on phenotypes, not genotypes directly. A mutation is "deleterious" or "beneficial" only in the context of the organism's environment, genetic background, and life history. A mutation that is neutral in one context may be strongly selected in another. Moreover, the same allele can be subject to purifying selection in one tissue and positive selection in another if it has pleiotropic effects.

The dN/dS Ratio as a Starting Point

The most widely used metric for quantifying selection at the sequence level is the dN/dS ratio (also denoted ω or Ka/Ks). This ratio compares the rate of nonsynonymous substitutions per nonsynonymous site (dN) to the rate of synonymous substitutions per synonymous site (dS). Synonymous substitutions do not alter the encoded amino acid and are generally assumed to be selectively neutral or nearly so, providing a molecular clock calibrated against the mutation rate. Nonsynonymous substitutions change the amino acid and are therefore the substrate for selection at the protein level.

The logic is straightforward:

  • dN/dS < 1: Nonsynonymous substitutions are being removed by purifying selection. Most protein-coding genes in most organisms show values between 0.05 and 0.3, reflecting strong constraint on amino acid sequence.
  • dN/dS = 1: Nonsynonymous substitutions accumulate at the same rate as synonymous ones, consistent with neutral evolution or a complete absence of selection at the protein level.
  • dN/dS > 1: Nonsynonymous substitutions are being fixed at a higher rate than synonymous ones, indicating positive selection favoring amino acid change.

While the dN/dS ratio is a powerful starting point, it is a genome-wide or gene-wide average. A gene with a dN/dS of 0.2 overall may harbor individual codons under strong positive selection—for example, the antigen-binding residues of major histocompatibility complex genes or the active sites of host-defense proteins. Detecting such episodic or site-specific selection requires more sophisticated codon-based models, discussed in the Statistical Methods section.

Molecular Mechanisms Underlying Each Mode of Selection

Purifying Selection: Removing Deleterious Alleles

Purifying selection operates at multiple molecular levels. At the protein level, most amino acid substitutions are deleterious because they destabilize the folded structure, disrupt catalytic residues, or interfere with protein-protein interactions. The magnitude of the fitness effect varies widely. Some substitutions are lethal—for example, mutations that eliminate the catalytic serine in serine proteases or that truncate a protein before its functional domain. Others are mildly deleterious, reducing fitness by a fraction of a percent—sufficient to be removed by selection in large populations but potentially fixed by drift in small ones.

Consider a concrete example: the human β-globin gene (HBB). The vast majority of nonsynonymous mutations in HBB are deleterious, causing hemoglobinopathies such as sickle-cell anemia (Glu6Val) or β-thalassemia (various frameshift and nonsense mutations). These mutations are maintained at low frequency in populations primarily because of heterozygote advantage in malaria-endemic regions—a special case of balancing selection—but the homozygous states are strongly deleterious. Purifying selection removes most HBB variants from the population within a few generations.

At the regulatory level, purifying selection acts on promoter sequences, enhancers, splice sites, and untranslated regions. Mutations that disrupt transcription factor binding sites or alter mRNA stability are generally deleterious. The extent of constraint on regulatory DNA is often comparable to that on protein-coding sequence, as revealed by conservation scores such as phyloP and phastCons (discussed below).

At the genomic level, purifying selection also acts on noncoding RNAs, Conserved Sequence elements of unknown function, and even on synonymous sites when they overlap splicing enhancers or RNA secondary structures. The latter point is important: the assumption that synonymous sites are neutral is an approximation. Codon usage bias—the preferential use of certain synonymous codons—reflects selection for translational efficiency and accuracy, particularly in highly expressed genes.

Positive Selection: Fixing Beneficial Alleles

Positive selection acts when a new mutation confers a fitness advantage. The molecular mechanisms are diverse. At the protein level, beneficial mutations may:

  • Alter substrate specificity of an enzyme (e.g., the evolution of new metabolic capabilities after Gene Duplication)
  • Improve binding affinity of a receptor for its ligand
  • Confer resistance to inhibitors or toxins
  • Enable evasion of host immune recognition by pathogens
  • Adapt proteins to new environmental conditions (temperature, pH, salinity)

A classic example is the evolution of antifreeze proteins in Arctic and Antarctic fish. These proteins bind to ice crystals and prevent their growth, allowing fish to survive in freezing waters. The antifreeze protein genes arose through duplication and rapid divergence, with positive selection acting on the duplicated copies to optimize ice-binding activity.

At the regulatory level, positive selection can act on promoter or enhancer sequences to alter gene expression patterns. The evolution of lactase persistence in human populations is a well-documented case: regulatory mutations that maintain lactase expression into adulthood were strongly favored in pastoralist populations. Similarly, changes in the expression of detoxification enzymes have accompanied the evolution of insecticide resistance in many insect species.

Positive selection can also act on genomic features such as Transposable Element insertions that bring new regulatory sequences under the control of existing genes, or on gene duplicates that acquire new functions (neofunctionalization). The key distinction from purifying selection is that the selected allele increases in frequency because it improves fitness, not merely because it avoids being deleterious.

Population Genetics Framework

Selection Coefficient and Fitness Effects

The population genetics framework formalizes selection in terms of fitness. Consider a diploid population with two alleles at a locus, A (ancestral) and a (derived). If the relative fitnesses of genotypes AA, Aa, and aa are 1, 1 + hs, and 1 + s, respectively, then s is the selection coefficient and h is the dominance coefficient. For a beneficial allele, s > 0; for a deleterious allele, s < 0. The dominance coefficient h describes whether the fitness effect is recessive (h = 0), additive (h = 0.5), or dominant (h = 1).

The selection coefficient determines the probability of fixation of a new mutation. For a newly arising beneficial allele with additive effects (h = 0.5) in a population of effective size Ne, the fixation probability is approximately 2s for small s. This is roughly twice the neutral fixation probability of 1/(2Ne). For a deleterious allele with selection coefficient -s, the fixation probability is approximately 2s/(e^(4Nes) - 1), which becomes vanishingly small when Nes >> 1.

The key parameter is not s alone but the product Nes. When Nes << 1, selection is ineffective and the allele behaves as effectively neutral. When Nes >> 1, selection dominates over drift. This is why Positive and Negative Selection are best understood in the context of population size: a mildly deleterious allele (s = -10⁻⁵) is effectively neutral in a population with Ne = 10³ but strongly selected against in a population with Ne = 10⁶.

Effective Population Size and Its Impact

The effective population size (Ne) is the size of an ideal Wright-Fisher population that would experience the same rate of genetic drift as the actual population. Ne is typically much smaller than the census population size due to fluctuations in population size, variance in reproductive success, and other factors.

Ne has profound consequences for the efficiency of selection. In large populations (high Ne), purifying selection is highly efficient: even mildly deleterious mutations are removed, and beneficial mutations are fixed rapidly. In small populations (low Ne), drift dominates, and slightly deleterious mutations can accumulate. This is why genome-wide patterns of selection differ between species with different Ne. For example, the human genome shows evidence of slightly deleterious nonsynonymous variants segregating at low frequency, consistent with the relatively small Ne of humans (~10⁴) compared to Drosophila (~10⁶).

The interaction between Ne and selection also affects the dN/dS ratio. In species with small Ne, the dN/dS ratio is inflated because slightly deleterious mutations are fixed by drift. Conversely, in species with large Ne, purifying selection is more efficient, and dN/dS ratios are lower. This must be considered when comparing selection pressures across species.

Genomic Signatures of Purifying Selection

Reduced Nucleotide Diversity

Purifying selection reduces nucleotide diversity at and around selected sites. When a deleterious allele is removed from the population, linked neutral variation is also removed—a process called background selection. The magnitude of this effect depends on the recombination rate: in regions of low recombination, background selection can reduce diversity over large genomic distances, whereas in regions of high recombination, the effect is localized to the immediate vicinity of the selected site.

The signature of background selection is a reduction in diversity (π, the average number of pairwise differences per site) that is correlated with the density of conserved elements and inversely correlated with recombination rate. Genomic regions with high gene density and strong constraint show reduced π compared to gene-poor regions. This is observable in essentially every eukaryotic genome examined.

Conservation Scores and PhyloP/PhastCons

Cross-species sequence conservation is a powerful signature of purifying selection. If a nucleotide site has remained identical across species that diverged hundreds of millions of years ago, the most parsimonious explanation is that mutations at that site are deleterious and have been repeatedly removed by purifying selection.

Two widely used conservation scores are phyloP and phastCons, both computed from multiple sequence alignments of related species. PhyloP scores are based on a phylogenetic hidden Markov model that estimates the probability of observing the pattern of substitutions at each site under a neutral model. Negative phyloP scores indicate faster-than-neutral evolution (possible positive selection), while positive scores indicate conservation (purifying selection). PhastCons uses a different approach, modeling conserved and non-conserved states along the alignment and computing the posterior probability that each site is in the conserved state.

These scores are essential for identifying functionally important elements in the genome. For example, the ENCODE project used phyloP and phastCons to identify candidate cis-regulatory elements, and the vast majority of sites with high conservation scores are under strong purifying selection. However, conservation scores have limitations: they cannot distinguish between constraint due to purifying selection and constraint due to low mutation rate, and they are sensitive to alignment errors and to the choice of species included in the alignment.

Genomic Signatures of Positive Selection

Selective Sweeps and Linkage Disequilibrium

When a beneficial allele is driven to fixation by positive selection, linked neutral variation is carried along—a phenomenon called a selective sweep. The signature of a completed sweep is a local reduction in nucleotide diversity, an excess of low-frequency derived alleles, and a specific pattern of linkage disequilibrium (LD) with high-frequency derived alleles.

The classic signature of a recent sweep is a "valley" of reduced heterozygosity flanked by regions of normal diversity. The width of this valley depends on the recombination rate and the strength of selection: strong selection and low recombination produce broad valleys, while weak selection and high recombination produce narrow valleys. In the extreme case of a complete sweep in a region of no recombination, all variation is eliminated except that introduced by new mutation since the sweep.

Detecting sweeps from population genomic data typically involves scanning for regions with:

  • Reduced π relative to the genomic background
  • Negative values of Tajima's D (excess of low-frequency variants)
  • High levels of LD among derived alleles
  • Extended haplotype homozygosity (EHH) in a subset of chromosomes

The latter signature, detected using statistics like iHS (integrated haplotype score) or XP-EHH (cross-population EHH), is particularly useful for detecting incomplete sweeps where the beneficial allele has not yet reached fixation.

High dN/dS and Fast-Evolving Sites

At the sequence level, positive selection is detected as an elevated rate of nonsynonymous substitution. The simplest approach is to compare dN/dS between lineages: a lineage that has undergone adaptive evolution will show a higher dN/dS than the background. For example, the dN/dS ratio of the human lineage is elevated for genes involved in immune defense, olfaction, and reproduction, consistent with ongoing positive selection in these functional categories.

Site-specific approaches identify individual codons under positive selection. The branch-site model implemented in PAML (Phylogenetic Analysis by Maximum Likelihood) allows dN/dS to vary both among sites and among lineages, enabling the detection of episodic positive selection affecting a few codons in a specific lineage. This approach has been used to identify positively selected sites in host-defense proteins, viral surface proteins, and reproductive proteins.

A cautionary note: high dN/dS can also arise from relaxation of purifying selection rather than positive selection. If a gene is no longer functionally constrained—for example, after Gene Duplication when one copy is free to accumulate mutations—the dN/dS ratio will rise toward 1 but rarely exceed it. Distinguishing relaxation from positive selection requires careful statistical testing, typically by comparing models that allow dN/dS > 1 (positive selection) against models that constrain dN/dS ≤ 1 (relaxation or neutrality).

Statistical Methods for Detecting Selection

Codon Substitution Models (PAML, HyPhy)

Codon substitution models are the gold standard for detecting selection in protein-coding sequences. These models describe the substitution process at the codon level, allowing dN/dS (ω) to vary among sites, among lineages, or both. The most widely used implementations are in the PAML package (codeml program) and HyPhy.

The standard approach in PAML involves comparing nested models using likelihood ratio tests:

  1. M0 (one ω for all sites) vs M3 (discrete distribution of ω across sites): tests for variation in selection pressure among sites.
  2. M1a (two site classes: ω < 1 and ω = 1) vs M2a (adds a third class with ω > 1): tests for the presence of positively selected sites.
  3. M7 (beta distribution of ω, constrained to 0-1) vs M8 (beta distribution plus an extra class with ω > 1): a more flexible test for positive selection.

When the likelihood ratio test is significant, the Bayes empirical Bayes (BEB) approach is used to identify individual sites with high posterior probability of belonging to the ω > 1 class. A typical analysis of a gene family of 20-50 sequences with 500-1000 codons requires an alignment of the coding sequences, a phylogenetic tree, and a few minutes of computation.

HyPhy offers complementary approaches, including the mixed effects model of evolution (MEME) for detecting episodic selection at individual sites and the branch-site unrestricted statistical test for variable selection among lineages. MEME is particularly useful for detecting selection that affects only a subset of lineages at a given site, which is common in host-pathogen coevolution.

McDonald-Kreitman Test

The McDonald-Kreitman (MK) test is a population genetics approach that compares polymorphism within species to divergence between species. The test is based on a 2×2 contingency table:

SynonymousNonsynonymous
Polymorphic (within species)PsPn
Fixed (between species)DsDn

Under neutrality, the ratio of nonsynonymous to synonymous polymorphism should equal the ratio of nonsynonymous to synonymous divergence: Pn/Ps = Dn/Ds. The neutrality index (NI) is calculated as (Pn/Ps)/(Dn/Ds). An NI < 1 indicates an excess of nonsynonymous divergence, consistent with positive selection fixing beneficial amino acid changes. An NI > 1 indicates an excess of nonsynonymous polymorphism, consistent with purifying selection removing deleterious alleles before they reach fixation.

The MK test is powerful because it does not require a species tree or an estimate of the mutation rate. However, it is sensitive to demography: population expansion creates an excess of rare variants, which can bias the polymorphism counts. The test also assumes that synonymous sites are neutral, which may be violated in genes with strong codon usage bias.

Site Frequency Spectrum-Based Tests

The site frequency spectrum (SFS) describes the distribution of allele frequencies across polymorphic sites. Purifying selection skews the SFS toward rare alleles because deleterious mutations are removed before they reach high frequency. Positive selection during a sweep also creates an excess of rare alleles, but the pattern differs in detail.

Tajima's D compares two estimators of the population mutation rate θ: π (average pairwise differences) and Watterson's estimator θW (based on the number of segregating sites). Under neutrality, D ≈ 0. Negative D indicates an excess of rare alleles, consistent with purifying selection, a recent selective sweep, or population expansion. Positive D indicates an excess of intermediate-frequency alleles, consistent with balancing selection or population contraction.

Fay and Wu's H is a more specific test for selective sweeps. It uses the frequency of derived alleles, weighting high-frequency derived alleles more heavily. A selective sweep produces an excess of high-frequency derived alleles (because they hitchhiked with the beneficial mutation), giving a significantly negative H. This test is less affected by demography than Tajima's D but requires accurate inference of the ancestral state, which can be problematic in regions of high mutation rate or alignment error.

A practical note: SFS-based tests are best applied to whole-genome or large genomic datasets where the genomic background distribution of the statistic can be estimated. A single Tajima's D value from a 1 kb locus is essentially uninterpretable; the same value from a 100 kb window in a whole-genome scan is meaningful in context.

Challenges and Pitfalls in Distinguishing Selection Modes

Demography vs Selection

The most pervasive challenge in detecting selection is distinguishing it from demographic history. Population expansions, contractions, bottlenecks, and migration all shape the SFS and patterns of diversity in ways that mimic selection.

For example, a population bottleneck followed by expansion produces an excess of rare variants (negative Tajima's D) across the entire genome, mimicking the signature of purifying selection or a selective sweep. Conversely, population structure can produce positive Tajima's D in some regions. The key diagnostic is genomic scale: demography affects all loci genome-wide, while selection affects specific loci or regions. This is why genome-wide scans are essential—a single locus with negative Tajima's D is uninformative, but a locus that is an extreme outlier in a genome-wide distribution is a candidate for selection.

Statistical approaches that jointly infer demography and selection, such as the composite likelihood method of SweepFinder or the ABC-based approach of ∂a∂i, can help disentangle these factors. However, they require assumptions about the demographic model that may not hold, and misspecification of the demographic model can lead to false positives.

Recombination and Linked Selection

Recombination breaks down linkage between selected and neutral sites, but the rate of recombination varies across the genome. In regions of low recombination, the effects of selection extend over large distances, creating broad genomic regions that appear to be under selection when they are merely linked to a selected site. This is the problem of linked selection, which includes both background selection (linked to deleterious alleles) and hitchhiking (linked to beneficial alleles).

The practical consequence is that a genomic region with reduced diversity may be interpreted as a selective sweep when it is actually a region of low recombination subject to background selection. Conversely, a region of high recombination may show no signature of selection even when a strongly selected site is present, because recombination has broken down the association.

Methods that account for recombination rate variation, such as the composite likelihood approach of SweepFinder2 or the machine learning approach of S/HIC, can help. However, recombination rate estimates are often imprecise, particularly in non-model organisms, and errors in recombination rate can propagate to errors in selection inference.

Gene Annotation and Alignment Errors

Selection inference is only as good as the underlying sequence data. Errors in gene annotation—incorrect exon-intron boundaries, missing exons, or misidentified start codons—can introduce frameshifts and premature stop codons that are misinterpreted as evidence of positive selection. Similarly, alignment errors in codon-based analyses can create spurious nonsynonymous substitutions.

A common failure mode is the inclusion of pseudogenes or misannotated gene fragments in a coding sequence alignment. Pseudogenes are not under purifying selection and show elevated dN/dS, which can be misinterpreted as positive selection. Careful quality control—checking for open reading frames, verifying splice sites, and using alignment programs that account for codon structure (e.g., MACSE, PRANK)—is essential.

Another subtle issue is the choice of outgroup for ancestral state inference. If the outgroup is too divergent, multiple substitutions can obscure the ancestral state, leading to errors in derived allele frequency estimation and biased SFS-based tests.

Practical Workflow for Analyzing Selection in Your Data

Data Requirements and Quality Control

Before any selection analysis, ensure your data meet the following requirements:

  1. Coding sequence alignments: For codon-based analyses, you need aligned coding sequences with correct reading frames. Use a codon-aware aligner (MACSE, PRANK, or PAL2NAL) rather than a protein aligner followed by back-translation.
  2. Phylogenetic tree: For PAML and HyPhy analyses, you need a phylogenetic tree. The tree can be estimated from the data (e.g., using RAxML or IQ-TREE on the concatenated alignment) or obtained from the literature.
  3. Population genomic data: For SFS-based tests and sweep detection, you need genotype calls for multiple individuals from one or more populations. The minimum is typically 10-20 individuals for reliable SFS estimation, though more is better.
  4. Ancestral state information: For Fay and Wu's H and related tests, you need to infer the ancestral allele at each polymorphic site, usually from an outgroup species.

Quality control steps include:

  • Remove sequences with premature stop codons or frameshifts (unless you are explicitly studying pseudogenes).
  • Check for saturation at synonymous sites (if dS is saturated, dN/dS estimates are unreliable).
  • Verify that the alignment does not contain misaligned codons, particularly in regions of low sequence similarity.

Choosing the Right Test

The choice of test depends on your question and data type:

QuestionData TypeRecommended Test
Is a gene under positive selection overall?Coding sequences from multiple speciesPAML M1a vs M2a, or M7 vs M8
Which codons are under positive selection?Coding sequences from multiple speciesPAML BEB, HyPhy MEME
Is a specific lineage under positive selection?Coding sequences with known treePAML branch-site model
Is there evidence of adaptive evolution within a species?Population polymorphism + divergenceMcDonald-Kreitman test
Has a recent selective sweep occurred?Population genomic dataiHS, XP-EHH, SweepFinder
Is a genomic region under purifying selection?Multiple species alignmentphyloP, phastCons

Interpreting Results with Caution

A significant result in any of these tests is not proof of selection. Consider the following:

  1. Effect size matters: A dN/dS of 1.2 with a p-value of 0.04 is less convincing than a dN/dS of 3.0 with a p-value of 10⁻⁶. Report effect sizes and confidence intervals, not just p-values.
  2. Replication across methods: If PAML identifies positive selection but the MK test does not, consider whether the discrepancy reflects different timescales (divergence vs polymorphism) or a methodological artifact.
  3. Biological plausibility: Does the positively selected gene have a function consistent with adaptation? A gene with no known function and no expression data is a weaker candidate than a gene with a clear role in host defense or reproduction.
  4. Genomic context: Is the signal in a region of low recombination? If so, linked selection may be responsible.

Common Pitfalls

  1. Using dN/dS on highly divergent sequences: When dS is saturated (typically > 2 substitutions per synonymous site), dN/dS estimates become unreliable. Consider using alternative approaches such as the ratio of radical to conservative amino acid changes (e.g., the PAML M0 model with a gamma distribution) or restricting the analysis to more closely related species.
  1. Ignoring the effects of recombination: In population genomic scans, treat regions of low recombination with caution. A selective sweep signature in a recombination cold spot may be a false positive.
  1. Applying the MK test to small samples: The MK test requires sufficient polymorphism data. With fewer than 10 alleles, the confidence intervals on the neutrality index are so wide that the test has almost no power.
  1. Interpreting conservation scores as evidence of function: A conserved sequence is likely functional, but conservation alone does not tell you what the function is. Conversely, a non-conserved sequence may still be functional if it has undergone lineage-specific adaptive evolution.
  1. Failing to account for demographic history: If your population has undergone a recent bottleneck or expansion, SFS-based tests will be biased. Use demographic-aware methods or compare your results to neutral simulations that incorporate the inferred demography.
  1. Overinterpreting branch-site results: The branch-site model can produce false positives when the foreground branch is very long or when there are alignment errors. Always check that the positively selected sites are in well-aligned regions and that the result is robust to the inclusion/exclusion of individual species.

Frequently Asked Questions

What is the difference between purifying and positive selection?

Purifying selection removes deleterious alleles because they reduce fitness, maintaining the functional integrity of genes and genomes. Positive selection increases the frequency of beneficial alleles because they improve fitness, driving adaptive evolutionary change. Purifying selection is the dominant mode of selection across most of the genome—the vast majority of new mutations are deleterious and are removed. Positive selection is rarer but is responsible for adaptations such as pathogen resistance, new metabolic capabilities, and environmental tolerance. The two modes are detected using different signatures: purifying selection leaves reduced diversity and cross-species conservation, while positive selection leaves elevated nonsynonymous substitution rates and selective sweep signatures.

How do you detect positive selection using dN/dS?

Positive selection is detected when the rate of nonsynonymous substitution per nonsynonymous site (dN) exceeds the rate of synonymous substitution per synonymous site (dS), giving dN/dS > 1. In practice, this is tested using codon substitution models implemented in PAML or HyPhy. The standard approach is to fit a null model that does not allow dN/dS > 1 (e.g., M1a or M7) and an alternative model that does (e.g., M2a or M8), then compare them with a likelihood ratio test. If the alternative model fits significantly better, individual sites with high posterior probability of dN/dS > 1 are identified using the Bayes empirical Bayes approach. The branch-site model extends this to detect positive selection affecting only a subset of lineages.

What does a dN/dS ratio less than 1 indicate?

A dN/dS ratio less than 1 indicates that nonsynonymous substitutions are being removed by purifying selection. The encoded amino acid sequence is under functional constraint, and most amino acid changes are deleterious. Typical values are 0.05-0.3 for most protein-coding genes, with lower values indicating stronger constraint. For example, histone genes have dN/dS values near 0.01 because even a single amino acid change disrupts nucleosome structure. A dN/dS of exactly 1 indicates neutral evolution, while values between 0.5 and 1 can indicate either weak purifying selection or a mixture of constrained and unconstrained sites within the gene.

Can demographic history mimic signals of selection?

Yes. Population bottlenecks, expansions, and structure can produce patterns of diversity that resemble selection. For example, a population bottleneck followed by expansion creates an excess of rare variants across the genome, mimicking the signature of purifying selection or a selective sweep. Population structure can create apparent positive selection when alleles are differentiated between subpopulations. The key diagnostic is genomic scale: demography affects all loci, while selection affects specific loci. Genome-wide scans that identify outlier loci against a demographic null model are the standard approach, but the demographic model must be correctly specified to avoid false positives.

What is a selective sweep?

A selective sweep is the process by which a beneficial allele increases in frequency and eventually reaches fixation, carrying linked neutral variation with it. The genomic signature of a completed sweep is a local reduction in nucleotide diversity, an excess of low-frequency derived alleles, and a specific pattern of linkage disequilibrium. The signature decays over time as new mutations accumulate and recombination breaks down the association between the selected site and nearby neutral variants. Detecting sweeps is a major goal of population genomics because they provide direct evidence of recent positive selection.

What is the McDonald-Kreitman test?

The McDonald-Kreitman test compares the ratio of nonsynonymous to synonymous polymorphism within a species to the ratio of nonsynonymous to synonymous divergence between species. Under neutrality, these ratios should be equal. An excess of nonsynonymous divergence (NI < 1) indicates that beneficial amino acid changes have been fixed by positive selection. An excess of nonsynonymous polymorphism (NI > 1) indicates that deleterious amino acid variants are segregating and being removed by purifying selection. The test is simple, requires only a single species pair, and does not need a phylogenetic tree, but it is sensitive to demography and assumes synonymous sites are neutral.

Why is it important to consider recombination when detecting selection?

Recombination breaks down the association between selected and neutral sites. In regions of low recombination, the genomic footprint of selection extends over large distances, creating broad regions that appear to be under selection when they are merely linked to a selected site. This can cause false positives in sweep detection and can inflate estimates of the number of selected sites. Conversely, in regions of high recombination, the footprint of selection is narrow and may be missed entirely. Accounting for recombination rate variation is essential for accurate inference of selection from population genomic data.

Key Takeaways

  • Purifying selection removes deleterious alleles and is the dominant mode of selection across most of the genome, leaving signatures of reduced diversity and cross-species conservation.
  • Positive selection fixes beneficial alleles and is detected through elevated dN/dS ratios, selective sweep signatures, and excess nonsynonymous divergence.
  • The dN/dS ratio is a powerful but imperfect metric; values < 1 indicate purifying selection, > 1 indicate positive selection, and = 1 indicate neutrality, but these interpretations require careful statistical testing.
  • Selection operates in the context of population genetics: the efficiency of selection depends on the product of the selection coefficient and the effective population size (Nes).
  • Detecting selection requires distinguishing it from demographic history, which affects all loci, and from linked selection, which affects regions near selected sites.
  • Multiple complementary methods—codon models, the McDonald-Kreitman test, and SFS-based tests—should be used together, and results should be interpreted with attention to effect size, biological plausibility, and genomic context.
  • Quality control of sequence data, including correct annotation, alignment, and ancestral state inference, is essential for reliable selection inference.

Further Reading

  • Bendall EE et al. Influenza A virus within-host evolution and positive selection in a densely sampled household cohort over three seasons. Virus evolution. 2024. PubMed 39444487
  • Bendall EE et al. Influenza A virus within-host evolution and positive selection in a densely sampled household cohort over three seasons. bioRxiv : the preprint server for biology. 2024. PubMed 39229225
  • Rousselle M et al. Hemizygosity Enhances Purifying Selection: Lack of Fast-Z Evolution in Two Satyrine Butterflies. Genome biology and evolution. 2016. PubMed 27590089
  • Johri P, Pfeifer SP, Jensen JD. Developing an evolutionary baseline model for humans: jointly inferring purifying selection with population history. bioRxiv : the preprint server for biology. 2023. PubMed 37090533
  • Mboumba Bouassa RS et al. Purifying Selection in Human Immunodeficiency Virus-1 pol Gene in Perinatally Human Immunodeficiency Virus-1-Infected Children Harboring Discordant Immunological Response and Virological Nonresponse to Long-Term Antiretroviral Therapy. Journal of clinical medicine research. 2020. PubMed 32587653
  • Civetta A. Positive selection within sperm-egg adhesion domains of fertilin: an ADAM gene with a potential role in fertilization. Molecular biology and evolution. 2003. PubMed 12519902

Related Clinical & Scientific Guides