Positive Selection Pressure: Mechanisms, Detection, and Pitfalls
By Dr. Zubair Khalid, DVM, MS, PhD ·

Introduction to Positive Selection Pressure
Definition and Conceptual Overview
Positive selection pressure is the evolutionary force that drives the spread of advantageous alleles through a population. When a mutation confers a fitness benefit—increased survival, reproductive success, or both—natural selection acts to increase its frequency. This process is the engine of adaptation, shaping everything from pathogen resistance to sensory perception.
The formal population-genetic definition centers on the selection coefficient, s. A beneficial allele with selection coefficient s > 0 increases in frequency at a rate determined by s, the effective population size (Nₑ), and the dominance coefficient. For a codominant allele, the fixation probability is approximately 2*s* for an allele present in a single copy in a diploid population, compared to the neutral expectation of 1/(2*Nₑ*). This means that even mildly beneficial mutations—with s on the order of 10⁻⁴ to 10⁻³—can reach fixation given sufficient time, whereas strictly neutral alleles are subject to the stochastic whims of genetic drift.
Positive selection operates at multiple timescales. On short timescales, it shapes allele frequencies within populations in response to environmental challenges. On longer timescales, it drives the fixation of substitutions that distinguish species. The signature of positive selection at the molecular level is therefore a combination of population-genetic patterns (within species) and comparative patterns (between species).
Positive vs. Purifying vs. Neutral Selection
The three regimes of selection form a continuum defined by the fitness effects of mutations. Purifying Selection vs Positive Selection are opposite ends of a spectrum. Purifying (negative) selection removes deleterious alleles from populations. Most nonsynonymous mutations that alter protein sequence are slightly deleterious and are eliminated or kept at low frequency. The signature of purifying selection is reduced diversity at functional sites relative to neutral expectations, and a dN/dS ratio (nonsynonymous substitutions per nonsynonymous site divided by synonymous substitutions per synonymous site) substantially below 1.
Neutral evolution, as formalized by the Neutral Theory of Molecular Evolution, posits that most observed polymorphism and divergence are selectively equivalent. Under strict neutrality, the rate of substitution equals the mutation rate, and dN/dS ≈ 1. Positive selection, by contrast, drives dN/dS > 1—an excess of amino-acid-altering changes relative to silent changes—because beneficial mutations are fixed more rapidly than neutral ones.
Critically, these categories are not static. A mutation that is neutral in one environment may become beneficial in another. The same gene can experience purifying selection in one lineage and positive selection in another. The distinction is operational: we infer the mode of selection from the patterns of variation and divergence, and those inferences are always conditional on the timescale and genomic context examined.
Mechanisms of Positive Selection
Beneficial Mutations and Fitness
Positive selection begins with a mutation that increases fitness. The molecular basis of such mutations varies: they may alter enzyme kinetics, protein stability, binding affinity, gene expression levels, or splicing patterns. A classic example is the evolution of β-lactamase enzymes in bacteria. A single amino acid substitution, such as the TEM-1 E104K mutation, can expand the substrate spectrum to include extended-spectrum cephalosporins, conferring resistance to clinically important antibiotics. The fitness benefit is context-dependent—in the presence of the antibiotic, the mutant allele has a large selective advantage; in its absence, the mutation may be neutral or slightly deleterious due to reduced catalytic efficiency on the original substrate.
The population-genetic fate of a beneficial mutation depends on its selection coefficient and the population's effective size. In large populations, selection is more efficient at discriminating between alleles of small effect. In small populations, drift dominates, and even beneficial mutations may be lost. The probability of fixation of a beneficial allele is approximately 2*s* for additive alleles, but this is reduced if the allele is recessive (fixation probability ≈ 2*s*·h, where h is the dominance coefficient) and increased if it is dominant.
Selective Sweeps and Hitchhiking
When a beneficial mutation sweeps to fixation, it carries with it all linked variation. This process, known as a selective sweep, leaves a distinctive genomic signature: a region of reduced heterozygosity surrounding the selected site. The size of the swept region depends on the recombination rate and the selection coefficient. For a hard sweep—one that starts from a single new mutation—the expected reduction in diversity extends over a distance proportional to s/r, where r is the recombination rate per base pair per generation. In Drosophila melanogaster, where recombination rates average ~2 cM/Mb, a strongly selected allele with s = 0.01 can reduce diversity over several hundred kilobases.
Linked neutral alleles that increase in frequency because of their association with a beneficial mutation are said to be "hitchhiking." This process can produce false signals of selection at loci that are merely linked to the true target. Conversely, the hitchhiking effect can be used to identify selected regions: a local reduction in diversity, an excess of rare variants, and a skewed site frequency spectrum are all expected signatures.
The distinction between hard and soft sweeps is important. Soft sweeps arise when selection acts on standing genetic variation or when multiple independent beneficial mutations arise at the same locus. Soft sweeps produce weaker reductions in diversity and are more difficult to detect with classical methods. They are increasingly recognized as common in natural populations, particularly for adaptations that involve pre-existing alleles, such as the spread of the EDAR V370A allele in East Asian human populations, which affects hair follicle density and tooth morphology.
Adaptive Introgression
Adaptive introgression is the movement of beneficial alleles between species through hybridization and repeated backcrossing. This mechanism can introduce alleles that have already been tested by selection in one species into a closely related species, bypassing the need for de novo mutation. The classic example is the introgression of Neanderthal alleles into modern humans. Several introgressed regions show evidence of positive selection, including the EPAS1 haplotype associated with high-altitude adaptation in Tibetans, which was introgressed from Denisovans.
The genomic signature of adaptive introgression is a haplotype that is more similar to the donor species than to the recipient species' own alleles, often with elevated divergence from the recipient's ancestral state. Detection methods compare patterns of allele sharing between species and look for regions where the divergence is inconsistent with the genome-wide background. The EPAS1 case is particularly striking: the introgressed haplotype is nearly fixed in Tibetan populations but absent from closely related lowland Han Chinese, and it contains multiple derived alleles that alter hypoxia response pathways.
Evidence of Positive Selection in Genomes
Reduced Nucleotide Diversity
A classic signature of a recent selective sweep is a local reduction in nucleotide diversity. The region surrounding a beneficial mutation that has recently fixed will show lower heterozygosity than the genomic background because all chromosomes trace back to the single ancestral chromosome that carried the mutation. The magnitude of this reduction depends on the time since fixation, the recombination rate, and the selection coefficient.
For example, in human populations, the LCT gene region—responsible for lactase persistence—shows a striking reduction in diversity in European and African pastoralist populations. The selective sweep is estimated to have occurred within the last 10,000 years, and the region of reduced diversity extends over approximately 1 Mb. The signal is absent in populations without a history of dairy farming, confirming the role of cultural practice in driving genetic adaptation.
However, reduced diversity alone is not diagnostic. Demographic events such as population bottlenecks or population expansion can produce similar patterns genome-wide. The key is to compare the target region to the genome-wide distribution: a region that shows a more extreme reduction than expected given the demographic history is a candidate for positive selection.
Elevated dN/dS Ratios
The dN/dS ratio (ω) compares the rate of nonsynonymous substitution (dN) to synonymous substitution (dS). Under purifying selection, ω < 1; under neutrality, ω ≈ 1; and under positive selection, ω > 1. This metric is the workhorse of comparative evolutionary analysis because it can be estimated from alignments of coding sequences from multiple species.
The classic example of elevated dN/dS is the MHC (major histocompatibility complex) genes in vertebrates. The antigen-binding groove of MHC molecules interacts directly with pathogen-derived peptides, and this region shows ω values exceeding 5 in some codon positions. The selection pressure is driven by the need to recognize a diverse and rapidly evolving array of pathogens; alleles that bind novel pathogen peptides are favored.
It is important to note that ω > 1 is a stringent criterion. Most genes under positive selection show ω values between 1 and 2 at only a subset of codons. Averaging over the entire gene—or over long evolutionary timescales—can obscure these signals because the signal of selection at a few sites is diluted by purifying selection at the majority of sites. This is why site-specific and branch-site models, discussed below, are essential.
Long Haplotypes and Linkage Disequilibrium
When a beneficial allele rises to high frequency rapidly, recombination has little time to break down the haplotype on which it arose. The result is a long haplotype with high frequency in the population, a pattern captured by statistics such as iHS (integrated haplotype score) and nSL (number of segregating sites by length). These statistics compare the length of haplotypes carrying the derived allele to those carrying the ancestral allele.
The ABCC11 gene in East Asian populations provides a clear example. A single nucleotide polymorphism (SNP) at position 538 (G→A) results in a truncated protein that reduces apocrine sweat secretion, leading to dry earwax and reduced body odor. The derived A allele is at high frequency in East Asians (~95%) but rare in Africans and Europeans. The haplotype carrying the A allele is long and shows high linkage disequilibrium, consistent with a recent selective sweep. The selective pressure is hypothesized to be related to reduced body odor, which may have been favored in colder climates or as a result of cultural preferences.
Statistical Methods for Detecting Positive Selection
Codon-Based Models (dN/dS)
The most widely used framework for detecting positive selection is the codon substitution model implemented in programs such as PAML (codeml) and HyPhy. These models use a maximum-likelihood framework to estimate ω from alignments of protein-coding sequences. The basic model (M0) assumes a single ω for all sites. More complex models allow ω to vary among sites (M1a: two categories, neutral and purifying; M2a: adds a third category for positive selection, ω > 1) or among branches (branch models).
The likelihood ratio test (LRT) compares the fit of nested models. For example, to test for positive selection at specific sites, one compares M1a (null: no sites with ω > 1) to M2a (alternative: a proportion of sites with ω > 1). A significant LRT indicates that the data are better explained by a model that allows positive selection. The Bayes empirical Bayes (BEB) approach is then used to identify the specific sites with high posterior probability of ω > 1.
Site-Specific and Branch-Site Tests
Site-specific models identify codons under positive selection across all lineages in the phylogeny. These are powerful for detecting conserved functional domains where a few key residues are under diversifying selection, such as the antigen-binding sites of MHC genes or the receptor-binding domains of viral surface proteins.
Branch-site models, by contrast, allow ω to vary both among sites and among branches. The "foreground" branch (the lineage of interest) is allowed to have a class of sites with ω > 1, while the "background" branches are constrained to ω ≤ 1. This is the appropriate test for detecting positive selection that occurred specifically in a lineage of interest, such as the evolution of the FOXP2 gene in the human lineage after divergence from chimpanzees. The branch-site test found significant evidence for positive selection in the human branch, with several amino acid changes that are likely related to the development of speech and language.
The branch-site test is more powerful than the site test for lineage-specific adaptations, but it is also more sensitive to assumptions about the phylogeny and alignment quality. False positives can arise from alignment errors that create spurious nonsynonymous differences, particularly in regions of low complexity or high divergence.
Population Genetic Tests
Population-genetic tests use polymorphism data from within a species to detect deviations from neutral expectations. These tests are based on the site frequency spectrum (SFS)—the distribution of allele frequencies across segregating sites.
Tajima's D compares two estimators of the population mutation parameter θ: θ_π (average pairwise nucleotide diversity) and θ_W (Watterson's estimator based on the number of segregating sites). Under neutrality, these are expected to be equal, so D ≈ 0. A selective sweep reduces diversity and creates an excess of rare variants, leading to negative D values. However, population expansion also produces negative D, and population bottlenecks produce positive D, so demographic confounding is a major issue.
Fay and Wu's H is a more specific test for selective sweeps. It uses the frequency of derived alleles, weighting high-frequency derived alleles more heavily. A selective sweep produces an excess of high-frequency derived alleles (because the beneficial allele and its linked variants are at high frequency), leading to negative H. This test is less sensitive to population expansion than Tajima's D but requires accurate inference of the ancestral state, which can be problematic in regions of high divergence.
Composite Likelihood and Machine Learning Approaches
Composite likelihood methods, such as SweepFinder and SweeD, scan the genome for regions where the local SFS is inconsistent with the genome-wide background. These methods calculate a composite likelihood ratio (CLR) that compares the probability of the data under a model with a sweep at a given position to the probability under a neutral model. They are more powerful than single-statistic tests because they integrate information across many sites.
Machine learning approaches, such as diHMM and SweepNet, have been developed to improve detection accuracy by learning the complex patterns associated with sweeps from simulated training data. These methods can incorporate multiple features—diversity, SFS, linkage disequilibrium, and haplotype structure—and can distinguish between hard sweeps, soft sweeps, and background selection more effectively than classical methods. However, they are computationally intensive and require careful validation to avoid overfitting to the simulation parameters.
Interpreting Positive Selection Results
Biological Validation of Candidate Genes
Statistical evidence for positive selection is a hypothesis, not a conclusion. The gold standard for validation is functional experimentation. For example, if a gene shows evidence of positive selection in the branch leading to high-altitude populations, one would test whether the derived alleles alter protein function in a way that improves hypoxia tolerance. This might involve expressing the ancestral and derived alleles in cell culture and measuring their effects on hypoxia-inducible factor (HIF) signaling, or generating knock-in mice carrying the derived allele and measuring physiological responses to low oxygen.
The EPAS1 gene in Tibetans is a model case. Statistical evidence identified the gene as a target of positive selection, and subsequent functional studies showed that the derived alleles reduce the expression of HIF-2α, leading to lower hemoglobin concentrations at high altitude. This is adaptive because excessive hemoglobin makes blood viscous and increases the risk of thrombosis; Tibetans maintain lower hemoglobin levels than Andean highlanders, who show a different pattern of adaptation.
Distinguishing Selection from Demography
The greatest challenge in positive selection detection is distinguishing selection from demography. Population bottlenecks, expansions, and migration can produce patterns of diversity and SFS that mimic selection. For example, a population bottleneck reduces diversity genome-wide, and a recent bottleneck followed by expansion can produce an excess of rare variants similar to a selective sweep.
Several strategies address this problem. One is to compare the candidate region to the genome-wide distribution: if the signal is more extreme than 99% of the genome, it is unlikely to be explained by demography alone. Another is to use demographic models inferred from neutral regions to generate null distributions for selection statistics. This is the approach taken by the ms and msprime simulation programs, which can simulate data under a specified demographic model and generate empirical null distributions for statistics like Tajima's D or iHS.
A third strategy is to use statistics that are robust to demography. The H12 statistic, which measures the sum of the squared frequencies of the two most common haplotypes, is less sensitive to demographic history than SFS-based tests. Similarly, the singleton density score (SDS) uses the decay of linkage disequilibrium around a focal allele to detect very recent selection, and it is relatively robust to older demographic events.
Common Pitfalls in Positive Selection Studies
Alignment and Saturation Issues
The most common source of false positives in dN/dS-based analyses is alignment error. Misaligned codons create spurious nonsynonymous differences, inflating dN and producing false signals of positive selection. This is particularly problematic in regions of low sequence complexity, such as proline-rich or glycine-rich domains, where gaps are difficult to place accurately.
Saturation is another issue. When sequences are highly diverged, multiple substitutions at the same site can obscure the true number of changes. Synonymous sites saturate faster than nonsynonymous sites because they are less constrained, which can deflate dS and inflate the dN/dS ratio. This is a particular concern for comparisons between deeply diverged lineages, such as mammals versus fish. The standard remedy is to check for saturation by plotting the number of transitions and transversions against divergence and to exclude third codon positions if saturation is evident.
Recombination and Its Effects
Recombination breaks down the association between alleles at different sites, which can severely affect the performance of selection detection methods. For dN/dS-based methods, recombination can create spurious signals because different regions of the gene may have different phylogenetic histories. The branch-site test is particularly sensitive to recombination: if recombination has occurred, the test may falsely reject the null hypothesis of no positive selection.
For population-genetic tests, recombination reduces the signal of a selective sweep because it breaks down the haplotype structure. A sweep that occurred 1,000 generations ago will have a much weaker signal in a high-recombination region than in a low-recombination region. This means that failure to detect a sweep does not rule out selection; it may simply reflect the local recombination rate.
The standard approach is to test for recombination before performing selection analyses. Programs like GARD (Genetic Algorithm Recombination Detection) can identify recombination breakpoints, and the analysis can be performed on the resulting non-recombining blocks. Alternatively, methods like omegaMap jointly infer recombination and selection, avoiding the need for a two-step approach.
Overinterpretation of dN/dS
The dN/dS ratio is a powerful statistic, but it is frequently misinterpreted. A dN/dS ratio greater than 1 is strong evidence for positive selection, but a ratio less than 1 does not rule out positive selection. If positive selection acts on a small subset of sites, the genome-wide average ω will be dominated by purifying selection at the majority of sites. For example, a gene with 5% of sites under strong positive selection (ω = 5) and 95% under strong purifying selection (ω = 0.05) will have an average ω of approximately 0.3—well below 1.
Similarly, ω can be inflated by relaxation of purifying selection. If a gene is no longer functionally constrained, nonsynonymous mutations accumulate at a rate closer to the synonymous rate, producing ω values approaching 1. This is not positive selection; it is the absence of negative selection. The distinction is critical: a pseudogene has ω ≈ 1 because it is unconstrained, not because it is adaptive.
Finally, dN/dS is sensitive to sampling. With few species or short branches, the estimates of dN and dS have large variances, and the LRT may lack power. Conversely, with very large datasets, even biologically insignificant differences can be statistically significant. The effect size—the proportion of sites under positive selection and their ω values—should always be reported alongside the p-value.
Case Studies of Positive Selection
Immune System Genes
The immune system is the most consistent target of positive selection in vertebrates. Pathogens evolve rapidly, and host immune genes must keep pace. The MHC genes are the classic example, with the antigen-binding groove showing some of the highest ω values known—often exceeding 5 at specific codons. The selection pressure is balancing selection: rare alleles are favored because pathogens are less likely to have evolved evasion strategies against them. This produces a pattern of trans-species polymorphism, where alleles are shared between species because they are older than the species themselves.
The TRIM5α gene in primates provides another example. This gene encodes a restriction factor that blocks retroviral infection by binding to the viral capsid. The gene shows strong signatures of positive selection, particularly in the B30.2/SPRY domain that interacts with the capsid. The selection pressure is driven by ancient retroviral epidemics: different primate lineages have different alleles that provide resistance to different retroviruses, and the pattern of amino acid change correlates with the history of retroviral exposure.
Sensory Genes
Sensory systems are frequent targets of positive selection because they mediate interactions with the environment. The opsin genes, which encode light-sensitive proteins in the retina, provide a textbook example. In primates, the duplication of the opsin gene on the X chromosome gave rise to the trichromatic color vision system. The M/L opsin genes show evidence of positive selection at sites that determine spectral sensitivity, with amino acid changes at positions 180, 277, and 285 shifting the absorption maximum of the photopigment.
The TRPV1 gene, which encodes a heat-activated ion channel, shows evidence of positive selection in vampire bats. The ancestral channel is activated at temperatures above 43°C, but the vampire bat channel has a lower activation threshold, allowing the bat to sense infrared radiation from its prey. This adaptation involved changes in the ankyrin repeat domain, which is thought to modulate the temperature sensitivity of the channel.
High-Altitude Adaptation
High-altitude adaptation is one of the best-studied examples of recent positive selection in humans. The EPAS1 gene in Tibetans, discussed above, shows one of the strongest signals of selection in the human genome. The derived haplotype is nearly fixed in Tibetans, shows a long haplotype structure, and is associated with reduced hemoglobin concentration at high altitude.
The EGLN1 gene, which encodes prolyl hydroxylase domain-containing protein 2 (PHD2), also shows evidence of positive selection in Tibetans. PHD2 hydroxylates HIF-1α and HIF-2α, targeting them for degradation under normoxic conditions. The derived alleles in Tibetans are associated with reduced PHD2 activity, leading to higher HIF levels and altered hypoxia response. The two genes—EPAS1 and EGLN1—act in the same pathway, and their combined effects produce the characteristic Tibetan phenotype of low hemoglobin and high oxygen saturation.
In Andean highlanders, the pattern is different. The EGLN1 gene also shows evidence of selection, but the specific alleles and their effects differ. Andeans have higher hemoglobin levels than Tibetans, and the selection pressure appears to have favored alleles that increase oxygen delivery rather than reduce hemoglobin concentration. This illustrates an important principle: the same environmental challenge can produce different genetic solutions in different populations.
Practical Guidelines for Detecting Positive Selection
Data Acquisition and Quality Control
The first step in any positive selection study is data acquisition. For comparative analyses, you need coding sequences from multiple species. The number of species required depends on the question: for site-specific tests, at least 10-20 species are recommended, and more are better. The sequences should be orthologous, not paralogous, and the species should span a range of divergence times appropriate for the question.
Quality control is essential. Check for frameshift mutations, premature stop codons, and sequencing errors. Align the sequences using a codon-aware aligner such as PRANK or MACSE, which account for the codon structure and reduce alignment errors. Inspect the alignment manually, particularly in regions of low complexity. Remove sequences that are clearly misaligned or contain obvious errors.
For population-genetic analyses, you need polymorphism data from a population. The sample size should be at least 20-50 individuals for reliable SFS estimates. The sequencing should be high coverage (≥20×) to minimize genotyping errors, and the variants should be filtered for quality. The ancestral state for each SNP should be inferred from an outgroup sequence, and sites where the ancestral state is ambiguous should be excluded.
Choosing Appropriate Methods
The choice of method depends on the question. If you want to know whether a specific gene has been under positive selection during the evolution of a lineage, use the branch-site test in PAML or HyPhy. If you want to identify individual codons under positive selection, use the site models (M1a vs. M2a) or the FUBAR (Fast Unconstrained Bayesian AppRoximation) method in HyPhy, which is faster and more robust for large datasets.
If you want to detect recent selective sweeps in a population, use a combination of SFS-based methods (SweepFinder, Tajima's D) and haplotype-based methods (iHS, nSL). The haplotype-based methods are more powerful for detecting incomplete sweeps, while the SFS-based methods are better for detecting completed sweeps. Use multiple methods and require concordance: a region that shows signals in both SFS and haplotype analyses is a stronger candidate than one that shows a signal in only a single test.
Reporting Results and Limitations
Report the effect sizes, not just the p-values. For dN/dS analyses, report the proportion of sites under positive selection and their ω values. For sweep detection, report the CLR score and the size of the region affected. Report the demographic model used for the null distribution and the simulation parameters. State clearly what the analysis can and cannot detect: a failure to detect positive selection does not mean that selection did not occur; it may reflect low power, recombination, or the action of soft sweeps.
Always validate with independent data. If you have a candidate gene from a comparative analysis, check whether it shows evidence of selection in population data. If you have a candidate region from a sweep scan, check whether it contains genes with relevant biological functions. The strongest studies combine multiple lines of evidence: comparative, population-genetic, and functional.
Frequently Asked Questions
What is positive selection pressure?
Positive selection pressure is the evolutionary force that increases the frequency of advantageous alleles in a population. It acts when a mutation confers a fitness benefit, such as increased survival or reproductive success, and drives the allele toward fixation. The strength of the pressure is quantified by the selection coefficient, s, which measures the relative fitness advantage of the beneficial allele.
How is positive selection pressure defined?
Positive selection pressure is defined operationally by its genomic signatures: reduced nucleotide diversity around the selected site, elevated dN/dS ratios at selected codons, and extended haplotype homozygosity. Statistically, it is defined as the rejection of the null hypothesis of neutrality in favor of a model that allows a class of sites or a lineage to have ω > 1, or as a deviation from the neutral site frequency spectrum in the direction expected for a selective sweep.
What is the difference between positive and purifying selection?
Positive selection favors beneficial mutations and drives them to fixation, while purifying selection removes deleterious mutations from the population. At the molecular level, positive selection produces an excess of nonsynonymous substitutions (dN/dS > 1), while purifying selection produces a deficit (dN/dS < 1). Positive selection is adaptive and lineage-specific, while purifying selection is conservative and maintains functional constraint. See __MASK_3__ for a more detailed comparison.
How do you detect positive selection in a gene?
The standard approach is to align coding sequences from multiple species and fit codon substitution models that allow ω to vary among sites or branches. A likelihood ratio test compares a null model that does not allow positive selection to an alternative model that does. Significant results identify the gene as a candidate, and the Bayes empirical Bayes approach identifies the specific codons under selection. Population-genetic tests, such as Tajima's D or iHS, can provide complementary evidence from within-species polymorphism data.
What does a dN/dS ratio greater than 1 indicate?
A dN/dS ratio greater than 1 indicates that nonsynonymous substitutions occur at a higher rate than synonymous substitutions, which is the signature of positive selection at the amino acid level. It means that amino acid changes are being fixed at a rate faster than expected under neutrality, implying that they are beneficial. However, ω > 1 is a stringent criterion, and most genes under positive selection show ω > 1 at only a small subset of codons.
Can positive selection be detected in non-coding regions?
Yes, but the methods differ. Non-coding regions do not have codons, so dN/dS-based methods do not apply. Instead, one uses population-genetic statistics that detect the signatures of selective sweeps: reduced diversity, skewed SFS, and extended haplotype homozygosity. Methods like SweepFinder and iHS can be applied to any genomic region. Functional validation is more challenging for non-coding regions, but regulatory changes can be tested using reporter assays or by measuring allele-specific expression.
What are common pitfalls when studying positive selection?
The most common pitfalls are alignment errors, which create spurious nonsynonymous differences; saturation, which deflates dS and inflates dN/dS; recombination, which breaks down haplotype structure and confounds branch-site tests; and demographic confounding, where population bottlenecks or expansions mimic the signatures of selection. Overinterpretation of dN/dS is also common: a ratio below 1 does not rule out positive selection, and a ratio near 1 may reflect relaxed constraint rather than adaptation.
Key Takeaways
- Positive selection pressure is the force that drives beneficial alleles to fixation, and it leaves distinctive genomic signatures: reduced diversity, elevated dN/dS, and extended haplotypes.
- The mechanisms include hard sweeps from new mutations, soft sweeps from standing variation, and adaptive introgression from hybridization between species.
- Detection methods fall into two broad categories: comparative methods based on dN/dS (PAML, HyPhy) and population-genetic methods based on the site frequency spectrum and haplotype structure (Tajima's D, SweepFinder, iHS).
- The branch-site test is the most powerful comparative method for detecting lineage-specific positive selection, but it is sensitive to alignment errors and recombination.
- Distinguishing selection from demography is the central challenge; use demographic null models and multiple independent statistics to build confidence.
- Statistical evidence is a hypothesis, not a conclusion; functional validation is essential to confirm that a candidate gene is truly adaptive.
- Common pitfalls include alignment errors, saturation, recombination, and overinterpretation of dN/dS; report effect sizes and limitations alongside p-values.
Further Reading
- Nijmeijer BM, Geijtenbeek TBH. Negative and Positive Selection Pressure During Sexual Transmission of Transmitted Founder HIV-1. Frontiers in immunology. 2019. PubMed 31354736
- Derbyshire MC. Bioinformatic Detection of Positive Selection Pressure in Plant Pathogens: The Neutral Theory of Molecular Sequence Evolution in Action. Frontiers in microbiology. 2020. PubMed 32328056
- Moir RD, Tanzi RE. Low Evolutionary Selection Pressure in Senescence Does Not Explain the Persistence of Aβ in the Vertebrate Genome. Frontiers in aging neuroscience. 2019. PubMed 30983989
- Coronado L et al. Positive selection pressure on E2 protein of classical swine fever virus drives variations in virulence, pathogenesis and antigenicity: Implication for epidemiological surveillance in endemic areas. Transboundary and emerging diseases. 2019. PubMed 31306567
- Pérez LJ et al. Positive selection pressure on the B/C domains of the E2-gene of classical swine fever virus in endemic areas under C-strain vaccination. Infection, genetics and evolution : journal of molecular epidemiology and evolutionary genetics in infectious diseases. 2012. PubMed 22580241
- Wachter J, Hill S. Positive Selection Pressure Drives Variation on the Surface-Exposed Variable Proteins of the Pathogenic Neisseria. PloS one. 2016. PubMed 27532335