# Positive Selection Definition and Mechanisms in Molecular Evolution

## Introduction to Positive Selection

Positive selection is the evolutionary process by which newly arising beneficial mutations increase in frequency and become fixed within a population because they confer a fitness advantage on their carriers. In molecular evolution, positive selection is detected when the rate of nonsynonymous substitution (amino acid–altering) exceeds the rate of synonymous substitution (silent), indicating that protein-coding changes are being actively retained by natural selection rather than removed or randomly drifting.

The concept is central to understanding how organisms adapt to new environments, resist pathogens, acquire novel functions, and diversify at the molecular level. Positive selection operates at the level of alleles and genotypes, but its effects are observed in the patterns of nucleotide substitution and polymorphism that accumulate in DNA sequences over evolutionary time.

### What Is Positive Selection?

Positive selection, also called directional selection or Darwinian selection, occurs when a new mutation increases the fitness of its bearer relative to individuals carrying the ancestral allele. The selected allele's frequency rises deterministically, driven by its selective advantage, rather than by random genetic drift. The strength of selection is quantified by the selection coefficient, *s*, which represents the fractional increase in fitness conferred by the mutant allele. A mutation with *s* = 0.01 increases relative fitness by 1% per generation.

For a diploid population, the fate of a beneficial mutation depends on both *s* and the effective population size (*Nₑ*). The probability of fixation of a newly arising beneficial mutation is approximately 2*s* for an autosomal locus in a diploid population, provided *s* is small and selection is additive. This approximation, derived from diffusion theory, shows that even weakly beneficial mutations (*s* ≈ 0.001) have a nonzero chance of fixation, though that chance is small in large populations. Strongly beneficial mutations (*s* > 0.05) are nearly guaranteed to fix if they are not lost by drift in the first few generations.

Positive selection is distinct from balancing selection, which maintains multiple alleles at a locus over long periods, and from diversifying selection, which favors different alleles in different subpopulations or ecological niches. In molecular evolution, the term "positive selection" is most often used to describe recurrent fixation of beneficial amino acid replacements at a locus, detected as an elevated rate of nonsynonymous substitution.

### Positive vs. Purifying vs. Neutral Evolution

The three regimes of molecular evolution are distinguished by the fate of new mutations and the ratio of nonsynonymous to synonymous substitution rates.

**Purifying selection** (also called negative selection) removes deleterious mutations from the population. Most nonsynonymous mutations that alter protein structure or function are harmful and are eliminated before they reach appreciable frequency. Purifying selection is the dominant force acting on protein-coding genes in most organisms, which is why amino acid sequences are highly conserved across deep evolutionary distances. The rate of nonsynonymous substitution (dN) is typically much lower than the synonymous rate (dS), yielding dN/dS ratios well below 1.

**Neutral evolution** describes the accumulation of mutations that have no effect on fitness. Synonymous substitutions and mutations in nonfunctional genomic regions evolve at the neutral rate, which equals the mutation rate. Under strict neutrality, dN/dS = 1, because nonsynonymous and synonymous mutations are fixed at the same rate.

**Positive selection** produces dN/dS > 1, because beneficial amino acid changes are fixed at a rate exceeding the neutral expectation. This signal is rare in genome-wide scans—most genes show dN/dS < 1—but it is enriched in specific functional categories such as immune response genes, reproductive proteins, and sensory receptors.

The relationship between these regimes is not static. A gene may experience purifying selection for most of its evolutionary history but undergo episodes of positive selection following environmental change or gene duplication. The __MASK_1__ framework captures this dynamic, where the same locus can shift between regimes over time. Understanding the distinction between __MASK_2__ is essential for interpreting dN/dS values correctly, as a gene-wide average near 1 can mask heterogeneous selection across sites.

## The Molecular Mechanisms of Positive Selection

### Beneficial Mutations and Fixation

The raw material for positive selection is the standing pool of new mutations generated by DNA replication errors, DNA damage, and repair processes. In nuclear genomes of multicellular eukaryotes, the spontaneous mutation rate is approximately 1 × 10⁻⁸ to 1 × 10⁻⁹ substitutions per site per generation. Most mutations are neutral or deleterious; beneficial mutations are rare, estimated at roughly 1 in 1000 to 1 in 10,000 new mutations in microbial systems.

A beneficial mutation must survive three phases to contribute to adaptation: (1) initial establishment, where it escapes stochastic loss while rare; (2) deterministic increase, where its frequency rises predictably according to its selective advantage; and (3) fixation, where it reaches 100% frequency in the population. The establishment phase is the most precarious. A new mutation exists in a single copy and has a high probability of being lost by drift, even if strongly beneficial. The probability of surviving this initial bottleneck is approximately 2*s* for additive alleles, as noted above.

Once established, the allele's frequency trajectory follows the logistic growth equation:

p(t) = p₀e^(st) / (1 + p₀(e^(st) − 1))

where p₀ is the initial frequency and t is time in generations. The time to fixation from a single copy is approximately (2/s)ln(2Nₑ) generations for a diploid population. For a population of 10⁵ individuals and *s* = 0.01, fixation takes roughly 2,400 generations. Stronger selection shortens this window, and weaker selection extends it.

The fixation of beneficial mutations is not an independent process at different loci. When a beneficial mutation sweeps to fixation, it carries with it all linked genetic variation—including neutral and deleterious alleles at nearby sites. This phenomenon, known as a selective sweep, has profound consequences for patterns of polymorphism in the genome.

### Selective Sweeps and Hitchhiking

A selective sweep occurs when a beneficial mutation rises to fixation so rapidly that recombination has little time to break up the association between the selected site and neighboring variants. The result is a reduction in genetic diversity flanking the selected site, because all chromosomes in the population descend from the single chromosome that carried the original beneficial mutation.

The size of the swept region depends on the ratio of the recombination rate (*r*) to the selection coefficient (*s*). The expected size of the region of reduced diversity is approximately *s*/r base pairs. For a typical mammalian recombination rate of 1 cM/Mb (approximately 10⁻⁸ per base pair per generation) and *s* = 0.01, the sweep region spans roughly 1 Mb. Stronger selection produces broader sweeps; higher recombination rates narrow them.

Hitchhiking describes the passive increase in frequency of neutral alleles that are linked to a beneficial mutation. A neutral allele that happens to reside on the same chromosome as a sweeping beneficial allele will rise in frequency along with it, even if the neutral allele has no fitness effect. This process generates characteristic distortions in the site frequency spectrum: an excess of rare variants (because new mutations arising after the sweep have not yet accumulated) and a deficit of intermediate-frequency variants (because the sweep removed pre-existing variation).

Two types of sweeps are distinguished: hard sweeps, where a single de novo beneficial mutation sweeps to fixation from a single copy, and soft sweeps, where selection acts on standing genetic variation or on recurrent mutations at the same locus. Soft sweeps preserve more diversity at linked sites and produce subtler genomic signatures. Population genomic studies increasingly recognize that soft sweeps may be more common than hard sweeps in large populations and in adaptation from standing variation, such as the spread of insecticide resistance alleles that existed at low frequency before insecticide exposure.

The __MASK_3__ framework encompasses both hard and soft sweeps, and the distinction matters for detection: methods that assume a hard sweep may miss soft sweep signatures. The __MASK_4__ acting on a locus determines which mode of sweep predominates, with strong, acute pressure favoring hard sweeps and chronic, moderate pressure favoring soft sweeps.

## Detecting Positive Selection: dN/dS Ratios

### The dN/dS Metric

The dN/dS ratio, also written as ω (omega) or Ka/Ks, is the most widely used metric for detecting positive selection in protein-coding sequences. It compares the rate of nonsynonymous substitutions per nonsynonymous site (dN) to the rate of synonymous substitutions per synonymous site (dS).

The computation requires an alignment of homologous coding sequences and a phylogenetic tree. Maximum likelihood methods estimate the number of synonymous and nonsynonymous substitutions along each branch using models of codon evolution, such as the Goldman-Yang or Muse-Gaut models. These models account for the genetic code structure, transition/transversion rate biases, and codon frequency biases.

The dN/dS ratio is estimated as:

ω = dN/dS

Under neutral evolution, ω = 1. Under purifying selection, ω < 1. Under positive selection, ω > 1.

The key advantage of dN/dS is that it controls for mutation rate and time: synonymous substitutions accumulate at the neutral rate and serve as an [internal calibration](/knowledge/diagnostics/molecular/internal-calibration) for the expected rate of nonsynonymous substitution in the absence of selection. This makes dN/dS comparable across genes, lineages, and timescales, provided the sequences are sufficiently diverged to accumulate a measurable number of substitutions.

### Interpreting dN/dS Values

The interpretation of dN/dS values requires attention to the timescale and the averaging window.

**Gene-wide averages** are informative but conservative. A gene-wide ω < 1 indicates that most amino acid changes are deleterious and removed by purifying selection. A gene-wide ω > 1 is strong evidence that a substantial fraction of amino acid changes are beneficial, but this signal is rare because most genes are predominantly conserved. For example, primate immune genes such as the major histocompatibility complex (MHC) class I genes show gene-wide ω values near or above 1, reflecting intense diversifying selection at antigen-binding sites.

**Site-specific estimates** are more sensitive. Within a gene, most codons may be under strong purifying selection (ω < 1) while a few codons are under positive selection (ω > 1). Averaging across all sites dilutes the signal. Codon models that allow ω to vary across sites, implemented in programs such as PAML (codeml) and HyPhy, can identify individual positively selected sites. These are the sites where amino acid changes are repeatedly fixed, often corresponding to functionally critical positions such as enzyme active sites, protein-protein interaction interfaces, or antigen recognition regions.

**Lineage-specific estimates** detect episodic selection. A gene may be conserved in most lineages but experience a burst of adaptive change in one lineage. Branch models allow ω to vary among branches of the phylogeny, identifying lineages where dN/dS > 1. The branch-site model combines both approaches, allowing specific sites to be under positive selection only on specific branches—a powerful framework for detecting adaptation that occurred once in a clade.

The [molecular clock definition](/knowledge/molecular-biology/molecular-clock-definition) is relevant here: dN/dS estimation assumes that synonymous substitutions accumulate at a relatively constant rate, providing the neutral baseline. Violations of the molecular clock, such as generation time effects or mutator phenotypes, can distort dN/dS estimates and produce false positives.

## Phylogenetic Methods for Positive Selection

### Site Models

Site models allow ω to vary among codon sites but assume the same ω for a given site across all branches of the phylogeny. The null model (M7) assumes ω follows a beta distribution bounded between 0 and 1, allowing only purifying or neutral evolution. The alternative model (M8) adds an additional class of sites with ω > 1, estimated from the data. A likelihood ratio test compares the fit of M7 and M8; a significant result indicates that some sites are under positive selection.

Bayes empirical Bayes (BEB) analysis, implemented in PAML, then calculates the posterior probability that each site belongs to the positively selected class. Sites with posterior probability > 0.95 are considered candidates for positive selection. Typical positively selected sites identified by this approach include positions 44, 45, 62, 63, 67, 69, 70, 71, 73, 74, 77, 80, 81, 84, 95, 97, 99, 113, 114, 116, 152, 156, 158, 159, 160, 163, 165, 167, 171 in MHC class I antigen-binding grooves across mammalian species.

Site models are most powerful when positive selection has acted repeatedly at the same sites across the phylogeny, as in host-pathogen arms races or ligand-receptor coevolution. They are less powerful for detecting episodic selection that affected only a few lineages.

### Branch Models

Branch models allow ω to vary among branches of the phylogeny. The simplest approach is the two-ratio model, where the branch of interest (the foreground branch) is allowed to have a different ω than the rest of the tree (the background branches). A likelihood ratio test compares this model to a one-ratio null model where all branches share the same ω.

Branch models detect lineage-specific shifts in selective pressure, such as the acceleration of dN/dS in the lineage leading to humans after the split from chimpanzees, or the elevated ω in duplicated genes that are undergoing neofunctionalization. However, branch models average over all sites in the gene, so they miss cases where only a few sites were positively selected on the foreground branch. A gene-wide ω of 1.2 on a branch could reflect weak positive selection at many sites or strong positive selection at a few sites diluted by purifying selection elsewhere.

### Branch-Site Models

Branch-site models are the most powerful phylogenetic approach for detecting episodic positive selection. They allow ω to vary both among sites and among branches. The alternative model (Model A) specifies four site classes: (1) sites conserved throughout the tree (0 < ω < 1), (2) sites neutral throughout (ω = 1), (3) sites conserved in the background but positively selected on the foreground branch (ω > 1), and (4) sites neutral in the background but positively selected on the foreground branch (ω > 1). The null model fixes ω = 1 for the foreground classes.

The branch-site test is implemented in PAML (codeml, model = 2, NSsites = 2) and in HyPhy's aBSREL (adaptive Branch-Site Random Effects Likelihood). It has been used to identify positively selected sites in many contexts, including the adaptation of HIV-1 to human hosts, the evolution of venom proteins in snakes, and the emergence of new functions in duplicated genes.

A key advantage of branch-site models is their ability to detect positive selection that occurred on a single branch or a few branches, even when the gene-wide average ω is well below 1. This makes them the method of choice for testing hypotheses about specific evolutionary events, such as the adaptation of a pathogen to a new host or the evolution of a novel trait in a particular lineage.

## Population Genetic Approaches

### McDonald-Kreitman Test

The McDonald-Kreitman (MK) test compares polymorphism within species to divergence between species at synonymous and nonsynonymous sites. It is based on the prediction that, under neutrality, the ratio of nonsynonymous to synonymous polymorphism within a species should equal the ratio of nonsynonymous to synonymous divergence between species.

The test uses a 2 × 2 contingency table:

| | Nonsynonymous | Synonymous |
|---|---|---|
| **Polymorphism (within species)** | Pn | Ps |
| **Divergence (between species)** | Dn | Ds |

Under neutrality, Pn/Ps = Dn/Ds. An excess of nonsynonymous divergence relative to polymorphism (Dn/Ds > Pn/Ps) indicates that beneficial nonsynonymous mutations have been fixed by positive selection. The direction of the test is important: an excess of nonsynonymous polymorphism relative to divergence suggests either mildly deleterious mutations segregating in the population or balancing selection maintaining amino acid variation.

The neutrality index (NI) quantifies the deviation:

NI = (Pn/Ps) / (Dn/Ds)

An NI < 1 indicates an excess of nonsynonymous divergence, consistent with positive selection. The proportion of amino acid substitutions fixed by positive selection (α) can be estimated as 1 − NI, though this estimate is biased by slightly deleterious mutations and demographic history.

The MK test requires polymorphism data from a single species and divergence data from a closely related outgroup. It is robust to many assumptions but sensitive to demography: population expansions and bottlenecks distort the polymorphism spectrum and can bias the test. Methods such as the asymptotic MK test and the DoFE (Direction of Fixation Effect) approach attempt to correct for these biases.

### Frequency Spectrum Tests

Frequency spectrum tests use the distribution of allele frequencies at polymorphic sites to detect the genomic signatures of positive selection.

**Tajima's D** compares two estimators of the population mutation rate θ = 4Nₑμ: the average number of pairwise differences (π) and the number of segregating sites (S). Under neutrality and equilibrium demography, both estimate the same θ, so D ≈ 0. A selective sweep reduces π more than it reduces S, because the sweep eliminates intermediate-frequency variants while new mutations accumulate as rare variants. This produces a negative Tajima's D. However, negative D is also produced by population expansion and purifying selection, so it is not specific to positive selection.

**Fay and Wu's H** is more specific. It uses the derived allele frequency spectrum, weighting high-frequency derived alleles. A selective sweep produces an excess of high-frequency derived alleles (because the beneficial mutation and its linked neutral hitchhikers are at high frequency) and a deficit of intermediate-frequency variants. Fay and Wu's H is negative when there is an excess of high-frequency derived alleles, which is a more specific signature of a recent sweep than Tajima's D alone.

**Composite likelihood methods** such as SweepFinder and SweeD combine information from multiple sites across a genomic window to identify regions with the characteristic sweep signature: reduced diversity, skewed frequency spectrum, and extended linkage disequilibrium. These methods are widely used in population genomic scans for recent adaptation, such as the identification of sweeps at the lactase gene (LCT) in European populations and at the EDAR gene in East Asian populations.

The __MASK_6__ distinction is critical for interpreting frequency spectrum tests: negative selection also reduces diversity and skews the frequency spectrum, but it produces a different pattern (excess of rare derived alleles, no high-frequency derived excess). Distinguishing the two requires comparing multiple summary statistics or using model-based inference.

## Evidence and Examples of Positive Selection

### Classic Examples

**Immune genes.** The MHC (HLA in humans) genes are among the best-characterized examples of positive selection. The antigen-binding groove of MHC class I and class II molecules must accommodate a diverse array of pathogen-derived peptides. Positively selected sites cluster in the peptide-binding region, where amino acid changes alter the repertoire of peptides that can be presented. Site-model analyses consistently identify ω > 1 at these positions across mammals, birds, and fish. The pattern reflects a long-term host-pathogen arms race: pathogens evolve to evade presentation, and hosts evolve new binding specificities.

**Color vision.** The opsin genes responsible for color vision show clear signatures of positive selection. In primates, the duplication of the X-linked opsin gene that gave rise to separate red- and green-sensitive pigments was followed by positive selection at sites that shift the spectral sensitivity of the pigments. In cichlid fish of the African Great Lakes, opsin genes have undergone repeated duplication and positive selection, contributing to the spectacular diversity of color vision and coloration in these adaptive radiations. The spectral tuning sites (e.g., positions 180, 277, 285 in the human red opsin) show elevated dN/dS in lineages where color vision diversified.

**High-altitude adaptation.** The EPAS1 gene, encoding hypoxia-inducible factor 2α, shows strong evidence of positive selection in Tibetan populations. The selected haplotype is associated with reduced hemoglobin concentration at high altitude, preventing the polycythemia that afflicts lowland populations. Population genetic analyses identified a 78-kb region of the EPAS1 gene with an extended haplotype homozygosity signal and a striking divergence between Tibetan and Han Chinese populations. The selected allele introgressed into Tibetans from Denisovan archaic humans, providing a striking example of adaptive introgression.

### Recent Genome-wide Studies

Genome-wide scans for positive selection have identified hundreds of candidate loci in humans and other species. In humans, [the 1000 Genomes Project](/knowledge/bioinformatics/the-1000-genomes-project-computational-insights) and the [Human Genome Diversity Project](/blog/guides/human-genome-diversity-project) have enabled scans for selective sweeps, polygenic adaptation, and local adaptation.

**Lactase persistence.** The LCT gene region shows one of the strongest signals of recent positive selection in humans. The −13910*T allele, located in an enhancer element upstream of LCT, confers lactase persistence into adulthood. This allele rose to high frequency in European and African pastoralist populations within the last 10,000 years, with selection coefficients estimated at 0.01–0.05 per generation. The selective sweep is detectable as an extended haplotype homozygosity signal and a high derived allele frequency.

**Skin pigmentation.** Genes involved in melanin synthesis, including SLC24A5, SLC45A2, and TYR, show evidence of positive selection in non-African populations. The SLC24A5 allele (A111T) that lightens skin pigmentation shows a strong sweep signal in European populations, with the derived allele at near-fixation. The selection coefficient has been estimated at approximately 0.04–0.10, among the strongest detected in human evolution.

**Pathogen resistance.** The glucose-6-phosphate dehydrogenase (G6PD) gene shows signatures of balancing selection and positive selection in malaria-endemic regions. The A- variant, which confers resistance to severe malaria but causes hemolytic anemia when homozygous, shows elevated frequency in African populations. The signature is complex because the selective advantage is heterozygote-specific, producing a pattern distinct from a classic selective sweep.

Genome-wide studies in non-model organisms have identified positively selected genes involved in [environmental adaptation](/blog/careers/environmental-adaptation-how-organisms-adjust-and-what-it-means-for-careers), reproduction, and immune defense. For example, scans in stickleback fish identified genes under positive selection in freshwater populations that colonized post-glacial lakes, including genes involved in osmoregulation and skeletal development. In Drosophila, scans have identified positively selected genes involved in xenobiotic detoxification, reproduction, and immune response.

## Common Pitfalls and Misinterpretations

### Recombination and Gene Conversion

Recombination breaks down the association between selected and linked sites, weakening the signatures of selective sweeps. More importantly for dN/dS analysis, recombination violates the assumption that all sites in a gene share the same genealogical history. When recombination is present, different regions of a gene may have different trees, and the likelihood models used in PAML and HyPhy assume a single tree. This can produce false positives for positive selection, particularly in genes with high recombination rates.

Gene conversion, a recombination-related process that transfers sequence information between paralogous genes, is a particularly insidious problem. If two duplicated genes undergo gene conversion, the dN/dS ratio can be inflated because the conversion events are misidentified as nonsynonymous substitutions. This is especially problematic in gene families with high sequence similarity, such as the MHC genes and the ribosomal RNA genes.

Mitigation strategies include: (1) testing for recombination using programs such as GARD (Genetic Algorithm Recombination Detection) or PhiPack before dN/dS analysis; (2) analyzing only the largest non-recombining block if recombination is detected; (3) using methods that account for recombination, such as the omegaMap approach that jointly estimates recombination and selection; and (4) being cautious when interpreting dN/dS results from gene families known to undergo gene conversion.

### [Statistical Significance](/blog/guides/statistical-significance) vs. Biological Relevance

A common error is equating [statistical significance](/blog/guides/statistical-significance) with biological importance. A likelihood ratio test may identify a gene as being under positive selection with high confidence, but the biological relevance depends on the magnitude of the effect and the functional context.

A dN/dS of 1.1 at a single site, while statistically significant in a large dataset, may have negligible biological impact. Conversely, a dN/dS of 5 at a site that alters an enzyme's substrate specificity can be highly consequential. The posterior probability from BEB analysis should be interpreted alongside the estimated ω value and the functional annotation of the site.

Another pitfall is the failure to account for multiple testing. Genome-wide scans test thousands of genes, and the expected number of false positives is high. A gene with a nominal p-value of 0.01 in a scan of 10,000 genes has an expected 100 false positives. Correcting for multiple testing using the false discovery rate (FDR) or Bonferroni correction is essential.

**Alignment errors** are a frequent source of false positives. Misaligned codons create spurious nonsynonymous differences, inflating dN. This is particularly problematic in regions of low sequence similarity or with indels. Using alignment programs that account for codon structure (e.g., PRANK, MACSE) and manually inspecting alignments at positively selected sites is recommended.

**Saturation** is another issue. When sequences are highly diverged, synonymous sites may be saturated—multiple substitutions have occurred at the same site—causing dS to be underestimated and dN/dS to be inflated. This is a particular concern for comparisons across deep evolutionary timescales. Methods that model multiple hits, such as the maximum likelihood approaches in PAML, partially address this, but saturation remains a limitation.

**Demography** confounds population genetic tests. Tajima's D and related statistics are sensitive to population expansion, bottlenecks, and population structure. A negative Tajima's D can result from population expansion rather than positive selection. Demographic inference methods, such as those implemented in ∂a∂i or fastsimcoal2, can be used to generate null distributions that account for demography.

## Practical Summary and Best Practices

### Study Design Checklist

1. **Define the biological question.** Are you testing a specific hypothesis about a gene or lineage, or scanning the genome for novel candidates? The approach differs: hypothesis-driven studies use targeted tests (branch-site models, MK test), while scans use genome-wide methods (SweeD, iHS, dN/dS across all genes).

2. **Assemble high-quality sequence data.** For dN/dS analysis, obtain coding sequences with correct reading frames. For population genetic tests, obtain polymorphism data from a well-sampled population and an appropriate outgroup. Check for paralogs and pseudogenes that could confound the analysis.

3. **Align sequences carefully.** Use codon-aware alignment programs. Remove poorly aligned regions. Verify that the alignment preserves the reading frame.

4. **Test for recombination.** Run GARD or PhiPack on your alignment. If recombination is detected, either analyze the largest non-recombining block or use recombination-aware methods.

5. **Build a reliable phylogeny.** Use maximum likelihood or Bayesian methods. Check for long-branch attraction and rooting issues. The tree topology affects dN/dS estimates, particularly for branch and branch-site models.

6. **Run appropriate models.** For hypothesis-driven tests, use PAML codeml or HyPhy. Compare nested models with likelihood ratio tests. For genome scans, use multiple complementary methods and require concordance.

7. **Correct for multiple testing.** Apply FDR correction when testing many genes or sites.

8. **Validate candidates.** Check that positively selected sites map to functionally important regions (active sites, binding interfaces, antigenic epitopes). Consider structural modeling or functional assays to confirm biological relevance.

9. **Report effect sizes.** Report ω values, not just p-values. Report the number of positively selected sites and their locations. Provide the alignment and tree as supplementary data.

### Reporting Results

When reporting positive selection results, include: (1) the alignment length and number of sequences; (2) the tree topology and branch lengths; (3) the models compared and the likelihood ratio test statistics; (4) the estimated ω values for each site class; (5) the list of positively selected sites with posterior probabilities; (6) the recombination test results; and (7) the multiple testing correction applied.

For population genetic tests, report: (1) the sample size and population; (2) the polymorphism and divergence counts; (3) the test statistic and p-value; (4) the estimated α (proportion of adaptive substitutions); and (5) the demographic model used for the null distribution.

## Frequently Asked Questions

### What is positive selection in biology?

Positive selection is the evolutionary process by which beneficial mutations increase in frequency and become fixed in a population because they improve fitness. In molecular evolution, it is detected when the rate of amino acid–altering (nonsynonymous) substitutions exceeds the rate of silent (synonymous) substitutions, indicating that protein changes are being actively favored rather than removed or randomly fixed.

### How is positive selection detected?

Positive selection is detected using two main classes of methods. Phylogenetic methods compare synonymous and nonsynonymous substitution rates (dN/dS) across sites and lineages using codon-based likelihood models. Population genetic methods use polymorphism data within species, such as the McDonald-Kreitman test, which compares polymorphism to divergence, and frequency spectrum tests like Tajima's D and Fay and Wu's H, which detect the genomic signatures of selective sweeps.

### What does dN/dS > 1 indicate?

A dN/dS ratio greater than 1 indicates that nonsynonymous substitutions (amino acid changes) are being fixed at a higher rate than synonymous substitutions (silent changes). Since synonymous substitutions are assumed to be neutral, dN/dS > 1 means that amino acid changes are being actively favored by natural selection—the signature of positive selection at the molecular level.

### What is the difference between positive and purifying selection?

Positive selection favors beneficial mutations and increases their frequency, producing dN/dS > 1. Purifying selection removes deleterious mutations and maintains conserved sequences, producing dN/dS < 1. Most protein-coding genes are under purifying selection for most of their length, with positive selection acting only at specific sites or during specific evolutionary episodes. The distinction is central to interpreting molecular evolution data, as detailed in __MASK_7__.

### What are branch-site models in positive selection analysis?

Branch-site models are phylogenetic likelihood models that allow the dN/dS ratio (ω) to vary both among codon sites and among branches of the phylogeny. They are designed to detect episodic positive selection: sites that are conserved in most lineages but experienced adaptive amino acid changes on specific foreground branches. The alternative model allows ω > 1 on the foreground branch at a subset of sites, and a likelihood ratio test compares this to a null model where ω = 1 on the foreground branch.

### Can positive selection be detected in non-coding regions?

Yes, but the standard dN/dS framework does not apply because there are no synonymous and nonsynonymous sites. For non-coding regions, positive selection is detected using population genetic methods that identify selective sweep signatures, such as reduced diversity, skewed allele frequency spectra, and extended haplotype homozygosity. [Comparative genomics](/blog/guides/comparative-genomics) approaches can also detect accelerated substitution rates in conserved non-coding elements, though distinguishing positive selection from relaxed constraint requires careful analysis.

### What is a selective sweep?

A selective sweep is the process by which a beneficial mutation rises to fixation and reduces genetic diversity at linked sites. Because the selected allele carries neighboring variants with it (hitchhiking), the region around the selected site shows reduced polymorphism, an excess of rare variants, and extended [linkage disequilibrium](/knowledge/bioinformatics/linkage-disequilibrium-and-haplotype-mapping). The size of the swept region depends on the strength of selection and the local recombination rate. Detecting selective sweeps is a primary goal of population genomic scans for recent adaptation.

## Key Takeaways

- Positive selection is the fixation of beneficial mutations, detected when the nonsynonymous substitution rate exceeds the synonymous rate (dN/dS > 1).
- Most protein-coding genes are under purifying selection; positive selection is concentrated at specific sites, such as antigen-binding grooves, enzyme active sites, and protein interaction interfaces.
- Phylogenetic methods (site, branch, and branch-site models) detect positive selection using codon-based likelihood frameworks, with branch-site models being the most powerful for episodic selection.
- Population genetic methods (McDonald-Kreitman test, Tajima's D, Fay and Wu's H, composite likelihood scans) detect recent or ongoing positive selection using polymorphism data.
- Recombination, alignment errors, saturation, and demographic history are major sources of false positives; these must be addressed with appropriate tests and corrections.
- Classic examples include MHC genes, opsin genes, lactase persistence, and high-altitude adaptation, illustrating the diversity of biological contexts where positive selection operates.
- Best practices require careful study design, multiple complementary methods, multiple testing correction, and validation of candidate sites in functional context.

## Further Reading

- Paroni G et al. *HER2-positive breast-cancer cell lines are sensitive to KDM5 inhibition: definition of a gene-expression model for the selection of sensitive cases*. Oncogene. 2019. [PubMed 30538297](https://doi.org/10.1038/s41388-018-0620-6)
- Li HF. *Optimal validation of accuracy in antibody assays and reasonable definition of antibody positive/negative subgroups in neuroimmune diseases: a narrative review*. Annals of translational medicine. 2023. [PubMed 37090045](https://doi.org/10.21037/atm-21-2384)
- Chiumello D et al. *Bedside selection of positive end-expiratory pressure in mild, moderate, and severe acute respiratory distress syndrome*. Critical care medicine. 2014. [PubMed 24196193](https://doi.org/10.1097/CCM.0b013e3182a6384f)
- Brac B et al. *Is There an Optimal Definition for a Positive Circumferential Resection Margin in Locally Advanced Esophageal Cancer?*. Annals of surgical oncology. 2021. [PubMed 34514523](https://doi.org/10.1245/s10434-021-10707-6)



<div data-calculator="molecular-cloning"></div>

## Related Clinical & Scientific Guides

* [MAPK Pathway: Mechanism, Function, and Clinical Relevance](/knowledge/molecular-biology/mapk-pathway)
* [Mammalian Cell Culture Bioreactors: A Practical Guide](/knowledge/molecular-biology/mammalian-cell-culture-bioreactor)
* [Nucleotide Formation: Biosynthesis and Assembly of DNA/RNA Building Blocks](/knowledge/molecular-biology/nucleotide-formation)