Molecular Evidence of Evolution: Mechanisms and Methods
By Dr. Zubair Khalid, DVM, MS, PhD ·

Introduction to Molecular Evidence of Evolution
What is Molecular Evidence?
Molecular evidence of evolution refers to the heritable changes in DNA, RNA, and protein sequences that document the shared ancestry of organisms. Unlike morphological evidence, which compares anatomical structures across species, molecular evidence operates at the level of nucleotide and amino acid sequences, revealing evolutionary relationships that may be invisible at the phenotypic level. The central premise is that all organisms share a common genetic code and homologous molecular machinery—DNA polymerase, ribosomes, and metabolic enzymes—that has been inherited and modified over billions of years.
The power of molecular evidence lies in its quantifiability. A DNA sequence is a linear string of four nucleotides (A, T, G, C); a protein is a linear string of twenty amino acids. These sequences can be aligned, compared, and subjected to statistical models that estimate the probability of observed differences under various evolutionary scenarios. This permits hypothesis testing in a way that morphological comparison rarely allows, because the mutational processes that generate sequence divergence are far better understood than the developmental processes that generate morphological variation.
Historical Context and Emergence
The molecular approach to evolution emerged in the mid-20th century, when protein sequencing became feasible. In the 1950s and 1960s, Frederick Sanger's methods for sequencing insulin and later proteins enabled direct comparison of homologous proteins across species. Emile Zuckerkandl and Linus Pauling's 1965 observation that hemoglobin sequences accumulate changes at a roughly constant rate gave rise to the molecular clock hypothesis. This was followed by the neutral theory of molecular evolution, proposed by Motoo Kimura in 1968, which argued that most molecular variation is selectively neutral and fixed by genetic drift rather than natural selection. The Neutral Theory of Molecular Evolution remains a foundational framework for interpreting molecular data, because it provides a null model against which selection can be detected.
The advent of Sanger DNA sequencing in 1977, polymerase chain reaction (PCR) amplification in the 1980s, and high-throughput sequencing in the 2000s transformed the field. Today, whole-genome sequencing of any organism is routine, and comparative genomics provides evolutionary evidence at a scale unimaginable to earlier researchers.
Sources of Molecular Evidence
DNA Sequence Data
DNA sequences are the most direct source of molecular evidence. Three categories of DNA sequence data are commonly used:
Protein-coding genes. These sequences are constrained by the genetic code and the functional requirements of the encoded protein. Coding sequences are divided into codons, and substitutions within a codon can be synonymous (silent, not changing the amino acid) or nonsynonymous (changing the amino acid). The ratio of nonsynonymous to synonymous substitution rates (dN/dS, often denoted ω) provides a powerful test for selection: ω = 1 indicates neutral evolution, ω < 1 indicates purifying selection, and ω > 1 indicates positive selection.
Non-coding DNA. Introns, intergenic regions, and regulatory sequences evolve under different constraints than coding regions. Because many non-coding regions are less functionally constrained, they accumulate substitutions more rapidly and are useful for resolving recent divergences. However, some non-coding regions—such as promoters, enhancers, and microRNA binding sites—are highly conserved and provide evidence of functional constraint across deep evolutionary timescales.
Mitochondrial and chloroplast DNA. Organellar genomes are typically maternally inherited, lack recombination (in most animals), and have elevated mutation rates relative to nuclear DNA. The mitochondrial gene cytochrome c oxidase subunit I (COI) is the standard barcode for animal species identification because it evolves fast enough to distinguish closely related species yet is conserved enough to be reliably amplified with universal primers. A typical PCR amplification of COI uses an annealing temperature of 50–55°C with primers LCO1490 and HCO2198, yielding a ~658 bp fragment.
Protein Sequence Data
Protein sequences provide evidence that is complementary to DNA data. Because the genetic code is degenerate, protein sequences are more conserved than the underlying DNA sequences over long evolutionary timescales. This makes proteins useful for resolving ancient divergences where DNA sequences have become saturated with multiple substitutions at the same site.
Cytochrome c, a small heme protein involved in electron transport, is a classic example. It is found in all eukaryotes and some prokaryotes. Human and chimpanzee cytochrome c are identical; human and rhesus monkey differ by one amino acid; human and yeast differ by approximately 45 amino acids. These differences correlate with phylogenetic distance, not with organismal complexity, demonstrating that molecular divergence tracks evolutionary time rather than phenotypic advancement.
Protein sequence data also reveal the effects of selection through the pattern of amino acid replacements. Conservative substitutions (e.g., leucine to isoleucine, both hydrophobic) are more common than radical substitutions (e.g., leucine to lysine, hydrophobic to charged), reflecting the functional constraints imposed by protein structure.
Genomic Structural Variations
Beyond point mutations, genomes carry structural variations that serve as evolutionary evidence. These include:
- Insertions and deletions (indels): These can shift reading frames when they occur in coding regions, usually with deleterious consequences. In non-coding regions, they accumulate more freely and can serve as phylogenetic markers.
- Gene duplications: Duplicate genes provide raw material for evolutionary innovation. After duplication, one copy may retain the original function while the other is free to acquire new functions (neofunctionalization) or partition the original functions (subfunctionalization).
- Chromosomal rearrangements: Inversions, translocations, and fusions change gene order. Shared derived rearrangements can be powerful phylogenetic markers because they are unlikely to arise independently.
- Transposable element insertions: The presence or absence of a transposable element at a specific genomic location is a nearly homoplasmy-free marker, because the probability of two independent insertions at exactly the same site is vanishingly small.
Mechanisms Generating Molecular Variation
Mutation and Substitution
Mutation is the ultimate source of all genetic variation. At the molecular level, mutations arise from replication errors, DNA damage, and the activity of mutagens. The raw mutation rate varies by organism and genomic region; in humans, the germline mutation rate is approximately 1.2 × 10⁻⁸ substitutions per site per generation.
A substitution is a mutation that has become fixed in a population. The distinction is crucial: most mutations are lost by drift or purifying selection, and only a fraction ever reach fixation. The substitution rate (often denoted k) is the rate at which new variants become fixed, and it is this rate that molecular evolutionists measure.
The mechanisms of mutation include:
- Base substitution: A single nucleotide is replaced by another. Transitions (purine-to-purine or pyrimidine-to-pyrimidine: A↔G, C↔T) occur more frequently than transversions (purine-to-pyrimidine) because of the chemistry of tautomeric shifts and deamination.
- Deamination: Cytosine deaminates to uracil at a measurable rate under physiological conditions (pH 7.4, 37°C). If unrepaired, this produces a C→T transition. Methylated cytosines (5-methylcytosine) deaminate to thymine, which is not recognized as a lesion, making CpG dinucleotides mutation hotspots.
- Oxidative damage: Reactive oxygen species generate 8-oxoguanine, which mispairs with adenine, causing G→T transversions.
- Replication slippage: In regions of tandem repeats, DNA polymerase can slip, leading to expansion or contraction of repeat number. This is the basis of microsatellite variation.
Genetic Drift and Effective Population Size
Genetic drift is the random fluctuation of allele frequencies due to sampling effects in finite populations. The rate of drift is inversely proportional to the effective population size (Nₑ), which is usually smaller than the census population size because of unequal sex ratios, fluctuating population sizes, and variance in reproductive success.
The fixation probability of a new neutral mutation is 1/(2Nₑ), and the rate of neutral substitution equals the mutation rate (k = μ), independent of population size. This counterintuitive result—that the neutral substitution rate does not depend on Nₑ—is a cornerstone of the Concept of Neutral Evolution. For selected mutations, the fixation probability depends on both the selection coefficient (s) and Nₑ. A mutation with a small beneficial effect (s > 1/Nₑ) has a fixation probability of approximately 2s, while a deleterious mutation (s < 0) is unlikely to fix unless |s| < 1/Nₑ.
The __MASK_3__ predicts that functionally constrained regions (e.g., protein-coding exons) evolve more slowly than unconstrained regions (e.g., pseudogenes), because most mutations in constrained regions are deleterious and removed by purifying selection. This prediction is strongly supported by empirical data: the substitution rate in pseudogenes is approximately equal to the mutation rate, while the rate in coding regions is typically 5–10 times lower.
Natural Selection and Molecular Adaptation
Natural selection acts on molecular variation when differences in fitness are associated with genetic differences. At the molecular level, selection can be detected through several signatures:
Purifying selection removes deleterious mutations. It is detected by a deficit of nonsynonymous substitutions relative to synonymous substitutions (dN/dS < 1) and by conservation of amino acid sequences across species.
Positive selection favors advantageous mutations. It is detected by an excess of nonsynonymous substitutions (dN/dS > 1), by the fixation of amino acid changes that alter protein function, and by population genetic tests such as the McDonald-Kreitman test, which compares polymorphism within species to divergence between species.
Balancing selection maintains multiple alleles in a population. It is detected by an excess of intermediate-frequency polymorphisms and by trans-species polymorphism, where identical alleles are shared across species because they have been maintained by selection since before speciation. The major histocompatibility complex (MHC) genes in vertebrates are a classic example.
Phylogenetic Inference from Molecular Data
Sequence Alignment
Phylogenetic inference begins with sequence alignment—the arrangement of homologous nucleotides or amino acids into columns that reflect common ancestry. Alignment is the most critical and error-prone step in molecular phylogenetics; errors here propagate through every subsequent analysis.
For protein-coding sequences, alignment is guided by the reading frame: codons must be aligned as units, and gaps should be placed between codons to preserve the frame. For protein sequences, alignment is guided by structural and biochemical properties of amino acids. Automated alignment programs such as MAFFT and MUSCLE use progressive or iterative algorithms. For example, MAFFT with the L-INS-i strategy is recommended for sequences with conserved motifs and variable regions; it uses a fast Fourier transform to identify homologous regions before refinement.
Manual curation is often necessary, especially for sequences with large insertions or deletions. A common rule of thumb is that the alignment should be scrutinized in regions where the automated program places gaps, because these are the regions where homology is most uncertain.
Substitution Models
Substitution models describe the relative rates of different nucleotide or amino acid changes. The simplest is the Jukes-Cantor (JC69) model, which assumes equal base frequencies and equal substitution rates among all nucleotides. More realistic models include:
- Kimura 2-parameter (K80): Distinguishes transition rate (α) from transversion rate (β).
- Hasegawa-Kishino-Yano (HKY85): Adds unequal base frequencies to K80.
- General time-reversible (GTR): Allows all six pairwise substitution rates to differ, plus unequal base frequencies.
For protein sequences, empirical amino acid substitution matrices such as JTT (Jones-Taylor-Thornton), WAG (Whelan and Goldman), and LG (Le and Gascuel) are used. These matrices are estimated from large databases of aligned protein sequences and provide the relative rates of each amino acid replacement.
Model selection is performed using information criteria such as the Akaike Information Criterion (AIC) or the Bayesian Information Criterion (BIC). The AIC is calculated as AIC = 2k − 2ln(L), where k is the number of parameters and L is the maximum likelihood. The model with the lowest AIC is preferred. Software such as jModelTest or ModelFinder automates this process, testing up to 88 candidate models.
Tree-Building Methods
Three main approaches are used to build phylogenetic trees from molecular data:
Maximum likelihood (ML). ML methods search for the tree and model parameters that maximize the probability of observing the data. For each site, the likelihood is calculated by summing over all possible ancestral states at internal nodes. The likelihood of the entire alignment is the product of site likelihoods, assuming sites evolve independently. ML is statistically consistent (converges to the true tree with sufficient data) and allows model-based correction for multiple substitutions. Programs such as RAxML and IQ-TREE implement efficient ML searches; IQ-TREE with the -m TEST option automatically selects the best model and performs 1000 ultrafast bootstrap replicates for branch support.
Bayesian inference. Bayesian methods combine the likelihood with prior probabilities on parameters and trees, producing a posterior distribution. The posterior probability of a tree is proportional to the likelihood times the prior. Because the space of possible trees is enormous, Bayesian inference uses Markov chain Monte Carlo (MCMC) sampling. MrBayes is the most widely used program; a typical run uses four chains (three heated, one cold) for 1–10 million generations, sampling every 1000 generations. Convergence is assessed by the standard deviation of split frequencies (< 0.01) and effective sample sizes (ESS > 200) for all parameters. The first 25% of samples are discarded as burn-in.
Maximum parsimony. Parsimony minimizes the total number of evolutionary changes required to explain the data. It makes no explicit model assumptions and can be statistically inconsistent under certain conditions—particularly when substitution rates vary among lineages—a phenomenon known as long-branch attraction. However, parsimony remains useful for analyzing morphological characters and for datasets with low divergence where multiple hits are rare.
The choice among methods depends on the research question and data characteristics. For most modern phylogenomic analyses, ML and Bayesian methods are preferred because they explicitly model the substitution process. For a detailed treatment of these methods, see Molecular Phylogenetics and Evolution.
Molecular Clocks and Divergence Time Estimation
Rate Variation and Relaxed Clocks
The molecular clock hypothesis posits that sequence divergence accumulates at a roughly constant rate over time. If true, the number of differences between two sequences is proportional to the time since their last common ancestor. The Molecular Clock Hypothesis was first proposed by Zuckerkandl and Pauling based on hemoglobin sequences, and it remains a central tool in evolutionary biology.
However, strict clock-like evolution is rare. Substitution rates vary among lineages due to differences in generation time, metabolic rate, DNA repair efficiency, and selective constraints. This has led to the development of relaxed clock models that allow rates to vary across the tree. The uncorrelated relaxed clock model, implemented in BEAST, assumes that each branch has its own rate drawn from a lognormal or exponential distribution. This model does not require rates to be autocorrelated between ancestor and descendant branches, making it flexible enough to accommodate diverse biological scenarios.
The Molecular Clock Model is not a single method but a family of models ranging from strict to fully relaxed. The choice of model is typically made using model comparison criteria such as the Bayes factor, which is the ratio of marginal likelihoods of two competing models.
Calibration with Fossils and Biogeography
A molecular clock provides relative times (branch lengths in substitutions per site) that must be converted to absolute times (millions of years) using calibrations. The most common calibrations come from the fossil record, which provides minimum ages for the divergence of lineages.
Calibration approaches include:
- Node calibration: A fossil is assigned to a specific node in the phylogeny, providing a minimum age for that node. For example, the oldest known fossil of the genus Homo is approximately 2.8 million years old, providing a minimum calibration for the split between Homo and Australopithecus.
- Tip calibration: Fossils are included as terminal taxa in the analysis, with their ages used to calibrate the tree.
- Total evidence dating: Morphological and molecular data are analyzed jointly, with fossils providing both phylogenetic placement and temporal constraints.
Biogeographic calibrations use known geological events, such as the separation of continents, to constrain divergence times. For example, the opening of the Atlantic Ocean (~100 million years ago) provides an upper bound for the divergence of taxa that are now separated by that ocean.
A typical BEAST analysis for divergence time estimation involves: (1) specifying the sequence alignment and substitution model, (2) specifying the clock model (e.g., uncorrelated lognormal relaxed clock), (3) specifying the tree prior (e.g., Yule process or birth-death), (4) adding calibration priors to nodes, and (5) running MCMC for 50–100 million generations. The output is a posterior distribution of trees with divergence times, summarized as a maximum clade credibility tree with 95% highest posterior density intervals on node ages.
For further reading on the theory and applications of molecular clocks, see Molecular Clock Studies and Molecular Clock Definition.
Comparative Genomics and Molecular Evidence
Conserved Non-coding Elements
Whole-genome comparisons have revealed that a substantial fraction of the genome is under purifying selection even though it does not code for proteins. In the human genome, approximately 5% of bases are estimated to be under purifying selection, but only ~1.5% code for proteins. The remaining constrained bases are in conserved non-coding elements (CNEs), which include promoters, enhancers, insulators, and non-coding RNA genes.
CNEs are identified by comparing genomes across species and identifying regions with unusually high sequence conservation. For example, the ultraconserved elements (UCEs) are regions of ≥200 bp that are 100% identical between human, mouse, and rat. Many UCEs function as enhancers during development, and their extreme conservation indicates strong purifying selection.
The evolutionary evidence from CNEs is twofold. First, their conservation across species separated by hundreds of millions of years demonstrates that they are functional and that purifying selection has maintained them. Second, their patterns of gain and loss across lineages provide phylogenetic information. For example, the loss of a conserved enhancer in a particular lineage can be associated with morphological changes in that lineage.
Pseudogenes as Evolutionary Fossils
Pseudogenes are non-functional copies of genes that have accumulated inactivating mutations—premature stop codons, frameshifts, or loss of regulatory elements. They are powerful evidence for evolution because they reveal the historical trajectory of genomes.
There are two main types:
Processed pseudogenes arise from retrotransposition: an mRNA is reverse-transcribed into cDNA and inserted into the genome. Because they lack introns and promoters, they are usually non-functional from birth. The human genome contains approximately 14,000 processed pseudogenes.
Non-processed (duplicated) pseudogenes arise from gene duplication followed by the accumulation of inactivating mutations in one copy. The globin gene family provides a classic example: the human β-globin cluster on chromosome 11 contains five functional genes (ε, Gγ, Aγ, δ, β) and one pseudogene (ψβ). The ψβ pseudogene has a premature stop codon and is not transcribed into a functional protein, yet its sequence is clearly homologous to the functional β-globin genes.
The evolutionary evidence from pseudogenes is compelling because they are "fossils" of genes that were once functional. Their presence in the genome can only be explained by common ancestry followed by loss of function. Moreover, pseudogenes evolve at the neutral mutation rate, making them useful for estimating mutation rates and calibrating molecular clocks.
Transposable Elements and Genome Dynamics
Transposable elements (TEs) are mobile genetic elements that comprise a substantial fraction of eukaryotic genomes—approximately 45% of the human genome and up to 85% of some plant genomes. TEs are classified into two major classes:
Class I (retrotransposons) move via an RNA intermediate. They include long interspersed nuclear elements (LINEs), short interspersed nuclear elements (SINEs), and long terminal repeat (LTR) retrotransposons. LINE-1 (L1) elements are the most abundant in mammals, comprising ~17% of the human genome.
Class II (DNA transposons) move directly as DNA. They are more common in bacteria and plants than in mammals.
TEs provide evolutionary evidence in several ways:
- Shared insertions: If a TE inserts at a specific genomic location in a common ancestor, all descendant species will share that insertion. The probability of independent insertions at the same site is negligible, making TE insertions nearly homoplasy-free phylogenetic markers.
- TE families as molecular fossils: Different TE families were active at different times in evolutionary history. The age distribution of TE copies within a genome reflects the history of TE activity and can be used to date genomic events.
- TE-mediated genome rearrangements: TEs can cause chromosomal rearrangements through unequal recombination between homologous elements at different locations, contributing to genome evolution.
Case Studies: Molecular Evidence in Action
Globin Gene Family Evolution
The globin gene family is one of the best-studied examples of molecular evolution. Hemoglobin is a tetramer of two α-like and two β-like globin chains, each bound to a heme group. In vertebrates, the α-like globin genes are clustered on one chromosome (chromosome 16 in humans) and the β-like genes on another (chromosome 11 in humans).
The evolutionary history of the globin family involves a series of gene duplications:
- The ancestral globin gene duplicated to give rise to myoglobin (muscle oxygen storage) and hemoglobin (blood oxygen transport) approximately 500–800 million years ago.
- The hemoglobin gene duplicated to give rise to the α and β lineages approximately 450–500 million years ago.
- Within each lineage, further duplications produced the developmental stage-specific genes: ζ and α in the α cluster; ε, Gγ, Aγ, δ, and β in the β cluster.
Each duplication event is documented by the sequence similarity of the duplicated genes. The α and β chains share approximately 45% amino acid identity, consistent with their ancient divergence. The developmental globins (ε, Gγ, Aγ) share higher identity with each other than with adult globins (δ, β), reflecting their more recent common ancestry.
The ψβ pseudogene in the β cluster provides additional evidence. Its sequence is most similar to the β gene, indicating that it arose from a duplication of an ancestral β-like gene. The pseudogene has accumulated mutations at the neutral rate, making it a "living fossil" that records the mutational history of the lineage.
Cytochrome c and Universal Conservation
Cytochrome c is a small (about 104 amino acids in vertebrates) heme-containing protein that transfers electrons between complex III and complex IV of the mitochondrial electron transport chain. It is found in all eukaryotes and in some prokaryotes, making it one of the most universally conserved proteins known.
The amino acid sequence of cytochrome c has been determined for hundreds of species. The pattern of conservation is striking:
- 35 of the 104 amino acids are identical in all species examined, from yeast to humans.
- The conserved residues include the heme-binding cysteines (positions 14 and 17), the histidine and methionine that coordinate the heme iron (positions 18 and 80), and residues involved in protein-protein interactions with cytochrome c oxidase and reductase.
- The variable residues are predominantly on the protein surface, where they are less constrained by function.
The evolutionary evidence from cytochrome c is twofold. First, the universal conservation of the protein across all eukaryotes demonstrates common ancestry of the eukaryotic lineage. Second, the pattern of variation—conserved core, variable surface—demonstrates the action of purifying selection maintaining functionally important residues while allowing neutral or nearly neutral changes elsewhere.
The number of amino acid differences between species correlates with phylogenetic distance. Human and chimpanzee cytochrome c are identical; human and rhesus monkey differ by 1 amino acid; human and dog differ by 10; human and yeast differ by 45. This correlation between sequence divergence and evolutionary distance is exactly what the molecular clock hypothesis predicts.
Human and Chimpanzee Genomes
The comparison of human and chimpanzee genomes provides molecular evidence at the genomic scale. The two genomes differ by approximately 1.2% at the nucleotide level (about 35 million single nucleotide differences), plus several million indels and larger structural variants.
Key findings include:
- Divergence time: The average autosomal divergence of 1.2% corresponds to a divergence time of approximately 6–8 million years, using a mutation rate of 1 × 10⁻⁹ substitutions per site per year.
- Heterogeneity of divergence: Divergence is not uniform across the genome. The Y chromosome shows lower divergence (~0.9%) than the autosomes, while the X chromosome shows intermediate divergence. This pattern reflects differences in mutation rates and effective population sizes between chromosomes.
- Selection differences: Genes involved in olfaction, immune response, and spermatogenesis show evidence of positive selection in the human lineage, while genes involved in keratin production show evidence of positive selection in the chimpanzee lineage.
- Gene loss: The human genome has lost functional copies of several genes that are intact in chimpanzees, including the MYH16 gene (myosin heavy chain 16), which is expressed in jaw muscles. The inactivation of MYH16 in humans is associated with the reduction of jaw muscle size, which may have facilitated the expansion of the braincase.
The human-chimpanzee comparison also reveals the action of purifying selection: functional regions of the genome show reduced divergence compared to non-functional regions, and the ratio of nonsynonymous to synonymous divergence (dN/dS) is less than 1 for most genes, indicating that most amino acid changes have been removed by selection.
Common Pitfalls and Best Practices in Molecular Evolutionary Analysis
Alignment Errors and Biases
Alignment errors are a major source of phylogenetic error. Misaligned sites are treated as homologous when they are not, leading to inflated divergence estimates and incorrect tree topologies. Common problems include:
- Gap placement: Automated aligners often place gaps incorrectly in regions of low similarity. Manual curation is essential, particularly for datasets with variable-length insertions or deletions.
- Saturation: In highly divergent sequences, multiple substitutions at the same site erase the phylogenetic signal. This is particularly problematic for third codon positions, which are often saturated for deep divergences.
- Over-alignment: Forcing alignment of non-homologous regions creates false homology. The solution is to remove ambiguously aligned regions using programs such as Gblocks or trimAl, which identify and exclude poorly aligned positions.
Best practice: Always inspect alignments visually. Use multiple alignment programs and compare results. Remove ambiguously aligned regions before phylogenetic analysis. For protein-coding genes, align codons, not individual nucleotides.
Model Misspecification
Using an incorrect substitution model can lead to systematic errors in phylogenetic inference. Common problems include:
- Ignoring rate heterogeneity: Most datasets have substantial variation in substitution rates among sites. Models that ignore this (e.g., JC69) underestimate divergence and can produce incorrect trees. The standard solution is to use a Γ distribution with a proportion of invariant sites (GTR+I+Γ).
- Ignoring base composition bias: If base frequencies differ among taxa, models that assume homogeneous base frequencies can be misled. Compositional bias can be detected with a chi-square test, and compositionally heterogeneous models (e.g., the non-stationary model in nhPhyML) should be used when bias is detected.
- Ignoring codon structure: For protein-coding sequences, codon models that account for the genetic code and selection on amino acid changes are more appropriate than nucleotide models.
Best practice: Use model selection criteria (AIC, BIC) to choose the best-fitting model. Test for compositional bias. Use codon models for protein-coding sequences when appropriate.
Long-Branch Attraction
Long-branch attraction (LBA) is a systematic error in which rapidly evolving lineages are incorrectly grouped together because they share many convergent changes. This is a particular problem for parsimony methods but can also affect ML and Bayesian methods when models are misspecified.
LBA is most severe when:
- Divergence times are deep and substitution rates are high.
- Base composition is biased in the long-branch taxa.
- The substitution model underestimates the probability of multiple hits.
Best practice: Use model-based methods (ML, Bayesian) with realistic models. Include taxa that break up long branches. Use methods that are robust to LBA, such as those that account for composition heterogeneity. Test for LBA by removing suspected long-branch taxa and seeing if the topology changes.
Gene Trees vs. Species Trees
Individual gene trees can differ from the species tree due to incomplete lineage sorting (ILS), gene duplication and loss, and horizontal gene transfer. ILS is particularly problematic for recently diverged species with large ancestral population sizes.
The multispecies coalescent model accounts for ILS by modeling the probability that gene trees differ from the species tree due to the random sorting of ancestral lineages. Methods such as ASTRAL and BEAST2 with the StarBEAST package implement coalescent-based species tree inference.
Best practice: When analyzing multiple genes, consider the possibility of gene tree discordance. Use coalescent-based methods for species tree inference. Test for the presence of ILS using programs such as D-statistics (ABBA-BABA test) or PhyloNet.
Frequently Asked Questions
What is molecular evidence of evolution?
Molecular evidence of evolution is the set of observations from DNA, RNA, and protein sequences that document the shared ancestry of organisms. It includes patterns of sequence similarity and divergence, the presence of pseudogenes and transposable elements, and the conservation of functional elements across species. This evidence supports the conclusion that all organisms share common ancestors and that the diversity of life has arisen through descent with modification.
How does molecular biology provide evidence for evolution?
Molecular biology provides evidence for evolution through several lines of observation: (1) the universal genetic code and shared molecular machinery (e.g., ribosomes, DNA polymerase) across all life; (2) the hierarchical pattern of sequence similarity that matches known phylogenetic relationships; (3) the presence of non-functional remnants of genes (pseudogenes) and mobile elements that record historical genomic events; (4) the observation that molecular divergence increases with time since common ancestry, as predicted by the molecular clock; and (5) the detection of natural selection at the molecular level through patterns of synonymous and nonsynonymous substitution.
What are some examples of molecular evidence for evolution?
Classic examples include: (1) the globin gene family, which shows a history of gene duplications followed by functional divergence; (2) cytochrome c, which is conserved across all eukaryotes with a pattern of variation that tracks phylogenetic distance; (3) the human-chimpanzee genome comparison, which shows ~1.2% nucleotide divergence and shared transposable element insertions; (4) processed pseudogenes, which are non-functional copies of genes created by retrotransposition; and (5) conserved non-coding elements, which show that purifying selection has maintained regulatory sequences across hundreds of millions of years.
How do scientists use DNA sequences to study evolution?
Scientists use DNA sequences to study evolution by: (1) aligning homologous sequences from different species; (2) fitting substitution models that describe the process of sequence change; (3) reconstructing phylogenetic trees using maximum likelihood, Bayesian, or parsimony methods; (4) estimating divergence times using molecular clocks calibrated with fossils or biogeographic events; and (5) detecting selection by comparing rates of synonymous and nonsynonymous substitution or by using population genetic tests.
What is a molecular clock and how is it used?
A molecular clock is a model that relates the amount of sequence divergence to the time since two lineages diverged from a common ancestor. It is based on the observation that substitutions accumulate at a roughly constant rate over time. Molecular clocks are used to estimate divergence times when fossil evidence is absent or incomplete. Modern relaxed clock models allow rates to vary among lineages, and calibration with fossils or known geological events converts relative times into absolute times. For more details, see Molecular Clock in Evolution.
Why are pseudogenes considered evidence of evolution?
Pseudogenes are considered evidence of evolution because they are non-functional copies of genes that can only be explained by common ancestry followed by loss of function. A pseudogene in the human genome that is also present in the chimpanzee genome at the same location must have been inherited from a common ancestor. The accumulation of inactivating mutations in pseudogenes at the neutral rate provides a molecular record of the time since the gene became non-functional.
What are common mistakes in interpreting molecular evidence?
Common mistakes include: (1) equating sequence similarity with functional similarity without considering homology; (2) ignoring the effects of selection and assuming all divergence is neutral; (3) using inappropriate substitution models that do not account for rate heterogeneity or composition bias; (4) overinterpreting gene trees as species trees without considering incomplete lineage sorting; (5) failing to calibrate molecular clocks properly, leading to unrealistic divergence time estimates; and (6) neglecting the possibility of horizontal gene transfer, particularly in prokaryotes.
Key Takeaways
- Molecular evidence of evolution derives from DNA, RNA, and protein sequences, and it provides quantifiable, testable support for common ancestry that complements morphological evidence.
- The Neutral Theory of Molecular Evolution provides the null model for interpreting molecular variation, with selection detected as deviations from neutral expectations.
- Mutation generates variation, genetic drift fixes neutral variants, and natural selection shapes the patterns of molecular divergence and conservation.
- Phylogenetic inference from molecular data requires careful sequence alignment, appropriate substitution models, and statistically rigorous tree-building methods such as maximum likelihood and Bayesian inference.
- Molecular clocks, including relaxed clock models, allow estimation of divergence times when properly calibrated with fossils or biogeographic events.
- Comparative genomics reveals evolutionary evidence in conserved non-coding elements, pseudogenes, and transposable element insertions, which are nearly homoplasy-free phylogenetic markers.
- Common pitfalls in molecular evolutionary analysis include alignment errors, model misspecification, long-branch attraction, and conflating gene trees with species trees; these can be mitigated with best practices such as model testing, coalescent-based methods, and careful data curation.
Further Reading
- Cobo-Simón M, Hart R, Ochman H. Escherichia Coli: What Is and Which Are?. Molecular biology and evolution. 2023. PubMed 36585846
- Lan D et al. Genetic Diversity, Molecular Phylogeny, and Selection Evidence of Jinchuan Yak Revealed by Whole-Genome Resequencing. G3 (Bethesda, Md.). 2018. PubMed 29339406
- Szudarek-Trepto N, Kazmierski A, Dabert J. Long-term stasis in acariform mites provides evidence for morphologically stable evolution: Molecular vs. morphological differentiation in Linopodes (Acariformes; Prostigmata). Molecular phylogenetics and evolution. 2021. PubMed 34147656
- Christaki E, Marcou M, Tofarides A. Antimicrobial Resistance in Bacteria: Mechanisms, Evolution, and Persistence. Journal of molecular evolution. 2020. PubMed 31659373
- Kimura M. The neutral theory of molecular evolution: a review of recent evidence. Idengaku zasshi. 1991. PubMed 1954033
- Goldberg J, Trewick SA, Paterson AM. Evolution of New Zealand's terrestrial fauna: a review of molecular evidence. Philosophical transactions of the Royal Society of London. Series B, Biological sciences. 2008. PubMed 18782728