Paralog and Ortholog: Mechanisms, Detection, and Pitfalls
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Orthologs, genes in different species descended from a single ancestral gene via speciation, are generally expected to retain ancestral functions, making them crucial for inferring gene function in novel genomes. Paralogs, arising from gene duplication within a genome, often diverge in function through neofunctionalization or subfunctionalization, serving as raw material for evolutionary innovation.
- Gene duplication mechanisms, including tandem duplication, whole-genome duplication (WGD), and retrotransposition, generate paralogs with varying retention rates and functional fates; WGD-derived ohnologs are more frequently retained and subfunctionalized than tandem duplicates.
- Identifying orthologs and paralogs relies on computational methods ranging from reciprocal best hits (RBH) in BLAST searches, which is efficient but struggles with gene duplication events, to phylogenetic reconciliation, considered the gold standard for accurately distinguishing orthologs from paralogs by comparing gene trees to species trees.
- Pitfalls in ortholog/paralog inference include gene loss leading to one-to-many relationships, hidden paralogs due to extreme divergence, and database errors from automated annotation pipelines, necessitating validation through multiple methods like synteny analysis and experimental complementation.
- The distinction between orthologs and paralogs is critical for comparative genomics, influencing genome annotation accuracy, functional inference in experimental systems (e.g., using model organisms), and evolutionary rate analyses (dN/dS ratios) to detect positive or purifying selection.
Introduction to Paralog and Ortholog
Definitions and Evolutionary Context
Orthologs and paralogs are two fundamental classes of homologous genes—genes that share a common evolutionary ancestor. The distinction between them is not based on sequence similarity per se, but on the type of event that created the gene copies. This distinction is critical because it carries different implications for function, evolutionary constraint, and experimental design.
Orthologs are genes in different species that descended from a single gene in the last common ancestor of those species. They arise through speciation events. For example, the α-globin gene in humans and the α-globin gene in mice are orthologs because they both trace back to the α-globin gene present in the last common ancestor of mammals.
Paralogs are genes within the same genome (or across genomes) that arose from a gene duplication event. The duplicated copies then evolve independently. Human α-globin and β-globin are paralogs; they arose from an ancient duplication of a primordial globin gene, followed by divergence into the α- and β-globin families.
The relationship between these terms is hierarchical. All orthologs and paralogs are homologs, but the specific evolutionary event—speciation versus duplication—determines the classification. A gene can be an ortholog of a gene in another species while simultaneously being a paralog of a gene in its own genome. For instance, human α-globin is an ortholog of mouse α-globin, but it is a paralog of human β-globin.
Why Distinguishing Matters
The ortholog/paralog distinction is not merely taxonomic pedantry. It has direct consequences for how we interpret gene function and evolutionary history. Orthologs, because they descend from a single ancestral gene through speciation, are more likely to retain the ancestral function. If you annotate a novel gene in a newly sequenced genome, the most reliable functional inference comes from its ortholog in a well-studied organism, not from a paralog that may have diverged in function after duplication.
Paralogs, by contrast, are the raw material for evolutionary innovation. After duplication, one copy is freed from purifying selection and can acquire new functions (neofunctionalization) or partition the ancestral functions (subfunctionalization). Consequently, paralogs frequently exhibit divergent functions, expression patterns, or regulatory regimes. Using a paralog to infer function can be actively misleading. Distinguishing the two is therefore a prerequisite for accurate functional annotation, for interpreting evolutionary rate analyses, and for understanding how gene families expand and contract across lineages.
Evolutionary Mechanisms Generating Paralogs and Orthologs
Gene Duplication Events
Paralogs are generated by gene duplication, which can occur through several distinct mechanisms. The scale and genomic context of the duplication profoundly influence the fate of the duplicated copies.
Tandem duplication arises from unequal crossing-over during meiosis, producing two adjacent copies of a gene. This is a common mechanism for generating small gene families, such as the HOX clusters or the olfactory receptor genes in vertebrates. Tandem duplicates often remain physically linked and can undergo further rounds of duplication, creating arrays of related genes.
Whole-genome duplication (WGD)—polyploidy—doubles the entire genome in a single event. Two rounds of WGD occurred early in the vertebrate lineage (the "2R" hypothesis), and a more recent WGD occurred in the teleost fish lineage (the "3R" event). WGD creates a genome-wide set of paralogs (ohnologs, named after Susumu Ohno). These duplicates are initially redundant in function, but they are subject to the same fates as any duplicated gene. The yeast Saccharomyces cerevisiae experienced a WGD approximately 100 million years ago, and the retained paralogs from that event have been extensively studied.
Retrotransposition involves the reverse transcription of an mRNA molecule and integration of the resulting cDNA into a new genomic location. The new copy is an intronless paralog (a processed pseudogene if nonfunctional, or a functional retrogene if it acquires regulatory elements). Because the retrogene lacks the original promoter, it is often expressed in a different tissue or under different conditions. The Drosophila gene jingwei is a classic example of a functional retrogene that evolved a new function after retrotransposition.
Transposable element-mediated duplication can duplicate large genomic segments. DNA transposons can excise and insert elsewhere, sometimes carrying flanking host genes. This mechanism is less common than tandem duplication or WGD but can contribute to gene family expansion in some lineages. The relationship between transposition and gene duplication is discussed further in the context of Transposable Element biology.
Speciation and Orthology
Orthologs arise when a single gene present in an ancestral population is inherited by two descendant species after speciation. The key point is that orthology is defined by the speciation event, not by the degree of sequence similarity. Two orthologs can have very different sequences if the lineages have diverged for a long time or if one lineage has experienced accelerated evolution. Conversely, two genes with high sequence similarity in different species are not necessarily orthologs—they could be paralogs that were present in the ancestor and have been retained in both lineages.
Consider the case of a gene family with three paralogs (A, B, C) in the last common ancestor of two species, X and Y. After speciation, species X and Y each inherit A, B, and C. The A gene in X is the ortholog of the A gene in Y. But the A gene in X is also a paralog of the B gene in Y, because their relationship is defined by the duplication event that created A and B in the ancestor. This is the "ortholog-conjecture" framework: orthology is a relationship between genes in different species connected through speciation, while paralogy is a relationship between genes connected through duplication, regardless of the species in which they currently reside.
Fates of Duplicated Genes: Neofunctionalization, Subfunctionalization, Pseudogenization
After duplication, the two copies are initially redundant. This redundancy is evolutionarily unstable because mutations in one copy are not deleterious—the other copy compensates. The most common fate is pseudogenization: one copy accumulates deleterious mutations and becomes a nonfunctional pseudogene. The rate of pseudogenization is high; most duplicated genes are lost within a few million years.
Neofunctionalization occurs when one copy acquires a new, beneficial function while the other retains the ancestral function. This requires that the new function be favored by Positive Selection Pressure. The classic example is the RNAse1 gene in colobine monkeys, where a duplicated copy evolved a new digestive function in the foregut, while the ancestral copy retained its role in pancreatic digestion.
Subfunctionalization (the duplication-degeneration-complementation model) occurs when the ancestral gene had multiple functions or regulatory domains, and each duplicate loses a subset of them. The two copies together now perform the full ancestral function, but neither does so alone. This is often seen in the partitioning of expression domains. The engrailed gene in zebrafish has two paralogs, eng1a and eng1b, which together recapitulate the expression pattern of the single engrailed gene in other vertebrates.
The fate of a duplicated gene is influenced by the mechanism of duplication. WGD-derived paralogs (ohnologs) are more likely to be retained and to undergo subfunctionalization, possibly because the initial redundancy is genome-wide and the dosage balance between interacting partners is maintained. Tandem duplicates, by contrast, are more frequently lost or neofunctionalized. The interplay between duplication and the Neutral Theory of Molecular Evolution is central here: most duplicates are fixed or lost through drift, and only rarely does positive selection drive their retention.
Functional Implications of Paralogs and Orthologs
Functional Divergence of Paralogs
Paralogs are expected to diverge in function over time, and this divergence is the basis for much of the functional diversity in genomes. The divergence can manifest at multiple levels: biochemical activity, substrate specificity, expression pattern, protein-protein interactions, and subcellular localization.
The human Hox gene family provides a clear example. The 39 Hox genes in humans are organized into four clusters (A, B, C, D), each cluster containing 13 paralogous groups. These paralogs arose through tandem duplications and two rounds of WGD. They share a highly conserved DNA-binding domain (the homeodomain) but have diverged in their N-terminal regions, which mediate interactions with cofactors. This divergence allows different Hox paralogs to regulate distinct sets of target genes during embryonic development, specifying positional identity along the anterior-posterior axis.
The cytochrome P450 (CYP) superfamily is another striking example. Humans have 57 CYP genes, all derived from a common ancestor. These paralogs have diverged to metabolize a remarkable range of substrates, from endogenous steroids and fatty acids to xenobiotics including drugs and environmental toxins. The substrate specificity of each CYP paralog is determined by the structure of the active site, which has diverged substantially between family members. CYP3A4, for instance, metabolizes approximately 50% of clinically used drugs, while CYP17A1 is specific to steroid biosynthesis.
Conservation of Ortholog Function
The expectation that orthologs retain the ancestral function is the basis for the "ortholog conjecture." This expectation is generally supported by data: orthologs tend to have more similar functions than paralogs at equivalent sequence divergence. The p53 tumor suppressor gene provides an instructive example. The human TP53 gene and the mouse Trp53 gene are orthologs; both encode a transcription factor that responds to DNA damage by inducing cell cycle arrest or apoptosis. Despite ~80% amino acid identity, the functional conservation is near-complete—human p53 can functionally replace mouse p53 in transgenic mice.
Similarly, the ATP7A gene, encoding a copper-transporting P-type ATPase, is functionally conserved between humans and zebrafish. Mutations in the human gene cause Menkes disease, and the zebrafish ortholog shows the same copper transport activity. This functional conservation is the reason why model organisms are useful for studying human disease genes: the ortholog in the model organism is likely to perform the same biochemical function.
Exceptions and Complexities
The ortholog conjecture is a heuristic, not a law. There are well-documented exceptions where orthologs have diverged in function. This can occur when the ancestral gene was multifunctional and different lineages have lost different subsets of functions. The p53 family itself illustrates this: the ancestral gene likely had functions related to both stress response and development, and the three human paralogs (TP53, TP63, TP73) have partitioned these functions. The TP63 and TP73 orthologs in different species show more functional variability than TP53.
Another complication is that a gene in one species can have multiple orthologs in another species due to lineage-specific duplications. If a gene duplicated in the mouse lineage after the human-mouse split, then the single human gene has two mouse orthologs (co-orthologs). These co-orthologs may have subfunctionalized, such that neither mouse gene alone performs the full ancestral function. Inferring function from a single co-ortholog would then be incomplete.
Finally, some genes have no true ortholog in other species. This can occur if the gene arose de novo from non-coding sequence, or if it has been lost in all other sequenced lineages. In such cases, functional inference must rely on paralogs or on experimental characterization.
Computational Methods for Identifying Paralogs and Orthologs
Sequence Similarity Methods
The simplest approach to identifying orthologs and paralogs is pairwise sequence comparison. The Basic Local Alignment Search Tool (BLAST) is the workhorse. The logic is that orthologs are more similar to each other than to any other gene in the respective genomes.
Reciprocal Best Hits (RBH) is the most widely used heuristic. For two genomes, A and B, a gene a in A and gene b in B are considered orthologs if a's best hit in B is b, and b's best hit in A is a. This method is simple, fast, and works well for one-to-one orthologs in genomes without extensive duplication. However, it fails when there are lineage-specific duplications: if gene a in A has two co-orthologs in B (b1 and b2), the RBH method will typically identify only one of them, depending on which has the higher score.
Reciprocal Best Hits with coverage filtering is an improvement. The alignment must cover a minimum fraction of both query and subject sequences (typically 50–70%) to avoid spurious hits from partial alignments or domain-only matches. This is particularly important for multi-domain proteins, where a shared domain can produce a high-scoring but misleading alignment.
Phylogenetic Approaches
Phylogenetic methods are the gold standard for distinguishing orthologs from paralogs. The logic is to reconstruct the gene tree and compare it to the species tree. A gene tree that matches the species tree indicates orthology; a gene tree that contains duplications indicates paralogy.
The reconciliation approach formalizes this. Given a gene tree and a species tree, reconciliation identifies the duplication and speciation nodes in the gene tree by mapping it onto the species tree. A node in the gene tree that corresponds to a speciation event in the species tree is a speciation node; a node that does not correspond to any speciation event is a duplication node. The descendants of a duplication node are paralogs; the descendants of a speciation node are orthologs.
The software Notung implements reconciliation and can also rearrange the gene tree to minimize the number of duplications and losses. TreeFam and PhylomeDB are databases that provide pre-computed gene trees and orthology assignments based on phylogenetic analysis. The Molecular Phylogenetics and Evolution approach is essential here, as the quality of the orthology inference depends on the quality of the gene tree and the accuracy of the species tree.
Clustering and Reconciliation Methods
For genome-wide analyses, clustering methods group genes into orthologous groups. These methods are more scalable than per-gene phylogenetic analysis.
OrthoMCL uses an all-against-all BLAST search, then applies a Markov clustering algorithm to group genes into orthologous groups. The clustering is based on the similarity scores, and the algorithm naturally separates paralogs into different clusters when their similarity is lower than the ortholog similarity. OrthoMCL groups contain both orthologs and in-paralogs (paralogs that arose after the speciation event separating the species in the group).
InParanoid focuses on pairwise orthology between two species. It identifies the best reciprocal hits as the "seed" ortholog pair, then adds in-paralogs—genes that are more similar to one of the seed orthologs than to any gene in the other species. This handles the co-ortholog problem better than simple RBH.
eggNOG (evolutionary genealogy of genes: Non-supervised Orthologous Groups) uses a clustering approach based on precomputed all-against-all similarities, followed by phylogenetic analysis of each cluster. It provides orthology assignments across thousands of genomes and includes functional annotations. The OMA (Orthologous MAtrix) database uses a different approach based on evolutionary distance and graph-based clustering, which is particularly robust for large-scale analyses.
A comparison of these methods is summarized in Table 1.
| Method | Approach | Handles Co-orthologs | Handles Multi-domain Proteins | Scalability | Output |
|---|---|---|---|---|---|
| Reciprocal Best Hits | Pairwise BLAST | Poor | Poor | Excellent | Pairwise orthologs |
| OrthoMCL | Markov clustering | Good | Moderate | Good | Orthologous groups |
| InParanoid | RBH + in-paralogs | Good | Moderate | Good | Pairwise orthologs + in-paralogs |
| eggNOG | Clustering + phylogeny | Good | Good | Excellent | Orthologous groups |
| Notung | Tree reconciliation | Good | Good | Moderate | Gene trees with duplication/speciation nodes |
Orthology and Paralogy in Comparative Genomics
Genome Annotation
Ortholog/paralog assignments are a cornerstone of genome annotation. When a new genome is sequenced, the first step in annotating protein-coding genes is typically to identify orthologs in well-annotated reference genomes. The functional annotation of the ortholog is then transferred to the new gene. This is the basis of the "ortholog-based annotation" used by pipelines such as Ensembl and NCBI.
The accuracy of this approach depends on the quality of the orthology assignment. If a gene is misidentified as an ortholog when it is actually a paralog, the functional annotation will be wrong. This is a particular risk for large gene families, where the boundaries between orthologs and paralogs can be blurred by domain shuffling and lineage-specific expansions.
Functional Inference
Beyond annotation, orthology is used to infer function in experimental systems. If a gene in a model organism has a known function, its ortholog in a non-model organism is presumed to have a similar function. This is the logic behind using yeast, worm, fly, and mouse as model systems for human biology. The BRCA1 gene provides an example: the human gene is involved in DNA double-strand break repair, and its mouse ortholog shows the same function. Knockout of the mouse ortholog recapitulates many features of human BRCA1-associated cancer predisposition.
However, functional inference from orthology has limits. The function of a gene is context-dependent—it depends on the cellular environment, the interaction partners, and the regulatory network in which it operates. Two orthologs may have identical biochemical activity but different biological roles because they are expressed in different tissues or regulated by different signals. The Conserved Sequence of the coding region is only part of the story; the regulatory regions, which are often less conserved, determine where and when the gene is expressed.
Evolutionary Rate Analyses
Orthologs and paralogs are used in evolutionary rate analyses to study the forces shaping gene evolution. The ratio of nonsynonymous to synonymous substitution rates (dN/dS, or ω) is a standard measure of selection. Orthologous sequences are used to estimate ω for a gene across species, providing evidence for Positive and Negative Selection. A ω significantly greater than 1 indicates positive selection; a ω less than 1 indicates purifying selection.
Paralogs are used to study the evolution of gene families. By comparing the ω values of paralogs that arose at different times, one can infer the strength and direction of selection after duplication. For example, the opsin gene family in vertebrates has undergone multiple duplications, and the paralogs have diverged in spectral sensitivity. The ω values for these paralogs show evidence of positive selection in the regions encoding the chromophore-binding pocket, consistent with adaptation to different light environments.
The Concept of Neutral Evolution is fundamental here: most sequence differences between orthologs are neutral, and the rate of neutral evolution is approximately constant (the molecular clock). Deviations from the clock—accelerated or decelerated evolution—indicate changes in selective pressure. Comparing the evolutionary rates of orthologs and paralogs can reveal whether a duplication was followed by neofunctionalization (accelerated evolution in one copy) or subfunctionalization (relaxed selection in both copies).
Challenges and Common Pitfalls in Ortholog/Paralog Inference
Gene Loss and One-to-Many Relationships
Gene loss is a major complication. If a gene is lost in one lineage, the orthology relationships become asymmetric. Consider three species: A, B, and C. If the gene is present in A and B but lost in C, then the A-B ortholog pair is straightforward. But if the gene duplicated in the A lineage after the A-B split, then A has two paralogs (A1 and A2) and B has one gene (B). Both A1 and A2 are co-orthologs of B. Simple RBH methods will identify only one of them, and which one is identified may depend on subtle differences in sequence evolution.
This is not a rare edge case. Gene loss is pervasive in eukaryotic genomes. The olfactory receptor gene family in humans has over 400 functional genes and hundreds of pseudogenes, while mice have over 1000 functional genes. Many of these are one-to-many or many-to-many relationships across species.
Hidden Paralogs
Hidden paralogs arise when a duplication is ancient and the paralogs have diverged so much that their similarity is no longer detectable by sequence search. This is particularly problematic for rapidly evolving genes or for genes that have undergone domain shuffling. A gene may share a domain with a paralog but have no detectable similarity outside that domain. BLAST searches may identify the domain-containing region but miss the rest of the gene, leading to incorrect orthology assignments.
The immunoglobulin superfamily is a prime example. These proteins share an immunoglobulin domain but have highly divergent sequences outside the domain. A BLAST search using the full-length protein may only detect the domain, and the resulting "hits" may include many non-orthologous proteins that happen to share the domain. Phylogenetic methods that use full-length alignments are more robust, but they require accurate alignment of divergent sequences, which is itself challenging.
Misannotation and Database Errors
Database errors are a pervasive problem. Many genes in public databases are annotated based on automated pipelines that rely on the very orthology assignments that may be incorrect. This creates a circularity problem: incorrect annotations propagate through databases, and downstream analyses based on these annotations inherit the errors.
A common error is the misannotation of paralogs as orthologs. For example, the human gene BRCA1 and the mouse gene Brca1 are orthologs, but there is a related gene, BARD1, that is a paralog of BRCA1. If a database incorrectly labels BARD1 as the ortholog of BRCA1, any functional inference will be wrong. This type of error is more common for large gene families and for genes with complex evolutionary histories.
Another issue is the use of inconsistent nomenclature. The same gene may have different names in different databases, or the same name may refer to different genes in different species. This is particularly problematic for gene families that have expanded independently in different lineages.
Practical Guidelines for Studying Paralogs and Orthologs
Step-by-Step Workflow
For a researcher who has a gene of interest and wants to identify its orthologs and paralogs, the following workflow is recommended:
- Obtain the protein sequence of your gene of interest. Use the longest isoform, as shorter isoforms may lack conserved domains.
- Perform a BLAST search against the genome of interest. Use BLASTP for protein sequences and TBLASTN for searching against genomic DNA. Set an E-value threshold of 1e-5 and require a minimum coverage of 50% of both query and subject.
- Identify candidate orthologs using reciprocal best hits. For each candidate hit in the target genome, perform a reverse BLAST against the query genome. If the original query is the best hit, the candidate is a likely ortholog.
- Expand to include co-orthologs. If the target genome has multiple genes that are all more similar to your query than to any other gene in the query genome, these are candidate co-orthologs.
- Build a phylogenetic tree. Align the candidate sequences using a multiple sequence alignment program (e.g., MAFFT, MUSCLE). Build a tree using maximum likelihood (e.g., RAxML, IQ-TREE) or Bayesian inference (e.g., MrBayes). Include a known outgroup sequence to root the tree.
- Reconcile the gene tree with the species tree. Use a program like Notung or compare the gene tree topology to the known species tree. Identify duplication and speciation nodes.
- Validate with a database. Cross-check your assignments with a curated database such as eggNOG, OMA, or TreeFam. Discrepancies should be investigated, not ignored.
- Confirm with synteny. If the genes are in conserved genomic neighborhoods (syntenic blocks), this provides independent evidence for orthology. Synteny is particularly useful for distinguishing orthologs from paralogs in recently diverged species.
Choosing the Right Tool
The choice of tool depends on the question. For a single gene in two species, RBH with manual inspection is often sufficient. For a genome-wide analysis, a clustering method like OrthoMCL or eggNOG is appropriate. For a detailed evolutionary analysis of a gene family, phylogenetic reconciliation is necessary.
Consider the trade-offs: RBH is fast but fails for complex scenarios. OrthoMCL is robust but requires careful parameter tuning. eggNOG is comprehensive but may not include your species of interest. Phylogenetic methods are the most accurate but are computationally intensive and require expertise.
Validating Results
Orthology assignments should always be validated. The most reliable validation is experimental: does the putative ortholog complement a mutant phenotype in the model organism? This is rarely feasible for large-scale analyses, but it is the gold standard for individual genes.
For computational validation, check the following:
- Sequence identity: Orthologs typically have higher sequence identity than paralogs at equivalent evolutionary distances. However, this is not absolute—some orthologs are highly divergent.
- Domain architecture: Orthologs should have the same domain architecture. Differences in domain composition suggest either a misannotation or a true functional divergence.
- Synteny: Conserved gene order is strong evidence for orthology.
- Expression data: If expression data are available, orthologs often show similar expression patterns, though this is not always the case.
Summary and Key Takeaways
The distinction between orthologs and paralogs is fundamental to molecular evolution and comparative genomics. Orthologs arise through speciation and are expected to retain ancestral functions; paralogs arise through duplication and are expected to diverge. This distinction underpins functional annotation, evolutionary rate analyses, and our understanding of how gene families evolve.
The key points to remember are:
- Orthology is defined by speciation; paralogy is defined by duplication. The same gene can be both an ortholog of a gene in another species and a paralog of a gene in its own genome.
- Gene duplication generates paralogs through multiple mechanisms, including tandem duplication, whole-genome duplication, and retrotransposition. The fate of duplicated genes—pseudogenization, neofunctionalization, or subfunctionalization—determines their functional divergence.
- Orthologs generally retain ancestral functions, but exceptions exist. Co-orthologs, lineage-specific duplications, and functional divergence can complicate functional inference.
- Computational methods for identifying orthologs and paralogs range from simple reciprocal best hits to sophisticated phylogenetic reconciliation. Each method has strengths and weaknesses, and the choice of method should match the question.
- Gene loss, hidden paralogs, and database errors are common pitfalls. Validation using multiple independent methods is essential.
- The ortholog/paralog distinction is not just a matter of terminology—it has practical consequences for experimental design, functional annotation, and evolutionary interpretation.
Frequently Asked Questions
What is the difference between orthologs and paralogs?
Orthologs are homologous genes in different species that arose from a single ancestral gene through a speciation event. Paralogs are homologous genes that arose through a gene duplication event. The key distinction is the type of event that created the gene copies: speciation for orthologs, duplication for paralogs.
Are orthologs always functionally equivalent?
No. The "ortholog conjecture" states that orthologs are more likely to retain the ancestral function than paralogs, and this is generally supported by data. However, there are exceptions. Orthologs can diverge in function if the ancestral gene was multifunctional and different lineages lost different functions, or if one lineage experienced positive selection for a new function.
How can I identify orthologs and paralogs in my genome of interest?
Start with reciprocal best hits using BLAST, then expand to include co-orthologs. Build a phylogenetic tree of the candidate sequences and reconcile it with the species tree to identify duplication and speciation nodes. Validate your assignments using a database like eggNOG or OMA and check for synteny.
Why do some genes have multiple paralogs but no orthologs?
This can occur if the gene family expanded independently in a lineage after it diverged from other species. If the ancestral gene was lost in other lineages, or if the paralogs arose through lineage-specific duplications, there may be no ortholog in other species. Alternatively, the gene may have arisen de novo from non-coding sequence.
What is the significance of paralogs in evolution?
Paralogs are the raw material for evolutionary innovation. After duplication, one copy can acquire a new function (neofunctionalization) or the copies can partition the ancestral functions (subfunctionalization). This process has generated much of the functional diversity in genomes, including the expansion of gene families involved in immunity, metabolism, and development.
Can a gene be both an ortholog and a paralog?
Yes. A gene can be an ortholog of a gene in another species while being a paralog of a gene in its own genome. For example, human α-globin is an ortholog of mouse α-globin, but it is a paralog of human β-globin. The classification depends on the specific pair of genes being compared.
What tools are commonly used for ortholog/paralog detection?
Common tools include BLAST for reciprocal best hits, OrthoMCL and InParanoid for clustering-based approaches, eggNOG and OMA for database-level assignments, and Notung for phylogenetic reconciliation. The choice of tool depends on the scale of the analysis and the complexity of the gene family.
Further Reading
- Kakularam KR et al. Paralog- and ortholog-specificity of inhibitors of human and mouse lipoxygenase-isoforms. Biomedicine & pharmacotherapy = Biomedecine & pharmacotherapie. 2022. PubMed 34801853
- Babini E et al. Solution structure of human beta-parvalbumin and structural comparison with its paralog alpha-parvalbumin and with their rat orthologs. Biochemistry. 2004. PubMed 15610002
- Koonin EV. Orthologs, paralogs, and evolutionary genomics. Annual review of genetics. 2005. PubMed 16285863
- Gao K, Miller J. Human-chimpanzee alignment: ortholog exponentials and paralog power laws. Computational biology and chemistry. 2014. PubMed 25443749
- Canard C, Chialvo P, Scott Chialvo C. Gene model for the ortholog of JhI-26 and a paralog in Drosophila dunni. microPublication biology. 2026. PubMed 42502753
- Chen T et al. L1.2, the zebrafish paralog of L1.1 and ortholog of the mammalian cell adhesion molecule L1 contributes to spinal cord regeneration in adult zebrafish. Restorative neurology and neuroscience. 2016. PubMed 26889968