How to Read a Phylogenetic Tree (and How It Differs From a Cladogram)
By Dr. Zubair Khalid, DVM, MS, PhD ·

A phylogenetic tree is a diagram that shows the lines of evolutionary descent connecting species, populations, genes, or other taxa to a common ancestor [3]. Every branching point on the tree represents an inferred speciation event, and the pattern of those branches, not the spacing of the labels, is what encodes relatedness.
You will meet these diagrams constantly: in a paper's Figure 1, in a lecture slide on viral variants, in a lab meeting where someone asks whether two isolates are "close." Reading them correctly changes what you can claim. A tree can tell you that two taxa share a recent common ancestor. It cannot tell you, by itself, which one is "more evolved," which one is ancestral, or how different they look.
Quick Answer
- A phylogenetic tree depicts evolutionary descent from a common ancestor; the root is the ancestral lineage, nodes are branching points, and tips are the labeled taxa [3].
- Two lineages that stem from the same branch point are sister taxa. They share an ancestor, but neither gave rise to the other [4].
- Relatedness is read from the most recent common ancestor (MRCA): trace each lineage back toward the root until they meet. Taxa sharing a younger MRCA are more closely related [1].
- The left-to-right order of tips carries no meaning. Rotating branches at a node changes the drawing, not the history [1][3].
- A cladogram shows only branching pattern, with edge lengths that mean nothing. A phylogram has edge lengths that represent time or genetic distance [5].
- On an unrooted tree you cannot say "A is more closely related to B than to C," because that statement depends on where the root falls [7].
Tree Anatomy and the Vocabulary You Need
Start with the parts. The root is the ancestral lineage from which everything else on the tree descends [3]. Nodes are the branching points, and each one corresponds to the last common ancestor of the lineages below it [3]. Internal branches connect two nodes; external branches connect a tip to a node [3]. Tips, also called terminals, carry the taxon labels, and those labels can be individual organisms, populations, species, or larger groups. Some trees show genes instead [3].
A clade, or monophyletic group, includes an ancestral lineage and all of its descendants [3]. There is a clean test for this: a clade can be cut away from the rest of the tree with a single cut, while a non-monophyletic set of taxa needs two or more cuts [3]. If you have to snip twice, you have left something out or pulled something in that does not belong.
Two more terms come up in every paper. Sister taxa are two lineages stemming from the same branch point [4]. A basal taxon is a lineage that evolved early from the root and remains unbranched [4]. A branch point with more than two lineages is a polytomy, which shows unresolved relationships [4].
The MRCA is the anchor for all relatedness claims. Any set of taxa has one: it is the youngest ancestor they all share, found by tracing each lineage back toward the root until the lineages meet [1]. Taxa that share a more recent common ancestor are more closely related than taxa whose MRCA is older [1]. That single rule resolves most arguments about "which is closer to which."
Rooted vs Unrooted Trees
A rooted tree has a single ancestral lineage at its base to which all organisms on the tree relate [4]. The root is the most recent common ancestor of every taxon in the tree, and it gives the direction of evolution [7]. Most tree-building methods do not estimate the root position, so rooting is a separate step you have to justify [7].
An unrooted tree does not show a common ancestor, but it does show relationships among species [4]. It represents all the rooted trees consistent with it, and the root would typically lie on one of its edges [5]. This is why the unrooted tree is weaker evidence than it looks. On an unrooted tree you cannot say "A is more closely related to B than to C," because that would be false if the root lay on the branches connecting A and B [7].
Two standard rooting methods are worth knowing. Outgroup rooting places the root where one or more taxa known to be more distantly related join the tree [7]. Midpoint rooting assumes equal evolutionary rates and places the root midway between the two longest branches [7]. Outgroup rooting is the method used in the worked example below.
What Branch Lengths Mean
Unless indicated otherwise, a tree depicts only branching history, and branch lengths are drawn however makes the tree tidy. In that case they carry no information [3]. When branch lengths are meaningful, the tree is often called a phylogram, and the lengths depict either the amount of evolution in a gene sequence or the estimated duration of branches [3].
In a phylogram built from sequence data, branch lengths indicate genetic change. Longer branches mean more divergence, usually measured as the average number of nucleotide or amino acid substitutions per site, and a scale bar is commonly shown [6]. The units matter: a branch length of 0.10 means 0.10 substitutions per site, not 10 substitutions and not 10 percent similarity.
A special case is the ultrametric tree, sometimes called a chronogram, where edge lengths represent time. In that case present-day taxa are equidistant from the root [5]. If you measure root-to-tip distance and it varies across tips, you are not looking at an ultrametric tree.
A branch length in substitutions per site mixes evolutionary rate and elapsed time, so a short branch is ambiguous. You cannot read rate off a branch length alone.
Cladogram vs Phylogram vs Chronogram
These three words get used loosely, so here is the practical distinction.
| Tree type | What edge lengths represent | What you can claim |
|---|---|---|
| Cladogram | Nothing; lengths are arbitrary [5] | Branching pattern and clade membership only |
| Phylogram | Time or genetic distance [5] | Pattern plus relative divergence or elapsed time |
| Ultrametric tree (chronogram) | Time, with present-day taxa equidistant from the root [5] | Pattern plus a time scale, if the calibration is sound |
A tree topology is the graph plus the leaf labels. It represents relationships but not time or genetic distance, and left/right orientation does not affect it [5]. Two trees have the same topology if you can turn one into the other by twisting, rotating, or bending branches without cutting and reattaching them. Nodes act like swiveling joints [3].
Counting helps you appreciate how much a topology claims. A rooted binary tree with $n$ leaves has $n - 1$ internal nodes, and an unrooted binary tree has $n - 2$ [5]. The number of possible topologies is
$$ (2n - 5)!! \quad \text{unrooted}, \qquad (2n - 3)!! \quad \text{rooted} $$
for $n > 2$, where $n$ is the number of leaves and $!!$ is the double factorial, the product of every other integer down to 1 or 2 [5]. For $n = 6$ that gives $7 \times 5 \times 3 \times 1 = 105$ unrooted and $9 \times 7 \times 5 \times 3 \times 1 = 945$ rooted binary topologies. With that many possible topologies, support values tell you how firmly the data favor each clade.
How to Read a Tree, Step by Step
- Identify the root and the direction of evolution. If the tree is unrooted, stop making directional claims [7].
- Check whether branch lengths are meaningful. Look for a scale bar or a note in the caption [6].
- For any pair of taxa, trace both lineages back to their first shared node. That node is the MRCA [1].
- Compare MRCAs across pairs. The pair with the younger MRCA is more closely related [1].
- Check support values on the internal branches that define the clades you care about [8].
- Ignore tip order. Rotate the tree mentally and confirm your reading survives the rotation [1][3].
Support values deserve their own note. Confidence values such as bootstrap percentages refer to the internal branch on which they are shown [8]. The phylogenetic bootstrap resamples characters, meaning alignment columns, with replacement while keeping all species, rebuilds the tree for each sample, and summarizes the results with a majority-rule consensus tree [10]. Felsenstein stated that a group appearing in 95% or more of bootstrap samples could be taken as statistically significant [10]. Experts generally accept branches with more than 80% or 90% bootstrap support, provided an appropriate evolutionary model was used [8].
Worked Example
The tree below is hypothetical and chosen for teaching. Numbers after a colon are branch lengths in substitutions per site. Numbers after a closing parenthesis are bootstrap support for that clade. Out is the outgroup used to root the tree.
(Out:0.50,((A:0.10,B:0.30)95:0.05,(C:0.15,(D:0.05,E:0.07)99:0.10)72:0.04)100:0.08);
You can reproduce every number with Biopython 1.88. Note that Bio.Phylo reads Newick with rooted = False by default, so the root here is the basal node joining Out and the ingroup.
from io import StringIO
from Bio import Phylo
nwk = "(Out:0.50,((A:0.10,B:0.30)95:0.05,(C:0.15,(D:0.05,E:0.07)99:0.10)72:0.04)100:0.08);"
t = Phylo.read(StringIO(nwk), "newick")
t.common_ancestor("A", "B") # clade A+B, support 95
t.distance("A", "B") # 0.40
t.is_monophyletic([t.find_any(n) for n in ["A", "B"]]) # returns the A+B clade (False if not monophyletic)
Phylo.draw_ascii(t)
MRCA results. A and B meet at the clade A+B, support 95. D and E meet at clade D+E, support 99. C and E meet at clade C+D+E, support 72. B and C meet at the ingroup clade A+B+C+D+E, support 100. A and E also meet at the ingroup, support 100.
Sister pairs. A and B are sisters. D and E are sisters. The clade (D,E) is sister to C. The clade (A,B) is sister to the clade (C,D,E).
Patristic distances, meaning the sum of branch lengths along the path between two tips: A to B is $0.10 + 0.30 = 0.40$. A to C is $0.10 + 0.05 + 0.04 + 0.15 = 0.34$. D to E is 0.12. C to E is 0.32. A to E is 0.36. B to C is 0.54. B to D is 0.54.
Here is the lesson that trips people up. A is genetically closer to C (0.34) than to its own sister B (0.40), because B sits on a long branch. Relatedness is still read from the MRCA: A and B share the youngest common ancestor on the tree. Relatedness and genetic distance answer different questions. One asks about shared ancestry, the other about accumulated change.
Root-to-tip distances. A is 0.23, B is 0.43, C is 0.27, D is 0.27, E is 0.29, and Out is 0.50. These are unequal, so this is a phylogram, not an ultrametric chronogram.
Monophyly checks. [A,B] is True. [C,D,E] is True. [B,C] is False. [A,B,C] is False.
Rotation. Running t.ladderize(reverse=True) gives the tip order D, E, C, A, B, Out. The B to C distance is still 0.54 and the MRCA is unchanged. Nothing about the biology moved.
Support. The clade C+D+E has 72% bootstrap, below the 80% guide from the EMBL-EBI course and below the 95% Felsenstein used, so treat that exact grouping cautiously. A+B at 95 and D+E at 99 are well supported.
Common Mistakes
- Reading adjacent tips as close relatives. Tip proximity is a drawing artifact. Relatedness comes from the branching pattern and the MRCA, not from who sits next to whom [2].
- Counting nodes between two taxa to measure relatedness. The number of nodes depends on which other taxa are included on the tree, so it is not a stable measure [2].
- Assuming overall similarity shows relatedness. Closely related lineages can look different, and distantly related ones can look alike through convergence. Snakes and earthworms are the classic illustration [2].
- Treating some tips as ancestors of others. Present-day tips are cousins, not ancestors. A tip on the left of an upright tree is not older than one on the right [2].
- Making directional claims from an unrooted tree. Without a root, "A is closer to B than to C" may be false [7].
- Reading branch lengths on a cladogram. If the caption does not say lengths are meaningful, they are not [3][5].
- Treating a bootstrap value as a probability that the clade is real. It refers to the internal branch it labels, and it depends on the model and the data [8].
Limitations
Bootstrap thresholds are a convention, not a law. Felsenstein used 95% [10], UC Berkeley's Tree Room says scientists look for values above 95% to feel very confident, and the EMBL-EBI course says more than 80% or 90% is generally accepted [8]. Values near 70% are also widely quoted in the literature. Pick a threshold, state it, and apply it consistently.
Terminology varies across authors. Some use "phylogram" only for branch lengths in substitutions and reserve "chronogram" or "timetree" for time-scaled trees, while Larget's outline allows phylogram edges to mean time or genetic distance [5]. Read the caption before you assume units.
A tree is a hypothesis, not a measurement. The topology depends on the alignment (check it before tree building, for example in OmniAlign), the substitution model, the rooting choice, and the taxa you included. Adding or removing a single taxon can change which clade looks well supported. The worked tree here is invented, and its numbers illustrate reading rules and do not describe real organisms.
Frequently Asked Questions
What is the difference between a phylogenetic tree and a cladogram?
A cladogram represents only a branching pattern, and its edge lengths represent nothing [5]. A phylogenetic tree in the broad sense includes cladograms, but when branch lengths carry information about time or genetic distance the diagram is often called a phylogram [3][5]. If you need to know which one you are looking at, check the caption and look for a scale bar.
How do I find sister taxa on a phylogenetic tree?
Sister taxa are two lineages that stem from the same branch point [4]. In the worked example, A and B are sisters, and D and E are sisters. They share an ancestor, but neither taxon gave rise to the other [4].
What does branch length tell me?
In a phylogram, branch length indicates genetic change, usually the average number of nucleotide or amino acid substitutions per site, and longer branches mean more divergence [6]. In an ultrametric tree, branch lengths represent time, so present-day taxa sit equidistant from the root [5]. In a cladogram, branch length means nothing [5].
Can I say A is more closely related to B than to C on an unrooted tree?
No. On an unrooted tree you cannot make that claim, because it would be false if the root lay on the branches connecting A and B [7]. Root the tree first, using an outgroup or midpoint rooting, then read the MRCAs [7].
What bootstrap value should I trust?
Felsenstein treated 95% or higher as statistically significant [10], and the EMBL-EBI course says more than 80% or 90% is generally accepted when an appropriate evolutionary model was used [8]. In the worked example, A+B at 95 and D+E at 99 are well supported, while C+D+E at 72 deserves caution.
References
- UC Berkeley Understanding Evolution, The Tree Room: Understanding evolutionary relationships
- UC Berkeley Understanding Evolution, The Tree Room: Misinterpretations about relatedness
- Baum D. Reading a Phylogenetic Tree: The Meaning of Monophyletic Groups. Nature Education 2008;1(1):190 (Scitable)
- OpenStax Biology 2e 20.1: Organizing Life on Earth
- Larget B. Lecture Outline: Trees (Genetics 629, University of Wisconsin-Madison)
- EMBL-EBI Training, Introduction to Phylogenetics: Branches
- EMBL-EBI Training, Introduction to Phylogenetics: Root
- EMBL-EBI Training, Introduction to Phylogenetics: Confidence
- Baum DA, Smith SD, Donovan SS. The tree-thinking challenge. Science. 2005;310(5750):979-980
- Felsenstein J. Confidence limits on phylogenies: an approach using the bootstrap. Evolution. 1985;39(4):783-791
Related Articles
- Phylogenetic Tree Workflow: From Aligned Sequences to a Defensible Figure
- Phylogenetic Tree of Life: Construction, Interpretation, and Pitfalls
- Multiple Sequence Alignment: Common Pitfalls and Quality Checks
- Building a Phylogenetic Tree from DNA Sequences
- How to Make a Phylogenetic Tree in MEGA from DNA or Protein Sequences
- How to Do a Multiple Sequence Alignment with Clustal Omega (and When to Use MAFFT or MUSCLE)