How to Make a Phylogenetic Tree in MEGA from DNA or Protein Sequences

By Dr. Zubair Khalid, DVM, MS, PhD ·

How to Make a Phylogenetic Tree in MEGA from DNA or Protein Sequences

MEGA is a free, standalone program for molecular evolutionary analysis. It handles the full path from raw sequences to a finished tree: alignment, substitution model selection, tree inference, and bootstrap testing. For a teaching lab or a small gene family, that single package removes the need to move files between four different tools.

By the end of this guide you will be able to take a FASTA file of DNA or protein sequences, align it, pick a substitution model with a statistical criterion, build a maximum likelihood or neighbor-joining tree, run a bootstrap test, root the tree, and export it as a Newick file or a PDF figure. The walkthrough uses MEGA 12.1, which was the current release as of October 2026; check the download page for the newest version before you start [1].

Quick Answer

  • Install MEGA 12.1 (GUI or command-line) from the MEGA home page, then open Align | Edit/Build Alignment, choose Create New Alignment, and pick DNA or protein [1][2].
  • Load your FASTA file with Data | Open | Retrieve sequences from File, select all sequences, and run Alignment | Align by ClustalW or Align by MUSCLE [2][3].
  • Save the session as a .mas file, then send the alignment to MEGA with Data | Phylogenetic Analysis [2][4].
  • Run Models | Find Best DNA/Protein Models (ML) and note the model with the lowest BIC score [5].
  • Build the tree with Phylogeny | Construct/Test Maximum Likelihood Tree, choosing your model and either Felsenstein's Bootstrap test or the Adaptive approach [6][7].
  • Root the tree on an outgroup in Tree Explorer, then export with File | Export Current Tree (Newick) or Image | Save as PDF File [6][8].

Step 1: Get Your Sequences into the Alignment Explorer

In this workflow, sequences first pass through the Alignment Explorer, which is where you inspect and align them. From the launch bar, choose Align | Edit/Build Alignment, select Create New Alignment, click Ok, and choose DNA or protein depending on your data [2]. The data type matters: it determines which substitution models MEGA will offer later, and a nucleotide alignment analyzed as protein will produce nonsense.

Load sequences with Data | Open | Retrieve sequences from File and point it at your FASTA file [2]. If your sequences are still in GenBank, you can pull them without leaving the program: use Web | Query Genbank, set Display Settings to FASTA (Text), and press the red plus sign to add each record to the alignment [2].

Check the sequence names before you go further. MEGA uses them as taxon labels, and a tree full of accession numbers is hard to read in a figure. Rename anything ambiguous now, while the alignment is still open.

Step 2: Align the Sequences

An unaligned set of sequences cannot be compared position by position, and every downstream step assumes homology column by column. Select everything with Edit | Select All, then run Alignment | Align by ClustalW and click Ok to accept the defaults [2]. MUSCLE is the other built-in option: with all sequences selected, click the flexing-arm toolbar icon, choose Align DNA, and click Compute [2].

For protein-coding DNA, the codon-aware options are usually better. Align by ClustalW (Codons) and Align by MUSCLE (Codons) translate the coding sequences to amino acids, align the proteins, then restore the original codons [3]. This keeps gaps in frame, which matters if you later want to analyze codons or check for frameshifts. You can also align through the translated view: open the Translated Protein Sequences tab, run Alignment | Align by ClustalW, then switch back to DNA Sequences [3].

Two housekeeping items before you leave. First, Delete Gap-Only Sites removes columns where every sequence has a gap; these columns carry no information and only slow the analysis [3]. Second, save your work with Data | Save Session to a .mas file [2]. When you are done, Data | Exit Aln Explorer closes the window [2].

If you need a second opinion on alignment quality, an external aligner such as Clustal Omega or MAFFT can be run outside MEGA and the result imported as FASTA. Our OmniAlign tool is another quick option for a small set of sequences.

Step 3: Choose a Substitution Model

The substitution model describes how likely each type of change is along a branch. Pick a model that fits the data poorly and the tree topology can be wrong, not merely imprecise. MEGA automates the choice: Models | Find Best DNA/Protein Models (ML) computes BIC, AICc, and log likelihood (lnL) for each candidate model, along with parameter estimates and the number of parameters [5].

MEGA treats the model with the lowest BIC score as optimal, and its guidance is to prefer a model with few parameters that still fits well, since extra parameters add variance [5]. The search covers up to 24 nucleotide models or 64 amino acid models. The filtered option evaluates derivative models (+G, +I, +F) only for promising primary models, using a default threshold of 5 [5]. That filter is not a shortcut with a hidden cost: the MEGA12 paper reports computation time reductions of as much as 70% when a complex model fits best, with 100% concordance against exhaustive testing on 240 simulated datasets [10].

Write down the winning model name. You will select it in the tree dialog in the next step.

Step 4: Build a Maximum Likelihood Tree

Open Phylogeny | Construct/Test Maximum Likelihood Tree. The Analysis Preferences dialog opens, where Substitution Type is set to Nucleotide or Amino Acid and you select the model you just identified [6]. Click Compute to start the search.

Several settings in that dialog deserve a look [7]:

  • Phylogeny Test: Felsenstein's Bootstrap test, or the Adaptive approach. The Adaptive method starts with 25 resampled alignments and adds replicates one at a time until every bootstrap value has a standard error at or below the threshold you set. The MEGA12 paper reports average speedups of 81% with this scheme, against the conventional 500 to 2,000 replicates [10].
  • Rates among Sites: enabled only when the model allows rate variation. Choosing gamma rates reveals a gamma category number field.
  • Gaps/Missing Data Treatment: Complete-deletion, Pairwise-deletion, or Partial-Deletion with a Site Coverage Cutoff percentage. Complete-deletion discards any column with a gap in any sequence, which can throw away most of a gappy alignment.
  • ML Heuristic Method: Nearest-Neighbor-Interchange or Subtree-Pruning-Regrafting. The second is a more thorough search and takes longer.
  • Number of Threads: set this to match your machine. MEGA 12 parallelizes ML inference at a fine grain [10].

If you only need a fast tree to check for obvious problems, Phylogeny | Construct/Test Neighbor-Joining Tree is much quicker. Choose a Model/Method such as p-distance and click Compute; the tree opens in Tree Explorer [8]. Neighbor-joining is not a substitute for ML when you are publishing a topology, but it is a reasonable first pass.

Step 5: Read and Root the Tree in Tree Explorer

The tree opens in Tree Explorer. Turn on branch lengths with View | Options, Branch tab, Display Branch Length, and switch among Traditional, Radiation, and Circle layouts with View | Tree/Branch Style [8].

An ML or NJ tree is unrooted by default, which means it shows relationships but not the direction of time. Two options exist. The standard approach is outgroup rooting: include a sequence known to fall outside your group of interest, then place the root on the branch leading to it in Tree Explorer. The alternative is View | Root on Midpoint, which roots on the midpoint of the longest path between two taxa. Midpoint rooting assumes a roughly constant rate across lineages, so treat it as a convenience, not a test.

Save your work with File | Save Current Session (a .mts session file), export the topology with File | Export Current Tree in Newick (.nwk) format, and close with File | Exit Tree Explorer [6]. For a figure, Image | Save as PDF File saves the tree as a PDF [8].

Worked Example

Two routes let you reproduce a full analysis.

Option A, bundled data. Open Drosophila_Adh.meg from the MEGA Examples folder. Run Models | Find Best DNA/Protein Models (ML) and note the model with the lowest BIC. Then run Phylogeny | Construct/Test Maximum Likelihood Tree with that model and either Bootstrap (500 or 1000 replicates) or the Adaptive option. Export with File | Export Current Tree [6].

Option B, protein data. Take five insulin precursors from UniProt: human P01308, mouse Ins2 P01326, chicken P67970, zebrafish O73727, and bovine P01317. Align them with Clustal Omega 1.2.4 defaults, which gives 112 columns. Biopython 1.88 identity distances (the proportion of differing columns) are 0.179 for human-mouse, 0.196 for human-cow, 0.250 for mouse-cow, 0.366 for human-chicken, 0.357 for cow-chicken, 0.393 for mouse-chicken, and 0.545 to 0.580 between zebrafish and the rest. A neighbor-joining tree built in Biopython from those distances with 1000 bootstrap replicates (random seed 1) and rooted on zebrafish gives ((human, mouse), cow) as the mammal clade, then chicken, then zebrafish. Bootstrap support is 100% for the mammal clade and 89.1% for human plus mouse (891 of 1000 replicates; 108 replicates instead grouped chicken, zebrafish, and mouse against the rest). Repeating this in MEGA with Phylogeny | Construct/Test Neighbor-Joining Tree, p-distance, and the Bootstrap method should return the same topology, though exact percentages will shift with the random resampling and distance options. This is a single short gene, so it is a gene tree, and modest support values are expected.

Common Mistakes and How to Fix Them

  • The tree has no meaningful root. Symptom: an unrooted starburst or a tree whose deepest split makes no biological sense. Cause: ML and NJ trees are unrooted by default. Fix: add an outgroup sequence and root on its branch, or use View | Root on Midpoint if you accept its assumptions.
  • Bootstrap values are all low. Symptom: low support on most interior branches. Cause: too few informative sites, a gappy alignment, or a model that does not fit. Fix: check the alignment for misaligned regions, revisit the model choice, and consider whether a single short gene can resolve the relationships you are asking about.
  • The alignment is mostly gaps after analysis. Symptom: a tree built from a fraction of your original columns. Cause: Complete-deletion drops every column with a gap in any sequence. Fix: switch to Pairwise-deletion or Partial-Deletion with a Site Coverage Cutoff [7].
  • The tree changes completely when you change the model. Symptom: different topologies from different substitution models. Cause: the data do not strongly support any one topology, often because the sequences are short or closely related. Fix: report the model selection result and treat the topology as provisional.
  • The analysis runs for hours. Symptom: no output after a long wait. Cause: Subtree-Pruning-Regrafting with a large alignment and 1000 bootstrap replicates. Fix: use the Adaptive bootstrap option, raise the thread count, or start with neighbor-joining to confirm the data are usable [7][10].
  • Sequence names are unreadable in the figure. Symptom: accession numbers or truncated labels. Cause: names were never edited in the Alignment Explorer. Fix: rename taxa in the Alignment Explorer before building the tree.

Limitations

MEGA is built for single-gene and small multi-gene datasets. For phylogenomic matrices with thousands of loci, IQ-TREE 2 is the more common choice: it integrates many substitution models and efficient methods designed for genomic-scale data [14]. If you move in that direction, note that IQ-TREE's ultrafast bootstrap (UFBoot) is not comparable to a standard bootstrap percentage, and its documentation warns against mixing the two [13]. The current IQ-TREE release shown on its site was 3.1.4 in September 2026, and the binary name may be iqtree2 or iqtree3 depending on how you installed it.

Bootstrap values themselves need careful reading. MEGA's help states that a branch with a bootstrap value of 95% or higher is considered correct, citing Felsenstein (1985), Efron (1982), and Nei and Kumar (2000) [9]. That is MEGA's convention, and it is stronger than the modern reading. A bootstrap value measures how consistently your data support a clade under resampling, not the probability that the clade is true [13]. Standard bootstrap resamples columns of a single alignment, so it cannot account for alignment uncertainty; Edgar 2022 showed that some topologies with high bootstrap support under standard methods are incorrect, and some with low bootstrap can be confidently resolved [12].

Platform support has shifted across releases. The MEGA12 paper describes the release as Windows with a macOS app in testing, while the current home page lists macOS and Linux downloads alongside Windows [1][10]. Check the download page, not the paper. MEGA 12.1 also added a redesigned Calibration Editor that connects to the TimeTree database for molecular time estimates, which is useful for divergence dating but outside the scope of a plain tree build [1].

Frequently Asked Questions

How do I make a phylogenetic tree in MEGA from DNA sequences?

Load your FASTA file into the Alignment Explorer as DNA, align with ClustalW or MUSCLE, run Models | Find Best DNA/Protein Models (ML), then build the tree with Phylogeny | Construct/Test Maximum Likelihood Tree using the best model [2][5][6]. Add a bootstrap test in the same dialog if you want support values. Export the result as Newick or PDF.

Can I make a phylogenetic tree from protein sequences instead?

Yes. Choose protein when you create the alignment, and MEGA will offer amino acid substitution models in the model selection step [2][5]. The workflow is otherwise identical. For coding DNA, aligning through the translated protein is often better because it keeps gaps in frame [3].

What is a bootstrap phylogenetic tree and what do the numbers mean?

Sites are resampled with replacement and the tree is rebuilt several hundred times; the percentage of replicates that recover each interior branch is the bootstrap value [9]. It reflects consistency under resampling of your alignment, not the probability that the clade is real [13]. Conventional practice uses 500 to 2,000 replicates, though MEGA's Adaptive option can stop earlier once the standard errors are small enough [10].

Which is better, maximum likelihood or neighbor-joining in MEGA?

Maximum likelihood is the standard for published trees because it uses an explicit substitution model and evaluates many candidate topologies. Neighbor-joining is far faster and useful for a first look at the data or for very large alignments [8]. If the two disagree, trust the ML result and investigate why.

Do I need MEGA 12, or will an older version work?

MEGA 12.1 is the current release as of October 2026, and the MEGA12 paper describes faster model selection, parallel ML inference, and an improved Tree Explorer [1][10]. Older versions will still build trees, but the menu paths and options may differ from this guide. Check the download page for the current release and your operating system.

References

  1. MEGA: Molecular Evolutionary Genetics Analysis (home page)
  2. MEGA 12 help: Aligning Sequences
  3. MEGA 12 help: Menu in Alignment Explorer
  4. MEGA 12 help: Data Menu (in Alignment Explorer)
  5. MEGA 12 help: Find Best DNA/Protein Models (ML)
  6. MEGA 12 help: Constructing Likelihood Trees
  7. MEGA 12 help: Analysis Preferences (Maximum Likelihood)
  8. MEGA 12 help: Building Trees From Sequence Data
  9. MEGA 12 help: Bootstrap Test of Phylogeny
  10. Kumar S et al. 2024. MEGA12: Molecular Evolutionary Genetic Analysis Version 12 for Adaptive and Green Computing. Mol Biol Evol 41:msae263
  11. Felsenstein J. 1985. Confidence limits on phylogenies: an approach using the bootstrap. Evolution 39:783-791
  12. Edgar RC. 2022. Muscle5: High-accuracy alignment ensembles enable unbiased assessments of sequence homology and phylogeny. Nat Commun 13:6968
  13. IQ-TREE Frequently Asked Questions (UFBoot interpretation)
  14. Minh BQ et al. 2020. IQ-TREE 2: New Models and Efficient Methods for Phylogenetic Inference in the Genomic Era. Mol Biol Evol 37:1530-1534

Related Articles