# How to Do a Multiple Sequence Alignment with Clustal Omega (and When to Use MAFFT or MUSCLE)

A multiple sequence alignment (MSA) places three or more sequences side by side so that homologous positions line up in the same column. EMBL-EBI describes the value of an MSA as highlighting areas of similarity that may be associated with specific features, and as the basis for phylogenetic analysis that models substitutions over evolution [1]. Clustal Omega is one of the most widely used tools for this job, and it is fast enough to handle large protein families on a laptop.

By the end of this guide you will be able to submit sequences to Clustal Omega through the web interface or the command line, read the output correctly, and recognize when MAFFT or MUSCLE will give you a better alignment. The worked example uses five insulin precursor proteins so you can check your own numbers against known values.

## Quick Answer

- Use Clustal Omega for routine protein alignments, especially larger sets where speed matters.
- Paste sequences in a recognized format (FASTA, EMBL/UniProt, NBRF/PIR, GDE, ALN/Clustal, GCG/MSF or RSF). Raw sequence without headers is rejected [1].
- The EBI web service accepts a maximum of 4000 sequences or a 4 MB file, whichever is smaller. For larger jobs, download Clustal Omega and run it locally [1].
- Set the output format to "Clustal w/ numbers" under "More options" if you want residue numbers in the alignment [1].
- Read the consensus line: `*` means a fully conserved residue, `:` means strongly similar properties, `.` means weakly similar properties [1].
- Switch to MAFFT L-INS-i for fewer than about 200 sequences when accuracy matters more than speed, or MUSCLE v5 when you need alignment ensembles to assess uncertainty [4][5][8].

## Step 1: Prepare Your Sequences

Clustal Omega accepts nucleic acid or protein sequences in NBRF/PIR, EMBL/UniProt, Pearson (FASTA), GDE, ALN/Clustal, GCG/MSF and RSF formats [1]. FASTA is the most common. Each sequence needs a header line starting with `>`, followed by the sequence on one or more lines.

The first word of each header line must be unique. If two sequences share the same identifier, the EBI service returns "Two sequences cannot share the same identifier" [1]. A header with no sequence data triggers "Entry found which does not contain a sequence" [1]. If the format is unsupported or malformed, for example sequences pasted without headers, you get "minimum 2 sequences required" even when you supplied more than two, because the parser cannot tell where one sequence ends and the next begins [1].

MSA tools are designed for three or more sequences. To align only two sequences, use a pairwise alignment tool instead [1]. The free [Pairwise Alignment Tool](/tools/pairwise-alignment) on this site handles that case.

Decide early whether you are aligning proteins or nucleotides. Mixing DNA and protein in one input is a common mistake. Clustal Omega auto-detects the type, but you can force it with `-t/--seqtype` on the command line or the `stype` parameter through the EBI REST API [2].

For protein-coding DNA, align the translated proteins and then map codons back (a codon-aware alignment) instead of aligning raw nucleotides. This keeps gaps in multiples of three and preserves the reading frame. Tools such as MEGA's Align by ClustalW/MUSCLE (Codons) or PAL2NAL do this mapping.

## Step 2: Run Clustal Omega Through the Web Interface

The EBI Clustal Omega page is the fastest route for a one-off alignment. For a quick check on a handful of sequences, the site's [OmniAlign](/tools/omnialign) alignment tool is another option. Paste your sequences into the input box, choose the output format, and submit. The default output format is listed as "ClustalW with character counts" in the REST API parameter details, but the FAQ describes the web tool as producing Clustal without numbering by default [1][2]. Check the "Output alignment format" option before you submit so you know what you will get.

To get residue numbers in the web tool, open "More options" and set the output alignment format to "Clustal w/ numbers" [1]. Clicking a parameter name opens its help text.

Results are stored for one week after submission. For large or long-running jobs, email submission is recommended. The exact Clustal Omega version that ran your job appears on the Submission details tab [1].

The default Clustal (ALN) output truncates sequence identifiers to 30 characters, and PHYLIP output truncates them to 10 characters [1]. Use short unique IDs to avoid name collisions in the output.

## Step 3: Run Clustal Omega from the Command Line

The standalone version removes the web service size limits and uses multiple processors where present. It can align hundreds of thousands of sequences in a few hours [3]. Clustal Omega 1.2.2 (2016-07-01) is listed on the project home page with source code for UNIX, precompiled binaries for Windows, UNIX and Mac, and packages for Gentoo, Slackware and Debian, as of October 2026 [3]. Check the download page for the current release.

A basic protein alignment:

```bash
clustalo -i ins5.fa -o ins5.aln --outfmt=clu
```

The `--outfmt` choices are `a2m=fa[sta]`, `clu[stal]`, `msf`, `phy[lip]`, `selex`, `st[ockholm]`, `vie[nna]`. The default output format is FASTA. Other useful options include `-t/--seqtype {Protein, RNA, DNA}` to force the sequence type, `--threads=<n>` to set the number of threads, `--force` to overwrite an existing output file, `--dealign` to remove existing gaps before aligning, `--full` to use the full distance matrix for the guide tree instead of mBed, `--distmat-out=<file>` to write the pairwise distance matrix, `--percent-id` to convert distances into percent identities, `--guidetree-out=<file>` to save the guide tree, `--iter=<n>` for combined guide-tree and HMM iterations, `--resno` to print residue numbers in Clustal format, `--wrap=<n>` to set line width, and `--output-order={input-order,tree-order}` to control sequence order.

To get a percent identity matrix alongside the alignment:

```bash
clustalo -i ins5.fa -o ins5.aln --outfmt=clu --full --percent-id --distmat-out=ins5.dist
```

Note that the EBI percent identity matrix is computed from the final alignment, while local `clustalo --percent-id` converts guide-tree distances, so the numbers may differ slightly from the web result.

## Step 4: Read the Output

The Clustal format alignment has one line per sequence, with a consensus line below each block. An asterisk (`*`) marks a column with a single, fully conserved residue. A colon (`:`) marks conservation between groups of strongly similar properties. A period (`.`) marks conservation between groups of weakly similar properties [1].

The strong (colon) groups are STA, NEQK, NHQK, NDEQ, QHRK, MILV, MILF, HY, FYW. The weak (period) groups are CSA, ATV, SAG, STNK, STPA, SGND, SNDEQK, NDEQHK, NEQHRK, FVLIM, HFY [1].

For DNA and RNA alignments the same symbols appear, but EBI advises that only the asterisk is meaningful and the colon and period should be ignored [1]. The similarity groups are defined for amino acid properties, not nucleotide chemistry.

Clustal Omega does not report ClustalW-style all-against-all pairwise scores because its guide tree is built differently. EBI instead computes a pairwise percent identity matrix from the finished alignment, downloadable from the Results Summary tab [1]. By default Clustal Omega uses the mBed sampling method to speed up guide tree calculation, so a traditional guide tree is not shown. EBI also provides a neighbor-joining phylogenetic tree in Newick format; its values are branch lengths, an indication of evolutionary distance [1].

Downloaded alignments are plain text files. EBI suggests specialist viewers such as Jalview or Genedoc for inspecting them [1].

## Step 5: Choose MAFFT or MUSCLE When Clustal Omega Is Not Enough

Clustal Omega uses the HHalign algorithm with its default settings as its core alignment engine, with a Gonnet transition matrix, gap opening penalty of 6 bits and gap extension of 1 bit [1]. It was published by Sievers et al. 2011 as a fast, scalable method for high-quality protein multiple sequence alignments [6]. Part of that speed comes from the mBed sampling method used for the guide tree [1].

MAFFT version 7 is described in Katoh and Standley 2013 [7]. Its accuracy-oriented methods are recommended for fewer than about 200 sequences [4]: L-INS-i for sequences with one alignable domain plus flanking regions, G-INS-i for globally alignable sequences, and E-INS-i when the nature of the sequences is unclear.

```bash
mafft --localpair --maxiterate 1000 input > output
mafft --globalpair --maxiterate 1000 input > output
mafft --ep 0 --genafpair --maxiterate 1000 input > output
```

The shortcuts `linsi`, `ginsi` and `einsi` do the same thing. For speed, FFT-NS-2 (`mafft --retree 2 --maxiterate 0`) and FFT-NS-i (`mafft --retree 2 --maxiterate 1000`) are available. `mafft --auto input > output` automatically chooses among L-INS-i, FFT-NS-i and FFT-NS-2 based on data size [4]. Options `--amino` and `--nuc` force amino acid or nucleotide input, and `--reorder` outputs sequences in aligned order instead of input order [4].

MUSCLE v5 takes a different approach. It builds an ensemble of high-accuracy alignments by perturbing an HMM and permuting its guide tree, and Edgar 2022 reports that some topologies with high bootstrap support are incorrect when alignment uncertainty is ignored [8].

```bash
muscle -align seqs.fa -output aln.afa
```

Output is aligned FASTA by default. `-threads` defaults to the number of CPU cores (or 20 if more than 20 cores). `-amino` or `-nt` force the alphabet. To generate alignment ensembles, use `-stratified` (default 4 replicates) or `-diversified` (default 100 replicates), and `-perm` to set the guide tree permutation (none, abc, acb, bca). For large sets where the default align (PPP) algorithm is too slow, the documented alternative is `muscle -super5 seqs.fa -output aln.afa` [5].

MUSCLE v5 documents the `-align`/`-output` syntax, while the older MUSCLE v3 manual used `muscle -in seqs.fa -out seqs.afa`. Check which version is installed before copying commands from older tutorials [5].

## Worked Example

Five insulin precursor proteins were fetched from UniProt REST on 2026-10-01: human INS P01308 (110 aa), mouse Ins2 P01326 (110 aa), chicken INS P67970 (107 aa), zebrafish ins O73727 (108 aa), bovine INS P01317 (105 aa). UniProt P01308 annotates signal peptide 1-24, insulin B chain 25-54, C-peptide 57-87 and insulin A chain 90-110.

```bash
clustalo -i ins5.fa -o ins5.aln --outfmt=clu
```

With Clustal Omega 1.2.4 and default settings, the result is a 112-column alignment. The consensus line has 43 `*`, 12 `:` and 6 `.` columns. The B-chain core (QHLCGSHLV.ALYLVCG..GFFY) and A chain (KRGIV.QCC...CS...L.NYCN, ending in fully conserved NYCN) are conserved. Zebrafish needs gaps in the signal peptide and C-peptide, which is where Clustal Omega and MAFFT L-INS-i (`mafft --localpair --maxiterate 1000 --clustalout`, v7.505) place gaps differently.

The percent identity matrix from `clustalo --full --percent-id --distmat-out` gives human vs mouse 81.8, human vs bovine 81.0, human vs chicken 62.6, human vs zebrafish 38.9, chicken vs zebrafish 44.9.

An independent check with Biopython 1.88 PairwiseAligner (global, BLOSUM62, open -10, extend -0.5), identities divided by alignment length, gives human-mouse 90/110 = 81.8%, human-bovine 88/110 = 80.0%, human-chicken 70/110 = 63.6%, human-zebrafish 52/127 = 40.9%, mouse-zebrafish 54/126 = 42.9%, chicken-zebrafish 61/121 = 50.4%.

Percent identity depends on the denominator (alignment length vs shorter sequence vs non-gap columns). Human-zebrafish is 40.9% over alignment length but 48.1% over the shorter sequence. Expect mammals at roughly 75-84% identity, chicken around 60-69% to mammals, and zebrafish around 39-50% to all others.

## Common Mistakes and How to Fix Them

- **"minimum 2 sequences required"**: The input format is unsupported or malformed. Check that every sequence has a header line starting with `>` and that sequences are not pasted as raw text without headers [1].
- **"Two sequences cannot share the same identifier"**: The first word of each header line must be unique. Rename duplicates before resubmitting [1].
- **"Entry found which does not contain a sequence"**: A header line has no sequence data below it. Remove empty entries [1].
- **"stream closed" error**: The upload was far over the EBI limits of 4000 sequences or 4 MB. Download Clustal Omega and run it locally [1].
- **Truncated sequence names in output**: The default Clustal (ALN) format truncates identifiers to 30 characters, and PHYLIP truncates to 10. Use short unique IDs [1].
- **Colon and period symbols in a DNA alignment**: These are only meaningful for protein alignments. For DNA and RNA, read only the asterisk [1].
- **Mixing DNA and protein in one input**: Force the sequence type with `-t/--seqtype` on the command line or `stype` through the EBI REST API [2].
- **Gaps not in multiples of three in a coding alignment**: Align translated proteins and map codons back instead of aligning raw nucleotides.
- **Unreliable regions in divergent alignments**: Very divergent sequences produce poorly aligned columns. Trimming those columns (for example with trimAl) before tree building is a common step.

## Limitations

Clustal Omega is a command-line tool at heart. The web servers may impose size limits that the standalone version does not [3]. The EBI service stores results for one week, so download anything you need to keep [1].

The default mBed sampling means a traditional guide tree is not shown, and ClustalW-style all-against-all pairwise scores are not reported [1]. If you need those scores, compute a percent identity matrix from the finished alignment or use a different tool.

Version numbers change. Clustal Omega 1.2.2 was listed on the project home page as of October 2026, but the command-line options described here come from the help text of version 1.2.4 (Ubuntu package), which is newer. Options can differ by build. MAFFT versions 7.463 through 7.486 had a serious bug that also affected `--auto`; upgrade to 7.487 or later. Check the current documentation for both tools before relying on a specific option.

The EBI pages conflict on the default output format: the FAQ says the web tool outputs Clustal without numbering by default, while the REST API parameter details list "ClustalW with character counts" as default [1][2]. Check the "Output alignment format" option in the web form.

## Frequently Asked Questions

### How do I use Clustal Omega for a protein alignment?

Paste your protein sequences in FASTA or another supported format into the EBI web form, or run `clustalo -i input.fa -o output.aln --outfmt=clu` locally. The tool auto-detects protein versus nucleotide, but you can force it with `-t Protein`. Check the output format option before submitting so you get the numbering and format you need.

### What is the difference between Clustal Omega and ClustalW?

ClustalW is the older program in the same family. Clustal Omega uses the HHalign algorithm as its core alignment engine and the mBed method for guide tree calculation, which makes it much faster on large sets [1][3]. Clustal Omega does not report ClustalW-style all-against-all pairwise scores because its guide tree is built differently [1].

### When should I use MAFFT instead of Clustal Omega?

Use MAFFT L-INS-i, G-INS-i or E-INS-i when you have fewer than about 200 sequences and accuracy matters more than speed [4]. L-INS-i suits sequences with one alignable domain plus flanking regions, G-INS-i suits globally alignable sequences, and E-INS-i is recommended when the nature of the sequences is unclear. For large sets, MAFFT's FFT-NS-2 and FFT-NS-i methods are faster.

### What is a MUSCLE alignment ensemble?

MUSCLE v5 can generate multiple alignments by perturbing an HMM and permuting its guide tree, using `-stratified` (default 4 replicates) or `-diversified` (default 100 replicates) [5]. Edgar 2022 reports that some topologies with high bootstrap support are incorrect when alignment uncertainty is ignored [8]. The ensemble lets you see which parts of the alignment are stable and which are not.

### Which multiple sequence alignment tool should I start with?

Start with Clustal Omega for routine protein work, especially when you have more than a few hundred sequences. Move to MAFFT for smaller, more divergent sets where accuracy matters. Use MUSCLE v5 when you need to assess alignment uncertainty. All three are free, and the choice often comes down to the size and divergence of your dataset.

## References

1. [EMBL-EBI Job Dispatcher: Clustal Omega FAQs](https://www.ebi.ac.uk/jdispatcher/docs/faqs/clustal/)
2. [EMBL-EBI REST API: Clustal Omega outfmt parameter details](https://www.ebi.ac.uk/Tools/services/rest/clustalo/parameterdetails/outfmt)
3. [Clustal Omega home page (clustal.org)](http://www.clustal.org/omega/)
4. [MAFFT manual page (mafft.cbrc.jp)](https://mafft.cbrc.jp/alignment/software/manual/manual.html)
5. [MUSCLE v5 manual: align command](https://drive5.com/muscle5/manual/cmd_align.html)
6. [Sievers F et al. 2011. Fast, scalable generation of high-quality protein multiple sequence alignments using Clustal Omega. Mol Syst Biol 7](https://doi.org/10.1038/msb.2011.75)
7. [Katoh K, Standley DM. 2013. MAFFT Multiple Sequence Alignment Software Version 7. Mol Biol Evol 30:772-780](https://doi.org/10.1093/molbev/mst010)
8. [Edgar RC. 2022. Muscle5: High-accuracy alignment ensembles. Nat Commun 13:6968](https://doi.org/10.1038/s41467-022-34630-w)
9. [Madeira F et al. 2024. The EMBL-EBI Job Dispatcher sequence analysis tools framework in 2024. Nucleic Acids Res 52:W521-W525](https://doi.org/10.1093/nar/gkae241)

## Related Articles

- [Multiple Sequence Alignment: Common Pitfalls and Quality Checks](/blog/guides/multiple-sequence-alignment-common-pitfalls-and-quality-checks)
- [Sequence Alignment: Choosing the Right Method for Your Biological Question](/blog/guides/sequence-alignment-choosing-the-right-method-for-your-biological-question)
- [Phylogenetic Tree Workflow: From Aligned Sequences to a Defensible Figure](/blog/guides/phylogenetic-tree-workflow-from-aligned-sequences-to-a-defensible-figure)
- [BLAST NCBI Search: A Practical Guide to Sequence Similarity](/knowledge/molecular-biology/blast-ncbi-search)
- [How to Make a Phylogenetic Tree in MEGA from DNA or Protein Sequences](/blog/research-skills/how-to-make-a-phylogenetic-tree-in-mega)
- [How to Use SnapGene Viewer: Open a Plasmid, Find Restriction Sites and Check Primers](/blog/research-skills/how-to-use-snapgene-viewer)